Copied!

How do different languages
and educational systems
organize the same knowledge?

A cross-lingual analysis of mathematics, physics and chemistry textbook knowledge — across China, Germany, the United Kingdom and the United States — measured through structural graph metrics and curriculum alignment.

[▶ Demo · 2:53]

Same 173s cut — switching keeps your position.

1,140+
Concepts · full project (=1143 = 556+367+220)
1,100+
Relations
3
Languages
ZH · EN · DE
3
Disciplines
Math · Physics · Chemistry
4
Education Systems
DE · UK · US · CN
19
Models19-model benchmark
0.881
Overall F1 (weighted)
SCOPE · Full project graph: Math 556 · Physics 367 (incl. 1 pending sensor node) · Chemistry 220 = 1,140+ concepts (README.md §dataset). The 3D view shows the mathematics subgraph (manifest.json: 556 nodes · 525 relations · 219 groups).

Research Questions

Three complementary lenses on how the same knowledge is organized differently

RQ1: Language

How do different languages organize the same knowledge?

LDS-K reveals heterogeneous cross-linguistic relationships: ZH-DE (0.519) textbooks converge substantially; ZH-EN (0.934) and DE-EN (0.938) are near the within-language noise floor. A Null Model falsifies the language-divergence interpretation — the core signal is ΔLDS (cognitive minus textbook).

LDS-K: 0.519–0.938

RQ2: Discipline

How do different STEM subjects organize knowledge?

All three disciplines follow an early-peak-later-decline pattern. Knowledge density peaks at middle or elementary school, then drops sharply. Maximum prerequisite depth is bounded at 8.

CDS: 0.042–0.271 HDS: max 8

RQ3: Education System

How do different curricula organize the same subject?

Coverage scores range from 12.7% (NRW) to 95.4% (China). Granularity (safest reading) and governance (hypothesis, paper §8.6) both fit these differences.

CS: 12.7%–95.4%

Research Contributions

What this project delivers to the field

1

A multilingual STEM knowledge graph

1,140+ concepts and 1,100+ direct relations across mathematics, physics and chemistry in Chinese, English and German — extracted from 180 textbooks and validated against 92 gold-standard annotations (social F1 = 0.939; weighted overall 0.881).

2

Three structural metrics for knowledge organization

CDS (concept density), HDS (hierarchy depth) and LDS (language drift) quantify how knowledge is connected, sequenced and structured across languages and disciplines.

3

Curriculum coverage analysis framework

The Coverage Score measures textbook–curriculum alignment across four education systems (Germany, UK, US, China), revealing how different educational philosophies produce systematically different alignment patterns.

4

Empirical findings across 4 education systems

12 findings spanning density universals (F1–F3, F6–F8), cross-language divergence (F4–F5), curriculum design philosophy (F9–F10), and human-validation revision (F11–F12) — all independently verifiable from the released data and pipeline.

Methodology

From textbook text to structural insights in five steps

%%{init:{'theme':'default','themeVariables':{'primaryColor':'#f8fafc','primaryTextColor':'#0a1f3d','primaryBorderColor':'#c6cdd6','lineColor':'#1e3a8a','secondaryColor':'#f1f3f5','tertiaryColor':'#e7ebef'}}}%% graph LR T1[180 Textbooks
ZH · EN · DE] --> MIMO[LLM API
Concept Extraction] MIMO --> KG[Knowledge Graph
1,140+ Concepts · 1,100+ Relations] GL[92 Gold Labels
F1 = 0.939 social / 0.881 weighted] -.->|Validate| KG KG --> CDS[CDS: Concept Density] KG --> HDS[HDS: Hierarchy Depth] KG --> LDS[LDS: Language Divergence] KG --> CURR[4 Curricula
NRW · UK · US · CN] CURR --> CS[Coverage Score] CDS --> INSIGHTS[Educational Insights
F1–F12] HDS --> INSIGHTS LDS --> INSIGHTS CS --> INSIGHTS style GL fill:#f8fafc,stroke:#1e7e4d,stroke-dasharray:4 4
CDS
2|E| / (|V|·(|V|−1))
How densely connected knowledge is at each education level. Higher = more integrated.
HDS
max BFS depth(prerequisite)
Longest prerequisite chain for each concept. Measures hierarchical depth.
LDS
1 − mean(J_node, J_edge)
Cross-language structural divergence. 0 = identical, 1 = completely different.
CS
|V_tb ∩ V_cur| / |V_cur|
Textbook–curriculum alignment. How much of the official curriculum textbooks actually cover.
Fig 2: LDS Calculation Flow
Fig 2 — How LDS is computed: two language graphs, alignment, two similarity components
Baseline-Glossar: 9 reference lines every finding is measured against
1. Structure Null
Grade-randomized graphs: the overlap expected by pure chance.
2. Size-matched bootstrap
k = 15/25/35: comparison at equal vocabulary size.
3. Within-language noise floor (0.97)
Same-language divergence: the baseline no cross-language claim may sit below. (textbook legacy line; human line 0.92-0.96; LLM line 0.85-0.87)
4. Wikipedia aligned control
Domain-pure social concepts in ZH/EN/DE, independently extracted.
5. Human N=15 floor
Between-subject margin +0.015: no separable language signal at this N.
6. LLM within-subject signal
LDS-C 0.93–0.96 (LLM-as-subject), permutation p < 0.01: language code dominates.
7. Permutation test
Label shuffling, e.g. ZH-DE 55/55, p<0.004 (500 perm.).
8. Heterogeneity injection (q-scan)
Consistency demonstration, not a causal proof of the human null.
9. Margin threshold (≥ 0.10)
Operational heuristic for audit use, not a validated boundary.

Sources

Every textbook cited by at least one graph concept. Node cards show the top-3 sources per concept; this is the full inventory (204 titles).

Local texts (2026-09-12): 89 open files + CN 11 books OCR done, fact-level (51 chapter txts). Text-grounded: physics ZH 66.5% / EN 62.1%, chemistry ZH 65.9% / EN 53.2% (semantic layer; substring 33.8/25.5) (college nodes out of high-school scope; DE link-only).

Mathematics · 32 titles
TitleLangRefs
A First Course in Probability (Sheldon Ross)en24
AP Calculus AB/BCen34
Abitur Mathematik Leistungskurs (Baden-Württemberg)de28
Cambridge IGCSE Mathematics (0580)en21
Differential Equations and Linear Algebra (Edwards & Penney)en14
Erwin Fischer: Lineare Algebrade10
Gewöhnliche und partielle Differentialgleichungende14
Gilbert Strang: Introduction to Linear Algebra (18.06)en28
IB Mathematics: Analysis and Approaches SL/HLen18
Khan Academy Mathematics 3-4en10
Khan Academy Mathematics 3rd-5th Gradeen21
Khan Academy Mathematics 5-6en15
Khan Academy Mathematics 6th-8th Gradeen37
Khan Academy Mathematics K-2en13
Lambacher Schweizer 5-8 Klassede17
Lambacher Schweizer Mathematik (erweitert)de30
Otto Forster: Analysis 1de21
Otto Forster: Analysis 2de7
Stewart Calculus (8th ed.)en31
Wahrscheinlichkeitsrechnung und Statistikde24
Westermann Mathematik Klasse 9-10de26
人教版《偏微分方程基础》zh8
人教版初中数学七年级zh15
人教版初中数学九年级zh18
人教版初中数学八年级zh17
人教版小学数学一年级上册zh10
人教版小学数学一年级下册zh11
人教版小学数学二年级上册zh9
人教版小学数学二年级下册zh9
人教版高中数学选修2-2zh85
同济大学《线性代数》zh64
浙江大学《概率论与数理统计》zh64
Physics · 83 titles
TitleLangRefs
A-Level Physics (AQA)en136
A-Level Physics (OCR)en136
AP Physics 1en136
AP Physics 2en136
AP Physics C: E&Men136
AP Physics C: Mechanicsen136
Arfken: Mathematical Methodsde127
Auer: Physikde93
BBC Bitesize: KS3 Physicsen93
Cambridge IGCSE Physicsen136
CK-12: Elementary Physical Scienceen10
CK-12: Middle School Physicsen93
Cornelsen: Physik aktivde93
Cornelsen: Physik entdeckende10
Cornelsen: Physik Oberstufede136
Demtröder: Experimentalphysikde381
Dorn-Bader: Physikde136
DUDEN: Physik Abiturde136
Duden: Physik Kompaktde103
Feynman Lectures on Physicsen254
GCSE AQA Physicsen136
GCSE Edexcel Physicsen136
Glencoe: Physics: Principles and Problemsen93
Greiner: Klassische Mechanikde127
Greiner: Quantenmechanikde127
Greiner: Thermodynamik und Statistische Mechanikde127
Griffiths: Introduction to Electrodynamicsen127
Halliday Resnick Walker: Fundamentals of Physicsen254
Holt: Middle School Scienceen93
IB Physics SL/HLen136
IGCSE Physics (0625)en136
Jackson: Klassische Elektrodynamikde127
Kern: Physikde136
Khan Academy Physicsen93
Khan Academy: Elementary Scienceen20
Kittel: Introduction to Solid State Physicsen127
Kleintolkroemer: Classical Electrodynamicsen127
Klett: Physikde93
Lambacher Schwere: Physikde186
Landau Lifshitz: Mechanicsen127
LEIFIphysik Induktionde1
National Geographic: Force, Motion, and Simple Machinesen10
NPTEL Sensor Technologiesen1
Pearson: Prentice Hall Science Exploreren93
Schiff: Quantenmechanikde127
Schrödinger: Quantum Mechanicsen127
Serway Jewett: Physics for Scientists and Engineersen254
Thieme: Physikde136
Tipler: Physikde254
Westermann: Physikde103
Westermann: Physik Oberstufede136
Young Freedman: University Physicsen254
人教版初中物理九年级全一册zh93
人教版初中物理八年级上zh93
人教版初中物理八年级下zh93
人教版小学科学三年级上zh10
人教版小学科学五年级上zh10
人教版小学科学四年级上zh10
人教版高中物理必修第一册zh136
人教版高中物理必修第三册zh136
人教版高中物理必修第二册zh136
人教版高中物理选择性必修第一册zh136
人教版高中物理选择性必修第三册zh136
人教版高中物理选择性必修第二册zh136
光学(赵凯华)zh127
力学(漆安慎)zh127
北师大版初中物理八年级zh93
北师大版小学科学四年级zh10
原子物理学(杨福家)zh127
大学物理(马文蔚)zh254
普通物理学(程守洙)zh254
沪教版小学科学三年级zh10
沪科版初中物理八年级zh93
沪科版高中物理zh136
热学(汪志诚)zh127
理论力学(梁昆淼)zh127
电动力学(郭硕鸿)zh127
电磁学(赵凯华)zh127
粤教版初中物理八年级zh93
粤教版高中物理zh136
苏科版初中物理八年级zh93
量子力学(曾谨言)zh127
鲁科版高中物理zh136
Chemistry · 89 titles
TitleLangRefs
A-Level Chemistry (AQA)en51
A-Level Chemistry (OCR)en51
AP Chemistryen51
Atkins: Physical Chemistryen123
Atkins: Physikalische Chemiede123
Auer Chemiede46
BBC Bitesize KS3 Chemistryen46
Brown: Chemistryen123
Bruice: Organic Chemistryen123
Bruice: Organische Chemiede123
Cambridge IGCSE Chemistryen51
Chang: Chemistryen123
CK-12 Chemistryen46
Clayden: Organic Chemistryen123
Clayden: Organische Chemiede123
Cornelsen Chemie entdeckende46
Cornelsen Chemie Oberstufede51
Dorn-Bader Chemiede46
Dorn-Bader Chemie Abiturde51
DUDEN Chemie Abiturde51
Duden Chemie Kompaktde46
Edexcel GCSE Chemistryen51
Engel: Physical Chemistryen123
Engel: Physikalische Chemiede123
GCSE AQA Chemistryen51
GCSE AQA Combined Scienceen51
GCSE Edexcel Chemistryen51
Glencoe Chemistryen46
Harris: Quantitative Chemical Analysisen123
Harris: Quantitative Chemische Analysede123
Holt Chemistryen46
Housecroft: Anorganische Chemiede123
Housecroft: Inorganic Chemistryen123
IB Chemistry SL/HLen51
IGCSE Chemistry (0620)en51
Kern Chemie Abitur LKde51
Khan Academy Chemistryen46
Klett Chemiede46
Laidler: Physikalische Chemiede123
Levine: Physical Chemistryen123
Levine: Physikalische Chemiede123
Mortimer: Physikalische Chemiede123
Pearson: Prentice Hall Chemistryen46
Petrucci: General Chemistryen123
Shriver & Atkins: Anorganische Chemiede123
Shriver & Atkins: Inorganic Chemistryen123
Silberberg: Chemiede123
Silberberg: Chemistryen123
Skoog: Analytical Chemistryen123
Skoog: Analytische Chemiede123
Solomons: Organic Chemistryen123
Thieme Chemiede51
Wade: Organic Chemistryen123
Westermann Chemiede46
Westermann Chemie Oberstufede51
Zumdahl: Chemistryen123
人教版初中化学九年级上zh46
人教版初中化学九年级下zh46
人教版高中化学必修第一册zh51
人教版高中化学必修第二册zh51
人教版高中化学选择性必修第一册(化学反应原理)zh51
人教版高中化学选择性必修第三册(有机化学基础)zh51
人教版高中化学选择性必修第二册(物质结构与性质)zh51
分析化学(华东理工)zh123
分析化学(武汉大学)zh123
化工原理(柴诚敬)zh123
北师大版初中化学zh46
无机化学(大连理工)zh123
无机化学(武汉大学)zh123
有机化学(汪小兰)zh123
有机化学(邢其毅)zh123
材料化学(曾兆华)zh123
沪教版初中化学zh46
沪科版高中化学zh51
湘教版初中化学zh46
湘教版高中化学zh51
物理化学(傅献彩)zh123
物理化学(南京大学)zh123
环境化学(戴树桂)zh123
生物化学(王镜岩)zh123
科粤版初中化学zh46
粤教版初中化学zh46
粤教版高中化学zh51
结构化学(周公度)zh123
苏教版初中化学zh46
苏教版高中化学zh51
高分子化学(潘祖仁)zh123
鲁教版初中化学zh46
鲁科版高中化学zh51

Refs = citations across concepts. New textbooks are appended here on ingestion (see docs/physics_sourcing.md for the gap list). Raw excerpts live in data/textbook/.

Findings

Five finding sections covering twelve numbered results (F1–F12): how knowledge is organized across disciplines, languages, education systems, and individuals

Finding A: Knowledge density peaks early

All three STEM disciplines concentrate their densest knowledge connections at introductory stages, not advanced ones.

Mathematics peaks at middle school (CDS = 0.271). Physics peaks at elementary school (CDS = 0.222). Chemistry shares the same middle school peak (CDS = 0.042). After the peak, density drops sharply — a 3.7× decline from middle to high school in mathematics.

This contradicts the natural assumption that "more advanced knowledge is more densely connected." Instead, curricula are designed to maximize connection density during foundational stages (early integration) before branching into specialization (later differentiation).

0.271
Math CDS (Middle)
0.222
Physics CDS (Elementary)
3.7×
Middle → High drop
ZH/EN/DE independent verification
Fig 3: CDS by Education Level
Fig 3 — Mathematics CDS peaks at middle school, then declines 3.7×
Fig 7: Three Subject CDS
Fig 7 — The same early-peak pattern holds across all three STEM disciplines
Details · how Fig 3 / Fig 7 are made

Source. CDS per level from the math graph (Fig 3) and the three-discipline matrix (outputs/physics_comparison.json + chemistry_comparison.json, Fig 7) — the same numbers as the CDS Terrain view.

Method. CDS = 2|E|/(|V|·(|V|−1)) per level cell; the math graph renders 238 of 525 relations. Nothing is rescaled.

Reading. Height = density; an early peak then decline in all three subjects (math middle 0.271, physics elementary 0.222, chemistry middle 0.042).

Finding B: Knowledge structures stay shallow

Educational knowledge has a universal upper bound on prerequisite depth, regardless of discipline.

Maximum prerequisite chain depth is HDS ≤ 8 for mathematics and HDS ≤ 6 for physics. Mean depth is even lower: 0.40 for math, 0.85 for physics. Over 83% of math concepts have no prerequisite chains at all (root concepts).

Physics has 2.1× more sequential depth than mathematics (60% roots in physics vs 83% in math), reflecting its more cumulative knowledge structure. But both disciplines respect the same shallow upper bound — suggesting a pedagogical or cognitive constraint on how knowledge can be sequenced.

8
Math HDS (max)
0.40
Math HDS (mean)
83%
Root concepts (Math)
2.1×
Physics depth vs Math
Fig 5: HDS Distribution
Fig 5 — HDS distribution: most concepts have zero depth; the longest chain reaches only 8

Finding C: Languages organize knowledge differently

Textbook LDS-K shows convergence (ZH-DE 0.519, ZH-EN 0.934, DE-EN 0.938) — an observation consistent with curriculum tradition; the core signal is ΔLDS.

Textbook structures look more convergent than random expectation — but decontamination falsifies this reading. ZH-DE (LDS-K = 0.519) collapses to 0.990 once de-labels containing Chinese text (167/219) are removed: the apparent convergence is a label artifact, not content convergence. ZH-EN (0.934) and DE-EN (0.938) sit near the within-language noise floor (0.97, textbook line) either way.

The pattern is also topic-dependent: within each language pair, LDS varies by up to 0.2 across the 5 social topics analyzed, suggesting that some knowledge domains are more culturally inflected than others.

0.519
ZH–DE (0.519)
0.938
DE–EN (0.938)
0.934
ZH–EN (0.934)
~0.2
Topic variance
Fig 4: Null Model Suite — Full vs Structure Null
Fig 4 — Null Model Suite: Full vs Structure Null. Textbook structures converge more than random expectation
Fig 8: LDS-K decontamination — convergence collapses after label cleaning
Fig 8 — Decontamination (T1): removing CJK-contaminated de-labels collapses ZH-DE 0.52 → 0.99. Verdict: artifact

Honesty note (T1 decontamination, 2026-09): ZH-DE J_node 0.556 → 0.020 after removing CJK-contaminated de-labels; size-matched bootstrap reverses the wiki comparison at every k. Verdict: label artifact (paper §3.8, Fig 8).

Finding D: Education systems emphasize different curriculum designs

Textbook–curriculum coverage varies dramatically — not necessarily quality differences: granularity (safest reading) and governance (hypothesis, paper §8.6) both fit the data.

Coverage scores range from 12.7% (NRW) to 95.4% (China). China shows near-perfect alignment under a centralized national curriculum. The UK (37.3%) shows moderate alignment through its national curriculum. The US (17.2%) and NRW (12.7%) reflect broad or detailed standards that afford local adaptation. These differences are consistent with educational governance (hypothesis, paper §8.6) — but granularity (A) remains the safest interpretation; the data fit all three explanations below.

95.4%
🇨🇳 China (highest)
37.3%
🇬🇧 UK
17.2%
🇺🇸 US
12.7%
🇩🇪 NRW Germany

Fig 8 — Overall Coverage Score (paper §8.5: NRW 12.7% · UK 37.3% · US 17.2% · China 95.4%)

Details · how Fig 8 is made

Source. Overall column of config/expert_graphs/coverage_all_curricula.json (paper §8.5) — the same table behind the coverage cards and Coverage Towers.

Method. Keyword bridge + 2-overlap threshold; overall = matched/total across all stages per system.

Reading. Bar = system overall; the gap runs China 95.4% down to NRW 12.7%.

F9 Coverage scores vary dramatically across education systems (12.7%–95.4%). Granularity (safest reading) and governance (hypothesis, paper §8.6) both fit these differences.
F10 China's near-perfect alignment (95.4%) versus NRW's low alignment (12.7%) reflects the spectrum from centralized curriculum to local adaptation.

Finding E: N=15 falsifies ΔLDS > 0 between-subject — language signal needs within-subject design

Extended human study (N=15) lands on the split-half floor: between-subject designs cannot separate language from participant variance. LLM within-subject (§5) shows the signal clearly.

Extended study (6 DE · 6 ZH · 3 EN): concept-level LDS-C (human) is 0.93–0.96 — indistinguishable from the within-language split-half floor (0.92–0.96) and label permutation (0.94). Concept-level ΔLDS ≈ 0 (−0.05…+0.05); relation-level Δ is not comparable across sparsity regimes. Pilot N=8 values (0.70–0.75) were not replicated and are reported for transparency only. The earlier human-vs-simulation comparison (p=0.05) is withdrawn (measurement-scale drift).

0.93–0.96
Human LDS-C (N=15)
0.92–0.96
Split-half floor
≈0
Concept ΔLDS
+0.08–0.09
LLM within-subject margin
Level DE–ZH DE–EN ZH–EN Rank Order
Human N=15 (Between) 0.93–0.96 0.93–0.96 0.93–0.96 ≈ floor
LLM Within-Subject ≫ floor (+0.08–0.09) ≫ floor ≫ floor signal
Textbook LDS-K 0.519 0.938 0.934
Simulation Baseline 0.646 0.655 0.640 withdrawn*

F11 N=15 falsifies ΔLDS > 0 between-subject: human LDS-C ≈ split-half floor — between-subject designs cannot separate language from participant variance.
F12 Concept-level ΔLDS ≈ 0; LLM within-subject margin (+0.08–0.09) shows the signal exists — earlier simulation comparison withdrawn.

Metric: LDS F11, F12 N=15 participants, 300 simulation (exploratory)

Interactive CognitiveSpace

Explore the mathematics subgraph in an interactive 3D knowledge graph — 556 nodes · 525 relations · 219 groups

Switch between zoom, rotate and fly modes. Explore the graph in three languages or switch discipline. Compare how different education systems structure the same mathematical knowledge.

Open Full Screen Gallery 556 nodes · 525 relations · 219 groups
Gallery Math 3D Physics 3D CDS Terrain Margin Proof Coverage Towers

Limitations

Transparent assessment of methodological boundaries and potential biases

Textbook Selection Bias

The corpus is not a comprehensive sample of all textbooks in each language. Chinese textbooks are predominantly one series (Renjiao), which may not represent the full diversity of Chinese mathematics education. German textbooks are weighted toward senior-level university texts (Forster, Fischer) with fewer primary-level resources.

Mitigation: Multi-source triangulation across 68 math textbooks (180 total, all subjects)

Limited Discipline Coverage

Only three STEM disciplines (mathematics, physics, chemistry) were analyzed. The framework has not been validated on humanities, social sciences, or professional fields. The universal claims (e.g., "knowledge density peaks early") are therefore limited to STEM education.

Mitigation: Framework is design-agnostic and extendable

Curriculum Granularity Mismatch

Curriculum documents vary enormously in granularity: the US math curriculum contains 2,124 concepts while the UK science curriculum has 397 (paper §8: NRW 299 · UK 397 · US 2,124 · CN 87). Higher granularity mechanically lowers coverage scores when matched against the same textbook graph. This is partially mitigated by normalizing per-curriculum, but cross-system comparisons remain approximate.

Mitigation: Analysis focuses on within-system trajectories, not absolute values

LLM Extraction Reliability

Concept extraction depends on a single production model. The 19-model benchmark (F1 range 0.55–0.67) confirms extraction consistency across architectures. While F1 = 0.939 on social concepts, the math domain achieves only 0.674 overall (German as low as 0.506). Extraction errors propagate through all downstream metrics. Relation extraction (prerequisite links) is less reliable than concept identification.

Mitigation: 92 gold labels ensure measurable quality bounds

Language Coverage Gap

Only three languages (ZH, EN, DE) were analyzed. These represent only two language families (Sino-Tibetan, Germanic) and one educational tradition cluster (Western Europe + East Asia). The framework's applicability to Romance, Slavic, Arabic, or other language families remains untested.

Mitigation: Existing three-language comparison already reveals systematic patterns

Gold Label Sample Size

92 gold labels, while sufficient for overall F1 estimation (especially for social concepts, n=72), are limited for fine-grained per-language, per-domain analysis. The mathematical domain has only 20 labels across 3 languages, and the German math F1 estimate (0.506) has high variance.

Mitigation: Social concept results (n=72) are statistically robust

Curriculum Analysis

How four education systems align their curricula with textbook content

🇩🇪 NRW (Germany)

12.7%

Coverage decreases toward upper-secondary specialization. Teachers exercise curricular freedom — they select from the curriculum rather than covering all of it. Reflects a specialization-driven philosophy.

Stages: Sek I → Sek II · Overall 12.7% (paper §8.5)

🇬🇧 UK National Curriculum

37.3%

The exam-driven system drives comprehensive textbook coverage — every curriculum point must be taught because it might be tested. Moderate overall alignment (37.3%, paper §8.5).

Stages: Key Stage 1–4 · Overall 37.3% (paper §8.5)

🇺🇸 US (CCSS + NGSS)

17.2%

Low overall coverage with stable trajectory across stages. The US standards define broad learning goals rather than detailed syllabi, giving textbook authors interpretive freedom while maintaining curriculum alignment.

Stable across stages · Standards: Common Core + NGSS · Overall 17.2% (paper §8.5)

🇨🇳 China National Curriculum

95.4%

Near-perfect alignment with a centralized national curriculum. Chinese textbooks closely follow the national syllabus, resulting in the highest coverage across all four systems. The highly selective, curriculum-driven textbook market aligns textbooks tightly with government standards.

Stable · centralized curriculum alignment · Overall 95.4% (paper §8.5)
Details · how the coverage cards are made

Source. Per-stage rows of config/expert_graphs/coverage_all_curricula.json — the same table behind Coverage Towers.

Method. Keyword bridge + 2-overlap threshold (≥2 shared keywords = covered); card headline = system overall.

Reading. Each card narrates one system; the numbers match the bright overall bars in Coverage Towers.

Three Competing Explanations

× A: Curriculum Granularity

Claims: Dense curricula produce higher coverage. Evidence: Counter-indicated — the US has the densest curriculum (2,124 concepts) but one of the lowest coverages (17.2%).

B: Educational Philosophy Hypothesis (paper §8.6)

Coverage patterns reflect educational governance: CN centralized (95.4%), UK national curriculum (37.3%), US broad standards (17.2%), NRW per-track specifications (12.7%).

! C: Division of Labor

Some concepts may be taught in other subjects (e.g., physics in chemistry). Plausible but unverified — cross-subject concept mapping could reveal systematic transfers.

Validation

Gold standard annotations and model benchmarks

Gold Dataset: 92 Labels

Human-annotated gold standard covering social concepts and mathematical concepts across all three languages. Validated against production-model extraction.

DomainZH F1DE F1EN F1Overalln
Social Concepts0.9740.9490.8820.93972
Mathematics0.8570.5060.7110.67420
All (weighted)0.9510.8420.8440.88192

Note: The low German math F1 (0.506) is a domain mismatch — the math gold labels use Chinese/English mathematical terminology not present in the German textbook corpus. Social concept extraction is uniformly strong across all three languages.

Model Benchmark: 19 Models

All free-quota Bailian API models tested on 92 gold labels. The production extraction model was selected by overall F1.

Production model: extraction model (F1=0.666) · Best overall: hy3-preview (0.674) · Best DeepSeek: v4-flash (0.608) · 19 models tested, F1 range 0.55–0.67
Details · how the benchmark chart is made

Source. 19 models scored on 92 gold labels (model-selection benchmark, Validation section).

Method. F1 on gold labels; the production extraction model was selected by overall F1. Range 0.55–0.67.

Reading. Bar = F1 per model. Used only for extractor selection — never for language claims (that is the 55-run replication below).

Replication roster: 55 full runs + 31 collecting (file truth 56/51/168 incl. qwen-max n=29/30)

CN run   western run (7)

Fused view — all 55 full-run ZH-DE margins (LDS-C minus split-half floor) under one P1 protocol (ZH/DE/EN × 10); all 55 ZH-DE p<0.01; 8 EN n.s. among 165 (500 permutations). Western runs in red. The 31 collecting runs (incl. qwen-max n=29/30, per paper §8.15) withhold margins until n=30. Source: data/lds_c/llm_subject/multi_model_replication_20260910.json + paper §5.10.

Details · how the margin chart is made

Source. 55 full runs from data/lds_c/llm_subject/multi_model_replication_20260910.json — the same rows as the Margin Proof view and the roster table below. 31 collecting runs (n<30) withhold margins.

Method. Margin = LDS-C minus split-half floor, one P1 protocol for all runs; 500 permutations, all 55 ZH-DE p<0.01; 8 EN n.s. among 165. Sorted descending; nothing rescaled.

Reading. Bar = margin, all 55 positive. Red = western (scattered = no origin clustering). The roster table below lists every row.

Full 55-row margin table (collecting runs listed below)
ModelOriginZH-DE marginSignif.
command-a-03-2025West · Cohere · CA+0.42p<0.01
gpt-5.6-lunaWest · origin undisclosed+0.22p<0.01
nemotron-3-super-120b-a12bWest · NVIDIA · US+0.17p<0.01
nemotron-3-ultraWest · NVIDIA · US+0.16p<0.01
gpt-oss-20bWest · OpenAI-weights · US+0.13p<0.01
laguna-s-2.1 · directWest · Poolside · US+0.10p<0.01
laguna-s-2.1 · kiloWest · Poolside · US+0.07p<0.01
deepseek-v3.1CN+0.29p<0.01
qwen-turboCN+0.25p<0.01
qwen3-maxCN+0.24p<0.01
deepseek-v3CN+0.24p<0.01
qwen3.8-maxCN+0.21p<0.01
glm-5CN+0.19p<0.01
qwen3.5-397b-a17bCN+0.18p<0.01
qwen3.7-maxCN+0.18p<0.01
glm-5.2-fast-previewCN+0.18p<0.01
qwen3.5-plusCN+0.18p<0.01
deepseek-v3.2CN+0.17p<0.01
deepseek-v4-flash · dashscopeCN+0.16p<0.01
MiniMax-M2.5CN+0.16p<0.01
qwen3.6-plusCN+0.16p<0.01
glm-5.2 · directCN+0.16p<0.01
qwen3.7-plusCN+0.15p<0.01
glm-4.6CN+0.15p<0.01
glm-5.1CN+0.14p<0.01
longcat-2.0CN+0.14p<0.01
kimi-k2.6 · directCN+0.14p<0.01
glm-4.7CN+0.13p<0.01
kimi-k2.6 · dashscopeCN+0.13p<0.01
MiniMax-M2.1CN+0.13p<0.01
kimi-k2.5CN+0.12p<0.01
kimi-k2.7-codeCN+0.12p<0.01
qwen3.5-27bCN+0.12p<0.01
glm-5.2 · dashscopeCN+0.12p<0.01
qwen3.6-27bCN+0.11p<0.01
deepseek-v4-pro · dashscopeCN+0.11p<0.01
qwen-flashCN+0.11p<0.01
Kimi-K2-InstructCN+0.10p<0.01
deepseek-r1-distill-qwen-14bCN+0.10p<0.01
qwen-plusCN+0.10p<0.01
qwen3.5-35b-a3bCN+0.10p<0.01
qwen3.5-flashCN+0.09p<0.01
qwen3.5-122b-a10bCN+0.09p<0.01
kimi-k2-thinkingCN+0.09p<0.01
qwen3.6-35b-a3bCN+0.08p<0.01
deepseek-v4-flash · directCN+0.08p<0.01
qwen3-235b-a22bCN+0.08p<0.01
mimo-v2.5CN+0.08p<0.01
deepseek-v4-pro · directCN+0.07p<0.01
qwen3.7-flashCN+0.07p<0.01
deepseek-r1-distill-qwen-32bCN+0.07p<0.01
qwen3.6-flashCN+0.05p<0.01
deepseek-r1CN+0.05p<0.01
deepseek-r1-distill-qwen-7bCN+0.04p<0.01
deepseek-r1-0528CN+0.03p<0.01
ModelOriginProgressStatus
Collecting — incomplete runs (n<30, margin withheld):
qwen-maxrepon=29/30collecting
gpt-oss-20brepon=20/30collecting
grok-4.6repon=18/30collecting
llama-3.3-70b-instruct-fp8-fastrepon=15/30collecting
ling-3.0-tinyrepon=12/30collecting
nemotron-3-ultra-550b-a55brepon=11/30collecting
mistral-medium-latestrepon=10/30collecting
nemotron-3-nano-30b-a3brepon=6/30collecting
laguna-s-2.1repon=5/30collecting
north-mini-coderepon=3/30collecting
glm-4.5repon=3/30collecting
glm-4.5-airrepon=3/30collecting
qwen3-14brepon=3/30collecting
qwen3-30b-a3brepon=3/30collecting
qwen3-32brepon=3/30collecting
qwen3-8brepon=3/30collecting
qwq-plusrepon=3/30collecting
gemma-4-26b-a4b-itrepon=3/30collecting
gemma-4-31b-itrepon=3/30collecting
ling-3.0-tinyrepon=3/30collecting
nemotron-3-super-120b-a12brepon=3/30collecting
nemotron-3-ultra-550b-a55brepon=3/30collecting
laguna-xs-2.1repon=3/30collecting
gemini-3.8-flashrepon=2/30collecting
mistral-largerepon=2/30collecting
gemma-4-31b-itrepon=2/30collecting
mistral-small-3.1-24b-instructrepon=1/30collecting
mistral-small-latestrepon=1/30collecting
nemotron-3-ultra-550b-a55brepon=1/30collecting
gpt-oss-120brepon=1/30collecting
muse-spark-1.3-contributorrepon=1/30collecting

Scope note: F1-benchmark (19 models, gold-label model selection) ≠ replication (55 runs, language signal) ≠ collecting (31 incomplete incl. qwen-max n=29/30, margin withheld). Families per paper §5.10: Qwen 20 · DeepSeek 12 · GLM 7 · Kimi 6 · MiniMax 2 · mimo · laguna · longcat · nemotron · gpt-oss · command-a · luna.

Extended replication (2026-09, paper §5.10): 55 measurements / 50 models — all 55 ZH-DE pairs significant (margin +0.03…+0.42); 8 EN-pair tests n.s.

Paper

Preprint manuscript available in the repository

Abstract

LinguaGraph asks whether language shapes the structural organization of knowledge. Two pipelines — textbook graphs (556 concepts, 525 relations, 219 groups) and LLM-as-subject (55 runs, 50 models) — isolate the language signal ΔLDS: N=15 falsifies it between-subject; within-subject shows it clearly (paper §5).

Sections

Three Conclusions · Introduction · Related Work · Methodology · Results · Discussion · Conclusion · Physics Results · LPA Analysis

9 files · ~161 KB md · 217 KB PDF · Read Methodology

Cite this work

@misc{linguaGraph2026,
  author = {Rong, Jiajun and Lan, Zhenxi},
  title = {LinguaGraph: Cross-Lingual Knowledge Structure Analysis Framework},
  year = {2026},
  publisher = {GitHub},
  journal = {BWKI 2026 — Bundeswettbewerb Künstliche Intelligenz},
  url = {https://github.com/jjjjjjjjnnjnn/BWKI-2026-LinguaGraph}
}

Open Science

All data, code and results are publicly available

Source Code

Complete pipeline scripts, extraction code, benchmark runners and evaluation tools.

Python · Bash · SQL

Datasets

All knowledge graphs, curriculum alignments, gold annotations and benchmark results.

1140+ concepts · 1100+ direct relations · 92 gold labels

Paper

Full manuscript with 3 research questions, 4 metrics, 12 findings and error analysis.

9 files · 54 references

Figures

8 publication-ready figures: CDS, HDS, LDS, null-model and coverage distributions.

14 figures · PNG · ~1.7 MB total (EN/DE/ZH; fig5 EN-only)

Benchmarks

19-model comparison across Bailian API. Full evaluation results with F1 scores per language and domain.

19 models · 92 gold labels · Full ranking

Contact

Questions, collaboration, or feedback? Open an issue on GitHub.

GitHub Issues