Original data study
How many extra Chinese words does each HSK level actually buy you, in everyday spoken coverage?
How this was collected
We take this site's own shipped word-frequency coverage dataset (every HSK word matched against the OpenSubtitles 2018 Chinese word-frequency corpus, CC BY-SA 4.0) and compute, for every cumulative HSK level in both the six-level HSK 2.0 scheme and the nine-level HSK 3.0 scheme, how many new words that level adds and how many percentage points of spoken-register coverage those new words buy. Dividing new words by the coverage gained gives a words-per-point figure: a direct, computed measure of diminishing returns as a learner climbs the ladder. Every row's coverage and word counts come straight from coverage.json; only the marginal-cost division is computed on this page.
This dataset was computed .
What the data shows
Levels analyzed
13
Cheapest level
4 words/pt
Median cost
156 words/pt
Priciest level
1,557 words/pt
Average cost
352 words/pt
HSK 2.0 Level 1 is the best value on the ladder: 150 new words buy 37.4 points of coverage, about 4 words per point. HSK 3.0 Level 7-9 is the steepest: 5,606 new words buy only 3.6 points, about 1,557 words per point, over 388x the cost of the cheapest level. Diminishing returns are steep and mostly monotonic within each scheme.
Full data table
| Item | Cost | Source |
|---|---|---|
| HSK 3.0 Level 1 506 new words, cumulative 506 words, 50.4% coverage (+50.4pt) | 10 words/pt | coverage.json (HSK words vs OpenSubtitles 2018 corpus) |
| HSK 3.0 Level 2 750 new words, cumulative 1,256 words, 60.7% coverage (+10.3pt) | 73 words/pt | coverage.json (HSK words vs OpenSubtitles 2018 corpus) |
| HSK 3.0 Level 3 953 new words, cumulative 2,209 words, 66.8% coverage (+6.1pt) | 156 words/pt | coverage.json (HSK words vs OpenSubtitles 2018 corpus) |
| HSK 3.0 Level 4 972 new words, cumulative 3,181 words, 69.8% coverage (+3.0pt) | 324 words/pt | coverage.json (HSK words vs OpenSubtitles 2018 corpus) |
| HSK 3.0 Level 5 1,059 new words, cumulative 4,240 words, 72% coverage (+2.2pt) | 481 words/pt | coverage.json (HSK words vs OpenSubtitles 2018 corpus) |
| HSK 3.0 Level 6 1,123 new words, cumulative 5,363 words, 73.9% coverage (+1.9pt) | 591 words/pt | coverage.json (HSK words vs OpenSubtitles 2018 corpus) |
| HSK 3.0 Level 7-9 5,606 new words, cumulative 10,969 words, 77.5% coverage (+3.6pt) | 1,557 words/pt | coverage.json (HSK words vs OpenSubtitles 2018 corpus) |
| HSK 2.0 Level 1 150 new words, cumulative 150 words, 37.4% coverage (+37.4pt) | 4 words/pt | coverage.json (HSK words vs OpenSubtitles 2018 corpus) |
| HSK 2.0 Level 2 147 new words, cumulative 297 words, 46.3% coverage (+8.9pt) | 17 words/pt | coverage.json (HSK words vs OpenSubtitles 2018 corpus) |
| HSK 2.0 Level 3 298 new words, cumulative 595 words, 52.5% coverage (+6.2pt) | 48 words/pt | coverage.json (HSK words vs OpenSubtitles 2018 corpus) |
| HSK 2.0 Level 4 598 new words, cumulative 1,193 words, 57.8% coverage (+5.3pt) | 113 words/pt | coverage.json (HSK words vs OpenSubtitles 2018 corpus) |
| HSK 2.0 Level 5 1,298 new words, cumulative 2,491 words, 61.6% coverage (+3.8pt) | 342 words/pt | coverage.json (HSK words vs OpenSubtitles 2018 corpus) |
| HSK 2.0 Level 6 2,500 new words, cumulative 4,991 words, 64.5% coverage (+2.9pt) | 862 words/pt | coverage.json (HSK words vs OpenSubtitles 2018 corpus) |
Limitations
- Token coverage is not comprehension: recognising a large share of the words spoken is not the same as understanding every sentence, and the least-frequent words are exactly the ones that carry new information.
- The corpus is film and TV subtitles (OpenSubtitles 2018), a standard, openly-licensed proxy for spoken Chinese, not a recording of spontaneous conversation.
- HSK 3.0 and HSK 2.0 are two different vocabulary lists (the 2021 restructure changed word counts and level boundaries), so a level number is not directly comparable across schemes, only the coverage and words-per-point figures are.
- A small share of HSK words (about 0.5% for HSK 3.0, 0.2% for HSK 2.0) did not appear in the corpus and are excluded from the token counts; see the methodology page for the full match-rate accounting.
Update history
- - HSK 3.0 Level 1 coverage and word counts computed from coverage.json
- - HSK 3.0 Level 2 coverage and word counts computed from coverage.json
- - HSK 3.0 Level 3 coverage and word counts computed from coverage.json
- - HSK 3.0 Level 4 coverage and word counts computed from coverage.json
- - HSK 3.0 Level 5 coverage and word counts computed from coverage.json
- - HSK 3.0 Level 6 coverage and word counts computed from coverage.json
- - HSK 3.0 Level 7-9 coverage and word counts computed from coverage.json
- - HSK 2.0 Level 1 coverage and word counts computed from coverage.json
- - HSK 2.0 Level 2 coverage and word counts computed from coverage.json
- - HSK 2.0 Level 3 coverage and word counts computed from coverage.json
- - HSK 2.0 Level 4 coverage and word counts computed from coverage.json
- - HSK 2.0 Level 5 coverage and word counts computed from coverage.json
- - HSK 2.0 Level 6 coverage and word counts computed from coverage.json
Read the methodology behind this data, or see the full per-level coverage explorer this study summarizes.