LSPHIL › Learning & teaching languages › How Many Words to Read a Newspaper? The Answer Depends on What You Count
How Many Words to Read a Newspaper? The Answer Depends on What You Count
The percentage of words in a text that a reader actually knows—what researchers call "lexical coverage"—is the number that determines whether reading flows or stutters. But ask "how many words do you need?" and the answer collapses into confusion unless you first specify the unit of counting. A word family bundles "run," "runs," "running," and "runner" together. A lemma keeps inflections but splits "running" the adjective from "runner" the noun. Word forms treat every variant as separate. The same vocabulary "size" can mean three different things.
English · 807 words
Why the number changes with the counting unit
Word families, the unit favoured by vocabulary researcher a leading vocabulary researcher, count a base form plus its regular inflections and transparent derivations as one item. "Run," "runs," "running," "runner," and "runners" live under a single roof. Lemmas are smaller: they group a headword with its inflections but typically distinguish parts of speech. "Running" as a gerund belongs with "run"; "running" as an adjective in "running water" may sit elsewhere. Word forms count every distinct spelling separately. The practical spread is enormous. A reader who knows 4,000 word families might command 6,000–8,000 lemmas or 12,000–15,000 word forms. When studies cite "8,000–9,000 word families for reading," they mean roughly 12,000–15,000 lemmas or double that in surface forms. The headline number is meaningless without its unit.

The coverage curve
Lexical coverage rises fast, then nearly plateaus. The most common 2,000 word families in the British National Corpus cover about 83% of running words. Double that to 4,000 word families, add proper nouns, and you hit roughly 95% coverage. But the final climb to 98% requires another 4,000–5,000 families—8,000–9,000 in total. The marginal return collapses. Each new thousand words buys less coverage than the last. Nation's calculations from a mini-corpus of five English novels trace this curve precisely: 4,000 families plus proper nouns for 95%, 8,000–9,000 for 98%. The gap between those percentages is not decorative. It marks the difference between tolerable friction and fluent comprehension.
What newspaper corpora show
Newspapers have their own frequency profile. In a 579,849-word newspaper corpus, 588 specialised "newspaper word families" accounted for 6.8% of running words—vocabulary absent from general lists. Combine the first 2,000 words of the General Service List with those 588 newspaper families, and coverage reached 86.5%. Add proper names, 2,521 GSL families, and the full newspaper list, and the figure rose to 92.5%. These figures sit below the 95% and 98% benchmarks because newspaper English is more restricted than literary fiction. The specialised vocabulary is narrow but dense. A reader prepared for novels may still stumble on institutional names, stock tickers, and regional references.
Why 98% matters more than 95%
Nation's work, replicated across corpora, treats 98% coverage as the threshold for unassisted reading. At 95%, unknown words arrive often enough to break rhythm; readers guess, skip, or stall. The final three percentage points of coverage are disproportionately expensive in vocabulary terms—thousands of additional families for fractional gain—but they carry the words that glue sentences together: low-frequency connectors, precise verbs, technical nouns. Frequency lists reliably deliver the high-coverage bulk. They do not reliably deliver the tail. A learner who stops at 95% has acquired the common words and missed the specific ones that distinguish "comprehensible" from "comfortable."
The counting problem behind the headline number
The question "how many words to read a newspaper?" presumes a stable answer. The research offers a conditional one. If you count word families, 8,000–9,000 is the cited range for general written English, with newspapers potentially lower due to formulaic structure. If you count lemmas, adjust upward. If you count word forms, double again. The figures also shift with corpus choice: the BNC, the newspaper corpus, or Nation's novel sample each weight differently. Proper nouns matter—newspapers are saturated with names that frequency lists ignore. The 98% threshold itself moves depending on whether "proper nouns" are included in the count or treated as given.
What this does not prove
The sources do not provide a single primary table mapping coverage to vocabulary size across all three counting units for newspaper text specifically. The steep-then-flat curve is inferred from scattered figures, not quoted verbatim from a unified study. The claim that the final percent "breaks comprehension" is a synthesis of the 98% threshold literature, not a direct source statement. And while word families, lemmas, and forms are defined in the research, no single authoritative source compares all three units side by side. The evidence supports the shape of the argument—coverage's diminishing returns, the unit problem, the 98% threshold—but readers should not treat any single figure as settled across all conditions.
The same vocabulary size can mean three different things. Until you know whether the count is in families, lemmas, or forms, and whether proper nouns are included, and which corpus supplied the frequency list, the number on the page is not information. It is decoration.
The register
What the rail used to do: newest, more on this desk, other languages.
- What You Can Actually Do at B1 Level (and What the CEFR Never Promises)English · 1074 words
- Why Spaced Repetition Stops Working: The Interval Is Not the ProblemEnglish · 724 words
- Mother-Tongue Instruction Is Not a Slogan: The DepEd Orders That Actually Define Philippine MTB-MLEEnglish · 1072 words
- Shadowing, Dictation and Chorusing: Three Drills for Three Different GapsEnglish · 766 words
In other languages
- Cómo se escriben los nombres extranjeros en español: transferencia, adaptación o tradiciónEspañol · 958 words
- Warum das Deutsche Substantive großschreibt – und was das im Satz leistetDeutsch · 874 words
- Hvorfor dansk staves anderledes end det udtalesDansk · 641 words
- Casino bonussenNederlands · 782 words
Newest on the site
- How Baybayin Letters Work: A Primer on the Philippines' Abugida ScriptEnglish · 1004 words
- How to Classify Any Script in Two Questions: What One Sign Means, and What Happens When the Vowel ChangesEnglish · 915 words
- Do Accent Marks Change Meaning? It Depends on the Language's RulesEnglish · 1009 words
- Why Spelling Reforms Succeed in Some Languages and Collapse in OthersEnglish · 987 words