Beeblio — read your way fluent · Home · Blog · Pricing · Start reading

Why 1,500 Words Cover About 80% of Everyday Text

By Beeblio ·

Pick up any Spanish novel, newspaper, or podcast transcript and count the words. You will almost certainly find that a small set of familiar words — que, ser, estar, de, no, un, en — appears again and again, while thousands of others appear once or twice and then vanish. This is not a quirk of Spanish or of any particular author: it is a mathematical law that governs every natural language ever studied. That law — and what it means for your study schedule — is the subject of this article. Understanding frequency vocabulary is not a motivational pep talk; it is the single most important piece of applied mathematics for language learners, and the numbers tell a clearer story than any textbook introduction ever will.

The Law Behind the Curve

In the 1930s and 1940s, the American linguist George Kingsley Zipf began counting words in large bodies of text and noticed something striking. Zipf's law — a fundamental statistical property first observed by J.B. Estoup and later systematically popularized by Zipf himself — is the power-law distribution of word frequencies in texts: the frequency of a word is inversely proportional to its rank, where rank is simply its position on a list sorted by decreasing frequency. In plain terms: the most frequent word occurs approximately twice as often as the second most frequent word, three times as often as the third, and so on.

When plotted on a log-log scale, the rank-frequency distribution of words in a language typically forms a straight line, indicating a power-law distribution. What makes this breathtaking for learners is the shape of the curve before the line straightens. The very top of the distribution is almost vertical: a tiny handful of words soaks up an enormous share of all the tokens (running words) you will ever encounter. Zipf's law suggests that focusing on a limited set of high-frequency vocabulary can be more efficient for language learners — and the corpus data bear that out with striking precision.

Crucially, Zipf's law is not an English-only phenomenon. Piecewise fitting of Zipf's curves to power-law functions in 50 languages shows R² values consistently above 0.96, with the middle segment of the curve covering the words ranked between 200 and 2,000 (or 3,000 in some languages) being especially smooth. Spanish, French, Mandarin, Finnish — the shape is universal. The research highlights Zipf's universal applicability across different types of language, which can inform teaching methodologies and vocabulary acquisition strategies.

What the Coverage Numbers Actually Say

Zipf's law predicts steep diminishing returns as you move down the frequency list. The corpus research confirms exactly that. The landmark figures on vocabulary size and text coverage come from Francis and Kucera's (1982) Brown Corpus — a diverse corpus of over 1,000,000 running words from 500 texts; because the corpus is highly diverse, the high-frequency words cover slightly less of the text, making these figures a conservative estimate.

Here is what the research shows for written English, band by band:

Frequency band (word families known) Cumulative text coverage Coverage added by this band
Top 1,000 ~78–81% ~78–81%
Top 2,000 ~86–90% ~8–9%
Top 3,000 ~90–95% ~3–5%
Top 5,000 ~95–96% ~1–3%
Top 8,000–9,000 ~98% ~2–3%

Vocabulary researcher Nation (2006) used British National Corpus data to measure how much each 1,000-word band covers written text: the first band (the most frequent 1,000 words) gives approximately 78–81%; the second band (words 1,001–2,000) adds approximately 8–9%; and the third band (2,001–3,000) adds approximately 3–5%, for a total of roughly 87–90% at 2,000 words and 90–95% at 3,000 words.

For Spanish specifically, coverage figures from comparable corpora are similar. Spanish data shows the top 1,000 words covering approximately 81.0% of text, the top 2,000 words covering 86.6%, and the top 5,000 words covering 92.5%. The shape of the curve is the same: enormous gains early, grinding diminishing returns later.

The headline number in the article title — "1,500 words cover about 80% of everyday text" — sits right in the steep part of this curve, interpolating between the 1,000-word and 2,000-word band figures. At 1,500 words you are past the first plateau but still in the richly rewarding zone. Here is what that looks like in real text. The following passage is a short Spanish excerpt at general everyday register:

Hoy es un buen día. Voy a la ciudad con mi amigo y después comemos en casa. Por la noche leo un libro.

Measured with Beeblio: a learner who knows the 1,500 most common Spanish words already knows 96% of the words in this passage (23 words) — just right. The new words: leo.

A learner at the 1,500-word band recognises roughly 96% of the tokens in that passage — enough for solid contextual reading. Notice what that means concretely: at 96% coverage, approximately one word in every 25 is unfamiliar, which is a workable density for reading with occasional guessing. The remaining 4% of tokens represent opportunities to pick up new vocabulary, not walls that block comprehension.

Why 80% Is Nowhere Near Enough

A common misconception is that 80% text coverage is "pretty good." It is not. In a text with 98% coverage, one out of 50 words is unknown to a reader. In a text with 80% coverage, one out of five words is unfamiliar. Reading with one unknown word in every five is exhausting, slow, and largely unproductive for incidental vocabulary acquisition, because context becomes too sparse to guess reliably.

The landmark empirical work here comes from Hu & Nation (2000). The comparison of comprehension scores with unknown-word densities showed a seemingly linear increase in comprehension as coverage of known words increased; Hu and Nation concluded that 80% coverage would lead to insufficient comprehension and that even knowing 95% of words in a text would generally not allow most learners "good comprehension." Hu and Nation suggest 98% as the probabilistic coverage at which most learners can read, and 95% as the coverage at which minimally acceptable comprehension can occur.

Seminal work by Laufer (1989) determined that 95% was the probabilistic lexical threshold above which adequate achievement would be reached by L2 learners on a reading test. A 2011 large-scale replication by Schmitt, Jiang, and Grabe — 661 participants from 8 countries completed a vocabulary measure based on words drawn from two texts and then completed a reading comprehension test; the results revealed a relatively linear relationship between percentage of vocabulary known and degree of reading comprehension, with no sharp threshold, but the 98% estimate was the more reasonable coverage target for readers of academic texts.

The practical upshot:

The Curve Is the Learner's Best Friend

Here is the optimistic reading of those numbers. The very steepness of Zipf's curve that makes the long tail so hard to conquer is also what makes the early stages so rewarding. Zipf's law reminds us that we shouldn't treat all vocabulary as equal. Some words are exponentially more useful than others — not just for comprehension, but also for production, task success, and confidence building. Prioritising high-frequency vocabulary early on gives learners the best chance of gaining access to comprehensible input and generating meaningful output from the start.

Think about what 1,000–2,000 words buys you. Knowing about 2,000 word families gives near to 80% coverage of written text and even greater coverage of informal spoken text. That is extraordinary leverage: 2,000 words — learnable in a year of consistent study — covers the lion's share of everything a language produces. By comparison, the jump from 8,000 to 9,000 words yields perhaps a single additional percentage point of coverage.

This is the learner's asymmetric advantage. The early frequency bands cost roughly the same effort per word as the later ones — arguably less, because high-frequency words appear so often that incidental reinforcement happens constantly — but they return vastly more coverage per unit of investment. A handful of high-frequency words do the heavy lifting, while thousands of others live in the shadows of occasional use.

The Mid-Frequency Gap

After the first 3,000 or so words, the curve flattens sharply and a problem emerges that researchers call the mid-frequency gap. By the third frequency band, even the entire 1,000-word band only increases coverage by 3–5%. Moreover, because the frequency of encountering words in this band decreases, it becomes difficult to learn them incidentally during natural reading or conversation. There is an estimate that while 200,000 word tokens of reading per year are needed to master the second band, nearly 3 million word tokens per year are required at the ninth-band level.

This is where intentional study — flashcards, deliberate review, spaced repetition — earns its keep. Low-frequency words in the long tail need intentional retrieval practice; spaced repetition systems (SRS) can help keep these words alive once learners have built a strong core lexicon. The spaced repetition system is founded on the theory that when faced with a large number of items to learn, we are far more likely to commit them permanently to memory if we review them not only repeatedly but at increasingly spaced intervals — as opposed to cramming, which is not very effective for long-term acquisition.

What Coverage Means for Reading Strategy

The coverage thresholds connect directly to a practical reading strategy endorsed by extensive reading researchers. Krashen's Input Hypothesis (1982) argues that comprehensible input is the sufficient condition for L2 acquisition, and his Reading Hypothesis postulates the facilitative effect of extensive reading on abilities such as reading comprehension, writing style, vocabulary, grammar, and spelling. Comprehension of text at the right level moves the learner's current competence from level 'i' to the next level 'i+1'; this is possible because the small proportion of problematic parts is resolved by resorting to cues from the extended text context and by activating the learner's world knowledge.

Day and Bamford (1998), in their foundational work on extensive reading, emphasised that the principle objective of an extensive reading plan is to give learners a circumstance to appreciate reading a foreign language and new real messages quietly at their own velocity and with satisfactory comprehension. "Satisfactory comprehension" maps directly onto the coverage thresholds: you want texts where you already know 95–98% of the words, not texts where you are drowning.

The implication is clear and counter-intuitive to many learners: reading slightly below your ceiling — texts you can almost fully understand — produces better acquisition outcomes than struggling with texts far above your level. One of the pedagogical implications of studies of lexical coverage has been that when L2 learners have the lexical coverage necessary to understand a text, they may have greater success at incidentally learning any unknown words that are encountered in that text.

"The top 1,000 words in any language should get you about 80% text coverage." — A useful rule of thumb that has held up across languages studied, as noted by corpus researchers following Nation's work on frequency bands.

Building Your Known-Words Count Strategically

Given all the above, a rational study strategy falls out naturally:

  1. Lock down the first 1,000–2,000 words before anything else. These give you ~80–90% coverage and are the foundation on which all reading comprehension rests. Curricula can be designed to prioritise frequently used words to enhance comprehension and retention.
  2. Read at your coverage level, not above it. Texts pitched at 95–98% known words are where reading fluency and incidental acquisition both happen. Choose materials accordingly — graded readers, frequency-levelled articles, or Beeblio-style measured passages.
  3. Use spaced repetition for the mid-frequency vocabulary. Spaced repetition is well suited to contexts in which a learner must acquire many items and retain them indefinitely in memory — it is therefore well suited for the problem of vocabulary acquisition in the course of second-language learning. Once natural reading encounters drop off (as they do past the 3,000-word band), SRS fills the gap.
  4. Track your known-words count. Because coverage is a direct function of how many high-frequency words you know, your word count is a genuine progress metric — not a vanity number. Every word added in the 1,000–5,000 range moves your coverage needle measurably.
  5. Set a realistic long-term goal. If 98% coverage of a text is needed for unassisted comprehension, an 8,000 to 9,000 word-family vocabulary is needed for written text. That is a multi-year target, but the curve means most of the benefit arrives long before you get there.

A Note on Honest Caveats

The coverage figures quoted above are based primarily on English corpora (the Brown Corpus, the BNC), though Spanish and French data point in the same direction. Exact percentages shift slightly depending on genre — newspaper text, fiction, and academic prose all have different vocabulary profiles. There are fairly stable figures for coverage within a genre such as newspapers, novels, or in the planned corpora such as LOB and Brown. The 95–98% comprehension threshold is robust across multiple studies, but recent work urges nuance: a recent partial replication of Hu and Nation's (2000) study by Kremmel et al. (2023) found that the earlier results could not be fully supported and they suggested caution in applying lexical coverage recommendations. Coverage is necessary but not sufficient — background knowledge, syntax, discourse structure, and reading purpose all play roles that pure vocabulary counts do not capture.

FAQ

Does the "1,500 words = ~80% coverage" figure apply to Spanish specifically?

Spanish corpus data shows the top 1,000 words covering approximately 81% and the top 2,000 words covering 86.6% of text. The 1,500-word estimate of roughly 80% coverage is therefore a reasonable interpolation for Spanish. The exact figure varies by genre and corpus, but the order of magnitude is solid.

Is 95% coverage really enough, or do I need 98%?

Both thresholds have research support and slightly different meanings. Research has concluded that 95% lexical coverage is the minimum percentage to understand general information in a text, but 98% is necessary for unassisted reading and detailed comprehension. For pleasurable independent reading, aim for 98%. For reading with occasional dictionary use, 95% is workable.

How many words do I need to reach 98% coverage in Spanish?

If 98% coverage is needed for unassisted comprehension, an 8,000 to 9,000 word-family vocabulary is needed for written text in English, and Spanish requires a similar range. However, you get most of the practical benefit — enough to read with light support — at around 3,000–5,000 word families.

Why does each new frequency band add less and less coverage?

This is Zipf's law in action. The frequency of a word is inversely proportional to its rank: the word ranked 1st appears far more often than the word ranked 1,000th, and enormously more often than the word ranked 10,000th. Because the high-frequency words have already soaked up most of the token "space" in any text, there is simply less room left for lower-frequency words to contribute to your coverage percentage. Each additional 1,000-word band covers a smaller slice of a pie that is already mostly accounted for.

Should I skip reading and just do flashcards until I hit 3,000 words?

No — and the research is fairly clear on this. Spaced repetition is effective for memorizing discrete facts, but language learning also requires developing an understanding of collocations, semantics, and culture; two key uses of SRS are to learn basic vocabulary needed for graded readers, and to memorize rare words. Reading and SRS work best in tandem: reading provides context, collocation, and incidental reinforcement; SRS consolidates words that appear too rarely to stick on their own.

Sources

In Beeblio, every passage you read is measured against your personal frequency band, so you always know your current coverage percentage before you start reading. When you tap an unknown word and it becomes a spaced-repetition card, you are not just studying a vocabulary item in isolation — you are moving your known-words count upward along the steepest, most rewarding part of Zipf's curve, where each new word genuinely shifts how much of the language you can read without effort.