Search this question and you will be handed 171,476, then 470,000, then 600,000, then “over a million,” often on the same page, each stated as fact. None of those figures is a mistake. They are answers to four different questions, and almost nobody says which question they answered.
This guide gives you the current figures from primary sources, shows exactly which counting rule produces which number, converts them all into a single common unit so you can compare them for the first time, and flags the widely repeated statistics that fall apart when you open the source.
Quick Answer
There is no official count of English words, because “word” has no single definition. The most defensible 2026 figures are: about 171,476 words in current use (Oxford English Dictionary, second edition), over 500,000 entries and roughly 600,000 word forms documented in the OED across 1,000 years of English, about 470,000 entries in Webster’s Third New International, and roughly 1,022,000 distinct word forms found in printed books as of 2000 (Michel et al., Science). The number that affects you personally is far smaller: an average 20-year-old native speaker recognises about 42,000 lemmas, and 8,000 to 9,000 word families cover 98% of an average novel.
If you need one sentence you can defend anywhere:
There is no exact total. English has roughly one million words under the broadest counting rule, the Oxford English Dictionary documents about 600,000 word forms across the language’s history, and about 171,476 of those were classed as in current use in the OED’s second edition.
The rest of this article explains why those figures differ, and shows that they disagree far less than they appear to.

The Numbers
| What is being counted | Figure | Source and date |
|---|---|---|
| Words in current use (OED2) | 171,476 | Oxford University Press, 1989 print edition |
| Obsolete words (OED2) | 47,156 | Oxford University Press |
| Derivative subentries (OED2) | ~9,500 | Oxford University Press |
| Main entries, OED2 print | 291,500 | Oxford University Press |
| Word forms defined or illustrated, OED | ~616,500 | OUP entry statistics |
| Total OED entries today | over 500,000 | Oxford University Press, 2026 |
| Oxford’s current-English dictionary (powers Google) | 350,000+ words and phrases | Oxford Languages |
| Webster’s Third New International, Unabridged | ~470,000 entries | Merriam-Webster (1961 base text, 1993 addenda) |
| Collins English Dictionary, 14th ed. | 732,000 words, meanings and phrases | Collins, 2023 |
| Chambers, 12th ed. | 620,000 references | Chambers, 2011 |
| Distinct English word forms, Wiktionary | ~1,380,000 | July 2026 dump |
| Word forms in printed books, year 2000 | ~1,022,000 (from 1,489,337 raw strings) | Michel et al., Science |
| Merriam-Webster’s broad vocabulary estimate | ~1 million, possibly off by a quarter-million | Merriam-Webster FAQ |
| “One millionth word” milestone | Not a real measurement | Global Language Monitor, 2009 (discredited) |
| Lemmas recognised by an average 20-year-old | ~42,000 | Brysbaert et al., 2016 (n = 221,268) |
| Word families for 98% coverage of written text | 8,000 to 9,000 | Nation, 2006 |
| Spoken word tokens per day, recent estimate | 12,792 | Pfeifer & Mehl, 2026 |
Key Takeaways
- The question is definition-dependent, not discovery-dependent. You pick the counting rule before you get a number.
- Comparing two dictionaries compares editorial policy, not language size.
- The famous figures are far closer than they look once you convert them into the same unit. Translated into lemmas, the “million words” corpus figure and the OED’s entry count land in the same band.
- The two most-quoted large numbers, 1,000,000 and 1,022,000, are the two least reliable on this page.
- Vocabulary research is much sounder than lexicon-size research. If you want a number you can act on, take it from vocabulary research.
- Several widely circulated statistics are misquoted at source, including Nation’s coverage thresholds, the CEFR vocabulary table, Brysbaert’s multiword figure, and almost everything about Shakespeare.
Why Nobody Can Count the Words in English?
Before you can count something, you have to decide what one of it looks like. English makes that harder than it sounds.
James Murray, the OED’s founding editor, set out the problem in 1888 and nobody has improved on it since:
“The vocabulary of English-speaking people presents, to the mind that endeavors to grasp it as a definite whole, the aspect of one of those nebulous masses familiar to the astronomer, in which a clear and unmistakable nucleus shades off on all sides, through zones of decreasing brightness, to a dim marginal film that seems to end nowhere, but to lose itself imperceptibly in the surrounding darkness.”
Two questions have no universal answer, and every count has to settle them privately.
Where does English stop? You must decide whether to include regional varieties, slang, jargon, brand names, proper nouns, abbreviations, loanwords, archaisms, and technical nomenclature.
What counts as one word? You must decide whether inflections, derivations, compounds, separate meanings, and fixed phrases count together or separately.
Here are the seven decisions that move the total by tens or hundreds of thousands each.

1. Inflected forms: is run one word or four?
Run, runs, running, ran. One lexeme, four surface forms. Counting word forms gives four; counting lemmas gives one. English is lightly inflected, so this multiplies a count by roughly 1.7 times. In Finnish or Turkish it would multiply it by orders of magnitude.
2. Derived forms: is runner a new word?
Mouse yields mousy, mousier, mousiest, mouser, mousetrap, fieldmouse. You have probably never seen mousily, and you understood it instantly. Word families group all of these under one head. Lemma counts separate them.
3. Parts of speech: is dog the noun the same word as dog the verb?
Oxford’s own FAQ raises this question and declines to answer it. Most dictionaries split them into separate entries, which inflates entry counts relative to “how many words do people actually know.”
4. Senses: run has over 600 documented meanings
The OED’s revision of run took lexicographer Peter Gilliver around nine months and produced 645 senses. Is that one word or 645? Oxford has suggested that counting distinct senses instead of headwords would push the OED’s total toward three-quarters of a million, although OUP publishes no official sense total.
A frequently mangled trivia point: run did not directly overtake set. Set is usually said to have led in 1928, with put overtaking it before run overtook both.
5. Compounds and hyphens: is hot dog one word?
Hot dog, hot-dog, hotdog. One concept, three orthographies. A space-based counter reads ice cream as two words even though speakers store it as one unit. Multiword expressions are exactly where the largest word-count claims quietly inflate.
6. Scientific and technical vocabulary
Merriam-Webster’s FAQ notes that the roughly one million estimate “includes the myriad names of chemicals and other scientific entities.” Chemical nomenclature is generative: systematic names are produced by rule, not coined by speakers. If they count, the total is not large, it is unbounded. There are also over a million described insect species, each with a binomial name.
The late Berkeley linguist Geoffrey Nunberg put the absurdity well: “It’s unlikely that the vision of a language with 5000 fish names probably would have ignited envy in the heart of Garcia-Lorca or Kafka or Flaubert (well, okay, maybe in Flaubert it would have).”
7. Productive rules mean the set is open
English lets any speaker generate numeral expressions and transparent compounds on demand. If every well-formed combination counted, the inventory would never close. Lexicographers require evidence that a form has become a conventional unit, not merely proof that grammar could produce it.
What Counts as One Word?
Most word-count arguments are two people using the same noun for different objects. A reliable figure names its unit. Here are the eight units in circulation.

One sentence, three correct counts
She runs daily, but she ran yesterday.
| Method | Count | Why |
|---|---|---|
| Tokens | 7 | Every occurrence counted |
| Types | 6 | she appears twice |
| Lemmas | 5 | runs and ran group under run; she counted once |
Nothing about the sentence changed. Only the definition changed. Now apply that same instability to the entire language.
The hardest borderline cases
Inflected forms. Should talk, talks, talked, talking be four words or one lemma? Corpus software counts four types, then a lemmatiser groups them. Dictionaries file them under one entry unless a form develops independent meaning.
Derived words. Should happy, unhappy, happiness, happily be one family or four lexemes? A vocabulary study groups them. A dictionary may give several of them separate headwords.
Homonyms. One spelling can represent unrelated words (bat the animal, bat the implement). One lexeme can also develop dozens of related senses. Counting spellings, entries, and senses gives three different totals.
Proper names. People, places, products, organisations, and fictional characters form an enormous open set. General dictionaries exclude most of them; corpora record them all as tokens.
Regional English. English includes established varieties in Canada, Australia, India, Ireland, Nigeria, New Zealand, Singapore, South Africa, the Philippines, the Caribbean, and elsewhere. A dictionary that documents more regional evidence will hold more entries by design, not because the language got bigger.
Slang and coinages. A term does not become a word when a dictionary adds it. Dictionaries document evidence of established use; they do not issue permits. Equally, a one-off joke typed once is not part of the shared vocabulary.
The Conversion Table That Reconciles Everything
This is the tool missing from every other article on this question, and it does more work than any single figure.
English averages roughly 1.7 word forms per lemma and roughly 3.8 lemmas per word family. Those ratios come straight from the vocabulary research: Brysbaert and colleagues report the same average 20-year-old as knowing 42,000 lemmas and 11,100 word families, which is 3.78 lemmas per family.
That means one word family is worth about 3.8 lemmas and about 6.4 word forms. Apply the ratios and the famous figures stop looking like rivals.

Ratios are population averages, and the underlying sets are not identical, so treat these as order-of-magnitude translations rather than exact equivalences.
What the conversion reveals?
Look at the lemma column. The OED’s “over 500,000 entries” and the Google Books “1,022,000 words” appear to disagree by a factor of two. Converted into the same unit they are approximately 500,000 versus 601,000 lemmas. The remaining gap is well explained: the corpus includes proper nouns, product names, variant spellings, technical strings, and residual scanning noise, while the OED excludes most proper nouns and requires editorial review.
The same move dissolves the other apparent contradictions. Once everything is in lemma-equivalents, the picture is nested rather than contradictory:
- ~42,000 lemmas is what one educated adult recognises.
- ~171,000 lemmas is the slice the OED’s second edition classed as current.
- ~500,000 entries is the full historical record, current plus obsolete plus regional.
- ~601,000 lemmas is what turns up in a very large sample of printed books.
Each set contains the one above it. The numbers were never fighting. They were measuring different depths of the same body of water in different units.
Rule of thumb: before you compare any two word counts, convert both to lemmas. Roughly 90% of the disagreement on this topic disappears at that step.
The Word Count Ladder: Five Standards, Five Numbers
Most articles say “it depends” and stop. Here is what “it depends” actually produces. Pick your rung, get your number.

The total shifts by more than two orders of magnitude across the ladder, and every rung is legitimate. Arguments about “the real number” are nearly always two people standing on different rungs.
When you see a word count quoted anywhere, ask which rung it sits on. If the answer is not stated, the number is decoration.
Which Number Should You Use?
Match the number to the job. Each row gives you a sentence you can paste directly into your work.
| Your situation | Use this figure | Copy-paste sentence |
|---|---|---|
| School essay, article, or presentation needing one citable figure | 171,476 | “The Oxford English Dictionary’s second edition contained full entries for 171,476 words then in current use.” |
| Making the point that English is very large | 1,022,000 | “A study in Science analysing 5.2 million digitised books estimated 1,022,000 distinct English word forms in print as of 2000.” |
| Describing the historical depth of English | 600,000 | “The Oxford English Dictionary documents roughly 600,000 word forms across more than 1,000 years of English.” |
| Writing about American English specifically | ~470,000 | “Webster’s Third New International, Unabridged, contains about 470,000 entries.” |
| Discussing what people actually know | 42,000 | “An average 20-year-old native speaker of American English recognises about 42,000 lemmas (Brysbaert et al., 2016).” |
| Teaching or learning English | 8,000 to 9,000 | “Around 8,000 to 9,000 word families provide 98% coverage of unsimplified written English (Nation, 2006).” |
| Building an NLP system or search index | 1.0 to 1.4 million | “Attested English word forms number in the low millions, depending on inflection handling and proper-noun policy.” |
Decision tree

What the Dictionaries Say (2026 Figures)
| Dictionary | Published figure | What it counts | Notes |
|---|---|---|---|
| Oxford English Dictionary | 600,000 words; over 500,000 entries; 3.5m quotations | Word forms and entries | The two figures are different units, not a contradiction. See below |
| OED, 2nd ed. (1989) | 171,476 current, 47,156 obsolete, ~9,500 subentries, 291,500 main entries | Print headwords | The source of the internet’s favourite number, and a 1989 snapshot |
| Oxford Languages (current English) | 350,000+ words and phrases | Present-day entries | The dataset behind Google’s dictionary results |
| Webster’s Third New International | ~470,000 entries | Entries | Base text from 1961, 1993 addenda |
| Merriam-Webster’s Collegiate, 12th ed. | No total published | – | Shipped November 2025, first new edition in 22 years. Added 5,000+ words including rizz, dumbphone, doomscroll |
| Collins English Dictionary, 14th ed. | 732,000 words, meanings and phrases | A composite metric | Not comparable to an entry count |
| Chambers, 12th ed. (2011) | 620,000 references | References | The 13th edition publishes no total |
| Cambridge Dictionary | No total published | – | Publishes corpus size (1.5bn words), not entry counts |
| Oxford Advanced Learner’s, 11th ed. | 180,000+ words, phrases and meanings | Learner-focused | |
| English Wiktionary | ~1,380,000 English word forms | Word forms including inflections | Highest figure traceable to a checkable dataset (July 2026 dump) |
Resolving the OED’s 500,000 versus 600,000 puzzle
Oxford appears to publish two different totals, and this has confused readers for years. It is not an error. They are different units.
- ~600,000 is a word-form count. OUP’s entry statistics record roughly 616,500 word forms defined or illustrated across the dictionary, which includes bold-type derivatives (around 157,000) and italicised phrases and combinations (around 169,000) sitting inside entries.
- 500,000+ is an entry count. These are the editorial units: headwords with their own definition structure.
An entry can contain many word forms. Both figures are accurate descriptions of the same dictionary. Once you know which unit each represents, they stop competing.
The trap in dictionary comparisons
Collins reports 732,000 “words, meanings and phrases.” Merriam-Webster reports 470,000 entries. These cannot be ranked against each other. One counts entries; the other counts entries plus senses plus phrases. Almost every article on this topic ranks them anyway, and both Chambers and Collins separately claim to be the largest single-volume English dictionary.
Is 171,476 the number of words in modern English?
No. Treat it as a historical reference point.
It comes from an Oxford summary of full entries judged to be in current use in the OED’s second edition, published in 1989. It is usually paired with 47,156 obsolete entries and about 9,500 derivative subentries. It predates decades of new vocabulary and revised evidence, and it excludes or groups many inflections, compounds, derivatives, technical terms, regional terms, phrases, and senses.
If you quote it, label it:
✅ “The OED’s second edition, published in 1989, reported 171,476 full entries for words then considered in current use.”
❌ “English has exactly 171,476 words.”
The “One Million Words” Claim
This is the largest accuracy failure in the search results for this question, and it is worth knowing in detail because the number is still quoted everywhere.

The event
On 10 June 2009, an organisation called the Global Language Monitor announced that English had acquired its one millionth word. Newspapers worldwide ran it. The figure entered general circulation and never left.
The millionth word was Web 2.0
Geoffrey Pullum of Language Log:
“First, that isn’t a word, it’s a phrase containing a noun (web) and one of those stylish postpositive decimal numeric quantifiers; and second, it is boring boring boring. If phrases containing numbers are allowed, no wonder there are a million words.”
Word 1,000,001 was financial tsunami. Other entries on the same list included cloud computing, shovel ready, zombie banks, and overseas contingency operations. All phrases.
The methodology problems
The date kept moving. GLM predicted the milestone for 2006, then 2007, then 2008, then 29 April 2009, then finally 10 June 2009 at “10:22 am, Stratford-on-Avon Time.”
The counter did not move. Ben Zimmer archived GLM’s public word counter and found it frozen for months: 991,833 on 31 December 2006 and still 991,833 on 2 April 2007; 995,116 on 23 October 2007 and still 995,116 on 23 December 2007.
The algorithm was never disclosed. GLM described its “Predictive Quantities Indicator” as simultaneously a matter of public record and proprietary. Pullum’s reply: “he can’t have it both ways.”
The timing was commercially convenient. The originally announced date, 29 April 2009, was the publication date of the paperback edition of GLM founder Paul Payack’s own book.
The counter script broke on the day. GLM’s site displayed: “The English Language WordClock: 1,000,001 / 0 words until the 1,000,000th Word.”
What linguists said
| Who | Verdict |
|---|---|
| David Crystal, on BBC Newsnight | “the biggest load of chicken droppings I’ve heard in a long time” |
| Geoffrey Nunberg, UC Berkeley | Called Payack “kind of a linguistic Madoff” and “a California marketing executive who has a gift for concocting appealing factoids about language trends” |
| Ben Zimmer | GLM “has managed to hoodwink unsuspecting journalists on a range of pseudoscientific claims” |
| Geoffrey Pullum | “The whole project is so bafflingly stupid it boggles the very few parts of one’s mind that remain to be boggled” |
The concession that settles it
Payack, quoted in the press at the time:
“You can’t be precise. We’re talking from a poet’s perspective. I’m a word lover, not a linguist.”
Verdict: do not cite “English has a million words” as a measured fact, and do not repeat GLM’s derived claims either, including the widely copied “a new English word is created every 98 minutes.” If you need a defensible large number, use the Google Books figure below with its limitations stated.
One clarification worth making, since it appears in several otherwise careful articles: the peer-reviewed Google Books study did not confirm GLM’s claim. It measured a different thing (word forms in print, including typos and numerals) using a disclosed method. The rough coincidence of scale is not corroboration.
The 1,022,000 Figure: The Best Large Number, and Its Limits
The most academically respectable large figure comes from Michel et al., “Quantitative Analysis of Culture Using Millions of Digitized Books,” published online in December 2010 and in print in Science in 2011. The team analysed 5,195,769 digitised books, roughly 4% of everything ever published, containing 361 billion English words.
Headline findings:
| Year | Estimated distinct English word forms |
|---|---|
| 1900 | 544,000 |
| 1950 | 597,000 |
| 2000 | 1,022,000 |
Growth averaged about 8,500 word forms per year, expanding the measured lexicon by over 70% in 50 years. The authors also estimated that 52% of the year-2000 lexicon was “lexical dark matter,” absent from standard reference works.
That is a serious paper. Here is what its own methods section says, which almost no article reporting the number mentions.
It counts strings, not words. A “1-gram” is any string uninterrupted by a space. The authors state that this “includes words (‘banana’, ‘SCUBA’) but also numbers (‘3.14159’) and typos (‘excesss’).”
It is an extrapolation, not a count. The raw figure was 1,489,337. The authors hand-annotated random samples and estimated that 31% of year-2000 strings were non-words. The famous 1,022,000 is approximately 1,489,337 × 0.69.
It counts word forms, not lemmas. Run, runs, running, ran are four entries. Convert to lemmas and you get roughly 601,000.
The corpus has documented problems. The authors’ own FAQ concedes: “the full text of the corpus contains trillions of letters. Even if one in a thousand is misrecognized, that still leaves you with a billion errors.” They identify husband being scanned as hufband as “the most prominent systemic OCR error in our data,” and acknowledge a “1 in 20 chance that the date is wrong.”
Independent critique. A 2015 PLOS ONE study by Pechenick, Danforth and Dodds found the corpus “is in effect a library, containing one of each book,” making it “more lexicon-like than text-like,” and concluded that its properties “call into question the vast majority of existing claims drawn from the Google Books corpus.”
Verdict: 1,022,000 is a real, published, peer-reviewed estimate of distinct English word forms appearing in digitised books as of 2000. It is not the number of words in English, its authors did not claim it was, and the growth rate should not be projected forward to manufacture a 2026 total.
How English Gains New Words?
Vocabulary grows through a small set of recurring processes. Recognising them makes new coinages far less mysterious.
| Process | How it works | Examples |
|---|---|---|
| Borrowing | Adopting a form from another language | schadenfreude, sushi, algorithm |
| Compounding | Joining existing words | smartphone, heat dome, doomscroll |
| Affixation | Adding prefixes or suffixes | unfriend, decarbonize, slopification |
| Conversion | Changing grammatical role without changing form | noun message becomes verb message |
| Blending | Merging parts of two words | brunch, smog, podcast |
| Clipping | Shortening a longer form | app from application, flannie from flannelette |
| Initialism and acronym | Abbreviation becomes a stable item | DM, AI, radar |
| Semantic change | An existing word gains a new sense | cloud in computing, stream for media, slop for AI output |
| Regional innovation | A variety of English develops vocabulary that spreads | yeah nah, agak-agak |
How Many New Words Enter English Each Year?
| Source | Additions | Period | What is counted |
|---|---|---|---|
| OED | 500 to 900+ per quarterly update, roughly 2,200 a year | 2025 to 2026 | New words, phrases and senses, plus revisions |
| Cambridge Dictionary | 6,212 | 12 months to Aug 2025 | Words, phrases and meanings |
| Merriam-Webster Collegiate, 12th ed. | 5,000+ words, 1,000+ phrases | Nov 2025 | Accumulated over 22 years |
| Dictionary.com | 1,235 entries, 1,798 senses | H1 2025 | Entries and senses |
| Michel et al. (books corpus) | ~8,500 per year | 1950 to 2000 | Distinct strings, contested |
Recent OED additions give a flavour of the range. The March 2026 update added over 500 new words, phrases and senses including doomscrolling, to touch grass, jelly (adjective, meaning jealous), and techno-futurist, plus a new sense of charismatic covering charismatic megafauna, with World English contributions from Hong Kong, the Philippines, Malaysia, Singapore, South Africa and Ireland. The June 2026 update added hundreds more including fiending, blingy, and the Australian flannie.
A correction worth making, since it appears widely: skibidi, delulu and tradwife were added by Cambridge in August 2025, not by the OED.
The gap between coinage and dictionary addition
There is a two-order-of-magnitude gap between corpus-derived “new word” rates and actual lexicographic additions, and the reason is instructive.
Michel and colleagues noted that the American Heritage Dictionary added only 2,077 single-word headwords in its entire year-2000 update, and that “over half the words added to AHD4 were part of the English lexicon a century ago.”
That is the real story. Most “new” dictionary words are not new. They are old words that finally crossed an evidence threshold. Mark Liberman made the statistical version of the point: given the very long tail of rare old words, the next string to cross any usage threshold is more likely to be an obscure eighteenth-century term that got digitised than a fresh coinage.
Two further cautions:
- Net growth is not gross additions. Cambridge actively removes words (younker, snollygoster, hodad).
- New words are not necessarily permanent. The American Dialect Society notes that “not all words chosen for a particular year are destined to become permanent additions,” citing Y2K (1999) and chad (2000).
Is English growing faster because of the internet and AI?
Diffusion is faster. Coinage is not demonstrably faster.
No authoritative lexicographic or corpus-linguistic finding shows that the rate of new word creation is rising. What is documented is that new words now spread far faster. Dictionary.com noted of its 2025 word of the year, 67, that it “shows the speed at which a new word can rocket around the world.”
The genuinely new development: AI is shifting word frequency
This is the most interesting current finding on the topic, and it concerns which words get used rather than how many exist.
Kobak, González-Márquez, Horvát and Lause (2025), Science Advances**.** The team analysed over 15 million biomedical abstracts from 2010 to 2024 and found that at least 13.5% of 2024 abstracts showed signs of LLM processing, reaching 40% in some subcorpora. They identified “excess vocabulary” including delve, intricate and underscore, whose frequency rose sharply after ChatGPT’s release in November 2022. The vocabulary shift they measured surpassed the effect of major world events including the Covid pandemic.
Yakura et al. (2025). An analysis of 740,249 hours of human speech from academic talks and podcasts found “a measurable and abrupt increase in the use of words preferentially generated by ChatGPT, such as delve, comprehend, boast, swift, and meticulous.” The authors describe a closed cultural feedback loop: models learned from us, and we are now learning from them.
There is a neat irony in the timing. The OED’s own 2026 workflow now includes an AI-assisted Quotations Finder.
2025 Words of the Year
| Body | 2025 Word of the Year | Definition |
|---|---|---|
| Merriam-Webster | slop | Low-quality digital content produced in quantity by AI |
| American Dialect Society | slop | 36th annual vote, January 2026 |
| Macquarie (Australia) | AI slop | Committee and People’s Choice |
| Collins | vibe coding | Using natural-language prompts to make AI write code |
| Cambridge | parasocial | Definition explicitly extended to cover AI |
| Oxford | rage bait | Content designed to provoke anger for engagement |
| Dictionary.com | 67 | “Impossible to define. Meaningless, ubiquitous, and nonsensical” |
Four of seven chose an AI-related term and two landed on the same word. Watch slop in particular: it is already generating compounds (sloppunk, slopification, workslop, slopaganda), which is the clearest sign a word has taken root.
How Many Words Does the Average Person Know?
This is where the research becomes solid and the answers become useful.
The definitive study
Brysbaert, Stevens, Mandera and Keuleers (2016), Frontiers in Psychology, based on 221,268 participants across 265,346 test sessions. Nothing has superseded it as of 2026.
| Age | 5th percentile | Median | 95th percentile |
|---|---|---|---|
| 20 years | 27,100 lemmas | 42,000 lemmas | 51,700 lemmas |
| 60 years | 35,100 lemmas | 48,200 lemmas | 56,400 lemmas |
In word families that median is 11,100 at age 20, rising to about 13,400 at age 60.
Further findings:
- Adults also recognise about 4,200 non-transparent multiword expressions, idioms such as kick the bucket whose meaning cannot be derived from the parts.
- Between ages 20 and 60 the average person adds about 6,200 lemmas, roughly one new word every two days.
- Growth continues steadily to around age 70.
- Scores vary widely with education and language exposure.
These are receptive figures. The test measured recognition, so knowing that a form exists could count even where the participant could not use it fluently.
⚠️ Common error to avoid: many sites report Brysbaert’s multiword-expression figure as ~11,000. It is 4,200. The 11,100 is the word-family count. This conflation appears on dozens of otherwise reputable pages.
Receptive versus productive vocabulary
- Receptive vocabulary is what you recognise while reading or listening.
- Productive vocabulary is what you can retrieve and use appropriately.
Receptive is always larger, and productive vocabulary is generally estimated at under half of it. This is why you can recognise a rare word in a novel and never once use it aloud.
Words known and words spoken are different measurements
A person can recognise 42,000 lemmas and still repeat a much smaller set during an ordinary day.
Pfeifer and Mehl (2026), Perspectives on Psychological Science, pooled 22 studies covering 2,197 participants aged 10 to 94 recorded between 2005 and 2019. Their headline result is a decline: average daily speech fell from 15,959 spoken word tokens per day in 2007 to 12,792 in the most recent data, a drop of roughly 338 words per day per year.
Three points matter here:
- These are tokens, every spoken occurrence including repetitions, not unique words.
- The decline is a finding about how much people talk, not about the size of English.
- Nobody has published a reliable figure for unique words used per day, which would require a separate type or lemma count.
Keep the measures separate. 42,000 lemmas known describes receptive vocabulary. 12,792 tokens spoken per day describes total spoken occurrences in one dataset.
Other figures you will see, and what they are worth
| Claim | Actual source | Verdict |
|---|---|---|
| “17,000 base words for educated adults” | Goulden, Nation & Read (1990) | Real, but n = 20 university graduates |
| “20,000 word families for a graduate” | Nation & Waring (1997) | A rounded rule of thumb glossing the 17,000 study, not an independent finding |
| “20,000 to 35,000 for native speakers” | TestYourVocab.com, 2010 to 2013 | Self-selected online sample; the site itself warns participants sit around the 98th percentile |
| “~10,000 words at age 8” | Various | Broadly consistent with developmental research |
| “4,000 to 5,000 word families at age 5” | Nation & Waring | Reasonable |
| “60,000 active / 80,000 passive” | No study | Folk statistic |
| “The average person knows only 800 to 1,000 words” | No study | Folk statistic, wrong by more than an order of magnitude |
Even a highly literate adult recognising 45,000 lemmas knows well under 10% of the word forms in the Google Books corpus. No one knows all of English. Everyone knows a working subset of it.
How Many Words Do You Need?
For most people searching this question, this is the real query hiding behind it. It has better answers than the headline question does.
The coverage research
Paul Nation’s 2006 study is the standard reference. His finding, stated conditionally:
“If 98% coverage of a text is needed for unassisted comprehension, then a 8,000 to 9,000 word-family vocabulary is needed for comprehension of written text and a vocabulary of 6,000 to 7,000 for spoken text.”
⚠️ The most common misquote online is dropping that “if.” Nation does not assert 8,000 to 9,000 unconditionally, and he adds two qualifiers almost nobody repeats: 98% coverage “does not make comprehension easy,” and even at 98%, “1 word in 50 will be unfamiliar.” Proper nouns are counted separately and add another 1% to 6% of running text depending on genre.
Cumulative text coverage by vocabulary size
Based on Nation’s analysis of a full-length novel (word families alone, then including proper nouns):
| Word families known | Written coverage | With proper nouns | What it feels like |
|---|---|---|---|
| 1,000 | 80.9% | 82.9% | 1 word in 5 unknown. Unreadable |
| 2,000 | 88.1% | 90.1% | 1 in 10 unknown. Exhausting |
| 3,000 | 91.2% | 93.3% | Gist only, heavy dictionary use |
| 4,000 | 93.0% | 95.1% | Readable with effort |
| 5,000 | 94.1% | 96.1% | Comfortable with occasional lookups |
| 7,000 | 95.4% | 97.4% | Fluent reading |
| 9,000 | 96.2% | 98.2% | Unassisted comprehension |
| 14,000 | 97.0% | 99.0% | Near-native |

Spoken English is less demanding: 6,000 word families plus proper nouns reaches 97.7% to 98.3%. A graded reader needs only 3,000 families plus proper nouns for 98.9%.
Newspapers are harder than fiction at the same vocabulary level, and the reason is proper nouns: they make up 4.6% to 6.1% of running words in news text. Reading newspapers unassisted takes roughly 8,000 families plus proper-noun recognition.
⚠️ A widely repeated oversimplification: “3,000 word families cover 95% of English.” That is roughly true for everyday conversation, where high-frequency words dominate. It is not true for written text, where 3,000 families reach about 91%. Always state which mode you mean.
The 95% versus 98% debate
The “you need 98%” claim is contested, and the dispute is useful to understand.
- Laufer (1989) identified 95% as the coverage level that best distinguished learners, but the comprehension criterion was only 55% on a test. That detail is rarely mentioned.
- Hu & Nation (2000), the study that established 98%, had n = 66 across four conditions using a single 673-word story.
- Schmitt, Jiang & Grabe (2011), with n = 661 across eight countries, found no threshold at all. The relationship was “essentially linear, from about 50% comprehension at 90% vocabulary coverage to about 75% comprehension at 100%.”
- Kremmel et al. (2023) attempted a replication of Hu & Nation and “could not fully replicate the original results.”
Schmitt’s practical calibration:
| If you want | You need roughly |
|---|---|
| 60% comprehension | 95% coverage |
| 70% comprehension | 98% to 99% coverage |
| 75% comprehension | To know essentially every word |
Notably, Schmitt and colleagues found no threshold yet still concluded that “the 98% estimate is a more reasonable coverage target for readers of academic texts” than 95%. The disagreement is about whether 98% is a cliff or simply a good target, not about whether it is the right thing to aim at.
Practical translation: there is no magic number where comprehension switches on. Every 1,000 word families you add buys measurable comprehension. Aim for 98%, and stop treating it as a gate.
Your vocabulary target ladder
Use this as a progress ladder, not a set of gates.
- 1,000 families – Survive basic interactions. About 80% of written text, 82% of speech
- 2,000 families – Handle everyday conversation
- 3,000 families – Read graded readers unassisted, follow simple television. This is what the Oxford 3000 targets
- 5,000 families – Comfortable general reading with a dictionary nearby
- 6,000 to 7,000 families – Understand unscripted native speech
- 8,000 to 9,000 families – Read novels and newspapers unassisted
- 10,000+ families – Academic and professional register
- 11,000 to 13,000 families – Native-speaker range (Brysbaert et al.)
About those CEFR vocabulary numbers
You have seen the table: A1 = 1,500 words, A2 = 2,500, B1 = 3,250, B2 = 3,750, C1 = 4,500, C2 = 5,000. It circulates as though it were official.
It is not. The CEFR framework specifies skills and can-do statements and deliberately contains no word lists. Milton and Alexiou, the researchers most associated with the table, state that “the scale of vocabulary knowledge which might reasonably be associated with the CEFR levels is now an unknown quantity.”
The numbers are score bands from the Swansea Levels Test (XLex), a specific measurement instrument. Here is the detail that makes them unusable as targets: XLex has a maximum possible score of 5,000. The “C2 = 5,000 words” figure is pressed against the ceiling of the test. It is an artifact of instrument design, not a finding. Milton and Alexiou’s own measured C2 average was 4,068, below the predicted band, and they concluded “it seems likely that there are no clear thresholds.”
The nearest thing to an official reference is the Cambridge English Vocabulary Profile, which contains 6,970 headwords across A1 to C2, with about 4,700 up to B2. Cambridge states explicitly that it is descriptive and “not intended to offer a ready-made lexical syllabus.”
An efficient vocabulary plan
- Start with a frequency list, not a dictionary. Work through the Oxford 3000 or the New General Service List. Frequency-ordered learning delivers more coverage per study hour than any other strategy.
- Learn word families, not isolated forms. When you learn decide, take decision, decisive and undecided in the same session.
- Learn each word with its common partners. Study make a decision, not decision alone. Collocations are where fluency actually lives.
- Study several senses of frequent words. A familiar form usually carries unfamiliar meanings; that is where intermediate learners stall.
- Meet each word repeatedly in real context. Reading and listening beat isolated lists for retention.
- Retrieve from memory rather than reread. Short recall practice with spaced repetition outperforms review.
- Push words from passive to active by writing or saying three original sentences with each new item within 48 hours.
- Re-test every three to six months. Native adults gain roughly 150 to 180 lemmas a year passively. A deliberate learner can beat that many times over.
- Count honestly. Do not claim four new words when an app counted four inflections.
Does English Have More Words Than Other Languages?
Almost certainly not in any meaningful sense, and the question is close to unanswerable as posed.
Here is what you are actually comparing when you compare word counts across languages.
| Language | Reference work | Figure | What it really means |
|---|---|---|---|
| English | OED | 500,000+ entries | Historical dictionary, includes obsolete words back to Old English |
| German | Duden, 29th ed. (2024) | ~151,000 headwords | An orthographic dictionary with different inclusion rules |
| French | Académie française, 9th ed. (2024) | ~53,000 words | A prescriptive academy dictionary, deliberately selective |
| Spanish | RAE DLE | Total not published | Also prescriptive and academy-governed |
| Korean | Standard Korean Language Dictionary | ~509,000 print / ~423,000 online | Comparable in scope to a large English dictionary |
| Korean | Urimalsaem | ~1.19 million | Crowd-sourced, deliberately includes dialect, jargon, archaisms, North Korean forms |
| Japanese | Nihon Kokugo Daijiten, 2nd ed. | ~503,000 entries | Historical dictionary, comparable to the OED |
| Chinese | Hanyu Da Cidian | ~370,000 words | Historical, most entries no longer current |
| Chinese | Xiandai Hanyu Cidian, 7th ed. | ~70,000 entries | The actual standard modern dictionary |
Three reasons the comparison fails
1. You are comparing editorial philosophies, not languages. The OED includes every word attested since roughly the year 1000. The Académie française includes what it considers proper French. One of these will inevitably produce a bigger number.
Victor Mair made the point precisely when a BBC article claimed a Chinese dictionary’s 370,000 entries beat the OED: that dictionary “is based on historical principles, and most of its entries are no longer current.” The genuinely comparable modern Chinese dictionary has around 70,000 entries, which sits in a similar band to modern English dictionaries for general use.
2. Morphology changes the shape of the question. In Finnish, Turkish, Hungarian or Inuktitut, a finite set of roots plus productive suffixation generates an effectively unbounded set of surface forms. A single Finnish noun has thousands of inflected forms. German can compound without limit; Rindfleischetikettierungsüberwachungsaufgabenübertragungsgesetz was a real law title. Asking “how many words does Turkish have” is not a harder version of the English question, it is a differently shaped question.
3. Writing conventions decide where words end. One language writes a compound as a single orthographic word while English writes the equivalent with spaces. Word-boundary conventions are a publishing decision, not a fact about expressive capacity.
Claims that should be retired
| Claim | Problem |
|---|---|
| “Korean has 1.1 million words” | Conflates two books. The Standard Korean Language Dictionary holds ~509,000 headwords; the million-plus figure belongs to Urimalsaem, a separate crowd-sourced project |
| “Arabic has 12 million words” | No traceable primary source. The usual explanation is that it counts theoretically generable root-and-pattern combinations rather than attested words; classical dictionaries record closer to 200,000 |
| “Tamil has over 1.5 million words” | Circulates without a citable methodology or inclusion criteria |
| “English has the most words of any language” | Not provable. English almost certainly has the most counted words, which is a fact about lexicography |
What English defensibly does have?
A very large documented lexicon, for four reasons that have nothing to do with expressive superiority:
- A double vocabulary inherited from Germanic and Latin or French sources after 1066, giving pairs like freedom/liberty and kingly/royal
- Centuries of heavy borrowing from hundreds of languages
- Its role as a global lingua franca, generating regional varieties that all feed the record
- The most extensively compiled dictionary tradition in the world
Nunberg’s summary is still the sharpest:
“Granted, our dictionaries can lick their dictionaries. Merriam Webster’s Third International clocks in at around half a million words, against a mere 150,000 for the biggest dictionaries the French or Russians can come up with. But that doesn’t mean we have more words for the things that matter. Once you get past fifty thousand words or so, you’re strictly in crossword puzzle territory.”
He also named the impulse behind the question: “There’s something bizarre about the satisfaction we take in our swollen wordbooks. It’s really the last residue of the imperial pride, all the bits of the map that were colored pink.”
How Many Words Did Shakespeare Use and Invent?
Two of the most-repeated Shakespeare statistics are wrong, and the corrections are more interesting than the myths.
Total words in his works: 884,647
From the Folger Shakespeare Library, attributed to Marvin Spevack’s concordances, across 118,406 lines. Folger adds its own caveat: “counts may differ based on the criteria of what is included.”
Different words: 31,534, and that number is inflated
The count comes from Spevack and underpins Efron & Thisted (1976) in Biometrika, a famous statistical paper that used it to estimate Shakespeare knew “at least 35,000 more words” he never wrote down.
The catch: 31,534 counts word types, not lemmas. Take, takes, taking, took, taken, takest counts as six. On a dictionary-headword basis, David Crystal puts Shakespeare’s vocabulary at 17,000 to 20,000. Lancaster University’s Encyclopedia of Shakespeare’s Language project puts the different-word count around 21,000 and explains why the raw number misleads: “the more you write, the more opportunities you have to use more words that are different.”
“Shakespeare invented 1,700 words” is largely a myth
The evidence-based figure is around 500.
Researchers at Lancaster University led by Jonathan Culpeper with Jonathan Hope used computational searching across millions of words of pre-Shakespearean text and reported that “only around 500 words do seem to first appear in Shakespeare.” As they note, 500 is still a large number: most writers coin nothing and produce no first recording at all.
Where 1,700 came from. David Crystal counted about 2,035 words for which Shakespeare is the earliest OED2 citation, then estimated about 1,392 plausible coinages after filtering. 1,700 is the midpoint of those two numbers. It is a midpoint of an estimate, not a count.
Why the OED overcredits him. Merriam-Webster’s own debunking is blunt: the myth “appears to have come about through a misreading of the data in the Oxford English Dictionary,” and “at no point did the editors of the OED say ‘We have X entries for which Shakespeare is the earliest known user; therefore he invented X number of words.’” The OED’s early citations came largely from volunteers, and “these volunteers preferred searching for words in Shakespeare, as opposed to legal documents, court memoranda, and turgid ecclesiastical screeds.”
Shakespeare received roughly 33,000 quotations in the OED’s first edition, against Walter Scott’s 15,000 and Milton’s and Chaucer’s 11,000 to 12,000. He is overrepresented by a factor of two or three.
Words he is credited with but did not coin
| Word | Actually attested |
|---|---|
| assassination | 1572 |
| uncomfortable | 1534, thirty years before his birth |
| bedazzle | 1572 |
| eyeball | 1575 |
| inaudible | 1589 |
| premeditated | 1564 |
| frugal | 1542 |
| hurry | 1582 |
| bold-faced | 1572 |
OED lexicographer Giles Goodland found that in one sample of 117 Shakespeare “firsts,” nearly half had already been superseded by earlier attestations.
Was he unusually inventive?
Hugh Craig’s 2011 study in Shakespeare Quarterly equalised samples at the first 10,000 words of every play and found that among 13 playwrights with three or more surviving titles, Shakespeare ranks 7th for vocabulary variety, behind Webster, Dekker and Jonson. Craig’s conclusion: “Shakespeare has a larger vocabulary because he has a larger canon.”
For scale: even if every word Shakespeare used had been his own coinage, all 21,000 or so, that would be about 4% of today’s documented English lexicon.
How to Verify Any Word-Count Claim?
Use this before you repeat a number anywhere.
The 7-question audit
- What is the unit? Tokens, types, lemmas, lexemes, word families, headwords, entries, phrases, or senses?
- What is the scope? Does it include historical, obsolete, dialectal, slang, scientific, and regional vocabulary?
- What is the source? A dictionary publisher, a peer-reviewed paper, a corpus project, or an unsourced blog?
- What is the date? A figure tied to OED2 in 1989 is not a 2026 census.
- What was the evidence threshold? Did a form appear once, clear a frequency cutoff, or receive manual review?
- How were forms normalised? Were capitalisation, variant spellings, hyphens, contractions and inflections merged?
- Is uncertainty reported? A broad linguistic estimate written with false precision is a warning sign in itself.
Source reliability tiers
| Tier | What it looks like | Examples | How to treat it |
|---|---|---|---|
| A | Primary publisher data or peer-reviewed research with a disclosed method | OUP entry statistics, Merriam-Webster FAQ, Brysbaert et al., Nation, Michel et al. | Cite directly, with the unit and date attached |
| B | Reputable secondary analysis or a checkable dataset | Language Log, Lancaster’s Shakespeare project, Wiktionary dumps | Usable, but verify the underlying figure |
| C | Aggregators, content farms, undisclosed algorithms, promotional counters | Global Language Monitor, most listicles, self-selected online tests | Do not cite. Trace to the original or drop the claim |
Red flags that a word-count statistic is unreliable
- 🚩 A suspiciously round number, especially exactly one million
- 🚩 No counting rule stated
- 🚩 Two dictionaries from different countries ranked as if they measured the same thing
- 🚩 The trail leads to the Global Language Monitor
- 🚩 It cites “a Harvard and Google study” without mentioning that the figure counts typos and numerals
- 🚩 It claims a specific date on which a word “entered the language”
- 🚩 The source is a language-learning company selling a course on the same page
- 🚩 A growth rate from an old study projected forward to produce a current total
A reproducible counting template
Researchers, editors and developers publishing their own count can use this:
We counted [unit] in [dictionary or corpus], version [date]. We included [scope] and excluded [scope]. We treated inflections as [same/separate], compounds as [rule], spelling variants as [rule], and multiword expressions as [rule]. The result was [number], subject to [known limitations].
Worked example:
We counted case-insensitive written types in a 2026 news corpus. We excluded punctuation and numerals, kept hyphenated forms intact, and counted inflections separately. The result describes unique spellings in that corpus, not all English words.
A smaller, well-defined number is more useful than a dramatic total with no counting rules behind it.
Frequently Asked Questions
How many words are in the English language?
There is no exact figure, and that is a property of the question rather than a gap in knowledge. Depending on the counting rule, defensible answers range from about 171,476 (words classed as in current use) through 500,000 to 600,000 (the OED’s documented record) to over 1 million (all attested word forms) to effectively unbounded (including generable chemical nomenclature).
How many words are in the Oxford English Dictionary?
Oxford describes the OED as documenting roughly 600,000 words across more than 1,000 years, alongside over 500,000 entries and 3.5 million quotations. The two figures are different units: 600,000 counts word forms defined or illustrated, while 500,000+ counts editorial entries. The widely quoted 171,476 refers to words in current use in the 20-volume second edition published in 1989.
Are there really a million words in English?
Not as a measured fact. The “millionth word” announcement of June 2009 came from the Global Language Monitor, whose method was never disclosed, whose predicted date moved from 2006 to 2007 to 2008 to April 2009 and finally to June 2009, and whose winning “word” was the phrase Web 2.0. Its founder later conceded, “You can’t be precise. I’m a word lover, not a linguist.” Merriam-Webster does cite a rough one-million estimate but warns it could be off by a quarter-million and includes chemical and scientific names.
How many words does the average person know?
About 42,000 lemmas for a 20-year-old native speaker of American English, rising to about 48,200 by age 60, based on a study of 221,268 participants. In word families that is roughly 11,100 to 13,400. Adults also recognise about 4,200 idiomatic multiword expressions. Productive vocabulary, the words you actively use, is generally estimated at under half of that.
How many words does a person say in a day?
A 2026 study in Perspectives on Psychological Science pooling 22 studies and 2,197 participants estimated 12,792 spoken word tokens per day in its most recent data, down from 15,959 in 2007. That counts every spoken occurrence including repetitions, not unique words. No reliable figure exists for unique words used per day.
How many words do you need to be fluent in English?
For unassisted reading of novels and newspapers, roughly 8,000 to 9,000 word families. For understanding unscripted speech, 6,000 to 7,000. For everyday conversation, 2,000 to 3,000 gets you a long way. There is no threshold at which fluency switches on; comprehension rises steadily with vocabulary, and fluency also depends on grammar, pronunciation, speed of retrieval, and interaction.
How many words cover 95% of everyday English?
About 4,000 word families plus proper nouns for written text, and around 3,000 for conversation. Handle this statistic carefully: 95% coverage still leaves one unknown word in every twenty, and research suggests it supports only around 60% comprehension.
How many words do you need for each CEFR level?
The CEFR does not officially specify vocabulary sizes. The circulated table (A1 = 1,500 through C2 = 5,000) comes from score bands on the Swansea Levels Test, whose maximum possible score is 5,000, so the top band is an artifact of the test ceiling. Use those numbers as rough orientation, not as targets. Cambridge’s English Vocabulary Profile, with 6,970 headwords across A1 to C2, is the nearest thing to an official reference, and Cambridge describes it as descriptive rather than a syllabus.
How many new words are added to English each year?
The OED adds 500 to 900+ new words, phrases and senses per quarterly update, roughly 2,200 a year. Cambridge added 6,212 in the twelve months to August 2025. Merriam-Webster’s 12th Collegiate added 5,000+ accumulated over 22 years. Most additions are not new coinages: over half the words added to one major dictionary update had already existed for a century.
Who decides whether something is an English word?
No academy controls English. Speakers establish words through repeated meaningful use, and dictionary editors document the forms that meet their publication criteria. A dictionary addition records an editorial decision supported by usage evidence; it does not grant permission or mark a birthday.
Does slang count as English?
Yes. Established slang is part of the language even when it is informal or limited to one community. Whether a particular term belongs in a given dictionary depends on that dictionary’s audience and on the strength, spread and duration of the evidence.
Which language has the most words?
Unanswerable as posed. Word counts compare dictionaries, and dictionaries differ in scope, era and editorial philosophy. Historical dictionaries including the OED, Japan’s Nihon Kokugo Daijiten and China’s Hanyu Da Cidian all report figures in the hundreds of thousands. Agglutinative languages such as Finnish and Turkish generate unbounded surface forms. English has the largest thoroughly documented lexicon, which is a fact about lexicography rather than about linguistic capacity.
Does English have more words than Spanish or French?
English dictionaries report larger totals, but that reflects editorial policy. The Académie française is prescriptive and deliberately selective, listing around 53,000 words in its ninth edition. The OED is historical and includes obsolete vocabulary going back a thousand years. They are not measuring the same thing.
What is the difference between a word, a lemma and a word family?
A word form is any surface form (run, runs, running, ran = four). A lemma is the dictionary headword grouping inflections (run = one). A word family groups a lemma with its transparent derivations (run, runner, rerun = one). English averages roughly 1.7 word forms per lemma and 3.8 lemmas per family. Nearly all confusion about word counts comes from mixing these three.
How many words did Shakespeare invent?
Around 500, based on computational searching of pre-Shakespearean texts by Lancaster University researchers Jonathan Culpeper and Jonathan Hope. The popular figure of 1,700 is the midpoint of an estimate derived from OED citation data, and the OED overcredits Shakespeare because its early volunteer contributors preferred reading him to reading legal records.
How many words are in the Merriam-Webster dictionary?
Webster’s Third New International, Unabridged, contains roughly 470,000 entries, though its base text dates from 1961 with a 1993 addenda section. The Collegiate, whose 12th edition shipped in November 2025 after a 22-year gap, does not publish a current total entry count.
Which English dictionary has the most words?
It depends on the metric, and both Chambers and Collins claim to be the largest single-volume English dictionary. Collins reports 732,000 “words, meanings and phrases,” a composite figure not comparable to an entry count. By entry count, the OED at over 500,000 is the largest.
What is “lexical dark matter”?
A term from the 2011 Science Google Books study for words that appear in real texts but are absent from standard dictionaries. The authors estimated that 52% of the year-2000 English lexicon fell into this category, though the estimate depends on their string-identification method and includes proper nouns, technical strings and variant spellings.
Why does every website give a different number?
Because they are answering different questions without saying so. Once you separate word forms, lemmas, word families, headwords, entries and senses, and convert everything into one unit, most of the apparent contradictions dissolve.
The Answer Worth Keeping
The honest answer to “how many words are in the English language” is that the question has five answers and you have to pick one.
If someone needs a single number, give them this: about 170,000 words are in current use, and over 500,000 entries covering roughly 600,000 word forms are recorded in the OED across the full history of the language. Both figures come from Oxford University Press.
If they want the number that matters, it is much smaller: you personally recognise about 42,000 lemmas, and about 9,000 word families cover 98% of everything you read.
If they mention a million, that figure came from a 2009 marketing exercise whose millionth word was Web 2.0, and whose author said on the record that he is “a word lover, not a linguist.” A peer-reviewed corpus study does put around a million distinct word forms in printed books, but converted into lemmas that lands close to the OED’s entry count rather than dwarfing it.
The counting problem is not a failure of lexicography. It is a feature of how language works. English does not come pre-divided into countable units any more than a coastline comes pre-divided into miles. What we can measure precisely is how much of the language you need to read a novel, follow a conversation, or pass an exam. Those numbers are smaller, better established, and considerably more useful.
Sources
Dictionaries and publishers
-
Merriam-Webster, “How many words are there in English?”
-
Merriam-Webster, “10 Words Shakespeare Never Invented”
-
Merriam-Webster, Collegiate Dictionary, Twelfth Edition · press release, 25 September 2025
-
Merriam-Webster, Word of the Year 2025: slop
-
Oxford English Dictionary, “The OED today” (entry, definition and quotation statistics)
-
Oxford English Dictionary, “Understanding entries: OED terminology”
-
Oxford English Dictionary, March 2026 update · new words, March 2026 · World English additions, March 2026
-
Oxford English Dictionary, new words, June 2026 update
-
Oxford University Press, “10 highlights from the March 2026 OED update”
-
Oxford Languages, current-English dictionary licensed to Google
-
Oxford, Word of the Year 2025: rage bait
-
Oxford Learner’s Dictionaries, The Oxford 3000 and Oxford 5000 · methodology
-
Collins, Collins English Dictionary · Word of the Year 2025: vibe coding
-
Cambridge, “Cambridge Dictionary adds skibidi, delulu and tradwife” (August 2025)
-
Cambridge, Word of the Year 2025: parasocial
-
Cambridge, English Vocabulary Profile
-
American Dialect Society, 2025 Word of the Year: slop
-
Macquarie Dictionary, Word of the Year 2025: AI slop
-
Chambers, The Chambers Dictionary
-
Folger Shakespeare Library, Frequently Asked Questions
-
kaikki.org, English Wiktionary extraction · raw data downloads
-
Michel, J.-B. et al. (2011), “Quantitative Analysis of Culture Using Millions of Digitized Books”, Science 331(6014):176–182. Free full text (PMC)
-
Pechenick, E.A., Danforth, C.M. & Dodds, P.S. (2015), “Characterizing the Google Books Corpus: Strong Limits to Inferences of Socio-Cultural and Linguistic Evolution”, PLOS ONE 10(10):e0137041
-
Brysbaert, M., Stevens, M., Mandera, P. & Keuleers, E. (2016), “How Many Words Do We Know?”, Frontiers in Psychology 7:1116
-
Nation, I.S.P. (2006), “How Large a Vocabulary Is Needed For Reading and Listening?”, Canadian Modern Language Review 63(1):59–82. Free PDF
-
Waring, R. & Nation, I.S.P. (1997), “Vocabulary Size, Text Coverage and Word Lists”, in Schmitt & McCarthy (eds), Vocabulary: Description, Acquisition and Pedagogy. HTML version
-
Goulden, R., Nation, P. & Read, J. (1990), “How Large Can a Receptive Vocabulary Be?”, Applied Linguistics 11(4):341–363. Free PDF
-
Hu, M. & Nation, I.S.P. (2000), “Unknown Vocabulary Density and Reading Comprehension”, Reading in a Foreign Language 13(1):403–430. Free full text
-
Laufer, B. (1989), “What Percentage of Text-Lexis Is Essential for Comprehension?”, in Lauren & Nordman (eds), Special Language: From Humans Thinking to Thinking Machines, 316–323
-
Schmitt, N., Jiang, X. & Grabe, W. (2011), “The Percentage of Words Known in a Text and Reading Comprehension”, Modern Language Journal 95(1):26–43. ERIC record
-
Kremmel, B., Indrarathne, B., Kormos, J. & Suzuki, S. (2023), “Unknown Vocabulary Density and Reading Comprehension: Replicating Hu and Nation (2000)”, Language Learning 73(4):1127–1163 (open access)
-
Milton, J. & Alexiou, T. (2009), “Vocabulary Size and the Common European Framework of Reference for Languages”, in Vocabulary Studies in First and Second Language Acquisition, 194–211
-
Capel, A. (2012), “Completing the English Vocabulary Profile: C1 and C2 vocabulary”, English Profile Journal 3:e1
-
Efron, B. & Thisted, R. (1976), “Estimating the Number of Unseen Species: How Many Words Did Shakespeare Know?”, Biometrika 63(3):435–447
-
Craig, H. (2011), “Shakespeare’s Vocabulary: Myth and Reality”, Shakespeare Quarterly 62(1):53–74
-
Kobak, D., González-Márquez, R., Horvát, E.-Á. & Lause, J. (2025), “Delving into LLM-assisted writing in biomedical publications through excess vocabulary”, Science Advances 11(27). Free preprint · data and code
-
Yakura, H. et al. (2025), “Empirical evidence of Large Language Model’s influence on human spoken communication”
-
Pfeifer, V.A. & Mehl, M.R. (2026), “Sliding Into Silence? We Are Speaking 300 Daily Words Fewer Every Year”, Perspectives on Psychological Science
Linguistic commentary and corrections
- Zimmer, B. (2009), “The ‘million word’ hoax rolls along”, Language Log. Includes the archived Global Language Monitor counter readings and Payack’s reply in the comment thread
- Pullum, G.K. (2009), “Millionth word story botched”, Language Log
- Pullum, G.K. (2009), “End times at hand”, Language Log
- Liberman, M., “Why estimating vocabulary size by counting words is (nearly) impossible”, Language Log
- Nunberg, G. (2006), “Size Doesn’t Matter”, NPR Fresh Air commentary
- Mair, V., “Hot words”, Language Log
- Crystal, D. (2009), “On the biggest load of rubbish”, DCblog
- Culpeper, J. & Hope, J. (2022), “Five myths about Shakespeare’s contribution to the English language”, The Conversation
- Encyclopedia of Shakespeare’s Language project, Lancaster University