Word counts 28 min read

How Word Length Affects Readability: Data, Formulas, Fixes

Word length moves readability scores more than sentence length, but it is only a proxy for familiarity. See the data, formulas, benchmarks and a 7-step fix.

Word length affects readability in two separate ways, and confusing them is why advice on this topic contradicts itself.

Inside a readability formula, word length is the most powerful variable there is. In Flesch Reading Ease, adding 0.1 syllables per word costs 8.46 points, the same penalty as adding 8.3 words to every sentence. Word length is weighted about 83 times more heavily than sentence length.

Inside a real reader’s head, the effect is smaller and mostly mechanical. Long words are skipped far less often (29.3% skip rate for short words against 10.7% for long ones) and cost roughly 10 milliseconds of extra gaze time per letter. When researchers hold word familiarity constant, most of the apparent length effect disappears.

The reason is that word length is not a cause of difficulty. It is a proxy for word frequency. MetaMetrics, the company behind the Lexile framework used to level most US school texts, says so in its own technical report: “Variables such as the average number of letters or syllables per word have been observed to be proxies for word frequency.”

What to do with that: shortening words reliably improves your score and often improves comprehension, because most long words in weak writing are also rare words. But the rule that works is prefer the familiar word, not prefer the short word. Those two rules agree about 80% of the time. The 20% where they diverge is where writers do real damage, cutting precise technical terms, breaking terminology consistency, and leaving short-but-obscure words like writ, jib, and ennui untouched because no formula flags them.

Practical target for general audiences: about 4.7 to 5.0 characters and 1.40 to 1.55 syllables per word, with 3+ syllable words under roughly 12% of your text.

Key Numbers

QuestionAnswerSource
Average English word length4.69 characters / 1.518 syllablesOriginal measurement, Brown Corpus (991,429 words)
Share of English words that are monosyllabic66.5%Original measurement, Brown Corpus
Cost of +0.1 syllables per word in Flesch Reading Ease−8.46 pointsFlesch (1948) coefficient
Cost of +0.1 syllables per word in Flesch-Kincaid+1.18 gradesKincaid et al. (1975)
Skip rate, short (4-letter) vs long (10–12 letter) words29.3% vs 10.7%Slattery & Yates (2018)
Extra gaze time, 3–4 letter vs 7–9 letter words+42 ms (about 10 ms per letter)Pollatsek et al. (2008)
Extra gaze time from low frequency, same materials+20 msPollatsek et al. (2008)
Score variation for the same text across toolsUp to 12.9 grade levels, using the same formulaMac et al. (2022), JAMA Network Open
US adults at the lowest literacy level (2023)28%, up from 19% in 2017NCES PIAAC Cycle 2
Is readability a confirmed Google ranking factor?NoGoogle (Mueller, 2018); Portent 5.8M-page study

Key Takeaways

  1. Word length is the most powerful variable inside readability formulas. In Flesch Reading Ease, +0.1 syllables per word costs 8.46 points, equal to adding 8.3 words to every sentence. Audit vocabulary before you touch sentence structure.
  2. The 5% Rule: swapping about 5% of your words for alternatives two syllables shorter moves Flesch Reading Ease about 8.5 points at any document length. Tractable, and therefore easy to game.
  3. Word length is a proxy for word frequency, not a cause of difficulty. MetaMetrics says so in its own Lexile report, and Zipf’s law of abbreviation explains why: languages shorten their common words.
  4. The best-validated approaches abandoned word length. Dale-Chall replaced it with a familiar-word list in 1948, Lexile uses corpus frequency, and in a 2025 CEFR levelling model word-length indices did not survive feature screening while Age of Acquisition ranked first.
  5. Long words genuinely cost eye-movement time, about 10 ms per letter, with skip rates falling from 29.3% to 10.7%. But the cleanest syllable-only experiment failed to replicate, and much of the neural effect disappears once multiple fixations are controlled for.
  6. Your tools are less accurate than you think. The same text can score 12.9 grade levels apart across calculators using the same formula. Acronyms and numbers count as one syllable, and one spelling convention shifted a test sentence by 18.8 Flesch points.
  7. The short-but-obscure word is the blind spot no tool covers. Writ, tort, jib, ennui, waive. Add a deliberate pass, and expect no score improvement for the most valuable edit you will make.
  8. Sometimes the longer word is correct. Terms of art, terminology consistency, translation memory, and factual precision outrank brevity when they conflict, and the plain-language guides say so themselves.
  9. Readability is not a Google ranking factor. Word length matters for search through demand. Add the common term, do not delete the precise one.
  10. The rule that survives all the evidence is Fowler’s, in the order he wrote it: prefer the familiar word, the concrete word, the single word, the short word. Short usually follows familiar. When it does not, familiar wins.

The Two-Track Model

Almost every argument about word length comes down to people measuring different things and assuming they measured the same thing.

Track 1 is the score track. It asks what happens to a number when you change your vocabulary. Here word length dominates, effects are large, and the math is exact because the formulas are linear.

Track 2 is the reader track. It asks what happens to a person when they meet your vocabulary. Here word length matters less, effects are measured in milliseconds, and much of what looks like a length effect turns out to be a familiarity effect in disguise.

Advice that only sees Track 1 tells you to swap every polysyllable and treats a Flesch score as proof of clarity. Advice that only sees Track 2 calls formulas useless and offers nothing to replace them. Both fail the same way.

The workable position: use Track 1 to find suspects, use Track 2 logic to judge them. A score tells you where to look. It cannot tell you what to do when you get there.

The Two Track Model

First, Separate Four Things That Get Confused

VariableWhat it meansExample changeWhat it affects
Word lengthCharacters or syllables per wordutilize (3 syllables) to use (1)Decoding effort, formula scores
Word countTotal words in the piece800 words to 400Reading time and depth, not scores
Sentence lengthWords per sentenceOne 40-word sentence to two 20-word onesWorking-memory load
Line lengthCharacters per displayed line140-character lines to 65Eye movement across the page

These move independently. You can halve word count without touching average word length. You can shorten every sentence and make average word length worse, because shorter sentences pack a higher density of content words. And short words set in 140-character lines still read badly; the screen target is 50 to 75 characters per line, which belongs to layout rather than vocabulary.

Readability formulas measure only two of the four. So “make it shorter” and “make it more readable” are different instructions, and no tool can tell you which one you followed.

One more distinction. Readability is how easily a reader understands content. Legibility is how easily they can see and distinguish characters. Clean typography cannot rescue unexplained jargon, and plain words fail in tiny low-contrast type.

How word length gets measured?

Tools quantify it four ways: letters per word (ARI, Coleman-Liau, LIX), syllables per word (Flesch, Flesch-Kincaid, FORCAST), polysyllabic count, meaning words of three or more syllables (Gunning Fog, SMOG), and long-word percentage above a letter threshold.

The calculation is simple: average word length = total letters ÷ total words. The four words in “We write clear text” contain 16 letters, so the average is 4.0.

What is not simple is what counts as a word. Contractions, hyphenated compounds, numerals, currency, and URLs are handled differently by different tools, so use identical settings when comparing drafts.

Every standard defines “long word” differently

Standard or formulaDefinition of a long or difficult wordUnit
LIX / RIX (Björnsson, 1968)More than 6 lettersLetters
Gunning Fog (1952)3+ syllables, excluding proper nouns and common inflectionsSyllables
SMOG (McLaughlin, 1969)3+ syllablesSyllables
Yoast SEO “word complexity”Over 7 characters and outside the top 5,000 words and not capitalisedLetters + frequency
Coleman-Liau, ARINo threshold, average characters per wordCharacters
Flesch, Flesch-KincaidNo threshold, average syllables per wordSyllables
Dale-Chall, SpacheNot on a list of about 3,000 familiar wordsFamiliarity
Lexile, ATOSCorpus word frequency, or a graded vocabulary listFrequency
ASD-STE100 (aerospace, Issue 9, 2025)Not in the ~900-word approved dictionary. No length rule.Approved vocabulary
WCAG 2.2 SC 3.1.3Unusual or restricted use, idioms, jargon. Length not mentioned.Familiarity

Two conclusions follow. The same word is hard or easy depending on the tool: interesting has three syllables, so Gunning Fog counts it as complex, and almost nobody would call it difficult. And the best-validated approaches abandoned word length entirely. Dale-Chall replaced it with a familiar-word list in 1948. Lexile replaced it with corpus frequency. Seventy-five years of research runs in one direction, away from length and toward familiarity.

Track 1: The Math of Word Length Inside Readability Formulas

Every mainstream formula except SMOG is linear in its word-length term, so the marginal effect is simply the coefficient. This is the most useful calculation in readability and almost nobody publishes it.

The Math of Word Length Inside Readability Formulas

ASL = average sentence length in words. ASW = average syllables per word. AWL = average characters per word. L = letters per 100 words. S = sentences per 100 words.

What +0.1 units of word length costs: −8.46 points in Flesch Reading Ease, +1.18 grades in Flesch-Kincaid, +0.59 in Coleman-Liau, +0.47 in ARI. Gunning Fog adds 0.4 grades per percentage point of hard words, and LIX adds 1.0 point.

SMOG is the exception. Its derivative falls as polysyllable count rises, so the first hard word in a passage costs far more than the fiftieth. SMOG punishes occasional jargon in otherwise plain text harder than any other formula.

The trade-off ratio

Divide the word-length coefficient by the sentence-length coefficient and you learn how much sentence length is worth the same as word length.

The trade off ratio

Two formulas by the same author disagree about your problem. Flesch Reading Ease weights word length 2.75 times more heavily, relative to sentence length, than Flesch-Kincaid does. Run both and one may point at your vocabulary while the other points at your sentences.

A controlled before-and-after, decomposed

One paragraph written twice, same sentence count, near-identical word count. Figures computed with textstat v0.7.13 and checked by hand.

Inflated (30 words, 3 sentences): “The board will utilize the strategy to facilitate personnel productivity. We will disseminate the revised regulations to each department subsequently. If you have an inquiry, consult your immediate supervisor initially.”

Plain (32 words, 3 sentences): “The board will use the plan to help staff work faster. We will send the new rules to each team next week. If you have a question, ask your line manager first.”

A controlled before and after, decomposed

Decomposing the 93.1-point Flesch gain: word length contributed +93.77 (99.3%), sentence length contributed −0.68 (0.7%). Word choice did essentially all the work while sentence length moved the wrong way. You can shift a readability score nearly a hundred points without touching a single sentence boundary.

The 5% Rule

How much editing does +0.1 syllables per word require? Total syllables must change by 0.1 × N, where N is your word count. A 1,000-word article needs 100 syllables removed, which is 50 swaps saving two syllables each, which is 5% of your words. The ratio holds at any length.

The 5% Rule: swapping roughly 5% of your words for alternatives two syllables shorter moves Flesch Reading Ease about 8.5 points and Flesch-Kincaid about 1.2 grades, regardless of document length.

Two things follow. Hitting a readability target is an afternoon’s work, not a rewrite. And it is tractable precisely because it is easy to game: the score moves 8.5 points whether the 50 words you changed were obscure or already perfectly clear. In short copy the effect is sharper, since one word choice in a 30-word product description can swing the score a full band.

Track 2: What Happens in a Real Reader’s Eyes

Long words get skipped far less often

Skilled readers skip roughly a quarter of the words on a page, and word length is the dominant factor in which ones. Slattery & Yates (2018) measured this across 92 adult readers on targets of 4 to 12 letters.

What Happens in a Real Reader's Eyes

Read the columns rather than the rows. Length moves skip rate by about 19 percentage points; predictability moves it by 2 to 3. The mechanism is physical: parafoveal vision can identify a 4-letter word without a direct fixation and cannot do the same for an 11-letter word.

Long words cost about 10 milliseconds per letter

Pollatsek et al. (2008) held sentence context constant and varied adjective length. Gaze duration was 42 ms longer on 7 to 9 letter adjectives than on 3 to 4 letter ones, roughly 10.5 ms per letter. On the same materials, the low-frequency effect was 20 ms. So length produced about twice the time cost of frequency there, and it is evidence about reading time, not comprehension.

The syllable finding that did not replicate

Here a claim repeated across the web needs correcting. Fitzsimmons & Drieghe (2011) ran the cleanest possible test: all target words exactly five letters, matched on frequency and orthographic neighbours, varying only in syllable count. Monosyllabic words were skipped 5.6% more often, with no difference in fixation or gaze durations.

That result failed to replicate. Drieghe and colleagues (2019, Psychonomic Bulletin & Review) ran a higher-powered version, found Bayesian evidence for a null effect, and concluded the original was probably a Type I error. The honest summary is that syllable count, isolated from letter count and frequency, has no demonstrated effect on eye movements at all. Yet Flesch, Flesch-Kincaid, Fog, SMOG, and FORCAST all measure syllables.

Memory, brain data, and the ceiling on what length explains

Baddeley, Thomson and Buchanan (1975) showed that people recall lists of short words better than long ones, because long words take longer to rehearse in the phonological loop. A sentence packed with polysyllables is therefore harder to hold in mind while parsing the rest of it, so word length and sentence length are not independent burdens; they multiply.

Schuster et al. (2016, Cerebral Cortex) found more occipital activation for longer words, but the linear length effect “ceased to be significant at the whole-brain level when controlling for multiple fixation cases.” Much of the apparent neural cost is the cost of needing a second fixation to see the word.

Nahatame & Uchida (2025) ranked lexical features by random-forest importance across six eye-tracking measures in second-language reading. Word length ranked first for skipping and total fixation duration, ahead of frequency, but total variance explained was R² = .11 to .22, leaving 78% to 89% of reading behaviour unexplained. The effect does travel across writing systems: Kuperman, Schroeder & Gnetov (2024), using the Multilingual Eye-movement Corpus, found length and frequency effects comparable across 12 alphabetic languages, though the thresholds are language-specific.

The synthesis: word length has a genuine, replicable effect on eye movements and much weaker evidence behind it about comprehension. A reader who fixates a word twice instead of once has spent 40 extra milliseconds and may have understood it perfectly.

The Frequency Confound: Length Is a Proxy, Not a Cause

The finding that reframes this topic comes from the vendor with the most commercial reason to say otherwise. MetaMetrics writes in its Lexile technical report:

“Variables such as the average number of letters or syllables per word have been observed to be proxies for word frequency. There is a high negative correlation between the length of words and the frequency of word usage. Polysyllabic words are used less frequently than monosyllabic words, making word length a good proxy for the likelihood that an individual will be exposed to a word.”

Lexile therefore drops word length and uses mean log word frequency instead.

The entanglement is structural, not coincidental. Zipf’s law of abbreviation holds that frequently used words get shorter over time. Petrini et al. (2023) found the pattern across 46 languages from 14 linguistic families, with word lengths systematically below chance. Languages compress their common words, so measuring length indirectly measures familiarity.

The Proxy Chain

word length → word frequency → reader familiarity → comprehension

Every arrow loses information. Length correlates with frequency, imperfectly. Frequency correlates with your specific reader’s familiarity, imperfectly, since a cardiologist knows anticoagulant better than jib. Familiarity correlates with comprehension, imperfectly, because you can know every word in a sentence and still miss the argument. Formulas measure the first link and report a number about the last. Three lossy conversions separate what the tool sees from what you care about, which is why the score is directionally useful and never authoritative.

Modern research keeps landing on familiarity

Word length does not survive feature selection. Zhang & Lu (2025, Studies in Second Language Acquisition) built a CEFR text-levelling model from 353 candidate indices across 1,181 texts. Twenty-four predictors survived screening, and none measures word length. The top feature by a wide margin was Kuperman Age of Acquisition, a familiarity measure.

Formulas predict measured reading ease poorly. Gruteke Klein et al. (2025) evaluated readability measures against eye-movement data from 360 adult readers across 30 Guardian articles in parallel Advanced and Elementary versions, over 2.1 million word tokens. Traditional formulas, machine-learning systems, frontier language models, and commercial education systems were “all poor predictors of reading ease in English,” and “existing methods are often outperformed by word properties commonly used in psycholinguistics.” Average per-word surprisal, meaning how unexpected a word is in context, performed best. The strongest traditional formula was Coleman-Liau, the character-based one.

Across five datasets, the best traditional metric measured familiarity. Belem et al. (2025) compared reference-free metrics on over 5,000 examples from five corpora. The strongest traditional performer was “number of difficult words,” the Dale-Chall familiar-list approach. Length-based formulas ranked well below it.

The words that break the proxy in both directions

WordLettersSyllablesLength-based formulas sayReality
writ41Very easyLegal term, opaque to most readers
tort41Very easyLegal term of art
jib31Very easySailing term
ennui52EasyRare French borrowing
waive51Very easyRoutinely confused with wave
lien41Very easyUnknown outside finance and law
grandmother113DifficultUnderstood by five-year-olds
information114Very difficultAmong the most common nouns in English
hippopotamus125Very difficultKnown to most preschoolers

Caroline Jarrett and Ginny Redish put it most sharply: “I wave my hand” and “I waive my rights” score identically on every formula, and one of them is a sentence most adults cannot reliably parse.

The practical rule: prefer the familiar word first, the short word second. Usually they are the same word. When they conflict, familiarity wins.

Benchmarks: What Word Length Looks Like in Real Writing

Every article on this topic quotes an average English word length, the numbers never agree, and none of the sources publishes a methodology. So I measured it.

Method: statistics computed over the Brown Corpus (500 samples of edited American English, 15 genres, 991,429 alphabetic tokens) using the CMU Pronouncing Dictionary for syllable counts, taking the minimum count across pronunciation variants. Tokens were restricted to alphabetic forms plus internal apostrophes, excluding numerals and symbols, and apostrophes are excluded from character counts. CMUdict coverage was 98.4% overall, and syllable statistics cover matched tokens only.

GenreWordsChars/wordSyllables/word% monosyllabic% 3+ syllables% over 6 letters
Adventure57,8674.261.31575.3%5.4%15.4%
Mystery47,7834.231.32774.7%5.9%15.3%
Romance58,1844.241.33474.5%6.1%15.5%
Fiction57,9084.311.34473.9%6.4%16.4%
Science fiction11,9394.481.41671.4%9.3%19.0%
Humor18,0444.501.44269.6%10.0%20.3%
Lore95,4064.701.52366.2%12.9%23.2%
Belles lettres150,1494.731.54665.6%13.7%24.2%
News84,4264.831.56062.9%13.5%25.0%
Editorial53,2074.761.56164.6%14.4%24.4%
Academic (“learned”)157,5264.991.64761.5%17.4%28.2%
Government60,1985.121.69759.1%19.8%30.7%
ALL GENRES991,4294.691.51866.5%12.7%23.1%

Read the top and bottom rows together. The gap between genre extremes is 0.86 characters and 0.382 syllables per word. Put that into Flesch Reading Ease and you get 32.3 points of difference from word choice alone, the distance from “Difficult” to “Easy.” Government prose is not hard because of long sentences. It is hard because it uses more than three times as many 3+ syllable words as adventure fiction (19.8% against 5.4%).

As an independent cross-check, Peter Norvig’s analysis of Google Books Ngrams (743.8 billion tokens) puts average word length at 4.79 letters. Wylie Communications reports the BBC at 4.7 characters per word, the Wall Street Journal at 4.8, and the New York Times at 4.9. The publications that live or die by being read sit right at the corpus average.

Famous texts

Famous texts, and the finding that ends the “short means simple” argument

TextWordsChars/wordSyllables/word% monosyllabic
Alice in Wonderland27,3333.941.26478.2%
Hamlet (Shakespeare)30,2664.041.19284.4%
King James Bible791,8424.071.25279.5%
Emma (Austen)161,6004.231.38271.9%
Moby-Dick218,3614.361.34774.6%
Gettysburg Address2724.221.34974.6%

Hamlet has the shortest average word length of any text measured here, at 1.192 syllables per word and 84.4% monosyllabic. Shorter than Alice in Wonderland. By any word-length metric, Hamlet is easier reading than a children’s book.

Shakespeare, the King James Bible, and Milton (1.312) cluster at the short end because they draw on the Germanic core of English: short, ancient, high-frequency words. Difficulty in these texts lives in syntax, ambiguity, allusion, and archaic senses of familiar forms, and word length cannot see any of it. One caveat the usual citation omits: CMUdict coverage for the Shakespeare texts is only about 84%, since forms such as hath and ‘tis are missing, so treat those syllable figures as indicative.

The Gettysburg Address is widely quoted as “74% monosyllabic,” and at 100% dictionary coverage this measurement confirms it at 74.6%. Inaugural addresses show the same drift over 230 years, from Washington in 1789 (1.641 syllables per word) to Biden in 2021 (1.401). That swing is worth 20.3 Flesch points, a measurable shift in American political register.

Why Your Readability Tools Disagree?

Mac et al. (2022, JAMA Network Open) tested 10 health web pages across 8 calculators and 16 calculator-formula combinations. Their finding: “the same text produced scores that varied by up to 12.9 grade reading levels even when using the same formula.” Only three combinations agreed within one grade of a manually calculated reference: Readability Studio (SMOG), SHeLL Editor (SMOG), and Microsoft Word (Flesch-Kincaid).

Syllable counters are hyphenation dictionaries in disguise

Most open-source syllable counters look a word up in CMUdict and count stressed vowel phonemes. On a miss, they fall back to a hyphenation dictionary and count break points plus one. Step two is the bug, because hyphenation points are not syllable boundaries. A hyphenation dictionary decides where a typesetter may break a line and deliberately refuses breaks that look bad, so it undercounts, and any word missing from CMUdict silently receives a typographic guess.

Verified against textstat v0.7.13:

Why Your Readability Tools Disagree

Four failure classes with real consequences. Acronyms count as one syllable, so “HTML” scores as easy as “cat” and dense technical documents get rewarded for initialisms. Numbers, currency and dates count as one syllable. Hyphenated compounds are inconsistent, with widely used libraries returning 2, 3, and 4 syllables for data-driven. Non-American spellings fall through the dictionary gap.

The spelling penalty, measured

SentenceSyllables/wordFlesch REFlesch-Kincaid
“We must utilise and organise and analyse the data.”1.55666.106.28
“We must utilize and organize and analyze the data.”1.77847.308.90

A spelling convention moves Flesch Reading Ease 18.8 points and Flesch-Kincaid 2.6 grades on a nine-word sentence of identical difficulty. Worth stating precisely: in this test only utilise fell through the dictionary, scored as one syllable against three for utilize, while organise and analyse counted correctly. The effect is per-word and dictionary-dependent rather than a blanket bias, which makes it harder to predict and easier to trip over.

How to game any readability score in ten seconds?

VersionFlesch REFlesch-KincaidARI
“The Food and Drug Administration reviewed the application on the first of March, two thousand twenty five.”55.229.7812.01
“The FDA reviewed the application on 2025-03-01.”66.795.689.66

Abbreviating gains 11.6 Flesch points and 4.1 grade levels, and produces the version that is objectively harder for a reader with low literacy or limited domain knowledge.

A second vector, verified here: changing utilise to use leaves Flesch Reading Ease completely unchanged while improving ARI by 3.8 grades and Coleman-Liau by 4.6. Any monosyllable-to-monosyllable shortening (large to big, purchase to buy) is invisible to Flesch, Fog, SMOG, and FORCAST, and fully visible to ARI, Coleman-Liau, LIX, and RIX. Run one syllable-based and one character-based formula, or you will not see half your own edits.

What each mainstream tool measures?

ToolUnderlying formulaWord-length ruleThreshold
Hemingway EditorARI (characters)Purple highlight for “words that have simpler exact synonyms”Grade 9, wordlist unpublished
Yoast SEOFlesch RE plus a “word complexity” checkOver 7 characters AND outside top 5,000 AND not capitalised10% complex words for green
GrammarlyFlesch Reading EaseAdvice onlyTarget 60+
Microsoft WordFlesch RE and Flesch-KincaidReports characters per word directlyFRE 60–70, FKGL 7.0–8.0
LanguageToolNoneNo word-length rule; flags sentences over 40 words40 words

Yoast’s rule is the best-designed word-length check in any mainstream tool, because it requires a word to be long, rare, and not a proper noun. That is a length-plus-familiarity test, closer to the research than to Flesch.

The Word Triage Framework: Keep, Cut, Define, or Hunt

Substitution lists give one instruction: replace the long word. That advice is wrong about a third of the time, because it collapses two independent dimensions into one. And the fifth cell of the matrix, the one no tool will ever flag, is where most real reader difficulty lives.

Familiar to your readerUnfamiliar to your reader
Long word① LEAVE IT ALONE information, grandmother, understand, community Your tool flags these. Ignore it. They cost a few milliseconds and nothing in comprehension.③ CUT IT if replaceable utilize, ascertain, notwithstanding ④ KEEP AND DEFINE if it is a term of art anticoagulant, res judicata, idempotent
Short word② IDEAL use, help, buy, ask, send No tool will ever tell you to change these.⑤ THE BLIND SPOT writ, tort, jib, ennui, waive, levy, cede, dram No length-based formula will flag these, and they are among the hardest words in your document.

The decision tree, per flagged word

  1. Would a typical member of your audience use this word in conversation? Yes → Quadrant ①, leave it whatever the tool says. No → continue.
  2. Does a shorter word carry the exact same meaning? Yes → Quadrant ③, cut it. No → continue.
  3. Is it a legal, medical, or technical term of art, or a term your organisation has standardised on? Yes → Quadrant ④, keep it, define it on first use, and never vary it with synonyms. No → it is probably showing off, so cut it.
  4. Then run the pass your tool cannot: scan for Quadrant ⑤, the short words that are rare, abstract, domain-specific, or confusable with a homophone.

The Quadrant ⑤ pass is the highest-value edit in this guide because it is the one no tool will prompt you to make. A quick heuristic: any word of 3 to 6 letters that you learned in a professional context, has a homophone, or you would hesitate to use with a stranger at a bus stop.

The four-signal friction check

When the matrix does not settle a word, score it on four signals. Anything scoring high on three or more needs attention regardless of length.

SignalQuestionHigher-risk pattern
LengthHow many letters or syllables?Long or polysyllabic
FamiliarityIs it common for this audience?Rare, specialist, or archaic
PredictabilityDoes the sentence prepare the reader for it?Unexpected term with little context
DensityHow many difficult words sit close together?Several unfamiliar terms in one sentence

This is an editorial model rather than a validated formula. Its value is that it captures the compounding effect averages hide: one unfamiliar word per paragraph is manageable, four in one sentence is not.

When the Longer Word Is the Right Word?

The strictest controlled-English standard sets no word-length limit

ASD-STE100 Simplified Technical English (Issue 9, January 2025) governs aerospace maintenance documentation, where a misread instruction can kill people. It is the most aggressive plain-language regime in existence: 53 writing rules, an approved dictionary of about 900 words each with one meaning and one part of speech, and roughly 1,200 unapproved words each with an approved alternative.

It contains no rule limiting word length or syllable count. Its mechanism is different: one word, one meaning. Rather than letting writers choose among begin, commence, initiate, originate, it approves start and prohibits the rest. The constraint is vocabulary size and ambiguity, not brevity. It also exempts technical terms, so “aural warning system” is fully compliant. When the standard governing aircraft manuals caps sentence length and vocabulary but places no cap on word length, the claim that shorter words are automatically clearer needs revisiting.

Consistency beats brevity, and the plain-language guides say so

Digital.gov’s page telling you to use short words also carries a section headed “Consistency counts”: “Use the same term throughout materials for the same concept. Don’t substitute synonyms that might confuse readers about whether you’re referencing the same item.” The Microsoft Writing Style Guide agrees, and sometimes prefers the longer word for precision, favouring “Because you created the table” over “Since you created the table” because since is temporally ambiguous.

If your document uses authenticate fourteen times, replacing three with log in and two with sign in to improve a score has made it worse. You lowered the syllable count and raised the question of whether these are three different operations.

Shorter synonyms break translation

Google’s developer documentation style guide states the case plainly: “If you use different names for the same thing, translators might think you’re referring to different concepts, and thus might use different translations.” Inconsistency also raises cost, “particularly when translation memory and machine translations are used as first steps.” If your content gets localised, synonym-swapping for a score has a price tag.

Simplification measurably deletes information

Devaraj et al. (ACL 2022) audited automatic text simplification for factual errors. Deletion errors appeared in 17.6% to 70.6% of outputs depending on model and dataset, and “deletion errors are far more common than insertion errors.” The human reference simplifications also contained substantial errors, so the failure mode is intrinsic to simplification itself. A 2025 analysis of health-text simplification adds the diagnosis: models “appear to have mastery over syntactic simplification and the primary hurdle is lexical complexity.” Machines shorten sentences safely; replacing words without losing meaning is the hard part. Medical-writing guidance names the specific risk, since swapping “associated with” for “causes” shortens the phrase and asserts something the evidence may not support.

The compression trap, and the rule that resolves it

Shortening words often makes documents longer. Anticoagulant (13 characters, one word) becomes “medicine that stops your blood from clotting” (43 characters, seven words). Your syllables per word drop, your Flesch score improves, and that passage grows by more than 200%. Both effects are real, and the formula sees only one.

The resolution is worth naming: define, do not delete. Write “anticoagulants (blood thinners)” on first use, then use the precise term consistently. You get the reader’s comprehension, the searcher’s vocabulary, and the clinician’s accuracy in one move.

Genuine terms of art follow the same logic. Hearsay, consideration, and res judicata carry settled judicial meanings no shorter synonym reproduces, which is why plain-language reformers in law target pseudo-legalese instead (shall, provided that, herein, and doublets like cease and desist). The Australian Government Style Manual gives the cleanest boundary rule anyone has written: “You can include technical or specialist terms if your research shows your audience uses them. But start with plain language words as the default.”

The 7-Step Word-Length Audit

1. Define the reader and the task. “Adults” is too broad. “First-time renters comparing lease terms” gives you a real basis for judging arrears, guarantor, and sublet.

2. Sample properly, then measure with two formula types. Analyse main content, not navigation or footers, and for a long page sample the introduction, a middle section, and the conclusion. Run one syllable-based formula and one character-based formula, and record raw average syllables and characters per word, which are more informative than any composite score. If both move after an edit, you made a real change.

3. Correct for the tool’s blind spots before trusting the number. Check acronyms and initialisms (counted as one syllable), numerals and currency (one syllable), URLs, non-American spellings, hyphenated compounds, headings and bullets where sentence detection fails, and proper nouns, which Gunning Fog excludes by definition but most implementations include.

4. Triage every flagged word through the matrix. Expect 30% to 50% of flagged words to belong in Quadrant ① and need no change at all.

5. Hunt Quadrant ⑤. Search for short, rare words. Check every homophone pair (waive/wave, cede/seed, principal/principle). Flag every word you first met in a professional context. Your score will not improve. Your reader’s comprehension will.

6. Check terminology consistency afterwards. This is the step that catches the damage: every concept still using one term, no technical term replaced with an approximate synonym, no causal claim strengthened by shortening, no hedge deleted rather than shortened, and the terms your audience searches for still present.

7. Validate with a reader, not a formula. Ask five people from your audience to read a passage aloud, since every stumble is a real difficulty signal regardless of word length. Ask one to explain the key point back in their own words. And check your site search and support tickets, which are your real familiar-word list and better than any published one.

ISO 24495-1:2023, the first international plain-language standard, builds its principles around whether readers can find, understand, and use information, stating that plain language “focuses on how successfully readers can use the document rather than on mechanical measures such as readability formulas.” Its word-choice clause is headed “Choose familiar words.”

Targets by Audience and Content Type

Treat these as ceilings worth noticing rather than scores to optimise.

Content typeFlesch-KincaidSyllables/wordChars/word3+ syllable words
Children and early language learners3–5≤1.35≤4.4<6%
Patient-facing health information6–8≤1.45≤4.8<8%
Consent forms, benefits, government services6–8≤1.45≤4.8<8%
Consumer web content and marketing7–9≤1.50≤5.0<12%
News and general journalism8–11~1.56~4.8~13%
B2B and professional content10–12~1.60~5.1<17%
Technical documentation10–13~1.65~5.2<20%
Academic and scientific writing13+~1.65+~5.3~17%+

Two caveats matter more than the numbers.

Grade levels were fitted to school textbook placements. Flesch’s original criterion was the grade level of a child who could answer three-quarters of the comprehension questions about a passage. Applying that scale to adult professionals is an extrapolation, so a “grade 12” score does not mean your reader needs a diploma.

The gap between adjacent targets is smaller than your measurement error. Moving from grade 8 to grade 6 requires roughly 0.17 syllables per word. Given that the same text can score 12.9 grades apart across tools, a mandate distinguishing grade 6 from grade 8 is measuring noise.

Verified Substitution Tables

All pairs below come from live government or standards sources, chiefly GOV.UK, digital.gov, and the European Commission.

Instead ofWriteInstead ofWrite
utilize / utiliseuseascertainfind out
assist / assistancehelpcommencestart, begin
approximatelyaboutpurchasebuy
terminateend, stopendeavourtry
additionalmore, extraobtainget
in order totoprior tobefore
subsequentlyafter, thennotwithstandingdespite
in the event thatifuntil such time asuntil
due to the fact thatbecausedespite the fact thatalthough
at this point in timenowpresentlynow
facilitatehelp, make easierimplementdo, start, apply
demonstrateshowanticipateexpect
pursuant tounderwith regard toabout
is unable tocannoton a monthly basismonthly
remunerationpaymethodologymethod

Phrase-level cuts do more work than single words. Conduct an analysis of the data becomes analyse the data. Are responsible for management of the program becomes manage the program. Carry out an evaluation of becomes evaluate. In view of the fact that becomes as. Within the framework of becomes under.

Kill doublets. Due and payable becomes due. Cease and desist becomes stop.

Cut empty modifiers. Absolutely, actually, completely, really, quite, totally, very.

Fix “shall.” US federal guidance calls it ambiguous and rare in everyday speech. Replace with must (obligation), must not (prohibition), may (discretion), should (recommendation).

Break noun strings. Digital.gov: “Readability suffers when three words that are ordinarily separate nouns appear one after the other.” Their example is the best in the genre: “Underground mine worker safety protection procedures development” becomes “Developing procedures to protect the safety of workers in underground mines.” Notice what happened. The fix made the text longer and reused the same long words, because word length was never the problem. No readability formula can detect a noun string.

Before and after, with the reason attached

Harder wordingClearer wordingWhy the revision helps
Utilize the portal to commence registration.Use the website to start signing up.Replaces formal verbs with familiar actions
Assistance is available subsequent to submission.Help is available after you submit the form.Names the action instead of nominalising it
Medication adherence is essential.Take your medicine exactly as prescribed.Converts health jargon into an instruction
The mutation is associated with the condition.The mutation is linked to the condition.Simplifies without asserting causation
The patient has hypertension.The patient has hypertension (high blood pressure).Keeps the precise term and defines it

Note the last two rows: neither revision shortens anything, and both improve the text. Now compare two 11-word sentences. “The organization will implement modifications to facilitate utilization of the application” runs 80 letters, or 7.27 per word. “The team will make changes to help people use the app” runs 43, or 3.91 per word. Same word count, 46% fewer characters. Word count alone cannot describe reading effort.

Accessibility: WCAG, Dyslexia, and Listening Audiences

What WCAG actually requires?

SC 3.1.5 Reading Level (Level AAA) requires a supplemental version when text demands reading ability beyond lower secondary education level, roughly nine years of schooling. The W3C pairs length with familiarity carefully: text using “short, common words and short sentences” is easier to decode than text using “long sentences and long or unfamiliar words.” W3C does not claim length alone.

SC 3.1.3 Unusual Words (Level AAA) requires a mechanism for identifying definitions of words used in unusual or restricted ways, including idioms and jargon. It does not mention word length at all.

Both are Level AAA, a tier almost no organisation claims conformance with, and the Working Group acknowledged it “could not find a way to test whether this had been achieved,” adopting reading level as “a way to introduce testability into a success criterion.”

The conformant action for a necessary long technical term is to define it, not delete it. In practice: define unusual terms at first use, expand abbreviations before using the short form, provide a glossary for specialist content, keep labels consistent, and offer a plain-language summary when advanced terminology is unavoidable.

Dyslexia: the finding that gets misreported

The best available decomposition comes from Rydel-Johnston & Kafkas (2025), analysing 322,776 word tokens from 57 participants, 19 of them dyslexic. The corpus is Danish, so treat magnitudes as indicative for English.

Factor (moving from Q1 to Q3)ControlsDyslexic readersAmplification
Word length+98.99 ms+108.87 ms1.08×
Word frequency−17.22 ms−25.66 ms1.34×
Surprisal (predictability)+10.65 ms+24.98 ms2.32×

Word length produces by far the largest raw cost for both groups, yet dyslexic readers are barely more sensitive to it than controls. What they are disproportionately sensitive to is unpredictability, and the authors’ counterfactual simplification closed only about a third of the dyslexic/control gap.

A further point substitution advice misses: dyslexia intervention treats long words as a decoding problem solved by teaching syllable and morpheme segmentation, not by choosing shorter synonyms. A short word with irregular spelling (yacht, choir, colonel) can be harder to decode than a long regular one (understanding). Regularity, not length, drives decoding.

Listening audiences and second-language readers

Screen-reader users hear text linearly and cannot skim, so a pile-up of polysyllables costs them time with no skipping available to recover it. The same constraint governs podcast scripts, voice interfaces, and text-to-speech, since listeners cannot re-read, which revives the Baddeley memory limit. Audio scripts should run shorter than page copy on both word and sentence length.

For second-language readers, English’s length-difficulty correlation is partly a historical accident. English pairs short Germanic words with long Latinate ones for similar concepts (ask/inquire, begin/commence, help/assist), so “short” and “everyday” tend to coincide, and that coincidence does not transfer; in Japanese the everyday word is often the longer native form. Research on learners keeps landing on age of acquisition and frequency rather than length, and the residual difficulty is usually syntax and idiom. A short idiom such as rule out can be harder than the longer literal remove from consideration.

Why this is worth the effort?

PIAAC Cycle 2 (2023) found 28% of US adults performing at literacy Level 1 or below, up from 19% in 2017, with the US mean at about 258 against an international average of 260. Across 31 participating countries, 26% of adults were at Level 1 or below, a tier at which readers can access a single piece of information in relatively short texts.

Does plain-language rewriting work? The most honest answer comes from Sayfi et al. (2024, Journal of Clinical Epidemiology), a randomised trial with 488 adults. A plain-language WHO recommendation improved correct understanding by 19.8% (p < 0.001), while a CDC recommendation improved by a non-significant 3.9% (p = 0.096). Secondary outcomes were positive across both. Read both results: the intervention bundled word choice with structure and framing, so anyone who tells you shortening words reliably improves comprehension is overstating the evidence, and anyone who says it never helps is understating it.

Word Length in Other Languages

Almost every guide on this topic is silently English-only. Three things break at a language boundary.

Compounding. German writes Geschwindigkeitsbegrenzung as one word where English uses two. Character-based formulas read this as extreme difficulty, while a German reader parses it by morpheme boundaries almost as easily as the phrase. Finnish, Turkish, and Hungarian agglutination produce the same artefact. English’s short-word, plain-register alignment is itself an accident of the Norman Conquest leaving two vocabularies side by side, so optimising length as a proxy for register only works where that coincidence holds.

The coefficients were refitted, and they changed a lot.

LanguageAdaptationBaseSentence coefficientWord-length coefficient
EnglishFlesch (1948)206.8351.01584.6
FrenchKandel & Moles2071.01573.6
SpanishFernández-Huerta206.841.0260.0
GermanAmstad1801.058.5

The word-length penalty in Spanish is 29% smaller than in English. Running an English formula on Spanish overstates difficulty, and running any of them on Japanese, Chinese, or Arabic is meaningless without a validated local adaptation.

Frequency and length decouple differently by language. Levshina (2022) found frequency beat informativity as a predictor of word length in Finnish and Hungarian, while informativity won in Arabic, Spanish, Turkish, and Indonesian.

Practical rule: for non-English text use LIX or RIX, which were designed to be language-neutral, or a properly refitted local adaptation. Never compare an English score to its translation’s score and conclude anything about the translation.

There is no credible evidence that readability or word length is a Google ranking factor. The confusion comes mostly from Yoast’s traffic-light interface, which sits inside an SEO plugin and therefore looks like an SEO signal.

Google’s John Mueller addressed this in a January 2018 Webmaster Central hangout:

“From an SEO point of view, it’s probably not something that you need to focus on, in the sense that, as far as I know, we don’t have kind of these basic algorithms that just count words and try to figure out what the reading level is based on these existing algorithms. But it is something that you should figure out for your audience.”

His example goes straight to word length. A medical site can be technically correct and use “medical words that are 20 characters long,” and “if nobody’s searching for those long words, then nobody’s going to find your content.” Word length matters for SEO through search demand, not through a readability score.

Two supporting data points. Google’s documentation on helpful content never mentions readability, reading level, or word length, and rebuts the related myth directly: “Are you writing to a particular word count because you’ve heard or read that Google has a preferred word count? (No, we don’t.)” And Portent’s study of 756,297 pages across 30,000 queries, extended to 5,813,565 pages, found no correlation between reading level and rank position, with top-30 content averaging an 11th-grade level. Yoast is candid about its own light too: improving readability will not make rankings “immediately soar,” and its thresholds were tuned so roughly 35% of a 75-article sample scored green, a distributional target rather than a discovered optimum.

Where readability does earn its keep?

Readability is not a ranking factor, and it drives what follows from ranking. Clear writing improves engagement signals such as time on page and task completion. Short, well-formed answer sentences are disproportionately what Google lifts into featured snippets. And AI answer engines retrieve and quote self-contained passages, so a clean definition sentence carrying both the common term and the precise term is quotable while a 45-word clause is not.

Treat claims that readability scores drive citation rates in AI answers with scepticism, since no methodologically transparent study has established that. What is defensible is Mueller’s logic: if your page never uses the words your audience uses, no amount of syllable-trimming helps. The practical rule: do not delete the technical term, add the common one.

Two related worries deserve a direct answer. Simplifying will not make you sound less professional, since plain-language research with expert readers, including scientists and judges, consistently finds experts rate the plainly written version higher. And short words do not create thin content, because word length and word count are separate variables.

Frequently Asked Questions

How does word length affect readability?

Two ways. Mechanically, word length is the highest-weighted variable in most readability formulas: Flesch Reading Ease carries a coefficient of −84.6 on syllables per word, so +0.1 syllables costs 8.46 points. Cognitively, longer words are skipped less often by the eye (29.3% against 10.7%) and cost roughly 10 ms of extra gaze time per letter. The cognitive cost is mostly eye movement rather than comprehension, and it largely reflects the fact that longer words tend to be rarer words.

Does word length affect comprehension, or just readability scores?

Mostly the scores. Experiments holding frequency and context constant find that word length affects where the eyes land far more than how long processing takes, and the best-known syllable-only experiment failed to replicate in a higher-powered 2019 study that found Bayesian evidence for a null effect. Comprehension gains from plain-language rewriting are real but inconsistent: a 2024 randomised trial found +19.8% understanding on one document and a non-significant +3.9% on another.

What is a good average word length, and how long is the average English word?

Edited general English prose runs 4.69 characters and 1.518 syllables per word, with 66.5% of words monosyllabic, measured here across 991,429 words of the Brown Corpus. Peter Norvig’s Google Books Ngrams analysis gives 4.79 letters per token, so the published 4.7 to 4.9 range reflects corpus differences rather than error. Aim for 1.40 to 1.50 syllables per word in consumer-facing content. Fiction sits at 1.33, news at 1.56, academic writing at 1.65, government prose at 1.70.

How do I calculate average word length?

Count the letters in your sample, excluding spaces and punctuation, and divide by the number of words. Check how your tool handles contractions, hyphens, numerals, and abbreviations, and use identical settings across drafts.

How many syllables make a word hard?

There is no agreed answer, which is itself informative. Gunning Fog and SMOG use 3+ syllables, some tools use 4+, LIX uses more than 6 letters, and Yoast requires a word to be over 7 characters, outside the top 5,000, and not capitalised. Dale-Chall, Spache, Lexile, and ASD-STE100 ignore length entirely. Treat 3+ syllables as a review trigger, not proof of difficulty.

Are long words always bad, and are short words always easier?

No to both, and the second is the more consequential error. Familiar long words such as information and understand read easily through whole-word recognition, while writ, tort, jib, ennui, and dram are short and opaque. Grandmother and hippopotamus are long and understood by children. Every length-based formula gets both cases backwards, and “I wave my hand” scores identically to “I waive my rights.”

Is word length or sentence length more important?

For your score, word length by a wide margin: +0.1 syllables per word equals adding 8.3 words to every sentence in Flesch Reading Ease. In the controlled test above, word choice accounted for 99.3% of a 93-point improvement while sentence length moved the wrong way. For comprehension the picture is less clear, but the order holds: audit vocabulary before splitting sentences. The two also compound, since long words consume the working memory a long sentence needs.

Why are shorter words usually easier to read?

Three mechanisms and one confound. Real: short words fit within the span peripheral vision can identify without a direct fixation, so they get skipped more; each extra letter adds roughly 10 ms of gaze time; and they load working memory less during rehearsal. The confound, larger than all three: short words are far more likely to be high-frequency words readers already know.

Why do readability tools give different scores for the same text?

Because they disagree about counting. A JAMA Network Open study of 8 calculators and 16 calculator-formula combinations found the same text scoring up to 12.9 grade levels apart even using the same formula. Causes include syllable counters falling back to hyphenation dictionaries, acronyms and numbers counting as one syllable, hyphenated compounds counted three different ways, and inconsistent sentence detection.

Why does my British English text score better than American?

The CMU Pronouncing Dictionary holds American pronunciations, so some -ise and -our spellings miss and fall through to a hyphenation fallback that returns one syllable. In testing here, one such word moved Flesch Reading Ease 18.8 points and Flesch-Kincaid 2.6 grades in a nine-word sentence with no change in actual difficulty. UK, Australian, Canadian, and Indian English pages are systematically over-credited.

What readability score should I aim for?

Less precisely than you have been told. Microsoft suggests Flesch Reading Ease 60 to 70 and Flesch-Kincaid 7.0 to 8.0. Hemingway defaults to grade 9. AMA recommends 6th grade for patient materials, NIH recommends 8th. But the gap between a grade 6 and grade 8 target is about 0.17 syllables per word, smaller than the disagreement between two tools measuring the same text. Score against your audience band, not a universal number.

Can I use AI to shorten my words?

With caution. Devaraj et al. (ACL 2022) found automatic simplification systems producing deletion errors in 17.6% to 70.6% of outputs depending on model and dataset, with substitution errors largely introduced by the models. A 2025 analysis concluded models have mastery over syntactic simplification while “the primary hurdle is lexical complexity.” Verify every factual claim, hedge, and technical term in the output.

When is it acceptable to use long words, and should academic writers avoid them?

Four cases justify a long word. It is long but familiar. It is a genuine term of art with no exact shorter equivalent, in which case WCAG asks for a definition rather than deletion. Consistency requires it, since federal plain-language guidance itself warns against substituting synonyms. Or your audience searches for it. Academic and technical writers should keep the accepted terms that carry needed meaning and simplify the verbs, transitions, and sentence structures around them: photosynthesis stays, it is important to note that goes. Worth remembering that ASD-STE100, used for aircraft maintenance manuals, caps sentence length and vocabulary size but sets no limit on word length.

Does word length matter for SEO, and do short words make content look thin?

Word length is not a ranking factor. Google’s John Mueller: “we don’t have kind of these basic algorithms that just count words and try to figure out what the reading level is.” Portent’s 5.8-million-page study found no correlation between reading level and rank. Word length matters through search demand instead: if nobody searches your 20-character technical term, nobody finds your page. And short words do not create thin content, since word length and word count are separate variables.

How do I check readability in Microsoft Word?

Word 365 and Word for the web: Home or Review, then Editor, then Document stats. Word 2016 to 2021 on Windows: File, Options, Proofing, tick “Check grammar with spelling” and “Show readability statistics,” then run spell check. Mac: Word menu, Preferences, Spelling & Grammar, enable both, then Review, Editor. Word’s Flesch-Kincaid was one of only three combinations in the JAMA Network Open comparison that agreed within one grade of manual scoring.

Do readability formulas work for languages other than English?

Not without refitting. The Flesch word-length coefficient is 84.6 for English, 73.6 for French, 60.0 for Spanish, and 58.5 for German. Compounding languages get penalised for morphology native readers parse easily. Use LIX or RIX, or a validated local adaptation, and never compare an English score to its translation’s score.

What should I use instead of a readability formula?

Not nothing, which is the answer that has made this critique easy to dismiss for forty years. Use formulas as a smoke alarm: a score far outside your genre benchmark means look at the text. Then add what formulas cannot see: a familiarity check rather than a length check, the Quadrant ⑤ pass, read-aloud testing with five real readers, a teach-back check, and a terminology consistency check after any word-level editing.