What Readability Tests Don't Tell You
Readability tests are perceived as objective, easy to implement, and inexpensive. Although widely used within organizations, they are often misunderstood and are unreliable and of little practical use.
What is the purpose of a readability test?
So-called " objective " readability tests are used to measure a text's level of difficulty. At least, in theory.
There are several variations. The best-known readability test is the Flesch-Kincaid test. It was developed by the U.S. military to provide a readability score that is easy to... read. For example, the test indicates that your text is readable and understandable to someone with reading skills at the 9th-grade level or higher.
This test is based on the idea that the longer the words, sentences, and paragraphs are, the harder the text is to understand. Using a formula, each text is assigned a readability score in a matter of seconds—a score that is supposed to represent the minimum level of education required to understand it. Sounds appealing, doesn't it?
The Trap of Ease
In practice, readability tests are of very limited use. To understand why, read the following two excerpts, one after the other:
Text 1
The calls of the two sexes are very different. Learn to recognize these signs of affection, as this will help you identify mated pairs and understand other interactions. Since there is no difference in their plumage, you must rely on their behavioral differences. Sometimes, when a pair flies overhead, you’ll hear one call followed by the other’s, and this will help you tell the two sounds apart. It’s also important to distinguish between the sexes. One of the most extraordinary aspects of the courtship display is that the male and female alternate their calls in such a well-orchestrated manner that their entire chorus sounds as if it were performed by a single bird. If you are near a pair of Canada geese during the breeding season, you are very likely to witness a remarkable courtship display. The male’s call is deep and consists of two syllables: ahonk; the female’s is higher-pitched and usually consists of a single syllable: hink. It is a series of visual and auditory signals that a pair exchanges whenever they reunite after being separated.
Text 2
If you’re near a pair of Canada geese during the breeding season, you’re likely to witness a remarkable greeting ritual. It’s a series of visual and auditory signals that a pair exchanges whenever they reunite after being separated. Learn to recognize these displays of affection, as this will help you identify mated pairs and understand other interactions. It’s also important to distinguish between the sexes. Since there are no differences in their plumage, you must rely on their behavioral differences. The calls of the two sexes are very different. The male’s call is low-pitched and consists of two syllables: ahonk; the female’s is higher-pitched and usually consists of a single syllable: hink. One of the most extraordinary aspects of the greeting ritual is that the male and female alternate their calls in such a well-orchestrated manner that their entire chorus sounds as if it were being performed by a single bird. Sometimes, when a pair flies overhead, you’ll hear one call followed by the other’s, and this will help you tell the two sounds apart.
Score achieved: 9th grade
In these two excerpts, the factors that influence the readability score are exactly the same:
9 sentences
An average of 21.4 words per sentence
An average of 5 characters per word
According to a Flesch-Kincaid readability test, texts 1 and 2 are equally readable. Any person with reading skills at the 9th-grade level will be able to understand both texts with equal ease—at least, that’s what the formula tells us. But to us humans, the first text is incoherent, while the second is easy to understand.
The clarity of a text does not depend solely on everyday words and short sentences. It also depends on semantic elements—that is, elements related to the text’s meaning. The order in which ideas are presented thus plays a crucial role in helping your reader understand the context. Providing clear context or placing the main idea first, for example, greatly facilitates understanding of the details that follow.
What Readability Tests Don't Tell You
Plain Language
From the 1940s to the late 1970s, advocates of plain language focused on constructing simple sentences. They developed writing guidelines that are still relevant today:
Eliminate complex sentences by simplifying the syntax,
Use the active voice whenever possible,
Use concrete, everyday words,
Replace difficult words with simpler ones, or explain them if it is impossible to replace them, or
Use short sentences.
All too often, plain language is (wrongly) reduced to simple syntax and vocabulary. Aside from word and sentence length, readability tests don’t even take these basic elements into account. The Flesch-Kincaid test cannot identify subjects, verbs, or objects, detect complex syntax or the passive voice, or recognize jargon.
Information Design
Readability tests do not take into account internal aspects of the text that pertain to information design:
The readability of the text, also known as reading ergonomics,
The use of meaningful titles,
The hierarchy of information, which is reflected in particular by clearly distinct headline levels,
Bulleted lists like the one you're reading right now, and finally
Visual elements such as diagrams, illustrations, and charts, which facilitate a more immediate understanding of the meaning (the reader does not have to mentally construct a spatio-temporal framework while trying to grasp the precise meaning conveyed by the text).
Semantic and Contextual Aspects
Readability tests also do not take into account the semantic and discursive aspects of a text:
Communication objectives,
The selection of messages in light of these objectives,
The structure of the text and the order of ideas, or
The cohesion and coherence of the text as a whole.
Finally, readability tests do not take into account contextual elements that convey meaning to the reader, such as:
The relationship between the sender and the receiver,
The sociocultural context,
Power or authority relationships,
And many other aspects that require empathy in order to communicate more effectively.
Are there any tests that should be avoided?
Several studies show that readability tests are sorely lacking in reliability.
Unintended Consequences
Essentially, writers who must achieve a target readability score tend to write for the test rather than for the reader. In a 1980 study that examined four other formulas in addition to the Flesch-Kincaid formula, a U.S. Army researcher named Kern reached the following conclusions:
Readability tests do not ensure that texts are tailored to the reading level of the target audience.
Rewriting texts to achieve a score that corresponds to a lower reading level does not improve readers' comprehension of the text.
Requiring that a text achieve a certain readability score causes writers to focus on that score rather than on organizing the material to meet the reader's information needs.
Lack of reliability
In 2007, Watson conducted a comprehensive analysis of online readability tests. His conclusions:
For the same text, the results could vary by as much as 7 grade levels (which is the difference between 7th grade and pre-college level).
According to him, the most reliable test is the one included with Microsoft Word (not available in French).
In 2017, Professor Zhou’s team demonstrated that, for the same text and using, in principle, the same formula, different readability tests available online yield widely varying results. The reason lies in how each tool defines paragraphs, sentences, and words. In particular, tests that are supposed to use the same formula do not process the text in the same way:
hyphens,
the slashes,
numbers,
abbreviations,
acronyms,
URLs,
the dates,
periods, semicolons, colons, etc.
For short texts, a single unusual aspect can have a major—but unrepresentative—impact on the readability score. In other words, online readability tests are particularly unreliable for most written communications today (web content, letters and notices, forms, etc.).
Despite all these shortcomings, readability tests have been highly successful in both the public and private sectors. Easy to implement and mistakenly perceived as objective, readability tests have thus become very widespread. Some laws even require a target reading level for certain regulated documents, which is far from a guarantee of the quality of the texts.
What should you take away from this?
As you can see, in our view, so-called “objective” readability tests should generally be avoided.
The readability of a text is not limited to sentence length and the number of characters in words. It also depends on contextual factors, such as the relationship between the writer and the reader; semantic factors, such as the apparent organization of the text and its internal logic; and ergonomic factors related to information design. Readability tests do not take all of these elements into account.
Readability tests are often misused. But even if they were used strictly to measure sentence and word length, different readability tests yield different results, making them unreliable for predicting the reading level required for comprehension.
Readability tests are particularly unreliable when it comes to evaluating short texts—that is, most everyday written communication.
To ensure the quality of your writing, you should test it with readers who are representative of your target audience. At the very least, have someone outside your field proofread it. If you wish to organize afocus group, always keep in mind that reading conditions are skewed in such a setting: readers’ attention is directed toward the document, they are confined for a certain period of time, and they are sometimes paid or compensated, which requires a sustained level of concentration that is not representative of what happens in everyday life. In practice, readers’ motivation and level of concentration vary widely.
Sources
This post is inspired by Karen A. Schriver’s excellent article , “Plain Language in the U.S. Gains Momentum: 1940–2015, ” IEEE Transactions on Professional Communications, vol. 60, no. 4, pp. 343–383, December 2017. (Award-winning article—highly recommended reading!)
R. P. Kern, “Usefulness of Readability Formulas for Achieving Army Readability Objectives: Research and State-of-the-Art Applied to the Army’s Problem, ” U.S. Army Research Institute of Behavioral and Social Sciences, Alexandria, No. TR 437, 1980.
C. Watson, “Write Better: A Comparison of Online Readability Testing Tools, ” Smiley Cat Web Design Blog, 2007.
S. Zhou, H. Jeong, and P. A. Green, “How consistent are the best-known readability equations in estimating the readability of design standards? ”, IEEE Transactions on Professional Communications, vol. 60, no. 1, pp. 97–111, Mar. 2017.
W. Lidwell, K. Holden, J. Butler, “Universal Principles of Design,” *Readability*, Rockport Publishers, 2010, .
Learn More
Discover our Clarity Assessment service for an expert analysis of the sources of complexity in your content.
Sign up for our training course, “The Art of Clear Communication and Information Design.”