What Readability Tests Don't Tell You

Readability tests are perceived as objective, easy to implement, and inexpensive. Although widely used within organizations, they are often misunderstood and are unreliable and of little practical use.


What is the purpose of a readability test?

So-called " objective " readability tests are used to measure a text's level of difficulty. At least, in theory.

There are several variations. The best-known readability test is the Flesch-Kincaid test. It was developed by the U.S. military to provide a readability score that is easy to... read. For example, the test indicates that your text is readable and understandable to someone with reading skills at the 9th-grade level or higher.

This test is based on the idea that the longer the words, sentences, and paragraphs are, the harder the text is to understand. Using a formula, each text is assigned a readability score in a matter of seconds—a score that is supposed to represent the minimum level of education required to understand it. Sounds appealing, doesn't it?

Did you know?

The Flesch-Kincaid test was designed to simplify military communications

The Flesch test, developed in 1948, produced scores ranging from 0 to 100. Its methodology has since been criticized. Flesch tested his formula on texts from the 1940s intended for an adult audience, such as excerpts from newspapers and magazines. However, his formula was based on a standardized reading comprehension test for children, developed in 1926. Furthermore, this standardized reading comprehension test relied on a multiple-choice questionnaire—a method of assessing text comprehension that may have been acceptable in the 1940s but is heavily criticized today. Nevertheless, this readability test gained traction in public administration, largely because it was perceived as objective.

In 1975, the U.S. Navy funded Professor Kincaid’s team to make the scores produced by the Flesch test even easier to interpret. Up to 30 percent of new recruits in the military had a reading level equivalent to 7th grade. Kincaid’s team therefore developed a new score, this time expressed as a grade level. Another improvement was that this test was based on a sample of texts produced by the military, which was more representative of the documents that service members were required to read. This test was later criticized, notably by a Navy researcher named Kern. In Kern’s view, giving writers a specific grade-level target led them to write for the test rather than with the target audience’s information needs in mind. The texts produced in this way were not better understood, even though they achieved the target readability scores.

The Trap of Ease

In practice, readability tests are of very limited use. To understand why, read the following two excerpts, one after the other:

Text 1

The calls of the two sexes are very different. Learn to recognize these signs of affection, as this will help you identify mated pairs and understand other interactions. Since there is no difference in their plumage, you must rely on their behavioral differences. Sometimes, when a pair flies overhead, you’ll hear one call followed by the other’s, and this will help you tell the two sounds apart. It’s also important to distinguish between the sexes. One of the most extraordinary aspects of the courtship display is that the male and female alternate their calls in such a well-orchestrated manner that their entire chorus sounds as if it were performed by a single bird. If you are near a pair of Canada geese during the breeding season, you are very likely to witness a remarkable courtship display. The male’s call is deep and consists of two syllables: ahonk; the female’s is higher-pitched and usually consists of a single syllable: hink. It is a series of visual and auditory signals that a pair exchanges whenever they reunite after being separated.

Text 2

If you’re near a pair of Canada geese during the breeding season, you’re likely to witness a remarkable greeting ritual. It’s a series of visual and auditory signals that a pair exchanges whenever they reunite after being separated. Learn to recognize these displays of affection, as this will help you identify mated pairs and understand other interactions. It’s also important to distinguish between the sexes. Since there are no differences in their plumage, you must rely on their behavioral differences. The calls of the two sexes are very different. The male’s call is low-pitched and consists of two syllables: ahonk; the female’s is higher-pitched and usually consists of a single syllable: hink. One of the most extraordinary aspects of the greeting ritual is that the male and female alternate their calls in such a well-orchestrated manner that their entire chorus sounds as if it were being performed by a single bird. Sometimes, when a pair flies overhead, you’ll hear one call followed by the other’s, and this will help you tell the two sounds apart.

Score achieved: 9th grade

In these two excerpts, the factors that influence the readability score are exactly the same:

  • 9 sentences

  • An average of 21.4 words per sentence

  • An average of 5 characters per word

According to a Flesch-Kincaid readability test, texts 1 and 2 are equally readable. Any person with reading skills at the 9th-grade level will be able to understand both texts with equal ease—at least, that’s what the formula tells us. But to us humans, the first text is incoherent, while the second is easy to understand. 

The clarity of a text does not depend solely on everyday words and short sentences. It also depends on semantic elements—that is, elements related to the text’s meaning. The order in which ideas are presented thus plays a crucial role in helping your reader understand the context. Providing clear context or placing the main idea first, for example, greatly facilitates understanding of the details that follow.

What Readability Tests Don't Tell You

Plain Language

From the 1940s to the late 1970s, advocates of plain language focused on constructing simple sentences. They developed writing guidelines that are still relevant today:

  • Eliminate complex sentences by simplifying the syntax,

  • Use the active voice whenever possible,

  • Use concrete, everyday words,

  • Replace difficult words with simpler ones, or explain them if it is impossible to replace them, or

  • Use short sentences.

All too often, plain language is (wrongly) reduced to simple syntax and vocabulary. Aside from word and sentence length, readability tests don’t even take these basic elements into account. The Flesch-Kincaid test cannot identify subjects, verbs, or objects, detect complex syntax or the passive voice, or recognize jargon.

Information Design

Readability tests do not take into account internal aspects of the text that pertain to information design:

  • The readability of the text, also known as reading ergonomics,

  • The use of meaningful titles,

  • The hierarchy of information, which is reflected in particular by clearly distinct headline levels,

  • Bulleted lists like the one you're reading right now, and finally

  • Visual elements such as diagrams, illustrations, and charts, which facilitate a more immediate understanding of the meaning (the reader does not have to mentally construct a spatio-temporal framework while trying to grasp the precise meaning conveyed by the text).

Semantic and Contextual Aspects

Readability tests also do not take into account the semantic and discursive aspects of a text:

  • Communication objectives,

  • The selection of messages in light of these objectives,

  • The structure of the text and the order of ideas, or

  • The cohesion and coherence of the text as a whole.

Finally, readability tests do not take into account contextual elements that convey meaning to the reader, such as:

  • The relationship between the sender and the receiver,

  • The sociocultural context,

  • Power or authority relationships,

  • And many other aspects that require empathy in order to communicate more effectively.

Are there any tests that should be avoided?

Several studies show that readability tests are sorely lacking in reliability. 

Unintended Consequences

Essentially, writers who must achieve a target readability score tend to write for the test rather than for the reader. In a 1980 study that examined four other formulas in addition to the Flesch-Kincaid formula, a U.S. Army researcher named Kern reached the following conclusions:

  • Readability tests do not ensure that texts are tailored to the reading level of the target audience.

  • Rewriting texts to achieve a score that corresponds to a lower reading level does not improve readers' comprehension of the text.

  • Requiring that a text achieve a certain readability score causes writers to focus on that score rather than on organizing the material to meet the reader's information needs.

Lack of reliability

In 2007, Watson conducted a comprehensive analysis of online readability tests. His conclusions:

  • For the same text, the results could vary by as much as 7 grade levels (which is the difference between 7th grade and pre-college level).

  • According to him, the most reliable test is the one included with Microsoft Word (not available in French).

In 2017, Professor Zhou’s team demonstrated that, for the same text and using, in principle, the same formula, different readability tests available online yield widely varying results. The reason lies in how each tool defines paragraphs, sentences, and words. In particular, tests that are supposed to use the same formula do not process the text in the same way:

  • hyphens,

  • the slashes,

  • numbers,

  • abbreviations,

  • acronyms,

  • URLs,

  • the dates,

  • periods, semicolons, colons, etc.

For short texts, a single unusual aspect can have a major—but unrepresentative—impact on the readability score. In other words, online readability tests are particularly unreliable for most written communications today (web content, letters and notices, forms, etc.).

Despite all these shortcomings, readability tests have been highly successful in both the public and private sectors. Easy to implement and mistakenly perceived as objective, readability tests have thus become very widespread. Some laws even require a target reading level for certain regulated documents, which is far from a guarantee of the quality of the texts.

What should you take away from this?

As you can see, in our view, so-called “objective” readability tests should generally be avoided.

  • The readability of a text is not limited to sentence length and the number of characters in words. It also depends on contextual factors, such as the relationship between the writer and the reader; semantic factors, such as the apparent organization of the text and its internal logic; and ergonomic factors related to information design. Readability tests do not take all of these elements into account.

  • Readability tests are often misused. But even if they were used strictly to measure sentence and word length, different readability tests yield different results, making them unreliable for predicting the reading level required for comprehension.

  • Readability tests are particularly unreliable when it comes to evaluating short texts—that is, most everyday written communication.

To ensure the quality of your writing, you should test it with readers who are representative of your target audience. At the very least, have someone outside your field proofread it. If you wish to organize afocus group, always keep in mind that reading conditions are skewed in such a setting: readers’ attention is directed toward the document, they are confined for a certain period of time, and they are sometimes paid or compensated, which requires a sustained level of concentration that is not representative of what happens in everyday life. In practice, readers’ motivation and level of concentration vary widely.

Sources

 
 

Learn More


Clément Camion, co-founder and partner

A member of the Quebec and New York bars and a graduate in political philosophy, Clément seeks to empower people by making their lives easier.

At En Clair, Clément brings his passion for innovation and his experience as a specialist in making legal concepts accessible and easy to understand.

In this capacity, he has made significant contributions to contract simplification projects for various organizations.

He co-authored the book *Clear and Useful Contracts for Consumers: Toward a New Standard* with Stéphanie Roy. He also co-authored a booklet andseveral articles on access to justice in the digital age as part of the research conducted by the Cyberjustice Laboratory in Montreal.

With his sharp and creative mind, Clément fascinates many with his many talents!

https://www.enclair.ca/equipe
Previous
Previous

A Clear and Simple Privacy Notice

Next
Next

Snow Removal Contract: First Draft