WhyHow Who it helps Reader Academy Apps & extension The scienceMagazine Quick startYour account Open the app

Testing your own comprehension

A reading comprehension test does two jobs at once: it tells you what stuck, and the act of being tested makes more of it stick. How our recall check works, and how to run an honest experiment on yourself.

Foundations desk 28 October 2025 7 min read 1,402 words
Generated abstract cover artwork for this article, drawn in the app's own geometry: grid.
Foundations · No. 06Generated artwork · grid
What this piece argues
  • Retrieval practice beats re-reading for retention — the testing effect is robust
  • Recognition tests flatter; plausible distractors are how we take the flattery out
  • A fair self-experiment needs comparable passages, honest answers and several sessions
  • One recall check is a probe of a session, not a measurement of you

The most useful thing about a comprehension test is not the score. It is that being tested changes what you remember. Retrieval practice — being made to pull something out of memory rather than look at it again — improves retention more than re-reading does, and the effect is one of the most dependable in the learning literature. So the recall check at the end of a session in our app is doing two jobs, and the score is the smaller one.

This is worth dwelling on, because it runs against a strong intuition. Re-reading feels productive: the second pass is smoother, the material feels familiar, and the feeling of familiarity gets mistaken for knowledge. Retrieval feels worse — effortful, halting, full of blanks — and produces better retention. The testing effect has been replicated across many laboratories, many kinds of material and many delays, and its practical translation is blunt: if you want to keep what you read, close the text and make yourself answer for it.

That is the honest reason a test sits at the end of our sessions at all. It would be tidy to say we built it purely as an instrument, a way to keep speed honest — and it does do that, as what words per minute actually measures argues at length. But an instrument you can game is half an instrument, and a test that also strengthens memory earns its place twice.

What the recall check actually does

The mechanism is simple to state. After a passage, the app shows you a set of items. Some are real targets — words that genuinely appeared in what you just read. The rest are plausible distractors: words that did not appear but comfortably could have, drawn to sit near the passage’s subject and register. Your job is to say which were there. The score you get, shown beside your speed, is how well you separated the two.

The distractors are the load-bearing part. A recognition test with lazy distractors is a machine for flattering its user: if the passage was about tides and the distractor is carburettor, you can reject it from the gist alone, without having read a single sentence carefully. Warm familiarity gets scored as comprehension, everyone feels good, nothing was measured. Make the distractor current, or estuary, and the gist stops helping. Now the only thing that separates target from distractor is whether you actually encoded the words on the screen, which is the thing we were trying to find out.

Why does the retrieval itself help? The most durable account is that memory is strengthened by the difficulty of the access, not by the number of exposures. Re-reading makes access easy and therefore teaches little; being made to decide, item by item, was this there? forces the memory system to do reconstructive work, and the work is what consolidates. This is also why the check is worth taking even when you expect to do badly. A failed retrieval followed by seeing the answer still beats a comfortable second pass — the effort was the treatment, and the score was only the receipt.

We should be candid about where this sits on the ladder of rigour. Recognition — did you see this? — is the easier end of memory testing; free recall, where you reconstruct the passage unprompted, is harder and tells you more. And knowing which words appeared is not the same as having followed the argument they were making; a full reading comprehension test would probe inferences, structure, the relations between claims. Ours probes encoding, with the flattery turned down as far as plausible distractors can turn it. That is a deliberate floor, not a ceiling — a check you will actually take after every session beats a better one you will not.

Two interleaved sets of similar marks, one set subtly distinguished, illustrating real targets mixed among plausible distractors in a recognition test.
Fig. 01 — targets among plausible distractorsGenerated · Present · absent · can you tell

How to test your reading comprehension fairly

Suppose you want to use this to answer a real question — does a faster setting cost me comprehension? does the paired imagery help me? does any of this beat the page? Those are experiments, and self-experiments fail in predictable ways. The failures are worth naming, because every one of them produces a result that feels like an answer.

The classic mistake is comparing incomparable passages. Read an easy passage fast and a dense one slow, and you will conclude that speed is free; read them the other way around and you will conclude it is ruinous. Both conclusions are about the passages. The second mistake is running one trial per condition and believing the difference, when a single recall score bounces around for reasons that have nothing to do with the setting — the passage suited you, your attention dipped, the distractors happened to be kind. The third is quietly wanting one answer, which shows up not in dishonest answers but in dishonest effort: attending harder in the condition you are rooting for.

  1. Fix the material

    Choose passages of similar length, difficulty and unfamiliarity — chapters of the same book work well. Unfamiliar matters: prior knowledge answers questions the reading was supposed to answer.

  2. Vary one thing

    One comparison at a time: this speed against that speed, or stream against page. Hold everything else — time of day, tiredness, material — as steady as you can manage.

  3. Alternate and repeat

    Run several sessions per condition, alternating rather than blocking, so that warming up or tiring out spreads across both sides instead of landing on one.

  4. Answer ruthlessly

    On the recall check, familiarity is not enough — if you could not say roughly where the word did its work, treat that as information, not as a technicality.

  5. Decide in advance what counts

    Pick the margin that would change your behaviour before you start. A difference you only decided was meaningful after seeing it is a story, not a result.

A note on what a result looks like when you get one. It is not a single triumphant comparison; it is a boring consistency. If a faster setting genuinely costs you nothing, that shows up as recall scores that refuse to separate across many alternated sessions, on material that gave the difference every chance to appear. If the imagery genuinely helps you, that shows up as a small persistent edge that survives your suspicion of it. Anything dramatic in a self-experiment is usually the passage, the hour or the mood — the real effects, where they exist, tend to be modest and stubborn rather than large and fragile.

The feeling of having understood is not evidence. It is the thing being tested.

The case for checking

What one score cannot tell you

Now the limits, stated as plainly as we can manage. A recall check after one passage is a probe, not a measurement of you. It samples one text, one sitting, one state of attention, through one narrow instrument. Treat any single score the way you would treat a single blood-pressure reading taken standing up in a corridor: real data, wrong occasion to draw conclusions from it.

There is also a subtler limit, which is that any fixed form of test teaches you to pass it. Practise against recognition checks for a month and some of your improvement will be test-craft: better memory for surface wording, sharper suspicion of distractors. That is not worthless — encoding words more deliberately is part of reading better — but it is narrower than comprehension, and we would rather tell you that here than have you discover it in your export file. The general problem, improvement that lives in the practised task and declines to travel, is the subject of the shape of a practice session.

The score is the by-product

Which returns us to where this started. If the recall check were a pure instrument, its value would stand or fall on measurement quality, and the caveats above would be damning. But the testing effect means the act itself is doing work regardless of what the number says: every check is a retrieval rep, and retrieval is the part of the session most likely to change what you keep. A reader who takes the check after every passage and never once looks at the score is still getting most of the benefit.

It also reframes what a bad score is for. In the instrument reading, a poor result is a verdict, and verdicts invite either discouragement or excuse-making. In the retrieval reading, a poor result is a session that did its job — it found the material that had not stuck, made you sweat over it, and marked it for the re-encounter that will now work better than the first pass did. The reader who never scores badly is either reading well within their capacity or being flattered by their test, and only one of those is worth congratulating.

So the advice lands somewhere unfashionable. Test yourself constantly and trust each result barely. The habit is where the retention comes from; the trend across many honest sessions is where the knowledge comes from; and the individual score, the number that feels like the point, is the least important thing on the screen.

A note on what this is. Signal is written in-house by the team that builds Reader Inc., so treat it as an argument rather than a review. Nothing here is medical, psychological or educational advice, and the app is not a treatment, therapy or diagnosis for any condition. Where we describe research we describe it in general terms; where we are reasoning past the evidence we say so. The app is free, runs entirely on your own device, and ships with a comprehension test switched on — which means you can check every claim we make against your own reading rather than taking our word for it.

About the artwork. Every image in Signal is generated — drawn by a program from the article it belongs to, using the same geometry, palette and stroke language as the app itself. Nothing is photographed and nobody is depicted. Each composition is deterministic: the same article always produces the same picture.