Reading, listening, and the third thing
Text-to-speech is the fastest-growing branch of the speed-reading family, and the fair thing to say about it is that it wins some arguments outright and loses others badly. Here is which — and what happens when you lock the two together.
What this piece argues
- Listening is bounded by speech rate, which starts below your silent reading speed.
- Comprehension parity holds for narrative; dense expository text still favours the page.
- Audio wins on hands, eyes, endurance and fatigue, and for some readers it is the only route.
- Locking voice to visual stream is our bet, not a finding — and it is off by default.
There is a version of speed reading that involves no reading at all. You put the words in your ears and get on with the washing-up. It works. It is the fastest-growing thing in this whole category, and for a great many people it is the only route to a book that has ever worked. It also has a ceiling, and the ceiling arrives sooner than most listeners think.
The category is broad and mostly honest about itself. Speechify and Voice Dream take a document and speak it. Every modern browser, phone and e-reader now has a read-aloud function built in and free. Audiobook publishers do the same job with a human narrator, a director and a studio, which is a different product at a different price. What they all share is a playback slider, and most people who use one for a few weeks drift upwards and settle somewhere between 1.5× and 2×.
That sounds like a doubling of reading speed, and it isn’t. Ordinary conversational speech runs somewhere near 150 words a minute, and narration is often slower still, because a narrator who rushes is a narrator you stop trusting. A competent adult reads continuous prose at roughly 200–300 wpm, a figure that has barely moved across a century of measurement. Listening therefore starts a lap behind. Double it and you arrive at something like a brisk read — a real gain over one-times playback, and nowhere near the multiple the interface implies.
The ceiling arrives sooner
Push further and two things go wrong at once. The first is that a speed slider does not only shorten the words, it shortens the gaps, and the gaps are load-bearing. Pauses at commas and clause boundaries are where a sentence is assembled into a proposition; the eye-movement literature finds the same integration cost on the page, where fixations lengthen at exactly those points. Squeeze the gaps out and you have not merely accelerated the prose. You have removed the moments in which the meaning was supposed to close. We think that is the main reason dense material degrades under compression faster than the listener notices.
The second problem is structural, and nobody markets against it. You cannot stop a stream without losing your place in it. A page holds still. When a sentence fails to resolve, your eye jumps back four words and repairs it in a fraction of a second — skilled readers do this on roughly one saccade in seven, and many are comprehension failures fixed in flight. In audio the equivalent move is a fifteen-second rewind that overshoots, then replays material you already had. The repair costs orders of magnitude more, so mostly you don’t make it. You let the sentence go and hope the next one explains it.
What listening wins
None of that is an argument against audio, and it would be dishonest to write one. Listening wins several things outright, and it wins one of them so completely that the rest of this article is a footnote to it. Reading requires your eyes and, in practice, your hands and your posture. Listening requires none of them. That is not a convenience feature. It is the difference between a commute that contains a book and a commute that doesn’t.
- Hands and eyes free. Walking, driving, cooking, queueing, exercising. Audio converts dead time into reading time, which is a larger effect than any rate increase we could offer you.
- Endurance. Most people can listen attentively for hours in a way they cannot read for hours. Ocular fatigue and neck ache have no equivalent in the ear.
- Lower effort per minute. Nothing has to be decoded. The words arrive already converted into sound, which is precisely the step that costs some readers the most.
- For readers with print difficulties it is often the only route. Dyslexia is a difficulty with decoding, not with thinking, and comprehension of spoken language is typically intact. Audio removes that bottleneck rather than easing it.
- Prosody carries structure. A good narrator’s intonation marks clause boundaries, irony and emphasis that a silent reader has to reconstruct without help.
You cannot stop a stream without losing your place in it.
The asymmetry between page and audio
Where the page still holds
The comparison research is less dramatic than either camp would like. For straightforward narrative — a story with characters, chronology and a plot that carries you — listening and reading land in roughly the same place. Comprehension is comparable, recall is comparable, and the suggestion that listening to a novel is a lesser form of having read it does not survive contact with the evidence. If your reading is mostly fiction and long-form journalism, the honest answer is that the format barely matters and you should choose whichever gets you through more of it.
The picture changes with dense expository prose. Text with structure — numbered arguments, tables, defined terms, a clause on page nine qualifying a claim on page four — favours the page, and the reason is not mysterious. It is the same asymmetry as before. Reading permits return; audio does not. A table read aloud is a list of numbers with no shape. A cross-reference read aloud is an instruction you cannot follow. Most of the advantage the page holds is simply the advantage of being able to go back.
| Dimension | Reading | Listening | Locked dual channel |
|---|---|---|---|
| Comfortable rate | 200–300 wpm; higher on familiar prose | Near 150 wpm at 1×; most settle at 1.5–2× | Set by the reader; voice leads or follows |
| Fatigue | Ocular and postural; limits long sessions | Low; sustainable for hours | Between the two, and untested over long sessions |
| Re-reading | Cheap — a backwards saccade costs a fraction of a second | Expensive — rewind overshoots, then replays | Removed by design; step-back keys only |
| Hands and eyes | Both occupied | Both free | Both occupied |
| Best-fit material | Dense, structured, argued, tabular | Narrative, interviews, familiar territory | Vocabulary-dense prose you intend to remember |
The third thing
Which brings us to the arrangement this app is actually built around, and it is neither of the two above. The app runs KittenTTS nano — a 15-million-parameter open-weight neural voice — entirely in the browser, with eight voices, no server and nothing uploaded. The voice can act as the master clock, with the display following the speech, or it can follow the reader, with your pace governing the speech. Both modes exist because which one helps depends entirely on why you turned the voice on.
What matters is the lock. Each word is spoken at the instant it appears, and the image paired with it appears at that instant too. This is not an audiobook with pictures, and the difference is not cosmetic. The bet — set out at length in the second channel — is that the phonological route and the imagery route can be made to arrive together instead of one queueing behind the other. Working memory is not one pool: there is a store for verbal-phonological material and a separate one for visual-spatial material. It is why a narrated diagram is learned better than the same diagram with the words printed beside it.
Now the part we would rather not write. That finding has a well-known partner which cuts the other way: presenting the same verbal content simultaneously as printed text and as speech can hurt rather than help, because the two have to be reconciled while competing for the same verbal resources. Our stream is printed text. Our voice speaks the same words. That is precisely the configuration the interference literature warns about, and our reason for thinking the lock might rescue it — perfect synchrony, one word, one instant, no reconciliation left to do — is our own reasoning rather than a result. Nobody has tested this combination, ourselves included.
Which one, when
The practical answer is not a ranking. Use audio for the material and the moments where audio is unbeatable: the novel on the train, the interview, the second pass through something you already half know, the hour when your eyes are needed elsewhere. Use the page for the contract, the paper, the filing, and anything with a table in it. Use both together when the material is vocabulary-dense and you intend to remember it — then check the result, because the only evidence that matters here is yours.
That check is not rhetorical. The app ends every passage with a recall test that mixes real targets with plausible distractors, and it puts the score next to the rate, which means it is perfectly capable of telling you the voice made you worse. Read a thousand words with it on and a thousand with it off, at the same pace, and take the test both times; the method in the comprehension tax covers how to do that without fooling yourself. We have no idea what your two numbers will say, and the whole thing was built so that you can find out without asking us.
A note on what this is. Signal is written in-house by the team that builds Reader Inc., so treat it as an argument rather than a review. Nothing here is medical, psychological or educational advice, and the app is not a treatment, therapy or diagnosis for any condition. Where we describe research we describe it in general terms; where we are reasoning past the evidence we say so. The app is free, runs entirely on your own device, and ships with a comprehension test switched on — which means you can check every claim we make against your own reading rather than taking our word for it.