Text to speech as a reading tool
Dyslexia is a difficulty with decoding, not with thinking — which is why a synthetic voice can hand a reader their own comprehension back. What text to speech reading does well, what it cannot practise, and why where the voice runs matters.
What this piece argues
- Comprehension by ear is typically intact in dyslexia; a voice routes around decoding
- Listening is not decoding practice, and an honest tool does not claim to be
- Cloud voices upload what you read; on-device voices keep it on the machine
- Speech can set the pace or follow the reader, and the difference changes the tool
There is a moment familiar to almost everyone who works with dyslexic adults: a document that had resisted twenty minutes of grinding effort is played aloud, and the person who could not get through it discusses it fluently, critically, in detail. Nothing about their understanding was ever missing. It was queued behind the decoding. Text to speech is the tool built on that one observation, and it deserves a more precise account than it usually gets.
Text to speech has existed for decades, mostly sounding like a fax machine with opinions. What changed recently is quality and location: neural voices now produce speech with usable prosody, and the smaller ones no longer need a data centre to do it. That combination — listenable and local — is what turned a niche accessibility feature into a serious general-purpose reading tool, and it is worth being precise about what the tool actually does.
Why it works: the bottleneck is decoding
Dyslexia is a specific learning difficulty affecting accurate and fluent word recognition, decoding and spelling. It is not a deficit of intelligence, and it is not a deficit of comprehension when language arrives by ear. Dyslexia is a difficulty with decoding, not with thinking. So a tool that converts print into speech does not fix the difficulty and does not need to: it routes the text around the narrow point and delivers it to machinery that was working all along. That is the whole mechanism, and its honesty is the best thing about it — no retraining story, no brain claims, just a change of input channel.
The practical consequences for adults are larger than the mechanism sounds. Reading volume stops being rationed by decoding stamina, which matters when the job produces forty pages a day. Long documents can run during a commute or a walk. And one under-advertised use may be the best of them: hearing your own writing read back is a formidable proofreading method, because the ear refuses to skip the missing word that the eye has been politely inserting for you.
It is also worth saying that the tool’s reach runs well past dyslexia, which is part of why it has finally been built properly. The same voice serves low vision, eyes that are simply finished at nine in the evening, and hands that are occupied with a steering wheel or a saucepan. Features built for accessibility have a habit of becoming features for everybody — kerb cuts, captions — and text to speech is following the same path. That mainstreaming matters for dyslexic adults in a quiet way: a tool everyone uses carries no flag when you use it.
What listening does not practise
Now the limit, stated as plainly as the strength. Listening is not decoding practice. The skill dyslexia affects — mapping letterforms to sounds, quickly and automatically — is exercised by doing it, ideally under structured literacy instruction, and a voice that does the decoding for you is by definition not exercising it. A decade of text to speech leaves print exactly as hard as it found it. That is not a flaw in the tool; a ramp is not walking practice either. It only becomes a flaw when someone sells the ramp as physiotherapy. The distinction between reaching comprehension and practising decoding is worth a whole article, and it has one: decoding is not comprehension.
Listening also has structural limits that have nothing to do with dyslexia. The ear is serial: you cannot skim a voice, glance back up a paragraph of it, or see the shape of an argument laid out on a page. Tables, equations, code and anything else that lives in two dimensions are served badly. And re-hearing a sentence costs more fumbling than re-reading one. For some material, on some days, print with generous spacing beats speech — the fuller comparison lives in reading versus listening. A good tool is one you can put down.
Speeds, voices and the ear
Playback rate is the control people learn first. Comfortable listening rates for new material sit near ordinary speaking pace, and practised listeners push well above it — screen-reader users famously run synthetic speech at rates that sound like static to everyone else, which is a skill built over years, not a setting. The honest caveat is that the speed–comprehension trade-off does not disappear because the channel changed. Push any channel fast enough on unfamiliar material and understanding thins. The rate worth having is the one that survives a comprehension check, not the largest number the slider offers.
Voice quality matters more than it is given credit for, and not as a luxury. A flat or garbled voice taxes the ear just enough that a long document becomes an endurance event; good prosody — pauses at commas, a falling line at sentence ends, a breath at paragraph breaks — is doing the same structural signalling for the ear that typography does for the eye. When you audition a voice, do not test it on a sentence. Test it on twenty minutes.
Where the voice runs
Here is the question almost nobody asks of a read aloud app: where is the voice? The most natural-sounding voices mostly run in the cloud, which means every document they speak is uploaded first. For a novel, perhaps that is tolerable. For the things adults actually need read — contracts, medical letters, HR documents, an unfinished thesis — it means the most sensitive reading you do is precisely the reading that travels. On-device voices invert this: the text is spoken where it already sits, and nothing leaves the machine.
Our own answer, declared plainly: the app’s voice is KittenTTS nano, a fifteen-million-parameter open-weight neural voice that runs in the browser itself, with eight voices to choose from. No audio server, no upload, working offline — the case for that architecture is made in a voice that never leaves the device. The honest trade is that a small local model does not match the polish of the best cloud voices; we think the privacy is worth the gap, but it is a gap, and pretending otherwise would be exactly the kind of claim this magazine exists to avoid.
The most sensitive reading you do is precisely the reading that travels.
The cloud-voice problem, in one line
Who leads, voice or reader
One design question changes what text to speech is for: who sets the pace? In the app the audio can be the master clock — speech-led, the words on screen keeping time with the voice, which turns reading into something closer to being read to — or it can follow the reader, voicing whatever pace you set. Speech-led is the accessibility posture: the voice carries the decoding and the eyes are free to confirm. Reader-led keeps you in charge and adds the voice as a second channel underneath. Neither is superior; they are different tools sharing a codebase, and knowing which one you need on a given document is most of the skill.
The dignity argument
A surprising amount of resistance to text to speech is not technical but moral — the suspicion that listening is cheating, that the document does not count unless the eyes did the work. It is worth saying clearly that this is superstition. The contract read aloud is the contract read; the ideas do not know which nerve delivered them, and comprehension, not decoding, was always the point of the exercise. Understanding is the product. Reading is the packaging.
The adult who queues a report for the commute, checks the two tables by eye afterwards, and proofs their reply by ear is not avoiding reading. They are running it properly — each channel doing the work it is fit for, none of them asked to prove anything. The tool disappears into the routine, which is what good tools do. The voice reads. The reader, at last, gets to think.
A note on what this is. Signal is written in-house by the team that builds Reader Inc., so treat it as an argument rather than a review. Nothing here is medical, psychological or educational advice, and the app is not a treatment, therapy or diagnosis for any condition. Where we describe research we describe it in general terms; where we are reasoning past the evidence we say so. The app is free, runs entirely on your own device, and ships with a comprehension test switched on — which means you can check every claim we make against your own reading rather than taking our word for it.