A reference on how books are made readableNot a library · no catalogue · nothing to borrow
Accessible Book Collection
In Audio

Synthetic Voices, Honestly Assessed

Text-to-speech has closed the gap faster than anyone predicted — and opened new ones nobody expected.

By the Register desk · Audio · 3 min read

TSI Speech+ talking calculator with red digital display and raised button keypad
Adequate for a thriller, still failing on a chemistry textbook.Photo: TSI Speech+ Talking Calculator 01 · Wikimedia Commons

01What the technology can actually do now

The neural voice engines that power today's screen readers and reading apps are not the robotic monotone of twenty years ago. Contemporary systems — trained on hours of real speech and refined through machine learning — handle sentence rhythm, rising and falling intonation, and even the difference between a question and a declarative sentence without explicit instruction. For a straightforward prose narrative, a well-configured synthetic voice is now genuinely listenable: the words are clear, the pace can be adjusted without distortion, and the listener's comprehension holds up across a long session.

This matters enormously for volume. A human narrator requires a recording booth, a director, multiple takes, editing and proofing — a process that runs to many hours of production time for every finished hour of audio. Synthesis produces a complete reading in minutes, from any text that can be fed to it. For periodicals, new releases, documents and correspondence, that speed is transformative. A magazine that could never justify the cost of a human recording can now be available to a blind reader on the day of publication.

The argument for synthetic audio is also the argument for reach: more titles, faster, at far lower cost. In the context of the book famine — the long-standing gap between what is published and what is accessible — that matters.

A transcription desk with a screen, headphones and a marked-up manuscript
The print original is the working document. Queries are raised on it, not on the proof.Photo: Mateusz Dach / Pexels

02Where it still fails, and why

The honest answer is that synthetic voices fail in layers. The first layer is linguistic: proper nouns, technical vocabulary and foreign words trip even the best engines. A thriller set in Kraków, with a cast of Polish characters, will have every name mangled the same confident, wrong way throughout. A human narrator can research pronunciation, ask a consultant, and mark the script; a synthesis engine applies its training data and moves on.

The second layer is structural. Chemistry, mathematics and music notation each carry symbols that are not ordinary words — they are instructions about relationships between quantities or sounds. A chemistry textbook is one of the hardest documents to render in any accessible format, and a synthetic voice given raw LaTeX or chemical notation produces either silence or nonsense, depending on what the engine does with characters it cannot resolve. Even when the text has been prepared carefully, reading "H two O" or "the integral from zero to infinity" is a pale substitute for a narrator who understands what the expression means and can pace it accordingly.

The third layer is emotional register. A thriller depends partly on tension — on a voice that builds urgency without melodrama, that holds back where the prose holds back. Current synthetic systems can vary pace and pitch but cannot make the interpretive decisions a recorded book demands as a performance. The result is a reading that is correct but inert: every sentence receives roughly the same weight. For functional reading — finding a fact, following an argument — this is usually fine. For fiction that depends on atmosphere, the flatness shows.

Where synthesis wins and loses — a quick reference

From the register
ItemWhat it means
Straight narrative prosecurrent engines handle well; listenable at normal and accelerated rates
Proper nouns and foreign wordsconsistent weak point; errors are confident and repeated
Mathematical and chemical notationrequires pre-processing; raw symbols produce garbled or absent output
Poetry and literary fictiontechnically correct; interpretive flatness is the real problem
Periodicals and documentsstrongest use case; speed and cost advantage is decisive here

03The practical split

The working distinction is roughly this: synthesis serves well where the reading is purposeful and the text is standard prose. Reference books, newspapers, business documents, most non-fiction — the text-to-speech version is often adequate and always faster to produce than a human recording. Genre fiction falls somewhere in the middle: acceptable for most readers, clearly inferior to a good narrator for readers who care about the experience.

Complex technical content, foreign-language text, poetry and literary fiction with heavy tonal work are where synthesis still asks for patience rather than delivering it. These are also the categories where human narration is hardest to fund, which makes the gap doubly frustrating.

None of this is fixed. The models improve, the training data grows, and voice engines released in the next few years will handle things that defeat current systems. But improvement is not uniformity: a voice engine that has closed the gap on prose fiction will still stumble on an equation or a Welsh place name. Understanding what the technology genuinely does, rather than what its marketing promises, is the only way to make sensible decisions about when to use it.

For periodicals, new releases, documents and correspondence, that speed is transformative.

A narration booth with foam walls, a music stand and a microphone, empty
Foam, a stand, and a chair that does not creak.Photo: Jessica Lewis 🦋 thepaintedsquare / Pexels
Hands resting on the open pages of a thick book in someone's lap
Experienced listeners run well above natural speaking rate.Photo: Gustavo Fring / Pexels

This is an independent publication about accessible book formats. It is not a library, publisher or lending service, and it does not provide access to books or documents.