The PDF Problem
A page image with letters on it is not a text. What tagging does, and why so little of it exists.

01Structure Is Not Decoration
A sighted reader looks at a page and reads a heading because it is large and bold, sitting above a block of smaller text. The visual cues are real and informative — but they live only in the eye. A screen reader, a refreshable display, a text-to-speech engine: none of these see a page. They need the structure to be declared explicitly, not performed visually.
The PDF format can carry that declaration. A tagged PDF embeds a logical tree alongside the visual layer: this span of text is a heading, this is a paragraph, this is a caption, this is decorative and should be skipped. A reading system can then navigate by heading, move between list items, and know the order in which content should flow — even when the visual layout arranges columns or sidebars that bear no relation to reading sequence.
Most PDFs carry none of this. A document exported from a word processor without accessibility settings produces a file that looks correct on screen and is, beneath the surface, a sequence of positioned drawing commands. A scanned page is worse: an image of text, not text at all. The letters are pixels. A screen reader finds the file and reads nothing, or reads a filename, and stops.

Optical character recognition can convert a scanned image to actual characters, but recognition alone does not produce structure. The output is a single undifferentiated string — no headings, no reading order, no indication of what is a column header and what is body copy. Structure must be added by hand, which means someone trained in PDF remediation working through a document element by element, tagging each one correctly, and checking the result with the tools a reader actually uses.
That work is slow and skilled. On a complex academic paper with figures, tables, equations and footnotes, it can take many hours. Publishers produce thousands of PDFs. The arithmetic explains why the catalogue of accessible titles is so thin: remediation is not being skipped out of carelessness, but because no one has yet built the cost of doing it into the ordinary price of making a document.
How tagging works
From the register| Item | What it means |
|---|---|
| Tagged PDF | a PDF carrying a logical structure tree that declares headings, paragraphs, reading order and decorative elements alongside the visual layer |
| Untagged PDF | a file of visual drawing commands; structurally opaque to assistive technology |
| Scanned PDF | a page photograph; text must be recovered by OCR before any structural work can begin |
| OCR (optical character recognition) | software that converts a page image to characters; produces text, not structure |
| Remediation | the manual process of adding tags and correcting reading order in an existing PDF |
A screen reader, a refreshable display, a text-to-speech engine: none of these see a page.

What remediation costs
From the register- 01A simple, single-column PDF: less than an hour per document
- 02A document with tables, equations, multi-column layouts and figures: potentially many hours per document
- 03The cost scales with complexity and is rarely built into standard publishing workflows

More in Structure
Related entriesThis is an independent publication about accessible book formats. It is not a library, publisher or lending service, and it does not provide access to books or documents.