Accessibility
PDFs on your website: the overlooked third
Almost every business website hands out PDFs.
12 min read
By Timo Wessels Published
Almost every business website hands out PDFs. Price lists, data sheets, registration forms, menus, maintenance contracts, annual reports. They sit on the page as a download link, and they almost never come up in the discussion about accessibility.
Yet the same applies to them as to the website: a PDF a screen reader cannot read out does not exist for some of your customers. And the share of inaccessible PDFs is clearly higher than the share of inaccessible web pages — because nobody ever looks at PDFs.
The first question: does it have to be a PDF at all?
Before it gets technical, the most important decision.
The rule of thumb is: wherever possible, HTML instead of PDF. A normal web page is more accessible, easier to find, readable on the phone and can be maintained without producing a document again every time.
What speaks against PDF on a website:
- It is a dead end. Once someone has landed in the PDF, it is hard to move on. Linking onward is awkward, and responsive behaviour practically does not exist — an A4 layout on a phone stays an A4 layout.
- It does not show up in the analytics. Common analytics tools cannot follow what happens inside the PDF.
- It is rarely updated. Because the change has to be made in the source document, exported again and uploaded again, many websites carry PDFs with prices from the year before last.
What PDF is right for, on the other hand:
- Documents that have to look exactly the same everywhere.
- Documents that are meant to stay unchanged — invoices, signed contracts, certificates.
- Anything meant for printing.
The menu, the service overview and the directions belong on a page. The maintenance contract may be a PDF.
(Rheinwerk, Barrierefreie Webseiten, ch. 11.1)
What makes a PDF accessible: tags
Without anything further, a PDF is just an arrangement of characters on a surface. A line sitting at the very top and set larger does not make it a heading — it is only larger.
The meaning comes through tags. These are structural declarations in the document that play exactly the role HTML elements play on a web page: this line is a first-level heading, this block is a paragraph, this is a list with four items, that is a table with a header row.
The most important tags, once you have seen them, are familiar:
| Area | Tags |
|---|---|
| Headings | <H1> to <H6> |
| Text | <P> paragraph, <Span> text section, <Quote> quotation, <Note> footnote |
| Lists | <L> list, <LI> item, <Lbl> label, <LBody> content |
| Tables | <Table>, <TR> row, <TH> header cell, <TD> data cell |
| Other | <Figure> graphic, <Link> link, <Form> form field |
The quickest test of all: open the PDF, go to the document properties and look for the entry “Tagged PDF”. If it says No, the document is not accessible. Full stop. Everything else is moot for now.
The second quickest test: search the PDF for a word you can see on the page. If it comes back with “No matches”, the text is not text but an image — a scan or a text area placed as a graphic. For a screen reader the document is then completely empty. This happens very often with scanned forms and old brochures.
(Rheinwerk, Barrierefreie Webseiten, ch. 11.2 and 11.4)
The standard behind it
What the WCAG are for web content, PDF/UA is for PDF documents — laid down in the standard ISO 14289. The current version, UA-2, dates from March 2024.
Four core requirements:
- Tagging. The document must contain structure tags.
- Navigation. It must offer navigation aids — bookmarks, links — so that nobody has to work through it linearly.
- Alternative texts. All non-text elements need a text alternative.
- Interoperability. The document must work with different software and hardware.
One point from EN 301 549 worth knowing: chapter 5 requires that accessibility information is preserved during a conversion. So if a website generates a PDF live — a data sheet from the product data, an invoice, a confirmation — then headings, tables and images in the generated PDF have to remain marked up.
That hits exactly the cases nobody thinks of: the automatically generated order confirmation, the dynamically generated quote. For separately created download PDFs this particular criterion does not apply — they should still be accessible.
(Rheinwerk, Barrierefreie Webseiten, ch. 11.1)
How an accessible PDF comes about
The most important sentence first: it comes about in the source document, not afterwards.
Word, PowerPoint and InDesign generate the fitting tags automatically — if the document was worked on cleanly. Repairing a completely unstructured PDF afterwards is laborious, and Acrobat's automatic tagging delivers worse results in practice than generating from the source program.
What that means for working in Word, concretely:
Use styles, do not format. A heading is set through the style “Heading 1”, not by making the text bold and bigger. Only the style later produces the <H1> tag. The keyboard shortcuts for this are Alt+1, Alt+2, Alt+3, and Ctrl+Shift+N for normal text.
As on the web: do not skip a level.
No empty paragraphs. A page break is set with Ctrl+Enter, not by pressing Enter fifteen times. Empty paragraphs are read out by the screen reader as “blank, blank, blank …” — in the worst case the listener assumes the document has ended. Existing empty paragraphs can be removed through Find and Replace: search for ^p^p, replace with ^p, replace all.
Nothing important in the header and footer. They are hidden from screen readers in the PDF conversion so they do not interrupt the reading flow on every page. Whatever is there does not arrive.
Alternative texts on images, short — at most about 150 characters, less is better. Decorative images can be excluded in Word 365 through “Mark as decorative”.
Meaningful link texts. Not “more”, not the bare URL. Word turns typed-in addresses into links automatically, and the screen reader then reads out the complete address character by character. Via right-click, Edit Link, the displayed text can be changed.
Tables with a header row. The header row has to be set explicitly as such, otherwise the screen reader cannot relate a value to its meaning. If the table also has a feature column on the left, “First Column” should be switched on as well. It also makes sense to switch off rows breaking across pages.
And one principle on this: No tables for layout. An address does not belong in a table just so that it lines up neatly.
Set the document title. The screen reader announces it first. In Word under File, Info. A document titled “Document1” is a missed opportunity — for findability as well.
Set the document language. Otherwise the screen reader may switch to the wrong pronunciation. Individual passages in another language can be marked separately — but not for single words, because every language switch creates a disturbing pause.
Colour not as the only means. The same principle as on the web. A good test for it: would the information survive a black-and-white printout?
Run the built-in check. Word has an accessibility checker under “Review” or “File, Info, Check for Issues”. It lists the most important points, jumps to the place on click and offers solutions.
(Rheinwerk, Barrierefreie Webseiten, ch. 11.3)
The export — this is where it breaks most often
This is the point where months of clean work are lost in one second.
When saving as PDF from Word there are two options:
- Standard (publishing online and printing) — right.
- Minimum size (publishing online) — wrong. This setting creates no tags. All structural information is gone.
The name of the wrong option sounds harmless and is therefore chosen often, especially when someone wants to keep the file small. To be safe, you can also check in the options whether “Document structure tags for accessibility” is switched on.
Rule to remember: Whoever chooses “minimum size” has just lost everything they built before.
How to check a PDF
PAC — the PDF Accessibility Checker. Free from axes4, for Windows. It checks against the criteria of the Matterhorn Protocol, and optionally against WCAG. About two thirds of the criteria are checked automatically, one third remains manual work.
PAC also has a screen reader preview — it shows what a screen reader makes of the document. That is the most revealing view of all, because you see immediately whether the order is right.
What PAC cannot do: fix errors. Corrections are made in the source document or in Adobe Acrobat Pro.
The online checker from axes4 at check.axes4.com works without installation. With confidential documents, though, uploading is something to consider.
Adobe Acrobat Pro has its own check and a wizard that guides you through the settings. It recognises text through character recognition, sets tags, adds alternative texts and above all allows you to change the reading order without touching the appearance. Acrobat's check does not cover the full scope of the Matterhorn Protocol, though.
What no machine can check: whether an alternative text makes sense, whether the reading order matches what was meant, whether a table has the right tags. Of the Matterhorn Protocol's 31 checkpoints and 136 failure conditions, 87 can be detected by machine and 47 need human judgement.
The share is remarkably similar to that for web pages — and it is the reason a green check result is no proof of an accessible document.
(Rheinwerk, Barrierefreie Webseiten, ch. 11.4)
Two side topics with practical consequences
Protection and screen readers. PDFs can be protected extensively — block printing, block copying, encryption. Important here: text access for screen readers must stay allowed. If it is blocked too, the document is completely inaccessible for blind users, no matter how cleanly it is tagged. It is a single setting that is easily switched off by accident along with the rest.
Scanned documents. A scan is an image. It only becomes text through character recognition. After that the recognition should be checked — character recognition produces errors, and a screen reader reads them out without anyone noticing. With forms and documents containing numbers that is a real problem.
What to do
Take stock. How many PDFs are on your website? For most clients the answer is clearly higher than estimated — they pile up over the years.
Sort them into three piles:
- Can go. Outdated price lists, brochures from 2019, forms nobody needs any more. The cheapest progress in this field.
- Belongs on a page. Everything that is really web content: service overviews, directions, menus, opening hours. This switch improves accessibility, findability and phone use in one go.
- Stays a PDF. Contracts, certificates, forms for printing, anything with a signature.
Check the third pile with PAC. For every document first: is it tagged? Is the text searchable?
Repair in the source document. If you still have the Word file, the way is clear: set styles, add alternative texts, set the table header, export correctly. If you no longer have it, recreating is often quicker than repairing.
And for the future: if your business publishes documents regularly, the Word template is the lever. A template set up cleanly once, with defined heading levels, makes sure every future document brings the structure along automatically. That is an hour's work and then done for good.
The German federal government provides a detailed guide to accessible documents that covers PowerPoint and InDesign besides Word.
Sources
- Rheinwerk, Barrierefreie Webseiten, ch. 11.1 (why PDFs have to be accessible) — PDF as ISO 32000 and ISO 32000-2, the advantages and disadvantages of PDF including the dead-end property and the lack of analytics, the rule of thumb HTML before PDF, the menu example, the note on character recognition for scans, PDF/UA as ISO 14289 with the version UA-2 from March 2024 and its four core requirements, and the requirement from EN 301 549 ch. 5 checkpoint 5.4 on preserving accessibility information during conversion, with the distinction from separately created download PDFs.
- Rheinwerk, Barrierefreie Webseiten, ch. 11.2 (PDF tags) — tags as the counterpart of HTML semantics, the complete table of standard tags, the test through the document property “Tagged PDF”, the option to set a language per tag, bookmarks as a navigation aid, the procedure for setting alternative texts in Acrobat and marking decorative images.
- Rheinwerk, Barrierefreie Webseiten, ch. 11.3 (step by step: an accessible PDF with Word) — the recommendation to generate from the source program rather than tag afterwards, styles and the shortcuts Alt+1 to Alt+3, dealing with breaks and empty paragraphs including the Find and Replace instructions, the note on hidden headers and footers, alternative texts with about 150 characters as the upper limit, meaningful link texts and the autocorrection of URLs, table header row and feature column, switching off row breaks across pages, document title and document language, dealing with passages in another language, the colour rule, Word's built-in accessibility check, and the export settings with the explicit warning against “Minimum size”.
- Rheinwerk, Barrierefreie Webseiten, ch. 11.4 (checking the accessibility of a PDF) — the search test as a sign of text as image, the Matterhorn Protocol with 31 checkpoints and 136 failure conditions, of which 87 can be detected by machine and 47 need human judgement, PAC as a free tool with about two thirds automatic coverage and the screen reader preview, the online checker from axes4, the capabilities of Adobe Acrobat Pro including changing the reading order without changing the appearance, the aspects of manual checking, and the note to allow text access for screen readers when protecting PDFs.
- Own audit practice — sorting the existing PDFs into three piles, the observation that the number of existing PDFs is regularly underestimated, the assessment that recreating without the source file is usually quicker than repairing, and the recommendation to set up the Word template as a lasting lever.