Digital Accessibility and Video Captions: A Practical Guide

Most video captions fail accessibility requirements not because of bad intent, but because auto-generated captions are good enough to feel like a solution when they are not. Here is what WCAG actually requires and why it matters beyond legal compliance.

Ghulam Mujtaba, Software Developer · Codingtron
GM

Ghulam Mujtaba

Software Developer at Codingtron · Vehari, Pakistan

Builder of Facebook to Transcript. Writes about AI transcription, video accessibility, and practical workflows for content creators and researchers.

Digital Accessibility and Video Captions: A Practical Guide

facebooktotranscript.com

Why auto-generated captions are not accessibility compliance

The most common accessibility mistake with video content is not skipping captions entirely — it is believing that auto-generated captions are sufficient. YouTube auto-captions exist. Facebook auto-captions exist. They are better than nothing. They are not better than an 88–93% accurate caption file when compliance and usability require close to 100%.

WCAG 2.1 Success Criterion 1.2.2 (Level A) requires captions for all pre-recorded synchronised media — any video that has both audio and visual content. The standard does not name a specific accuracy percentage, but guidance from the BBC, FCC, and DCMS consistently points to 99% accuracy as the required standard for pre-recorded content.

At 88% accuracy — the low end of YouTube's auto-caption range — roughly one word in nine is wrong. For a viewer relying on captions to understand the content, that error rate is not a minor inconvenience. It is the difference between following the content and not following it. The standard exists because it matters to the people it is written for.

Who actually uses captions: it is not who you expect

Accessibility requirements for video captions are written primarily for deaf and hard-of-hearing viewers. The people who actually use captions most often are hearing viewers. Studies in the UK indicate approximately 80% of caption users are hearing.

The reasons are varied and some of them are mundane: watching video on a phone in a noisy environment without headphones. Watching at work or in a public space where audio is not appropriate. Being a non-native speaker watching content in a second language where reading alongside the audio improves comprehension. Having auditory processing difficulties that make rapid speech harder to follow, even with normal hearing sensitivity. Watching at night without waking a partner.

This is useful context when making the business case for caption quality. The audience for well-captioned video is much larger than the audience for specifically accessibility-focused content. It includes a significant portion of everyone who watches video in non-ideal conditions, which covers most of the internet most of the time. When I've reviewed caption engagement analytics on published content, the numbers consistently confirm this — hearing viewers use captions far more than most creators expect.

What WCAG actually requires versus what it suggests

WCAG 2.1 SC 1.2.2 (Level A) requires captions for pre-recorded synchronised media. SC 1.2.4 (Level AA) extends this to live content. Both require captions synchronised with the audio — appearing in time with the speech, not just provided as a static text alternative elsewhere on the page.

The distinction that trips most organisations up: a transcript linked below a video satisfies SC 1.2.3 (audio description alternative) at Level A, but it does not satisfy SC 1.2.2, which specifically requires synchronised captions. If a video has spoken content and a linked transcript, it meets one criterion and fails another. Many organisations believe a transcript alone fulfils the requirement. It does not.

Legal obligation depends on jurisdiction and organisation type. US courts have increasingly interpreted ADA requirements to include websites. The UK Equality Act and EU Web Accessibility Directive impose similar requirements on public sector organisations. For any public-facing video content, treat WCAG AA compliance as the baseline regardless of specific legal exposure.

The process for adding compliant captions to existing video

The fastest path to WCAG 1.2.2 compliance for pre-recorded video is three steps: generate an AI transcript to get a base SRT file with timing already attached; correct the output for accuracy (a human review pass for a 20-minute video takes 20–30 minutes and typically raises accuracy from 88–93% to 99%+); upload the corrected SRT file to the video platform.

YouTube, Facebook, LinkedIn, and Vimeo all accept SRT file uploads for existing videos through creator/admin settings. Captions are stored separately and displayed by the player — the video does not need to be re-uploaded or re-encoded. The full process for generating an SRT file and uploading it is covered in the subtitle upload guide, including platform-specific steps for each major video platform.

Speaker identification and non-speech audio

AI transcription captures speech. It does not capture the non-speech audio content that matters to viewers who cannot hear: relevant sound effects, music with emotional significance, speaker identification when multiple speakers are not visually distinguishable. Standard notation for non-speech elements uses square brackets in the caption stream: [upbeat music], [applause], [door opening]. These are added manually to the SRT file.

Speaker identification is the most commonly missed element in caption compliance audits. Format: [Speaker Name:] or [Role:] at the start of that speaker's first caption block in each new section. For a typical 20-minute multi-speaker video, adding speaker labels takes 15–20 minutes and is the single most impactful accessibility improvement beyond basic accuracy correction.

Retroactive captioning: prioritising a large backlog

Organisations with a large archive of uncaptioned video face a common problem: where to start. Not all uncaptioned content carries equal risk or equal audience. A practical triage framework: prioritise public-facing over internal, high-traffic over low-traffic, recently published over old content that may be superseded. Build a list of the top 20 videos by traffic or visibility and caption those first.

A systematic programme that captions new content within one week of publication and works through the backlog by priority is sustainable. A one-time effort to caption everything simultaneously usually is not. For social platform captioning specifics, the social video accessibility guide covers platform-specific captioning features for Facebook, Instagram, TikTok, LinkedIn, and YouTube.

The fastest path to AA compliance

Generate an AI transcript to get a base SRT file with timing. Human-review the output for accuracy (20–30 minutes for a 20-minute video). Add speaker identification labels and non-speech annotations. Upload the corrected SRT to the video platform. This meets WCAG 1.2.2 Level A without re-uploading or re-encoding the video.

Frequently Asked Questions

What is the WCAG requirement for video captions?

WCAG 2.1 Success Criterion 1.2.2 (Level A) requires captions for all pre-recorded synchronised media — any video that has both audio and visual content. Success Criterion 1.2.4 (Level AA) extends this to live audio content like live-streamed events. Both require captions that are synchronised with the audio. A transcript linked below a video meets 1.2.3 (audio description alternative) at Level A but does not satisfy 1.2.2, which specifically requires captions that appear in sync with speech. Many organisations get this wrong and believe a linked transcript is sufficient.

Does WCAG compliance apply to social media videos?

The WCAG specification is a technical standard, not a law — legal obligations depend on jurisdiction and organisation type. In the US, the ADA applies to places of public accommodation that operate websites, which courts have increasingly interpreted to include social media content by covered entities. The UK Equality Act and EU Web Accessibility Directive impose similar requirements on public sector organisations. For private businesses, legal exposure depends on organisation size and the nature of content published. When in doubt, captions are the conservative choice.

Who benefits from captions beyond deaf users?

Non-native speakers watching in their second language, users in loud environments (commuting, open offices), users with auditory processing difficulties who can hear but struggle with fast speech, users watching without headphones in public, and users with ADHD who benefit from having audio and text simultaneously. Studies indicate about 80% of UK caption users are hearing viewers. Captions are not a niche accessibility feature — they improve the experience for a significant majority of video viewers.

Are auto-generated captions sufficient for accessibility compliance?

No. Auto-generated captions must meet accuracy requirements to be considered accessible. WCAG does not specify a percentage threshold, but guidance from bodies like the BBC and FCC points to 99% accuracy for pre-recorded content. YouTube auto-captions typically reach 88-93% accuracy, which means roughly one word in ten is wrong — not acceptable for compliance without correction. Platforms that offer only auto-captions are generally not considered compliant.

What is the difference between open and closed captions?

Closed captions can be turned on or off by the viewer. Open captions (burned in) are always visible. Both meet the WCAG requirement — the distinction is a user preference question, not a compliance question. Closed captions are generally preferred for web video because viewers can turn them off. Open captions are useful for content published on platforms where closed caption support is inconsistent, like some social media autoplay environments.

How do I add captions to a video that is already published?

Generate a transcript, convert it to SRT format with timing, review and correct the accuracy, then upload the SRT file to the platform. YouTube, Facebook, LinkedIn, and Vimeo all accept SRT file uploads for existing videos in the creator/admin settings. This is the standard retroactive captioning process and it changes nothing about the video itself — captions are stored separately and displayed by the player. The process for generating an SRT is covered in the transcript-to-caption guide.

Ready to Convert Your Facebook Videos to Text?

Use our free AI-powered tool to transcribe any Facebook video in seconds.

Try the Free Transcription Tool