Auto captions for short clips: what actually matters, and the one error nobody tests for

Word-timed, burned in, in the safe zone, spelled right. Where auto-captions fail on clips, and why a 98%-accurate transcript still misspells the guest's name.

5 min read

Captions on a short clip do a different job from subtitles on a film. Subtitles help you follow dialogue. Captions on a clip are the reason someone scrolling with the sound off stops at all — they're the hook, carried in text. That changes what "good captions" means, and most of what's written about captions is written for the other job.

Here's what actually matters on a 30-second vertical clip, in rough order, and the one failure that gets through every automated check.

1. Word-timed, not line-timed

TV subtitles appear a line at a time, two seconds ahead of the speech. Clip captions appear a word at a time, in step with the voice, often with the current word highlighted. This isn't a style preference; it's what makes a muted clip watchable. The eye follows the highlight the way the ear would follow the voice, and the rhythm of the speech survives without the sound.

That needs a transcript with a timestamp on every word, which is the same thing that makes clean clip boundaries possible. A tool that only knows where sentences start can't do word-by-word captions properly; it fakes it by spacing words evenly across the sentence, and the words drift out of step with the voice by the end of a long one.

How to check: watch a clip with sound and see whether the highlighted word is the word being said. On a slow, clear sentence it'll always match. Check a fast one.

2. Burned in, not a sidecar file

Two ways to deliver captions: burn them into the pixels, or attach a caption file (SRT) the platform renders. For long-form, sidecar captions are better: searchable, translatable, toggleable. For clips, burned-in wins for two reasons.

Muted autoplay in the feed shows the video before any platform caption UI loads, and on some surfaces never shows it. Burned-in text is there from frame one. And burned-in captions are yours — styled, positioned, sized for a phone — where the platform's rendering is small and sits exactly where the platform's own UI already is.

The trade-off is that a burned-in error can't be fixed after posting. Which is why the rest of this post is about catching errors first.

3. In the safe zone

Each app paints its own UI over the bottom 20–30% of the frame and the right 15%. Captions placed where TV subtitles go — along the bottom — sit under TikTok's caption line and sound attribution, and are unreadable for everyone who didn't hide the UI, which is nobody.

The visible region across all three platforms is roughly the middle 55–60% of the height. Captions want the lower half of that band: below the face, above the UI. The safe-zone numbers per platform.

4. Big enough, and few enough words

A phone in a feed is viewed at arm's length with divided attention. Text that's comfortable on a laptop is too small by half. The working rule is two to four words on screen at once, in a heavy sans-serif, tall enough that a single word spans a fifth of the frame width. Six-word lines in a thin font are the most common self-inflicted caption failure after placement.

5. The error nobody tests for: proper nouns

Here is the failure that survives every accuracy figure.

Speech recognition is now very good on ordinary words; a modern model transcribes clean speech at around 95–98% word accuracy. The 2–5% it misses is not spread evenly. It concentrates on the words the model has never heard: your guest's surname, your product's name, your town, the acronym your industry uses. A transcription model has a vocabulary, and a name that isn't in it gets replaced by the nearest thing that is.

So a transcript can be 98% accurate and still spell your guest's name wrong every single time it appears, because it's the same unknown word each time. And it's the word that matters most: a clip introducing someone whose name is misspelled in the caption reads as careless in a way that a dropped "the" never does.

We hand the transcriber the video's title as a hint, which fixes the cases where the name is in it — a podcast episode titled with the guest's name transcribes that name correctly. It does nothing for a name that only appears in the speech. No tool solves this fully; the ones that claim to are guessing well.

How to check: search the caption text for every proper noun in the clip before posting. Names, companies, places, products. It takes twenty seconds and it's the only check that catches this class of error, because the model is confident about its wrong answer, so no "low confidence" flag fires.

6. Languages

Captions are only useful in the language the clip is in, which sounds obvious until a tool that supports "captions" turns out to support English captions. Our transcription auto-detects the language and runs in all 99 that the speech model knows, on every plan including free, and the captions come out in the spoken language with fonts that cover every script (the part most tools skip: a tool that transcribes Japanese but ships only a Latin font burns empty boxes into the clip). Translated captions (Spanish audio, English captions) are a different feature and a much harder one, because the translation has to be re-timed to the original speech word by word. If you need that, check specifically, and check the timing on a fast passage.

What to do before posting, every clip

  • Play it muted. Can you follow it from the captions alone? That's the whole test of whether they work.
  • Play it with sound. Does the highlight track the voice on the fastest sentence?
  • Read every proper noun. Names, brands, places. Fix by re-rendering, not by hoping.
  • Check the placement against the platform UI, on a phone, not in the tool's preview.

Captions are the one part of a clip the viewer reads before deciding to listen. They deserve the twenty seconds.

Try it on your own video

3 videos a month free, no card. Paste a link and see what comes back.