Auto captions: the styles that hold attention
Word-by-word captions became the norm for short video because a large share of the audience watches without sound: in the Verizon Media / Publicis Media study (5,616 adults, April 2019), 69% of consumers watch video without sound in public places and 80% say they are more likely to watch a captioned video to the end. Five style families cover most of 2026: karaoke (word highlighted as it is spoken), boxed, outline, glow and minimal. Ground rules: 3 to 5 words per line, bold sans-serif font, centered low inside the safe zone, one style for the whole account.
Why are auto captions essential on TikTok, Reels and Shorts?
Because sound is often off and attention is decided in three seconds. The most solid figures date from 2019: the Verizon Media / Publicis Media study reports that 69% of consumers watch without sound in public, 25% even in private, and that 80% are more likely to finish a captioned video (Forbes, July 31, 2019). Platforms have not published a more recent official figure for shorts; the "85% watched without sound" numbers that circulate come from a 2016 Facebook feed-video study, not from vertical formats. Be wary of undated statistics.
Second reason: captions are read by algorithms. On-screen text and automatic transcription help TikTok and YouTube classify content. Third reason: accessibility, which is not incidental for deaf and hard-of-hearing viewers.
Classic captions or word-by-word captions?
A classic caption shows a full sentence at the bottom of the frame for two to four seconds. A word-by-word (or "dynamic") caption shows a group of 3 to 5 words and highlights the spoken word, in sync with the voice. The second dominates shorts for a simple reason: it forces the eye to follow the rhythm of speech, which holds attention even without sound. It does require a transcript with a timestamp for every word; current models (OpenAI's Whisper, for example) provide word-level timestamps (OpenAI repository).
The 5 caption styles that dominate in 2026
| Style | What it looks like | When to use it | Risk |
|---|---|---|---|
| Karaoke | Word group displayed, the spoken word changes color or grows | Podcasts, interviews, any talking content; the default style for shorts | Eye strain if the highlight color is too saturated |
| Boxed | White text on a solid colored rectangle, sometimes one word at a time | Fast content, punchlines, comedy; very readable on busy backgrounds | Hides part of the image; avoid on faces |
| Outline | Text with a thick black border, no background | Tutorials, demos, when the image must stay visible | Less punchy; average readability on light, textured backgrounds |
| Glow | Text with a colored luminous halo | Gaming, music, night-time mood | Dated on B2B or educational content |
| Minimal | Thin text, no strong animation, one line | Luxury, design, documentary, LinkedIn | Loses attention on TikTok against more animated styles |
Clipping and captioning tools sell these families under brand names ("Hormozi", "MrBeast", "Viral"...); behind the names are the same five mechanics. ClipOnIt offers them under their generic names (karaoke, boxed, outline, glow, minimal) with 7 colors, 3 sizes, 3 positions and 3 to 8 words per line.
How many words per line?
Three to five. Below that, text jumps too fast and becomes exhausting; above, the line spills onto two rows and the highlighted word gets lost. For fast speech (over 160 words per minute), stay at 3 or 4 words. For slow, measured speech, 5 works. One line at a time; two simultaneous lines fall back into classic captioning.
Which font and size?
- Font: bold, sans-serif, wide (Montserrat, Inter, Poppins and equivalents). Thin or handwritten fonts vanish on mobile.
- Size: text should take 6 to 9% of the frame height (roughly 110 to 170 px on 1080×1920). Test on a phone, not a desktop monitor.
- Color: white or yellow for text, a single accent color for the highlighted word. Avoid pure red and pure green, which vibrate on OLED screens.
- Caps: fine for punchlines, tiring over 45 seconds of speech. Use capitals for the hook and normal case for captions.
Where to place captions?
Centered low, inside the safe zone shared by the three platforms (900×1400 px centered in the 1080×1920 frame), never in the bottom 400 pixels where the handle and description appear, nor in the right 120 pixels taken by the buttons. If the face is framed low, move captions up to the center rather than letting them overlap the mouth. Per-platform zones are detailed in short video length and format in 2026.
Are automatic captions reliable?
On clean audio, yes, at 95% and above in major languages. Errors cluster on four points: proper nouns, numbers ("twenty twenty-six" versus "2026"), loanwords, and punctuation. Plan a two-minute proofread per clip; serious tools let you edit a word without re-rendering the video. Two upstream tips: one mic per person and no background music in the source recording, music is added afterwards with automatic ducking.
Should captions include emojis?
Sparingly. One emoji on a key word (a number, an emotion) catches the eye; one emoji per sentence turns the clip into visual noise. The rule that holds: at most three emojis per 45-second clip, never on two consecutive word groups, and never on serious content (finance, health, legal).
Translated captions: one version per language
To reach a second language, export a separate clip with translated captions rather than a bilingual clip. Two lines in two languages crowd the safe zone and split attention. Most paid tools translate into a handful of languages in one click; proofread, idioms travel badly. The full method is in turning a one-hour podcast into ten shorts.
How to pick a style and stick to it
Run a two-week test rather than a taste decision. Export the same five clips in two styles (karaoke and boxed, say), post them on alternate days at the same time, then compare average completion rate per style, not views. Pick the winner, write down its settings (font, size, accent color, words per line, position) and apply them to every future clip. Revisit once a quarter at most; a caption style is part of the account's identity, and changing it every month erases the recognition you built.
The mistakes that make viewers leave
- Changing caption style from one clip to the next: the account loses its visual identity.
- Highlighting too slow or too fast against the voice: check sync on the first three seconds.
- Captions over the face: move them up or down, never over the mouth.
- Typos in the hook: it is the one text everyone reads.
- Forgetting captions on the title card: the hook is read without sound too.
FAQ
Do captions really increase retention?
The dated primary data says so (Verizon/Publicis 2019: 80% more likely to finish a captioned video). The more recent figures quoted by tool vendors ("+38% retention", "+40% views") are unsourced; treat them as orders of magnitude, not facts.
Which style for a B2B podcast?
Karaoke with a sober accent color, or outline. Avoid glow and solid colored boxes.
Can I auto-caption a video that is already edited?
Yes, that is Submagic's and CapCut's use case; clipping tools also do it on the clips they generate. The 2026 comparison says who does what.
Should a 16:9 short for YouTube be captioned?
Yes, with the same style but a smaller size (4 to 6% of the height); the horizontal frame leaves more room for text.