SRT or VTT: which format should you choose for real-time subtitles?
Updated October 8, 2026
These are the two formats everyone comes across first, and they are often confused. Yet they encode different intentions, and live broadcasting widens the gap. Here is what each one looks like, what really sets them apart, and how to choose.
SRT: the simplest, and that is its strength
SubRip (SRT) is just a sequence of numbered blocks: an index, two timestamps separated by an arrow, one or two lines of text, then a blank line. Milliseconds are separated by a comma.
1
00:00:01,000 --> 00:00:03,500
Good evening and welcome
to the news.
2
00:00:03,600 --> 00:00:06,000
Our top story tonight…
No styling, no positioning, no metadata. SRT has in fact never been formally standardised: it was born from a piece of software, and its character encoding is not fixed. Some players interpret tags such as italics or bold, with no guarantee. That very poverty is what makes it universal: editing software, media players, YouTube and Vimeo all accept it without question. It is a format for transport, not for presentation.
WebVTT: built for the web
WebVTT, specified by the W3C, is the native format of the HTML5 <track> element. It keeps SRT's structure and adds what the web was missing. A file must start with the line WEBVTT, milliseconds are separated by a period, and each subtitle can carry placement settings.
WEBVTT
00:00:01.000 --> 00:00:03.500 line:85% align:center
<v Anchor>Good evening and welcome to the news.
00:00:03.600 --> 00:00:06.000
Our top story tonight…
The line, position, size and align settings place the text on screen; the voice tag identifies the speaker; classes let you style the text in CSS, through the ::cue pseudo-element. The format also supports comments, chapter tracks and metadata tracks usable from JavaScript. Finally, the encoding is mandated: UTF-8, which settles accent problems from the outset.
For a web player or a streaming platform, this richness is not decorative: it lets you place a subtitle above a news ticker, or tell two speakers apart. An SRT file mechanically converted to VTT is still an SRT file: it gains none of these capabilities.
SRT and VTT, point by point
| Criterion | SRT | WebVTT |
|---|---|---|
| Specification | No formal standard | W3C |
| Header | None | WEBVTT line required |
| Milliseconds | Separated by a comma | Separated by a period |
| Encoding | Not fixed | UTF-8 required |
| Positioning | Not standard | line, position, size, align settings |
| Styling | Tags tolerated by some players | Classes and CSS (::cue), voices |
| Playback in a browser | Not native | Native, via <track> |
| HLS streaming | Not supported | WebVTT segments |
What live broadcasting changes
In real-time transcription, the file is never “finished”: it grows while the programme airs. In HLS, subtitles are cut into WebVTT segments, like the video, and listed in their own playlist. Each segment carries an X-TIMESTAMP-MAP header (RFC 8216) that aligns its timestamps with those of the video stream; if it is missing or wrong, the subtitles drift.
Two practical consequences. Segment duration adds to display delay: six-second segments hold the text back on screen by as much. And a subtitle that spans two segments must be repeated in each to remain visible. In DASH and CMAF, subtitles travel in MP4 fragments, as WebVTT or IMSC1 depending on the delivery specification.
SRT has no place in that delivery chain, but it remains ideal for archiving the programme once it is over, or for importing it into editing software. The choice depends less on the format than on when the text is consumed.
Converting SRT to VTT
The conversion is simple, as long as nothing is forgotten:
- Add the line WEBVTT at the top of the file, followed by a blank line.
- Replace the comma before the milliseconds with a period in every timestamp.
- Save the file as UTF-8: SRT files are often encoded otherwise, and accented characters are the first to suffer.
- Keep or remove the numbers: WebVTT accepts them as cue identifiers.
- Replace SRT's non-standard tags, such as font tags, with WebVTT classes.
The result is a valid VTT file, but without positioning or styling: those can only be added from a source that carries them.
Which format for which destination
- Editing and post-production: SRT, accepted everywhere.
- YouTube and Vimeo: both are accepted; SRT is enough if no placement is needed.
- HTML5 web player: WebVTT, the only format the <track> element reads natively.
- Live HLS streaming: WebVTT segments, or IMSC1 if the delivery specification requires it.
- Archiving a programme: SRT for simplicity, TTML if styling must be preserved.
- Broadcast delivery: neither; see professional formats such as EBU-TT-D, STL or SCC.
Frequently asked questions
Can a browser play an SRT file?
Not natively: the HTML5 track element expects WebVTT. JavaScript video players can often convert SRT on the fly, but the safest option is to deliver VTT directly.
Why do the accented characters in my SRT file display incorrectly?
Because SRT does not fix an encoding and the file was probably saved in a legacy encoding, such as Windows-1252, then read as UTF-8. Save it again as UTF-8, or switch to WebVTT, which requires it.
Can a subtitle be placed at the top of the screen in SRT?
Not in a standard way. Some players recognise position tags inherited from other formats, with no guarantee at all. If placement matters, for instance to avoid covering a news ticker, use WebVTT.
Does OnAirText deliver SRT and VTT during the broadcast?
Yes. OnAirText produces both continuously from the same stream, from the Starter plan, as well as professional broadcast formats on the higher plans.
Both, live
OnAirText produces SRT and VTT continuously during the broadcast, from the same stream. Early Access →