Free Online Subtitle Generator

Generate Subtitles Online

The free online subtitle generator turns speech in a video or audio file into timed SRT and VTT captions, using the OpenAI Whisper model in the browser. Because the model runs on the device, the audio is never uploaded.

Caption and transcription work by OnlinePCApps since 2013

No upload SRT and VTT 99 languages Free
0
Bytes uploaded
99
Languages
0
Files stored
$0
Free, always
Why generate subtitles online

Three Reasons to Add Subtitles

A video without captions loses viewers who watch on mute, miss the audio or speak another language. The free online subtitle generator writes timed captions the moment the speech is transcribed.

Reach People Who Watch on Mute

Most social video plays silently until a viewer taps to unmute. Captions carry the message in those first seconds, which holds attention and lifts watch time on a feed that scrolls fast.

A video that lands without sound.

Skip Typing Every Line by Hand

Writing captions by hand for a long clip takes hours of pausing and typing. The Whisper model transcribes the whole track with timestamps in one pass, leaving only a quick review of names and terms.

Minutes of review, not hours of typing.

Keep Unreleased Footage Private

An interview, a course lesson or an unreleased cut should not sit on an outside server for captioning. The model runs in the browser, so the audio stays on the device while the captions are written.

Captions without exposing the audio.
How it works

Generate Subtitles Online in Three Steps

1

Add the Video or Audio

Drag a video or audio file onto the panel above, or browse to it on the device. An MP4, MOV, MP3 or WAV all work. The audio is read out of a video on its own.

MP4 SRT language
2

Let the Model Transcribe

Pick the spoken language or let it auto-detect, then start. The Whisper model transcribes the speech into timed segments on the device, and a translation to English can be turned on for another language.

captions.srt
3

Review and Download

Read the caption text, fix any name or term, then download it as SRT, VTT or plain text. The file is ready for YouTube, an editor or an HTML5 video player.

Timed captions SRT, VTT, TXT
Pick language Or auto-detect
One or a batch Download or ZIP
Need it faster?

Caption a Long Video With a Larger Model

A short clip is captioned in one pass on this page. For a feature-length film, a set of lectures or the largest Whisper model at its best accuracy, the desktop edition runs straight from disk with the graphics card behind it and writes the whole batch in one run.

Get Desktop Version Free trial · Windows 7 to 11
What the free tool does

What the Subtitle Generator Does

To generate subtitles is to turn spoken words into timed lines of text. The tool does that on the device with the Whisper model.

Auto-Generate Timed Captions

The Whisper model listens to the track and writes each phrase with a start and end time. The result is a caption file that lines up with the speech, ready to sit under a video without any manual timing.

Export SRT, VTT or Text

The captions save as SRT for YouTube and video editors, VTT for an HTML5 web player or plain text without timing for notes and search. Each file downloads straight to the device.

99 Languages and Translation

Whisper recognises speech in 99 languages and detects the language on its own. A non-English track can be captioned in its own tongue, or translated to English captions in the same pass.

Review and Fix Before Export

The transcript shows in the browser to read and correct, since a name, an acronym or a brand can trip any model. Word-level timing can be turned on for karaoke-style captions where each word is aligned.

To generate subtitles, to make an SRT file and to add captions to a video all name the same task. A search for subtitle generator or auto captions reaches this tool, and the source video is left as it is.

Reference

What Converts Cleanly and What to Watch

The model writes captions from the speech it hears. These points decide how the result lands.

The caseResultWhat happens and why
Input formatsvideo or audioMP4, MOV, MKV, WebM, MP3, WAV, M4A and FLAC are read, with audio pulled from a video.
SRT and VTTsubtitle filesSRT suits YouTube and editors, VTT suits an HTML5 web player, both with timing.
Plain text or JSONflexibleA plain TXT transcript suits notes and search, while JSON holds per-segment times.
Languages99 supportedSpeech is recognised in 99 languages, with the language detected on its own.
Translate to EnglishoptionalA non-English track can be turned into English captions in the same pass.
Timestampssegment or wordSegment timing suits normal captions, word-level suits karaoke-style highlighting.
Model sizetiny to largeA larger model reads speech more accurately but downloads more and runs slower.
Model downloadonce from a CDNThe model downloads once from a public CDN, then is cached and works offline.
Accuracyreview advisedClean audio reads well, though a name, an accent or background music warrants a check.
Where it runson the deviceThe audio is transcribed in the browser, so the file itself is not uploaded.

How On-Device Captioning Works

A subtitle generator runs automatic speech recognition, or ASR, the technology that turns spoken audio into text. This tool uses OpenAI Whisper, an open-source model released in 2022 and trained on 680,000 hours of audio, run in the browser through Transformers.js with WebAssembly and WebGPU. The audio is decoded and read on the device, and one honest point sits at the centre of the privacy story: the audio file itself is never uploaded, while the model does download once from a public CDN, after which it is cached and the tool works offline. Whisper ships in sizes from tiny to large, and a larger model reads speech more accurately but weighs more to download and runs slower, so the choice trades accuracy against speed. It reaches strong accuracy on clean audio, often around 95 percent and near human on many benchmarks. A proper noun, a homophone or a heavy accent can still throw it, so a quick review before publishing is wise. The captions save as SRT or VTT, and the standards below define both.

Honest comparison

In the Browser vs a Cloud Service

Both write captions from speech. The trade is real, and an interview or a cut is often confidential.

Point of comparison This tool Captioned in the browser On the device Cloud service On a server
Where the audio goes Stays on the device Uploaded to a server
Price and caps Free with no file cap Free tier often capped
A two-hour file, fastest Bound by the device Server farms run faster
Label who is speaking Plain captions only Speaker labels included
Certified accuracy A draft to review Human-checked service

Those last three rows favour a cloud service, since a two-hour file at top speed, speaker labels and a human-checked guarantee each call for more than a browser tab offers. For an everyday video the first two rows are what count. The audio stays on the machine, and captioning is free with no per-minute charge.

Why no upload

The Audio Never Leaves the Device

The audio is decoded and read inside the browser by the Whisper model through Transformers.js, so the transcription is client-side and the file stays on the machine that opened it. The model downloads once from a public CDN, and no audio is passed to a server for the job.

A recording can be a private thing, an interview under embargo, a legal deposition or an unreleased cut never meant for an outside server. A cloud service has to upload the whole file to caption it, far more exposure than a subtitle track is worth.

1. Open the browser tools at the Network tab
2. Clear the log and let it record
3. Caption a video with the panel above
The audio stays put. Only the one-time model file is fetched, then the work is local.
0
Bytes uploaded
0
Files stored
0
Accounts required
0
Watermarks added
Alternatives

Other Ways to Add Subtitles

Each of these writes captions for a video. They differ in effort, in cost and in where the file goes.

YouTube Studio

Auto-captions are added after a video is uploaded to YouTube.
They are free and align to the speech on their own.
The video must be uploaded first and captions stay on YouTube.

CapCut or an Editor

Auto-caption tools in an editor caption and style the text.
They tie in with the rest of an edit in one place.
An install is needed and many upload the audio to caption it.

Cloud Caption APIs

Services such as Whisper API or Deepgram caption on a server.
They add speaker labels and scale to huge files.
The audio is uploaded and the cost is charged by the minute.

Type by Hand

A caption editor lets each line be typed and timed by hand.
It gives full control over every word and cue.
It is slow, an hour of work for a few minutes of video.

YouTube captions, editor tools and cloud APIs all caption well, yet each either uploads the audio or ties into one platform. This page keeps the work local while writing SRT and VTT files with nothing to install and nothing sent to a server.

Before converting

Three Things to Know Before Captioning

A little context sets what to expect.

Pick SRT, VTT or Text

SRT suits YouTube and video editors. VTT suits an HTML5 web player. Plain text suits a script or notes without timing, and the destination usually points to one of the three.

Review Names and Terms

The model reads clean speech well but can trip on a proper noun, a brand or a technical term. A quick read of the transcript before export catches these, which matters most for published work.

Naming the Language Helps

Auto-detect works, though naming the spoken language up front speeds the model and sharpens the result. A larger model reads harder audio better at the cost of a bigger download and a slower run.

The desktop edition

When a Long Film Needs Captioning

The browser handles a clip in memory, which fits an everyday video. A feature-length film, a set of lectures or the largest Whisper model at its best accuracy belongs on the desktop edition, which reads from disk with the graphics card behind it and writes the whole batch far faster.

In a browser taban everyday clip
On the desktoplong films and batches, faster
Whole Folders

Point it at a folder of videos and every file is captioned to its own SRT in a single run, saved beside the original.

Largest Model

The largest Whisper model runs with the graphics card behind it, the highest accuracy on hard audio without a long wait in a tab.

Caption Presets

Saved presets set the model, the language and the output format in one click, ready for a repeated captioning workflow.

Common questions

Subtitle Generator Questions

It turns the speech in a video or audio file into timed captions, straight in the browser. The OpenAI Whisper model transcribes the track into segments with start and end times, which then save as an SRT, a VTT or a plain text file, in any of 99 languages.
Drop the video on the panel, pick the spoken language or leave it on auto-detect, then start. The model extracts the audio from the video and writes timed captions. Review the text for any name or term, then download it as SRT or VTT.
No. The audio is transcribed by the Whisper model running in the browser, so the file itself never leaves the device. The one network request is a one-time download of the model from a public CDN, after which the tool works offline.
Yes. On the first run the Whisper model is fetched once from a public CDN, from around 75 MB for a small model to a few hundred MB for a larger one. It is then cached in the browser, so later runs are quick and work without a connection.
SRT, the SubRip format, is numbered cues with start and end times and is read almost everywhere, including YouTube and video editors. VTT, or WebVTT, is the caption format for HTML5 video and web players. Both hold the same timed lines, and either serves as closed captions on a player that shows them.
On clean speech the Whisper model is strong, often around 95 percent and near human on many benchmarks. Accuracy drops with background music, overlapping voices, heavy accents or low-quality audio. A quick review before publishing is wise, above all for proper nouns, brand names, numbers and homophones such as their and there.
Whisper recognises speech in 99 languages and can detect the language on its own. Naming the language up front speeds the run and sharpens the result, though auto-detect handles a track when the language is unknown.
Yes. For a non-English track, turning on the translate option writes English captions in the same pass, since Whisper was trained on translation as well as transcription. Quality is strong for common languages and weaker for rarer ones.
Yes. The transcript shows in the browser to read and correct before export. This is the moment to fix a misheard name or an acronym, since no model is perfect, especially on specialist terms it has not met before.
Segment timestamps give each phrase a start and end time, which suits normal captions. Word-level timestamps assign a time to every word, which suits karaoke-style highlighting or aligning text tightly to speech. The mode is chosen before transcribing.
This tool writes a subtitle file that a player or an editor lays over the video, which keeps the text editable. To burn captions permanently into the picture as hardsubs, the video has to be re-encoded with the text baked in, which a desktop video editor handles.
Export the captions as an SRT file, then open the video in YouTube Studio, go to Subtitles and upload the file. YouTube reads the timing from the SRT and shows the captions, which can still be edited inside YouTube afterwards.
Yes. The tool runs in a mobile browser as well as on a desktop. A short clip captions on a phone without an app, though a long video and a larger model lean on memory, so a computer is steadier for a feature-length file.
There is no fixed limit set by the tool. The practical ceiling is the memory of the device, since the audio is held there while it transcribes. A short clip finishes quickly, while an hour-long file is steadier on a desktop.
After the model has downloaded once, yes. The first run fetches the model from a CDN, and from then on it is cached in the browser, so captions can be generated with no connection at all on later visits.
Hosting the page costs little, and because the model runs on the device nothing is received or stored here. The desktop edition, built for long films, batches and the largest model, is what earns. Some services charge by the minute, while this one sets no such meter.

OnlinePCApps Media Group

Written and reviewed by Tomas Vidal, who has worked on speech recognition and captioning here since 2013

Last reviewed August 2026
13
Years on media tools
0
Files uploaded
99
Languages
0
Watermarks

Automatic captioning used to mean uploading a file and waiting on a server, and the shift that changed it was moving the model into the browser itself. This tool runs OpenAI Whisper through Transformers.js with WebAssembly and WebGPU, so the audio is read on the device and the transcription happens there too. The honest detail worth stating plainly is that the audio never leaves the machine, while the model does download once from a public CDN and is then cached for offline use. A larger model reads harder speech more accurately but weighs more and runs slower, which is the real trade to weigh. No model is flawless either, so a proper noun, a brand or a heavy accent can slip through. A quick review before publishing is time well spent. Because a recording is often confidential, keeping the audio on the device is the point of the whole design.

Standards followed

Whisper · OpenAI WebVTT · W3C SubRip · SRT Transformers.js · WebGPU
Built on the same shared design system as every OnlinePCApps tool. onlinepcapps.com

Generate Subtitles Online for Free

Free and online, with no sign-up and no upload. Caption a video in the browser with the audio kept on the device.

Generate Subtitles Free, no account Try Desktop Edition For long films and batches
Generate Subtitles