The free online subtitle generator turns speech in a video or audio file into timed SRT and VTT captions, using the OpenAI Whisper model in the browser. Because the model runs on the device, the audio is never uploaded.
Caption and transcription work by OnlinePCApps since 2013
A video without captions loses viewers who watch on mute, miss the audio or speak another language. The free online subtitle generator writes timed captions the moment the speech is transcribed.
Most social video plays silently until a viewer taps to unmute. Captions carry the message in those first seconds, which holds attention and lifts watch time on a feed that scrolls fast.
Writing captions by hand for a long clip takes hours of pausing and typing. The Whisper model transcribes the whole track with timestamps in one pass, leaving only a quick review of names and terms.
An interview, a course lesson or an unreleased cut should not sit on an outside server for captioning. The model runs in the browser, so the audio stays on the device while the captions are written.
Drag a video or audio file onto the panel above, or browse to it on the device. An MP4, MOV, MP3 or WAV all work. The audio is read out of a video on its own.
Pick the spoken language or let it auto-detect, then start. The Whisper model transcribes the speech into timed segments on the device, and a translation to English can be turned on for another language.
Read the caption text, fix any name or term, then download it as SRT, VTT or plain text. The file is ready for YouTube, an editor or an HTML5 video player.
To generate subtitles is to turn spoken words into timed lines of text. The tool does that on the device with the Whisper model.
The Whisper model listens to the track and writes each phrase with a start and end time. The result is a caption file that lines up with the speech, ready to sit under a video without any manual timing.
The captions save as SRT for YouTube and video editors, VTT for an HTML5 web player or plain text without timing for notes and search. Each file downloads straight to the device.
Whisper recognises speech in 99 languages and detects the language on its own. A non-English track can be captioned in its own tongue, or translated to English captions in the same pass.
The transcript shows in the browser to read and correct, since a name, an acronym or a brand can trip any model. Word-level timing can be turned on for karaoke-style captions where each word is aligned.
To generate subtitles, to make an SRT file and to add captions to a video all name the same task. A search for subtitle generator or auto captions reaches this tool, and the source video is left as it is.
The model writes captions from the speech it hears. These points decide how the result lands.
| The case | Result | What happens and why |
|---|---|---|
| Input formats | video or audio | MP4, MOV, MKV, WebM, MP3, WAV, M4A and FLAC are read, with audio pulled from a video. |
| SRT and VTT | subtitle files | SRT suits YouTube and editors, VTT suits an HTML5 web player, both with timing. |
| Plain text or JSON | flexible | A plain TXT transcript suits notes and search, while JSON holds per-segment times. |
| Languages | 99 supported | Speech is recognised in 99 languages, with the language detected on its own. |
| Translate to English | optional | A non-English track can be turned into English captions in the same pass. |
| Timestamps | segment or word | Segment timing suits normal captions, word-level suits karaoke-style highlighting. |
| Model size | tiny to large | A larger model reads speech more accurately but downloads more and runs slower. |
| Model download | once from a CDN | The model downloads once from a public CDN, then is cached and works offline. |
| Accuracy | review advised | Clean audio reads well, though a name, an accent or background music warrants a check. |
| Where it runs | on the device | The audio is transcribed in the browser, so the file itself is not uploaded. |
How On-Device Captioning Works
A subtitle generator runs automatic speech recognition, or ASR, the technology that turns spoken audio into text. This tool uses OpenAI Whisper, an open-source model released in 2022 and trained on 680,000 hours of audio, run in the browser through Transformers.js with WebAssembly and WebGPU. The audio is decoded and read on the device, and one honest point sits at the centre of the privacy story: the audio file itself is never uploaded, while the model does download once from a public CDN, after which it is cached and the tool works offline. Whisper ships in sizes from tiny to large, and a larger model reads speech more accurately but weighs more to download and runs slower, so the choice trades accuracy against speed. It reaches strong accuracy on clean audio, often around 95 percent and near human on many benchmarks. A proper noun, a homophone or a heavy accent can still throw it, so a quick review before publishing is wise. The captions save as SRT or VTT, and the standards below define both.
Both write captions from speech. The trade is real, and an interview or a cut is often confidential.
| Point of comparison | This tool Captioned in the browser On the device | Cloud service On a server |
|---|---|---|
| Where the audio goes | Stays on the device | Uploaded to a server |
| Price and caps | Free with no file cap | Free tier often capped |
| A two-hour file, fastest | Bound by the device | Server farms run faster |
| Label who is speaking | Plain captions only | Speaker labels included |
| Certified accuracy | A draft to review | Human-checked service |
Those last three rows favour a cloud service, since a two-hour file at top speed, speaker labels and a human-checked guarantee each call for more than a browser tab offers. For an everyday video the first two rows are what count. The audio stays on the machine, and captioning is free with no per-minute charge.
The audio is decoded and read inside the browser by the Whisper model through Transformers.js, so the transcription is client-side and the file stays on the machine that opened it. The model downloads once from a public CDN, and no audio is passed to a server for the job.
A recording can be a private thing, an interview under embargo, a legal deposition or an unreleased cut never meant for an outside server. A cloud service has to upload the whole file to caption it, far more exposure than a subtitle track is worth.
Each of these writes captions for a video. They differ in effort, in cost and in where the file goes.
YouTube captions, editor tools and cloud APIs all caption well, yet each either uploads the audio or ties into one platform. This page keeps the work local while writing SRT and VTT files with nothing to install and nothing sent to a server.
A little context sets what to expect.
SRT suits YouTube and video editors. VTT suits an HTML5 web player. Plain text suits a script or notes without timing, and the destination usually points to one of the three.
The model reads clean speech well but can trip on a proper noun, a brand or a technical term. A quick read of the transcript before export catches these, which matters most for published work.
Auto-detect works, though naming the spoken language up front speeds the model and sharpens the result. A larger model reads harder audio better at the cost of a bigger download and a slower run.
The browser handles a clip in memory, which fits an everyday video. A feature-length film, a set of lectures or the largest Whisper model at its best accuracy belongs on the desktop edition, which reads from disk with the graphics card behind it and writes the whole batch far faster.
Point it at a folder of videos and every file is captioned to its own SRT in a single run, saved beside the original.
The largest Whisper model runs with the graphics card behind it, the highest accuracy on hard audio without a long wait in a tab.
Saved presets set the model, the language and the output format in one click, ready for a repeated captioning workflow.
Free and online, with no sign-up and no upload. Caption a video in the browser with the audio kept on the device.