Drop any audio or video file, pick a language, and get a full transcript with timestamps - all in your browser. The model is downloaded once (~75 MB) and cached locally, so subsequent runs are instant and work offline.
Recognition runs locally. Your audio never leaves your device.
The model is cached in your browser. Subsequent runs work offline.
No account, no credits, no usage cap. Open and use.
Speech recognition normally means uploading your audio to someone's GPU. Here the model runs on your own machine instead, which is what makes the tool free and private at the same time.
The trade-off is model size. On-device weights are a fraction of what a hosted service runs, so accuracy drops on noisy audio, strong accents, and overlapping speakers - that is where the cloud option below earns its keep.
The free in-browser tool runs a small open-source model - fine for clean audio, but it can struggle with heavy accents, noise, or overlapping speakers. For broadcast-grade accuracy and SRT/VTT you can drop straight into a video editor, run your file through Subformer's cloud transcription in Subtitles-only mode.
Open Subtitles-only mode