Drop any audio or video file, pick a language, and get a full transcript with timestamps - all in your browser. The model is downloaded once (~75 MB) and cached locally, so subsequent runs are instant and work offline.
Recognition runs locally. Your audio never leaves your device.
The model is cached in your browser. Subsequent runs work offline.
No account, no credits, no usage cap. Open and use.
Speech recognition normally means uploading your audio to someone's GPU. Here the model runs on your own machine instead, which is what makes the tool free and private at the same time.
The trade-off is model size. On-device weights are a fraction of what a hosted service runs, so accuracy drops on noisy audio, strong accents, and overlapping speakers - that is where the cloud option below earns its keep.
The free in-browser tool runs a small open-source model - fine for clean audio, but it can struggle with heavy accents, noise, or overlapping speakers, and only covers a limited set of languages. For broadcast-grade accuracy, far broader language support, and SRT/VTT/TXT downloads, use our paid cloud-based Speech to Text.
Try cloud-based Speech to Text