From spoken words
to a file you can keep.
Open an audio file, transcribe it in your browser, then review the words and timestamps. Download plain text or SRT subtitles.
Start with your audio
Up to 10 minutes and 50 MB per file. WAV and MP3 are good starting points; other formats depend on your browser. A desktop browser is recommended.
Small English model
The first run downloads about 55 MB of model and runtime files. The model comes from Hugging Face, which receives your IP address and download requests. Your audio and transcript stay on this device.
Only model files are cached (about 43 MB), never your audio or transcript. Browser storage may be cleared automatically. Hugging Face privacy.
Open a file to begin. No model downloads until you press Transcribe.
Listen. Check. Keep.
- Run locally. The model runs in a background worker on your CPU. Longer recordings can take several minutes.
- Review the result. This small model can miss words or produce text that was not spoken, especially with noise or silence. Names and timestamps need checking.
- Download your edits. TXT contains your edited words; SRT includes your edited segment times.
English speech only. No speaker identification, meeting bot or automatic summary. Refreshing clears your audio and transcript; download before leaving.
Example: a short extract from John F. Kennedy’s inaugural address, supplied with the Transformers.js examples.
Review your transcript
Times are seconds from the start of your audio. Use Play to check a segment. Clear a segment’s text to leave it out of both downloads.
Some missing or overlapping model timestamps were adjusted to fit the audio. Review them before exporting subtitles.