Text Tools

Audio to Text

Turn an audio or video file into a timestamped transcript, and download it as subtitles.

Rate this tool

How to use the Audio to Text

  1. Drop in an audio or video file — MP3, WAV, M4A, OGG, MP4 and more all work.
  2. Choose the spoken language, or leave it on auto-detect.
  3. Click Transcribe and wait while the model runs.
  4. Read the timestamped lines, click any of them to hear that moment, then download TXT, SRT or VTT.

About the Audio to Text

This tool turns an audio or video file into text with a timestamp on every line. Drop in a recording, choose the language it is spoken in, and you get back a transcript split into short timed segments — click any line and the player jumps to that moment, so checking a word against the recording takes a second rather than a hunt.

The timestamps are the reason this is separate from Voice to Text, which gives you one block of text and can also record straight from your microphone. Use that one for dictation and quick notes. Use this one when the timing matters: subtitling a video, quoting an interview with the position in the recording, or splitting a long podcast into sections. Alongside plain text you can download SRT and VTT subtitle files, which is what YouTube, Vimeo, VLC and every video editor expect.

Everything runs inside your browser using the Whisper speech model. The file is decoded and transcribed on your own device — it is never uploaded, which is the point when the recording is a client call, a medical note or an interview under embargo. Only the model itself is downloaded, once, and after that repeat files start immediately. Longer recordings take longer and a phone will be slower than a laptop, so start with a few minutes if you want to see the speed before committing an hour of audio.

Frequently asked questions

Which audio and video formats work?

Anything your browser can decode, which covers MP3, WAV, M4A, AAC, OGG, FLAC, WebM and MP4. For a video file the soundtrack is used and the picture is ignored, so you can subtitle a video without exporting the audio first.

What is the difference between SRT and VTT?

They hold the same thing in almost the same way. SRT numbers each subtitle and separates seconds from milliseconds with a comma; VTT starts with a WEBVTT line and uses a dot. Use SRT for YouTube and most video editors, VTT for HTML5 video on a web page. Both are offered so you do not have to convert one into the other.

Is my audio uploaded anywhere?

No. The transcription runs in your browser on your own machine. The only thing fetched over the network is the speech model, once, and it is served from this site rather than a third party.

Do I have to pick the language?

No — auto-detect handles most recordings. Choosing the language explicitly helps when the audio is noisy, when the speaker has a strong accent, or when a few English words appear in otherwise non-English speech, which is the usual cause of a transcript that switches language halfway.

How accurate is it, and how long can the file be?

Clear speech transcribes well; heavy background noise, several people talking over each other, or a very quiet recording will all cost accuracy. There is no fixed length limit beyond 1 GB, but the work happens on your device, so a long file on a phone can take a while. Treat the result as a first draft you skim rather than a finished document.