Audio to Text Converter
Upload a recording and get an editable transcript in seconds. Free, no sign-up, and your file is never stored.
Your recording
Drag and drop an audio or video file
or click to choose one
mp3, m4a, wav, webm, mp4 — up to 25MB
Privacy: Your file is sent to our transcription provider, transcribed, and discarded. We never store the audio.
Transcript
Choose a file on the left and hit Transcribe. The text appears here, ready to edit, copy or download.
How automatic audio to text transcription works
When you upload a file here, it is sent to a speech recognition model that listens to the audio in small slices, predicts the most likely sequence of words for each slice, and stitches the predictions into running text. The model has been trained on many thousands of hours of transcribed speech, so it has a strong sense of which word sequences are probable in a given language. That is also why it occasionally writes a plausible phrase that is not what was said: it is guessing from sound and context, not understanding.
Punctuation is added by a second pass that looks at pauses, intonation and sentence structure. It is good at full stops and commas, less reliable with question marks and quotation. Language is detected automatically from the first few seconds of audio, which is why a file that starts with a long silence or music can occasionally be assigned the wrong language; trimming the beginning fixes it.
The whole process is fast because nothing is done by hand. A typical three-minute voice memo comes back in 10 to 30 seconds. That speed is the main advantage over human transcription, which costs money per minute and takes hours or days. The trade-off is accuracy on difficult audio, covered below.
Supported formats and the 25MB limit
The converter accepts the common audio and video containers you are likely to have: mp3, m4a (iPhone voice memos, Apple Music exports), wav, webm (browser recordings), mp4 and mov (video; the audio track is extracted), plus ogg, opus, flac and aac. If your file plays in a normal media player, it will almost certainly transcribe.
The limit is 25MB per file and five files per hour per visitor. What 25MB means in minutes depends on how the audio was encoded. As a rough guide: a 128kbps mp3 is about 1MB per minute, so 25MB is around 25 minutes; a 64kbps voice recording is twice that; an uncompressed wav at CD quality is about 10MB per minute, so you would get only two or three minutes. Video is the least efficient, since most of the bytes are pictures you do not need.
- Long wav file? Convert it to mp3 first with any free audio editor and it will shrink by around 90 percent with no audible loss for speech.
- Long video? Export the audio track only (most editors have an "export audio" option) before uploading.
- Recording over 30 to 40 minutes? Split it into parts. Each part counts as one of your five hourly uploads.
How accurate is audio to text, and what hurts it
On a clean recording of one person speaking clearly into a decent microphone, automatic transcription gets the large majority of words right; the mistakes tend to cluster around names, product terms, numbers and homophones. On harder audio the error rate rises quickly, and it is worth knowing which conditions cause the most trouble so you can either fix the recording or budget time for clean-up.
- Crosstalk. Two people speaking at once is the single biggest problem. The model has to pick one voice and usually mangles both. Meetings and lively interviews suffer most.
- Background noise. Traffic, cafe chatter, keyboard clatter and air conditioning all mask consonants, which are what distinguishes most words. Even a quiet hum lowers accuracy.
- Phone and conference audio. Phone lines cut off the high frequencies that carry s, f and th sounds, and video-call compression adds artefacts. Expect more errors than from a local recording.
- Accents and dialects. Models are trained mostly on standard varieties. Strong regional accents, non-native speakers and code-switching between languages all reduce accuracy.
- Distance from the microphone. A laptop mic across the table picks up room echo. A headset or phone held close to the mouth is dramatically better.
- Specialist vocabulary. Medical, legal and technical terms, company names and acronyms are often replaced with a common word that sounds similar.
Getting a better recording next time
Most accuracy problems are solved before you press record. Use a headset or hold the phone close, pick a quiet room, ask people not to talk over each other, and say names and acronyms clearly the first time they come up. A minute of care at the start saves ten minutes of editing at the end.
How to clean up a transcript
Raw machine transcripts are usable for search and reference as they are, but if anyone else is going to read the text, give it a quick edit. The box above is editable, so you can do this in place and then copy or download the result.
- Read with the audio playing. Fixing from memory misses the errors that look plausible; hearing the original catches them.
- Fix names and terms first with find and replace. A misheard product name usually appears the same wrong way every time.
- Add paragraph breaks where the topic changes. Machine output comes back as one long block, which is hard to read.
- Decide on verbatim or clean. For a quote, keep the exact words. For notes, remove filler words, false starts and repetitions so the text reads naturally.
- Mark anything you cannot make out with [inaudible] rather than guessing.
- Add speaker names at the start of each turn if more than one person is talking.
What people use audio to text for
The same converter serves very different jobs, and what counts as a good transcript changes with the job.
- Lectures and classes. Students record a lecture, transcribe it and search the text for the part they missed. Accuracy on technical terms matters more than punctuation.
- Interviews. Journalists and researchers need quotable text. Verbatim accuracy matters, and crosstalk is the enemy, so a quiet room and one person speaking at a time pays off.
- Meetings. Notes from a call that nobody wrote down. The transcript is usually skimmed, then summarised, so a rough version is fine.
- Voice memos. Ideas dictated while walking or driving, turned into text for a document or an email. These are short, single-speaker and transcribe very well.
- Sales and practice calls. Reps transcribe their own calls to count how much they talked, which questions they asked and where the conversation stalled. The transcript is the raw material for a review.
- Speeches and presentations. Record a rehearsal, transcribe it, and you can see your actual word count, your filler words and how far you drifted from the script.
Privacy, limits, and when you need something bigger
Your file travels over an encrypted connection to our transcription provider, is transcribed, and is deleted; we do not keep the audio and we do not save the transcript on our servers. The text exists only in your browser until you copy or download it. If you close the tab, it is gone. For confidential recordings, that is the behaviour you want, but it also means you should download the transcript before leaving the page.
The free tool is deliberately small: 25MB per file, five files an hour, plain text output, no speaker labels, no timestamps and no storage. For a two-hour interview, a batch of recorded meetings, or anything you need to keep and come back to, you will need a full account or a dedicated transcription service. We would rather say that plainly than have you discover it on the third upload.
If what you really want to know is how you sound, a transcript is only step one. Pavone turns a recording of your practice call or presentation into a scored review of pace, filler words, clarity and structure, and lets you rehearse the same conversation against an AI partner until it lands. Start with the transcript here, then take it further when you are ready.
Frequently asked questions
Is there a free way to convert audio to text?
Yes. This page transcribes audio and video files up to 25MB for free, with no account. Upload the file, wait a few seconds, and copy or download the transcript. The limit is five files per hour per visitor.
How accurate is audio to text?
For a clear recording of one or two speakers in a quiet room, modern speech recognition is typically in the mid-to-high 90s percent for word accuracy. Background noise, several people talking over each other, strong accents, or compressed phone audio pull that down noticeably, so always read the transcript through before using it.
Can I transcribe a 2-hour recording?
Not with this free tool. It is capped at 25MB per file, which is roughly 25 to 50 minutes of mp3 depending on bitrate. A two-hour recording needs a full account or a dedicated transcription service. You can also split the file into parts with a free audio editor and upload them one at a time.
What is the best audio to text converter?
It depends on what you need. For a quick, free transcript of a short file, a browser tool like this one is the fastest. For long recordings, speaker labels, timestamps, or team collaboration, a paid transcription service or a full speaking-practice account will serve you better.
How do I convert mp3 to text?
Drag the mp3 into the upload zone above (or click to choose it), then press Transcribe. The transcript appears in an editable box when it finishes. Mp3 is the most common format we see and works without any conversion.
Does audio to text work with video files?
Yes. Mp4, webm and mov files are accepted; the audio track is extracted and transcribed. Video files are larger than audio for the same length, so a 25MB video is usually only a few minutes long. If you have a long video, export the audio as mp3 first.
How long does it take to transcribe audio?
Usually 10 to 30 seconds for a typical file. Longer files take proportionally longer; a 25MB file can take up to a minute or two. The page keeps polling until the transcript is ready, so leave the tab open.
Is my audio file stored anywhere?
No. The file is sent to our transcription provider over an encrypted connection, transcribed, and then deleted. We do not keep the audio, and the transcript is only shown in your browser; it is not saved on our side.
Can I transcribe audio in languages other than English?
Yes. Language is detected automatically, and most major European and Asian languages are supported. Accuracy is highest for English and other widely spoken languages, and lower for smaller languages and heavy dialects.
Why is my transcript missing punctuation or speaker names?
This tool returns plain text with automatic punctuation but no speaker labels or timestamps, to keep it simple and fast. If you need to know who said what, add the names yourself in the editable box, or use a service that offers speaker diarization.
Can I transcribe a voice memo from my phone?
Yes. iPhone voice memos export as m4a and Android recorders usually produce m4a or mp3; both are supported. Share the memo to your computer or upload it straight from your phone browser.
How do I fix mistakes in the transcript?
The transcript box is editable. Play the recording alongside it, fix misheard names and technical terms, and add paragraph breaks where the topic changes. Then copy or download the cleaned-up text.
You might also like
Online Voice Recorder
Record audio from your microphone in the browser, play it back, and download it as a file. Nothing is uploaded.
Words to Minutes Converter
Paste a speech or script and see exactly how long it takes to say at slow, average and fast speaking rates.
Online Teleprompter
Free browser teleprompter: paste your script, set the speed and font size, mirror it, and read while you record.
Start practicing in minutes
AI Roleplays for any scenario
- Build a roleplay from your scenario
- Practice with AI, voice to voice
- Get instant, structured feedback after practice
- Free to start — no credit card required
Start practicing in minutes
AI Roleplays for any scenario
- Build a roleplay from your scenario
- Practice with AI, voice to voice
- Get instant, structured feedback after practice
- Free to start — no credit card required
