One file, one Transcribe button
Drop an MP3, WAV, M4A, or similar clip and press Transcribe. The transcript stays on this page.
On this device · No upload · No account
Convert a local audio file to text in this browser. The speech model runs on your device and the file is not uploaded.
Drop an audio file or click to select
MP3, WAV, M4A, OGG, FLAC, or WebM. One file, up to 25 MB. Nothing is uploaded.
Audio to text
Speech is converted on this device with Whisper. The first visit downloads a small English model and caches it in this browser. Safari and iPhone use WASM when WebGPU is missing.
First visit downloads the speech model. It stays cached in this browser so later files start faster.
0%
Features
The homepage is the working tool. You convert one local file to text, see a progress bar while the speech model caches, then copy or download the transcript without leaving the tab.
Drop an MP3, WAV, M4A, or similar clip and press Transcribe. The transcript stays on this page.
The English speech model runs in this tab. WebGPU is used when the browser has it; Safari and iPhone fall back to WASM.
Audio stays in this browser. Nothing is posted to our server, and you do not create a login.
After the pass finishes you can copy the text or download a .txt file. View opens a dialog on this page.

What it is
Audio to Text Online is a browser tool that converts a file you already have into written text. It loads onnx-community/whisper-tiny.en through Transformers.js. That is an on-device speech model, not a public CDN image host and not a server queue. The audio never leaves this tab.
We do not fetch YouTube, WhatsApp, or cloud storage. If you need those jobs, this page is only for a local file you can drop yourself.
How to use
Three on-page steps: add a file, press Transcribe, then copy or download the txt. Step 1 is choosing the clip. There is no paste-a-link path on this first English homepage.
Step 1
Drop one clip or click the card. Keep it at 25 MB or smaller so the tab can finish.

Step 2
The first visit downloads a small English model and caches it. Later files skip that wait.

Step 3
Confirm the filename, read the transcript here, then copy it or save a text file. Nothing is stored after you leave.

Why this page
Some converters upload the recording to a cluster. This page keeps the file in the tab so a voice memo or interview clip does not become a server-side copy.
No account, no upload, and the transcript is not stored after you leave.
Safari and iPhone still run the same English model through WASM, just more slowly on the first visit.
One file, 25 MB, English Whisper-tiny. We do not claim realtime, 4K, or server-grade accuracy.
Tips
Practical notes so the first model download and the transcript stay predictable.
Quiet speech beats music or overlapping talk. This is a tiny English model, not a studio caption desk.
The first visit downloads the model and caches it. A progress bar says so. Later files reuse that cache.
Clear the file with the × on the card if you picked the wrong clip. There is only one primary Transcribe button.
If the tab runs out of memory, use a shorter WAV or MP3. We do not offer unlimited cloud hours here.
Who uses this
People who already have a recording and need text they can edit. This page is only for that local file, not for a YouTube URL or a meeting bot.
Turn a phone memo into lines you can paste into notes without sending the clip to a host.
Get a first-pass transcript of a short interview you recorded yourself, then clean names by hand.
Convert a lecture excerpt you already downloaded so you can search the words on the same page.
FAQ
Open this page, drop an MP3, WAV, M4A, or similar file, then press Transcribe. The transcript appears on the same page so you can copy or download a txt file.
Yes. The file is read in this tab and sent only to the in-browser Whisper model. AudioToText.im does not upload the recording to our server.
No account and no login. Choose a file and press Transcribe. The first visit downloads a speech model and caches it in this browser.
Yes. When WebGPU is missing the page falls back to WASM. The first model download can take a minute on a phone; later files reuse the cache.
Yes. After Transcribe finishes, use Download for a .txt file or Copy to put the same text on the clipboard. Nothing is stored after you leave.
MP3, WAV, M4A, AAC, OGG, FLAC, and WebM work when the browser can decode them. Keep the file at 25 MB or smaller so the tab can finish.
The first visit downloads onnx-community/whisper-tiny.en and caches it. Later transcriptions on the same browser skip that wait unless the cache was cleared.
No. This first page only converts a file you already have. It does not fetch YouTube, WhatsApp, Word, or Google Docs audio.
It is an on-device English Whisper-tiny pass. Quiet speech works better than music or heavy noise. We do not claim server-grade or realtime accuracy.
Yes. One file up to 25 MB. Longer or larger clips can exhaust the tab. We do not offer unlimited cloud hours on this page.
After the model is cached, the page can transcribe without sending audio out. You still need the site files themselves if the tab is fresh.
This first English homepage uses an English-only model. Other languages are not claimed here and are not silently guessed as extras.
Drop a file in the first viewport and press Transcribe. This link only scrolls there.
Convert audio to text