Dictation without a server listening
Tick Live draft while you speak under the Record button, press Record, and speak — then watch a draft of your words appear on this page while the tape runs, transcribed by Whisper, an AI speech-recognition model, on your own device. Popular dictation tools send your voice to a server to be recognised; this page has no server to send it to, so voice to text happens here or it does not happen at all.
The draft follows about half a minute behind your voice and settles when you press Stop; then correct it in place, label speakers, and export text, Word, a numbered PDF or subtitles. On a slower machine the draft runs further behind the tape and the page says so plainly — the recording itself is never at risk, whatever the model’s pace, because the tape does not depend on the draft.
This recording
- Source
- Length
- Model windows
- Engine
- Checked when the AI model starts
- Language
- Not known until it is transcribed
For a recording longer than five minutes, the quick preview is selected first so you can check the language and draft quality before committing to the whole run. The language is worked out from the first thirty seconds, which is all the model reads before it decides — a quiet opening, a strong accent or a burst of music is enough to get it wrong. Whatever it decides is printed above. If it is wrong, choose the right language here and the recording is transcribed again with that language forced, replacing the transcript and any corrections in it. The larger model is more accurate outside English and downloads an extra 80 MB from this site the first time you choose it; nothing is downloaded until you do.
The page reports whether Whisper AI is using graphics acceleration, multiple CPU threads or a slow one-thread fallback as soon as the model starts. That measured engine is what the progress estimate uses.
This is really a laptop tool. The first Transcribe on any device downloads about 71 MB of speech model and then runs it here; a phone will manage a short recording and will be slow, hot and occasionally out of memory on a long one. An hour-long interview belongs on a computer. Everything else on this page — recording, playing back, correcting, exporting — is light enough for any device.
The time remaining will appear after the first 30-second window finishes on this device.
Keeping writes this recording into this browser's storage on this computer. It stays there after you close the tab and after you shut the machine down, it is never sent anywhere, and anyone who can use this computer can play it. It is not kept until you press that button, and you have to press it again for the next recording.
Marked passages
These become an index on the first pages of the PDF, each one with the page and line it landed on, and a parenthetical on the line itself. The Word file lists the same marks by their position in the recording — a word processor repaginates the document when it opens it, so only the PDF can promise a page and a line.
Subtitles No transcript yet
Document Numbered lines · a timestamp on every line · US Letter
Speech models read a fixed thirty-second window, so a long recording is planned as many overlapping windows rather than one job. Neighbouring windows share five seconds, and the shared tail is dropped once — otherwise a word straddling a cut would be lost, which in a transcript people cite from is the sentence that matters.
Tapes kept on this device
Nothing is kept on this device. A recording is written here only when you press “Keep on this device”, once for each recording — there is no default and no setting that changes that.
These are on this computer, in this browser, and nowhere else: they were never uploaded and there is nothing to upload them to. They survive closing the tab, and they survive restarting the machine. Anyone who can use this computer — anyone who borrows it, repairs it or takes it — can play whatever is on this list. Delete a tape the moment it is no longer worth that.
Free, no account, and nothing to install — this page is the whole audio to text converter. It records, it plays the recording back, it exports the transcript as SRT, WebVTT or plain text, and you are the only person who ever holds a copy.
Quiet Transcript uses Whisper, an AI speech-recognition model. Use the dedicated speech-to-text tool for recorded audio or the video-to-text tool for a video's audio track. Both routes keep the same local-processing and draft-review limits as this page.
Recording here
Press Record and the browser asks you for the microphone at that moment. Nothing on this page touches it before you press, and no device list is read at load: until that press there is nothing to allow. While the tape runs the deck says so in three ways at once — the word REC, a lamp, and a counter — and the level bar beside them moves with what the microphone is actually hearing, which is how you catch a muted input before the interview rather than after it.
Pause holds the tape without ending it; Stop ends it and hands you the recording. From there
you can play it, transcribe it, save it to a file, or keep it on this device. Recording needs
MediaRecorder, which not every browser has; where it is missing the record
controls do not appear and loading a file still works.
What is kept on this device, and how to delete it
Nothing is kept automatically. Not the audio, not the transcript, not your corrections, not a setting, not a draft. A recording is written to this device only when you press Keep on this device, and you press it once for each recording you want to keep. There is no autosave to switch off because there is no autosave.
What you do keep is written into this browser's own storage on this computer. It stays there after the tab closes and after the machine is restarted, it is never sent anywhere, and it is readable by anyone who can use this computer. The shelf above lists every one of them with its name, date, length and size, tells you how much space they take, deletes any single one, and deletes all of them at once. Clearing this site's data in your browser removes them too.
Checking a draft against the tape
Press play and the transcript follows: the line under the playhead is highlighted as it runs. Press the time on any line to jump the recording to that moment, correct the words in place, and type who was speaking. Whatever is on screen is what the download contains.
Under the deck the recording is drawn as sound — the loud parts tall, the quiet parts flat, a pause a flat line — so a silence, an interruption or the moment somebody raised their voice is a shape you can see before you have listened to any of it. Press anywhere on it and the recording plays from there, and drag along it to move through the recording while you look. The part already played is oxide and the part still to come is grey, so how far in you are is a glance rather than a reading of the counter. Once there is a transcript, each segment boundary is a short mark along the top edge: the wave and the transcript are the same timeline, measured the same way.
That drawing is the precise control, and the reels are the coarse one. Drag a reel and the tape winds: the pack on the left thins, the pack on the right thickens, and the counter, the playhead and the highlighted line all move with them — one turn of a reel is a sixth of the recording. Everything reads and moves the same position; there is nothing here that can be in two places at once. Give the wave keyboard focus and the arrow keys move it a second at a time, Home and End go to the ends, and the position is announced as a timecode, so nothing about it needs a mouse or a thumb.
What you hear is the recording after it has been reduced to the 16 kHz mono the model is given — not your original file, which is left where it is. That is the honest way round: when a word comes out wrong, the useful question is what the model heard.
What this tool does, and what it will not do
Audio in
Below is what this browser, on this device, says it can read — asked of it as the page loaded rather than copied from a compatibility table. Support differs between browsers and between builds of the same browser, so a list written by hand would be wrong for somebody: Chromium without proprietary codecs answers no to AAC while Chrome on the same machine answers yes.
- MP3not checked
- WAVnot checked
- M4A / AACnot checked
- FLACnot checked
- Opus in OGGnot checked
- Vorbis in OGGnot checked
- Opus in WebMnot checked
- AIFFnot checked
- MP4 video, its audio tracknot checked
- WebM video, its audio tracknot checked
Reading and playing go through the same decoder here, so a file that transcribes will also play. The list above is the browser's own report and not a promise: the only certain answer is the file itself, and the tool gives it the moment you load one.
Languages
The model understands ninety-nine languages — among them Russian, Ukrainian, Spanish, German, French, Portuguese, Arabic, Hindi, Mandarin, Japanese, Korean, Turkish and Polish — and it works out which one it is hearing from the first thirty seconds of the recording.
It tells you what it decided, and you can overrule it. The language is printed with the transcript, not hidden inside a menu, because a wrong guess here does not look like an error: it looks like a transcript. Thirty seconds of a quiet room, a strong accent or music before anybody speaks are each enough to mislead it. Pick the right language and the recording is transcribed again with that language forced.
It transcribes; it never translates. What was said in Russian comes back in Russian. A tool that quietly returned English for a Russian recording would be changing the words of the person on the tape, which is the one thing this site exists not to do.
Accuracy outside English is the honest limit. whisper-tiny is the smallest model
of its family, and it is weakest exactly where the training data is thinnest. Where the
result is not good enough, the picker beside it offers whisper-base, which is
materially better outside English and downloads an extra 80 MB from this site the first
time you choose it.
Text out
- SubRip,
.srt— the subtitle file most editors and players open. - WebVTT,
.vtt— the web's own caption format. - Plain text with timestamps,
.txt— for quoting and for citing a position. - Markdown,
.md— for a notes app or a repository. - PDF,
.pdf— the document you can file and cite out of. - Word,
.docx— the same document, editable. - Copy — the timestamped text straight to the clipboard.
Every one of them carries your speaker labels, and each file is written from what is on screen at the moment you press it, so your corrections are in it. Nothing is written until you press one.
What the PDF and the Word file actually are
Not a printout of the screen. The PDF is a numbered transcript: every printed line carries a number down the left margin and the numbering restarts on each page, so a passage can be cited the way transcripts are cited — page 7, line 12. Each page repeats the matter and the recording it came from, so a page separated from the bundle still identifies itself; page numbers read N of M, so a missing page is visible; and the line saying a machine wrote it is on every page rather than on a cover sheet that gets dropped. Timestamps run down their own column, on every line or once a minute or not at all, as you choose.
Anything you marked is indexed on the first pages, with the page and line it landed on, and appears again as a parenthetical on the line itself. Mark an exhibit while you listen and it is findable without reading the transcript.
Line numbers are in the PDF and not in the Word file, and that is a limit rather than an oversight: a word processor repaginates the document when it opens it, so any page and line printed here would describe page breaks it is not showing. The Word file is the copy to edit, it says so in its own first paragraph, and its index cites the recording's timestamps instead.
Every file is opened again and measured before you get it
When a file is written, this page reads it back the way somebody receiving it would — it counts the pages, reads the numbered lines off them, checks that each marked passage is where its index says, and confirms the words in the finished file are the words that went in. The result is listed under the buttons with the measurement beside each check, because a line saying Passed and nothing else is a promise rather than a check.
If a check fails there is no download. Not a warning next to a working button — no file at all, and the reason in plain words. A document that failed its own check is exactly the one that gets forwarded to somebody else. That is also why a transcript in Cyrillic, Greek, Hebrew or Chinese is refused as a PDF and offered as Word, Markdown, plain text or subtitles instead: the fonts built into a PDF cover Western European characters, and a file full of unreadable marks is worse than an honest refusal.
The PDF and the Word file need a document library, so pressing one of those two for the first
time downloads about 1.5 MB — 420 KB for the PDF, 1.08 MB for Word, both
served from this site. A visitor who only ever wants a .txt never downloads
either: the button on the page tells you the cost before you press it.
Limits, stated as plainly as the promises
-
Weaker outside English. The model is
whisper-tiny, which understands ninety-nine languages but is the smallest of its family in all of them. English is its strongest; Russian, Spanish, German, French and the other well-represented languages are usable and visibly rougher; a language with little training data may come back as something you would not file. Choosing the larger model helps most exactly here. - The language is a guess until you check it. It is decided from the first thirty seconds, printed with the transcript, and changed by the picker — which transcribes again with your answer forced. Nothing here ever translates: whatever language was spoken is the language you get back.
- No automatic speaker separation. Nothing here works out who is talking, and this build will not gain it. Every speaker label in your transcript is one you typed.
-
The smallest model, by default.
whisper-tinyis chosen so it can run on an ordinary laptop or phone.whisper-baseis offered beside it and is better, particularly outside English; it costs an extra 80 MB, downloaded from this site only if you pick it, and it is slower to run. - The first run downloads about 71 MB. The model and its runtime come once, from this site, and your browser keeps them afterwards.
- 2 GB per file. Larger than that is refused rather than half-read; split it and run the parts.
- Nothing is kept unless you ask. The transcript, your corrections and the speaker labels are never saved anywhere: reload and they are gone, so download the file before you close the tab. A recording can be kept, and only by pressing Keep on this device for that recording — after which anyone who can use this computer can play it, until you delete it from the shelf.
- The transcript is a draft. Not a certified, sworn or verbatim record.
Why it works this way
Attorney–client privilege, medical confidentiality and a journalist's obligation to a source are not preferences. A cloud transcription service asks you to set them aside; a tool that runs in your browser does not have to ask, because there is no server here to receive the file.
You do not have to take that on trust. Open your browser's network panel, then record something or load a recording, and watch: no request carries your audio, because none is made. That is true of a recording you keep on this device as well — keeping writes it into your own browser, and there is no step after that.
The first run downloads the model once
The speech model and its runtime are about 71 MB, served from this site and cached by your browser afterwards. That download is the price of not sending your recording to anyone: the work has to happen somewhere, and here it happens on your machine.
A transcript produced here is a draft. Automatic transcription is worst exactly where the audio is worst — crosstalk, accents, distance from the microphone. Check it against the recording before you quote it, file it or enter it into evidence.