A private AI video transcriber

Quiet Transcript reads the audio track from a video your browser can decode, then uses Whisper AI to produce an editable, timestamped transcript. It does not analyse the pictures and it does not upload the video.

MP4 and WebM support depends on the codecs built into your browser. Start with the first five minutes of a long video, check the detected language and draft quality, then choose the whole recording.

A recorded meeting is a video file like any other: an MP4 that Zoom, Google Meet or Teams saved to your computer transcribes here the same way, without the meeting leaving your device.

Press record, or load a recording.

Record here

The microphone is asked for at the moment you press Record, and not before — nothing on this page is listening until then. What you record is held in this page, is not sent anywhere, and is not kept when you leave unless you press Keep on this device.

Or load a recording

Free, no account, and nothing to install — this page is the whole audio to text converter. It records, it plays the recording back, it exports the transcript as SRT, WebVTT or plain text, and you are the only person who ever holds a copy.

Quiet Transcript uses Whisper, an AI speech-recognition model. Use the dedicated speech-to-text tool for recorded audio or the video-to-text tool for a video's audio track. Both routes keep the same local-processing and draft-review limits as this page.

Recording here

Press Record and the browser asks you for the microphone at that moment. Nothing on this page touches it before you press, and no device list is read at load: until that press there is nothing to allow. While the tape runs the deck says so in three ways at once — the word REC, a lamp, and a counter — and the level bar beside them moves with what the microphone is actually hearing, which is how you catch a muted input before the interview rather than after it.

Pause holds the tape without ending it; Stop ends it and hands you the recording. From there you can play it, transcribe it, save it to a file, or keep it on this device. Recording needs MediaRecorder, which not every browser has; where it is missing the record controls do not appear and loading a file still works.

What is kept on this device, and how to delete it

Nothing is kept automatically. Not the audio, not the transcript, not your corrections, not a setting, not a draft. A recording is written to this device only when you press Keep on this device, and you press it once for each recording you want to keep. There is no autosave to switch off because there is no autosave.

What you do keep is written into this browser's own storage on this computer. It stays there after the tab closes and after the machine is restarted, it is never sent anywhere, and it is readable by anyone who can use this computer. The shelf above lists every one of them with its name, date, length and size, tells you how much space they take, deletes any single one, and deletes all of them at once. Clearing this site's data in your browser removes them too.

Checking a draft against the tape

Press play and the transcript follows: the line under the playhead is highlighted as it runs. Press the time on any line to jump the recording to that moment, correct the words in place, and type who was speaking. Whatever is on screen is what the download contains.

Under the deck the recording is drawn as sound — the loud parts tall, the quiet parts flat, a pause a flat line — so a silence, an interruption or the moment somebody raised their voice is a shape you can see before you have listened to any of it. Press anywhere on it and the recording plays from there, and drag along it to move through the recording while you look. The part already played is oxide and the part still to come is grey, so how far in you are is a glance rather than a reading of the counter. Once there is a transcript, each segment boundary is a short mark along the top edge: the wave and the transcript are the same timeline, measured the same way.

That drawing is the precise control, and the reels are the coarse one. Drag a reel and the tape winds: the pack on the left thins, the pack on the right thickens, and the counter, the playhead and the highlighted line all move with them — one turn of a reel is a sixth of the recording. Everything reads and moves the same position; there is nothing here that can be in two places at once. Give the wave keyboard focus and the arrow keys move it a second at a time, Home and End go to the ends, and the position is announced as a timecode, so nothing about it needs a mouse or a thumb.

What you hear is the recording after it has been reduced to the 16 kHz mono the model is given — not your original file, which is left where it is. That is the honest way round: when a word comes out wrong, the useful question is what the model heard.

What this tool does, and what it will not do

Audio in

Below is what this browser, on this device, says it can read — asked of it as the page loaded rather than copied from a compatibility table. Support differs between browsers and between builds of the same browser, so a list written by hand would be wrong for somebody: Chromium without proprietary codecs answers no to AAC while Chrome on the same machine answers yes.

Reading and playing go through the same decoder here, so a file that transcribes will also play. The list above is the browser's own report and not a promise: the only certain answer is the file itself, and the tool gives it the moment you load one.

Languages

The model understands ninety-nine languages — among them Russian, Ukrainian, Spanish, German, French, Portuguese, Arabic, Hindi, Mandarin, Japanese, Korean, Turkish and Polish — and it works out which one it is hearing from the first thirty seconds of the recording.

It tells you what it decided, and you can overrule it. The language is printed with the transcript, not hidden inside a menu, because a wrong guess here does not look like an error: it looks like a transcript. Thirty seconds of a quiet room, a strong accent or music before anybody speaks are each enough to mislead it. Pick the right language and the recording is transcribed again with that language forced.

It transcribes; it never translates. What was said in Russian comes back in Russian. A tool that quietly returned English for a Russian recording would be changing the words of the person on the tape, which is the one thing this site exists not to do.

Accuracy outside English is the honest limit. whisper-tiny is the smallest model of its family, and it is weakest exactly where the training data is thinnest. Where the result is not good enough, the picker beside it offers whisper-base, which is materially better outside English and downloads an extra 80 MB from this site the first time you choose it.

Text out

Every one of them carries your speaker labels, and each file is written from what is on screen at the moment you press it, so your corrections are in it. Nothing is written until you press one.

What the PDF and the Word file actually are

Not a printout of the screen. The PDF is a numbered transcript: every printed line carries a number down the left margin and the numbering restarts on each page, so a passage can be cited the way transcripts are cited — page 7, line 12. Each page repeats the matter and the recording it came from, so a page separated from the bundle still identifies itself; page numbers read N of M, so a missing page is visible; and the line saying a machine wrote it is on every page rather than on a cover sheet that gets dropped. Timestamps run down their own column, on every line or once a minute or not at all, as you choose.

Anything you marked is indexed on the first pages, with the page and line it landed on, and appears again as a parenthetical on the line itself. Mark an exhibit while you listen and it is findable without reading the transcript.

Line numbers are in the PDF and not in the Word file, and that is a limit rather than an oversight: a word processor repaginates the document when it opens it, so any page and line printed here would describe page breaks it is not showing. The Word file is the copy to edit, it says so in its own first paragraph, and its index cites the recording's timestamps instead.

Every file is opened again and measured before you get it

When a file is written, this page reads it back the way somebody receiving it would — it counts the pages, reads the numbered lines off them, checks that each marked passage is where its index says, and confirms the words in the finished file are the words that went in. The result is listed under the buttons with the measurement beside each check, because a line saying Passed and nothing else is a promise rather than a check.

If a check fails there is no download. Not a warning next to a working button — no file at all, and the reason in plain words. A document that failed its own check is exactly the one that gets forwarded to somebody else. That is also why a transcript in Cyrillic, Greek, Hebrew or Chinese is refused as a PDF and offered as Word, Markdown, plain text or subtitles instead: the fonts built into a PDF cover Western European characters, and a file full of unreadable marks is worse than an honest refusal.

The PDF and the Word file need a document library, so pressing one of those two for the first time downloads about 1.5 MB — 420 KB for the PDF, 1.08 MB for Word, both served from this site. A visitor who only ever wants a .txt never downloads either: the button on the page tells you the cost before you press it.

Limits, stated as plainly as the promises

Why it works this way

Attorney–client privilege, medical confidentiality and a journalist's obligation to a source are not preferences. A cloud transcription service asks you to set them aside; a tool that runs in your browser does not have to ask, because there is no server here to receive the file.

You do not have to take that on trust. Open your browser's network panel, then record something or load a recording, and watch: no request carries your audio, because none is made. That is true of a recording you keep on this device as well — keeping writes it into your own browser, and there is no step after that.

The first run downloads the model once

The speech model and its runtime are about 71 MB, served from this site and cached by your browser afterwards. That download is the price of not sending your recording to anyone: the work has to happen somewhere, and here it happens on your machine.

A transcript produced here is a draft. Automatic transcription is worst exactly where the audio is worst — crosstalk, accents, distance from the microphone. Check it against the recording before you quote it, file it or enter it into evidence.