How Quiet Transcript works
A practical guide to recording, transcribing, checking and exporting on your own device — the models, languages, long files, subtitles, documents, accounts and limits.
Record
Press Record. The browser asks for the microphone at that moment and not before.
- While the tape runs you see the word REC, a lamp, a running counter and a level bar that moves with what the microphone hears — so a muted input shows up in the first seconds, not after the interview.
- Pause holds the tape; Stop ends it and hands you the recording.
- Then play it, transcribe it, save it to a file, or keep it in this browser.
- Recording needs
MediaRecorder. Where a browser lacks it, the record controls do not appear and opening a file still works.
Open a file
Press Open a file and choose audio or video. It is read on your device the moment you choose it.
- Which formats open depends on your browser. The tool asks the browser as the page loads and lists the answer under Which files open here? — MP3, WAV and M4A open almost everywhere.
- From a video, only the audio track is read.
- Files up to 2 GB on a computer. On a phone, shorter recordings work best.
Transcribe
Press Transcribe on this device. Whisper, an AI speech-recognition model, runs inside your browser.
- The first run downloads the model and its runtime, about 71 MB, from this site. Your browser keeps them, so later runs start quickly.
- The progress bar counts real work: window N of the total, the audio covered, and a time estimate measured on your device after the first window.
- The result card names the engine: graphics acceleration where the browser offers it, several CPU threads, or a single thread on the slowest path.
Models and languages
Ninety-nine languages. The language is worked out from the first 30 seconds, printed with the transcript, and you can change it.
- Small is
whisper-tiny, part of the first download, and right for most English. - Larger is noticeably better outside English and downloads an extra 80 MB the first time you choose it. It is slower, especially on a phone.
- A wrong language guess does not look like an error — it looks like a transcript. If the language shown is wrong, choose the right one and the recording is transcribed again with it.
- It transcribes and never translates: what was said in Russian comes back in Russian.
Long recordings
A long recording is transcribed in 30-second windows that overlap by five seconds. Start with the first five minutes to check the language and quality, then run the whole file.
Live draft
Tick Live draft while you speak before pressing Record, and text appears about 30 seconds behind your voice.
- It downloads the model when you press Record, so the first text appears once that is done.
- On a slower device the draft falls behind and the rest arrives after you press Stop.
- The recording never depends on the draft: if the draft stops, the tape keeps running.
Check and correct
Press Play and the transcript follows the recording. Press a line’s time to jump there, fix the words in place, and type who was speaking.
- The recording is drawn as sound: loud parts tall, silence flat. Press anywhere on it and the recording plays from there, or drag along it. Once there is a transcript, each segment boundary is a short mark along the top edge.
- That drawing is the precise control, and the reels are the coarse one: one turn of a reel is a sixth of the recording. With the wave focused, arrow keys move a second at a time and Home and End go to the ends.
- Speaker labels are typed by hand, never guessed; Carry each speaker label down fills the following lines until you type a new one.
- Press a line’s mark to add it to the index of marked passages.
- What you hear is the 16 kHz mono the model was given — useful when you want to know why a word came out wrong.
Subtitles
Open the Subtitles panel under the transcript: it checks every cue against reading-speed guides and offers repairs.
- The guides: 17 characters a second, at most 84 characters, one to seven seconds on screen, no overlap.
- Split divides a long cue; Merge joins a short one with the next; the timing buttons hold a cue a tenth of a second longer.
- Repairs change timing and layout only — never a word or a speaker.
Exports and documents
Download SubRip (.srt), WebVTT (.vtt), plain text (.txt), Markdown (.md), a numbered PDF (.pdf) or a Word file (.docx) — each written from what is on screen, corrections and labels included.
- The PDF is a numbered transcript. Every line is numbered down the margin and the numbering restarts on each page, so a passage can be cited as page 7, line 12. Each page repeats the matter and reads page N of M.
- Anything you marked is indexed on the first pages, with the page and line each mark landed on.
- The Word file is the copy to edit. Line numbers are in the PDF and not in the Word file, because a word processor repaginates; its index cites timestamps instead.
- Every file is opened again and measured before you get it. If a check fails there is no download. That is why a transcript in Cyrillic, Greek, Hebrew or Chinese is refused as a PDF and offered as Word, Markdown, text or subtitles instead.
- The PDF and Word engines, about 1.5 MB together, download only when first used: 420 KB and 1.08 MB.
Keeping a recording on this device
A recording made here is kept only when you press Keep on this device, once for each recording.
- Kept recordings stay in this browser after you close the tab, and anyone who can use this device can play them.
- The Kept on this device list shows every one, with the space used. Deleting takes two presses; Delete everything removes them all.
- A file you opened is not offered for keeping: it is already on your disk.
Account, saving and sharing
An account is optional. Sign in with Google to save transcripts and share read-only links; transcribing never needs one.
- Save to my account sends the text, times, speaker labels, marks, language, model and length of that transcript. The audio is never sent.
- Your library holds up to 200 saved transcripts. Open one on any device; to play along, attach the recording from your own files.
- A share link is read-only and expires after 7, 30 or 90 days. Turn it off at any time from your account page.
- Deleting your account removes your profile, every saved transcript and every share link, immediately.
Privacy in practice
Your audio is transcribed on your device and never uploaded. You can check it yourself.
- Open your browser’s developer tools and the Network panel.
- Record something, or open a file, and press Transcribe.
- Watch: no request carries your audio. The only requests are the page and the model, from this site.
Analytics on this site is cookieless and counts named steps, never your files or words. The one advertisement loads only when you scroll to it. The details are on the privacy page.
Limits
- Weaker outside English. Usable in the well-represented languages and rougher in others; the larger model helps most there.
- The language is a guess until you check it. It is printed with the transcript and can be changed.
- No automatic speaker separation. Every label is one you typed.
- The first run downloads about 71 MB.
- 2 GB per file on a computer; less on a phone.
- Nothing is kept unless you ask. Reload the page and an unsaved transcript is gone: download it or save it to your account first.
- The transcript is a draft — not certified, sworn or verbatim. Check it against the recording before you quote it or file it.
Troubleshooting
- “This browser would not start the speech engine”
- Reload the page and try again. If it repeats, update the browser, or use a computer with Chrome, Edge or Safari.
- “The speech engine stopped unexpectedly”
- Usually the device ran out of memory. Try the first five minutes, the small model, or a computer.
- It is slow on my iPhone or iPad
- Phones transcribe on the CPU and are several times slower than a laptop. Keep the screen on, keep the tab in front, and use the small model for long recordings.
- “This browser could not decode that recording”
- The browser cannot read that format. Check Which files open here? and convert the file to MP3 or WAV.
- The model downloads every time
- A private or incognito window forgets it when closed. Use a normal window to keep it.