The convenience pipeline for interview transcription runs through somebody else's server: record the call, upload the audio, get back text. Most reporters have run it without thinking twice. It's worth thinking twice once — because the upload creates something that didn't exist when transcription was a cassette and a foot pedal: a copy of your source's voice on infrastructure you don't control. Local transcription is the version where that copy never comes into existence. This guide covers the threat model honestly, then the mechanics in QuickSpeak (which we make).
What an upload actually creates
When meeting audio lands on a transcription vendor's servers, three new exposures come with it:
- Legal process you're not party to. A subpoena to the vendor can reach materials you were never asked to produce — and your source finds out from the vendor's compliance department, not from you fighting the request.
- Somebody else's breach surface. Your source's identity is now one row in whatever database gets breached next year. You did the careful work; the vendor's incident response decides the outcome.
- A retention policy you didn't write. "De-identified," "deleted after 30 days," "used to improve our models" — each vendor documents different terms, and terms change. Otter's training on de-identified user meetings is the current example (details here); the pattern predates and will outlast them.
On-device transcription makes none of these exposures. The recording is captured by your browser, transcribed by a model running in WebAssembly on your machine, and saved to your disk. There is no vendor to subpoena, breach, or read a policy from, because there is no vendor in the pipeline.
The honest limits: local transcription is one exposure removed, not safety. Your laptop can be seized or stolen; your cloud backups sync what you keep; your email metadata identifies sources all by itself. Source protection is a discipline — and "the transcription vendor" is one fewer party in that threat model, not an all-clear.
The workflow
- Install QuickSpeak — no account, no sign-up, nothing that knows who you are. The recognition model downloads once and runs locally from then on.
- For phone or video interviews, join in a Chrome tab and set Audio source to Browser tab. For desktop-app calls, use Desktop (Zoom app, Teams); for in-person or phone-on-speaker, Microphone.
- Record the interview and let it run. When you stop, Transcribing → Identifying speakers → Saving all happen locally — you can watch it work with Wi-Fi switched off, which is the proof.
- Name the speakers. Click each label and type the name or the protecting pseudonym ("SOURCE A") — whatever belongs in your notes is what the transcript says.
- Verify quotes against audio. The recording is kept locally next to the transcript; click any line's moment and listen. No transcription model — local or cloud — is reliable enough to quote without this step, and anyone who tells you otherwise hasn't been corrected by an editor yet.
- Export to your CMS as .txt (free) or Markdown (Pro), and delete the local files when the piece runs. Deletion is a file operation, not a support request.
Field details that matter
- Offline works, on purpose. Recording and transcription both run disconnected — useful on a source's site with no trusted network, or on a story where you'd rather the laptop not phone home at all.
- Speaker profiles help series reporting. On Pro, QuickSpeak learns each voice on your device and applies your naming across future interviews — useful when the same three sources appear across months of coverage. The voice models live in your browser profile; delete them from the Speakers page any time.
- Consent still applies. Recording laws bind journalists exactly like everyone else — many jurisdictions require all-party consent. Disclosure is standard practice in the profession anyway; local processing changes the storage story, not the consent one.
What this costs
Nothing, for the loop a reporter actually needs: unlimited interviews recorded and transcribed, speakers named, full history, .txt export. There's no minutes meter to blow through during a heavy news week. Pro ($12/month or $99/year) adds cross-meeting search — "which source mentioned the memo?" across your whole archive — plus Markdown export and on-device AI summaries for long transcripts. Every install starts with 14 days of full Pro, no card required, which is also one more thing a newsgatherer appreciates: no card.
FAQ
Why avoid cloud transcription for interviews?
Uploads create a vendor-held copy of your source's voice — exposed to legal process against the vendor, vendor breaches, and vendor retention policies. Local transcription leaves no third party in the pipeline.
Does local transcription protect sources completely?
No. Your laptop, email, and backups are still the bigger surface. It removes one exposure, and it's the one you'd otherwise be adding blindly.
How accurate is it for interviews?
Strong on clean audio, weaker in noisy rooms — the same physics as every model. Keep the audio (QuickSpeak does, locally) and verify quotes by ear before print.
Can I redact a source's name from a transcript?
You label speakers whatever the story requires, including pseudonyms, and the label applies throughout. The underlying audio stays local, so what leaves your machine is only what you export.