How on-device speech-to-text actually works

October 5, 2026·engineering·by the QuickSpeak team — this is the system we built

"Your audio never leaves your machine" is easy to say and, until recently, was hard to deliver — speech recognition meant rack-mounted GPUs, and laptops transcribed nothing. That changed. A modern speech model is a few hundred million parameters, WebAssembly executes compiled model code in a browser sandbox at useful speed, and the whole pipeline — recognition, speaker identification, even the AI that writes your meeting summary — fits on the machine the meeting happened on. This is how that pipeline works in QuickSpeak, part by part: the models we run, why the browser is a good runtime for them, and what on-device still can't do.

The pipeline, end to end

Five stages turn a meeting into a transcript, and every one of them executes locally:

  1. Capture. The extension records the meeting tab's audio, your desktop audio, or your microphone. Nothing is transmitted — capture writes into memory on your machine.
  2. Recognition. The audio goes to the speech model, running in WebAssembly inside your browser. Words come out with timestamps.
  3. Speaker separation. In parallel, the diarization step computes a voice embedding — a numeric fingerprint of how each voice sounds — for every utterance, and clusters utterances that share a fingerprint. The clusters become "Speaker 1, Speaker 2"; you supply the names.
  4. Saving. Transcript, audio, and speaker labels are written to your browser profile, on your disk.
  5. Optionally, understanding. On Pro, a small language model — also local — reads the transcript and drafts a meeting brief: summary, key decisions, open questions, action items.

The models

For transcription, the batch pipeline runs NVIDIA's Parakeet TDT 0.6B v3 — a current-generation, multilingual speech model in the 0.6-billion-parameter class, executed through sherpa-onnx compiled to WebAssembly. Batch means the model sees the complete recording once the meeting ends, which lets it produce its best pass; a separate streaming model generates live captions while you record, trading a little accuracy for words appearing as they're spoken. Recognition covers 36 languages, with a per-meeting auto-detect in the language selector.

For speakers, the same idea that powers speaker recognition everywhere applies: utterances become embeddings, and embeddings that sit close together belong to one voice. The embedding model runs locally like everything else — which has a consequence worth noticing for anyone thinking about biometric data: the voice models that let QuickSpeak recognize "this is Priya again" are computed on your machine and stored in your browser profile. There is no server copy, because the server never received the audio that would make one.

For the AI briefs, a small language model runs through llama.cpp compiled to WebAssembly (the wllama runtime). The weights come as GGUF files, downloaded once from a model host and cached by the browser alongside the speech models. When you ask a transcript "what did we decide?", the completion is generated by your CPU, from a model that lives in your cache — the summary of your confidential meeting is written on the machine the meeting happened on.

Why the browser is a good runtime for this

The sandbox is the privacy model. WebAssembly code executes inside the browser's security sandbox with no arbitrary filesystem or network access beyond what the app explicitly does. A transcription app whose only network calls are one-time model downloads and optional anonymous telemetry can be inspected at that boundary — and the offline test makes the claim physical: disconnect your network and record. Everything works, including transcription, because nothing in the pipeline needs a network. The app shows an offline tag when there's no connection; it's a status, not an error.

Threading has to be polite. WASM inference is CPU-hungry, and a transcript that fans every core to 100% while your meeting runs is a transcript nobody keeps. QuickSpeak's local LLM caps itself at four threads and always leaves one core for the interface, so the machine stays responsive during the meeting you're supposed to be paying attention to.

Zero install is the distribution. A Chrome extension plus a web app means no binaries to trust, no admin rights, no platform builds — the same compiled model runs on any desktop browser, and updates ship like any web page. The cost is that first run: the model weights download once (you'll see a progress bar), and the browser caches them from then on.

What on-device still can't do

An engineering post that only lists wins is a brochure, so:

  • Frontier cloud models still win on hard audio. A whisper under a ceiling fan in a fourth language, six voices deep in crosstalk — the largest cloud models have more headroom there. On clean to decent audio the gap is small enough that your own meetings should decide it, not benchmarks.
  • Your hardware is the variance. A modern laptop transcribes an hour-long meeting in minutes; an old machine takes longer and heats up. Cloud vendors hide this spread behind server fleets; local tools wear it.
  • The first run costs a download. A few hundred megabytes of model weights, once. On a metered connection, that's a real cost a cloud tool hides in its bill.
  • Small local LLMs write plainer summaries than frontier cloud ones. For meeting briefs, plain and private usually beats eloquent and uploaded — but it's a trade, and we'd rather name it than bury it.

Verify it yourself

None of this requires trusting us. Install the extension, record two minutes, then switch on airplane mode and record again — transcription works, the offline tag appears, and your network monitor stays quiet. The offline walkthrough covers the setup. The airplane test is the only privacy demo in this category that proves anything, and it works because of how the pipeline is built: there's no upload step to turn off, because there's no server to upload to.

FAQ

How can a browser run a speech recognition model?

WebAssembly: compiled model code executes in a sandbox at near-native speed. Weights download once, the browser caches them, and every run happens on your hardware.

Which models run locally?

Parakeet TDT 0.6B v3 for batch transcription, a streaming model for live captions, on-device voice embeddings for speaker separation, and a llama.cpp-based local LLM for the AI briefs — all executed on your machine.

Is it as accurate as cloud transcription?

Competitive on clean audio; frontier cloud models still win on the hardest audio. The free tier exists precisely so you can run that comparison on your own meetings.

Does the local LLM send my transcript anywhere?

No. The brief is generated by a small model in your browser cache, executing on your CPU. The summary of your meeting is written on the machine the meeting happened on.