Meeting transcription: the complete guide

October 5, 2026·by the QuickSpeak team·a field guide, not a brochure — we make a transcription tool and link every claim

Every meeting produces the same raw material — an hour of speech — and leaves you with the same problem: speech is useless to search, quote, or remember by. Transcription is the fix, and since speech-recognition models got good enough to run anywhere from a data center to a laptop fan, the market has exploded into a dozen architectures, three capture methods, and a wall of pricing pages. This guide is the map: how the technology actually works, what moves accuracy, what the law expects of you, and how to pick a tool you won't regret.

How a transcript gets made

Two questions define every tool on the market, and vendors prefer you ask neither directly.

Question one: how does the audio get captured?

  • A bot joins the meeting. Otter, Fireflies, and tl;dv's default mode send a visible participant that records from inside the call. Maximum convenience (it happens whether you remember or not), with costs: clients ask what it is, some organizations block unknown participants outright, and the recorder's presence changes how people talk.
  • An extension or desktop app captures audio your machine already plays. QuickSpeak, Tactiq, Granola, Krisp, Fathom's bot-free mode. Nothing joins the call; the tool listens to the meeting tab's audio or your system audio or your microphone. Capture becomes your decision, made per meeting.
  • File upload. You record however you like and hand the tool an audio file. Slowest loop, most control.

The honest subtlety: capture method and privacy are separate axes. A bot that announces itself is a disclosure mechanism; a bot-free tool shifts disclosure duty to you. We've written a whole comparison of the nine major tools by capture and processing if you want the full field.

Question two: where does processing happen?

A cloud tool uploads the audio and transcribes it on the vendor's servers; the transcript — and usually the audio — then lives in the vendor's storage, governed by their retention policy, reachable through your account. An on-device tool runs the recognition model on your own hardware: the model is downloaded once, cached, and executes in your browser, so there is no upload step and nothing works without your machine. Cloud buys frontier-model polish and team features. On-device buys custody, offline capability, and free tiers that never run out — because the compute bill is yours.

What a good transcript needs

A wall of raw text is half a transcript. The features that separate useful from tedious:

  • Speaker labels. "We should move the deadline" means different things from a client and from your engineer. Separation is done by diarization — the model clusters voices — and the names are yours to assign. How speaker naming works, and how to make it stick.
  • Search across meetings, not just within one. "When did we discuss the pricing change?" is a question about your archive, not one call. Most tools gate this; check before you build a library.
  • Exports that don't expire. Your transcript should outlive your subscription. Tools that hold exports behind paywalls are holding your meeting history hostage — we compared every free tier on exactly this.
  • The audio itself, kept. Transcription models are good, not perfect. When a quote matters, you play the moment. A tool that discards audio (Tactiq's model, for instance) can't support that.

What actually moves accuracy

Model choice matters less than vendors admit; audio quality matters more. In rough order of impact:

  1. Distance to the microphone. A laptop mic in a conference room hears the first two rows. Everything downstream — word accuracy, speaker separation — degrades with distance.
  2. Crosstalk. Overlapping speech produces blended audio no model separates cleanly. Turn-taking transcribes well; interruptions don't, in every tool.
  3. Jargon and names. Models trained broadly fumble your product codename. Tools with custom vocabulary (QuickSpeak Pro has it) let you teach the correct spelling once.
  4. Accents and noise. Modern multilingual models (NVIDIA's Parakeet family, Whisper) handle accents far better than their ancestors, but the gap between clean and messy audio is still the largest single factor.
  5. The model itself. Mostly a tie among current-generation options on clean audio — which is why the free-tier test beats the marketing page: run your real meetings through two tools and compare.

Consent and law, in plain terms

Recording law attaches to you, not your tool. Many jurisdictions require one party's consent (yours suffices); others — California, Illinois, and much of the EU in practice — expect everyone on the call to know. A visible bot doubles as disclosure; a bot-free tool makes disclosure your job. The one-sentence version: tell people you're recording, and check the rule where your participants are, not just where you sit. For professions with stricter duties, we've written specifically about therapy sessions and journalism — both generalize: consent first, custody second, tool last.

Where transcripts should live

The custody question decides more purchases than feature lists. A cloud transcript is convenient, shareable, and searchable from anywhere — and it exists on a vendor's infrastructure under the vendor's policies, exposed to whatever legal process or breach reaches that vendor. A local transcript is a file: portable, deletable, offline-capable, and private by absence of a server. Neither is wrong. But the choice should be made consciously — most people make it by default, because a bot joined their calendar one day. If you're evaluating the trade, the per-tool write-ups — Otter, Fireflies, Fathom, Granola, Tactiq, tl;dv — document each vendor's data flow with links to their own pages.

Choosing a tool, in five questions

  1. Does a bot join calls — and is auto-join default-on? (That's the tool deciding when other people get recorded.)
  2. Where does processing happen, and where does the recording live afterwards? Which region, what retention, whose keys?
  3. Is my content used for training, and is that default-off or default-on with an opt-out? Read the privacy page, not the pricing page.
  4. What does the free tier include when your real calendar hits it — minutes, history, exports?
  5. Does it work with the network off? The only demo that proves the privacy claim. Try it in airplane mode.

A short glossary, because the marketing pages won't define them

ASR / STT
Automatic speech recognition; speech-to-text. The model that turns audio into words.
Diarization
Segmenting a transcript by who was speaking. Done with voice embeddings — numeric fingerprints of how voices sound — clustered across the recording.
Live captions vs. batch transcription
Streaming models output words as they're spoken (lower peak accuracy); batch models process the complete recording after it ends (higher accuracy, no live view). Good tools offer both paths.
WASM / WebAssembly
A way to run compiled model code inside a browser sandbox — the mechanism that makes browser-based on-device transcription possible without installing anything.

Where to go next

If you want the shortlist, start with the nine-tool comparison. If you know your constraint is privacy, the Otter and Fireflies write-ups cover the cloud defaults worth knowing. If it's a specific meeting platform, the Meet, Zoom, and Teams guides walk the actual clicks. And if it's languages — 36 of them work, in the browser.