How to get speaker names in your transcripts

October 4, 2026·guide·the difference between a record and a wall of text

A transcript without names is half a transcript. "We should move the deadline" means one thing from the client and another from your engineer; "I'll own that" is only actionable if you know who said it. The technical name for separating voices is diarization, and every serious transcription tool does some version of it. The interesting differences are in what happens next: whether you can fix the labels, and whether the fixes stick. This guide covers both, in QuickSpeak (which we make).

How speaker separation works

While a recording transcribes, QuickSpeak also computes a voice embedding — a numeric fingerprint of how each voice sounds — for every utterance, and clusters utterances that share a fingerprint. Those clusters become the labels you see: Speaker 1, Speaker 2, Speaker 3. The whole computation runs on your machine; the fingerprints are as local as the transcript itself.

What no software can do is know which voice is which person. A voice isn't a name tag. So the design is: the machine separates, you name.

Naming speakers in a meeting (free)

  1. Record a meeting as usual — the Identifying speakers stage after you stop runs the diarization. It works per meeting on the free tier.
  2. Open the transcript and click a speaker label. Type the real name: "Speaker 2" becomes "Priya" everywhere she spoke in that meeting.
  3. Wrong split or merge? Use Re-process speakers on the meeting page to re-run detection from the stored audio — useful after a mis-grouped speaker or a change in the room.
  4. Export as usual — the names travel with the .txt, Markdown, or JSON.

Keeping names across meetings (Pro)

Per-meeting labels are disposable. Pro adds speaker profiles: when QuickSpeak hears a voice it has heard before, it matches it to the existing profile and applies the name you already chose.

  1. Record a second meeting with the same people. The Speakers page shows each profile with how many voice samples it has learned from.
  2. Rename once, on the profile. The correction propagates to every future meeting — and because the profile lives on your device, the correction happens locally too.
  3. Delete a profile any time from the Speakers page. That's the whole story of your biometric data here: computed on your machine, stored on your machine, deleted by you. No support ticket, because there's no server holding it.

The compounding effect is the point. The first meeting with a new team costs you four renames; every meeting after costs zero. Compare that with cloud tools, where reviewers of one well-known notetaker describe speaker attribution on a four-person call as a coin flip — the label exists, but you can't trust it or permanently fix it.

Getting good separation in practice

  • Audio quality is the whole game. Speaker separation runs on the same audio as transcription: a real microphone beats a laptop mic, and a room echo beats both.
  • Crosstalk is the enemy of clusters. Two people talking over each other produce blended embeddings. Diarization handles turn-taking well and interruptions badly, everywhere, in every tool.
  • Similar voices need more samples. Two deep voices in a big room may need a rename pass. That's what the rename and re-process affordances are for — a ten-second fix, not a flaw.
  • Say names once, early. If you're recording a call where names matter, greeting people by name gives you an anchor when you label the transcript later.

FAQ

What is speaker identification in a transcript?

Diarization: segmenting a transcript by who was speaking, so it reads as a labeled conversation. QuickSpeak does it with voice embeddings computed on your device; you assign the names.

Can it detect real names automatically?

No tool can — a voice doesn't carry its name. QuickSpeak separates the voices automatically and you name them once. On Pro, that name persists across every future meeting via speaker profiles.

Is it free?

Per-meeting speaker separation and labeling: free, unlimited. Cross-meeting profiles: Pro, with 14 days free on every install.

Where do the voice models live?

In your browser profile on your disk. They're computed locally from the meeting audio and never uploaded; deleting a profile deletes them.