Blog

How a transcript knows who said what.

Voice activity detection, segmentation, clustering, and the extra step that turns Speaker 2 into a real name. A plain-language tour of speaker diarization, honest about where it still struggles.

A transcript that reads as one unbroken wall of text is nearly useless. What people actually want to know afterwards is not just what was said, but who said it. Did the finance lead commit to that deadline, or did someone merely suggest it? The technology that answers this is called speaker diarization: taking a recording of several people talking and splitting it into "who spoke when". Here is how it works, in plain language, and where it honestly still struggles.

Diarization answers one question about every second of audio: which voice is speaking right now? Everything else, including putting real names on those voices, is built on top of that answer.

Step one: finding the speech at all

Before a system can decide who is speaking, it has to decide whether anyone is speaking. This is voice activity detection: scanning the audio, a few hundredths of a second at a time, and labeling each slice as speech or not-speech. Keyboard clatter, the air conditioner, a chair scraping, a long thoughtful pause: all of that gets marked as silence as far as the transcript is concerned. What survives is a series of speech regions, the raw material for everything that follows. Get this step wrong and everything downstream inherits the error, because a cough classified as speech will eventually be assigned to somebody.

Step two: cutting speech into single-voice pieces

The speech regions now get chopped into segments, stretches of audio that (ideally) contain exactly one voice. The system listens for change points: moments where the acoustic character of the voice shifts because one person stopped and another started. Think of it as drawing cut lines in the audio every time the voice changes, without yet knowing or caring who any of the voices belong to. In a polite, turn-taking meeting these cut lines are clean. In a lively one, as we will get to, they are anything but.

Step three: sorting the pieces into voices

Each segment is then converted into a voice fingerprint: a compact numerical summary of what that voice sounds like, capturing pitch, timbre, and speaking style. Segments from the same person produce similar fingerprints, so the system can cluster them: pile all the similar-sounding segments together, and each pile becomes a speaker. Often the system does not even know in advance how many people were in the room; it has to discover that four piles fit the audio better than three or five.

The output at this point is a fully structured transcript with one catch: the speakers are anonymous. Every line is attributed, but to labels like Speaker 1, Speaker 2, Speaker 3. Useful, and for a recording of strangers it is the best anyone can do. But for a meeting, it is one step short of what you actually want.

Step four: turning "Speaker 2" into a real name

The step most people never see described is the mapping from anonymous voices to actual meeting participants. A meeting, unlike a random recording, comes with context: a calendar invite with an attendee list, and a meeting platform that knows who is present and, crucially, whose microphone is active at any moment. Platforms like Zoom, Teams, Meet, and Webex expose active-speaker and caption signals that say, in effect, "the audio right now belongs to the participant named Omar."

How the matching works

Suppose diarization produced four anonymous voices, and the platform reported four participants. Whenever the platform's active-speaker signal overlaps in time with a diarized segment, that is a vote: Speaker 2's segments keep lining up with moments the platform attributed to Omar Al-Haddad, so Speaker 2 is almost certainly Omar. Accumulate enough votes across the whole meeting and every anonymous voice acquires a name, including for the stretches where the platform signal was missing or ambiguous. This alignment of diarization with platform participant data is how MeetriX puts real names on every line instead of Speaker 1 and Speaker 2.

This is also why a purpose-built meeting assistant can do something a generic transcription API structurally cannot: the API only ever receives an audio file, while the assistant was in the meeting, watching who was talking.

Where diarization honestly struggles

No vendor should tell you diarization is solved, and we will not either. The failure modes are well known, and worth knowing:

  • Overlapping speech. The hardest case in the field. When two people talk at once, the audio is literally a mixture of both voices, and a segment-then-cluster pipeline must either pick one speaker or split the moment imperfectly. Heated discussions and enthusiastic agreement are exactly where overlap spikes.
  • Short interjections. A one-second "تمام" or "yes, agreed" barely contains enough acoustic evidence to fingerprint. Very short segments are the ones most often attributed to the wrong voice, or absorbed into a neighbor's turn.
  • Similar voices. Two speakers with close pitch and timbre, common among colleagues of similar age recorded on the same speakerphone, can blur into one cluster, or one person can split into two.
  • Shared microphones. A conference room where five people share one device removes the per-participant signal a platform would normally provide, so the system leans entirely on acoustics. Uploaded room recordings work, but attribution is genuinely harder there than on a call where everyone has their own microphone.

Good systems mitigate all of these, and platform alignment in particular rescues many of the marginal cases, but "mitigated" is not "eliminated". A fair way to evaluate any tool, ours included, is to run it on one of your own noisy, bilingual, interruption-heavy meetings and read the result with the failure modes above in mind.

Why attribution is worth all this machinery

Everything valuable that comes after the transcript depends on attribution being right. Action items need owners: "I will send the redlines by Thursday" is only a commitment if you know who "I" is. Minutes need decisions traced to the people who made them. Talk-time analytics need each second assigned to the correct person. A diarization error does not just misfile a sentence; it can hand someone else's deadline to the wrong colleague. That is why the whole pipeline, from voice activity detection to name matching, exists: so the record of the meeting reflects not just the words, but the people behind them. You can see what that unlocks downstream in summaries and action items, where every extracted commitment carries an owner pulled straight from the attributed transcript.

Stay in the conversation. MeetriX takes the notes.

It joins the call, transcribes who said what in Arabic or English, and sends the summary and action items before you are back at your desk. 600 free minutes when you connect your calendar.

Arabic & English · 32 Arabic dialects · No credit card required