Almost any service today turns a call into text, and with very good accuracy. That is why transcription stopped being the problem. The problem is that running text, on its own, lets you audit nothing.
Take a real sentence from a collections call: «if you don’t pay today you’ll be reported». Read in flat text it says almost nothing. It depends on who said it and when they said it within the conversation. That is where the two pieces of this article come in.
Diarization: who is speaking
Diarization — speaker diarization — is the process that splits a recording into turns and assigns each one to a speaker. It does not identify the person by name: it distinguishes that there are two voices and marks which is which throughout the audio.
Without it, the sentence in the example could have been said by the agent — which would be an improper threat — or by the customer repeating what they were told somewhere else. The text is identical; the meaning, opposite. An audit that cannot tell those two apart is not an audit.
Diarization also enables an observable that says a great deal about a call on its own: how speaking time is split. An agent taking ninety percent of a service call is not serving anyone, they are reciting.
Timestamp: when they said it
The timestamp is each turn's position within the audio: start and end, in seconds. It looks like a technical detail, and it is what separates an opinion from proof.
- It makes every observation verifiable. "The agent did not identify themselves" is a claim; "the agent did not identify themselves, and here are the first twenty seconds" is evidence anyone can check in ten seconds.
- It lets you audit sequence, not just content. Many protocols require certain things to happen BEFORE others: identify yourself before asking for data, disclose the recording before continuing. Without times, sequence is invisible.
- It makes disagreement reviewable. When a supervisor and an agent disagree, the discussion is settled by replaying the minute, not by comparing memories.
How they fit into Qualidot
In Qualidot both are part of the same path, and neither is the product: they are the condition for the product to be worth anything.
- The audio is transcribed and diarized: you get turns, each with its speaker and its times.
- Observables are computed over those turns — who spoke how much, whether there were interruptions, at what point the phrases the rubric requires or forbids appeared.
- The company rubric is applied to those observables, not to loose text.
- Every observation stays anchored to its minute, so the review is reproducible by anyone who wants to check it.
It is the same idea that holds up class auditing: an evaluation that cannot be checked again is not an evaluation, it is an opinion in the shape of a report.


