AI contact centers: how to measure quality and tone at scale
When AI handles thousands of conversations, measuring quality and tone at scale becomes the challenge. How to do it without reviewing everything by hand.

The contact center was one of the first places where AI stopped being an experiment. In 2026, voice AI already handles 19% of the inbound volume of contact centers —up from just 6% in 2024—, with banking and telecommunications leading adoption.[1] But that leap in scale brings a new problem: when a human handled 50 conversations a day, a supervisor could listen to a few and make corrections. When AI handles tens of thousands, how do you know whether quality holds up —or whether it’s degrading without anyone noticing?
This article is for the CX and QA teams that operate conversational AI at scale: what to measure, which metrics mislead, and how to maintain quality when volume makes manual review impossible.
The metric that misleads: beware of the “deflection rate”
Let’s start with the most common mistake. For years, the star metric for chatbots was the deflection rate: what percentage of queries the AI resolved without passing them to a human. It sounds good —fewer queries to the human, less cost— but it measures what the AI did, not what the customer achieved.[2]
The example is telling: a customer asks about the returns policy and receives an automatic link to a FAQ page. The system logs that interaction as “successfully deflected.” But did the customer find their answer? Did they complete their return? Will they buy again? The metric stays silent, because deflection measures the bot’s activity, not the customer’s outcome.[2:1]
The consequence of optimizing the wrong metric is expensive: 61% of customers who end up talking to a human agent do so out of dissatisfaction with the chatbot.[3] A bot that “deflects” a lot but resolves little doesn’t save costs —it shifts them, along with an angrier customer, onto the human who handles the case afterward.
The metric that does matter is resolution: was the problem resolved without the need for a follow-up contact or an escalation? That’s the one that separates an AI that informs from one that resolves.[2:2]
What to measure when quality is at scale
In a contact center, quality is not a single thing: it’s a set of dimensions that must be monitored in every conversation, not in a sample a supervisor manages to listen to.
| Dimension | Why it matters in a contact center |
|---|---|
| Resolution / completeness | Whether the customer resolved their problem, not just whether the bot replied |
| Tone and empathy | In a service interaction, the how weighs as much as the what |
| Accuracy and hallucinations | An invented policy or fact generates a later complaint |
| Correct escalation | When and how it hands control to a human, without trapping the customer |
| Consistency across channels | The same correct answer by voice, chat and text |
Tone deserves a separate mention, especially in Spanish-speaking markets. AI that feels helpful is welcome; AI that feels like a wall to avoid speaking with a person generates rejection. Ensuring the agent maintains an empathetic register —especially in complaints or difficult situations— is not an aesthetic luxury: it’s an experience differentiator that directly impacts satisfaction and retention.
The problem scale makes critical: silent degradation
There’s a risk inherent to operating AI at high volume that the “test before launch and you’re done” model doesn’t cover: silent degradation. An agent that worked well at launch can start failing weeks later —because the provider updated the model underneath, because the type of queries changed, because the context became outdated. With a volume of tens of thousands of conversations, that degradation isn’t detected by a supervisor listening to random calls: by the time someone notices, it has already affected thousands of customers.
The only defense is continuous monitoring: automatically evaluating a significant portion of the real conversations, not an anecdotal sample, and having alerts when a metric drops. Testing an AI contact center doesn’t end at go-live; that’s exactly where the part that matters most begins.
How ArtificialQA solves it
Reviewing the quality of an AI contact center by hand is unfeasible at scale, and looking only at the deflection rate is fooling yourself. ArtificialQA connects to your agent —by URL or by API, no code— and evaluates it systematically on the dimensions that matter here: resolution, tone and empathy, accuracy, correct escalation and consistency. And because it allows continuous evaluation, it detects degradation before the customer notices it —not after the complaints pile up.
For a CX or QA team, that means going from “I think the bot is working fine” —based on the few conversations someone managed to review— to having a reliable number over the whole volume, one you can track over time. Because at scale, quality that isn’t measured is not quality: it’s an assumption waiting to turn into a complaint.
Frequently asked questions
Why is the deflection rate a misleading metric? Because it measures what the bot did (how many queries it avoided passing to a human), not what the customer achieved. A query “deflected” to a FAQ can leave the customer without resolving their problem. The metric that matters is real resolution.
Which metrics should you measure in an AI contact center? Resolution/completeness, tone and empathy, accuracy and hallucinations, correct escalation to a human, and consistency across channels —in every conversation, not in a small sample.
How do you maintain quality when AI handles tens of thousands of conversations? With continuous, automatic monitoring of a significant portion of the real conversations, instead of manual review of samples, and with alerts for metric drops to detect silent degradation.
What is silent degradation in an AI contact center? It’s the drop in quality that occurs after launch —due to model updates, changes in queries or outdated context— and that, because of the volume, isn’t detected by a supervisor listening to random calls until it has already affected many customers.
Why does tone matter so much in a contact center? Because in a service interaction the how weighs as much as the what. An AI that feels like a wall to avoid speaking with a person generates rejection and affects satisfaction and retention.
Product Manager of ArtificialQA at QAlified, with 15+ years in software testing and automation. She works at the intersection of quality and AI: designing and evaluating approaches to test non-deterministic systems and ensure their behavior in production.


