AI contact centers: how to measure quality and tone at scale

When AI handles thousands of conversations, measuring quality and tone at scale becomes the challenge. How to do it without reviewing everything by hand.

Natalia Nario
Natalia Nario
· Product Manager · ArtificialQA
AI contact centers: how to measure quality and tone at scale

The contact center was one of the first places where AI stopped being an experiment. In 2026, voice AI already handles 19% of the inbound volume of contact centers —up from just 6% in 2024—, with banking and telecommunications leading adoption.[1] But that leap in scale brings a new problem: when a human handled 50 conversations a day, a supervisor could listen to a few and make corrections. When AI handles tens of thousands, how do you know whether quality holds up —or whether it’s degrading without anyone noticing?

This article is for the CX and QA teams that operate conversational AI at scale: what to measure, which metrics mislead, and how to maintain quality when volume makes manual review impossible.

The metric that misleads: beware of the “deflection rate”

Let’s start with the most common mistake. For years, the star metric for chatbots was the deflection rate: what percentage of queries the AI resolved without passing them to a human. It sounds good —fewer queries to the human, less cost— but it measures what the AI did, not what the customer achieved.[2]

The example is telling: a customer asks about the returns policy and receives an automatic link to a FAQ page. The system logs that interaction as “successfully deflected.” But did the customer find their answer? Did they complete their return? Will they buy again? The metric stays silent, because deflection measures the bot’s activity, not the customer’s outcome.[2:1]

The consequence of optimizing the wrong metric is expensive: 61% of customers who end up talking to a human agent do so out of dissatisfaction with the chatbot.[3] A bot that “deflects” a lot but resolves little doesn’t save costs —it shifts them, along with an angrier customer, onto the human who handles the case afterward.

The metric that does matter is resolution: was the problem resolved without the need for a follow-up contact or an escalation? That’s the one that separates an AI that informs from one that resolves.[2:2]

What to measure when quality is at scale

In a contact center, quality is not a single thing: it’s a set of dimensions that must be monitored in every conversation, not in a sample a supervisor manages to listen to.

Dimension Why it matters in a contact center
Resolution / completeness Whether the customer resolved their problem, not just whether the bot replied
Tone and empathy In a service interaction, the how weighs as much as the what
Accuracy and hallucinations An invented policy or fact generates a later complaint
Correct escalation When and how it hands control to a human, without trapping the customer
Consistency across channels The same correct answer by voice, chat and text

Tone deserves a separate mention, especially in Spanish-speaking markets. AI that feels helpful is welcome; AI that feels like a wall to avoid speaking with a person generates rejection. Ensuring the agent maintains an empathetic register —especially in complaints or difficult situations— is not an aesthetic luxury: it’s an experience differentiator that directly impacts satisfaction and retention.

The problem scale makes critical: silent degradation

There’s a risk inherent to operating AI at high volume that the “test before launch and you’re done” model doesn’t cover: silent degradation. An agent that worked well at launch can start failing weeks later —because the provider updated the model underneath, because the type of queries changed, because the context became outdated. With a volume of tens of thousands of conversations, that degradation isn’t detected by a supervisor listening to random calls: by the time someone notices, it has already affected thousands of customers.

The only defense is continuous monitoring: automatically evaluating a significant portion of the real conversations, not an anecdotal sample, and having alerts when a metric drops. Testing an AI contact center doesn’t end at go-live; that’s exactly where the part that matters most begins.

How ArtificialQA solves it

Reviewing the quality of an AI contact center by hand is unfeasible at scale, and looking only at the deflection rate is fooling yourself. ArtificialQA connects to your agent —by URL or by API, no code— and evaluates it systematically on the dimensions that matter here: resolution, tone and empathy, accuracy, correct escalation and consistency. And because it allows continuous evaluation, it detects degradation before the customer notices it —not after the complaints pile up.

For a CX or QA team, that means going from “I think the bot is working fine” —based on the few conversations someone managed to review— to having a reliable number over the whole volume, one you can track over time. Because at scale, quality that isn’t measured is not quality: it’s an assumption waiting to turn into a complaint.


Frequently asked questions

Why is the deflection rate a misleading metric? Because it measures what the bot did (how many queries it avoided passing to a human), not what the customer achieved. A query “deflected” to a FAQ can leave the customer without resolving their problem. The metric that matters is real resolution.

Which metrics should you measure in an AI contact center? Resolution/completeness, tone and empathy, accuracy and hallucinations, correct escalation to a human, and consistency across channels —in every conversation, not in a small sample.

How do you maintain quality when AI handles tens of thousands of conversations? With continuous, automatic monitoring of a significant portion of the real conversations, instead of manual review of samples, and with alerts for metric drops to detect silent degradation.

What is silent degradation in an AI contact center? It’s the drop in quality that occurs after launch —due to model updates, changes in queries or outdated context— and that, because of the volume, isn’t detected by a supervisor listening to random calls until it has already affected many customers.

Why does tone matter so much in a contact center? Because in a service interaction the how weighs as much as the what. An AI that feels like a wall to avoid speaking with a person generates rejection and affects satisfaction and retention.

#contact-center#industries#tone
Natalia Nario
Natalia Nario
Product Manager · ArtificialQA

Product Manager of ArtificialQA at QAlified, with 15+ years in software testing and automation. She works at the intersection of quality and AI: designing and evaluating approaches to test non-deterministic systems and ensure their behavior in production.

Put these ideas into practice

We'll help you apply AI testing to your own agent. Leave your details and we'll set up a demo.

  1. DigitalApplied, Customer Service AI Agent Statistics 2026, citing Forrester Wave: voice AI handles 19% of contact center inbound volume in 2026 vs 6% in 2024; banking and telco lead. ↩︎

  2. Notch.cx, 2026 Customer Service AI Metrics: why deflection rate measures the bot’s activity and not the customer’s outcome; the FAQ example; resolution (First Contact Resolution / no follow-up contact) as the metric that matters. ↩︎ ↩︎ ↩︎

  3. Galileo-ft, How Banks Can Improve Customer Support with Conversational AI (2025): 61% of customers who contact a human agent do so out of dissatisfaction with the chatbot; reduced abandonment with intelligent escalation. ↩︎