What a QA Scorecard Should Look Like When It's Built for Machine Scoring, Not Human Reviewers

Published on:
September 15, 2026

A QA scorecard built for machine scoring needs objective, transcript-verifiable criteria (did the agent say X, did the agent do Y within Z turns) instead of subjective judgment calls, because a model can only score what it can locate and verify in the text of a conversation. Most scorecards in use today were written for a human reviewer sitting with a checklist and a handful of tickets. When you point an AutoQA engine at that same scorecard, the ambiguous criteria (tone was "professional," agent showed "empathy") produce inconsistent scores, not because the model is weak, but because the QA scorecard was never designed to be machine-checkable. At Revelir AI, we score 100% of customer service conversations for companies like Xendit and Tiket.com, and rebuilding scorecards for machine scoring, not translating them, is the single biggest driver of scoring accuracy we see in production.

TL;DR

  • Human-built scorecards lean on subjective judgment ("sounded empathetic"); machine-scorable scorecards need binary, transcript-verifiable criteria the model can locate in the text.
  • Industry frameworks like the 4Cs still apply, but AutoQA works best when each criterion is rewritten as a yes/no or scaled check tied to specific policy language, not a vibe [verint.com][balto.ai].
  • Manual QA reviews only 1 to 5 percent of conversations; a scorecard designed for machine scoring is what makes 100 percent coverage possible.
  • AI eliminates reviewer fatigue and inter-rater variance but needs the QA scorecard to compensate for what it doesn't do well, like reading emotional subtext, through explicit sub-criteria.
  • Scorecards should be retrieved against the company's actual SOPs at scoring time, not hardcoded once, so the AI is judging against current policy, not a stale template.

About the Author: This article is written by the team at Revelir AI, which builds RevelirQA, an AI quality assurance platform that scores 100% of customer service conversations against a company's own SOPs. Revelir's scorecards run in production across fintech and travel platforms handling thousands of tickets a week in English, Indonesian, Thai, and Tagalog.

Why Doesn't a Human QA Scorecard Just Work for AI Scoring?

A scorecard designed for a human reviewer assumes the reviewer brings context the document doesn't spell out. A criterion like "agent handled the customer professionally" works fine when a QA lead has years of calls in their head to calibrate against. A large language model has no such tacit calibration; it only has the words in front of it and the instructions in the scorecard. Ask it to judge "professionalism" without a definition, and it will produce a plausible-sounding score that may not mean the same thing from ticket to ticket. Industry research on scorecard design confirms this split directly: human-reviewed scorecards can assess subjective emotional nuance, while automated scorecards need to be built around objective, transcript-based criteria such as binary compliance checks and script adherence. That's not a limitation to work around. It's the design constraint that should shape the entire scorecard from the first line.

What Makes a Criterion "Machine-Scorable"?

A machine-scorable criterion is one where the correct score can be derived entirely from evidence present in the transcript, without requiring the model to guess at intent or infer something never stated. This is the core design test: could two different reviewers, both reading only the transcript and the criterion text, agree on the score? If the answer depends on tone of voice, unstated company norms, or "reading between the lines," the criterion isn't ready for automated scoring yet. Traits of a well-formed machine-scorable criterion:

  • Binary where possible. "Did the agent confirm the customer's account ID before processing a refund? Yes/No" beats "Agent verified customer identity appropriately."
  • Anchored to specific policy language. The criterion references the exact refund window, disclosure line, or escalation trigger from the SOP, not a paraphrase.
  • Scoped to a locatable moment. "Within the first two agent turns" or "before the ticket was closed" gives the model a place to look, rather than asking it to judge the conversation as a whole.
  • Free of compound judgments. One criterion, one thing being checked. "Agent was polite and resolved the issue efficiently" is two different judgments smuggled into one line item.
Industry guidance on scorecard design points at the same categories human teams already use, like greeting quality, communication skills, process adherence, and problem resolution [balto.ai]. The shift for machine scoring isn't the categories themselves, it's rewriting each one from a description of a quality into a description of evidence.

How Should QA Scorecard Categories Change for Automated Scoring?

Building on the criterion-level rewrite above, the harder question is whether entire categories need to be restructured, not just reworded. Standard QA frameworks tend to group criteria under headers like greeting and introduction, communication skills, process and compliance, and problem-solving [balto.ai], echoing balanced-scorecard thinking that groups metrics by strategic dimension rather than scoring them all the same way [balancedscorecard.org]. That grouping still holds for auto QA, but each category needs a different scoring mechanism depending on how verifiable it is:

Category Human-reviewer version Machine-scorable version
Greeting/opening "Warm, professional opening" Binary check for required disclosures or brand-mandated opening phrase
Process/compliance "Followed procedure appropriately" Checklist of specific, named policy steps pulled from the SOP, each scored independently
Communication skills "Communicated clearly and empathetically" Split into separate sub-checks: jargon avoided, customer's question restated, apology present where policy requires one
Resolution "Resolved the issue satisfactorily" Did the ticket reach a defined resolution state per the SOP, within policy-defined turn or time limits
This is where scoring against the company's actual policy documents, rather than a generic industry template, matters most. A criterion like "followed refund procedure" is only machine-scorable if the model can retrieve the specific refund SOP for that company at the moment of scoring. This is the design principle behind RevelirQA's approach: policies and SOPs are ingested into a vector database, and the AI retrieves the relevant SOP before scoring each conversation, so the scorecard isn't checking against a static, generic QA scorecard that goes stale the moment a policy changes.

What Do Machine-Scored Scorecards Get Right That Human Sampling Can't?

Stepping back from scorecard mechanics, a separate but related question is what this design work actually buys a QA team. The honest answer starts with coverage. Manual QA sampling typically reviews only 1 to 5 percent of total conversations, which means a scorecard problem, a policy drift, or a training gap can live undetected in the other 95 to 99 percent for weeks. A scorecard engineered for machine scoring is what makes reviewing 100 percent of conversations possible in the first place, because the criteria are cheap and consistent to check at scale. The second gain is consistency itself: machine learning models apply the same QA scorecard without fatigue, mood, or drift between reviewers, delivering uniform scoring across every agent and every shift. Zendesk's own auto QA platform, built into Zendesk QA, is a direct industry acknowledgment of this shift, using AI to score all interactions against predefined criteria rather than a sample. Salesforce Service Cloud generally reaches the same outcome through AppExchange integrations like Leaptree or In-gage rather than a native tool. The pattern across the industry is the same: automated quality assurance scoring is becoming the default expectation, not an experimental add-on.

What Should a Scorecard Still Leave to Human Judgment?

None of this means the scorecard should try to make the model do everything. Industry studies are consistent on this point: AI eliminates human variability in applying a QA scorecard, but it often lacks the contextual reasoning and emotional intuition a human reviewer brings to a genuinely ambiguous case. The practical response isn't to abandon automated scoring for those moments, it's to design the scorecard so those moments get flagged for a human, rather than silently mis-scored. A well-built machine scorecard should include:

  • Escalation flags for conversations where sentiment drops sharply or a criterion can't be resolved from the transcript.
  • A reasoning trace attached to every score, showing which policy document was retrieved and why the model scored the way it did, so a QA lead can audit disagreements instead of guessing at the model's logic.
  • Sentiment arc tracking (start versus end of conversation), which catches a customer who is technically "resolved" per policy but leaves frustrated, a pattern a binary resolution checkbox alone would miss entirely.
This is also where auditability stops being a nice-to-have. In regulated industries like fintech, a scorecard that produces a score with no visible reasoning is a compliance liability, not just a QA gap. RevelirQA attaches a full trace, the model used, the SOP retrieved, and the reasoning applied, to every single score, which is part of why the platform is already running in production QA workflows at a fintech company like Xendit.

Frequently Asked Questions

Can I just reuse my existing QA scorecard for AutoQA?
You can start from it, but expect to rewrite most criteria. Subjective descriptions like "handled professionally" need to become specific, transcript-checkable statements before a model can score them consistently.

Does automated quality assurance replace human QA reviewers entirely?
No. AutoQA replaces manual sampling, reviewing every conversation instead of 1 to 5 percent, but human reviewers still matter for calibrating the QA scorecard, handling flagged edge cases, and coaching agents based on what the scores surface.

What's the difference between AutoQA and auto QA?
They're the same thing, just written differently. AutoQA (one word) and auto QA (two words) both refer to automated quality assurance that scores conversations against a QA scorecard without manual sampling.

How many criteria should a machine-scorable scorecard have?
Fewer, more specific criteria generally outperform a long list of broad ones, since each criterion needs to be independently verifiable from the transcript.

Can a scorecard built for machine scoring evaluate AI chatbots as well as human agents?
Yes, if the criteria are policy-based rather than tied to human behaviors like tone of voice. RevelirQA, for example, scores AI agents and human agents on the same QA scorecard, giving CX leaders one consistent view of quality across both.

Do QA scorecards need to change when company policy changes?
Yes, and this is a common failure point. A scorecard hardcoded to a policy version from months ago will keep scoring against outdated rules. Retrieving the current SOP at scoring time avoids this drift.

About Revelir AI

Revelir AI builds RevelirQA, an AI quality assurance platform for customer service that scores 100 percent of support conversations against a company's own policies and SOPs, retrieved via RAG before every evaluation. Founded in 2025 by Rasmus Chow and headquartered in Singapore, Revelir runs in production at companies like Xendit and Tiket.com, scoring thousands of tickets a week across English, Indonesian, Thai, and Tagalog. Every score carries a full reasoning trace, model, retrieved documents, and reasoning, giving QA and compliance teams an auditable record behind every evaluation. The platform scores both human agents and AI chatbots on the same scorecard, and integrates with any helpdesk, including Zendesk and Salesforce, via API.

If your QA scorecard was built for a reviewer with a clipboard and needs rebuilding for a model that scores every ticket, talk to Revelir AI about what that looks like for your policies.

References

  1. How to Build Call Center QA Scorecards for Better CX | Verint (verint.com)
  2. Call Center Quality Monitoring Scorecard Best Practices | Balto (balto.ai)
  3. Balanced Scorecard Basics (balancedscorecard.org)