A quality reviewer with headphones listening to a recorded call and marking a printed scorecard with a pen

/Free tool · AI QA scorecard

Call center QA scorecard for human and AI agents.

This call center QA scorecard grades any customer interaction on one rubric, whether a person or an AI agent handled it. Score a call, chat, email or WhatsApp thread against 13 weighted criteria, see whether it passes, and keep a running log you can export to a spreadsheet. It runs in your browser and nothing is sent to us.

Interaction details

Opening

Verified identity as policy requires before discussing the account. 8 pts Critical

Mark not applicable if the request didn’t touch account data.

Confirmed the customer’s actual problem before solving it. 10 pts

Restated it or asked one good question. Solving the wrong problem fast counts as missed.

Resolution

Every fact, price and policy given was correct. 15 pts Critical

One wrong commitment is a miss. For AI agents, check anything specific: dates, amounts, policy names.

Resolved the issue, or set a clear next step with a time. 15 pts

“Someone will get back to you” without a when is partial at best.

No needless holds, transfers or repeated questions. 5 pts

Asking for information the customer already gave counts against this.

Stayed in scope and said so when it didn’t know. 7 pts

No guessing, and no legal, medical or financial advice. This is where AI agents most often fail.

Communication

Tone fit the situation and acknowledged frustration where there was some. 8 pts

Scripted empathy repeated three times is partial.

Plain language, no internal jargon, and checked the customer understood. 5 pts

Long AI answers that bury the step the customer needs count as partial.

Answered in the customer’s language at a professional level. 5 pts

Includes switching language when the customer did.

Compliance and handoff

Gave every required disclosure. 8 pts Critical

Recording notice, payment terms, and telling the customer when they’re talking to an AI where that’s required.

Asked for and repeated only the data the task needed. 4 pts

Full card numbers read back aloud or pasted into chat are a miss.

Escalated or handed off when it should have, with context passed along. 5 pts

If the customer had to repeat themselves after a transfer, it’s partial at best.

The record states the issue, the outcome and the next step. 5 pts

Could the next agent pick this up without asking the customer anything?

Your log

Interactions you add are kept in this browser. Download the CSV to share them.

Nothing logged yet.

Saved in this browser only. Nothing is sent to us.

Why one scorecard for people and AI agents

Most contact centers now run some mix of human agents, AI chat and voice agents, and AI that drafts replies for people to send. If each runs on its own scorecard, you can’t answer the question leadership will ask: is the AI as good as the team, and where isn’t it?

The failures that hurt customers are the same either way. A wrong price, a missing disclosure or a handoff where the customer has to repeat everything is a failure whoever caused it. So the rubric scores the outcome the customer got, and you tag who handled it.

How scoring works

Each criterion carries points, and the total across all 13 is 100. Met scores the full points, partial scores half and missed scores zero. Not applicable drops the criterion out of the calculation, so an email isn’t marked down for skipping a recording notice.

Three criteria are critical: identity verification, accuracy and required disclosures. Missing any of them fails the interaction whatever the total, because a 92 with a privacy breach isn’t a good call. Partial on a critical item doesn’t trigger the fail, but it should trigger a conversation.

The bands are our defaults: 90 and above is excellent, 80 to 89 meets the standard, 70 to 79 needs coaching, and anything lower is below standard. Adjust them to your program and keep them fixed for at least a quarter, or the trend line means nothing.

Scoring AI agents fairly

Score AI agents on the transcript the customer saw or heard, not on the logs behind it. Two criteria do most of the work. Accuracy catches invented answers, which in AI agents usually look confident and specific. Staying in scope catches the agent answering questions it shouldn’t, such as legal or medical advice.

Sample AI interactions on purpose, not at random. Pull the conversations where the customer asked for a person, where the agent handed off, and where the conversation ran long. Random samples of AI traffic are mostly easy questions handled well, which tells you little.

Where AI drafts and a person sends, tag the interaction hybrid. If hybrid scores drop on accuracy, agents may be approving drafts without reading them.

Reading QA scores next to CSAT, FCR and handle time

A QA score tells you whether an interaction met your standard. It doesn’t tell you whether the customer was happy or whether the problem stayed solved. Put it beside the numbers you already track by agent and queue: customer satisfaction (CSAT), first contact resolution (FCR) and average handle time.

Read them together. High QA with low CSAT often means the standard is wrong, or the policy behind it frustrates customers. Low QA with short handle times can mean agents are rushing. If FCR drops while QA holds steady, check how reviewers mark the resolution criterion.

Adapting the rubric to your program

Start from the outcomes your program is judged on, then work back to behaviors a reviewer can hear or read. Each criterion should pass one test: would two reviewers mark the same interaction the same way? If not, rewrite it.

Keep critical items for failures that cause harm on their own, such as a privacy breach or a wrong commitment. Give the most points to what drives your main outcome, which is usually accuracy and resolution. A sales queue might add a criterion for recommending the right product. A collections queue would add the disclosures its regulator requires.

Review the rubric every quarter. A criterion that’s met in nearly every interaction has stopped telling you anything, so drop it or fold it into another.

Running calibration with this sheet

Pick five interactions. Have every reviewer score them alone, then compare. Wherever two reviewers differ by more than one rating on a criterion, rewrite the guidance for that criterion until they’d agree. Do it monthly, and include whoever reviews AI transcripts in the same session as the human-call reviewers.

The log keeps every interaction you score in this browser, with the average. Download it as CSV before calibration and paste it into your team’s sheet.

Frequently asked questions

What is a call center QA scorecard?

It’s the rubric a quality team uses to grade customer interactions. Each criterion has points, the total gives a score, and critical items such as compliance can fail the interaction on their own. Scores are used for coaching, calibration and spotting trends by agent, queue or channel.

What is an AI QA scorecard?

It’s a QA scorecard used to grade interactions handled by AI agents, or a scorecard applied automatically by AI to every interaction. This one is the first kind. It uses the same criteria for people and AI so you can compare them directly, which is the usual starting point before automating the scoring.

Which criteria are critical?

Identity verification, accuracy of facts and policies, and required disclosures such as the recording notice or telling the customer they’re talking to an AI. A miss on any of them fails the interaction. Some programs add data handling to the critical list. Add it if your regulator or your clients would.

How many interactions should we score per agent?

Enough to see a pattern, not a bad day. Many teams start with a handful per agent per week and adjust from there. Automated QA tools can score every interaction, but a manual sample like this one is still how you check the automation is right.

How do you turn QA scores into coaching?

Pick one behavior per session, not the whole sheet. Review the interaction together, let the agent score it first, then compare. Agree on one thing to keep doing and one thing to change, and write both in the notes. Check that behavior in the next few reviews, so the agent sees the coaching led somewhere and the score reads as feedback rather than a verdict.

Is my data stored or sent anywhere?

No. Scores and the log are saved in your browser’s local storage only. Nothing reaches OTRO. Don’t paste customer personal data into the notes. Use the interaction ID from your own system instead.

Can I change the criteria or points?

Not in the tool itself. Print it or export the log and adapt the rubric in your own sheet. The criteria are a sound default for most support and sales queues, and the bands are labelled as defaults for that reason.

Want every interaction scored, not a sample?

The bilingual teams we manage through partner centers in Mexico and Colombia have every call, chat and email scored by AI, on a scorecard we calibrate with you each month.

See the call center service