Skip to main content

ROI Analysis

Per-Call QA Scoring: Every Call Instead of a Sample of Five

Traditional call QA scored five calls a month per person because listening was expensive. Per-call QA scoring at full coverage changes what QA is even for.

8 min read

Per-call QA scoring sounds like more of what practices already do, and it is closer to the opposite. Traditional call quality assurance existed to sample. A supervisor listened to five calls per person per month, scored them against a rubric, and used the result in a performance conversation. Every part of that design, including its purpose, was determined by the fact that listening to a call cost twenty minutes of a supervisor’s day.

When the cost of reviewing a call approaches zero, keeping the old design produces something worse than nothing. A five-call sample scored out of a thousand calls has an enormous margin of error, so it cannot reliably tell a good month from a bad one. It is also, at that sample size, mostly a measurement of which five calls the supervisor happened to pull. Practices then attach performance consequences to it, which is how QA becomes something staff manage around rather than something the practice learns from.

Full coverage changes what QA is for

This is the part that gets missed. Moving from five calls to every call is not a precision upgrade to the same activity. It changes the function.

Sampling-based QA is an audit. Its unit of analysis is a person, its output is a score, and its use is performance management. That is what a small sample can support and roughly all it can support.

Full-coverage scoring is process detection. Its unit of analysis is the call type, its output is a rate, and its use is finding which parts of the front-office process fail and how often. A practice can now say that insurance verification was skipped on 12 percent of new-patient calls this month, which is a workflow finding with an owner. It could never say that from five calls, and if it tried, the number would be noise.

Practices that carry the audit framing into full coverage get the worst of it. Scoring every call and attaching each score to an individual produces a surveillance system, predictable resentment, and staff who optimize for the rubric. Scoring every call to find where the process leaks produces the thing everyone actually wanted.

Score only what two reviewers would score the same way

Full coverage is only meaningful if the scoring is mechanical, which forces a useful discipline on the rubric.

The scoreable items are the checkable ones. Were the practice’s identifiers confirmed. Was insurance verified or flagged for verification. Were the required intake fields collected. Was the appointment read back. Was the escalation rule followed when a listed phrase came up. Was a callback commitment given and recorded. Each is a yes or no that any two people would answer identically from the transcript.

What cannot be scored this way is everything about tone, warmth, and whether the caller felt looked after. Those matter, and practices that try to automate them end up with a number that collapses the first time somebody disputes it. Keep them, but keep them in a separate, small, human-reviewed sample. Two instruments with two purposes beats one instrument that quietly does neither.

The rubric should also be short. A twenty-item scorecard produces a composite score that averages away exactly the signal you needed, because a call can fail the one item that mattered and still score 90 percent.

Rates by call type, not averages by person

The reporting shift matters as much as the collection shift, and it is where most implementations quietly revert to old habits.

An average quality score per staff member is the natural output of an audit and nearly useless as a process instrument. It blends call types with different difficulty, it moves slowly, and it invites arguments about attribution instead of about the process.

What to report instead is a pass rate per rubric item, cut by call type. Verification captured on new-patient calls. Appointment read back on reschedules. Callback commitment recorded on message-taking calls. Each of those is a specific line in the process, and a low number points at one thing rather than at a person.

The cut by call type is what makes it actionable. Verification failing on 4 percent of established-patient calls and 30 percent of new-patient calls is not a training problem, it is a workflow that was designed for one path and not the other. An average across both would show something mediocre and unexplained.

When an individual genuinely is an outlier, full coverage makes that unambiguous and easy to handle on its own terms. It just should not be what the reporting is built around.

Can you prove the program did anything

QA programs get cut in budget reviews for the same reason follow-up programs do, and the broader evidence gives finance a reason to push back. MGMA found that among practices using AI in patient visits, about 44% said it had not reduced staff workload, 39% said it had, and 17% were unsure.

That spread is mostly about instrumentation rather than about the tools themselves. A practice that defined its measures before turning something on can say which bucket it is in and show the arithmetic. A practice that did not is offering an impression.

For scoring, the measures that survive scrutiny are the ones tied to an outcome rather than to the score. Pass rate on verification against downstream eligibility-related denials. Pass rate on appointment read-back against no-show rate for those appointments. Callback commitments recorded against callback commitments met.

Each of those is a claim the practice can check against a system it already runs, which is a considerably stronger position than defending a composite quality score that only exists inside the QA tool.

Roll it out as measurement first

The sequencing decision determines whether staff cooperate, and it is worth being deliberate about.

Run the first stretch as measurement only. No individual attribution, no performance use, results reported at the process level, and say so plainly at the start. This is not primarily a trust gesture, although it is that. It is that the first month of full-coverage data is mostly a report on the rubric itself, and a rubric that has never met real calls always contains two or three items that are ambiguous, unfair, or measuring something other than what was intended.

Fix those with the people whose calls are being scored. They will find the problems faster than anyone reviewing transcripts, and the rubric that comes out of that process is one they helped write.

After that, publish the process-level rates openly and keep them there. The practices where this works well are the ones where the front office sees the same numbers management sees, because the numbers are about the process rather than about them, and improving them becomes a shared piece of work rather than a thing being done to somebody.

Key Takeaways

  • Full coverage is not a bigger sample, it is a different function. Sampling QA audits people; full-coverage scoring detects process failures and reports rates.
  • A five-call monthly sample cannot distinguish a good month from a bad one. Attaching performance consequences to it teaches staff to manage the sample.
  • Score only items two reviewers would score identically. Tone and warmth belong in a separate small human-reviewed sample, not in an automated composite.
  • Keep the rubric short. A twenty-item scorecard averages away the one failure that mattered and still reports 90 percent.
  • Report pass rate per rubric item cut by call type, not an average score per person. Verification failing on new-patient calls but not established ones is a workflow gap, not a training gap.
  • Launch as measurement only, with no individual attribution, and fix the rubric with the people being scored before it carries any consequences.

The reason to score every call is not thoroughness for its own sake. It is that a rate calculated across a thousand calls supports conclusions a five-call sample never could, and those conclusions are about the process rather than about the person who happened to answer. Practices that make that shift stop having monthly conversations about individual scores and start having weekly conversations about which step of which call type is failing, which is both more useful and considerably easier for everyone involved to be in the room for.

Sources

Ready to See It in Action?

See how PGA scores every front-office call instead of a monthly sample

Schedule a Demo →

Written by Kevin Henrikson