Skip to main content

ROI Analysis

Building a QA Rubric Your Front Office Will Not Game

Any call QA rubric that carries consequences will be optimized against. What a health center should measure so that gaming the rubric means doing the job.

8 min read

Any QA rubric that carries consequences will be optimized against. That is not a statement about the character of front-office staff, it is what happens to every measured target in every organization, and health centers are unusually well placed to know it because so much of the work already runs against reported measures. The design question is therefore not how to stop people optimizing. It is how to build a rubric where optimizing for the score means doing the job properly.

Most call rubrics fail that test in a specific way. They score things that are easy to observe and loosely connected to the outcome: whether a greeting was used, whether the call stayed under a target length, whether the caller was thanked. Each is checkable, which is why it ends up in the rubric, and each can be satisfied on a call that helped nobody. A rubric built from those items produces rising scores, a front office that has learned the pattern, and no change at all in whether patients got what they called for.

Score the outcome, and the steps that cause it

The rubric items that resist gaming are the ones tied to something that either happened or did not, downstream of the call.

Did the caller who wanted an appointment leave with one on the calendar. Was eligibility captured or explicitly flagged for follow-up. Was sliding fee scale eligibility raised where it applied. Was the interpreter need identified and recorded. Was the callback that was promised actually placed. These are not judgments about how the call was handled. They are facts about what exists in the record afterwards.

The useful property is that you cannot satisfy them without doing the work. A staff member optimizing for appointment-booked rate books more appointments. A staff member optimizing for greeting compliance produces nothing.

Call length is the classic item to leave out. Scoring against a target handle time reliably produces shorter calls and worse outcomes, because the fastest way to end a call is to not resolve it. If length is worth watching at all, watch it as context next to resolution rate rather than as a scored item, and expect the good performers to sit above the average on some call types.

Health centers score against a wider job

A community health center front office is doing several things at once that a specialty practice front desk is not, and a rubric borrowed from elsewhere will miss all of them.

Coverage conversations are the bulk of it. When MGMA asked practices which phone tasks consumed the most staff time, eligibility and prior authorization led at 45%, with scheduling second at 31%. At a health center that first number understates the work, because the coverage conversation frequently is not about a plan at all. It is about whether the patient qualifies for the sliding fee scale, what documentation that takes, and what they will owe in the meantime.

That conversation belongs in the rubric explicitly. Was sliding fee eligibility raised when the patient had no coverage on file. Was the documentation requirement stated clearly. Was the patient told what happens if they cannot bring it. Those items are checkable from the transcript and they map directly to whether the patient shows up.

Language access belongs there too, as a recorded fact rather than an impression. Was a language preference already on file applied without the patient having to ask for it. Health centers generally know this matters and frequently do not measure it call by call, which means nobody can say how often it is missed.

Keep the scored set separate from the reported set

Health centers live with grant and program reporting, and the instinct is to make the QA rubric feed it. That instinct is worth resisting for a while.

The moment a call score becomes an input to something reported externally, its purpose changes. People manage the number rather than the process, ambiguous items get resolved in the flattering direction, and the practice loses the honest internal signal that made scoring worth doing. This is the ordinary behavior of measured targets and it does not require anyone to act in bad faith.

Keep two layers. An internal scored set, used to find where the process leaks, reported at the process level and not tied to individual consequences at the start. And whatever the external reporting requires, drawn from the systems of record, which is where it should come from anyway.

They can inform each other. If internal scoring shows eligibility capture failing on a third of new-patient calls, that is a plausible explanation for a downstream number that looks wrong, and it points at a fix. That is a much better relationship between the two than making one a feeder for the other.

Write it with the people it scores

A rubric handed down finished is a rubric that will be worked around, and it will also be wrong in two or three places that were obvious to everyone except the person who wrote it.

Run the first month as measurement only, at the process level, and say clearly that no individual consequence attaches yet. Then sit with the front office and go through the items that produced surprising numbers. Some will be genuine findings. Others will be a rubric item that does not match how the work actually happens, and the staff will identify those immediately.

That conversation is where the item about escalating to a nurse gets clarified, where somebody points out that eligibility cannot be captured for a walk-in scheduled by a partner agency, and where an item everyone agreed to on paper turns out to be unmeasurable in practice.

The rubric that comes out the other side is shorter, better, and considerably harder to game, mostly because the people who would have gamed it helped decide what it measures.

Report rates, review weekly, change one thing

The output should be a small number of process rates, not a leaderboard.

Appointment-booked rate on calls where the caller wanted one, by call type. Eligibility or sliding fee captured on new-patient calls. Language preference applied where one was on file. Callback commitments met. Escalations connected to a person, with the time on each.

Review them weekly, on a fixed day, and end the review with one change and one owner. Not a list. A rubric review that generates six action items generates zero completed ones, and after two months of that the front office correctly concludes the scoring does not lead anywhere.

One more habit worth building in early: revisit the rubric itself quarterly. Items that have sat at 99 percent for two quarters have stopped carrying information and can come out. Items that are always contested need rewriting or removing. A rubric that never changes is one nobody is really using.

Key Takeaways

  • Score outcomes and the steps that cause them. Appointment booked, eligibility captured, callback placed, interpreter need recorded. You cannot satisfy those without doing the work.
  • Leave call length out of the scored set. Scoring against handle time produces shorter calls and worse resolution, because the fastest way to end a call is to not resolve it.
  • Put sliding fee eligibility and language access in the rubric explicitly. A rubric borrowed from a specialty practice will not contain either, and both decide whether the patient shows up.
  • Keep the internal scored set separate from external reporting. The moment a score feeds a reported number, people manage the number instead of the process.
  • Run the first month as measurement only, then rewrite the ambiguous items with the front office. They will find the broken ones faster than any reviewer.
  • Report a handful of process rates weekly and end each review with exactly one change and one owner. Revisit the rubric quarterly and retire items stuck at 99 percent.

Every scored measure gets optimized against eventually, so the only durable protection is to score things that cannot be satisfied without doing the work. For a health center front office that means appointments actually on the calendar, eligibility and sliding fee conversations actually had, interpreter needs actually recorded, and callbacks actually placed. A rubric built from those items will still be gamed, and that is fine, because gaming it and doing the job well turn out to be the same activity.

Sources

Ready to See It in Action?

See how PGA scores health center calls against criteria that hold up

Schedule a Demo →

Written by Kevin Henrikson