ROI Analysis
Reviewing Thousands of Automated Calls Without Listening to Them All
Nobody can listen to a month of automated calls across twelve sites. How an MSO reviews by exception, using what athenaOne already records about each one.
Once automated calls run across a dozen practices, the review problem inverts. The old worry was that nobody was listening to enough calls. The new one is that there are forty thousand of them a month and any listening plan a person can actually sustain covers a rounding error of the total.
Most MSOs respond by sampling randomly. Pick ten a week, listen, score them, file the scores. It feels rigorous and it is close to useless, because a random sample of a population that is mostly fine returns calls that are mostly fine.
The calls worth an operations director’s time are rare by definition. They are the ones where something unusual happened, and random sampling is the least efficient way to find rare things.
The alternative is to stop treating review as listening and start treating it as exception handling. Let the system tell you which interactions look different, and spend the review hour on those.
Review the record the call produced, not the audio
Reviewing is worth the hour because the results across the sector are genuinely uneven. An August 5, 2025, MGMA Stat poll of 244 applicable responses found 71% of practice leaders reported some use of AI in patient visits, but among those using it, 44% said it had not reduced staff workload, 39% said it had, and 17% were unsure.
That spread is what a review program exists to explain. Something separates the groups where automation removed hours from the groups where it did not, and it is visible in the work, not in the sales material.
Every automated interaction that matters leaves something behind in athenaOne. It might be a patient case, an appointment, a document, or a task in somebody’s queue. That artifact is far more reviewable than the audio, because it is structured and it can be queried in bulk.
The practical shift is to review outcomes first and pull audio only when an outcome looks wrong. A case closed with a reason that does not fit the request. A case reopened within a day. A case that changed hands three times. A booking made and cancelled within an hour.
All four of those are queries rather than listening sessions. They run across every site at once and they take a manager to the twenty interactions actually worth opening.
On athenaOne the changed-case feed is what makes this practical at volume. Instead of pulling every case at every practice on a schedule, an oversight process subscribes to what changed and works only that, which is the difference between an overnight job and a real-time view.
Close reasons are the highest-value field nobody configures
The reason a case was closed is the cheapest quality signal an MSO has and it is almost always left at defaults.
A short, deliberate list of close reasons turns a pile of finished work into a distribution you can read. Booked as requested. Handed to staff for a decision. Patient declined. Could not reach. Wrong department, rerouted. Each of those means something different about the front office.
When one site shows a rising share of handed to staff and its neighbours do not, that is a configuration problem at that site rather than a vendor problem. When could not reach spikes in a single department, that is usually a phone number or an hours problem.
The automation writes those reasons consistently, which is the part humans never manage across twelve locations. What it does not do is decide what the pattern means. A distribution that has shifted is a question for the operations director, not an answer.
Sites diverge, and the divergence is the finding
The most useful review an MSO can run is not vertical, it is horizontal. Same request type, same measure, twelve columns.
The reason this works is that the sites are supposed to be doing the same thing and are not. The provider roster on a practice website does not match the roster in the system. One practice retired an entire appointment type catalog overnight and folded it into a single short follow-up type. A new provider’s intake cap was set two months ago and nobody removed it.
None of those show up in a call recording. All of them show up as one column behaving differently from eleven others.
So the standing review is a comparison table, not a listening list. Completion by request type per site, handed-to-staff share per site, oldest open item per queue per site. Then audio, and only for the outliers.
What the evidence says about expecting headcount changes
Be careful about what this program is sold as internally. A June 2, 2026, MGMA Stat poll found that most practice leaders, 68%, say their organizations have not redesigned a role or adjusted staffing with the help of AI in the past year. Only about one in four, 26%, say they have, and another 5% were unsure, across 260 applicable responses.
Most groups added capability and kept their people. For an MSO that is the right outcome anyway, because the constraint was never that the sites were overstaffed.
What changes is what the staff spend the day on. The person at the desk stops being interrupted forty times an hour and starts finishing the work only they can do.
So report the review program that way. Not seats removed. Open items aged down, exceptions caught before the patient called back, and sites brought into line with each other.
Keep the review administrative
An exception review across thousands of interactions is a quality process for front-office work, and it has a boundary worth stating in the program document itself.
What gets scored is whether the request was captured correctly, routed to the right queue, completed in the right record, and handed off with enough context. Whether the automation stayed inside its scope. Whether the patient got a straight answer about logistics.
What does not get scored is anything about the care the patient needs. A reviewer is not grading the automation on how it handled a description of a problem, because it should not have been handling that at all. It should have routed it to staff and stopped.
That is also the review’s most important check. Sample specifically for interactions where a patient raised something clinical and confirm that every one of them was handed to a person promptly and with the details attached.
What a review hour should actually look like
Twenty minutes on the comparison table across sites. Twenty minutes on the exception queue, meaning cases reopened, reassigned repeatedly, or closed with a reason that does not match the request. Twenty minutes on a deliberate sample: every interaction that got handed to staff at one site, read end to end.
That is a repeatable weekly hour and it beats any listening plan, because it is aimed at the things that vary rather than at the things that do not.
The practices worth spending it on will be obvious within two weeks. In most groups two or three sites generate the majority of the exceptions, and they are rarely the sites people expected.
None of this works if the automation cannot write structured outcomes into the systems the group already runs. PGA works across 440+ of athenahealth’s roughly 800 endpoints, which is why the review can be a query against real records rather than a pile of recordings and a spreadsheet.
Key Takeaways
- Review the record each interaction produced rather than the audio, because structured outcomes can be queried across every site at once.
- Configure a short, deliberate list of close reasons, since the default list turns finished work into a pile instead of a readable distribution.
- Subscribe to what changed rather than polling every case on a schedule, which is what makes real-time oversight practical at volume.
- Run the standing review horizontally, comparing the same measure across all sites, because divergence between locations is usually the finding.
- Sample deliberately for interactions handed to staff and confirm each one carried enough context to act on without a callback.
- Report the program as aged-down queues and exceptions caught early, not as headcount, which the evidence does not support anyway.
Nobody is going to listen to forty thousand calls, and nobody needs to. Query the outcomes, configure close reasons that mean something, compare sites against each other, and spend the listening hour only on the interactions that already look wrong.
Related reading
- the manager call review workflow, a week in twenty minutes
- closing a patient case: reasons, rules, and what gets measured
- subscribing to changes instead of polling for them
Sources
Ready to See It in Action?
See how PGA surfaces the calls worth reviewing across every site, using what landed in athenaOne rather than a recording pile
Schedule a Demo →Written by Kevin Henrikson