Skip to main content

ROI Analysis

The ROI Case for Internal Medicine Front Office KPIs

Internal medicine front office KPIs that show whether automation worked: third next available by type, aged orders, repeat status calls, refill first-contact.

9 min read

Ask an internal medicine practice how the last automation project went and you get a shrug. Something got bought. A dashboard exists. Nobody can say whether the chronic care panel is being worked better than it was a year ago. That gap is the whole problem with internal medicine front office KPIs: the numbers on the vendor dashboard and the numbers that decide whether your practice makes money are rarely the same ones.

Most practices are three or four tools into this by now. Each one demoed well. Each one solved a slice and handed the awkward remainder back to the same two people. Staff ended up doing the original work plus checking on the software, and the reporting that came with it counted the software’s activity rather than the practice’s.

Internal medicine feels this harder than most specialties. The panel is large, chronic, and old, which means the work is not a stream of discrete visits. It is refill volume, results callbacks to schedule, referrals going out that need chasing, and wellness visits competing with sick visits for the same slot. None of it produces a clean metric on its own.

So the honest question is not whether the tool is working. It is which four or five numbers would tell you, and whether anything in your practice currently produces them.

Most front office dashboards measure the vendor, not the practice

Calls handled. Containment rate. Minutes saved. Those are vendor metrics. They go up when the software is busy, which is not the same as your practice being better off, and they are unfalsifiable in the way that matters: no number on that list can ever come back and say this did not work.

The field data says the skepticism is earned. An August 2025 MGMA Stat poll found 71% of practice leaders reported some use of AI in patient visits, but among those using it, 44% said it had not reduced staff workload, 39% said it had, and 17% were unsure. A tool can be in use, on a dashboard, and invisible in the actual workload at the same time.

The fix is not a better dashboard. It is picking metrics that already exist in your practice, that you were tracking or should have been before anyone sold you anything, and holding the automation to those. If the number a vendor reports cannot be reproduced from your own athenahealth data, it is marketing.

Four numbers an internal medicine practice can act on

Start with third next available appointment by appointment type and provider, not overall. An average across all types hides the thing you need to see, which is that a sick visit is bookable Tuesday and a wellness visit is eleven weeks out.

Second, open orders and follow-up tasks aged past fourteen days. In a chronic panel this is the closest thing to a backlog dollar figure you have. Every aged order is a visit that was clinically ordered, never scheduled, and never billed.

Third, repeat status calls per outbound referral. Internal medicine sends a lot of referrals out, and each one generates calls asking whether it went, whether it was received, and when the patient will hear something. One call per referral is the workflow. Three is a reporting failure.

Fourth, first contact resolution on refill request intake. MGMA Stat polling on phone time puts prescription refills at 6% of the most time consuming call work, well behind eligibility and prior authorization at 45% and scheduling at 31%. Refills are not the biggest line, but they are the most repetitive one, which makes them the cleanest read on whether intake is capturing enough on the first call.

All four come out of athenahealth data you already own: appointment types and templates, the order and tickler queue, referral orders, and the department buckets requests land in.

Your utilization number lies because of the generic slot

Here is the complication that breaks most measurement projects before they start. Practices run generic template slots, the Any 15 and Any 30 varieties, and the schedule returns one of those when someone searches for a specific appointment type. An Any 15 might legitimately need to be an hour for a new patient with a complicated history. Which specific types a generic slot is actually eligible for, per provider and per department, is the entire mapping problem.

Until that mapping is written down, utilization is a number about slots rather than a number about work. It looks healthy while a provider caps new patients at four a day and a thousand established patients sit in the follow-up queue that nobody has hours to work.

This is where an automation partner earns the term partner rather than vendor. Writing the generic slot mapping is a project: appointment type by appointment type, provider by provider, department by department, with the practice deciding the rules and the automation enforcing them on every call and every outbound campaign. Nobody sells that as a SKU. It is the work.

The measurement payoff arrives after the mapping, not before. Once a generic slot resolves to a known type, fill rate by type becomes real, and so does every number built on top of it.

Wellness visits versus sick visits is a measurement problem first

Medicare covers an annual wellness visit once every 12 months per patient, and the eligibility window is a date, not an opinion. That makes it one of the few pieces of internal medicine volume that is fully countable in advance. You know exactly who is eligible, exactly when, and exactly how many of them you booked.

Almost nobody measures it that way. The common metric is wellness visits completed, which rises and falls with how much room the schedule happened to have. The metric worth running is completion against the eligible panel, by month, with the misses named. That number tells you whether recall outreach is working or whether wellness visits are simply losing the slot fight to same day sick demand every week.

Automated outreach fits here cleanly because the work is calendar arithmetic and phone calls. Pull the eligible list, confirm the window has opened, call the patient, book the visit against the correct athenahealth appointment type, and report back who was reached and who was not.

What the automation does not do is any part of the visit. It books the appointment and closes the loop on outreach. Everything inside the room stays with the practice, and any caller who raises something a licensed clinician needs to hear gets routed to staff.

Per call QA is the measurement nobody actually runs

Most practices sample a handful of calls a month, if that, and the sample gets skipped the first week anyone is out sick. The result is that call quality is managed by complaint. You hear about the calls that went badly enough for a patient to say something.

Scoring every call changes what the number means. The criteria are administrative and checkable: was the callback number captured, was the request filed to the correct department bucket, was the patient given a specific next step and date, was the follow-up task opened. Those are pass or fail, they apply to human and automated calls alike, and they produce a trend rather than an anecdote.

Flagged calls go to a person. Low confidence handling, a caller who repeated themselves three times, anything where the request did not map to a known path. A supervisor listens to those and nothing else, which is the only version of QA that survives a busy month.

The organizational side is slower than the technical side, and the data shows it. A June 2026 MGMA Stat poll found 68% of practice leaders said their organization had not redesigned a role or adjusted staffing with the help of AI in the past year, while 26% had. Measurement is what moves a practice from the first group to the second, because you cannot redesign a job you have never scored.

Key takeaways

  • Reject any metric you cannot reproduce from your own athenahealth data. Calls handled and containment rate measure the software, not the practice.
  • Track third next available by appointment type and provider, never as a single blended average, because the blend hides the wellness visit backlog.
  • Treat open orders and follow-up tasks aged past fourteen days as your backlog number. Each one is ordered work that was never scheduled and never billed.
  • Fix the generic Any 15 and Any 30 slot mapping before you trust utilization. Until a generic slot resolves to a known type, fill rate is a number about slots.
  • Measure annual wellness visits against the eligible panel by month rather than raw completions, and name the misses.
  • Score every call on administrative criteria, callback number captured, correct bucket, specific next step, task opened, and route only flagged calls to a supervisor.
  • Baseline all of it before deploying anything. A number you started collecting after go live cannot prove go live changed it.

None of these numbers require new software to define. They require somebody to decide which five matter, write down how each is calculated, and pull the baseline before the next tool arrives. That is a two week project and it is the highest return work an internal medicine practice can do this quarter.

What automation adds afterward is the ability to produce them continuously, on every call and every outbound campaign, out of the athenahealth data already sitting there. Measurement first, then the tool. Doing it in the other order is how practices end up four vendors deep with nothing to show a board.

Related reading: scheduling automation for independent primary care, the hidden cost of manual scheduling, and primary care patient communication.

Sources:

Ready to See It in Action?

See the front office scorecard Pretty Good AI reports back from your athenahealth data, call by call.

Schedule a Demo →

Written by Kevin Henrikson