Your AI receptionist answered 412 calls last month. How many did it handle badly? Most firms can’t answer that, which is a strange place to be with a system that’s the first voice a prospective client hears. The vendor dashboard shows calls answered and appointments booked. It doesn’t show the caller who hung up at forty seconds because the assistant kept asking her to repeat a street name.
Key Takeaways
- Vendor dashboards measure completion, not quality. Sample real transcripts weekly or you’re flying blind.
- Five to ten calls a week, scored against a short rubric, catches most problems within a month of them starting.
- The expensive failures are silent: wrong information delivered confidently, and escalations that never happened.
- Every failure should produce a knowledge base fix, a prompt change, or a new escalation trigger. Otherwise it repeats.
- Give your staff a one-click way to flag a bad call. They notice things a sampling process never will.
Why the dashboard isn’t enough
Vendor reporting answers operational questions. Did the call connect, how long did it run, was an appointment created. Useful, and not the same as whether the caller got what they needed.
Consider a call marked “completed, appointment booked.” Sounds like a win. Listen to it and you find the assistant booked a bankruptcy consultation for someone describing a landlord dispute, because it latched onto the word “eviction” and ran with it. That appointment is going to waste an attorney’s hour and annoy a person who needed a different kind of help. On the dashboard it’s a green checkmark.
The gap between “the system did something” and “the system did the right thing” is where quality assurance lives.
The failures worth hunting for
Confident wrong answers
The worst category by a distance. The assistant tells a caller you accept a particular insurance plan when you dropped it in January, or quotes a consultation fee from last year’s pricing. It sounds authoritative, the caller believes it, and you find out three weeks later during an uncomfortable conversation at the front desk. These almost always trace back to a stale knowledge base rather than to the model itself.
Missed escalations
Some calls should reach a person immediately. A patient describing chest pain. A caller with a filing deadline tomorrow. Someone who’s already said “I want to speak to a human” twice. When the assistant keeps trying to resolve those itself, you get a small operational problem and, occasionally, a serious one.
Data captured wrong
Phone numbers with a transposed digit, names spelled phonetically, callback times recorded in the wrong time zone. Quiet, cumulative, and expensive. A lead you can’t call back is a lead you paid for and lost.
Tone that doesn’t fit the moment
Harder to score, still real. A cheerful assistant handling a call about a family member’s death reads as awful. Most platforms let you shape tone by call type; few firms bother.
A review process that takes twenty minutes
You don’t need to listen to everything. Pull a sample each week and be deliberate about which calls you pull.
- Two or three random calls, so you see typical performance
- Every call under 45 seconds that didn’t book anything, since short calls usually mean an abandon
- Every call over five minutes, which usually means the assistant got stuck in a loop
- Any call staff flagged
Score each on five things, pass or fail, no scale: Did the assistant identify why the person called? Was every piece of information it gave accurate? Was contact data captured correctly? Should this call have been escalated, and was it? Would you be comfortable if the caller were your best referral source?
Binary scoring is the point. A 1-to-5 scale invites people to write “3” and move on. Pass or fail forces a decision, and the failures are what you’re looking for.
Turning a failure into a fix
Reviewing without changing anything is just a new form of paperwork. Every failed call should route to one of four actions, and it’s usually obvious which.
Wrong information means the knowledge base is out of date. Fix the source, not the transcript. Missed escalation means an escalation rule needs adding, and these should be explicit rather than left to the model’s judgment. Confusion about intent usually means the opening question is too open-ended, and tightening it helps more than any amount of prompt tuning. Repeated failures on the same scenario mean that scenario shouldn’t be automated at all, at least not yet.
That last one deserves emphasis. Not every call type should be handled by AI. If a category keeps failing after two rounds of fixes, route it straight to a person and stop fighting it. There’s no prize for full automation.
The knowledge base is the real system
Most AI receptionist problems aren’t AI problems. They’re documentation problems wearing a costume. The assistant can only work from what you’ve given it, and what most firms give it is a document written during onboarding and never touched again.
Make someone responsible for it. Then tie updates to the events that actually change things: a fee schedule change, a new attorney or provider joining, an insurance contract ending, a holiday closure, a new service line. Quarterly review of the whole document, with a line-by-line check of anything involving a price or a date. It’s a half-hour of work that prevents the most damaging failure category you have.
Let your staff report problems
Your front desk hears the aftermath of every bad call. “The robot told me you were open Saturday.” “I already gave all this information to the computer.” They know where the system fails, and in most firms nobody’s asked them.
Give them a channel that costs nothing to use: a Slack channel, a shared form, a sticky note on the monitor if that’s what works. Then close the loop by telling them what got fixed. Staff stop reporting when reports vanish into nothing, and that silence gets misread as everything working fine.
What good looks like
After a few months of steady review, firms we work with tend to land around 90% of sampled calls passing all five checks, escalation accuracy above 95% on the categories that matter, and contact data errors under 2%. Getting there takes a quarter, not a week.
The firms that get real value from these tools aren’t the ones with the best vendor. They’re the ones who treat the assistant like a new hire: watch the early calls closely, correct patterns instead of incidents, and keep checking after the novelty wears off.
