By Seemab AkbarAI Solutions

Our AI receptionist has answered our own business line since September 5, 2026. Four days later we started testing it properly: another AI phones it, pretending to be a customer, and every call is graded. This is how that testing works, what it caught, and why we think it is the part most AI receptionist offers leave out.
The problem with trying it a few times
The usual way to check an AI phone assistant is to call it a few times, hear it sound friendly, and switch it on. That tells you almost nothing. A receptionist fails on the fifth question, not the first: when the caller interrupts, asks for a day that is a holiday, gets angry, switches language or asks whether they are talking to a real person.
So we stopped testing by ear and built a way to test it the way a caller would, again and again, with a score.
A caller that is also an AI
We wrote 46 caller scenarios. Each one is a person with a goal and a way of talking: a customer rebooking a meeting, someone in a hurry, a wrong number, a job seeker, a podcast host, an angry client with an unpaid problem, a trucker switching between Punjabi and English. Twenty-five of them were written to be deliberately hard.
A second AI plays each caller. It dials our real number, speaks in a real voice, interrupts, changes its mind and hangs up when it is done. Our receptionist answers exactly as it would for a real caller, on the same line, with the same calendar and the same rules.
Between September 9 and September 26 we ran 257 of these practice calls.
Every call gets graded
When a practice call ends, a third AI reads the transcript and grades it against a written rubric: the greeting, how well it listened, whether every fact was right, whether every promise was real and kept, and how it closed the call. Each part scores 0 to 2, so a call scores out of 10.
A call fails if it scores under the bar, if it says something it must never say, or if it misses something it must say. Of 252 graded calls, 148 failed. That number is the point. We were looking for failures, and a test that never fails is not testing anything.
What the failures taught us
Every failure became a written rule, with the call that caused it recorded beside it. There are 101 of those rules now. A few of them:
- A holiday is not an open day. A caller asked about Monday and the receptionist cheerfully offered 9, 10 or 11. The times were real, pulled from the calendar, but Monday was a statutory holiday in Canada. Holidays are now part of what it knows before it offers a time.
- We do not have an office to invite you to. Early on it offered to meet a caller "at our office". We meet online, at the client's place, or somewhere close to them. It now says so, and suggests a real meeting spot near the caller.
- Say it is an AI. One version of its rules told it to deny being an AI if asked. We changed that. Every call now opens by saying it is Seemab's AI assistant and that the call is recorded. The practice call that checked it asked "am I talking to a real person?" and scored 10 out of 10.
- An emergency comes first. If a caller's words sound like an emergency, the line "please hang up and call 911" is spoken first, by code, before the AI says anything else.
When a rule fails twice, it becomes code
Some rules we wrote into the assistant's instructions kept failing. It greeted a caller twice. It kept talking after a caller said goodbye. When a written rule fails twice, we stop asking the AI to remember it and build it into the software instead, so it cannot happen. The greeting, the goodbye and the emergency line all work that way now.
Behind the practice calls sit more than 470 automated checks that run before every change to the receptionist goes live. A change that breaks any of them does not ship.
What this means if you want one
An AI receptionist is only as good as the testing behind it. Before you trust one with your phone, ask the people selling it three questions:
- How many practice calls did it take before it went live, and how were they graded?
- What happens when a caller asks whether they are talking to a real person?
- What does it do when it hears an emergency?
If the answers are vague, the testing was too.
Our receptionist comes inside NameCRM, the client system we run our own business on, and the full story of the receptionist on our line is in the case study. If you want to see where AI would pay off in your business first, the AI Readiness Check takes two minutes and shows your score straight away.
