Phone it and ask it something it should refuse. That one test outperforms every feature list.
A buyer's checklist, including the questions that are uncomfortable for us to answer.
Feature lists in this category are close to useless, because every vendor lists the same features and none of them describe judgement. The evaluation that works is a four-minute phone call in which you deliberately try to make the receptionist do something it should not, followed by a short list of questions whose answers are hard to fake. Both are below, including the ones we would rather you did not ask.
The four-minute phone test
Call the number and try four things in order. Each takes under a minute and each probes a layer the demo did not show you.
Do this from a mobile, outdoors, with some background noise, and give a name that is not common in your region. Those three conditions describe a large share of real calls and none of a demo's.
- 1. Ask something it should refuse. A clinical, legal or coverage question in your industry. Listen for whether the refusal names what happens next, or is just an apology.
- 2. Describe something urgent, oddly. Not the textbook phrasing. Real emergencies are described in unusual words, and that is exactly where model judgement fails and a deterministic rule holds.
- 3. Interrupt it mid-sentence. Then change the subject. See whether it follows you or finishes the sentence it had planned.
- 4. Ask for a human, twice. Whether there is a path, and what happens when nobody picks up, is a decision someone made or failed to make.
The vendor questions that are hard to fake
Eleven, grouped by what they reveal. Vagueness on the second group is the strongest negative signal available to a buyer.
| Ask | A weak answer sounds like |
|---|---|
| What will it refuse to answer, specifically, for my business? | "It's trained to be helpful and safe." |
| What triggers an immediate escalation, and to whom? | "It detects urgency automatically." |
| What happens if the escalation is not answered? | Silence, or a first-time consideration of the question. |
| What constrains a booking besides the calendar? | "It syncs with your calendar." |
| What can it write into my systems, versus propose? | "Full integration." |
| What is retained (audio, transcript, fields) and for how long? | One number for all three, or "the standard retention." |
| Can I change retention per line? | "That's a platform setting." |
| Can you delete one caller's data on request? | Anything other than a described process. |
| What does it do when it is not sure? | "It's very accurate." |
| What happens on a caller you cannot understand? | "It asks them to repeat." (How many times?) |
| Who wrote the refusal rules, and can I see them? | "They're built in." |
The pattern in every weak answer is a capability where a decision was asked for.
Where we would score badly on this test
Three places, named because a buyer's checklist written by a vendor is worthless unless it cuts both ways. If these matter more to you than the judgement layer, we are the wrong choice.
There is a fourth, softer one: a business whose calls genuinely are all one shape does not need what we build, and would be paying for judgement it does not have to exercise.
- Speed to launch. A configurable product can be live this afternoon. We take about 14 days, because the decisions are the work.
- Data residency. We cannot currently promise call data stays in Canada. Our platform documents SOC 2 Type 1 and 2 and HIPAA under a BAA, retention configurable per agent, and PII excludable per agent, but residency is not something we will claim.
- Published pricing. AnswerAI publishes no price, which does make us harder to compare on a spreadsheet. If a number on a page before a conversation is what you need, that is a fair reason to look elsewhere.
Every weak vendor answer in this category has the same shape: a capability offered where a decision was asked for. "It detects urgency automatically" is not an answer to the question of what happens when somebody says the wrong thing at 2am.
Run the four-minute test on our line before you run it on anyone else's. It is the fastest way to know what we are claiming.
Start free pilotHow to use this against us
A buyer's guide written by a seller is a marketing document unless it names its own weak spots.
- We sell the thing this page argues for. AnswerAI builds custom lines rather than selling a configurable product, so this cluster describes our own commercial position. The arguments are made from mechanism and you should check them against a vendor who disagrees.
- The test is designed around what we are good at. It weights judgement heavily because that is what we build. A buyer optimising for speed or price should weight differently and will reach a different answer.
- No comparative testing. We have not run this test across competitors and published results. That would be a stronger page and it is not what this is.
Questions this raises
- What is the single best test of an AI receptionist?
- Phone it and ask something it should refuse. A built line has a specific refusal that names what happens next; a templated one apologises generically or answers.
- What should I ask about data handling?
- What is retained (audio, transcript and structured fields separately) for how long, whether that is configurable per line, and what the process is for deleting one caller's data.
- How do I test escalation without causing a real emergency?
- Describe a plausible urgent situation in unusual phrasing and see what happens. You are testing whether a rule fires, not whether a human responds.
- Should I trust a buyer's guide written by a vendor?
- Only as far as it names where it is weak. This one lists three places we would score badly. Apply the same standard to every other guide you read, including the ones that flatter us.
- Does a longer setup time mean a better line?
- Not inherently. It means the decisions were made deliberately rather than defaulted, which is worth something only if your calls actually vary.
Sources
- Recording of Customer Telephone Calls: guidance for organizationsOffice of the Privacy Commissioner of Canadaupdated 2018-04-18, consulted 2026-08-09
- PIPEDA fair information principlesOffice of the Privacy Commissioner of Canadaconsulted 2026-08-09
- Racial disparities in automated speech recognitionKoenecke et al., PNAS 117 (14)2020, consulted 2026-08-09
The only real test is calling one yourself.
We build a working line on your calls, your pricing, your booking rules and your escalation rules, then you phone it and tell us what does not sound like you. About 14 days. No call limit during the pilot.
Start free pilot
