The caller is standing next to a running excavator. That is not an edge case, that is the trades.
Noise is not evenly distributed across industries, and the industries with the worst audio are often the ones with the most valuable calls.
A phone call is a narrow, lossy channel before any noise arrives. Traditional telephony carries roughly 300Hz to 3.4kHz (enough for intelligible speech and not much more) and the codecs that compress it are tuned to preserve what a human ear needs, not what a recogniser needs. Add an excavator, a hairdryer or a dining room, and the signal competes with noise inside an already narrow band. This is why a caller who sounds perfectly clear to a person can be difficult for a system.
Why does noise hurt more on a phone?
Because the useful signal has already been reduced. The consonants that distinguish similar words live largely in the higher frequencies, and those are exactly what a narrowband channel attenuates. Noise then masks what is left, and there is no headroom to recover from.
Human listeners compensate for this constantly and invisibly. We reconstruct missing consonants from context, from expectation, and from knowing the person. A recogniser does some of this (that is what the language model is doing) but it has less to work with and it fails differently, producing a confident wrong word where a human would ask.
The other factor is how the phone is being held. Speakerphone in a truck, a handset under a chin, a headset in a salon: each changes the distance and angle between mouth and microphone, and far-field audio in a reverberant room is a substantially harder problem than a handset held properly.
Which businesses have the hardest audio?
The ones where the caller or the callee is in the environment the business operates in. Trades calling from sites, salons with dryers running, restaurants at service, and vet clinics with animals are all systematically noisier than a law office.
| Context | What is in the background | What it costs |
|---|---|---|
| Trades on site | Equipment, traffic, wind, speakerphone in a vehicle | Addresses and callback numbers, exactly the fields that must be right |
| Salons | Dryers, music, close conversation | Names and service names, often said quickly |
| Restaurants at service | Dining room, kitchen, poor handset placement | Party size and times, misheard as similar numbers |
| Veterinary | Distressed animals, echoing rooms | Urgency cues in a call where urgency is the decision |
The pattern is unkind: the noisiest contexts are frequently the ones where getting the detail wrong costs the most.
What does a line do about it?
Assume the audio will sometimes be bad and make the important fields survivable. Confirm the callback number every time, keep critical questions short and closed, and route to a person after two failures rather than grinding through a third.
There is a temptation to solve this with noise suppression alone. It helps, and aggressive suppression can also remove parts of the speech signal along with the noise, which trades one failure for a subtler one. Design that tolerates imperfect audio is more robust than processing that assumes it can be cleaned up.
- Confirm the callback number on every call; it is what makes every other error recoverable.
- Prefer short closed questions in noisy contexts. "Morning or afternoon" survives noise that an open question does not.
- Cap retries. A third attempt in a noisy environment is where callers hang up.
- Do not ask a caller to move somewhere quieter unless the alternative is failing.
A phone call throws away most of the frequency range before any noise arrives, which is why a caller who sounds perfectly clear to a person can be genuinely hard for a recogniser standing on the same job site.
For a trades line we design the questions around bad audio from the start, because the site is where the valuable calls come from.
Start free pilotWhat this page does not claim
This describes a mechanism. No measurements of our own lines are published here.
- No accuracy figures. No recognition-accuracy figures by noise condition appear here. A number given without the exact audio conditions it was measured under means nothing.
- No vendor comparison. Noise robustness varies between providers and changes with model releases. Not tested here.
Questions this raises
- Why does a caller who sounds clear to me confuse the AI?
- You are reconstructing missing sound from context and familiarity. The phone channel discards much of the higher frequency range where distinguishing consonants live, and noise masks what remains.
- Does noise cancellation fix it?
- Partly. Aggressive suppression can also remove parts of the speech signal, trading an obvious failure for a subtler one. Designing questions that survive imperfect audio is more reliable.
- Which industries have the hardest phone audio?
- Trades calling from sites, salons with equipment running, restaurants during service, and veterinary clinics. Unhelpfully, several of these are also where a misheard detail costs the most.
- Is speakerphone worse?
- Usually. It increases the distance between mouth and microphone and picks up more of the room, which is a substantially harder problem than a handset held normally.
- What is the single most important field to protect?
- The callback number. If it is right, every other error on the call can be fixed later. If it is wrong, none of them can.
Sources
- Racial disparities in automated speech recognitionKoenecke et al., PNAS 117 (14)2020, consulted 2026-08-09
Want an AI receptionist built around these rules rather than despite them?
We build a working receptionist on your calls, your booking rules and your escalation rules, and you call it yourself before committing to anything. About 14 days, no call limit during the pilot.
Start free pilot
