Five commercial speech recognisers misread 35% of words from Black speakers and 19% from white speakers, in the same study.

That gap is a measured finding, not a suspicion, and any business putting a voice agent on a public phone number should know it exists.

Speech recognition does not work equally well for everyone, and there is peer-reviewed evidence of how unequally. Koenecke and colleagues ran audio from sociolinguistic interviews through five leading commercial systems and found an average word error rate of 0.35 for Black speakers against 0.19 for white speakers, roughly a two-fold gap. Errors were highest for African American men and rose with heavier use of African American Vernacular English features. The authors attributed the gap primarily to the acoustic model, and to training data that under-represented those speakers.

What exactly did the study find?

A consistent disparity across every system tested, not a quirk of one vendor. All five exhibited substantial gaps, which points at a shared cause in how the models were trained rather than at any single company's engineering.

The study is from 2020 and recognition has improved considerably since, so the specific numbers should not be quoted as today's performance. What has not changed is the mechanism: a model's accuracy on a voice depends on how much of that kind of voice it was trained on, and training corpora are not evenly distributed across the people who will call a business.

  • Average word error rate: 0.35 for Black speakers, 0.19 for white speakers.
  • All five commercial systems showed the disparity: Amazon, Apple, Google, IBM, Microsoft.
  • Error rates were highest for African American men.
  • The gap widened with denser use of African American Vernacular English features.
  • The authors located the problem in the acoustic model, tied to insufficient training data.

Why does this matter for a voice agent on a business phone?

Because a business phone does not get to choose its callers. A recognition gap on a consumer product is an annoyance the user can route around. The same gap on the only number a business publishes means some callers systematically get a worse service, and the business will not see it happening.

That invisibility is the operational problem. A caller whose name is misheard three times does not file a report; they hang up and call a competitor. The business sees a shorter call and no booking, which looks like a caller who changed their mind. Nothing in the metrics distinguishes that from a recognition failure.

Any large North American city makes this concrete rather than theoretical. A line taking calls in Calgary, Toronto, Houston or Los Angeles is taking calls from a genuinely multilingual population, and a system tuned on general North American English will not be equally good at all of it. That is a design input, not a footnote.

What can a line actually do about it?

Design for recovery rather than for accuracy. A system that confirms critical fields, offers a human early, and never requires a caller to repeat themselves more than twice degrades gracefully for the callers it handles worst, which is the only control a business genuinely has.

  • Cap the retries. Two failed attempts at the same field should route to a person, not a third attempt. The third attempt is where callers hang up.
  • Confirm rather than re-ask. Reading back what was heard lets a caller correct one field instead of repeating a whole sentence.
  • Make the human path obvious. A caller who is not being understood should not have to guess the magic word.
  • Watch short calls without outcomes. They are the signature of this failure, and the only place it shows up in the data.

A recognition gap on a consumer gadget is an annoyance the user routes around; the same gap on the only number a business publishes means some callers systematically get worse service, and it shows up in the data as people who changed their mind.

Nick Lovett, Founder, AnswerAI

Retry caps and an obvious route to a person are design decisions, and they are the ones that decide how a line treats the callers it understands least well.

Start free pilot

What this page does not claim

The cited study is from 2020 and tested systems available then. It is presented as evidence of a structural problem, not as a current benchmark.

  1. Not current performance. Recognition has improved substantially since 2020. The 0.35 and 0.19 figures describe that study's systems at that time.
  2. No figures for our own lines. We do not measure or publish recognition accuracy broken down by caller population, so this page reports the research rather than claiming to have replicated it.
  3. Accent is not the only variable. Audio quality, background noise and connection all interact with speaker characteristics. The study controlled for its own conditions, not for a phone line's.

Questions this raises

Is speech recognition worse for some accents?
Peer-reviewed evidence says yes. Koenecke et al. (PNAS, 2020) measured an average word error rate of 0.35 for Black speakers against 0.19 for white speakers across five commercial systems.
Was the problem specific to one vendor?
No. All five systems tested (Amazon, Apple, Google, IBM and Microsoft) showed substantial disparities, which points to a shared cause in training data rather than one company's implementation.
Has this been fixed since 2020?
Recognition has improved considerably, and the specific numbers should not be treated as current. The underlying mechanism, accuracy depending on training data representation, has not gone away.
How would a business notice this happening on its own line?
Mostly it would not, which is the problem. The signature is short calls that end without an outcome, which is indistinguishable from a caller who changed their mind unless someone goes looking.
What is the single most useful mitigation?
Capping retries and routing to a human after two failures on the same field. It does not fix recognition, but it stops the failure mode that makes callers give up.

Sources

  1. Racial disparities in automated speech recognitionKoenecke et al., PNAS 117 (14)2020, consulted 2026-08-09
Nick Lovett

Nick Lovett

Founder, AnswerAI

Nick Lovett builds AI receptionists for service businesses across North America, and writes these from the call data they produce. Lovett Ventures Inc., Calgary.

Want an AI receptionist built around these rules rather than despite them?

We build a working receptionist on your calls, your booking rules and your escalation rules, and you call it yourself before committing to anything. About 14 days, no call limit during the pilot.

Start free pilot