When Speech Recognition Becomes a Clinical Safety Risk
- 6 hours ago
- 5 min read

AI-powered speech recognition is increasingly being used across healthcare in the form ambient voice technology, clinical documentation to triage, call handling and patient-facing tools.
Our customers are learning that converting spoken language into structured or written information can save considerable time and reduce administrative burden. But from a clinical safety perspective, we support our customers to ask
Who does it work well for, under what conditions, and what happens when it gets something wrong?
A transcription error can become a clinical hazard
Speech recognition systems do not necessarily perform equally across every accent, dialect or speech pattern.
Performance may be affected by factors including:
regional or non-native accents;
speech impairments or atypical speech patterns;
age-related differences in speech;
background noise;
poor microphones;
telephone or compressed audio; and
whether particular speech groups were adequately represented in the data used to develop and test the model.
In many settings, an occasional transcription error may simply be inconvenient but in health or care settings, it can be harmful.
Imagine a system incorrectly transcribing a patient's symptoms, medication, dosage or description of urgency. If that information subsequently informs triage, clinical review or another decision, the original recognition error may propagate further into the care pathway.
The potential consequences include inappropriate treatment, missed or delayed diagnosis, delayed treatment and psychological distress.
That makes speech recognition performance a clinical safety issue, not simply a technical accuracy issue.
Fairness is also a legal and organisational risk
There is another important dimension: equality.
Under the Equality Act 2010, organisations have duties in relation to discrimination across both the provision of services and employment.
The Act can capture apparently neutral practices which, in practice, place people sharing a protected characteristic at a particular disadvantage, known as indirect discrimination.
That matters when an organisation introduces an AI tool into a healthcare pathway.
If a speech-recognition system consistently performs less well for particular groups, the issue may extend beyond model accuracy. Depending on the circumstances, poorer performance could contribute to unequal access to a service, poorer quality of care or other disadvantage.
Disability also carries additional considerations, including protection from discrimination arising from disability and duties around reasonable adjustments.
And we should not only think about patients.
The same technology is often used by employees - for example clinicians dictating notes, administrative staff using voice tools or staff interacting with AI-enabled systems as part of their work. Equality Act protections extend to employment and work-related activities as well as access to services.
So an organisation evaluating an AI system should be asking two related questions:
Does this system work safely and equitably for the people receiving our services?
And:
Does it work equitably for the people we expect to use it as part of their job?
A system that routinely struggles with a clinician's accent, for example, may create additional workload, require repeated correction and potentially affect their ability to use a tool which colleagues can use effectively. Likewise, a voice system that performs poorly for a patient with a speech impairment may create both a clinical safety risk and an equality concern.
This is one reason that fairness testing should not sit in an isolated “AI ethics” exercise. It can form part of the evidence supporting clinical safety, equality compliance and organisational risk management.
“The supplier has tested it” isn't quite enough
One of the challenges for organisations implementing AI is working out what evidence they should actually ask suppliers for as we know that an accuracy figure is not enough and could relate to limited testing or testing that was not in real-world, messy conditions.
For an AI system containing speech recognition, clinical safety assurance might reasonably explore:
What performance metric is being used?
For speech-to-text systems this might include Word Error Rate (WER), but the metric should be meaningful in the context in which the system will actually be used.
Where did the test evidence come from?
Was testing performed by the supplier, an independent organisation, or against a recognised benchmark? What population and audio conditions were represented?
Were important user groups considered separately?
An overall average can conceal materially poorer performance for particular accents, languages, demographic groups or speech patterns.
Were both service users and employees considered?
The populations interacting with the system may be very different. Supplier testing focused on patient speech may tell you very little about how the system performs for a diverse clinical or administrative workforce — and vice versa.
Was testing representative of the intended healthcare environment?
A model tested using clear studio-quality audio may behave differently when presented with a hurried patient on a noisy telephone line or a clinician dictating in a busy consultation room.
These questions move assurance away from simply recording that “the AI was tested” and towards understanding whether the evidence actually supports its intended use.
Fairness testing belongs in clinical safety too
Two useful approaches are group testing and counterfactual testing.
Group testing compares system performance across relevant populations — for example, whether recognition accuracy differs materially between different patient groups.
Counterfactual testing asks whether identical inputs produce different results when a relevant characteristic is changed (such as gender). This can be particularly useful where controlled testing is possible.
The objective does not necessarily have to be mathematically identical performance across every population - the important thing is identifying meaningful differences in performance, understanding whether those differences could contribute to clinical harm or unequal treatment, and deciding whether further controls are required.
Supplier evidence is only half the picture
Even strong supplier testing does not necessarily demonstrate that the system will perform safely or fairly in every organisation, because local conditions matter.
The population served by one GP practice may be very different from the population represented in the supplier's original testing. The workforce may also have a different demographic or linguistic profile. Audio equipment, workflows, clinical pathways and the way staff use the output may all differ.
Clinical safety work therefore needs to connect supplier assurance with local assurance.
A supplier might demonstrate that its speech recognition model has been tested across a range of accents and achieved an acceptable performance threshold.
The deploying organisation can then consider whether those populations and conditions reflect its own patients and workforce, and whether additional local testing or monitoring is required
This is exactly the kind of problem LENS is designed to surface
One of the reasons we developed LENS — Lawful, Explainable, Necessary and Safe — was that AI assurance information is often fragmented.
LENS helps organisations identify hazards arising from the specific AI capabilities within a system and connects those hazards to practical assurance tasks.
For a speech-recognition hazard such as the one described above, LENS can prompt teams to:
identify the relevant performance metrics for the AI models in the pipeline;
establish the supplier's testing evidence;
undertake or review AI system behaviour and performance testing;
undertake appropriate group-based or counterfactual fairness testing;
consider performance across both service users and employees;
assess whether supplier testing reflects the local population, workforce and operating environment; and
record the resulting evidence alongside the clinical safety assessment.
This means the clinical safety record demonstrate what evidence was examined, which groups were considered, what testing was undertaken, what the results showed and why the residual risk was considered acceptable.
Are you an existing myKafico customer? You can ask to have your Clinical Safety Officer added to LENS today





Comments