Who's accountable when the AI gets it wrong?


Dave Hughes
Who's accountable for when AI goes wrong ?

Who’s accountable when the AI gets it wrong?

In September, our CTO Alastair Williamson-Pound wrote in Computer Weekly about the UK government’s AI opportunities action plan. His question: when an AI system hallucinates, shows bias, or suffers a security breach, who takes responsibility? Right now, he wrote, the answer is often “it depends.” That uncertainty is the biggest threat to the whole programme.

The same question sits underneath every conversational AI deployment we work on. Before a single line of dialogue gets built, someone needs to decide who owns the decision when the system triages a citizen, escalates a case, or gets it wrong. Human, system, or some hybrid of the two. This has to be settled before deployment, not worked out afterwards when something has already gone wrong.

Accountability that survives contact with procurement

Alastair’s argument in Computer Weekly starts earlier than most people expect: at the point of purchase. Procurement teams often commit to AI tools without knowing what data trained them, how the system reaches its decisions, or whether AI was the right answer to begin with. Suppliers treat training data and model architecture as trade secrets. Procurement staff, not trained to interrogate AI-specific risk, don’t ask the questions that would surface it.

His proposed fix is a GDPR-style model: liability follows control. A supplier selling a closed system with no visibility into its workings carries the risk. A buyer who takes a transparent, configurable tool and mishandles it, say by feeding it sensitive data it wasn’t built for, carries the risk instead. Every likely failure gets named in the contract, with responsibility assigned to whoever had control over that part of the system.

This is where an Independent Certified Implementation Partner (ICIP) such as Mercator Digital earn their place. Government departments rarely have the in-house expertise to translate an ambiguous policy requirement into a working system, and fewer still have the expertise to evaluate a supplier’s claims about bias, explainability, or data provenance before signing. Bringing a tested framework for that evaluation, rather than reinventing it per project, is one of the clearest ways a partner reduces risk before a system goes anywhere near a citizen.

Layered on top of the contract: a human checking outputs, with the threshold for intervention set high at first and relaxed only as the system proves itself. Not a permanent fixture, but not something switched off on day one either.

Telling people they’re talking to a machine

None of this works if the citizen doesn’t know what they’re dealing with. Someone interacting with a conversational AI system needs to know it’s AI, not a person, and increasingly that’s not just good practice but a legal requirement. Trust breaks the moment someone realises they’ve been talking to a machine that didn’t say so.

What HMRC taught us

Before Mercator Digital, several of our team worked inside HMRC, building and running digital services used by millions of people. A few things held true there that hold true now.

Trust is built slowly and lost quickly. One bad experience with a service can undo a long run of good ones, and that math is worse in government than almost anywhere else, because people don’t get to choose a different provider.

Saying what a system can’t do builds more trust than pretending it can do everything. Overpromising an AI system’s capability is the fastest way to burn the goodwill you’re trying to build.

Design for the person least confident using digital services first. If the service works for them, it works for everyone above them too. Designing for the average user and hoping the edges hold rarely does.

Staff trust matters as much as citizen trust. A caseworker who doesn’t trust the AI won’t back it in front of the person sitting across from them, no matter how well it performs.

Trust is organisational, not just technical

Alastair’s Government Transformation Summit piece points at where this actually breaks down: data. He cites research suggesting 80% of AI projects fail, largely due to inadequate training data, and notes that 69% of data and AI roles in government are classified as hard to fill.

Departments are being asked to run AI-enabled services on data estates that were never built for it, using teams that don’t yet have the people to close the gap. That’s not a technology problem. It’s an organisational one, and it’s why trust in public sector AI has to be built at the level of the institution, not just the model.

en_GB
Scroll to Top