---
title: "AI Accountability"
date: 2026-07-12
description: Agency can move to the machine; accountability cannot, and answering for what AI decides now takes capability that policy alone does not supply.
author: Mario Thomas
canonical: https://mariothomas.com/briefings/ai-accountability/
---

## Start here

Two short reads that set the terms before going deeper: what AI accountability actually is, and the order to take the briefing in.

### The Line That Does Not Move

AI accountability is the question of who answers when the machine
decides, and courts, regulators, and clients have answered
consistently. 'The AI did it' fails as a defence. An airline was
held liable for its chatbot's incorrect bereavement-fare advice.
Lawyers have been sanctioned for citing cases that never existed.
A major consultancy partly refunded a government client for a
report built on fabricated references. The work moved to the
machine; the accountability did not move with it.

Most of the Boards I meet treat this as a policy problem: write
the AI policy, add a human-in-the-loop clause, approve the
deployment on accuracy metrics and business case. The gap is not
in the paperwork. It is in capability. Can the organisation
verify outputs at the volume the system generates? Can the system
explain a decision to the person it has just decided about, well
enough to survive the four safeguards UK law has required since
5 February 2026? Is the assurance beneath the Board paper
sampling, or proof? And whose values does the model apply to the
case nobody designed for?

The position this briefing takes is that accountability is
engineered, not documented. Verification capability grows in
people, which is why replacing the junior roles that build
expertise quietly destroys the pipeline that produces tomorrow's
verifiers. Reasoning capability is designed into systems, and the
law now assumes it is there. Assurance can move from probable to
provable in domains a Board can formally specify. And the ethical
standard in force is either chosen deliberately or inherited from
the model's provider by default.

A reader who finishes this briefing should be able to look at any
material AI deployment and answer four questions: who verifies it,
whether it can explain itself, whether its assurance rests on
probability or proof, and whose standard it runs. Where those
answers are missing, the judgement to make is not how to mitigate.
It is whether to approve the system at all.

### From Principle to Proof

The core pieces in [the articles](#core-reading) are numbered
because they build one argument. Start with The Accountability
Gap, which establishes the principle everything else rests on:
agency for the work can be transferred to a machine,
accountability for the outcome cannot, and verification is the
capability that closes the space between the two.

The Reasoning Gap takes that principle into law. Since 5 February
2026 the UK regime has required four safeguards for solely
automated decisions, and the article shows why those safeguards
are capability tests rather than policy positions, and why
probabilistic systems do not carry the capability by default.
From Probable to Provable then turns to assurance: automated
reasoning, and the reasoned indicators that let a Board prove
what it once had to estimate. Ethical AI closes the sequence at
the deepest layer, the value system a model carries into
deployment, and the choice between accepting it, rejecting it,
or building alignment the organisation owns.

The further pieces put the argument to work. Maximum Fidelity
assembles all four indicator types into a single decision
instrument. AI and the CFO follows the line to the first
executive who signs, and The Balancing Item prices the oversight
labour that verification requires, the hours most business cases
never count. The Verification Premium shows what that labour is
made of, expertise, which delegating the writing of code to AI
consumes rather than replaces. Agentic AI explains what an agent
actually is and where delegated agency sits, the transfer The
Accountability Gap builds on, and AI and the CEO carries the line
to its executive end, the chief executive who can hand over the
building and the drafting but still answers for the bets. The
Board in the machine is the 2022 piece where this thinking began,
and Minimum Lovable Governance is the operating principle through
which the duty actually gets delivered.

From there, [Remake](#remake-assets) holds the mechanisms
beneath the thinking, and [the questions](#faqs) are the ones I
would take into the next Board meeting. None of it requires
technical depth. All of it requires a willingness to ask what
the organisation could prove about the decisions it has already
automated.

## Core reading

One argument, built in order: accountability cannot transfer, the law now assumes the capability, proof can replace probability, and the values in force must be chosen.

1. [The Accountability Gap: When AI Delegation Meets Human Responsibility](https://mariothomas.com/blog/ai-agency-accountability/) (15 minute read, 16 November 2025): Organisations are transferring decision-making agency to AI while accountability stays with people, and approving deployments without the verification capability that accountability needs.
2. [The Reasoning Gap: The Capability the Law Now Demands of Boards](https://mariothomas.com/blog/the-reasoning-gap/) (11 minute read, 3 May 2026): UK law now requires four safeguards for solely automated decisions. Most Boards have approved probabilistic systems that cannot deliver them in operation. Podcast edition: 12 minute listen.
3. [From Probable to Provable: What Automated Reasoning Means for the Board](https://mariothomas.com/blog/automated-reasoning-explainer/) (13 minute read, 5 April 2026): Automated reasoning gives Boards access to proof, not probability. This article explains what it is, where it already operates, and why it changes governance. Podcast edition: 16 minute listen.
4. [Ethical AI: When the Model Imposes Values Your Organisation Did Not Choose](https://mariothomas.com/blog/ethical-ai-inherited-values/) (14 minute read, 17 May 2026): A foundation model arrives with a value system its provider built and the Board did not choose. The decision: accept it, reject it, or build. Podcast edition: 15 minute listen.

## Further reading

- [Maximum Fidelity: How Four Indicator Types Strengthen Board Decisions](https://mariothomas.com/blog/maximum-fidelity-four-indicators/) (13 minute read, 12 April 2026): Four indicator types give boards progressively higher decision fidelity: lagging, leading, predictive, and reasoned. Together they represent the most accountable governance instrument available. Podcast edition: 15 minute listen.
- [AI and the CFO: Standing Behind the Numbers the Machine Produces](https://mariothomas.com/blog/ai-board-director-cfo/) (12 minute read, 7 June 2026): AI can run the close, sharpen the forecast, and operate out of sight, yet the CFO still signs. Accountability for the numbers does not move. Podcast edition: 13 minute listen.
- [The Board in the machine](https://mariothomas.com/blog/the-board-in-the-machine/) (10 minute read, 17 October 2022): AI and machine learning are becoming ubiquitous in business decisions, and Boards need to know what is deployed and how it is governed.
- [Minimum Lovable Governance: The AI Operating Principle Boards Should Use](https://mariothomas.com/blog/minimum-lovable-governance/) (13 minute read, 30 November 2025): Minimum lovable governance replaces episodic compliance with continuous, embedded oversight people actually want to use: guardrails that earn adoption rather than enforce it. Podcast edition: 17 minute listen.
- [The Balancing Item: The AI Oversight Cost Your Business Case Never Priced](https://mariothomas.com/blog/unpriced-cost-ai-oversight/) (10 minute read, 26 July 2026): Every AI business case counts the hours saved. Almost none counts the oversight hours added, and people are silently absorbing the difference. Podcast edition: 13 minute listen.
- [The Verification Premium: What Classical Training Reveals About AI Coding Costs](https://mariothomas.com/blog/vibe-coding-vs-classical-training/) (13 minute read, 25 January 2026): AI coding tools amplify the expertise gap rather than closing it: senior developers capture twice the gains. The verification premium is the cost nobody budgets. Podcast edition: 18 minute listen.
- [Agentic AI: Strip Away the Hype and Understand the Real Strategic Choice](https://mariothomas.com/blog/agentic-ai-explainer/) (17 minute read, 2 November 2025): Agentic AI is this year's poster child, and most of the confusion is about what agents actually do. The Board's decision is strategic, not technical.
- [AI and the CEO: Choosing the Bets That Matter](https://mariothomas.com/blog/ai-board-director-ceo/) (12 minute read, 21 June 2026): AI can build, deliver, and draft, yet the chief executive still chooses and still answers. Accountability for the bets does not move. Podcast edition: 12 minute listen.

## Remake

The mechanisms beneath the thinking: the model, diagnostic, methodology, and principle from the Remake Library that turn this briefing into apparatus a Board can use.

- **Model: Maximum Fidelity**. The evidence discipline of grading every reading a Board relies on by one of four indicator types: lagging indicators of past outcomes, leading indicators of early signals, predictive indicators of future value, and reasoned indicators that prove what must hold true. [Remake Library](https://mariothomas.com/remake/library/#maximum-fidelity)
- **Diagnostic: Process Audit**. The per-process evaluation applied one process at a time: what the work is for, how it is done today and whether it is still needed, and why it needs the technology at all, returning a verdict on the Stop, Keep, Remake scale with every reading graded by indicator type. [Remake Library](https://mariothomas.com/remake/library/#process-audit)
- **Methodology: AI Business Case**. The integrated decision framework that crystallises across an ADAPT engagement rather than at a single stage: strategic alignment established at Align, cost and readiness evidenced at Diagnose, value shaped at Advise, and execution designed at Plan. [Remake Library](https://mariothomas.com/remake/library/#ai-business-case)
- **Principle: Minimum Lovable Governance**. Governance embedded in how work happens: proportionate to risk, continuous rather than episodic, and used because it works. [Remake Library](https://mariothomas.com/remake/library/minimum-lovable-governance/)

## Questions

The questions a Board should be asking about its own position on accountability, answered from the work in this briefing.

### If the AI produced the work, are we still accountable when it is wrong?

Yes, and the record is clear. Courts, regulators, and clients have consistently rejected 'the AI did it' as a defence: an airline held liable for its chatbot's advice, lawyers sanctioned for fabricated citations, a major consultancy partly refunding a government client. Agency for the work can move to a machine; accountability for the outcome stays with the humans who deployed it. I set out the evidence, and the strategic choice it forces, in [The Accountability Gap](/blog/ai-agency-accountability/).

### What does UK law actually require of our automated decisions?

Since 5 February 2026, the UK GDPR as amended by the Data (Use and Access) Act 2025 has required four safeguards for any significant decision taken solely by automated processing: information about the decision, the ability to make representations, human intervention, and the right to contest. On the page these are procedural rights. In operation they are capability tests, and most probabilistic systems cannot pass them unless the capability was engineered in at design time. [The Reasoning Gap](/blog/the-reasoning-gap/) works through what that means for the systems a Board has already approved.

### Does deploying AI reduce our need for human expertise?

The evidence points the other way: the more an organisation delegates to AI, the more expertise it needs to verify the outputs. Verification capability comes from years of doing the work, which is why replacing junior roles to fund AI efficiency quietly destroys the pipeline that produces the senior experts who check the machine. The choice between augmentation and replacement, and its ten-year consequences, is the second half of [The Accountability Gap](/blog/ai-agency-accountability/).

### Is there stronger assurance available than testing and sampling?

In domains that can be formally specified, yes. Automated reasoning proves properties across every possible state of a system rather than the scenarios someone thought to test, which is why civil aviation and the TLS 1.3 protocol already use it, and why financial institutions are beginning to apply it to capital adequacy calculations. For a Board this arrives as a fourth indicator type: reasoned indicators, which prove what must hold true rather than estimate what is likely. [From Probable to Provable](/blog/automated-reasoning-explainer/) explains the discipline, and [Maximum Fidelity](/blog/maximum-fidelity-four-indicators/) shows all four types applied to real decisions.

### Whose values is our AI actually applying?

Unless the Board has done deliberate work, the provider's. Every foundation model arrives with a value system built upstream in pre-training and alignment, and system prompts, retrieval, and guardrails constrain that standard without re-authoring it. The real decision, taken deployment by deployment, is to accept the provider's standard, reject the deployments where it bears directly on people, or build alignment the organisation owns. [Ethical AI](/blog/ethical-ai-inherited-values/) sets out how to make that choice deliberately rather than inherit it by default.

### Who performs the oversight our AI business cases assume?

In most business cases I see, nobody is named. The case counts the hours AI saves and leaves out the hours it adds: the reviewing, correcting, and deciding-whether-to-trust its outputs demand. That labour lands on existing people on top of existing jobs, and BCG Henderson Institute research published in 2026 found the strain attaching to oversight load, not to AI use itself. A right to human intervention on demand needs standing capacity to answer it. The credible answer names the roles, states the capacity removed from their workload, and prices it fully loaded. [The Balancing Item](/blog/unpriced-cost-ai-oversight/) sets out the three questions that test whether it has been priced.

### Can our Board pause or reverse an AI system whose behaviour proves unacceptable?

It should be able to, and the Institute of Directors' 2025 paper on AI governance in the Boardroom expects the Board to retain exactly that authority. The authority is only as real as the mechanism beneath it. A foundation model's value standard moves with every new version, without a fresh approval, so the Board needs to know which deployments it has accepted, which it has rejected, and what would trigger intervention on each. [Agentic AI](/blog/agentic-ai-explainer/) names the kill-switch as a governance minimum for any system running its own loop, and [Ethical AI](/blog/ethical-ai-inherited-values/) sets out the accept, reject, or build choice that gives the authority something to act on.

### Which of our approved decision systems are probabilistic rather than rule-based?

Few of the Boards I meet can say, and the inventory is the first priority The Reasoning Gap sets. A rule-based system carries its reasoning on the surface: the rule applied, the facts it operated on, the output that followed. A probabilistic system, of the kind approved for credit decisioning and fraud detection, does not, and the four safeguards UK law has required since 5 February 2026 have to be engineered in at design time. Most production systems are hybrid, so the question belongs at the point where the consequential output is produced, with the engineering function in the room. [The Reasoning Gap](/blog/the-reasoning-gap/) sets out the inventory.

## References

The statute, the regulator, the courts, and the research this briefing draws on across its articles.

- **UK Government** (2025): [Data (Use and Access) Act 2025](https://www.legislation.gov.uk/ukpga/2025/18/contents). The statute that rewrote the UK's automated decision-making regime and introduced the four safeguards.
- **Information Commissioner's Office** (17 December 2024): [Rights related to automated decision making including profiling](https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/individual-rights/individual-rights/rights-related-to-automated-decision-making-including-profiling/). The regulator's guidance on rights related to automated decision-making and profiling.
- **Court of Justice of the European Union** (7 December 2023): [Case C-634/21 SCHUFA Holding (Scoring)](https://infocuria.curia.europa.eu/tabs/affair?sort=AFF_NUM-DESC&searchTerm=%2522C%252D634%252F21%2522&publishedId=C-634%2F21). The European court ruling that refined what counts as a solely automated decision.
- **Institute of Directors** (2025): [AI Governance in the Boardroom](https://web.archive.org/web/20251121075957/https://www.iod.com/app/uploads/2025/09/AI-Governance-in-the-Boardroom-1c7612e872fa3fce3f9d6cad78b0b4ba.pdf). The Institute of Directors business paper on how AI oversight obligations attach to the Board as a whole.
- **Stanford HAI** (23 May 2024): [AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries](https://hai.stanford.edu/news/ai-trial-legal-models-hallucinate-1-out-6-or-more-benchmarking-queries). Benchmarking research finding general-purpose LLMs hallucinate on legal queries 58-82% of the time.
- **PwC** (2025): [The Fearless Future: 2025 Global AI Jobs Barometer](https://www.pwc.com/gx/en/issues/artificial-intelligence/job-barometer/2025/report.pdf). The 2025 Global AI Jobs Barometer: a 56% wage premium for AI skills and three times the revenue growth per employee in the most AI-exposed industries.
- **Stanford CRFM** (December 2025): [The Foundation Model Transparency Index](https://crfm.stanford.edu/fmti/). Scores major model providers at roughly 40 out of 100 on disclosure of how their models are built and aligned.
- **EUR-Lex** (12 July 2024): [Regulation (EU) 2024/1689 (Artificial Intelligence Act)](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689). Articles 9 to 15 impose documentation, transparency, and human oversight obligations on high-risk AI systems.
- **RTCA** (December 2011): [DO-178C: Software Considerations in Airborne Systems and Equipment Certification](https://www.rtca.org/do-178/). The civil-aviation software certification standard; its formal-methods supplement DO-333 (2011) recognises formal verification as assurance evidence.
- **CBC** (15 February 2024): [How can I mislead you? Air Canada found liable for chatbot's bad advice on bereavement rates](https://www.cbc.ca/news/canada/british-columbia/air-canada-chatbot-lawsuit-1.7116416). The 2024 tribunal ruling that held Air Canada liable for its chatbot's incorrect bereavement-fare advice, the first of the three cases the accountability argument rests on.
- **Harvard Business Review** (16 September 2025): [The Perils of Using AI to Replace Entry-Level Jobs](https://hbr.org/2025/09/the-perils-of-using-ai-to-replace-entry-level-jobs). Payroll-data evidence of a 13% decline in AI-exposed entry-level roles, and the expertise pipeline risk it signals.
- **Legal Dive** (26 June 2023): [Judge in ChatGPT case most troubled by attorneys’ lack of candor](https://web.archive.org/web/20260725171955/https://www.legaldive.com/news/chatgpt-lawyer-fake-cases-lawyer-uses-chatgpt-sanctions-generative-ai/653925/). The 2023 sanctions case in which lawyers filed ChatGPT-fabricated citations, one of the three rulings the accountability argument rests on.
- **Fortune** (7 October 2025): [Deloitte was caught using AI in $290,000 report to help the Australian government crack down on welfare after a researcher flagged hallucinations](https://fortune.com/2025/10/07/deloitte-ai-australia-government-report-hallucinations-technology-290000-refund/). Fortune's 2025 report of Deloitte partly refunding an Australian government client after AI-generated errors in a commissioned report.
- **arXiv** (4 February 2026): [Alignment Drift in Multimodal LLMs: A Two-Phase, Longitudinal Evaluation of Harm Across Eight Model Releases](https://arxiv.org/abs/2602.04739). The 2026 pre-print on alignment drift in multimodal models, the evidence that a model's value standard moves between versions without a fresh approval.
- **IEEE** (26 June 2017): [Verified Models and Reference Implementations for the TLS 1.3 Standard Candidate](https://ieeexplore.ieee.org/document/7958594). The 2017 IEEE paper on verified models and reference implementations for TLS 1.3, the protocol case the automated-reasoning argument cites.

## The ideas beneath this briefing

Ideas I've named and matured writing about AI Accountability: what each one means, and where it started.

- **Accountability Gap**: When an organisation delegates work to AI without building the capability to verify it, leaving people answerable for outputs no one has actually checked. For a Board, no delegation to AI should be approved without also approving who checks the output and how, because accountability without a verification step is accountability in name only. [Read more](https://mariothomas.com/blog/ai-agency-accountability/)
- **Reasoning Gap**: The gap between the four legal safeguards required for solely automated decisions and a system's actual ability to interrogate and explain its own decisions, a capability built into rule-based systems but absent by default in probabilistic ones. [Read more](https://mariothomas.com/blog/the-reasoning-gap/)
- **Reasoned Indicator**: A fourth indicator type alongside lagging, leading, and predictive indicators; where those estimate or forecast, a reasoned indicator proves what is possible, impossible, or must hold true under any combination of inputs. [Read more](https://mariothomas.com/blog/automated-reasoning-explainer/)
- **Decision Fluency**: A chief executive's hands-on familiarity with AI tools sufficient to judge what a bet is worth, distinguished from coding skill and treated as a duty rather than a nicety. [Read more](https://mariothomas.com/blog/ai-board-director-ceo/)

All concepts: https://mariothomas.com/glossary/concepts/

## More Board Briefings

More complete resources on AI and emerging technology for the Boards that need the full picture.

- [AI Governance](https://mariothomas.com/briefings/ai-governance/): Governance people route around fails to govern; the task is governing AI the Board cannot fully see without strangling adoption.
- [AI Regulation](https://mariothomas.com/briefings/ai-regulation/): AI regulation has split into incompatible regimes, and the Board's task is not compliance with each but a deliberate position across all of them.
- [AI & the Board](https://mariothomas.com/briefings/ai-and-the-board/): AI changes how every Board duty is discharged, from director to Company Secretary, and moves none of the accountability.
- [AI Risk](https://mariothomas.com/briefings/ai-risk/): The AI risks that bite are seldom on the register, and the bill for sovereignty shocks, readiness gaps, verification costs, and model risk arrives later.
