Governments need an AI exception ledger before pilots scale

In an article responding to Global Government Forum’s recent AI interviews and features, Dr Gleb Tsipursky – a behavioural scientist and chief executive of Disaster Avoidance Experts – explains why keeping a formal record of the incidents and interventions that occur during pilots will help governments scale artificial intelligence safely
Global Government Forum recently highlighted a central problem facing public-sector AI: governments no longer lack promising ideas, but they still struggle to turn scattered pilots into services that work consistently at scale. In a separate interview, Phil Swan, director for government digital enablement at the UK’s Government Digital Service, described AI as a service and leadership challenge rather than simply a technology project. That distinction should shape the next phase of government adoption.
Now that the EU’s new transparency obligations are in force, public bodies have stronger reasons to tell people when certain AI systems or synthetic content are involved. Disclosure matters. Yet a label cannot reveal whether a flawed output was caught, whether a frontline worker had authority to challenge it, or whether the same mistake is recurring across several agencies.
Governments therefore need an internal companion to public transparency: an AI exception ledger.
An exception ledger is a structured record of the moments when an AI-assisted workflow departs from its expected path. It captures not only serious incidents, but also near misses, human overrides, awkward handoffs, hidden rework and cases in which a system technically completed a task while making the service worse. Those moments are where government learns what a pilot can and cannot safely do.
Consider a system that drafts a benefits letter, summarises a planning file, helps triage permit applications or analyses procurement documents. A conventional dashboard may show response times, usage and average accuracy. It may not show that staff repeatedly rewrite the same confusing paragraph, that one type of case always requires escalation, or that time saved in one team creates verification work for another.
The ledger should be brief enough to use under pressure. For each exception, it should record the public service or task involved, the input and its age, what departed from expectations, the human decision or override, who or what could have been affected, the correction made, and the condition that should trigger future review. The purpose is not to document every prompt. It is to identify patterns that ordinary performance measures miss.
Event: Public Service Data.AI is the UK’s flagship annual event for civil servants working to unlock the power of data and artificial intelligence across government. Brought to you by Global Government Forum and hosted by HM Government, the event will take place in London on 15 October 2026 and is free to attend for UK and international public servants. Find out more about the conference and register to attend
Human in the loop, shared learning, and improving decisions at scale
This approach improves accountability in three ways.
First, it makes human responsibility visible. ‘Human in the loop’ is often treated as a sufficient safeguard, even when the reviewer has little time, limited authority or no clear standard for intervention. A ledger shows when people actually challenge the system, what they change and whether their intervention prevents a problem. It turns human oversight from a slogan into observable work.
Second, it helps governments share learning. One department may discover that an AI summary becomes unreliable when source documents are older than a certain date. Another may find that a citizen-facing assistant handles routine questions well but fails when a person expresses distress or disputes an official record. Those lessons should not remain buried in local email chains. Common exception categories can reveal reusable design patterns, procurement requirements and training needs across government.
Third, it improves decisions about scale. Pilots often look successful because they are evaluated on average performance in controlled conditions. Public services operate in the tails: unusual cases, incomplete records, conflicting rules, language barriers and people whose circumstances do not fit a standard template. A pilot should graduate only when leaders understand those exceptions and can show that the service has a reliable way to detect, escalate and learn from them.
The ledger must not become a blame register. Employees will hide workarounds and near misses if every entry is treated as evidence of personal failure. Managers should reward early reporting, especially when a worker pauses a process before harm occurs. Aggregated patterns usually matter more than naming the individual who found the problem.
Nor should the ledger become surveillance. It should focus on the performance of the workflow, not on counting which employee used a tool most often. Public servants need space to exercise judgment, explain uncertainty and test safer alternatives. The goal is to make the system more accountable, not to reduce professional discretion to another metric.
Central digital teams can support this practice by publishing a small common taxonomy. Useful categories might include stale data, unsupported inference, missing context, privacy concern, inaccessible output, inconsistent treatment, failed handoff, excessive verification work and unclear ownership. Agencies could adapt the categories to their services while contributing anonymised lessons to a shared evidence base.
The measures should also change. Alongside time saved and service volume, leaders should track the share of exceptions caught before they affected the public, correction time, recurrence, workload transferred between teams, consistency of overrides and whether staff understand when to stop the process. These indicators reveal whether adoption is becoming safer as it grows.
Government AI will not scale through central rules alone. It will scale through thousands of frontline decisions about what to trust, what to question and when to intervene. Public transparency tells citizens that AI is present. An exception ledger gives government the operational memory to show that someone is still paying attention.
The strongest public sector AI systems will not be those that never produce an exception. They will be those whose institutions notice exceptions early, learn across organisational boundaries and improve before a near miss becomes a public failure.
Author
Gleb Tsipursky PhD is a behavioural scientist, CEO of Disaster Avoidance Experts, and author of The Psychology of AI Adoption at Work: From Resistance to Results (Georgetown University Press, 2026).