To Err is Human. What is it to be AI?
Part 1 of 2: The Diagnosis
To Err is Human. What is it to be AI?
Part 1 of 2: The Diagnosis
An essay on AI, trust, and the changing social contract of software
---
Artificial intelligence is advancing rapidly into domains that have, until now, been the exclusive territory of human judgement. This shift exposes a profound expectations gap: one built by sixty years of conditioning that computers give correct answers. The novelty here is not simply that AI systems are probabilistic — statistical and scoring systems have existed for decades, and their probabilistic nature has generally been understood by those deploying them. The distinctive challenge of conversational AI is that it delivers probabilistic judgement through an answer-shaped interface: fluent, personalised, and authoritative in register, in a way that actively obscures the uncertainty beneath. This essay argues that unless this gap is honestly acknowledged and systematically addressed — by technology leaders, regulators, and organisations deploying AI — it will not merely produce friction. It will generate a crisis of trust sufficient to stall the adoption of AI at the scale its economic projections require. This is not an argument against artificial intelligence. It is an argument for a more rigorous and culturally honest contract with it.
---
I. The Oracle Problem
For most of its history, the computer has been an oracle. Ask it a question and it gives you an answer. The answer may depend on the quality of the data you fed it, the logic of the programme written for it, or the competence of the person who configured it — but the machine itself does not equivocate. It does not have a bad day. It does not misread the file because it stayed up too late. Given the same inputs, it will produce the same output every single time. That is the foundational promise of the digital computer, and it is a promise that has been kept with extraordinary reliability for six decades.
This reliability has shaped more than our technical infrastructure. It has shaped our psychology. Three generations of workers have entered organisations in which the computer was the authoritative arbiter of record. The ledger balances because the system says it does. The compliance report is correct because the software generated it. The contract is valid because the database shows it signed. We did not arrive at this trust naively — it was earned, transaction by transaction, across billions of interactions. But in earning it, we created something more durable and more dangerous than a useful working assumption. We created a cultural reflex. It is worth noting that many of these systems were never truly oracular in a strict philosophical sense. They were sometimes wrong, sometimes poorly configured, sometimes fed bad data. But they were deterministic enough, repeatable enough, and embedded deeply enough in organisational process that the experience of using them came to feel authoritative. The reflex was not logically derived. It was socially constructed, through decades of reliable-enough behaviour, until it hardened into expectation.
That reflex is this: when a computer produces an output, it is correct. Not probably correct. Not correct within a margin of error. Correct. Any deviation from this is not a feature of the system — it is a fault in the system, and faults are fixed. This is the Oracle Problem. And it is about to collide with a technology that is constitutionally incapable of satisfying it. The obvious objection is that organisations have always known computers can be misconfigured, fed bad data, or programmed with flawed logic. That is true. But those failures were understood as failures of the people who built or operated the system, not of the system itself. The machine remained correct in principle; the humans around it were the variable. What AI introduces is a system that is uncertain by design — and that distinction, as this essay will argue, changes everything that follows from it.
---
II. The Judgement Threshold
The work that computers have traditionally automated shares a common characteristic: it is, in principle, fully specifiable. A payroll system can be told, precisely and completely, every rule it needs to follow. A sorting algorithm can be given an unambiguous definition of order. A compliance check can be given every regulation it must test against. This is not to say such systems are simple — they can be extraordinarily complex — but complexity is not the same thing as judgement. Judgement is what is required when the rules run out; when context matters; when two defensible interpretations of the same evidence lead to different conclusions.
For decades, that threshold was the informal boundary of automation. Computers handled the specifiable. Humans handled everything else. This was not merely a technical limitation — it was an epistemological one. Nobody seriously attempted to encode human judgement into software because the project was obviously doomed. Judgement, by its nature, resists complete specification.
Large language models and the agentic AI systems built upon them represent the first serious attempt to cross that threshold. They do not work by following explicit rules. They work by developing, through exposure to vast quantities of human-generated text, something that functions operationally like judgement. They can read an ambiguous regulatory disclosure and produce a plausible interpretation. They can review a contract and identify clauses that might be problematic in a particular context. They can assess a company’s ESG data and generate a narrative report. None of these tasks was automatable before. All of them are, to varying degrees of reliability, automatable now. But there is a further distinction that has received insufficient attention. Previous probabilistic systems — statistical models, scoring engines, actuarial tables — generally presented their outputs in a form that signalled their own uncertainty. A risk score is a number with a range. A probability table is legibly an estimate. The system’s nature was visible in its output. Conversational AI presents its probabilistic judgement as a fluent, personalised answer — syntactically confident, contextually specific, and structurally indistinguishable from the output of a knowledgeable human expert. The interface erases the uncertainty that the underlying model has not resolved. This is where the expectations gap becomes acute: not because AI is probabilistic, but because it presents probability in the form of certainty.
The move into agentic AI is not a quantitative advance in automation. It is a qualitative one. And the qualitative difference is that the system can now be wrong in the way that humans are wrong — through misjudgement rather than malfunction.
This distinction matters enormously. A malfunction is a failure of engineering. It can be identified, isolated, and repaired. A misjudgement is something else entirely. It may be reasonable given the available information. It may even be defensible. But it is still wrong, and the systems we have built to manage human error — verification, appeal, accountability, oversight — are not the same systems we have built to manage software failures. We are now operating in a space where neither set of systems is fully adequate.
---
III. The Error We Don’t Have Words For
Language is one of the most reliable indicators of cultural readiness. When a society has developed the vocabulary to describe a phenomenon, it has usually also developed the conceptual framework to manage it. And the reverse is equally telling: the absence of language for something is evidence that the cultural framework for it does not yet exist.
Consider the proverb: “To err is human; to forgive, divine.” Alexander Pope wrote those words in 1711, but the sentiment they encode is far older. It represents millennia of accumulated cultural wisdom about the nature of human fallibility — the recognition that those who exercise judgement will sometimes exercise it badly, and that the appropriate response is not systemic rejection but proportionate accountability, correction, and yes, forgiveness. We have built this wisdom into our legal systems, our employment practices, our personal relationships, and our organisational structures. We expect error. We have processes for it.
It is important to be precise about what the proverb is actually doing. It is not merely acknowledging that errors happen. It is licensing tolerance of them — encoding, at the level of shared cultural understanding, that imperfect judgement is an inherent feature of minds rather than a correctable defect. And crucially, it extends that licence to all human judgement, including judgement that has been made highly abstract, statistical, or systematised. The actuary sits at the furthest end of this spectrum: their outputs are probability tables, not conversations. An actuarial output arrives in a form that announces its own nature — a range, a confidence interval, a modelled estimate subject to stated assumptions. It does not present itself as an answer. It presents itself as a calibrated probability, and the institutional frameworks built around it — professional standards, regulatory oversight, peer review — are calibrated accordingly. And yet we still extend the proverb’s permission to the actuary, because a human being constructed the model, selected the assumptions, and stands behind the number. The humanity, at that point, is almost vestigial — but it still matters. It is doing invisible cultural work that we have never had to make explicit, because we have never before encountered judgement without a human being behind it. What is striking about conversational AI is that it collapses both of these properties simultaneously. It removes the human author and it removes the visible uncertainty. It delivers a probabilistic judgement in the form of a confident answer, authored by no one, hedged by nothing. That combination has no institutional or cultural precedent.
What artificial intelligence removes is not the error. It removes the humanity that licenses tolerance of the error. An AI system producing a regulatory interpretation, a risk assessment, or a clinical recommendation is doing something that looks, functionally, like what a human professional does. But it exists outside the moral and cultural category that the proverb governs. It is neither human enough to inherit the permission slip, nor mechanical enough to be held to the older standard of deterministic correctness. It occupies a novel category for which we have no prepared cultural response — and the absence of that response is not merely a linguistic curiosity. It is a diagnostic signal that the institutional infrastructure for managing AI-generated error does not yet exist.
The legal system makes this concrete. If a human professional — a solicitor, a financial adviser, a sustainability consultant — makes a misjudgement that causes harm, we have well-developed mechanisms for accountability: negligence claims, regulatory sanctions, professional indemnity. These mechanisms work because there is a human being at the end of the chain who can be held responsible, retrained, or removed. If an AI system makes the same misjudgement, the accountability chain does not simply transfer — it becomes fragmented and insufficiently legible. Is the developer of the model responsible? The organisation that deployed it? The individual who did not verify its output? The prompt engineer who designed the interaction? Each party can point to another. Current law provides no clear answer, and the absence of that answer is not a technical gap. It is a symptom of a cultural one — we have not yet worked out what we think about a judgement that has no human author.
The negligence framework is further undermined by a subtler problem. Negligence law is not simply about getting the wrong answer. It is built around the concept of the reasonable professional — did the person exercise the standard of care that a competent practitioner would have applied in the same circumstances? Courts do not ask whether the outcome was correct. They ask whether the process was sound. This is why many negligence claims fail: the defendant considered the available evidence, applied professional standards, and reached a conclusion that a reasonable peer might also have reached. The law has, in effect, already institutionalised the proverb. It accepts that good-faith judgement sometimes produces wrong outcomes, and it does not punish that. It punishes the failure to exercise judgement properly.
An AI system cannot be assessed against this standard at all. A probabilistic system does not exercise judgement in good faith and occasionally fall short. It produces outputs that are, by design and by definition, sometimes incorrect — and the frequency of that incorrectness is a function of its architecture, not a deviation from its intended behaviour. The error is not a failure of the system. It is the expected output of the system operating normally. You cannot hold a system negligent for doing precisely what it was designed to do. The existing negligence concepts do not map cleanly onto a probabilistic system operating as designed. The framework is not merely strained — it addresses a fundamentally different kind of failure than the one AI produces.
In open-ended judgement domains, error is not a temporary bug on the way to zero — it is bound up with the flexibility that makes the system useful. A model precise enough to eliminate uncertainty in ambiguous interpretation would no longer be interpreting ambiguity. It would be resolving it through rules — which is, once again, determinism by another name.
This is the hardest truth in this paper, and the one most consistently avoided in public discourse about AI reliability. The flexibility that allows a language model to interpret an ambiguous regulatory clause, to read contextual nuance in a corporate disclosure, to generate a plausible analysis from incomplete data — all of this emerges from the same probabilistic architecture that makes error inevitable. Remove the probability and you remove the capability. The question in open-ended judgement domains is therefore not how we engineer AI systems to a zero-error standard — in those domains, that standard is not achievable without sacrificing the contextual flexibility that makes the systems valuable. The question is instead how we build the cultural, legal, and institutional frameworks to accommodate a permanently imperfect but genuinely useful tool. That is a considerably harder and more honest problem than the one most organisations think they are solving.
---
IV. The Arithmetic of Scale
To understand why the error problem is not merely philosophical, it is useful to examine the arithmetic of scale. Consider an organisation that employs one hundred professionals to perform a knowledge-intensive task — reviewing regulatory disclosures, assessing supplier risk, producing compliance reports. Each of those professionals makes mistakes. The error rate varies by individual, by complexity of the task, by time of day and state of mind — but let us posit an average error rate of two percent. In a large organisation, that is an accepted operational reality. Training is provided. Spot checks are run. A senior reviewer handles edge cases. The errors are distributed, individual, and correctable.
Now consider replacing those one hundred professionals with an AI system operating at the same two percent error rate. The economic case looks straightforward — same output, dramatically reduced cost. But the error profile is completely different, and the difference is not benign.
Human errors at scale are, in statistical terms, largely uncorrelated. Worker A’s mistake on a Monday morning is unlikely to be the same mistake as Worker B’s on a Thursday afternoon. The errors scatter. This scattering means they are relatively unlikely to compound and relatively likely to be caught by the natural diversity of the team’s working patterns.
AI errors at scale are, by contrast, highly correlated. If the model has a systematic bias in its interpretation of a particular regulatory clause — perhaps because that clause was ambiguously represented in its training data — then that bias will manifest identically across every single instance where the clause appears. The error does not scatter. It replicates. A two percent human error rate produces scattered individual mistakes. A two percent AI error rate can produce a systematic, organisation-wide misinterpretation that propagates invisibly until it surfaces as a regulatory failure, a legal liability, or a reputational crisis. It is worth acknowledging that human organisations can also produce correlated error — through shared training, shared incentive structures, shared hierarchical pressure, or shared ideological assumptions. The groupthink failures of financial institutions before 2008 are one example. But human-generated correlation of this kind typically builds slowly, can be disrupted by individual dissent, and tends to leave visible traces in internal communications and decision records. AI-generated correlation propagates instantly, at scale, across every deployment of the same model, with no internal dissent and no audit trail of the reasoning behind it. The speed, breadth, and invisibility of AI correlation are qualitatively different from anything that human organisations ordinarily produce.
The same error rate means something entirely different depending on whether errors are distributed or systematic. This is being almost universally underestimated by organisations currently evaluating AI for high-consequence tasks.
There is a further dimension, best illustrated by the kind of manual processing work that was common before automation arrived in public sector administration. In benefits assessment, for instance, teams of clerks would process large volumes of claims, and every manager was judged against an error rate — typically a threshold of no more than five percent of assessments containing a material mistake. Individual performance varied considerably. One poor performer could drag an entire team’s rate above the acceptable threshold. But the response was well understood and relatively precise: identify the individual, provide retraining, monitor their work more closely, and if the errors persisted, remove them from the assessment process. The intervention was targeted, low-cost, and left the rest of the team’s performance undisturbed. The system was, in that sense, self-correcting at the individual level. An AI system operating across the same task does not offer this kind of remediation. If the system develops a systematic error pattern — selecting the wrong category, misapplying a rule, consistently misreading a class of input — the available responses are limited and none of them are surgical. Prompts can be revised. Guardrails can be added. In extreme cases the model can be replaced or fine-tuned. But each of these interventions is expensive, technically disruptive, and liable to introduce new problems while solving the original one. You cannot take the AI aside for a quiet conversation and retrain it on the specific case it keeps getting wrong. The asymmetry between human and AI remediation is as significant as the asymmetry in error correlation — and it is almost entirely absent from the commercial assessments organisations make when evaluating AI for high-consequence tasks. A reasonable response to all of this is that AI reliability is improving rapidly, and that the error rates of frontier models today are materially lower than they were two years ago. That is true, and it matters. The argument here is not that AI systems cannot become more reliable — they can and will. It is that in open-ended judgement domains, improving reliability and eliminating uncertainty are different things, and that the cultural and institutional frameworks for living with residual uncertainty need to be built regardless of where the reliability curve eventually plateaus.
---
V. The Human-in-the-Loop Illusion
The standard organisational response to concerns about AI reliability is the human-in-the-loop. A subject matter expert reviews AI outputs before they are acted upon. Exceptions are escalated. Edge cases are flagged. This is presented as a responsible middle path between full automation and the status quo. In many contexts, it is a reasonable one. In many others, it is an illusion — and an economically costly illusion at that.
The illusion rests on a false assumption: that verification is substantially cheaper than creation. It is worth distinguishing three cases, because they have very different implications. The first is review that is genuinely cheap: for simple, high-volume tasks — categorising invoices, routing customer enquiries, flagging anomalies in structured data — a trained reviewer can glance at an AI’s output and confirm or correct it in seconds. Here the economics of human-in-the-loop are sound, and the model works as advertised. The second case is review that is nearly as hard as creation: in knowledge-intensive professional work, verifying an AI-generated output requires essentially the same domain expertise, and much of the same cognitive effort, as producing it. The productivity gain is real but considerably more modest than the headline billing of AI-driven automation. The third case is the most dangerous: performative review, where output volume, time pressure, and the authoritative presentation of AI-generated content combine to produce a nominally supervised process that is, in practice, unsupervised. The human reviewer is present but not genuinely reviewing. The oversight exists on paper but not in substance.
The ESG disclosure context illustrates the second and third cases in combination. An AI system can produce a plausible-looking disclosure from structured data inputs. But for a qualified sustainability professional to verify that disclosure — to confirm that the materiality assessment is defensible, that the metrics have been correctly calculated, that the narrative is consistent with the underlying data, that the interpretation of regulatory requirements is sound — requires essentially the same domain knowledge, and much of the same cognitive effort, as producing the disclosure in the first place.
In these domains, the human reviewer is not checking a box. They are exercising the same judgement the AI was supposed to supply. The system has not replaced their work. It has, at best, produced a first draft for them to validate — which is a productivity gain, but a considerably more modest one than the headline billing of AI-driven automation. And at worst, it has produced an authoritative-looking output that the reviewer skims rather than scrutinises, because the volume of outputs and the pressure to process them makes deep review economically impractical. This is where the human-in-the-loop becomes not just ineffective, but actively dangerous: it creates the appearance of oversight without its substance.
---
VI. The Compounding Problem: When Agents Coordinate
The carbon calculation system described earlier in this essay is, in the architecture of modern AI deployment, a relatively simple construction. Three sequential probabilistic steps, each with a defined input and output, operating within a bounded domain. And yet its error profile is already sufficiently complex that the same inputs can produce materially different outputs on different runs, with no single stage appearing obviously broken and no straightforward mechanism for a human reviewer to identify where in the chain a given error originated.
This complexity is about to become considerably more serious. The mechanism by which it will do so is the emergence of multi-agent AI systems — architectures in which an orchestrating AI does not simply execute a fixed sequence of steps, but dynamically coordinates a network of specialised sub-agents, deciding in real time which to invoke, in what order, with what parameters, and how to interpret their outputs before issuing the next instruction.
In a single probabilistic system, the error rate is a property of that system. In a chained system of three steps, errors interact — a drift in stage one influences what stage two surfaces, which shapes what stage three has to work with. But in a multi-agent architecture, the problem changes character entirely. The orchestrating agent is itself probabilistic. Its decision about which sub-agents to invoke, and how to frame their tasks, is a judgement — one that may be correct most of the time and wrong in ways that are contextually specific, difficult to predict, and deeply consequential for everything that follows. Each sub-agent it coordinates is also probabilistic. And the orchestrating agent’s interpretation of each sub-agent’s response — before it decides what to do next — introduces a further layer of uncertainty. You do not have a chain of probabilities. You have a dynamic graph of probabilistic interactions, where errors do not simply accumulate but compound in ways that are highly dependent on the specific path the orchestration takes.
There is a further property of these systems that distinguishes them sharply from anything discussed so far: the intermediate steps may be entirely invisible. In a three-stage pipeline, a developer can at least inspect the output of each stage in sequence. In a mature multi-agent system, an orchestrating agent may invoke sub-agents, receive their outputs, make decisions based on those outputs, and act on those decisions — all before any human operator has the opportunity to observe what has occurred. The error does not merely compound. It compounds in the dark.
The commercial logic driving the adoption of multi-agent systems is self-reinforcing. The productivity case rests precisely on minimal human intervention. The features that make them economically attractive are the same features that make their error profiles hardest to audit. Organisations deploying these systems for the efficiency gains they offer are, in a meaningful sense, purchasing opacity along with the productivity.
The parallel with complex financial instruments in the years before 2008 is instructive, if uncomfortable. Individual mortgage-backed securities were imperfectly understood by many of the institutions that held them. The collateralised debt obligations constructed from those securities were considerably less understood. The synthetic instruments built from those were, for most practical purposes, opaque to the people trading them. Each layer of abstraction added not merely more complexity but more distance between the people making decisions and the underlying risk they were actually carrying. The system did not appear to be failing until the moment it catastrophically did — because the opacity that was a feature of the architecture was also the reason its fragility could not be seen until it was too late.
The governance frameworks currently being developed for AI — to the extent that they exist at all — are largely calibrated to single-model deployments: a system takes an input, produces an output, and a human reviews it. That framework was already inadequate for a three-stage pipeline. For multi-agent orchestration at scale it is not even a starting point. The question of where accountability sits when an orchestrating agent makes a poor decision that propagates through five sub-agents before surfacing as a wrong answer to the end user is one that no existing legal, regulatory, or organisational framework is equipped to answer. The chain of custody for the error does not merely dissolve, as it does in a single-model deployment. It was never visible in the first place.
This is not a distant prospect. The architecture is already in commercial deployment across multiple sectors. The trajectory is clear. And the window for developing adequate governance frameworks — before multi-agent systems are embedded deeply enough in critical processes that their error profiles become a systemic rather than an operational risk — is narrowing faster than the pace of institutional adaptation would suggest is comfortable. It is worth acknowledging that thoughtful engineering can mitigate some of this. Logging intermediate steps, designing for auditability, building in human confirmation at high-stakes decision points — these are real responses to the opacity problem. The concern is not that engineering solutions are unavailable, but that commercial pressure typically pushes against their adoption: logging and checkpoints add latency and cost, and the organisations most eager to deploy agentic systems are often those most motivated to maximise their autonomy.
---
Part 2 — The Consequences and the Response — follows next week, covering the adoption ceiling, what history tells us about societies learning to live with probabilistic judgement, a practical framework for deployment, and a conclusion on what it will take to build the cultural and institutional infrastructure AI actually needs.
---
Alasdair Mangham is Executive Director EMEA at Convene, a governance technology company. His career in technology started in 1994, building systems to enable sign language users to connect with remote interpreters via video telephony — bridging the gap between citizens and government services that were designed around voice — and has since taken in the first local government websites, open source content management, the semantic web, digital democracy, voice interfaces, and what the industry used to call human-computer interaction, then UX, and is now apparently calling an AI agent. Throughout, he has worked on the requirements, seen them built into systems, and lived with the adoption. The questions in this essay have followed him every step of the way.
