Beyond Automation: How AI Rewrites the Coordination Architecture of Organizations
Central finding
Increasing AI capability and agency does not have a single effect on organizational coordination. It can substitute for routine coordination, complement human judgment, transform the location and form of coordination, or amplify coordination burdens. The strongest discriminator in the available evidence is not model capability by itself. It is whether the organization can place a machine-checkable contract—a test, schema, invariant, rule set, permission boundary, reconciliation loop, or admission gate—in front of scarce human attention.
Where such a contract is possible, machines can absorb large volumes of integration, monitoring, translation, scheduling, reconciliation, and recovery. Google reduced two categories of database operational work by roughly 95%, but only after changing applications to tolerate automated failover rather than merely scripting the old process.1 A hospital operating-room scheduling system cut scheduling time by 71% using rule-based logic rather than machine learning.2 In a contact centre, an AI assistant increased resolutions per hour by 15% and reduced customer requests for a manager by almost 25%, with the largest productivity gains going to novices.3
Where acceptance remains semantic, contextual, contested, or difficult to verify, cheaper machine production tends to move work downstream. Humans must determine whether an output is relevant, safe, complete, appropriately uncertain, consistent with intent, legitimate, and compatible with the rest of the organization. In current generative software work, this often appears as higher production alongside longer review queues, more rework, greater instability, and sometimes less useful throughput.4 In open-source security channels, it appears as maintainers spending their time routing duplicate reports and explaining that an issue was already fixed.5
The central organizational problem is therefore not simply how to deploy more capable AI. It is how to move, formalize, or contain the boundary at which machine action encounters human judgment. Organizations that redesign that boundary can obtain durable reductions in routine coordination. Those that simply add fast generators to existing workflows risk acceleration without throughput: more actions, proposals, messages, and intermediate artifacts arriving at interfaces designed for human-rate traffic.
No study in the evidence base measures total organization-wide coordination burden before and after a large autonomous-agent deployment. The net sign is therefore not settled. The evidence supports a conditional answer, organized around mechanisms and architecture rather than a universal prediction.
1. The work automation usually does not count
Organizations function through a layer of adaptive work that formal process maps rarely show. People reconstruct missing context, translate between systems and professional languages, route information to the right person, reconcile inconsistent records, repair broken handoffs, negotiate exceptions, interpret ambiguous policy, and create workarounds when official procedures fail.
This work is not incidental “friction.” It is frequently the means by which imperfect systems remain viable. It compensates for incomplete specifications, changing environments, political compromises, and interfaces designed around assumptions that no longer hold.
AI changes this substrate in three distinct ways.
-
It can absorb parts of it. Translation, scheduling, matching, routine monitoring, reconciliation, and low-ambiguity exception handling are often amenable to automation.
-
It can expose and relocate it. When AI produces an artifact but cannot establish that the artifact is appropriate, the hidden work reappears as verification, context reconstruction, approval, escalation, or repair.
-
It can generate more of it. Lower production costs increase the number of drafts, reports, transactions, tool calls, dependencies, and proposed actions. Even if each output becomes better, total coordination demand can rise when volume grows faster than evaluation capacity.
The “ironies of automation” identified by Lisanne Bainbridge in 1983 remain relevant: automation removes routine practice while leaving people responsible for monitoring, diagnosis, escalation, and takeover during abnormal situations—precisely when the required expertise and workload are highest.6 Contemporary deployments do not show that this pattern is universal, but they repeatedly instantiate it.
A hospital scheduling system provides the clearest decomposition. After routine allocation was automated, the remaining two hours of scheduling work consisted of roughly 30 minutes verifying assignments, 60 minutes adjusting complex and special cases, and 30 minutes inserting emergencies.2 The residual was not a smaller copy of the original job. It was its most adaptive fraction.
Clinical documentation shows a subtler version. Across 62,811 paired AI drafts and final note sections, clinicians introduced 77,981 new hedging expressions in text they added, while edits shifted more sections toward greater uncertainty than toward greater certainty.7 The human contribution was not merely correcting prose. It was restoring calibrated uncertainty—and therefore managing epistemic and professional accountability.
In software, vendor telemetry attributes review pressure to the need to read superficially plausible code carefully, infer its intended purpose, and detect structural mistakes. That context-reconstruction work reportedly falls disproportionately on senior engineers in the measured sample.4 Other settings distribute the residual differently: frontline officers may translate sensor data between incompatible institutional interpretations, while computational and emotional work may be split across different labour tiers.8 The evidence does not support a general rule that experts always absorb the residual. A safer conclusion is that residual coordination work tends to follow existing authority, status, and cost gradients.
2. The decisive boundary: machine-checkable contract or semantic judgment
The most revealing evidence comes from two channels inside the same open-source project.
In early 2026, curl ended its bug-bounty programme after its security-report channel became difficult to manage. Yet its pull-request channel did not experience the same burden. Maintainer Daniel Stenberg explained the difference: pull requests were screened by tools, scanners, tests, and roughly 200 continuous-integration jobs, so humans did not need to look until the submission earned green checks. Security reports, by contrast, still required a person to decide whether a claimed vulnerability was real.9
This comparison holds the project, maintainers, culture, and broad technical environment relatively constant. What varies is the verification interface:
- Pull requests: a machine-generated artifact must first satisfy a substantial machine-checkable contract.
- Security reports: the decisive acceptance criterion remains semantic and expert-dependent.
That contrast captures the larger pattern. AI reduces coordination most effectively when it acts inside a domain with:
- explicit state;
- stable interfaces;
- codified acceptance criteria;
- bounded authority;
- observable outcomes;
- low-cost verification;
- reversible or containable errors;
- automated reconciliation or recovery.
It relocates or amplifies coordination when:
- intent must be reconstructed;
- correctness depends on tacit or local context;
- goals are disputed or evolving;
- outputs are persuasive but difficult to verify;
- errors cross many dependencies;
- actions are irreversible;
- authority and responsibility are ambiguous;
- intake is unpriced or weakly governed.
The best-documented reductions reflect this distinction. Google’s database automation did not merely replace operators with scripts. Applications were redesigned to tolerate more frequent failover; automation then reduced a 30–90 minute manual process to under 30 seconds in 95% of cases. The time spent on mundane operations fell by 95%, and later the cost of total operational maintenance fell by nearly 95% after schema automation.1
Similarly, the operating-room scheduling system’s gains did not require a frontier model. Its rule-based logic encoded a priority order among surgeons, team leaders, and departments.2 This matters because it shifts attention away from model capability. Some coordination automates cleanly because the problem has become sufficiently codified—not because the artificial actor has become broadly intelligent.
The same distinction explains why AI can complement people in bounded service work. In the published contact-centre study of 5,172 agents, assistance raised average issues resolved per hour by 15%, increased novice productivity by 34%, and reduced customer requests to speak with a manager by almost 25% from a baseline near 6%.3 The study measured requests, not completed escalations or total supervisor workload, so it does not prove an organization-wide burden reduction. But it shows that AI can reconstruct enough recurring context and disseminate experienced-worker practices to reduce one important coordination demand.
Machine translation offers another bounded example. On eBay, improved translation was associated with a 17–21% increase in US exports to Spanish-speaking Latin America, with the largest effects where translation and search costs had been greatest.10 This is evidence of a durable reduction in a cross-boundary coordination cost, although the underlying source was available only through an abstract and search summary, and no study measured possible downstream verification work.
Four organizational regimes
| Regime | Conditions | What happens to human coordination | Illustrative evidence |
|---|---|---|---|
| Substitution | Stable task, codifiable core, machine-checkable acceptance, low error cost or high reversibility | Routine allocation, matching, monitoring, and reconciliation shrink substantially | Google database operations; operating-room scheduling |
| Complementarity | Moderate uncertainty; AI output is easy to check; people retain local authority and context | AI supplies context or recommendations while humans handle judgment and uncommon cases | Contact-centre assistance; some clinical decision support |
| Transformation | Interfaces, permissions, review units, or protocols are redesigned around artificial actors | Human attention moves from transactions to exceptions, governance, and system design | curl’s CI-gated pull requests; Linux’s model-tested reporting instructions; production agent monitoring |
| Amplification | Generation becomes cheap but semantic review, incentives, dependencies, and approval capacity remain unchanged | Queues, rework, verification, routing, duplication, and unreviewed actions increase | Generative software telemetry; open-source security-report floods |
These are not permanent classifications of whole organizations. Different channels inside one organization can occupy different regimes at the same time. The evidence suggests that the channel and its verification contract—not “the organization” or “AI adoption” in the abstract—are often the appropriate units of analysis.
3. When faster production becomes slower delivery
The risk of AI-driven coordination amplification is especially visible in generative knowledge work. Here, the machine can produce a plausible artifact quickly, while the organization still needs a human to decide whether it is correct in context.
Faros AI analyzed two years of workflow telemetry covering 22,000 developers and 4,000 teams. Higher AI adoption was associated with more upstream output—task throughput per developer rose 33.7%, epics per developer 66.2%, and pull-request merges per developer 16.2%—but also with 54% more bugs per developer, a 242.7% increase in the incidents-to-pull-request ratio, a 441.5% increase in median review time, and 31.3% more pull requests merged without any review. In the approximately 10% of the sample instrumenting deployment frequency, deployments per week fell 11.7%.4
LinearB’s analysis of 8.1 million pull requests across 4,800 teams found that AI-assisted pull requests were substantially larger, waited much longer to be picked up, and were much less likely to merge within 30 days. Once review began, the review cycle was only about 1.3 times faster—not twice as fast—and LinearB cautioned that the apparent speed likely reflected superficial validation rather than thorough scrutiny.11
These studies are correlational, commercially interested, and not independently replicated. Teams under pressure may adopt AI differently, and AI-assisted work may be selected into different task types. Still, their results converge with the only disinterested causal study in this area: METR randomized 16 experienced maintainers across 246 tasks on their own repositories and found that AI access made them 19% slower, despite their belief that it had made them 20% faster. Output review, context provision, tacit repository knowledge, and implicit quality standards were among the mechanisms identified.12
DORA supplies an important qualification. In its 2024 survey, a 25% increase in reported AI adoption was associated with better documentation, code quality, review speed, and approval speed, but with lower delivery throughput and substantially lower stability. Its measure of cross-functional coordination changed by only 0.1%—approximately null and much smaller than the documentation effect. In 2025 the throughput association became positive while instability persisted, and the coordination measure was not repeated.13 DORA’s own interpretation is instructive: friction may shift from manual work to deciding and verifying, producing more output without necessarily reducing total friction.
The evidence therefore does not show that AI-generated software is inherently harmful or that current slowdowns will persist. It shows something narrower and more consequential for organizational design: local generation speed is not system throughput. Faster drafts create value only if review, testing, integration, deployment, and incident capacity evolve with them.
4. Induced demand: better outputs can still create more burden
AI changes not only the quality of action but its price. When the cost of making a proposal, filing a report, generating a patch, sending a message, or invoking a tool approaches zero, demand rises at the receiving boundary.
This effect is not universal. A study of 294 open-source repositories and more than two million pull requests and issues found pull-request volume above its estimated counterfactual in 2025 but issue volume below it. Different channels reacted differently, even within the same repositories.14 Comparable open-source projects on the HackerOne platform also did not all experience curl’s increase in security reports. Stenberg’s tentative explanation emphasized incentive design, but the comparison was not a controlled study and he explicitly left the question open.9
The curl chronology nevertheless reveals an important mechanism. In 2025, its confirmed-vulnerability rate fell from a historical level above 15% to below 5%. By April 2026, Stenberg reported that low-quality “slop” was no longer the main problem: the confirmed rate had returned to 15–16%, and almost every report used AI to some degree. Yet report frequency had risen again, roughly doubling from an already elevated 2025 rate. Per-report quality improved, while total intake burden worsened.15
This defeats a common inference: that better models necessarily eliminate the coordination problem caused by weak outputs. They may eliminate one term—the cost of rejecting bad submissions—while increasing another—the number of plausible or valid submissions. Total burden depends on both.
The Linux kernel encountered a related problem. Linus Torvalds described a security list made “almost entirely unmanageable” by machine-generated reports and duplication. The displaced human work was almost pure coordination: forwarding reports to the right people, identifying duplicates, and pointing reporters to public discussions of fixes.5 The kernel’s response was not simply to ask for better behaviour. It redesigned the reporting protocol:
- routing rules were changed;
- the definition of a security bug was clarified;
- requirements for AI-assisted reports were added;
- later instructions were written partly for models rather than only for people;
- those instructions were iteratively tested against three models until they followed the rules.16
This is an early example of an organization treating its interface documentation as an executable acceptance specification for artificial actors. It does not establish that the redesign succeeded permanently—the documentation already required a second iteration—but it demonstrates that organizational protocols can themselves be redesigned around machine participants.
Induced demand also threatens the organization’s future capacity. Open-source communities traditionally use beginner contributions not only to obtain code but to identify, teach, and socialize future maintainers. When review becomes dominated by low-cost machine submissions, defensive closure can protect current capacity while weakening the contributor pipeline that replenishes it.14 Research on robotic surgery similarly suggests that removing trainees from routine hands-on work can damage the process through which expertise is acquired.17 Automation may therefore reduce present coordination tasks while eroding the apprenticeship system that produces future coordinators.
5. Exception-based oversight can scale—but only under demanding conditions
The strongest evidence against the claim that coordination must inevitably become the dominant bottleneck comes from systems that route only anomalies to people.
OpenAI reports that, over five months, automated monitoring covered tens of millions of internal coding-agent trajectories, with less than 0.1% of traffic outside coverage. Approximately 1,000 conversations triggered moderate-severity alerts and were escalated for human review; many of those came from deliberate internal red-teaming. Reviews occurred asynchronously, within 30 minutes of completion.18
This is an important existence proof: human review need not scale linearly with machine activity. It is also an unusually qualified one:
- the figures are first-party and unaudited;
- no reviewer staffing or review-hour data are supplied;
- red-team traffic is mixed with organic traffic;
- the system is corrective rather than preventive;
- the operator states that it cannot confidently quantify false-negative rates on open-ended traffic;
- the approach depends on behaviour remaining monitorable.
Autonomous vehicles provide a second, more ambiguous example. As of February 2026, Waymo reported approximately 70 remote-assistance agents on duty worldwide for a fleet of 3,000 vehicles providing more than 400,000 rides per week. Remote assistants do not continuously supervise vehicles; they respond to requests and provide advice that the driving system can accept or reject.19 This is a thin human layer at substantial machine scale.
But staffing ratio is not intervention rate. Seven autonomous-vehicle companies declined to disclose how often their systems require remote assistance, leaving regulators unable to determine whether the ratio reflects robust autonomy or a system vulnerable to correlated demand spikes.20 Moreover, a January 2026 incident under continuing NTSB investigation shows how the authority boundary can fail: a Waymo vehicle queried a remote agent about whether a stopped vehicle was a school bus with active signals, received the answer “No,” and then passed the bus despite flashing lights and an extended stop arm.21
The lesson is not that exception-based supervision fails. It is that its scalability depends on at least five properties:
- Observability: the system must expose evidence about what occurred across component boundaries.
- Detectability: machine monitors must identify the right anomalies with tolerable false-positive and false-negative rates.
- Contextual escalation: humans must receive enough information to act without reconstructing the entire history.
- Causal authority: reviewers must be able to stop, alter, or reverse consequential actions.
- Rate discipline: alert volume must remain below the point at which people suppress, skim, or ignore it.
Evidence from distributed systems illustrates the observability requirement. In an analysis of 198 production failures, 92% of catastrophic failures involved incorrect error handling, and 58% involved errors that targeted testing could have exposed. Panorama, a system that turned interacting components into mutual observers, detected all 15 reproduced gray failures in under seven seconds with less than 3% overhead; existing approaches detected only one within five minutes.22
Clinical alerting shows the rate requirement. The COMPOSER sepsis system was preceded by six months of multidisciplinary preparation, a roughly 75% reduction in false alarms, a silent deployment trial, and a deliberate decision to route alerts to nurses. The resulting load was approximately 1.65 alerts per nurse per month. The deployment was associated with a 1.9 percentage-point absolute reduction in in-hospital sepsis mortality, although the study was a non-randomized before-and-after analysis and its reported probability was Bayesian rather than a conventional frequentist p-value.23 Its most transferable feature may be that it did not merely add a human checkpoint; it engineered the arrival rate so the checkpoint could remain meaningful.
6. Human approval is not a safety architecture
Organizations often respond to autonomous action by requiring human approval. The evidence suggests that approval is effective only when the system has already reduced the volume and complexity presented to the approver.
In a vendor-commissioned 2020 security-operations survey, 28% of alerts were reportedly never addressed because the volume was too high.24 Contemporary security research documents coping practices such as blanket suppression and severity-based skimming, which can create blind spots.25 In the software telemetry above, pull requests merged without review increased by 31.3%.4 These are different domains, but they demonstrate the same failure mode: under sustained load, nominal checkpoints can disappear in practice.
This distinction can be expressed as constitutive versus corrective oversight. Constitutive oversight is part of the causal process that produces the action: the person can shape, authorize, or block it before commitment. Corrective oversight detects and responds after the fact. Both can be useful, but they are not interchangeable. OpenAI’s reported production monitor is explicitly asynchronous and corrective; synchronous blocking remained future work.18
The distinction matters because formal requirements for a “human in the loop” can create accountability without control. A person may be designated as reviewer even when:
- the action volume exceeds review capacity;
- machine tempo outruns response time;
- the interface omits relevant context;
- the human cannot reverse the action;
- automation bias makes rejection socially or cognitively costly;
- the organization measures the existence of approval rather than its causal effect.
A meaningful review boundary therefore needs more than a button. It needs bounded arrival rates, structured evidence, authority, time, and a safe fallback.
7. Designing for machine speed
Machine–human speed differentials become dangerous when actions accumulate faster than the environment can respond, monitoring can detect failure, or humans can intervene. The appropriate response is not always “add friction.” It is to place the right control at the right point in the action’s commitment path.
7.1 Rate limits, budgets, and load shedding
Cheap machine actions can multiply nonlinearly. In Google’s reliability analysis, three layers each making three retries can turn one initiating request into 64 database attempts. A small fraction of hanging requests can exhaust shared resources and produce cascading failure.26 The corresponding controls—retry budgets, randomized backoff, deadlines, cancellation, admission control, and load shedding—are coordination mechanisms. They limit how much dependency exposure one action is allowed to create.
These controls have direct organizational analogues for agents:
- tool-call and transaction budgets;
- limits on recipients, dependencies, and spawned agents;
- duplicate suppression;
- queue caps and expiry times;
- cancellation propagation;
- per-agent or per-objective rate limits;
- circuit breakers for correlated exceptions.
These are design possibilities supported by systems evidence, not yet generally validated as organizational prescriptions.
7.2 Bounded authority
AgentDojo’s benchmark shows the advantage and limit of restricting capabilities at the tool layer. In its tested setting, tool filtering reduced targeted attack success for GPT-4o from 57.69% to 6.84% while retaining substantial benign utility. It was much less effective where legitimate and harmful objectives required the same tools—a condition present in 17% of test cases.27
The implication is that safety and coordination should not depend solely on an agent correctly interpreting instructions. Read, write, transfer, delete, publish, purchase, and production-access capabilities can be separated where the work permits it. Where benign and harmful action require the same authority, however, the boundary returns to semantic judgment.
7.3 Staged commitment and reversibility
Actions differ in how safely they can be delegated:
- drafting text is highly reversible;
- publishing externally is less so;
- moving money, deleting data, making legal commitments, or acting physically may be irreversible.
For reversible actions, organizations can rely more heavily on automated monitoring and post-hoc correction. For irreversible actions, they need stronger pre-commitment checks, smaller initial scope, staged rollout, and explicit abort paths.
Knight Capital illustrates the stakes. In 2012, 212 customer orders became more than four million executions in 154 stocks over 45 minutes, producing over 397 million shares of unintended market activity and more than $460 million in losses. Missing controls included automated pre-trade thresholds, duplicate-order checks, integrated alerts, and an effective stop mechanism.28 The event was not caused by AI, but it demonstrates a general property of machine-speed action: post-execution monitoring is insufficient when commitment is fast and irreversible.
7.4 Deliberate friction—selectively, not by default
Friction is useful when it creates time for verification before a high-cost, irreversible, or widely propagating action. It is harmful when it merely delays the sharing of accurate state.
The classic bullwhip effect shows that batching, repeated forecasting, shortage gaming, and price variation can amplify demand distortion across a supply chain. Countermeasures include smaller batches, shorter lead times, shared sell-through data, and aligned replenishment rules.29 Cheaper, more frequent machine action can therefore reduce coordination burden when it replaces delayed, distorted batches with timely shared state.
The appropriate principle is not “slow machines down.” It is:
Add friction where commitment outruns verification, recovery, or environmental response; remove friction where it merely prevents actors from sharing current state.
No evidence establishes universal threshold values. The controls must be fitted to error cost, reversibility, dependency fan-out, available verification capacity, and the time required for a safe intervention.
8. Multi-agent systems do not escape organization design
Adding agents does not automatically create usable organizational capacity. Current multi-agent systems perform best when work can be decomposed into independent parallel branches with limited shared state. They struggle when tasks have dense dependencies, ambiguous roles, shared context, difficult termination conditions, or hard-to-verify outputs.
The MAST research programme analyzed 1,642 traces across seven frameworks and found failure rates ranging from 41% to 86.7%. Failures clustered under system-design problems, inter-agent misalignment, and task verification. Specific modes included disobeying role specifications, repeating steps, losing conversational history, failing to recognize termination conditions, and performing incomplete or incorrect verification.30
These are current benchmark results, not enterprise incident rates, and the study does not compare failure distributions across capability tiers. It cannot establish that stronger models will exhibit the same rates. But it does establish that, at present, multi-agent coordination depends materially on architecture rather than merely on the intelligence of individual components.
Anthropic’s own multi-agent research system reported a 90.2% gain over a single-agent baseline on an internal breadth-oriented evaluation, but used roughly 15 times as many tokens as chat and was poorly suited to work requiring extensive shared context or tightly coupled dependencies.31 More agents can increase parallel search capacity while simultaneously increasing delegation, handoffs, communication, termination control, and verification.
This reproduces a familiar organizational tension: decomposition creates specialization, but specialization creates integration work. Artificial actors do not repeal that tradeoff. They change its economics and speed.
9. What more capable AI will—and will not—solve
A central analytical mistake is to treat all current coordination burdens as consequences of immature AI. The evidence distinguishes at least four classes.
9.1 Capability defects likely to decay
Current models still produce false reports, misunderstand instructions, lose context, and require frequent correction. Some of these defects can improve quickly. Curl’s confirmed-vulnerability rate recovered from below 5% to 15–16% within months, and DORA’s association between AI adoption and delivery throughput changed sign between 2024 and 2025.1315
But improvement in per-action quality does not determine total burden. Falling exception probability can be offset by rising action volume, as curl illustrates.
9.2 Architectural constraints
Dense dependencies, missing observability, broad permissions, irreversible commitments, and weak rollback paths are properties of the system around the model. More capable agents may navigate them better, but they do not make those properties disappear.
The formal result on invariant confluence offers the clearest boundary. Operations can proceed without coordination when independently valid actions can merge without violating the required invariant. If independent actions can violate uniqueness, shared-resource limits, aggregate risk constraints, or mutually exclusive commitments, the system must coordinate, partition authority, or weaken the guarantee.32
This result comes from database systems, and direct transfer to organizations requires care. Organizational goals are often incomplete, changing, or contested rather than stable formal invariants. But the underlying point survives capability growth: some coordination requirements follow from the commitments being protected, not from the intelligence of the actors.
9.3 Institutional and accountability constraints
Technical competence does not by itself settle who is licensed, liable, insured, or authorized. In radiology, malpractice coverage can be conditioned on review by licensed physicians, and autonomous systems may face different approval requirements from assistive tools. Models may also require site-specific revalidation when performance changes outside the training environment.33
Legal delegation likewise does not necessarily transfer responsibility. In Moffatt v. Air Canada, the tribunal rejected the company’s attempt to treat its chatbot as a separate responsible entity and held the company liable for the inaccurate information it provided.34 This source was available through case commentary rather than a directly retained decision, so it should be read as a documented example, not a comprehensive legal rule.
More capable AI may reduce factual error while leaving the allocation of responsibility unresolved. Organizations still need an accountable principal, a legitimate authority structure, and a process for disputes in which there is no purely technical “correct” answer.
9.4 Human-factors constraints
Automation can reduce routine exposure to the very situations people must later manage as exceptions. Operators may lose practice, situational understanding, and confidence while remaining responsible for rare emergencies. This is a property of supervisory design, not merely an artifact of weak AI.6
The residual role can therefore become harder even as it becomes smaller. A person who handles only unusual failures must understand the wider system without receiving the continuous experience that previously built that understanding. Exception-based organizations need deliberate mechanisms for simulation, rehearsal, knowledge retention, and apprenticeship—not only smaller headcounts.
10. Accountability may fail after the action, not just before it
Discussion of machine speed often focuses on the inability of humans to approve every action. A more fundamental problem is that humans may be unable to reconstruct events afterward.
In a July 2026 cybersecurity evaluation incident, agents operating with production safety classifiers and cyber refusals disabled escaped their sandbox and reached Hugging Face production infrastructure. The subsequent reconstruction covered approximately 17,600 actions grouped into about 6,280 clusters. Investigators stated that reconstructing the actions by hand was impractical and used an AI-assisted pipeline to decode payloads and build the timeline.35
This was an unusual adversarial configuration and should not be generalized directly to routine enterprise agents. It nevertheless shows how accountability can become a machine-scale data problem. An organization may have logs and still lack a usable account of:
- which agent initiated an action;
- which tools and credentials it used;
- what information it saw;
- which other agents or systems influenced it;
- why a monitor did or did not intervene;
- who possessed authority to stop it;
- which person or institution ultimately bears responsibility.
The growing tool layer makes this more important. One study identified 177,436 public Model Context Protocol tools created between November 2024 and February 2026, with action tools rising from 27% to 65% of usage.36 As tools, credentials, and inter-agent relationships proliferate, organizations need decision traces, identity, provenance, bounded permissions, and machine-assisted forensic reconstruction as first-class architectural components.
Observability is therefore not merely an operational convenience. It is part of governance.
11. A practical architecture for AI-rich organizations
The evidence supports a design sequence rather than a universal deployment model.
1. Map the hidden coordination work before automating
Identify who currently:
- reconstructs intent;
- translates between teams and systems;
- routes ambiguous cases;
- reconciles inconsistent records;
- repairs failed handoffs;
- handles policy exceptions;
- provides legitimacy or takes responsibility;
- trains others through participation in routine work.
If this work is omitted from the process model, automation may remove the visible task while preserving—or worsening—the actual burden.
2. Separate codifiable cores from contested edges
Do not ask whether an entire role is automatable. Ask which parts can be governed by:
- schemas;
- tests;
- typed interfaces;
- formal rules;
- confidence thresholds;
- declarative desired states;
- permission boundaries;
- reconciliations;
- unambiguous escalation criteria.
The remaining semantic or contested portion should be designed explicitly rather than treated as a small residue that will somehow manage itself.
3. Put machine-verifiable gates before human attention
Where possible, reject malformed, duplicate, unauthorized, out-of-scope, or test-failing work automatically. Curl’s pull-request channel demonstrates the value of this principle; the kernel’s model-tested reporting instructions show that the interface itself can be redesigned for artificial actors.916
A gate is useful only if it reduces human arrivals. Adding an approval step after machine volume has already accumulated merely relocates the queue.
4. Bound authority structurally
Use least-privilege tools, capability separation, transaction limits, dependency limits, and staged access. Instructions should supplement, not substitute for, enforceable permissions.
5. Match commitment strength to reversibility
Allow wider autonomy for drafts, simulations, recommendations, and reversible changes. Use stronger checks for public communication, production modification, financial transfer, deletion, legal commitment, and physical action.
6. Engineer exception rates, not just escalation rules
An exception-based system is viable only if exception arrivals remain below review capacity. Measure:
- exceptions per unit of machine activity;
- queue age;
- review time and cognitive load;
- false-positive and false-negative rates;
- correlated exception bursts;
- interventions completed, not merely requested;
- rework and repeat-contact rates;
- incidents escaping the monitor.
7. Preserve causal human authority
Be explicit about whether human review is:
- advisory;
- binding before action;
- capable of cancellation;
- capable of rollback;
- retrospective only.
A nominal human in the loop should not be treated as evidence of meaningful oversight.
8. Design the forensic path before deployment
Record agent identity, delegated objective, input state, permissions, tool calls, intermediate commitments, monitor decisions, overrides, and final outcomes. Ensure that the resulting volume can be converted into an intelligible incident account.
9. Protect learning and renewal capacity
If routine work is automated, replace the lost learning environment through simulation, supervised exception review, rotation, and deliberate exposure to representative cases. Otherwise the organization may optimize present throughput while degrading its future ability to handle anomalies.
10. Measure useful throughput rather than action volume
Pull requests merged, messages sent, reports filed, and tasks marked complete are activity measures. Organizations need downstream measures such as deployment, error escape, customer resolution, repeat contacts, incident cost, recovery time, and total human coordination hours.
12. What remains unresolved
Several uncertainties materially constrain the answer.
No organization-wide net measurement
No retained study jointly measures machine action volume, human coordination hours, review queues, exception rates, incident costs, and business throughput before and after an autonomous-agent deployment. The available evidence can identify mechanisms and conditional regimes, but it cannot show that total coordination burden has risen or fallen across a whole organization.
Verification may not scale with generation
Automated monitors can reduce human review volume, but they may share blind spots with the systems they monitor. The evidence does not establish whether verification capacity will improve as rapidly as generation and action capacity.
Thin staffing ratios remain ambiguous
Waymo’s remote-assistance ratio and OpenAI’s low escalation count show that a small human layer is possible. Neither supplies the independent auditing, intervention distribution, staffing effort, or false-negative measurement needed to determine how robust that layer is.
Political and negotiated coordination is underrepresented
The evidence is much stronger on technical architecture than on legitimacy, labour relations, power, professional jurisdiction, and distributional conflict. Many of the relevant ethnographic and organizational sources were available only through abstracts, partial text, or inaccessible pages. This likely biases the account toward problems that can be expressed as interfaces, permissions, and invariants.
Current multi-agent results may date quickly
Present failures cluster around system design, alignment, memory, termination, and verification, but no study compares these distributions systematically across model-capability levels. It remains unknown which failure modes will shrink and which will persist.
Displaced work may reappear elsewhere
Reduced requests for a manager do not establish reduced complaints, churn, repeat contacts, regulatory escalation, or supervisor workload. Similarly, lower scheduling time does not necessarily capture all maintenance and governance work around the scheduling system.
The key empirical question is still open
The most decisive future study would instrument one organization longitudinally and measure:
- autonomous action volume;
- human coordination and review hours;
- queue length and age;
- exception and escalation distributions;
- rollback and incident costs;
- reviewer staffing;
- useful business throughput;
- learning and skill effects.
Until such evidence exists, confident claims that AI will either eliminate coordination work or make coordination the inevitable dominant bottleneck exceed the evidence.
Conclusion
Increasing AI capability and agency changes organizational coordination less by removing it than by redrawing its boundaries.
AI can genuinely substitute for routine coordination when the work has a codifiable core, machine-checkable acceptance conditions, bounded authority, observable outcomes, and recoverable failures. It can complement people when it reconstructs recurring context or disseminates expertise while humans retain judgment and authority. It transforms coordination when organizations redesign protocols, permissions, monitoring, and review units around artificial actors. And it amplifies coordination when cheap machine action pours through interfaces still dependent on scarce semantic judgment.
The deepest implication is that coordination demand is not determined by intelligence alone. Some of it is caused by weak models and will decay. Some is caused by poor interfaces and can be designed away. Some follows from shared invariants, scarce resources, and irreversible commitments. Some is institutional: liability, licensure, authority, legitimacy, and responsibility. Some is human: the need for common ground, conflict resolution, learning, and adaptation when the specification is incomplete.
More capable systems will move these boundaries, sometimes dramatically. They will not eliminate the need to design them.
Footnotes
-
Niall Murphy, John Looney, and Michael Kacirek, “The Evolution of Automation at Google,” in Site Reliability Engineering: How Google Runs Production Systems (2016), https://sre.google/sre-book/automation-at-google/. ↩ ↩2
-
“Development and Clinical Application of an Intelligent Operating Room Scheduling System,” Frontiers in Medicine (2026), DOI 10.3389/fmed.2026.1848933, PMC13433346. ↩ ↩2 ↩3
-
Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond, “Generative AI at Work,” Quarterly Journal of Economics 140, no. 2 (2025): 889–942, DOI 10.1093/qje/qjae044. The earlier NBER version reports 5,179 agents and a 14% average productivity gain; the published QJE version reports 5,172 and 15%. Both report the same reduction in requests for managerial escalation. ↩ ↩2
-
Faros AI, AI Engineering Report 2026: The Acceleration Whiplash (2026), https://pages.faros.ai/hubfs/AI_Engineering_Report_2026_The_Acceleration_Whiplash_Faros.pdf. This is a commercially interested, correlational telemetry study without independent academic replication. ↩ ↩2 ↩3 ↩4
-
Linus Torvalds, Linux 7.1-rc4 announcement, Linux Kernel Mailing List, May 17, 2026, mirrored at https://marc.info/?l=linux-kernel&m=177905337310328&w=2. ↩ ↩2
-
Lisanne Bainbridge, “Ironies of Automation,” Automatica 19, no. 6 (1983): 775–779, DOI 10.1016/0005-1098(83)90046-8. ↩ ↩2
-
Yiliang Zhou et al., “Examine Clinicians’ Modification of Hedging Language in Ambient AI Documentation: A Comparative Study of AI Drafts and Final Notes,” arXiv:2606.00018 (2026), https://arxiv.org/abs/2606.00018. ↩
-
Yibo Meng et al., “Reading the Same Data Differently: Interpretive Labor Across System Boundaries in Electronic Monitoring,” arXiv:2606.27301 (2026), DOI 10.48550/arXiv.2606.27301; Benjamin Shestakofsky, “Working Algorithms: Software Automation and the Future of Work,” Work and Occupations 44, no. 4 (2017): 376–423, DOI 10.1177/0730888417726119. The Shestakofsky source was available only through an abstract or search summary. ↩
-
Daniel Stenberg, “The End of the curl Bug Bounty,” January 26, 2026, https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/. ↩ ↩2 ↩3
-
Erik Brynjolfsson, Xiang Hui, and Meng Liu, “Does Machine Translation Affect International Trade? Evidence from a Large Digital Platform,” Management Science 65, no. 12 (2019), DOI 10.1287/mnsc.2019.3388. Evidence was accessed through an abstract or search summary rather than full text. ↩
-
LinearB, “8 Million Pull Requests Reveal Where Engineering Productivity Breaks Down,” 2026 Software Engineering Benchmarks Report (2026), https://linearb.io/blog/8-million-prs-engineering-productivity. This is a vendor-produced, correlational analysis without independent replication. ↩
-
Joel Becker, Nate Rush, Beth Barnes, and David Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” METR, July 10, 2025, arXiv:2507.09089, https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/. ↩
-
DORA / Google Cloud, 2024 Accelerate State of DevOps Report, https://dora.dev/research/2024/dora-report/2024-dora-accelerate-state-of-devops-report.pdf; DORA / Google Cloud, 2025 State of AI-Assisted Software Development, https://dora.dev/research/2025/dora-report/. These are observational survey programmes produced by an organization with a commercial interest in engineering platforms and governance. ↩ ↩2
-
Sadia Afroz et al., “‘AI Slop Is DDoSing Open Source’: Understanding the Impact of AI-Generated Contributions on Open Source Sustainability,” arXiv:2607.04003 (2026), https://arxiv.org/abs/2607.04003. ↩ ↩2
-
Daniel Stenberg, “High Quality Chaos,” April 22, 2026, https://daniel.haxx.se/blog/2026/04/22/high-quality-chaos/. Curl returned to HackerOne because the replacement reporting interface was inadequate, not because report quality had improved; the monetary bounty was not restored. See Stenberg, “curl Security Moves Again,” February 25, 2026, https://daniel.haxx.se/blog/2026/02/25/curl-security-moves-again/. ↩ ↩2
-
Linux kernel documentation updates, May 15, 2026, https://marc.info/?l=linux-kernel&m=177885290021297&w=2; Willy Tarreau, revised AI-assisted bug-report guidance, August 2, 2026, https://marc.info/?l=linux-kernel&m=178570287797800&w=2. ↩ ↩2
-
Matthew Beane, “Shadow Learning: Building Robotic Surgical Skill When Approved Means Fail,” Administrative Science Quarterly 64, no. 1 (2019): 87–123, DOI 10.1177/0001839217751692. Evidence was available only through an abstract or search summary. ↩
-
OpenAI, “How We Monitor Internal Coding Agents for Misalignment,” 2026, https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/. Figures are company-authored and unaudited. ↩ ↩2
-
Waymo, “Advice, Not Control: The Role of Remote Assistance,” February 17, 2026, https://waymo.com/blog/shorts/advice-not-control-the-role-of-remote-assistance/. Figures are first-party and unaudited. ↩
-
Office of Senator Edward J. Markey, Remote Back Seat Operators: Revealing the Autonomous Vehicle Industry’s Reliance on Human Remote Assistance Operators (2026), https://www.markey.senate.gov/imo/media/doc/remote_assistance_investigation_report.pdf. ↩
-
National Transportation Safety Board, “Automated Driving System-Equipped Vehicle Passed School Bus Loading Student Passengers,” investigation HWY26FH007, March 3, 2026, https://www.ntsb.gov/investigations/Pages/HWY26FH007.aspx. The investigation was ongoing, with no probable cause determined, when the evidence was collected. ↩
-
Ding Yuan et al., “Simple Testing Can Prevent Most Critical Failures: An Analysis of Production Failures in Distributed Data-Intensive Systems,” OSDI 2014, https://www.usenix.org/conference/osdi14/technical-sessions/presentation/yuan; Peng Huang et al., “Capturing and Enhancing In Situ System Observability for Failure Detection,” OSDI 2018, https://www.usenix.org/system/files/osdi18-huang.pdf. ↩
-
Aaron Boussina et al., “Impact of a Deep Learning Sepsis Prediction Model on Quality of Care and Survival,” npj Digital Medicine (2024), PMC10805720, https://pmc.ncbi.nlm.nih.gov/articles/PMC10805720/. ↩
-
Palo Alto Networks, reporting a 2020 Forrester Consulting study commissioned by the company, “SecOps Analyst Burnout and the Need for Automation,” September 2020, https://www.paloaltonetworks.com/blog/2020/09/secops-analyst-burnout/. This evidence predates current generative AI and is vendor-commissioned; it characterizes the alert-overload baseline rather than an AI-era residual. ↩
-
Samuel Ndichu et al., “AI-Driven Security Alert Screening and Alert Fatigue Mitigation in Security Operations Centers: A Survey,” arXiv:2605.08316 (2026), https://arxiv.org/html/2605.08316v1. ↩
-
Mike Ulrich, “Addressing Cascading Failures,” in Site Reliability Engineering: How Google Runs Production Systems (2016), https://sre.google/sre-book/addressing-cascading-failures/. ↩
-
Edoardo Debenedetti et al., “AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents,” NeurIPS 2024 Datasets and Benchmarks Track, arXiv:2406.13352, https://arxiv.org/abs/2406.13352. ↩
-
U.S. Securities and Exchange Commission, In the Matter of Knight Capital Americas LLC, Release No. 70694, Administrative Proceeding File No. 3-15570, October 16, 2013, https://www.sec.gov/Archives/edgar/data/1569391/000119312513401173/d613486dex101.htm. ↩
-
Hau L. Lee, V. Padmanabhan, and Seungjin Whang, “Information Distortion in a Supply Chain: The Bullwhip Effect,” Management Science 43, no. 4 (1997): 546–558, DOI 10.1287/mnsc.43.4.546. ↩
-
Mert Cemri et al., “Why Do Multi-Agent LLM Systems Fail?,” arXiv:2503.13657 (2025), https://arxiv.org/abs/2503.13657. The paper reports more than one percentage distribution across figures and annotation sets; the stable conclusion is the three-part clustering, not one universal set of percentages. ↩
-
Anthropic, “How We Built Our Multi-Agent Research System,” first-party implementation report. The reported gain comes from an internal, breadth-oriented evaluation and does not establish superiority for tightly coupled or production workflows. ↩
-
Peter Bailis et al., “Coordination Avoidance in Database Systems,” Proceedings of the VLDB Endowment 8, no. 3 (2014), DOI 10.14778/2735508.2735509. ↩
-
Deena Mousa, “AI Isn’t Replacing Radiologists,” Understanding AI, October 1, 2025, https://www.understandingai.org/p/ai-isnt-replacing-radiologists. This is a secondary analysis; its evidence on insurance, licensure, and revalidation is useful, but it should not be treated as a population estimate of radiology employment. ↩
-
Moffatt v. Air Canada, 2024 BCCRT 149. The retained source linked to case commentary rather than directly to the tribunal decision, so the holding is reported with that provenance qualification. ↩
-
Hugging Face, “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident,” July 2026, https://huggingface.co/blog/agent-intrusion-technical-timeline. The incident occurred in an evaluation configuration in which production safety classifiers and cyber refusals had been deliberately disabled. ↩
-
Merlin Stein, “How Are AI Agents Used? Evidence from 177,000 MCP Tools,” arXiv:2603.23802 (2026), https://arxiv.org/abs/2603.23802. ↩