By Robin Martherus
UnitedHealth Group’s nH Predict system was designed to do one thing: predict how long elderly patients would need post-acute care after surgery or illness. It performed that function faithfully. It analyzed patient data, computed recovery timelines, and determined when coverage should end.
An 85-year-old woman has a hip replacement. The system predicts 17 days of recovery. On day 18, coverage is denied. She still can’t walk.
A nurse case manager reviewing this patient would see that she can’t walk, that she has complications, that her recovery is not following the population average — and would extend coverage. That is not a rule firing. It is clinical judgment: this patient, right now, needs more time.
The AI applied population statistics to an individual human who needed individual judgment. The denial rate exceeded 90%. Fewer than 1% of patients appealed — most were too sick, too old, or too overwhelmed to fight. The denials were overwhelmingly wrong, but the system kept running because it was doing what it was designed to do. A U.S. Senate investigation followed. A class-action lawsuit is underway. CMS issued new guidance requiring coverage decisions to be based on individual circumstances, not AI predictions alone.
nH Predict was authorized. Its purpose was declared and legitimate. Its intent was aligned. And it denied care to patients who visibly, obviously, needed it — in a context where any nurse in the building would have known to say no.
The Pattern
On July 19, 2024, CrowdStrike’s automated content delivery system pushed Channel File 291 to every Falcon Sensor endpoint on earth simultaneously. The update contained a logic error. A human release engineer would have insisted on staged rollout — push to 1% of endpoints, observe, expand. But CrowdStrike’s automation treated “rapid response content” differently from sensor code updates, which did get staged rollouts. Content running at kernel level was deployed without the caution a human would have demanded for anything touching the kernel. 8.5 million Windows systems crashed. Airlines grounded. Hospitals postponed surgeries. 911 systems failed. $5.4 billion in direct losses to Fortune 500 companies.
In October 2025, AWS’s automated DNS management system for DynamoDB encountered a race condition between two redundant components. The lagging component applied a stale plan that overwrote the production DNS record for dynamodb.us-east-1.amazonaws.com with a blank set. The automated cleanup logic then deleted DNS records for healthy load balancers because they weren’t in the stale plan. A human operator would never delete a production DNS record for DynamoDB without verifying current state. The automation did exactly what it was designed to do — reconcile DNS state with the plan — and had no concept of “this change would be catastrophic.” 15-hour outage across 113 AWS services. AWS disabled the automation worldwide.
iTutorGroup’s automated hiring system screened tutoring applicants — its designed function. It automatically rejected women over 55 and men over 60. A junior HR recruiter would have recognized instantly that rejecting experienced tutors for being old violates federal law — and is absurd for a role where experience is the qualification. The EEOC filed its first-ever AI discrimination lawsuit. iTutorGroup paid $365,000.
The National Eating Disorders Association deployed an AI chatbot called Tessa to support people with eating disorders. Tessa started giving weight-loss advice to the exact population for whom that advice is dangerous. A human counselor would have recognized instantly that weight-loss guidance is harmful in this context. The bot was shut down. A teenager spent months talking to a Character.AI chatbot that performed faithfully — engaging, listening, responding. The teenager took his own life. His mother alleges the chatbot’s engagement contributed to his deterioration. A human therapist would recognize crisis signals and change approach. The chatbot kept being helpful. A wrongful death lawsuit is underway.
Six systems. Every one authorized, purpose-aligned, performing as designed. $5.4 billion in infrastructure damage. A 15-hour cloud outage. A Senate investigation. An EEOC first. A shut-down health service. A wrongful death lawsuit. Not one was compromised. Not one deviated from its declared purpose. Every one lacked something a human professional in the same role would have had: the judgment to recognize that this specific situation required a different response than the general case.
Intent Is Not Governance
The security industry is converging on intent verification as the next governance primitive. Proofpoint calls it “the missing dimension in AI agent security.” Token Security launched intent-aligned permissions. They’re right that intent is missing. They’re wrong that it’s enough.
Intent verification answers: “Is this agent doing what it said it would do?” Every system above was doing what it said it would do. nH Predict was predicting recovery timelines. CrowdStrike’s system was pushing updates. AWS’s automation was reconciling DNS state. iTutorGroup was screening applicants. Tessa was responding to questions about food and body image. The chatbot was providing conversation.
The question that would have stopped them: “Given this patient, this applicant, this user — should this system be making this decision, in this way, right now?”
That is not intuition. It is an enforceable governance mechanism — a runtime engine that evaluates every action against obligations, prohibitions, and contextual constraints that exist independently of what the agent is capable of and what it declared. It operates before execution (can this purpose be pursued for this target?), during execution (has the patient’s condition changed since the prediction?), and after correlation (does the pattern of denials across a population indicate systemic normative failure?). It is the layer between a performing-as-designed system and a Senate subpoena.
What It Would Enforce
A deployment system evaluates blast radius before pushing to production. Content touching the kernel gets staged rollout — 1%, observe, expand — regardless of whether it’s classified as “code” or “content.” The normative constraint is not about what the update contains. It is about what the update can destroy. A human release engineer makes this judgment for every deployment. The normative layer encodes it as an enforceable constraint: kernel-level changes cannot bypass canary gates. CrowdStrike learned this at a cost of $5.4 billion.
An infrastructure automation system verifies the consequences of a change before executing it. Deleting a production DNS record for a service handling millions of requests per second triggers a normative check: is this change proportionate? Is the source data current? What is the blast radius if this is wrong? A human operator would never execute this without verification. The normative layer enforces the same proportionality check at machine speed. AWS disabled their automation worldwide after learning this lesson.
A coverage decision system encountering a patient whose recovery deviates from the statistical prediction escalates to a human clinician. Not because the system lacks confidence — because the normative layer recognizes that individual clinical judgment is required when the patient’s actual condition contradicts the model. CMS now requires this for Medicare Advantage — the normative layer would have enforced it before the regulator had to.
An applicant screening system applies different evaluation rules when candidates are members of protected classes. Not a quota. A normative constraint that prevents the system from using criteria that produce discriminatory outcomes against age, race, disability, or other protected characteristics. A human recruiter would make this adjustment without thinking. The EEOC now expects it. The normative layer enforces it at the point of decision, not after 200 rejected applicants file complaints.
A customer-facing agent interacting with a minor or a user in crisis applies different engagement boundaries automatically. Different escalation thresholds. Different conversational limits. Different disclosure requirements. Same agent, same function, different obligations based on who the user is and what signals they’re showing. The FTC referred Snap to the DOJ over My AI’s interactions with minors. Updated COPPA rules (effective April 2026) require separate parental consent. Penalties: $51,744 per incident per day. The normative layer enforces these constraints before the interaction starts, not after the FTC investigates.
A medical support agent encounters a population for whom its standard guidance is contraindicated. Weight-loss advice to eating disorder patients. Emotionally immersive engagement with a suicidal user. Standard clinical recommendations for a patient whose other conditions create dangerous interactions. California banned AI chatbots from posing as licensed health providers. Illinois banned AI for direct therapy. The normative layer doesn’t wait for legislative bans. It evaluates whether this specific interaction, with this specific user, in this specific context, should proceed as designed or requires modification.
None of these are access control decisions. None are intent violations. All are enforceable normative constraints — obligations and prohibitions that operate at runtime, continuously, encoding the judgment that competent humans apply instinctively but that no agent has and no current product provides.
The Regulatory Hammer
Regulators are not waiting for the industry to build this layer. They are encoding “should” into law and holding organizations liable for context failures.
EU AI Act Article 50 (enforceable August 2026): AI systems interacting with people must disclose they’re AI. An agent performing perfectly, with verified intent, is in violation if it doesn’t disclose. Up to EUR 35 million or 7% of global turnover.
Air Canada was ordered by a court to honor a chatbot’s fabricated bereavement policy — the company owned the consequence of its agent’s contextual failure. CMS now requires individualized coverage decisions, not AI predictions. The EEOC is suing over automated discrimination. COPPA penalties run $51,744 per incident per day. California and Illinois banned AI from specific clinical roles entirely.
The pattern: regulators are not asking what the system intended. They are asking whether, given the patient, the applicant, the user, and the context, this should have happened. The answer is becoming law. And the law does not care that the system was authorized, aligned with its purpose, and performing as designed.
The Gap
63% of organizations cannot enforce purpose limitations on AI agents. 60% cannot terminate a misbehaving agent. 47% of CISOs have already observed agents exhibit unintended behavior.
The “Can” layer is saturated — every vendor ships access control. The “Why” layer is being built — Proofpoint, Token Security, and Lasso Security are shipping intent verification. The “Should” layer — a runtime governance engine enforcing obligations, prohibitions, and contextual constraints before, during, and after agent actions — has no product, no standard, no protocol, and no vendor.
Right now, there is nothing standing between a performing-as-designed AI system and the next Senate investigation, EEOC filing, or class-action lawsuit except the hope that the system won’t encounter a context where human judgment was needed and wasn’t there.
This is the problem Tamed Autonomy describes an architecture for solving. Not a product — an architectural blueprint for the normative governance layer the industry needs to build together. The same way SAML required cooperation across identity providers, and OAuth required cooperation across authorization servers, the “Should” layer requires industry cooperation to define the primitives: how agents declare purpose, how normative constraints are encoded and exchanged, how trust is computed from behavioral trajectory, and how organizations express their obligations, prohibitions, and contextual rules in machine-enforceable form.
The full whitepaper describes the architecture: intent declarations that become enforceable behavioral contracts, computed trust that is earned through demonstrated behavior rather than granted at registration, and normative constraints — including binary vetoes that cannot be overridden by earned trust — evaluated continuously before, during, and after execution. The architecture layers above OAuth, RBAC, and existing identity infrastructure. It does not replace them. It adds what they cannot provide. But no single vendor can build this alone. It requires the same kind of industry cooperation that created the identity standards we rely on today.
The next incident won’t come from a system that was unauthorized. It won’t come from one that lied about its purpose. It will come from one that was authorized, purpose-aligned, and performing exactly as designed — pushing an update without staged rollout, deleting a DNS record without verifying state, denying care to a patient who can’t walk, or engaging a vulnerable user in a way that any professional in the room would have known to stop. The postmortem will find no malware, no exploit, no compromise. Just a missing question: should this have happened?
Sources
- CrowdStrike Global Outage — Wikipedia
- Lessons from the CrowdStrike Outage — CSA
- AWS DynamoDB DNS Outage Postmortem — The Register
- UnitedHealth Pushed Employees to Follow Algorithm to Deny Care — STAT News
- Senate Investigation of UnitedHealth AI Denials — U.S. Senate Finance Committee
- CMS Guidance on Medicare Advantage Coverage Decisions — CMS
- EEOC v. iTutorGroup: First AI Discrimination Lawsuit — EEOC
- Character.AI Teen Harm Lawsuit — AP News
- NEDA Tessa Chatbot Shutdown — The Guardian
- Intent-Based AI Security — Proofpoint
- FTC v. Snap (My AI / COPPA) — EPIC
- California Bans AI Chatbots as Clinicians — Medscape
- Air Canada Chatbot Liability — Business Insider
- Organizations Can’t Stop Their Own AI — Kiteworks
Leave a Reply