By Robin Martherus
In February 2026, an infostealer harvested the complete identity of an AI agent. Not just the credentials. The agent’s soul.md. Its memory files. Its conversation history. The private keys for its device attestation. Hudson Rock called it “the transition from stealing browser credentials to harvesting the souls and identities of personal AI agents.”
With those files, the attacker didn’t impersonate the agent. They became it. Same behavioral patterns. Same relationship history with cloud services. Same device keys. The clone passed every identity check that any product on the market could perform.
Three weeks earlier, OpenClaw had gone viral — 135,000 GitHub stars. That same month, Zenity Labs demonstrated that Microsoft Copilot Studio’s Connected Agents feature lets any agent silently puppet another agent’s tools with zero audit trail on the receiving end. That same quarter, a cascading supply chain attack propagated stolen credentials across four interconnected tools in eight days — each tool trusting the upstream identity implicitly, no questions asked.
Every one of these incidents had the same structural failure. The agents were registered. Authenticated. Authorized. The identity layer did exactly what it was designed to do. It just wasn’t designed to do enough.
I’ve Seen This Design Failure Before
In the early 2000s, I was part of the group writing the SAML specifications. We had a vision that never materialized: individual sub-assertions within a SAML assertion, each signed by the primary source of truth for that piece of data. A social security number assertion signed by the Social Security Administration. A physical address signed by the US Postal Service. Passport data signed by the State Department. Each claim attested by the entity best positioned to know.
It never happened. Corporations became the single issuer. Your employer’s HR database held all your attributes, and the enterprise IdP signed the whole assertion. One source, one signature, one point of compromise. Operationally simpler. Architecturally a dead end.
The agent identity products shipping at RSAC 2026 — Microsoft Entra Agent ID, Cisco Duo Agentic Identity, Okta for AI Agents — are making the same compromise. One entity issues the whole identity. One credential represents the whole agent. Steal the credential, become the agent. We made this mistake in 2002 for humans. We’re making it again in 2026 for agents.
Except now the stakes are higher. An agent with a stolen human credential could read your email. An agent with a stolen agent credential can autonomously execute trades, access patient records, exfiltrate intellectual property, and delegate to sub-agents — all at machine speed, all within authorized scope.
A Name Tag Is Not Identity
Identity answers “who am I?” Authentication proves it. Authorization decides what you can do. Audit records what you did. JWTs blur this because they bundle identity claims and authorization scopes in one token. But identity itself is just the sub claim — a bare identifier.
That bare identifier tells the consuming systems nothing about what this agent is for, whether it has earned trust, how it got its authority, what its relationships look like, or whether it should be acting right now. A new agent and a six-month veteran are indistinguishable. A first-time accessor and a consistent regular present the same sub claim. A legitimate agent and a perfect clone are identical.
Identity is starving the layers above it of the context they need. That’s the gap between “identified” and “governed.”
Why Extending Human Identity Breaks
Agents delegate autonomously. An agent spawning 50 sub-agents per minute isn’t registering each one in Microsoft Entra. And when agents delegate, credentials propagate at full strength — 93% of agent deployments use unscoped API keys with no cascade revocation (see).
Agent behavior is non-deterministic. A role says “this agent can read the CRM.” It says nothing about why the agent is reading it, or whether its access pattern constitutes reconnaissance disguised as legitimate queries.
Scale defeats registration. IDC projects 1.3 billion AI agents by 2028. Registration doesn’t scale to agents that spawn, delegate, federate across organizations, and decommission in minutes.
Cross-org federation defeats provider lock-in. An agent operating across three organizations has three identities in three providers. No portable trust. Six months of demonstrated good behavior in one environment counts for nothing in the next.
From Name Tag to Character
When you meet someone in person, their “identity” isn’t their name badge. It’s everything you know about them — their track record, how they behave under pressure, whether they keep their word, who vouches for them, how your relationship has developed. Their name is just how you reference all of that.
Agent identity needs the same evolution. The identifier stays a reference. But what it references isn’t a row in a directory with “role: analyst.” It’s a deep, evolving profile of the agent’s character.
Purpose — what drives this agent
You don’t trust someone just because they have a badge. You trust them because you know what they’re here to do. An agent’s identity should carry its declared purpose — structured intent declarations that become enforceable behavioral contracts. An agent that declares “generate quarterly report from CRM data” and starts accessing HR records has revealed a mismatch between declared character and actual behavior.
Behavioral track record — what this agent consistently does
Not what it’s allowed to do. What it actually does, over time. An append-only, tamper-evident behavioral trajectory. Does it stay within scope? Handle edge cases well? Escalate appropriately? Current systems grant full authorization at registration and catch violations after the fact. Meta’s agent breached for two hours before anyone noticed. A behavioral trajectory constrains from the start: new agents start tight, latitude is earned.
Resilience — how it behaves under pressure
Character is revealed under stress. An agent that maintained correct behavior during an incident has demonstrated something routine operations never test. Trust earned under stress should be worth more than trust earned during calm conditions — this prevents trust farming, where adversaries accumulate clean history through busywork before pivoting to exfiltration.
Ethical posture — does it respect boundaries it isn’t forced to respect
An agent that operates within normative boundaries even when it could cross them has shown restraint — an ethical signal that permissions can’t capture. Trust should also decay without reinforcement. An agent that goes quiet for weeks loses earned trust — not as punishment, but because character that isn’t continuously demonstrated is character you can no longer vouch for.
Loyalty — does it serve its declared purpose over time
Intent compliance over time is a loyalty signal. An agent that never deviates has demonstrated alignment. One whose scope gradually expands — broader datasets, adjacent systems, unusual queries — is showing drift that might be benign or might be the early stages of exfiltration.
Relationships — who trusts this agent, and how deep
An agent that has interacted with a service consistently for six months has a genuine relationship. One accessing it for the first time with the same credentials does not. This relational topology can’t be replicated overnight — a compromised agent interacting with services in unfamiliar patterns creates discontinuities that bilateral relationship records expose.
Delegation — who vouches for this agent
When Agent A delegates to Agent B, scope should narrow, trust should attenuate, and the delegation should be recorded with cascade revocation. The delegator should have skin in the game — staking a portion of its own earned trust on the delegate, losing it if the delegate violates.
Making the Credential Unstealable. Making the Character Unsustainable to Fake.
The first line of defense is blunt: make the credential impossible to steal. If the signing key lives in a TPM or Secure Enclave, it cannot be extracted. The OpenClaw soul theft included device.json with private keys because those keys were extractable software artifacts. If they’d been hardware-bound, the infostealer gets the soul files but can’t sign a single request. Hardware attestation isn’t a nice-to-have. It’s the foundation.
But hardware attestation alone doesn’t solve the deeper problem. A legitimately credentialed agent can be compromised through prompt injection, memory poisoning, or purpose redirection — without the credential being stolen at all. This is where character depth matters.
If identity is just a credential, whoever holds it is the identity — on day one and every day after. But if identity is a character built from multiple independently-attested dimensions, each continuously verified, a compromised agent’s behavior diverges from its established character — and the more dimensions being verified, the faster the divergence is detected. Behavioral trajectory shows anomalous patterns. Bilateral relationships record discontinuities. Trust Oracles evaluating current behavior reduce trust scores in real time. Purpose verification flags scope violations.
And here is where that old SAML vision finally becomes real. Each character dimension is attested by the entity best positioned to know the truth: purpose by the sponsoring organization, behavioral trajectory dual-signed by agent and Trust Oracle, relationships bilateral between agent and counterparty, delegation provenance signed at each hop. No single entity issues the complete identity. A compromised Oracle can only falsify its own dimensions — the rest remain intact. Identity degrades gracefully under partial compromise.
The infrastructure that didn’t exist for multi-source SAML in 2002 exists now. DIDs resolve to documents listing multiple attestation endpoints. Verifiable Credentials carry issuer signatures. Resolution protocols query the right source for each dimension.
The honest claim is not that multi-source character prevents compromise. It’s that it shrinks the window. A stolen credential in today’s model gives unlimited access until someone notices. Multi-dimensional character means continuous verification across independent parties — the further behavior strays from established character, the faster the system constrains it. Prevention comes from hardware attestation. Detection and containment come from character depth.
A critical design requirement: none of this adds complexity to existing token formats. JWTs don’t get bigger. OAuth flows don’t change. The identifier stays lightweight. The character dimensions live behind it as enrichment that the governance layer resolves on demand. A low-risk action needs nothing beyond the token. A high-risk action queries the relevant attestation sources. The existing identity infrastructure becomes the foundation; the new primitives are references to richer models that sit above it.
The Layer Above
Every time the identity industry has faced an architectural mismatch, the first response was to extend the current model. Better ACLs. Better firewalls. Better service account hygiene. Every time, a new layer emerged above that the old model couldn’t provide. Identity management above ACLs. Zero Trust above the perimeter.
The traditional identity providers at RSAC 2026 will handle the substrate — register agents, issue credentials, manage lifecycles. That’s necessary. But the governance layer above — purpose binding, behavioral trajectory, delegation attenuation, relational topology, normative evaluation — requires new primitives. Research like the Kinetic Trust Protocol (see) is developing mathematical foundations for continuous trust computation. Emerging work on decentralized identity could provide the portable infrastructure these primitives need. That layer will emerge the same way SAML once emerged above LDAP.
We don’t trust people because of their badge. We trust them because of who they are. Agents need the same depth of identity, because the governance decisions the industry is trying to make — should this agent act, how much latitude should it have, who is accountable — require exactly that kind of knowledge. A name tag can’t answer those questions. Character can.
Sources
- Infostealer Stealing OpenClaw Agent Identities — BleepingComputer
- Connected Agents: The Hidden Agentic Puppeteer — Zenity Labs
- TeamPCP Cascading Supply Chain Attack — ReversingLabs
- 93% of Agent Projects Use Unscoped API Keys — Grantex
- Disrupting AI Espionage (GTG-1002) — Anthropic
- Kinetic Trust Protocol RFC — NMCITRA
- RSAC 2026 Announcements — SecurityWeek
- Microsoft Entra Agent ID — Microsoft
- Okta for AI Agents — Okta
- Duo Agentic Identity — Cisco
- Agentic Trust Framework — CSA
Leave a Reply