Agentic security

I Built the Detector. Then I Built What Defeats It.

By Robin Martherus


I built an AI agent detector. 100 behavioral features. Machine learning ensemble — Random Forest, Gradient Boosting, neural network. 99.37% accuracy on a three-class problem: human, legitimate AI agent, rogue AI agent. Adaptive attack detection that catches swarm coordination, agents probing the detection boundary, and evolutionary strategy mutation across sessions. It works.

Then I built an autonomous AI attack swarm. Nine specialized agents coordinating through a message bus. Two competing language models debating attack hypotheses and testing them live against targets. Evolutionary algorithms breeding strategy genomes that improve every generation. And I realized: my detector wouldn’t catch it.

Not because the detector is bad. Because the assumptions underneath it — and underneath every agent detector in the industry — are about to break.

What the Detector Gets Right

The core insight that makes agent detection work today is simple: AI agents are too precise. The coefficient of variation (CV) in request timing — how much the interval between actions varies — is the primary discriminator. Humans are messy. We pause to read. We get distracted. We speed up and slow down unpredictably. AI agents, even when they add artificial delays, exhibit mechanical consistency. CV below 0.25 is almost certainly an agent. CV above 0.60 is almost certainly human.

Layer 100 features on top — navigation entropy, request method diversity, circadian rhythm patterns, JWT identity analysis, impersonation scoring — and you get a detector that classifies with over 99% accuracy on current-generation agents. Add adaptive attack detection for swarm signatures, agents clustering in the uncertain zone between 0.25 and 0.60, and evolutionary strategy mutation, and you can catch sophisticated multi-agent campaigns.

This is the state of the art. The Cloud Security Alliance reported this week that 68% of organizations cannot clearly distinguish between human and AI agent activity in their systems (see). A detector that solves this is valuable today.

But “today” is a short window.

Three Things That Break It

1. Reasoning agents don’t look like agents

My detector fingerprints 66 automation frameworks — Selenium, Puppeteer, python-requests, LangChain, and dozens more. It catches their behavioral signatures. Traditional bots and script-based agents trip these detections reliably.

My attack swarm doesn’t use any of them. Its debate layer reasons about endpoints through competing LLMs. A Builder model hypothesizes that an admin config endpoint might expose user data. The system makes a single, surgical HTTP request to test the hypothesis. A Breaker model challenges the finding or proposes escalation. Another request. Evidence feeds back. Up to five rounds of reasoned probing, with every request informed by the response to the last one.

The timing is naturally irregular — LLMs “think” for variable durations between requests. The request volume is low. The navigation pattern looks like a developer exploring an API, not a scanner brute-forcing endpoints. The CV falls comfortably in the human range.

This isn’t theoretical. XBOW, the autonomous pentest platform that hit #1 on HackerOne’s U.S. leaderboard in 90 days, handles 85% of custom-built, never-before-seen vulnerabilities. Google’s Big Sleep found a zero-day in SQLite through LLM reasoning (see). These systems don’t match framework signatures because they’re not using frameworks. They’re using judgment.

2. Evolutionary agents adapt to your detection boundary

My attack swarm evolves. Strategy genomes — configurations of tools, parameters, and priorities — are scored on fitness (vulnerability discovery, novelty, efficiency, adaptability). Top performers breed. Mutations are proposed by the debate layer with evidence and reasoning. Each generation is different from the last.

My detector has adaptive attack detection that tracks evolutionary strategy mutation across sessions. But here’s the structural problem: the detector’s ML models define a decision boundary. An evolutionary system will find that boundary. That’s what genetic algorithms do — they explore the fitness landscape until they find the optimum, and in this case the optimum is the point just inside your detection threshold.

76% of detected malware now exhibits AI-driven polymorphism (see). The same evolutionary principle applied to agent behavior means no two sessions need look the same. The malware that hits your endpoint has never existed before. The agent probing your API at 4 PM has different behavioral characteristics than the one at 2 PM.

3. Swarm agents look complementary, not similar

My detector catches swarm coordination by identifying sessions with similar behavioral fingerprints and correlated timing. It also detects metadata leakage — swarm agents that accidentally share generation IDs or strategy identifiers.

But my attack swarm’s nine agents don’t look similar to each other. Recon looks different from Exploitation looks different from Lateral Strategy. They exhibit complementary behavior — their actions assemble into a coherent attack chain when viewed together, but individually each agent looks like a different, unremarkable actor.

The GTG-1002 campaign disclosed by Anthropic in November 2025 worked the same way: ~30 agents across ~30 targets, each conducting a different phase of the attack lifecycle. The coordination is in the campaign logic, not in behavioral similarity. Clustering similar sessions would find nothing.

The Industry’s Approach Isn’t Solving This Either

The agent detection industry is building three categories of solutions, none of which address these problems:

Bot detection evolved. Arkose Labs, DataDome, HUMAN Security — they’ve expanded from bot detection to “agent trust management.” But the core approach is still behavioral classification. DataDome processes 5 trillion signals daily. HUMAN Security analyzes navigation paths and escalation curves. These are better versions of what my detector does — and subject to the same limitations against reasoning agents that pass as human.

Identity registration. Microsoft Entra Agent ID, Cisco Duo Agentic Identity, Okta for AI Agents, CyberArk — they give every agent a registered identity. This solves the cooperative case: legitimate agents that want to be identified can be. But adversarial agents won’t register. And the 80% of agents that DataDome found don’t self-identify at all won’t start because Microsoft ships a registration protocol.

Hardware attestation. Raypher Labs binds agent identity to TPM 2.0 chips. Beyond Identity uses hardware-bound device credentials with proof-of-possession tokens. Yubico and Delinea demonstrated hardware-attested Role Delegation Tokens at RSAC 2026. This is the strongest signal available — provably unforgeable without physical hardware access. But it only works on controlled infrastructure. Nation-state swarms operating from adversary-controlled hardware will not present TPM attestations.

Each of these solves part of the problem. None addresses the core challenge: detecting sophisticated autonomous agents that reason, evolve, and coordinate while deliberately mimicking human behavior.

What I Think Detection Actually Needs

Building both the detector and what defeats it taught me that detection alone is a losing game. Open-world behavioral classification — distinguishing a frontier AI agent from a human based on behavior alone — is eroding. GPT-4 has been judged “more human than humans” in controlled Turing tests (see). A 2025 PNAS study showed autonomous synthetic respondents passing standard attention checks 99.8% of the time (see).

That doesn’t mean all detection is failing. Domain-specific detectors still work in constrained environments — Prolific reported 100% accuracy on its February 2026 authenticity checks (see). Cooperative attestation works for enrolled agents. But generic open-world classification is too brittle to be the architectural hinge for governance decisions.

What I learned: detection has to stop asking “is this an agent?” as the primary governance gate. That question gets harder every month. The questions that remain answerable are different.

“Is this actor doing what it declared?”

If agents are required to declare their purpose — structured intent declarations specifying what they intend to do and why — detection shifts from classification to verification. “Is this consistent with declared scope?” is a more tractable problem than “is this human or agent?” because you’re checking behavior against a concrete contract, not trying to win an adversarial classification game.

It’s not a perfect solution. Intent declarations are still adversarially gameable — if the declaration is broad enough to encompass malicious actions, or if the semantic gap between declaration and behavior is hard to measure, verification weakens. My own analysis of the “Honest Liar” scenario — a malicious actor who declares plausible intent and stays within contract while still causing harm — shows this can’t be fully solved, only made progressively more expensive through continuous behavioral friction.

But it’s a better foundation than classification. My detector already has 15 impersonation features comparing declared identity against observed behavior. Generalizing that from identity mismatch to purpose mismatch is the natural next step. An agent that declares “generate quarterly report from CRM data” and starts accessing HR records is detectable with high confidence regardless of how human-like its timing is.

The human problem

There’s a catch. Humans can’t declare structured intent. They browse, explore, change their mind mid-session. If you require intent declarations from agents but not humans, adversaries just mimic humans to avoid the requirement. If you require them from everyone, you break human workflows.

Two paths through this. For sensitive systems, route access through agent intermediaries that declare intent on behalf of their human principals. When a user asks Copilot to “pull the Q4 numbers,” that natural language request is an intent declaration — it’s just not being treated as one by security infrastructure today. This concentrates access through a governed channel, but it also concentrates the attack surface — the intermediary itself is prompt-injectable and overprivileged. OpenAI’s March 2026 guidance explicitly warns that systems must be designed so manipulation impact stays constrained even when attacks succeed (see).

For general systems where agent mediation isn’t practical, use a layered authorization stack — identity, per-action authorization, trust computation, behavioral monitoring — that evaluates each action in context without requiring actor classification as a gate. This is not a new idea. NIST’s ABAC framework (see), Zero Trust architecture (see), and CSA’s Context-Based Access Control (see) all evaluate actions in context, not just actor identity. The direction is established. What’s missing is the additional structure on top.

“Do these actions form a reasoning chain?”

Reasoning agents evade timing analysis but can’t hide what they’re reasoning about. A sequence like GET /admin/config → GET /api/Users (with user-role token) → PUT /api/Users/1 {role: admin} isn’t random browsing. It’s inferential escalation — each request informed by the previous response. The timing looks human. The reasoning chain doesn’t.

Detecting this requires modeling request sequences as hypothesis chains: does each request logically follow from the information gained in the previous response? Is the access pattern converging on something? Information-theoretic measures — how much each request narrows the agent’s uncertainty about the target — could provide signals robust against timing evasion.

“Does this session fit into a larger campaign?”

Swarm agents that look complementary rather than similar require a different detection paradigm: graph-based campaign detection. Model sessions not by behavioral similarity but by information flow. When session A discovers an endpoint and session B (different source, different timing, different behavioral fingerprint) subsequently targets what that endpoint revealed, the informational correlation is a campaign signal even though no individual session looks anomalous.

“How much trust has this actor earned?”

Binary classification — agent or human, legitimate or rogue — is brittle. A better primitive is continuous trust. An actor that has behaved consistently within declared scope for months has earned operational latitude. An actor with no history or erratic behavior gets tight constraints. Trust decays without reinforcement. Trust collapses fast on violation.

The security industry has been moving in this direction for years. CrowdStrike’s acquisition of SGNL for continuous dynamic authorization (see), Okta’s per-action risk evaluation, and the ABAC/RAdAC literature all treat authorization as context-dependent and risk-informed rather than static. What I think is still missing is the mathematical structure: trust as a continuous signal with entropic decay, asymmetric earn/collapse dynamics, and anti-accumulation resistance that increases as access volume grows. Whether that additional structure is worth the calibration overhead is an open empirical question — but the direction is right.

The Real Architecture

What I ended up with is not a better detector. It’s a layered stack.

Identity binds observations to a stable principal over time — without this, everything else collapses. Attestation lets cooperative agents prove who they are cryptographically. Per-action authorization evaluates each action against resource classification and context. Trust computation adds a continuous behavioral trust signal informed by detection history and demonstrated consistency. Intent verification checks whether behavior matches declared purpose. Normative evaluation asks whether the action should happen given organizational context.

Detection feeds this stack as a load-bearing input — enriching trust computation, detecting swarm coordination, informing monitoring intensity. But the system is designed so that detection error degrades governance quality rather than failing it entirely. If detection misclassifies an actor, the same trust physics, action risk evaluation, and normative constraints still apply. That’s a meaningful improvement over an architecture where detection error applies the wrong governance model entirely.

And then: offensive AI research reveals how agents adapt to those constraints, and the loop repeats.

The detection question everyone is trying to answer — “is this an agent or a human?” — is a useful input but the wrong question to build a governance architecture around. The question that actually governs: “has this actor earned the right to perform this action on this data in this context?”

That question is answerable regardless of whether the actor is human, agent, or something in between. It gets more answerable over time as behavioral history accumulates. And it doesn’t hit the Turing test ceiling because you’re not trying to see through a disguise — you’re evaluating a track record.

That’s the architecture behind Tamed Autonomy. Identity and attestation as the foundation. Detection as the sensory layer. Trust computation as the continuous signal. Intent verification as the purpose check. Normative reasoning as the governance layer. Each one insufficient alone. Together: a system that gets harder to evade over time, not easier.


Sources

Leave a Reply