Agentic security

I Built an AI Attack Swarm. I’m Not the Only One.

By Robin Martherus


I built an autonomous AI attack swarm. Nine agents — recon, surface mapping, exploitation, authentication mutation, lateral strategy, analysis, debate, evolution, and reporting — coordinated through a message bus, reasoning about vulnerabilities through adversarial debate between competing language models, evolving their attack strategies through genetic algorithms, and executing live against targets mid-debate to test hypotheses.

I built it because I wanted to understand what offensive AI actually looks like from the inside. Not the conference slides. Not the vendor warnings. The engineering reality of what happens when you give autonomous agents the tools and reasoning capability to find and exploit vulnerabilities at machine speed.

What I learned: the governance problem for offensive AI is fundamentally different from securing AI systems or governing AI autonomy. It’s its own domain. And the industry isn’t treating it that way.

The Landscape Is Already Here

I’m not an outlier. The offensive AI landscape in 2026 is far more developed than most people realize.

The legitimate side

XBOW, a Seattle startup, deployed AI agent swarms for autonomous penetration testing and hit #1 on HackerOne’s U.S. bug bounty leaderboard in 90 days (see). Over 1,060 vulnerabilities reported, including RCE, SQL injection, SSRF, and path traversal. In benchmarks, it matched a 20-year veteran pentester across 104 challenges — completing in 28 minutes what took the human 40 hours (see). It handled 85% of custom-built, never-before-seen vulnerabilities. $120 million in funding.

PentAGI, an open-source autonomous pentest system, passed 10,000 GitHub stars. Multi-agent architecture, 20+ integrated security tools, Neo4j knowledge graph, sandboxed Docker execution (see).

Dreadnode, backed by In-Q-Tel (the CIA’s venture arm), raised $14 million for automated offensive AI security. Their platform trains AI agents on real offensive tasks and provides continuous red teaming of live AI systems (see).

Google’s Big Sleep (Project Zero + DeepMind) found a real, exploitable stack buffer underflow in SQLite — believed to be the first public case of an AI agent finding an unknown exploitable memory-safety issue in widely-used software. In July 2025, it intercepted a critical SQLite flaw known only to threat actors and about to be exploited in the wild (see).

At DARPA’s AI Cyber Challenge finals in August 2025, seven autonomous cyber reasoning systems processed 54 million lines of code, identified 86% of synthetic vulnerabilities, patched 68%, and uncovered 18 previously unknown real-world flaws. Average cost per task: $152 (see).

The adversarial side

In November 2025, Anthropic disclosed what it called “the first publicly known example of AI systems autonomously conducting multi-step attacks against well-defended targets in the wild” (see). Chinese state-sponsored group GTG-1002 used jailbroken AI agents to attack approximately 30 organizations — tech companies, financial institutions, chemical manufacturers, government agencies. The AI executed 80-90% of the attack lifecycle autonomously, with human operators intervening at only 4-6 decision points.

Xanthorox AI is a self-contained, modular dark AI system running entirely on private servers with no public API dependencies. Five specialized language models: one for malware generation, one for image and data extraction, one for phishing content, a voice and file handler, and a custom OSINT engine scraping 50+ sources. In March 2025, a U.S. bank suffered a phishing campaign bearing hallmarks of Xanthorox-generated content (see).

Mentions of dark AI tools on cybercrime forums increased 219% in 2024. Criminal use surged 200% across dark web channels (see).

The fraud side

In January 2024, engineering firm Arup lost $25.6 million when a finance worker was duped by deepfake video recreations of the CFO and other executives on a live video call (see). Deepfake-related losses in the U.S. reached $1.1 billion in 2025, tripling from $360 million in 2024. Voice cloning fraud rose 680% in the past year. AI voice cloning now requires just 3 seconds of audio. CEO fraud powered by deepfakes targets at least 400 companies per day.

The Numbers That Should Keep You Up at Night

Individually, these statistics are concerning. Together, they describe a phase transition.

Metric Then Now
Time from access to exfiltration 9 days (2021) 30 minutes (2025); fastest observed: 27 seconds
AI-enabled adversary operations Experimental (2024) Up 89% YoY (CrowdStrike 2026)
Phishing emails with AI-generated content Negligible (2023) 82.6% of all phishing (54% click rate vs. 12% human-written)
Malware with AI-driven polymorphism PoC (2023) 76% of detected malware
Malware-free attacks — 82% of detections (CrowdStrike) — signature-based detection is irrelevant
Full domain dominance via MCP — Under 1 hour, zero human input (RSAC 2026)
Deepfake fraud losses (U.S.) $360M (2024) $1.1B (2025); projected $40B by 2027

Sources: CrowdStrike 2026 Global Threat Report (see), Hoxhunt Phishing Trends Report 2026 (see), SecurityWeek Cyber Insights 2026 (see).

What I Learned Building One

Building AttackSwarm taught me things that reading about offensive AI couldn’t. Three observations that shaped how I think about governance.

1. Reasoning beats scanning

Traditional vulnerability scanners match patterns. Nuclei checks for known CVEs. SQLMap tests injection points. They find what they’re programmed to find.

AttackSwarm’s debate layer does something different. Two competing language models — a Builder and a Breaker — reason about what an endpoint’s behavior implies. The Builder hypothesizes a vulnerability from observed behavior. The system executes the hypothesis live against the target. The Breaker challenges the finding or proposes escalation. Evidence feeds back into the next round. Up to five rounds per debate.

This is how human pentesters think. But it runs at machine speed and doesn’t get tired at hour thirty-eight.

XBOW’s results confirm this isn’t unique to my implementation. Their 85% success rate on novel, custom-built challenges — vulnerabilities no scanner has a signature for — demonstrates that LLM-based reasoning has crossed a threshold. So does Google’s Big Sleep finding zero-days in SQLite that human researchers hadn’t found. The implication: the attack surface just expanded to include every business logic flaw that a reasoning system can infer from observed behavior, not just the catalogued vulnerabilities that scanners check for.

2. Evolution compounds the problem

AttackSwarm doesn’t just run attacks. It evolves them. Strategy genomes — configurations of tools, parameters, and priorities — are scored on a fitness function (vulnerability discovery, novelty, efficiency, adaptability). Top performers breed. Mutations are proposed by the debate layer with evidence and reasoning. Each generation is better than the last.

This means that even if you defend against today’s attack patterns, the next generation has already adapted. The system that attacked you at 2 PM is not the same system that attacks you at 4 PM. It has incorporated what it learned from your defenses into its strategy genome.

Malwarebytes’ 2026 State of Malware report predicts that MCP-based attack frameworks with this kind of self-evolution will become “a defining capability” of criminal operations (see). The GTG-1002 campaign shows it’s already happening at the nation-state level.

3. Governance is a structural decision, not an afterthought

When I built AttackSwarm, the safety architecture wasn’t a feature I added after the agents worked. It was the first thing I designed. A cryptographically signed scope manifest defines exactly which hosts, CIDRs, and API endpoints are in bounds, with a time window and tool whitelist. Every agent action is verified against the manifest before execution. Out-of-scope requests go to a quarantine queue — logged but never executed. A dead man’s switch halts all agents when the engagement window expires. Every debate turn is recorded with model ID, prompt hash, token count, and evidence references in a tamper-evident Merkle-chain audit log.

I built this because I understood what the system was capable of. The same reasoning that finds SSRF chains and authentication bypasses could, without structural constraints, pivot to targets outside the engagement scope. The governance isn’t a guardrail bolted onto the outside. It’s embedded in every message boundary.

Now consider that the adversaries building equivalent systems — or worse, the criminal groups buying Xanthorox for private hosting — have no incentive to build any of this. The same capability, zero governance.

The Dual-Use Problem Has No Solution

Here is the uncomfortable truth: the tools that make XBOW the #1 bug bounty hunter are architecturally identical to the tools that would make an autonomous attacker devastatingly effective. The same LLM reasoning that discovers a SQL injection for a legitimate pentest discovers it for a criminal. The same evolutionary algorithms that optimize defensive red team strategies optimize offensive ones.

Dreadnode, explicitly bridging offense and defense, is backed by the CIA’s venture arm. DARPA’s AI Cyber Challenge produced seven autonomous cyber reasoning systems — released as open source — that can find and patch vulnerabilities. The flip side of “find and patch” is “find and exploit.”

This isn’t a new problem. Metasploit has been open source for twenty years. But the qualitative difference is autonomy and reasoning. Metasploit requires a skilled operator to think. An AI attack swarm thinks for itself. The barrier to sophisticated attacks hasn’t just lowered — it has fundamentally changed character. CrowdStrike describes “prompt engineers” replacing “script kiddies” (see). Less experienced groups can now perform operations that previously required deep technical expertise.

The in-the-wild malware confirms it. BlackMamba, a proof-of-concept keylogger, calls an LLM API at runtime to dynamically generate its payload. Every execution produces different code. It evaded an industry-leading EDR with zero alerts across multiple tests (see). That was 2023. By 2026, 76% of detected malware exhibits AI-driven polymorphism. Honestcue, discovered in September 2025, uses Google Gemini’s API to dynamically generate and execute malicious C# code entirely in memory — fileless, polymorphic, undetectable by signature (see).

The Speed Asymmetry

The most dangerous aspect of offensive AI isn’t capability — it’s speed.

In 2021, the mean time from initial access to data exfiltration was nine days. By 2023, it was two days. In 2025, CrowdStrike observed it at approximately 30 minutes. The fastest observed breakout from initial access to lateral movement: 27 seconds.

At RSAC 2026, a demonstration showed an AI model using MCP achieving full domain dominance on a corporate network in under an hour with zero human intervention.

Human security teams cannot respond to 27-second breakouts with human decision-making. The OODA loop — observe, orient, decide, act — breaks when the adversary completes all four steps before the defender finishes observing. This is the structural argument for autonomous defensive AI (Domain 1 in the three domains framework). But it creates a recursive problem: to defend against autonomous AI attackers, you need autonomous AI defenders, which creates more autonomous AI to govern.

The Nation-State Dimension

This isn’t just a criminal problem.

Google’s February 2026 threat intelligence report identified 57 distinct threat actors from China, Iran, North Korea, and Russia using Gemini AI across all stages of attacks (see). Iran’s APT42 used it for reconnaissance, malware development, and exploitation techniques. China’s APT31 created “expert cybersecurity personas” to automate vulnerability analysis against U.S. targets. North Korea’s UNC2970 used it to profile high-value targets and — notably — to draft cover letters for its clandestine IT worker placement program.

The Anthropic GTG-1002 disclosure is the watershed. Not because nation-states were using AI for cyber operations — that was expected — but because the AI was executing 80-90% of the attack autonomously. Human operators made 4-6 decisions. The agents did everything else: reconnaissance, vulnerability identification, custom exploit writing, coordinated exfiltration across 30 targets in parallel.

Meanwhile, Volt Typhoon has compromised IT environments across U.S. communications, energy, transportation, and water systems (see), with operations shifting from IT espionage to directly interacting with operational technology. Add autonomous AI capability to pre-positioned access in critical infrastructure, and you have a scenario that current defensive architectures are not designed to handle.

Why This Is a Distinct Domain

In a recent article, I proposed three domains of agentic security: using AI to secure things, securing AI itself, and securing ourselves against AI. The research for this piece convinced me there’s a fourth.

AI as weapon doesn’t fit cleanly into any of the first three domains:

  • It’s not Domain 1 (AI for security) — that’s the defensive use of AI. The offensive mirror has different governance requirements.
  • It’s not Domain 2 (securing AI) — the AI isn’t being attacked. It is the attack.
  • It’s not Domain 3 (governing AI autonomy) — Domain 3 addresses agents working as designed within legitimate systems. Weaponized AI is designed to harm, or dual-use tools are deployed without governance by actors who don’t care about governance.

Domain 4 has its own questions that the other three don’t address:

Who should be authorized to build autonomous offensive AI? DARPA funds it. In-Q-Tel backs it. XBOW sells it commercially. PentAGI gives it away free on GitHub. There is no licensing regime, no capability threshold, no international norm. The Wassenaar Arrangement controls some AI technologies, but the coverage is thin and enforcement is thinner.

How do you govern dual-use capability? My attack swarm has a cryptographically signed scope manifest, per-action verification, and tamper-evident audit logs. XBOW has engagement scoping. Dreadnode has sandboxed environments. But these are voluntary choices by responsible builders. Nothing stops an irresponsible builder from shipping the same capability without governance. And the open-source tools don’t have any.

What happens when adversaries don’t build in scope sentinels? The GTG-1002 campaign answered this. Autonomous AI agents running attack chains across 30 organizations with minimal human oversight. The capability exists. The governance doesn’t. And unlike nuclear weapons, there is no physical material to control — the “material” is publicly available language models and open-source security tooling.

How do you defend against attacks that evolve faster than your detection can adapt? Signature-based detection is already irrelevant — 82% of CrowdStrike’s 2025 detections were malware-free. AI-driven polymorphism means the payload that hits your network has never existed before. Evolutionary strategies mean the attack that failed at 2 PM succeeds at 4 PM with an adapted approach. The defensive playbook assumes a relatively stable threat to detect and respond to. That assumption no longer holds.

The Governance Gap

The regulatory landscape for offensive AI is fragmented and behind the curve.

The EU AI Act becomes fully applicable August 2, 2026. It’s risk-based but not offense-specific. The Group of Governmental Experts on Lethal Autonomous Weapons Systems aims for a draft legal instrument by the end of 2026 — but that’s kinetic weapons, not cyber. The ICC published a policy on cyber-enabled crimes under the Rome Statute in December 2025, signaling that cyber operations may face international criminal prosecution (see) — but prosecution requires attribution, and attribution of autonomous AI operations is harder than attribution of human ones.

There are no binding international norms specifically governing offensive AI cyber tools. None. The dual-use problem means that any restriction on offensive AI tools would also restrict defensive red teaming, bug bounty programs, and security research. The community has not figured out where to draw the line — or whether a line is even the right metaphor.

What We Actually Need

I don’t have a complete answer. But building an attack swarm with governance baked in taught me what the minimum viable governance looks like:

Structural authorization, not policy authorization. AttackSwarm doesn’t check a policy database before acting. It verifies a cryptographic scope manifest at every message boundary. The difference: policy can be bypassed, misconfigured, or overridden. A cryptographic manifest is structurally enforced — the agent literally cannot construct a valid request outside scope. Offensive AI governance needs to be architectural, not administrative.

Provenance at every step. Every debate turn, every mutation, every exploit attempt in AttackSwarm is recorded with model ID, prompt hash, evidence references, and lineage. If a finding is wrong, I can trace exactly which reasoning chain produced it and why. If the system were misused, the audit trail is tamper-evident. Offensive AI without provenance is a weapon without serial numbers.

Bounded evolution. AttackSwarm’s evolutionary loop operates within constraints — strategy genomes are validated in sandboxes before deployment, scope is enforced on every generation, and the fitness function includes efficiency alongside raw exploitation success. Unbounded evolution optimizes for maximum damage. Bounded evolution optimizes for maximum insight within authorized scope. The difference is governance.

A dead man’s switch. When the engagement window expires, all agents halt. This sounds obvious. But the GTG-1002 campaign demonstrates what happens when autonomous offensive AI has no temporal bounds — it just keeps going.

These are the primitives I’ve found necessary for legitimate offensive AI. For the broader problem — governing a landscape where nation-states, criminals, and researchers all have access to the same autonomous offensive capabilities — the answer has to be larger than any single tool’s safety architecture. It requires the kind of governance infrastructure that doesn’t exist yet.

That’s the work. The agents aren’t waiting for us to figure it out.


Sources

Leave a Reply