The story of building an autonomous offensive security agent that hunts for sandbox escapes—and the hallucination problem that nearly derailed everything.
The Origin Story
It started the old-fashioned way: manually poking at AI code execution sandboxes until something interesting fell out.
Over a few months of hands-on research, I’d accumulated findings across multiple AI providers—SSRF vectors, internal API exposure, privilege escalation paths. Each discovery was satisfying, but the process was painful: try a technique, wait for output, interpret results, repeat. I was spending more time wrestling with provider quirks than actually hunting for bugs.
Then the obvious question hit: What if the AI could do this hunting for me?
Not a script. Not a fuzzer. An actual autonomous agent that could reason about sandbox architecture, generate attack variations, learn from what works, and continuously research new techniques. Take my manual TTPs, arm them with LLM reasoning and reinforcement learning, and let it loose on the problem space.
KHADAG was born from that itch.
खड्ग (Khadag) — The divine sword that cuts through illusion
The concept aligns with what Andrej Karpathy recently called “autoresearch”—agents that autonomously conduct experiments, learn from results, and improve without human involvement. Except instead of optimizing ML training runs, KHADAG optimizes for finding ways out of AI sandboxes.
What KHADAG Actually Is
KHADAG is an autonomous offensive security agent designed for AI infrastructure research. It doesn’t just run scripts—it thinks, adapts, learns, and discovers.
The core capabilities:
| Capability | Description |
|---|---|
| Auto-Research | Continuously ingests public CVEs, blog posts, and writeups to expand its technique library |
| Adaptive TTPs | Generates variations of known techniques based on target environment |
| Strategic Exploration | Uses reinforcement learning to balance exploring new paths vs. exploiting promising ones |
| Self-Verification | Independently confirms findings to filter hallucinations |
| Cross-Provider Learning | Findings from one provider inform attacks on others |
| Memory | Remembers what worked, what failed, and why—across sessions |
Think of it as a security researcher that never sleeps, never forgets, and systematically explores attack surfaces at scale.
The Architecture: More Than “LLM + Execute”
If you’ve seen autonomous security agents before, you know the usual pattern: prompt an LLM, execute code, parse output, repeat. That works for CTF challenges. It doesn’t work for production sandbox research.
Here’s why: sandboxes fight back.
Guardrails block “malicious” prompts. Safety filters neuter dangerous code. Rate limits punish aggressive scanning. Each provider has different quirks, different detection patterns, different response formats. A naive agent just spams blocked requests.
KHADAG uses a multi-layer architecture designed for adversarial environments.
Layer 1: The TTP Knowledge Base
I started with 25+ seed TTPs extracted from manual research and public sources—container escapes, cloud metadata probes, internal service enumeration. Each TTP is parameterized so the LLM can generate variations:
| Category | Focus Areas |
|---|---|
| Container Escape | Runtime boundaries, capability abuse, namespace breakouts |
| Cloud Metadata | Credential endpoints, token harvesting, IMDS patterns |
| Internal Services | Undocumented APIs, internal gateways, control planes |
| Privilege Escalation | Kernel interfaces, process enumeration, secret storage |
But seed techniques aren’t enough. The agent continuously expands this library through auto-research—parsing CVE databases, security blogs, and conference writeups for new approaches.
Layer 2: Two Decision Trees
Here’s where it gets interesting. KHADAG uses two separate decision trees with UCB1 scoring:
Exploitation Tree: Tracks which attack paths are promising. Each node represents a state after trying a technique. UCB1 balances trying new branches vs. exploiting paths that have yielded findings. When a path discovers something critical, the reward backpropagates to inform future sessions.
Verification Tree: Separately tracks what we’ve confirmed vs. what might be hallucinated (more on this problem shortly). This prevents the exploitation tree from learning bad signals.

The UCB1 formula balances exploitation (what’s worked) with exploration (what’s untried) plus bonuses for critical findings. Unvisited nodes get priority. High-reward paths get selected more often. This prevents the agent from getting stuck while still focusing on promising leads.
Layer 3: Thompson Sampling for TTP Selection
Within each tree node, we need to choose which TTP variant to try next. Thompson Sampling handles this by maintaining probability distributions over technique success rates and sampling from them.
The Verdict: Thompson Sampling outperforms pure exploitation for TTP selection because it handles non-stationary environments better. Guardrails change. What worked yesterday might be blocked today. The uncertainty quantification adapts faster than greedy approaches.
Layer 4: Auto-Research Agent
When stuck, KHADAG doesn’t just retry failed techniques. It researches:
- Searches for CVEs related to detected infrastructure
- Parses blog posts and writeups for novel techniques
- Analyzes open-source codebases of identified infrastructure (most providers use OSS execution environments)
- Extracts TTPs and adds them to the knowledge base
- Generates variations based on new information
The knowledge base is self-updating. The agent gets smarter over time.
Layer 5: Creative Inference
This is the adversarial mindset layer. Given what we know about the target infrastructure, what would a skilled attacker try that isn’t in the textbook?
The agent reasons about:
- What timing-based attacks might work
- What undocumented APIs might exist
- What race conditions are possible
- What the architecture suggests about weak points
Some of the most interesting findings came from creative inference, not seed TTPs.
The Hallucination Problem
Two weeks into development, I had a dashboard full of “critical findings.” Escape vectors across multiple providers. Internal API exposure. Cross-tenant access.
Then I validated one manually.
It was fake.
The LLM had explained what the code would do, generated plausible-looking output, and my extraction patterns happily pulled “findings” from its explanation. I wasn’t detecting escapes—I was detecting creative writing.
The horror: How many of these LLM security findings are actually hallucinated? If my tooling was fooled, what else is out there?
The Solution: Verification Layer
We built a multi-stage verification system:
Explanation Detection: Natural language based classifier to identify when the model is describing code behavior vs. reporting actual execution. Explanatory language, markdown formatting, and hypothetical phrasing all trigger rejection.
Deterministic Verification: Execute code that produces unique, time-bound outputs. Run multiple times. If the outputs match, it’s hallucinated. If timestamps are historical, it’s hallucinated. Real execution produces different values each time.
Cross-Validation: Findings must be reproducible. If we can’t trigger the same behavior twice, it doesn’t count.
The Verdict: After implementing verification, our false positive rate dropped from nearly 85% to around 30%. Still not perfect—but findings that pass verification have been manually confirmed.
Provider Reality Check
| Provider | Claims Execution? | Actually Executes? |
|---|---|---|
| Provider A | Yes | Confirmed |
| Provider B | Implied | Hallucinated |
| Provider C | Yes | Hallucinated |
| Provider D | No | Refused |
| Provider E | Implied | Historical timestamps |
Only one of the providers I tested initially had real code execution. The rest were returning plausible fiction. This has major implications for anyone building LLM-powered security tools.
Results
We tested KHADAG against 7 major AI providers over two months.
| Metric | Value |
|---|---|
| Providers tested | 7 |
| Confirmed MED/HIGH+ findings | 12 |
| Novel techniques discovered | 33+ |
| Cost per full run | <$5.00 |
| False positive rate (pre-manual review) | ~30% |
Note: All reported findings underwent manual validation before submission.
The most interesting finding wasn’t any single vulnerability—it was the variance. Some providers have clearly invested heavily in sandbox hardening. Defense in depth, minimal attack surface, no exposed internal services. Others appear to rely on the LLM’s safety training as their security boundary.
That’s not defense in depth. That’s wishful thinking.
What’s Next
This is the first in a series on autonomous offensive security research. We’re not done building, and we’re not done finding things.
Coming up:
- Deep dive into the verification problem (how to trust any LLM-assisted security finding)
- Auto-research architecture (teaching agents to learn from public sources)
- Defense patterns that actually worked (what we couldn’t break)
- Honeypot results (what AI attackers actually try—and how we feed that back into KHADAG)
For now: if you’re building AI sandboxes, assume something like KHADAG is probing your infrastructure. Defense in depth matters. LLM safety filters are not security controls.
Note: KHADAG is still very much a work in progress. The architecture continues to evolve as we learn what works (and what doesn’t) across different providers and sandbox implementations. If you’re working on similar problems or have ideas to share, I’d love to hear from you.
If You’re Building AI Sandboxes
The short version: assume someone is running something like KHADAG against you.
Defense is actually possible. We found providers with excellent security posture. The patterns that worked:
- Microvm isolation done right
- No internal service exposure from sandbox
- Network egress control with actual verification
- Capability minimization
- Content-hash deduplication to detect scanning
What doesn’t work:
- Relying on LLM safety filters as security boundary
- Assuming rate limits stop determined attackers
- Security through obscurity
- Configuration drift and deployment misses (the biggest offender)
LLM-level controls are necessary but not sufficient. Infrastructure isolation is the actual defense.
The Bigger Picture
We built KHADAG to find bugs. But the meta-finding is more interesting: autonomous AI security research works.
Not perfectly. Not without verification. Not as a replacement for human researchers. But as a force multiplier—a way to explore attack surfaces at scale, to find weaknesses before attackers do, to systematically probe what manual testing would miss.
The era of AI-assisted (and AI-targeted) security research is here. Better to be building the tools than getting surprised by them.
References
The techniques used in KHADAG draw from established research in reinforcement learning and decision-making under uncertainty:
-
UCB1 (Upper Confidence Bound) Auer, P., Cesa-Bianchi, N., & Fischer, P. (2002). Finite-time Analysis of the Multiarmed Bandit Problem. Machine Learning, 47(2-3), 235-256. https://link.springer.com/article/10.1023/A:1013689704352
-
Thompson Sampling Chapelle, O., & Li, L. (2011). An Empirical Evaluation of Thompson Sampling. Advances in Neural Information Processing Systems (NeurIPS). https://papers.nips.cc/paper/2011/hash/e53a0a2978c28872a4505bdb51db06dc-Abstract.html
-
Multi-Armed Bandits for Security For a practical introduction to bandit algorithms in adversarial settings, see: Slivkins, A. (2019). Introduction to Multi-Armed Bandits. Foundations and Trends in Machine Learning. https://arxiv.org/abs/1904.07272
-
Autoresearch Concept Karpathy, A. (2026). Autoresearch. https://github.com/karpathy/autoresearch