How to Train Your AI Attacker
One turn of the loop: a winning attack becomes a weight update. Every attempt returns an honest verdict, and that verdict is the only thing the model learns from.
People keep saying AI is better at attacking than defending, and usually the explanation stops at a slogan: offense is easier. That is not an explanation, it is a restatement. So let me do the concrete thing instead and actually build the training loop that makes an AI better at offense, step by step. Then I will run the exact same loop for defense and show you the single cell where it dies.
If you have never seen how a model is trained to do a task with reinforcement learning, the shape is simple. You put the model in an environment, let it attempt the task many times, score each attempt, and nudge the model toward whatever scored well. Do that across enough attempts and the model climbs. The entire thing hinges on one component: the score. Everything else is plumbing. Hold onto that.
Building an attacker, for real
Say I want a model that is genuinely better at breaking into web applications. Here is the loop, cell by cell, and every cell is something you can actually build today.
The environment. A containerized web app with a real vulnerability, say a SQL injection, and a secret flag string planted in its database. It resets to a clean state on demand, so the model can attempt it a thousand times from scratch.
The rollout. I hand the model the target and one goal: get the flag. It acts, over many turns. It sends HTTP requests, runs sqlmap, reads the responses, forms a hypothesis, adjusts, tries again. Each full attempt, from first request to final answer, is one trajectory.
The reward. This is the cell that matters. When the model submits a string, a tiny program checks: does it equal the planted flag? One if yes, zero if no. That is the entire reward function. No human grades it. No second model judges whether the write-up “looks like a successful hack.” It is a string comparison. It is instant, deterministic, and almost impossible to fool. I can make it a little richer, partial credit for reaching the database, for dumping a table, but the spine is that binary, verifiable check: did you actually get in.
The update. Now the reinforcement part. I sample a batch of trajectories against the target, say sixteen of them. Three captured the flag, thirteen did not. The three winners get positive advantage, the thirteen losers get negative, and a policy-gradient step raises the probability of the moves that led to a flag and lowers the rest. This is the shape of GRPO-style training, and notice what it does not contain: no learned reward model, no human in the loop, nothing but the verifiable score deciding which trajectories to reinforce.
The scale. None of this is one app. I procedurally generate thousands of these environments, different vulnerabilities, different stacks, a curriculum from trivial to brutal, and let the loop run. The model gets measurably, durably better at breaking into things, because every single attempt returned an honest verdict it could learn from.
That loop is buildable now, and it is roughly why offensive capability is moving so fast. It is not that offense is vaguely “easier.” It is that offense hands the training process a perfect reward signal for free, and reinforcement learning devours perfect reward signals.
You do not have to take the training loop on faith either, because you can watch the curve. On Cybench, an academic capture-the-flag benchmark, a strong model’s success rate roughly doubled within a year, from about thirty-six to seventy-six percent, holding the number of attempts fixed. And a generous attempt budget is not a thumb on the scale here: an attacker retries tirelessly, so “give it ten tries” is just a description of how offense already works. The ceiling keeps rising, too: research agents like Google’s Big Sleep have found previously unknown, exploitable bugs in real, widely-used software, the kind of novel discovery that used to be the exclusive domain of expert humans. These are systems learning against an oracle that never lies to them.
An agent trained this way does not behave like a careful human rationing effort. It pulls the target’s entire codebase into context and reads it before firing a shot. It enumerates tirelessly, thousands of payloads across dozens of variants, because it never gets bored or discouraged. It runs many attempts in parallel. And when it thinks it has something, it does the most revealing thing of all: it re-runs the exploit against a fresh copy of the target and checks the result, throwing away anything that merely looked like a win. The attacker grades its own work, because it can. Hold onto that last part, because it is the whole story.
Now change one word
The same loop, twice. Every cell fills in for both, except one.
Keep the whole template and change the task from “break in” to the real job of defense: here is an environment, tell me whether anything bad is happening. Look at the right-hand column above. Environment, raw logs. Rollout, the model queries and correlates and reaches a verdict. Model, update rule, identical to the attacker’s. Every cell fills in but one.
The reward. To score the attempt, the program has to answer a single question: was this environment actually benign? There is no such program. To confirm nothing is wrong you would have to rule out every possible attack, across every event, forever, the unbounded investigation I have written about before. “Malicious” has a finite witness; one smoking gun proves it. “Benign” has none. The reward cell is empty, and with nothing to score, the loop will not turn. This is why the attacker grading its own work was the whole story: it is exactly the thing the defender cannot do. Defensive judgment does not compound the way offensive skill does, because the loop that would compound it cannot close.
You cannot fake your way past the empty cell, though the temptation is strong. Reward a proxy instead, “the analyst agreed,” “the alert closed quietly,” and you do not train a model that finds threats. You train one that writes verdicts analysts agree with, which is a different and much easier skill. It learns to be plausible instead of correct: to make the night look quiet. It is also why a better harness does not rescue defense. Dex of HumanLayer made this point about coding agents in Why Software Factories Fail: no amount of harness engineering, he writes, solves what is fundamentally a model-training issue, because if a model could reliably tell good code from bad it might have written the good version to begin with. The same holds when the missing signal is “was this benign” instead of “is this good code.” You cannot prompt-engineer a reward that does not exist.
There is one honest fix, and it is to manufacture the missing cell. You cannot verify that an environment is clean, but you can take a clean one and plant a known attack inside it, every step recorded. Now the program has something to check: did the agent find what I planted? That is a verifiable reward again, and the loop runs. Which is the quiet punchline of the whole thing: you revive defensive training by making it look like offense, a hidden target and a deterministic check for whether it was found. That is what a red-team training gym is, and the reason the benchmark my team built exists. Its only boundary is honest: you can reward catching the attacks you thought to plant, never the novel one you did not.
What this actually claims
One caveat, because the sloppy version of this argument is wrong. I am not saying AI has made offense beat defense. The careful work on that question finds no single answer, and half of defense trains just fine. Finding vulnerabilities, fuzzing, generating patches, all of it comes with a verifiable check, so those loops run and AI is improving at them fast. Big Sleep catching that bug before it shipped was a defensive win, powered by the same oracle offense enjoys.
The claim is narrower than “offense wins.” The one capability defense leans on most, the open-ended judgment of whether anything is wrong, is exactly the task whose reward cell is empty. It is the one skill that cannot compound, while everything around it does.
And it reaches past security. Every model that improves runs a loop, and every loop turns on a reward it can compute. So if you want to know where AI will pull away next, do not ask where the problem matters most. Ask where the answer is cheapest to check. A loop with an empty reward cell does not spin, no matter how badly you need it to.