· ai security-operations self-improvement red-teaming machine-learning

The Sycophantic SOC

Every AI SOC pitch ends the same way. The system triages alerts, your analysts correct it, the corrections flow back into the model, and the loop improves forever. Human feedback is the gold-standard signal, the ingredient that makes the whole thing trustworthy. More feedback, better system. Who could object to learning from your best people?

Instead of arguing with that in the abstract, let me tell you a story about a SOC where everything works. The analysts are right every time. The model learns from them faithfully. And the loop still ends somewhere bad, because of a number that never appears anywhere in it.

Three alerts, three correct dismissals

Week one. A detection fires: the service account svc-inventory enumerated four thousand computer objects over LDAP at 2 a.m. An analyst looks, recognizes it, and closes it: “Asset inventory job. Runs every Tuesday night. Benign.” The same alert has fired every Tuesday for months, and every closure note says the same thing.

Week three. Encoded PowerShell spawns on an IT admin’s workstation. The analyst pulls the process tree, sees the RMM agent as the parent, and confirms with IT that the vendor encodes its update bootstrap. “Known updater behaviour. Verified with IT. Benign.”

Week five. An impossible-travel alert: the VP of sales logged in from Boston, then from Frankfurt forty minutes later. The analyst checks and finds the corporate VPN egresses in Frankfurt. “VPN egress point. Benign.”

That sequence is the best-case version of the loop. Three skilled analysts, three correct verdicts, three tickets closed with clear reasoning. This is the feedback every vendor wants and every SOC leader hopes their team produces. There is no failure anywhere in this section.

What the model learned

The self-improving system absorbs those closures, and here is where it pays to be precise about what “learning” means. The model does not store the analyst’s reasoning. It stores a boundary.

The analyst wrote “the Tuesday inventory job is benign.” The model, generalizing across dozens of similar closures, learns something wider: service accounts that enumerate the directory are expected here. The analyst wrote “the RMM updater encodes its bootstrap.” The model learns: encoded PowerShell on admin workstations is expected here. The analyst wrote “the VPN egresses in Frankfurt.” The model learns: Frankfurt logins are expected here.

That widening is generalization, and it is the entire point of using a model. If the system only memorized exact past alerts, it would suppress nothing new and save nobody any time. The value proposition is the widening. So is the danger: the analyst labeled an instance, and the model learned a region. Nobody decided where the edges of that region are. Nobody has even seen them. The edges are wherever gradient descent put them, inferred from closure notes that were never written to be training data.

A cluster of labeled benign events sits inside a much larger dashed region the model learned to suppress; a red X marking the Thursday attack falls far from the cluster but still inside the region, auto-closed as schedule variance The analysts labeled the dots. The model drew the boundary, and nobody chose where its edges are.

The obvious objection is that a modern system stores the reasoning too, and the better ones do: keep the closure note, check each new alert against it. That narrows the region, and it is worth doing. But look at what the check itself is. “Is a Thursday afternoon consistent with a note about a Tuesday night job?” has no mechanical answer; it is a judgment call, and the model making it learns from which of its judgments analysts accepted. Enforced strictly, the stored rationale is an allow-list rule scoped to Tuesdays at 2 a.m., which suppresses almost nothing and saves no one any time. Read generously, it is the widening again, now with a rationale attached. Where the system lands between those poles is set by feedback, and feedback rewards the generous reading, because the generous reading is the one that closes tickets.

Each retrain, the boundary fits your environment a little more snugly. False positives drop. The dashboard improves every quarter. Everyone involved is doing their job well.

The Thursday that looks like a Tuesday

Eight months in, someone phishes the VP of sales.

The attacker logs in with the stolen credentials. Geo-velocity fires, and the system closes it on its own: consistent with corporate VPN egress pattern, see 14 prior closures. The attacker lands on an IT admin workstation and pulls tooling down with encoded PowerShell: matches known updater behaviour class, verified benign by analyst review. From the admin host they harvest the credential for svc-inventory, and on a Thursday afternoon they enumerate the entire directory with it.

Thursday, 2 p.m., is not Tuesday, 2 a.m. The deviation is right there in the telemetry. The system notes it and explains it away: schedule variance within historical norms for this account class. Confidence 0.97. That is the generous reading, working as trained. Every closure note is fluent, cites precedent, and is wrong.

Twenty-five days later, which is the median dwell time for externally notified victims, the company learns about the breach the way nearly half of victims still do: from an outsider. In the post-incident review, someone pulls the auto-closure notes and finds the whole intrusion, narrated step by step, in the calm voice of a system explaining why none of it mattered. The kill chain was assembled entirely from behaviours your own analysts had marked benign. Each individual verdict was right. The region each verdict grew into was not.

We have seen the human version of this. In 2013, Target’s FireEye deployment flagged the exfiltration malware in time to stop it, and the security team took no action; Bloomberg’s reconstruction could say only that the sirens went off and, for some reason, nothing happened. A human non-response like that happens once. A trained model turns non-response into policy, applies it at machine speed, and writes you a justification.

Overfitting, in the technical sense

There is an exact name for what happened to that model, and it helps to build it from the ground up.

A student preparing for an exam can do two different things with a stack of past papers: learn the subject, or learn the papers. Both produce excellent scores on practice tests. The difference only shows up when the real exam asks something the papers never did. We call the second strategy overfitting: the model has minimized error on its training distribution rather than learning the concept the training data was supposed to point at. In security the situation is sharper, because the real exam is written by an adversary, and the adversary has, in effect, read your practice answers.

Now look at what the triage model’s training distribution is. It is your organization’s benign traffic, labeled by analysts whose job is to explain it. The malicious class is missing. Months of feedback might contain thousands of benign closures and a handful of true positives, most of them commodity malware the EDR caught anyway. You asked for a model of malicious versus benign and supplied data that can only teach normal for this org versus abnormal, and the closure notes push it one step further, toward explainable versus unexplainable. The concept it converges on is not “attack.” It is “thing my analysts wouldn’t bother with.”

Three properties of the loop make this worse over time rather than better:

The model curates its own training data. Once it suppresses a class of alert, that class stops generating analyst labels at all. Sculley et al. flagged this as a first-class hazard, the hidden feedback loop, back in 2015: a model that influences its own future inputs. A suppressed detection is beyond correction. Its error rate has become unmeasurable, and unmeasurable reads as zero on every dashboard you own.

Agreement is the objective. The reward the loop optimizes is analyst approval, and Goodhart’s law does the rest: the system learns to produce verdicts analysts will accept, which means verdicts matching their priors. Anthropic measured this dynamic in language models: both humans and the preference models trained on their judgments prefer convincingly written answers that match the reader’s beliefs over correct ones, a non-negligible fraction of the time. Optimizing against human feedback produced sycophants. A sycophantic assistant tells you your essay is great. A sycophantic SOC tells you your environment is clean.

The tails go first. When models retrain on their own accepted outputs, the rare events drop out of the distribution before anything else. In a SOC, the rare events are the attacks. The follow-up research also located the exit: collapse is avoided when training keeps drawing on data the model did not generate itself. Hold onto that result, because it is the prescription this post ends on.

One scope note before the payoff: everything above scales with autonomy. A deployment that gates every suppression behind human review blunts the loop, but only by as much as reviewers disagree with a model trained to be agreed with; a queue of confident, fluent, precedent-citing closure notes turns review into ratification. The story ran on full auto-close because that is where the failure runs fastest, not because it is the only place it runs.

The number that never appears in the loop

Here is the meta point the story has been building to, and it is the part I think the AI SOC conversation keeps missing.

Every dismissal the loop learns from records a benefit: one more false positive that will never waste an analyst’s minute again. That benefit is visible, measurable, and compounds on a dashboard. What the loop never records is the cost side of the ledger, because the cost is a counterfactual: how many attack chains was this detection the unique tripwire for?

Take the noisy LDAP alert from week one. Its precision on historical data was abysmal; it fired benign a hundred times in a row. But directory enumeration is how an attacker orients before moving; it sits on the path of nearly every Active Directory lateral-movement chain an attacker can run. For a large family of intrusions, that alert was the only observable between initial access and impact, the one place the chain had to cross a sensor. Judged by history, the detection was worthless noise. Judged by the attacks that route through it, it was load-bearing.

A threat-informed analyst can see the second judgment, and some encode it. The week-one closure note could have read “noisy, but this is our only tripwire for directory recon,” and that sentence carries real coverage information. The failure is what the loop’s objective does with it. The reward is agreement, delivered through closures, so a coverage argument survives training only when it happens to coincide with a verdict the model is reinforced for. Escalating a hundred benign firings because of the attacks they might someday represent looks, to the objective, like a hundred errors. Coverage reasoning exists in analysts; the loop just never prices it, and an unpriced signal gets optimized away as surely as an absent one.

Four intrusion families all route through one noisy directory-enumeration detection before impact; feedback measures its hundred benign firings but cannot measure that the four chain families have no other tripwire Feedback prices the noise. It has no way to price the coverage, because coverage is a claim about attacks that haven’t happened.

I made the human-scale version of this argument in The Overfitting Problem in Detection Engineering: every exception you hand-tune into a rule carves a blind spot, and attackers live in the exclusions. The learned version is categorically worse. The exceptions are now carved automatically and continuously, their edges are wherever the model generalized to, and no human reviews a boundary nobody can see. And each dismissal was already a liability awaiting future context on its own; the loop bundles those liabilities into policy at scale.

Make the counterfactual factual

If the essential number is a counterfactual, the only way to measure it is to make it factual: run the attacks. This is narrower than the standing advice to test your defenses. In this architecture the adversarial exercise is a measurement instrument, the one input that forces the loop to price coverage, because a planted chain is caught or missed and neither outcome can be reinterpreted as agreement.

A suppression has to be paid for with evidence. Before the loop retires a detection class, run the attack chains that route through that observable and show the system still catches them somewhere else. If it can’t, the noise was the price of the only sensor you had, and the correct move is triage automation on top of the alert, never suppression of it.

Adversarial findings are training data, at the loop’s own cadence. A red team miss is the correction your environment would never generate on its own, and it has to arrive as often as the model retrains, which agentic red teams have made affordable.

Measure planted-unknown recall, not agreement. Agreement rate is the signal being optimized, so it is the metric Goodhart eats first. The metric that resists is the fraction of deliberately injected attack chains the system catches, scored the way verifiable domains score everything: against known answers. You planted the chain; you know what should have been found. If your AI SOC vendor reports agreement rate and time-to-close but can’t give you a planted-unknown recall number, you are measuring the sycophancy, not the security.

One objection deserves to be met head-on, because this essay has no right to wave it off: planted chains are also a distribution. A pipeline trained on red-team findings can overfit to them the way it overfit to the environment, learning to catch whatever resembles your red team while live adversaries adapt away from it. This is the model-collapse result wearing a different coat, and its escape clause applies here too: the loop stays healthy only on adversarial data it did not generate and cannot predict. In practice, refresh the attack distribution from outside, from live intrusion tradecraft and threat intelligence rather than replays of last quarter’s exercise. Point the red team at the current suppression boundary instead of a fixed playbook, so the target moves whenever the model moves. And keep the split honest: chains that enter training and chains that score planted-unknown recall must be disjoint, or the recall metric Goodharts the way agreement did. A red team that becomes predictable stops being a source of novelty and becomes one more environment to overfit.

This extends The Detection Bias Trap, where I argued AI SOCs inherit their analysts’ conservative drift and need a sparring partner. The story above is one level deeper than bias. Even with threat-informed analysts writing coverage arguments into their closure notes, the loop’s objective never prices the one quantity a defense exists to protect: coverage of attacks that haven’t happened yet.

A self-improving defense that listens only to its analysts is not learning to defend your environment. It is learning to agree with it, one correct dismissal at a time.