· security-operations soc verifiability detection-engineering ai

Triage Doesn't Produce a Verdict. It Produces Residual Risk.

Anton Chuvakin wants to kill triage, and he’s right. For twenty-five years the SOC pipeline has been Detect → Triage → Investigate: a human at a queue, spending a few minutes per alert deciding whether it’s worth a real look. That filter was never a good idea. It was a rationing scheme. We built it because humans don’t scale and we couldn’t afford to investigate everything, so we invented a cheap step to decide what the expensive step would skip. When a machine can investigate 100% of alerts, the filter has no reason to exist. The pipeline collapses to Detect → Investigate.

Concede all of it. Then push one step further.

The filter is dead. The verdict isn’t.

At the end of every one of those deep machine investigations is still a call: true positive or false positive, escalate or close. And the verdict is a much older artifact than the triage queue. It’s a lie about how investigations actually end.

”Benign” is where you stopped, not what you found

An investigation has no natural stopping point. I’ve argued this at length: you don’t finish one, you abandon it. The budget runs out, the clock says move on, the case looks boring enough. And every event you haven’t pulled yet could overturn the verdict you’re about to write. There is no query after which “benign” becomes safe to declare.

So “false positive” isn’t a finding. It’s a place you stopped. A real attack that got investigated and missed produces the same artifact as a correct close: same green check, same tidy narrative, same closed ticket. The verdict records where the investigation ran out of money. It says nothing about what was there.

That’s the dishonest move buried in the binary. It takes “we stopped looking” and prints “there was nothing to find.” Every closed alert ships with a certificate the investigation never earned.

Ask for the risk, not the answer

The fix is small. Stop asking the investigation for a verdict. Ask it for the risk it couldn’t rule out.

Call it residual risk: the chance a real threat survived your search, weighted by what that threat could reach. It comes from two things.

The first is the timeline itself: how much risk the case carries before you touch it. A login from a new device into a crown-jewel system, with three plausible stories that all fit the evidence, is inherently risky. A well-understood alert on a low-value box is not. Breadth, novelty, blast radius, the number of open hypotheses. This is the same list he reaches for when he says human reviewers should oversample “crown-jewel assets, first-time-seen behaviors, and anything closed suspiciously fast.” He’s eyeballing a score we should be computing.

The second is depth. Every rung of investigation you climb retires some of that risk: rules out a hypothesis, widens the evidence window, kills a story that would have hurt. A shallow look retires little. A deep one retires more. Neither reaches zero, because the search was unbounded to begin with. You drive residual risk down toward a threshold you can live with, and you report what’s left.

Triage stops being a classifier. It becomes a measurement.

No, it isn’t “one minus confidence”

Start with the fair objection: isn’t residual risk just one-minus-confidence in a new coat of paint? Every competent agent already emits a confidence score. Subtract it from one, call the remainder risk, and go home.

No. Confidence and residual risk answer different questions, and they routinely point in opposite directions.

Confidence is a claim about the verdict. It’s computed over the evidence the agent actually gathered: given what I looked at, how sure am I of the call? Residual risk is a claim about the world. It’s the risk that survives because the agent stopped: given everything I didn’t gather, how much could still be wrong?

The difference shows up the moment you dig. Confidence often rises on a shallow look, because you haven’t yet surfaced the fact that ruins the story. The event that adds nothing to your confidence and everything to your risk is the one sitting in a window you never queried. Confidence measures your certainty about the map you drew. Residual risk measures the territory you never walked.

So the two numbers have to be built differently. Confidence is calibrated against the evidence in hand, which means a system can be perfectly calibrated and still blind by design: calibrated to the wrong reference class, what it saw rather than what was there. Residual risk has to be calibrated against ground truth you inject from outside. And it moves differently. The one thing that only grows with depth is your search, so the chance a real threat slipped past it can only fall. But residual risk is that shrinking miss-chance times the odds a threat is there at all, and those odds aren’t monotonic: hit the impossible-travel event and they jump, and residual risk jumps with them. So it doesn’t glide down. It decays in expectation, which is why another rung is worth running, but any rung that surfaces something spikes it. That spike is the point, not a defect: a number that only ever fell would be telling you that investigating a live breach makes it safer.

Two curves plotted against depth of investigation. Confidence starts high on a shallow look, dips when a complicating fact surfaces, then recovers. Residual risk trends downward as investigation deepens, but spikes upward at that same complicating fact before resuming its fall toward a risk-appetite line, and never reaches zero. A highlighted zone at shallow depth marks where confidence is already high but so is residual risk: the close that hides a miss.

Watch the two numbers move on Dana’s sign-in, the same alert from the last post.

Rung one, location history: she’s used this city before. Confidence climbs to 0.9 and closes benign. Residual risk stays high, because you’ve read two days of an account that reaches payroll and none of what the session actually did.

Rung two, session actions: no new OAuth grant, no forwarding rule, no second MFA device. Confidence barely moves. Residual risk drops hard, because the stories that would have hurt are now dead.

Rung three, the full nine-day window: day nine holds an impossible-travel event the first looks never queried. Confidence doesn’t ease off. It collapses. Residual risk spikes past where it started.

Same alert, three rungs. Confidence rose then cratered; residual risk fell then spiked. They were never the same number, and only the second one was tracking whether you should stop.

Attack-chain hypotheses fanning out from Dana's sign-in, with depth increasing left to right across three investigation rungs. Credential stuffing is ruled out at rung one, takeover persistence at rung two, and the benign-travel leading story is disproven at rung three, at the same moment a dormant prior-compromise chain ignites into an impossible-travel finding. One faint chain, a more patient adversary, is never ruled out and runs off the right edge. A readout beneath the rungs reads residual high, then falling, then spikes. Every rung retires a hypothesis, so residual risk falls. But a chain you’d stopped pruning can reignite, and one you can’t reach never dies, which is why the number spikes and never hits zero.

Where residual risk pays rent

This isn’t just a more honest label. It gives his next argument the instrument it needs.

Kill the filter and deep investigation gets expensive, so someone has to decide how deep the machine goes on which alerts. He names this precisely: the new triage in a suit, where the decision “moved from a human clicking a queue to a policy sitting in a config file,” an unexamined control “that quietly decides what you never look at.” His prescription is governance: give the policy an owner, version it, “write the shadow down,” review it quarterly. All of that is right, and all of it needs something to act on. A static config can’t write its own shadow down. If the rule is “phishing gets depth two, off-hours gets the cheap path,” nobody sees what that costs until it’s a breach.

Residual risk turns the static rule into a loop. You don’t set depth by alert class. You keep investigating until residual risk drops below your appetite, then you stop. The shadow he says to write down is no longer something you have to go hunting for. It’s the number on every case you closed early. Under-investigation stops being invisible, because leftover risk is the output.

Be precise about what that appetite is set on, because there are two candidates and they behave nothing alike. Residual risk is a probability times an impact: the chance a threat survived, times what it could reach. Now run Dana’s exact login two places. On a decommissioned test box, the miss-chance is just as real, but the blast radius is nil, so expected loss clears your appetite after one rung and you stop, honestly. On her payroll-reaching account, the impact term is so large that even a small surviving miss-chance holds expected loss above the line no matter how deep you dig. You investigate to the bottom of your budget and the number still reads elevated. That isn’t the model failing. It’s the model refusing to certify “clear” where being wrong is catastrophic, and handing you back the decision that was always yours: carry this risk knowingly, or spend more to shrink it. The binary verdict never showed you that choice. It just closed the ticket.

It also settles the budget fight he sees coming, “cost per investigation” against detection coverage, two teams arguing in two currencies. Residual risk collapses them into one question: is another rung worth it? Only if the risk it retires is worth more than the tokens it burns. That is a decision you can actually make.

One honest caveat, because the argument dies without it: this doesn’t make the control ungameable, and built carelessly it can be worse. An adversary who understands the model will try to make their activity produce a low-risk-looking timeline, the way they learn to land in a low-severity bucket today. A residual-risk score frozen into a fixed function hands them a stationary target they can probe offline until they find the path that reads as benign, the way evasion attacks farm a static classifier. So the defense isn’t to pretend the gate can’t be gamed. It’s to keep it moving: scoring tied to freshly injected canaries and ground truth that shifts under the attacker’s feet, rescored as the case develops. A live gate driven by evidence about the actual case is a harder mark than a guessable rule about its metadata. But “harder to game than the thing it replaces” is the honest claim, not “impossible.”

Calibrating the number

All of this rests on one figure I’ve been hand-waving: how much risk does a given depth actually retire? Say “residual risk: elevated” and you’ve dressed a guess in a lab coat.

You calibrate it the only way the domain allows. You can’t prove benign from inside the data, so you smuggle a positive in from outside: plant a threat you control, run the investigation, and see whether that depth of search catches it. He has the same instinct, calling them canary cases: “known-bad cases with deliberately wrong agent verdicts” seeded into the review queue to measure the reviewers. Run that at every depth and you get what the model needs: an empirical map from how far you looked to how much you actually caught. That is what gives residual risk error bars instead of vibes.

And keep the error bars honest. The planted ground truth is itself an approximation, and residual risk is an estimate of a quantity that is formally unbounded. Reporting it as a crisp 0.37 would be its own kind of lie, the false-precision cousin of the binary verdict. A band is the honest shape: low, elevated, high, with a stated confidence in the band.

Then confront the sharpest objection to all of this, the one that turns the argument back on me. Canary cases can only test threats you can imagine and synthesize. The residual risk you most want to catch is the day-nine event you never thought to query, the novel move outside your model of attack. So the empirical map from depth to catch-rate is calibrated on the reference class of plantable threats, “what I could think to plant,” and says nothing about the one that was actually there. That is structurally the same blindness I just pinned on confidence, and I won’t pretend otherwise. I’ve conceded exactly this before: a planted flag measures whether you catch the attack you already imagined, not the one outside your imagination. But the two blindnesses differ in the way that matters. Confidence hides its reference class and reports as truth. The canary map declares its reference class and reports as a bound. One is a blind spot you don’t know you have; the other is a coverage boundary you can see the edge of and widen on purpose. Residual risk doesn’t escape the limit. It makes the limit legible, which is the only kind of progress this domain allows.

The output that admits what it doesn’t know

The binary verdict survived this long because it feels like an answer. Residual risk feels like an admission, which is exactly why it’s the better output. It keeps on the record the one thing every closed alert hides today: how much you didn’t rule out when you decided to stop.

A fair challenge here: haven’t I just moved the verdict, not killed it? Something still has to happen. You page or you don’t, you contain or you don’t, and “stop when residual risk is under appetite” is itself a decision. True. The action that ends an investigation is irreducibly binary, and no framing dissolves that. But look at what the decision is made of now. The old verdict was improvised per alert, sealed inside a closed ticket, and self-certifying: it declared “nothing here” on its own authority. The appetite line is set once, in policy you can version and an owner can defend, applied to a measured number, with the risk you chose to carry written down beside the close. The decision didn’t disappear. It came up out of the ticket and into the open, which is the whole content of Chuvakin’s “write the shadow down,” and the opposite of smuggling the call one level deeper.

He’s right that triage must die. But the queue was only ever the symptom. The disease is a SOC that mistakes running out of budget for finding nothing. So kill the verdict that hides the mistake, and keep the one honest thing left underneath it: the risk you decided to live with, on the record, where someone can finally argue with it.