← Operating Log
AI SafetyAgent Architecture

A Second Net Is Not a Floor

The agent-safety debate is stuck at the perimeter, arguing about whose net is bigger. But a perimeter bolted on after the fact is a second net — and a second net is not a floor.

Fire hoses racked in tight, mirrored rows inside a warship's steel compartment, pale couplings facing out above a polished deck: standing damage-control gear, staged and ready on a floor built to hold.

There's a picture making the rounds: a "regulated agent" behind a compliance firewall, a filter in the middle, an "unleashed agent" on the other side firing off attacks. Agent wars. Pick a side.

The instinct behind it isn't wrong. The visible AI race, whose chatbot is smartest and whose copilot is fastest, really is the wrong half to be watching. The half that decides whether you can put an agent anywhere near money, records, or customers is somewhere else entirely. Credit where it's due: that part lands.

But the picture itself is the wrong model, and it's worth saying why, because a lot of budget is about to be spent on the wrong thing.

A firewall is a perimeter. A filter is a perimeter. Both assume the game is detection: catch the bad action on its way through, block it, stay one step ahead. That's an arms race, and the defender loses arms races eventually; you only have to be wrong once. A perimeter bolted on after the fact is a second net. A second net is not a floor. If the only thing standing between an agent and an irreversible action is a classifier's judgment call, you don't have a control. You have a probability.

So the question isn't "how do we block bad agents." It's narrower and harder: what can this agent actually reach, and what stops it when it goes wrong, not if?

That reframes the whole problem from detection to structure. You stop asking the model to behave and start removing its ability to misbehave in the first place. In practice that's a short list of properties, and the point of each is that you can check it, not just claim it:

Agents propose; people commit. The irreversible step is not the agent's to take: the send, the spend, the write to the system of record. It drafts; a person commits.

Fail-closed. On error, on timeout, on ambiguity, the agent stops. It does not proceed on a guess.

Least privilege on the irreversible. Destructive tools are withheld until explicitly granted and scoped to the task, not left standing open because they were convenient during the demo.

A deterministic floor under the probabilistic layer. The model is allowed to be clever, and occasionally wrong, on top. Underneath it sits a layer that is neither: rules that hold regardless of what the model decides. The classifier is the second net. The floor is the floor.

An append-only record tied to a verifiable identity. After the fact, you can name the identity behind every action and read what it did, from a log that is added to, never quietly edited.

None of that is a manifesto. It's how I run my own operation. The system that drafts and files my work halts when it's handed an order that isn't stamped for the session in front of it. It won't claim its own identity; I register it. It cannot commit or ship its own work. Those aren't aspirations on a slide; they're the conditions under which it is allowed to run at all.

So when the feed fills with agent wars, I read it as a sign the conversation is still standing at the perimeter, arguing about whose net is bigger. The work that decides whether agents are deployable into anything that matters is happening one layer down, at the control plane. It's quieter because it's less cinematic. Just an agent that can't reach the thing it shouldn't, and a record that says so.

That's the half worth watching. It's the half I build.

Steady on. ⚓


This is the operator's-desk version of a Blue Jacket Consultancy position. © 2026 Blue Jacket Businesses LLC, d/b/a Blue Jacket Consultancy; shareable in unmodified form with attribution, not for use in machine-learning training without written permission.

Subscribe to the Operating Log

Field notes on building governed, production AI systems — new issues to your inbox.

Subscribe on Substack →

Handled by Substack. Unsubscribe anytime.