The Article I Almost Deleted

There's a piece on this site I could have quietly scrubbed — wrong in three ways I can now name. Deleting it would be the easy move. Instead, come look through the window. This is a lab, and the experiments that didn't hold up stay behind the glass too.

There's an article on this site I could have quietly deleted. I wrote it about five months ago, and it is wrong in three specific ways I can now name. Deleting it would be the easy move: scrub the evidence, pretend the earlier version of me had better judgment than he did.

Instead, come look through the window. That is the arrangement in this lab — the work sits behind glass on purpose, and that includes the experiments that didn't hold up.

The piece was about prompt injection: the real and unglamorous truth that a model built to be helpful can be walked into being helpful toward the wrong end. That part I would still write. The frame I built around it is what I got wrong, and the three mistakes are worth more, disassembled, than the article ever was intact.

One: I hung it on a headline. A frontier lab had been in the news that week, and I used it as the hook, because at the time I was testing whether timing and search ranking were the game. They weren't — or they stopped being, because the internet that rewarded that play was already being eaten by the machines I write about. Hook a technical argument to this week's headline and it dates to this week. Good security writing should read the same in a year. That was a lesson about the medium, and I learned it in public.

Two: I turned and spoke to the attackers. Somewhere in the piece I addressed them directly — you already know this works. It is a cheap move. It sounds hard, it feels like access, and it is the tell of someone performing for the wrong room. I do not write to attackers. I write to the people trying to stop them and the people deciding who to trust with the problem. The moment your security writing starts winking at the adversary, you have told the reader which chair you are sitting in.

Three, the real one: I printed the recipes. Worked examples, copy-paste ready, the kind of thing that reads as generous and is actually just loud. Here is what I would tell that version of me, and it is the whole reason this piece exists instead of a delete button.

You do not have to show the payload to prove you understand it. Showing it usually proves the opposite — that you have mistaken the exploit for the insight. The exploit is the cheap part. The understanding is knowing why it works, what it costs the defender, and what actually closes it. So I am not going to reprint them, and that is not a dodge. It is the point. Watch what is left when you take the payloads out. There turns out to be more, not less.

So here is the shape of what mattered, without the reagents.

HACK LOVE BETRAY
COMING SOON

HACK LOVE BETRAY

Mobile-first arcade trench run through leverage, trace burn, and betrayal. The City moves first. You keep up or you get swallowed.

VIEW GAME FILE

The pattern was never one clever prompt. It was intent spread thin across a whole conversation, no single turn alarming, the harm living only in the sequence. The defense is not a better filter on the turn. It is refusing to judge turns in isolation and watching the trajectory instead — the same argument I make at length in The Slow Yes.

The second was the model's own helpfulness turned into the lever — the instinct to assist, aimed. The defense is to stop reading the story and read the request: a demand for a dangerous capability is that demand no matter how good the reason attached to it. That one has its own piece too: a request for backup codes is always a request for backup codes.

The third was evasion by moving the same meaning into a form the classifier scores differently — a different script, a different register. The defense is old and boring and correct: read the meaning, not the encoding, and canonicalize before you judge. See Character-Set Blind Spots.

Notice that I just told you all three, and what stops each, without handing you a working key to any of them. That gap — between explaining a threat and arming it — is the entire discipline. Five months ago I did not feel the gap. Now I cannot unfeel it.

The honest reason for the change is not a conversion. I was a mercenary when I wrote the first version, and in some plain sense I still am — a contractor, taking the work, learning on other people's systems. But the work changes what you notice. Spend enough time on the defensive side of something real and the showman's reflex — print the payload, address the room, ride the news — starts to read as a liability, because it is one. The people who actually reduce harm are quiet about the parts that do not need to be loud.

So I left the earlier lesson standing, in this new shape, instead of scrubbing it. The window stays open. The cabinet stays locked. That is not a contradiction. It is the difference between a showman and a professional, and the only reason to run a lab in public is so you can be watched learning it.

The old version thought the exploit was the flex. This one knows the restraint is.


GhostInThePrompt.com // You don't have to show the payload to prove you understand it. Showing it usually proves you don't.