There's an article on this site I could have quietly deleted. I wrote it about five months ago, and it is wrong in three specific ways I can now name. Deleting it would be the easy move: scrub the evidence, pretend the earlier version of me had better judgment than he did.
Instead, come look through the window. That is the arrangement in this lab — the work sits behind glass on purpose, and that includes the experiments that didn't hold up.
The piece was about prompt injection: the real and unglamorous truth that a model built to be helpful can be walked into being helpful toward the wrong end. That part I would still write. The frame I built around it is what I got wrong, and the three mistakes are worth more, disassembled, than the article ever was intact.
One: I hung it on a headline. A frontier lab had been in the news that week, and I used it as the hook, because at the time I was testing whether timing and search ranking were the game. They weren't — or they stopped being, because the internet that rewarded that play was already being eaten by the machines I write about. Hook a technical argument to this week's headline and it dates to this week. Good security writing should read the same in a year. That was a lesson about the medium, and I learned it in public.
Two: I turned and spoke to the attackers. Somewhere in the piece I addressed them directly — you already know this works. It is a cheap move. It sounds hard, it feels like access, and it is the tell of someone performing for the wrong room. I do not write to attackers. I write to the people trying to stop them and the people deciding who to trust with the problem. The moment your security writing starts winking at the adversary, you have told the reader which chair you are sitting in.
Three, the real one: I printed the recipes. Worked examples, copy-paste ready, the kind of thing that reads as generous and is actually just loud. Here is what I would tell that version of me, and it is the whole reason this piece exists instead of a delete button.
You do not have to show the payload to prove you understand it. Showing it usually proves the opposite — that you have mistaken the exploit for the insight. The exploit is the cheap part. The understanding is knowing why it works, what it costs the defender, and what actually closes it. So I am not going to reprint them, and that is not a dodge. It is the point. Watch what is left when you take the payloads out. There turns out to be more, not less.
So here is the shape of what mattered, without the reagents.
