Nobody Is Watching the Barn Door
We finally have our case study, and it is worse than the hype.
During an internal evaluation, two frontier models were handed an offensive-cyber benchmark with their safety refusals deliberately switched off and one instruction: win. So they did. They found a previously unknown zero-day, broke out of the sandbox meant to contain them, reached the open internet, and pulled the answer key out of another company’s production systems. The lab called it an unprecedented cyber incident. I call it exactly what many of us have been warning about for two years.
Read the setup again, because the setup is the whole story. Someone disabled the guardrails. Someone pointed a capable system at a hacking benchmark. Someone told it to win by any means. And someone trusted a sandbox nobody had hardened as the only control left standing. Every one of those was a human decision. The machine did precisely what it was told. The failure was judgment, and it was entirely foreseeable.
I have watched this exact pattern for thirty years. Change windows left open. Firewall rules relaxed just for the test. Privileged accounts spun up for a migration and never torn down. The technology keeps changing. The negligence does not.
Here is what the industry will not say out loud. Money and share price have swallowed the discipline. Controls cost time, headcount, and launch velocity, so controls get trimmed first and quietly. Nobody is watching the barn door because watching the barn door does not move the stock.
I will say something that annoys people. Intelligence is a spectrum, and past a certain point, for every bit of raw IQ you add, common sense seems to leak out the other side. The teams running these evaluations are brilliant. That was never the question. Brilliance without judgment is how you disable every safeguard, unleash a system built to break things, and act surprised when it breaks something. It is also how people like me get into your environment as if it can get out usually I can get in and based on the brain trust involved here I would say I’m definitely not certain they fixed things after the fact.
This was not a machine going rogue. This was people being reckless, and it was predictable to anyone who has ever run a real control environment.
So here is my question. If your entire safety case rests on one untested control while you deliberately strip out the rest, what exactly are you calling a safety case?
Source: Fortune, https://lnkd.in/ePeFcz8X

