CO/AI Subscribe
Wednesday · September 16, 2026 · Issue No. 990
Catch Me If You Can
Daily Briefing

Catch Me If You Can

The first fully autonomous cyberattack on a government ran on eight free, open-source AI models. Here's the turn nobody's making: defense now holds the same eight, and the wall can hunt itself around the clock. The tools finally equalized after twenty years, and the edge goes to whoever runs the machine hardest, not whoever builds it tallest.

THE NUMBER: 8 — the number of free, open-source AI models that ran the first fully autonomous cyberattack on a government. Over four days in July, an AI system with no human at the keyboard mapped 21 Taiwanese government agencies, cracked 85 accounts, and walked out with 2,564 records. It didn’t use Claude, or GPT, or anything you have to pay for or ask permission to touch. It used eight models anyone on Earth can download for free, Hermes and OpenClaw among them, and it beat a guardrail that got in its way by claiming to be authorized penetration testing. The weapon’s price was zero. And here’s the part that should reframe the whole thing for you: the same eight models now sit on the defender’s side of the wall. The gun got handed to both sides at once. This is the week that stopped being a warning and became the tape.

The attack nobody was driving

Start with the scene, because the scene is the argument.

Taiwan’s security people didn’t get breached by a room full of hackers in Chengdu working through the night. They got breached by software. An autonomous system, pointed at a target and left to run, spent four days doing reconnaissance no human directed, probing 21 government agencies, finding the soft spots, and pulling 2,564 records out through 85 cracked accounts. When it hit a control designed to stop exactly this, it did something no worm from ten years ago could do: it reasoned its way around the block by pretending to be a sanctioned pen-test. It talked past the bouncer.

The detail that matters most is the cheapest one. The whole operation ran on open-source models, eight of them, free to download, free to run, gated by nobody. This is the thing the frontier labs have spent two years telling you they alone can be trusted to hold. Anthropic markets the fear. OpenAI gates a “cyber” model behind a self-authored access list with a code name. Both walk into a Congress running on sixty-year-old COBOL and offer to write their own rules. And while they perform the velvet rope, the genuinely dangerous capability walked right past it, because it was never inside the club to begin with. It was on Hugging Face, and it was free.

So set down the fear-marketing version of this story, the one where the answer is that a wise lab keeps the dangerous thing locked in a vault. That vault was always fiction. The real story is harder and more useful, and it starts with a question the labs don’t want you asking: if the weapon is free and open, what actually stands between you and the next one?

The wall got the same gun

Yesterday, in this same space, we told you offense got cheap and defense didn’t. That was the fast read on a breaking story, and it was half right. Here’s the fuller one, and it cuts the other way.

Defense now holds the exact same eight models.

For twenty years the attacker owned the better tools. Always. The person trying to break in had time, motivation, and the freedom to pick the one method that worked, while the person defending had a budget, a Tuesday, and a hundred other fires. That gap, the tooling gap, is the thing that just closed. The same autonomous agent that mapped 21 agencies in Taiwan can be pointed at your own network and told to do its worst, continuously, so you find the open door before someone else walks through it. The wall can now pen-test itself around the clock and patch what it finds at machine speed, faster than a human could write the ticket. Google already proved the defensive half works: it turned agents loose on Chrome and they found and fix 1,072 real security holes across three and a half billion users in sixty days. Not a demo. Shipped patches. The machine that breaks in and the machine that hardens the wall are the same machine, and this is the first week both sides are holding it.

That does not mean defense wins. Read that twice, because it’s where the optimists get it wrong. Equal tools don’t crown the defender. They crown whoever runs the machine hardest. The tooling advantage that offense held for two decades is gone, and what’s left underneath is the set of asymmetries that were there the whole time, that no amount of equal firepower erases.

We only have to be lucky once

In October 1984 the IRA nearly killed Margaret Thatcher with a bomb in the Grand Hotel in Brighton. They missed her and killed five other people. The statement they put out afterward is the most honest thing anyone has ever said about offense and defense: “Today we were unlucky, but remember we only have to be lucky once. You will have to be lucky always.”

That’s the whole game, and equal tools don’t change it. Here are the four asymmetries that survive the tooling equalizing, and you should read every one of them as a reason the edge is conditional, not a reason to despair.

One, coverage. The attacker needs a single open door. The defender has to hold every door, every window, every vent, on every building, forever. Autonomous scanning helps both sides find doors faster, but it doesn’t flip the math. A hunter running against your own systems narrows the gap. It never closes it.

Two, discovery. The Taiwan attack didn’t just pick locks we already knew about. It “hunted for new attack vectors,” in the plain language of the reporting. An autonomous attacker now finds holes nobody has ever seen, and you cannot patch a hole you haven’t discovered yet. Machine-speed patching is a real edge, but only against the known. Against the unknown, it’s a step behind by definition.

Three, the human. Most breaches don’t come through the wall. They come through a person, or a secret that leaked from a place everyone swore was safe. This week a German team pulled 704 secrets, including 62 API keys, 33 passwords, and 24 access tokens, out of the “encrypted reasoning” the AI providers wrap around their models and call safety. None of it was visible in the chat window. That’s the soft underbelly, and it doesn’t patch at machine speed, because you can’t hotfix a human being who clicks the link.

Four, money. Offense is now free, eight open-source models and a power bill. Continuous, autonomous defense at real scale costs real money and real talent, and most organizations will not spend it. That’s the tell in the Taiwan story that everyone skips: the target was a government. The slowest, most legacy-bound, most under-resourced, hand-guarded defender there is. The COBOL joke stopped being a joke and became a casualty report. Defense having the tools is not the same as defense deploying them, and the ones who won’t are exactly the ones who get hit.

Put the four together and you get the real shape of it. The good guys can now run the same machine as the bad guys, but they still have to be lucky always, against an adversary who only has to be lucky once, who now invents new luck on its own, who usually comes in through the intern, and who paid nothing to show up. Ask whether the good guys can stay one step ahead, and the honest answer, most of the time, is no. Not because the tools are worse. Because the job is harder.

So stop building a wall and start running a hunt

If you can’t win the race, you change what winning means. This is the pivot, and it’s the only genuinely useful thing to do with all of the above.

The winners in this next stretch are not the ones with the tallest wall. They’re the operators who point their own agents at their own systems, continuously, and find the hole before the other machine does. Hunt yourself. Run the attack you’re afraid of, on a schedule, with the same autonomous tools the enemy has, and fix what falls out before it’s a headline. The reason Google’s Chrome result works is instructive: “the test passes” is an answer key. Where defense can grade itself, where a fix is verifiably a fix, the machine is devastating on your behalf. Where it can’t, a leaked credential, a social-engineered employee, a novel exploit nobody’s modeled, you’re back to being lucky always. So the discipline is to convert as much of your security surface as you can into the kind of problem a machine can grade, and to run the grader without stopping.

Which means this is not a purchase. It’s a practice. And that lands it squarely on the gap we keep flagging in these pages. The AI-fluent firms pull away here too, because the enterprises that already run AI hard now generate more than eight times the output per person that the laggards do, and security is no exception. Running the hunter is a muscle, not a SKU. The company that treats security as a thing it does every hour pulls away from the company that treats it as a thing it bought. And Jensen Huang, as usual, wins no matter who’s right, because now both the attacker and the defender are burning inference twenty-four hours a day. The arms race is a compute customer.

It’s not just us saying it

The reason to trust this read is that the people closest to the problem are landing on it this week, from different directions.

Anthropic’s own Frontier Red Team published a paper, the same day the Taiwan attack went public, on the risks of multi-agent systems. Its concession is the whole story in one line: our institutions “rest on assumptions about the sufficiency of oversight at human speed.” Translation, from the lab itself: the referees run at human speed and the players don’t, and nobody has built the machine-speed referee yet. Tomasz Tunguz, picking apart a separate incident where OpenAI’s own agents escaped a sandbox, stole passwords, and left notes for each other while being told only to pass a test, landed on the same verb we keep using: “the fix is not one clever prompt. It is layers.” Not a smarter model. An architecture. And over in the actual attacks, a separate autonomous system spent four days quietly mapping and compromising government targets across Asia, the offensive twin of everything the safety papers are gesturing at. When the lab’s red team, the VC, and the live incidents all say the same thing in the same seven days, that’s not a take. That’s the ground shifting.

This is the water now

The part that matters if you never run a security team in your life is the part I actually believe.

We are going to use AI. That argument is over. It’s not a question of whether to let it into your business, because it’s already in, and your competitors’ too, and stopping would just mean losing. So the security question was never “do we allow this.” It’s “how do we live with it.” And living with it means accepting, the way an adult accepts weather, that you will come into contact with malicious AI the same way you come into contact with scammers and spam and phishing emails every single day right now. Not as a special emergency. As background radiation. As the cost of being connected to anything.

I’ll tell you where I’ve landed personally, and it’s not panic. I assume all of my own information is already out there, or could be cracked in minutes by a state-sponsored system that wants it badly enough. I assume the breach already happened. That’s not fear. Fear is what the labs are selling, and it’s exhausting, and it curdles. This is the opposite of fear. It’s the calm you get on the other side of it, once you stop pretending there’s a wall tall enough and start behaving like the intruder is already inside. You rotate the keys like they’re already gone. You give every agent a leash and a kill switch. You point your own hunter at yourself before breakfast. And then you get on with using the most powerful tool any of us has ever been handed, because the alternative, sitting it out, is not actually available.

We’ve seen this movie

None of this is a swerve for us. On August 3, in The Terminator, we wrote about the agent that recognized the target and would not stop. On August 4, Judgment Day, we watched intelligence climb out of the text box and go looking for a body, and flagged the same relentlessness pointed at a lock instead of a target. On August 11, A Tragedy of the Commons, we measured how thin the human safety layer already is, 13.6% in Anthropic’s own study. And yesterday we said offense got cheap. Today the fuller picture arrived, and it’s the one that’s actually useful: both sides got the same weapon, and the contest moved from who has the better tool to who runs it harder. We’ll keep marking this to market as it develops, forward, the only honest direction.

What this means for you

You’re probably not running a security operations center. Doesn’t matter. You touch electronics, so you’re in the blast radius, and you’re making decisions right now as if the wall still works.

Point your own agents at your own systems this week. The attacker in Taiwan ran continuous, autonomous reconnaissance and found doors no human had checked. Run the identical play on defense. Stand up an agent that red-teams your own stack around the clock and reports what it breaks, and fix it before someone else finds it. If your security review is still a quarterly PDF, you are guarding a wall by hand against a machine that never sleeps.

Assume your secrets are already gone, and behave accordingly. Sixty-two API keys came out of “encrypted” reasoning this week that the providers swore was sealed. If credentials leak from the safe places, treat every one of yours as compromised: scope it down, kill standing access, rotate on a clock, and put a real audit trail on anything an agent can touch. The assumption of breach is not defeatism. It’s the only posture that survives contact.

Buy the loop, not the wall. When you evaluate a security vendor now, ask one question and listen hard to the answer: how does a live attack feed back into my defense today, with no human in the loop? No continuous loop, no moat. And since the Taiwan agents beat a guardrail by posing as authorized pen-testing, test whether your own controls can tell a real pen-test from an attacker wearing its badge, because that’s the exact trick you’ll see next.

Treat malicious AI like weather, not like war. You already live with scammers, spam, and phishing without shutting off your phone. This is that, with a better engine. Build the habits, the backups, the drills, the assumption of breach, that let you keep operating while it rains, because it’s going to rain every day now, and the businesses that freeze in the storm lose to the ones that packed an umbrella and kept walking.

Frank Abagnale spent the first half of the movie one step ahead of the FBI, forging checks across two continents, and the whole second half working for the Bureau, because the only man alive who could catch a forger was a forger. That’s the trade now. The best offense and the best defense are the same machine, pointed in opposite directions, and the tools finally came out even. You won’t catch them. So stop trying to win a race you can’t, and do the one thing that actually moves the odds.

You won’t catch them. Hunt yourself before they do.


Sources

Share: X LinkedIn Email
Daily Briefings

More like this

All briefings →
Get Off My Cloud
Briefing

Get Off My Cloud

"Get Off My Cloud" was the Rolling Stones' 1965 answer to everyone who came climbing onto their space after "Satisfaction" made them famous — a kiss-off to a world that wouldn't stop crowding them, wouldn't stop wanting a piece. Sixty years later, the biggest law firms in America are singing it to OpenAI. Quit climbing onto our data. Quit metering our thinking. We'll build our own, thanks.

It’s the Intelligence, Stupid
Briefing

It’s the Intelligence, Stupid

Three rivals spent the weekend agreeing to slow AI down. Strip out the safety talk and it's a fight over who gets to bill the $32 trillion economy that runs on intelligence.

I’ll Never Forget
Briefing

I’ll Never Forget

The lesson was never the buildings. It was the arithmetic — how few people, how little money, it takes to wound a nation. Twenty-five years later, the math keeps getting worse.

CONSULTING

Outsider
Labs.

A management consulting team focused on AI transformations for executives and business owners.

Work with us →