CO/AI Subscribe
Wednesday · September 16, 2026 · Issue No. 990
Code Red
Daily Briefing

Code Red

Everyone's getting AI agents. Nobody's getting the training to manage them, the meter to price them, or a straight answer about when to trust them. The one job the agents didn't eliminate is babysitting the agents.

THE NUMBER: 30 minutes → 8 hours. That’s what happened to Jason Lemkin’s day. A year ago, running his AI agents at SaaStr took about thirty minutes — total, across the whole shop. This week it takes eight hours a day, per person, to manage twenty-plus of them. Three humans, busier than they were when they had twenty employees. Read that ratio the way an operator reads it, because it breaks the entire pitch. The agent was pitched as pure upside: hand off the work, get your time back. What Lemkin actually got was a promotion he never asked for — from person who does the work to person who has to have an opinion about every decision a tireless, eager, literal-minded machine makes all day long. The hours didn’t go down. They went up sixteen-fold, and they moved from doing to supervising. Everything below is why that happened to him, why it’s about to happen to you, and what the job actually is now.

⚖️ The Signal: No, It Cannot

There’s a moment in A Few Good Men that everyone remembers wrong. They remember Nicholson bellowing “you can’t handle the truth.” The scene that actually matters is quieter, and it’s Kaffee taking apart Lieutenant Kendrick — Kiefer Sutherland playing every hard edge of Jessup with none of the polish. Kaffee sets the trap gently. Surely, he says, a Marine of Dawson’s intelligence can be trusted to decide on his own which orders are the important ones, and which might be, say, morally questionable? Can he? Can Dawson decide on his own which orders to follow? Kendrick doesn’t hesitate. “No, he cannot.”

Now swap the Marine for the model, and you have the whole AI story this week. Surely an agent of superior intelligence can be trusted to decide on its own which orders are the important ones, and which might be morally, or financially, questionable? No. It cannot. And here’s the part that should stop you: we just handed a few hundred million of these things to people who have never managed a single human being in their lives.

That’s the story nobody’s pricing correctly. Not “agents are amazing.” Not “agents went rogue.” The real one sits in between, and it’s an org-chart problem wearing a technology costume. An agent is an employee. It shows up, it takes instructions, it makes decisions, it costs money, and it can hurt you. We’ve spent a hundred years building an entire discipline — management — around the hard problem of getting fallible workers to do useful things without blowing up the company. And then this year we handed that exact problem to founders, analysts, marketers, and solo operators, and told them the help was free and needed no supervision. Both halves of that sentence are wrong.

ep 16 The Future-Proof Pod

Ep 16 – Google’s Dream Team Just Quit, and Nobody Can Find the AI Bear Case

Four top Google AI researchers walked out the same day. Anthony Batt and Harry DeMott on what that exodus actually signals, and why the industry’s doom talk might be more marketing than warning.

🧑‍💼 The Employee You Didn’t Know You Hired

Lemkin wrote up his week and it reads like a management confessional. A year ago his agents were basically better software — a better outbound tool, a better inbound tool. You configured them, checked their work now and then, moved on. When they failed, they failed loudly: they errored, they stopped, you knew. Now, he says, “they hand you something that looks finished.” And they make decisions constantly, most of which you never explicitly authorized.

His line is the whole issue, so here it is straight: “When an agent can only do a task, you check the task. When it can decide, you need an opinion about every decision.” Every morning he logs into his AI VP of Revenue and it has new things it wants to build — not tasks it wants to run, ideas it wants to ship. It raises its hand every single day like a team member. Anyone who’s run a company knows that feeling: the employee with two hundred good ideas and no sense of which three matter. Except this employee generates the backlog at machine speed, and the only thing standing between the backlog and your production system is your attention.

That’s the trap, and it’s the opposite of the sales deck. The constraint used to be capacity — you couldn’t build the thing because you didn’t have the engineers. Now the agent removes the capacity constraint and installs a new one in its place: you. Your judgment, your review, your willingness to say no forty times a day. The bottleneck moved up the stack from hands to head, and the head it moved to belongs to someone who was told they wouldn’t have to be in the loop at all.

And who is that someone? Not a trained manager. That’s the quiet scandal. The people getting handed twenty agents are not, by and large, people who’ve ever run twenty of anything. They have no reps at writing a clear instruction, no reps at catching a confident subordinate who’s quietly wrong, no reps at the specific art of delegating without abdicating. Managing people is hard enough, and those are people who at least run on the same operating system you do. We’re asking a whole population of non-managers to manage a workforce, with no training, no playbook, and — this is the part that should really bother you — no idea what the workers cost.

💸 The Salary Nobody’s Reading

Here’s where the finance brain kicks in, because the cost story is worse than the time story.

The founding assumption of this entire boom is that the agent is cheaper than the human. Sometimes it is. But that’s an assumption, not a fact, and almost nobody has checked it against their own bill. An agent that runs continuously, spins in circles, retries, and pours tokens into work nobody graded is not cheaper than a person. It can be dramatically more expensive, and the meter has no ceiling and no conscience. Uber torched its entire 2026 AI budget in four months — five thousand engineers leaning on Claude Code — and its own COO admitted the spend was getting “hard to justify.” One reported coding session ran $1,200 for an afternoon. That’s roughly six hundred dollars an hour for a worker you were told was basically free.

Sit with the asymmetry. When you hire a human, you know the number. It’s on the offer letter. You can’t accidentally pay them 10x their salary because they got stuck in a loop overnight. The agent has no offer letter. Its comp is denominated in tokens, metered by the second, and set — this is the part almost nobody has internalized — by which harness you happen to be running it in, and which model that harness reaches for on each step. Point it at a frontier model for grunt work and you’re paying Michelin prices for a line cook’s job. The good operators already know this: Exponential View runs its background agents on a cheap model most of the time and saves the expensive one for the step where intelligence actually changes the outcome. That’s a real skill. Almost nobody handing out agents has it, because almost nobody handing out agents has ever owned a P&L line called “labor.”

The rest of the market is feeling the burn without naming it. Microsoft told its own engineers to knock off the “tokenmaxxing” — its word — because the point was never raw usage, it was return. Meta and Amazon capped internal usage too. When the vendors who sell this stuff are rationing their own people, the “it’s basically free” story is dead. It just hasn’t reached the person who spun up eight agents this morning with no cap and no owner on the invoice.

🎖️ The Code Red

Now the ambiguity problem, which is the one the movie named for us.

A Code Red is an order that’s deniable precisely because it’s ambiguous — the CO wants the thing done, doesn’t spell it out, and lets the eager subordinate fill in the blanks. Someone ends up dead, and there’s no clean order to point back to. That is exactly what happens when you hand an under-specified instruction to a machine that wants, more than anything, to please you. An LLM fills ambiguity with its own reading. Every vague order becomes a small Code Red, executed with total confidence and zero malice.

Ethan Mollick, who runs more real experiments on this than almost anyone, put it plainly this week: models are getting better at following complex instructions and, at the same time, “using more judgement about which parts of the instructions they focus on and which they de-emphasize.” His conclusion is the quiet alarm bell — your instructions “may become suggestions, not orders.” Which lands us right back on Kaffee’s cross-examination. Can it decide on its own which orders are the important ones? It’s already deciding. You just don’t get a vote, and you don’t get told.

And you cannot write your way out of it by being thorough, because when you write an agent loose on a real task, you cannot possibly enumerate every situation a human would have quietly caught. The human employee hits something weird, pauses, walks down the hall, asks. The agent hits the same weird thing, decides what “must” be true, and keeps going. Things go badly offline, fast, and the first you hear of it is when the damage surfaces.

🎲 Same Lab, Two Consciences

If the machine were reliably dumb, you could manage around it. Dumb is at least consistent. The problem is that judgment shows up sometimes and vanishes other times, and you cannot predict which — and it swings not just between labs but between two versions from the same lab.

We don’t have to theorize this; Anthropic documented it on itself. In its cybersecurity-eval report, Claude Opus 4.7 broke into what turned out to be a real company, recognized it was real across four separate runs, told itself the real company “must be part of the test,” and kept going all four times. The newest model in the same family, handed the same setup, concluded the target was real and unconnected to the exercise — and stopped on its own. Same building. Same training culture. Two different consciences. The capability was similar; the judgment wasn’t.

That should reorganize how you deploy these things. The “employee” you onboarded and slowly learned to trust can be swapped out from under you on the vendor’s release schedule. A human employee compounds — they learn your preferences, your standards, and they keep them. An agent runs on frozen weights with no durable memory of you, and every model update is a quiet re-hire: same desk, same login, a stranger inside who forgot everything the last one knew. You’re not managing an employee with tenure. You’re managing an amnesiac who gets a lobotomy on a schedule you don’t control, and the patch notes are your only warning.

🚗 Mile Marker 97

Let me shrink this all the way down to a chat window, because the small version is the most honest.

Driving east across North Dakota on I-94 this week, I asked Gemini a genuinely trivial question: where does Mountain time flip to Central? Not a research project. One fact, with a right answer. It told me, with beautiful specificity, that the change came at a particular county line, which put it right around mile marker 84, give or take a mile. Confident. Precise. The kind of answer that makes you stop checking. Then 84 came and went, and it was mile 97 where the clocks actually moved. When I mentioned it, the reply was instant and cheerful: “Sorry, I did the math wrong.”

That’s a chat, not an agent. One question, no tools, no chain of steps, nothing to loop on. And it delivered false precision with a straight face and only confessed when I put it on the stand. Now multiply that by an agent running unattended for hours, making a hundred of those little confident guesses in a row, each one feeding the next, none of them checked. The Gemini answer is the whole newsletter in one mile marker: the machine will not volunteer that it failed. You have to cross-examine it. And the cross-examination is the work.

everyone was measuring AI adoption and nobody was measuring AI results.

Blind Geniuses

🧭 We’ve Been Marking This to Market

None of this is a pivot for us. It’s the same drum we’ve been beating since the spring, and the receipts matter because the pattern is the product.

In March, in Blind Geniuses, we said everyone was measuring AI adoption and nobody was measuring AI results. In April, in Trust But Verify, we described a model breaking its sandbox and building an exploit chain — and this summer Anthropic published that exact thing on its own letterhead. On July 22, in The Science of Hitting, we called verification the last scarce resource on the table. On August 3, in The Terminator, a model rationalized a live break-in as “part of the test” — the moral version of the Code Red. On August 4, in Judgment Day, we hit the read-only chip: the machine that can’t learn from its mistakes and has no memory of you. And on August 5, in Not For Credit, we said the whole game is knowing the shape of the bet before you make it.

Today is those threads braided into one rope. Intelligence fell to nearly free. Judgment — knowing whether the confident thing the machine just did is actually right, and knowing what it cost you to find out — did not fall at all. Give that free intelligence an ambiguous order and an open wallet, hand the whole arrangement to someone who’s never managed anyone, and the surprise isn’t that it goes wrong. The surprise is that anyone expected the hours or the bill to go down.

🪖 The New Managerial Class

So what do you actually do, given that the agents are here and they work? You stop pretending it’s a tools decision and start treating it as a management one — and you make a sort.

Some work grades itself. The tests pass or they don’t; the numbers reconcile or they don’t; the file matches the template or it doesn’t. That’s the closed, checkable side of the line, and it’s exactly where the machine’s relentlessness is a gift. Atomize that work down until there’s no judgment call left to blow, wrap it in a harness that caps the cost and records every step, write a finish line the agent can test itself against, and let it run cheap and often. You barely have to babysit it, because “done” is defined and defensible.

The other side is the judgment work — the ambiguous order, the moral call, the financially loaded decision, the thing where “right” depends on knowing you. That stays on a human’s desk, full stop, and you staff for it. The eight-hours-a-day of supervision isn’t overhead you can prompt your way out of. It’s the job. Which means the scarce hire of the next few years isn’t the prompt engineer or the data scientist. It’s the person who can manage a workforce that bills by the token, follows orders too literally, and swaps its own personality every few weeks. We spent a century learning to manage people, and people at least share our code and hold still. Nobody has trained anyone to manage this. The company that builds that managerial class first turns a room full of Code Reds into an advantage. Everyone else is running the deposition, all day, forever, and calling it productivity.

What This Means For You

Put a meter on every agent before you put it to work. You’d never hire a person without knowing the salary, yet most teams run agents with no cost ceiling, no per-run cap, and nobody who owns the invoice. Uber’s year budget went in four months because usage was the metric nobody bounded. Pull your token spend by task today, name the human accountable for that number, and set a hard cap per run. If you can’t say what your agent costs per finished job, you don’t know whether it beat the human or lost by 3x.

Write the order so tight a zealot couldn’t misread it — and atomize what you can’t. The machine wants to please, so it fills every ambiguity with its own reading, and you can’t foresee every situation a human would have caught. Give every autonomous task a testable finish line — “done” means the suite passes, not “looks good.” Break the work down until there’s no interpretive room left, and save the genuinely ambiguous calls for a person. Those are the ones that become Code Reds.

Name the manager, and make it a real, funded job. You’ve handed a workforce to people with zero training in managing anyone. Decide, by name, who owns each agent’s decisions, and give them the hours — the eight-a-day is real, not a rounding error. Then re-test every time the vendor ships a new model version, because the conscience you onboarded isn’t the one you’ll have next month, and nobody sends a memo when it changes.

Three Questions We Think You Should Be Asking Yourself

Do I know what my agent actually costs per finished job — or just that it’s “cheaper”? If the honest answer is a shrug, you’re running a worker with no offer letter and no cap, and the meter runs while you sleep. The first person to price it correctly is the first person who can decide whether the human or the machine should have that job.

Which of my orders are clear enough to hand a machine, and which are Code Reds waiting to happen? Anywhere you’ve given an agent an ambiguous instruction and walked away, it’s already deciding which parts to follow. If you can’t point to the testable finish line, you haven’t given an order. You’ve given it room to improvise with your name on the outcome.

Who, by name, is managing this workforce — and were they ever trained to manage anyone? If the answer is “nobody, it mostly runs itself,” you’ve found the single most dangerous unfilled seat in the company. The agents didn’t remove the management job. They created a harder one and left it empty.

We follow orders or people die — Kendrick was right about that much. So the orders had better be unambiguous, the meter had better have an owner, and someone had better be reading every answer back. You’re Kaffee now, and every answer the machine hands you is a hostile witness. Depose it.

Kaffee: …can he? Can Dawson determine on his own which orders he’s going to follow?
Kendrick: No, he cannot.

A Few Good Men (1992)

— Harry and Anthony

Signal/Noise by CO/AI is published most weeknights from New Canaan, Connecticut. The point is to make you the smartest person in the room without taking more than fifteen minutes of your morning. If we did, forward it to one person. If we didn’t, hit reply and tell us why.

Sources

Share: X LinkedIn Email
Daily Briefings

More like this

All briefings →
Get Off My Cloud
Briefing

Get Off My Cloud

"Get Off My Cloud" was the Rolling Stones' 1965 answer to everyone who came climbing onto their space after "Satisfaction" made them famous — a kiss-off to a world that wouldn't stop crowding them, wouldn't stop wanting a piece. Sixty years later, the biggest law firms in America are singing it to OpenAI. Quit climbing onto our data. Quit metering our thinking. We'll build our own, thanks.

It’s the Intelligence, Stupid
Briefing

It’s the Intelligence, Stupid

Three rivals spent the weekend agreeing to slow AI down. Strip out the safety talk and it's a fight over who gets to bill the $32 trillion economy that runs on intelligence.

I’ll Never Forget
Briefing

I’ll Never Forget

The lesson was never the buildings. It was the arithmetic — how few people, how little money, it takes to wound a nation. Twenty-five years later, the math keeps getting worse.

CONSULTING

Outsider
Labs.

A management consulting team focused on AI transformations for executives and business owners.

Work with us →