Meta Is Finally Playing Offense in AI. It Still Isn’t Winning
A year ago, Meta paid roughly $14.3 billion to acquihire Scale AI and put its 28-year-old founder, Alexandr Wang, in charge of fixing the company’s AI strategy. That was a big, embarrassing bet for a company that size to have to make, an admission that Meta’s internal AI effort wasn’t good enough on its own. We give Wang real credit for what’s happened since. Meta shipped Muse Spark 1.1, its first paid model, on July 9, the same week OpenAI’s GPT-5.6 family went GA and xAI took Grok 4.5 public. Less than a month later, on August 5, Meta shipped Muse Code, a terminal-based coding agent built on the newer Muse Spark 1.2, that installs with one command and takes on full software engineering tasks: planning changes, writing the code, checking whether the results actually work. That’s a company shipping on the same cadence as the frontier labs, not trailing them by a generation.
The benchmark result is the part worth taking seriously. On Terminal-Bench 2.1, Muse Spark 1.2 running through Muse Code scored 82.9 percent. That’s ahead of GPT-5.6 Terra on Codex, which came in at 81.8. It’s genuinely close behind Claude Opus 5 at 86.7. A year after getting laughed at for paying a fortune to poach one guy’s data-labeling startup, Meta built a coding agent that beats OpenAI’s on a real benchmark and is within striking distance of Anthropic’s, which is still the best in the category. That’s a real punch from a company that was nowhere in this specific fight a year ago, not a rounding error.
Wang’s own framing of the opportunity, from the same Y Combinator conversation, is worth quoting directly: “You can have a swarm of agents accomplish more than like a team of 100 engineers… very very handily actually, very very easily.” Tan had asked him for something obvious and near-term that people haven’t figured out yet. Wang’s answer was, essentially, his own company’s bet: that the constraint on building things stops being headcount and starts being vision. That’s a real, coherent thesis, and Meta backing it with an actual shipped product instead of just a keynote slide is the difference between this launch and a year of Zuckerberg press tours promising superintelligence was coming.
The Contributor Tier is where the good story stops. Muse Code’s pricing has a catch that isn’t buried in fine print so much as it’s the entire strategy. The heavily discounted Contributor Tier, the one Wang told CNBC is “more than 10 times cheaper than even the pay-as-you-go tier,” requires handing Meta permission to train future models on your prompts and your completions. Muse Code defaults to that tier on install. You have to actively switch to Standard pricing if you don’t want your code feeding Meta’s next model. That’s not a transparent discount for early adopters, it’s a data-collection funnel wearing a pricing page as a costume, and it’s exactly the kind of move that earns Meta the reputation it already has around privacy.
Wang’s own words undercut the confidence he’s selling. During the same podcast, Tan joked that tools like Hermes Agent are “Ferraris that break down on the side of the road all the time,” then asked Wang point blank whether Muse Code would be different. Wang’s actual answer: “Yeah, hopefully hopefully it doesn’t it doesn’t break down.” That’s the language of someone who knows he shipped something real but isn’t fully sure it holds up in the field yet, not a company that’s closed the gap. It lines up with Verdent’s read on the launch: a first-time entrant worth testing, not yet worth standardizing your team on.
The xAI comparison is where “catching up to the frontier” stops being true. Grok isn’t trying to win the coding-agent fight Meta just entered. Its edge is real-time X data access, fewer content restrictions, and raw compute scale, a completely different lane with a completely different moat. Meta closing the gap with Codex on Terminal-Bench says nothing about whether it can compete with what Grok is actually building toward. Beating one competitor on one benchmark is real progress. Calling it catching up to the whole frontier is a category error, and it’s one plenty of the coverage this week made without blinking.
We want this rivalry. A Meta that ships real products and gets uncomfortably close to Claude on a real benchmark makes every other lab move faster, and that’s good for anyone building on top of this stuff. Wang’s bet that ambition and vision are the scarce resource, not compute or headcount, is an interesting thesis, and Meta is the first of the giants to back it with a shipped product instead of a promise. A coding agent that’s still behind Anthropic, funded in part by a pricing tier that quietly harvests your code, and irrelevant to whatever xAI is actually building is real progress. It’s not the frontier yet.