Meta launched Muse Code in beta on August 5, the company's first terminal coding agent and a direct challenge to OpenAI's Codex and Anthropic's Claude Code. Built on Muse Spark 1.2, the agent scored 54 on Artificial Analysis's Intelligence Index, placing Meta in a tie for third among U.S. labs with SpaceX AI.
The benchmark story
Muse Spark 1.2 jumped 260 Elo points in agentic knowledge work since April's Muse Spark 1.0, its former weak spot, and now ranks just behind frontier models. On Terminal-Bench 2.1, Muse Code scored 82.9%, narrowly behind Claude Opus 5 at 86.7% but ahead of GPT-5.6 Terra. Meta claims the agent built six game features in parallel without collisions, using background sub-agents in isolated git worktrees.
The pricing trap: Standard API pricing remains $1.25 input / $4.25 output per million tokens, matching Muse Spark 1.1. But the new contributor tier runs approximately 21x cheaper — in exchange for letting Meta train on your prompts and code. The default onramp funnels developers' proprietary codebases into Meta's training pipeline; opting out requires switching to standard pricing.
What Meta isn't saying
Zuckerberg teased open-sourcing these systems but gave no timeline. Muse Spark 1.2 remains closed-weight, accessible only via Meta's API. The model's coding scores were produced inside the harness it was trained against — standard practice for agentic evaluation, but it means performance through other scaffolds remains unknown. For a lab that built its AI reputation on open-weight Llama, Muse is the opposite bet: proprietary inference, contributor economics, and a direct monetization play.