Meta Says Muse Spark 1.3 Has Frontier Performance — But Its Best Results Come From a Model Developers Can’t Broadly Use Yet
Carl Franzen
9:19 am, PT, September 3, 2026
Credit: VentureBeat made with OpenAI ChatGPT-Images-2.0
Meta’s newest AI model, Muse Spark 1.3, unveiled yesterday, is faster and more performant on third-party benchmarks than its predecessor — with a caveat.
"Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter," Meta co-founder and CEO Mark Zuckerberg wrote on X, calling it Meta’s “biggest jump” yet in coding and agentic work.
There is substance behind both parts of that claim. Muse Spark 1.3 makes significant gains over last month’s 1.2 release, particularly on long-running agent tasks. The version developers can access now is also one of the strongest price-performance offerings near the top of independent model rankings.
Meta’s strongest Muse Spark 1.3 benchmark results come from its max reasoning configuration. Meta says that version is still completing additional safety testing and will arrive “shortly”; the third-party benchmarking firm Artificial Analysis says it evaluated max in a limited partner preview, and currently lists no API provider at all for the configuration.
The version broadly rolling out this week through its Muse Code harness and the Meta Model API uses Meta’s previously available reasoning settings, including xhigh.
That makes the more relevant enterprise question not whether Muse Spark 1.3 can reach frontier territory, but how close the model companies can actually deploy today gets — and at what real cost.
The shipping model is very good, but not the benchmark leader.
Meta does disclose results for both configurations in its underlying evaluation report, so this is not a case of the company hiding the deployable model. But its launch materials prominently showcase the max variant, and some of the largest scores belong to that configuration.
Meta Muse Spark 1.3 benchmarks. Credit: VentureBeat made with OpenAI ChatGPT-Images-2.0
For example, Meta reports GDPval-AA v2 scores of 1,754 Elo for max versus 1,709 for xhigh, OSWorld 2.0 scores of 66.9 versus 57.2, and JobBench scores of 64.9 versus 61.2.
On some tests the distinction is negligible or reversed: DeepSearchQA is tied at 89.4, while xhigh scores 89.2 on Terminal-Bench 2.1 versus max at 88.8.
Artificial Analysis scores Muse Spark 1.3 max at 62 on its Intelligence Index and the shipping xhigh version at 61. The latter ties GPT-5.6 Sol max, Grok 4.6 high, and Claude Opus 5 high. But Anthropic still occupies the top of the leaderboard: Claude Fable 5.1 reaches 66 at max and 65 at xhigh, while Claude Opus 5 reaches 63 at max and xhigh.
In other words, Muse Spark 1.3 xhigh is legitimately in the frontier cluster, but it is not the model currently setting the frontier.
That is still a substantial change from Muse Spark 1.2. VentureBeat’s coverage of last month’s launch found Meta fielding a credible coding challenger that nevertheless generally trailed Anthropic’s best model. Muse Spark 1.2 scored 82.9% on Terminal-Bench 2.1 versus Opus 5’s 86.7%, and also finished behind Opus on the other main coding comparisons Meta presented.
With 1.3, Meta is no longer merely showing up in that contest. On several coding and agentic evaluations, it is trading wins with OpenAI and Anthropic.
Meta says the underlying model has also become easier to operate. Muse Spark 1.3 is trained to maintain multiple workflows in a long thread, gather context with tools, detect gaps in its own plans, ask users for clarification when necessary, and confirm before consequential actions. In Meta engineers’ internal comparisons, it used roughly 20% fewer tool calls and 25% fewer tokens.