Meta just did something it swore it would never do. On July 9, 2026, Meta Superintelligence Labs opened its first metered, pay per token API, and the model behind it, Muse Spark 1.1, is turning heads for a different reason too: on JobBench, a benchmark built to measure real agentic tool use, it scored 54.7, ahead of Claude Opus 4.8 at 48.4 and GPT-5.5 at 38.3. For a company that built its AI reputation on giving Llama weights away for free, that is a significant pivot.
Meta Muse Spark 1.1 AI agent capability now includes genuine computer use: the model can look at a screenshot, decide what to click, open an app, and complete a multi step task on a real desktop with almost no hand holding. In this article you will learn what Muse Spark 1.1 actually does, how it stacks up against Claude and GPT on agentic benchmarks, what the new Meta Model API costs, and what this closed, paid shift signals for anyone building or buying AI agents in 2026.
What Is Meta’s New Agentic AI Model
Muse Spark 1.1 is Meta’s second release from Superintelligence Labs and its first major update since the original Muse Spark launched in April 2026 as the company’s first closed, proprietary frontier model. The new version keeps that closed approach but adds a 1 million token context window and a real computer use capability, meaning the model can drive a desktop environment from a single plain language goal.
Give it an instruction like find the spreadsheet, pull last quarter’s numbers, and draft a summary email, and Muse Spark 1.1 will look at the screen, locate the right app, click through the interface, and reason about what it sees before acting. That is a meaningfully different skill from a chatbot that only answers questions in text.
The model also introduces a main agent and sub agent structure. As the main agent, Muse Spark 1.1 gathers context, builds a plan, and delegates pieces of the job to parallel sub agents. As a sub agent, it stays inside its assigned task, understands which tools it can use, and escalates back to the main agent when something falls outside its lane. That kind of built in task delegation is exactly the multi agent architecture enterprises have been assembling by hand with frameworks like LangGraph and CrewAI, now baked directly into a single model.
Consumers can try the model in Thinking mode inside the Meta AI app or at meta.ai, while developers can call it through the new Meta Model API, an OpenAI compatible endpoint now in public preview.
How Muse Spark 1.1 Performs on AI Agent Benchmarks
Meta built Muse Spark 1.1 around tool use, not raw chat quality, and the benchmark scores back that framing up. On JobBench, which tests whether a model can finish real, professional multi step work rather than just produce a good sounding answer, Muse Spark 1.1 leads the field at 54.7 against Claude Opus 4.8’s 48.4 and GPT-5.5’s 38.3.
On Finance Agent v2, an agentic financial analysis benchmark, Muse Spark 1.1 again comes out ahead at 57.2, compared to 53.9 for Opus 4.8 and 51.8 for GPT-5.5. It also leads on MCP Atlas, Humanity’s Last Exam with tools enabled, and HealthBench Professional, a spread of results pointing to real strength in multi step, tool calling tasks rather than one cherry picked number.
The picture flips in other categories. On coding benchmarks like SWE-Bench Pro and Terminal-Bench 2.1, and on multimodal tests like CharXiv, Muse Spark 1.1 lands in a solid third place behind Claude and GPT. That is worth remembering before assuming Muse Spark 1.1 is now the default choice for every agentic workload: it is a specialist in tool use and orchestration, not an across the board leader.
Pricing plays into the comparison too. At $1.25 per million input tokens and $4.25 per million output tokens, Muse Spark 1.1 costs roughly a quarter of what Anthropic and OpenAI charge for comparable frontier models, a gap worth weighing against the introductory pricing Anthropic set for Claude Sonnet 5.
What This Means for Businesses Building AI Agents
For a developer or small business evaluating tools right now, Muse Spark 1.1 adds a genuinely new option to the field, not just another chatbot wrapper. Its combination of a 1 million token context window, native computer use, and built in sub agent delegation means teams that previously stitched together a framework on top of a base model can now get orchestration behavior out of the box.
The pricing gap matters just as much. At roughly a quarter of the cost of Claude or GPT, Muse Spark 1.1 makes agentic workflows, the kind that call a model dozens of times to complete one task, meaningfully cheaper to run at scale. That is especially relevant for smaller teams that have been priced out of the most capable agentic models, a group we cover in our guide to AI agents for small business.
Practically, teams should treat Muse Spark 1.1 as a strong option for tasks built around tool calling, task delegation, and screen level automation, while still leaning on Claude or GPT for the heaviest coding or multimodal work. Testing a workload against JobBench style tasks before committing is a reasonable way to decide which model earns the job.
Access is already open: consumers can try it today inside the Meta AI app, and developers can start building against the Meta Model API in public preview without waiting on a longer rollout.
The Quiet End of Meta’s Open Weight Era
Muse Spark 1.1’s paid API is a bigger strategic signal than any single benchmark. Meta built its AI reputation on giving Llama weights away for free while OpenAI and Anthropic charged by the token. A metered, pay per use API, detailed in Meta’s own announcement, is the same revenue model Meta spent years positioning itself against, and the company says an open source variant is still in development with no release date attached.
That shift lands inside a pricing war that keeps getting more crowded, a trend The New Stack has also been tracking closely. DeepSeek’s roughly $0.44 per million output tokens already sets a price floor for the industry, and open weight models arriving elsewhere mean some organizations can self host a capable coding model with no per token cost at all. Meta is betting that computer use and agentic orchestration, not raw price, will be what convinces developers to pay for Muse Spark 1.1 instead of routing around it.
Key Takeaways
Three things stand out from this launch. First, Muse Spark 1.1 is a genuine agentic specialist, leading Claude Opus 4.8 and GPT-5.5 on JobBench and Finance Agent v2 while trailing both on coding and multimodal tasks. Second, its computer use and sub agent delegation are built into the model itself, not bolted on through an external framework. Third, the paid Meta Model API marks a real strategic pivot away from Meta’s free weight legacy, priced at roughly a quarter of what Claude and GPT charge.
For anyone building with AI agents, this is a model worth testing against your own workflows rather than taking benchmark charts at face value. Explore more breakdowns of the tools, pricing shifts, and platform launches shaping agentic AI at BigAIAgent.
Would you trust a general purpose agent to run your desktop unsupervised, or do you still want a human checking every click?








