
An AI workflow detects a problem, corrects its work and delivers a verified result for human review. AI-generated image via ChatGPT (OpenAI)
Claude Opus 5 Launch Aims to Finish AI Work With Fewer Human Fixes
Anthropic has released Claude Opus 5 across its platforms, promising a model that can verify its work, correct problems and persist through complex tasks. Businesses now have to decide whether Opus 5 can complete their tasks from start to finish, reduce enough human correction and save enough time to justify the cost.
Strong first attempts still leave people responsible for catching mistakes, repairing incomplete work and pushing difficult tasks across the finish line. Anthropic reports that Opus 5 completes more business, knowledge-work and software-engineering tasks while charging the same base price as Opus 4.8.
Opus 5 is designed to inspect its progress, create workarounds when a task is missing something and build ways to test its own results. Customers can also adjust its effort setting to balance performance, speed, token use and task cost.
The release primarily affects teams evaluating AI for multi-step business processes, knowledge work and software development. Their results will depend on the work being assigned, the effort setting selected and how reliably the model produces a usable result.
In short, Opus 5 is Anthropic’s attempt to move AI closer to completing complex work with less human correction. Its real-world value will depend on how much dependable work it finishes before a person has to step in.
Agentic work involves an AI carrying out a series of connected steps, using tools and checking its progress as it works toward a completed result.
Key Takeaways: Claude Opus 5 Capabilities, Cost and Business Value
Claude Opus 5 is Anthropic’s agentic AI model designed to check, revise and test its work as it completes complex, multi-step tasks with less human correction.
Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens, the same base price as Opus 4.8 despite Anthropic’s reported performance improvement.
Claude Opus 5 can create workarounds, investigate underlying problems and build tests to verify its results when a task lacks a necessary tool or source of validation.
Claude Opus 5 reached a 26% maximum pass rate on AutomationBench, compared with 18.1% for GPT-5.6 Sol, 17.4% for Fable 5 and 17% for Opus 4.8 in Anthropic’s results.
Claude Opus 5 lets customers adjust its effort setting to favor stronger reasoning or reduce token use, response time and cost, while Fast mode provides approximately 2.5 times the default speed at twice the base price.
Anthropic ran or selected the published benchmark comparisons, and early-access companies reported results from their own workflows, so the evidence does not establish how much time or human correction every organization will save.
Claude Opus 5 will provide business value when it can complete an organization’s actual tasks from start to finish and reduce enough review, repair work and failed attempts to justify its total cost.
Claude Opus 5 is designed to complete complex work with less human correction
Anthropic has released Claude Opus 5 across all of its platforms, describing it as a high-capability model designed for everyday use. It is now the default model for Claude Max subscribers and the strongest model available through Claude Pro.
Anthropic says Opus 5 approaches the advanced capabilities of its Fable 5 model at half the price while retaining the same base rates as Opus 4.8: $5 per million input tokens and $25 per million output tokens through the Claude API. Tokens are the small units of text the model processes as it reads a request and produces a response, so the unchanged rates give customers what Anthropic describes as stronger performance without a higher base price.
The company says Opus 5 is more capable of checking its own work, correcting problems and continuing until it completes a task, rather than stopping after a strong first answer. If that behavior proves consistent, teams could spend less time reviewing partial results, identifying mistakes and telling the model how to proceed.
That promise depends on what Opus 5 does when a task becomes difficult or its first approach falls short.
How Claude Opus 5 checks, revises and verifies its work
In one Frontier-Bench task, Opus 5 was asked to write code that would recreate a machine part as a three-dimensional model in FreeCAD, a computer-aided design program. The test deliberately gave the model no direct way to view the original drawing, so Opus 5 created its own computer-vision pipeline to extract the part’s shape and dimensions from the image’s raw pixels. It then reconstructed the complete part and, according to Anthropic, succeeded repeatedly. Anthropic did not report the number of successful attempts, but said no competing model working under the same conditions solved the task in five attempts.
Opus 5 showed a similar ability to look past an incomplete solution when Anthropic gave it a real bug from a widely used open-source package manager. The model traced the problem to its underlying cause and repaired an edge case that an existing community patch had missed. A competing model corrected the visible symptom, left the underlying problem unresolved and still reported the task as complete.
An engineer at a trading firm also used Opus 5 to build a market-data feed for a new exchange in a single session, a task earlier models had failed to complete even with detailed plans from the engineer. When the model found that no live feed was available to validate its work, it built a test harness, a set of checks used to verify that the code was reading the exchange’s data correctly.
Each example follows the same pattern: Opus 5 encountered a missing capability or source of verification, created a way around it and continued working toward a checked result.
Those examples show what finishing the job can look like, but how consistently does Opus 5 display that behavior across business processes, knowledge tasks, computer use and software engineering?
Claude Opus 5 completes more multi-step tasks in Anthropic’s benchmarks
AutomationBench tests whether an AI model can complete business tasks from beginning to end, including the individual actions required to finish the work. Opus 5 reached a maximum pass rate of 26% in Anthropic’s results, compared with 18.1% for GPT-5.6 Sol, 17.4% for Fable 5 and 17% for Opus 4.8. Across the effort levels shown in Anthropic’s chart, Opus 5 completed approximately 22% to 26% of the tasks, while the other models peaked at roughly 17% to 18%.
Anthropic also compared the models at similar task costs, since giving a model more time and computing power can improve its results while making each completed job more expensive. At the same cost per task, the company reported that Opus 5 passed approximately 1.5 times as many tasks as the next-best model. Its lowest tested effort setting still produced a higher pass rate than any of the other models reached at their best.
Zapier tested that capability with an account-health workbook, asking Opus 5 to identify customers at risk of leaving, notify the appropriate account owners and prepare a summary for the company’s retention team. Previous models failed the tested workflow; Opus 5 reached 100%. Zapier CEO Wade Foster described the result:
"Claude Opus 5 topped Zapier's AutomationBench leaderboard without spending more tokens than prior Claude models. It took a raw account-health workbook and ran a full churn-prevention sequence end to end: flagging at-risk accounts, alerting the right owner, and summarizing for retention ops. Previous models didn't pass; Opus 5 hit 100%."
-- Wade Foster, CEO Zapier
Anthropic reported a similar advantage on GDPval-AA v2, which evaluates knowledge work across different effort and cost levels. The benchmark uses an Elo score, a relative rating in which a higher number indicates stronger performance on the tested tasks. Opus 5 rose from approximately 1,450 at its lowest effort setting to 1,861 at its highest. At maximum effort, it scored above Fable 5 at 1,747, GPT-5.6 Sol at 1,736 and Opus 4.8 at 1,593.
Box also reported gains in its own early-access testing, including an 8% improvement over Opus 4.8 overall, an 11% improvement in data analysis and a 17% improvement in due-diligence workflows.
"Claude Opus 5 delivers the industry intelligence and accuracy that is essential for the analysis of specialized enterprise content. Box found that Opus 5 outperforms Opus 4.8 by 8% and delivers notable performance gains in the data analysis (11% improvement) and due diligence (17% improvement) workflows that technology, healthcare, and public sector organizations rely on daily."
-- Ben Kus, CTO Box
On OSWorld 2.0, which tests a model’s ability to operate a computer, Anthropic reported that Opus 5 outperformed the other tested models at comparable costs and exceeded Fable 5’s best result at slightly more than one-third of the cost. On Frontier-Bench, an evaluation of valuable software-engineering work, Opus 5 more than doubled Opus 4.8’s performance while costing less per task.
Anthropic ran the Frontier-Bench comparison internally using mini-SWE-agent, the software framework that gave each model its instructions and tools, with computing infrastructure hosted on Google Kubernetes Engine. Scores were based on the average result across five attempts at each task. When Opus 5 or Fable 5 refused a request because of a safety classifier, Anthropic substituted Opus 4.8 for that attempt.
The benchmark comparisons were selected and run by Anthropic, while the Zapier and Box results came from early-access customers testing their own workflows. They provide evidence across several kinds of work, but the published results do not establish that every organization will see the same completion rates or improvements.
Higher completion rates are valuable only if obtaining them does not make routine work too expensive or slow, so what does Opus 5’s advantage cost?
Claude Opus 5 effort settings balance performance, speed and cost
Opus 5 costs $5 per million input tokens and $25 per million output tokens, the same base rates Anthropic charged for Opus 4.8, despite what the company describes as greatly improved performance. Customers can adjust the model’s effort setting according to the job: higher settings prioritize its strongest reasoning, while lower settings use fewer tokens to produce faster, less expensive results.
In one early-access legal-work evaluation, Opus 5 produced similar performance while using an average of 26% fewer tokens than Opus 4.8 at its maximum reasoning setting
Anthropic says Fable 5 offers higher peak capability at approximately twice the price of Opus 5, although the performance gap can narrow depending on the task and configuration. On CursorBench 3.2, Opus 5 at maximum effort came within 0.5% of Fable 5’s highest score while costing half as much per task.
Teams can also pay for speed through Opus 5’s Fast mode, which runs approximately 2.5 times as fast as the default version and costs twice the model’s base price. Buyers therefore have to balance the quality of the result against token consumption, response time and the total cost of completing the task.
Anthropic’s benchmark curves and early-customer reports help compare those tradeoffs, but token prices do not reveal the full cost of the work. The company did not provide a general measurement of how much employee time Opus 5 saves by avoiding failed attempts or reducing the corrections needed before a result can be used. Without those figures, buyers cannot fully determine whether a higher-performing configuration will save enough review and repair work to justify its cost in their own workflows.
If the benefit depends on the task, effort setting and amount of human correction avoided, how should a team determine whether Opus 5 has crossed the threshold from impressive to useful?
Claude Opus 5’s business value depends on reducing human correction
Lovable’s early-access testing focused on whether Opus 5 could produce dependable results across repeated difficult coding tasks. The company reported a 22% improvement over Opus 4.7 on its hardest agentic coding evaluations, which test whether a model can independently carry out a series of steps, along with much less variation from one run to the next.
"Claude Opus 5 came out ahead of every model in its family on our internal evals. It isn't just better on our hardest agentic coding tasks, up 22% over Opus 4.7, it's steadier, with far less variance run to run. For the millions of builders on Lovable, that consistency is the whole game. Reliable results, build after build."
-- Fabian Hedin, Co-Founder Lovable
Another early-access evaluation tested Opus 5 on financial modeling and reported an average accuracy gain of nine percentage points. The model also completed the work with one-third fewer conversational turns and tool calls, reducing the total time required by 60%. Both results came from participating companies testing their own internal workflows.
Reliability also depends on how a model behaves when it encounters risky or conflicting instructions. Anthropic’s automated pre-deployment audit rated Opus 5 as its most aligned model to date, meaning it followed Claude’s Constitution, the principles intended to guide its behavior, more closely than the recent Anthropic models included in the comparison. The company also reported lower rates of deceptive behavior, less susceptibility to misuse and fewer reckless actions that could have difficult-to-reverse consequences. Opus 5 received an overall misaligned-behavior score of 2.3, the lowest among the models Anthropic compared, although these results came from the company’s automated audit rather than independent safety testing.
Anthropic’s cybersecurity safeguards allow Opus 5 to find vulnerabilities in source code while blocking binary-based vulnerability scanning, penetration testing and exploit generation. That means it can help developers identify weaknesses in readable code, but it will block requests to probe compiled software, test whether a system can be broken into or create a way to exploit a discovered flaw. When those safeguards flag a request in Claude.ai, Claude Code or Claude Cowork, the system falls back to Opus 4.8 by default. Developers using the API can also enable automatic fallbacks that send a flagged request to another available model instead of blocking it.
Developers can change the tools available to Opus 5 during a conversation without invalidating the prompt cache, the stored prompt information that prevents the system from processing the same context again. Anthropic also says Opus 5 carries no special model-specific data-retention requirement for general access, although the existing terms for each product and account still determine how data is retained.
For businesses, the adoption test therefore belongs in their own workflows: Can Opus 5 carry the organization’s business tasks from start to finish and verify its own work? How much correction does it still require, which effort setting produces an acceptable result, what does each completed task cost and how often does a safety fallback send the request to another model?
Opus 5’s real-world value will depend on whether its ability to verify its work and persist through problems reduces the work between an initial response and a dependable completed task. If it consistently removes that last mile of human correction, teams may be able to hand over more of a complex job while focusing their review on the decisions and risks that still require it.
Q&A: Claude Opus 5 for Business Explained
Q: What is Claude Opus 5?
A: Claude Opus 5 is Anthropic’s agentic AI model for complex, multi-step work. It is designed to check its progress, correct problems, use tools and continue working toward a completed result with less human intervention.
Q: What can Claude Opus 5 do that earlier models could not?
A: Opus 5 can recognize when a task is missing a capability or source of verification, create a workaround and test whether its result is correct. In Anthropic’s examples, it built a computer-vision pipeline to interpret an image, traced a software bug to its underlying cause and created a test harness when no live data was available for checking its code.
Q: Can Claude Opus 5 complete business tasks from start to finish?
A: Opus 5 reached a maximum pass rate of 26% on AutomationBench, compared with approximately 17% to 18% for the other models in Anthropic’s results. Zapier also reported that Opus 5 completed one tested customer-retention workflow with a 100% result after previous models failed it. Businesses still need to test that performance in their own workflows.
Q: How much does Claude Opus 5 cost?
A: Opus 5 costs $5 per million input tokens and $25 per million output tokens, the same base rates as Opus 4.8. Customers can lower the effort setting to reduce token use, response time and cost, or pay twice the base price for Fast mode, which runs approximately 2.5 times as fast as the default version.
Q: Have Claude Opus 5’s performance claims been independently verified?
A: Anthropic selected and ran the published benchmark comparisons, while early-access customers reported results from their own internal workflows. The evidence covers several kinds of work, but it does not establish that every organization will achieve the same completion rates, accuracy gains or time savings.
Q: What happens if Claude Opus 5 blocks a request?
A: Requests flagged by Anthropic’s safety safeguards fall back to Opus 4.8 by default in Claude.ai, Claude Code and Claude Cowork. Developers using Anthropic’s API can enable automatic fallbacks that route a flagged request to another available model instead of blocking it.
Q: How can a business decide whether Claude Opus 5 is worth the cost?
A: A business should test whether Opus 5 can complete its actual tasks from start to finish, verify its work and reduce the time employees spend correcting mistakes or recovering from failed attempts. That benefit should then be compared with the effort setting, response speed, cost per task and frequency of model fallbacks.
What This Means: How Claude Opus 5 Could Reduce Human Correction in Complex Work
Claude Opus 5 is Anthropic’s attempt to make AI responsible for checking, revising and testing complex work until it reaches a usable result with less human correction. If that behavior holds consistently, teams could spend less time directing each intermediate step and focus more of their attention on reviewing the completed work.
Opus 5 demonstrated self-verification when it created a workaround or test after encountering missing information or no direct way to check its work. That behavior allowed it to proceed when earlier models stopped or reported incomplete solutions as finished.
Business leaders, developers and knowledge-work teams should care most when they assign AI tasks involving several connected steps. These teams have the clearest reason to test whether Opus 5 can reduce the time employees spend checking and repairing incomplete AI-generated work.
Fewer failed attempts and shorter correction cycles could improve the speed and total cost of completing a task. Opus 5 also carries the same base price as Opus 4.8, with adjustable effort settings that let teams balance performance, response time, token consumption and cost for regular use.
Organizations should test Opus 5 on a representative workflow and measure the completed result, including the employee intervention required and the cost at each effort setting. That comparison will show whether the model can reliably remove enough work to produce a genuine operating benefit.
In short, businesses should judge Opus 5 by the cost of a usable completed task, including the employee work required to get there.
AI earns more responsibility when teams can rely on it to carry difficult work across the finish line.
Sources:
Anthropic: Introducing Claude Opus 5
https://www.anthropic.com/news/claude-opus-5
Claude Help Center: Data retention practices for Covered Models
https://support.claude.com/en/articles/15425996-data-retention-practices-for-covered-models
Claude Help Center: Covered Models
https://support.claude.com/en/articles/15425695-covered-models


