
AI value depends on whether a model can complete dependable work with the required review, capacity, and total cost—not on its token price alone. AI-generated image via ChatGPT (OpenAI)
OpenAI’s New Scorecard Measures AI’s Business Value by Completed Work
OpenAI has introduced Useful Intelligence per Dollar, a proposed scorecard that asks businesses to measure AI’s value by the dependable work it completes, not by token prices, subscription fees, or usage limits alone.
For business buyers, the decision is whether a model, plan, and level of access give their team dependable capacity to complete the workflow they need at a cost the company can sustain. A low token price cannot answer that question on its own because it does not show whether the model can complete the work on the first pass without extra employee time, review, retries, and rework.
OpenAI’s scorecard evaluates whether AI produces useful work, what each successful task costs, how dependable its results are, and whether each AI dollar produces more value as use grows. It asks businesses to count the full resources required to complete a workflow and compare that total with the number of results that meet the required quality bar.
In short, a low token price or plan fee does not show whether AI is worth the investment. Businesses need to measure what it costs to complete dependable work, including the human effort it still requires and whether the selected model is available when the work needs to be done.
Useful Intelligence per Dollar is OpenAI’s proposed measure of the useful, dependable work AI produces for each dollar a business spends.
Key Takeaways: OpenAI’s Useful Intelligence per Dollar Scorecard
Useful Intelligence per Dollar is OpenAI’s proposed scorecard for assessing AI’s business value through useful work, the full cost of successful tasks, dependable results, and value that grows as use expands.
OpenAI says AI creates business value when it completes work that serves a real need, such as resolving a customer issue, shipping a code change, reviewing a contract, or preparing a forecast
The cost per successful task includes the model price, compute, employee time, human review, retries, and rework required to reach a result that meets the quality bar
A lower token price can produce a higher cost per successful task when employees must repeatedly correct, review, or finish the AI’s work
Dependability measures whether AI results are ready to use, need correction, or require a person to step in and complete the work
Value at scale increases when the number of tasks meeting the required quality bar grows faster than the total cost of completing them
Anthropic’s changing Fable 5 access shows that plan terms, shared weekly capacity, and usage-credit requirements can affect whether a team can keep using the model needed for its workflow
OpenAI’s AI Value Scorecard Measures the Full Cost of Completed Work
For years, software value has often been measured through seats purchased, active users, and license renewals. OpenAI argues that AI requires a different measure: the work it actually accomplishes.
The company says that judgment requires looking beyond a metric such as cost per token. A lower-cost model may require more attempts, more employee time, or more human review before it produces a usable result, while a more capable model may complete the same task in one pass despite having more expensive tokens. OpenAI’s measure is the full cost of producing a successful outcome compared with the value that outcome creates.
The company calls its proposed scorecard “Useful Intelligence per Dollar.” It asks whether AI completes work that serves a real business need, what each successful task costs, whether people can depend on the result, and whether the value produced by each dollar grows as use expands.
OpenAI’s economic test is whether the value of AI-completed work grows faster than the cost of producing it. A token rate or subscription price can contribute to that calculation, but neither can show the full cost of getting dependable work completed.
That scorecard becomes harder to apply when a team cannot count on continued access to the model it selected for an important workflow.
Anthropic’s Fable 5 Access Changes Affect AI Cost and Availability
Anthropic released Claude Fable 5 on June 9, then suspended access for all users three days later after export controls took effect immediately and the company could not verify users’ nationalities in real time. Anthropic said those controls were lifted on June 30 and began restoring access on July 1.
When Fable 5 returned, Anthropic said it could use up to 50% of the weekly limits on Pro, Max, Team, and select Enterprise plans through July 7. The promotion was later extended and ended July 19 at 11:59:59 p.m. Pacific Time, according to Anthropic’s current support guidance.
Starting July 20, Max plans and premium Team and seat-based Enterprise seats can use Fable 5 for up to 50% of their regular weekly limits without an additional charge. That does not add 50% more capacity to a plan. Fable 5 draws from the plan’s existing weekly limit, and Anthropic says it consumes that limit faster than its other models, leaving less of the weekly allowance available for a customer’s remaining Claude use. Once Fable 5 use reaches 50% of that weekly limit, those customers can pay with usage credits to keep using it or switch to another Claude model. Customers whose plans do not include Fable 5 need usage credits from the start.
In the rules that took effect July 20, Fable 5 is not included in the plan usage limits for Pro, standard Team, or standard seat-based Enterprise customers. Pro and standard Team customers need usage credits to use it, while standard Enterprise customers can use it only if their organization has enabled credits. Eligible Pro customers can claim $100 in usage credits. Eligible Team organizations can claim $100 for each purchased standard seat, up to $2,500 per organization, with the credits placed in one shared balance for the team. Anthropic says the credits will gradually become available starting July 20, must be claimed by August 2 at 11:59 p.m. Pacific Time, and expire September 17 regardless of when they are claimed. Standard seat-based Enterprise customers do not qualify. Usage-based Enterprise and API customers are billed at Anthropic’s standard API rates: $10 per million input tokens and $50 per million output tokens.
Keeping Fable 5 in an important workflow now depends on the customer’s plan, how much shared weekly capacity remains, and whether the customer can or will pay with usage credits. If price, plan terms, shared capacity, and credit requirements can each determine whether a workflow continues, a business needs a way to measure the cost of one successful task.
OpenAI’s Scorecard Measures the Full Cost of Successful Workflows
OpenAI’s “Useful Intelligence per Dollar” scorecard begins with two questions: Is AI completing work that matters, and what does each successful task cost? The company recommends starting with one workflow and defining what “done” means in the system where that work happens.
OpenAI asks businesses to look at how many customer issues AI helped resolve, code changes it helped ship, contracts it reviewed, and decisions improved because the right context was available at the right moment.
OpenAI argues that tokens create value only when they become work people can use. As models become more capable, they can handle longer and more complex tasks by maintaining context, reasoning through multiple steps, working across tools, and adapting as they go.
For a support team, “done” might mean a customer issue resolved. For engineering, it might mean a code change that passes its tests. For a legal team, it might mean a contract reviewed accurately and on time.
OpenAI uses forecast preparation as an example of how much work can sit behind one completed task. Before a finance team reaches a review meeting, it may need to find the latest forecast, move information into Excel or Sheets, identify changes, reconcile tabs, rebuild slides, and check that the totals add up. OpenAI says ChatGPT Work can take on much of that process, giving the team more time to focus on the questions that require its judgment: What changed? Why? What should happen next?
The cost of a successful task also depends on the work required to complete it well. A quick answer may require little compute, while a coding, research, or financial workflow can involve deeper reasoning, tool use, and many actions. Those workflows can require more compute, though they may also create more value.
OpenAI says the cost of a successful task includes the model price, the compute used to run it, and the likelihood that it reaches the right result. A business also needs to count employee time, human review, retries, and rework.
Its calculation is:
Add the full cost of completing the work.
Count the tasks that met the required quality bar.
Divide the total cost by the number of successful tasks.
A lower token price may therefore produce a higher cost per outcome if the work requires repeated attempts, delays, extra review, or human cleanup. OpenAI argues that a more capable model can justify a higher price when it produces the right result in one pass and reduces those costs.
The company presents GPT‑5.6 as a model family for making that choice by workflow: Sol is its flagship tier, Terra balances performance and cost, and Luna is its fastest and most affordable tier. A high-volume task that needs quick answers may fit Luna, while work requiring more depth may fit Terra or Sol. OpenAI says the economics of the full task should determine the model choice rather than its price alone.
OpenAI says GPT‑5.6 Sol with maximum reasoning set a new state of the art on the Artificial Analysis Coding Agent Index while using 54% fewer output tokens than another leading model. The company presents that comparison as evidence that GPT‑5.6 was trained to get more useful work from each token.
Across the GPT‑5.6 family, OpenAI’s stated goal is more successful work per dollar. It says greater efficiency can make existing tasks more affordable, while greater capability can make new kinds of work possible.
A task lowers its total cost only when people can use the result without repeatedly correcting it or finishing the work themselves.
AI Dependability Shows Whether Results Need Correction or Escalation
The third part of OpenAI’s scorecard is dependability: how often AI gets the work right. OpenAI describes adoption as a progression from drafting, to finding context and reasoning across tools and data, to taking action, handling exceptions, and completing workflows with people providing judgment and control where needed.
Each stage can create more value, though it also asks more of the AI system. OpenAI says dependability has direct economic value: when results are accurate, well-sourced, consistent, and escalated appropriately, people spend less time reviewing, correcting, and repeating the work.
Teams can track three outcomes:
Ready to use: The result met the quality bar as delivered.
Needs correction: Another attempt or human edits were required.
Needs escalation: A person had to step in and finish the work.
OpenAI says those measures offer a fuller picture than model accuracy alone because they show whether AI is reducing the total work required to complete a project. A result that frequently needs correction or escalation may save little time, even if the model performs well on an isolated benchmark.
Dependability also requires clear boundaries before AI moves from drafting toward taking action. OpenAI says organizations should define the data AI can access, the systems it can use or change, and when a person must review or approve an action. Those controls determine how much context and responsibility a company can safely give an AI system.
OpenAI says safety, security, privacy, and control provide the foundation for deeper AI use. People need to understand how the system behaves, how it handles their data, and how its actions are governed.
The company says ChatGPT Work builds on ChatGPT Enterprise’s security, privacy, compliance, and workspace-management features. That foundation is meant to let organizations give AI more context and access to more valuable workflows while maintaining appropriate oversight. Capability earns first use, OpenAI argues. Dependability makes AI part of how work gets done.
A more capable model may justify a higher price when it produces usable work more often while reducing human cleanup and delays. That value still depends on whether the model remains available within a team’s plan when the work needs to be completed.
Value at Scale Measures Whether Each AI Dollar Produces More Value
OpenAI’s final scorecard measure asks whether each AI dollar produces more value as use grows. To measure that, a business should follow the same workflow over time: track how many tasks met the quality bar, the full cost of completing those tasks, and the resulting cost per successful task. If completed work grows faster than total cost while quality holds or improves, each AI dollar is producing more value.
OpenAI identifies compute as a main driver of that improvement. Compute is the processing capacity used to train and run AI models. Training compute builds future model capability, while inference, the computing used when a trained model answers a request or completes a task, delivers useful work today. Because compute affects product quality, speed, dependability, availability, and cost, it can change both how much work AI completes and what it costs to complete it.
That is why OpenAI links better models, more efficient inference, purpose-built hardware, higher use of available computing infrastructure, smarter routing of requests, and stronger product design to the return on compute. Each generation of infrastructure can help train more capable models, while better algorithms, hardware, and software can make those models more efficient to run. Customers should then see better answers, faster results, fewer corrections, more dependable products, and lower costs for the work they need completed.
OpenAI calls the system connecting those improvements a shared intelligence platform. Its promise is that people can move work from a goal to a finished result more efficiently in one connected workflow, rather than restarting the task across separate tools. People can use it through ChatGPT, including ChatGPT Work, which can gather information from connected apps, files, and the web to carry a task forward. In the new desktop app, OpenAI has combined Chat, Work, and Codex in one app, while developers can use Codex or the API to build with the same underlying models. Enterprises can also connect those capabilities to the systems where their work happens. OpenAI’s argument is that improvements to its models, infrastructure, or software can benefit customers across those products, creating value that a token price alone cannot capture.
Model Price, Capacity, and Access Affect the Cost of Completed Work
Anthropic’s Fable 5 changes show why a business cannot separate a model’s listed price from its plan terms, capacity, usage credits, and availability. Those conditions determine whether the team can use the model needed to complete the workflow it selected.
A lower token rate offers limited value if a team cannot use the model when the work needs to be done, while a higher-priced model may still reduce total cost if it completes dependable work with less review and rework.
Businesses Should Choose AI Models by Cost, Dependability, and Access
OpenAI’s new framework gives businesses four measures for judging whether AI is creating value through completed work.
Useful work shows what AI produces. Cost per successful task shows what it takes to reach the outcome. Dependability shows how much of that work people can confidently use, while value at scale shows whether each dollar and each unit of compute accomplish more over time. As a scorecard, those four measures offer a more complete way to assess AI’s value and whether it is delivering a return on a business’s investment.
For business buyers, the decision is whether a provider’s model, plan, capacity, and price give the team dependable access to complete the workflow it needs. That—not token price alone—determines the value AI can bring to the business.
Q&A: OpenAI’s Useful Intelligence per Dollar Scorecard
Q: What is OpenAI’s Useful Intelligence per Dollar scorecard?
A: Useful Intelligence per Dollar is OpenAI’s proposed way for businesses to measure AI value through the useful work it completes, the cost of successful tasks, the dependability of its results, and whether each AI dollar produces more value as use grows.
Q: How does OpenAI calculate the cost of an AI task?
A: A business adds the full cost of completing a workflow, including the model price, compute, employee time, human review, retries, and rework. It then divides that total by the number of tasks that met the required quality bar to find the cost per successful task.
Q: What makes AI work dependable?
A: Dependable AI work meets the required quality bar often enough that people can use it with appropriate oversight. OpenAI suggests tracking whether results are ready to use, need correction, or need a person to step in and finish the work.
Q: What does value at scale mean for AI?
A: Value at scale means the number of tasks meeting the quality bar grows faster than the total cost of completing them. When quality holds or improves while completed work grows faster than cost, each AI dollar is producing more value.
Q: Why is a low token price not enough to show AI value?
A: A lower-priced model can cost more overall if people must repeatedly correct its work, review answers, retry the task, or complete the remaining work themselves. A higher-priced model may lower the total cost when it produces usable results more reliably and with less human cleanup.
Q: How do Anthropic’s Fable 5 access changes affect businesses?
A: Fable 5 shows that the listed model price does not fully determine a workflow’s cost or availability. A team’s ability to keep using the model can depend on its plan, remaining shared weekly capacity, and whether it can use or purchase usage credits.
Q: What should a business measure before choosing an AI model and plan?
A: Businesses should measure whether the model, plan, price, and available capacity give their team dependable access to complete the workflow it needs. The decision should account for the full cost of successful work, including the human effort required when AI results need correction or escalation.
What This Means: AI Value Depends on Completed Work
AI becomes a good investment when it helps a team complete dependable work at a sustainable total cost. Token prices, subscription fees, and usage limits are part of that cost, though none of these can show the value AI creates on its own.
OpenAI’s Useful Intelligence per Dollar scorecard gives businesses a more complete way to measure the value of AI by evaluating whether it produces useful work, what each successful task costs, how dependable its results are, and whether each AI dollar produces more value as use grows. It asks businesses to define what “done” means for a workflow, add the full cost of reaching that outcome, and compare it with the number of results that meet the required quality bar.
Business buyers and teams using AI for customer support, coding, finance, legal review, and other important workflows should care because they must choose a model, plan, and level of access that can complete the work they need at a price the company can sustain.
Anthropic’s Fable 5 access changes show why a listed model price cannot settle that choice alone. Plan terms, shared weekly capacity, and usage-credit requirements can determine whether a team has dependable access to the model it selected. Losing access can make a workflow slower and more expensive when people must switch models, pay for additional use, or take over work themselves.
Teams now need to choose AI models and plans based on the workflow they must complete, the quality that workflow requires, and the access they can rely on. A lower token price may produce a higher cost per outcome if the work requires repeated attempts, delays, extra review, or human cleanup. OpenAI argues that a more capable model can justify a higher price when it produces the right result in one pass and reduces those costs.
In short, AI value comes from dependable work completed at a sustainable total cost, including the model access a team needs to finish the workflow.
The model that creates the most value is the one a team can depend on to complete the work when it needs to be done.
Sources:
OpenAI: A scorecard for the AI age
https://openai.com/index/a-scorecard-for-the-ai-age/OpenAI: ChatGPT is now a partner for your most ambitious work
https://openai.com/index/chatgpt-for-your-most-ambitious-work/Anthropic: Claude Fable 5 on your plan
https://support.claude.com/en/articles/15424964-claude-fable-5-on-your-planClaude: Fable 5 plan access update
https://x.com/claudeai/status/2078302415804379218Anthropic: Pricing
https://platform.claude.com/docs/en/about-claude/pricingAnthropic: Claude Fable 5 one-time free credits promotion
https://support.claude.com/en/articles/15862783-claude-fable-5-one-time-free-credits-promotionAnthropic: Claude Fable
https://www.anthropic.com/claude/fableAnthropic: Redeploying Claude Fable 5
https://www.anthropic.com/news/redeploying-fable-5Anthropic: Introducing Claude Fable 5 and Claude Mythos 5
https://platform.claude.com/docs/en/about-claude/models/introducing-claude-fable-5-and-claude-mythos-5Claude: Claude Fable 5 promotional access
https://x.com/claudeai/status/2072402639644766602Anthropic: Claude Status - Incident History
https://status.claude.com/historyClaude: Claude Fable 5 promotional access
https://x.com/claudeai/status/2074548243971604641Claude: Claude Fable 5 promotional access
https://x.com/claudeai/status/2076351401006154204
Editor’s Note: This article was created by Alicia Shapiro, CMO of AiNews.com, with writing support, AEO/GEO/SEO optimization, image concept development, and editorial structuring support from ChatGPT, an AI assistant. All final editorial decisions, perspectives, and publishing choices were made by Alicia Shapiro.

