CMOtech Canada - Technology news for CMOs & marketing decision-makers
Canada
OpenAI cuts GPT-5.6 API prices & adds faster Sol mode

OpenAI cuts GPT-5.6 API prices & adds faster Sol mode

Thu, 30th Jul 2026 (Yesterday)
Mark Tarre
MARK TARRE News Chief

OpenAI has cut API prices for its GPT-5.6 Luna and Terra models and introduced a faster processing option for GPT-5.6 Sol. The changes also reduce how Luna and Terra usage is counted in ChatGPT Work and Codex.

The steepest cut applies to Luna, which now costs 80% less. Terra prices are down 20%. Sol pricing is unchanged, but the API now includes a Fast mode that offers quicker responses at a premium.

Under the revised API pricing, Terra costs USD $2 per million input tokens and USD $12 per million output tokens. Luna costs USD $0.20 per million input tokens and USD $1.20 per million output tokens.

Fast mode replaces OpenAI's Priority Processing option for Sol. The service delivers up to 2.5 times faster speeds than standard processing at twice the price. Existing API requests tagged as priority will move automatically to the new mode.

The update affects several parts of OpenAI's commercial product range. Terra and Luna remain available in ChatGPT Work, Codex and the OpenAI API. Free and Go tier users in ChatGPT Work and Codex can access Terra, while paid users, including Plus, Pro, Business and Enterprise, can choose both Terra and Luna.

OpenAI presented the changes as part of a broader effort to lower the cost of using large language models for routine business tasks. Lower prices make high-volume workloads such as document analysis, customer-interaction classification and routine implementation more economical to run at scale.

Model split

OpenAI positions the GPT-5.6 family around different levels of complexity and cost. Sol is aimed at harder problems, Terra at everyday production work and Luna at high-volume workflows.

The split reflects a broader AI market trend as suppliers offer corporate customers a menu of trade-offs between speed, quality and price, rather than a single flagship system for every task. In practice, businesses may use one model to plan or resolve uncertainty and another to carry out defined steps more cheaply.

OpenAI said Luna can use tools and complete multi-step workflows, expanding the range of jobs that can be handled at lower cost. It added that Luna delivers performance comparable to models that were at the frontier a year ago, at a fraction of the cost and at much higher speed.

One benchmark cited by OpenAI compared Luna with a rival model called Fable 5 on professional work measured by Agents' Last Exam. According to the company, Luna outperformed that model at an estimated cost per task nearly 99% lower.

Efficiency drive

The price cuts follow efficiency improvements in how OpenAI builds and runs its models. Gains came from model design, inference systems and the software layer that links models with tools and context.

According to OpenAI, GPT-5.6 models now take a more direct route through work, with better routing of jobs across hardware, more efficient token generation and context management that avoids repeating completed work. Those changes reduce the time, tokens and cost required for each result.

OpenAI also said GPT-5.6 Sol helped identify some of those gains within a human-led process. Sol rewrote and optimised production kernels, designed and ran hundreds of experiments to improve token generation, and monitored training runs while intervening when problems arose.

That work helped reduce the end-to-end cost of serving the model by 20%, according to OpenAI, while experiments increased token-generation efficiency by more than 15%. The company argued that this creates a feedback loop in which stronger models help improve the systems that run them.

Customer impact

For customers, the immediate effect is likely to be most visible in budgeting and workload allocation. Companies that had reserved larger models for a narrow set of tasks may now re-evaluate where lower-cost models can produce acceptable results, especially in coding, back-office workflows and classification tasks.

OpenAI also said subscription prices and quota budgets for ChatGPT and Codex will not change, even though Terra and Luna usage will consume fewer credits. That means existing subscribers may be able to stretch current budgets further without moving to a higher plan.

The broader market context is intensifying competition on model economics as providers try to persuade businesses to move AI from pilots into routine operations. While top-end model performance still draws attention, pricing, latency and operational efficiency are becoming more important in enterprise purchasing decisions.

OpenAI included one customer endorsement in the announcement. "GPT-5.6 Luna is the closest we've come to intelligence too cheap to meter. I've never seen a model this affordable be this powerful - it's unlocking use cases for Replit we didn't expect to build for a long time," said Michele Catasta, President & Head of AI, Replit.