News

OpenAI Cuts GPT-5.6 Luna Prices by 80% in AI Cost War

2 Aug 2026 By OfficeForge's AI team · human-reviewed 6 min read
OpenAI Cuts GPT-5.6 Luna Price 80% in AI Cost War

The AI price war has officially escalated. OpenAI has announced substantial price cuts for its GPT-5.6 model series, led by an 80% reduction on its smallest and fastest model, GPT-5.6 Luna. This aggressive move, described by CEO Sam Altman as "major price cuts today," directly targets competitors like Google and Anthropic, shifting the frontier model competition decisively toward cost efficiency. The source details the changes, which position Luna as a disruptive force in the low-cost inference tier.

What Changed: The Price Cuts

OpenAI's pricing overhaul affects two of the three models in its frontier GPT-5.6 family and introduces a new performance tier for its flagship.

* GPT-5.6 Luna (80% cut): The most significant change. Previously priced at $1 per million input tokens and $6 per million output tokens (combined $7), Luna now costs $0.20 per million input tokens and $1.20 per million output tokens. This brings its combined price down to $1.40 per million tokens. * GPT-5.6 Terra (20% cut): The mid-tier model's price dropped from a combined $17.50 to $14 per million tokens ($2 input, $12 output). This now matches the pricing of Google's Gemini 3.1 Pro Preview for context windows of 200,000 tokens or less. * GPT-5.6 Sol Fast Mode (new): The flagship model now offers a premium "Fast" mode. Priced at $10 per million input tokens and $60 per million output tokens (combined $70), it promises up to 2.5 times the throughput compared to the Standard mode's $5/$30 pricing, without altering the model's underlying intelligence.

The price for GPT-5.6 Sol Standard remains unchanged at $5/$30 ($35 combined).

The Competitive Landscape: A Race to the Bottom

These cuts are a direct strategic response to recent moves by other providers. Just days before, Anthropic released its Claude Opus 5 at the same price point as its predecessor, and Google launched Gemini 3.6 Flash and Gemini 3.5 Flash-Lite, models optimized for lower costs and faster agent workloads.

OpenAI's Luna cut is particularly aggressive. While not the absolute cheapest model on the market—Xiaomi's MiMo-V2.5 Flash and DeepSeek's flash models remain less expensive—Luna's new price brings a "frontier-series" model from OpenAI into direct competition with the budget tier. According to a pricing comparison table, Luna now undercuts several models, including: * Google's Gemini 3.5 Flash-Lite ($2.80 total) * MiniMax-M3 ($1.50 total) * OpenAI's own older GPT-5.4 model ($17.50 total)

As noted by Krea AI's Nic Dunz, the Terra cut effectively offers the same intelligence as GPT-5.4 for roughly one-thirteenth the cost, illustrating the pace of value improvement.

Why This Matters for Self-Hosted and Agent-Based Teams

For businesses and technical teams, especially those building AI-powered workflows or agents, this price war is overwhelmingly positive. However, the biggest winners are not teams consuming AI through all-inclusive SaaS subscriptions. The greatest value accrues to teams operating with a self-hosted architecture and a "bring your own key" (BYO) model.

When you pay an API provider like OpenAI, Anthropic, or Google directly via your own key, you capture 100% of any price reduction immediately. There is no middleman markup on tokens, no "platform fee" that absorbs the savings, and no waiting for a SaaS vendor to pass the discount along. A self-hosted system that routes requests to the most cost-effective model for each task becomes exponentially more powerful as the price-performance ratio of frontier models improves dramatically.

The OfficeForge Advantage: This dynamic is core to the self-hosted AI team model. With OfficeForge, you connect your own OpenRouter, OpenAI, or Anthropic key. You pay the provider's direct API rate—like these new Luna prices—without any token markups. You can also assign the optimal model to each AI "employee": use a powerful, costlier model for your Coder's complex tasks and a super-efficient, cheap model like Luna for your Secretary's email summaries and routing. This granular cost control is only possible with full ownership.

Get OfficeForge — $199

This strategic flexibility is crucial. An agent workflow might call the flagship Sol model for a complex planning step, switch to the mid-tier Terra for execution, and use dozens of calls to the ultra-cheap Luna for verification, formatting, or data extraction. With BYO keys, you pay the blended rate for all these calls directly. A SaaS platform would typically charge a flat, higher rate per "seat" or "action," obscuring the underlying cost and limiting your ability to optimize.

The Broader Shift: Cost as the New Battleground

The introduction of Sol Fast mode at a significant premium also reveals a market segmentation strategy. Providers are now offering clear tiers not just for intelligence, but for speed and throughput, catering to different workload needs within a single model family. The competition is no longer just about which model is smarter, but about offering the right trade-off between intelligence, latency, *and* cost for a specific task.

This price war signals a maturation of the AI infrastructure market. As models become more commoditized at various capability tiers, the battleground shifts to economics. For businesses, this reduces a major barrier to adoption and makes large-scale, agent-driven automation financially viable.

Conclusion: Ownership Maximizes the Advantage

OpenAI's dramatic price cuts on GPT-5.6 Luna and Terra redefine the cost landscape for frontier AI models. The move intensifies competition, putting pressure on all providers to deliver better value. For technical teams and businesses building AI into their operations, the message is clear: the tools are becoming vastly cheaper.

The ultimate beneficiary of this trend is the team that controls its own stack. By owning the infrastructure and holding the API keys, you ensure that every dollar saved by a provider flows directly to your bottom line. This is the core philosophy behind building a permanent, cost-effective AI workforce on your own terms—a principle that makes the current price war not just news, but a strategic opportunity.

FAQ

What is the new price for GPT-5.6 Luna after OpenAI's cut?

GPT-5.6 Luna now costs $0.20 per million input tokens and $1.20 per million output tokens, for a combined total of $1.40 per million tokens.

How much did OpenAI reduce the price of GPT-5.6 Terra?

GPT-5.6 Terra received a 20% price cut. Its combined cost decreased from $17.50 to $14 per million tokens.

What is the new pricing for GPT-5.6 Sol Fast mode?

The new Sol Fast mode is priced at $10 per million input tokens and $60 per million output tokens, for a combined $70 per million tokens, offering up to 2.5x the throughput.

How does GPT-5.6 Luna's new price compare to Google's Gemini models?

Luna's combined $1.40 per million tokens is now cheaper than Google's Gemini 3.5 Flash-Lite ($2.80) and far below Gemini 3.6 Flash ($9).

Why are self-hosted AI teams a primary beneficiary of these API price wars?

Teams using a self-hosted setup with their own API key (BYO) pay the provider directly at these new, lower rates without any intermediary markup, maximizing the cost savings.

🛠

This article was researched, written and illustrated by OfficeForge's own AI team — Andrey (research), Kirill (writing), Alla (design) — the same five AI employees the product ships with. Founder-directed, human-reviewed. The blog is our product, doing real work.

This article was produced by the same AI team you can put on your own task board. Build your team →
On sale now

Run your own AI team

One-time purchase, your server, your data. The license key is emailed instantly.

Get OfficeForge — $199