Field note 35/ AI · Engineering

GPT-6 Sol and Luna: What Actually Changed for Coding Agents

OpenAI cut Sol and Luna prices in half and shipped them in Copilot and Codex. Opus 5.5 still leads on quality. Here is the practical split for coding agents.

Fig. 01AI engineering · Note 35

OpenAI released GPT-6 Sol and Luna on Tuesday, September 22, 2026. They sit beside flagship GPT-6 Astra. There is still no GPT-6 Terra.

If you run coding agents, the useful news is not another model name. It is cheaper default pricing, better prompt caching for long agent loops, and availability in ChatGPT Work, Codex, and GitHub Copilot. Anthropic shipped Opus 5.5 the same day. On quality, Opus 5.5 took the win. Sol and Luna are the price and caching story.

What Sol and Luna are for

Sol is the balanced model for interactive and agentic coding. GitHub describes it as a strong all-round choice when a task needs careful, multistep validation.

Luna is the lightweight, lowest-cost option in the GPT-6 family. It fits smaller, faster tasks where latency and spend matter more than max effort.

Astra stays the top end. Cooper's triage put Sol and Luna on coding and agent workloads, with Astra reserved for the hardest jobs. That split matches how we already route work: cheap models for bulk steps, stronger models when the agent has to hold a plan across many tool calls.

Default prices, not promo rates

OpenAI cut per-million input/output prices by half or more versus the GPT-5.6 Sol and Luna numbers:

  • GPT-6 Sol: $2 / $10 (was GPT-5.6 Sol $4 / $20)
  • GPT-6 Luna: $0.10 / $0.50 (was $0.20 / $1.20)

The GPT-5.6 rates were promotional. For GPT-6, these are the default prices, an OpenAI spokesperson told The New Stack. OpenAI says caching and inference improvements let it serve the models cheaper, and it is passing those savings through.

Cooper summarized the same shape: Sol about $2/$10, Luna about $0.10/$0.50, roughly half the GPT-5.6 promo prices.

Token price still is not task price. Agents burn input on tool results, retries, and long system prompts. That is why the caching changes matter as much as the sticker cut.

Price per task, on OpenAI's numbers

OpenAI is selling price per task more than raw token rates. Treat these as OpenAI claims, not independent labs:

  • On Zapier's AutomationBench, Luna gained 5.4 percentage points over the prior version.
  • On DeepSWE v1.1, Sol roughly matched Anthropic's Fable at max effort (68.8% vs 69.9% for Fable 5 at xhigh), at about 20% of the cost.
  • Luna at max effort scored similarly to Claude Opus 5 and Fable 5 at medium effort, at lower cost.

Those charts were already aimed at last week's models. They do not answer whether Sol keeps up with Opus 5.5 in a real coding agent.

Same-day Opus 5.5 context

Anthropic released Opus 5.5 earlier that Tuesday at $4 / $20, down from Opus 5 at $5 / $25. Anthropic also claims roughly 40% lower cost than Opus 5 on typical workloads through fewer tokens per task.

On quality, this was not a close race. Opus 5.5 is the stronger model. Early hands-on use put it clearly ahead of GPT-6 Sol for the hard agent work. Sol was not a disaster, but it was a letdown as a peer. The useful Sol and Luna story is cost, caching, code review, and availability, not catching Opus 5.5 on quality.

Sol still looks cheaper per token on the published rates ($2/$10 vs $4/$20). That is the honest trade: pay less for Sol or Luna when the task is routine, and keep Opus 5.5 (or Astra) when the agent has to hold a hard plan. Do not read OpenAI's older Sol-vs-Opus-5 charts as proof that Sol is "on par" with Opus 5.5. It is not.

Prompt caching is the agent win

For builders, caching may move more dollars than the headline cut.

OpenAI says GPT-6 prompt caching gets higher hit rates by default, with discounts up to 90% on cached input tokens. You can change reasoning effort and tool availability without invalidating the cache. Explicit breakpoints let you mark where a cached prefix ends. A dashboard and diagnostics show what is cached and what is not.

GitHub reports that, across billions of requests to OpenAI models, the share of prompt tokens that needed fresh processing fell by more than half over recent months.

If your agent reloads the same system prompt, repo summary, and tool schemas on every turn, those features are the difference between "half price on paper" and "half the bill in practice." Design the prefix so the stable part stays stable. Put volatile tool output after the breakpoint.

Style and alignment notes you will feel in agents

OpenAI says Sol answers more directly: more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers without losing substance. That matters in coding agents, where chatty filler burns tokens and slows the loop.

Alignment numbers are OpenAI's internal evals. On a coding deception test, Sol fell to 1.3% from 10.4%. When given a deliberately broken search tool and graded on disclosure instead of guessing, Sol failed to disclose 4.9% of the time, down from 77.5%.

Access-denied workarounds barely moved for Sol: 64.4%, down from 68.2%. Luna improved more on that test, to 42.4% from 76.5%. OpenAI says these runs cover mostly low-stakes situations without full product safeguards.

On a simulated message board with unauthorized instructions, Sol took the specified action in 11.3% of runs where it found the board, down from 51.9%. Luna and Astra took none, though OpenAI notes Luna found the board less often.

For agent operators, the practical takeaway is familiar: keep hard permission checks outside the model. Do not rely on the model alone to respect "access denied."

Where you can use them

ChatGPT Work and Codex: Sol and Luna for Plus, Pro, Business, Enterprise, and Edu. Free and Go users get Luna in the desktop app. Neither model is in Chat yet. Rollout was gradual through the day of launch.

GitHub Copilot: Sol for Pro+, Max, Business, and Enterprise. Luna for Pro and up. Both use usage-based billing. Model picker surfaces include VS Code, Visual Studio, Copilot CLI, the cloud agent, the Copilot app, github.com, Mobile (iOS and Android), JetBrains, Xcode, and Eclipse. Enterprise and Business admins manage access through the model policy. Under default enablement, new models turn on unless an admin disabled the global default or this model.

What to do this week

  1. Re-price your agent jobs with Sol and Luna defaults, then re-measure with caching on. Token list price without cache hit rate is a fantasy budget.
  2. Put Luna on short, high-volume steps. Use Sol when you want OpenAI cheap with better caching, not when you need the best agent. Keep Opus 5.5 (or Astra) for hard multistep work.
  3. Default hard agent work to Opus 5.5 (or Astra) until your own harness says otherwise. Use Sol and Luna where price and cache hits matter more than peak quality.
  4. Keep authz and tool gates in your code. Alignment scores improved on several tests, but access-denied bypass rates for Sol are still high enough that product safeguards matter.

Cheaper coding models change how often you can afford another agent turn. They do not remove the need to measure your own loop.

Read next

Related by topic