GPT-5.6 Luna
Officialopenai/gpt-5.6-luna
The efficient GPT-5.6 model for high-volume assistants, extraction, routing, and coding tasks. It keeps the 1.05M-token context window, vision, tools, configurable reasoning, persisted reasoning, and explicit prompt caching at the lowest price in the family.
Pricing
Source: openai_x0.10 · Verified 2026-08-02
Fast $0.04 / $0.24 /M
Cache read $0.002/M · Cache write $0.025/M
| Tier | Input | Cache read | Cache write | Output |
|---|---|---|---|---|
| Standard ≤272K | $0.02 | $0.002 | $0.025 | $0.12 |
| Standard >272K | $0.04 | $0.004 | $0.05 | $0.18 |
| Priority ≤272K | $0.04 | $0.004 | $0.05 | $0.24 |
Estimate cost
Protocols
- OpenAI Chat CompletionsStatus: Available: Streaminghttps://api.routemux.com/v1
- OpenAI ResponsesStatus: Available: Streaminghttps://api.routemux.com/v1
- Anthropic MessagesStatus: Available: Streaminghttps://api.routemux.com/anthropic
- Google Vertex AIStatus: Not supported—
- OpenAI ImagesStatus: Not supported—
- OpenAI SoraStatus: Not supported—
Call IDs
Use either ID to call this model via the API.
openai/gpt-5.6-lunaTry it
Replace the ROUTEMUX_KEY placeholder with your API key. Create one →
from openai import OpenAI
client = OpenAI(
base_url="https://api.routemux.com/v1",
api_key="<ROUTEMUX_KEY>",
)
completion = client.chat.completions.create(
model="openai/gpt-5.6-luna",
messages=[{"role": "user", "content": "What is the meaning of life?"}],
)
print(completion.choices[0].message.content)Other variants in this family
OpenAI's flagship GPT-5.6 for the hardest reasoning, coding, and agentic work.
$0.5/M↓ · $3/M↑
Balanced GPT-5.6 quality and cost for production workloads.
$0.2/M↓ · $1.2/M↑
OpenAI GPT-5.4 mini — fast, cheapest GPT-5 text route.
$0.075/M↓ · $0.45/M↑
OpenAI GPT-5.4 — strong general model, 1M context, lower cost than 5.5.
$0.25/M↓ · $1.5/M↑
OpenAI GPT-5.5 — flagship, 1M context, top reasoning & agentic coding.
$0.5/M↓ · $3/M↑
Frequently asked questions
›What is GPT-5.6 Luna?
The most efficient GPT-5.6 for high-volume and latency-sensitive work.
›How large is the context window of GPT-5.6 Luna?
GPT-5.6 Luna supports a context window of up to 1.1M tokens, with up to 128K output tokens per request.
›How much does the GPT-5.6 Luna API cost?
Through RouteMux, GPT-5.6 Luna costs $0.02 for input and $0.12 for output (USD per 1M tokens), below the official list price.
›How do I call GPT-5.6 Luna via API?
GPT-5.6 Luna is OpenAI-compatible: point your base URL at RouteMux and set the model field to openai/gpt-5.6-luna — no code changes needed.
›Which input and output modalities does GPT-5.6 Luna support?
GPT-5.6 Luna accepts text, image as input and produces text as output.
›When was GPT-5.6 Luna released?
GPT-5.6 Luna was released on July 9, 2026.