All Things AI
Intermediate

Models in ChatGPT

ChatGPT gives you access to a family of models spanning instant responses, deep reasoning, and long-context document analysis. Understanding the differences between them - and which plan unlocks each - lets you choose the right tool for each task rather than defaulting to the most expensive option.

Status update (Sep 2026): OpenAI released GPT-6 Astra on September 3, 2026 - a flagship agentic model that operates software through screens (fills forms, edits spreadsheets, drives engineering tools), initially to a limited set of organisations. The API-facing generation has moved through GPT-5.5 and GPT-5.6 (tiers renamed Sol/Terra/Luna) since this page's GPT-5 Instant/Thinking/Pro breakdown below was written. The consumer ChatGPT app tier names may have changed further by the time you read this - verify current tier names at chatgpt.com before relying on the specific labels below. The underlying pattern (a fast tier, a reasoning tier, a maximum-compute tier) has held across every generation so far.

Quick Reference - Model Comparison

ModelContextAPI Input ($/1M)API Output ($/1M)AccessBest For
GPT-5 Instant128K--Free / Go / Plus / ProEveryday tasks, drafting, fast responses
GPT-5 Thinking196K--Plus / Pro / Team / EnterpriseComplex reasoning, analysis, detailed tasks
GPT-5 Pro1M--Pro only ($200/mo)Maximum quality, research-grade tasks
GPT-6 Astra (API only)1.05M$2.00$10.00API onlyHighest reasoning quality, autonomous agentic pipelines
GPT-6 Luna (API only)1.05M$0.10$0.50API onlyCost-efficient reasoning at scale

ChatGPT interface pricing is subscription-based (no per-token billing). API pricing as of Sep 2026 - verify at openai.com/api/pricing. Consumer-tier prices are not separately published for the API; the relevant API models are listed separately at openai.com/api/pricing.

A Note on Parameters

OpenAI does not publicly disclose parameter counts for any model in the GPT-4 or GPT-5 generation. Reported figures in the press are estimates or leaks and should not be treated as authoritative. What OpenAI does characterise is relative capability, speed, and cost - and those are the dimensions that matter for practical use.

The GPT-5 Family

GPT-5 is the current foundation model generation powering ChatGPT. OpenAI offers it in three tiers, each representing a different compute budget:

GPT-5 Instant

The fastest and most cost-efficient variant. Targets a 128K token context window. Available on Free and Go tiers. Handles everyday tasks - drafting, summarisation, Q&A, light coding - with low latency. Not suitable for tasks requiring deep multi-step reasoning or exhaustive analysis.

GPT-5 Thinking

The extended reasoning variant. Targets a 196K token context window. Available on Plus and above. Before producing its response, the model performs an internal chain-of-thought process - breaking down problems, exploring approaches, and self-checking conclusions. This produces noticeably better results on complex reasoning, maths, legal analysis, and technical writing, at the cost of higher latency and token consumption.

GPT-5 Pro

The maximum compute variant, exclusive to the Pro plan. Used for the hardest research-grade tasks where accuracy and depth matter more than speed. Powers the most demanding Deep Research sessions and the highest-quality outputs across all task types. Not available on any other plan.

Reasoning Models: From o-Series to Built-In Adaptive Thinking

Status update (Sep 2026): o3 and o4-mini, and the GPT-5-era "Thinking" tier that briefly replaced them in the ChatGPT interface, have both been superseded. As of the GPT-6 generation, adaptive thinking is built directly into the flagship model (GPT-6 Astra) rather than shipped as a separate reasoning-only SKU - the same shift Anthropic made with Claude Fable 5.1's always-on adaptive thinking. A dedicated, separately-branded reasoning tier is no longer how either lab ships this.

The older o-series models were architecturally distinct from standard GPT models - rather than generating a response immediately, they allocated additional "thinking time" before arriving at an answer. That same behaviour is now a built-in mode of the flagship model rather than a separate model family.

GPT-6 Astra

The flagship, with adaptive thinking built in. 1.05M context window. Via API at $2.00/$10.00 per 1M tokens. Excels at olympiad-level mathematics, complex code debugging, scientific reasoning, and multi-step logical proofs, and can operate software autonomously (fills forms, edits spreadsheets, drives engineering tools).

GPT-6 Luna

A faster, much cheaper tier of the same generation. 1.05M context window. Via API at $0.10/$0.50 per 1M tokens - roughly 20ร— cheaper than Astra. The default choice for high-volume tasks that don't need Astra's full agentic capability.

Previous Generation: GPT-4o, GPT-4.1, and GPT-5.6

GPT-4o and GPT-4.1 were OpenAI's multimodal and long-context specialists through 2025, and GPT-5.5/GPT-5.6 (with Sol/Terra/Luna tiers) carried the API generation through most of 2026 before GPT-6 Astra's September 2026 launch. GPT-5.6 Terra remains available at roughly $2.00/$12.00 per 1M tokens and is still broadly deployed - the mid-tier previous generation is often the more cost-effective choice when Astra's full agentic capability isn't needed.

Plan Availability Summary

ModelContextAvailable OnPrimary Use
GPT-5 Instant128KFree, Go, Plus, ProEveryday tasks, fast responses
GPT-5 Thinking196KPlus, Pro, Team, EnterpriseComplex reasoning, detailed analysis
GPT-5 ProLargePro onlyMaximum quality, research-grade output
GPT-6 Astra1.05MAPI (limited rollout); ChatGPT (rolling out)Formal logic, maths, code correctness, autonomous computer use
GPT-6 Luna1.05MAPIFast, cheap reasoning at scale
GPT-5.6 Terra1MAPI (all); ChatGPT (previous-gen)Multimodal tasks, mid-tier cost efficiency

How Reasoning Differs From Standard Generation

A model without adaptive thinking responds by predicting the most likely next token given the conversation history - a fast, fluent process. A model with adaptive thinking (GPT-6 Astra, Claude Fable 5.1) inserts a deliberate thinking phase before generating output. This internal monologue is hidden from users but visible in the API response's reasoning_tokens count. The model may explore multiple solution paths, catch errors in its own reasoning, and revise its approach before producing the final answer.

The practical implication: use a fast tier (GPT-6 Luna, Claude Haiku 4.5) for speed-sensitive tasks (chat, drafting, summarisation), switch to the adaptive-thinking flagship (GPT-6 Astra) when accuracy matters more than latency (debugging, proofs, complex analysis), and reserve the highest tier (Claude Fable 5.1) for the most demanding research-grade outputs.

Checklist

  • Why did OpenAI and Anthropic both move away from a separate reasoning-only model family?
  • How does adaptive/built-in thinking differ architecturally from a standard fast-response model?
  • Which OpenAI model is the current flagship, and what distinguishes its agentic capability?
  • What makes GPT-5.6 Terra still useful despite being superseded by GPT-6 in most tasks?
  • When would you choose GPT-6 Luna over full GPT-6 Astra for a reasoning task?