Foundation Models
Model releases, benchmark results, pricing changes, and open-weight developments. Curated for builders who need to track what is available and what it costs.
Claude Fable 5.1 and Opus 5.5 - Anthropic's September Refresh
Fable 5.1 became the new flagship with always-on adaptive thinking and a 1M token context window; Opus 5.5 followed as a near-Fable-quality model at roughly 40% lower cost than the prior Opus generation, also extended to 1M context.
Why it matters: The Opus 5.5 pricing move is the notable part: Anthropic is explicitly offering a 'near-flagship for less' tier rather than only pushing the top end higher, reinforcing the 2026 shift toward cost-aware model selection covered in Model Economics.
xAI Ships Grok 4.7
Grok 4.7 reached OpenRouter at $1.60/$4.80 per 1M tokens, xAI's latest release in a rapid year of Grok iteration (3 to 4.3 to 4.5 to 4.7).
Why it matters: Grok continues to compete primarily on cost and X/Twitter-grounded real-time knowledge rather than raw benchmark leadership - a distinct positioning worth knowing when picking a model for social-data-aware applications.
OpenAI Launches GPT-6 Astra - Flagship Agentic Model Operates Software Autonomously
GPT-6 Astra can operate software through the screen the way a person would - filling forms, updating CRM records, editing spreadsheets, and driving engineering tools - initially released to a limited set of organisations, with cheaper Sol and Luna API tiers following later in September.
Why it matters: This is the clearest signal yet that the frontier race has moved from 'better chat answers' to 'autonomous computer operation.' Teams building computer-use agents (see Computer Use Agents) now have a first-party flagship model built for exactly that, not a general chat model repurposed for it.
Google Ships Gemini 3.8 Flash to Stable GA
Gemini 3.8 Flash reached stable general availability, the fourth Flash point release in the 3.x generation this year (3.5 to 3.6 to 3.7 to 3.8), at introductory pricing through the end of 2026.
Why it matters: Google's cadence of frequent, GA-stable Flash point releases within a single generation is now a distinct competitive strategy - aggressive cost reduction and iteration speed rather than big-bang generational jumps.
Gemini 3.6 Flash GA - Better Coding, Computer Use, and Lower Output Pricing
Google releases Gemini 3.6 Flash with output pricing cut to $7.50/M tokens (from $9.00), 49% DeepSWE coding score (up from 37% for 3.5 Flash), and 83% on OSWorld computer-use benchmark (up from 78.4%). Google also launches Gemini 3.5 Flash Cyber - a dedicated model for software vulnerability detection. Gemini 4 was teased during the announcement.
Why it matters: 3.6 Flash is a meaningful step up from 3.5 Flash on both coding and computer use, at a lower price point. The cybersecurity-specific model (Flash Cyber) signals specialised models for security workflows as a product category - following OpenAI Sol's cybersecurity emphasis. The Gemini 4 tease sets up the next capability leap for the Google ecosystem.
GPT-5.6 Family GA - Sol Tops Coding Agent Index at $5/M Input
OpenAI launches GPT-5.6 Sol, Terra, and Luna to general availability. Sol (flagship) leads the Artificial Analysis Coding Agent Index at 80 points, scores 59 on the Intelligence Index (1 point below Claude Fable 5), and is 54% more token-efficient on AI coding tasks than its predecessor. All three models support 1M token context windows. Sol becomes the new ChatGPT default.
Why it matters: Sol resets expectations on coding agent cost-efficiency - frontier-class performance with materially fewer output tokens means lower real-world cost than headline pricing suggests. The three-tier family (Sol/Terra/Luna) gives developers a clear cost/capability ladder, matching Anthropic's and Google's tiered structures and finally completing OpenAI's product lineup for agentic use cases.
Grok 4.5 Released - 4th on Intelligence Index at 60% Lower Cost Than Competing Flagships
xAI releases Grok 4.5, built on a 1.5 trillion-parameter V9 foundation model with Cursor coding data mixed into training. Priced at $2/$6 per 1M tokens - more than 60% cheaper than Claude Opus 4.8 or GPT-5.5. Ranks 4th on the Artificial Analysis Intelligence Index, above all Gemini models and all open-weight models. Available in Grok Build, Cursor, and the SpaceXAI console.
Why it matters: Grok 4.5 breaks the pattern of frontier-quality implying frontier-tier pricing. At $6/M output tokens vs $25-50 for competing flagships, it becomes the default cost/performance choice for coding-heavy agentic workloads that don't require Fable 5 or Sol's peak capability. xAI is pricing aggressively to grow developer mindshare.
Mistral Releases Robostral Navigate - Robotics Navigation Model Trained Entirely in Simulation
Mistral releases Robostral Navigate, a robotics navigation model that lets robots navigate complex environments using a single camera and natural language prompts. Hardware-agnostic, trained entirely in simulation with no real-world data collection. Part of a broader Mistral push into physical AI, alongside a new open-weight Mixture-of-Experts family entering early access (July 6).
Why it matters: Simulation-trained navigation that transfers to real hardware removes the most expensive bottleneck in physical AI development - large-scale real-world data collection. Hardware-agnosticism means it deploys on existing robot platforms without custom training. Mistral is staking out a position in physical AI at the same moment the model market is commoditising on the software side.
Claude Fable 5 Export Controls Lifted - Restored Globally After Cybersecurity Classifier Deployed
The US Commerce Department lifted the emergency export-control order on Claude Fable 5, which had been forced offline June 12 after a prompt was found to bypass its safeguards and identify software vulnerabilities. Anthropic deployed a new cybersecurity safety classifier and restored full global access on July 1 across Claude.ai, the Claude Platform, Claude Code, and Claude Cowork.
Why it matters: The full saga - launch, government shutdown, classifier fix, reinstatement - played out in under three weeks. It established a new precedent: US authorities can and will act in near-real-time to restrict a frontier model's global availability on national security grounds. The classifier-and-reinstate resolution shows a viable pathway, but the speed of the intervention signals permanent real-time government oversight of the most capable models.
Claude Sonnet 5 - Near-Opus Performance, Now Default Across All Claude Plans
Anthropic releases Claude Sonnet 5 at $2/$10 per 1M tokens (introductory through August 2026), with near-Opus 4.8 performance. Adaptive thinking is on by default; a new effort-level dial (low to xhigh) lets developers tune reasoning depth per request. Prompt caching delivers up to 90% cost savings. Now the default model for Free, Pro, Max, Team, and Enterprise plans.
Why it matters: Sonnet 5 makes frontier-class reasoning the default, not an expensive option. The effort-level dial is a practical new tool for cost-optimising agentic pipelines - cheap reasoning for simple tasks, expensive reasoning only where it earns its cost. Introductory pricing makes this the most cost-efficient frontier model available at launch.
OpenAI Launches GPT-5.6 Sol, Terra, Luna - Limited to Trusted Partners at US Government Request
OpenAI announces three new frontier models - GPT-5.6 Sol (strongest model yet, with top performance on coding and cybersecurity), Terra, and Luna - but limits the launch to a small group of trusted partners at the US government's request. OpenAI previewed capabilities with the government ahead of launch and says it is developing a repeatable process for future model releases under this framework.
Why it matters: Government-gated frontier model launches are becoming a structural feature of the AI landscape. Following the brief suspension of Claude Fable 5 weeks earlier, GPT-5.6's restricted launch confirms the US is establishing a national security review process for the most capable AI systems. For enterprises and developers, this means top-tier model access will increasingly depend on trust tier and jurisdiction - not just API keys.
Claude Fable 5 - Anthropic's Frontier Model Launched Then Briefly Pulled by US Export Order
Anthropic releases Claude Fable 5 - state-of-the-art across coding, scientific research, vision, and knowledge work at $10/$50 per 1M tokens (less than half Claude Mythos Preview pricing). Within three days of launch, a US government export directive temporarily forced it offline; it was reinstated June 22 for Pro, Max, Team, and Enterprise subscribers.
Why it matters: Fable 5 sets a new capability ceiling for commercially available models. The temporary government-directed suspension - the first of its kind for a major AI model - signals the US is moving toward active national security oversight of frontier model deployments, not just policy guidance. This precedent will shape how Anthropic and other labs release their most capable models going forward.
Microsoft Build 2026 - MAI-Thinking-1 and MAI-Code-1-Flash, First In-House Models
Microsoft announces 7 MAI models at Build 2026. MAI-Thinking-1 (35B active params, 256K context) is Microsoft's first in-house reasoning model - trained without OpenAI distillation, preferred over Claude Sonnet 4.6 in blind evals. MAI-Code-1-Flash (5B params) integrates deep into GitHub Copilot and VS Code for all plans.
Why it matters: Microsoft's first home-built frontier-class models mark a strategic shift away from pure OpenAI dependency. MAI-Code-1-Flash becoming the backbone of GitHub Copilot affects every developer using Copilot - Microsoft controls the model and pricing, not OpenAI.
Claude Opus 4.8 - #1 on AI Intelligence Index, Dynamic Workflows for Claude Code
Anthropic releases Claude Opus 4.8 at $5/$25 per 1M tokens - the only model to complete every case end-to-end on the Super-Agent benchmark. Scores 84% on Online-Mind2Web (browser agent eval). Introduces Dynamic Workflows in Claude Code for very large-scale tasks. Fast mode now 3ร cheaper than for previous Opus.
Why it matters: Opus 4.8 is the new benchmark leader for agentic and computer-use tasks. For builders running autonomous coding agents or browser agents, it outperforms GPT-5.5 on key evals while matching it on cost. The 4ร reduction in unremarked code flaws vs Opus 4.7 matters for production code review workflows.
Gemini 3.5 Flash GA - Google's Speed/Quality Leader for Agents and Coding
Gemini 3.5 Flash reaches general availability following its Google I/O announcement. Achieves 284 tokens/second and an Intelligence Index score of 55 - strong coding and agentic performance at Flash-tier pricing. Gemini 3.5 Pro (2M token context, Deep Think reasoning mode) committed for June GA.
Why it matters: Gemini 3.5 Flash is now the natural default for Google-stack builders who need fast agentic throughput. The 2M context window on Pro will make whole-codebase and large document workflows more practical than any current model.
Qwen3 - Alibaba's New Model Family with Thinking Mode
Alibaba releases Qwen3 family: 0.6B to 235B MoE. All models support a 'thinking mode' toggle (chain-of-thought on/off per request). Qwen3-Coder 480B MoE targets software engineering. Apache 2.0 licensed. Strong in 29 languages.
Why it matters: Thinking-mode toggle is a practical innovation - use fast mode for simple tasks, reasoning mode for complex ones, in the same model. Qwen3 Coder 480B directly challenges frontier coding models at open-weight prices. Extends Chinese AI labs' impact on the open-weight ecosystem.
Grok 4.3 - xAI's Updated Flagship with 1M Context
xAI releases Grok 4.3 - current flagship at $1.25/$2.50 per 1M tokens with a 1M token context window. Positioned between Claude Sonnet and Opus on price/quality. $150/month in free developer credits available via data-sharing program.
Why it matters: 1M context window at competitive pricing makes Grok 4.3 a viable choice for long-document tasks. The free $150/month credit program is the most generous developer offer from any major AI lab in 2026 - lowering the barrier to experiment with xAI's API significantly.
GPT-5.5 Launches - OpenAI's Most Capable General Model, $5/$30 per 1M Tokens
GPT-5.5 launches across ChatGPT and API, priced at $5/$30 per 1M tokens. It matches GPT-5.4 latency while performing at a significantly higher level - excels at multi-step agentic tasks, coding, and tool use, using fewer tokens per Codex task than its predecessor.
Why it matters: Sets a new capability/efficiency bar for paid frontier models. The combination of higher intelligence and lower token usage makes it more economical for agentic pipelines than it first appears from headline pricing. GPT-5.5 Pro ($30/$180) targets research-grade complexity.
Llama 4 Scout & Maverick - Meta Ships Open-Weight Multimodal MoE
Meta releases Llama 4 Scout (109B MoE, 10M token context) and Maverick (400B MoE, 17B active parameters). Both are natively multimodal, Apache 2.0 licensed, and match or beat GPT-4o on major benchmarks at a fraction of the inference cost.
Why it matters: First open-weight models to seriously challenge frontier closed-source models on quality. The 10M context window on Scout is the largest of any openly available model. MoE architecture means inference cost scales with active parameters (17B), not total (400B). Shifts the open vs closed model debate significantly.
DeepSeek Releases R2 - Open-Weight Reasoning Model
DeepSeek R2 achieves competitive reasoning performance with an open-weight license, making advanced reasoning accessible to self-hosted deployments.
Why it matters: Open-weight reasoning models reduce dependency on closed APIs for complex tasks. Important for enterprises with data residency requirements.
GPT-5 Launches - OpenAI Frontier Model with 400K Token Context
GPT-5 launches as OpenAI's new flagship with a 400K token context window, strong AIME 2025 maths performance, and significantly improved multi-step project execution and autonomous coding capability.
Why it matters: Sets a new capability baseline for closed frontier models. The 400K context window makes whole-codebase and large document reasoning practical via API. Forces pricing and capability recalibration across all competing providers.
Anthropic Releases Claude Opus 4 - Most Capable Model Yet
Claude Opus 4 sets new benchmarks across coding, reasoning, and extended thinking tasks, with improved tool use and agentic capabilities.
Why it matters: Represents a significant step in model capability for builders relying on agentic workflows and complex multi-step reasoning.
OpenAI Releases o3 and o4-mini - Reasoning Models with Native Tool Use
o3 and o4-mini combine chain-of-thought reasoning with native tool use, enabling models to search the web, run code, and call APIs mid-reasoning.
Why it matters: Reasoning + tool use in a single model removes the need to orchestrate separate search and reasoning steps, simplifying agentic pipeline design.
Meta Releases Llama 4 - Natively Multimodal Open-Weight MoE Models
Meta releases Llama 4 Scout (17B active params, 10M token context, runs on a single H100) and Maverick (17B active/400B total, 1M context) - the first natively multimodal Llama models trained on text, images, and video data.
Why it matters: Llama 4 is the new open-weight baseline for self-hosted multimodal deployments. Enterprises with data residency requirements now have a competitive open alternative to closed frontier models at a fraction of the API cost.
Google Gemini 2.5 Pro - 1M Token Context and Thinking Mode Released
Gemini 2.5 Pro adds a thinking mode (extended reasoning) alongside its 1M token context window, topping key benchmarks including coding and maths.
Why it matters: 1M context makes whole-codebase and whole-document analysis practical. Thinking mode brings reasoning capability to Google's ecosystem.