All Things AI
Beginner

What's New in AI

A curated timeline of the milestones that actually changed how AI is built and used - from early 2024 through mid-2026. Not every model release; the ones that shifted the industry or changed what practitioners could do.

Last updated: Jul 2026

Model ReleaseProtocol / StandardTool / PlatformIndustry Shift

2026 (so far)

Sep 22, 2026Model Release

Claude Fable 5.1 and Opus 5.5 - Anthropic's flagship refresh

Anthropic shipped Fable 5.1, extending the flagship's context window to 1M tokens with always-on adaptive thinking, and Opus 5.5, a near-Fable-quality model at roughly 40% lower cost than the prior Opus generation - also bumped to 1M context. The Opus move is the notable one: rather than only pushing the ceiling higher, Anthropic is explicitly selling a 'near-flagship for less' tier, reinforcing 2026's broader shift toward cost-aware model selection over reflexively reaching for the top-of-line model.

Sep 21, 2026Model Release

Grok 4.7 lands on OpenRouter

xAI's Grok 4.7 reached OpenRouter at $1.60/$4.80 per 1M tokens - the latest in a rapid year of Grok releases (3 โ†’ 4.3 โ†’ 4.5 โ†’ 4.7). xAI continues to compete primarily on cost and real-time X/Twitter-grounded knowledge rather than chasing benchmark-leaderboard supremacy outright.

Sep 3, 2026Model Release

GPT-6 Astra - OpenAI's flagship moves from chat to autonomous computer use

OpenAI released GPT-6 Astra to a limited set of organisations. Astra operates software through the screen the way a person would - filling out forms, updating CRM records, editing spreadsheets, driving engineering tools - rather than just answering questions about them. Cheaper API tiers (Sol, Luna) followed later in the month at $2/$10 and $0.10/$0.50 per 1M tokens. The positioning is explicit: this is not a better chatbot, it's an operator.

Sep 2, 2026Model Release

Gemini 3.8 Flash reaches stable GA

Google shipped Gemini 3.8 Flash to stable general availability - the fourth Flash point release within the 3.x generation in under a year (3.5 โ†’ 3.6 โ†’ 3.7 โ†’ 3.8). Introductory pricing ($0.75/$3.75 per 1M tokens) holds through the end of 2026 before standard rates apply. The cadence itself is the story: Google is iterating within a generation rather than waiting for big-bang jumps.

Jul 21, 2026Model Release

Gemini 3.6 Flash - stronger coding, computer use, and a teaser of Gemini 4

Google released Gemini 3.6 Flash with improved coding (49% on DeepSWE, up from 37%) and computer use (83% on OSWorld, up from 78.4%), plus a cut in output pricing. Alongside it: Gemini 3.5 Flash Cyber - a specialised model for software vulnerability detection. Google teased Gemini 4 during the announcement, signalling another major capability jump is imminent.

Jul 18, 2026Tool / Platform

Alibaba Agent Native Cloud - first cloud platform purpose-built for multi-agent deployments

Alibaba unveiled Agent Native Cloud at WAIC 2026: a cloud architecture redesigned for multi-agent coordination at enterprise scale. Internal metrics reported: 15 agents handle 85% of developer support requests, 90% reduction in operational support time, software release cycles compressed to one day. Positions Alibaba Cloud as the first hyperscaler to build agent orchestration infrastructure from the ground up rather than bolting it onto existing compute.

Jul 10, 2026Protocol / Standard

EU AI Act chatbot disclosure rules take effect - Article 50 now enforceable

Article 50 of the EU AI Act became live and enforceable on July 10. Any business deploying AI chatbots interacting with EU users must now clearly disclose that users are talking to an AI. Separately, the EU's Digital Omnibus package (approved June 29) deferred the hard deadline for high-risk AI systems from August 2026 to December 2027, easing the immediate compliance pressure for regulated industries.

Jul 8-9, 2026Model Release

GPT-5.6 GA and Grok 4.5 - pricing war hits the frontier

OpenAI launched GPT-5.6 Sol/Terra/Luna to general availability on July 9 - Sol tops coding agent benchmarks and is 54% more token-efficient than its predecessor. xAI released Grok 4.5 on July 8 at $2/$6 per 1M tokens, 4th on the Intelligence Index, more than 60% cheaper than competing flagships. Two top-tier model releases in 24 hours - pricing competition at the frontier is now structural.

Jul 2, 2026Industry Shift

OpenAI files for IPO at $730B, proposes 5% US government equity stake

OpenAI filed a confidential S-1 with the SEC targeting a September 2026 IPO at a $730B valuation. It simultaneously proposed giving the US government a 5% equity stake - roughly $42.6B - modelled on the Alaska Permanent Fund. The proposal would require Congressional approval and would make the US government a co-owner of the world's most prominent AI company.

Jul 1, 2026Industry Shift

Fable 5 export controls lifted - full global access restored after 19-day shutdown

The US Commerce Department lifted its emergency export-control order on Claude Fable 5. Anthropic deployed a cybersecurity safety classifier addressing the vulnerability that triggered the June 12 shutdown and restored global access on July 1. The full cycle - launch, government order, classifier fix, reinstatement - completed in under three weeks, establishing the first real-time government intervention and reinstatement precedent for a frontier AI model.

Jun 30, 2026Model Release

Claude Sonnet 5 - adaptive thinking as default, near-Opus quality at Sonnet price

Anthropic releases Claude Sonnet 5 at $2/$10 per 1M tokens (introductory). Adaptive thinking is on by default - no separate toggle needed. A new effort-level dial (low to xhigh) lets developers tune reasoning depth per request, opening up fine-grained cost control for agentic pipelines. Becomes the default model across all Claude plans.

Jun 2026Protocol / Standard

Great American AI Act - first bipartisan draft of a US federal AI framework

Representatives Obernolte and Trahan release a 269-page bipartisan discussion draft of the Great American AI Act - the most credible federal AI legislation proposed to date. It would impose binding obligations on $500M+ revenue frontier AI developers and preempt state AI laws for three years. If enacted, it would replace the current state-by-state patchwork with a single federal compliance regime.

Jun 2026Industry Shift

Government-gated frontier AI - a new release paradigm

Claude Fable 5 (Anthropic) and GPT-5.6 Sol/Terra/Luna (OpenAI) both launched under US government review restrictions in June 2026, within weeks of each other. OpenAI previewed GPT-5.6 with the government before launch and said it is developing a 'repeatable process' for future releases. The most capable AI models are now subject to national security vetting before broad public access - a structural shift in how frontier AI reaches developers and enterprises.

Jun 2026Model Release

Claude Fable 5 - new capability ceiling, briefly pulled by US export order

Anthropic releases Claude Fable 5 - state-of-the-art across coding, science, vision, and knowledge work at $10/$50 per 1M tokens. Within three days, a US government export directive temporarily forced it offline; it was reinstated June 22. The first government-directed suspension of a major AI model established a new precedent for how frontier models reach the public.

Jun 2026Industry Shift

Apple WWDC - Gemini-powered Siri, Claude and Gemini as iPhone extensions

Apple announces Siri AI powered by a custom Google Gemini model (1.2T params), with multi-AI Extensions in iOS 27 - users route queries to Claude or Gemini directly from Siri settings. Consumer AI becomes platform-level infrastructure: Anthropic and Google gain access to the iPhone installed base without building their own native iOS integration.

Jun 2026Industry Shift

Microsoft Build - MAI models, first in-house frontier AI

Microsoft announces 7 MAI models at Build 2026. MAI-Thinking-1 (35B active params, 256K context) is Microsoft's first in-house reasoning model - trained without OpenAI distillation, preferred over Claude Sonnet 4.6 in blind evals. MAI-Code-1-Flash becomes the backbone of GitHub Copilot for all plans, giving Microsoft direct control over its most-used developer AI tool.

May 2026Tool / Platform

Devstral Small - agentic coding at $0.07/1M input

Mistral releases Devstral Small - an agentic coding specialist at $0.07/$0.28 per 1M tokens. Outperforms Codestral Small on SWE-bench agentic tasks, making economically viable coding agents accessible for high-volume loops.

May 2026Model Release

Qwen3 - thinking-mode toggle across open-weight models

Alibaba releases Qwen3 family (0.6Bโ€“235B MoE, Apache 2.0). All models support a thinking/non-thinking mode toggle per request. Qwen3 Coder 480B MoE targets software engineering. Strong multilingual support (29 languages). Extends Chinese AI labs' influence on open-weight ecosystem.

Apr 2026Model Release

GPT-5.5 and DeepSeek V4 Preview

GPT-5.5 (April 23) and DeepSeek V4-Pro (April 24, 1.6T total params, MIT license) arrived within 24 hours of each other - emblematic of 2026's pace. V4-Pro's permissive licensing continued the open-weight momentum DeepSeek started in January 2025.

Apr 2026Model Release

Claude 4 family: Haiku 4.5, Sonnet 4.6, Opus 4.7

Anthropic's fourth generation. Opus 4.7 launched April 16 - the most capable Claude to date. Sonnet 4.6 became the go-to production model: mid-tier cost, frontier-class reasoning for most tasks. The 4.x family raised the bar on agentic and multi-step tasks.

Apr 2026Tool / Platform

Graphify - codebase knowledge graphs go open-source

An MIT-licensed tool that turns any codebase into a queryable knowledge graph using local Tree-sitter parsing (no source code sent to external servers). 71.5ร— token reduction on real codebases. Crossed 22,000 GitHub stars in under ten days.

Apr 2026Tool / Platform

Claude Code Routines - scheduled AI automation

Anthropic added three trigger types to Claude Code: scheduled (cron-like), API-triggered, and GitHub event-triggered. AI-assisted automation moved from manual prompting to persistent background workflows.

Apr 2026Model Release

Grok 4.3 - xAI flagship with 1M context window

xAI releases Grok 4.3 at $1.25/$2.50/1M tokens with a 1M token context window - one of the largest of any commercial model. $150/month free developer credits available. Live X/Twitter search integration remains Grok's unique differentiator.

Apr 2026Model Release

Llama 4 Scout & Maverick - open-weight multimodal MoE

Meta releases Llama 4 Scout (109B MoE, 10M context) and Maverick (400B MoE, 17B active). Apache 2.0 licensed, natively multimodal, matching GPT-4o on major benchmarks at a fraction of inference cost. Scout fits a single H100. The open vs closed gap narrowed significantly.

Feb 2026Model Release

Gemini 3.1 Pro leads scientific reasoning

94.3% on GPQA Diamond - a PhD-level science benchmark. 77.1% on ARC-AGI-2. Google's Gemini family regained benchmark leadership in the first quarter, intensifying the rotation between labs at the top of the leaderboards.

2025

Aug 2025Model Release

GPT-5 integrated into ChatGPT

OpenAI's most capable model yet, with major improvements in multi-step reasoning, scientific problem-solving, and multi-modal understanding. Shipped directly into ChatGPT, making frontier capability accessible to all subscribers.

Jun 2025Tool / Platform

LazyGraphRAG - knowledge graphs become affordable

Microsoft's LazyGraphRAG reduced knowledge graph indexing cost from $20โ€“500 to under $5 by deferring community summaries to query time. Graph RAG moved from a research technique to a practical production tool.

Apr 2025Model Release

GPT-4.1 family - OpenAI's mid-range tier

GPT-4.1, Mini, and Nano gave developers a coherent OpenAI lineup matching the tiered structure Anthropic popularised. The cost gap between tiers widened: GPT-4.1 Nano at sub-$0.10/1M tokens vs GPT-4.1 at $2.00/1M.

Apr 2025Model Release

Meta Llama 4 - natively multimodal MoE

Llama 4 Scout and Maverick used Mixture-of-Experts architecture (activating only a fraction of parameters per token) to deliver competitive performance at dramatically lower inference cost. Maverick beat GPT-4o on multiple benchmarks. Behemoth (2T total params) signalled the raw scale Meta was willing to deploy.

Mar 2025Model Release

Gemini 2.5 Pro leads benchmarks

Google's Gemini 2.5 Pro with 'thinking budget' took top positions on coding and reasoning leaderboards. The 'thinking budget' feature let developers trade cost against reasoning depth - a new knob for production optimization.

Feb 2025Model Release

Claude 3.7 Sonnet - hybrid reasoning

Anthropic's first model with an explicit 'extended thinking' mode - letting users toggle the reasoning depth. Topped SWE-bench for autonomous software engineering tasks. The practitioner shift toward reasoning models accelerated.

Jan 2025Industry Shift

DeepSeek R1 - the 'DeepSeek shock'

A Chinese lab released an open-weight reasoning model trained for approximately $6M - matching OpenAI o1 on coding and maths. It became the #1 free app on the US iOS App Store within days. The narrative that frontier AI required billions in training compute collapsed overnight.

2024

Dec 2024Model Release

OpenAI o3, Gemini 2.0, DeepSeek V3

A landmark month: o3 achieved near-human scores on ARC-AGI; Gemini 2.0 Flash matched o1 performance at lower cost; DeepSeek V3 quietly shipped as a top-tier open-weight model. The frontier moved dramatically in 30 days.

Nov 2024Protocol / Standard

Anthropic open-sources MCP

The Model Context Protocol defined how AI assistants connect to external tools and data sources - files, databases, APIs, services. Within weeks, hundreds of MCP servers appeared. MCP became the de facto standard for AI tool integration in 2025.

Oct 2024Model Release

Llama 3.2 adds vision and mobile

Meta added visual understanding to Llama and released models small enough to run on smartphones. On-device AI became practical for the first time on consumer hardware.

Sep 2024Model Release

OpenAI o1 - reasoning models arrive

o1 allocated compute to 'thinking before answering' - generating an internal chain of reasoning invisible to users. It scored above PhD level on maths and science benchmarks. Test-time compute emerged as a new design dimension.

Jun 2024Model Release

Claude 3.5 Sonnet surpasses Opus

A mid-tier model beating the flagship at lower cost. This broke the assumption that 'best quality = most expensive.' Cost-performance optimisation became a primary engineering concern from this point forward.

May 2024Model Release

GPT-4o released

Faster, cheaper, and natively multimodal - audio, text, and images in a single model. Real-time voice mode with natural interruptions previewed a new tier of human-AI interaction. Speed and cost dropped 2ร— vs GPT-4 Turbo.

Apr 2024Model Release

Meta releases Llama 3

Llama 3 70B matched proprietary models on most benchmarks. The era of 'open-weight models can compete with closed models' began in earnest. Developers gained a credible alternative to API-only providers.

Mar 2024Model Release

Anthropic launches Claude 3 (Haiku, Sonnet, Opus)

The first coherent model family with clear capability/cost/speed tiers. Opus topped GPT-4 on multiple benchmarks. The tiered release became the template every lab copied in 2024โ€“2025.

Feb 2024Model Release

OpenAI debuts Sora

Sora produced photorealistic video from text prompts at a quality that shocked the industry. It signalled that generative AI's next frontier was temporal media, not just static images.