All Things AI
Beginner

Claude Model Family

Anthropic's Claude models are organised into four tiers - Haiku (fast and cheap), Sonnet (the balanced workhorse), Opus (near-frontier capability at a lower cost than the flagship), and Fable (the flagship, added September 2026 above Opus). Each generation raises the ceiling on all tiers. Understanding the tier system and current model IDs is the first step to using Claude effectively.

A Note on Parameters

Anthropic does not publicly disclose parameter counts for any Claude model. No official figure has ever been published for Haiku, Sonnet, or Opus. Estimates that circulate online come from third-party inference-speed analysis and should not be treated as authoritative. The dimensions that matter for practical use are context window, speed, cost, and benchmark performance - all of which Anthropic does publish.

The Naming Convention

Claude model names follow a consistent pattern: family tier + version number. The tier names are inspired by Japanese poetry forms - reflecting their positioning from concise-and-fast to expansive-and-powerful:

  • Haiku - Fastest, cheapest. Designed for high-throughput tasks where latency and cost per token are the priority: classification, routing, extraction, simple Q&A.
  • Sonnet - Balanced. The default choice for most production workloads: coding, analysis, writing, agentic tasks. Best value in the family.
  • Opus - Near-flagship capability at meaningfully lower cost than Fable. Used for demanding tasks that don't need the absolute ceiling: complex research, multi-step agentic reasoning, extended thinking.
  • Fable - The flagship, added September 2026. Most capable Claude yet, with always-on adaptive thinking. Reserve for the hardest tasks where Fable's extra cost over Opus is clearly justified.

Version numbers (3, 3.5, 3.7, 4, 4.5, 5, 5.1, 5.5) indicate the generation. A higher version is almost always strictly better than a lower version within the same tier. The API model ID (e.g. claude-sonnet-5) is what you use in code; the display name is what appears in Claude.ai.

Current Models (as of Sep 2026)

ModelAPI Model IDContextBest For
Claude Haiku 4.5claude-haiku-4-5-20251001200KClassification, routing, extraction, high-volume pipelines
Claude Sonnet 5claude-sonnet-5200KCoding, analysis, writing, agents - the default production choice, most agentic Sonnet yet
Claude Opus 5.5claude-opus-5-51MNear-Fable quality at ~40% lower cost; complex research, hard agentic tasks
Claude Fable 5.1claude-fable-5-11MFlagship - hardest coding and knowledge-work tasks, always-on adaptive thinking

Previous Generation (still available)

ModelContextNotes
Claude Opus 4.7 / Sonnet 4.6200KPrior generation flagship pair; superseded by Fable 5.1 / Opus 5.5 / Sonnet 5
Claude 3.7 Sonnet200K (128K with extended thinking)First hybrid reasoning model; extended thinking mode for hard problems
Claude 3.5 Sonnet200KStrong coding + instruction following; widely adopted before 3.7
Claude 3 Opus / Haiku200KOriginal 3-series; long since superseded, still referenced for historical context

Context Windows

Sonnet 5 and Haiku 4.5 carry a 200,000-token context window; Opus 5.5 and Fable 5.1 extend that to 1 million tokens - among the largest in the industry. In practical terms, 200K is already:

  • ~150,000 words of text (~500 pages)
  • Entire medium-sized codebases
  • Hours of meeting transcripts
  • Multiple long documents simultaneously

Claude is particularly strong at accurately using its full context - it doesn't lose track of information near the middle or start of a long document (a known failure mode for some competing models, sometimes called the "lost in the middle" problem). This makes it a leading choice for long-document analysis tasks.

Extended Thinking

Starting with Claude 3.7 Sonnet and now always-on in Fable 5.1, Claude supports extended thinking mode - where the model spends additional tokens on internal reasoning before producing its final response. This is Claude's equivalent of the adaptive thinking built into OpenAI's GPT-6 Astra.

  • Activated via the API by setting a thinking parameter with a token budget
  • The internal thinking is returned as a separate thinking block in the response
  • Most useful for: complex maths, formal logic, multi-step agentic tasks, hard coding problems
  • Adds cost (thinking tokens are billed) and latency - use selectively, not by default

Multimodal Input

All Claude models accept text and images as input. Claude can read and reason about PDFs, screenshots, charts, diagrams, and mixed text+image documents. File upload limits:

  • Claude.ai Free/Plus: up to 50MB per file
  • Claude.ai Pro/Enterprise: up to 1GB per upload session
  • API: image input via base64 or URL; file analysis via the Files API

Claude does not generate images or audio - it is text-in, text-out (with image/document input support).

Pricing Tier Logic

Anthropic pricing follows the tier hierarchy: Haiku is cheapest, Sonnet is mid, Opus is near-flagship, Fable is the flagship. For the current generation (check anthropic.com/pricing for live rates):

Haiku 4.5

$1 / $5 per 1M tokens. Best for pipelines running millions of calls per day.

Sonnet 5

$2 / $10 per 1M tokens. The best balance of cost and capability for most production workloads.

Opus 5.5

$4 / $20 per 1M tokens. ~40% cheaper than the prior Opus generation for near-flagship quality.

Fable 5.1

$10 / $50 per 1M tokens. Highest cost per token - reserve for tasks where output quality directly drives business value.

Input tokens are cheaper than output tokens across all models. Extended thinking tokens are billed at the input token rate but consumed before the response - factor this into cost estimates for Opus-heavy agentic workflows.