AI Foundations, Part 3: Choosing the Right Model for the Job
Most providers give you a flagship, a balanced workhorse, and a fast cheap option. Knowing which one to reach for matters more than people think.

Ask ten people which AI model to use and eight of them will name whichever one is loudest in the news that month. That’s not a strategy. It’s brand recall.
The actual answer depends on the job.
Not one model — a family
Every serious AI provider ships a small family of models, not one, and each tier trades capability for speed and cost. Anthropic’s current lineup is a clean example of the pattern:
- Claude Opus 5 — the flagship. Most capable, most expensive, slowest. Reach for it on genuinely hard problems: multi-step reasoning, ambiguous requirements, anything where getting it right matters more than getting it fast.
- Claude Sonnet 5 — the balanced, everyday model. Strong capability at a fraction of the cost and latency, and the one doing most of the actual work in a well-run setup. Most tasks belong here.
- Claude Haiku 4.5 — small and fast. Cheap and quick, built for high-volume or simple operations: classifying a support ticket, extracting a field from a form, answering something that doesn’t need deep reasoning.
OpenAI’s GPT line and Google’s Gemini line follow the same three-tier shape under different names. This isn’t an Anthropic quirk — it’s how the whole industry prices intelligence now.
Rule of thumb: flagship for hard problems where quality matters most, the balanced tier for most tasks, the fast tier for high-frequency simple work. Routing every request to the flagship model is the single most common way people overspend on AI for no quality benefit. Most of what you ask it to do doesn’t need that much horsepower.
Where the model actually runs
There’s a second axis to this decision, orthogonal to which tier you pick: does the model run on the provider’s servers, or on your own hardware?

A frontier model — Claude Opus 5, or its equivalent from any provider — runs on the provider’s infrastructure and gets called over an API. It’s the largest, most capable tier available at any given moment, billed per token, always the current version. The provider upgrades it out from under you, for better or worse, and your request leaves your machine and travels to someone else’s servers.
A locally-hosted model is an open-weight model you download and run yourself, on your own GPU. Meta’s Llama, Mistral, and Alibaba’s Qwen are the families people actually run at home, usually through a tool like Ollama that handles the model-serving plumbing for you. The economics flip: instead of a per-token bill, you’re paying a fixed hardware cost up front, and every request stays on your own machine. Nothing leaves, it works with no internet connection at all, and you decide if and when to update it. The trade-off is capability — the largest models you can realistically run on consumer or prosumer hardware still trail the frontier tier on genuinely hard problems.
In practice, local models are a strong substitute for the fast, cheap tier — high-volume, low-stakes, routine work — not a replacement for the flagship tier on anything that actually needs to be right. Same tier philosophy as before, just a different decision about where it physically runs.
What these models are actually good at
Across the current generation of frontier models, a consistent set of core capabilities has emerged:
- Code generation and review — writing, explaining, and debugging code across virtually any language
- Document analysis — reading PDFs, spreadsheets, and transcripts, then extracting and synthesizing what matters
- Long-context reasoning — holding an entire codebase or a stack of long documents in view at once, rather than working from fragments
- Tool use — calling functions, running code, searching the web, managing files, when it’s been given the ability to do so
- Agentic behavior — planning and executing a multi-step task with minimal hand-holding, which is the whole subject of a later post in this series
The second dial: effort
Picking a model tier answers one question: how capable does this need to be. There’s a completely separate question sitting right next to it: how hard should the model work on this specific request. That’s effort, and it’s easy to conflate with model choice because both show up as a setting on the same API call.

Anthropic’s API exposes it directly as output_config.effort, with five levels: low, medium, high, xhigh, and max.
response = client.messages.create(
model="claude-opus-5",
max_tokens=16000,
output_config={"effort": "high"},
messages=[{"role": "user", "content": "..."}]
)
Here’s what effort actually does, mechanically. Before writing its final answer, a model can generate a private scratch-work chain of intermediate reasoning — checking its own steps, trying an approach, backing out of it, trying another — and only then commit to the response you actually see. Turning effort up gives it more room to do that. Turning it down skips most of it, and the model answers off something closer to first-pass intuition.
That extra reasoning isn’t free. It’s billed like any other tokens and it adds latency, so effort is a real lever, exactly like model tier is a lever, except it controls depth of thought on one request rather than raw capability. Researchers demonstrated just how directly controllable this is by literally forcing a model to keep reasoning past where it wanted to stop, simply by appending “Wait” to its output — and watched it catch and fix its own mistakes as a direct result of the extra thinking time, with no change to the model itself.
The two dials are genuinely independent. A quick, low-effort call to the flagship model is often faster and just as good as a maxed-out-effort call to a smaller one. A genuinely hard problem run at low effort will disappoint no matter how capable the model is on paper — it never got the room to think it through.
Rule of thumb: raise effort when the task is hard enough to be worth the wait. Keep it low for anything routine or latency-sensitive, regardless of which tier you’re running it on.
Talking to the model programmatically
Every provider exposes an API for building the model into your own applications. The shape is nearly universal regardless of vendor: you send a list of messages, optionally a system instruction setting the model’s role, and get a response back. This example uses Anthropic’s Python SDK — OpenAI and Google’s SDKs follow the same request/response shape with different field names.
response = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
system="You are a precise technical assistant. Cite sources where relevant.",
messages=[
{"role": "user", "content": "Explain what a P-trap is in plumbing."}
]
)
print(response.content[0].text)
Multi-turn conversation works by simply appending to that message list as it goes — the model has no memory of its own, so the full running history is what gets sent, every time:
messages = [
{"role": "user", "content": "What is a P-trap?"},
{"role": "assistant", "content": "A P-trap is a curved pipe section..."},
{"role": "user", "content": "Why is it required?"},
]
response = client.messages.create(model="claude-sonnet-5", max_tokens=1024, messages=messages)
The API doesn’t remember your last conversation. You do — and you prove it by sending the transcript back every single time.
The brain isn’t the whole system
Pick the right tier, dial in the right effort, and you’ve still only described the brain — raw pattern-recognition and reasoning capacity, fixed the moment training ended. A brand new call to that model, with nothing else attached, is a bit like a newborn’s brain: fully wired, capable of startling things almost immediately, and yet it has no idea what your project is, what you talked about an hour ago, or where anything actually lives. Not because it’s incapable — because it hasn’t had the chance to have any of that happen to it yet.
One caveat to that analogy: it’s not a blank slate in the usual sense of the phrase. The wiring is all there, fully formed, the same way the weights are fully trained before you ever send a message. What’s missing isn’t capacity. It’s lived experience.
That’s what the framework wrapped around the model actually supplies — the coding agent, the assistant, whatever application is making the call. It’s the one deciding what goes into the model’s context window before the request is ever sent: project files, prior conversation, notes, the results of whatever it just looked up. The model itself doesn’t remember any of it between calls. The framework does the remembering, and hands the right pieces back in, every single time.
The model doesn’t remember your last conversation. The framework around it does — and it proves it by handing the transcript back in, every single time.
Same shape as the code example a few paragraphs up, one level higher on the stack: there, it was you assembling the message list by hand. Out in the world, it’s the framework doing that assembly automatically. Either way, nothing gets remembered that wasn’t deliberately handed back in. How a real system decides what to hand back in — what belongs in context and what gets left out — gets a full post of its own, later in this series.
Where to go if you want the real thing
- “s1: Simple test-time scaling” (Muennighoff et al., 2025) — the paper behind the “Wait” example above: a small, cheap technique for directly controlling how long a model reasons before answering, and the measurable accuracy gain from letting it think longer.
- Inference-Time Scaling: How Modern AI Models Think Longer to Perform Better — a plain-language walkthrough of why spending more compute at answer-time, not just at training-time, turned into one of the biggest recent gains in model quality.
Stop picking by brand
Picking a model isn’t a brand question. It’s a triage question: how hard is the task, how much does getting it exactly right matter, and how many times a day will you run it. Answer those three honestly and the right tier is usually obvious.
Send everything to the flagship and you’re paying flagship prices for fast-tier work. Send hard problems to the fast tier and you’ll get answers that are wrong with total confidence. Neither mistake is subtle once you’re looking for it.