AI Foundations, Part 6: Agents Are Delegation, Not Just Automation
An agent doesn't just respond, it perceives, decides, and acts in a loop until a goal is actually done. That loop is why delegating to one feels completely different from running a script.

Here’s a word that’s been rattling around in my head lately: agent. I kept running into it, using it myself even, without being totally sure I could explain what actually makes something an agent versus just a clever script. So I went looking for the real answer, and once it clicked, it changed how I think about almost everything I build.
The short version: an agent decides what to do next. A script has that already decided for it, by you, before it ever ran. That one distinction turns out to explain a lot.
What automation actually is
Start with the boring half, because you need it as a contrast. Automation, a shell script, a CI pipeline, an n8n workflow, a Zapier chain, runs a sequence of steps that a human author wrote down in advance. Trigger fires, step one runs, step two runs, step three runs, done. The order is fixed. The steps are fixed. Given the same input, you get the same sequence every single time, because the sequence isn’t being decided at runtime, it was decided once, by a person, at design time.
There’s nothing wrong with that. Fixed sequences are fast, cheap, and completely predictable, which is exactly what you want for a nightly backup job or a deploy pipeline. You don’t want your CI system “reasoning” about whether to run your test suite. You want it to just run the tests, the same way, every time.
The limitation shows up the moment the task doesn’t fit a sequence you can fully specify in advance, when the right next step depends on what you find out along the way, and there are too many branches to hand-code all of them. That’s the gap agents are built to fill.
What actually makes something an agent
Anthropic’s engineering team drew this line clean in a piece I keep coming back to, Building Effective Agents: a workflow is a system where the LLM and tools are orchestrated through code paths a human defined in advance. An agent is a system where the model dynamically directs its own process and tool use, staying in control of how it accomplishes a task. Same underlying model, same tools available, the difference is who’s deciding the path, and when that decision gets made.
Concretely, in the context of a large language model, an agent is the model paired with three things:
- A set of tools it can call — functions that let it act on the world (search a database, write a file, hit an API) instead of just producing text
- A goal or task to work toward — stated in plain language, not a fixed procedure
- The ability to run multiple steps in a loop — observing what each action returned, and deciding the next action based on that, until the task is actually finished
The verb is what separates this from ordinary question-and-answer. An agent acts. It doesn’t respond once and stop, it keeps going, checking its own progress, until the goal is met or it hits a stopping condition you’ve set.
A workflow is a path someone already walked and wrote down. An agent is being told where to end up and trusted to find the way.
Inside the loop, mechanically
If you read the tools post one entry back, you already know these mechanics: the model gets the conversation so far, requests a tool call instead of writing prose, your harness runs it and hands the result back, and the model reasons again in light of what it just learned. Same cycle, no changes underneath. What’s different about an agent isn’t the machinery, it’s what tells that machinery to stop.
In the tools post, the loop had a natural finish line: answer one bounded question, maybe chase down one follow-up fact along the way, then write the answer, done. An agent’s loop doesn’t get that luxury. It’s chasing a goal, not a single question, so there’s no moment where “enough tool calls happened” is obviously true. The model has to decide for itself, turn by turn, whether the goal is actually satisfied yet or whether it needs to try something else, until it either produces a final answer or an iteration, time, or cost limit built by whoever’s running the loop cuts it off.
That’s what makes “the model decides” more than a slogan. Every one of those turns is a fresh inference call that has just seen the result of the last one and gets to change its mind, not a human-authored “step 2” fixed in code ahead of time.
This “reason, act, observe, repeat” pattern traces back to research like ReAct: Synergizing Reasoning and Acting in Language Models (Yao et al., ICLR 2023), which showed that interleaving explicit reasoning steps with actions, rather than either pure reasoning or pure acting, measurably improved task success on benchmarks like multi-hop question answering and interactive decision-making environments. The model doing a visible “here’s what I’m thinking, here’s what I’ll try” before each action turned out to make it better at recovering from a bad guess mid-goal, which is exactly the situation an open-ended loop keeps putting it in.
A website chat widget is the easiest place to feel this distinction firsthand, because you’ve almost certainly used one. You type a question, it answers, maybe by quietly pulling your order status from a real system behind the scenes, and then it stops and waits for you to type the next thing. You’re still the one driving: you decide what the next question is, it just answers each one well. That’s the bounded-question loop from the tools post, running once per message you happen to send. An agent takes you out of the middle of that. Point it at “get my refund processed and confirm it hit my card” instead of asking it one question at a time, and it decides on its own which questions it needs answered along the way, works through them, and only comes back to you once the goal itself is actually done, not once per message.

Why delegate to a subagent at all
Once you accept that an agent can run a loop on its own, the next fun question is why you’d ever want more than one running at a time. Good agentic tooling, Claude Code is the one I use daily, lets you spawn subagents from inside a session: separate instances handling a specific subtask in their own isolated context, reporting back only what matters to the parent. A few reasons I reach for one instead of doing everything in a single long conversation:
- A research task would flood your main working session with noise you don’t need to see, fifteen search results and three dead ends, when you only care about the one paragraph of conclusion
- You want to hand off a long-running task and keep working on something else while it runs, instead of sitting there watching it
- You need something more restricted, a reviewer with read-only access, for instance, so a “check this code” task literally cannot also modify the code it’s reviewing
- Keeping a subtask’s mess out of your main session’s context is a real budget saving, not just a tidiness preference. Every token that subagent burns chasing a wrong turn never touches your main conversation
A simple subagent definition, this is close to the actual format Claude Code reads, and what you’re really doing here is writing a narrower job description, with its own name, its own model, and its own restricted toolset:
name: code-reviewer
description: Reviews code changes for security, quality, and standards compliance
model: claude-opus-4
tools: [Read, Glob, Grep]
system: |
You are a senior code reviewer focused on safety-critical systems.
Check for: security issues, safety issues (race conditions, unchecked
errors), standards compliance, and adequate test coverage.
Cite specific file and line number for every finding.
Notice what’s not in that tool list: nothing that writes or executes. That’s deliberate, the reviewer’s whole value is that it can’t fix what it finds, only report it. Restriction is a feature here, not a limitation you’re working around.
Three shapes agent systems tend to take
Once you’re delegating, patterns emerge fast. Three come up over and over, and none of them are exotic, they’re just org-chart thinking, applied to a system that can actually staff the chart itself.
Sequential — each agent’s output becomes the next agent’s input, a relay race:
Analyzer → Architect → Developer → Reviewer → Deployer
Good for work that has a genuine pipeline shape, where each stage needs the previous stage’s finished output before it can start, not just a rough draft of it.
Parallel — several agents work the same problem from different angles at once, then a final step reconciles what they found:
┌── Specialist A
Requirements ── ├── Specialist B → Integrator
└── Specialist C
Good when the sub-questions are genuinely independent, three different research angles on the same topic, say, and reconciling disagreement at the end is cheaper than forcing one agent to hold all three angles in its head simultaneously.
Hierarchical — an orchestrator breaks the goal into subtasks and delegates each to a specialist, some of which delegate further:
Orchestrator
├── Research Agent
├── Design Agent
└── Implementation Agent
├── Backend Agent
└── Frontend Agent
Good for genuinely large, multi-part goals, where the orchestrator’s real job isn’t doing the work, it’s deciding how to break the work up and knowing enough about each piece to judge whether what came back is actually done.
None of these shapes are exotic. They’re org-chart thinking, applied to a system that can actually staff the chart itself.
Where the loop can bite you
I don’t want to only sell you the upside here, because the honest picture includes the failure modes too, and they’re worth knowing before you hand something real to a loop. A system that decides its own next step can also decide wrong, repeatedly, with total confidence. Left unchecked, an agent can:
- Loop past the point of usefulness, retrying a failing approach with small variations instead of recognizing it’s stuck
- Call a destructive tool it technically had permission to call, because nothing in its goal statement told it not to
- Burn real time and real API cost chasing a subtask that was never actually necessary for the goal
This is why the tool list in that subagent definition above matters as much as the system prompt does, and why serious agentic frameworks build in hard stops: maximum iteration counts, cost ceilings, required human approval before anything irreversible. Delegating a goal isn’t the same as delegating unlimited trust. You still decide what the agent is allowed to do, you’re just no longer deciding exactly how it does it.
Where this actually pays off
What gets me excited about this isn’t novelty for its own sake. It’s what happens once you stop imagining agents as a future feature and start feeding them goals you’d normally have gritted your teeth and done yourself.
Getting real-world errands actually done. “Get quotes from three contractors for the driveway repair, compare them, and book whichever one can start soonest within budget.” That’s not one lookup, it’s reaching out to each one, following up when somebody doesn’t reply, comparing what actually came back, and checking it against your own calendar before committing. Handing that whole goal to an agent and getting back “booked for Thursday morning, here’s why I picked that one over the other two” is a genuinely different experience than doing the coordinating yourself.
Running my own home lab. This is the one I live with daily. I’ve got a whole cluster of servers running my Brain vault and a dozen other services, and instead of me noticing something’s gone sideways, my own agent, I call her Gemma, watches for it, reads the logs, forms a real hypothesis about what broke, restarts the right thing, and tells me afterward what actually happened, not “did you try turning it off and on again.” That’s not a script cycling a fixed command on a schedule. That’s a system deciding what’s wrong before it decides what to do about it.
Working through a genuinely hard decision. “My car’s getting old, should I replace it with something like an EV or just get another one like it?” isn’t a single search, it’s reliability data, real total cost of ownership, incentives, and how the tradeoffs actually play out for the way you drive, all weighed against each other rather than looked up one at a time. An agent that does that comparison for real and hands you one grounded recommendation, not ten open tabs, is worth more than the raw search results ever were.
Trying to hand-hold every one of those steps yourself defeats the entire point of having a system that can reason about what it finds along the way. Once you’ve felt the difference between doing it all in one long thread and handing off a real goal to something that works it end to end and reports back only what matters, going back to the old way feels like working with one hand tied behind your back.
Where to go if you want the real thing
- Building Effective Agents — Anthropic’s own field notes on the workflow-vs-agent distinction and the patterns that actually hold up in production
- ReAct: Synergizing Reasoning and Acting in Language Models — the research (Yao et al., ICLR 2023) behind the reason-act-observe loop most agent frameworks still run today
- AI Agents vs. Automation: Understand the Difference & Choose the Right Solution — a plain-language walkthrough if you want the business-side framing rather than the technical one