Agentic AI · August 17, 2026

AI Foundations, Part 2: How These Models Actually Work

The model isn't understanding you — it's predicting the next word with startling sophistication. Here's the mechanism, the vocabulary, and the research it's built on.


A small glossy white robot with glowing orange eyes turning a hand-crank wheel on a mechanical machine

Here’s a sentence that sounds like it should be disqualifying: the model doesn’t understand your question. It predicts the most probable next word, then the next, then the next.

Say that to most people and they assume it means the output must be shallow. It isn’t. Now that’s worth thinking about.

The actual sequence, end to end

Strip away the mystique and a single response is generated like this:

  1. Your prompt gets broken into tokens — chunks of text, roughly word- or syllable-sized.
  2. Tokens become numerical vectors — a process called embedding.
  3. Attention layers process how those tokens relate to each other, across many stacked layers.
  4. The model outputs a probability distribution over every possible next token.
  5. One token is sampled from that distribution.
  6. That token gets appended to the sequence, and the whole process repeats.
  7. It stops when a designated “stop” signal is generated.

That’s it. Every essay, every function, every explanation you’ve ever gotten from one of these models was built one predicted token at a time, in a loop.

Diagram: prompt tokenized, embedded, passed through transformer layers, sampled into one token, then looped back in

It feels like thinking. Technically it’s sequential prediction. What’s strange — and worth pondering — is that at sufficient scale, those two things stop being clearly distinguishable in the output.

Two words worth actually defining

Most explanations of this technology throw around “Transformer” and “attention” as if they’re self-evident. They’re not jargon for jargon’s sake — they’re the two specific ideas that made modern AI possible, and each has a real, precise meaning.

The Transformer is a neural network architecture — a blueprint for how the layers of the model are arranged and how information flows through them. It was introduced in a single, now-famous 2017 research paper from a team at Google: “Attention Is All You Need” (Vaswani et al.). Before this paper, the dominant approach to language processing read text sequentially, one word at a time, in order — which made it slow to train and bad at connecting ideas that were far apart in a passage. The Transformer’s core proposal was to drop that sequential requirement entirely and let the model look at the whole input at once, weighing every part against every other part in parallel. That single architectural change is the reason today’s models can be trained on the scale they are, and it’s the common ancestor of essentially every major language model built since.

Attention is the specific mechanism inside that architecture that does the weighing. For every token, the model calculates how relevant every other token is to it, and uses that to build a representation that blends in the right context. It’s what lets a model connect a requirement stated on line one to code written five hundred lines later — the two don’t have to be near each other for the model to notice they’re related. Distance in the text is irrelevant to attention; relevance is what matters.

Self-attention, specifically, is the model applying that mechanism within a single passage — learning which words and phrases matter to each other in what you actually wrote. Feed it “the reactor tripped because the sensor failed,” and it learns that reactor, tripped, sensor, and failed are all bound together — not because someone told it that relationship exists, but because it inferred it from the statistical pattern of billions of similar sentences during training.

Diagram: seven words in a sentence with curved lines of varying thickness connecting related words like "reactor" to "tripped" and "sensor" to "failed", regardless of distance between them

After an attention step, each token’s representation runs through an additional set of layers that apply further learned transformations — this combination of “attend, then transform” is one Transformer block, and a real model stacks dozens of them on top of each other. Stack enough of these blocks, at enough scale, and what emerges is something that reads, unmistakably, as reasoning.

Attention doesn’t care how far apart two ideas are. It only cares whether they’re related.

Where to go if you want the real thing

None of this needs to stay abstract — the actual paper and some excellent plain-language breakdowns of it are freely available:

  • “Attention Is All You Need” — the original 2017 paper (Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, Polosukhin). Dense, but the abstract and introduction are readable without a machine learning background.
  • The Illustrated Transformer, by Jay Alammar — the most-cited plain-language, diagram-heavy walkthrough of exactly how attention and the Transformer block work, step by step.
  • The Illustrated GPT-2, by the same author — covers the autoregressive loop in the “actual sequence” section above in much more visual detail, including how the probability distribution and sampling step actually work.

Why context is the whole game

There’s no persistent memory between separate conversations, unless something is deliberately built to provide it. Everything the model knows about your specific situation, in a given session, comes from what’s sitting in its context window right then — the block of text (your prompt, any files, the running conversation history) it can actually see when it generates the next token.

That single fact explains three things people often get wrong:

  • Vague prompts get vague answers. A well-structured prompt with real, specific context outperforms a vague question by a wide margin — not because the model is being difficult, but because it genuinely has nothing else to go on.
  • Persistent project-context files matter enormously. Loading essential background into the prompt automatically, at the start of every session, is what separates an assistant that “gets it” from one that has to be re-briefed every time.
  • A structured notes system — a “second brain” — is powerful precisely because it puts the right knowledge in front of the model exactly when it’s needed. More on that later in this series.

Once you internalize that the model is only as good as what’s currently in front of it, a lot of “why did it get that wrong” moments stop being mysterious. It didn’t get it wrong. It didn’t have it.