Agentic AI · August 15, 2026

AI Foundations, Part 1: What Are We Actually Talking About?

Cutting through the AI/ML/deep learning/LLM stack — and the handful of terms that actually matter once you start using this stuff for real work.


A cute modern cartoon robot working at a laptop

Everyone throws around “AI” like it means one thing. It doesn’t, and the fuzziness is exactly why so many smart, capable professionals feel like they can’t get a foothold — every article assumes you already know what a “model” is, what “training” means, or why a “context window” matters.

It’s a stack, not a single thing.

Four nested layers

Start here, because it clears up more confusion than anything else in this series:

Artificial Intelligence (broadest)
  └── Machine Learning (AI that learns from data)
        └── Deep Learning (ML using neural networks)
              └── Large Language Models (the thing you actually talk to)

Artificial Intelligence is any technique that gets a machine to mimic human behavior — a broad umbrella going back to the 1950s.

Machine Learning is AI that learns from examples instead of explicit rules. Rather than programming “if X then Y,” you show the system thousands of X→Y pairs and it learns the pattern itself.

Deep Learning is machine learning using layered neural networks, loosely inspired by the brain. Each layer extracts a more abstract feature from whatever came before it.

Large Language Models are deep learning models trained on enormous amounts of text — so much that they learn the statistical shape of language well enough to generate coherent answers, hold a conversation, reason through a problem, and write working code. This is the layer that actually shows up on your screen.

The handful of terms worth actually knowing

You don’t need the math. You need these, because they explain behavior you’ll otherwise find mysterious or frustrating:

Training — not memorization, prediction-and-correction at enormous scale. Researchers assemble a massive corpus of real text — licensed books, public websites, articles, code repositories, whatever they can legally gather at scale — and feed it through the model piece by piece. For each piece, the model tries to predict the next token, checks that guess against what the source text actually says comes next, and nudges its billions of parameters slightly toward whatever would have made the correct token more probable — and every alternative slightly less probable. Do that trillions of times across trillions of tokens of that corpus, and the parameters gradually settle into a shape that encodes the statistical structure of language itself: which words, phrases, and patterns tend to follow which others. Nothing is stored as a lookup table — a fact or phrase is “learned” in the sense that the token sequence expressing it becomes highly probable to produce, not in the sense of a saved record somewhere you could point to. This happens once, ahead of time, over weeks to months, on thousands of GPUs running in parallel.

Inference — using those now-fixed parameters to actually generate something. At each step, the model outputs a full probability distribution over every token it knows — not one answer, a ranked likelihood across all of them — picks one (see Temperature, below, for how), appends it to the sequence, and repeats for the next token. This is what happens every time you send a message: the “response” is really thousands of these next-token probability calculations happening in sequence, fast enough to look instantaneous.

Parameters — the actual object being trained. A large language model is built from a specific neural network architecture called a transformer (the “T” in GPT) — stacked layers of interconnected mathematical functions that take a token’s numerical representation and transform it, layer by layer, into a prediction. Every connection between those layers carries a weight: a single number controlling how much influence one value has on the next. Parameters are all of those weights taken together — billions or trillions of individual numbers, each one nudged slightly during training toward whatever made correct predictions more likely. There’s no single parameter that means “dog” or “Tuesday” — meaning is distributed across millions of these numbers acting in combination, which is a real part of why even the researchers who build these models can’t fully explain what any one part is doing. Roughly, more parameters means more capacity to represent complex patterns — though not the whole story; the training data and process matter just as much.

Token — the basic chunk of text a model processes, roughly 3-4 characters or about three-quarters of a word. Every model has a maximum number of tokens it can hold in view at once — that’s the context window.

Temperature — a dial on randomness, and it’s worth being precise about what it actually turns: not the parameters. Those are fixed the moment training ends — temperature can’t touch what the model “knows.” What it changes is the very last step of inference, the moment the model has to actually pick a word. By that point the model has already computed a full probability for every possible next token — 70% chance of “the,” 12% chance of “a,” 3% chance of “an,” and so on across its whole vocabulary. Temperature reshapes that distribution right before a token gets sampled from it. Low temperature sharpens it — the already-most-likely word gets pushed even further ahead, so the model reliably grabs the same safe, predictable choice almost every time. High temperature flattens it — the lower-probability words become real contenders, so the model is willing to gamble on something less obvious. That’s the real trade: low temperature gives repeatable, “boring” but reliable output, which is exactly what you want for code or data extraction; high temperature gives more varied, surprising output, good for brainstorming, at the cost of occasionally picking a word that’s plausible but worse purely because the dial let it take the chance. Technical work wants it low. Brainstorming wants it higher.

Hallucination — the model generating something plausible-sounding but false, stated with exactly the same confidence as something true. It’s a real, documented limitation, not a rare edge case, and it’s caused real damage. In 2023, two New York lawyers filed a legal brief written with ChatGPT’s help that cited six court cases — complete with quotes and docket numbers — that simply didn’t exist; the model had invented them wholesale, and the lawyers were sanctioned for it (Mata v. Avianca). That same year, Google’s own promotional demo for its Bard chatbot confidently stated the James Webb Space Telescope took the first-ever image of a planet outside our solar system — false, that actually happened in 2004 via a ground-based telescope — and the error wiped roughly $100 billion off Alphabet’s market value in a single day. Neither model was being deceptive. Both were doing exactly what they’re built to do: producing the most statistically likely next words, which is not the same operation as knowing what’s true. Verify factual claims against a real source before you act on them — especially the ones stated confidently.

The model isn’t lying to you when it hallucinates — it’s doing exactly what it was built to do, predicting the most statistically likely next words, and sometimes the most likely words are wrong.

Why this matters before you touch a tool

Every confusing moment I’ve hit using this stuff for real engineering and writing work has traced back to one of these terms — running into the context window limit, getting a hallucinated fact I didn’t catch, seeing wildly different output because temperature wasn’t what I expected. None of it is mysterious once you have the vocabulary. It’s just new machinery, and like any new machinery, the fear goes away once you know what the dials actually do.

That’s the whole point of this series — not the hype, not the doom, just the actual mechanism, one piece at a time.