What Is AI Inference? Context Engineering Explained

What is AI inference? The using stage, where the model acts on its context window. The four context misses and the context engineering fixes.

· By Guilherme Salgueiro

In Memento, Leonard Shelby cannot form new memories, so every scene begins with him waking up somewhere he does not recognise and piecing together the present from what he carries: a handful of Polaroids, a few notes in his pocket, and the facts he has tattooed across his body so he cannot lose them. He reads them, decides what they mean, and acts with complete confidence, because from where he stands that is all the truth there is.

What makes the film work is that Leonard is not incapable. He can still drive, read people, and reason his way through a problem, because everything he learned before the injury is intact. What he lost is the ability to carry the last few minutes forward, which means the quality of every decision he makes depends entirely on the quality of what he wrote down.

That is a near-perfect description of AI inference. The model learned everything it knows during training, before you ever met it, and every turn since then is a Leonard scene: it wakes up, reads whatever is in front of it, and acts as if that were the whole world. When the Polaroid is right, the result looks like intelligence. When the Polaroid is stale or missing, the model acts just as confidently and simply goes somewhere wrong.

If you work with AI, you probably already sense that context matters. What usually goes unexplained is how to set that context up so the model acts on the right Polaroid, and that gap is what this piece is about.

What is AI inference?

Inference is the stage where a trained model uses what it learned to produce an answer, a prediction, or an image from new input, as opposed to training, which is the stage where it learned in the first place. Google Cloud compresses the idea into a single line:

AI inference is the "doing" part of artificial intelligence.

That definition is accurate, but it stops exactly where things get interesting for anyone who uses AI rather than hosts it, because the rest of that page moves on to deployment: cloud versus edge, serving infrastructure, and Google Cloud's own products. For everyone else, the useful question is not how inference runs but what it runs on.

Seen that way, you already spend most of your AI time inside the using stage. Every chat turn, every agent edit, and every generated image is an act of inference, which is also why ChatGPT is best understood as one inference product rather than the concept itself.

It also changes what briefing an agent actually means. You are not teaching the model anything when you brief it, since its learning finished long ago; you are handing Leonard a fresh set of Polaroids. And what the model sees in that moment is not "the project" as it exists in your head, but only what actually landed in its window this turn:

OpenAI's API docs describe this plainly, noting that each request stands on its own and that continuity only exists because the earlier messages are sent back in:

While each text generation request is independent and stateless, you can still implement multi-turn conversations by providing additional messages as parameters to your text generation request.

The consequence is simple but easy to forget: if something is not in this turn's input, it is not in the model's head. The constraint you meant to add does not count, and the current spec does not count either if last month's version is the one sitting in the window.

That is why two common misreadings of inference both lead people astray. The first treats it as magic, because the chat interface hides the using stage so well that the answer seems to come from nowhere, and it becomes easy to forget that the model was working from a specific set of inputs you gave it. The second treats it as a serving problem of GPUs, latency, and context-window size, which matters to the people running the infrastructure but rarely explains why your agent did the wrong thing. In practice, the using stage fails far more often because the context was outdated, missing, wrong, unread, or gone by the next session, and a bigger window does not fix that, since a bigger pocket simply holds more stale Polaroids.

What does good context look like at inference?

Once you accept that the model only acts on what is in its window, the next question is practical: what should be in that window, and how would you know whether the model actually used it? The answer gets much easier once you notice that "context" is not one thing. It arrives in three layers, each with a different lifespan, and each one breaks in its own way.

Leonard's system in Memento happens to map onto those layers almost exactly:

That last failure is not just a matter of tidiness. Anthropic's engineering team points to research on what they call context rot, which finds that "as the number of tokens in the context window increases, the model's ability to accurately recall information from that context decreases."

That is why their guidance for every layer points in the same direction: give the model less but better, aiming for "the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome." In practice that often means handing the agent a pointer rather than a payload. Instead of pasting a whole document into the window, you give it the file path and let it load what it needs at the moment it needs it, an approach Anthropic calls "just in time" context. A pocket with the three Polaroids that matter is easier to act on than a pocket with three hundred, in the same way that a compass is useful because it points in one clear direction rather than every possible one.

Setting up the layers well is only half the job, though, because the window can hold exactly the right information and the agent can still act on something else. So it helps to have a simple test for whether context was actually used: for any action that matters, can you trace it back to a specific line of context that justified it? If the agent can name the file, the rule, or the version of the spec it relied on, you can check that against what is current. If it cannot, the context was decoration, however carefully you prepared it.

Good context at inference is layered, lean, and traceable: standing rules that stay current, task material that is the right version, live results that do not pile up, and actions you can trace back to the line that justified them.

The four misses that waste inference

Knowing the three layers is one thing. Watching them fail is more instructive, and they rarely fail all at once. They fail one small miss at a time, in scenes that look harmless on their own, until you are convinced the model is the problem. So picture a single job, something simple and real like updating the pricing page for your product, and watch it play out in four scenes.

Scene one: the Polaroid

You open a session, ask the agent to refresh the pricing page, and twenty minutes later it hands you a clean, well-structured page with clear copy and tidy components. It also has the wrong prices on it. The agent is not guessing: it found the numbers in a rules file you wrote in the spring, back when the plans were different, and it applied them with the calm certainty of someone who has done their homework.

That is Leonard and the Polaroid of Teddy with "Don't believe his lies" scrawled across it. He trusts the note completely, without knowing when he wrote it or what he knew at the time, and the whole of Memento turns on that one unexamined piece of paper. Your agent did exactly the same thing with your standing context, and it had no way to tell you the Polaroid was out of date, because from inside the window an old fact looks identical to a current one.

The miss: you don't know what context you gave them. The tell: a decision that would have been perfect a month ago.

Scene two: the speech

You fix the prices by hand and move on. Then you open a fresh session, and because the agent remembers nothing, you explain the product again: the three plans, the annual discount, the tone you want, the components it should not touch. By the third session the speech is slightly shorter, and by the fourth you are giving it with the weary rhythm of a flight attendant doing the safety demonstration, and you have started forgetting bits, so that session's agent never hears about the annual discount at all.

In 50 First Dates, Henry faces the same problem every morning with Lucy, who wakes up without any memory of the past year. His first instinct is to win her over again from scratch each day, which is charming for about a week and exhausting after that. What actually works is a videotape she watches over breakfast, recorded once and kept up to date, that tells her what happened and who he is. The briefing stops depending on how well he performs it that morning.

The miss: you keep repeating yourself, which means your standing context lives in your head instead of a file. The tell: you could recite your opening briefing from memory.

Scene three: the radio room

Halfway through the job, you do the responsible thing and update the pricing document with the new plans. You even point the agent at it. And it still ships a change based on the old numbers, citing the previous version of the file with complete confidence, as if the new one were never there.

On the night the Titanic sank, other ships had sent ice warnings, and at least one, from the Mesaba, was acknowledged in the radio room and never taken to the bridge while the operators worked through a backlog of passenger telegrams. The warning was on board. It just never reached the decision. LangChain's guide to LLM evaluation makes the same point about agents: plenty of failures do not trace back to the model at all, but to stale instructions, a missing example, or context that no longer matches how the team works. If that feels personal, it may help to know that Andrej Karpathy has admitted that his agents do not listen to the instructions in his AGENTS.md files either.

The miss: the context is there and they still don't act on it. The tell: the right file is in the window, but the agent cannot tell you which line it followed.

Scene four: the elevator

Somewhere in the middle of all this, one session worked out that the checkout test fails randomly and is safe to rerun. It was a genuinely useful discovery. A few sessions later, a new one hits the same test, panics, and spends forty minutes investigating it from scratch, arriving at the same conclusion with the same quiet pride.

In Severance, the employees at Lumon do real work all day on the severed floor, then step into the elevator and walk out knowing none of it. Whatever they figured out stays behind, and the next morning their work selves start over with the same blank expression. A session without write-back is that elevator: the lesson stays on the floor and the next session starts from zero.

The miss: nothing compounds. The tell: the agent rediscovers something it already figured out.

The twist: it was never the model

By the end of the job, the natural verdict is that the model is not good enough, and the natural next step is to try a different one. But look back at what actually happened. The model wrote good copy, structured the page well, reasoned correctly from what it was given, and diagnosed a flaky test on its own. Every failure came from what was in its pocket when it woke up: an old note, a briefing that changed every session, a warning it never acted on, and a lesson nobody wrote down.

A new model would have played the same four scenes. The good news is that every one of those four misses has a fix, and none of them requires a better model.

How do you make agents act on context? Shoot the director's cut

Same agent, same model, same pricing page. This time, each scene gets one small change to what goes into the agent's pocket, and the story ends very differently.

Rewrite the Polaroid: look before you leap

Before asking for anything, you spend two minutes checking what the agent can actually see. Claude Code's /context command lists what is occupying the window, and Cursor's docs explain that it can automatically pull in your open files, terminal output, and linter errors, so it is worth knowing what came along for the ride. The spring rules file is right there, prices and all, so you delete the stale lines and date the ones that remain, which means next time anyone can tell at a glance how old a fact is.

This habit matters more than it sounds, because loading is less permanent than it looks. Rules scoped to certain folders only load once a matching file is read, and compaction summarises them away with the rest of the conversation, so "we have a rule for that" is a claim worth checking rather than assuming. Shopify's Tobi Lütke set a useful bar for this step: give the model enough context that the task is plausibly solvable. If you cannot say what the agent has, you cannot say whether that bar is met.

This time the agent ships the right prices on the first try.

Record the tape: write the speech once

Instead of giving the speech again, you write it down once, in the file every session already loads. CLAUDE.md, AGENTS.md, and Cursor rules all serve as Henry's videotape, and Claude Code's memory docs now read AGENTS.md alongside CLAUDE.md, so one well-kept file can brief several tools at once. If you are unsure what to put in it, GitHub's analysis of more than 2,500 agents.md files found that the useful ones cover six areas: commands, testing, project structure, code style, git workflow, and boundaries.

Keep the tape short. Claude Code's docs suggest staying under 200 lines per file, because a longer briefing takes more of the window and gets followed less reliably, which is the context-rot problem from earlier showing up in your own instructions. Your test from now on is simple: if you explain something twice, it belongs in the file.

Every session now hears about the annual discount, because nobody has to remember to mention it.

Get the warning to the bridge: make it show its work

Updating the pricing document was the right move. What was missing was any way to tell whether the agent used it, so this time you ask it to name the file and the section it relied on before it changes anything. That one request turns the traceability test from earlier into a habit, and it makes a stale citation obvious the moment it appears, because you can see the old filename in the answer. OpenAI's agent-improvement cookbook takes the same idea further, requiring every material claim to cite a source file and running a script that checks those citations against real files.

For the few things that must never happen, such as publishing prices without a review, instructions are not enough. Claude Code's memory docs are candid that CLAUDE.md is "context, not enforced configuration", and they point to hooks, small scripts that run before a tool is used and can block an action outright. Instructions ask; hooks enforce. Deeds, not words: the point of all this is not that the agent read the right file, but that the right file changed what it did.

This time the agent quotes the new document, and you can see that it did.

Stop the elevator: write it down before the session ends

The moment the agent discovers that the checkout test is flaky, you ask it to add one line to a short progress file before the session ends. Dex Horthy of HumanLayer calls the habit behind this frequent intentional compaction, and his YC Root Access talk makes the case for deliberately writing progress down rather than letting it vanish with the session. Anthropic's context-management tools build the same loop into the platform: context editing clears stale tool results as the window fills, and a memory tool keeps what should survive in files outside the window.

The next session reads the note, reruns the test, and moves on in thirty seconds instead of forty minutes.

The new ending

None of these changes involved a new model, a bigger window, or a new tool subscription. They took a few minutes each, and every one of them has a safety net, because a context file in version control can always be rolled back if an edit makes things worse. The model in the first scene was always capable of this ending. It just needed the right things in its pocket when it woke up.

Why a live source of truth beats a longer memory

Fix one job and you will want to fix them all, which usually leads to the obvious idea: keep everything, so the agent never forgets anything again. It sounds sensible right up until the pocket is full of Polaroids again.

Harry Potter has a better idea. The Marauder's Map is brilliant for one reason: it shows where everyone in Hogwarts is right now. It is not a diary of where they wandered last term, and if it were, Harry would have been caught in the first chapter. Your project's source of truth should work the same way, as a shared working library that tells every session what is true today, not a log of everything anyone ever said.

!Append-only log versus live source of truth: keep project sources current so every session sees what exists now

There is one catch, and Leonard is the cautionary tale. His real problem in Memento is not that he forgets. It is that he writes notes to his future self without checking them, and one bad note is enough to send him after the wrong man. Agents that write to memory can make the same mistake, saving a stale result that every later session then trusts. Redis calls this context poisoning, and the fix is pleasantly boring: let agents update the map, but only after checking that what they are about to save is still true. Date your entries, keep the files in version control, and a bad note becomes easy to spot and easy to undo.

Where to start: pick the depth you will keep

You do not need all of this at once. Pick the step you will actually keep doing:

You do not need a new model for any of it. That is AI-native leverage at the moment of inference: the same model, with a better pocket.

Leonard never gets his memory back, and your agents will not either. What you control is what is in their pocket when they wake up.

Frequently asked questions

What is context engineering?

Context engineering is deliberately shaping what sits in a model's context window at inference time (instructions, files, examples, and state) so the model acts on the right information.

What is inference in AI?

Inference is the using stage: a trained model producing an answer, a prediction, or an image from new input. Training is the learning stage. ChatGPT is one inference product, not the concept. The practical job at inference is actioning context: setting it up, then making sure agents act on it.

Why do AI agents ignore context?

Because presence is not use. Context can sit in a file, a rule, or the window and still not govern the next action. The four common misses: you don't know what you actually gave the agent, you re-brief it every session, it cites a stale document instead of the current one, and nothing compounds because there is no live source of truth.

How do I set up context for AI agents?

Inventory what is actually in the window. Keep one project-scoped store that answers what exists now, and delete or date anything stale. Put ways of working in files every session already loads. Check transcripts for whether the agent acted on the current file. Let agents update the store only after a freshness check.