Context engineering is memory management
The context window is not a text box. It is scarce, contended memory, and the old disciplines of managing scarce memory apply almost unchanged.
The phrase "context engineering" started as a rebrand of prompt writing, but the work it names is older than prompts. You have a fixed budget of fast, expensive memory. Multiple subsystems want a share of it. Whatever you load into it displaces something else, and the cost of a bad allocation shows up as degraded behavior somewhere downstream, usually not where you made the mistake. Systems programmers will recognize the shape. This is memory management.
Taking the analogy seriously turns most context decisions from taste into engineering.
What is actually competing for the window
In any serious agent or assistant, the window holds five tenants, and they fight:
- The system prompt. Instructions, policies, persona. Rent is paid on every single request.
- Tool definitions. Every schema you register is resident whether or not the model uses it. Twenty tools with verbose descriptions can quietly outweigh your instructions.
- Retrieved material. Documents, search results, database rows. The tenant with the most variable appetite.
- Conversation history. Only grows unless something evicts it.
- Working room. The model's output has to fit too, and long reasoning needs slack.
A budget nobody wrote down is still a budget. The only question is whether allocations happen by design or by whatever got appended last.
More window is not a pardon. Million-token contexts change the constant, not the discipline. Models attend unevenly across very long contexts, cost scales with what you load, and every irrelevant token is a small tax on attention that relevance has to pay back. Bigger RAM never made careless allocation a good idea either.
The prefix is a cache line
The budget decides what gets into the window; order decides what it costs. Inference providers price cached input tokens at a steep discount, and the mechanism underneath is KV cache reuse: if the opening bytes of your request exactly match a previous request, the saved attention state stands in for recomputing prefill. The word doing the work is "exactly." One changed character early in the prompt invalidates everything after it.
This makes prompt ordering a cache strategy rather than a style preference. Stable content goes first: system prompt, tool schemas, standing examples. Variable content goes last: the user's message, retrieval results, timestamps. Put a "current date" line at the top of your system prompt and you have invalidated your entire prefix once a day for nothing.
Cache-hostile layout:
[today's date and time] [user profile] [system prompt] [tools] [history]
Cache-friendly layout, same information:
[system prompt] [tools] [standing examples] [user profile] [history] [today's date and time]
Same tokens, same behavior, and the second layout can be dramatically cheaper and faster on every request after the first. Very few optimizations in this field are this free.
Eviction is the hard part, here too
History grows until something gives. The three honest strategies map cleanly onto old ideas.
Truncation drops the oldest turns. Cheap, predictable, and it forgets commitments made early in the conversation, which users experience as a broken promise.
Summarization compresses old turns into a digest. This is lossy compression, and the loss is invisible until the missing detail matters. A summary that says "user provided account details" has destroyed the account details.
Externalization moves information out of the window entirely, into files, a database, or a scratchpad, and retrieves it on demand. This is paging, and it inherits paging's failure mode: the page fault. If the agent does not know the information exists, it never asks for it back.
Production systems end up with all three, plus the judgment call about which information is load-bearing. What to evict is the whole game, and no general-purpose policy relieves you of knowing what your application actually needs to remember. Every token in the window should be able to say what it is doing there.
Retrieval is allocation, not decoration
The retrieve-versus-include decision is the same calculation as heap versus stack. Content that every request needs, and that fits, belongs resident in the prompt. Content that requests need rarely, or that dwarfs the budget, belongs external with a retrieval path. Teams get this wrong in both directions: shipping a fifty-page policy manual on every request because retrieval felt like work, or building a vector database to store twelve FAQ entries that would fit in the system prompt forty times over.
The sizing questions are boring and decisive: How big is the corpus? What fraction is relevant to a typical request? How often does it change? A corpus that fits in a fraction of the window and changes monthly does not need infrastructure. It needs to be pasted in.
Where this leaves you
Write the budget down. Literally: a table of tenants, their token allocations, and who is allowed to grow. Measure the real sizes; tool schemas in particular are always bigger than anyone guesses. Order the prompt for prefix stability, evict deliberately, and remember that every addition to the resident set pushes something else out.