The agent loop in production
The loop itself is trivial. Everything that makes an agent good or bad in production lives in the parts the diagram hides.
Strip away the framework branding and every agent is the same control loop: send the conversation to the model, and the model either answers, which ends the loop, or requests a tool. Your code runs the tool, appends the result to the conversation, and goes around again.
your task
|
v
+-------------+
| the model | <---------------+
+-------------+ |
| | |
answers asks for a tool |
| | |
v v |
done your code runs it |
| |
+--- result joins ---+
the conversation
You can write this in forty lines and it will work in a demo. Production is different, and the difference is never the loop. It is four things the diagram hides: what the tools look like, when the loop stops, what goes back into the conversation when a tool fails, and where state lives.
Tool schemas are the real API surface
The model chooses actions by reading your tool definitions, not your documentation or your intentions. A tool's name, description, and parameters are a prompt, and they get the same quality of attention as any other prompt, which is to say: the model does what the schema makes easy, not what you meant.
A schema the model will misuse:
search(query: string)
"Searches."
The same tool, specified like an interface:
search_orders(
customer_email: string,
status?: "open" | "shipped" | "refunded",
placed_after?: ISO8601 date
)
"Search this customer's orders. Returns at most 20, newest
first. Use placed_after to narrow instead of paging."
The second version does three jobs the first ignores. It scopes the tool so the model cannot fantasize about what it searches. It constrains parameters with types and enums, so bad calls fail at the schema instead of at your backend. And it states the return contract, so the model can plan its next step without guessing. Most of what people call agent unreliability is interface ambiguity, and the fix looks like API design because it is API design.
Stop conditions are a budget
The loop's only natural exit is "the model answered without asking for a tool." That is one exit for the happy path and zero for everything else. Production loops need explicit budgets: a maximum number of iterations, a wall-clock ceiling, a token spend cap, or all three. Without them, a confused model will happily alternate between two tools forever, and the failure you observe is a bill.
The subtler version is the loop that terminates but should not have. The model declares success after a tool error it silently absorbed. Which brings up the ugliest gap in the diagram.
Tool errors are model input, not exceptions
When a tool fails, the result still goes back into the conversation. The model reads it and decides what to do next. That makes your error messages prompts. A tool that returns an empty string on failure, or a stack trace, or an HTTP 500 with no body, is telling the model nothing usable, and the model will fill the vacuum with confidence.
The production pattern is to catch everything at the tool boundary and return errors written for the model: what failed, why, and what a reasonable retry would look like. "Rate limited, retry after 30 seconds" gives the loop somewhere to go. A silent empty result teaches the model that the database contains nothing, and it will report that to your user as a fact.
Log the transcript. The conversation is the only complete record of what the agent believed at each step. When a run goes wrong, the transcript is your debugger. Persist it, all of it, including tool results, or you will be reconstructing failures from guesswork.
Where state lives decides what your agent is
A loop whose whole state is the conversation is stateless in the way that matters: kill it, replay the transcript, and you are back where you were. The moment you bolt on external memory, scratch files, a database the model writes to, a plan object it updates, you have distributed state, and the classic questions arrive on schedule. What happens when the tool succeeded but the process died before the result was appended? Can two runs touch the same memory? Which side is authoritative when the transcript and the store disagree?
None of these questions are novel and none of the answers are AI-shaped. Idempotent tools, transactional writes, and a single source of truth are the same medicine they have always been. The trap is believing the agent is special enough to skip them.
The loop is the easy part; the engineering is in the arrows. If you are starting today, my honest suggestion is to write the loop yourself once, without a framework, against real tools. Not because frameworks are bad, but because every production failure you eventually debug will live in one of these four places, and you want to have seen them plainly, in code you wrote, before a framework wraps them in abstraction.