bhuvanesh.co

Ideas

Short and unpolished. Working notes that might grow into full pieces. Each one has a stable link if you want to point at it.

TTFT is a product decision

Time to first token lives on infrastructure dashboards, but it is set mostly by choices product teams make: how long the system prompt is, how many tools are registered, how much retrieval gets stuffed in front of the user's question. The prompt decides perceived speed before the first token renders, and everything after that is throughput. If the app feels sluggish, the fix is more often a shorter prompt than a faster model. TTFT belongs in product reviews, next to the funnel metrics.

Tool schemas are API design

Most of what gets filed as agent unreliability is interface ambiguity. The model chooses actions by reading tool names, descriptions, and parameter types, and it does what the schema makes easy rather than what the author meant. Vague name, vague behavior. Untyped parameter, invented values. The fixes are the ordinary ones: scoped names, enums instead of free strings, documented return contracts, error messages written for the caller. We have known how to design interfaces for decades. The new caller just reads more literally than the old ones.

Stable prefixes are free money

Prompt caching discounts only apply when the opening bytes of a request exactly match a previous one, because the mechanism underneath is KV cache reuse. Which means the ordering of a prompt is a cache strategy. Stable content first (system prompt, tool schemas, standing examples), variable content last (user message, retrieval, timestamps). A date string near the top of a system prompt silently invalidates the whole prefix every day. Reordering costs nothing, changes no behavior, and is probably the cheapest inference optimization most teams have not made.

Agent memory is a filing problem

Most agent memory features fail at retrieval, not storage. Writing things down is easy; the hard question is whether the agent can find the right note at the right moment, which is a problem of naming, structure, and retrieval conventions rather than embedding quality. A plain directory of well-named files with an index often beats a vector store here, because the agent can list, scan, and reason about what exists instead of hoping a similarity search surfaces it. Before reaching for infrastructure, ask whether the real gap is a filing convention.

The context window is not a database

A pattern that keeps appearing: teams treat the context window as the system of record, accumulating state in the transcript until it overflows, then acting surprised when summarization loses the one detail that mattered. The window is working memory. It is small, expensive, and lossy under compression. Durable state belongs outside it, in something transactional, with the transcript holding pointers rather than payloads. The test is simple: if losing the conversation would lose data, the architecture is wrong.