What the Query Plan Knew: Foreknowledge for LLM KV-Cache Management
Agent runtimes know what KV state will be needed, and when.
The KV cache has become a tiered memory hierarchy. The scarce resource that runs it well is no longer only memory capacity or bandwidth — it is timely information about future demand: foreknowledge, by which I mean workflow-derived knowledge of future KV demand that exists before the corresponding cache access occurs. This piece argues that agent runtimes are a particularly promising early source of it, and that its value reduces to three empirical questions: does it convert into physical reuse, does it arrive early enough to act on, and does acting on it pay?
In an earlier article, I argued that the KV cache is escaping the boundary of GPU memory: when requests share an exact token prefix, the key and value state computed for that prefix can be reused to skip prefill work, and state that valuable now lives in a real hierarchy spanning GPU HBM, CPU RAM, local NVMe, and shared network storage. This piece is about the decision that hierarchy has to make continuously, and the information it should make it with.
The decision is residency. For each reusable prefix, which tiers should currently hold a copy, and when should that change? Some modern tiered systems are write-through — NVIDIA’s KV block manager, for example, offloads blocks from GPU to host memory as they are produced — so copies can coexist at multiple levels and be evicted independently. Under memory pressure, the real choices are to preserve an expensive residency, rely on a copy in a lower tier, or discard the last copy entirely. For a correct cache implementation, discarding is safe: whatever was discarded can be recomputed by the next request that needs it, and nothing becomes incorrect. The stakes are economic, not semantic. (A separate line of research removes tokens from the context of a live sequence to shrink its footprint; that changes the model’s computation and is a different subject, which I set aside here.)
The canonical local baseline for the residency decision is recency. vLLM’s automatic prefix cache evicts reusable blocks in LRU order, and recency earned that default position honestly: it is low-overhead, and it makes its decision entirely from information local to the cache. The question of this essay is what better information the decision could run on — and what that information costs.
Two structural patterns recency cannot see
Agent workloads produce two situations in which recency is not merely suboptimal but blind, and the two situations carry different kinds of knowledge. Keeping them distinct matters for everything that follows.
The first is the tool-call pause. A model reasons over a conversation, emits a tool call, and stops. From the serving engine’s point of view, the request has ceased touching its KV state, and its recency begins to decay. A policy driven only by cache accesses cannot see why; the agent runtime can: an outstanding dependency. When the tool returns, another model call will likely resume the same computation with much of the same prefix. This is a prediction, and it should be treated as one — tools fail, branches get cancelled, routing may choose a different model, and the resumed prompt may not reproduce the exact token prefix that caching requires. But it is a prediction grounded in a fact the access stream does not contain: this particular session is waiting on this particular call.
The second is branch fan-out. A research agent instantiates five branches from a shared parent context. The cache has observed zero accesses from those branches, yet five consumers of the parent prefix already exist. That five branches were instantiated is a fact about workflow state, not a forecast about model behavior — though the KV demand it implies is still conditional, since branches can be cancelled before they ever issue a request. Evidence of this kind sits much closer to certainty than a statistical forecast does, and the difference between the two will do real work later in this essay. One honesty flag belongs here as well: fan-out implies physical reuse only when each child request preserves the parent’s token prefix exactly — append-only context construction is the usual way this happens, and frameworks that compact, summarize, or reorder context between steps break it. The final section returns to this.
Database systems met the same wall four decades ago, and the episode is worth one compressed paragraph because it names the pattern. LRU behaves badly on a looping scan over a relation slightly larger than memory: by the time the scan wraps around, LRU has retained exactly the pages whose next use is farthest away, and MRU — absurd in a general-purpose cache — can win in exactly this regime. The reason is that the execution layer knew it was running a loop. Chou and DeWitt’s DBMIN work built on that observation, exploiting the database’s knowledge of the access patterns its own operators generate — the knowledge embodied in the query plan — rather than asking a replacement policy to rediscover those patterns from page faults. Operating-systems research learned the same lesson independently with informed prefetching, where applications disclosed future file accesses to the storage layer. The recurring pattern deserves a name for the rest of this piece: foreknowledge — knowledge about future computation that exists in a layer above the cache before the corresponding access occurs.
The strong opponent: history
Before crediting foreknowledge with anything, it is worth building the strongest case that it is unnecessary, because the historical record contains that case too.
The database story has a second act. After DBMIN, increasingly capable history-only policies — LRU-K, which tracked multiple past accesses per page, and ARC, which balanced recency against frequency adaptively and resisted scans — recovered a substantial share of what semantic interfaces had promised, at a fraction of the integration cost. A cache that watches its own references carefully can infer a great deal about a workload without being told anything.
Modern serving data extends the same argument. A 2025 study of production traces at a large cloud provider, KVCache Cache in the Wild, found KV reuse both highly skewed — roughly 10% of cache blocks accounted for 77% of all reuses — and substantially predictable from history once requests were grouped into workload categories, and the authors built an effective workload-aware eviction policy on that finding alone. Note what this means for prediction timing: history is not confined to reacting after an access. A category-level model could learn, for example, that sessions of a given class tend to resume within 400 milliseconds — without knowing that any particular tool call exists.
So the bar for foreknowledge has to be stated carefully, and it is not “beats LRU.” The right question is: what facts does the runtime possess that are not recoverable from the access stream, and what predictions can it make earlier or more accurately because it holds those facts? The two patterns above are the concrete instances. The existence of an outstanding tool call for a specific session is a fact; the category-level resume-time estimate from the previous paragraph is the closest history can come to it, and it cannot say whether this session’s call is still outstanding, already returned, or failed. The identity and count of already-instantiated branches are facts that history discovers only after the accesses arrive — at which point the eviction decision may already have been made. And a point prediction tied to a specific dependency can arrive with enough lead time to schedule data movement, where a category distribution only describes the aggregate. Everything foreknowledge claims from here on must clear this bar.
The machinery is ready; the missing input is timely demand information
What do shipping systems actually do today? More than is commonly assumed, which makes the gap easier to state precisely.
The tiering machinery is production-real and partially proactive. A request arriving with a prefix cached in a slower tier does not stall the engine: vLLM’s connector interface lets the metadata lookup itself run asynchronously — the scheduler holds the request unadmitted while the answer is pending — and then holds it in a waiting state while KV streams into GPU memory, with the rest of the batch decoding throughout. A layer-wise loading interface allows computation of early layers to overlap the transfer of later ones. Demotion runs off the critical path: write-through offload moves completed blocks downward asynchronously as they are produced. None of this is exotic; it is the documented behavior of current vLLM, and the ecosystem’s shared tiers plug into exactly these interfaces.
The economics of the reactive path are also better than intuition suggests. For a Llama-3.1-70B-class model at BF16, KV state runs roughly 320 KB per token, so a 10,000-token cached prefix is about 3.2 GB. Recomputing that prefill on a single H100 costs on the order of a second; moving 3.2 GB at an effective 400 Gb/s takes roughly 64 milliseconds before contention, and proportionally less on faster links. Fetch-on-arrival beats recomputation by a wide margin for long prefixes, which is why the deep tiers are viable with no prediction at all. The trade should be stated honestly, though: a fetch exchanges prefill compute for data movement, time-to-first-token, and shared network and PCIe bandwidth, and a wrongly speculated fetch additionally occupies fast memory that something else needed. Nothing here is free; it is merely cheaper than the alternative when the prediction is right.
Here is the gap, stated at the precision the current landscape requires. The conventional serving path discovers future demand through request arrivals and cache events — a lookup fires because a request showed up, a block is retained because it was hit. What the path does not receive is workflow semantics supplied early enough to schedule placement around future execution. The mechanisms can execute an early decision; they are not given the early input.
Agent workloads are where that input exists, and they are also where acting on it pays, for a reason that has to be stated more carefully than “agents can wait.” Who absorbs fetch latency differs by workload. An interactive chat product is optimized aggressively for time-to-first-token, so added blocking latency spends a contested budget. In an agent workflow, the consumers of a model call’s output are a parser and the next model call, and the honest unit of analysis is the idle session — a computation awaiting a tool, or paused between turns — rather than agent traffic wholesale. An interactive coding agent with a person watching a forty-step chain compounds every per-step delay and is a poor deep-tier candidate; an idle session with a predicted resume time is a good one, under whatever service-level objective governs it.
What upgrades this from tolerance to opportunity is the distinction between a blocking fetch and a background prefetch. Without foreknowledge, the fetch begins when the resumed request arrives, and its latency lands on that request. With foreknowledge — the runtime knows the session is waiting on a tool call that typically takes 800 milliseconds — the transfer can run during the stall itself, and — assuming bandwidth is available and the prediction holds — a 100-millisecond fetch from remote NVMe is fully absorbed before the model call is even issued. Measured against the previous section’s bar: the tool-gap prefetch is a lead-time win, driven by a fact about a specific session that category history can only approximate; retention under fan-out is a known-consumers win, driven by facts the access stream has not yet seen.
The synthesis follows. Agentic traffic supplies both the slack — time in which data can move without extending the workflow’s critical path — and the signal: the demand information needed to schedule promotion back before the state is needed. Latency tolerance alone is not new; batch systems have always had it, and they built their own rich traditions of locality and prefetching around dependency structure. What distinguishes these workloads is that their idle intervals coincide with runtime-visible continuation structure: the system can know not only that a session is idle, but why it is idle and what computation is likely to follow. That coincidence is what makes agent workflows a particularly promising early user of foreknowledge in the KV hierarchy.
The interface is being designed right now
Research systems have been carrying workflow knowledge into serving decisions since 2024, and within the past few months the mainstream engines have begun drafting interfaces for it in public. Proposals can fail, but the shape of this interface is now a live design question in mainstream engines rather than a speculation — which moves the useful discussion from whether to what.
The research systems arrived first, each exploiting a specific kind of foreknowledge. Parrot established the general problem in 2024 — application-level dataflow among LLM calls is invisible behind request APIs — and InferCept quantified the serving cost of computations interrupted by external actions. The current wave is more specific: Continuum predicts tool-call duration and uses tool-aware time-to-live retention together with program-level scheduling; Tokencake proactively offloads KV during function-call stalls and uploads it predictively before the agent resumes; KVFlow derives each agent’s distance from execution in a workflow graph and uses it to drive both eviction and prefetch; PBKV extends the idea to dynamic workflows by predicting the next several invocations. The engines are absorbing the idea in real time, at varying maturity. A May 2026 SGLang RFC proposed and prototyped an agent-hints metadata path from the API layer into its radix-tree nodes and eviction policy, and its June 2026 Programmatic KV Cache RFC proposes a router-initiated control path with Retain, Share, Prefetch, and Demote hints — with the engine explicitly retaining authority to accept, clip, defer, or reject them. vLLM’s current roadmap lists agent hints such as session and correlation identifiers, with targeted prefix-cache prefetching, eviction, and selective offloading built on them. NVIDIA Dynamo already exposes agent hints such as priority, with priority-aware cache eviction supported on SGLang backends, and versioned experimental work has explored TTL-based prefix pinning. Independent teams reaching the same boundary from different directions is itself the strongest evidence that the boundary is real.
The contested question is what should cross it, and the proposals mix three categorically different things. Some fields are facts: five branches exist; this session is awaiting a tool; this future request has already been constructed. Some are predictions: the tool will likely return within 300 milliseconds; this prefix has an estimated 80% chance of reuse. And some are physical directives: pin this state for thirty seconds; keep this in GPU memory. My position is that the interface should carry facts and calibrated predictions and leave physical actions to the serving-side resource manager, which alone sees memory pressure, bandwidth contention, and competing work. The right analogy is a logical plan with cost estimates handed to an optimizer — the query plan was valuable because it exposed structure from which resource decisions could be derived, not because it contained pin instructions.
The strongest argument for that position is not aesthetic but organizational. For a vertically integrated deployment, the runtime-to-cache interface is internal plumbing and the distinction matters less. For a serving provider, the “runtime” is the customer’s orchestration code — software outside the trust boundary — and the three categories behave very differently there. Facts are attractive because they are verifiable after the fact: the provider can observe whether the announced branches actually issued requests, and score each source’s reliability over time. Predictions from untrusted code are a priority-gaming vector — every tenant’s tools always return in 100 milliseconds, every tenant’s prefix deserves to stay hot — and need exactly that per-source calibration to be worth anything. Directives from outside the boundary are bids for shared resources; they can be made honest by pricing them, and the engine-side proposals’ reservation of the right to reject them points in the same direction. Stated generally, the boundary that matters is not agent runtime versus serving provider but the layer that possesses workflow information versus the layer that makes residency decisions — sometimes two companies, sometimes two processes in the same stack. The production caching APIs fit this frame once read correctly: Anthropic’s prompt caching and Google’s context caching are retention controls, and their commercial existence demonstrates that applications value explicit reuse-and-lifetime controls across exactly this boundary. The open question is whether the boundary should carry the richer information — why and when the state will be needed, rather than only what should remain reusable — so that the provider’s optimizer can choose among HBM, RAM, NVMe, and recomputation with the customer’s structure in view.
What would settle it
The current systems report end-to-end speedups on their evaluated workloads, and those results establish feasibility. What they do not by themselves establish is which signals carry the value and whether a general interface is justified. Three measurements would, and they follow directly from the bar set earlier.
The first is conversion. When the runtime signals future reuse, how often — and weighted by how many tokens — does an identical, reachable prefix actually materialize? Ten correct predictions of tiny prefixes matter less than one correct prediction of a 20,000-token context. Logical dependence is not physical reuse: the future call must serialize the same information in the same order, against a compatible model, on hardware that can reach the cached state, and frameworks that compact, summarize, or reorder context break the chain silently. Conversion rates, separated by signal type, reveal where the apparent opportunity actually survives.
The second is lead time. The same correct prediction enables different actions at different distances: a few milliseconds of notice improves an eviction ranking; hundreds of milliseconds enable a remote prefetch that hides entirely inside a tool call. An instantiated branch, an outstanding tool call, and a multi-step workflow forecast are different evidence, and pooling them obscures exactly what a system designer needs to know. The useful output is a lead-time distribution per signal, measured against the duration of the action it would enable — a hint helps only if it arrives at least one transfer-time before demand — with the caveat that earlier is not monotonically better, since acting early reserves scarce state sooner.
The third is economic yield, because a signal can convert perfectly and still lose money. The accounting is GPU compute saved, minus transfer cost, added residency in fast tiers, and the pollution cost of false positives, under a fixed service-level objective. A signal that reliably protects tiny prefixes, or that displaces state more valuable than what it saves, fails this test while passing the other two.
The baselines matter as much as the measurements. The comparison that means something is not plain LRU but the strong history policy — reuse statistics grouped by workload category, in the style the production trace study demonstrated — together with an offline future-aware placement policy computed on the same trace under the same capacity and transfer constraints, which bounds how much improvement exists to be captured at all. If category-level history already sits near that bound on a workload, no interface can add much there, and knowing that cheaply is worth as much as any speedup.
The database lesson was never that MRU sometimes beats LRU. It was that the cache is not necessarily the layer with the best information about future accesses — and that noticing this was the easy part. The hard part was deciding which portions of the upper layer’s knowledge were worth turning into machinery, and the field answered that question with decades of measurement, keeping the semantic paths that paid and letting history-based policies take the rest. The same question is now open one layer up from a much larger memory hierarchy. The tiers are built. The interfaces are being drafted in public. The scarce input is timely knowledge of future demand, and the agent runtime is where it already lives. What remains is to measure whether that knowledge converts into real prefixes, arrives early enough to act on, and pays for its own transmission — which is to say, whether the query plan of the agent era deserves the influence its ancestor earned.
This article follows my earlier piece, “Why LLM Inference Is Disaggregating Its Memory”, written in May 2026. I have since left Aerospike. The views here are an independent technical exploration; no serving, storage, or accelerator system discussed here is being recommended.
References
- L. A. Bélády, A Study of Replacement Algorithms for a Virtual-Storage Computer, IBM Systems Journal, 1966.
- Hong-Tai Chou and David J. DeWitt, An Evaluation of Buffer Management Strategies for Relational Database Systems, VLDB 1985.
- Elizabeth J. O’Neil, Patrick E. O’Neil, and Gerhard Weikum, The LRU-K Page Replacement Algorithm for Database Disk Buffering, SIGMOD 1993.
- R. Hugo Patterson et al., Informed Prefetching and Caching, SOSP 1995.
- Nimrod Megiddo and Dharmendra Modha, ARC: A Self-Tuning, Low Overhead Replacement Cache, FAST 2003.
- Woosuk Kwon et al., Efficient Memory Management for Large Language Model Serving with PagedAttention, SOSP 2023.
- Jiahao Wang et al., KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider, USENIX ATC 2025.
- Chaofan Lin et al., Parrot: Efficient Serving of LLM-based Applications with Semantic Variable, OSDI 2024.
- Reyna Abhyankar et al., InferCept: Efficient Intercept Support for Augmented Large Language Model Inference, ICML 2024.
- Hanchen Li et al., Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live, 2025.
- Tokencake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications, 2025.
- Zaifeng Pan et al., KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows, 2025.
- Haoyu Zheng et al., Efficient Serving for Dynamic Agent Workflows with Prediction-based KV-Cache Management (PBKV), 2026.
- SGLang, Agent-Aware KV Cache Phase 1 for Agentic Workloads, RFC, 2026.
- SGLang, Programmatic KV Cache for Agentic Workloads, RFC, 2026.
- vLLM, Automatic Prefix Caching, KV Offloading, and Roadmap Q3 2026.
- NVIDIA, Dynamo KV Block Manager and SGLang for Agentic Workloads.
- Anthropic, Prompt Caching; Google, Gemini Context Caching.