Prefix caching
What this demonstrates: reusing pre-computed KV cache across requests with shared prompt prefixes, including a second-tier CPU pool shared across instances.
For workloads where many requests share a system prompt, RAG context, or a long instruction (e.g., agent traces), recomputing prefill for each request is wasted work. Prefix caching keeps prefix KV blocks around (in NPU memory by default; optionally also in CPU or CXL) and reuses them on hits.
LLMServingSim ships block-hash prefix caching (ported from vLLM v0.19.0's block pool) with three flavors:
- Per-instance NPU pool (default, always on).
- Cross-instance shared CPU pool: second-tier prefix cache, which
models attaching LMCache or vLLM's
OffloadingConnector. - CXL-backed pool: same as above but in CXL memory.
Prerequisites
- Simulator container set up
- A workload with shared prefixes (the bundled
example_trace.jsonlhas some, real ShareGPT or agentic traces have a lot)
Cluster config
The simplest setup uses
configs/cluster/single_node_multi_instance.json (two instances on
one node, no special memory config). The shared CPU pool is enabled
at runtime via CLI flags, not the config:
{
"num_nodes": 1,
"link_bw": 16,
"link_latency": 20000,
"nodes": [
{
"num_instances": 2,
"cpu_mem": {"mem_size": 512, "mem_bw": 256, "mem_latency": 0},
"instances": [
{
"model_name": "meta-llama/Llama-3.1-8B",
"hardware": "RTXPRO6000",
"npu_mem": {"mem_size": 96, "mem_bw": 1597, "mem_latency": 0},
"pd_type": null,
"tp_size": 1
},
{ "...": "second instance, identical" }
]
}
]
}
The cpu_mem.mem_size (512 GB here) caps how big the CPU prefix pool
can grow.
Run
Per-instance prefix caching (default)
python -m serving \
--cluster-config 'configs/cluster/single_node_multi_instance.json' \
--dtype bfloat16 --block-size 16 \
--dataset 'workloads/example_trace.jsonl' \
--output 'outputs/prefix_default_run.csv'
--enable-prefix-caching is on by default. Prefix blocks are kept
in each instance's own NPU memory; if a request lands on instance A
that prefixes-into a cached block on instance B, no reuse happens.
Shared CPU prefix pool
python -m serving \
--cluster-config 'configs/cluster/single_node_multi_instance.json' \
--dtype bfloat16 --block-size 16 \
--enable-prefix-caching --enable-prefix-sharing --prefix-storage CPU \
--dataset 'workloads/example_trace.jsonl' \
--output 'outputs/prefix_cpu_pool_run.csv'
The two extra flags:
--enable-prefix-sharing: turn on the second-tier pool.--prefix-storage CPU: pool lives incpu_mem. Other options:CXL(requires acxl_memconfig block),None(NPU-only).
When an NPU prefix is evicted, it spills to the CPU pool instead of disappearing. Requests on any instance can now hit the CPU pool on lookup.
Expected output
With the shared CPU pool enabled, the throughput log gains prefix hit-rate counters:
[20.0s] Avg prompt throughput: 2412.0 tokens/s, Avg generation throughput: 820.0 tokens/s
├─Running Instance[0]: 6 reqs, Waiting: 0 reqs, Total # 1 NPUs, Each NPU Memory Usage 63218.51 MB (64.301 % Used), Prefix Cache Hit ratio 41.82 %, (31024 / 74208)
├─Running Instance[1]: 4 reqs, Waiting: 0 reqs, Total # 1 NPUs, Each NPU Memory Usage 63204.51 MB (64.287 % Used), Prefix Cache Hit ratio 42.11 %, (30298 / 71950)
└─Node[0]: Total CPU Memory Usage 8192.00 MB, 1.600 % Used, Prefix Cache Hit ratio 36.04 %, (52664 / 146158)
The per-instance ratios are NPU-tier hits; the ratio on the Node
line is the shared CPU pool's, and it replaces the per-instance CPU
split because --enable-prefix-sharing --prefix-storage CPU makes the
pool node-wide.
The prompt_t (prompt throughput) counts all input tokens,
including those served from cache, matching vLLM's reporting
convention.
What's interesting
- NPU memory pressure stays bounded even on workloads with huge shared prefixes. The CPU pool absorbs eviction.
- Cross-instance reuse is the killer feature for multi-replica deployments. Without prefix sharing, a 90% prefix-overlap workload effectively sees prefix caching as 1/N as effective with N instances.
- CXL pool is an option when CPU memory is the bottleneck. Set
--prefix-storage CXLand add acxl_memblock to the cluster config (see CXL memory tier). The pool then lives in CXL memory at CXL latency. - Block-aware tracking. The simulator's
prompt_taccumulator includes prefix-cache-hit tokens, so its reported prompt throughput matches vLLM's (which also counts cached tokens).
Related examples
- CXL memory tier: backing the prefix pool with CXL memory.
- Multi-instance LOAD routing - the multi-instance baseline this builds on.