AI & HPC Data Centers
Fault Tolerant Solutions
Integrated Memory

This benchmark whitepaper explores how expanding effective KV cache capacity with a CXL-backed external memory tier can reduce latency, increase throughput, and improve GPU efficiency.

Modern AI inference systems increasingly encounter a memory wall before they reach a compute limit. As context windows grow and concurrency increases, KV cache demand can exceed available GPU memory, creating scheduler queues, increasing latency, and limiting throughput.
This whitepaper evaluates how a CXL-backed external KV memory tier on the Penguin Solutions MemoryAI™ KV Cache Server can expand effective KV cache capacity under sustained inference load. Using a representative enterprise RAG workload, the benchmark shows how preserving and restoring KV cache data can reduce latency, increase throughput, and improve GPU efficiency.
Readers will learn:
Compared to a baseline configuration using GPU-resident KV cache only, the report measured:
Download the whitepaper to see the methodology, benchmark results, and practical guidance for evaluating external KV memory in production AI environments.

This benchmark whitepaper explores how expanding effective KV cache capacity with a CXL-backed external memory tier can reduce latency, increase throughput, and improve GPU efficiency.