vLLM 0.7.0 Ships Chunked Prefill, PagedAttention v3 & 3.2x Throughput Boost on H100s
High-throughput LLM serving engine vLLM introduces unified chunked prefill, FlashInfer integration, and specialized DeepSeek MLA CUDA kernels.
Source: vLLM Engineering Blog ↗ Published: 2026-08-30
The open-source vLLM project has released version 0.7.0, introducing massive optimizations for multi-tenant inference, continuous batching, and reasoning models.
Performance Benchmarks
- Chunked Prefill by Default: Interleaves token prefill chunks with decode steps to eliminate Time-to-First-Token (TTFT) latency spikes.
- DeepSeek MLA Triton Kernels: Custom GPU kernels accelerating Multi-Head Latent Attention decoding by up to 3.2x on NVIDIA H100/H200 SXM5 nodes.
- FlashInfer & FP8 FlashAttention-3: Native support for Ada Lovelace and Hopper FP8 tensor cores with zero quantization perplexity loss.
- Automated Speculative Decoding: Built-in Eagle and draft-model speculative decoding boosting generation throughput by up to 2.4x.