Stack Hive HQ
Sponsored Partner
advertisement
Deploying DeepSeek R1 671B on Kubernetes: vLLM FP8 Cluster Architecture Blueprint
Infrastructure & DevOps

Deploying DeepSeek R1 671B on Kubernetes: vLLM FP8 Cluster Architecture Blueprint

Architectural blueprint for running self-hosted DeepSeek R1 across NVIDIA H100 nodes: tensor parallelism, KV cache memory calculations, and production Kubernetes manifests.

Maya Lin
Maya Lin
Lead MLOps Engineer
Published: 2026-08-31 • 11 min read

DeepSeek R1 671B is the first open-weights reasoning model to match proprietary closed models on mathematical reasoning and code synthesis. However, running a 671B parameter Mixture of Experts model requires careful GPU memory budgeting and optimized tensor parallelism across high-speed NVLink fabrics.

1. VRAM Mathematics & FP8 Quantization

At BF16 precision, 671B parameters require ~1.34 TB of pure weight storage. By utilizing native FP8 block-wise quantization, the model weights consume ~680 GB. On a standard 8x NVIDIA H100 SXM5 node (8 x 80 GB = 640 GB), weights plus KV cache exceed a single node. Consequently, production deployments utilize either a dual-node 16x H100 setup or 8x H200 (141 GB HBM3e) nodes.

# vLLM DeepSeek R1 Deployment Args\n- "--model=deepseek-ai/DeepSeek-R1"\n- "--tensor-parallel-size=8"\n- "--max-model-len=32768"\n- "--enable-chunked-prefill=true"\n- "--trust-remote-code"\n- "--dtype=fp8"\n- "--kv-cache-dtype=fp8"

2. Multi-Head Latent Attention (MLA) Decoding

Unlike standard multi-head attention where KV cache explodes linearly with context length, DeepSeek R1 employs MLA. KV cache is projected into a low-dimensional latent space, reducing memory consumption to just 512 bytes per token. In vLLM 0.7.0, specialized Triton kernels bypass decompression during decode, allowing up to 128 concurrent active streams per cluster.

INTERACTIVE SAAS CALCULATOR

LLM Token & Prompt Caching Cost Estimator

Monthly API Invocations50,000 requests
Avg. Input Tokens per Request1,500 tokens
Avg. Output Tokens per Response500 tokens
Standard API Cost:$600.00 / mo
Cost with Prompt Caching:$458.25 / mo
Estimated Monthly Savings
$141.75
Save ~24%

Frequently Asked Questions

Can DeepSeek R1 671B run on a single 8x H100 (80GB) node?

In FP8 precision, weight storage alone requires ~680 GB, leaving insufficient space for KV cache on 640 GB nodes. It requires either 8x H200 (141GB) or a 2-node 16x H100 cluster.

Why is chunked prefill essential for reasoning models?

Reasoning models produce long generation traces. Chunked prefill prevents long incoming prompts from starving active token generation batches, keeping TTFT consistent.

Related Tutorials