DeepSeek Open-Sources 671B DeepSeek-V3 and R1 with Multi-Head Latent Attention Architecture
DeepSeek releases full model weights for DeepSeek-V3 and R1, deploying a 671B Mixture-of-Experts network with 37B active parameters and FP8 DualPipe training.
Source: DeepSeek AI Research ↗ Published: 2026-08-29
DeepSeek has open-sourced the complete weights and technical reports for DeepSeek-V3 and its reasoning counterpart DeepSeek-R1, fundamentally altering open-weights infrastructure economics.
Engineering Innovations
- Multi-Head Latent Attention (MLA): Low-rank joint key-value compression that slashes KV cache memory footprint by 93% compared to conventional MHA.
- Fine-Grained MoE Routing: Activates only 37 billion parameters out of 671 billion across 256 expert modules with auxiliary-loss-free load balancing.
- DualPipe FP8 Mixed-Precision: Overlaps cross-node computation and communication, achieving 180 TFLOPS sustained throughput on training clusters.
- MIT Licensed Open Weights: Unrestricted commercial weights published on Hugging Face for local quantization and fine-tuning.