DeepSeek V4.1 Flash

DeepSeek · September 10, 2026

● activeOpen Weightmixture of expertsmultimodal
Parameters680B (38B active)
Context Window256K tokens

Description

Open-weight release introducing extreme KV-cache 4x memory compression and hyper-sparse MoE activation, dramatically lowering token serving barriers.

Benchmark Scores

MMLUMassive Multitask Language Understanding — 57 subjects
90.2%
HumanEvalCode generation pass@1 — Python problems
92.8%
MATHMATH benchmark — competition-level problems
96.1%
SWE-benchReal-world software engineering
65.4%

Key Innovations

Open Weight
Open WeightModel weights are publicly released but training data/code may not be. Enables fine-tuning but not full reproduction.
MoE
MoEArchitecture where only a fraction of the model's parameters are active for each input, allowing massive scale with lower compute.
Reasoning
ReasoningStructured step-by-step problem solving, often using chain-of-thought or tree-of-thought approaches.

Family Tree

Lineage