DeepSeek V4.1 Flash
DeepSeek · September 10, 2026
● activeOpen Weightmixture of expertsmultimodal
Parameters680B (38B active)
Context Window256K tokens
Description
Open-weight release introducing extreme KV-cache 4x memory compression and hyper-sparse MoE activation, dramatically lowering token serving barriers.
Benchmark Scores
MMLUMassive Multitask Language Understanding — 57 subjects
90.2%HumanEvalCode generation pass@1 — Python problems
92.8%MATHMATH benchmark — competition-level problems
96.1%SWE-benchReal-world software engineering
65.4%Key Innovations
Open Weight
Open WeightModel weights are publicly released but training data/code may not be. Enables fine-tuning but not full reproduction.
MoE
MoEArchitecture where only a fraction of the model's parameters are active for each input, allowing massive scale with lower compute.
Reasoning
ReasoningStructured step-by-step problem solving, often using chain-of-thought or tree-of-thought approaches.