RELEASED
DeepSeek releases V4.1 Flash, priced to undercut the frontier
2026-09-11
DeepSeek launched V4.1 Flash, a 552B mixture-of-experts model (only ~8B parameters active while reading, ~16B while writing) with aggressive KV-cache compression that pushes API prices far below Western frontier models. DeepSeek is routing flagship V4 Pro traffic to it. Independent coding benchmarks put it roughly on par with GPT-5.6 Sol — Matthew Berman covered the launch the same week.