DeepSeek Releases V4.1-Flash, Claims It Beats Its Own Flagship V4-Pro
DeepSeek released V4.1-Flash on September 10 — the smallest model in a new architecture family, but one the company says outperforms its own larger flagship, V4-Pro, on performance, cost, speed, and total runtime.
The mixture-of-experts model has 552 billion total parameters and uses a new causal encoder-decoder design that keeps just 8 billion parameters active while reading a prompt and 16 billion while generating a response. It was trained from scratch on 45 trillion tokens with a 1M-token context window and adds native visual understanding. A major focus was shrinking the key-value cache: entries are stored in 4-bit floating point, bringing the footprint to about 890 bytes per token — roughly a quarter of what the previous V4-Flash needed.
On Terminal-Bench 2.1, V4.1-Flash scored 90.6, narrowly ahead of Claude Opus 5 (89.1) and GPT-5.6 Sol (88.8). Starting September 14, API requests sent to V4-Pro will be silently answered by V4.1-Flash instead, billed at the smaller model's rates, until a full V4.1-Pro ships.
Photo via Wikimedia Commons, © Yuri Samoilov, licensed under CC BY 3.0.
Comments
No comments yet. Be the first to share your thoughts.
Leave a comment