DeepSeek's 552B V4.1-Flash adds native vision while shrinking memory and prices

DeepSeek's 552B V4.1-Flash adds native vision while shrinking memory and prices

DeepSeek has released V4.1-Flash, a 552-billion-parameter mixture-of-experts model with native image understanding. Its causal encoder-decoder design activates only 8 billion parameters per input token and 16 billion on output. The company says the model's cache needs a quarter of the previous generation's HBM and an eighth of its SSD storage. API prices fall, with off-peak rates half of peak.

3 sources

  1. 01MarkTechPost — DeepSeek AI Released DeepSeek-V4.1-Flash with 1M Context, FP4 KV Cache, and Cross-Layer Attention Reusemarktechpost.com
  2. 02Neowin — DeepSeek launches V4.1-Flash multimodal reasoning modelneowin.net
  3. 03TechBriefly — DeepSeek launches V4.1-Flash with ultra-cheap off-peak API pricingtechbriefly.com