Skip to content
AI

DeepSeek Launches V4.1-Flash AI Model With Faster Inference and Lower Costs

DeepSeek Launches V4.1-Flash AI Model With Faster Inference

Chinese AI startup DeepSeek has launched DeepSeek-V4.1-Flash, introducing a new AI model designed to deliver stronger performance while using fewer resources. The company describes it as the smallest model in its new architecture family, with native multimodal visual understanding and a focus on faster inference, higher throughput and improved scalability.

The launch comes as DeepSeek continues to compete with leading AI developers in China and abroad. The company is also preparing for a potential initial public offering on Shanghai's technology-focused STAR Market, making the latest model release an important step as it works to strengthen both its technology and commercial position.

New Architecture Targets Higher Efficiency

DeepSeek-V4.1-Flash uses a 552-billion-parameter Mixture-of-Experts architecture based on a new Causal Encoder-Decoder design. However, only around 8 billion parameters are activated for input processing and 16 billion for output generation, allowing the model to handle workloads with a smaller active computational footprint.

DeepSeek says the model combines new pretraining methods with larger-scale reinforcement learning after training. The company also claims that its benchmark results put V4.1-Flash ahead of several flagship models, including DeepSeek-V4-Pro, although these performance comparisons are based on the company's own evaluations.

Native Multimodal Understanding Comes Built In

V4.1-Flash introduces native visual understanding, allowing it to process images alongside text. This capability is built directly into the model architecture, giving developers a single system for handling both text and visual inputs.

The model is now available through DeepSeek's API under the deepseek-flash designation. Official partners including WorkBuddy and OpenCode have also integrated V4.1-Flash, expanding access for developers using the model in coding, automation and agent-based workflows.

Smaller Cache Reduces Infrastructure Needs

A major focus of the new model is its reduced KV cache requirement, which determines how much memory and storage an AI system needs to retain contextual information during processing. DeepSeek says V4.1-Flash requires only one-quarter of the HBM and one-eighth of the SSD storage needed by the previous generation for its cache.

The reduction could be particularly useful for AI agents and long-running workloads, where maintaining large amounts of context can add significantly to infrastructure costs. By reducing the cache footprint, DeepSeek says V4.1-Flash can support more users while requiring fewer resources.

DeepSeek Cuts API Costs

The efficiency improvements are also reflected in the model's pricing. DeepSeek has reduced API costs for V4.1-Flash and continues to use separate peak and off-peak rates, with off-peak pricing set at 50% of peak prices.

The new pricing took effect on September 10. DeepSeek is also phasing out older V4-Flash models, with compatible requests temporarily routed to V4.1-Flash. From September 14, requests using V4-Pro will also be routed to V4.1-Flash at the new model's pricing until a future V4.1-Pro model becomes available.

V4.1-Flash Expands Open-Source Deployment

DeepSeek is also working with the open-source community to expand inference support and deployment options for the new model. The company has made V4.1-Flash available to developers and is encouraging large-scale users with substantial GPU and storage infrastructure to explore deployments.

This approach continues DeepSeek's strategy of combining advanced model capabilities with relatively accessible deployment options. The company is putting particular emphasis on efficiency, aiming to make powerful AI systems more practical for developers and organisations that need to run them at scale.

DeepSeek Strengthens Its Position in the AI Race

The V4.1-Flash launch reflects a broader shift in the AI industry toward improving inference efficiency and operating costs, rather than relying only on larger models and more computing power. DeepSeek's new architecture, reduced cache requirements and native multimodal capabilities are all aimed at making advanced AI more efficient to operate.

For DeepSeek, the release also arrives at an important point in its expansion. As competition among Chinese and global AI companies intensifies, the company is seeking to demonstrate that its models can combine capability, speed and affordability while it moves toward a potential public listing.