Today (31/07/2026), the global artificial intelligence market was shaken once again as DeepSeek V4 Flash officially released its major update, codenamed 0731. More than just a routine technical upgrade, this new version delivers a remarkable performance breakthrough through post-training optimization. Many experts believe that OpenAI and Anthropic will have a “headache” as a low-cost Flash model can now compete directly with frontier models.
The destructive power of DeepSeek V4 Flash's post-training technique
The most surprising aspect of this 31/07 update is that the model's architecture and size remain completely unchanged from the previous preview version. DeepSeek continues to use the Mixture-of-Experts (MoE) architecture with a total of 284 billion parameters (284B), of which only 13 billion parameters are active (13B active) for each token. However, the results achieved after rerunning the entire post-training process using Reinforcement Learning (RL) have truly astonished the technology community.
Looking at the actual figures, version 0731 has improved scores in most programming benchmarks and agentic tasks:
- Terminal Bench 2.1: Achieved 82.7 points, soaring from the preview version's 61.8. This surpasses GLM-5.2 (81.0 points) and comes close to OpenAI's heavyweight GPT-5.6 Luna (84.7 points).
- DeepSWE (evaluating real-world software problem-solving capability): Recorded an unbelievable leap from 7.3 points to 54.4 points, far ahead of GLM-5.2 (46.2 points).
- NL2Repo: Achieved 54.2 points (compared to 39.4 points in the preview version).
- DSBench-FullStack: Achieved 68.7 points, nearly doubling the previous version's score (37.0 points).

Hardware optimization: 10x smaller size, massive KV Cache
In practical deployment, the GLM-5.2 model occupies up to 1.5TB, while the weight file of DeepSeek V4 Flash is only around 160GB. Being nearly 10 times smaller makes this model extremely easy to run on mainstream hardware. For developers or businesses wanting to self-host locally, equipping two professional graphics cards such as the RTX 6000 Pro or using the DSpark solution is enough to achieve smooth operation with impressive processing speed.
The next standout feature is its exceptional context compression capability. According to benchmark measurements, the KV Cache memory per token is proportional to the number of active parameters. With a 1 million token context (1M context), GLM-5.2 consumes up to 80GB of VRAM, while this new Flash model uses only 6GB VRAM (even using less memory than Qwen 27B at the same context length). This is a decisive advantage for long-document processing applications or continuous conversational automation systems.
Not only does it save resources, but this model's API pricing is also unbelievably low. The current price in the official DeepSeek updated documentation is only $0.14 per 1 million input tokens và $0.28 per 1 million output tokens. Compared to GPT-5.6 Luna's pricing of $0.20 input / $1.20 output, the Chinese solution is four times cheaper on output tokens, creating enormous pricing pressure on major US developers.

Real-world business applications with solutions from DPS.MEDIA
The arrival of a smart, ultra-fast, and affordable model like this new Flash version opens up a golden opportunity for Vietnamese businesses. Instead of paying thousands of USD each month for expensive APIs, businesses can integrate this model into automated customer service workflows, AI coding assistants, or internal data management systems.
To optimize operating costs and apply AI technology to real business operations, businesses can explore the Digital Marketing services and comprehensive automation solutions from DPS.MEDIA. With experience implementing n8n Workflows automation systems and integrating advanced AI tools, we help SMEs free up labor, reduce manual errors, and maximize operational efficiency.
If you want to begin your digital transformation journey, integrate next-generation AI assistants, or optimize your brand's omnichannel marketing operations, contact DPS.MEDIA today via Hotline/Zalo: 0961545445 or visit our office at: 56 Nguyen Dinh Chieu, Tan Dinh Ward, District 1, Ho Chi Minh City for the fastest support.
Ý kiến bạn đọc (9)
Chưa có bình luận nào. Hãy là người đầu tiên chia sẻ cảm nghĩ!
Đăng nhập / Tạo tài khoản
Đăng nhập với Google để gửi bình luận và tương tác cùng cộng đồng.
Tiếp tục là đồng ý với điều khoản sử dụng và chính sách bảo mật của DPS Media.