{"id":39979,"date":"2026-07-31T16:27:11","date_gmt":"2026-07-31T09:27:11","guid":{"rendered":"https:\/\/dps.media\/deepseek-v4-flash-0731\/"},"modified":"2026-07-31T16:27:11","modified_gmt":"2026-07-31T09:27:11","slug":"deepseek-v4-flash-734","status":"publish","type":"post","link":"https:\/\/dps.media\/en\/deepseek-v4-flash-734\/","title":{"rendered":"DeepSeek V4 Flash: The \u201cUnbelievable\u201d Leap That Crushes Trillion-Dollar Rivals"},"content":{"rendered":"<?xml encoding=\"utf-8\" ?><p>Today (31\/07\/2026), the global artificial intelligence market was shaken once again as <a href=\"https:\/\/dps.media\/en\/deepseek-v4-flash-734\/\">DeepSeek V4 Flash<\/a> officially released its major update, codenamed 0731. More than just a routine technical upgrade, this new version delivers a remarkable performance breakthrough through post-training optimization. Many experts believe that OpenAI and Anthropic will have a \u201cheadache\u201d as a low-cost Flash model can now compete directly with frontier models.<\/p><h2>The destructive power of DeepSeek V4 Flash's post-training technique<\/h2><p>The most surprising aspect of this 31\/07 update is that the model's architecture and size remain completely unchanged from the previous preview version. DeepSeek continues to use the Mixture-of-Experts (MoE) architecture with a total of 284 billion parameters (284B), of which only 13 billion parameters are active (13B active) for each token. However, the results achieved after rerunning the entire post-training process using Reinforcement Learning (RL) have truly astonished the technology community.<\/p><p>Looking at the actual figures, version 0731 has improved scores in most programming benchmarks and agentic tasks:<\/p><ul>\n<li><strong>Terminal Bench 2.1:<\/strong> Achieved <strong>82.7 points<\/strong>, soaring from the preview version's 61.8. This surpasses GLM-5.2 (81.0 points) and comes close to OpenAI's heavyweight GPT-5.6 Luna (84.7 points).<\/li>\n<li><strong>DeepSWE (evaluating real-world software problem-solving capability):<\/strong> Recorded an unbelievable leap from 7.3 points to <strong>54.4 points<\/strong>, far ahead of GLM-5.2 (46.2 points).<\/li>\n<li><strong>NL2Repo:<\/strong> Achieved <strong>54.2 points<\/strong> (compared to 39.4 points in the preview version).<\/li>\n<li><strong>DSBench-FullStack:<\/strong> Achieved <strong>68.7 points<\/strong>, nearly doubling the previous version's score (37.0 points).<\/li>\n<\/ul><p><img decoding=\"async\" src=\"https:\/\/dps.media\/wp-content\/uploads\/2026\/07\/deepseek_v4_flash_buoc_nhay_vot_khong_tu_img2-scaled.png\" alt=\"DeepSeek V4 Flash\" style=\"display:block;margin:20px auto;max-width:100%;height:auto\" title=\"\"><\/p><h2>Hardware optimization: 10x smaller size, massive KV Cache<\/h2><p>In practical deployment, the GLM-5.2 model occupies up to 1.5TB, while the weight file of <a href=\"https:\/\/dps.media\/en\/deepseek-v4-flash-734\/\">DeepSeek V4 Flash<\/a> is only around 160GB. Being nearly 10 times smaller makes this model extremely easy to run on mainstream hardware. For developers or businesses wanting to self-host locally, equipping two professional graphics cards such as the RTX 6000 Pro or using the DSpark solution is enough to achieve smooth operation with impressive processing speed.<\/p><p>The next standout feature is its exceptional context compression capability. According to benchmark measurements, the KV Cache memory per token is proportional to the number of active parameters. With a 1 million token context (1M context), GLM-5.2 consumes up to 80GB of VRAM, while this new Flash model uses only <strong>6GB VRAM<\/strong> (even using less memory than Qwen 27B at the same context length). This is a decisive advantage for long-document processing applications or continuous conversational automation systems.<\/p><p>Not only does it save resources, but this model's API pricing is also unbelievably low. The current price in the <a href=\"https:\/\/api-docs.deepseek.com\/updates\/\" rel=\"nofollow noopener\" target=\"_blank\">official DeepSeek updated documentation<\/a> is only <strong>$0.14 per 1 million input tokens<\/strong> v\u00e0 <strong>$0.28 per 1 million output tokens<\/strong>. Compared to GPT-5.6 Luna's pricing of $0.20 input \/ $1.20 output, the Chinese solution is four times cheaper on output tokens, creating enormous pricing pressure on major US developers.<\/p><p><img decoding=\"async\" src=\"https:\/\/dps.media\/wp-content\/uploads\/2026\/07\/deepseek_v4_flash_buoc_nhay_vot_khong_tu_img3-scaled.png\" alt=\"DeepSeek V4 Flash\" style=\"display:block;margin:20px auto;max-width:100%;height:auto\" title=\"\"><\/p><h2>Real-world business applications with solutions from DPS.MEDIA<\/h2><p>The arrival of a smart, ultra-fast, and affordable model like this new Flash version opens up a golden opportunity for Vietnamese businesses. Instead of paying thousands of USD each month for expensive APIs, businesses can integrate this model into automated customer service workflows, AI coding assistants, or internal data management systems.<\/p><p>To optimize operating costs and apply AI technology to real business operations, businesses can explore the <a href=\"https:\/\/dps.media\/en\/\">Digital Marketing services<\/a> and comprehensive automation solutions from DPS.MEDIA. With experience implementing n8n Workflows automation systems and integrating advanced AI tools, we help SMEs free up labor, reduce manual errors, and maximize operational efficiency.<\/p><p>If you want to begin your digital transformation journey, integrate next-generation AI assistants, or optimize your brand's omnichannel marketing operations, contact DPS.MEDIA today via Hotline\/Zalo: <strong>0961545445<\/strong> or visit our office at: <strong>56 Nguyen Dinh Chieu, Tan Dinh Ward, District 1, Ho Chi Minh City<\/strong> for the fastest support.<\/p>","protected":false},"excerpt":{"rendered":"<p>Detailed review of the DeepSeek V4 Flash (0731) update, featuring outstanding coding performance that surpasses GLM-5.2 and extremely cost-effective API pricing.<\/p>","protected":false},"author":0,"featured_media":39976,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"footnotes":""},"categories":[1933,70],"tags":[675],"class_list":["post-39979","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-tech","category-tin-tuc","tag-dps-media"],"acf":[],"rankmath_keywords":{"primary":"DeepSeek V4 Flash","secondary":[""]},"yoast_keywords":{"primary":"","secondary":[]},"yoast_focuskw":"","rankmath_focuskw":"DeepSeek V4 Flash","seo_keywords":{"primary":"DeepSeek V4 Flash","secondary":[""]},"_links":{"self":[{"href":"https:\/\/dps.media\/en\/wp-json\/wp\/v2\/posts\/39979","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dps.media\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/dps.media\/en\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/dps.media\/en\/wp-json\/wp\/v2\/comments?post=39979"}],"version-history":[{"count":0,"href":"https:\/\/dps.media\/en\/wp-json\/wp\/v2\/posts\/39979\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dps.media\/en\/wp-json\/wp\/v2\/media\/39976"}],"wp:attachment":[{"href":"https:\/\/dps.media\/en\/wp-json\/wp\/v2\/media?parent=39979"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/dps.media\/en\/wp-json\/wp\/v2\/categories?post=39979"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/dps.media\/en\/wp-json\/wp\/v2\/tags?post=39979"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}