Gemini 3.6 Flash along with two specialized variants, 3.5 Flash-Lite and 3.5 Flash Cyber, has just been officially introduced by Google. This upgrade focuses on solving the two biggest challenges for developers: optimizing agent operating costs and increasing task processing speed without compromising output quality.

According to official announcement from the Google Blog, Flash 3.6 not only excels in programming benchmarks but also significantly reduces wasted tokens in multi-step workflows.

Gemini 3.6 Flash
Google has just launched the trio of new Gemini models.

Gemini 3.6 Flash: Optimized costs and reduced token consumption

Anyone building AI Agents knows the headache of continuous reasoning loops causing API bills to skyrocket. Gemini 3.6 Flash was created to address this issue directly. According to statistics from the Artificial Analysis Index, 3.6 Flash consumes 17% fewer output tokens than its predecessor, 3.5 Flash.

Google has set highly competitive pricing: $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. By reducing unnecessary reasoning steps and limiting unnecessary tool calls, the actual cost per agent task has dropped significantly.

Gemini 3.6 Flash token consumption performance on OSWorld
Comparison chart of the token efficiency of 3.6 Flash versus 3.5 Flash on the OSWorld benchmark.

Real-world performance across benchmarks

Despite being more cost-efficient, 3.6 Flash delivers clear improvements across many complex scenarios:

  • Advanced programming: DeepSWE results reached 49% (compared to 37% for 3.5 Flash), reducing erroneous code edits and shortening execution loops. Machine learning research results on MLE Bench increased to 63.9%.
  • Computer Use: The OSWorld-Verified score increased from 78.4% to 83.0%. This feature is now built in as a client-side tool in the Gemini API and Gemini Enterprise.
  • Knowledge processing & multimodal capabilities: Achieved a score of 1421 on GDPval-AA v2. Enterprises such as Hebbia and Harvey reported that the model extracts tables, analyzes documents, and prepares financial reports with high accuracy.
Output quality comparison between Gemini 3.6 Flash and previous versions
The output quality and accuracy of 3.6 Flash excel across many agentic tasks.

Gemini 3.5 Flash-Lite: 350 tokens per second for high-load systems

If you need an ultra-fast response model for powering agent search or processing millions of invoices every day, 3.5 Flash-Lite is the answer. With processing speeds of up to 350 output tokens per second, it is the fastest model in the 3.5 series.

Pricing for 3.5 Flash-Lite is $0.3 per 1M input tokens and $2.5 per 1M output tokens. Despite its low cost, this model still outperforms the older Gemini 3 Flash in SWE-Bench Pro tests (54.2% versus 49.6%) and Terminal-Bench 2.1 (54% versus 31%).

Gemini 3.5 Flash-Lite model performance
Gemini 3.5 Flash-Lite is optimized for workflows requiring low latency and high throughput.

Gemini 3.5 Flash Cyber: A security vulnerability hunting assistant

Alongside the two public releases, Google introduced Gemini 3.5 Flash Cyber, a model fine-tuned specifically for cybersecurity. When combined with the CodeMender security agent, the system can automatically detect, test, and patch software vulnerabilities before attackers can exploit them.

Due to the sensitive nature of cybersecurity technology, Google stated that 3.5 Flash Cyber is currently available only on a limited basis to government agencies and trusted partners through the CodeMender testing program.

Gemini 3.5 Flash Cyber security model evaluation
Gemini 3.5 Flash Cyber works with CodeMender to efficiently find and fix source code vulnerabilities.

Try and deploy the new AI models in real-world applications

Starting today, developers can try Gemini 3.6 Flash and 3.5 Flash-Lite on Google AI Studio, Android Studio, and the Google Antigravity platform. In addition, Google confirmed that it has begun pre-training for the Gemini 4 generation and will soon expand the 3.5 Pro release.

If your business wants to apply the latest AI models to automate processes or optimize customer growth campaigns, please refer to high-quality Digital Marketing services of DPS.MEDIA for guidance on the most suitable implementation roadmap.

Ý kiến bạn đọc (0)