Gemini 3.6 Flash and Gemini 3.5 Flash-Lite Are Now Available on RouterHub
RouterHub Team · Updated 2026-07-22
Gemini 3.6 Flash (google/gemini-3.6-flash) and Gemini 3.5 Flash-Lite (google/gemini-3.5-flash-lite) are now available on RouterHub, starting from the day of Google’s July 21, 2026 launch. Developers can access both models through RouterHub’s Chat API.
The two releases address different needs. Gemini 3.6 Flash brings stronger performance to coding, AI agents, multimodal reasoning, and knowledge work. Gemini 3.5 Flash-Lite focuses on speed and efficiency for document processing, extraction, search, classification, translation, and other high-volume tasks.
Availability on RouterHub
| Gemini 3.6 Flash | Gemini 3.5 Flash-Lite | |
|---|---|---|
| Provider | ||
| RouterHub model ID | google/gemini-3.6-flash | google/gemini-3.5-flash-lite |
| Availability | Available now | Available now |
| RouterHub API type | Chat API | Chat API |
| Inputs | Text, image, file, audio, and video | Text, image, file, audio, and video |
| Output | Text | Text |
| Maximum input | 1,048,576 tokens | 1,048,576 tokens |
| Maximum output | 65,536 tokens | 65,536 tokens |
| Best starting fit | Complex coding, agents, multimodal analysis, and demanding knowledge work | High-volume processing, extraction, search, and focused automation |
Although the models share the same context limits and input formats, they are built for different jobs: Gemini 3.6 Flash raises the capability ceiling, while Gemini 3.5 Flash-Lite is designed to make large-scale processing faster and more economical.
Gemini 3.6 Flash: Coding, Agentic, and Multimodal Capability
Gemini 3.6 Flash is Google’s new workhorse model for tasks that combine code, tools, long context, and visual information. It is a strong fit for AI coding products, complex agent workflows, document-heavy research, technical charts, and multimodal applications.
For developers, the important shift is that more demanding work can remain on a Flash-class model. A coding agent can analyze a repository, plan changes, use tools, and iterate on results; a document application can connect text with charts, images, audio, or video; and a knowledge assistant can work across much larger source collections.
Its one-million-token input limit also opens the door to large repositories and document collections. As with any long-context model, real-world retrieval quality still depends on how information is structured and how much context an application sends at once.
Gemini 3.5 Flash-Lite: High-Throughput, Cost-Efficient Execution
Gemini 3.5 Flash-Lite is built for scale. It is designed for applications that process large numbers of requests and need to keep response time and per-request cost under control. Strong use cases include document parsing, structured extraction, classification, translation, search support, and repetitive background processing.
It can also handle specialized steps inside a multi-agent system—for example, extracting fields from incoming documents, translating receipts, classifying search results, or preparing structured data before a more capable model takes over.
Flash-Lite still supports multimodal input and long context, so its role is broader than simple text transformation. It gives teams a lighter option for bringing image, audio, video, and document understanding into products that operate at high volume.
Provider-Published Benchmark Comparison
Google’s published benchmark results show how the two models differ across coding, agentic, knowledge-work, chart-reasoning, and long-context tasks. These are Google and Google DeepMind results, not RouterHub-run benchmarks, and they should not be combined into a single overall score.
| Evaluation | Gemini 3.6 Flash | Gemini 3.5 Flash-Lite | Focus |
|---|---|---|---|
| SWE-Bench Pro (Public) | 58.7% | 54.2% | Diverse agentic coding tasks |
| Terminal-bench 2.1 | 78.0% | 54.0% | Agentic terminal coding, Terminus-2 harness |
| MLE-Bench | 63.9% | 39.2% | Machine-learning engineering |
| GDPVal-AA v2 | 1,421 Elo | 1,140 Elo | Economically valuable knowledge work |
| CharXiv Reasoning, no tools | 85.2% | 74.5% | Information synthesis from complex charts |
| CharXiv Reasoning, with tools | 89.4% | 76.5% | Chart reasoning with search and code execution |
| GDM-MRCR v2, 128K average | 91.8% | 72.2% | Cumulative long-context retrieval |
| GDM-MRCR v2, 1M pointwise | 54.0% | 21.3% | Retrieval at the full context length |
Benchmark settings vary. Google reports the Flash-Lite results with high thinking unless otherwise noted, while some evaluations use different tool or reasoning conditions. The 128K long-context result is cumulative and the 1M result is measured at the full context length, so those two rows are not a direct like-for-like comparison.
RouterHub Pricing
RouterHub pricing checked July 22, 2026, per one million tokens:
| Model | Input | Output |
|---|---|---|
| Gemini 3.6 Flash | $1.50 | $7.50 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 |
At these rates, Gemini 3.6 Flash is positioned for quality-sensitive work that still benefits from Flash-class economics, while Gemini 3.5 Flash-Lite gives high-volume products a substantially lighter cost profile. The final cost of a task will also depend on how many tokens and tool calls it takes to reach a usable result.
Implications for AI Product Teams
Together, these models expand both ends of Google’s Flash portfolio. Gemini 3.6 Flash pushes a fast model further into complex coding, agents, and professional work. Gemini 3.5 Flash-Lite makes multimodal and agentic capabilities more accessible for products where every millisecond and every token matter.
They are not simply a premium model and a budget model. They represent two useful ways to build: one optimized for harder tasks, and one optimized for scale. Many AI products may benefit from using both—Gemini 3.6 Flash where deeper capability changes the outcome, and Gemini 3.5 Flash-Lite where fast, focused execution matters most.
With both available on RouterHub from launch day, developers can start bringing these new capabilities into applications immediately. Gemini 3.6 Flash raises what teams can expect from a fast general-purpose model; Gemini 3.5 Flash-Lite broadens how economically those capabilities can be used at scale.