Most cost-efficient model in the Gemini 3 family, offering excellent performance at a fraction of the cost. Optimized for high throughput and cost-sensitive deployments with 1M token context window and configurable thinking capabilities.
- Most cost-efficient model in the Gemini 3 family at $0.25 input, $1.50 output per 1M tokens
- 1M token context window with full multimodal support for text, audio, images, video, and documents
- Optimized for high throughput and low-latency applications requiring speed and cost efficiency
- Configurable thinking capabilities for flexible reasoning depth based on task complexity
- Full tool support including Search Grounding, Code Execution, and Function Calling
Web Search
Search the web for current info.
Code Execution
Run Python code in a sandbox.
JSON Mode
Output responses in valid JSON format.
URL Context
Fetch and process content from URLs.
Google Maps
Ground responses in real-time Google Maps data for location-aware queries.
Image Generator
Generate images from text prompts.
Model Information
Supported Formats
Google's most intelligent stable model for coding tasks, released July 21, 2026 as the successor to Gemini 3.5 Flash. Combines frontier-class intelligence with Flash-tier speed, with full support for thinking, function calling, structured outputs, code execution, URL context, and search/Maps grounding.
The fastest and cheapest model in Google's 3.5 line, released July 21, 2026 as the successor to Gemini 3.1 Flash Lite. Built for high-volume, low-latency jobs like agentic search, document processing, and translation.
Google's most advanced reasoning model (Gemini 3.1 Pro) with configurable thinking levels (minimal/low/medium/high) for optimal cost-performance balance. Successor to Gemini 3 Pro with enhanced reasoning and tool capabilities including Google Maps grounding.