Key Capabilities
- SSE streaming — Real-time delivery of thinking chunks and image chunks
- Thinking mode — Internal reasoning chunks (
thought: true) streamed before the image - Text-to-image — Generate images from text descriptions
- Image editing — Pass a reference image via
inline_datacombined with text instructions - Aspect ratio control —
1:1,4:3,3:4,16:9,9:16 - Resolution control —
1K(~1024px),2K(~2048px),4K(~4096px, by longest side)
SSE Response Format
The streaming endpoint returns newline-delimited SSE data lines, each starting withdata: followed by a JSON object. There are three chunk types:
- Thinking chunk — Arrives first;
parts[0].thoughtistrue - Image chunk — Contains
parts[0].inlineDatawithmimeTypeand base64data(note: camelCase in streaming responses) - Final usage chunk — Contains top-level
usageMetadatawiththoughtsTokenCountand per-modality token details
In streaming responses, the image field is
inlineData (camelCase), while in the request body it is inline_data (snake_case). This is native Gemini API behavior.Text-to-Image Example
Image Editing Example (with Reference Image)
Pass both atext instruction and an inline_data reference image in the same parts array.
Parameters
API Reference
View the interactive API Playground for Gemini 2.5 Flash Image (Streaming).

