Gemini Omni Flash — Google's multimodal video model

Generate cinematic video from text, images and visual references with conversational editing and flexible creative control.

Published May 19, 2026 · API preview June 30, 2026 · Last verified September 7, 2026

A multimodal filmmaking studio combining image, text and video references

  Officially disclosed

Core Gemini Omni Flash capabilities

These are vendor-published capabilities, not Gomni benchmark claims.

Conversational video editing

Refine and edit video using natural-language follow-ups.

Multimodal referencing

Use combinations of text, image and video inputs for scene control.

Real-world knowledge

Google describes the model as using Gemini knowledge for history, biology and narrative logic.

Official access and published preview price

Google AI Studio and the Gemini API list gemini-omni-flash-preview. Google published $0.10 per output second in June 2026. Preview terms, regions and price can change; re-check Google's current docs.

Developer access

Google AI Studio and Gemini API

Gomni workflow

Pay-as-you-go generation workspace

Evidence policy for examples

Decorative visuals on this page illustrate workflow concepts. They are not presented as Gemini Omni Flash outputs. A model-output gallery will publish only after prompt, parameters and endpoint are recorded.

Prompt

Store the exact prompt and reference inputs.

Parameters

Record endpoint, version, duration, ratio and resolution.

Result

Keep successes and failures in the benchmark sample.

How to use Gemini Omni Flash

01

Choose an input

Start with a text prompt, reference image, or both.

02

Generate a clip

Choose the duration, resolution and aspect ratio for your project.

03

Review and refine

Check consistency, motion and framing before downloading or iterating.

  FAQ

Gemini Omni Flash questions

Try Gemini Omni Flash on Gomni

Create with pay-as-you-go credits and no monthly commitment.