Conversational video editing
Refine and edit video using natural-language follow-ups.
Generate cinematic video from text, images and visual references with conversational editing and flexible creative control.
Published May 19, 2026 · API preview June 30, 2026 · Last verified September 7, 2026

● Officially disclosed
These are vendor-published capabilities, not Gomni benchmark claims.
Refine and edit video using natural-language follow-ups.
Use combinations of text, image and video inputs for scene control.
Google describes the model as using Gemini knowledge for history, biology and narrative logic.
Google AI Studio and the Gemini API list gemini-omni-flash-preview. Google published $0.10 per output second in June 2026. Preview terms, regions and price can change; re-check Google's current docs.
Google AI Studio and Gemini API
Pay-as-you-go generation workspace
Decorative visuals on this page illustrate workflow concepts. They are not presented as Gemini Omni Flash outputs. A model-output gallery will publish only after prompt, parameters and endpoint are recorded.
Store the exact prompt and reference inputs.
Record endpoint, version, duration, ratio and resolution.
Keep successes and failures in the benchmark sample.
Start with a text prompt, reference image, or both.
Choose the duration, resolution and aspect ratio for your project.
Check consistency, motion and framing before downloading or iterating.
● FAQ
Create with pay-as-you-go credits and no monthly commitment.