Multimodal references
Combine text, image and short video inputs where the API supports them.
● Video Generator
Type a prompt, drop an image, or do both. Gemini Omni Flash returns a cinematic clip in seconds.
Turn text, images, audio, and video into cinematic clips — powered by Google's Gemini Omni Flash model.
No Google subscription, no waitlist. Sign in with email and start generating.
Multimodal input · Conversational editing · Real-world physics

● Explainer
Google's multimodal video model combines text, image and video references with conversational editing. The public preview model identifier is gemini-omni-flash-preview; Google disclosed a price of $0.10 per output second and a current 10-second generation limit.
Combine text, image and short video inputs where the API supports them.
Iterate on a generated video using natural-language editing instructions.
Audio references and scene extension were not supported in the preview API disclosure; verify current docs before production use.
Dates move only when the underlying capability, access or evidence changes.
New comparison guides, localized pages and dedicated visual examples are now available.
Google announced Omni 1.1 Flash with first/last frame control, scene extension and 4K upscaling.
Gemini Omni Flash became available through Google AI Studio and the Gemini API.
● Decision guides
Each page separates vendor-published facts from Gomni measurements. Where the benchmark dataset is missing, the score remains Not tested.
Compare access, reference controls, first-frame testing and disclosed costs.
Compare Google and OpenAI workflows without inventing a permanent winner.
A status-led comparison with unverified fields left explicitly blank.
● FAQ
Sign up with email, claim your starter credits, and start creating in under a minute.