Text to video
Describe the subject, action, camera, environment, lighting, and sound to generate a new short scene.
Turn a written scene, a still image, or defined first and last frames into a short AI video with the Veo 3.1 workflow available on this platform.
Choose text-to-video, image-to-video, or first-and-last-frame mode. Set the platform options and review the point estimate before generating.
Generated content will appear here
Facts reviewed:
Veo 3.1 is Google’s video generation model family for turning text and images into short videos. Google documents native audio, landscape and portrait output, image-to-video generation, and first-and-last-frame control. The official model family includes standard and Fast variants.
This page connects those capabilities to Nano Banana Pro’s current video workflow. Veo 3.1 is selected by default, and the model menu also lets you switch to another available video family. The supported mode, routing, duration, resolution, aspect ratio, audio controls, point costs, and routing labels are platform settings rather than a complete reproduction of the Google API.
Use a workflow that matches the source material and the amount of control your shot needs.
Describe the subject, action, camera, environment, lighting, and sound to generate a new short scene.
Upload a still as the opening visual, then guide motion, camera movement, atmosphere, and audio with a prompt.
Define both ends of a shot so the generated motion has a clearer visual start and destination.
Include dialogue, ambience, effects, and music cues in the prompt when using a route that generates audio.
Choose a 16:9 or 9:16 platform option for widescreen scenes or vertical social formats.
Use 720p for flexible durations; the workspace automatically limits 1080p and 4K selections to 8 seconds.
Use a workflow that matches the source material and the amount of control your shot needs.
Prototype hero shots, product reveals, mood films, and visual directions before a full production.
Explore short landscape and vertical clips for feeds, stories, ads, and campaign testing.
Turn written scenes or approved key frames into moving previsualization for creative review.
Test camera language, light changes, environmental motion, ambience, and sound cues.
Explore time-of-day changes, spatial atmosphere, weather, and controlled camera paths while preserving a location brief.
Test short spoken exchanges, room tone, effects, and visual timing before committing to production.
Load an example into the workspace, then revise the subject, action, shot language, timing, and audio cues.
A text-to-video brief with camera movement, materials, lighting, and sound.
Cinematic macro shot of a brushed silver travel bottle on dark volcanic stone at dawn. The camera makes a slow clockwise orbit as a narrow beam of warm sunlight reveals condensation and realistic reflections. Quiet wind, distant seabirds, subtle metallic resonance, no logos, premium commercial lighting.
Upload a suitable opening frame before using this motion-focused prompt.
Preserve the subject and composition of the uploaded image. Add a gentle forward camera push, natural movement in hair and fabric, soft drifting dust in the backlight, and subtle environmental motion. Keep facial features stable. Add quiet room tone and distant city ambience.
Upload the intended opening and closing frames, then describe the transition between them.
Create one continuous, physically plausible transition from the first frame to the last. The camera slowly cranes upward while the scene changes from blue hour to sunrise. Preserve architecture and perspective, introduce warm light gradually, and avoid cuts, sudden warping, or new foreground objects.
A portrait-format shot plan with blocking, pacing, and environmental audio.
9:16 portrait fashion film in one continuous shot. A model in a cobalt raincoat steps from a dim tram into a wet neon street, pauses under a transparent umbrella, then looks toward a passing light. Slow handheld follow, realistic reflections, restrained movement, soft rain, tram bell fading behind, no logos or on-screen text.
A first-and-last-frame brief that protects geometry while changing light and activity.
Interpolate smoothly from the uploaded daytime first frame to the blue-hour final frame. Preserve the building geometry, camera position, lens perspective, and signage placement. Let shadows lengthen, interior lights turn on floor by floor, and a few pedestrians enter naturally. One uninterrupted shot, no warping, no new architecture.
Start from text, upload one opening image, or provide both the first and last frames.
Specify subject, action, camera movement, environment, visual style, and desired audio.
Choose Fast or Pro, duration, resolution, ratio, and audio, then review the point estimate.
Check continuity, motion, details, dialogue, and sound before using or iterating on the result.
Generated video may contain unstable identities, distorted motion, inaccurate dialogue, inconsistent objects, or audio that does not perfectly match the scene. Review every frame and the full soundtrack before publication. Results should not be treated as documentary evidence.
Nano Banana Pro uses platform points and model-routing aliases. Fast and Pro labels, the optional audio control, exposed resolutions, and the estimate beside the button describe this integration. Google’s API availability, model IDs, quotas, and billing are separate. The workspace restricts 1080p and 4K to 8-second tasks in line with the reviewed official specifications.
Only upload images you are authorized to use. Do not create deceptive impersonation, violate privacy or intellectual-property rights, or use generated material in ways that breach platform policies or applicable law.
Video model capabilities and availability can change. Official Google documentation and Nano Banana Pro platform information are listed separately below.
Return to the workspace, choose the right input mode, and review the point estimate before generating.