Text to video
Build a new scene from a chronological prompt covering subject, action, camera, lighting, shot changes, and pacing.
Create a short video from a written scene, animate one approved still, or guide a continuous transition with first and last frames.
Choose text, one image, or first and last frames, then set duration, resolution, and aspect ratio before reviewing the platform point estimate.
Generated content will appear here
Facts reviewed:
Wan 2.7 is a video generation model family available through Alibaba Cloud Model Studio. Official documentation lists text-to-video with multi-shot narrative and audio synchronization, plus image-to-video for first-frame generation, first-and-last-frame generation, video continuation, and audio-driven workflows. The documented output is 720p or 1080p, 2–15 seconds, 30 fps MP4 for the core text and image routes.
Nano Banana Pro currently exposes a focused subset: text-to-video, one-image animation, and first-and-last-frame generation. This workspace does not accept custom audio or video reference files and does not expose Wan’s separate reference-to-video or video-editing routes. Its aspect ratios, point estimates, and submission requirements are platform settings.
For text-to-video, treat the prompt like a compact edit sheet. Establish the main subject and visual grammar once, then give each shot a distinct purpose: orientation, action, evidence, reaction, or reveal. Repeat only the continuity constraints that matter across cuts—identity, wardrobe, prop geometry, screen direction, weather, and color—so the model has room to vary composition without losing the story.
For image-to-video, separate what must remain fixed from what may move. Preserve identity, object proportions, architecture, material colors, and camera axis; then specify motion amplitude, timing, camera path, wind, reflections, shadows, or focus changes. A narrow motion brief usually protects a prepared key visual better than asking every element to animate at once.
Before delivery, review the generated clip at normal speed and frame by frame. Check the first and last two seconds, transitions between shot sizes, hands and contact points, edge stability, product geometry, signage, repeated background figures, exposure jumps, and unintended cuts. Keep the prompt, source permissions, settings, point estimate, version, reviewer, and intended channel with the approved output.
Choose the simplest input route that gives the shot enough structure, while keeping broader official capabilities separate from this integration.
Build a new scene from a chronological prompt covering subject, action, camera, lighting, shot changes, and pacing.
Animate one permitted image while preserving its important subject, style, composition, and spatial relationships.
Upload two compatible images to define the starting frame and visual destination of a continuous transition.
Official documentation highlights connected multi-shot storytelling for the Wan 2.7 text-to-video route.
Wan 2.7 officially supports audio-enabled workflows, although this page does not currently expose audio input or an audio toggle.
Choose 720p for lower platform cost or 1080p when the concept needs more output detail.
Choose the simplest input route that gives the shot enough structure, while keeping broader official capabilities separate from this integration.
Draft a short sequence with connected locations, shot sizes, actions, and a consistent principal subject.
Animate approved visualizations with controlled camera movement, daylight changes, and restrained environmental motion.
Turn an approved key visual into a short horizontal, square, or vertical motion concept.
Use prepared first and last frames to explore seasons, lighting changes, transformations, and reveals.
Design mobile-first reveals with controlled material movement, negative space, stable product geometry, and a deliberate final hold.
Explore time-of-day, lighting, crowd density, and environmental changes between compatible architectural keyframes.
Each example loads the matching platform mode so you can revise it directly in the live workspace.
A multi-shot text brief with a consistent subject, location progression, and clear transitions.
A quiet cinematic travel sequence following the same woman in a rust-red coat through a coastal town at sunrise. Shot 1: wide view as she walks down stone steps toward the harbor. Shot 2: medium tracking shot beside colorful fishing boats while gulls cross the background. Shot 3: close-up as she unfolds an old postcard and looks toward the lighthouse. Natural transitions, consistent face and clothing, soft sea haze, realistic movement, no captions or logos.
Turn an approved interior image into a restrained camera and lighting study.
Preserve the room layout, furniture design, materials, window placement, and perspective from the uploaded image. Add a slow lateral camera move, sunlight gradually traveling across the wooden floor, subtle curtain movement, and a small shift in plant shadows. Photorealistic, stable geometry, no new objects, no text.
Connect matching compositions while controlling the environmental transformation between them.
Create one continuous transition from the uploaded summer opening frame to the uploaded winter ending frame. Keep the building, camera position, and main tree aligned. Leaves turn gold and drift away, the sky cools, light snow begins, and the ground gradually becomes white. No cuts, smooth time progression, stable architecture, physically coherent wind and snowfall.
Build a mobile-first product story with a clear opening hook, material motion, and a stable final hold.
9:16 vertical product film for an unbranded amber glass perfume bottle. 0–3s: extreme close-up as a narrow light streak travels across the glass and condensation beads. 4–8s: camera pulls back while the bottle rises slowly through low white mist. 9–12s: orbit thirty degrees and settle on a centered hero frame with clean negative space above. Preserve bottle shape, cap alignment, liquid level, and amber color; no labels, text, hands, or extra objects.
Use compatible first and last frames to test crowd continuity, lighting progression, and a deliberate camera path.
Connect the uploaded daytime plaza frame to the uploaded nighttime festival frame in one continuous forward glide. Keep building lines, fountain position, paving pattern, and camera height aligned. Pedestrians gradually increase without duplication, shop lights turn on in sequence, the sky shifts naturally to blue hour, and warm lanterns appear around the fountain. No cuts, no warped architecture, stable travel direction.
Start from text, upload one approved image, or provide two visually compatible first and last frames.
Describe each shot or movement chronologically, including subject continuity and the intended ending.
Choose duration, 720p or 1080p, and an aspect ratio that matches the intended publishing surface.
Inspect the current Nano Banana Pro point estimate and all selected settings before submitting the task.
Wan 2.7 can still produce identity drift, inconsistent objects, unstable hands, accidental cuts, or imperfect transitions. Review every frame before publishing, presenting, or delivering generated work.
Official Wan 2.7 documentation includes custom audio, video continuation, reference-to-video, and video editing capabilities that are not all available here. Nano Banana Pro currently accepts text and up to two frame images on this page. The 2–15 second range and 720p/1080p options align with core official routes; aspect ratios and point estimates describe this platform integration.
Only upload content you are authorized to use. Avoid deceptive impersonation, privacy violations, intellectual-property infringement, or presenting generated footage as documentary evidence. Preserve applicable provenance and disclosure information.
Model capabilities and regional availability can change. Official Alibaba Cloud documentation and Nano Banana Pro platform settings are listed separately below.
Return to the live workspace, choose the right source mode, and confirm resolution, duration, and points before generating.