Text to video
Describe subjects, interaction, camera language, lighting, shot order, dialogue, effects, and ambience in one production brief.
Turn a scene description or approved reference images into a short AI video with an optional generated soundtrack and platform controls.
Choose text, one image, or first and last frames, then set duration, resolution, aspect ratio, and generated audio before reviewing the point estimate.
Generated content will appear here
Facts reviewed:
Seedance 2.0 is a video creation model from ByteDance Seed. The official launch describes a unified multimodal audio-video architecture that can use text, images, audio, and video as references. It emphasizes complex motion, physical plausibility, reference consistency, editing, extension, multi-shot storytelling, and synchronized stereo audio.
The full model can accept richer mixed references, but Nano Banana Pro currently exposes a narrower live workflow: text-to-video, one-image animation, or a platform first-and-last-frame route. This page does not offer audio or video reference uploads. Its 2–15 second duration range, 480p–1080p resolutions, aspect ratios, and point estimates are platform settings.
Use the standard route when a direction has moved beyond rough exploration and needs a deliberate production pass. Build a continuity brief before generating: lock wardrobe, prop dimensions, screen direction, eyelines, lens character, depth of field, motivated key light, color palette, weather, dialogue pronunciation, and ambient texture. Arrange these constraints around chronological story beats so visual and acoustic decisions reinforce the same scene.
Review a candidate at its intended playback speed, then inspect frame grabs at action peaks and shot boundaries. Check faces, hands, object edges, product markings, spatial geography, exposure shifts, and transition logic. Listen on headphones for clipping, channel imbalance, abrupt ambience, masked dialogue, and effects that arrive early or late. A polished-looking clip still needs this delivery review before approval.
For commercial handoff, package a provenance sheet with source licenses, performer consent, brand approvals, music clearance, generation timestamp, reviewer initials, version identifier, intended channels, distribution territory, and retention date. Mark unresolved defects and ownership conditions explicitly. This record keeps approved frames attached to their usage terms and gives legal, marketing, and production teams one auditable release reference.
Use the model for motion-rich ideas while keeping official capabilities separate from the controls available in this integration.
Describe subjects, interaction, camera language, lighting, shot order, dialogue, effects, and ambience in one production brief.
Animate one permitted still image while directing motion, performance, camera movement, timing, and sound.
Upload first and last images in the workspace to guide the beginning and destination of a continuous transition.
Enable generated audio and write dialogue, music, effects, and environmental sound alongside the visual timeline.
The official release highlights improved stability and physical plausibility for multi-subject interaction and demanding movement.
Plan a compact sequence of connected shots; the official model supports high-quality audio-video output up to 15 seconds.
Use the model for motion-rich ideas while keeping official capabilities separate from the controls available in this integration.
Explore staging, connected shots, performance, dialogue, and sound before committing to full production.
Prototype sports, dance, interaction, and camera choreography that require coherent timing and physical motion.
Animate an approved character or portrait for short horizontal, square, or vertical campaign concepts.
Test controlled movement, lighting transitions, sound cues, and compact visual storytelling around approved assets.
Block short exchanges with explicit speaker order, camera timing, lip synchronization, room tone, and reaction beats.
Use compatible keyframes to plan camera travel through interiors, exteriors, changing light, and connected spaces.
Each example loads the matching platform mode so you can revise the prompt in the live workspace.
A text-to-video brief with interacting subjects, camera continuity, physical detail, and synchronized sound.
Cinematic indoor fencing practice between two experienced athletes. Begin with a low tracking shot beside their feet, rise into a medium two-shot as they exchange three controlled attacks and parries, then arc behind the winning athlete during the final touch. Accurate footwork, realistic blade flex, consistent uniforms and faces, cool overhead light, shoe squeaks, blade impacts, restrained room reverb, no titles or logos.
Animate an approved portrait while preserving appearance, clothing, and the original composition.
Preserve the person’s facial identity, hairstyle, clothing, and background composition from the uploaded image. Add a natural breath, a brief glance toward the window, one small smile, and a slow camera push-in. Keep hands outside the frame, maintain realistic skin texture and lighting, soft city ambience, no dialogue, no added text.
Connect two compatible prepared frames with a clear transformation and camera path.
Transition continuously from the uploaded opening frame to the uploaded ending frame. The camera moves forward through a curtain of drifting paper shapes as the lighting changes from cool dawn to warm sunset. Preserve the central subject and scene geometry, avoid cuts, keep movement smooth and physically coherent, add gentle paper rustle and distant wind.
Coordinate restrained acting, shot progression, lip timing, and environmental sound in a compact scene.
Naturalistic two-character scene in a late-night laundromat. 0–5s: locked medium two-shot while one character folds a blue shirt and says, “I kept your key.” 6–10s: slow push toward the second character as they answer, “Then you knew I would return.” 11–15s: hold eye contact while a washer stops behind them. Stable faces and wardrobe, accurate lip sync, fluorescent hum and machine click, no music or subtitles.
Test a designed camera path and physically coherent lighting change between prepared architectural keyframes.
Connect the uploaded daylight lobby frame to the uploaded evening rooftop frame in one continuous architectural move. Camera glides forward, rises through the central atrium, and clears the roofline as daylight shifts gradually to blue hour. Preserve structural lines, material colors, and travel direction; no cuts, no warped columns, subtle room tone changing into rooftop wind.
Start from text, upload one still image, or provide two compatible frames for the platform transition workflow.
Describe the subject, action, shot order, camera movement, lighting changes, and ending in chronological order.
If generated audio is enabled, place dialogue, effects, music, and ambience beside the visual event they should match.
Choose duration, resolution, and aspect ratio, then inspect the current point estimate before submitting.
Seedance 2.0 can still produce identity drift, unstable hands or objects, unintended cuts, inaccurate dialogue, or imperfect sound timing. Review the complete video and audio track before publication or client delivery.
ByteDance describes broader multimodal reference and editing capabilities than this page currently exposes. Nano Banana Pro accepts text and up to two frame images here; it does not accept audio or video reference files. The available 2–15 second range, resolutions, aspect ratios, generated-audio switch, and points belong to this integration rather than the official model product.
Only upload material you are authorized to use. Do not impersonate people, mislead viewers, violate privacy, infringe intellectual property, or remove required provenance and disclosure information from generated content.
Model capabilities and availability can change. Official ByteDance Seed sources and Nano Banana Pro platform settings are listed separately below.
Return to the live workspace, choose the right input mode, and review the current point estimate before generating.