Wan AI Video GeneratorCreate with the latest model, then compare the complete Wan family

Create with Wan 3.0 using text, images, video, and audio references. Compare every available Wan model below and choose the right workflow for your next video.

Reference media
Model
Resolution
Aspect ratio
5s

A light, dreamy live-action commercial: a young European-American girl blows bubbles alone on a sunny beach. The visual is bright and breezy in vertical format, built from clear sea-sky blue, golden side backlight, pale beige sand, and iridescent highlights flowing across soap-bubble surfaces. A six- or seven-year-old girl stands at the edge of a windy seaside boardwalk, pale blonde hair continuously lifted by the wind, wearing a light summer dress and white sporty sandals, holding a small bubble solution bottle and a round children's wand. She looks excited and focused: first dipping carefully, then looking up and blowing gently, countless bubbles surrounding her as she laughs and reaches to pop them, continuously interacting with bubbles and sea breeze. Bubble formation is explicit: the clear bottle's solution gleams wet in the sun, the wand lifted from the bottle carries a complete iridescent film with a droplet hanging from its rim; the child exhales evenly, the film first tightens into a tiny mirror with rainbow interference patterns, then slowly bulges into a hemisphere, and finally a large main bubble detaches followed by two or three smaller ones carried away by the sea wind. Once free, bubbles are immediately pushed sideways and forward by the wind, their paths longer, driftier, and more unstable; bubbles of different sizes drift past in layered depth, some grazing the girl's cheek, others rising above her head. The 15-second mini-story is complete: opening three seconds in close-up, wind lifting her bangs and skirt as she opens the bottle, dips the wand, and lifts it, the camera clearly showing the full iridescent film. Then she raises the wand and blows the first big bubble and a few small ones; as they leave the wand the wind carries them upward, sea sparkles and distant wave lines reflected on their surfaces, and she laughs with widened surprised eyes. Mid-section she reaches for the largest bubble ahead, misses, laughs, turns to watch a lower cluster float past her waist, then half-turns to chase again. The climax from 10-13 seconds: a larger bubble is blown back near her face, she touches it, and it bursts at her fingertip into tiny water droplets and a brief flash of membrane light; she pauses, then breaks into happy laughter. The environment emphasizes the seaside wind and transparent air: sparkling sea, distant low waves, wooden boardwalk railings, wind-blown beach grass, and occasional seagulls. Sunlight creates strong iridescent highlights on the bubbles, wind continuously redirects their paths and lifts her hair, skirt, and wand, the air salty and bright with moisture. The camera language is designed for vertical 15 seconds: opening with close-ups of the film and pursed lips to grab attention, mid-section with medium and full-body follow shots of her chasing bubbles sideways, using the upper half of the vertical frame for floating bubbles and the lower half for running, inserting a brief slow-motion of the fingertip pop, and ending on an upward tilt that gathers sea breeze, bubbles, and the child's smiling face. Bright, rhythmic, and far more fluid than standing still to blow bubbles.

Demo 1 of 5 · Click the video to view its full prompt and settings.

Compare models

Which Wan AI model should you use?

Choose by workflow, source material, and final delivery needs. This table shows where each version fits best before you start creating.

Compare
3.0Current
2.7
2.6
2.5
Best fitMaximum Wan controlVersatile creation and editingReference-led storytellingAccessible audio-first creation
InputsText, image, video, audioText, image, reference video, source videoText, image, up to 2 reference videosText, image, optional audio
Upload resourcesUp to 20 mixed references5 references · 3 edit images1 image or 1–2 videos1 image + optional audio
AudioNative synchronized audioGenerated or uploaded audioNative dialogue and soundUpload or generate audio
Output2–30s · 480p–1080p5–15s · 720p–1080p5–15s · 720p–1080pUp to 10s · 480p–1080p
Key advantageLong clips + multimodal directionVideo Edit workflowMulti-shot character consistencySimple one-pass lip sync
Model pageTry 3.0Try 2.7Try 2.6Try 2.5

Available settings can vary by generation mode. Open a model page to review its current controls before creating.

Model collection

Explore every Wan AI model

Each version is tuned for a different creative job. Review its practical advantages, supported inputs, and output options, then try the model that fits your project.

3.0
Model 01 · Wan AI archive
Current model

Wan 3.0

Best for

Longer campaign clips, reference-led productions, and complete audio-visual scenes.

The most flexible Wan workflow, combining long-form generation, native synchronized audio, and multimodal reference direction in one creator.

Try Wan 3.0
Clip length
2–30 seconds
Output
480p–1080p
Inputs
Text · image · video · audio
References
Up to 20 files

Upload resources

Reference mode accepts up to 10 images, 5 videos, and 5 audio files. Image mode also supports first and optional last frames.

Build longer complete scenes

Generate clips up to 30 seconds so an idea has room for action, pacing, and a clear finish.

Direct with mixed references

Combine visual, motion, and audio sources to define subjects, style, timing, and sound.

Generate picture and sound together

Native synchronized audio reduces the need to assemble dialogue, ambience, and visuals afterward.

2.7
Model 02 · Wan AI archive

Wan 2.7

Best for

Creative iteration, video restyling, reference-guided motion, and social production.

A versatile creation and editing model with text, image, reference-video, and video-edit workflows for controlled cinematic output.

Try Wan 2.7
Clip length
5–15 seconds
Output
720p–1080p
Modes
Text · image · reference · edit
References
Up to 5 media files

Upload resources

Reference mode supports up to 5 videos and images in total. Video Edit accepts one source video plus up to 3 reference images; optional audio is available in supported modes.

Move from generation to editing

Create a new scene or reshape existing footage without leaving the same model workflow.

Guide identity and movement

Reference media gives the model clearer signals for subjects, motion, and visual direction.

Choose speed or delivery quality

Work in 720p for iteration or generate 1080p output when the result is ready to share.

2.6
Model 03 · Wan AI archive

Wan 2.6

Best for

Character stories, multi-person dialogue, product narratives, and reference video control.

A production-oriented model for multi-shot storytelling, character reference, and synchronized dialogue across text- and image-led scenes.

Try Wan 2.6
Clip length
5–15 seconds
Output
720p–1080p
Inputs
Text · image · reference video
References
Up to 2 videos

Upload resources

Upload a source image for image-to-video or one to two MP4/MOV reference videos up to 30MB each. Reference generations support 5- or 10-second clips.

Tell a story across shots

Multi-shot generation connects actions and scenes with a clearer narrative rhythm.

Carry characters across scenes

Reference videos help preserve recognizable appearance, behavior, and voice cues.

Handle dialogue natively

Audio-visual synchronization supports expressive speech and multi-person interactions.

2.5
Model 04 · Wan AI archive

Wan 2.5

Best for

Short ads, multilingual talking scenes, product clips, and fast social content.

A straightforward audio-synced generator for turning a prompt or image into polished short videos in common social and campaign formats.

Try Wan 2.5
Clip length
Up to 10 seconds
Output
480p–1080p
Inputs
Text · image · audio
Formats
Multiple aspect ratios

Upload resources

Start with a prompt or a source image, then optionally upload voice, music, or sound effects to guide pacing and lip sync.

Create audio and video in one pass

Generate a complete short clip with dialogue, sound effects, and visuals already aligned.

Support multilingual speech

Build voice-led content for different audiences without a separate dubbing workflow.

Deliver for different channels

Choose common resolutions and aspect ratios for social, marketing, and presentation use.