Click or Drag-n-Drop
PNG, JPG or GIF, Up-to 2048 x 2048 px
Click or Drag-n-Drop
PNG, JPG or GIF, Up-to 2048 x 2048 px
Click or Drag-n-Drop
PNG, JPG or GIF, Up-to 2048 x 2048 px
Click or Drag-n-Drop
PNG, JPG or GIF, Up-to 2048 x 2048 px
Click or Drag-n-Drop
PNG, JPG or GIF, Up-to 2048 x 2048 px
Generates a synced soundtrack in the same pass. No price change.
Wan 3.0 Video is Alibaba's all-in-one video generation model, the newest release in the Wan family from Tongyi Lab. Where earlier versions split the work across separate models, Wan 3.0 folds text-to-video, image-to-video, reference-to-video, and video editing into a single endpoint. You give it a prompt, a first and last frame, reference images, reference clips, reference audio, or even a document, and the model infers what you are asking for and renders it.
The headline change is length: Wan 3.0 generates up to 30 seconds of video in a single pass at 30fps, double its predecessor. That is long enough for continuous camera moves and one-take shot language instead of the short fragments most models cap out at. It also generates a native audio track alongside the video, in the same pass, with multilingual voice output.
-1 for smart duration that matches your prompt and source.Wan 3.0 fits filmmaking and short drama, social media and advertising, e-commerce product demos, character animation, and document-to-video workflows. The 30-second window lets a single generation hold a complete exchange, camera move, or transition without stitching. Reference-to-video keeps characters, props, spaces, and style consistent across shots, which suits branded content and recurring characters. Document input turns a slide deck, spreadsheet, or PDF into an explainer or promo in one upload, making it a practical tool for marketing and educational video.
Write cinematic, specific prompts: name the subject, the motion, the camera move, the lighting, and the ambient sound you want in the audio track. Keep prompt_extend on for short prompts so the model enriches detail automatically. Use first-and-last-frame control when you need an exact start and finish, and reference images for consistent identity. Output realism is strong on faces, micro-expressions, and reference fidelity; audio texture and on-screen text rendering are the areas Alibaba flags as still improving, so verify any in-frame typography.
Can Wan 3.0 make videos longer than 30 seconds? A single pass tops out at 30 seconds. Longer narratives are built with the video extension capability rather than one call.
Does Wan 3.0 generate audio? Yes. Audio is on by default and produced in the same pass as the video, with multilingual voice output. Toggling audio does not change pricing.
What resolutions and aspect ratios are supported? 480P, 720P, and 1080P, with 16:9, 4:3, 1:1, 3:4, 9:16, and an adaptive ratio that follows your input.
What inputs can it take? Text, images (first frame, last frame, up to 10 reference images), reference video and audio clips, and documents such as PPT, PDF, DOC, and XLS.
Is there a 4K tier or open weights? No. Output caps at 1080P, and Alibaba has not published open weights for Wan 3.0; earlier Wan releases were the open ones.
How do I keep a character consistent across a shot? Supply reference images and, where possible, first-and-last-frame control, then check frame one against the final frame for identity drift.