Pick Text or Image
Choose Text to Video to start from a prompt, or Image to Video to upload a first frame, a first and last frame, or reference images.
One continuous take up to 30 seconds, with sound generated in the same pass
Wan 3.0 is the latest video model from Alibaba's Tongyi Wan team. Start from a text prompt, a first frame (plus an optional last frame), or up to nine reference images, then choose 2–30 seconds at 480p, 720p or 1080p. Audio is generated with the picture, and the credit cost is shown before you submit.
Create a Wan 3.0 Video in 4 Steps
Choose Text to Video to start from a prompt, or Image to Video to upload a first frame, a first and last frame, or reference images.
Describe the subject, action, camera movement, setting, mood and any dialogue or sound you want to hear.
Choose 2–30 seconds, 480p/720p/1080p and an aspect ratio. The credit estimate updates as you change each setting.
Submit, wait a few minutes while Wan 3.0 renders, then watch the full clip before downloading or refining the prompt.
Longer Takes, Native Sound and Three Ways to Start
Wan 3.0 renders a single continuous clip of 2 to 30 seconds — twice the 15-second ceiling of Wan 2.7 — so a scene can build, turn and resolve without stitching short fragments together.
Dialogue, ambience and sound effects are generated together with the picture, so clips come back scored instead of silent. Audio is on by default and switching it off does not change the credit cost.
Write a prompt from scratch, animate a single image as the opening frame, or upload two images to lock both the first and the last frame and let Wan 3.0 fill in the motion between them.
Switch Image Use to Reference Images to guide characters, products, outfits and locations with up to nine images, and point to them in the prompt as @image1, @image2 and so on.
Alibaba built Wan 3.0 around realism: more expressive faces and micro-expressions, steadier subjects across the take and cleaner on-screen text than earlier Wan versions.
Render at 480p for quick drafts, 720p for everyday content or 1080p for final delivery, in 16:9, 9:16, 1:1, 4:3, 3:4 or Adaptive framing. Output is a 30 fps MP4.
How to Make a Wan 3.0 Video Online
The Wan 3.0 generator on this page runs Alibaba's wan3.0-video model in your browser — nothing to install, no API key or cloud console required. Follow these steps to go from an idea or a still image to a finished clip.
Use the playground at the top of this page; Wan 3.0 is already selected. Sign in when you are ready to generate — Wan 3.0 runs on a paid plan or purchased credits.
Stay on Text to Video for prompt-only generation. For Image to Video, set Image Use to First / Last Frame (image 1 opens the clip, an optional image 2 closes it) or Reference Images (up to nine).
Write the scene as a short shot list: who is in frame, what they do, how the camera moves, and what we hear. With reference images, name them as @image1, @image2 so the model knows which is which.
Pick 2–30 seconds, then 480p, 720p or 1080p. Use 9:16 for Reels and TikTok, 16:9 for YouTube, 1:1 for feeds, or Adaptive to let the model frame the shot.
Submit and let Wan 3.0 render — longer and higher-resolution clips take longer. Check faces, hands, text and audio sync, then download or adjust one direction at a time.
480p costs about a quarter of 1080p per second. Nail the motion and timing with a short 480p draft, then re-run the winning prompt at full resolution.
A 30-second clip needs more than one beat. Describe the scene in order — opening, action, turn, ending — so the model has something to do for the full length.
When the final pose or product shot matters, upload it as the last frame. Wan 3.0 plans the motion so the clip lands on that image.
Because audio is generated with the picture, name it: a line of dialogue in quotes, rain on a window, a crowd cheering. Specific cues give more convincing sound.
Long-Take AI Video for Every Creator
Stage a complete 30-second moment with consistent characters, dialogue and sound — useful for short dramas, story teasers and character tests.
Turn product shots into moving showcases, lock the hero frame as the last frame, and check labels and geometry before the clip goes live.
Produce vertical 9:16 clips for TikTok, Reels and Shorts with sound already in place, then iterate quickly at 480p before final renders.
Visualize a concept, a process or a place as a narrated clip for lessons, onboarding and presentations.
Everything About the Wan 3.0 Video Generator
Wan 3.0 (wan3.0-video) is Alibaba Tongyi Lab's newest video model, released in public beta on August 6, 2026. It is an all-in-one model: one engine handles text-to-video, first/last-frame image-to-video and reference-guided generation, with clips up to 30 seconds and audio generated in the same pass.
You can generate from a text prompt, animate a first frame, set both a first and a last frame, or guide the video with up to nine reference images. Wan 3.0 can also read reference video, audio, documents and webpages through Alibaba's API; those inputs are not available on this page yet.
Any whole number of seconds from 2 to 30, at 480p, 720p or 1080p, delivered as a 30 fps MP4. Wan 3.0 has no 4K tier — claims of 4K output are not accurate.
Wan 3.0 is priced per second of output: about 2.74 credits per second at 480p, 5.48 at 720p and 10.96 at 1080p, rounded up per video. A 5-second clip costs 14 credits at 480p, 28 at 720p or 55 at 1080p; a 30-second 1080p clip costs 329 credits. Turning audio on or off does not change the price, and the exact cost is shown before you generate.
First / Last Frame treats your images as exact frames: image 1 opens the clip and an optional image 2 closes it. Reference Images treats up to nine images as guidance for characters, objects and places that can appear anywhere in the video. The two can't be combined in one generation.
Yes. Every paid plan includes a commercial license, so you can use your Wan 3.0 videos in ads, client projects, social media, and products you sell. The free plan is for personal use. You stay responsible for what goes in: clear any people, brands, music, or source assets in your prompts and reference images before publishing.
Yes. Audio is generated together with the picture and is on by default, so dialogue, ambience and effects arrive in sync. You can turn it off for a silent clip; the credit cost is the same either way.
No. Unlike the open-weight Wan 2.1 and 2.2 releases, Wan 3.0 is a closed model available only through cloud APIs, so there are no official weights to download. Treat any 'free Wan 3.0 download' as unofficial.
You can explore the settings and see the exact credit cost without paying. Generating with Wan 3.0 requires a paid plan or purchased credits; credits are charged when you submit and refunded automatically if a generation fails.
Most clips finish in about one to five minutes. Longer durations and 1080p take longer, so start with a short 480p draft to check the motion before rendering the final version.