Alibaba's Tongyi Wanxiang Wan 3.0, on OVOV. Start from a prompt, from a first and last frame, or from a prompt backed by your own reference images, videos and audio — at 480P, 720P or 1080P, 5 to 30 seconds, with a native soundtrack.
Real generations from our video models — Veo 3.1, Kling and Seedance. Hover any clip to play, click the speaker icon to hear the audio.
Hover to play · Click speaker to unmute
Real generations from our video models — Veo 3.1, Kling and Seedance. Hover any clip to play, click the speaker icon to hear the audio.
Alien Fantasy World
Luxury Perfume Ad
Product Launch Teaser
Maldives Travel Commercial
Eastern Wuxia Duel
Hollywood Racing Movie
Spaceship to the Stars
Golden Retriever & Butterflies
Sushi Chef at Work
Red Dress on the Cliffs
Midnight Duel
Handheld Selfie Vlog
Mouth-watering Food Close-up
Urban Street Dance
Luxury Lipstick Ad
One model that takes text, keyframes or your own footage, renders up to 1080P, runs to half a minute, and writes its own soundtrack while it works.
In text-to-video you can attach up to 10 reference images, up to 5 reference videos (15 seconds total) and up to 5 reference audio clips (15 seconds total). Wan 3.0 pulls character, style, motion and tone from your own material instead of guessing.
Give it a starting image, and optionally an ending image, and Wan 3.0 renders the motion between them. The straightest route to product reveals, transformations and before-and-after arcs.
Three resolution tiers, including a true 1080P output. Rough the idea out at 480P for a few credits, then render the version you actually ship at 1080P.
Six fixed lengths: 5, 10, 15, 20, 25 or 30 seconds. A five-second test costs pocket change; a full 30-second spot fits inside one generation.
Wan 3.0 writes ambience, effects and motion-matched sound alongside the picture. The switch is yours to flip, and turning audio off does not change the price.
16:9, 9:16, 1:1, 4:3 and 3:4. Credits are charged per second — 3 at 480P, 6 at 720P, 12 at 1080P — and a failed generation is refunded in full.
Pick the input that matches what you already have. Wan 3.0 handles picture and sound in the same pass, so there is nothing to assemble afterwards.
Write a prompt on its own, upload a first frame (and optionally a last frame), or attach reference images, videos and audio alongside your prompt. Then set length, aspect ratio and resolution.
The model plans motion, lighting and pacing across the whole duration, then renders the video with its matching audio track. Most jobs land in one to five minutes; longer clips take longer.
Take the finished video with audio baked in — no watermark, commercial rights included. If a generation fails, your credits come straight back.
Three input modes and a 1080P tier cover a lot of ground: finished ads, vertical social cuts, continuity work built on your own footage, and long atmospheric films.
A 30-second commercial at full HD, rendered in one pass. Setup, demonstration, benefit and sign-off fit inside a single generation, sound included.
TikTok, Reels and Shorts in 9:16, from a 5-second hook to a 30-second story. Audio arrives attached, so nothing needs a soundtrack pass afterwards.
Attach the same reference images, video and audio to every generation and the character, product and tone carry from clip to clip instead of drifting.
Aerial sweeps, coastal drives and street-level wandering that need time to breathe — plus the wind, water and city sound that sells the place.
Wan 3.0 accepts prompts up to 20,000 characters, which is far more room than most clips need. These habits spend that room on the things the model actually acts on.
Name your materials, then describe the finished shot
The three modes take different prompts. Text-to-video has to carry the whole scene in words. First-and-last-frame already has the composition, so the prompt is about movement. Reference mode is about what to borrow from the material you attached.
Example
"First frame: a sealed cardboard box on a table. Last frame: the box open with a camera inside. Prompt: "hands lift the flaps in one smooth motion, soft window light"."
In reference mode you can point at the material directly — image 1, video 1, audio 1 — so the model knows which asset governs which part of the shot instead of blending everything together.
Example
"Keep the woman from image 1 and the jacket from image 2, follow the slow dolly-in of video 1, and match the room tone of audio 1."
Thirty seconds is a lot of screen time. Write the progression in the order it should happen, or the model will hold a single idea for half a minute.
Example
"The room starts dim and empty, morning light climbs the far wall, dust drifts through the beam, and the camera slowly settles on a cup of coffee going cold."
Audio is generated by default, so anything you leave unsaid gets invented for you. One or two lines of sound direction usually keeps it in character.
Example
"Sound: steady rain on a tin roof, an occasional low roll of thunder in the distance, no music."
Stunning, epic and cinematic tell the model nothing. Materials, colours, light direction and strong verbs do.
Example
"Avoid: "a stunning epic mountain shot". Prefer: "granite ridges under low side light, snow spilling off the crest in a thin plume"."
Wan 3.0 offers 5, 10, 15, 20, 25 and 30 seconds, five aspect ratios and three resolutions. Deciding before you generate saves credits and keeps the composition right.
Example
"A 15-second 9:16 vertical clip: a barista pulling an espresso shot, steam rising into warm café light, shallow depth of field."