Powered by Alibaba Tongyi Wanxiang

Wan 3.0 · One Video Model, Three Ways In

Alibaba's Tongyi Wanxiang Wan 3.0, on OVOV. Start from a prompt, from a first and last frame, or from a prompt backed by your own reference images, videos and audio — at 480P, 720P or 1080P, 5 to 30 seconds, with a native soundtrack.

Duration
Aspect
Resolution

AI video examples on OVOV

Real generations from our video models — Veo 3.1, Kling and Seedance. Hover any clip to play, click the speaker icon to hear the audio.

Hover to play · Click speaker to unmute

AI video examples on OVOV

Real generations from our video models — Veo 3.1, Kling and Seedance. Hover any clip to play, click the speaker icon to hear the audio.

Cinematic

Alien Fantasy World

Fragrance

Luxury Perfume Ad

Product

Product Launch Teaser

Travel

Maldives Travel Commercial

Action

Eastern Wuxia Duel

Film

Hollywood Racing Movie

Sci-Fi

Spaceship to the Stars

Lifestyle

Golden Retriever & Butterflies

Food

Sushi Chef at Work

Fashion

Red Dress on the Cliffs

Game Trailer

Midnight Duel

Vlog

Handheld Selfie Vlog

Food

Mouth-watering Food Close-up

Dance

Urban Street Dance

Makeup

Luxury Lipstick Ad

Why Wan 3.0 is worth a look

One model that takes text, keyframes or your own footage, renders up to 1080P, runs to half a minute, and writes its own soundtrack while it works.

Multimodal References

In text-to-video you can attach up to 10 reference images, up to 5 reference videos (15 seconds total) and up to 5 reference audio clips (15 seconds total). Wan 3.0 pulls character, style, motion and tone from your own material instead of guessing.

First & Last Frame Control

Give it a starting image, and optionally an ending image, and Wan 3.0 renders the motion between them. The straightest route to product reveals, transformations and before-and-after arcs.

480P, 720P and 1080P

Three resolution tiers, including a true 1080P output. Rough the idea out at 480P for a few credits, then render the version you actually ship at 1080P.

5 to 30 Seconds

Six fixed lengths: 5, 10, 15, 20, 25 or 30 seconds. A five-second test costs pocket change; a full 30-second spot fits inside one generation.

Native Audio, On by Default

Wan 3.0 writes ambience, effects and motion-matched sound alongside the picture. The switch is yours to flip, and turning audio off does not change the price.

Five Aspect Ratios, Priced by the Second

16:9, 9:16, 1:1, 4:3 and 3:4. Credits are charged per second — 3 at 480P, 6 at 720P, 12 at 1080P — and a failed generation is refunded in full.

From prompt, keyframes or footage to a finished clip

Pick the input that matches what you already have. Wan 3.0 handles picture and sound in the same pass, so there is nothing to assemble afterwards.

1

Choose your way in

Write a prompt on its own, upload a first frame (and optionally a last frame), or attach reference images, videos and audio alongside your prompt. Then set length, aspect ratio and resolution.

2

Wan 3.0 renders picture and sound

The model plans motion, lighting and pacing across the whole duration, then renders the video with its matching audio track. Most jobs land in one to five minutes; longer clips take longer.

3

Download and use it

Take the finished video with audio baked in — no watermark, commercial rights included. If a generation fails, your credits come straight back.

What Wan 3.0 is good for

Three input modes and a 1080P tier cover a lot of ground: finished ads, vertical social cuts, continuity work built on your own footage, and long atmospheric films.

1080P Product & Brand Spots

A 30-second commercial at full HD, rendered in one pass. Setup, demonstration, benefit and sign-off fit inside a single generation, sound included.

  • 1080P output for pre-roll, landing pages and client delivery
  • 16:9 for widescreen, 4:3 for classic broadcast framing
  • Product sound and ambience generated with the picture

Vertical Social Shorts

TikTok, Reels and Shorts in 9:16, from a 5-second hook to a 30-second story. Audio arrives attached, so nothing needs a soundtrack pass afterwards.

  • 9:16 vertical, or 1:1 for square feed posts
  • Draft at 480P for 3 credits a second, publish at 720P or 1080P
  • Native audio instead of a licensed music track

Consistency Across a Series

Attach the same reference images, video and audio to every generation and the character, product and tone carry from clip to clip instead of drifting.

  • Up to 10 reference images to hold character and style
  • Up to 5 reference videos for motion and camera behaviour
  • Up to 5 reference audio clips to keep the sound in family

Travel & Atmosphere Films

Aerial sweeps, coastal drives and street-level wandering that need time to breathe — plus the wind, water and city sound that sells the place.

  • Long drifting camera moves that need the full 30 seconds
  • Environmental audio rendered alongside the visuals
  • 1080P for hero videos that sit at the top of a page

Prompt writing guide for Wan 3.0

Wan 3.0 accepts prompts up to 20,000 characters, which is far more room than most clips need. These habits spend that room on the things the model actually acts on.

Name your materials, then describe the finished shot

Pick the input mode before you write

The three modes take different prompts. Text-to-video has to carry the whole scene in words. First-and-last-frame already has the composition, so the prompt is about movement. Reference mode is about what to borrow from the material you attached.

  • Text only: describe setting, subject, camera and sound from scratch
  • First and last frame: describe the motion between them, not the frames themselves
  • References (attached in text-to-video): say what each asset contributes — face, outfit, camera move, tone

Example

"First frame: a sealed cardboard box on a table. Last frame: the box open with a camera inside. Prompt: "hands lift the flaps in one smooth motion, soft window light"."

Address your references by number

In reference mode you can point at the material directly — image 1, video 1, audio 1 — so the model knows which asset governs which part of the shot instead of blending everything together.

  • Up to 10 reference images, 5 reference videos and 5 reference audio clips
  • Reference video totals 15 seconds; reference audio totals 15 seconds
  • Give each asset one job: identity from one image, camera move from one video

Example

"Keep the woman from image 1 and the jacket from image 2, follow the slow dolly-in of video 1, and match the room tone of audio 1."

Plan the arc across the length you chose

Thirty seconds is a lot of screen time. Write the progression in the order it should happen, or the model will hold a single idea for half a minute.

  • Open with the setting, bring in the subject, then the movement
  • Name what changes over time: light, weather, distance, speed
  • Shorter lengths (5-10s) suit one continuous action

Example

"The room starts dim and empty, morning light climbs the far wall, dust drifts through the beam, and the camera slowly settles on a cup of coffee going cold."

Write the sound, not just the picture

Audio is generated by default, so anything you leave unsaid gets invented for you. One or two lines of sound direction usually keeps it in character.

  • Name the ambience: rain on glass, distant traffic, room tone
  • Name the sounds that match the motion: footsteps, engine, wind
  • Asking for quiet is valid direction — say so if you want it sparse
  • Attach reference audio when you already have a tone in mind

Example

"Sound: steady rain on a tin roof, an occasional low roll of thunder in the distance, no music."

Trade praise words for visual facts

Stunning, epic and cinematic tell the model nothing. Materials, colours, light direction and strong verbs do.

  • Skip: stunning, epic, breathtaking, gorgeous, masterpiece
  • Use verbs with force: slams, drifts, ignites, whips
  • Name the light source and the time of day

Example

"Avoid: "a stunning epic mountain shot". Prefer: "granite ridges under low side light, snow spilling off the crest in a thin plume"."

Choose length, ratio and resolution upfront

Wan 3.0 offers 5, 10, 15, 20, 25 and 30 seconds, five aspect ratios and three resolutions. Deciding before you generate saves credits and keeps the composition right.

  • 16:9 and 4:3 for landscape, 9:16 and 3:4 for vertical, 1:1 for feeds
  • Draft at 480P (3 credits per second), finish at 720P (6) or 1080P (12)
  • Write for the frame you picked — vertical needs the subject close and centred
  • Length is fixed to 5, 10, 15, 20, 25 or 30 seconds — round your storyboard to one of them

Example

"A 15-second 9:16 vertical clip: a barista pulling an espresso shot, steam rising into warm café light, shallow depth of field."

Frequently asked questions









Start Your AI Creative Journey Today

Join OVOV to generate images, videos, music, and voice with AI—unleash limitless creativity.
Sign up now and get 10 free credits instantly. No waiting, start creating right away.