Black Forest Labs — the team behind the FLUX image models — built a video model. On OVOV it renders 5 to 20 seconds at HD or FHD, writes its own synchronised audio, and lets you pin up to ten keyframes along the timeline instead of settling for a first and last frame.
Real generations from our video models — Veo 3.1, Kling and Seedance. Hover any clip to play, click the speaker icon to hear the audio.
Hover to play · Click speaker to unmute
Real generations from our video models — Veo 3.1, Kling and Seedance. Hover any clip to play, click the speaker icon to hear the audio.
Alien Fantasy World
Luxury Perfume Ad
Product Launch Teaser
Maldives Travel Commercial
Eastern Wuxia Duel
Hollywood Racing Movie
Spaceship to the Stars
Golden Retriever & Butterflies
Sushi Chef at Work
Red Dress on the Cliffs
Midnight Duel
Handheld Selfie Vlog
Mouth-watering Food Close-up
Urban Street Dance
Luxury Lipstick Ad
Up to ten ordered keyframes, a full HD tier, twenty seconds of runway and a soundtrack written in the same pass — from the team that made the FLUX image models.
Upload one image and it becomes the starting frame. Upload two and they become the first and last frame. Upload up to ten and the extra frames are distributed evenly along the timeline, so you are choreographing the whole clip rather than only its ends.
FLUX 3 Video writes ambience, effects and motion-matched sound alongside the picture, synchronised to what is on screen. You can switch audio off if you plan to score the clip yourself — the price stays the same either way.
Two resolution tiers: HD at roughly 1280×704 and FHD at roughly 1920×1088. Rough the idea out at HD, then spend FHD credits once the prompt and the keyframes are settled.
Four fixed lengths: 5, 10, 15 or 20 seconds. Most of our video models cap out at 15, so the top tier here leaves room for a setup, a turn and a payoff inside one generation.
16:9, 9:16, 1:1, 4:3, 3:4, plus 21:9 and 2:1 for cinematic openers and letterboxed title cards. Set the frame before you generate — composition is planned for the ratio you pick.
Black Forest Labs built the FLUX image models that already run on OVOV. FLUX 3 Video is the same lab moving into motion, so the prompt habits that work on their stills carry over.
Write the shot, optionally pin the frames it has to hit, and FLUX 3 Video renders picture and sound together — nothing to assemble afterwards.
Start from a prompt on its own, or upload between one and ten keyframe images in the order they should appear. Then choose length (5, 10, 15 or 20 seconds), aspect ratio and HD or FHD.
The model plans motion, lighting and pacing across the whole duration — passing through your keyframes in order if you supplied them — then renders the video with its matching audio track.
Take the finished video with audio baked in — no watermark, commercial rights included. If a generation fails or is rejected by content moderation, your credits come straight back in full.
Ordered keyframes, an FHD tier and two ultrawide ratios cover the work where the shot has to land on specific images at specific moments — and still sound like something.
When the clip has to pass through images you already have — a logo, a product state, a character pose — pin them as keyframes in order and let the model fill in the motion between them.
A brand spot at roughly 1920×1088, or a 21:9 title sequence that reads as film rather than feed. Sound arrives attached, so there is no scoring pass before you can show it to anyone.
TikTok, Reels and Shorts in 9:16 or 3:4, from a five-second hook to a twenty-second story, with native audio instead of a licensed music bed.
Coastal drives, rain on a window, a market waking up — shots that need time to breathe, plus the wind, water and street noise that makes the place believable.
The model reads your prompt and your keyframes together. These habits keep the two pulling in the same direction instead of arguing.
Pin the frames that matter, describe the motion between them
Text on its own gives the model the whole scene to invent. Keyframes take that freedom away image by image. Neither is better — they just need different prompts.
Example
"With two keyframes (closed box, open box): "hands lift the flaps in one smooth motion, soft window light, cardboard scrape and a soft thud"."
Count matters as much as content. One image is a starting point. Two are the ends of a journey. Anything more is a schedule the model has to keep, spaced evenly across the length you chose.
Example
"Four keyframes across 20 seconds land at roughly 0s, 6.7s, 13.3s and 20s — so pick images that make sense five to seven seconds apart."
Twenty seconds is a long time to hold one idea. Write the progression in the order it should happen, and the model has something to spend the duration on.
Example
"The room starts dim and empty, morning light climbs the far wall, dust drifts through the beam, and the camera settles on a cup of coffee going cold."
Audio is generated by default and synchronised to the action, so anything you leave unsaid gets invented for you. One or two lines of sound direction usually keeps it in character.
Example
"Sound: steady rain on a tin roof, a low roll of thunder in the distance, no music."
Stunning, epic and cinematic tell the model nothing. Materials, colours, light direction and strong verbs do — the same discipline that works on the FLUX image models.
Example
"Avoid: "a stunning epic mountain shot". Prefer: "granite ridges under low side light, snow spilling off the crest in a thin plume"."
FLUX 3 Video offers 5, 10, 15 and 20 seconds, seven aspect ratios and two resolution tiers. Deciding before you generate saves credits and keeps the composition right.
Example
"A 10-second 21:9 opener: headlights sweeping across a wet runway at dusk, low camera, tyre spray catching the light."