Fast low-detail preview rendered at hd. Enable to iterate quickly, disable for final quality.
Generates synchronized speech, effects and ambience. Keep true for sound, false for a silent clip.
FLUX 3 Text to Video is Black Forest Labs' first video model, turning a single text prompt into a cinematic clip with native, synchronized audio. Instead of stitching sound on in a second pass, FLUX 3 is one multimodal model trained jointly on image, video, and audio, so dialogue, sound effects, and ambience are generated in the same pass as the picture. From plain text you get clips up to 20 seconds at HD or Full HD, 24 fps, across cinematic and social aspect ratios, no reference image or footage required.
Use FLUX 3 Text to Video for social ads, product teasers, cinematic B-roll, explainer clips, music-video moments, and dialogue scenes that need talking characters with lip-sync. In testing, a golden-hour coastal prompt with a short spoken line produced a photoreal character, a smooth push-in camera move, and layered wave, wind, and voice audio, reliably across repeated runs.
Write the shot like a director: name the subject, camera move, lighting, dialogue in quotes, and the sounds you want. Keep spoken lines short for 5-second clips, and pick 16:9 for cinematic framing or 9:16 for verticals. Turn on audio for talking scenes; enable draft to iterate quickly, then render the keeper at full quality.
Does FLUX 3 Text to Video generate sound? Yes. Audio is on by default, including speech, effects, and ambience synced to the picture.
How long can clips be? Up to 20 seconds per generation, from 5 seconds up.
Does it do lip-sync and other languages? Yes, it renders multilingual speech with strong lip-sync.
What resolutions are supported? HD and Full HD, at 24 fps.
Do I need an input image? No. This is pure text-to-video; just describe the scene.
Can I preview cheaply first? Yes, use draft mode for a fast low-detail preview before a full render.