
Conceptual collage of a prompt turning into a short video frame with audio waveform
Text-to-video is usually the first thing people try when a new model drops. You type a scene and hope for a clip. FLUX 3 from Black Forest Labs includes that path, and the clip can ship with native audio in the same pass.
Note: FLUX 3 Video is live on TheFluxTrain. Clips are 5–20 seconds, 720p or 1080p, with optional native audio. Cheaper 720p drafts can be enhanced to 1080p.
Quick answer: FLUX 3 text-to-video generates a clip up to 20 seconds with native audio from a text prompt. On TheFluxTrain you run it from Video from Text or the FLUX 3 model page. Image generation, BFL's own public API list, and open-weight FLUX 3 Dev are still on the roadmap.
For the full release map, read FLUX 3: Image, Video, Editing Features and Release Status.
FLUX 3 is a multimodal model. Text-to-video is one input mode: describe a scene, get motion plus sound together.
A lot of older stacks generate silent video and add audio later. BFL's pitch is that dialogue and impacts can stay timed because audio is part of the same generation.
Text-to-video is also the baseline people use for comparisons. BFL's early preference numbers used 10-second, 720p text-to-video clips with audio. Treat those as maker-reported signals until independent tests land.
From the announced capabilities, a useful prompt usually covers:
A barista pours latte art in a small Tokyo cafe at golden hour. Camera slow push-in over the counter. Soft jazz in the background. She says in Japanese that the drink is ready. 8 seconds. No cutaways.
You will refine this once you see how the model weights speech versus ambience.
Text-to-video sits next to:
If you need a locked look before motion, start with a still once FLUX 3 image generation opens, then animate it.
Time: Prompt writing is minutes. Generation time depends on the model and queue.
You can also start from the FLUX 3 model page.
Yes. BFL lists text-to-video among FLUX 3 Video capabilities, with native audio in the same generation.
Up to 20 seconds in one pass. Longer sequences need chained clips. See the video length guide.
BFL says native audio is part of FLUX 3 Video output, including multilingual dialogue and event-tied sounds.
No. On TheFluxTrain it uses credits by duration and resolution. See the Rate Card.
Yes. Use Video from Text or the FLUX 3 model page.
BFL published early preference margins on short text-to-video clips. Independent reviews will matter more once access expands. Start with the release overview for context.
FLUX 3 also supports image-to-video and visual references. If your job starts from a still, use the image-to-video guide.