
Conceptual collage of a prompt turning into a short video frame with audio waveform
Text-to-video is usually the first thing people try when a new model drops. You type a scene and hope for a clip. FLUX 3 from Black Forest Labs includes that path, and the clip can ship with native audio in the same pass.
Note: FLUX 3 Video is in gated early access from BFL as of July 2026. It is coming soon on TheFluxTrain. This guide follows the public announcement, not general product access.
Quick answer: FLUX 3 text-to-video generates a clip up to 20 seconds with native audio from a text prompt. Early access is gated through Black Forest Labs. Image generation, public pricing, and open-weight FLUX 3 Dev are still on the roadmap. Track TheFluxTrain support on the FLUX 3 model page.
For the full release map, read FLUX 3: Image, Video, Editing Features and Release Status.
FLUX 3 is a multimodal model. Text-to-video is one input mode: describe a scene, get motion plus sound together.
A lot of older stacks generate silent video and add audio later. BFL's pitch is that dialogue and impacts can stay timed because audio is part of the same generation.
Text-to-video is also the baseline people use for comparisons. BFL's early preference numbers used 10-second, 720p text-to-video clips with audio. Treat those as maker-reported signals until independent tests land.
From the announced capabilities, a useful prompt usually covers:
A barista pours latte art in a small Tokyo cafe at golden hour. Camera slow push-in over the counter. Soft jazz in the background. She says in Japanese that the drink is ready. 8 seconds. No cutaways.
You will refine this once you have real access and see how the model weights speech versus ambience.
Text-to-video sits next to:
If you need a locked look before motion, start with a still once FLUX 3 image generation opens, then animate it.
Time: Prompt writing is minutes. Generation time and cost are unpublished for general users.
When FLUX 3 is available here, text-to-video should fit the same pattern as other video models: prompt in, clip out, then optional editor or multi-shot assembly. Until then, use the video paths already on the platform and watch the FLUX 3 model page for status.
Yes. BFL lists text-to-video among FLUX 3 Video capabilities, with native audio in the same generation.
Up to 20 seconds in one pass. Longer sequences need chained clips. See the video length guide.
BFL says native audio is part of FLUX 3 Video output, including multilingual dialogue and event-tied sounds.
Public pricing has not been announced. Early access is gated.
No. It is coming soon.
BFL published early preference margins on short text-to-video clips. Independent reviews will matter more once access expands. Start with the release overview for context.
FLUX 3 also supports image-to-video and visual references. If your job starts from a still, use the image-to-video guide.