TheFluxTrain
Guide·

FLUX 3 Text to Video: What It Can Do and What Is Still Early Access

FLUX 3 text-to-video can generate up to 20 seconds with native audio from a prompt. See what BFL announced and what is still gated early access.
Conceptual collage of a prompt turning into a short video frame with audio waveform

Conceptual collage of a prompt turning into a short video frame with audio waveform

Text-to-video is usually the first thing people try when a new model drops. You type a scene and hope for a clip. FLUX 3 from Black Forest Labs includes that path, and the clip can ship with native audio in the same pass.

Note: FLUX 3 Video is in gated early access from BFL as of July 2026. It is coming soon on TheFluxTrain. This guide follows the public announcement, not general product access.

Quick answer: FLUX 3 text-to-video generates a clip up to 20 seconds with native audio from a text prompt. Early access is gated through Black Forest Labs. Image generation, public pricing, and open-weight FLUX 3 Dev are still on the roadmap. Track TheFluxTrain support on the FLUX 3 model page.

For the full release map, read FLUX 3: Image, Video, Editing Features and Release Status.

What is FLUX 3 text-to-video?

FLUX 3 is a multimodal model. Text-to-video is one input mode: describe a scene, get motion plus sound together.

A lot of older stacks generate silent video and add audio later. BFL's pitch is that dialogue and impacts can stay timed because audio is part of the same generation.

Text-to-video is also the baseline people use for comparisons. BFL's early preference numbers used 10-second, 720p text-to-video clips with audio. Treat those as maker-reported signals until independent tests land.

What can you ask for in a FLUX 3 text-to-video prompt?

From the announced capabilities, a useful prompt usually covers:

  • Subject and action: who or what moves, and what they do
  • Camera: static, push-in, pan, or cut language if you plan multi-shot later
  • Setting and light: place, time of day, atmosphere
  • Audio intent: dialogue language, ambient sound, or event sounds
  • Duration feel: keep it inside the 20-second one-pass limit

Example

A barista pours latte art in a small Tokyo cafe at golden hour. Camera slow push-in over the counter. Soft jazz in the background. She says in Japanese that the drink is ready. 8 seconds. No cutaways.

You will refine this once you have real access and see how the model weights speech versus ambience.

How does FLUX 3 text-to-video fit the wider model?

Text-to-video sits next to:

If you need a locked look before motion, start with a still once FLUX 3 image generation opens, then animate it.

What do you need before you try FLUX 3 text-to-video?

  • Access through BFL's gated early-access program, or a later public API
  • A clear scene brief, not only mood words
  • A plan for clip length at or under 20 seconds
  • Reference notes if you will chain more clips later

Time: Prompt writing is minutes. Generation time and cost are unpublished for general users.

What are the limits right now?

  • Early access only; not generally available
  • No public price or rate limits
  • Longer stories need multiple clips, not one endless take
  • Preference scores vs other models are preliminary and BFL-reported
  • TheFluxTrain integration is coming soon

How will this show up on TheFluxTrain?

When FLUX 3 is available here, text-to-video should fit the same pattern as other video models: prompt in, clip out, then optional editor or multi-shot assembly. Until then, use the video paths already on the platform and watch the FLUX 3 model page for status.

Frequently asked questions

Does FLUX 3 support text-to-video?

Yes. BFL lists text-to-video among FLUX 3 Video capabilities, with native audio in the same generation.

How long can a FLUX 3 text-to-video clip be?

Up to 20 seconds in one pass. Longer sequences need chained clips. See the video length guide.

Does text-to-video include sound?

BFL says native audio is part of FLUX 3 Video output, including multilingual dialogue and event-tied sounds.

Is FLUX 3 text-to-video free?

Public pricing has not been announced. Early access is gated.

Is FLUX 3 text-to-video on TheFluxTrain yet?

No. It is coming soon.

How does it compare to Runway or Luma?

BFL published early preference margins on short text-to-video clips. Independent reviews will matter more once access expands. Start with the release overview for context.

Can I use image references in text-to-video?

FLUX 3 also supports image-to-video and visual references. If your job starts from a still, use the image-to-video guide.

Where is the official announcement?

Black Forest Labs' FLUX 3 post.