TheFluxTrain
Guide·

FLUX 3: Image, Video, Editing Features and Release Status

FLUX 3 brings image generation, editing, and 20-second video with native audio into one model. See what is available now and what is still planned.
Multimodal collage of a still frame, video film strip, and audio waveform representing FLUX 3 image and video generation

Multimodal collage of a still frame, video film strip, and audio waveform representing FLUX 3 image and video generation

Black Forest Labs announced FLUX 3 on July 23, 2026. It is the first FLUX model designed to learn from images, video, and audio together, with action prediction extending the same foundation into robotics. The announcement covers a lot of ground, but only FLUX 3 Video is in gated early access right now. Image generation and editing are coming later.

Quick answer: FLUX 3 is Black Forest Labs' new multimodal model for image generation, image editing, and video with native audio. FLUX 3 Video entered gated early access on July 23, 2026. FLUX 3 Image is expected in the coming weeks, while public pricing, broad API access, and the FLUX 3 Dev release date remain unannounced.

This article is based mainly on Black Forest Labs' original FLUX 3 announcement. The model is still in early access, so it helps to separate what people can apply to use now from the rest of the roadmap.

What is FLUX 3?

FLUX 3 is a multimodal foundation model. In plain language, BFL trained one architecture across images, video, and audio instead of treating each medium as a separate problem.

An image shows where objects are at one moment. Video adds movement and time. Audio gives another clue about what caused an event, such as an impact or an engine starting. Training across those signals lets the model use one medium to check another. A ball should move according to its apparent weight, and its impact sound should arrive at the right moment.

The model builds on Self-Flow, BFL's method for aligning multimodal generation and understanding in one architecture. BFL also increased the training compute and data used across the three media types. More technical details are promised later, so there is not enough public information yet to judge the full training recipe or reproduce it.

What can FLUX 3 Video do?

FLUX 3 Video can generate a clip up to 20 seconds long with native audio in one pass. BFL lists support for:

  • text-to-video generation
  • image-to-video from a starting frame or visual reference
  • video-to-video generation using a source clip
  • continuation from existing video and audio
  • keyframe-controlled transitions
  • multilingual dialogue
  • connected clips for longer, multi-shot sequences
  • animated typography across different styles and aspect ratios

Native audio is the part worth watching. The model generates the visual sequence and its sound together, which should help dialogue and physical events line up. BFL says facial expressions, multilingual output, and sounds tied to on-screen events are current strengths.

Reference images and clips are central to the pitch too. You could carry a character from one source video into another scene, animate a still, or define key moments and ask the model to create the movement between them. BFL says visual references can help keep characters consistent when separate clips are chained into a longer sequence.

These are early-access capabilities. They do not guarantee that every prompt will stay consistent across several minutes. Long sequences still depend on multiple generated clips, careful reference management, and the tools BFL provides around the model.

How does FLUX 3 handle image generation and editing?

FLUX 3 Image is designed for synthesis and editing. BFL says it supports varied styles, aspect ratios, and resolutions, with better handling of complex prompts than earlier FLUX versions.

Text rendering is one of the more useful announced improvements. The company says FLUX 3 can render accurate text in multiple languages. If that holds up outside selected examples, it could help with posters, packaging mockups, and signs where readable words need to be part of the generated image.

There is an important availability catch. FLUX 3 Image was not released alongside the video early-access program. BFL says image early access will open in the following weeks. Until people can test it, there is no solid public answer on editing precision, maximum resolution, consistency across repeated edits, or how well typography survives local changes.

Is FLUX 3 an AI image and video editor?

Parts of FLUX 3 fit normal editing tasks, though it is too early to describe it as a complete editor.

For video, reference-based generation can carry elements from an existing clip into a new scene. Continuation extends footage, and keyframes offer control over transitions. Video-to-video generation can change context or style while preserving a central subject. BFL also plans video and audio generation and editing through APIs and private-weight access.

For images, BFL has explicitly announced synthesis and editing. The company has not published a full editing interface, a list of controls, or public hands-on results. Its post also names interactive image and video editing as part of the broader future, alongside simulation and computer use.

Editing is part of the FLUX 3 model design and roadmap. A generally available, fully specified editing product is not here yet.

When will FLUX 3 be released?

The release is happening in stages.

FLUX 3 versionRelease status on July 24, 2026
FLUX 3 VideoGated early access is open
FLUX 3 ImageEarly access expected in the following weeks
FLUX 3 ActionAccess through selected research and commercial partners
FLUX 3 DevOpen-weight release planned, with no date announced

BFL says video and audio generation and editing will come through APIs and private-weight access. Image synthesis and editing are planned for the same access routes. The company has not announced general availability dates, pricing, rate limits, model size, or hardware requirements.

If you see a page claiming that FLUX 3 is fully public, check its date and the exact product it refers to. “FLUX 3 is released” currently means the model family has been announced and selected capabilities have entered early access.

How do FLUX 3's early video comparisons look?

BFL reports that FLUX 3 was preferred over several competing video models in its preliminary comparisons. The largest stated margins were against Luma Ray 3.2, where FLUX 3 was preferred in 93% of comparisons, and Runway Gen-4.5 at 77%. Reported preference against Grok Imagine Video reached 69%, while results against other named models ranged from 52% to 60%.

Those figures need context. The evaluations used 10-second, 720p text-to-video clips with audio. BFL says the model and evaluation harness are still under development. A complete methodology and full benchmark report have not been published.

Treat the numbers as an early signal from the model maker. Independent tests will matter more once access expands, especially for prompt adherence, audio quality, editing control, consistency, generation time, and cost.

Why does FLUX 3 include action prediction?

FLUX 3's training is meant to capture how scenes change over time. That makes the video backbone relevant to action prediction: a system sees the current state, predicts what may happen next, and connects an action with its physical result.

BFL is testing two routes. One adds action prediction directly to FLUX 3. The other fine-tunes the pretrained video backbone with a smaller amount of task-specific data. The first named partnership is with mimic robotics on FLUX-mimic, a model for dexterous manipulation and production tasks.

This does not turn the early-access video generator into a general robot controller. It explains why BFL considers content creation and physical AI related research problems. Both require a model to learn how the world changes.

What could FLUX 3 make possible?

The near-term uses are easier to see than the longer-term claims. Creators could start from a still image, animate it with sound, continue a source clip, or move a recurring character into a different setting. Keyframes could give directors more control over transitions. Multilingual speech may also reduce the number of separate tools needed for localized video.

Image generation and editing may be useful for layouts that include readable words, especially if multilingual typography works reliably. A shared model across stills and motion also raises the possibility of moving an idea from image concept to animated sequence without rebuilding its visual identity at each step.

Further out, BFL points to interactive editing, simulation, computer use, and physical AI. Those are research directions. The product details currently support a tighter conclusion: FLUX 3 is an early attempt to put image creation, image editing, video, and native audio on one multimodal foundation.

Frequently asked questions

Is FLUX 3 available now?

FLUX 3 Video is available through a gated early-access application as of July 24, 2026. FLUX 3 Image has been announced for early access in the following weeks. The full model family is not generally available.

Does FLUX 3 generate video with sound?

Yes. BFL says FLUX 3 Video generates native audio with its video output, including multilingual dialogue and sounds connected to events in the scene. A single generation can be up to 20 seconds long.

Can FLUX 3 edit images?

Image editing is an announced FLUX 3 Image capability. Its early-access release is still pending, and BFL has not published complete editing controls, pricing, or independent results.

Can FLUX 3 edit existing videos?

The announced capabilities include video-to-video generation, continuation from existing video and audio, visual references, and keyframe transitions. BFL also plans video and audio editing access through APIs and private weights, but broad access is not available yet.

Does FLUX 3 support image-to-video?

Yes. It can animate a starting frame or use images as visual references for a new video. BFL also describes using references to help preserve characters across separate scenes.

How long can a FLUX 3 video be?

FLUX 3 can generate up to 20 seconds of video with audio in one pass. Longer, multi-shot sequences require chaining individual clips.

Is FLUX 3 open source?

BFL has announced a future open-weight model called FLUX 3 Dev. The license, release date, model size, and hardware requirements have not been published, so “open-weight” is the accurate description for now.

How much does FLUX 3 cost?

Public pricing has not been announced. The early-access status also means there is no general API price or standard cost per video available to compare yet.

TheFluxTrain is a visual creation platform for building repeatable image and video workflows. As FLUX 3 access broadens, we will evaluate where its image, video, and editing capabilities fit into practical creative workflows rather than treating the announcement alone as proof of production readiness.