FLUX 3: All You Need to Know About Black Forest Labs' First Video Model

Share article

FLUX 3 is the first video model and first multimodal foundation model from Black Forest Labs, the company behind the FLUX image models. Trained jointly on images, video, and audio in one architecture, it turns a text prompt, or a starting image, into a clip of up to 20 seconds with native, synchronized sound. On getimg.ai, the video model is available now for Text to Video and Image to Video generation, with the image model coming soon.

What FLUX 3 is

FLUX 3 is a multimodal foundation model that learns from images, video, and audio at the same time, inside a single architecture, rather than treating each as a separate problem. In practice that means the same model that generates a picture also understands motion and sound, and it uses one to inform the others.

Black Forest Labs describes the reasoning behind this directly. Each modality, it says, "is a projection of the same underlying reality," and learning from all of them at once means their constraints reinforce each other: "the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past." The company calls FLUX 3 "our first model built entirely on that principle, and a checkpoint on our mission to develop real-world visual intelligence."

Under the hood, FLUX 3 builds on Self-Flow, the company's method for aligning multimodal generation and understanding within the same architecture, on top of the flow matching techniques the earlier FLUX models are built on. For a creator, the takeaway is simpler than the theory: FLUX 3 was trained to make video and its sound together, so what you hear tends to line up with what you see.

The first capability out of the gate is video, and that is what is live on getimg.ai now.

AI video generated with FLUX 3.

The backstory: Black Forest Labs moves into video

Black Forest Labs is best known for the FLUX family of image models, including FLUX.1 for generation and FLUX.1 Kontext for prompt-based editing, which became some of the most widely used open and commercial image models available. Until now, everything the company shipped produced still images.

FLUX 3 is the shift into motion. It is both the company's first video model and its first model designed as a single multimodal backbone, meant to serve content creation as well as longer-term "physical AI" work like action prediction for robotics. Rather than bolting a video model onto an image model, Black Forest Labs trained one system across images, video, and audio from the start, then scaled up the compute and data behind it.

That framing matters because it puts FLUX 3 in the same broad direction as models like Google's Gemini Omni, which also fold video generation into a larger multimodal model instead of keeping it separate. The bet across the category is that a model which understands the world more completely will generate motion and sound that behave more believably.

What FLUX 3 brings

Black Forest Labs says FLUX 3 Video is "already particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities." Those three strengths, plus native audio and long clip length, are the headline of the release.

Native, synchronized audio

Every FLUX 3 clip can carry its own sound, generated together with the picture rather than added afterward. Because the model learns audio and video jointly, effects and dialogue tend to land on the moment they belong to instead of drifting out of sync. You can ask for specific sound effects, ambient background, spoken lines, or a musical mood.

AI video generated with FLUX 3.

Human faces and performance

The model's stated strength is human expression, which is one of the hardest things for video models to get right. Subtle changes in a face, a shift from neutral to a smile, a flicker of hesitation, are where synthetic video usually breaks down. FLUX 3 is tuned to hold those micro-expressions together across a shot.

AI video generated with FLUX 3.

Multilingual dialogue

FLUX 3 supports dialogue in multiple languages, with sound generated to match the picture, which makes it useful for spoken lines and characters that are not limited to one language. As with any video model, keep individual lines short and generate a couple of takes when the exact delivery matters.

AI video generated with FLUX 3.

Length that most models cannot match

FLUX 3 generates clips of up to 20 seconds, currently the longest single-generation clip length of any major AI video model. The next tier reaches up to 15 seconds, including MiniMax H3, Seedance 2.0, and HappyHorse 1.1, while others such as Gemini Omni cap at 10.

AI video generated with FLUX 3.

The extra length means a full beat, a small scene, or a complete spoken sentence can play out in one clip instead of being stitched together from several.

How FLUX 3 benchmarks

Black Forest Labs ran preliminary human-preference evaluations using 10-second, 720p text-to-video clips with audio, and reported that raters preferred FLUX 3 Video over a range of current models. The results, as published by the company, are below.

Compared With

FLUX 3 Video Preferred In

Luma Ray 3.2

93% of comparisons

Runway Gen-4.5

77% of comparisons

Grok Imagine Video

Up to 69% of comparisons

Kling v3 Pro

60% of comparisons

Happy Horse v1

59% of comparisons

Seedance 2.0 & Gemini Omni Flash

52% of comparisons

These are the model maker's own preliminary numbers, run on its chosen prompts and settings, so treat them as a strong early signal rather than a settled verdict. The most reliable way to judge a video model is still to run your own prompts against the alternatives, which is exactly what having several models in one place makes easy.

How to use FLUX 3 on getimg.ai

FLUX 3 is available now on getimg.ai as a video model. You can generate from a text prompt or from a starting image, and each clip arrives with native audio. Here is exactly what the model supports today.

Capability

What You Get with FLUX 3 on getimg.ai

Generation Modes

Text-to-video, or image-to-video from a starting image

Clip Length

5 to 20 seconds

Resolution

HD (720p) or Full HD (1080p)

Aspect Ratios

21:9, 16:9, 4:3, 1:1, 3:4, 9:16

Audio

Native, generated in sync with the video

To generate a clip, open the AI video generation Action, pick FLUX 3, describe your scene and the sound you want, add a First Frame/First + Last Frame image(s) if you wish, then set your length, resolution, and aspect ratio and run it.

AI video generated with FLUX 3.

A couple of things are still on the roadmap. FLUX 3 on getimg.ai currently supports Text to Video and Image to Video only: video continuation, which extends a clip you already have, is coming soon. The FLUX 3 image model will follow as a separate capability further down the line.

If you need a longer or higher-resolution final piece, you can take any clip further with a video upscaler after generating.

The Bottom Line

FLUX 3 is Black Forest Labs' first video model and its first multimodal foundation model, built to generate video and its audio together from the same understanding of the world. Its creators highlight human expression, sound tied to physical events, and multilingual dialogue as its strengths, and their early benchmarks put it ahead of a broad set of current video models, with the usual caveat that vendor numbers are a starting point, not the last word.

Its standout practical feature is length: up to 20 seconds per clip, longer than the rest of the field. On getimg.ai you can use it now for text-to-video and image-to-video at up to FHD across seven aspect ratios, with video continuation and a FLUX 3 image model both on the way. The clearest way to see where it fits is to run it against the other video models on your own prompts, in the same place.

Create a video with FLUX 3 on getimg.ai, where you can run it against every other video model side by side.

Frequently Asked Questions

Get started with getimg.ai

Create an account and start creating AI content for free. Work smarter, not harder.

Love creating with getimg.ai?

Invite a friend using your referral link. When they subscribe, you both get rewarded.

Start earning

Have questions or feedback?

We're here to help.

Contact us