FLUX 3: All You Need to Know About Black Forest Labs' First Video Model
FLUX 3 is the first video model and first multimodal foundation model from Black Forest Labs, the company behind the FLUX image models. Trained jointly on images, video, and audio in one architecture, it turns a text prompt, or a starting image, into a clip of up to 20 seconds with native, synchronized sound. On getimg.ai, the video model is available now for Text to Video and Image to Video generation, with the image model coming soon.
What FLUX 3 is
FLUX 3 is a multimodal foundation model that learns from images, video, and audio at the same time, inside a single architecture, rather than treating each as a separate problem. In practice that means the same model that generates a picture also understands motion and sound, and it uses one to inform the others.
Black Forest Labs describes the reasoning behind this directly. Each modality, it says, "is a projection of the same underlying reality," and learning from all of them at once means their constraints reinforce each other: "the sound has to match the impact, the motion has to obey the mass, the future has to follow from the past." The company calls FLUX 3 "our first model built entirely on that principle, and a checkpoint on our mission to develop real-world visual intelligence."
Under the hood, FLUX 3 builds on Self-Flow, the company's method for aligning multimodal generation and understanding within the same architecture, on top of the flow matching techniques the earlier FLUX models are built on. For a creator, the takeaway is simpler than the theory: FLUX 3 was trained to make video and its sound together, so what you hear tends to line up with what you see.
The first capability out of the gate is video, and that is what is live on getimg.ai now.
AI video generated with FLUX 3.
The backstory: Black Forest Labs moves into video
Black Forest Labs is best known for the FLUX family of image models, including FLUX.1 for generation and FLUX.1 Kontext for prompt-based editing, which became some of the most widely used open and commercial image models available. Until now, everything the company shipped produced still images.
FLUX 3 is the shift into motion. It is both the company's first video model and its first model designed as a single multimodal backbone, meant to serve content creation as well as longer-term "physical AI" work like action prediction for robotics. Rather than bolting a video model onto an image model, Black Forest Labs trained one system across images, video, and audio from the start, then scaled up the compute and data behind it.
That framing matters because it puts FLUX 3 in the same broad direction as models like Google's Gemini Omni, which also fold video generation into a larger multimodal model instead of keeping it separate. The bet across the category is that a model which understands the world more completely will generate motion and sound that behave more believably.
What FLUX 3 brings
Black Forest Labs says FLUX 3 Video is "already particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities." Those three strengths, plus native audio and long clip length, are the headline of the release.
Native, synchronized audio
Every FLUX 3 clip can carry its own sound, generated together with the picture rather than added afterward. Because the model learns audio and video jointly, effects and dialogue tend to land on the moment they belong to instead of drifting out of sync. You can ask for specific sound effects, ambient background, spoken lines, or a musical mood.
AI video generated with FLUX 3.
Human faces and performance
The model's stated strength is human expression, which is one of the hardest things for video models to get right. Subtle changes in a face, a shift from neutral to a smile, a flicker of hesitation, are where synthetic video usually breaks down. FLUX 3 is tuned to hold those micro-expressions together across a shot.
AI video generated with FLUX 3.
Multilingual dialogue
FLUX 3 supports dialogue in multiple languages, with sound generated to match the picture, which makes it useful for spoken lines and characters that are not limited to one language. As with any video model, keep individual lines short and generate a couple of takes when the exact delivery matters.
AI video generated with FLUX 3.
Length that most models cannot match
FLUX 3 generates clips of up to 20 seconds, currently the longest single-generation clip length of any major AI video model. The next tier reaches up to 15 seconds, including MiniMax H3, Seedance 2.0, and HappyHorse 1.1, while others such as Gemini Omni cap at 10.
AI video generated with FLUX 3.
The extra length means a full beat, a small scene, or a complete spoken sentence can play out in one clip instead of being stitched together from several.
How FLUX 3 benchmarks
Black Forest Labs ran preliminary human-preference evaluations using 10-second, 720p text-to-video clips with audio, and reported that raters preferred FLUX 3 Video over a range of current models. The results, as published by the company, are below.
Compared With | FLUX 3 Video Preferred In |
Luma Ray 3.2 | 93% of comparisons |
Runway Gen-4.5 | 77% of comparisons |
Grok Imagine Video | Up to 69% of comparisons |
Kling v3 Pro | 60% of comparisons |
Happy Horse v1 | 59% of comparisons |
Seedance 2.0 & Gemini Omni Flash | 52% of comparisons |
These are the model maker's own preliminary numbers, run on its chosen prompts and settings, so treat them as a strong early signal rather than a settled verdict. The most reliable way to judge a video model is still to run your own prompts against the alternatives, which is exactly what having several models in one place makes easy.
How to use FLUX 3 on getimg.ai
FLUX 3 is available now on getimg.ai as a video model. You can generate from a text prompt or from a starting image, and each clip arrives with native audio. Here is exactly what the model supports today.
Capability | What You Get with FLUX 3 on getimg.ai |
Generation Modes | Text-to-video, or image-to-video from a starting image |
Clip Length | 5 to 20 seconds |
Resolution | HD (720p) or Full HD (1080p) |
Aspect Ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 |
Audio | Native, generated in sync with the video |
To generate a clip, open the AI video generation Action, pick FLUX 3, describe your scene and the sound you want, add a First Frame/First + Last Frame image(s) if you wish, then set your length, resolution, and aspect ratio and run it.
AI video generated with FLUX 3.
A couple of things are still on the roadmap. FLUX 3 on getimg.ai currently supports Text to Video and Image to Video only: video continuation, which extends a clip you already have, is coming soon. The FLUX 3 image model will follow as a separate capability further down the line.
If you need a longer or higher-resolution final piece, you can take any clip further with a video upscaler after generating.
The Bottom Line
FLUX 3 is Black Forest Labs' first video model and its first multimodal foundation model, built to generate video and its audio together from the same understanding of the world. Its creators highlight human expression, sound tied to physical events, and multilingual dialogue as its strengths, and their early benchmarks put it ahead of a broad set of current video models, with the usual caveat that vendor numbers are a starting point, not the last word.
Its standout practical feature is length: up to 20 seconds per clip, longer than the rest of the field. On getimg.ai you can use it now for text-to-video and image-to-video at up to FHD across seven aspect ratios, with video continuation and a FLUX 3 image model both on the way. The clearest way to see where it fits is to run it against the other video models on your own prompts, in the same place.
Create a video with FLUX 3 on getimg.ai, where you can run it against every other video model side by side.
Frequently Asked Questions
FLUX 3 is Black Forest Labs' first multimodal foundation model, trained on images, video, and audio in a single architecture. Its first available capability is video generation, which turns a text prompt or a starting image into a clip of up to 20 seconds with native, synchronized sound.
Black Forest Labs, the company behind the FLUX family of image models such as FLUX.2, FLUX.1 and FLUX.1 Kontext. FLUX 3 is its first video model and its first model built as a single multimodal backbone rather than an image-only system.
Both, eventually. FLUX 3 is a multimodal model, but video is the first capability to launch and the one available now on getimg.ai. FLUX 3 image generation is a separate capability that is coming soon.
FLUX 3 generates clips of 5 to 20 seconds on getimg.ai. The 20-second maximum is currently the longest single-generation length among major AI video models; the next longest, including Minimax H3, Seedance 2.0, and HappyHorse 1.1, reach up to 15 seconds.
Yes. FLUX 3 generates sound together with the video, including effects, ambient background, spoken dialogue, and music. Because the audio and picture are made jointly, sounds tend to line up with the action that causes them.
FLUX 3 outputs HD (720p) or FHD (1080p), in six aspect ratios: 21:9, 2:1, 16:9, 4:3, 1:1, 3:4, and 9:16, covering wide cinematic, landscape, square, and vertical formats.
Yes. With Image to Video, you provide a starting image and the model animates forward from it, which is the reliable way to keep a specific character, product, or look consistent. Video continuation is coming soon.
In Black Forest Labs' own preliminary evaluations, human raters preferred FLUX 3 Video over models including Luma Ray 3.2, Runway Gen-4.5, Kling v3 Pro, Seedance 2.0, and Gemini Omni Flash. Those are vendor-run numbers, so the most reliable comparison is to run the same prompt across models yourself, which you can do on getimg.ai.
Yes. FLUX 3 is available now on getimg.ai for Text to Video and Image to Video generation. Video continuation and a FLUX 3 image model are both coming soon.




