July 24, 2026 ChainGPT

FLUX 3: Multimodal 20s Video + Audio Sparks New NFT, Metaverse Possibilities

FLUX 3: Multimodal 20s Video + Audio Sparks New NFT, Metaverse Possibilities
Black Forest Labs has taken its FLUX family from stills to motion: the German lab unveiled FLUX 3 on Thursday, and the headline feature is video. Trained simultaneously on images, video and audio inside a single shared model (true multimodality), FLUX 3 can produce up to 20-second clips with synchronized sound — dialogue, effects and ambient audio that line up with what’s happening on screen. How it performs - In human preference tests run by BFL, reviewers chose FLUX 3’s clips over Runway Gen-4.5 in 77% of pairwise comparisons and over Luma Ray 3.2 in 93%. FLUX 3 edged out Gemini Omni and Seedance in 52% of comparisons. These are subjective preference tests — viewers simply pick the clip they find more convincing. - Still-image capability remains strong: BFL shared samples that demonstrate a wide stylistic range beyond photorealism, continuing the model’s legacy as a versatile image generator. Why multimodality matters According to co-founder and CEO Robin Rombach, a model trained to predict video must learn underlying physical dynamics — weight, contact and timing — not just pixels. BFL bets that this deeper understanding makes the model useful beyond content creation, enabling applications that require a notion of real-world motion. Robots that “understand” movement BFL and Zurich-based mimic robotics turned that idea into a product called FLUX-mimic: FLUX 3’s video-prediction engine plus a lightweight decoder that maps the model’s internal motion predictions into robot actions. Audi is already testing FLUX-mimic for tasks like fitting flexible door seals — soft-body manipulation that conventional industrial automation finds difficult. Mimic says the full system reacts in roughly 101 milliseconds, in the ballpark of human visual reflexes. Company background and trajectory - Black Forest Labs was founded in August 2024 by researchers who helped build the original Stable Diffusion models at Stability AI. - Early FLUX releases (Flux Dev and Schnell) won praise in the open-source community, at one point besting MidJourney and Stable Diffusion 3 in artist comparisons. FLUX 1.1 Pro later topped the Artificial Analysis image rankings. - FLUX.2 (November 2025) failed to gain the same traction, and the open-source lead eventually shifted when Alibaba’s Z-Image Turbo matched Flux-quality on lower-end GPUs in late 2025. - FLUX 3 is positioned as a comeback, though it’s not fully open yet. Availability and roadmap - Video and Action features are in early access via APIs and select partners (mimic robotics is a named partner). - Image generation will follow “in the coming weeks,” BFL says. - The company plans to release an open-weight Dev tier for local use, but that release is slated for later in 2026. What this means for crypto and web3 creators Generative video with built-in audio opens new doors for NFT creators, metaverse content, animated avatars and multimedia minting workflows. Early-access APIs could be integrated into web3 platforms that offer tokenized media tools or on-chain provenance, though BFL’s current access model is still gated and the fully open weights won’t arrive until next year. Takeaway FLUX 3 pushes multimodal generative AI into video and ties that capability to robotic control in a way that could matter for both content creators and industrial partners. The preference-test wins are eye-catching, but subjective; broader benchmarking and wider access will be needed to see how FLUX 3 stacks up long term — and how quickly it filters into creator tooling and web3 ecosystems. Read more AI-generated news on: undefined/news