The FLUX 3 prompt generator gives you 20 free, copy-ready prompts for Black Forest Labs' new multimodal AI — text-to-video up to 20 seconds with native synchronized audio. First dedicated FLUX 3 prompt library. No signup needed.
The FLUX 3 prompt generator on this page gives you 20 professionally written, copy-ready prompts for Black Forest Labs' first multimodal AI model — a unified architecture that generates images, video (up to 20 seconds), and native synchronized audio from a single prompt. FLUX 3 launched on July 23, 2026, and is currently in early access at bfl.ai/models/flux-3.
FLUX 3 is a significant shift for Black Forest Labs, known for the widely used FLUX 2 image model. Rather than a separate video product, FLUX 3 is a single architecture trained jointly on images, video, and audio — meaning you describe motion and sound together in one prompt and receive a clip where the audio is generated in sync with what you see, not dubbed on afterward.
Every prompt below is written for FLUX 3's multimodal input format — motion direction, timing, audio description, and visual style are all combined into a single coherent instruction. These prompts also work with other leading video AI tools like Sora 2, Kling 3, and Grok Imagine Video.
Click any prompt to copy — optimised for FLUX 3, also works with Sora 2, Kling 3, Runway Gen-4.5, and Grok Imagine Video
A sweeping cinematic time-lapse of jagged mountain peaks at golden hour — the sun descending behind a ridgeline, long amber shadows stretching across alpine meadows, wisps of cloud drifting in fast motion below the peaks, a lone eagle circling in slow motion against the accelerated sky. Native ambient audio: wind through grass, distant thunder rolling in from the west, birds calling at dusk. 20 seconds, 4K, wide panoramic frame, no camera shake.
A slow-motion street-level video following a person walking through a Tokyo back-alley in the rain at night — neon kanji signs reflecting in wet asphalt, rain droplets catching the pink and blue light, steam rising from a ramen stall, the subject's silhouette reflected in a puddle. Native audio: rain on pavement, distant traffic, the hiss of a wok from inside the stall. 20 seconds, shallow depth of field, cinematic colour grade.
A luxury watch product reveal video — the watch emerging from soft blue-grey smoke on a black mirrored surface, close-up rotating slowly to reveal dial details (hour markers, lume hands, sapphire crystal), light raking across the bracelet links, ending on a wide reveal shot with the brand name appearing in gold serif text. Native audio: a single clear piano note, soft whoosh on the smoke, ambient silence. 20 seconds, macro lens feel, perfect studio lighting.
A dramatic coastal video of a storm rolling toward shore — dark cumulonimbus clouds building on the horizon, the sea turning from grey-green to steel-black, white foam crashing against sea-stacked rocks in slow motion, rain beginning to streak the frame in the final seconds, lightning flickering in the far clouds. Native audio: waves crashing, deep rumble of distant thunder, wind rising, seagulls fleeing. 20 seconds, wide cinematic frame, desaturated dramatic grade.
A ballet dancer warming up alone in a sunlit studio — side-light from tall windows casting long shadows, the dancer moving through slow arabesque sequences reflected in a full-wall mirror, dust motes drifting in the light, ending in a held pose with the mirror showing both front and back simultaneously. Native audio: piano rehearsal tape playing softly, the dancer's breath, fabric swishing, the soft thud of pointe shoes. 20 seconds, 4K, documentary realism.
A close-up, slow-motion video of onions caramelising in a cast-iron pan over a flame — the onions softening from white to golden to deep amber over the 20 seconds, steam rising, butter melting along the edges, the pan surface glistening as the sugars develop, a hand occasionally stirring with a wooden spoon. Native audio: the sustained sizzle, the crackle of butter, a subtle low hiss of the flame. Warm 4K, overhead plus side angle cuts.
First-person POV walking through a dimly lit space-station corridor — hexagonal wall panels glowing faint blue, emergency red lights pulsing at junctions, the visor of the suit reflected in the helmet display overlay, a door sliding open to reveal a vast viewport overlooking a ringed planet. Native audio: the low hum of life support systems, visor breathing, the magnetic click of boots on metal, distant machinery. 20 seconds, cinematic sci-fi colour palette.
A serene underwater video drifting through a California kelp forest — towering amber kelp fronds swaying gently in the current, shafts of sunlight refracting and dancing on the sandy floor, a leopard shark gliding in the mid-ground, a school of garibaldi fish flashing orange near the surface. Native audio: the muffled underwater ambience, soft bubbles, distant whale-call echo. 20 seconds, wide angle, cool green-gold palette.
A fashion editorial video with a model in a floor-length ivory silk dress standing before a clean white cyclorama — a wind machine creating dramatic fabric movement in waves, strands of hair catching the light, the fabric billowing in slow motion against the white background, the model's pose holding still while fabric wraps and releases. Native audio: the soft rush of wind, silk fabric swooshing, ambient studio silence. 20 seconds, high-key lighting, editorial clean grade.
An intimate campfire video in a pine forest at night — orange and amber flames dancing low in a stone ring, embers rising and cooling in the dark, silhouettes of pine trees fading into the black background, a pair of hands reaching toward the warmth, sparks floating upward. Native audio: crackling fire, the pop of sap, a distant owl call, wind moving through pines. 20 seconds, low exposure, warm amber grain, real-fire natural light only.
A macro timelapse of a red ranunculus bud slowly unfurling to full bloom against a soft sage-green background — petals peeling back one by one in delicate spirals, stamens revealing themselves, the entire bloom completing in 20 seconds, a final close-up on the fully open flower with morning light catching the translucent petals. Native audio: a soft ambient tone, the faintest paper-thin rustle of petals. 4K macro, clean studio light, no music.
A slow-motion video inside a boxing gym at dusk — a fighter working a heavy bag, gloves thudding in rhythmic combinations, sweat catching the warm overhead light, the bag swaying back and then the fighter stepping in again, dust motes in the shafts of light from high windows, ending on a held breathing moment. Native audio: leather on canvas thuds, the chain rattling, laboured breathing, the murmur of other fighters in the background. 20 seconds, desaturated gritty grade.
An overhead slow-motion video of a barista pouring steamed milk into a double espresso — the white milk meeting the crema in swirling patterns, a tulip latte art design forming in the final seconds, steam rising from the ceramic cup, the barista's hands steady and precise. Native audio: the gentle pour, a soft sizzle of steam, the quiet ambience of a cafe. 20 seconds, overhead 4K, warm natural light, shallow depth of field on the cup surface.
A 20-second video from a Brooklyn rooftop at magic hour — the Manhattan skyline glowing amber and coral behind water towers and rooftop chimneys, a pigeon flock wheeling in a wide arc against the peach sky, the last sliver of sun disappearing behind a glass tower, city lights beginning to appear in the windows below. Native audio: distant traffic hum, the wing-beat rush of pigeons passing, a faraway siren, ambient city drone. Cinematic wide frame, warm sunset grade.
A slow push through an ancient stone library — towering wooden bookshelves lined with leather-bound volumes, shafts of dusty light falling from high arched windows, a wooden ladder on rails, the camera drifting past an open tome on a reading stand with illuminated manuscript pages, candlelight flickering in sconces. Native audio: the creak of the wooden floor, distant wind through stone, the soft flutter of a page turning. 20 seconds, warm amber and sepia grade, no people.
A cinematic product video of a black electric sports car on a wet midnight runway — rain-soaked asphalt reflecting the underlit car body, the car launching from standstill in near-silence (only a high-pitched electric whine), headlights cutting through rain ahead, the rear end kicking out a spray of water in the final seconds. Native audio: the EV electric whine rising, rain on the runway, water spray. 20 seconds, low camera angle, deep black colour grade, anamorphic lens flare.
A misty morning video at a Tibetan monastery — monks in saffron robes walking single file across a stone courtyard as dawn light breaks over the mountains, butter lamps flickering in a row along a wall, prayer flags undulating in the mountain breeze, a monk ringing a bronze bell at the edge of the frame. Native audio: the deep resonant bell tone sustaining, fabric of the robes in the wind, distant chanting from inside, gravel underfoot. 20 seconds, muted cool-warm palette, documentary realism.
A dramatic aerial-style video drifting over a pine forest at night where controlled burn embers glow in the dark — thousands of small orange points of light among black silhouetted trees, smoke drifting in visible layers catching the ember-light below, the camera moving slowly forward over the glowing landscape. Native audio: the low roar of fire in the distance, crackling, wind moving through the smoke. 20 seconds, dark and atmospheric, no sky visible, ember-lit palette only.
A close-up video of a vinyl record spinning on a turntable — the needle sitting in the groove, the label rotating in the centre, ambient room light catching the grooves in rainbow diffraction patterns, a hand carefully placing the tonearm down in the first seconds, the crackle of the record beginning. Native audio: the crackle and hiss of the vinyl start, a warm jazz melody beginning (diegetic, from the record), turntable motor hum. 20 seconds, warm low-lit interior, shallow DOF on the needle.
An ultra slow-motion video of a crystal wine glass shattering from a single tap — the fracture beginning at the point of impact and radiating outward in fine lines before the glass explodes into hundreds of spinning shards suspended in mid-air, light diffracting through the fragments, the shard cloud expanding outward and beginning to fall. Native audio: the single high tap, then the deep resonant crack extending into the ringing harmonic sustained across the entire 20 seconds. Clean white background, 4K, macro lens.
How FLUX 3 compares to the leading video AI models on the key dimensions that matter for prompt creators:
| Model | Max Length | Native Audio | Best For |
|---|---|---|---|
| FLUX 3 (BFL) | 20s (chainable) | ✅ Unified — generated in the same pass as the video | Sound-critical scenes, multi-reference compositing, open-weight workflow |
| Sora 2 (OpenAI) | Up to 2 min | ⚠️ Added post-generation | Long narrative sequences, ChatGPT integration, story-driven content |
| Runway Gen-4.5 | 60s | ✅ Native (added June 2026) | Fastest generation time, 20+ cinematic camera controls, real-time previews |
| Kling 3 | 3 min | ✅ Native | Longest clips, action/sport sequences, consistent character motion |
| HappyHorse 1.1 | 30s | ✅ Native + multilingual lip-sync | #2 video leaderboard, dialogue scenes, multilingual content |
| Veo 4 (Google) | 60s+ | ✅ Native | Highest photorealism, 4K output, Google Workspace integration |
FLUX 3 is a multimodal AI foundation model developed by Black Forest Labs (BFL), the team behind the original FLUX.1 and FLUX 2 image models. Released on July 23, 2026, FLUX 3 is BFL's first model to move beyond still images — it generates video clips up to 20 seconds long with native, synchronized audio from text prompts, uploaded images, or existing video clips. FLUX 3 Image and FLUX 3 Action (robotics) are planned to follow in the coming weeks under the same unified architecture.
FLUX 3 is currently in early access at bfl.ai/models/flux-3. To use the prompts on this page: open the FLUX 3 interface, select 'Text to Video' or 'Image to Video' depending on your starting point, copy any prompt from this page, paste it into the prompt field, and generate. For image-to-video, upload your reference image first and then paste the prompt to describe the motion and audio you want. Each prompt on this page is written to work best with FLUX 3's multimodal understanding of movement, timing, and audio description.
Yes — native synchronized audio is one of FLUX 3's headline features. Unlike most video AI models that generate silent clips and require a separate audio step, FLUX 3 learns the relationship between visual content and sound within the same unified architecture. This means you can describe audio alongside motion in a single prompt and receive a clip where the sound is generated to match what's happening on screen — rain when it's raining, steel-on-steel when metal strikes, music where music belongs.
FLUX 3 Video currently generates clips up to 20 seconds in early access. Clips can be extended and chained together for longer sequences. The model also supports video continuation from an existing clip, letting you extend a generated scene past the initial 20-second window into longer composite videos.
FLUX 3 competes directly with Sora 2 (OpenAI) and Runway Gen-4.5 in the premium multimodal video segment. FLUX 3's primary differentiators are its unified audio generation (no separate TTS or sound-design step), its image-to-video and video continuation capabilities, and BFL's open-weight commitment (a Dev tier is planned). Sora 2 currently supports longer clips (up to 2 minutes) and has deeper ChatGPT integration. Runway Gen-4.5 leads on real-time generation speed and cinematic camera controls. FLUX 3's edge is audio fidelity-to-visual and multi-reference compositing inherited from the FLUX family.
FLUX 3 is currently in paid early access at bfl.ai. BFL has confirmed a free open-weight Dev tier is planned for later in 2026 — the same release strategy they used for FLUX.1 Dev, which became one of the most-used open-weight image models globally. Until the open-weight release, FLUX 3 Video is available via the bfl.ai API and through early-access partners integrating it into their platforms.
Black Forest Labs' image model — 4MP, 10 reference images, open-weight Dev
OpenAI video AI — up to 2-minute clips, best temporal coherence
3-minute AI video clips, native audio, #1 action sequences
Fastest cinematic video generation — 20+ camera controls, 4K
#1 image-to-video leaderboard, native audio + lip-sync
Multi-shot video from a single prompt, 1080p native audio output