The MiniMax H3 prompt generator gives you 20 free, copy-ready video prompts for MiniMax's new omni-modal 2K video AI. Native stereo audio, 9-reference omni input, multi-shot storytelling — all structured for H3's breakthrough capabilities.
The MiniMax H3 prompt generator on this page provides 20 free, professionally crafted prompts for MiniMax H3 — also known as Hailuo 3.0 — the latest AI video model from MiniMax, launched July 31, 2026. H3 is a generation-level upgrade: a single omni-modal transformer that understands text, images, video, and audio simultaneously, and returns 2K video with native synchronized stereo audio in one pass.
MiniMax H3 currently ranks #2 globally on the Image-to-Video leaderboard (Elo 1351 on Arena) and introduces the most flexible multi-reference input system of any public video model — accepting up to 9 reference images, 3 video clips, and 3 audio clips simultaneously to lock character identity, motion, and voice. It also supports motion transfer, instruction-based editing, and multi-shot storytelling in a single generation.
Every prompt below is structured for H3's distinct capabilities — covering cinematic portraits, omni-reference fashion, multi-shot sequences, native audio integration, and physics showcases. Paste directly into hailuoai.video, the MiniMax API, kie.ai, or Runware.
H3 parses full narrative descriptions, not keyword tags. Use this structure:
Click any prompt to copy — paste into hailuoai.video, MiniMax API, kie.ai, or Runware
A medium close-up of a steelworker exiting a mill at dusk: molten sparks drift across the background in slow arcs as the worker pulls off heavy gloves, exhales slowly, and looks toward the horizon. Sweat on the brow catches the orange foundry light. The camera holds steady at chest height, slightly below eye level, for 12 seconds before a slow pull-back reveals the full industrial skyline. 2K resolution, golden-hour light, HBO documentary quality. Audio: the low percussion of industrial machinery fading as ambient evening wind takes over.
An editorial fashion sequence: a model in a structured cream linen blazer walks from shadow into a pool of gallery light — three seconds of approach, a pause in the light as she turns to face camera with a composed expression, then resumes walking out of frame right. Reference her face, outfit, and posture from provided images for consistency. Camera angle: low 3/4 tracking shot transitioning to static mid-shot at the pause. Vogue editorial quality, 2K, 14 seconds. Audio: the soft tap of heels on polished concrete.
A drone shot rising from ground level through morning mist to reveal a stone temple complex at dawn: the camera begins tight on lichen-covered steps, then rises slowly to reveal colonnades, then the full facade, then the surrounding jungle canopy still in pre-dawn shadow. Golden light catches the upper reliefs first, then descends. The reveal takes 12 seconds with a steady vertical ascent. 2K aerial photography quality, warm ambient audio: cicadas, distant birds, the rustle of wind through canopy. No music.
A three-shot cooking sequence for a ramen documentary: Shot 1 (4 seconds) — extreme close-up of tare being ladled into a ceramic bowl, steam rising immediately; Shot 2 (5 seconds) — mid-shot of noodles lifted from boiling water and placed with chopsticks; Shot 3 (5 seconds) — wide shot of the completed bowl placed before a diner who lifts their chopsticks. Continuous lighting across all three shots — warm under-counter LEDs, restaurant ambience. Warm ambient audio: kitchen sounds layered throughout. 2K, restaurant documentary quality.
A motion-transfer sequence: the sweeping, rotational arm movement of a flamenco dancer is mapped onto the structural elements of a Gothic cathedral interior — the camera sweeps through the nave with the same arc and rhythm as the dancer's arm, from the floor tiles up to the vaulted ceiling and back. The motion should feel like architecture dancing. Visual style: dramatic chiaroscuro, deep shadows with single God-ray illumination. 2K, 12 seconds. Audio: the echo of a flamenco heel tap reverberating through stone, fading to silence.
A commercial-style imagination sequence: a child in a hand-made cardboard astronaut helmet sits in a cardboard rocket in a backyard at night — in the first 4 seconds, the rocket is clearly cardboard with crayon-drawn flames. Then the scene slowly blends into a real rocket launch sequence — the same child now in full NASA gear, looking out at actual stars — then cuts back to the backyard as the sprinklers come on. 12 seconds total. 2K, warm colour grade. Audio: the child's laughter morphing into rocket engine sound then back to sprinkler noise and laughter.
A close documentary sequence of a jazz quartet in a dim basement venue: Shot 1 — close-up of the pianist's left hand moving through a walking bass line on the keys; Shot 2 — the drummer's brush strokes on the snare, barely audible but rhythmically precise; Shot 3 — the saxophonist raising the instrument to their lips before the first breath. 14 seconds total. The native audio should capture the acoustic space — slightly reverberant, intimate, no ambient noise overlay. 2K, low-light documentary. Audio-visual sync is the primary quality target.
A luxury perfume launch hero video: the bottle rises slowly from black — illuminated from below by a thin strip of amber light — and rotates one full revolution over 10 seconds before the brand name appears in translucent white serif text at the centre of frame. The liquid inside shifts visibly as the bottle turns, catching and refracting the light. Background: absolute matte black. The rotation ends with the bottle perfectly face-on. 2K, luxury production quality. Audio: a single low note on a cello, sustaining through the full sequence.
A dynamic tracking sequence of a parkour athlete at blue hour across city rooftops: the camera follows from behind as the athlete leaps a 2-metre gap in full stride, rolls through the landing on the opposite roof without breaking pace, and continues into a sprint along a parapet with the city lights emerging below. The camera keeps pace at shoulder height, slightly shaky, tracking through the roll and accelerating out of it. 2K, blue-hour ambient light — no artificial lighting. Audio: the impact and roll footfalls, wind rushing past the microphone, receding city ambient.
A vertical ink wash scroll animation in classical Chinese brush style: the camera pans slowly upward from a misty river at the base of the composition to a snow-capped mountain peak — passing through pine forest, cliff faces, and a waterfall mid-composition. Each layer is painted in diluted ink washes with dry-brush detail on the trees and rock faces. The animation is slow — 15 seconds for the full vertical pan. No colour, monochrome ink on warm paper texture. Audio: complete silence except for a single water drop at the three-second mark.
A sci-fi character study: a space station technician in a worn jumpsuit floats through a narrow corridor checking conduit panels, pauses at a porthole window where a gas giant fills the view, then pushes off and continues. Reference character appearance and jumpsuit from provided images for consistency. Camera: handheld-style following shot transitioning to static at the porthole. 14 seconds. 2K, Arrival-level production design. Audio: the hum of life support systems, the soft metallic tap of the technician's hand on the porthole frame.
A wildlife close-up of an Arctic fox in its white winter coat at dusk: the fox sits on a snow-covered rock, ears angled forward, watching something off-frame. After 6 seconds it turns its head directly to camera — holds the gaze for 3 seconds — then trots off frame right with its brush trailing low. Camera: static telephoto, slightly compressed depth of field, snow bokeh in foreground. 2K, BBC Planet Earth natural history quality. Audio: wind over tundra, the soft crunch of the fox's departure.
A generative editing sequence: begin with a wide shot of a landmark plaza filled with tourists. Over 8 seconds, apply a progressive instruction-based edit — the crowd fades and disappears person by person from the edges of frame inward, leaving the landmark entirely clear and quiet in the final frame. The transition should feel like time-lapse emptying, not a hard cut. Final 4 seconds: hold on the empty plaza with ambient wind audio. 2K. Audio: crowd noise in the first half fading to silence and birdsong in the second half.
A vintage lookbook sequence in a 1970s film aesthetic: a model in a wide-lapel caramel suede jacket walks slowly through a sunlit courtyard, pausing to glance back over one shoulder at camera. Film grain texture, slight colour fade, warm golden-brown palette, mild lens vignette. Shot on what appears to be 16mm film — grain visible, slight gate weave. 12 seconds. 2K with authentic film look applied in-model. Audio: scratchy lo-fi vinyl playing distant bossa nova, ambient outdoor sound underneath.
A consistent-location sequence for a real-estate walkthrough: reference provided images of the interior space to maintain exact room proportions and furnishing positions across three shots — Shot 1 (4s): door opens to reveal living room; Shot 2 (5s): 180-degree pan of the space from the centre; Shot 3 (5s): floor-to-ceiling window reveal with city view. Lighting consistent across all shots: diffused noon daylight through the windows. 2K, architectural photography quality. Audio: ambient room tone only — no music.
An anime-style close-quarters combat sequence in heavy rain under neon signs: two fighters exchange a 4-hit combination — the first strike blocked, the second landing, a counter-grab, a throw that ends with one fighter braced against a neon-lit wall. Raindrops freeze momentarily at each impact in bullet-time, then resume. Neon colours reflect in the wet street beneath them. Camera cuts between angles — wide to mid to close — across the sequence. 12 seconds. Tatsuki Fujimoto visual style, saturated neon palette, 2K. Audio: rain, impact sounds, a distant police siren.
A controlled-light studio portrait sequence for a professional headshot: subject enters from frame right, positions themselves at a marked spot before a seamless grey background, turns to face camera with a composed expression, and holds for 6 seconds — subtle micro-expressions. A second angle: the key light shifts from a 3/4 position to side-lighting over 3 seconds, dramatically changing the mood from approachable to authoritative. 12 seconds total. 2K, Peter Lindbergh-quality studio lighting. Audio: studio silence — no ambient noise.
A sweeping exterior reveal of a contemporary glass tower at night: the camera begins at street level looking up, then slowly ascends in a vertical pull rising alongside the facade — each floor's office lights creating a lit grid pattern that reflects and refracts in the glass. At the top, the camera pulls back to show the full tower against a deep blue night sky. 15 seconds. 2K, architectural CGI-level quality. Audio: low city ambient — distant traffic, wind — giving way to a single sustained low note at the apex.
A morning routine short-form narrative across 5 shots: Shot 1 (2s) — hand turns off alarm at 5:45 AM; Shot 2 (3s) — close-up of coffee pouring into a white cup; Shot 3 (3s) — shoes lacing up in the entryway; Shot 4 (3s) — front door opening to a cold dawn; Shot 5 (4s) — wide shot of the figure walking into morning light. Consistent colour grade across all shots: cool desaturated, blue-grey dawn light. 2K. Audio: continuous ambient audio threading through all shots — alarm fading, coffee sounds, footsteps, door close, birdsong emerging.
A physics simulation showcase: a glass sphere rolls along the edge of a table in extreme slow motion — approaching the edge, tipping, then falling and shattering on a marble floor below, each fragment moving in physically accurate trajectories from the impact point. The camera is positioned at table height for the approach and drop, then cuts to floor level for the shatter. 12 seconds total at 50% playback speed. 2K, product photography studio lighting. Audio: the crack and cascade of glass rendered at full fidelity, silence before and after.
MiniMax H3 leads for multi-reference input flexibility and audio-visual synchronization — the only top-ranked model that accepts 9 images + 3 videos + 3 audio clips simultaneously:
| Model | Max Resolution | Native Audio | Max Refs | Leaderboard |
|---|---|---|---|---|
| MiniMax H3 / Hailuo 3.0 ★ | 2K (1440p) | Yes — in-model | 9 img + 3 vid + 3 audio | #2 I2V (Elo 1351) |
| Kling v3 (Kuaishou) | 1080p | Limited | 5 images | #1 T2V (Elo 1934) |
| Seedance 2.5 (ByteDance) | 4K native | Yes — co-generated | 50 multi-modal refs | Top 5 T2V |
| HappyHorse-1.0 (Alibaba) | 1080p | No | Image only | #2 T2V (Elo 1816) |
| FLUX 3 (Black Forest Labs) | 720p → 1080p | Yes — synchronized | Image + audio | Early access |
| Hailuo 2.3 (MiniMax) | 1080p | No | Image only | #2 overall (May 2026) |
★ MiniMax H3 / Hailuo 3.0: #2 globally on Image-to-Video Arena leaderboard (Elo 1351, August 2026). Available via hailuoai.video, MiniMax API, kie.ai, and Runware. Open weights release pending.
The MiniMax H3 prompt generator on this page gives you 20 free, copy-ready video prompts for MiniMax H3 — also known as Hailuo 3.0 — the latest AI video model from MiniMax, launched July 31, 2026. MiniMax H3 is an omni-modal generation model that produces 2K video with native stereo audio from a single prompt, accepting up to 9 reference images, 3 video clips, and 3 audio clips simultaneously. It currently ranks #2 globally on the Image-to-Video leaderboard (Elo 1351).
MiniMax H3 is an AI video generation model developed by MiniMax, a Shanghai-based AI company. Its consumer app brand is Hailuo 3.0, available at hailuoai.video. Launched July 31, 2026, H3 is MiniMax's most powerful model to date — a general-purpose omni-modal transformer that understands text, images, video, and audio together and returns 2K video with native stereo audio. It supports clips from 5 to 15 seconds (extendable to ~30 seconds), aspect ratios from 21:9 to 9:16, and features including motion transfer, multi-shot storytelling, omni-reference creation, and instruction-based video editing. MiniMax has indicated it will release H3 as open weights.
MiniMax H3 is a generation-level upgrade over Hailuo 2.3, not an incremental update. H3 introduces: (1) 2K native resolution at 24fps — vs. Hailuo 2.3's 1080p maximum; (2) Native stereo audio generated in the same pass as the video — Hailuo 2.3 had no audio; (3) Omni-reference input — you can now lock in character identity, motion, and voice simultaneously using up to 9 images, 3 videos, and 3 audio clips; (4) Motion transfer — apply the movement pattern from one subject to another; (5) Multi-shot storytelling — the model maintains scene and character consistency across multiple shots in a single generation. H3 also replaces the per-style mode approach (anime/ink wash/game CG) with a unified model that handles any style.
MiniMax H3 responds to detailed narrative descriptions. Use this structure: (1) Shot type and camera movement — 'static telephoto,' 'low tracking shot,' 'slow vertical ascent'; (2) Subject, action, and timing — describe what happens in sequence, including specific durations; (3) Multi-shot structure — if generating multiple shots, label each as 'Shot 1 (Xs),' 'Shot 2 (Xs),' etc.; (4) Audio intent — describe the intended sounds since H3 generates synchronized audio natively; (5) Resolution and quality reference — state '2K' and a quality benchmark like 'BBC documentary' or 'Vogue editorial'; (6) Duration — specify total clip length within the 5–15 second range (up to ~30 seconds with extension). Unlike earlier Hailuo models, H3 handles full narrative structure well — write in complete sentences, not keyword tags.
MiniMax H3 (Hailuo 3.0) is accessible via the official Hailuo AI web app at hailuoai.video, where a free tier with limited daily generations is available. The MiniMax API also provides H3 access for developers — model IDs are published on the MiniMax developer portal. Third-party platforms including kie.ai and Runware have added H3 support. MiniMax has also committed to releasing H3 as open weights, which will allow local deployment. For omni-reference generation (using multiple image, video, and audio inputs simultaneously), the API provides the most flexible access.
MiniMax H3, Kling v3, and Seedance 2.5 are the three strongest AI video models as of August 2026. Kling v3 leads on text-to-video in the Arena leaderboard and excels at cinematic camera motion and long-subject consistency. MiniMax H3 leads on image-to-video (Elo 1351, #2 I2V) and has the strongest omni-reference input system — 9 images + 3 videos + 3 audio clips in one prompt. Seedance 2.5 (ByteDance) leads on native 4K output and 30-second clip length. For multi-reference character consistency and audio-visual synchronization in one pass, H3 has no peer. For pure text-to-video cinematic quality, Kling v3 still leads. For the longest clips at highest resolution, Seedance 2.5 is the choice.
#1 T2V globally — cinematic camera motion & long sequences
MiniMax's previous model — anime, physics, ink wash styles
ByteDance — 4K native, 30-second clips, synchronized audio
Alibaba's #2 T2V — photorealistic cinematic quality
Black Forest Labs — unified image + video + audio generation
Google's flagship video model — photorealistic with dialogue