The HappyHorse prompt generator gives you 25 free, copy-paste prompts for Alibaba's #1-ranked AI video model. HappyHorse 1.1 adds 9-image reference-to-video, multilingual lip-sync, emotional audio, and natural skin rendering — prompts updated for the June 2026 release.
The HappyHorse prompt generator on this page provides 25 free, copy-paste prompts for HappyHorse 1.1, Alibaba's AI video model that has held the #1 position on the Artificial Analysis Video Arena leaderboard since its anonymous debut in April 2026. Version 1.1, released June 22 2026, brings major upgrades: 9-image reference-to-video for multi-character consistency, naturalized skin rendering, emotional audio with lip-sync, and improved motion physics.
HappyHorse originally went viral before anyone knew who built it. It appeared anonymously at the top of the global leaderboard, sparking days of speculation on X and Reddit. When Alibaba was confirmed as the creator on April 10, CNBC, The Verge, and the wider AI press covered the story. With version 1.1, Alibaba has addressed the top user complaints from 1.0 — stiff motion, oily facial rendering, and limited multi-character control — while adding multilingual dialogue lip-sync.
Every prompt below is copy-ready for HappyHorse 1.1. The first 20 prompts work on both 1.0 and 1.1. The final 5 are specifically designed for 1.1's new features: reference-to-video with character tagging, emotional dialogue with natural speech pacing, and the improved facial close-up rendering.
Less than two months after 1.0's anonymous debut, HappyHorse 1.1 targets the real production pain points users reported — stiff motion, oily faces, ignored prompts, and limited multi-character control:
Upload up to 9 reference images and tag them as character1, character2, etc. HappyHorse fuses identity, wardrobe, and brand elements into a coherent clip — essential for brand campaigns and narrative shorts.
Eliminates the "oily facial appearance" and over-sharpening from 1.0. Naturalized skin textures preserve pores and soft skin qualities. Large facial close-ups now render with high fidelity instead of the uncanny plastic look.
Dialogue delivery now features natural variations in speech rate, pauses, and emotional tone. Background music can be explicitly controlled or disabled via prompts. Audio-visual synchronization is vastly improved.
Runs, jumps, object collisions, and fabric movement feel physically plausible. Frame-to-frame coherence is significantly improved, reducing the jitter and temporal artifacts that occasionally appeared in 1.0.
HappyHorse 1.1 is live on the official HappyHorse site, Alibaba Cloud Model Studio (Bailian), and the Qwen app.
HappyHorse debuted at #1 on the Artificial Analysis Video Arena in April 2026 and maintains top-tier rankings with version 1.1:
| Rank | Model | T2V Elo | I2V Elo |
|---|---|---|---|
| #1 | HappyHorse 1.1 (Alibaba) | 1,125 | 1,313 |
| #2 | Seedance 2.0 (ByteDance) | 1,219 | 1,344 |
| #3 | Kling 3.0 (Kuaishou) | 1,105 | 1,262 |
| #4 | SkyReels V4 | 1,104 | — |
| #5 | Veo 3.1 (Google) | 1,099 | 1,298 |
| #6 | Grok Imagine Video 1.5 | — | 1,326 |
Source: Artificial Analysis Video Arena, June 2026. Elo scores from T2V-with-audio and I2V-without-audio leaderboards.
HappyHorse 1.1 is built for cinematic realism with multi-character control. Use this structure:
Click any prompt to copy it — paste directly into HappyHorse 1.1 (prompts 1–20 also work on 1.0)
Time-lapse of a thunderstorm rolling across the Canadian Rockies: a clear summit at 6 AM, clouds building from the valley floor over 4 hours, lightning strobing behind dark peaks as the storm arrives, snow beginning to fall in the final frame, drone pullback revealing the full mountain range in chaotic beauty, BBC documentary quality, native ambient audio of wind and thunder building, 15 seconds.
Continuous tracking shot following a parkour athlete through downtown Seoul at dusk: he vaults a barrier onto a parked car, rolls off the hood, sprints across a glass skybridge, leaps to grab a fire escape railing, pulls up, turns to look at camera — all in one smooth gimbal move without a cut. Neon city lights, motion blur on the city below, authentic urban sound design, 12 seconds, Nike campaign aesthetic.
Single-shot luxury watch reveal: a gloved hand places a rose gold timepiece face-up on a backlit frosted glass surface, the camera starts at the crown, drifts slowly in a continuous macro arc over the dial — catching the guilloché engraving, the sapphire crystal reflection, the hand-polished case edge — ends on a wide showing the whole watch, dramatic spotlight from above, deep shadow, ASMR-style quiet ambient sound, 10 seconds. Rolex or Patek campaign quality.
A father and teenage daughter reunite at an airport arrivals hall after three years apart. He spots her first — medium shot of his face showing disbelief, relief, then tears. She sees him — wide shot of her running, bag bouncing. They collide — tight hug, her face buried in his shoulder. His hand on the back of her head. Cut to a slow-motion wide of the embrace, other travellers flowing past. Handheld, warm colour grade, quiet ambient hall noise, 15 seconds. BAFTA-quality human drama.
Slow drift through a pristine coral reef in the Coral Sea, camera at 6 metres depth, dappled sunlight shimmering through crystal water, a school of anthias parting around the lens, a sea turtle hovering ahead, gentle current moving coral fans, vibrant reds, oranges, and purples in extreme clarity, zero bubbles — freediver shot aesthetic, National Geographic Blue Planet quality, natural hydrophone audio, 15 seconds.
Night sky timelapse from the Atacama Desert: stars wheel overhead in an arc centred on the Milky Way core, the Magellanic Clouds visible low on the southern horizon, occasional meteor streaks, a silhouetted radio telescope slowly rotates in the foreground, pre-dawn light beginning to warm the right horizon in the final 3 seconds. European Southern Observatory production quality, ultra-dark sky, no artificial light pollution, 12 seconds.
A model in a sculptural red couture gown walks through an empty Louvre gallery at night — shot from behind as she approaches a painting, camera dollying slowly toward her, she stops, turns three-quarters, fabric sweeping across the marble floor. No people, only the two figures: her and the art. Single source spotlight from above, deep shadow, complete silence except for echoing footsteps. Vogue Paris aesthetic, 10 seconds.
Authentic documentary b-roll of a startup team preparing for an investor pitch: a founder rehearsing alone in a glass-walled conference room at 11 PM, city lights behind her, marker on whiteboard, a moment of doubt — she pauses, looks at her notes, breathes, nods to herself. Cut to the team arriving next morning, coffee in hand, setting up laptops. Hand-held observational camera, natural lighting only, no staged moments, Netflix documentary warmth, 12 seconds.
A remotely operated vehicle's camera captures a Humboldt squid at 500m depth: the animal hovers motionless, chromatophores pulsing in hypnotic colour patterns — deep red, then white, then iridescent — its siphon slowly venting, two dinner-plate eyes reflecting the ROV floodlights, then in an instant it jets backward into the black. Real submersible footage aesthetic, natural ROV lighting, genuine scientific discovery weight, 10 seconds.
A jazz pianist alone in a dim recording studio at 1 AM: single lamp over the Steinway, sheet music yellowed at the edges, he begins to play — starting with one hand finding a motif, then both hands join. Camera starts on his face (eyes closed, concentration), drifts slowly to his hands, then pulls wide to show the whole empty room around him. Intimate and raw, natural piano acoustics with room tone, 15 seconds.
A female ice climber ascending a 40-metre frozen waterfall in the Canadian Rockies: her ice axe swings and bites, crampons kicking in, breath visible in the -20°C air. Camera on a rope above her — looking down as she approaches, pauses to place a screw, looks up at the camera with fierce composure. Rope team partner visible small below. Adrenaline photography meets documentary, natural wind and axe sound, 12 seconds.
A master glassblower shaping a molten vessel: furnace glow illuminating her face orange, she works the blowpipe, rotating rhythmically, a glowing orange bubble extending and thinning at the end. Close-up on the molten glass itself — translucent, internal heat making it glow like a sun — then wide to show her silhouette against the furnace. ASMR aesthetics: natural flame hiss, no music, 10 seconds. PBS craftsmanship documentary quality.
K-pop concert opening sequence: darkness, a heartbeat on the speakers, a single spotlight materialises on the centre stage floor, a performer steps into it — sharp choreography freeze pose — then the lights explode as the beat drops: 6 performers in synchronised formation, stadium crowd erupting. Multi-camera cut sequence: wide → close-up of faces → overhead drone → ground-level tracking shot, all matching perfectly to the beat, 10 seconds.
Photorealistic POV from the ISS Cupola window: an astronaut's hands rest on the glass, Earth fills the entire viewport below — the Sahara visible at the top, the Atlantic ocean cloud systems in the middle, South America's coast emerging from the right. The station drifts imperceptibly. Complete silence. Then a voice through the comms: 'Houston, you copying this?' Ultra-photorealistic, no CGI visible, 15 seconds.
Street level during Mumbai monsoon season: a chaiwallah under a tin roof shelter serves tea to office workers waiting out the downpour, steam from the chai kettle, sheets of rain visible in the background, umbrella handles sticking out of bags, a taxi honking distantly. Intimate reportage photography brought to motion, natural monsoon audio, warm incandescent light from the stall, 10 seconds. Magnum Photos quality.
Smooth drone reveal of a parametric architecture museum: beginning from ground level close to the fluid white concrete form, the camera slowly rises and pulls back, revealing the full sweeping curves of the building against a moody overcast sky, a reflecting pool at its base catching the grey light. No people. Architecture speaks alone. Abstract ambient score implied, professional architectural photography movement, 12 seconds.
3D animated short opening: a tiny mouse in an oversized explorer hat stands at the entrance to a hole in a skirting board — beyond it, a vast luminous world of giant fungi and glowing beetles. She adjusts her hat, consults a hand-drawn map, and steps in. The camera follows at mouse-height. Pixar production quality, warm magical lighting, whimsical orchestral score building, charming and adventurous, 10 seconds.
Tight observational shot of a politician at a press conference before cameras go live: she's alone at the podium adjusting her notes, a media adviser whispers something urgent, she nods once, straightens, the press room begins to fill behind her, camera flash tests pop. Not a word is said. The tension of what's about to happen is entirely visual. Handheld, natural fluorescent lighting, ambient crowd murmur, 10 seconds.
An F1 pit stop at 24 fps with 120fps slow-motion spliced in: the car enters the box at speed — wide shot — the crew descends like a mechanical organism — front jack up, all four guns firing simultaneously, tyres off in under 1.5 seconds, tyres on — then the cut to ultra-slow-motion of a single lug nut spinning off through the air, the tyre being slammed on, the gun tightening, the front jack dropping — back to real speed as the car screams away. 10 seconds, pure technical theatre.
A solarpunk village in 2080: rounded adobe buildings with living roofs of wildflowers, wind turbines integrated into giant sculptural tree forms, children playing near a communal garden harvesting tomatoes, solar panels shaped like giant leaves on market stall awnings, crystal-clear river running through the village centre. Bright afternoon light, genuine human warmth, no dystopia, pure hopeful future vision, painterly realism, 12 seconds.
Reference-to-video: use character1 (female founder, blazer, dark hair) and character2 (male engineer, hoodie, glasses) from uploaded reference images. Scene: both walk through a sunlit open-plan office, character1 explains something on a whiteboard, character2 nods, then they turn to camera together and smile. Camera: smooth steadicam follow, golden-hour window light, shallow depth of field on faces. Dialogue with natural pacing and pauses. Brand anthem music fading in at the end, 12 seconds. Apple keynote campaign quality.
Close-up of a woman in her 30s sitting on a park bench at dusk, speaking directly to camera: 'I spent years pretending everything was fine, and then one morning I just... stopped.' Her voice cracks slightly on 'stopped,' she pauses, looks away, then back with a half-smile. Shallow depth of field, bokeh city lights behind her, handheld intimacy, natural skin texture with visible pores and freckles. Audio: her voice with natural speech-rate variation and emotional tone shifts, distant city ambient, no background music. 15 seconds. A24 film quality.
Reference-to-video: use product1 (wireless earbuds in charging case) from 4 uploaded reference angles. A pair of hands lifts the case from a minimalist desk, opens it slowly — macro shot of the hinge mechanism, then the earbuds nestled inside. One earbud is lifted out, the LED pulses blue. Cut to the earbud placed in an ear — profile shot, natural skin rendering, soft studio sidelight. ASMR audio: the satisfying click of the case, the gentle snap of the earbud seating. No music. 12 seconds. Samsung Galaxy campaign quality.
Split-screen documentary: three people on a busy Tokyo street each answer the same question in different languages — Japanese, English, Portuguese. Each speaker's lips sync perfectly to their words. Left panel: elderly man in a cap, warm smile. Centre: young woman with headphones. Right: teenage boy mid-laugh. Handheld, natural street lighting, background pedestrians blurred. Each speaker has natural speech pacing — the elderly man speaks slowly, the woman is animated, the boy is rapid-fire. Authentic street ambient audio layered beneath speech, 15 seconds.
Extreme macro portrait of a 60-year-old craftsman's face: camera drifts slowly from his weathered left hand holding a chisel up to his face — every pore visible, crow's feet deepening as he squints at his work, salt-and-pepper stubble catching the raking afternoon light from a workshop window. No oily sheen, no over-sharpening — naturalized skin with subsurface scattering. He exhales, sawdust particles drift through the light beam. Complete silence except wood shavings falling, 10 seconds. Magnum Photos intimacy.
The HappyHorse prompt generator on this page provides 25 free, professionally crafted prompts for HappyHorse 1.1, the AI video model developed by Alibaba (Taotian Group) that has held the #1 position on the Artificial Analysis Video Arena leaderboard since April 2026. Version 1.1, released June 22 2026, adds 9-image reference-to-video, multilingual lip-sync, and naturalized skin rendering. Each prompt is built for HappyHorse's motion realism, native audio generation, and cinematic narrative capability.
HappyHorse is an AI video generation model developed by Alibaba's Taotian Group. Version 1.0 launched in April 2026 and immediately ranked #1 on the Artificial Analysis Video Arena leaderboard with an Elo of 1,341 for text-to-video and 1,402 for image-to-video. It initially appeared anonymously, sparking viral speculation about which company built it. Version 1.1 was released on June 22 2026 with major upgrades to visual quality, audio, and multi-character reference capabilities. HappyHorse is available on the official HappyHorse site, Alibaba Cloud Model Studio (Bailian), and the Qwen app.
HappyHorse 1.1, released June 22 2026, brings five major upgrades: (1) Improved motion modeling — runs, jumps, and collisions feel physically plausible; (2) Natural skin rendering — eliminates the oily facial look and over-sharpening from 1.0, preserving pores and soft skin qualities; (3) Reference-to-Video (R2V) — upload up to 9 reference images and tag them as character1, character2, etc., fusing identity, wardrobe, and brand elements into a coherent clip; (4) Enhanced audio — natural speech-rate variation, emotional tone, controllable background music, and improved lip-sync; (5) Multilingual dialogue lip-sync across multiple languages.
HappyHorse went viral because it appeared anonymously at the top of the global AI video benchmark leaderboard, outperforming every known model overnight. The mystery of who built it — combined with the unusual name — drove massive speculation on X and Reddit's AI communities. When Alibaba was revealed as the creator on April 10 2026, it confirmed that the race for AI video supremacy was far more competitive than many assumed, with Chinese companies matching or exceeding Western models across key benchmarks.
HappyHorse 1.1 responds best to structured, cinematic language. For best results: (1) Open with a shot description — camera type, angle, and movement; (2) Describe the subject and action with specific physical detail; (3) Specify the environment, time of day, and lighting; (4) Include audio direction — HappyHorse generates native sound with emotional tone, so describing ambient audio and speech pacing significantly improves the result; (5) For multi-character scenes, use the R2V tag format: 'character1 (description)' referencing your uploaded images; (6) End with a production quality reference and duration hint (8–15 seconds).
As of June 2026, HappyHorse 1.1 remains #1 on the Artificial Analysis Video Arena leaderboard, ahead of Kling 3.0, Seedance 2.0, and Veo 3.1. HappyHorse's unique advantages are 9-image reference-to-video for multi-character consistency, naturalized skin rendering, and emotional audio with lip-sync. Kling 3.0 leads the arena.ai text-to-video leaderboard with strong structured multi-shot sequencing. Veo 3.1 remains Google's safest overall pick for native-audio video. Seedance 2.0 (via Dreamina) leads text-to-video-with-audio on the Artificial Analysis benchmark.
HappyHorse 1.1 is accessible via the official HappyHorse site, Alibaba Cloud Model Studio (Bailian), and the Qwen app — all went live on June 22 2026. The previous 1.0 version remains available via fal.ai and Alibaba Cloud. Pricing and access tiers vary by platform — check each directly for current options. Creative Fabrica has also announced early access to HappyHorse 1.1.
Kuaishou's top AI video model — multi-shot sequences
ByteDance's quad-modal video model with native audio
Open-source AI video with synchronized audio
Google's AI video model — photorealistic with native audio
xAI's #1 image-to-video model — 15s clips with audio
ByteDance's consumer AI video tool in CapCut