The Wan 3.0 prompt generator gives you 20 free, copy-paste video prompts for Alibaba's new 30-second AI video model. Text, image, audio, video, and document-to-video — structured for Wan 3.0's multimodal input and single-shot long-form generation.
The Wan 3.0 prompt generator on this page provides 20 free, professionally crafted video prompts for Wan 3.0 — Alibaba Cloud's AI video model that entered public beta on August 6, 2026. Wan 3.0 is the successor to Wan 2.7 and makes a significant leap: single-shot videos up to 30 seconds long, a broader multimodal input stack, and a unique document-to-video feature that accepts PDFs, PowerPoint files, and spreadsheets as creative references.
Where Wan 2.7 was known for its open-source model weights and Thinking Mode, Wan 3.0 focuses on production-grade hosted output via the DashScope API at $0.05–$0.20 per second of generated video. The 30-second single-shot duration puts Wan 3.0 ahead of most competing models for long-form content — documentary sequences, product demo videos, and data visualisations — without the need to stitch shorter clips together.
Every prompt below is structured for Wan 3.0's multimodal strengths: clear camera direction, scene sequencing, audio description, and explicit duration within the 30-second maximum. Paste directly into the DashScope API or any integrated platform. No signup required for the prompts on this page.
Wan 3.0 generates best from structured, descriptive prompts. Use this formula:
Click any prompt to copy — paste directly into the DashScope API or any Wan 3.0 integrated platform
A cinematic product reveal sequence for a sleek matte-black smart device: the camera begins in darkness, then a single dramatic light source — cool white, surgical — illuminates the device on a black plinth from above. Over 15 seconds the camera slowly orbits 90 degrees around the device at eye level, the surface reflecting the light source cleanly. In the final 5 seconds, a hand reaches in from frame left and lifts the device to reveal its screen glowing with a minimal interface. 4K, ultra-premium product aesthetic, 20 seconds total. No text.
A compressed time-lapse of a major city intersection at morning rush hour: pedestrians stream across crosswalks in overlapping flows, taxi doors open and close in rhythmic succession, a street cart vendor serves coffee, cyclists weave through standing traffic. The camera holds at mid-height on a wide angle, capturing the full intersection for 25 seconds at 10x speed. Natural urban audio — horns, footsteps, the mechanical click of crossing lights. 16:9.
A slow drift through a bioluminescent deep-sea environment at 1,000 metres: the camera moves forward at 0.3 mph through absolute darkness broken only by pulsing blue-green organisms — jellyfish trailing 2-metre tendrils, a school of lanternfish spiraling in formation, a deep-sea anglerfish hovering motionless with its lure swaying gently. No sound except the ambient low-frequency pressure of deep water. National Geographic underwater quality, 25 seconds, 16:9.
A series of close-up shots edited in slow rhythm — hands rinsing a chawan with a bamboo ladle, matcha powder scooped from a lacquer caddy, a whisk working the tea into a fine green froth, steam rising from the finished bowl, the bowl passed between two sets of hands. The camera cuts methodically between these macro moments over 25 seconds. Soft diffused natural light through shoji screens, complete quiet except for water sounds. 16:9. No text.
An aerial drone perspective at 200 metres altitude looking down on a wildfire advancing through dry pine forest: the fire line moves left to right across the frame as winds push it, smoke columns rising in twisting spirals, retardant aircraft making a drop run from upper left to lower right dropping a red chemical line that steams against the fireground. The drone holds position for 25 seconds, camera tilted 45 degrees down. Broadcast documentary quality, 16:9. No text overlays.
A slow first-person walk through a monumental brutalist university campus: concrete columns rise 20 metres on either side, the camera moves at walking pace through a covered colonnade, the geometry shifting as each column passes — light and shadow alternating in rhythmic repetition. In the distance, a courtyard opens with a reflecting pool. Students are present but blurred in soft background. 25 seconds, 16:9. Grey morning light, slight overcast. No text.
A 25-second sequence inside a professional kitchen during dinner service: a chef's knife breaks down a whole fish with clean, rapid strokes — the camera starts on extreme close-up of the blade at work, then cuts to a wide of the prep station where four chefs work in synchronised silence, then close again on a sauce being finished with cold butter, swirled until glossy. Stainless steel, harsh directional work lighting, the rhythm of knives and ladles as ambient sound. 16:9. No text.
A 30-second slow-motion storm sequence over a flat wheat prairie at dusk: a wall of cumulonimbus cloud advances from the horizon, lightning bolts striking the plain at 5-second intervals in slow motion — each strike illuminating the full prairie in white for a half-second. The wheat ripples violently as the gust front arrives. Thunder as low-frequency ambient audio. The camera holds static at 16:9 wide, tripod-mounted, ground level, capturing the full sky-to-ground drama. No text.
A cinematic exploration of an abandoned 1960s factory interior: the camera moves slowly through the production floor, shafts of light cutting through broken skylights and illuminating floating dust particles, corroded machinery still positioned at workstations, a clock stopped at 4:17. The camera passes a row of lockers, one hanging open with a calendar still on the wall. 25 seconds, slow tracking shot at hip height, natural ambient audio — wind through broken windows, distant birds. 16:9. No text.
A real-time star-count sequence in the Atacama Desert: the camera looks straight up at the Milky Way core from a dark-sky site, the galactic centre arching diagonally across the frame. Over 25 seconds, stars visibly move in slow drift as Earth rotates — a meteor trail crosses the upper right corner at the 15-second mark. The Milky Way's dust lanes are visible. Natural silence except for sparse desert insect sounds. 4K, 16:9. No artificial light.
A dancer rehearses alone in a large studio: the camera is placed low in the corner, a wide shot capturing the dancer in mid-floor performing a grand allegro combination — a sequence of jumps, turns, and a long diagonal finishing in arabesque. Natural rehearsal studio lighting, wooden floor, mirror wall visible at far end showing the dancer's reflection. The camera does not cut or move during the 25-second sequence. Ambient piano from an adjacent room. 16:9. No text.
A fixed wide-angle shot at the end of a subway platform during rush hour, simulating a long-exposure photographic effect: commuters stream past the camera, their bodies motion-blurred into flowing trails, only the platform pillars and the train itself sharp. A train arrives from the far end of the tunnel at the 10-second mark, its headlights growing in the distance, then flooding the platform with light. 25 seconds, 16:9. Ambient station audio — announcements, brakes, crowd murmur. No text.
A potter centres clay on a spinning wheel in a traditional Japanese workshop: the camera begins in a wide of the workshop — wooden beams, tools hanging on walls, a kiln door at the far end — then slowly pushes in over 20 seconds to a close-up of the hands shaping a cylinder from the spinning clay. Water glistens on the clay surface as it rises and is pulled. Ambient audio: the electric hum of the wheel, clay-on-clay sound. 16:9. Warm tungsten workshop light. No text.
An expedition team of three figures in heavy polar gear walks across a sea-ice plain toward camera against an approaching whiteout blizzard: the horizon behind them is a wall of horizontal snow reducing visibility with each second. The camera holds static, the figures growing larger as they approach, the blizzard swallowing the background entirely by the 20-second mark. Howling wind audio that builds in intensity throughout the 25 seconds. 16:9, flat arctic light. No text.
A handheld walk through a packed night market in Bangkok: the camera weaves between vendors — grills smoking, woks tossing flames, coloured string lights overhead, a woman hand-pressing sugarcane juice, children running in the periphery. The audio is live market ambience — Thai pop from a speaker, sizzling fat, the vendor's calls. The camera settles briefly on each scene before moving on. 25 seconds, 16:9. Warm, slightly overexposed market lighting. No text.
Interior B-roll of an Antarctic research station during a winter storm: scientists in fleece layers work at instrument stations, frost visible on the inner window panes, the structure creaking subtly under wind load. The camera moves slowly from the main room through a connecting tunnel to a secondary lab where someone monitors atmospheric data on screens. Natural interior light — fluorescent and monitor glow. 25 seconds, 16:9. Ambient: wind roar outside, ventilation hum inside. No text.
A first-person walk across a glass-bottomed suspension bridge spanning a 300-metre gorge at dusk: each footstep onto the transparent panels reveals the forested gorge floor far below through the glass. The camera looks straight ahead for the first 10 seconds as the bridge sways subtly, then tilts down to show the transparent floor panel and the abyss through it for 5 seconds, then continues forward. Wind audio, distant river far below. 25 seconds, 16:9. No text.
A black-and-white high fashion film sequence: a model in a sculptural black coat moves through a pure white studio in a series of poses that flow from one to the next over 25 seconds — not a runway walk but a choreographed movement study. The camera is stationary but switches between three angles: full body, waist up, close on face and collar. Hard directional light, deep shadows. Complete silence. 16:9. No text.
A pick-up basketball game on an outdoor court at golden hour: the camera tracks a player driving the baseline from a low angle, the hoop and backboard silhouetted against the orange sky. The player rises for a finger-roll layup and the ball banks in — the camera holds on the net's ripple for 2 seconds before cutting to wide of the full court, players calling the next play. 25 seconds, 16:9. Urban court ambient audio — sneaker squeaks on concrete, the ball's hollow bounce. No text.
A camera drifts slowly through a hyperscale data centre at operational speed: row after row of server racks, blue and white indicator lights pulsing in irregular rhythms, cold-aisle/hot-aisle containment panels alternating, the low-frequency roar of cooling infrastructure as ambient audio. At the 15-second mark, a robotic maintenance unit moves along a track at the top of a rack row. The camera holds from the central aisle, a long perspective shot through the centre of the facility. 25 seconds, 16:9. No text.
Wan 3.0 leads for long-form 30-second single-shot video and unique document-to-video capability:
| Model | Max Length | Inputs | Best For | Access |
|---|---|---|---|---|
| Wan 3.0 (Alibaba) ★ | 30 seconds | Text, image, audio, video, documents | Long-form single-shot, data viz, documentary | DashScope API — $0.05–$0.20/sec |
| Kling v3 (Kuaishou) | 30 seconds | Text, image, video reference | Cinematic narrative, photorealism | klingai.com — paid plans |
| Seedance 2.5 (ByteDance) | 30 seconds | Up to 50 multimodal references | Complex multi-ref, synchronized audio | Dreamina / CapCut / API |
| FLUX 3 Video (BFL) | 20 seconds | Text, image, audio reference | Cinematic clips with native audio, BFL image quality | fal.ai / BFL API |
| HappyHorse (Alibaba ATH) | 30 seconds | Text, image | #1 T2V leaderboard, photorealistic motion | happyhorse.ai / API |
| Wan 2.7 (Alibaba) | 10 seconds | Text, image | Open-source, self-hosted, Thinking Mode | Open-source weights / DashScope |
★ Wan 3.0: public beta August 6, 2026. Hosted via Alibaba Cloud DashScope. No open-source weights at time of writing — production API only. Pricing subject to change.
The Wan 3.0 prompt generator on this page gives you 20 free, copy-paste video prompts for Wan 3.0 — Alibaba Cloud's AI video model that entered public beta on August 6, 2026. Each prompt is written to leverage Wan 3.0's 30-second single-shot generation capability and its broad multimodal input — text, image, audio, video, and document references. Paste any prompt directly into the DashScope API or any platform that has integrated Wan 3.0.
Wan 3.0 is the latest AI video generation model from Alibaba Cloud's Tongyi Lab, successor to Wan 2.7. It entered public beta on August 6, 2026, via the DashScope API endpoint. Wan 3.0's headline capabilities include: 30-second single-shot video generation from a single prompt pass, multimodal reference input (text, image, audio, video clips), and a unique document-to-video feature that accepts PDFs, PowerPoint presentations, and spreadsheets as creative references — a capability no other major video model currently offers. Pricing runs from $0.05 to $0.20 per second of generated video depending on output resolution.
Wan 3.0 makes three significant upgrades over Wan 2.7: (1) Extended single-shot duration — Wan 3.0 generates up to 30 seconds in one pass vs. Wan 2.7's 10-second clips, eliminating the need to stitch sequences; (2) Broader multimodal input — Wan 3.0 accepts PDFs, PPTs, and spreadsheets as reference material alongside image and audio inputs, allowing it to generate visualisations of business data and documents; (3) Public API access via DashScope with metered pricing, giving developers and creators direct programmatic access without a waitlist. Wan 2.7 was notable for its open-source model weights and Thinking Mode; Wan 3.0 focuses on hosted, production-grade output quality.
Wan 3.0 responds well to structured, scene-description-first prompts. Use this approach: [Camera position and movement] + [Subject and setting] + [Action sequence in chronological order] + [Lighting and atmosphere] + [Audio description] + [Duration and aspect ratio]. Example: 'A drone shot descending from 100 metres to ground level above a harbour at dawn — fishing boats leave their moorings in sequence, wake patterns expanding in the still water. Amber sunrise light. Ambient: low engine hum, gulls. 25 seconds, 16:9.' Key tips: (1) Specify duration up to 30 seconds; (2) Name camera movement explicitly; (3) Describe audio for scenes where sound matters — Wan 3.0 supports audio input references; (4) For document-to-video use, describe the visual metaphor you want the document data visualised as.
Each model has a distinct strength. Wan 3.0 leads for long-form single-shot video (30 seconds) and is the only model that accepts business documents as creative references — making it uniquely suited for data visualisation and corporate content. Kling v3 (Kuaishou) is the current #1 global T2V model for cinematic narrative quality and photorealism — the benchmark for drama and film-like output. Seedance 2.5 (ByteDance) handles up to 50 simultaneous multimodal references and leads for complex multi-reference video with synchronized audio. FLUX 3 Video (Black Forest Labs) is the newest entrant with native audio, strong for 20-second cinematic clips with BFL's image-quality DNA. Wan 3.0's DashScope API pricing ($0.05–$0.20/sec) is competitive for production volume.
Wan 3.0 is available via the Alibaba Cloud DashScope API — the same platform that hosts Wan 2.7 and other Alibaba AI services. As of August 2026, the public beta endpoint is live with pay-per-second pricing ($0.05–$0.20/sec depending on resolution). Access requires a DashScope account and API key, available at dashscope.aliyuncs.com. Third-party platforms that previously integrated Wan 2.7 (ComfyUI nodes, video generation aggregators) are in the process of adding Wan 3.0 support. Wan 3.0 model weights have not been published as open-source — unlike Wan 2.7's 27B open weights — so self-hosted deployment is not currently available.
Wan 2.7 — open-source 27B, Thinking Mode, self-hosted
#1 T2V globally — cinematic photorealistic narrative
ByteDance — 50 multimodal references, synchronized audio
Black Forest Labs — 20-second clips, native audio
Alibaba ATH — #1 T2V leaderboard, photorealistic motion
Open-source video with synchronized audio in one pass