- HappyHorse
- AI Video
- 1.1 Update
- Tutorial
Alibaba's "HappyHorse" Gets Faster: HappyHorse 1.1 Officially Released
HappyHorse 1.1 launched June 22 with upgrades across motion, consistency, prompt following, visual quality, and audio. Here's what changed and how to try it.

On June 22, Alibaba officially released HappyHorse 1.1, its native multimodal AI video generation model. Less than two months after the 1.0 gray launch, this update targets real production pain points: stiff motion, face drift, prompts that get ignored, and over-sharpened “oily” frames.
If you’ve been searching for a happyhorse tutorial or comparing happyhorse video tools, this guide walks through what 1.1 actually changed.
Where 1.0 left off
Since HappyHorse 1.0 opened for gray testing on April 27, the model has been adopted across short drama, e-commerce ads, brand marketing, and game CG pipelines.
The core architecture stays the same: a 15-billion-parameter single-stream Transformer that encodes text, image, video, and audio in one sequence for native audio-video generation. Version 1.0 topped Artificial Analysis Video Arena text-to-video and image-to-video rankings — a high bar for 1.1 to beat.
What 1.1 upgrades: five dimensions
| Dimension | Key changes |
|---|---|
| Dynamic expression | Better motion modeling and temporal consistency |
| Subject consistency | Up to 9 reference images at once; stronger R2V fusion |
| Instruction following | Long-context understanding; multi-shot and N-grid references |
| Visual quality | Less oily look and over-sharpening; more skin texture detail |
| Audio expression | Improved native joint audio-video output |
Motion: less “stop-motion” feel
Stiff character motion and disconnected camera rhythm were among the most common 1.0 complaints. Version 1.1 improves motion modeling and frame-to-frame coherence — runs, jumps, and object collisions feel more physically plausible.
For happyhorse video workflows, that means fewer re-rolls just to get usable action.
Visual fidelity: less plastic skin
Creators often noted 1.0 frames could look overly smooth or harshly sharpened. Version 1.1 addresses this directly:
- Reduced oily sheen and excessive sharpening
- Preserved pores, fine lines, and natural skin texture
- Better micro-details: sweat, lighting shifts, debris on impact
Short drama and ad workflows benefit most from these changes.
Multi-image R2V: lock identity with up to 9 refs
Reference-to-Video (R2V) is a headline upgrade. Upload up to 9 reference images and tag them in your prompt as character1, character2, etc. The model fuses identity, wardrobe, and brand elements into a coherent clip.
| Use case | What 1.1 R2V delivers |
|---|---|
| Product ads | Stable logos and packaging details |
| Short drama | Same character across multiple shots |
| Multi-shot scripts | Better N-grid reference parsing and scene planning |
Prompts that actually stick
Long prompts with multiple scenes and characters used to confuse the model. Version 1.1 improves instruction following and shot planning — the output tracks your intent more reliably.
Where to access it
HappyHorse 1.1 is live on the official site, Alibaba Cloud Model Studio (Bailian), and the Qwen app. Enterprise teams and developers can call the full capability set via API.
Try it yourself
Ready to test HappyHorse 1.1 happyhorse video generation? Use the button below to open the official app and pick the video generation tool.
New users typically get free trial credits — run the same prompt on 1.0 and 1.1 to compare.
Bottom line
HappyHorse 1.1 is a production-driven upgrade, not a vanity version bump. Smoother motion, stabler faces, and stronger multi-image R2V make it worth a side-by-side test if you’re already on 1.0 — and a good entry point if you’re new.