HappyHorse Multi-Reference Video

Upload 1 to 9 reference images and assign each to a character token in your prompt. HappyHorse maintains identity consistency for every referenced person across the full clip.

Multi-reference generation lets you supply multiple source images — one per character — so the AI model can maintain each person's distinct appearance throughout a video. The model maps each reference to a named token (e.g., [person1], [person2]) in the prompt, binding facial features, body proportions, and clothing to that token. This is fundamentally different from single-reference models that can only preserve one identity, forcing multi-person scenes to hallucinate secondary characters.

What you can do

Up to 9 reference images per generation

HappyHorse accepts 1 to 9 reference images — the highest count in the current AI video model landscape. Each reference binds to a separate character token, so a group scene with 9 distinct people is possible in a single generation.

Character token binding in prompts

Reference images are assigned to tokens like [person1] through [person9]. Use these tokens in your prompt to position and direct each character independently: "[person1] hands a coffee cup to [person3] while [person2] waves from the background."

Cross-character interaction

Because all references are loaded in the same generation pass, characters can interact naturally — handshakes, conversations, passing objects. Single-reference models require compositing separate clips to achieve this.

Consistent identity across duration

Facial features, skin tone, hairstyle, and clothing remain stable from frame 1 through the end of the clip. No mid-clip identity drift, even with camera angle changes or partial occlusions.

Mixed reference types

References can be headshots, full-body photos, or stylized illustrations. HappyHorse extracts identity features regardless of the source image format, though front-facing photos with neutral expressions produce the most accurate results.

Get started

How to use

1

Open Sunra Video and select HappyHorse

Go to Sunra Video and select HappyHorse from the model dropdown.

2

Upload your reference images

Click the reference image upload area and add 1 to 9 images. Each image should clearly show one person's face — front-facing, well-lit, minimal occlusion. Label or note the order (person1, person2, etc.).

3

Write your prompt using character tokens

Reference each uploaded image using its token: [person1], [person2], etc. Describe the scene with specific actions for each: *"[person1] sits at a desk typing while [person2] stands behind pointing at the screen. [person3] enters through the door carrying a folder."*

4

Set duration and aspect ratio

Choose your clip length and aspect ratio. For multi-character scenes, 16:9 widescreen gives more room for character positioning. Longer durations (8–10s) allow more complex interactions.

5

Generate and review identity consistency

Click Generate and check that each character matches their reference throughout the clip. If one character drifts, try a clearer reference photo with better lighting or a more frontal angle.

Built for creators

Whether you're a solo creator, an agency, or a brand — every model adapts to how you work.

F1 pit stop sequence

[person1] in a red racing suit leaps over the pit wall and changes the front-right tire while [person2] in a blue suit handles the rear-left tire. [person3] holds the lollipop sign and drops it as the car launches forward. Overhead camera, 16:9, 8 seconds.

Rollercoaster reality warp

[person1] and [person2] sit side by side in the front row of a rollercoaster. [person1] screams with arms raised while [person2] grips the bar and laughs. The track twists through a surreal neon portal. POV from the seat behind, 16:9, 6 seconds.

Bedroom explosion chaos

[person1] jumps on the bed launching pillows into the air while [person2] ducks behind the door and [person3] catches a flying blanket. Feathers drift everywhere, warm lamplight, handheld camera feel. 16:9, 10 seconds.

Copy & use

Prompt templates

Two-person conversation

[person1] and [person2] sit across from each other at a coffee shop table. [person1] gestures while speaking, [person2] nods and smiles. Warm afternoon light through the window. Shallow depth of field. 16:9, 8 seconds.

Model: HappyHorse · References: 2 images · Duration: 8s · Aspect: 16:9

Team introduction

[person1], [person2], [person3], and [person4] stand in a row in a modern office lobby. Each waves at the camera in sequence from left to right. Clean white background, professional attire. 16:9, 10 seconds.

Model: HappyHorse · References: 4 images · Duration: 10s · Aspect: 16:9

Family dinner scene

[person1] sits at the head of a dining table, [person2] and [person3] on either side, [person4] at the far end. [person1] raises a glass for a toast, others follow. Warm candlelight, rustic wooden table. 16:9, 10 seconds.

Model: HappyHorse · References: 4 images · Duration: 10s · Aspect: 16:9

Product handoff demo

[person1] in a lab coat hands a product box to [person2] in business casual. [person2] inspects the box and nods approvingly. Clean studio background, soft key light. 16:9, 6 seconds.

Model: HappyHorse · References: 2 images · Duration: 6s · Aspect: 16:9

Who it's for

Use cases

Multi-character narrative videos

Short films, web series, and explainer videos with a recurring cast. Upload your character references once and generate consistent scenes across episodes — no continuity errors between shots.

Team and group photo animations

Turn a corporate team photo into an animated introduction video. Upload individual headshots for each team member and prompt a scene where they interact — waving, shaking hands, or presenting together.

Family and event videos

Generate personalized family videos for holidays or celebrations. Upload photos of family members and create scenes like a family dinner, a birthday party, or a group walk in a park — each person recognizable.

E-commerce model consistency

Fashion and lifestyle brands can maintain the same model identity across multiple product videos. Upload the model's reference and generate them wearing different outfits in different settings without rebooking a shoot.

Compare

HappyHorse Multi-Reference vs Other Models

HappyHorse (1-9 references)Other Models
Max reference images9 images per generation — each bound to a separate character tokenKling 3.0: 1 reference. Veo 3.1: up to 3 assets. Seedance 2.0: 1–2 references
Multi-character interactionAll characters rendered in one pass — natural interactions between referenced peopleSingle-reference models require generating characters separately and compositing
Identity binding methodNamed tokens ([person1]–[person9]) in prompts — explicit control per characterMost models use a single implicit reference — no way to direct multiple identities
Group scene qualityEach person maintains their reference identity — no face blending between charactersModels with 1 reference often blend secondary characters' features with the primary
Use case fitBest for multi-person narratives, team videos, family contentBetter suited for single-subject content: portraits, solo product demos, monologues
Get the best results

Tips & best practices

Use clear, front-facing reference photos

Identity extraction works best with well-lit, front-facing headshots or waist-up photos. Side profiles, sunglasses, or heavy shadows reduce matching accuracy. One person per reference image.

Start with 2–3 references before scaling up

More references increase generation complexity. Start with 2–3 characters to validate your prompt structure, then add more. Beyond 5 characters in a single scene, positioning becomes harder to control precisely.

Be explicit about character positions

With multiple characters, vague spatial descriptions lead to crowded or ambiguous compositions. Specify positions: "[person1] on the left, [person2] in the center, [person3] on the right."

Expect diminishing returns above 6 references

While HappyHorse supports up to 9 references, scenes with 7–9 characters leave less visual space per person. Identity accuracy remains high, but individual character detail decreases as the frame gets more crowded.

Community

Loved by creators worldwide

Join thousands of creators, agencies, and brands who use Sunra every day.

The quality jumped overnight

We switched our product video pipeline to Sunra last month. Kling 3.0 with native audio is genuinely usable for social ads now. Our team ships 30+ variations a week without touching After Effects.

MJ
Marcus Johansson
Head of Content, DTC Brand

Finally a tool my whole team can use

I'm technical, my co-founder isn't. She hops into Sunra, types a prompt, and gets a polished video in minutes. Canvas is the killer feature — we brainstorm visually and export straight to pitch decks.

PK
Priya Kapoor
Startup Founder

Cut our pre-production costs in half

We prototype every scene in Sunra before we shoot. Directors see framing, pacing, and mood before a single camera rolls. It's become essential to our pre-vis workflow.

JW
James Whitfield
Production Supervisor

Built our TikTok presence from zero

Brand new account, three videos a day, all on Sunra. Hit 50k followers in four months. The variety lets us test hooks constantly without burning out on production.

SC
Sofia Castellanos
TikTok Creator

Audio quality matches the visuals

Kling 3.0 audio finally feels coherent with the footage. No more awkward mismatched foley. I haven't opened my DAW for social cuts in a month.

TN
Theo Nakamura
Sound Designer

Documentary pre-vis breakthrough

Pre-visualizing reenactments and archival sequences used to cost us 15% of every doc budget. Sunra lets me block scenes for free, then shoot only what matters.

PV
Priya Venkatesan
Documentary Producer
FAQ

Questions & answers

How many reference images can HappyHorse use?
HappyHorse supports 1 to 9 reference images per generation. Each image is bound to a character token ([person1] through [person9]) that you use in your prompt to control each character independently.
How do character tokens work in HappyHorse prompts?
When you upload reference images, each is assigned a token like [person1], [person2], etc. Use these tokens in your prompt to refer to specific characters: "[person1] shakes hands with [person2]." The model maps each token to the corresponding reference face and appearance.
How does HappyHorse compare to Kling 3.0 for multi-character scenes?
Kling 3.0 supports only 1 reference image, making it ideal for single-subject videos. HappyHorse supports up to 9 references, making it the better choice for scenes with multiple identifiable characters.
What kind of reference photos work best?
Front-facing, well-lit photos with a clear view of the face. Headshots or waist-up photos work best. Avoid sunglasses, heavy shadows, or photos with multiple people. See best practices for detailed guidance.
Do all characters maintain consistency for the full clip?
Yes. Each referenced character maintains their facial features, skin tone, and hairstyle from start to finish. Clothing from the reference is also preserved unless you explicitly describe different attire in the prompt.
Can I use multi-reference with video-to-video editing?
Yes. HappyHorse also supports video-to-video editing with up to 5 reference images. You can swap characters in an existing clip while maintaining the original motion and timing.
Is HappyHorse multi-reference free to use?
Free daily credits on Sunra cover HappyHorse generations including multi-reference. Multi-reference does not have a separate surcharge. See pricing for subscription credit amounts.
What happens if two reference images look very similar?
If two references share similar features (e.g., siblings), the model may occasionally blend them. Use references with distinct hairstyles, face shapes, or clothing to help the model differentiate. Explicitly describe distinguishing features in the prompt.

Ready to create?

Start with free daily credits. No credit card required.