AI Hug Prompt Guide: Copy-Paste Templates for the Embrace

AI hug prompt templates that work: the base two-photo embrace recipe, every variable worth swapping, and the fixes for broken hands and melted faces.
Oct 2, 2026

The difference between a hug clip that lands and one that gets scrolled past is rarely the model — it is the prompt. Image-to-video models stage an embrace reliably when the prompt tells them four things: who is walking toward whom, where it happens, how the camera behaves, and what the moment feels like. Skip one and the model invents it, which is where melted faces and extra people come from.

This page is the prompt cookbook for the AI hug video trend. Every template here is original and built to be pasted straight into the generator — switch to Custom Scene, upload your two photos, paste, render.

The base template

Vertical 9:16 image-to-video of a heartfelt embrace.
Subject A from the first photo walks in from the left, Subject B from the second photo walks in from the right.
Both keep the exact face, hairstyle, clothing, and skin tone from their input photos; no face blending, no identity swap.
Setting: a sunlit park path with soft morning light.
Camera: steady medium shot, gentle push-in as they close the distance, no cuts.
Action: they meet at center and share a warm, lingering embrace, natural body language, genuine emotion.
Mood: heartwarming and tender.
Avoid extra people, background movement, or camera shake.

Three lines in that recipe do the heavy lifting, and they are the ones people usually leave out:

  • The identity lock. "Keep the exact face... no face blending" is what stops the model from averaging two faces into a stranger. Keep it even though it feels redundant.
  • The approach, not just the hug. "Walks in from the left... meet at center" gives the model a story beat. Prompts that only say "two people hugging" produce a static merge; prompts with an approach produce the moment that makes viewers feel something.
  • The negative line. "Avoid extra people" earns its place — open settings are where the model likes to add bystanders.

Swap one variable at a time

Change a single bracket per generation so you can see what each edit does.

Hug style — the emotional dial:

  • warm gentle embrace (the default; works for every pairing)
  • tearful reunion (add "eyes closed, holding on tight")
  • big bear hug (add "lifting slightly off the ground")
  • forehead-to-forehead pause before the embrace (slower, quieter, extremely effective for memorial clips)

Setting — the context dial:

  • sunlit park path, family living room, train platform, hospital garden
  • airport arrivals hall (the reunion classic; add "arrivals board soft in the background")
  • rainy street under one umbrella, autumn sidewalk with falling leaves

Camera — the craft dial:

  • steady medium shot with gentle push-in (the default; keeps both faces readable)
  • over-the-shoulder as the gap closes (more cinematic, riskier for faces)
  • slow lateral track (elegant but doubles the chance of hand artifacts)

Mood — the finishing dial: heartwarming, bittersweet, joyful, quiet. One word is enough; stacking mood words makes the model overact.

Two classic failure modes, and the prompt-level fix

Broken hands. Hands are the hardest thing in the frame during an embrace. Three fixes, in order of effectiveness: ask for "a gentle embrace, arms around each other's back" (hands hidden, nothing to break); shorten the clip; or crop the photos so the subjects enter as upper bodies with arms already visible and symmetrical.

Melted or blended faces. Almost always a photo problem wearing a prompt costume: mismatched scale, profile shots, or heavy filters. The prompt-side patch is to strengthen the identity lock ("identical to the input photos, photorealistic"), but if faces still drift, re-shoot the input — front-facing, well-lit, both subjects at similar distance from the camera. The trend guide covers photo-picking in detail.

The one-photo variant

Working from a single photo that already has both people in frame? Same template, minus the two-subject lines:

Vertical 9:16 image-to-video of a heartfelt embrace.
The two people in the photo turn toward each other and share a warm embrace, keeping their exact faces and clothing from the input photo.
Setting: the place in the photo, soft natural light.
Camera: steady medium shot, gentle push-in, no cuts.
Mood: heartwarming and tender.
Avoid extra people, background movement, or camera shake.

The approach beat is shorter — a turn instead of a walk — but it is still there, and it still matters.

Red lines that live in the prompt stage

The prompt is where compliance starts, not where it ends. Keep these rules from the first render: only use photos you have the right to use, of people who would consent to the result; never prompt a celebrity or public figure into an embrace clip — likeness is the fast lane to a takedown; label the output as AI-generated when you post. The full list lives in the trend guide's red-lines section.

Frequently asked questions

How long should an AI hug prompt be?

Long enough to cover identity, approach, setting, camera and mood — the base template above is about 90 words and that is the sweet spot. Shorter prompts outsource decisions to the model; much longer prompts dilute the lines that matter.

Do I need to write "vertical 9:16" in the prompt?

The generator renders vertical 9:16 by default, so it is belt-and-suspenders — but harmless, and it stops the model from composing a wide cinematic frame if you ever run the prompt elsewhere.

Can I use these prompts in other tools?

Yes — the recipe is tool-agnostic. It works in any image-to-video model that accepts two reference images. The generator is simply the shortest path from the template to a finished, watermark-free clip.

My embrace comes out stiff and robotic. What do I change?

Add motion words to the approach and the contact: "steps quickly, arms open", "wraps around and sways slightly". Hugs read as real when something moves after contact — a pat on the back, a slight sway — so give the model one small post-embrace action.