The difference between a hug clip that lands and one that gets scrolled past is rarely the model — it is the prompt. Image-to-video models stage an embrace reliably when the prompt tells them four things: who is walking toward whom, where it happens, how the camera behaves, and what the moment feels like. Skip one and the model invents it, which is where melted faces and extra people come from.
This page is the prompt cookbook for the AI hug video trend. Every template here is original and built to be pasted straight into the generator — switch to Custom Scene, upload your two photos, paste, render.
Vertical 9:16 image-to-video of a heartfelt embrace.
Subject A from the first photo walks in from the left, Subject B from the second photo walks in from the right.
Both keep the exact face, hairstyle, clothing, and skin tone from their input photos; no face blending, no identity swap.
Setting: a sunlit park path with soft morning light.
Camera: steady medium shot, gentle push-in as they close the distance, no cuts.
Action: they meet at center and share a warm, lingering embrace, natural body language, genuine emotion.
Mood: heartwarming and tender.
Avoid extra people, background movement, or camera shake.Three lines in that recipe do the heavy lifting, and they are the ones people usually leave out:
Change a single bracket per generation so you can see what each edit does.
Hug style — the emotional dial:
Setting — the context dial:
Camera — the craft dial:
Mood — the finishing dial: heartwarming, bittersweet, joyful, quiet. One word is enough; stacking mood words makes the model overact.
Broken hands. Hands are the hardest thing in the frame during an embrace. Three fixes, in order of effectiveness: ask for "a gentle embrace, arms around each other's back" (hands hidden, nothing to break); shorten the clip; or crop the photos so the subjects enter as upper bodies with arms already visible and symmetrical.
Melted or blended faces. Almost always a photo problem wearing a prompt costume: mismatched scale, profile shots, or heavy filters. The prompt-side patch is to strengthen the identity lock ("identical to the input photos, photorealistic"), but if faces still drift, re-shoot the input — front-facing, well-lit, both subjects at similar distance from the camera. The trend guide covers photo-picking in detail.
Working from a single photo that already has both people in frame? Same template, minus the two-subject lines:
Vertical 9:16 image-to-video of a heartfelt embrace.
The two people in the photo turn toward each other and share a warm embrace, keeping their exact faces and clothing from the input photo.
Setting: the place in the photo, soft natural light.
Camera: steady medium shot, gentle push-in, no cuts.
Mood: heartwarming and tender.
Avoid extra people, background movement, or camera shake.The approach beat is shorter — a turn instead of a walk — but it is still there, and it still matters.
The prompt is where compliance starts, not where it ends. Keep these rules from the first render: only use photos you have the right to use, of people who would consent to the result; never prompt a celebrity or public figure into an embrace clip — likeness is the fast lane to a takedown; label the output as AI-generated when you post. The full list lives in the trend guide's red-lines section.
Long enough to cover identity, approach, setting, camera and mood — the base template above is about 90 words and that is the sweet spot. Shorter prompts outsource decisions to the model; much longer prompts dilute the lines that matter.
The generator renders vertical 9:16 by default, so it is belt-and-suspenders — but harmless, and it stops the model from composing a wide cinematic frame if you ever run the prompt elsewhere.
Yes — the recipe is tool-agnostic. It works in any image-to-video model that accepts two reference images. The generator is simply the shortest path from the template to a finished, watermark-free clip.
Add motion words to the approach and the contact: "steps quickly, arms open", "wraps around and sways slightly". Hugs read as real when something moves after contact — a pat on the back, a slight sway — so give the model one small post-embrace action.