Why your character's face keeps changing between shots
A teardown of the drift every AI filmmaker hits by shot four, and the written anchors that stop it.
6 September 2026 · Written by Fares Walid Mansour, FWM
Shot one is the woman you saw in your head. By shot four she has a different jaw, a different age and a different nose, and the sequence stops being one film.
The symptom
You write a woman on a rooftop. Shot one lands and it is her. You change the angle for shot two and she is somebody else: the jaw is narrower, she has aged six years, the nose is new.
By shot four you are not directing a film, you are auditioning strangers. You raise the weight on your reference image, you re-roll, you burn another twenty generations, and the face still slides.
The model is not failing at rendering. It is doing exactly what your words asked. Your words described a type of woman, so it cast a new one every take.
The prompt that drifted
a beautiful young woman standing on a rooftop at night, city lights behind her, cinematic, moody, highly detailed, 8k, dramatic lighting, film still
Read it as a casting notice rather than a shot. It asks for a category, not a person. "Beautiful", "young" and "woman" describe a set with millions of members, and each generation picks a different one.
The four words at the end make it worse. "Cinematic", "moody", "highly detailed" and "8k" spend your prompt's attention on adjectives that every image already claims, and leave nothing for the face.
Why it drifts
A diffusion model has no memory of shot three. Each generation starts from noise and reads your text as the only description of the world. Nothing carries over unless you write it again.
So identity has to be re-established every shot, in language specific enough to name one person. A face is held by structure, not by beauty: bone, proportion, hairline, the details that survive a change of angle and light.
Age is the most common leak. "Young" spans fifteen years, and the model will wander across all of it. A number pins the face where a mood word cannot.
The second leak is the light. A face lit from below at 3200K and the same face in cold moonlight read as two different people, even when the bones are identical. Continuity of light is continuity of face.
The prompt that holds
Medium shot, 85mm, eye level, shallow depth of field. SUBJECT: Mira, 34, Lebanese, oval face, strong straight eyebrows, a small vertical scar through the left eyebrow, wide-set dark brown eyes, aquiline nose, black hair pulled back tight with a centre part, two loose strands at the left temple. No makeup except a matte dark red lip. WARDROBE: oversized charcoal wool overcoat, collar up, black crew-neck beneath, one silver ring on the right index finger. BLOCKING: standing three-quarters to camera, weight on the left leg, looking off-frame right, chin level. LIGHT: single warm practical from frame left at 3200K, raking across the left cheek, the right side of the face falling into shadow with no fill. Cold blue city haze behind her, 6000K, well out of focus. STOCK: Kodak Vision3 500T, visible grain, gentle halation on the practical. NEGATIVE: no smiling, no makeup change, no jewellery beyond the ring, no second light source on the face, no wind in the hair, no age change, no beauty retouching.
Nothing here is decoration. Every line is a decision that would otherwise be made for you, differently, on every generation.
The anchors, named
A continuity anchor is a written detail that has to be identical in every shot of the sequence. Pick few, pick specific, and repeat them word for word. Six is usually enough for a face.
Hold these
- Name and exact age: Mira, 34. Never "young", never a range.
- One irregular mark: the scar through the left eyebrow. Irregularity is what the model cannot average away.
- Face structure in three terms: oval face, wide-set eyes, aquiline nose.
- Hair as a mechanism, not a colour: pulled back tight, centre part, two loose strands at the left temple.
- One wardrobe object that recurs: the charcoal overcoat with the collar up.
- Key light direction and temperature: frame left, 3200K, no fill.
Copy that block into every shot in the sequence unchanged. Then write only what changes: the angle, the lens, the blocking, the beat. The face stops being a variable.
One irregular detail is worth ten flattering ones. Models are trained toward the average face, and an average has no identity. The scar, the crooked tooth, the uneven hairline are what hold.
The same anchors, four dialects
The anchors do not change between models. The way you deliver them does, because each model reads a prompt differently.
- Midjourney
- Front-load the identity clause, keep the whole prompt under about 60 words, and lock the face with a character reference and a fixed seed. Style words go last, where they cannot outrank the bones.
- Sora
- Write it as a shot description in full sentences, subject before camera before light. Repeat the identity clause verbatim in every shot of the sequence, and never paraphrase it, because a paraphrase is a new person.
- Veo
- Use labelled fields, one line each: subject, wardrobe, action, camera, lighting, ambience. The label structure is what Veo holds onto; prose in one long block loses the wardrobe first.
- Kling
- Anchor the first frame, then write only the motion. Kling drifts most across long moves, so keep each generation short and cut, rather than asking one clip to hold a face for ten seconds.
Reference pages change every few months. Treat the flags and field names your model documents today as the current dialect, and the anchors above as the part that outlives them.
What to do tonight
- Write the identity clause once, in six anchors, before you generate anything.
- Give the character a name and a number for an age.
- Add one irregular mark that no average face has.
- Fix the key light direction and its colour temperature for the whole sequence.
- Paste the clause into every shot unchanged, and write only the difference beneath it.
- Put drift in the negative list by name: no age change, no makeup change, no second key light.
That is the whole fix. It is not a setting, it is a document, and it is why a sequence written as a sequence holds when eleven separate prompts do not.