Building a Character Zo Can Remember Across Videos

A character who looks slightly different in every clip breaks a series fast. The face drifts a little, the jacket changes color a shade at a time, and by the fourth video nobody would recognize clip one as the same character. Fixing it comes down to two habits: locking the wording of a description, and handing over one clean reference photo instead of re-describing a character from memory every time.
Describe once, reuse verbatim
The wording of a character description matters almost as much as a reference image does. A phrase like a fox with one white ear and a red scarf typed slightly differently each time, a fox, white-tipped ear, wearing red versus a red-scarfed fox with a white ear, reads as two different characters to a model with no memory of the first one. Save the exact phrase that worked and paste it in whole every time, instead of re-describing the character loosely from memory.
One clean reference beats a paragraph
The same logic covered in a closer look at reference images applies here even more directly: a single clear photo of the character locks the face and outfit in a way no amount of careful wording matches. For a character appearing across many clips, that one reference is worth setting up properly once, rather than re-explaining the same details in every new chat.
Tag the character the same way you’d tag a reference
The tagging trick that works for locking a face or a product in a single shot works even better for a character meant to recur. Upload the reference once, give it a short tag, fox, and call the character by that tag in every new chat instead of retyping the full description: the fox from before chases a paper boat down a gutter. The tag carries the visual details forward on its own, so the wording only has to cover what’s different about that particular clip.
Vary the shot, not the character
Keep the character’s description fixed across a series and change only the action, setting, and camera around them. A fox with one white ear and a red scarf trots through a market square, a fox with one white ear and a red scarf naps under a bridge, same fixed block of words, different scene attached to it. Rewriting the character description even slightly between clips is usually where a series starts to drift.
When the face still drifts
Even with a reference image, fine facial geometry is the detail most likely to shift a little between generations, since it’s the hardest thing in the frame to pin down exactly. A bold, easy-to-render trait, a specific scarf color, a distinctive ear marking, a signature accessory, survives the trip between clips far better than subtle bone structure does. Give a character one loud, simple visual anchor and it stays recognizable even on the takes where the face itself isn’t a perfect match.
A series survives cuts better than one long story
A running character doesn’t need a continuous plot to feel like a series. Six independent clips of the same fox in six different situations reads as a series just as well as one long connected story would, and each clip still only has to hold one clear beat rather than carrying its share of a bigger plot. Treat a character series the same way as any other clip on its own, one action, one setting, one small change from start to end, and let the character’s consistency be the thread that ties the separate pieces together.
Starting a series? Send me the reference photo once and just say 'same fox as before' in every new chat. I'll pull it back up instead of you re-uploading it each time.