Back to blog
Prompt engineeringMay 5, 2026 · 8 min read

Reference Images: The Cheat Code Almost Nobody Uses

Type the same detailed prompt twice and the output still won’t match exactly. A face shifts a little, a logo redraws itself slightly wrong, a jacket picks up a different collar. None of that is a flaw in the wording. A pure text prompt has no memory between generations, so every run is the model’s best fresh guess at your words. A reference image breaks that cycle, and most people never touch the upload button long enough to find out.

What a reference image actually locks

Four things benefit the most from a reference instead of a description: a specific face, an exact outfit or pattern, a product’s real shape, and a brand’s precise colors. All four share the same problem in words alone, easy to describe approximately and almost impossible to describe exactly. A photo sidesteps the whole exercise.

Text only, described three times

Take 1
Take 2
Take 3

Text + one reference image

Take 1
Take 2
Take 3

Same words, three takes, three slightly different results. Add one reference photo and the target stops moving.

One reference beats three adjectives

Chasing consistency through wording tends to mean piling on adjectives: same red jacket, same exact red, cherry red, not orange-red. None of it reliably narrows the color the way a single photo of the actual jacket does. The model isn’t short on descriptive vocabulary. It’s short on a fixed target to aim at, and a reference image is that target.

Multiple references, one scene

A single reference locks one thing. Real shots often need more than one thing locked at once, a specific face and a specific product in the same frame, say. Most tools that support references let you hand over a small handful at once, each one tagged with a short name, then call each one out by that name inside the prompt instead of re-describing it: the woman from @face holds up the bottle from @bottle. The tag does the pointing, so the prompt stays about the action and the framing instead of turning into a redundant physical description of something already sitting right there in the reference.

Skip this

a woman with brown hair holds up a green glass bottle with a white label

Try this

the woman from @face holds up the bottle from @bottle, tilting it toward the light

Where reference images fail

A reference photo won’t fix physics, and it won’t settle a fight between two conflicting instructions in the same prompt. It also needs to be a real, cleanly hosted photo, not a screenshot forwarded through three messaging apps and compressed into mush along the way. And it’s not unlimited: hand over one clear reference for the one detail that actually needs to be exact, rather than several loosely related images and hoping the model reconciles all of them into a coherent scene.

Tag it instead of describing it

Once a reference is in the chat, refer back to it by what it is instead of re-describing it in the prompt: this jacket, the same face, keep the logo as shown. Re-typing a full description alongside an uploaded photo just gives the model two slightly different instructions to reconcile, which reintroduces the exact drift the reference was supposed to remove.

A reference is a floor, not a ceiling

None of this means the wording around a reference stops mattering. A reference locks what a photo can show, a face, a color, a shape, but it says nothing about motion, mood, or camera work, which still need describing the way they would in any other prompt. Treat the image as the answer to one specific question and the prompt as the answer to everything else, rather than expecting one photo to carry a whole scene on its own.

Say it, and Zo makes it.

Chat with Zo