A selfie reference is the photo you provide for a model to generate portraits from. In reference-based tools it is used directly at generation time, rather than being consumed in a training run beforehand.
It is the highest-leverage input in the whole process. Everything downstream inherits its limitations.
What makes a good reference
- Face clearly visible and in focus. Blur cannot be recovered.
- Even, soft lighting. Harsh shadows across the face get interpreted as facial structure.
- Neutral or slight expression. An extreme expression propagates into every output.
- Facing roughly forward. Steep angles hide geometry the model then has to invent.
- No sunglasses, no heavy filters. Filters have already altered the features being read.
- Recent. The output will look like the photo, not like you now.
What does not matter
Background, clothing, and setting are all replaced during generation. A reference taken in a messy room in a t-shirt is completely fine — those are exactly the elements being changed.
This surprises people, who often hunt for their most flattering existing photo. The best reference is usually a plain, well-lit, forward-facing shot taken thirty seconds ago against a blank wall.
Why one good reference beats several mediocre ones
With reference-based generation, a single sharp, evenly lit photo outperforms a handful of dim or angled ones. Ambiguity in the input becomes invention in the output — and invention is where identity retention breaks down.
Practical advice
Stand facing a window, hold the phone at eye level, do not smile hard, and take it now. That beats searching your camera roll.
Related: what is an AI headshot and what is a headshot pack.