The fastest way to get repeatable results out of an image model is to stop writing sentences and start filling slots. A prompt is not prose; it is a small structured record with fields. Once you see the fields, editing an image becomes editing a field instead of rewriting a paragraph.
The slots
A complete image prompt has roughly seven slots. You will rarely fill all of them, but knowing they exist is what lets you diagnose a disappointing output.
- Subject: the one thing the image is about.
- Descriptors: attributes of that subject (clothing, material, age, expression, action).
- Style or medium: the visual language (photograph, oil painting, vector, specific film stock, art movement).
- Composition and framing: shot type, camera angle, subject placement.
- Lighting: source, direction, hardness, time of day.
- Color: palette or mood ("muted earth tones," "high-saturation neon," "monochrome").
- Quality and parameters: rendering cues ("sharp focus," "shallow depth of field") and any model flags (aspect ratio, style strength).
A worked before and after
Start with the prompt most people actually write:
a woman in a coffee shop, nice lighting, high quality, 4k
Every word here is either vague or filler. "Nice lighting" is not a visual fact. "High quality, 4k" are cargo phrases that do little on modern models. The result is a generic stock-photo look with no point of view.
Now fill the slots deliberately:
Subject: a woman in her sixties reading a paperback
Descriptors: gray cardigan, reading glasses low on her nose, half-smiling
Style: candid documentary photograph, muted 1970s color film
Composition: medium shot, seated by a window, rule-of-thirds, subject on the left
Lighting: soft window light from the right, gentle shadows, warm
Color: warm muted palette, faded reds and browns
Quality: 35mm, shallow depth of field, subtle film grain
Flattened into a single line for the tool:
A woman in her sixties reading a paperback, gray cardigan, reading glasses
low on her nose, half-smiling, candid documentary photograph in muted 1970s
color film, medium shot seated by a window, rule-of-thirds with the subject
on the left, soft warm window light from the right, faded reds and browns,
35mm, shallow depth of field, subtle film grain.
The second prompt is longer, but not because it is padded. Every clause is a decision. If the coat comes out the wrong color, you change one descriptor. If the framing is too tight, you change the composition slot. Nothing else moves.
Why order and specificity matter
Two levers do most of the work.
Order. Several diffusion pipelines weight earlier tokens more heavily and cut prompts off at a token limit, so a subject buried at the end of a long prompt loses influence, and anything past the cutoff can be dropped silently. Lead with the subject and the handful of attributes you most want honored. Details that are nice-to-have go later.
Specificity. Vague adjectives ("beautiful," "professional," "epic") do not correspond to consistent visual features, so the model fills the gap with its own average. Concrete nouns and named techniques ("chiaroscuro," "tilt-shift," "backlit," "risograph") correspond to features the model actually learned. The more you replace mood words with visual facts, the more the output becomes something you chose rather than something you got.
Comments 0
No comments yet. Be the first to share your thoughts.