Subject

Prompting for AI Images

Getting what you actually pictured out of a text-to-image model. How to structure an image prompt, when negative prompts help, and how Midjourney, Stable Diffusion, Flux, and DALL-E differ in what they want from you.

A dark workbench with a laptop showing an abstract generated-image grid and a notebook of prompt drafts beside it

The habit most people bring to image models is the chatbot habit: type a request, read the reply, ask for a fix. That works with a language model because there is something on the other end reasoning about what you meant. An image model is closer to a translation function. It takes your words and turns them into pixels. There is no colleague to infer intent, so "make it better" has nothing to grab onto.

Once you internalize that, the whole craft shifts. You stop giving instructions and start describing a scene.

Describing versus commanding

Commanding sounds like: "add a dog," "make it more dramatic," "fix the lighting." These assume a listener who tracks state and understands goals. Most image models do not work that way; each generation starts from your full text description, not from a running conversation.

Describing sounds like: "a golden retriever sitting in the foreground, left of frame, looking up at the subject," "high-contrast lighting with deep shadows, single hard light source from the right," "soft overcast daylight, no harsh shadows." Every one of those is a visual fact the model can render. The mental switch is to ask, for anything you want changed, what does it actually look like, and then write that down.

The six slots

A dependable image prompt covers six things. You do not need all six every time, but naming them is what separates a controllable prompt from a lucky one.

  • Subject: the main thing. "An elderly fisherman," "a ceramic teapot," "a mountain village."
  • Context: where it is and what surrounds it. "On a wooden dock at dawn," "on a marble kitchen counter beside fresh lemons."
  • Style: the visual language. "In the style of a 1970s film photograph," "flat vector illustration," "oil painting." Naming a medium or a movement is more reliable than naming a living artist, and many hosted tools restrict or ignore living-artist names anyway.
  • Composition: framing and layout. "Wide shot, subject small in frame," "extreme close-up," "rule-of-thirds, subject on the left."
  • Lighting: often the biggest lever on mood. "Golden-hour backlight," "flat studio lighting," "single candle, rest of the room dark."
  • Medium and quality: "35mm photograph," "watercolor on rough paper," "sharp focus," "shallow depth of field."

A prompt that touches most of these:

An elderly fisherman mending a net, on a weathered wooden dock at dawn,
in the style of a muted 1970s color film photograph, wide shot with the
subject small on the right third, soft golden backlight and long shadows,
35mm, shallow depth of field, gentle film grain.

Compare that to "a fisherman, cinematic, high quality." Both run. Only one gives you a scene you can predict and then adjust.

Iterate one slot at a time

Image generation is stochastic, so treat the first batch as a survey, not an answer. Generate several, pick the closest, and if your tool lets you reuse the seed, do that so the composition holds steady while you change one thing. Swap "dawn" for "overcast noon" and you learn exactly what lighting does. Change five words at once and you learn nothing, because you cannot attribute the difference.

The workflow that holds up: describe a full scene, generate a batch, lock onto the best one, then move a single slot per iteration until it lands. Vague commands get you a slot machine. Specific descriptions plus disciplined iteration get you the image you had in your head.

Where to go next

A short editorial reading list. Pick whichever fits how you like to learn.

  • NerdSip: 5-minute AI micro-course on almost any topic, on iOS and Android

Comments 0

Add a comment

Comments are reviewed by an editor before they appear.

No comments yet. Be the first to share your thoughts.