A working method for writing an AI image prompt: what to put first, which words carry weight, what negatives are for, and how to iterate without guessing.
A good image prompt names the subject first, then how it should look, then how it should be framed. Most disappointing results come from prompts that are either too short to steer the model or so long that everything in them gets diluted. The fix is always structure, not length of prompt.
If you take one thing from this guide: change one element at a time, and generate more than one image before you judge whether a change worked.
What should go in a prompt, and in what order?
Put the subject first, the treatment second, the framing third. Models weight the beginning of a prompt more heavily, so the first few words should be the thing you would be most upset to lose. Everything after that shapes it.
A working skeleton:
Subject. What the picture is of. "A wooden fishing boat."
Detail on the subject. What makes it specific. "Peeling green paint, nets piled at the stern."
Setting. Where it is. "Moored at a stone harbor wall at low tide."
Light. Often the element that changes the picture most. "Overcast morning, flat gray light."
Treatment. Photograph, oil painting, pencil sketch, 3D render.
Framing. Wide shot, close crop, shot from above.
Final Prompt: "A wooden fishing boat with peeling green paint and nets piled at the stern, moored at a stone harbor wall at low tide, overcast morning with flat gray light, photograph, wide shot."
That prompt does not use special syntax and it is not long. It is specific in the places that decide what the image looks like.
Which words carry the most weight?
Concrete nouns and lighting terms move an image more than adjectives do. "Beautiful," "amazing," and "high quality" are close to noise, because they do not describe anything the model can point at. "Backlit," "overcast," "harsh midday sun," and "single lamp in a dark room" change the picture immediately.
The same is true of composition. "Close-up," "wide shot," "shot from below," and "centered" give the model a frame to work in. Without one it will pick, and it tends to pick a safe middle distance.
That the wording matters this much is a measurable property rather than folklore. Investigating Prompt Engineering in Diffusion Models by Witteveen and Andrews sets out techniques for measuring the effect specific words and phrases have on the output of text-to-image models, and concludes with guidance on choosing prompts to produce a desired effect. The useful takeaway is the method: change one term, hold everything else still, and look at what moved.
A useful test on any word in your prompt: could two people draw noticeably different pictures from it? If not, the word is doing work. "Stunning" fails this test. "Rain-slicked" passes.
Style references deserve care. Naming a well-known living artist to copy their style is a legal and ethical question, not only a technique question, and several services restrict it. Describing the qualities you want, "thick visible brushstrokes, muted palette, hard shadows," gets you most of the way and does not put you in that position.
What are negative prompts for?
A negative prompt lists what you do not want, and it is for removing recurring defects rather than for excluding subjects. It is most useful once you have seen the failure, not as a preemptive wall of terms.
The way people misuse it is to paste a long generic list before generating anything. That list constrains the model in ways you have not observed and cannot evaluate. Generate first, see what keeps going wrong, then name that specific thing.
Negatives also do not reliably remove an object. Asking for a street with "no cars" in the negative often produces cars anyway, because the model is being steered rather than filtered. Describing the scene so cars have no place in it, "a pedestrianized alley too narrow for vehicles," works better than forbidding them.
Not every service exposes negative prompts, and for many people they are not the missing piece.
How do you iterate without guessing?
Change one thing, keep everything else fixed, and generate a batch each time. This is the whole method, and it is the difference between learning your model and rerolling until something happens.
A workable loop:
Write the shortest prompt that names your subject and treatment.
Generate several at once. You are sampling a range, not testing a hypothesis.
Pick the closest result and decide the single biggest gap between it and what you want.
Change only the part of the prompt addressing that gap.
Generate several again and compare against the previous batch.
If your tool exposes a seed, fix it during step four so the only variable is your edit. Without a seed, compare batches rather than single images, because the variation between two random draws will otherwise swamp the effect of your change.
Keep the prompts that worked. People rebuild the same prompt from scratch weeks later because they never saved it. A text file is enough.
When is the prompt not the problem?
Some failures are structural and no rewording fixes them. Recognizing these saves hours.
Hands, teeth, and text. These fail for reasons built into how the models work rather than because your description was insufficient. Compose around them, crop, generate more and pick, or add text afterward in an editor.
A specific person, product, or logo. A description cannot summon your specific product. That needs photography, or a model trained on it.
Small edits to an existing image. Regenerating with "the same but blue" gives you a different image. This is what inpainting and image-to-image are for, and they are a different tool rather than a better prompt.
Consistency across a series. Getting the same character in six pictures is a project, not a prompt. It typically needs fixed seeds, reference images, or fine-tuning.
Counting. Models are unreliable at "exactly five apples." If a count matters, plan to compose it yourself.
If you find yourself on the tenth rewording of a prompt with no progress, check this list before writing an eleventh.
Does prompt length help?
Up to a point. Past it, extra words dilute the ones that matter, because attention is being spread across everything you wrote. A prompt that names subject, light, treatment, and framing precisely will usually beat a much longer one padded with quality adjectives.
There is also a hard limit. Prompts are truncated after a certain length, and how much fits varies by model. Words past that cutoff have no effect at all, which is one reason a long prompt can behave as though its ending was ignored. It was.
Write the short version first. Add only what is missing.
How should you evaluate a prompting workflow?
If you are picking a tool to work in rather than experimenting casually, judge the workflow:
Can you generate a batch from one prompt? Single-image generation makes iteration expensive and slow.
Can you see the prompt and settings that produced a past image? Without this you cannot reproduce your own good results.
Is there a seed or any reproducibility?
Can you retrieve your history, and for how long? Retention windows are common. Assume the good one you made last month may not be there.
What is stored, and who can see it? Prompts are generally logged and reviewed. This is normal and worth knowing rather than assuming otherwise.
Write the prompt you would give a photographer or an illustrator, then cut the words that would not change what they did. That instinct produces better prompts than any list of magic terms, and it survives model changes, which lists of magic terms do not.
What Carpathian Provides
Carpathian AI runs our own image models on hardware we own in the United States, with batch generation from a single prompt so the iterate-and-compare loop above works, and a session view that keeps each generation next to the prompt and settings that produced it. It needs a free account, and each account carries an image allowance before anything is billed.
How to Write an AI Image Prompt That Works | Carpathian