A plain-English guide that takes you from one-word prompts to pro results. The nine ingredients, a before and after, per-model tips, and a formula you can copy and paste.
Every AI image generator works the same way underneath. You give it words, and it fills in everything you did not say with the average of the millions of images it learned from. Average is exactly why lazy prompts look generic. Type "a dog" and the model has to invent the breed, the light, the angle, the background and the mood, so it reaches for the blandest safe version of all of them.
The fix is not a fancier model. It is telling the model the things it would otherwise have to guess. A detailed prompt on a two-year-old model beats a sloppy prompt on the newest one almost every time. Once you learn the ingredients below, your images stop looking like everyone else's and start looking like the picture in your head.
Think of a prompt like ordering a custom sandwich. If you only say "a sandwich," you get whatever the kitchen feels like. Name the bread, the filling, the sauce and how it is toasted, and you get exactly what you wanted. These nine ingredients are the bread and filling of every strong image prompt.
| Ingredient | What it answers | Weak | Strong |
|---|---|---|---|
| Subject | Who or what is in the frame, and what they are doing | a woman | an elderly fisherwoman mending a net |
| Style | The look, era, or artist feel | (none) | 1970s film photography, muted |
| Medium | Photo, oil painting, 3D render, pencil, pixel art | (none) | 35mm photograph |
| Lighting | How light hits the scene | (none) | soft golden-hour backlight |
| Camera | Lens, angle, distance, shot type | (none) | close-up, 85mm, shallow depth of field |
| Composition | How things are arranged in the frame | (none) | rule of thirds, subject left, negative space |
| Color | Palette and grading | (none) | warm amber and teal, low saturation |
| Detail | Texture and level of finish | (none) | weathered skin, visible net fibers, sharp focus |
| Mood | The feeling the image should carry | (none) | quiet, patient, hard-won calm |
You do not need all nine every time. But the more of them you name, the less the model guesses, and guessing is what makes images look cheap. If you only have room for four, pick subject, lighting, camera and mood. Those four carry the most weight.
Here is the same idea written two ways. The first is what most people type. The second layers in the nine ingredients. Read them out loud and you can already picture how different the results will be.
Weak prompta cat by a window
Strong promptA fluffy orange tabby cat curled asleep on a sunlit wooden windowsill,
35mm film photograph, soft morning light streaming through the glass,
close-up shot on an 85mm lens with shallow depth of field,
composed off-center with warm negative space,
cozy amber and honey color palette, low saturation,
highly detailed fur with visible individual strands, sharp focus on the face,
calm and peaceful early-morning mood
Same cat, same window. The weak prompt gives the model nine decisions to make. The strong prompt makes eight of them for it and leaves only the pose to chance. That is the difference between a snapshot and a photograph you would frame.
When you are not sure where to start, fill in this template. Drop any part you do not need, keep the order, and put the most important words first because most models pay more attention to early words.
The formula[subject + what they are doing], [style], [medium],
[lighting], [camera / shot], [composition],
[color palette], [detail / texture], [mood] [+ settings]
Worked example you can adapt right now:
Fill-in examplea lone hiker reaching a mountain summit at dawn, epic cinematic style,
digital painting, dramatic rim lighting from the rising sun,
wide establishing shot from below, subject small against a vast sky,
cold blue shadows and warm gold highlights, crisp detailed rock and snow,
triumphant and awe-struck mood, 16:9
Want the tool to build this for you and score it before you spend a single credit? Use our free helpers:
The nine ingredients transfer everywhere. How you format them does not. Here is how to phrase a prompt for each of the big five, plus a link to a free copy-paste library tuned to each one.
Midjourney loves short, comma-separated phrases rather than long sentences, and it responds strongly to art and mood words. Put your parameters at the very end. Use --ar 16:9 for aspect ratio, --style raw when you want it to follow you more literally, and --stylize (or --s) to dial its artistic flair up or down. A style reference image with --sref keeps a consistent look across a set.
FLUX reads full natural-language sentences beautifully and is the best of the group at rendering real, readable text inside an image. Write it like you are describing a photograph to a friend, in complete sentences, and be specific about materials and lighting. It rarely needs a negative prompt, so do not clutter it with one.
DALL-E is the most conversational. Just describe what you want in plain sentences, then refine by chatting: "make the lighting warmer," "move the subject left," "add a second character." It is strong at following instructions and at short bits of text, but it ignores Midjourney-style parameters, so keep everything in plain words.
Nano Banana shines at editing an image you already have and at keeping a character or face consistent across shots. Upload a reference and describe the change in plain language: "keep this exact person, change the background to a snowy street at night, keep the warm lighting on their face." Precise, instruction-style sentences work better here than a pile of keywords.
Stable Diffusion gives the most control and runs on your own machine for free. Many models respond to comma-separated tags, and you get a separate negative prompt to push away what you do not want. Emphasize a word by weighting it, like (golden hour:1.3), and lock a result with a fixed seed. It has the steepest learning curve but the highest ceiling.
You get good at this the same way you get good at anything: reps with feedback. Start with a bare subject, generate once, then add exactly one ingredient and generate again. Watch what each addition does. In ten minutes you will feel which words move the image most, and you will stop guessing. To skip the trial and error, paste your prompt into Rate My Prompt and it will score you 0 to 100 and tell you which ingredient is missing.
The Vault is our full library of pro, copy-paste prompts for every major model, organized by subject and style so you can grab one and go. Or score your own free first.
Open The Vault → Rate my prompt free