GPT-Image-2 reads a prompt literally, so you prompt it by naming things instead of hinting at them: the subject, what it is doing, the light, the framing, and the exact words you want printed inside the picture. Pick it in the model menu of the bot, then write one message that states all of it. Everything you leave out, the model fills in on its own, which is the single reason most first attempts come back looking like someone else’s idea.
The GPT-Image-2 prompt formula
Six slots cover almost every prompt: subject, action, setting, light and lens, framing, style. Write them in that order, in plain sentences:
A cream colored A3 poster taped by its four corners to a weathered wooden wall, warm side light from the left, shot straight on, the poster reads “FRESH BREW” in bold black letters with a smaller line under it reading “2 FOR 1 UNTIL FRIDAY”, with an illustrated latte cup below the text
Subject is the one thing the image is about: a noun with two or three adjectives, not a category. Action is what it is doing right now, which is what stops a portrait from looking like a mugshot. Setting places it and carries the props: a wooden wall and tape read very differently from a gallery frame.
Light and lens does most of the visual work: warm side light, overcast daylight, a single hard lamp, 35mm wide, 85mm portrait with shallow depth of field. GPT-Image-2 applies them literally rather than averaging them into a house style. Framing decides close-up, waist up, or wide, plus the camera angle. Style closes the prompt: photorealistic, editorial, flat illustration, 3D render.
Literalism cuts both ways, which is why this formula pays off faster here than elsewhere. The model will not invent an atmosphere to cover a vague prompt, and it will not quietly drop an odd instruction either. Add a slot, and you see that slot in the next render.
The two posters further down the page are the same idea run twice. The first came from one line, “a poster about a coffee sale”, and the model wrote its own copy: a headline, a brand name, a discount and an end date, none of which were asked for. The second names the exact wording, the surface and the light. Naming the words is what puts your words in the picture instead of the model’s.
Attaching photos: @image1 and @image2
For editing, send up to 4 photos in one message and write the instruction as the caption of that same message. Photos in one message and text in the next are two separate messages, and they will not be paired. The step-by-step guide covers that one-message rule in full.
Once several photos are attached, address them by number. GPT-Image-2 follows these references closely, which is what makes it the editing specialist:
Put the face from @image1 onto the person in @image2, keep the hairstyle and the jacket from @image2 unchanged, match the skin tone to @image1, studio light, waist up
Say what to keep as explicitly as what to change. “Keep the background unchanged” and “do not alter the pose” are instructions the model actually follows, and without them an edit redraws more of the frame than you wanted. Leave the numbers out and the model decides for itself which reference is for what.
Where GPT-Image-2 is strong, and where it is not
| Strong at | Struggles with |
|---|---|
| Readable text inside the image | Long paragraphs of text |
| Photo edits from your own attachments | No 21:9 ultrawide option, which only Seedream offers in the bot |
| Following literal, multi-part instructions | Loose, evocative one-line prompts |
| Product shots, signage, posters, packaging | Anything the strict filter flags |
| Cheap iteration at 1 credit per 1K image | Heavily stylized illustration and painterly art |
For stylized art, start with Seedream; for maximum quality on a complex scene, NanoBanana Pro.
Getting text right inside an image
Text is why most people pick this model, so read this section twice.
Put the exact wording in quotation marks. Anything you paraphrase, the model rewrites. “A sign that mentions the opening hours” gets you invented hours; the sign reads "OPEN 8 TO 6" gets you those words.
Name the surface the text is printed on: a poster taped to a wall, a paper coffee cup, a shop awning, a book cover, a phone screen. The surface determines the typography, the perspective and how the letters wrap, and the model handles all three better when it knows what it is writing on.
State the language when it is not English, for example Cyrillic text reading "ОТКРЫТО". Without it, mixed-script prompts drift back to Latin letters.
Keep it short. One headline plus at most one supporting line is the reliable ceiling. Reliability drops off sharply past a short headline’s worth of words, and a full paragraph of body copy will not come out clean on any model available today.
When a line does come back garbled, re-run the same prompt rather than adding more words. Lettering varies between runs, and at 1 credit per 1K image a second attempt is cheaper than a rewrite. If two runs both fail, shorten the string.
This is the clearest gap in the lineup: GPT-Image-2 lands short strings that NanoBanana 2 tends to scramble. If you need text plus a more painterly look, see the NanoBanana guide for how far that model can be pushed.
What the filter refuses
GPT-Image-2 has the strictest content filter of the models in the bot. Four things account for nearly every refusal: real named people, brand marks and logos, violence, and explicit content. The refusal is a flat no, not a degraded image, so the fix is always in the prompt.
Rewrite by describing the type rather than naming the instance:
Instead of “Keanu Reeves in a black suit on a rainy street”, write “a man in his fifties with long dark hair and a beard, black suit, rainy neon lit street at night, cinematic”.
Instead of “a can of Coca-Cola on a table”, write “a red aluminium soda can with an unbranded white script logo, on a wooden table, studio light”.
Both keep the idea and pass. Seedream has a more permissive filter, and the Seedream guide covers what it is genuinely better at, mainly stylized and ultrawide work. The rules still apply there, so a prompt refused on principle rather than on wording is not worth re-routing.
Common fixes
| Symptom | Change to the prompt |
|---|---|
| Lettering comes back garbled | Shorten the string, put it in quotes, name the surface, then re-run rather than rewrite |
| The model wrote its own headline | Give the exact wording in quotation marks instead of describing the message |
| The wrong attached photo got edited | Number your references: @image1, @image2, and say which is the source and which is the target |
| One instruction was ignored | Split the sentence in two, put the instruction in its own clause, and state what must stay unchanged |
| The result looks flat and generic | Add light and lens: direction, quality, focal length, depth of field |
| The edit changed more than you asked | Add “keep everything else unchanged” and name the parts to preserve |
Ready-made prompts
The fastest way to learn a model is to run a prompt that already works, then change one slot at a time. The GPT-Image-2 prompt library has ready-to-paste prompts with the images they produced, and the model page lists resolutions, aspect ratios and credit costs.
Open the bot, pick GPT-Image-2, and paste one in.