How to prompt NanoBanana

A practical NanoBanana prompt guide: the prompt formula, referencing attached photos with @image1, when NanoBanana 2 is enough and when Pro is worth it, and how each handles text in an image.

NanoBanana is Google’s Gemini image line, and the bot offers it as NanoBanana 2 for everyday work and NanoBanana Pro for the renders that have to be right the first time. Both take the same kind of prompt, so learning one teaches you both and the only real decision is which tier you spend credits on. What sets the family apart is that it reads a prompt for material, surface and mood instead of working through it like a checklist.

The NanoBanana prompt formula

Six slots cover almost every prompt: subject, action, setting, light and lens, framing, style. Here is one prompt carrying all six:

A woman in her thirties in a grey wool coat sitting at a cafe window on a rainy afternoon, one hand wrapped around a ceramic mug with steam rising off it, the other forearm resting on the wooden counter, soft daylight from the left falling through rain streaked glass, 85mm close portrait with the background dropping out of focus, natural documentary photography

Note that it reads as a description, not an inventory. Tags stacked with commas (woman, cafe, rain, moody, 85mm) get averaged into a mood; connected clauses keep the relationships, light falling through the glass, steam rising off the mug. Say it out loud, then check all six slots are there.

The failure mode that follows does not look like failure. GPT-Image-2 misses literally, so the gap is visible and you know which slot to add. NanoBanana lands in the middle: a thin prompt still returns a competent, plausible image, the average of everything your words could have meant, with nothing pointing at what you never wrote.

Texture and material is the fastest way out of that middle. Ribbed knit, glazed ceramic, weathered brass, condensation on glass: each names a surface the model can resolve, and each narrows the average. Adjectives about quality (“beautiful”, “masterpiece”, “8k”) narrow nothing, since every render was already trying to be those.

The two pictures further down the page show the difference. The first came from three words, “a woman in a cafe”, and the model answered with a pleasant wide shot: a smiling stranger, a pastry, a whole cafe of people behind her. The second names her age, the coat, the mug, the rain on the glass, the light and the focal length, and the crowd, the pastry and the smile all disappear, because none were asked for.

Attaching photos: @image1 and @image2

Up to 4 photos go in one message with the instruction as its caption, and you then address them in the prompt as @image1, @image2 and so on; the step-by-step guide covers the mechanics with screenshots.

Keep the face and hair from @image1 exactly as they are, dress her in the grey wool coat from @image2, place her at a rainy cafe window, soft light from the left, keep the background of @image2 unchanged

What matters here is that NanoBanana does not paste a reference face, it re-renders it, so a likeness survives only as far as your prompt protects it. “Keep the face from @image1 unchanged, same bone structure, same nose, do not restyle the skin” holds up across a change of clothes or light where “keep the face” does not. Ask for new light and it relights the face too, and a relit face is a slightly different person. A dim or heavily filtered source also gives it less to hold onto.

Blending is where the family is genuinely strong. A face from one attachment, wardrobe from another, a location from a third comes back as one scene rather than a collage, without the cut-out edges such edits usually leave. Give every attachment a job: a reference with no stated role is absorbed into the mood instead of used. The more you attach, the more the tier matters.

NanoBanana 2 or NanoBanana Pro

The whole nano banana 2 vs pro question comes down to this: both are good, and only one of them is cheap.

NanoBanana 2NanoBanana Pro
Resolutions1K, 2K, 4K2K, 4K
Credits2 at 1K, 3 at 2K, 4 at 4K4 at 2K, 5 at 4K
Cheapest iteration2 credits at 1K4 credits at 2K, no 1K tier
Built forEveryday images, best quality per creditMaximum quality, fewest artifacts
Where it shinesSingle clear subject, quick iterationMany elements held together in one frame

Pro’s cheapest render costs the same as a 4K image on 2, and that pricing is the whole workflow. Iterate on NanoBanana 2, render the keeper on Pro. Run the prompt at 1K for 2 credits, change one slot, run it again, and spend Pro credits only once the wording works.

The difference you are paying for shows up on complexity. One subject and clean light looks near identical on both tiers, so reaching for Pro there wastes credits. Load the prompt up with several people, a busy background, specific wardrobe and a particular light, and NanoBanana 2 starts quietly dropping details: a described garment turns generic, a face at the back of the frame smears. Pro holds the list together and keeps its edges clean at 4K, which is why it is also the one for hero images and close-range skin. The same gap shows on attachments: with three or four references in play, 2 serves one well and lets the rest blur, while Pro keeps each doing its own job.

Text inside images is the weak spot

Every other section here is about getting more out of the model. This one is about knowing its ceiling, which is the most useful gemini image model prompt tip on the page. NanoBanana 2 manages a short label often enough to be tempting and misspells it often enough to be frustrating, and it is the least reliable model here once a string runs past a word or two. NanoBanana Pro is markedly better, and on a short headline usually fine, but “usually” is the operative word.

To give nano banana text in image its best chance: put the exact wording in quotation marks, keep it short, name the surface it is printed on, and re-run rather than rewrite when a letter comes out wrong, since lettering varies between runs of the same prompt.

When the wording has to be exactly right, switch models. GPT-Image-2 is the strongest for text and the cheapest to iterate on, and the GPT-Image-2 guide covers how to get a string to land the first time.

What the filter refuses

Expect roughly GPT-Image-2 levels of strictness, over the same four categories: real named people, brand marks and logos, violence, and explicit content. The way out is identical, so rather than repeat it here, the GPT-Image-2 guide works the rewrite pattern through with examples.

What is particular to this family is the refusal that arrives while editing rather than generating. Once you attach a photo of a real, identifiable person, the limit stops being your wording and becomes the face in the frame: no phrasing unlocks an edit of a public figure, and none unlocks a picture placing someone in a situation they were never in. Re-running it or softening the sentence changes nothing, which is worth knowing before you spend credits finding out.

The workflow the reference tools exist for stays wide open: your own face, friends who agreed to it, restyling, relighting, a change of wardrobe or setting. Expect the strictest reading on children, on anything intimate, and on edits that would read as a claim about something that really happened. Seedream is looser on style, not on subject, so the Seedream guide is worth reading for stylized and ultrawide work, not as a detour around a refusal you earned.

Common fixes

SymptomChange to the prompt
Details went missing on a busy promptCut the scene down to what matters, or render it on NanoBanana Pro, which holds many elements together
Lettering came out muddy or misspelledShorten the string, put it in quotes, name the surface, re-run rather than rewrite, and move to GPT-Image-2 if it must be exact
The attached photo was ignoredNumber your references as @image1 and @image2 and state which one the face, the wardrobe and the background each come from
Faces look plastic and over-smoothAsk for skin texture, pores and fine flyaway hair, add a real light direction, and name a photographic style instead of “beautiful”
The result is pleasant but genericAdd materials and surfaces: knit, wool, glazed ceramic, wet glass, worn leather
The mood is flatSet the light before anything else: direction, softness, time of day, and whether it is coming through something

Ready-made prompts

Reading about materials teaches you less than watching a render change when glazed ceramic becomes hammered copper. The NanoBanana prompt library has ready-to-paste prompts next to the images they produced, and the model pages for NanoBanana 2 and NanoBanana Pro list the resolutions, aspect ratios and credit costs.

Open the bot, pick a NanoBanana tier, and paste one in.

Same idea, two prompts

Prompt: 'a woman in a cafe'. Generated with NanoBanana Pro.
The same subject with light, lens, wardrobe and focus specified. Generated with NanoBanana Pro.

Frequently asked questions

NanoBanana 2 or NanoBanana Pro?

NanoBanana 2 covers 1K to 4K at 2 to 4 credits and is the everyday choice. NanoBanana Pro runs at 2K and 4K for 4 to 5 credits and earns the difference on complex, multi-part prompts and on anything where artifacts would be obvious. Start on 2, move up when a render has to be flawless.

Is NanoBanana good at text inside images?

NanoBanana 2 is the weakest of the models here at lettering, and NanoBanana Pro is markedly better. When the words have to be exactly right, GPT-Image-2 is the model to use.

How do I tell the model which attached photo to use?

Send up to 4 photos in one message with your instruction as the caption, then refer to them by number in the prompt: keep the face from @image1, use the background of @image2. Numbering removes the guesswork.

Why was my NanoBanana prompt refused?

Its content filter is strict, comparable to GPT-Image-2. Named real people, brand marks, violence and explicit content are the usual triggers. Describing a type of person rather than naming one usually passes.

zosee is an independent product. Not affiliated with, endorsed by, or sponsored by Telegram, Google, ByteDance, OpenAI, or Midjourney. Model names are trademarks of their respective owners.

Try free in Telegram

Subscribe to our prompt channel @zosee_channel and claim 5 free credits, enough for up to 5 images. Credits never expire.