A confession first. I spent 2023 writing image prompts the way everyone did - a comma soup of tags: "cat, 4k, cinematic, masterpiece, trending on artstation, highly detailed." That style is dead in 2026. nano-banana-2 doesn't reward tag-stacking. It rewards direction.
1. Stop Stacking Tags. Direct.
nano-banana-2 (Google's Gemini 3.1 Flash Image) reads prompts like a creative director reads a brief. It parses intent, physics, and composition - it is a thinking model, not a keyword matcher. The old "dog, park, 4k, realistic" signals nothing to it. Write the scene the way you'd describe a shot to a photographer: subject, action, environment, light, lens. Specificity beats vocabulary every time.
2. Lock Lens, Light, and Material for Photorealism
This is where nano-banana-2 earns its #1 spot on Artificial Analysis Image Arena (1280 Elo as of July 2026, ahead of GPT Image 2 and the older Nano Banana Pro). It renders real light and material - but only if you hand it verifiable shooting parameters. State the lens ("85mm portrait," "35mm film," "macro"), the light direction ("soft light from front-left at 45°," "noon overhead sun"), and the material feedback ("linen creases visible," "slight refraction at the glass rim").
Vague input gives vague output. "Nice clothes" fails. "Distressed denim jacket, faint seam wear on the right shoulder, warm-toned brass buttons reflecting light" lands.
3. Multi-Object: Lock Features on First Mention
nano-banana-2 natively binds up to 5 characters and 14 objects across a workflow. The trick is to nail the unchangeable traits the first time a subject appears. Wrap unique identifiers in parentheses: "(man with round glasses, mole on left brow, dark green cargo pants)", "(silver thermos with three parallel scratches on the surface)". On later references, just say "the man" or "the thermos" - do not re-describe, or the model regenerates them from scratch. This is gold for family portraits, storyboards, and product scenes that need cross-image consistency.
4. Text Rendering: English Double Quotes Are the Switch
In-image text used to be the embarrassment of AI image gen. nano-banana-2 fixed most of it - official accuracy around 96%, climbing past 98% on English layouts - but only if you trigger its text-control mode.
Wrong: "write 'new arrivals' on it." The model ignores or warps it.
Right: "the words 'NEW ARRIVALS' centered at the bottom, Source Han Sans Bold, font size 12% of frame height, white outline, semi-transparent dark grey mask behind."
Rule: every piece of text you want shown goes inside English double quotes, with font, size, position, and layout all specified. Mix Chinese, English, digits, and symbols freely - but the quotes are non-negotiable.
5. Go JSON When Text Gets Dense
For posters, infographics, and UI mockups, the flat prompt runs out of steam. Switch to structured (JSON) prompting with fields like text_overlays, font, kerning, language, and render_mode. Set render_mode to crisp or vector-like - never the default soft. Community tests push OCR recognition from under 60% to 98%+ once JSON is on. It is reproducible, batchable, and A/B-testable, three things flat prompts will never give you.
6. Knolling and Flat-Lay Need Trigger Words
Product flat-lays (knolling), exploded diagrams, and disassembly shots have dedicated trigger words baked into the model. Use knolling, flat lay, and disassemble deliberately. Miss them and you get a generic arrangement; include them and the model snaps into the layout you actually wanted.
7. Edit Prompts Like a Brief, Not a Description
When you're editing an existing image - one of nano-banana-2's strongest modes, it accepts up to 14 reference images - do not re-describe the whole frame. Write an editor's brief with three parts: what to change, what to preserve, what to reference. "Swap the background to a sunset sky, keep the subject's face and pose, match the color grade to reference image B." Skip the "preserve" line and the model redraws everything - your product details, face, and logo all get rolled into the repaint.
8. Subject Segmentation Skip for Product Work
For e-commerce and product shots, turn on subject segmentation skip: the model identifies the product subject and leaves it untouched, repainting only background and lighting. Pair it with a reference image to lock appearance and a prompt that marks no-go zones. Engraved text, logos, and textures survive intact. This is the difference between a usable SKU shot and a refund.
9. Pick Your Thinking Level
nano-banana-2 exposes a thinking level - Minimal (default, fastest) or High/Dynamic (slower, sharper). Simple tasks: Minimal, save the latency. Complex prompts with many objects, dense text, or spatial reasoning: High, let the model reason before it renders. The cost is a few extra seconds; the payoff is fewer rerolls. The choice is yours, not the model's.
Example Transformation
Tag-soup (2023 style): "cat, 4k, cinematic, masterpiece, photorealistic, bokeh, trending"
Director-style (2026, nano-banana-2): "A Siamese cat with piercing blue eyes, sitting on a velvet cushion, shot on an 85mm portrait lens at f/1.8, soft window light from camera-left at 40°, shallow depth of field with creamy bokeh, fine art photography, the words 'STUDIO 9' engraved on a brass plate at the lower-right corner in Source Han Serif, 4K detail"
The first gets you a cat. The second gets you a deliverable.
How x-rush Fits In
You don't have to memorize which model tops which leaderboard this week. x-rush integrates the top models in the industry - Claude Fable 5 and GPT-5.6 for text and reasoning, Qwen3.7 Max for Chinese-heavy prompts, nano-banana-2 for images, Kling 3.0 and Veo 3.1 for video, Suno V5 and ElevenLabs for audio - and routes each prompt to the model that fits the task. When the leaderboard shifts, the routing shifts with it. You write the brief; the platform picks the right hands. That is the whole point of not marrying one model.