Upgrade
Back to blog

Text to Image vs Image to Image: Which One Are You Doing?

Learn when to start from a text prompt or an existing picture, what a reference image controls, and how to move between both AI image workflows.

Jul 10, 2026Image Generator TeamImage Generator Team

There are two composers here, and the difference between them is smaller than it looks: both take a prompt, both bill the same credits, both save to the same library. The only difference is whether the model starts from your words alone or from a picture you hand it.

That difference decides which one you should be in.

Text to image: when the picture doesn't exist yet

Text to image is generation from a blank slate. You describe the scene, pick a model and an aspect ratio, and get a picture back in seconds. Use it when there is nothing to start from — a hero image, a concept, an illustration for a post, a product shot of something that hasn't been photographed.

All three models work here, including FLUX Schnell, which is text-only and by far the cheapest way to explore.

The cost of the blank slate is control: everything the model wasn't told, it invents. Which is exactly the problem the other composer solves.

Image to image: when something should stay

Image to image starts from a reference picture. You upload it, write what should change, and the model renders a new version that keeps what you didn't ask it to change. Upload a JPG, PNG or WebP up to 10MB; you can give up to three references on either model that supports them.

Typical jobs:

  • Restyle a photo — keep the composition, change the medium. Watercolour, 3D render, film still, line art.
  • Change colour and light — midday to golden hour, or a product repainted in a different colourway.
  • Keep a character consistent — give Nano Banana Pro three shots of the same subject and put them in a new scene without redrawing the face.
  • Blend references — a subject, a background and a style reference in one render; the prompt says how the three should meet.
  • Rewrite text in place — a new headline on a poster whose layout already worked.
  • Iterate on your own output — feed a generated image straight back in and push it one step further.

Three models take references: Nano Banana Pro at 16 credits, GPT Image 2 at 24, and Seedream 5.0 Pro at 8. FLUX Schnell is text-only and doesn't appear in that composer.

Write the difference, not the description

This is the mistake that costs people the most credits: on image to image, describing the whole picture again. The model can already see it. "Same pose, watercolour style, warmer light" outperforms a paragraph re-describing a photo you just uploaded — the long version competes with the reference instead of guiding it.

Say what should be different. Everything else is the reference's job.

Moving between the two

The two composers work best as one loop:

  1. Draft in text to image on FLUX Schnell — four images a run, a credit each.
  2. Pick the frame that's closest and open it in image to image.
  3. Change one thing at a time — the light, the style, the wording on the sign.
  4. Feed the result back in as the new reference when you want to go further.

Every render lands in your library with its prompt and model, so any image you made last week is a valid starting point this week.

Not sure which door to walk through? If you can point at the picture you want to change, start there. If you can only describe it, start with words.