How I Used GPT Image 2.5: Reference Images, Local Edits, and Multilingual Layouts

Input-and-output examples of GPT Image 2.5: white-background compositions, reference-based scenes, local character edits, text-rich graphics, and English-to-Japanese image editing.

How I Used GPT Image 2.5: Reference Images, Local Edits, and Multilingual Layouts

I used GPT Image 2.5 for several specific image tasks: creating white-background compositions, generating scenes from reference photos, editing parts of an existing image, and generating or translating complete graphics with text.

The examples below are organized by function. Each shows the input and result, with a focus on what changed and which visual elements carried through. Click any image to enlarge it and inspect the details.

01 · White-background composition: adjust the background and framing

The first task was to turn a reference photo into a more focused white-background image. The original includes the surroundings from the photo shoot. The result replaces that background with white and adjusts the subject’s position and scale within a square frame.

This is a composition task based on a reference image. Compare the subject’s outline, pleated structure, and details at the top, alongside the changes to the background and empty space.

Input for white-background composition: a photo of a teal filter
Input: a product photo with its original surroundings.
White-background result: a filter centered in a square image
Result: a white background and a new composition within a square frame.

02 · Scene generation from a reference: put an object into use

The second task started with a flat-lay photo of a red packaging bag. I used it to generate a complete indoor scene with a person and a packing action.

In the result, the bag sits on a table while a woman places packing material inside it. Lighting, tabletop props, and a room setting give the object a specific context. The useful details to examine are how the reference object fits into the action and how naturally the hands meet the bag.

Scene-generation input: a flat-lay photo of a red packaging bag
Input: the original red packaging bag photo.
Scene-generation result: a woman packing a red bag indoors
Result: a person, a packing action, and an indoor scene built around the reference.

03 · Local editing: change the person in an existing scene

Next, I used the scene from the previous section as the input for a local edit, changing the person in the image from a woman to a man.

This pair shows how successive edits can build on a useful image. Once the bag, table, lighting, and room were in place, I could focus the next change on the person. The result changes the person’s face and hands while carrying through the red bag, tabletop props, and indoor composition.

Before local editing: the packing scene with a woman generated in the previous step
Editing input: the scene with a woman generated in the previous step.
After local editing: a scene with a man, retaining the bag and indoor composition
Local-edit result: a different person, with the bag and scene composition carried through.

The sequence is “original product photo → scene with a woman → locally edited scene with a man.” Enlarging both images lets you inspect the changed area and the surrounding content together, then judge how well the edit preserved the parts you wanted to keep.

04 · Generating text and images together: compose a complete infographic

Another task was to generate a complete infographic from a product reference. In the sunglasses example, the input is a white-background product image. The result brings a main visual, headings, color options, icons, a dimensions panel, and a scene with people into one image.

The interesting part here is how the model organizes the information. A large heading introduces the subject, the main product image draws attention, and grouped panels hold the details. Color and background connect the elements, with the text generated as part of the image.

Text-and-image generation input: multicolor sunglasses on a white background
Input: a product reference image.
Generated infographic with English headings, products, icons, and a lifestyle scene
Result: products, English copy, icons, dimensions, and a scene composed into a complete infographic.

Zoom in to examine the text and icons, then return to the page view to assess the hierarchy and reading order. Numbers, dimensions, and feature claims still need to be checked against the actual product information.

05 · Multilingual image editing: turn a complete English graphic into Japanese

This example edits the text and layout of an existing image. The input is already a complete graphic with English headings, feature copy, a person, and detail panels. I used that same English graphic as the input for two separate Japanese results.

Small white-background reference image of a crossbody bag
Small product reference. The language edits below use the complete English graphic as their input.
Shared input for multilingual editing: the complete English feature graphic
Shared input: a complete graphic with English copy, a person, and feature panels.
First multilingual result: a four-panel version with the main headings and features in Japanese
Result one: the main headings and feature copy in Japanese, retaining a panel-based structure.

The first Japanese result continues the combination of a person as the main visual and separate feature panels. The main headings and descriptions change to Japanese. Some decorative English text remains, so the result is best described as a language change for the main copy.

06 · Layout restructuring: a different direction from the same input

The second Japanese result was also edited from the complete English graphic above. It shifts the visual emphasis toward travel and the person wearing the bag, reorganizing the Japanese copy, icons, and product details around that direction.

Layout-restructuring input: the same complete English graphic
Input: the same English graphic.
Second multilingual result: a Japanese travel layout with reorganized copy and scenes
Result two: a Japanese travel version from the English input, changing both language and layout.

These results show two editing directions from one input. One keeps the information in distinct panels; the other gives more prominence to a travel scene. Their respective paths are “English graphic → Japanese panel layout” and “English graphic → Japanese travel layout.”

When comparing versions across languages, I pay particular attention to headings, numbers, units, and the strength of feature claims. For example, the travel version uses 「撥水」, meaning water-repellent, where the English input says “Waterproof.” Once text is part of the generated image, its meaning still needs to be checked.

How I would continue using these functions

These examples suggest a clear working sequence: start with a reference image to establish the subject, then generate the background, action, and layout it needs. When an existing image needs a change to one part, use that result as the next editing input. To explore another language or visual direction, create separate versions from the same complete image.

For general guidance on image generation, editing, and successive revisions, see the OpenAI image generation guide.

Share