← Writing
Alumbriva

What If AI-Generated Images Could Still Be Edited?

I’ve been thinking about this lately: in the future, could ChatGPT Work’s image generation maybe give users an editable version after generating an image, instead of only giving us one fixed image with everything flattened together?

I felt this need very strongly again today while using ChatGPT to make a YouTube banner.

The first image it generated was actually already pretty close to what I wanted. After that, the changes I wanted were small: move the astronaut and the little sun slightly, lower the title a bit, adjust the text size and spacing, and tweak the position of the glowing arc around the text so the whole composition would fit YouTube’s crop area better.

If the elements in the image could be selected individually, I probably could’ve finished those edits myself just by dragging things around a little. But because the final result was a non-editable image, I couldn’t directly adjust the text, characters, or decorative elements. The only thing I could do was keep telling the model, in words, what needed to change.

So I’d say, “move the astronaut a little further inward,” and it would generate another version. Then I’d say, “move the title slightly lower,” and it would generate another one again… But each time it regenerated the image, it wasn’t only the part I wanted to change that might change. The font, the arc, the proportions of the characters, and other visual details I was already happy with could all shift too.

In the end, we went back and forth through 7 versions. Version 4 was the closest overall to what I wanted, but because I felt like the details could still be handled a bit better, we tried three more times. The result… maybe my expectations were too high 😂 Version 4 was actually pretty good already.

This reminded me of a very similar problem I ran into a long time ago when I tried Replit.

At the time, Replit could generate a webpage or product interface from a text prompt. But when I wanted to adjust the size, color, layout, or exact position of some text, I couldn’t directly edit those things in the interface. I could only keep describing the effect I wanted to the agent.

The problem is that visual feeling is sometimes really hard to fully explain through language. There’s always some distance between the words you give the AI and the image you have in your head. Even when I tried to be as specific as possible, what it made still didn’t fully match what I was imagining. So I kept revising, kept explaining, and in the end spent a lot of time and credits on it, while getting more and more frustrated. Eventually I just downloaded the whole project, moved it into VS Code, and handed it off to another coding agent to continue from there.

Then in March this year, I saw Replit Agent 4 start making some generated interfaces directly editable. Even though, when I tried it again, the experience still didn’t feel fully mature, I still thought it was a really good direction.

Because it was no longer asking users to control every part of the creative process only indirectly through an agent. Instead, it let AI generate a rough structure and visual direction first, and then allowed the person to directly step into some of the more specific parts.

I think this is also a very important part of human in the loop.

Especially for people who are not professional designers or illustrators, language is probably still the most direct tool for expressing intention. When we look at an image, we go by instinct and judgment and say things like: “this feels a bit crowded,” “the title could go a little lower,” or “the astronaut might look more balanced if it moved inward a little.”

We may not instinctively open a drawing tool and sketch out the ideal position or effect, and we may not know the right professional terms either. For a lot of ordinary users, describing our feeling and rough direction in language is probably the easiest and most common way to work.

But once the model has already generated something that is broadly satisfying, continuing to do every tiny adjustment through language no longer feels like the best method.

Language is great for describing intention, feeling, and direction. But it’s not necessarily the best interface for things like “move this a little to the left,” “make the font size a bit smaller,” or “only change this element and keep everything else the same.”

So, echoing what I mentioned at the beginning: could ChatGPT Work maybe offer some kind of editable format after generating an image?

For example, the text could remain as editable text layers; the characters, arc, background, and decorative elements could be selected, moved, and resized separately; users could lock the parts they’re already happy with; and when regeneration is needed, maybe only the selected element would be regenerated instead of the whole image.

Maybe it could also provide safe-area guides for platforms like YouTube, X, and Instagram, or let users export the result into Figma, Canva, SVG, PSD, or other formats that can still be edited.

I’m not sure how difficult any of this would be technically, and I also don’t know whether the ChatGPT Work team is already thinking about something along these lines. But from my own experience, I would really need this kind of capability.

AI can help me generate a visual direction, atmosphere, and first version that would be hard for me to create on my own. But when the result starts getting close to what I want, I also want to be able to directly take over those very specific adjustments. Then, when I need new ideas, elements, or options again, I can hand that part back to the AI.

To me, that feels much closer to what I think human–AI collaboration could be, compared with having every single change continue to go through language prompts. And I also feel that uniquely human taste and judgment don’t just show up in reviewing and approving what AI makes — they also show up in that tiny exploratory instinct of “if I nudge this just a little, wow, this suddenly feels much more right.”

In other words, what AI ultimately gives people maybe doesn’t have to be only a finished image. It could also be a creative state that the user can still understand, take over, and keep modifying.

So… could something like this maybe become part of ChatGPT Work in the future? 👀

Type to search writing and voice notes.