OpenAI Adds Image Generation Capabilities to GPT-4o

OpenAI has announced the rollout of image generation capabilities in GPT-4o, its latest and most advanced multimodal model. The new feature is now available to ChatGPT Free, Plus, Pro, and Team users, with access for Enterprise and Education users expected soon. API access for developers is also slated to launch in the coming weeks.

The image generation tool is designed to produce photorealistic and visually detailed outputs based on user prompts. It supports the creation of both creative and functional images, such as conceptual art, diagrams, posters, and more. Users can also upload reference images to guide style, content, or composition.

“At OpenAI, we have long believed image generation should be a primary capability of our language models,” the company said in a statement. “That’s why we’ve built our most advanced image generator yet into GPT-4o. The result—image generation that is not only beautiful, but useful.”

Among its standout features is GPT-4o’s ability to render clear and legible text within images. It also offers strong performance in blending visual and linguistic elements seamlessly.

OpenAI acknowledges current limitations, including image cropping issues, hallucinated details in vague prompts, difficulty rendering multilingual text, and lack of precision in complex edits. To promote safety, the company adds provenance metadata, blocks unsafe prompts, and uses an LLM trained to enforce policy compliance.