Masked Image Regeneration for Text-Guided Photorealistic Editing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image generation systems face challenges in accurately capturing realistic images, maintaining image features during editing, and efficiently generating images in various styles while requiring significant computational resources.

Innovation Solution

A machine learning model is used to regenerate or edit images based on text inputs, involving a masked region removal, text-to-image generation, and pixel value replication to extend image dimensions, utilizing a deep learning model with a large language model and sub-models for enhanced image creation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional image generation systems are used to generate realistic images, then image quality can be improved, but computational power and memory requirements increase significantly

Engineering Contradiction:
Improveimage qualityVSAvoidcomputational power
Core Design Contradiction:
Manufacturing precisionVSPower

Solution Approach 1:

The image generation task is divided into two separate models: a text-to-image generation model that creates images from text descriptions, and an image editing model that performs localized edits on existing images. This segmentation allows each model to be optimized for its specific function, reducing the computational burden compared to using a single comprehensive system for all image generation and editing tasks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts the image generation capability from the image editing function. By using a pre-trained text-to-image model separately from the editing model, the system can leverage the generation model's expertise in creating realistic images while the editing model focuses on localized modifications, improving overall efficiency and reducing computational requirements for editing operations.

Inventive Principle:
Principle #2Taking out (Extraction)

2Length of stationary object

If conventional image editing systems are used to expand images beyond original borders, then image dimensions can be increased, but maintaining image semantics and style becomes computationally difficult and expensive

Engineering Contradiction:
Improveimage dimensionsVSAvoidcomputational complexity
Core Design Contradiction:
Length of stationary objectVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by using a pre-trained text-to-image model to generate the initial image content before the editing phase. This preliminary generation creates a high-quality base image that maintains proper semantics and style, reducing the computational complexity required in subsequent editing and expansion operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary approach where the text-to-image model serves as a mediator between the text description and the final edited image. This intermediary model ensures that generated and expanded regions maintain consistency with the original image's semantics and style without requiring complex computational analysis of the entire image.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If conventional systems attempt to generate images in various styles while maintaining photorealism, then style versatility can be improved, but computational efficiency decreases

Engineering Contradiction:
Improvestyle versatilityVSAvoidcomputational efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system applies local quality by allowing different styles to be applied to different regions of the image through the editing model. The text-to-image model generates photorealistic base images, while the editing model can locally modify specific regions with different styles based on text descriptions, achieving style versatility without sacrificing overall computational efficiency.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20260087598A1Systems and methods for image generation with machine learning models
Publication Date: 2026.03.26 OPENAI OPCO LLC
  • US20260087598A1 patent drawing
  • US20260087598A1 patent drawing
  • US20260087598A1 patent drawing

AI summary

Disclosed herein are methods, systems, and computer-readable media for regenerating a region of an image with a machine learning model based on a text input. Disclosed embodiments involve accessing a digital input image. Disclosed embodiments involve generating a masked image by removing a masked region from the input image. Disclosed embodiments involve accessing a text input corresponding to an image enhancement prompt. Disclosed embodiments include providing at least one of the input image, the masked region, or the text input to a machine learning model configured to generate an enhanced image. Disclosed embodiments involve generating, with the machine learning model, the enhanced image based on at least one of the input image, the masked region, or the text input.