Masked Image Regeneration for Text-Guided Border Expansion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image generation systems face challenges in accurately capturing realistic images, maintaining image features during editing, and efficiently generating images in various styles while requiring significant computational resources.

Innovation Solution

A machine learning model is employed to regenerate or edit images based on text inputs, utilizing a masked region and text-to-image model to extend image dimensions and maintain semantic features, with sub-models for image embedding and generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional image generation systems are used to generate realistic images, then image quality may be improved, but computational power and memory requirements increase significantly

Engineering Contradiction:
Improveimage qualityVSAvoidcomputational power
Core Design Contradiction:
Manufacturing precisionVSPower

Solution Approach 1:

The system segments the image generation task into multiple components: a text-to-image model for generating base images, an image expansion model for extending borders, and a fusion model for combining results. This segmentation allows each component to be optimized independently, reducing overall computational requirements while maintaining image quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by pre-processing input images and text prompts before main generation. The text-to-image model generates initial images that are then refined by the expansion model, avoiding the need for single high-computation passes and enabling more efficient resource utilization.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If image editing techniques are used to expand images beyond original borders, then image versatility is improved, but maintaining semantic features becomes computationally difficult

Engineering Contradiction:
Improveimage expansion capabilityVSAvoidsemantic features
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system implements feedback mechanisms where the expansion model receives both the original image and generated content, processes them together, and produces refined output that maintains semantic consistency. The model learns from training data to preserve semantic features while expanding image boundaries, reducing information loss.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If conventional systems are used for image generation and editing, then basic functionality is achieved, but the ability to change aspect ratio while maintaining image features is limited

Engineering Contradiction:
Improveaspect ratio flexibilityVSAvoidimage feature maintenance
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The system dynamically adjusts image dimensions and aspect ratios during the generation and expansion processes. The model can handle variable input sizes and produce outputs in multiple aspect ratios while maintaining semantic features through learned transformations and adaptive processing.

Inventive Principle:
Principle #15Dynamics

4Productivity

If conventional image generation methods are used, then generation speed may be maintained, but the ability to generate images in various styles while maintaining photorealism is limited

Engineering Contradiction:
Improvegeneration speedVSAvoidstyle variety
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system employs universal models that can generate images in multiple styles while maintaining photorealism. The text-to-image model and expansion model are designed to handle diverse stylistic requirements through learned representations, enabling single models to perform multiple generation tasks efficiently.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12518352B2Systems and methods for image generation with machine learning models
Publication Date: 2026.01.06 OPENAI OPCO LLC
  • US12518352B2 patent drawing
  • US12518352B2 patent drawing
  • US12518352B2 patent drawing

AI summary

Disclosed herein are methods, systems, and computer-readable media for regenerating a region of an image with a machine learning model based on a text input. Disclosed embodiments involve accessing a digital input image. Disclosed embodiments involve generating a masked image by removing a masked region from the input image. Disclosed embodiments involve accessing a text input corresponding to an image enhancement prompt. Disclosed embodiments include providing at least one of the input image, the masked region, or the text input to a machine learning model configured to generate an enhanced image. Disclosed embodiments involve generating, with the machine learning model, the enhanced image based on at least one of the input image, the masked region, or the text input.