Masked Image Regeneration for Text-Guided Photorealistic Editing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image generation systems face challenges in accurately capturing realistic images, maintaining image features during editing, and efficiently generating images in various styles while requiring significant computational resources.
Innovation Solution
A machine learning model is used to regenerate or edit images based on text inputs, involving a masked region removal, text-to-image generation, and pixel value replication to extend image dimensions, utilizing a deep learning model with a large language model and sub-models for enhanced image creation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If conventional image generation systems are used to generate realistic images, then image quality can be improved, but computational power and memory requirements increase significantly
Solution Approach 1:
The image generation task is divided into two separate models: a text-to-image generation model that creates images from text descriptions, and an image editing model that performs localized edits on existing images. This segmentation allows each model to be optimized for its specific function, reducing the computational burden compared to using a single comprehensive system for all image generation and editing tasks.
Solution Approach 2:
The patent extracts the image generation capability from the image editing function. By using a pre-trained text-to-image model separately from the editing model, the system can leverage the generation model's expertise in creating realistic images while the editing model focuses on localized modifications, improving overall efficiency and reducing computational requirements for editing operations.
2Length of stationary object
If conventional image editing systems are used to expand images beyond original borders, then image dimensions can be increased, but maintaining image semantics and style becomes computationally difficult and expensive
Solution Approach 1:
The system performs preliminary actions by using a pre-trained text-to-image model to generate the initial image content before the editing phase. This preliminary generation creates a high-quality base image that maintains proper semantics and style, reducing the computational complexity required in subsequent editing and expansion operations.
Solution Approach 2:
The patent introduces an intermediary approach where the text-to-image model serves as a mediator between the text description and the final edited image. This intermediary model ensures that generated and expanded regions maintain consistency with the original image's semantics and style without requiring complex computational analysis of the entire image.
3Adaptability or versatility
If conventional systems attempt to generate images in various styles while maintaining photorealism, then style versatility can be improved, but computational efficiency decreases
Solution Approach 1:
The system applies local quality by allowing different styles to be applied to different regions of the image through the editing model. The text-to-image model generates photorealistic base images, while the editing model can locally modify specific regions with different styles based on text descriptions, achieving style versatility without sacrificing overall computational efficiency.
Data Source
AI summary
Disclosed herein are methods, systems, and computer-readable media for regenerating a region of an image with a machine learning model based on a text input. Disclosed embodiments involve accessing a digital input image. Disclosed embodiments involve generating a masked image by removing a masked region from the input image. Disclosed embodiments involve accessing a text input corresponding to an image enhancement prompt. Disclosed embodiments include providing at least one of the input image, the masked region, or the text input to a machine learning model configured to generate an enhanced image. Disclosed embodiments involve generating, with the machine learning model, the enhanced image based on at least one of the input image, the masked region, or the text input.


