Text-Based Real Image Editing Through Disentangled Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image editing systems face inefficiencies in generating modified images due to the need for numerous diffusion steps, low image quality, and challenges in disentangling image elements, often resulting in undesired changes to non-targeted elements and artifacts.

Innovation Solution

A system utilizing an inversion model to generate intermediate features and an image generation model to create synthetic images efficiently, allowing precise modifications while maintaining the integrity of other image elements, using a combination of inversion and diffusion-based techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If conventional diffusion-based image editing is used, then image modification can be achieved, but numerous diffusion steps are required resulting in low processing speed

Engineering Contradiction:
Improveprocessing speedVSAvoidnumber of diffusion steps
Core Design Contradiction:
SpeedVSProductivity

Solution Approach 1:

The inversion model performs preliminary encoding of the input image into intermediate features before the diffusion process begins. This pre-processing step captures the essential structure and content of the image, allowing the subsequent diffusion steps to focus only on the modifications needed, thereby reducing the total number of diffusion steps required and improving processing speed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The image editing process is segmented into two distinct stages: (1) inversion stage where the input image is transformed into intermediate features using the inversion model, and (2) diffusion stage where modifications are applied to these features. This segmentation allows each component to be optimized independently, reducing overall processing time while maintaining image quality.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If conventional image editing is used, then modifications can be made, but image quality is low due to artifacts and undesired changes

Engineering Contradiction:
Improveimage qualityVSAvoidartifacts and undesired changes
Core Design Contradiction:
Manufacturing precisionVSObject-generated harmful factors

Solution Approach 1:

Intermediate features serve as an intermediary representation between the input image and the final edited image. These features capture the essential structure and content while being more amenable to controlled modification. By operating in this intermediate feature space rather than directly on pixel values, the system achieves higher fidelity edits with fewer artifacts and undesired changes.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The inversion model performs a preliminary transformation of the input image into intermediate features that preserve structural information. This pre-processing step creates a more stable foundation for subsequent editing operations, reducing the likelihood of artifacts and maintaining higher image quality throughout the diffusion process.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If conventional diffusion steps are used, then image modification is possible, but processing time is excessive

Engineering Contradiction:
Improveediting efficiencyVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The inversion model performs a rapid preliminary encoding of the input image into intermediate features before the diffusion process begins. This pre-processing step captures the essential structure and content of the image in compressed form, allowing the subsequent diffusion steps to operate more efficiently and reduce overall processing time while maintaining edit quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The editing process is divided into a fast inversion stage that creates intermediate features, followed by a more targeted diffusion stage that applies only necessary modifications. This segmentation eliminates redundant processing steps and focuses computational resources on the essential editing operations, significantly reducing total processing time.

Inventive Principle:
Principle #1Segmentation

4Measurement precision

If standard image editing is used, then changes can be made, but disentanglement of image elements is challenging causing undesired changes to non-targeted elements

Engineering Contradiction:
Improveediting precisionVSAvoiddisentanglement complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Intermediate features act as an intermediary representation that naturally separates different image elements and attributes. By transforming the input image into this intermediate space, the system achieves automatic disentanglement of image elements, making it easier to modify specific targets without affecting non-targeted elements, thereby improving editing precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The inversion model performs a preliminary transformation that organizes image information into intermediate features with inherent disentanglement properties. This pre-organization of image data before editing simplifies the subsequent modification process, allowing precise control over which elements are changed while protecting non-targeted elements from undesired alterations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250292463A1Real-time text-based disentangled real image editing
Publication Date: 2025.09.18 ADOBE INC
  • US20250292463A1 patent drawing
  • US20250292463A1 patent drawing
  • US20250292463A1 patent drawing

AI summary

A method, apparatus, non-transitory computer readable medium, and system for image processing include obtaining an input image depicting a first element, a text description of the input image, and a modification prompt describing a second element different from the first element, generating an intermediate output based on the input image and the text description, where the intermediate output represents the first element, and generating a synthetic image based on the intermediate output and the modification prompt, where the synthetic image replaces the first element from the input image with the second element from the modification prompt.