Augmented Diffusion Inversion Latent Trajectory Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for text-guided image generation and editing suffer from instability, distortion, and inaccuracy, particularly in inversion processes, leading to failed image manipulation attempts.

Innovation Solution

The method employs augmented diffusion inversion using latent trajectory optimization, which involves generating an augmented noise vector by deterministically applying noise to the image, encoding it with a bias correction variable, and using latent trajectory optimization to determine diffusion trajectories for the augmented noise vector.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional inversion processes are used for text-guided diffusion, then the process can be performed, but the results are unstable, distorted, and inaccurate

Engineering Contradiction:
Improvestability of inversion processVSAvoidaccuracy of image manipulation
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent applies preliminary action by performing deterministic noise application and bias correction before the main inversion process. The method deterministically applies noise to the input image to generate an initial noise vector, then corrects biases in the diffusion trajectories before inversion, ensuring more stable and accurate results throughout the generation process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes parameters by introducing bias correction variables that adjust the diffusion trajectories. By modifying the noise schedule and applying corrective biases to the trajectory parameters, the method transforms the inversion process to achieve both stability and accuracy simultaneously

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If inversion is performed with text perturbations, then text-guided generation can occur, but inaccuracy increases

Engineering Contradiction:
Improvetext-guided generation capabilityVSAvoidinversion accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies feedback by using the biased diffusion trajectories to guide the inversion process. The method calculates bias corrections based on the difference between conditional and unconditional trajectories, then feeds this correction back into the inversion process to maintain accuracy despite text perturbations

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces bias correction variables as intermediaries between the text prompt and the inversion process. These corrective terms mediate the interaction between text conditioning and image generation, allowing text-guided generation while maintaining inversion accuracy through the intermediary correction layer

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If existing inversion techniques are used, then image manipulation can be attempted, but resource usage is high

Engineering Contradiction:
Improveimage manipulation capabilityVSAvoidcomputational resource usage
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the essential corrective components from the full diffusion process. By separating the bias correction from the complete inversion procedure and applying only the necessary corrective terms, the method reduces computational resource usage while maintaining image manipulation capability

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12236559B2Augmented diffusion inversion using latent trajectory optimization
Publication Date: 2025.02.25 INTUIT INC
  • US12236559B2 patent drawing
  • US12236559B2 patent drawing
  • US12236559B2 patent drawing

AI summary

Augmented Denoising Diffusion Implicit Models (“DDIMs”) using a latent trajectory optimization process can be used for image generation and manipulation using text input and one or more source images to create an output image. Noise bias and textual bias inherent in the model representing the image and text input is corrected by correcting trajectories previously determined by the model at each step of a diffusion inversion process by iterating multiple starts the trajectories to find determine augmented trajectories that minimizes loss at each step. The trajectories can be used to determine an augmented noise vector, enabling use of an augmented DDIM and resulting in more accurate, stable, and responsive text-based image manipulation.