Diffusion Model Noise Representation for Text-Image Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-image synthesis systems suffer from computational inaccuracies and operational inflexibility, particularly in generating images with multiple text concepts, leading to missing concepts and limited operational flexibility.

Innovation Solution

The system utilizes a text-image alignment loss to train a diffusion model, implementing lightweight finetuning and generating noise representations for multiple text concepts to enhance the generation of high-quality text-conditioned images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing generative models are used for text-to-image synthesis, then image generation capability is achieved, but computational accuracy deteriorates leading to missing text concepts

Engineering Contradiction:
Improvecomputational accuracyVSAvoidconcept preservation
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the text prompt into individual text concepts and generates separate noise representations for each concept. This segmentation allows the model to process and align each text concept independently with its corresponding image region, thereby improving computational accuracy and preventing concept omission.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces noise representations as intermediary elements between text concepts and image generation. By generating noise representations for individual text concepts and combining them to form a prompt noise representation, the model creates a mediating layer that facilitates accurate alignment and prevents direct mapping errors.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If existing generative models process text prompts, then image generation is produced, but operational flexibility deteriorates due to limited concept handling

Engineering Contradiction:
Improveoperational flexibilityVSAvoidconcept preservation
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

By segmenting the text prompt into individual text concepts and generating separate noise representations for each, the model gains operational flexibility in handling multiple concepts simultaneously. This segmentation enables the model to adapt to various text prompts while maintaining ease of operation through systematic processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal processing mechanism that can handle multiple text concepts through a single unified approach. The noise representation generation and combination process serves multiple functions: individual concept processing, concept alignment, and image generation guidance, thereby enhancing operational flexibility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If text prompts with multiple concepts are processed, then comprehensive concept coverage is achieved, but system complexity increases

Engineering Contradiction:
Improvenumber of text conceptsVSAvoidprocessing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments the complex task of processing multi-concept text prompts into simpler sub-tasks: generating noise representations for individual concepts and combining them. This segmentation reduces processing complexity by breaking down the overall complexity into manageable steps.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges individual concept noise representations to form a unified prompt noise representation. This merging process simplifies the overall processing by consolidating multiple concept representations into a single structure that can be efficiently processed by the diffusion model.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250078327A1Utilizing individual-concept text-image alignment to enhance compositional capacity of text-to-image models
Publication Date: 2025.03.06 ADOBE INC
  • US20250078327A1 patent drawing
  • US20250078327A1 patent drawing
  • US20250078327A1 patent drawing

AI summary

The present disclosure relates to systems, methods, and non-transitory computer-readable media that utilize a text-image alignment loss to train a diffusion model to generate digital images from input text. In particular, in some embodiments, the disclosed systems generate a prompt noise representation form a text prompt with a first text concept and a second text concept using a denoising step of a diffusion neural network. Further, in some embodiments, the disclosed systems generate a first concept noise representation from the first text concept and a second concept noise representation from the second text concept. Moreover, the disclosed systems combine the first and second concept noise representation to generate a combined concept noise representation. Accordingly, in some embodiments, by comparing the combined concept noise representation and the prompt noise representation, the disclosed systems modify parameters of the diffusion neural network.