Jointly Trained Text Encoder for Image Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional text-based image generation systems use fixed text encoders that were not specifically trained for generative models, leading to sub-optimal text-image alignments.

Innovation Solution

An image generation system that jointly trains a text encoder and an image generation model to improve text-image alignment by generating images based on text embeddings, where the text encoder is trained to produce embeddings that accurately capture semantic information from text prompts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a fixed text encoder is used that was not trained for generative models, then the system complexity is reduced and ease of manufacture is improved, but text-image alignment deteriorates

Engineering Contradiction:
Improveease of manufactureVSAvoidtext-image alignment
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent merges the text encoder and image generation model into a jointly trained system. The text encoder is specifically trained alongside the image generation model to produce text embeddings that are optimally aligned with the generative process, resolving the contradiction between system simplicity and alignment quality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The text encoder undergoes preliminary joint training with the image generation model before deployment. This preliminary action ensures that the text encoder learns to produce embeddings that are specifically tailored for the generative model, achieving high text-image alignment before the system is manufactured.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If a fixed text encoder is used, then device complexity is reduced, but text-image alignment deteriorates

Engineering Contradiction:
Improvedevice complexityVSAvoidtext-image alignment
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent combines the text encoder and image generation model into an integrated jointly-trained system. This merging allows the text encoder to be optimized specifically for the generative model's requirements, achieving high text-image alignment without requiring a completely separate complex system.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system transitions from a static fixed text encoder to a dynamic jointly-trained encoder that adapts its parameters during training with the image generation model. This dynamic training process enables the text encoder to optimize its output embeddings for the specific generative model, improving alignment without permanent structural complexity.

Inventive Principle:
Principle #15Dynamics

3Manufacturing precision

If jointly trained text encoder and image generation model are used, then text-image alignment is improved, but device complexity increases

Engineering Contradiction:
Improvetext-image alignmentVSAvoiddevice complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent merges the text encoder and image generation model into a unified jointly-trained system. This consolidation achieves high text-image alignment by optimizing both components together, while the merged structure can be more efficient than maintaining separate independently-trained systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The jointly-trained system serves multiple functions: the text encoder produces embeddings optimized for generation, and the image generation model creates images from these embeddings. This multi-functionality achieved through joint training improves text-image alignment while avoiding the complexity of multiple separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240320873A1Text-based image generation using an image-trained text
Publication Date: 2024.09.26 ADOBE INC
  • US20240320873A1 patent drawing
  • US20240320873A1 patent drawing
  • US20240320873A1 patent drawing

AI summary

A method, apparatus, non-transitory computer readable medium, and system for image generation include obtaining a text prompt and encoding, using a text encoder jointly trained with an image generation model, the text prompt to obtain a text embedding. Some embodiments generate, using the image generation model, a synthetic image based on the text embedding.