Joint Embedding Image Generation for Text and Style Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional text-to-image models are limited in their ability to accurately generate images that reflect both textual descriptions and stylistic inputs, often failing to effectively combine and interpret information from diverse domains.

Innovation Solution

The proposed system conditions an image generation model on a joint embedding space that maps both text and image embeddings, allowing it to generate images that accurately reflect textual descriptions and stylistic inputs by learning from high-level textual descriptions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional text-to-image models are used, then image generation is simpler, but the accuracy of reflecting both textual descriptions and stylistic inputs deteriorates

Engineering Contradiction:
Improveimage generation accuracyVSAvoidmodel structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges text embedding space and image embedding space into a joint embedding space, allowing the model to process both textual descriptions and stylistic inputs in a unified representation. This combination enables the model to accurately reflect both types of inputs simultaneously without requiring separate processing pipelines.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces mapping networks as intermediary components that transform text embeddings and image embeddings into the joint embedding space. These mapping networks serve as mediators that enable seamless integration of different input types while maintaining their respective characteristics.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If separate text and image processing is used, then model simplicity is maintained, but the ability to combine information from diverse domains deteriorates

Engineering Contradiction:
Improvedomain information integrationVSAvoidembedding space structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The joint embedding space serves as a universal representation that can accommodate both text and image embeddings. This multi-functional space allows the model to process diverse domain information (textual descriptions and stylistic inputs) through a single unified mechanism, enhancing adaptability across different input types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If text-based image generation is used, then ease of use for laypersons is improved, but the ability to incorporate stylistic input from original images deteriorates

Engineering Contradiction:
Improveuser accessibilityVSAvoidinput modality flexibility
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system accepts multiple input modalities (text prompts and image prompts) through a unified architecture. Users can provide textual descriptions, image examples, or both, and the model processes all inputs through the joint embedding space, maintaining ease of use while enhancing input flexibility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12586259B2Image generation using a text and image conditioned machine learning model
Publication Date: 2026.03.24 ADOBE INC
  • US12586259B2 patent drawing
  • US12586259B2 patent drawing
  • US12586259B2 patent drawing

AI summary

A method, apparatus, non-transitory computer readable medium, and system for image generation include obtaining a text embedding of a text prompt and an image embedding of an image prompt. Some embodiments map the text embedding into a joint embedding space to obtain a joint text embedding and map the image embedding into the joint embedding space to obtain a joint image embedding. Some embodiments generate a synthetic image based on the joint text embedding and the joint image embedding.