Single-Stream Transformer for Text-Aligned Image and Video Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems suffer from deficiencies in accuracy, efficiency, and operational flexibility in generating high-quality images and videos, often failing to simultaneously preserve image and video reconstruction, lacking strong text and image/video semantic alignment, and requiring excessive computational resources and time.

Innovation Solution

The generative AI digital visual system employs a dual-variational autoencoder model and improved positional encoding strategies, using a single-stream transformer to unify diverse inputs and enable seamless knowledge transfer across modalities, optimizing the diffusion transformer model with a mixed training strategy that includes image, key-frame, and video clip training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If distinct architectures are used for generating content in different modalities, then operational flexibility is improved, but device complexity increases

Engineering Contradiction:
Improveoperational flexibilityVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges distinct architectures for image and video generation into a unified transformer model that processes both modalities through a single architecture, eliminating the need for separate models while maintaining operational flexibility across different content types

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The transformer model is designed with universal multi-functionality to handle both image and video generation tasks within a single system, allowing the same architecture to serve multiple generative purposes without requiring modality-specific designs

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If conventional systems generate images and videos, then media content is produced, but accuracy and reconstruction quality deteriorate

Engineering Contradiction:
Improvemedia generation capabilityVSAvoidreconstruction quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent changes key parameters including the use of cosine attention instead of standard attention mechanisms, modified positional encodings, and adjusted diffusion sampling parameters to simultaneously improve both generation capability and reconstruction quality across image and video modalities

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If semantic alignment between text and image/video is enhanced, then generation accuracy is improved, but computational resources increase

Engineering Contradiction:
Improvesemantic alignment accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary action by pre-computing and storing cosine attention matrices and positional encoding transformations during the training phase, allowing the model to achieve high semantic alignment accuracy during inference without requiring excessive computational resources at generation time

Inventive Principle:
Principle #10Preliminary action

4Manufacturing precision

If training strategies are optimized for image generation, then image quality is improved, but video generation capability deteriorates

Engineering Contradiction:
Improveimage generation qualityVSAvoidvideo generation capability
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The transformer model is designed with universal multi-functionality to handle both image and video generation tasks within a single system, allowing the same architecture to serve multiple generative purposes without requiring modality-specific designs

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes key parameters including the use of cosine attention instead of standard attention mechanisms, modified positional encodings, and adjusted diffusion sampling parameters to simultaneously improve both generation capability and reconstruction quality across image and video modalities

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20260073580A1Single stream transformer for text-to-image/video synthesis
Publication Date: 2026.03.12 ADOBE INC
  • US20260073580A1 patent drawing
  • US20260073580A1 patent drawing
  • US20260073580A1 patent drawing

AI summary

The present disclosure relates to systems, methods, and non-transitory computer-readable media that generates an image or a video from a text prompt. For example, the disclosed systems receive a text prompt and generates text tokens from the text prompt. Moreover, the disclosed systems generate combined tokens by combining the text tokens with noised tokens. Further, the disclosed systems generate denoised tokens by removing noise from noised tokens in a manner that incorporates a context indicated by the text tokens and further generates an image or video from the denoised tokens.