Dual-VAE Latent Training for Semantically Aligned Image and Video Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems suffer from inaccuracies in generating high-quality images and videos, inefficiencies in resource consumption, and operational inflexibilities in responding to media generation requests, particularly in maintaining semantic alignment and spatial-temporal relationships.

Innovation Solution

A dual-variational autoencoder model is used to reconstruct frames, combined with a single-stream transformer that incorporates improved positional encoding strategies and a diffusion transformer model for seamless knowledge transfer across modalities, enabling efficient and accurate generation of images and videos.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional diffusion models are used for media generation, then semantic alignment can be achieved, but manufacturing precision and reliability deteriorate due to inaccuracies in generating high-quality images and videos

Engineering Contradiction:
Improveimage and video generation qualityVSAvoidsemantic alignment maintenance
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent segments the diffusion model into two distinct components: a 2D VAE for processing individual frame images and a 3D VAE for processing temporal sequences. This segmentation allows each component to specialize in its specific function, with the 2D VAE achieving high-fidelity image reconstruction and the 3D VAE maintaining temporal coherence, thereby resolving the contradiction between generation quality and semantic alignment reliability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces latent representations as an intermediary between the input images and the generated output. The 2D VAE encodes input images into latent representations, which are then processed by the 3D VAE to generate temporal sequences. This intermediary latent space enables precise control over both spatial quality and temporal semantics, resolving the contradiction between manufacturing precision and reliability

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If comprehensive training on both images and videos is performed, then generation quality improves, but resource consumption increases

Engineering Contradiction:
Improvemedia generation accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The training process is segmented into two distinct stages: first training the 2D VAE on individual frames to achieve high-quality image reconstruction, then training the 3D VAE on temporal sequences to achieve coherent video generation. This segmentation allows efficient resource utilization by focusing computational power on specific tasks at each stage rather than simultaneously training a monolithic model on all data

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The 2D VAE is trained in advance on image data before the 3D VAE training begins. This preliminary action prepares high-quality latent representations that serve as the foundation for subsequent video generation training, enabling the system to achieve comprehensive training quality while managing computational resources efficiently through staged preparation

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If separate models are used for image and video processing, then operational flexibility improves, but device complexity increases

Engineering Contradiction:
Improveoperational flexibilityVSAvoidmodel architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a nested architecture where the 2D VAE is embedded within the larger 3D VAE framework. The 2D VAE processes individual frames and produces latent representations that are then fed into the 3D VAE for temporal processing. This nesting allows the system to maintain operational flexibility by treating image and video processing as hierarchical levels rather than completely separate models, reducing overall system complexity while preserving adaptability

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS20260073579A1Dual-VAE for more efficient and effective diffusion model training
Publication Date: 2026.03.12 ADOBE INC
  • US20260073579A1 patent drawing
  • US20260073579A1 patent drawing
  • US20260073579A1 patent drawing

AI summary

The present disclosure relates to systems, methods, and non-transitory computer-readable media that leverages a dual-variational autoencoder model. For example, the disclosed systems generate an image embedding from a first frame of a sequence of frames by using a two-dimensional variational autoencoder. Moreover, the disclosed systems generate motion embeddings from motion within a video by using a three-dimensional variational autoencoder. Further, the disclosed systems generate a reconstructed image from the image embedding and a reconstructed video from the motion embeddings and the image embedding. Additionally, the disclosed systems modify parameters of a dual-variational autoencoder model based on a measure of accuracy of the reconstructed image and the reconstructed video.