Single-Stream Transformer for Text-Aligned Image and Video Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems suffer from deficiencies in accuracy, efficiency, and operational flexibility in generating high-quality images and videos, often failing to simultaneously preserve image and video reconstruction, lacking strong text and image/video semantic alignment, and requiring excessive computational resources and time.
Innovation Solution
The generative AI digital visual system employs a dual-variational autoencoder model and improved positional encoding strategies, using a single-stream transformer to unify diverse inputs and enable seamless knowledge transfer across modalities, optimizing the diffusion transformer model with a mixed training strategy that includes image, key-frame, and video clip training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If distinct architectures are used for generating content in different modalities, then operational flexibility is improved, but device complexity increases
Solution Approach 1:
The patent merges distinct architectures for image and video generation into a unified transformer model that processes both modalities through a single architecture, eliminating the need for separate models while maintaining operational flexibility across different content types
Solution Approach 2:
The transformer model is designed with universal multi-functionality to handle both image and video generation tasks within a single system, allowing the same architecture to serve multiple generative purposes without requiring modality-specific designs
2Productivity
If conventional systems generate images and videos, then media content is produced, but accuracy and reconstruction quality deteriorate
Solution Approach 1:
The patent changes key parameters including the use of cosine attention instead of standard attention mechanisms, modified positional encodings, and adjusted diffusion sampling parameters to simultaneously improve both generation capability and reconstruction quality across image and video modalities
3Measurement precision
If semantic alignment between text and image/video is enhanced, then generation accuracy is improved, but computational resources increase
Solution Approach 1:
The patent performs preliminary action by pre-computing and storing cosine attention matrices and positional encoding transformations during the training phase, allowing the model to achieve high semantic alignment accuracy during inference without requiring excessive computational resources at generation time
4Manufacturing precision
If training strategies are optimized for image generation, then image quality is improved, but video generation capability deteriorates
Solution Approach 1:
The transformer model is designed with universal multi-functionality to handle both image and video generation tasks within a single system, allowing the same architecture to serve multiple generative purposes without requiring modality-specific designs
Solution Approach 2:
The patent changes key parameters including the use of cosine attention instead of standard attention mechanisms, modified positional encodings, and adjusted diffusion sampling parameters to simultaneously improve both generation capability and reconstruction quality across image and video modalities
Data Source
AI summary
The present disclosure relates to systems, methods, and non-transitory computer-readable media that generates an image or a video from a text prompt. For example, the disclosed systems receive a text prompt and generates text tokens from the text prompt. Moreover, the disclosed systems generate combined tokens by combining the text tokens with noised tokens. Further, the disclosed systems generate denoised tokens by removing noise from noised tokens in a manner that incorporates a context indicated by the text tokens and further generates an image or video from the denoised tokens.


