Anchor Token Video Generation for Accurate Frame Inclusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for generating digital videos suffer from inaccuracies, inefficiencies, and operational inflexibilities, failing to accurately include specified content, maintaining image/video and text semantic alignment, and consuming excessive computational resources.

Innovation Solution

A generative AI digital visual system employs a frame anchoring technique using a diffusion transformer model to transform image tokens into anchor tokens, denoising noised tokens, and ensuring the digital image is included as a coherent frame in the generated video, with improved training methods and spatial-temporal positional encodings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing systems generate digital videos from input prompts, then video generation capability is provided, but accuracy in including specified content and maintaining semantic alignment deteriorates

Engineering Contradiction:
Improveaccuracy of including specified contentVSAvoidvideo generation quality
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the video generation process into distinct components: generating image tokens from the input image, creating anchor tokens with timestep information, generating noised tokens, and using a diffusion transformer model to combine these into final video tokens. This segmentation allows each component to be optimized independently, improving both accuracy of content inclusion and overall generation quality

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces anchor tokens as an intermediary element that bridges the input image and the final video output. These anchor tokens contain embedded timestep information and serve as a reference guide for the diffusion transformer model, ensuring that the generated video accurately reflects the specified content while maintaining semantic alignment throughout the generation process

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If existing systems process video generation requests, then generative tasks are performed, but computational resource consumption increases

Engineering Contradiction:
Improvegenerative task capabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary processing by generating image tokens from the input image before the main video generation process. These image tokens are then transformed into anchor tokens with timestep information embedded, preparing the data in advance for the diffusion transformer model. This preliminary action reduces the computational burden during the actual video generation by having the data ready in the appropriate format

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the input image into a different parameter space by converting it into image tokens, then into anchor tokens with timestep parameters embedded. This parameter transformation allows the diffusion transformer model to work more efficiently by operating on token representations rather than raw pixel data, reducing computational resource consumption while maintaining generative task capability

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20260073484A1Generating a digital video utilizing a set of anchor tokens to denoise noised tokens
Publication Date: 2026.03.12 ADOBE INC
  • US20260073484A1 patent drawing
  • US20260073484A1 patent drawing
  • US20260073484A1 patent drawing

AI summary

The present disclosure relates to systems, methods, and non-transitory computer-readable media that generate a digital video based on denoised tokens. In particular, the disclosed systems generate a set of image tokens from a digital image that is part of an image-to-video request. Furthermore, the disclosed systems generate a set of anchor tokens from the set of image tokens by adding a timestep embedding to the set of image tokens that indicates that the set of anchor tokens are fully denoised. Further, the disclosed systems generate combined tokens from the set of anchor tokens and noised tokens that are generated from noise. Moreover, the disclosed systems generate denoised tokens by using a diffusion transformer model to process the combined tokens. Further, from the denoised tokens, the disclosed systems generate the digital video that includes the digital image.