Unified Visual Tokenizer With Window and Causal Attention

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current generative models for visual processing, such as language model (LM)-based and diffusion models, face limitations in application flexibility and data scalability due to tokenizers being specifically designed for either images or videos, lacking compatibility between them.

Innovation Solution

A visual encoder architecture that employs a decoupled approach with window and causal attention mechanisms in spatial and temporal dimensions, respectively, to process both static images and dynamic videos, using a unified tokenizer to generate encoding representations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If separate tokenizers are designed for images and videos, then processing specialization is improved, but compatibility between image and video processing deteriorates

Engineering Contradiction:
Improveprocessing specializationVSAvoidcompatibility between image and video processing
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by designing a unified tokenizer that can process both image and video data. The tokenizer uses the same architecture and parameters for both modalities, enabling a single model to handle multiple types of visual data without requiring separate specialized components, thus achieving compatibility while maintaining processing effectiveness

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If different processing architectures are used for images and videos, then modality-specific performance is improved, but model complexity increases

Engineering Contradiction:
Improvemodality-specific performanceVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the video processing into temporal segments (frames) that are processed independently through the same tokenizer architecture. This allows the model to handle video data efficiently by breaking it down into manageable units while using the same processing components, avoiding the need for entirely separate architectures for different modalities

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250384680A1Visual processing
Publication Date: 2025.12.18 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US20250384680A1 patent drawing
  • US20250384680A1 patent drawing
  • US20250384680A1 patent drawing

AI summary

According to embodiments of the disclosure, a method, an apparatus, a device, and a storage medium for visual processing are provided. A method includes: converting a plurality of image blocks divided from visual data into a plurality of embedding representations respectively, where the visual data includes an image or a video; extracting, by using a first processing block in a trained visual encoder, first feature information from the plurality of embedding representations according to a first attention mechanism; extracting, by using a second processing block in the visual encoder, second feature information from the first feature information according to a second attention mechanism; and generating, by using a tokenizer in the visual encoder, an encoding representation corresponding to the visual data based on the second feature information. In this manner, the encoding efficiency can be improved, and better universality and scalability can be achieved.