Partial Media Decompression for Compact Generative AI Tokens

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Training generative AI models with uncompressed media data, such as high-resolution images and videos, is challenging due to the vast volume of information and redundant content, which complicates the training process.

Innovation Solution

Using partially decompressed data as input for generative AI models, where compressed data is partially decompressed to extract syntax elements, converted into tokens, and encoded for the model, allowing efficient training and inference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If uncompressed media data is used for training generative AI models, then the model can access complete information content, but the data volume becomes excessively large and contains significant redundant information

Engineering Contradiction:
Improveinformation content representationVSAvoiddata volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent extracts syntax elements from compressed media data through partial decompression, converting them into tokens that represent the essential information content. This extraction process removes redundant information while preserving the core semantic meaning, thereby reducing data volume without significant loss of information content.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms the representation parameters of media data by converting raw sample values into syntax elements and then into tokens. This parameter transformation changes the data from a high-volume, redundant format to a compact, information-dense format that is more suitable for AI model training.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If uncompressed media data is used for training, then all detail information is available, but the training process becomes more complex and less efficient

Engineering Contradiction:
Improvetraining effectivenessVSAvoidtraining process complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent introduces syntax elements and tokens as intermediary representations between the original compressed data and the AI model. These intermediaries simplify the input data structure, making it more amenable to processing while preserving the essential information needed for effective training.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the media data processing into distinct stages: compression, partial decompression to extract syntax elements, token conversion, and model training. This segmentation allows each stage to be optimized independently, reducing overall training complexity while maintaining effectiveness.

Inventive Principle:
Principle #1Segmentation

3Quantity of substance

If compressed data is used directly as input, then data volume is reduced, but the data becomes less accessible and harder to process

Engineering Contradiction:
Improvedata volumeVSAvoiddata accessibility
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The patent applies partial decompression rather than full decompression, extracting only the syntax elements needed for training while leaving the data in a compressed-like token format. This partial action maintains data compactness while improving accessibility and processability for the AI model.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250240440A1Partial decompression of compressed media for input to a generative artificial intelligence model
Publication Date: 2025.07.24 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250240440A1 patent drawing
  • US20250240440A1 patent drawing
  • US20250240440A1 patent drawing

AI summary

A computer system performs operations to prepare input to a generative artificial intelligence (“AI”) model. The system receives compressed data for media, which has been compressed according to a media compression format to produce the compressed data. The system partially decompresses the compressed data (e.g., performing parsing and entropy decoding operations). This produces syntax elements of the compressed data according to the media compression format. The system converts the syntax elements into tokens that represent the syntax elements, respectively. Unlike the syntax elements (in the media compression format), the tokens are encoded in an input format for the generative AI model. The system stores the tokens in memory or storage, from which the system can provide the tokens to the generative AI model for use in a training process or inference process for media synthesis, media compression, media decompression, or another purpose.