Generative AI Media Synthesis from Partially Decompressed Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Training generative AI models with uncompressed media data, such as high-resolution images and videos, is challenging due to the vast volume of information and redundant content, which complicates the training process.

Innovation Solution

Utilizing partially decompressed data as input for generative AI models, where syntax elements from compressed media are converted into tokens and used for training and inference, reducing data volume and simplifying data organization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If uncompressed media data is used for training generative AI models, then the information content is complete and accurate, but the data volume is extremely large and contains significant redundant information

Engineering Contradiction:
Improveinformation completenessVSAvoiddata volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential syntax elements from compressed media data that are necessary for training generative AI models. Instead of using complete uncompressed data, the system identifies and extracts key syntax elements (such as motion vectors, prediction modes, and other structural parameters) that capture the important information while eliminating redundant pixel-level data. This extraction process resolves the contradiction by maintaining information completeness at the structural level while dramatically reducing data volume.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies preliminary compression and syntax element extraction before feeding data to the generative AI model. By performing media compression and syntax element identification in advance, the system prepares the data in a form that is both information-rich and computationally efficient. This preliminary action transforms the raw media data into a condensed representation that retains essential information while reducing redundancy, thus resolving the contradiction between information completeness and data volume.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If uncompressed high-resolution images and videos are used for training, then the training data contains all necessary details, but the training process becomes computationally complex and time-consuming

Engineering Contradiction:
Improvetraining accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent extracts essential syntax elements from compressed media data that capture the structural and semantic information necessary for training. By focusing on key parameters such as motion vectors, block partitioning modes, and prediction information, the system achieves high training accuracy without processing the full uncompressed pixel data. This extraction approach significantly reduces computational complexity while maintaining the essential training signals needed for learning effective representations.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms the input data representation from raw pixel values to compressed syntax elements. This parameter change fundamentally alters the data format from a high-dimensional pixel grid to a structured set of compression parameters. The transformation enables the model to learn from the organizational structure and semantic information in the data while avoiding the computational burden of processing millions of pixel values, thus resolving the contradiction between training accuracy and computational complexity.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If syntax elements from compressed data are used as input, then the data volume is reduced and processing is simplified, but the vocabulary diversity for the generative AI model is limited

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidvocabulary diversity
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent segments the compressed media data into multiple distinct syntax element categories, including motion vectors, prediction modes, block partitioning information, and other structural parameters. By dividing the data into these meaningful segments, the system creates a diverse set of input features that provide the generative AI model with varied information types. This segmentation approach maintains processing efficiency while enhancing vocabulary diversity, as each syntax element type contributes unique patterns and information for model learning.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the one-dimensional syntax element streams into multi-dimensional feature representations by organizing syntax elements according to their semantic categories and hierarchical relationships. This dimensional transformation enriches the vocabulary available to the generative AI model by creating structured feature spaces that capture different aspects of the media content. The system maintains processing efficiency through this transformation by using efficient encoding schemes while significantly expanding the effective vocabulary and feature diversity available for model training.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20250238968A1Media synthesis using a generative artificial intelligence model that accepts partially decompressed data as input
Publication Date: 2025.07.24 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250238968A1 patent drawing
  • US20250238968A1 patent drawing
  • US20250238968A1 patent drawing

AI summary

A computer system performs operations to synthesize media using a generative artificial intelligence (“AI”) model. The system receives input tokens that represent input syntax elements, respectively, of compressed data for input media, which has been compressed according to a media compression format. The system provides the input tokens to the generative AI model and receives predicted tokens from the generative AI model. The predicted tokens represent output syntax elements, respectively, of compressed data for output media. Finally, the system reconstructs the output media (e.g., converting the predicted tokens to the output syntax elements, and then decompressing the output syntax elements using a media decoder). The generative AI model is trained for media synthesis using a set of training data. In training, the system can measure loss in terms of conformity of the predicted tokens to syntax of the media compression format and/or based on ratings of the output media.