Machine Learning Video Encoder Quantization Parameter Transformation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern video codecs face increased computational complexity due to improved coding efficiency, leading to longer processing times, despite advancements in coding techniques like rate-distortion optimization and quantization parameter management.

Innovation Solution

A method and apparatus that utilize a machine-learning model to encode image blocks by presenting a derived value from a non-linear quantization parameter, which is used to calculate a Lagrange multiplier for rate-distortion calculations, reducing computational complexity while maintaining coding efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If improved coding efficiency techniques are used in video encoders, then coding efficiency is improved, but computational complexity increases

Engineering Contradiction:
Improvecoding efficiencyVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent transforms the quantization parameter through a non-linear function to create a derived value that better correlates with rate-distortion characteristics. This parameter transformation enables the machine learning model to make more accurate mode decisions with reduced computational iterations, thereby maintaining coding efficiency while lowering computational complexity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional rate-distortion optimization algorithms with a machine learning model that has been trained on video data. The ML model directly predicts optimal encoding modes based on the derived quantization parameter, substituting complex iterative optimization mechanics with a trained prediction system that achieves similar or better coding efficiency with reduced computational overhead

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If more computation time is allocated to achieve improved coding efficiency, then coding efficiency improves, but processing time increases

Engineering Contradiction:
Improvecoding efficiencyVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary training of the machine learning model offline using extensive video data and rate-distortion optimization examples. This preliminary action embeds learned patterns into the model, enabling real-time encoding to use pre-learned knowledge rather than performing complex optimization calculations during actual video processing, thus reducing processing time while maintaining coding efficiency

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a simplified computational model that copies the essential decision-making patterns from complex rate-distortion optimization algorithms. The machine learning model replicates the behavior of thorough optimization processes through trained weights and biases, achieving similar coding efficiency outcomes with significantly reduced computational time during actual encoding operations

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11310501B2Efficient use of quantization parameters in machine-learning models for video coding
Publication Date: 2022.04.19 GOOGLE LLC
  • US11310501B2 patent drawing
  • US11310501B2 patent drawing
  • US11310501B2 patent drawing

AI summary

Encoding an image block using a quantization parameter includes presenting, to an encoder that includes a machine-learning model, the image block and a value derived from the quantization parameter, where the value is a result of a non-linear function using the quantization parameter as input, where the non-linear function relates to a second function used to calculate, using the quantization parameter, a Lagrange multiplier that is used in a rate-distortion calculation, and where the machine-learning model is trained to output mode decision parameters for encoding the image block; obtaining the mode decision parameters from the encoder; and encoding, in a compressed bitstream, the image block using the mode decision parameters.