Video Coding Rate Control Using Machine Learning and Game Theory

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video coding technologies, particularly HEVC, face challenges in accurately handling bit allocation for Coding Tree Units (CTUs) due to drastic motions and scene changes, leading to suboptimal Rate-Distortion (R-D) model prediction and quality smoothness issues.

Innovation Solution

A joint machine learning and game theory-based framework for video coding rate control (RC) is introduced, utilizing a machine learning-based R-D model classification scheme and mixed R-D model cooperative bargaining game theory for improved bit allocation optimization, specifically for inter frame CTUs, to enhance prediction accuracy and quality smoothness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If spatial-temporal prediction and regression methods are used for predicting R-D model parameters, then the prediction process can be performed efficiently, but the prediction accuracy deteriorates especially for CTUs with drastic motions and scene changes

Engineering Contradiction:
Improveprediction process efficiencyVSAvoidR-D model parameter prediction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent changes the approach from using fixed spatial-temporal prediction parameters to using adaptive deep neural network parameters that are trained on historical R-D data. The network dynamically adjusts its predictions based on input features such as motion vectors, prediction mode, and quantization parameters, allowing accurate R-D modeling even under drastic motion and scene change conditions where traditional parameter methods fail.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the traditional mechanical/spatial-temporal prediction system with an intelligent deep neural network system. Instead of relying on fixed mathematical models that assume smooth temporal evolution, the neural network learns complex non-linear relationships from training data, substituting the rigid prediction mechanism with a flexible data-driven approach that adapts to varying video conditions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If R-λ model based RC methods are used, then buffer control and quality smoothness can be maintained, but bit allocation optimization for CTUs deteriorates due to inability to handle drastic motions and scene changes

Engineering Contradiction:
Improvebuffer control and quality smoothnessVSAvoidCTU level bit allocation optimization
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the video coding process into independent CTU-level units, each processed separately through the deep neural network R-D model. This segmentation allows the system to handle each CTU independently, adapting bit allocation to local motion characteristics and scene changes without being constrained by global R-λ model assumptions, thereby improving CTU-level optimization while maintaining overall buffer control.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dynamic adaptability by using a deep neural network that can adjust its R-D model parameters in real-time based on the specific characteristics of each CTU, such as motion vectors, prediction mode, and quantization parameters. This dynamic approach replaces the static R-λ model with a flexible system that adapts to changing conditions, improving bit allocation optimization for CTUs with drastic motions and scene changes.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If fixed QP method is used, then R-D performance achieves optimal results, but device complexity and computational overhead increase due to multiple encoding attempts

Engineering Contradiction:
ImproveR-D performanceVSAvoidcomputational overhead
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by training a deep neural network on historical R-D data before actual video encoding. This pre-training phase allows the network to capture the relationship between coding parameters and R-D performance without requiring multiple encoding attempts during production. The trained network then provides accurate R-D predictions directly during encoding, achieving optimal R-D performance with reduced computational complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a virtual copy of the R-D relationship through the deep neural network, which learns to replicate the complex relationship between coding parameters and performance from training data. Instead of physically performing multiple encoding attempts to determine optimal parameters, the system uses the trained network copy to predict R-D performance, achieving the same accuracy with significantly reduced computational overhead.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10542262B2Systems and methods for rate control in video coding using joint machine learning and game theory
Publication Date: 2020.01.21 CITY UNIVERSITY OF HONG KONG
  • US10542262B2 patent drawing
  • US10542262B2 patent drawing
  • US10542262B2 patent drawing

AI summary

Systems and methods which provide a joint machine learning and game theory modeling (MLGT) framework for video coding rate control (RC) are described. A machine learning based R-D model classification scheme may be provided to facilitate improved R-D model prediction accuracy and a mixed R-D model based game theory approach may be implemented to facilitate improved RC performance. For example, embodiments may provide inter frame Coding Tree Units (CTUs) level bit allocation and RC optimization in HEVC. Embodiments provide for the CTUs being classified into a plurality of categories, such as by using a support vector machine (SVM) based multi-classification scheme. An iterative solution search method may be implemented for the mixed R-D models based bit allocation method. Embodiments may additionally or alternatively refine the intra frame QP determination and the adaptive bit ratios among frames to facilitate improving the coding quality smoothness.