Video Coding Rate Control Using Machine Learning and Game Theory
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding technologies, particularly HEVC, face challenges in accurately handling bit allocation for Coding Tree Units (CTUs) due to drastic motions and scene changes, leading to suboptimal Rate-Distortion (R-D) model prediction and quality smoothness issues.
Innovation Solution
A joint machine learning and game theory-based framework for video coding rate control (RC) is introduced, utilizing a machine learning-based R-D model classification scheme and mixed R-D model cooperative bargaining game theory for improved bit allocation optimization, specifically for inter frame CTUs, to enhance prediction accuracy and quality smoothness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If spatial-temporal prediction and regression methods are used for predicting R-D model parameters, then the prediction process can be performed efficiently, but the prediction accuracy deteriorates especially for CTUs with drastic motions and scene changes
Solution Approach 1:
The patent changes the approach from using fixed spatial-temporal prediction parameters to using adaptive deep neural network parameters that are trained on historical R-D data. The network dynamically adjusts its predictions based on input features such as motion vectors, prediction mode, and quantization parameters, allowing accurate R-D modeling even under drastic motion and scene change conditions where traditional parameter methods fail.
Solution Approach 2:
The patent replaces the traditional mechanical/spatial-temporal prediction system with an intelligent deep neural network system. Instead of relying on fixed mathematical models that assume smooth temporal evolution, the neural network learns complex non-linear relationships from training data, substituting the rigid prediction mechanism with a flexible data-driven approach that adapts to varying video conditions.
2Reliability
If R-λ model based RC methods are used, then buffer control and quality smoothness can be maintained, but bit allocation optimization for CTUs deteriorates due to inability to handle drastic motions and scene changes
Solution Approach 1:
The patent segments the video coding process into independent CTU-level units, each processed separately through the deep neural network R-D model. This segmentation allows the system to handle each CTU independently, adapting bit allocation to local motion characteristics and scene changes without being constrained by global R-λ model assumptions, thereby improving CTU-level optimization while maintaining overall buffer control.
Solution Approach 2:
The patent introduces dynamic adaptability by using a deep neural network that can adjust its R-D model parameters in real-time based on the specific characteristics of each CTU, such as motion vectors, prediction mode, and quantization parameters. This dynamic approach replaces the static R-λ model with a flexible system that adapts to changing conditions, improving bit allocation optimization for CTUs with drastic motions and scene changes.
3Measurement precision
If fixed QP method is used, then R-D performance achieves optimal results, but device complexity and computational overhead increase due to multiple encoding attempts
Solution Approach 1:
The patent performs preliminary action by training a deep neural network on historical R-D data before actual video encoding. This pre-training phase allows the network to capture the relationship between coding parameters and R-D performance without requiring multiple encoding attempts during production. The trained network then provides accurate R-D predictions directly during encoding, achieving optimal R-D performance with reduced computational complexity.
Solution Approach 2:
The patent creates a virtual copy of the R-D relationship through the deep neural network, which learns to replicate the complex relationship between coding parameters and performance from training data. Instead of physically performing multiple encoding attempts to determine optimal parameters, the system uses the trained network copy to predict R-D performance, achieving the same accuracy with significantly reduced computational overhead.
Data Source
AI summary
Systems and methods which provide a joint machine learning and game theory modeling (MLGT) framework for video coding rate control (RC) are described. A machine learning based R-D model classification scheme may be provided to facilitate improved R-D model prediction accuracy and a mixed R-D model based game theory approach may be implemented to facilitate improved RC performance. For example, embodiments may provide inter frame Coding Tree Units (CTUs) level bit allocation and RC optimization in HEVC. Embodiments provide for the CTUs being classified into a plurality of categories, such as by using a support vector machine (SVM) based multi-classification scheme. An iterative solution search method may be implemented for the mixed R-D models based bit allocation method. Embodiments may additionally or alternatively refine the intra frame QP determination and the adaptive bit ratios among frames to facilitate improving the coding quality smoothness.


