ML-Based Bitrate Allocation for Non-Backward Compatible Video
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional methods for allocating bitrates to base and enhancement layers in non-backward compatible dual or multi-layer codec video systems fail to effectively prioritize visual attention areas, leading to suboptimal video quality, as they do not accurately correlate peak signal-to-noise ratio (PSNR) with the human visual system's perception of dynamic range.
Innovation Solution
A machine learning-based method is employed to determine bitrate allocations for base and enhancement layers by identifying and prioritizing features that attract human visual attention, using a supervised classification approach and feature extraction techniques to optimize bitrate distribution based on the importance of visual content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If traditional bitrate allocation methods are used in non-backward compatible multi-layer codec, then the encoding process is simple and fast, but the video quality does not effectively preserve highlight areas and does not correlate with human visual perception
Solution Approach 1:
The patent applies preliminary action by training a machine learning model in advance on a dataset of videos with various bitrate allocations and human quality assessments. The trained model is then used during encoding to predict optimal bitrate allocations for base and enhancement layers based on video content features, avoiding the need for complex real-time optimization during the encoding process itself.
Solution Approach 2:
The patent introduces a machine learning model as an intermediary between the video content and the bitrate allocation decision. The model takes video features as input and outputs predicted optimal bitrate allocations, serving as a mediator that translates content characteristics into quality-optimized encoding parameters without requiring direct complex optimization algorithms.
2Adaptability or versatility
If more bits are allocated to base layer, then backward compatibility is improved, but the preservation of highlight areas and overall visual quality deteriorates
Solution Approach 1:
The patent applies dynamics by making the bitrate allocation adaptive rather than static. The machine learning model dynamically determines the optimal split between base and enhancement layer bitrates based on the specific video content's characteristics, such as the presence and importance of highlight areas. This allows the system to adjust allocations scene-by-scene or frame-by-frame to optimize both compatibility and quality.
Solution Approach 2:
The patent applies local quality by differentiating bitrate allocation based on local content characteristics. The machine learning model analyzes video features to identify regions with important visual information (such as highlight areas) and adjusts the enhancement layer bitrate allocation accordingly, allocating more bits to preserve these critical regions while maintaining adequate base layer quality for compatibility.
3Measurement precision
If bitrate allocation is optimized for PSNR, then signal quality is improved, but the correlation with human visual system perception of dynamic range deteriorates
Solution Approach 1:
The patent applies parameter changes by shifting the optimization criterion from PSNR (peak signal-to-noise ratio) to a human-perception-based quality metric. The machine learning model is trained using subjective quality assessments from human observers as ground truth, enabling it to learn bitrate allocations that prioritize visual perception relevance over traditional signal quality measures. This fundamentally changes the optimization parameter from engineering-centric PSNR to perception-centric quality.
Data Source
Figure 1~2
Figure 3A~4
Figure 5
AI summary
Novel methods and systems for non-backward compatible video encoding are disclosed. The bitrates of the base layer and enhancement layer are dynamically assigned based on features found in scenes in the video compared to a machine learned quality classifier.