Adaptive Video Frame Region Encoding for Bandwidth-Constrained PSNR Stability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video encoding methods face challenges in maintaining a high peak signal-to-noise ratio (PSNR) and minimizing distortion, especially when transmission bandwidth is limited or variable, leading to large fluctuations in PSNR due to the use of a single resolution for encoding and decoding.
Innovation Solution
A method and apparatus for video decoding and encoding that partition a video frame into regions, each encoded and decoded using a corresponding resolution from a set of at least two different resolutions, with syntax elements indicating the resolution used for each region, allowing adaptive encoding and decoding based on transmission bandwidth.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If a high resolution is used for encoding different blocks in a video frame, then the video quality is improved, but the transmission bandwidth consumption increases significantly
Solution Approach 1:
The patent applies local quality by partitioning the video frame into multiple blocks and assigning different resolution levels to different blocks based on their importance. Important blocks (e.g., containing key objects or motion) are encoded at high resolution, while less important blocks are encoded at lower resolution. This resolves the contradiction by maintaining high video quality where needed while reducing overall bandwidth consumption.
Solution Approach 2:
The patent segments the video frame into multiple blocks that can be independently encoded at different resolutions. Each block is evaluated for its importance, and the encoding resolution is adapted accordingly. This segmentation approach allows the system to optimize the trade-off between video quality and bandwidth usage on a block-by-block basis rather than applying a uniform resolution to the entire frame.
2Quantity of substance
If a low resolution is used for encoding different blocks in a video frame, then the transmission bandwidth consumption is reduced, but the video quality deteriorates
Solution Approach 1:
The patent applies local quality by partitioning the video frame into multiple blocks and assigning different resolution levels to different blocks based on their importance. Important blocks (e.g., containing key objects or motion) are encoded at high resolution, while less important blocks are encoded at lower resolution. This resolves the contradiction by maintaining high video quality where needed while reducing overall bandwidth consumption.
Solution Approach 2:
The patent segments the video frame into multiple blocks that can be independently encoded at different resolutions. Each block is evaluated for its importance, and the encoding resolution is adapted accordingly. This segmentation approach allows the system to optimize the trade-off between video quality and bandwidth usage on a block-by-block basis rather than applying a uniform resolution to the entire frame.
3Device complexity
If a single resolution is used for encoding all blocks in a video frame, then the encoding process is simplified, but the PSNR fluctuates significantly with varying bandwidth conditions
Solution Approach 1:
The patent applies dynamics by making the encoding resolution adaptive rather than fixed. The system dynamically selects the resolution for each block based on importance metrics and current bandwidth conditions. This dynamic approach allows the encoding process to respond to varying conditions, maintaining stable PSNR performance across different bandwidth scenarios while accepting increased encoding complexity.
Solution Approach 2:
The patent changes the resolution parameter on a block-by-block basis rather than using a single fixed resolution for the entire frame. By adjusting the resolution parameter according to block importance and bandwidth availability, the system achieves stable PSNR performance across varying conditions. This parameter change approach transforms the encoding process from static to adaptive.
Data Source
AI summary
Disclosed is a video decoding method, including: obtaining a current video frame, the current video frame being partitioned into a plurality of regions; obtaining a syntax element carried in syntax data corresponding to each of the plurality of regions, the syntax element being used for indicating a resolution used to decode the region, and a plurality of resolutions used to decode the plurality of regions including at least two different resolutions; and decoding the each of the plurality of regions by using the resolution corresponding to the region. The plurality of resolutions are determined according to a transmission bandwidth of a video stream including the current video frame from a source to a destination, e.g., by comparing the transmission bandwidth with a preset bandwidth threshold.


