Bi-Directional Motion Refinement for Non-Equal Reference Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding standards like VVC face challenges in achieving optimal coding efficiency due to limitations in decoder-side motion vector refinement and bi-directional optical flow techniques, leading to suboptimal video quality and increased computational complexity.
Innovation Solution
Implementing bi-directional prediction methods for motion vector refinement, including decoder-side motion vector refinement (DMVR) and bi-directional optical flow (BDOF), with techniques such as bilateral template matching, multi-pass refinement, and weighted sum approaches to enhance accuracy and flexibility, allowing for non-equal distance reference pictures and combining sample-based and block-based refinements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If decoder-side motion vector refinement (DMVR) and bi-directional optical flow (BDOF) are implemented to improve motion vector accuracy, then video quality and coding efficiency are improved, but computational complexity increases
Solution Approach 1:
The patent divides the current block into multiple sub-blocks and performs motion vector refinement separately for each sub-block. This segmentation allows the complex BDOF calculations to be distributed across smaller regions, improving motion vector accuracy locally while managing overall computational complexity through parallel processing of sub-blocks
Solution Approach 2:
The patent applies BDOF refinement selectively rather than uniformly across all blocks. By using criteria such as motion magnitude thresholds and reference picture availability, the method performs partial refinement only where it provides significant benefit, thereby improving motion vector accuracy for critical blocks while reducing unnecessary computational complexity for blocks where refinement is not needed
2Adaptability or versatility
If bi-directional prediction with non-equal distance reference pictures is used to enhance flexibility in video coding, then adaptability to various video content is improved, but coding complexity increases
Solution Approach 1:
The patent enables dynamic selection of reference pictures from both list 0 and list 1 with non-equal distances to the current picture. The reference picture selection and BDOF application are dynamically adjusted based on motion characteristics, picture availability, and coding conditions, providing adaptability to various video content while managing complexity through conditional application
Solution Approach 2:
The patent changes the parameter of reference picture distance equality by allowing non-equal distance reference pictures in bi-directional prediction. This parameter change enables greater flexibility in handling diverse motion patterns and video content, while the associated complexity is managed through selective application based on coding conditions and motion thresholds
3Manufacturing precision
If multiple refinement passes (sample-based and subblock-based) are applied to improve motion compensation accuracy, then video quality is improved, but processing time increases
Solution Approach 1:
The patent performs preliminary motion estimation to obtain initial motion vectors before applying refinement passes. This preliminary action provides a good starting point for the subsequent sample-based and subblock-based refinements, reducing the number of iterations needed and thereby improving motion compensation accuracy while minimizing additional processing time
Solution Approach 2:
The patent implements multi-pass refinement where sample-based and subblock-based refinements are applied in periodic stages. Each pass builds upon the previous refinement results, progressively improving motion compensation accuracy. The periodic application of different refinement techniques allows the system to achieve high precision while managing processing time through staged refinement
Data Source
AI summary
Method and apparatus of using bi-directional prediction to refine MV are disclosed. According to one method, a sample-based refinement and a subblock-based refinement are determined for the current block. A final refinement for the current block is determined based on the sample-based refinement and the subblock-based refinement. According to another method, one or more high-level syntaxes are signalled or parsed, where the high-level syntaxes indicate whether non-equal distance reference pictures are allowed for bi-directional motion refinement. In response to the high-level syntaxes indicating the non-equal distance reference pictures being allowed, a refined MV is determined for at least one block in the current picture based on a reference picture in list 0 and a reference picture in list 1, where the picture distance between the first reference picture and the current picture and the picture distance between the second reference picture and the current picture are different.


