Motion Vector Precision Adjustment for Parallel Merge Candidate Derivation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video encoding/decoding methods face limitations in encoding efficiency due to dependency between temporal and bi-prediction merge candidate derivation processes, increased memory access bandwidth, and complex hardware logic in motion compensation, particularly in high-resolution and high-quality image data transmission and storage.
Innovation Solution
The method involves deriving spatial and temporal merge candidates from co-located blocks, using a reference picture list to select a reference picture for temporal merge candidates, and performing parallelization of merge candidate derivation processes to enhance encoding/decoding efficiency, including uni-directional, bi-directional, tri-directional, and quad-directional predictions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional merge mode is used with spatial, temporal, and bi-prediction merge candidates, then motion compensation is achieved, but encoding efficiency is limited due to dependency between derivation processes
Solution Approach 1:
The patent segments the merge candidate derivation by introducing a combined merge candidate that separates spatial and temporal components. The combined merge candidate is derived independently from co-located blocks in reference pictures, allowing parallel processing with spatial merge candidates. This segmentation eliminates the dependency between temporal and bi-prediction merge candidate derivation processes, enabling improved encoding efficiency through parallelization.
2Reliability
If bi-prediction merge candidate is used, then motion compensation is achieved, but memory access bandwidth increases
Solution Approach 1:
The patent extracts the temporal prediction component into a separate combined merge candidate derived from co-located blocks. By taking out the temporal component from the bi-prediction process, the patent reduces memory access requirements. The combined merge candidate uses reference picture information that is already available in the reference picture buffer, eliminating the need for additional bi-prediction merge candidate derivation that would require increased memory bandwidth.
3Adaptability or versatility
If zero merge candidate derivation is performed differently according to slice type, then bi-prediction zero merge candidate is generated, but hardware logic becomes complex
Solution Approach 1:
The patent creates a universal combined merge candidate derivation process that works across all slice types. The combined merge candidate is derived from co-located blocks using the same reference picture list mechanisms regardless of slice type. This universal approach replaces the complex slice-type-specific zero merge candidate derivation, simplifying hardware logic while maintaining adaptability through the reference picture list structure that already handles different slice configurations.
4Productivity
If multiple prediction directions are used, then encoding efficiency is enhanced, but merge candidate derivation complexity increases
Solution Approach 1:
The patent merges spatial and temporal prediction into a combined merge candidate that represents a unified motion information source. By combining these prediction types into a single candidate structure derived from co-located blocks, the patent reduces the overall complexity of managing multiple separate merge candidate lists. The combined merge candidate consolidates motion information from multiple directions while maintaining a streamlined derivation process that improves encoding efficiency without proportionally increasing complexity.
Data Source
AI summary
The present invention relates to a method for encoding/decoding a video. To this end, the method for decoding a video may include: deriving a spatial merge candidate from at least one of spatial candidate blocks of a current block, deriving a temporal merge candidate from a co-located block of the current block, and generating a prediction block of the current block based on at least one of the derived spatial merge candidate and the derived temporal merge candidate, wherein a reference picture for the temporal merge candidate is selected based on a reference picture list of a current picture including the current block and a reference picture list of a co-located picture including the co-located block.


