Temporal Motion Vector Prediction Derivation in Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding systems face inefficiencies in transmitting motion vector information due to the bandwidth requirements of side information for spatial and temporal prediction, particularly in High-Efficiency Video Coding (HEVC), where explicit predictor signalling for Advanced Motion Vector Prediction (AMVP) can be cumbersome and inefficient.
Innovation Solution
The method involves selecting at least two collocated pictures from a set of reference pictures to derive temporal motion vector predictors (TMVPs) based on reference motion vectors, scaling these vectors, and combining them to form new predictors for improved motion vector prediction, which can be signaled in various video parameter sets or headers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If explicit predictor signalling is used for Advanced Motion Vector Prediction (AMVP), then motion vector prediction accuracy is improved, but bandwidth consumption and device complexity increase
Solution Approach 1:
The patent extracts only the necessary motion vector information from collocated blocks and uses it to derive temporal motion vector predictors. By selecting specific collocated pictures and blocks based on availability and reliability criteria, the system obtains sufficient prediction accuracy without transmitting all motion vector data explicitly, thus reducing complexity.
Solution Approach 2:
The derived temporal motion vector predictors serve multiple functions: they act as candidates for AMVP, provide fallback options when spatial predictors are unavailable, and can be used in merge mode. This multi-functionality reduces the need for separate explicit signalling mechanisms, thereby reducing device complexity while maintaining prediction accuracy.
2Measurement precision
If motion vectors are transmitted for temporal prediction, then prediction accuracy is improved, but bandwidth consumption increases
Solution Approach 1:
The system uses motion vector information from previously decoded collocated blocks to derive temporal motion vector predictors automatically. This self-service mechanism eliminates the need for explicit transmission of predictor information, as the predictors are generated from already-available decoded data, thereby reducing bandwidth consumption while maintaining prediction accuracy.
Solution Approach 2:
The patent performs preliminary derivation of temporal motion vector predictors from collocated blocks before they are needed for the current block prediction. By preparing these predictors in advance from reference pictures that are already decoded, the system avoids the need for additional transmission, reducing bandwidth consumption while ensuring prediction accuracy is maintained.
3Measurement precision
If multiple collocated pictures are used to derive temporal motion vector predictors, then prediction accuracy and flexibility are improved, but processing complexity increases
Solution Approach 1:
The patent implements a dynamic selection process where the system adaptsively chooses which collocated pictures and blocks to use based on their availability and reliability. The derivation process dynamically adjusts the number of TMVPs generated (one or two) based on the number of available collocated pictures, optimizing the balance between prediction accuracy and processing complexity for each specific coding scenario.
Solution Approach 2:
The system applies different quality criteria to different collocated blocks based on their local characteristics. Each collocated block is evaluated individually for motion vector availability and reliability, and only suitable blocks are used to derive TMVPs. This local quality assessment ensures high prediction accuracy while avoiding unnecessary processing of unsuitable blocks, thereby managing processing complexity effectively.
Data Source
AI summary
A method and apparatus for encoding or decoding a motion vector (MV) of a current block of a current picture using advanced temporal motion vector prediction are disclosed. At least two collocated pictures are selected from a set of reference pictures of the current picture. One or more TMVPs are derived based on reference motion vectors (MVs) associated with collocated reference blocks of the collocated pictures. A motion vector prediction candidate set including one or more TMVPs is then determined. The current block is encoded or decoding using the motion vector prediction candidate set. The reference motion vectors (MVs) are scaled before the reference motion vectors (MVs) are used to derive the TMVPs.


