Inter Prediction Mode for Video Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video processing technologies face challenges in efficiently encoding and decoding next-generation video contents with high spatial resolution, high frame rate, and high dimensionality of scene representation, leading to increased memory storage, memory access rate, and processing power requirements.
Innovation Solution
A method is proposed for deriving a temporal motion vector from one reference picture, selecting a reference picture for this derivation using a signaled syntax, and applying Advanced Temporal Motion Vector Prediction (ATMVP) based on spatial candidates, which reduces memory bandwidth and addresses the additional line buffer problem.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional inter prediction modes are used for next-generation video contents, then processing capability is maintained, but memory bandwidth and processing power requirements increase drastically
Solution Approach 1:
The patent extracts and utilizes only the necessary motion information from reference pictures by deriving temporal motion vectors from a single reference picture rather than processing multiple reference pictures. This extraction approach reduces memory bandwidth requirements while maintaining prediction accuracy for high-resolution video content.
Solution Approach 2:
The patent changes the parameter of reference picture selection by using a signaled syntax to select one reference picture for temporal motion vector derivation. This parameter change optimizes memory access patterns and reduces the quantity of data that needs to be processed, thereby improving video processing efficiency without drastically increasing memory bandwidth requirements.
2Measurement precision
If multiple reference pictures are processed for motion vector prediction, then prediction accuracy is improved, but additional line buffer problems occur
Solution Approach 1:
The patent extracts motion information from a single selected reference picture rather than processing multiple reference pictures simultaneously. This extraction strategy maintains prediction accuracy by focusing computational resources on one high-quality reference while avoiding the line buffer complexity associated with managing multiple reference pictures.
Solution Approach 2:
Instead of processing multiple reference pictures and then selecting the best one, the patent inverts the approach by using signaled syntax to pre-select one reference picture for temporal motion vector derivation. This inversion simplifies the buffering requirements while maintaining prediction accuracy.
Data Source
AI summary
In the present disclosure, a method of decoding a video signal and a device therefor are disclosed. Specifically, a method of decoding an image based on an inter prediction mode includes deriving a motion vector of an available spatial neighboring block around a current block; deriving a collocated block of the current block based on the motion vector of the spatial neighboring block; deriving a motion vector in a sub-block unit in the current block based on a motion vector of the collocated block; and generating a prediction block of the current block using the motion vector derived in the sub-block unit, wherein the collocated block may be specified by the motion vector of the spatial neighboring block in one pre-defined reference picture.


