Inter Prediction Decoding With Merge Fallback and MMVD Samples
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing demand for high-resolution, high-quality images and videos, particularly in virtual and augmented reality, requires more efficient compression techniques to reduce transmission and storage costs, as existing methods struggle with the increased data volume.
Innovation Solution
Implementing a method and apparatus that utilize a default merge mode when a merge mode is not selected, applying a regular merge mode, and deriving prediction samples based on motion vector difference (MMVD) and combined inter-picture merge and intra-picture prediction (CIIP) modes to enhance inter prediction efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high resolution and high quality image/video are transmitted or stored using existing methods, then image quality is maintained, but transmission and storage costs increase significantly
Solution Approach 1:
The patent extracts and transmits only the essential difference information (residuals) between original and predicted blocks rather than transmitting complete high-resolution image data. By separating the prediction component from the residual component, the system transmits minimal data while maintaining reconstruction quality.
Solution Approach 2:
The patent transforms image data from spatial domain to frequency domain using transform techniques, and applies quantization to change the precision parameters of transform coefficients. This parameter transformation reduces the amount of data needed to represent high-quality images.
2Quantity of substance
If conventional compression techniques are used for high resolution images, then data volume is reduced, but compression efficiency is insufficient for VR/AR and immersive media
Solution Approach 1:
The patent divides the image into multiple blocks and further partitions them into sub-blocks for independent processing. This segmentation allows different prediction and transformation techniques to be applied to different regions, improving overall compression efficiency for complex VR/AR content.
Solution Approach 2:
The patent extends conventional 2D block processing to 3D volume processing for immersive media, and introduces temporal dimension processing for video sequences. This multi-dimensional approach enables more efficient compression of VR/AR content with higher spatial and temporal resolutions.
3Adaptability or versatility
If a merge mode is not selected for a current block, then prediction flexibility is reduced, but derivation of prediction samples becomes uncertain
Solution Approach 1:
The patent prepares multiple candidate prediction blocks in advance (merge candidates from neighboring blocks, intra prediction candidates, inter prediction candidates) before the actual prediction is needed. When a merge mode is not selected, these pre-prepared candidates provide reliable fallback options for prediction sample derivation.
Solution Approach 2:
The patent creates a universal prediction framework that can handle both merge mode and non-merge mode cases through a unified candidate list structure. The same candidate list serves multiple purposes: providing merge candidates when merge mode is selected, and providing fallback candidates when merge mode is not selected, ensuring reliability in all scenarios.
Data Source
AI summary
According to the disclosure of the present document, when the inter prediction type of a current block indicates biprediction, weight index information for candidates in a merge candidate list or a sub-block merge candidate list can be derived, and thus coding efficiency can be increased.


