MMVD Video Coding with Template Matching for Repetitive Patterns
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing MMVD (Merge with MVD Mode) technique in video coding is complex and lacks flexibility, which affects coding efficiency, particularly for video content with repetitive patterns.
Innovation Solution
The method introduces a flexible MMVD design that determines expanded merge motion vectors by adding offsets to a base motion vector, with the application to reference pictures in different lists determined by decoder-side matching costs and bi-prediction weights, allowing for separate MVDs and adaptive reordering of candidates based on template matching.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If the existing MMVD technique is used, then motion vector coding is achieved, but the complexity increases and flexibility is reduced
Solution Approach 1:
The patent segments the motion vector prediction process into distinct candidate lists (first merge candidate list and second merge candidate list), where each list contains specific types of motion vector candidates. This segmentation allows the encoder to selectively use different candidate lists based on video content characteristics, reducing overall complexity while maintaining flexibility.
Solution Approach 2:
The patent introduces dynamic selection mechanisms where the encoder can adaptively choose between different MMVD configurations based on template matching results and video content. The flexibility to switch between candidate lists and adjust motion vector offsets dynamically resolves the contradiction by making the system adaptable rather than fixed.
2Productivity
If template matching and adaptive reordering are implemented, then coding performance improves, but encoder complexity increases
Solution Approach 1:
The patent applies template matching and adaptive reordering selectively rather than universally. The encoder performs these operations only when beneficial for specific video content types (e.g., screen content with repetitive patterns), avoiding unnecessary computational overhead for other content types. This partial application maintains coding efficiency improvements while controlling complexity increases.
Solution Approach 2:
The patent changes key parameters such as motion vector offsets, candidate list compositions, and reference picture selections based on template matching results. By dynamically adjusting these parameters rather than using fixed values, the system achieves better coding efficiency while the parameter changes are guided by predefined rules that limit complexity growth.
3Measurement precision
If separate MVDs are used for different reference pictures, then coding accuracy improves, but the number of parameters increases
Solution Approach 1:
The patent segments motion vector prediction into separate candidate lists for different reference pictures (L0 and L1 lists). Each list contains tailored motion vector candidates appropriate for its reference picture type, improving accuracy by providing specialized candidates rather than generic ones. This segmentation organizes parameters systematically, making the increased quantity manageable.
Solution Approach 2:
The patent creates a universal framework where the same MMVD mechanism can handle multiple reference pictures with different characteristics. The merge candidate lists are designed to be multi-functional, serving both L0 and L1 reference pictures while allowing separate MVD adjustments, thus achieving accuracy improvement without proportionally increasing parameter complexity.
Data Source
AI summary
A method and apparatus for video coding using MMVD mode are disclosed. According to this method, a first expanded merge MV is determined for the current block, where the first expanded merge MV is derived by adding a first selected offset from a first set of offsets to a base MV, and whether the first expanded merge MV is applied to a first reference picture in L0 or a second reference picture in L1 is determined implicitly by the decoder side, or the first expanded merge MV is applied to the first reference picture in the L0 and a second expanded merge MV is applied to the second reference picture in the L1. The current block is encoded or decoded by using motion information comprising the first expanded merge MV. According to another method, separate MVDs are used for reference pictures in different reference lists.


