Matrix Based Intra Prediction Multiple Transform Blocks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The current version of Versatile Video Coding (VVC) restricts matrix-based intra prediction (MIP) to coding blocks with dimensions equal or less than the maximum transform size, limiting coding efficiency when the maximum coding block size exceeds the maximum transform size.
Innovation Solution
A method and apparatus that enable MIP prediction for coding blocks with dimensions greater than the maximum transform size by allowing multiple transform blocks and using a MIP weight matrix to decode these blocks, even when the coding block height or width exceeds the maximum transform size.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If MIP prediction is restricted to coding blocks with dimensions equal or less than the maximum transform size, then the decoding process remains simple, but coding efficiency is limited when maximum coding block size exceeds maximum transform size
Solution Approach 1:
The patent divides the coding block into multiple transform blocks, each with dimensions that do not exceed the maximum transform size. This segmentation allows MIP prediction to be applied to larger coding blocks while maintaining compatibility with the transform size constraints, thereby improving coding efficiency without overwhelming complexity in the decoding process.
Solution Approach 2:
The patent introduces a new dimension of organization by allowing a coding block to be composed of multiple transform blocks arranged in a specific configuration. This dimensional restructuring enables the system to handle larger blocks through MIP prediction by breaking them down into manageable transform block units, resolving the contradiction between coding efficiency and decoding simplicity.
2Productivity
If multiple transform blocks are allowed for MIP prediction, then coding efficiency improves for large coding blocks, but the structure becomes more complex
Solution Approach 1:
The coding block is segmented into multiple transform blocks that can be independently processed. This segmentation strategy improves coding efficiency for large blocks while managing structural complexity through a systematic division into standardized units, each following the same transformation rules.
Solution Approach 2:
The patent creates a universal structure where multiple transform blocks can be used for MIP prediction regardless of the specific coding block size. This multi-functional approach allows the same transform block mechanism to handle various block dimensions, improving efficiency without proportionally increasing complexity.
3Adaptability or versatility
If maximum coding block size is increased beyond maximum transform size, then more flexibility is achieved, but MIP prediction cannot be applied without modification
Solution Approach 1:
The patent applies segmentation to divide large coding blocks into multiple transform blocks, each compliant with the maximum transform size constraint. This allows MIP prediction to be implemented for larger blocks by breaking them down into manageable units, maintaining ease of implementation while achieving greater flexibility.
Solution Approach 2:
The patent introduces dynamic adaptability by allowing the number and arrangement of transform blocks to vary based on the coding block dimensions. This dynamic structure enables MIP prediction to work with various block sizes while maintaining the constraint that individual transform blocks do not exceed the maximum size, balancing flexibility with implementation simplicity.
Data Source
AI summary
A method, decoder, and apparatus are provided. Responsive to a current block being a MIP predicted block, it is determined whether it has one or multiple transform blocks. A MIP weight matrix to be used to decode the current block is determined based on a MIP prediction mode. Responsive to the MIP predicted block having one transform block, the MIP predicted block is derived based on the MIP weight matrix and previously decoded elements in the bitstream. Responsive to the MIP predicted block having multiple transform blocks: deriving a first MIP predicted block is derived based on the MIP weight matrix and previously decoded elements in the bitstream and remaining MIP predicted blocks are derived based further on decoded elements in at least one decoded transform block of the current block. The MIP predicted block(s) are output for subsequent processing.


