Single Matrix Intra Prediction for Video Coding Memory Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding schemes face inefficiencies in encoding and decoding processes due to the complexity of intra prediction methods, particularly in Versatile Video Coding (VVC) VTM 5.0, which results in high memory footprints and computational burdens.
Innovation Solution
The proposal introduces a simplified affine transformation for intra prediction, denoted as Single Matrix Intra Prediction (SMIP), which reduces the memory footprint by using a single affine transformation instead of the multiple transformations in VVC, and modifies the signaling method to encode this transformation efficiently in the bitstream.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple affine transformations are used for intra prediction in VVC, then prediction accuracy is improved, but memory footprint and computational complexity increase significantly
Solution Approach 1:
The patent segments the set of multiple affine transformations into a single representative transformation that captures the essential prediction characteristics. Instead of applying multiple full affine transformations, the invention uses a single transformation matrix that is selected or derived to represent the most probable prediction direction, thereby reducing memory storage requirements and computational operations while maintaining adequate prediction accuracy.
Solution Approach 2:
The patent changes the parameters of the affine transformation by reducing the number of transformation matrices from multiple to single. This parameter change involves selecting a single transformation that optimizes the trade-off between prediction accuracy and complexity, using techniques such as mode selection based on block characteristics or deriving a single transformation from multiple candidates through optimization criteria.
2Productivity
If complex intra prediction methods are employed, then compression efficiency is improved, but encoding and decoding time increase
Solution Approach 1:
The patent segments the complex intra prediction process into a simplified workflow that uses a single affine transformation. This segmentation reduces the number of processing steps required during encoding and decoding, thereby decreasing computational time while preserving the essential functionality needed for effective compression.
Solution Approach 2:
The patent extracts only the most critical components of the multiple affine transformations needed for effective prediction, discarding redundant transformations. By taking out only the essential prediction characteristics and representing them through a single transformation, the invention reduces processing time while maintaining compression efficiency.
3Adaptability or versatility
If multiple transformations are stored for different block sizes, then prediction flexibility is improved, but memory requirements increase by a factor of 7.92
Solution Approach 1:
The patent makes the single affine transformation universal by designing it to be applicable across different block sizes through scaling or adaptive selection mechanisms. Instead of storing separate transformation matrices for each block size, the invention uses a single transformation that can be adapted to various block dimensions, thereby providing prediction flexibility while dramatically reducing memory requirements.
Solution Approach 2:
The patent changes the parameters of the single affine transformation dynamically based on block size characteristics. Rather than storing multiple fixed transformations, the system adjusts parameters of the single transformation (such as scaling factors or selection criteria) to match different block sizes, maintaining adaptability while using minimal memory.
Data Source
AI summary
The disclosure relates to a method for encoding image data, the method including intra-predicting, or predicting by combining inter-prediction and intra-prediction, a first block of the image data by using an intra-prediction mode using a first single transformation obtained by taking account of the first block size. The disclosure also relates to a method for encoding image data, a variable coding length being used for signaling a plurality of Prediction Modes by the encoding, the method comprising:—intra-predicting a first block of the image data by using an intra-prediction mode using a first transformation, the first transformation being obtained by taking account of the first block size, —encoding information signaling a use of the intra-prediction mode in a bitstream, the information being encoded as one of the plurality of Prediction Modes. The disclosure further relates to the corresponding decoding methods, devices and media.


