Neural Network Intra Prediction Transform Selection for Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding methods face limitations in choosing the appropriate transform(s) to apply to the residue of an image block when predicted by a neural network-based intra prediction mode, leading to inefficiencies in encoding and decoding processes.
Innovation Solution
Adapt the transform process by determining an intra prediction of an image block using a neural network from its surrounding context, obtaining information on the transform method, and applying it to the residue, with the transform information being signaled or inferred by the neural network, allowing for flexible and efficient transform adaptation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a neural network-based intra prediction mode is used, then prediction accuracy is improved, but the choice of transform method becomes limited and less adaptable
Solution Approach 1:
The patent applies dynamics by making the transform method adaptive rather than fixed. The system dynamically selects transform methods based on the neural network prediction mode and block characteristics, allowing the transform process to adapt to different prediction scenarios and improve overall coding efficiency.
Solution Approach 2:
The patent changes the parameter of transform method selection from a static, limited choice to a dynamic, adaptable selection. By introducing multiple transform methods and selecting them based on neural network prediction modes and block characteristics, the system optimizes the transform process for different content types and prediction accuracies.
2Device complexity
If a fixed transform method is applied to neural network predicted blocks, then processing simplicity is maintained, but encoding and decoding efficiency deteriorates
Solution Approach 1:
The system transitions from a static, fixed transform approach to a dynamic, adaptive one. Multiple transform methods are available, and the selection is made based on neural network prediction modes and block characteristics, optimizing encoding efficiency without significantly increasing processing complexity.
Solution Approach 2:
The patent performs preliminary classification of blocks based on neural network prediction modes and characteristics before selecting the transform method. This preliminary action allows the system to prepare and select the most appropriate transform method in advance, improving encoding efficiency while maintaining manageable processing complexity.
3Productivity
If adaptive transform selection is implemented for neural network predictions, then compression performance is improved, but computational overhead increases
Solution Approach 1:
The patent optimizes the balance between compression performance and computational overhead by changing the parameter of transform method selection. Instead of exhaustive search, the system uses criteria based on neural network prediction modes and block characteristics to select from a limited set of transform methods, achieving good compression performance with controlled computational overhead.
Solution Approach 2:
The system performs preliminary analysis of block characteristics and neural network prediction modes to pre-select appropriate transform methods. This preliminary action reduces the computational overhead by avoiding exhaustive search while still achieving adaptive optimization for improved compression performance.
Data Source
AI summary
At least a method and an apparatus are presented for efficiently encoding or decoding video. For example, an intra prediction of an image block using at least one neural network from a context comprising pixels surrounding the image block is determined and an information relative to a transform method to apply for decoding the image block is also determined. The transform method is adapted to the neural network intra prediction mode of the block to encode or decode. The information relative to the transform method is inferred from the at least one neural network used in intra prediction of the image block at the encoding and either signaled or also inferred at the decoding.


