Neural Network Intra Prediction Transform Adaptation for Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video encoding and decoding methods face limitations in choosing the appropriate transforms for image blocks predicted by neural network-based intra prediction modes, leading to inefficiencies in compression and increased computational burden.
Innovation Solution
Adapt the transform process by determining an intra prediction of an image block using a neural network, obtaining information on a transform method tailored to the neural network-based intra prediction mode, and applying inverse transforms to residue blocks for efficient decoding and encoding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a neural network-based intra prediction mode is used, then prediction accuracy is improved, but the choice of transform method becomes limited and less adaptable
Solution Approach 1:
The patent introduces dynamic adaptability by allowing the transform method to be selected or inferred based on the neural network intra prediction mode. The system transitions from a static transform selection to a dynamic one where the transform method adapts to the specific prediction mode used, resolving the contradiction between improved prediction accuracy and limited transform adaptability.
Solution Approach 2:
The patent changes the parameter of transform method selection from a fixed state to a variable state that depends on the neural network prediction mode. By making the transform method parameter changeable and adaptive to different prediction modes, the system maintains both high prediction accuracy and transform method versatility.
2Device complexity
If traditional transform methods are applied to neural network predicted blocks, then processing is simplified, but compression efficiency deteriorates
Solution Approach 1:
The patent applies preliminary action by pre-defining mappings between neural network intra prediction modes and transform methods. This allows the system to maintain simple processing (by using pre-established mappings) while achieving improved compression efficiency (by selecting appropriate transforms for each prediction mode). The adaptability is built in advance rather than requiring complex real-time decisions.
3Adaptability or versatility
If multiple transform methods are supported for neural network prediction, then adaptability is improved, but computational overhead increases
Solution Approach 1:
The patent uses preliminary action by pre-establishing mappings between prediction modes and transform methods, avoiding the need for complex real-time optimization. This reduces computational overhead while maintaining transform method versatility, as the optimal transforms are determined in advance during standard development rather than during actual video processing.
Solution Approach 2:
The system achieves self-service by using the neural network prediction mode itself to indicate or infer the appropriate transform method. The prediction mode 'services' the transform selection process, eliminating the need for separate complex transform selection logic and reducing overall computational overhead while maintaining adaptability.
Data Source
AI summary
At least a method and an apparatus are presented for efficiently encoding or decoding video. For example, an intra prediction of an image block using at least one neural network from a context comprising pixels surrounding the image block is determined and an information relative to a transform method to apply for decoding the image block is also determined. The transform method is adapted to the neural network intra prediction mode of the block to encode or decode. The information relative to the transform method is inferred from the at least one neural network used in intra prediction of the image block at the encoding and either signaled or also inferred at the decoding.


