Neural Network Post-Filter Input Tensor Identification in Video Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding and decoding schemes, such as H.264/AVC and H.265/HEVC, face challenges in determining the complexity of neural network models for post-filtering processes, particularly in analyzing the topology and complexity of neural networks without explicit information, and in defining input and output formats for neural network filters, which hinders efficient processing and image generation.
Innovation Solution
The implementation of header decoding and encoding circuitry to specify input tensor identification parameters for neural network post-filters, allowing for the determination of processing capabilities and image transformation without analyzing the neural network model's URI, and defining syntax tables for neural network filter supplemental enhancement information to clarify input and output formats.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If neural network model information is transmitted using URI or explicit topology definition, then the neural network filter can be specified, but the complexity and processing capability cannot be determined without analyzing the model
Solution Approach 1:
The patent extracts only the essential complexity information (number of layers, filters, and parameters) from the complete neural network model, transmitting this extracted metadata separately from the full model definition. This allows the decoder to determine processing capability without analyzing the entire model topology.
Solution Approach 2:
The patent performs preliminary analysis of the neural network model during encoding to extract complexity metrics before transmission. These pre-computed complexity parameters are included in the bitstream, allowing the decoder to immediately determine processing capability without performing model analysis during decoding.
2Manufacturing precision
If neural network topology is explicitly defined, then the filter structure can be specified, but the processing capability determination still requires model analysis
Solution Approach 1:
The patent introduces complexity metadata as an intermediary between the neural network model definition and the processing capability determination. This metadata layer provides a simplified representation that directly indicates processing requirements without requiring analysis of the full model structure.
3Device complexity
If color space and chroma sampling are not indicated in SEI, then the supplemental enhancement information remains simple, but the output format cannot be determined
Solution Approach 1:
The patent segments the SEI message structure into distinct fields for color space indication and chroma sampling specification. This segmentation allows the addition of necessary format information while maintaining a structured and organized message format that is easier to parse and process.
4Adaptability or versatility
If tensor channel and color component relationships are not defined, then the neural network processing can be flexible, but the input and output processing cannot be identified
Solution Approach 1:
The patent introduces explicit parameters in the SEI message that define the mapping between tensor channels and color components. These parameters allow the system to maintain processing flexibility while providing clear instructions for input preparation and output generation, resolving the ambiguity in channel-color relationships.
Data Source
AI summary
According to an aspect of the present disclosure, an image decoding apparatus for decoding information specifying a neural network includes: header decoding circuitry that decodes an input tensor identification parameter specifying a process for deriving an input tensor input for a post filter of the neural network. The input tensor identification parameter is a parameter related to a color component channel.


