Neural Video Bitstream Signaling for Decoder Compatibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video encoders and decoders, particularly those using artificial intelligence approaches like auto-encoders, face compatibility issues due to differing constraints and syntax elements, making it difficult for standard decoders to interpret data encoded by neural networks.
Innovation Solution
An audio and/or video encoder and decoder system that uses artificial intelligence to transmit decoding features, allowing standard decoders to identify and support necessary configurations for decoding neural network-encoded data, using structured signaling to reduce information overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If encoding neural networks use completely different constraints and syntax elements from conventional video encoders, then the encoding efficiency and adaptability of neural networks are improved, but the compatibility with standard decoders deteriorates
Solution Approach 1:
The patent introduces syntax elements as an intermediary mechanism that translates between the neural network encoder's internal representations and the decoder's expected format. These syntax elements serve as a communication bridge, allowing the encoder to convey necessary decoding information (such as layer configurations, activation functions, and processing parameters) without requiring the decoder to understand neural network internals, thus resolving the compatibility issue while preserving encoding adaptability
Solution Approach 2:
The patent segments the encoding information into distinct syntax element categories (sequence-level, stream-level, and packet-level syntax elements). This segmentation allows different levels of decoding requirements to be addressed independently, enabling decoders to selectively process only the syntax elements they need while maintaining compatibility with neural network encoded content
2Measurement precision
If VVC syntax elements are defined over limited ranges of values, then the precision and control of conventional encoding are improved, but the ability to represent broader encoding indications required by neural networks deteriorates
Solution Approach 1:
The patent extends the value representation by introducing additional dimensions to the syntax element structure. Instead of relying solely on limited scalar values, the patent uses structured syntax elements that can represent multi-dimensional information (such as arrays of values, hierarchical structures, and composite parameters), thereby expanding the representational capacity while maintaining the precision control inherent in conventional syntax element definitions
3Reliability
If an encoder transmits detailed decoding configuration information, then the decoder compatibility is improved, but the data overhead and signal complexity increase
Solution Approach 1:
The patent implements partial signaling by transmitting only the essential syntax elements required for decoding compatibility, rather than all possible configuration details. The syntax elements are designed to convey the minimum necessary information (such as layer count, input/output dimensions, and critical processing parameters) while omitting redundant or implementation-specific details, thus achieving decoder compatibility with minimal data overhead
Data Source
AI summary
A method for encoding audio and/or video data performed by an encoding device configured to perform at least one step of encoding audio and/or video data using an encoding artificial neural network. The encoding method includes: encoding the data, generating a data signal containing the encoded data, encoding information representing a decoding configuration to be had by a decoding device in order to decode the encoded data, inserting the encoded information into the signal.


