Multi-Channel Audio Encoding with Flexible Downmix Configuration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional multi-channel audio signal encoding methods fail to allow for variable downmix combinations per frame and require all information to be present in the header, leading to inefficiencies in encoding and decoding, especially in streaming services where header information may be missing.
Innovation Solution
The method involves encoding spatial information and generating additional configuration information based on selected header information, which is then inserted into the bitstream, enabling flexible retransmission and varying downmix combinations per frame.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If SAC configuration information is included only in the header of a bitstream, then the header structure is simple, but decoding cannot be performed when the header is not received (e.g., in streaming services)
Solution Approach 1:
The patent segments the SAC configuration information by separating it into two parts: essential parameters (sampling frequency, frame length) remain in the header, while flexible parameters (tree configuration information specifying downmix combination) are retransmitted in the frame data. This segmentation allows decoding to proceed with available information while maintaining the option to enhance quality with additional data.
Solution Approach 2:
The patent performs preliminary action by pre-identifying which configuration parameters should be retransmitted in the frame data. The encoder determines in advance that tree configuration information will be included in the frame, ensuring that decoding can proceed even if the header is missing, while still allowing for optimal downmix combinations.
2Adaptability or versatility
If tree configuration information is included only in SAC configuration information, then the configuration structure is simple, but the same downmix combination must be used throughout the entire multi-channel audio signal
Solution Approach 1:
The patent applies dynamics by making the downmix combination configurable on a per-frame basis rather than being fixed throughout the entire audio signal. The tree configuration information is retransmitted in the frame data, allowing the decoder to adapt the downmix combination dynamically to match the characteristics of each specific frame, thereby achieving optimal encoding/decoding efficiency for varying audio content.
3Loss of information
If all configuration information is placed in the header, then the header contains complete information, but the bitstream cannot be decoded in streaming services where the header may not be received
Solution Approach 1:
The patent extracts the tree configuration information from the header and places it in the frame data. This extraction ensures that critical configuration information needed for decoding is available in the frame itself, allowing streaming services to decode audio even when the header is not received, while still maintaining the option to use header information when available.
Data Source
AI summary
Methods and apparatuses for encoding and decoding a multi-channel audio signal are provided. In the encoding method, spatial information that is calculated based on a multi-channel audio signal and a downmix signal is encoded, and additional configuration information is generated based on information that is selected from the encoded spatial information. The downmix signal is encoded, and then, a bitstream is generated by combining the encoded downmix signal with the encoded spatial information. Thereafter, the additional configuration information is inserted into the bitstream. Therefore, it is possible to configure an optimum bitstream according to the circumstances by retransmitting all or part of information included in a header.


