Audio Encoder Dynamic Multichannel Coding Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codecs have fixed multichannel coding techniques that are not adaptable to signal characteristics, limiting their efficiency in encoding diverse audio content such as speech and music, as they cannot dynamically switch between different coding methods based on the core codec configuration.
Innovation Solution
An audio encoder and decoder system that incorporates a switchable core codec with joint multichannel coding and parametric spatial audio coding, allowing for dynamic selection of multichannel coding techniques based on the core coder choice, using a combination of linear prediction domain and frequency domain encoding to optimize encoding for different signal types.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If fixed multichannel coding techniques are used, then device complexity is reduced, but adaptability to different signal characteristics deteriorates
Solution Approach 1:
The patent implements dynamic switching between different multichannel coding techniques (joint multichannel coding and parametric spatial audio coding) based on signal characteristics. The system can adaptively select between M/S-Stereo, spatial audio coding, and parametric stereo modes depending on the core codec configuration and signal type, transforming a static system into a dynamic one that optimizes performance for different audio content.
Solution Approach 2:
The audio encoder is designed to perform multiple multichannel coding functions within a single device. It can implement joint multichannel coding for speech signals, spatial audio coding for music content, and parametric stereo for various signal types. This multi-functionality allows one device to handle diverse audio formats and characteristics without requiring separate dedicated systems for each coding technique.
2Productivity
If signal adaptive core coder is used, then encoding efficiency is improved, but compatibility with fixed multichannel coding techniques deteriorates
Solution Approach 1:
The system dynamically adjusts the multichannel coding technique based on the core codec configuration. When ACELP or TCX is used for speech signals, the system selects appropriate joint multichannel or spatial audio coding modes. When AAC is used for music content, it switches to frequency-domain based multichannel coding. This dynamic adaptation ensures both high encoding efficiency and full compatibility across different core coder configurations.
Solution Approach 2:
The audio encoder incorporates feedback mechanisms that monitor signal characteristics and core codec performance in real-time. Based on this feedback, the system automatically selects the most suitable multichannel coding technique, ensuring optimal encoding efficiency while maintaining compatibility with the chosen core coder. The feedback loop enables continuous optimization of the coding process.
3Productivity
If parametric coding techniques are used for bandwidth extension, then encoding efficiency is improved, but compatibility with time domain and frequency domain processing deteriorates
Solution Approach 1:
The patent implements a universal bandwidth extension framework that can operate in both time domain and frequency domain. The system includes time-domain based parametric coding for speech signals and frequency-domain based coding for music content. This dual-domain capability allows the same bandwidth extension mechanism to adapt to different signal types and processing requirements, eliminating compatibility issues between different processing domains.
Solution Approach 2:
The system changes the fundamental parameters of the coding process based on the signal type and processing domain. For time-domain processing, it uses linear prediction and temporal envelope parameters. For frequency-domain processing, it employs spectral characteristics and frequency-band based parameters. This parameter adaptation enables efficient parametric coding while maintaining full compatibility with both time and frequency domain processing techniques.
Data Source
AI summary
Audio encoder for encoding a multichannel signal is shown. The audio encoder includes a downmixer for downmixing the multichannel signal to obtain a downmix signal, a linear prediction domain core encoder for encoding the downmix signal, wherein the downmix signal has a low band and a high band, wherein the linear prediction domain core encoder is configured to apply a bandwidth extension processing for parametrically encoding the high band, a filterbank for generating a spectral representation of the multichannel signal, and a joint multichannel encoder configured to process the spectral representation including the low band and the high band of the multichannel signal to generate multichannel information.


