Parametric Audio Decorrelator Structure for Bandwidth Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio coding systems face challenges in efficiently encoding and reconstructing multiple audio signals due to bandwidth limitations and storage constraints, particularly in systems with multiple loudspeakers, where existing parametric coding methods require significant metadata transmission and storage.
Innovation Solution
The proposed method involves receiving a downmix signal with associated wet and dry upmix coefficients, computing intermediate and decorrelated signals through linear mappings, and combining these to reconstruct multiple audio signals, reducing the need for metadata transmission and storage by deriving coefficients based on upmix coefficients, thereby enhancing fidelity and reducing bandwidth requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If parametric coding methods are used to reduce bandwidth, then bandwidth requirements are reduced, but metadata transmission and storage requirements increase significantly
Solution Approach 1:
The patent extracts only the essential information needed for audio reconstruction by using a downmix signal combined with a small set of upmix coefficients. Instead of transmitting complete audio signals or extensive metadata, the system extracts the core spatial information and reconstructs the full audio scene from this compressed representation, significantly reducing data requirements while maintaining audio quality.
Solution Approach 2:
The patent transforms the audio representation from multiple independent channels to a lower-dimensional downmix signal plus coefficient parameters. By changing the dimensional space from transmitting full multi-channel audio to transmitting a compressed representation with mathematical parameters, the system achieves efficient bandwidth utilization while preserving the essential audio characteristics through the relationship between the downmix and upmix coefficients.
2Manufacturing precision
If multiple audio signals are transmitted to maintain audio quality, then reconstruction fidelity is improved, but bandwidth and storage requirements increase
Solution Approach 1:
The patent employs upmix coefficients that encode the spatial relationships and phase information needed to reconstruct the original audio signals from the downmix. These coefficients act as feedback information that enables the decoder to regenerate the full audio scene by applying the coefficient transformation to the downmix signal, achieving high fidelity reconstruction without transmitting the complete original signals.
Solution Approach 2:
The patent changes the representation parameters from raw multi-channel audio data to a compressed form consisting of a downmix signal and a set of upmix coefficients. This parameter transformation allows the system to maintain audio reconstruction fidelity by preserving the essential spatial and temporal characteristics through the mathematical relationship between the downmix and coefficients, rather than transmitting all original data.
3Manufacturing precision
If decorrelators are added to increase dimensionality for better reconstruction, then audio fidelity is improved, but device complexity increases
Solution Approach 1:
The patent implements a universal audio reconstruction system where the same downmix signal and upmix coefficients can be processed by standard audio equipment to reconstruct multiple audio signals. The decorrelation process is integrated into the existing audio processing pipeline, allowing the system to handle various audio configurations and playback scenarios without requiring specialized complex hardware for each specific case.
Solution Approach 2:
The patent performs the decorrelation and spatial separation calculations in advance during the encoding process, pre-computing the upmix coefficients that capture the essential spatial information. This preliminary action eliminates the need for complex real-time decorrelation operations at the decoding stage, reducing the computational burden and device complexity while maintaining high reconstruction fidelity.
Data Source
Figure 1~2
Figure 3~4
AI summary
An encoding system encodes multiple audio signals (X) as a downmix signal (Y) together with wet and dry upmix coefficients (P, C). In a decoding system, a pre-multiplier (101) computes an intermediate signal (W) by mapping the downmix signal linearly in accordance with a first set of coefficients (Q); a decorrelating section (102) outputs a decorrelated signal (Z) based on the intermediate signal; a wet upmix section (103) computes a wet upmix signal by mapping the decorrelated signal linearly in accordance with the wet upmix coefficients; a dry upmix section (104) computes a dry upmix signal by mapping the downmix signal linearly in accordance with the dry upmix coefficients; a combining section (105) provides a multidimensional reconstructed signal (X) by combining the wet and dry upmix signals; and a converter (106) computes the first set of coefficients based on the wet and dry upmix coefficients and supplies this to the pre-multiplier.