Multi-Channel Audio Decoder Extrapolating Spatial Covariance Matrix
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current parametric multichannel decoders for mobile devices face challenges with unwanted channel cancellation or amplification due to limited bit-rate and processing power, and existing technologies fail to provide high-quality 3D audio rendering efficiently.
Innovation Solution
The method involves extrapolating a partially known covariance matrix to a complete spatial covariance matrix, allowing for the synthesis of arbitrary linear combinations of multi-channel surround audio signals, which enhances the quality and efficiency of 3D audio rendering by reducing computational complexity and eliminating the need for buffering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If complete 3D audio rendering with head related filtering is used, then 3D audio quality is improved, but computational complexity increases significantly
Solution Approach 1:
The audio signal is divided into multiple frequency sub-bands, and the head related filtering operation is performed independently in each sub-band. This segmentation allows the complex filtering operation to be broken down into simpler parallel operations, reducing overall computational complexity while maintaining 3D audio quality
Solution Approach 2:
The patent replaces the traditional time-domain convolution operation with an efficient frequency-domain implementation using Fast Fourier Transform (FFT). This substitution transforms the computationally intensive mechanical filtering process into more efficient mathematical operations in the frequency domain
2Measurement precision
If high bit-rate is used for multi-channel audio coding, then audio quality is improved, but bandwidth consumption increases
Solution Approach 1:
The patent extracts and transmits only the essential spatial parameters (inter-channel correlations and channel level differences) separately from the audio signal itself. This extraction allows the audio to be coded at lower bit-rates while still preserving the spatial information needed for high-quality 3D audio rendering
Solution Approach 2:
The patent transforms the audio signal representation from transmitting full multi-channel waveforms to transmitting compact spatial parameters (correlations and level differences) that describe the spatial relationships. This parameter change enables efficient bandwidth utilization while maintaining audio quality
3Adaptability or versatility
If arbitrary channel combinations are synthesized directly, then flexibility is improved, but channel cancellation or amplification artifacts occur
Solution Approach 1:
The patent performs preliminary computation of the covariance matrix based on the transmitted spatial parameters before synthesizing the arbitrary channel combinations. This preliminary action ensures that the statistical relationships between channels are properly accounted for, preventing cancellation or amplification artifacts in the final synthesized output
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
The basic concept of the present invention is to extrapolate a partially known spatial covariance matrix of a multi-channel signal in the parameter domain. The extrapolated covariance matrix is used with the downcoded downmix signal in order to efficiently generate an estimate of a linear combination of the multi-channel signals.