Parametric Audio Mixing for Multichannel Signal Reconstruction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio encoding and decoding systems face challenges in efficiently reducing bandwidth and storage requirements while maintaining high fidelity for multichannel audio signals, especially when played back on systems with fewer speakers than the original format, and in efficiently reconstructing the sound field.
Innovation Solution
The proposed solution involves encoding an M-channel audio signal into a two-channel downmix signal and associated metadata, where the downmix signal is formed as a linear combination of two groups of channels, and using mixing coefficients to generate a two-channel output signal that approximates different partitions of the M-channel audio signal, allowing for improved playback quality and reduced computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multichannel audio signals are transmitted or stored directly, then high fidelity is maintained, but bandwidth and storage requirements increase significantly
Solution Approach 1:
The patent extracts only the essential side information (metadata) describing the multichannel audio signal properties from the full signal. This metadata contains parameters like level differences and cross-correlation values that are sufficient for reconstruction, while the actual audio data is downmixed to fewer channels, significantly reducing bandwidth and storage while preserving perceptual quality
Solution Approach 2:
The patent transforms the multichannel audio signal into a compact representation by changing the parameter set used to describe the audio. Instead of storing all channel data, it stores transformed parameters (side information) that capture the essential spatial and temporal characteristics, enabling efficient compression and reconstruction
2Quantity of substance
If multichannel audio signals are downmixed to fewer channels, then bandwidth and storage requirements are reduced, but reconstruction fidelity deteriorates
Solution Approach 1:
The patent uses side information (metadata) extracted from the original multichannel signal as feedback to guide the reconstruction process. This metadata contains critical spatial parameters that are used by the decoder to synthesize the missing channels, ensuring that the reconstructed signal maintains high fidelity despite the reduced channel count
Solution Approach 2:
The patent performs preliminary extraction and encoding of side information (metadata) during the encoding phase. This preprocessing step captures essential spatial characteristics before downmixing, enabling the decoder to accurately reconstruct the multichannel signal from the downmixed audio and metadata without losing fidelity
3Adaptability or versatility
If multichannel audio signals are encoded for playback on systems with fewer speakers, then compatibility is improved, but sound field reconstruction accuracy deteriorates
Solution Approach 1:
The patent creates a universal encoding format that works across different playback configurations. The side information (metadata) contains spatial parameters that enable the same encoded signal to be accurately reconstructed on systems with varying numbers of speakers, making the solution universally compatible while maintaining sound field accuracy through parameter-based reconstruction
Data Source
Figure 1~2
Figure 3~5
Figure 6~8
AI summary
In an encoding section (100), a downmix section (110) forms first and second channels (L 1 , L 2 ) of a downmix signal as linear combinations of first and second groups (401, 402) of channels, respectively, of an M-channel audio signal; and an analysis section (120) determines upmix parameters (α LU ) for parametric reconstruction of the audio signal, and mixing parameters (α LM ). In a decoding section (1200), a decorrelating section (1210) outputs a decorrelated signal (D) based on the downmix signal; and a mixing section (1220) determines mixing coefficients based on the mixing parameters or the upmix parameters, and forms a K-channel output signal (L͂ 1 ,...,L͂ K ) as a linear combination of the downmix signal and the decorrelated signal in accordance with the mixing coefficients. The channels of the output signal approximate linear combinations of K groups (501-502, 1301-1303) of channels, respectively, of the audio signal. The K groups constitute a different partition of the audio signal than the first and second groups, and 2 ≤ K < M.