Ambisonic Decoding with Compressed Sensing for Wider Sweet Spots
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Ambisonic decoding technologies face challenges in accurately reconstructing complex sound fields with a large number of speakers, leading to spectral distortion and reduced sweet spots due to the under-determined nature of the linear equation used to derive speaker signals from ambisonic signals.
Innovation Solution
The implementation of compressed sensing techniques, combined with auditory masking and independent component analysis, to upscale lower-order ambisonics and increase sparsity, allowing for more accurate and efficient decoding of sound fields by optimizing speaker signals using L1-norm minimization and perceptual sparsity methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If compressed sensing techniques with L1-norm minimization are applied to upscale lower-order ambisonics, then the sparsity of sound fields increases and decoding accuracy improves, but the computational complexity increases
Solution Approach 1:
The patent changes the norm parameter from L2-norm to L1-norm in the minimization process. This parameter change enables sparsity promotion in the decoded speaker signals, allowing accurate reconstruction of sound fields with fewer speakers while maintaining computational tractability through convex optimization
Solution Approach 2:
The patent segments the ambisonic decoding problem into two stages: first upscaling lower-order ambisonics to higher-order using compressed sensing, then decoding with L1-norm minimization. This segmentation allows each stage to be optimized independently, managing computational complexity while improving accuracy
2Reliability
If the number of speakers is increased to improve sound field reconstruction quality, then the sweet spot expands and spatial resolution improves, but the system complexity and power consumption increase
Solution Approach 1:
The patent extracts and removes inaudible spectral components from the ambisonic signals using auditory masking models. By eliminating these redundant components before decoding, the system achieves accurate sound field reconstruction with fewer speakers, reducing system complexity while maintaining quality
Solution Approach 2:
The patent changes the approach from using more speakers to using sparsity-promoting L1-norm minimization. This parameter change in the optimization objective allows the system to achieve the same reconstruction quality with fewer active speakers by concentrating energy on the most important spatial components
3Productivity
If auditory masking is applied to remove inaudible sounds, then the sparsity of sound fields increases and decoding efficiency improves, but the processing complexity increases
Solution Approach 1:
The patent applies auditory masking and removes inaudible spectral components as a preliminary step before the main decoding process. This preliminary action reduces the dimensionality and complexity of the subsequent L1-norm minimization problem, improving overall decoding efficiency while the masking model itself uses standard psychoacoustic algorithms
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An embodiment of this disclosure provides an audio receiver. The audio receiver includes a memory configured to store an audio signal and processing circuitry coupled to the memory. The processing circuitry is configured to receive the audio signal. The audio signal comprises a plurality of ambisonic components. The processing circuitry is also configured to separate the audio signal into a plurality of independent subcomponents. Each of the independent subcomponents is from a different source. Each of the plurality of ambisonic components is split into the independent subcomponents. The processing circuitry is also configured to decode each of the independent subcomponents. The processing circuitry is also configured to combine each of the decoded independent subcomponents into a speaker signal.