Ambisonic Decoding Using Compressed Sensing and Auditory Masking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Ambisonic decoding technologies face challenges in accurately reconstructing complex sound fields with a large number of speakers, leading to spectral distortion and reduced sweet spots, especially when the number of speakers exceeds the number of ambisonic signals.

Innovation Solution

The implementation of compressed sensing techniques, such as L1-norm optimization and auditory masking, to separate and decode ambisonic signals into independent subcomponents, increasing sparsity and improving the accuracy of sound field reconstruction by reducing the number of required speakers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If the number of speakers is increased to improve sound field reconstruction quality, then the coverage and listening area are improved, but the system complexity and driving power requirements increase

Engineering Contradiction:
Improvesweet spot areaVSAvoidspeaker configuration complexity
Core Design Contradiction:
Area of stationary objectVSDevice complexity

Solution Approach 1:

The patent extracts and removes inaudible components from the ambisonic signal using auditory masking techniques. By identifying and eliminating frequency components that are below the masking threshold (inaudible sounds), the system reduces the complexity of the signal that needs to be reproduced by speakers, while maintaining the perception of the original sound field. This allows for simpler speaker configurations to achieve the same perceptual quality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies L1-norm optimization to change the parameter representation of the speaker signals from traditional L2-norm (least squares) to L1-norm (least absolute deviation). This parameter change promotes sparsity in the speaker signal distribution, meaning fewer speakers need to be actively driven at high power levels. The L1-norm optimization fundamentally changes how the signal energy is distributed across the speaker array, reducing overall system complexity and power requirements.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If traditional ambisonic decoding is used with many speakers, then complete sound field coverage is achieved, but spectral distortion increases and sweet spots become smaller

Engineering Contradiction:
Improvesound field reconstruction accuracyVSAvoidspectral distortion
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent converts the harmful effect of inaudible sound components into a benefit by using auditory masking to identify and remove these components. The inaudible frequencies, which would otherwise cause spectral distortion when reproduced by speakers, are eliminated from the signal. This transformation turns a potential source of distortion into an opportunity for cleaner signal reproduction with reduced spectral artifacts.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The patent segments the ambisonic signal into different frequency components and processes each segment independently through L1-norm optimization and auditory masking. By dividing the frequency spectrum into manageable segments and applying sparse decoding to each, the system maintains accurate sound field reconstruction for audible frequencies while eliminating the harmful spectral distortion that would result from reproducing all frequency components with many speakers.

Inventive Principle:
Principle #1Segmentation

3Loss of energy

If L1-norm optimization is applied to create sparse speaker signals, then the number of required speakers is reduced and driving power is decreased, but the processing complexity increases

Engineering Contradiction:
Improvedriving powerVSAvoidsignal processing complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent applies preliminary auditory masking to the ambisonic signal before the L1-norm optimization step. By pre-processing the signal to remove inaudible components that would otherwise require processing, the system reduces the complexity of the subsequent L1-norm optimization. This preliminary action eliminates unnecessary computational work on frequencies that cannot be heard, making the overall process more efficient despite the inherent complexity of L1-norm optimization.

Inventive Principle:
Principle #10Preliminary action

4Quantity of substance

If inaudible sounds are masked and removed, then the sparsity of speaker signals is increased and driving power is reduced, but the processing requirements increase

Engineering Contradiction:
Improvespeaker signal sparsityVSAvoidmasking threshold detection
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The patent employs auditory masking models that use feedback mechanisms to dynamically determine masking thresholds based on the actual signal content. The system continuously analyzes the spectral content of the ambisonic signal and adjusts the masking thresholds accordingly, using feedback from the signal itself to identify which frequencies are audible and which are masked. This feedback-based approach automates the detection process, reducing the difficulty of manually determining masking thresholds while accurately increasing speaker signal sparsity.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10020000B2Method and apparatus for improved ambisonic decoding
Publication Date: 2018.07.10 SAMSUNG ELECTRONICS CO LTD
  • US10020000B2 patent drawing
  • US10020000B2 patent drawing
  • US10020000B2 patent drawing

AI summary

An embodiment of this disclosure provides an audio receiver. The audio receiver includes a memory configured to store an audio signal and processing circuitry coupled to the memory. The processing circuitry is configured to receive the audio signal. The audio signal comprises a plurality of ambisonic components. The processing circuitry is also configured to separate the audio signal into a plurality of independent ambisonic subcomponents such that each of the independent ambisonic subcomponents is from a different source. The processing circuitry is also configured to decode each of the independent ambisonic subcomponents. The processing circuitry is also configured to combine each of the decoded independent ambisonic subcomponents into speaker signals.