Sparse Audio Signal Processing for Spatial Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional inter-channel time difference estimation in spatial audio coding requires high computational load and significant bandwidth, especially at high audio sampling rates, making it inefficient for multi-channel spatial audio encoding and reproduction.

Innovation Solution

Transforming audio signals into a sparse domain and re-sampling them to reduce bandwidth requirements for accurate audio reproduction while retaining essential information for spatial audio encoding, allowing for efficient data transmission and processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If high audio sampling rates (48 kHz or higher) are employed for spatial audio coding, then audio quality is improved, but computational load increases significantly

Engineering Contradiction:
Improveaudio qualityVSAvoidcomputational load
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential spatial audio parameters (inter-channel time difference, level difference, and coherence) from the full audio signal, rather than processing the complete high-rate audio data. This selective extraction of critical information reduces computational complexity while preserving audio quality for spatial encoding purposes.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by processing only the necessary portions of the audio signal required for spatial parameter estimation. Instead of analyzing the entire high-resolution audio waveform, the method focuses computational resources on extracting spatial cues, thereby reducing overall computational load while maintaining audio quality where it matters most.

Inventive Principle:
Principle #16Partial or excessive action

2Measurement precision

If conventional inter-channel analysis mechanisms are used for spatial audio encoding, then spatial audio quality is maintained, but bandwidth requirements increase significantly

Engineering Contradiction:
Improvespatial audio qualityVSAvoidbandwidth
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential spatial parameters (inter-channel time difference, level difference, and coherence) from the multi-channel audio signal, transmitting these compact parameter representations instead of the full audio data. This extraction approach maintains spatial audio quality while dramatically reducing the bandwidth required for transmission between sensors and servers.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If full audio data is transmitted from distributed sensors to central server, then spatial audio encoding accuracy is improved, but transmission bandwidth increases

Engineering Contradiction:
Improvespatial audio encoding accuracyVSAvoidtransmission bandwidth
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the critical spatial audio parameters (inter-channel time difference, level difference, and coherence) at the sensor level before transmission. This local extraction and transmission of essential information maintains spatial encoding accuracy at the central server while minimizing the bandwidth consumed in data transmission across the network.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS9042560B2Sparse audio
Publication Date: 2015.05.26 NOKIA TECHNOLOGIES OY
  • US9042560B2 patent drawing
  • US9042560B2 patent drawing
  • US9042560B2 patent drawing

AI summary

A method comprising: sampling received audio at a first rate to produce a first audio signal; transforming the first audio signal into a sparse domain to produce a sparse audio signal; re-sampling of the sparse audio signal to produce a re-sampled sparse audio signal; and providing the re-sampled sparse audio signal, wherein bandwidth required for accurate audio reproduction is removed but bandwidth required for spatial audio encoding is retained AND/OR a method comprising: receiving a first sparse audio signal for a first channel; receiving a second sparse audio signal for a second channel; and processing the first sparse audio signal and the second sparse audio signal to produce one or more inter-channel spatial audio parameters.