Spatial Audio Voice Activity Detection Using Inter-Channel Cues

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing spatial audio coding technologies face challenges in accurately modeling complex audio signals at low bitrates, particularly in distinguishing between foreground and background signals, which affects the intelligibility and quality of spatial audio playback.

Innovation Solution

The proposed method employs spatial cues such as inter-channel correlation (ICC), inter-channel time difference (ICTD), and inter-channel level difference (ICLD) to improve voice or sound activity detection (VAD/SAD) in spatial audio coding. These cues are used to classify signal types and adapt encoding strategies, enhancing the accuracy of signal representation and rendering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If spatial audio coding is used to reduce bandwidth, then transmission efficiency is improved, but signal accuracy deteriorates

Engineering Contradiction:
Improvetransmission efficiencyVSAvoidsignal accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The audio signal is segmented into foreground and background components using spatial cues. The foreground signal (direct sound) is encoded with higher precision while the background signal (reverberation) is encoded with lower precision, allowing differential allocation of bitrate resources to maintain overall signal accuracy while improving transmission efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different quality levels are applied to different parts of the audio signal based on their spatial characteristics. Direct sound paths receive higher encoding quality while reverberant fields receive lower quality encoding, optimizing the trade-off between bandwidth consumption and perceived audio fidelity.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If low bitrate encoding is used for background noise, then bandwidth consumption is reduced, but noise modeling accuracy deteriorates

Engineering Contradiction:
Improvebandwidth consumptionVSAvoidnoise modeling accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

Instead of encoding the complete background noise signal, only the essential spatial characteristics and parameters of the noise field are encoded at low bitrate. The receiver synthesizes the background noise using these parameters, achieving acceptable noise modeling accuracy with minimal bandwidth consumption.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The encoding approach changes from time-domain noise encoding to parameter-based spatial characterization. By encoding spatial parameters (such as spatial cues, correlation coefficients) rather than the noise signal itself, the system achieves efficient bandwidth usage while maintaining sufficient noise modeling accuracy for spatial audio reproduction.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3465681B1Method and apparatus for voice or sound activity detection for spatial audio
Publication Date: 2025.02.12 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • EP3465681B1 patent drawingFigure 1
  • EP3465681B1 patent drawingFigure 2~3
  • EP3465681B1 patent drawingFigure 4a~5a

AI summary

A method and apparatus for voice or sound activity detection for spatial audio. The method comprises receiving direct source detection decision and a primary voice/sound activity decision, and producing a spatial voice/sound activity decision based on the direct source detection decision and the primary voice/sound activity decision.