ML-Based Sound Source Localization Confidence Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional sound source localization techniques face challenges in accurately estimating confidence metrics due to long settling times and robustness issues, especially in environments with multiple switching sources or when sound reflections have higher energy than direct paths, relying on a single feature which can lead to inaccurate direction estimation.

Innovation Solution

The implementation of machine learning-based confidence estimation systems that utilize multi-channel representations of sound, including features like sound classification and environmental characteristics, to provide updated and more accurate confidence metrics by aggregating multiple features and integrating new features efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Stability of the object's composition

If conventional SSL techniques use time-averaging filter to obtain robust estimate, then stability is improved, but settling time increases

Engineering Contradiction:
Improvestability of SSL angle estimateVSAvoidsettling time
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The patent changes the parameter used for confidence estimation from traditional single-feature metrics to multiple features including inter-microphone time differences, signal energies, and spectral characteristics. This allows the system to achieve stability without relying heavily on time-averaging, thus reducing settling time while maintaining robustness

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If SSL techniques rely on single feature (beam strength or correlation) for confidence metric, then device complexity is reduced, but measurement precision deteriorates

Engineering Contradiction:
Improvecomplexity of confidence estimationVSAvoidaccuracy of confidence metric
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent merges multiple features including inter-microphone time differences, signal energies, spectral characteristics, and spatial patterns into a unified confidence estimation framework. This combination approach improves measurement precision by considering multiple aspects of the audio signal simultaneously, while the features are efficiently processed to avoid excessive complexity

Inventive Principle:
Principle #5Merging (Combining)

3Ease of operation

If SSL techniques use traditional SB or TDOA methods, then ease of operation is maintained, but reliability deteriorates in challenging conditions

Engineering Contradiction:
Improveease of SSL implementationVSAvoidrobustness of SSL angle estimate
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent segments the confidence estimation process into multiple independent feature extractions (time differences, energies, spectral characteristics) that are then combined. This segmentation allows each feature to be computed using simple, well-established methods while the combination provides enhanced reliability in challenging acoustic conditions

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3639051B1Sound source localization confidence estimation using machine learning
Publication Date: 2023.07.26 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3639051B1 patent drawingFigure 1
  • EP3639051B1 patent drawingFigure 2
  • EP3639051B1 patent drawingFigure 3

AI summary

Techniques are described herein that are capable of performing sound source localization (SSL) confidence estimation using machine learning. An SSL operation is performed with regard to a sound to determine an SSL direction estimate and an SSL-based confidence associated with the SSL direction estimate based at least in part on a multi-channel representation of the sound. The SSL direction estimate indicates an estimated direction from which the sound is received. The SSL-based confidence indicates an estimated probability that the sound is received from the estimated direction. The multi-channel representation includes representations of the sound that are detected by respective sensors (e.g., microphones). Additional characteristic(s) of the sound are automatically determined. A machine learning (ML) operation is performed based at least in part on the SSL direction estimate, the SSL-based confidence, and the additional characteristic(s) to determine an ML-based confidence associated with the SSL direction estimate.