Acoustic Ranging via RTF Features and Gradient Boosting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing acoustic localization techniques for indoor environments, such as those used in smart home devices, face challenges in accurately estimating distance due to poor time of flight estimation and reliance on complex infrastructure or computationally expensive methods like computer vision, while also being affected by environmental conditions.

Innovation Solution

A novel fully acoustic ranging system based on acoustic features extracted from the acoustic relative transfer function (RTF) using an optimized distributed-gradient-boosting algorithm with regression trees, combined with the signal-to-reverberation ratio and sparseness coefficient, estimates distance between a fixed smart speaker and a wearable device, eliminating the need for room impulse response measurements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If time of flight measurement is used for distance estimation, then the estimation process is simple, but the accuracy is poor due to poor ToF estimation

Engineering Contradiction:
Improvedistance estimation accuracyVSAvoidacoustic feature extraction and machine learning model
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces acoustic features (direct-to-reverberant ratio, clarity index, reverberation time) as intermediary variables that mediate between the raw acoustic signals and distance estimation. These features serve as intermediate representations that capture environmental characteristics, enabling more accurate distance prediction than direct ToF measurement while avoiding the complexity of full acoustic environment modeling.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical/time-based measurement approach (ToF) with an acoustic field-based approach using machine learning models. Instead of measuring signal propagation time directly, the system substitutes this with analysis of acoustic feature relationships, achieving better accuracy by leveraging the statistical patterns in acoustic environments rather than relying on precise timing measurements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If computer vision is used for localization, then visual information can be obtained, but it requires huge image data sets, is affected by light conditions, raises privacy issues, and has high computational cost

Engineering Contradiction:
Improvelocalization accuracyVSAvoidcomputational cost
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent substitutes the optical/computational vision system with an acoustic-based system. Instead of using cameras and processing visual data, the system uses microphones and acoustic signal processing to achieve localization. This replacement eliminates the need for image datasets, removes light condition dependencies, addresses privacy concerns by using audio instead of video, and significantly reduces computational requirements while maintaining localization accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If acoustic features from room impulse response are used, then distance estimation can be achieved, but it requires knowledge of or blind estimation of the acoustic environment

Engineering Contradiction:
Improvedistance estimation accuracyVSAvoidacoustic environment modeling
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential acoustic features (direct-to-reverberant ratio, clarity index, reverberation time) needed for distance estimation from the acoustic signals, rather than attempting to model the entire acoustic environment. This extraction approach isolates the critical parameters that correlate with distance while ignoring unnecessary environmental details, achieving accurate distance estimation without full acoustic environment knowledge or complex blind estimation procedures.

Inventive Principle:
Principle #2Taking out (Extraction)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This method provides accurate distance estimation and proximity detection, enabling applications like automatic speech recognition, beamformer tuning, and scene analysis without requiring additional infrastructure or high computational costs, and is effective with both white noise and speech signals.

Implementation Method 1

The RTF represents a filter that maps the signal record from the fixed smart speaker to the signal recorded from the wearable device

Methodology Applied
Scientific EffectAcoustic relative transfer function:

Implementation Method 2

an improved proportionate normalized least mean square (IPNLMS) filter, which is an adaptive filter, is applied for the estimation of the acoustic RTF

Methodology Applied
Scientific EffectAdaptive filtering:

Implementation Method 3

When a sound signal from an acoustic source is captured by one or more microphone(s), the sound signal often gets corrupted by the acoustic environment, i.e., background noise and convolutive effects of room reverberation

Methodology Applied
Scientific EffectAcoustic propagation: Sound

Implementation Method 4

background noise and convolutive effects of room reverberation

Methodology Applied
Scientific EffectReverberation: Reverberation

Data Source

PatentUS11835625B2Acoustic-environment mismatch and proximity detection with a novel set of acoustic relative features and adaptive filtering
Publication Date: 2023.12.05 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11835625B2 patent drawing
  • US11835625B2 patent drawing
  • US11835625B2 patent drawing

AI summary

A method of performing distance estimation between a first recording device at a first location and a second recording device at a second location includes: estimating acoustic relative transfer function (RTF) between the first recording device and the second recording device for a sound signal, e.g., by applying an improved proportionate normalized least mean square (IPNLMS) filter; and estimating the distance between the first recording device and the second recording device based on the RTF. The at least one acoustic feature extracted from the RTF estimated between the first recording device and the second recording device includes at least one of clarity index, direct-to-reverberant ratio (DRR), and reverberation time. A distributed-gradient-boosting algorithm with regression trees is used in combination with signal-to-reverberation ratio (SRR) and the at least one acoustic feature extracted from the RTF to estimate the distance between the first recording device and the second recording device.