Acoustic Ranging via RTF Features and Gradient Boosting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing acoustic localization techniques for indoor environments, such as those used in smart home devices, face challenges in accurately estimating distance due to poor time of flight estimation and reliance on complex infrastructure or computationally expensive methods like computer vision, while also being affected by environmental conditions.
Innovation Solution
A novel fully acoustic ranging system based on acoustic features extracted from the acoustic relative transfer function (RTF) using an optimized distributed-gradient-boosting algorithm with regression trees, combined with the signal-to-reverberation ratio and sparseness coefficient, estimates distance between a fixed smart speaker and a wearable device, eliminating the need for room impulse response measurements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If time of flight measurement is used for distance estimation, then the estimation process is simple, but the accuracy is poor due to poor ToF estimation
Solution Approach 1:
The patent introduces acoustic features (direct-to-reverberant ratio, clarity index, reverberation time) as intermediary variables that mediate between the raw acoustic signals and distance estimation. These features serve as intermediate representations that capture environmental characteristics, enabling more accurate distance prediction than direct ToF measurement while avoiding the complexity of full acoustic environment modeling.
Solution Approach 2:
The patent replaces the mechanical/time-based measurement approach (ToF) with an acoustic field-based approach using machine learning models. Instead of measuring signal propagation time directly, the system substitutes this with analysis of acoustic feature relationships, achieving better accuracy by leveraging the statistical patterns in acoustic environments rather than relying on precise timing measurements.
2Measurement precision
If computer vision is used for localization, then visual information can be obtained, but it requires huge image data sets, is affected by light conditions, raises privacy issues, and has high computational cost
Solution Approach 1:
The patent substitutes the optical/computational vision system with an acoustic-based system. Instead of using cameras and processing visual data, the system uses microphones and acoustic signal processing to achieve localization. This replacement eliminates the need for image datasets, removes light condition dependencies, addresses privacy concerns by using audio instead of video, and significantly reduces computational requirements while maintaining localization accuracy.
3Measurement precision
If acoustic features from room impulse response are used, then distance estimation can be achieved, but it requires knowledge of or blind estimation of the acoustic environment
Solution Approach 1:
The patent extracts only the essential acoustic features (direct-to-reverberant ratio, clarity index, reverberation time) needed for distance estimation from the acoustic signals, rather than attempting to model the entire acoustic environment. This extraction approach isolates the critical parameters that correlate with distance while ignoring unnecessary environmental details, achieving accurate distance estimation without full acoustic environment knowledge or complex blind estimation procedures.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This method provides accurate distance estimation and proximity detection, enabling applications like automatic speech recognition, beamformer tuning, and scene analysis without requiring additional infrastructure or high computational costs, and is effective with both white noise and speech signals.
Implementation Method 1
The RTF represents a filter that maps the signal record from the fixed smart speaker to the signal recorded from the wearable device
Implementation Method 2
an improved proportionate normalized least mean square (IPNLMS) filter, which is an adaptive filter, is applied for the estimation of the acoustic RTF
Implementation Method 3
When a sound signal from an acoustic source is captured by one or more microphone(s), the sound signal often gets corrupted by the acoustic environment, i.e., background noise and convolutive effects of room reverberation
Implementation Method 4
background noise and convolutive effects of room reverberation
Data Source
AI summary
A method of performing distance estimation between a first recording device at a first location and a second recording device at a second location includes: estimating acoustic relative transfer function (RTF) between the first recording device and the second recording device for a sound signal, e.g., by applying an improved proportionate normalized least mean square (IPNLMS) filter; and estimating the distance between the first recording device and the second recording device based on the RTF. The at least one acoustic feature extracted from the RTF estimated between the first recording device and the second recording device includes at least one of clarity index, direct-to-reverberant ratio (DRR), and reverberation time. A distributed-gradient-boosting algorithm with regression trees is used in combination with signal-to-reverberation ratio (SRR) and the at least one acoustic feature extracted from the RTF to estimate the distance between the first recording device and the second recording device.


