Sound Source Localization via Acoustic Wave Decomposition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Electronic devices face performance degradation due to acoustically reflective surfaces, which hinder accurate sound source localization and speech recognition by confusing sound reflections, leading to poor user experience.

Innovation Solution

Implementing sound source localization using acoustic wave decomposition (AWD) with Multi-Channel Linear Prediction Coding (MCLPC) coefficients to decompose sound fields into disjoint acoustic plane waves, allowing devices to identify the direction of speech by calculating noise statistics and signal quality metrics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional sound localization methods are used in environments with acoustically reflective surfaces, then the device can operate with simple processing, but the localization accuracy deteriorates due to confusion from sound reflections

Engineering Contradiction:
Improvesound source localization accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the acoustic signal into distinct components using acoustic wave decomposition (AWD), separating the direct sound path from reflected sound paths. This allows the system to process and analyze each wave component independently, improving localization accuracy by identifying the direct path signal while filtering out reflections from acoustically reflective surfaces.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces Multi-Channel Linear Prediction Coding (MCLPC) coefficients as an intermediary tool to model and predict the acoustic environment. These coefficients serve as a mediator between the raw microphone signals and the final localization decision, enabling the system to average out room impulse responses and distinguish direct sound from reflections through statistical analysis.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the device processes audio data to improve speech recognition in noisy environments, then speech recognition accuracy improves, but computational resources and processing time increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary acoustic wave decomposition and noise statistics calculation before the main speech recognition process. By pre-processing the audio data to separate direct sound from reflections and pre-calculating noise characteristics, the system reduces the computational burden on subsequent speech recognition stages, improving overall efficiency while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses the microphone array's own captured audio data to generate noise statistics and signal quality metrics without requiring external reference signals. This self-service approach allows the device to adapt to its specific acoustic environment autonomously, improving speech recognition in varying conditions without additional computational overhead from external calibration sources.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11425495B1Sound source localization using wave decomposition
Publication Date: 2022.08.23 AMAZON TECH INC
  • US11425495B1 patent drawing
  • US11425495B1 patent drawing
  • US11425495B1 patent drawing

AI summary

A system that performs sound source localization (SSL) using acoustic wave decomposition (AWD) or an approximation. When a device detects a wakeword represented in audio data, the device performs SSL processing in order to determine a position of the user relative to the device (e.g., estimate angle of the user). The device calculates noise statistics based on first audio data representing the wakeword and second audio data preceding the wakeword. Thus, upon detecting the wakeword, the device calculates the noise statistics and a signal quality metric corresponding to the wakeword. In addition, the device uses Multi-Channel Linear Prediction Coding (MCLPC) coefficients to average out the room impulse response. Using the noise statistics, the MCLPC coefficients, and the audio data, the device performs AWD processing to decompose the sound field to disjoint acoustic plane waves, enabling the device to identify the most likely direction for the line-of-sight component of speech.