Processor Onset Detection for Speaker Localization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing DOA estimation methods face challenges in accurately localizing a target speaker in noisy, reverberant environments with low Signal to Noise Ratios (SNRs), particularly when the target speaker is not the dominant sound source, leading to reduced DOA resolution and increased detection delays.

Innovation Solution

A processor that applies localization algorithms to determine source directions and onset detection to attribute scores, providing a direction-output-signal for a beamformer to focus on the target speaker, even if they are not dominant, using a combination of localization and onset detection algorithms to enhance wake word detection in challenging acoustic conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional DOA estimation methods are used to localize the dominant sound source, then the detection is robust in noisy environments, but the target speaker cannot be localized accurately when they are not the dominant source

Engineering Contradiction:
ImproveDOA localization accuracyVSAvoiddetection reliability in low SNR
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent segments the sound source detection process into two independent components: (1) onset detection that identifies sudden acoustic events regardless of dominance, and (2) DOA estimation that localizes identified onsets. This segmentation allows the system to detect non-dominant speakers by focusing on onset characteristics rather than overall sound energy, resolving the contradiction between accuracy for non-dominant sources and reliability in noisy environments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary onset detection mechanism that acts as a bridge between raw microphone signals and DOA estimation. The onset detector identifies acoustic events with sudden energy increases, and only these detected onsets are passed to the DOA estimation algorithm. This intermediary filtering enables accurate localization of non-dominant speakers by selecting relevant acoustic events while rejecting continuous background noise, thus improving measurement precision without sacrificing reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the system waits for sufficient speech statistics to improve DOA resolution, then the measurement precision improves, but the detection delay increases

Engineering Contradiction:
ImproveDOA resolutionVSAvoiddetection delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing onset detection before DOA estimation. The onset detector proactively identifies acoustic events with sudden energy increases in real-time, and immediately triggers DOA estimation for these detected onsets. This preliminary identification eliminates the need to accumulate extended speech statistics, achieving high DOA resolution with minimal detection delay by acting on identified onsets rather than waiting for statistical convergence.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If Voice Activity Detection (VAD) and dereverberation are applied to improve speaker localization, then the measurement precision improves, but the device complexity increases

Engineering Contradiction:
Improvespeaker localization accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes the VAD and dereverberation processing steps from the speaker localization pipeline. Instead of applying these complex preprocessing operations, the system directly uses onset detection based on sudden energy increases followed by DOA estimation. This extraction of unnecessary components maintains speaker localization accuracy by focusing on onset characteristics while significantly reducing device complexity and processing requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12185059B2Processor
Publication Date: 2024.12.31 NXP USA INC
  • US12185059B2 patent drawing
  • US12185059B2 patent drawing
  • US12185059B2 patent drawing

AI summary

A processor is configured to receive a plurality of sounds signals from a respective plurality of microphones. One or more localization algorithms is applied to the received plurality of sound signals to determine a plurality of source directions. A plurality of tracking directions is determined based on current values for the determined plurality of source directions and previous values for the determined plurality of source directions. An onset detection algorithm is applied to the received plurality of sound signals, and in response to detecting an onset, the processor is configured to determine an onset direction that represents the direction from the microphones to the source of the detected onset. A score is attributed to each of the determined plurality of tracking directions based on any determined onset directions and the one of the determined plurality of tracking directions that has the highest score is provided as a direction-output-signal.