Acoustic Source Detection Using Phase Delay Variance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Far-field audio processing in smart home devices and similar consumer-level devices faces challenges due to noise interference, where noise sources can drown out speech and be misinterpreted as commands, leading to incorrect processing of voice inputs.

Innovation Solution

The use of multiple microphones to process audio signals, analyzing phase delay variance to identify noise sources, and employing a beamformer to filter out interference, allowing for improved speech recognition and differentiation between talkers and noise sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple microphones are used to capture far-field audio, then the ability to detect and filter noise sources is improved, but the device complexity increases

Engineering Contradiction:
Improvenoise filtering capabilityVSAvoidmicrophone array complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The audio signal processing is segmented into distinct stages: noise source detection using phase delay variance analysis, talker identification, and beamforming filtering. Each stage processes specific aspects of the audio signal independently, allowing complex noise filtering to be achieved through modular processing rather than a single complex system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Phase delay variance serves as an intermediary metric that mediates between the raw microphone signals and the final noise filtering decision. By analyzing the phase differences and their variance across multiple microphones, the system identifies noise sources without directly processing the complex spatial audio field, simplifying the detection mechanism.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If phase delay variance analysis is performed to identify noise sources, then speech recognition accuracy is improved, but processing time increases

Engineering Contradiction:
Improvenoise source identification accuracyVSAvoidaudio processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of phase delay variance to identify and classify noise sources before the main speech recognition processing begins. By pre-characterizing the acoustic environment and identifying persistent noise sources, the system can filter them out in subsequent processing stages, reducing the computational burden on speech recognition algorithms.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The beamforming process applies partial filtering based on the identified noise sources rather than attempting to process all audio components equally. By focusing computational resources only on filtering the identified noise sources and their spatial components, the system achieves effective noise reduction with reduced overall processing time compared to exhaustive spectral analysis.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If beamforming is used to filter noise sources, then speech intelligibility is improved, but the device complexity increases

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidsignal processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The beamforming filter applies different processing characteristics to different spatial directions and frequency components. By calculating the phase delay variance locally for each microphone pair and using this information to determine filtering strength and directionality, the system achieves adaptive noise filtering without requiring a monolithic complex processing system.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically changes processing parameters based on the identified noise sources and their spatial characteristics. The beamforming weights, filter coefficients, and processing intensity are adjusted as parameters rather than fixed values, allowing the system to adapt to varying acoustic environments while maintaining manageable computational complexity through parameter optimization rather than structural complexity.

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enhances speech recognition accuracy by reducing noise interference and correctly identifying talkers, thereby improving the reliability of voice command processing in noisy environments.

Implementation Method 1

analyzing phase delay variance to identify noise sources

Methodology Applied
Scientific EffectPhase delay:

Implementation Method 2

employing a beamformer to filter out interference

Methodology Applied
Scientific EffectBeamforming:

Data Source

PatentUS10142730B1Temporal and spatial detection of acoustic sources
Publication Date: 2018.11.27 CIRRUS LOGIC INC
  • US10142730B1 patent drawing
  • US10142730B1 patent drawing
  • US10142730B1 patent drawing

AI summary

Noise sources may be identified as either an interference source, such as a television, or a talker source by analyzing phase information of the microphone signals. A phase delay variance may be computed from pairs of microphone signals. A profile of an interference source may be learned over time by updating a stored profile when the phase delay variance is below a threshold. The stored profile may be used to identify interference sources received by the microphones by determining a correlation between the microphone signals and the stored profile. When an interference source is detected, control parameters may be generated to control a beamformer to reduce contribution of the interference source to an output audio signal. The output audio signal may be used for speech processing, such as in a smart home device.