Adaptive Filter Coherence Estimation for Speech Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In reverberating environments, such as vehicle cabins, the reliable estimation of signal coherence between microphone signals is challenging due to acoustic reflections, leading to poor speech detection performance with conventional methods that rely on coarse spectral resolution and temporal smoothing, resulting in delayed reaction times and misdetection of speech onsets and pauses.

Innovation Solution

The use of adaptive finite impulse response filters to compensate for the difference in acoustic transfer functions between microphones, combined with the Normalized Least Mean Square algorithm and frequency-dependent noise weighting, allows for improved estimation of signal coherence by minimizing error signal power density and temporally smoothing power density spectra based on signal-to-noise ratio.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If coarse spectral resolution (30-50 Hz per frequency band) is used for coherence estimation, then computational complexity is reduced, but speech detection reliability deteriorates due to phase discontinuities at nulls

Engineering Contradiction:
Improvecomputational complexityVSAvoidspeech detection reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The frequency spectrum is divided into multiple finer frequency bins instead of using coarse 30-50 Hz bands. This segmentation allows phase relationships to be evaluated at higher resolution, avoiding the phase discontinuities that occur at nulls in coarse bins, thereby improving speech detection reliability while maintaining manageable computational complexity through efficient algorithms.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from temporal smoothing (time domain) to frequency-domain analysis with fine spectral resolution. By examining phase relationships across multiple fine frequency bins and combining results, the system achieves more reliable speech detection without relying on temporal smoothing that suppresses fast changes.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If temporal smoothing with constant parameters is applied to coherence estimation, then reliability of speech detection is improved, but reaction time to speech onsets and offsets deteriorates

Engineering Contradiction:
Improvespeech detection reliabilityVSAvoidreaction time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces constant temporal smoothing parameters with dynamic, frequency-dependent smoothing parameters. Each frequency bin can have its own smoothing characteristic based on the local signal characteristics, allowing fast reaction to speech onsets and offsets in high-energy regions while maintaining reliability in low-energy regions through appropriate smoothing.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Different smoothing parameters are applied to different frequency regions based on local signal characteristics. This local quality approach allows the system to maintain high temporal resolution where speech energy is concentrated while applying stronger smoothing only where necessary for reliability, avoiding the uniform suppression of fast temporal changes.

Inventive Principle:
Principle #3Local quality

3Reliability

If fine spectral resolution is used for coherence estimation, then speech detection reliability is improved, but computational complexity increases

Engineering Contradiction:
Improvespeech detection reliabilityVSAvoidcomputational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary processing of the microphone signals including pre-filtering and pre-whitening before the fine spectral analysis. This preliminary action prepares the signals in advance, reducing the computational burden of the subsequent fine spectral resolution coherence estimation and making the overall system computationally feasible while maintaining high reliability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the problem from time-domain correlation to frequency-domain analysis using the Fast Fourier Transform. This parameter change in the analysis domain allows fine spectral resolution to be achieved efficiently, reducing the computational complexity compared to direct time-domain methods while maintaining or improving speech detection reliability.

Inventive Principle:
Principle #35Parameter changes

4Device complexity

If normalized signal correlation is used for coherence estimation, then computational simplicity is maintained, but accuracy deteriorates in reverberating environments due to ignoring signal spectra

Engineering Contradiction:
Improvecomputational simplicityVSAvoidcoherence estimation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent applies pre-whitening filtering as a preliminary step before coherence estimation. This preliminary action equalizes the power spectrum of the input signals, removing the effect of reverberation and background noise from the spectral shape. This allows subsequent correlation-based or spectral-based coherence estimation to be more accurate without significantly increasing computational complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces simple time-domain correlation with frequency-domain spectral analysis. By transforming the signals to the frequency domain and analyzing the spectral characteristics, the system achieves more accurate coherence estimation that accounts for signal spectra, while using efficient FFT-based methods to maintain computational feasibility.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS8238575B2Determination of the coherence of audio signals
Publication Date: 2012.08.07 CERENCE OPERATING CO
  • US8238575B2 patent drawing
  • US8238575B2 patent drawing
  • US8238575B2 patent drawing

AI summary

Embodiments of the invention disclose computer-implemented methods, systems, and computer program products for estimating signal coherence. First, a sound generated by a sound source is detected by a first microphone to obtain a first microphone signal and by a second microphone to obtain a second microphone signal. The first microphone signal is filtered by a first adaptive finite impulse response filter to obtain a first filtered signal. The second microphone signal is filtered by a second adaptive finite impulse response filter, to obtain a second filtered signal. The coherence of the first filtered signal and the second filtered signal is determined based upon the filtered signals. The first and the second microphone signals are filtered such that the difference between the acoustic transfer function for the transfer of the sound from the sound source to the first microphone and the transfer of the sound from the sound source to the second microphone is compensated in the first and second filtered signals.