Near-Field Speech Detection Using Spatial Statistics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice activity detectors fail to reliably distinguish desired speech from interfering sounds, especially when the microphone array is positioned off-axis or in the presence of diffuse-field sounds, leading to ineffective speech enhancement and noise reduction.

Innovation Solution

A telephone system with at least two microphones and a circuit that processes audio signals using statistics such as maximum normalized cross-correlation, inter-microphone level difference, and direction of arrival to enhance near-field speech detection, combining these statistics to improve reliability and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional voice activity detectors are used, then the system is simple, but they fail to reliably distinguish desired speech from interfering sounds

Engineering Contradiction:
Improvespeech detection reliabilityVSAvoiddetector complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines multiple spatial statistics (inter-microphone level difference, normalized cross-correlation, direction of arrival, and diffuse-field gain) into a unified near-field detection system. This merging of multiple detection mechanisms enables reliable distinction between near-field speech and far-field interfering sounds, resolving the contradiction between detection reliability and system simplicity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system changes the detection parameters by using multiple spatial statistics instead of a single voice activity detection parameter. By computing and combining multiple spatial parameters (level difference, cross-correlation, direction of arrival), the system achieves reliable near-field speech detection in complex acoustic environments.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If the microphone array is positioned off-axis, then the device is more adaptable to user positioning, but speech enhancement effectiveness deteriorates

Engineering Contradiction:
Improvemicrophone positioning adaptabilityVSAvoidspeech enhancement reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system dynamically adapts to different microphone array orientations by computing direction of arrival and using it to control speech enhancement algorithms. This dynamic adjustment allows the system to maintain effective speech enhancement regardless of whether the microphone array is positioned on-axis or off-axis relative to the user's mouth.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes operational parameters based on detected sound field characteristics. By computing spatial statistics and using them to control the speech enhancement algorithm, the system adapts its processing parameters to maintain effectiveness across different microphone positioning scenarios.

Inventive Principle:
Principle #35Parameter changes

3Object-affected harmful factors

If speech enhancement algorithms are applied, then noise reduction is improved, but interference with desired speech increases

Engineering Contradiction:
Improvenoise reduction effectivenessVSAvoiddesired speech quality
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent applies local quality by using near-field detection to selectively control speech enhancement only when near-field speech is present. The spatial statistics enable the system to identify the local acoustic environment (near-field vs. far-field) and adjust enhancement intensity accordingly, reducing harmful interference with desired speech while maintaining noise reduction effectiveness.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system uses feedback from spatial statistics computation to control the speech enhancement algorithm. By continuously monitoring inter-microphone level difference, cross-correlation, and direction of arrival, the system provides feedback control that adjusts enhancement parameters to prevent over-enhancement and preserve desired speech quality.

Inventive Principle:
Principle #23Feedback

4Measurement precision

If multiple microphones are used, then spatial statistics accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvespatial statistics precisionVSAvoidmicrophone array complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the spatial statistics computation into distinct, manageable components: inter-microphone level difference calculation, normalized cross-correlation computation, direction of arrival estimation, and diffuse-field gain calculation. This segmentation of the measurement process enables accurate spatial statistics computation using multiple microphones while maintaining manageable system complexity through modular processing.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10015589B1Controlling speech enhancement algorithms using near-field spatial statistics
Publication Date: 2018.07.03 CIRRUS LOGIC INC
  • US10015589B1 patent drawing
  • US10015589B1 patent drawing
  • US10015589B1 patent drawing

AI summary

A telephone includes at least two microphones and a circuit for processing audio signals coupled to the microphones. The circuit processes the signals, in part, by providing at least one statistic representing maximum normalized cross-correlation of the signals from the microphones, doaEst, dirGain, or diffGain and comparing the at least one statistic with a threshold for that statistic. At least one of noise reduction and speech enhancement is controlled by an indication of near-field sounds in accordance with the comparison. Indication of near-field speech can be further enhanced by combining statistics, including a statistic representing inter-microphone level difference, each of which have their own threshold. dirGain and diffGain are derived from signals incident upon the microphones such that the desired near-field signal is not suppressed.