Small Array Microphone Noise Suppression Using Dual Voice Activity Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional array microphone systems for speech recognition are limited by their inability to provide voice activity detection for in-beam and out-of-beam signals, require a minimum microphone spacing, lack noise suppression control, and are ineffective for diffuse noise, making them inadequate for adverse environments.

Innovation Solution

A small array microphone system with multiple microphones that includes a first voice activity detector for in-beam speech, a second for out-of-beam noise, a reference signal generator, a beamformer, and a multi-channel noise suppressor, which uses time-domain and frequency-domain processing to enhance speech recognition by suppressing noise and providing reliability detection signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional array microphone systems use multiple microphones with minimum spacing D, then noise suppression capability is improved, but device size and microphone spacing requirements increase

Engineering Contradiction:
Improvenoise suppression capabilityVSAvoidmicrophone spacing
Core Design Contradiction:
ReliabilityVSLength of moving object

Solution Approach 1:

The patent changes the operational parameters of the microphone system by using digital signal processing techniques to achieve noise suppression without requiring physical spacing between microphones. The system processes signals from closely-spaced microphones using adaptive filtering and beamforming algorithms, transforming the problem from a spatial solution to a signal processing solution.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical spacing requirement with digital signal processing mechanisms. Instead of relying on physical distance between microphones to achieve noise suppression, the system uses digital beamforming, adaptive noise cancellation, and spectral processing to achieve the same effect with much smaller physical dimensions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If conventional array microphone systems process multi-channel signals, then speech recognition accuracy is improved, but processing complexity and computational requirements increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the noise suppression process into distinct stages: voice activity detection, spectral estimation, noise modeling, and signal reconstruction. By dividing the complex processing into modular stages, the system achieves high speech recognition accuracy while managing computational complexity through organized signal processing pipelines.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary voice activity detection and spectral estimation before full noise suppression processing. By pre-processing the signals to identify speech regions and estimate noise characteristics in advance, the system reduces the computational burden of subsequent noise cancellation operations while maintaining high recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If conventional systems use fixed beamforming, then implementation simplicity is improved, but adaptability to different noise environments and source positions deteriorates

Engineering Contradiction:
Improveimplementation simplicityVSAvoidadaptability to noise environments
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent implements adaptive beamforming where the beam pattern dynamically adjusts based on the estimated positions of speech sources and noise. The system continuously updates beamformer weights using signal processing algorithms that track source movements and environmental changes, providing both simplicity through automated adaptation and versatility through real-time reconfiguration.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent incorporates feedback mechanisms where the system continuously monitors the input signals, estimates noise and speech characteristics, and adjusts beamforming parameters accordingly. This closed-loop approach maintains implementation simplicity through automated control while achieving high adaptability to different noise environments and source configurations.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8068619B2Method and apparatus for noise suppression in a small array microphone system
Publication Date: 2011.11.29 FORTEMEDIA INC
  • US8068619B2 patent drawing
  • US8068619B2 patent drawing
  • US8068619B2 patent drawing

AI summary

A small array microphone system includes an array microphone having a plurality of microphones and operative to provide a plurality of received signals, each microphone providing one received signal. A first voice activity detector (VAD) provides a first voice detection signal generated using the plurality of received signals to indicate the presence or absence of in-beam desired speech. A second VAD provides a second voice detection signal generated using the plurality of received signals to indicate the presence or absence of out-of-beam noise when in-beam desired speech is absent. A reference signal generator provides a reference signal based on the first voice detection signal, the plurality of received signals, and a beamformed signal, wherein the reference signal has the desired speech suppressed. A beamformer provides the beamformed signal based on the second voice detection signal, the reference signal, and the plurality of received signals, wherein the beamformed signal has noise suppressed. A multi-channel noise suppressor operative to further suppress noise in the beamformed signal and provide an output signal. A speech reliability detector provides a reliability detection signal indicating the reliability of each frequency subband. The first voice detection signal, the second voice detection signal, the reliability detection signal and the output signal are provided to the speech recognition engine.