Multi-Microphone Linear Filtering for Wake Word Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional beamforming techniques for suppressing noise in audio recordings require a specific array configuration of microphones, which is not feasible in all network devices due to hardware or design constraints, leading to challenges in obtaining high-quality recordings of voice commands in noisy environments.
Innovation Solution
Implementing multi-microphone noise suppression techniques using linear time-invariant filtering to estimate noise from one audio signal and filter it out in another, enhancing voice detection accuracy without relying on geometrical microphone arrangements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If conventional beamforming techniques are used for noise suppression, then noise can be suppressed to some extent, but the technique requires specific microphone array configurations and has suboptimal performance in detecting voice inputs in noisy environments
Solution Approach 1:
The patent transforms the noise suppression approach from spatial-domain beamforming to frequency-domain Wiener filtering. By changing the parameter space from spatial coordinates to frequency components, the system can effectively suppress noise without requiring specific microphone geometries. The Wiener filter operates on frequency-transformed audio signals, allowing noise suppression to be applied universally across different microphone configurations.
Solution Approach 2:
The patent replaces the mechanical/spatial beamforming system with a signal processing-based Wiener filtering system. Instead of relying on the physical arrangement of microphones and spatial filtering, the invention uses frequency-domain signal processing to achieve noise suppression. This substitution allows the same noise suppression effect to be achieved regardless of the physical microphone configuration.
2Object-affected harmful factors
If conventional beamforming techniques are used for noise suppression, then noise can be suppressed, but the detection accuracy of voice inputs in noisy environments remains suboptimal
Solution Approach 1:
The patent applies Wiener filtering in the frequency domain to preserve speech components while suppressing noise. By operating in the frequency domain rather than the time domain, the filter can more effectively distinguish between noise and speech components. This parameter transformation enables more accurate voice input detection by maintaining the integrity of speech signals while removing noise interference.
Solution Approach 2:
The Wiener filter uses feedback from the estimated noise spectrum to continuously adjust the filtering process. The system estimates the noise components from the audio signal, uses this estimation to compute the optimal filter coefficients, and applies these coefficients to suppress the identified noise. This feedback mechanism ensures that speech components are preserved while noise is effectively removed, improving voice detection accuracy.
3Measurement precision
If linear time-invariant filtering with Wiener filtering is used, then voice detection accuracy is enhanced and noise is effectively suppressed, but the computational processing complexity increases
Solution Approach 1:
The patent applies Fast Fourier Transform (FFT) to convert the audio signal from the time domain to the frequency domain before applying Wiener filtering. This preliminary transformation allows the filtering operation to be performed more efficiently in the frequency domain. By pre-transforming the signal, the system can apply simple multiplicative filtering operations rather than complex convolution operations, reducing the overall computational complexity.
Solution Approach 2:
The patent replaces complex adaptive filtering algorithms with the simpler linear time-invariant Wiener filter. While Wiener filtering operates in the frequency domain requiring FFT operations, the filter coefficients are determined analytically from the signal statistics rather than requiring complex adaptive optimization. This substitution provides a good balance between performance and computational complexity, especially for real-time processing.
Data Source
AI summary
Systems and methods for suppressing noise and detecting voice input in a multi-channel audio signal captured by a plurality of microphones include (i) capturing a first audio signal via a first microphone and a second audio signal via a second microphone, wherein the first and second audio signals respectively comprises first and second noise content from a noise source; (ii) identifying the first noise content in the first audio signal; (iii) using the identified first noise content to determine an estimated noise content captured by the plurality of microphones; (iv) using the estimated noise content to suppress the first and second noise content in the first and second audio signals; (v) combining the suppressed first and second audio signals into a third audio signal; and (vi) determining that the third audio signal includes a voice input comprising a wake word.


