Dual Beamforming for Multichannel Audio Voice Activity Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multichannel audio processing techniques in voice user interface platforms require refinement to improve noise suppression and voice activity detection, particularly in environments with varying background noise and speech patterns.
Innovation Solution
Implementing a dual beamforming approach with a slowly-adapting and quickly-adapting beamformer, where the slowly-adapting beamformer adapts to background noise and the quickly-adapting beamformer adapts to speech, using a short-time Fourier transform to process multichannel audio streams and selectively update covariance matrices based on voice activity detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single beamformer is used for both noise and speech adaptation, then the system is simpler to implement, but the accuracy of voice activity detection and noise suppression deteriorates
Solution Approach 1:
The patent divides the single beamformer into two separate beamformers: a slowly-adapting beamformer for noise estimation and a quickly-adapting beamformer for speech enhancement. This segmentation allows each beamformer to specialize in its respective function, improving voice activity detection accuracy while maintaining manageable system complexity through clear functional separation.
Solution Approach 2:
The patent implements dynamic adaptation rates for the two beamformers, with the slowly-adapting beamformer using a smaller step size for noise estimation and the quickly-adapting beamformer using a larger step size for speech tracking. This dynamic approach allows the system to adapt to changing acoustic environments effectively, resolving the contradiction between simplicity and accuracy.
2Speed
If the quickly-adapting beamformer continuously updates its covariance matrix, then the response to speech changes is faster, but the reliability of noise estimation deteriorates
Solution Approach 1:
The patent applies different adaptation speeds (step sizes) to the two beamformers: the quickly-adapting beamformer uses a larger step size for rapid speech tracking, while the slowly-adapting beamformer uses a smaller step size for reliable noise estimation. This dynamic parameter adjustment resolves the contradiction between fast response and reliable estimation.
Solution Approach 2:
The patent assigns different local properties to the two beamformers: the slowly-adapting beamformer is optimized for stable noise estimation with conservative updates, while the quickly-adapting beamformer is optimized for rapid speech tracking. This local quality differentiation allows each component to excel at its specific function without compromising the other.
3Object-affected harmful factors
If noise suppression is aggressively applied, then the noise reduction effect is stronger, but the distortion of speech signal increases
Solution Approach 1:
The patent uses the slowly-adapting beamformer as an intermediary to provide accurate noise estimates to the quickly-adapting beamformer. This intermediary noise estimation mechanism allows the quickly-adapting beamformer to suppress noise effectively while preserving speech signal quality, as the noise subtraction is based on reliable statistical estimates rather than aggressive filtering.
Solution Approach 2:
The patent implements feedback through voice activity detection that monitors the output of the quickly-adapting beamformer and uses this information to control the adaptation process. This feedback mechanism ensures that noise suppression is applied appropriately only when speech is present, preventing speech distortion while maintaining noise reduction effectiveness.
Data Source
AI summary
Aspects of the present disclosure provided a method for voice control that includes transforming, using a short-time Fourier transform (STFT) applied to data in each window aligned across each input channel of the multichannel audio stream, the multichannel audio stream into a complex valued frequency-domain representation. For a current window, the method further includes: updating a first complex-valued covariance matrix corresponding to a slowly-adapting beamformer and forming a single-channel denoised estimate for each frequency band in the STFT; calculating a voice activity detection (VAD) estimate for each frequency band in the STFT by comparing a magnitude of the single-channel denoised estimate to a magnitude of each input channel of the multichannel audio stream; and selectively updating or refraining from updating, responsive to the VAD estimate respectively indicating a presence or an absence of speech, a second complex-valued covariance matrix corresponding to a quickly-adapting beamformer.


