Microphone Array Speech Enhancement With Per-Band Interference Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio conferencing technologies struggle to accurately enhance speech of interest while suppressing interfering noise and competing talkers, often leading to over-suppression or under-suppression of audio signals.
Innovation Solution
The method involves determining and applying per-band gains based on the angle of arrival and covariance of microphone signals, clustering audio objects into regions of interest and non-interest, and using beam forming techniques to enhance speech within the region of interest and suppress noise outside it.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional audio conferencing technologies use multiple microphones to pick up audio from multiple speakers, then the system can capture more speech signals, but it becomes difficult to isolate and emphasize a speaker of interest relative to noise from other regions and microphones
Solution Approach 1:
The audio spectrum is divided into multiple frequency bands, and the spatial domain is segmented into regions of interest and non-interest. Per-band gains are applied to different spatial regions, allowing precise control over speech and noise from different directions without affecting the entire audio spectrum uniformly.
Solution Approach 2:
Different processing strategies are applied to different spatial regions. Regions of interest receive enhancement processing while regions of non-interest receive suppression processing. The per-band gain structure allows different frequency components to be treated differently based on their spatial origin, achieving local optimization of speech quality and noise suppression.
2Object-affected harmful factors
If the system applies suppression to noise signals, then noise levels decrease, but speech signals may also be inadvertently suppressed leading to loss of useful audio information
Solution Approach 1:
The system changes the gain parameter differently for different frequency bands and spatial regions. By adjusting per-band gains based on angle of arrival and covariance analysis, the system can suppress noise in certain bands while preserving speech in other bands, preventing information loss while reducing noise.
Solution Approach 2:
The patent replaces conventional mechanical audio mixing with a sophisticated digital signal processing system that uses covariance matrices and angle of arrival estimation. This substitution enables intelligent differentiation between speech and noise based on their spatial and spectral characteristics, allowing selective suppression without losing useful speech information.
3Device complexity
If the system uses broad-band suppression techniques, then noise suppression is simplified, but the accuracy of speech enhancement decreases leading to over-suppression or under-suppression
Solution Approach 1:
The audio signal is segmented into multiple frequency bands, each processed with its own gain structure. This segmentation allows the system to achieve precise speech enhancement by treating different frequency components separately, capturing the spectral characteristics of speech and noise more accurately than broad-band methods.
Solution Approach 2:
The gain structure is made dynamic by continuously estimating angle of arrival and covariance matrices for each time frame. The per-band gains are adjusted in real-time based on the current spatial and spectral conditions, allowing the system to adapt to changing speech and noise environments and maintain high enhancement accuracy.
Data Source
AI summary
Methods, systems, and media for processing audio are provided. In some embodiments, a method involves receiving, from a plurality of microphones, an input audio signal. The method may involve identifying an angle of arrival associated with the input audio signal. The method may involve determining a plurality of gains corresponding to a plurality of bands of the input audio signal based on a combination of at least: 1) a representation of a covariance of signals associated with microphones of the plurality of microphones on a per-band basis; and 2) the angle of arrival. The method may involve applying the plurality of gains to the plurality of bands of the input audio signal such that at least a portion of the input audio signal is suppressed to form an enhanced audio signal.


