Multiple-Microphone Speech Enhancement via Adaptive Noise Cancellation and Neural Blending
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional single-microphone and dual-microphone speech systems struggle with noise reduction, especially when noise power exceeds speech power, and are prone to errors due to stationary and non-stationary noise scenarios, leading to limited operational effectiveness.
Innovation Solution
A multiple-microphone speech enhancement apparatus combining an adaptive noise cancellation (ANC) circuit, a blending circuit, a noise suppressor, and a control module, which uses a neural network-based noise suppressor and beamformer to classify noise and speech components, and adjust blending gains to optimize noise suppression across various environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional single-microphone or dual-microphone noise reduction approaches are used, then the system operates under limited circumstances (stationary noise with noise power less than speech power), but it fails when noise power exceeds speech power or in non-stationary noise scenarios
Solution Approach 1:
The system dynamically adapts its processing mode based on real-time noise classification. The control module switches between different noise reduction strategies (spectral subtraction, Wiener filtering, adaptive noise cancellation) depending on whether noise is stationary or non-stationary, and whether noise power exceeds or is less than speech power. This dynamic adaptation resolves the contradiction by making the system versatile across different operational conditions while maintaining reliable noise suppression in each specific scenario.
Solution Approach 2:
The system changes key parameters including noise power spectrum estimates, filtering coefficients, and blending ratios based on the detected noise characteristics. When noise power exceeds speech power, the system adjusts the noise power spectrum estimation and modifies the blending ratio between noise-reduced and original signals to prevent over-suppression. This parameter adaptation enables the system to maintain reliability across varying noise conditions.
2Reliability
If Voice Activity Detector (VAD) is used to control adaptive filter in Adaptive Noise Cancellation (ANC), then speech self-cancellation is prevented during voice active periods, but the system fails when high-level background noise causes VAD to make wrong decisions or when sudden noise is mistaken for speech
Solution Approach 1:
The system introduces a noise classification module as an intermediary between the VAD and the adaptive filter control. This intermediary classifies noise characteristics (stationary vs. non-stationary, noise power relative to speech power) and uses this classification to intelligently control when to trust VAD decisions and when to override them. For example, when non-stationary noise is detected, the system may disable VAD-based control to prevent false speech detection. This intermediary layer resolves the contradiction by enabling the system to maintain speech preservation accuracy while adapting to various noise conditions.
Solution Approach 2:
The control strategy dynamically adjusts based on noise classification results. The system transitions from rigid VAD-based control to more flexible control mechanisms when noise characteristics change. During stationary noise periods, VAD control is effective; during non-stationary noise or when noise power exceeds speech power, the system switches to alternative control strategies that don't rely solely on VAD decisions. This dynamic control adaptation resolves the contradiction between speech preservation and operational versatility.
3Reliability
If adaptive filter training is stopped when speech is present to prevent self-cancellation, then speech integrity is maintained, but the adaptive filter cannot converge and ANC stops operating effectively
Solution Approach 1:
The system dynamically controls the adaptive filter training based on noise classification and speech detection results. During non-stationary noise periods or when noise power exceeds speech power, the system disables adaptive filter training to prevent divergence from false speech detection. During stationary noise periods with clear speech absence detection, the system enables training to achieve convergence. This dynamic control resolves the contradiction by maintaining speech integrity when needed while enabling ANC operation when conditions are favorable.
Solution Approach 2:
The system performs preliminary noise classification before deciding whether to enable adaptive filter training. By classifying noise characteristics in advance, the system can proactively enable training only when conditions are suitable (stationary noise, speech absent), preventing convergence issues before they occur. This preliminary assessment resolves the contradiction by ensuring speech integrity is maintained while maximizing ANC operational effectiveness.
4Reliability
If multiple microphones and complex processing (classifying sections, neural network noise suppressor) are used to suppress noise regardless of power level, then noise suppression performance is improved, but device complexity increases
Solution Approach 1:
The system segments the noise suppression task into multiple specialized components: a noise classification module that identifies noise characteristics, a neural network-based noise suppressor that processes different noise types differently, and a control module that coordinates their operation. Each component handles a specific aspect of the problem, allowing the system to achieve high noise suppression performance across all conditions without requiring a monolithic complex structure. This segmentation resolves the contradiction by distributing complexity across modular, specialized units.
Solution Approach 2:
The control module serves multiple functions: it classifies noise characteristics, controls the adaptive filter training, manages the blending ratio between noise-reduced and original signals, and coordinates the neural network noise suppressor. This multi-functionality reduces overall system complexity by consolidating control logic into a single intelligent module rather than requiring separate dedicated components for each function. The universal control module resolves the contradiction by achieving high noise suppression performance through intelligent coordination rather than through brute-force complexity.
Data Source
AI summary
A speech enhancement apparatus is disclosed and comprises an adaptive noise cancellation circuit, a blending circuit, a noise suppressor and a control module. The ANC circuit filters a reference signal to generate a noise estimate and subtracts a noise estimate from a primary signal to generate a signal estimate based on a control signal. The blending circuit blends the primary signal and the signal estimate to produce a blended signal. The noise suppressor suppresses noise over the blended signal using a first trained model to generate an enhanced signal and a main spectral representation from a main microphone and M auxiliary spectral representations from M auxiliary microphones using (M+1) second trained models to generate a main score and M auxiliary scores. The ANC circuit, the noise suppressor and the trained models are well combined to maximize the performance of the speech enhancement apparatus.


