Adaptive Audio Source Classification for Noise Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current noise suppression systems in adverse audio environments face challenges due to fixed noise suppression levels, which can lead to speech distortion and inadequate handling of varying noise types and speech fluctuations, as they rely on signal-to-noise ratios and fixed classification thresholds that are not robust against changes in the audio environment.
Innovation Solution
The system adaptively classifies audio sources by deriving acoustic features, determining global summaries, and using inter-microphone level differences to differentiate between speech and noise, generating control signals for an enhancement filter to adjust noise suppression dynamically and minimize speech degradation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If fixed noise suppression level (12-13 dB) is used, then speech distortion is avoided, but noise suppression is insufficient in adverse audio environments
Solution Approach 1:
The patent applies dynamics by transitioning from fixed noise suppression levels to dynamic, adaptive noise suppression. The system continuously adjusts suppression levels based on real-time analysis of acoustic features, speech energy fluctuations, and noise characteristics, allowing the noise suppression to vary between frames and adapt to changing audio conditions while maintaining speech quality.
Solution Approach 2:
The patent changes the parameter of noise suppression level from a fixed value to a variable parameter that adapts to different audio conditions. By analyzing acoustic features and speech energy over time, the system modifies the suppression parameter dynamically, enabling higher suppression when needed while avoiding distortion when speech is present.
2Object-affected harmful factors
If dynamic noise suppression based on SNR is used, then noise suppression level increases, but speech distortion increases due to SNR being adversely impacted by speech energy fluctuations
Solution Approach 1:
The patent implements feedback mechanisms that continuously monitor acoustic features, speech energy, and noise characteristics. This feedback information is used to adjust the noise suppression levels in real-time, allowing the system to respond to speech energy fluctuations and adapt suppression to maintain both high noise reduction and speech quality.
Solution Approach 2:
The patent segments the audio signal processing into multiple frames and analyzes acoustic features within each frame. This segmentation allows the system to handle speech energy fluctuations on a per-frame basis, applying appropriate suppression levels to each segment independently, thereby maintaining speech quality while achieving high noise suppression.
3Device complexity
If fixed classification threshold discrimination system is used, then classification is simple, but classification robustness decreases when conditions change
Solution Approach 1:
The patent applies dynamics to the classification system by transitioning from fixed thresholds to adaptive, data-driven classification. The system learns from the audio data itself, adjusting classification criteria based on observed acoustic features, speech energy patterns, and noise characteristics, thereby maintaining high classification accuracy under varying conditions.
Solution Approach 2:
The classification system performs self-service by automatically adapting to changing conditions through analysis of the incoming audio data. Rather than relying on pre-set thresholds, the system uses the audio signal characteristics themselves to determine appropriate classification, making it self-adjusting and robust to environmental changes.
Data Source
AI summary
Systems and methods for adaptively classifying audio sources are provided. In exemplary embodiments, at least one acoustic signal is received. One or more acoustic features based on the at least one acoustic signal are derived. A global summary of acoustic features based, at least in part, on the derived one or more acoustic features is determined. Further, an instantaneous global classification based on a global running estimate and the global summary of acoustic features is determined. The global running estimates may be updated and an instantaneous local classification based, at least in part, on the one or more acoustic features may be derived. One or more spectral energy classifications based, at least in part, on the instantaneous local classification and the one or more acoustic features may be determined. In some embodiments, the spectral energy classification is provided to a noise suppression system.


