Speech Enhancement via Machine Learning Noise Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods and systems for stationary noise reduction are inadequate as they fail to effectively reduce non-stationary noise sources such as dog barking, keyboard clicking, baby crying, music, and reverberation.
Innovation Solution
A machine learning-based approach that involves a computing device receiving sound inputs, converting them to time-frequency samples, determining time-frequency losses based on signal-to-noise ratios and speech probability estimates, and applying these losses to reduce non-speech portions of the sound inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional stationary noise reduction methods are used, then stationary noise can be reduced, but non-stationary noise cannot be effectively reduced
Solution Approach 1:
The patent applies dynamics by transitioning from static stationary noise reduction methods to dynamic non-stationary noise reduction. The system continuously adapts to changing noise characteristics by processing audio in overlapping frames and updating noise profiles in real-time, allowing the noise reduction algorithm to respond to temporal variations in noise properties while maintaining effectiveness against both stationary and non-stationary noise types
Solution Approach 2:
The patent implements parameter changes by modifying key parameters such as frame size, hop size, and noise profile update rates to optimize performance for different noise types. The system dynamically adjusts spectral subtraction parameters and applies different reduction factors based on detected noise characteristics, enabling effective reduction of both stationary and non-stationary noise through adaptive parameter modification
2Object-affected harmful factors
If aggressive noise reduction is applied, then noise levels decrease, but speech quality and naturalness deteriorate
Solution Approach 1:
The patent applies local quality by performing noise reduction operations at different levels of the audio signal processing hierarchy. Instead of uniformly processing the entire audio signal, the system operates on individual frequency bins and time frames, applying localized spectral subtraction and gain adjustment only where noise is detected, thereby preserving speech quality in clean regions while reducing noise in contaminated regions
Solution Approach 2:
The patent implements feedback mechanisms by continuously monitoring the reduced audio output and using it to update noise profiles for subsequent processing stages. The system employs feedback loops that adjust reduction factors based on detected speech activity and noise characteristics, preventing over-reduction of speech components while maintaining effective noise suppression through iterative refinement
Data Source
AI summary
Methods, systems, apparatuses for speech enhancement are described. A computing device may receive sound inputs and reduce non-speech portions of the sound inputs based on a machine learning model.


