Situation-Dependent Transient Noise Suppression in Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing noise suppression technologies often distort speech signals while attempting to remove transient noises like key clicks, as they apply uniform suppression strategies that are ineffective during voiced and unvoiced segments, leading to suboptimal user experience in audio/video calls.
Innovation Solution
A method and system for situation-dependent transient noise suppression that estimates voice probabilities to apply different levels of suppression, using a 'soft' approach during voiced segments to minimize distortion and a 'hard' approach during unvoiced segments, thereby maintaining signal intelligibility and reducing noise annoyance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If uniform noise suppression is applied to all audio segments, then transient noise is suppressed, but speech signal is distorted
Solution Approach 1:
The audio signal is segmented into voiced and unvoiced segments using voice activity detection. Different suppression strategies are applied to each segment type: soft suppression for voiced segments to preserve speech quality, and hard suppression for unvoiced segments to maximize noise removal. This segmentation resolves the contradiction by allowing targeted suppression that adapts to the local signal characteristics.
Solution Approach 2:
The suppression strategy dynamically adapts based on the detected voice state. The system transitions between soft and hard suppression modes depending on whether the current segment is voiced or unvoiced. This dynamic adaptation allows the system to optimize the balance between noise suppression and speech preservation in real-time, resolving the contradiction of uniform suppression.
2Object-affected harmful factors
If aggressive noise suppression is applied, then transient noise is reduced, but speech intelligibility deteriorates
Solution Approach 1:
The system applies different quality levels of suppression to different local segments of the audio signal. Voiced segments receive soft suppression that preserves speech quality, while unvoiced segments receive hard suppression that maximizes noise removal. This local differentiation resolves the contradiction by ensuring that speech intelligibility is maintained where it matters most (in voiced segments) while still achieving noise reduction overall.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided are methods and systems for providing situation-dependent transient noise suppression for audio signals. Different strategies (e.g., levels of aggressiveness) of transient suppression and signal restoration are applied to audio signals associated with participants in a video/audio conference depending on whether or not each participant is speaking (e.g., whether a voiced segment or an unvoiced/non-speech segment of audio is present). If no participants are speaking or there is an unvoiced/non-speech sound present, a more aggressive strategy for transient suppression and signal restoration is utilized. On the other hand, where voiced audio is detected (e.g., a participant is speaking), the methods and systems apply a softer, less aggressive suppression and restoration process.