PCEN Mask Thresholding for DNN Speech Enhancement Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional DNN-based speech enhancement models struggle with removing stationary noise from studio-quality clean speech signals and require efficient mask thresholding and voice activity detection to improve performance.
Innovation Solution
Implement per-channel energy normalization (PCEN) to determine masks and a speech-aware loss function for training DNN-based models, incorporating PCEN-based thresholding and voice activity detection to differentiate between speech and non-speech frames.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If ideal ratio mask (IRM) is used for training DNN-based speech enhancement models, then the model can be trained to remove artifacts from noisy waveforms, but the model cannot effectively remove stationary noise present in studio-quality clean speech signals
Solution Approach 1:
The patent applies preliminary action by performing PCEN processing on clean speech signals before generating training masks. This preprocessing step normalizes the energy distribution across frequency channels, making stationary noise components more distinguishable from speech components in advance of the main training process
Solution Approach 2:
The patent changes parameters by applying PCEN transformation to the clean speech signal, which modifies the energy representation in the time-frequency domain. This parameter transformation enhances the contrast between speech and stationary noise, enabling the model to learn better separation characteristics
2Productivity
If a threshold is applied to zero out small energy time-frequency bands in clean speech, then computational complexity is reduced and processing efficiency is improved, but meaningful speech content may be incorrectly removed
Solution Approach 1:
The patent transforms the energy representation using PCEN, which normalizes the dynamic range of energy values across frequency channels. This parameter change allows for more accurate thresholding decisions, as the normalized energy values better reflect the true speech presence probability
Solution Approach 2:
The patent uses PCEN-processed energy values as feedback to dynamically adjust thresholding decisions. The normalized energy representation provides continuous feedback about speech presence, enabling adaptive threshold selection that preserves meaningful speech while removing noise
3Reliability
If the DNN model is trained to preserve speech in speech regions, then speech quality is maintained, but the model cannot aggressively remove unwanted artifacts in non-speech regions
Solution Approach 1:
The patent applies local quality by computing PCEN separately for each frequency channel and using channel-specific thresholds. This allows the model to apply different processing characteristics to different frequency regions, preserving speech where present and aggressively removing artifacts where absent
Solution Approach 2:
The patent segments the speech signal into multiple frequency channels and processes each channel independently with its own PCEN computation and thresholding. This segmentation enables region-specific optimization, where speech preservation and artifact removal can be controlled independently in different frequency bands
Data Source
AI summary
Described herein is a method of determining at least one mask for use in training a deep neural network (DNN)—based mask-based audio processing model. In particular, the method may comprise obtaining a time-frequency representation of a target audio signal for use in the training. The method may further comprise determining a per-channel energy normalization (PCEN) measure for the target audio signal. The method may yet further comprise determining the at least one mask based on the PCEN measure.


