PCEN Mask Thresholding for DNN Speech Enhancement Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional DNN-based speech enhancement models struggle with removing stationary noise from studio-quality clean speech signals and require efficient mask thresholding and voice activity detection to improve performance.

Innovation Solution

Implement per-channel energy normalization (PCEN) to determine masks and a speech-aware loss function for training DNN-based models, incorporating PCEN-based thresholding and voice activity detection to differentiate between speech and non-speech frames.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If ideal ratio mask (IRM) is used for training DNN-based speech enhancement models, then the model can be trained to remove artifacts from noisy waveforms, but the model cannot effectively remove stationary noise present in studio-quality clean speech signals

Engineering Contradiction:
Improvespeech enhancement performanceVSAvoidstationary noise removal capability
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent applies preliminary action by performing PCEN processing on clean speech signals before generating training masks. This preprocessing step normalizes the energy distribution across frequency channels, making stationary noise components more distinguishable from speech components in advance of the main training process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes parameters by applying PCEN transformation to the clean speech signal, which modifies the energy representation in the time-frequency domain. This parameter transformation enhances the contrast between speech and stationary noise, enabling the model to learn better separation characteristics

Inventive Principle:
Principle #35Parameter changes

2Productivity

If a threshold is applied to zero out small energy time-frequency bands in clean speech, then computational complexity is reduced and processing efficiency is improved, but meaningful speech content may be incorrectly removed

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidspeech content preservation accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent transforms the energy representation using PCEN, which normalizes the dynamic range of energy values across frequency channels. This parameter change allows for more accurate thresholding decisions, as the normalized energy values better reflect the true speech presence probability

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent uses PCEN-processed energy values as feedback to dynamically adjust thresholding decisions. The normalized energy representation provides continuous feedback about speech presence, enabling adaptive threshold selection that preserves meaningful speech while removing noise

Inventive Principle:
Principle #23Feedback

3Reliability

If the DNN model is trained to preserve speech in speech regions, then speech quality is maintained, but the model cannot aggressively remove unwanted artifacts in non-speech regions

Engineering Contradiction:
Improvespeech preservation qualityVSAvoidartifact suppression capability
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent applies local quality by computing PCEN separately for each frequency channel and using channel-specific thresholds. This allows the model to apply different processing characteristics to different frequency regions, preserving speech where present and aggressively removing artifacts where absent

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the speech signal into multiple frequency channels and processes each channel independently with its own PCEN computation and thresholding. This segmentation enables region-specific optimization, where speech preservation and artifact removal can be controlled independently in different frequency bands

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250273226A1PCEN-based mask thresholding and voice activity detection for training DNN-based speech enhancement models
Publication Date: 2025.08.28 DOLBY LABORATORIES LICENSING CORP
  • US20250273226A1 patent drawing
  • US20250273226A1 patent drawing
  • US20250273226A1 patent drawing

AI summary

Described herein is a method of determining at least one mask for use in training a deep neural network (DNN)—based mask-based audio processing model. In particular, the method may comprise obtaining a time-frequency representation of a target audio signal for use in the training. The method may further comprise determining a per-channel energy normalization (PCEN) measure for the target audio signal. The method may yet further comprise determining the at least one mask based on the PCEN measure.