Speech Enhancement Network Training for Residual Noise Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech enhancement networks generate residual noises in non-speech segments, reducing the quality of speech enhancement due to ineffective noise reduction in noisy speech signals.

Innovation Solution

A method for training a speech enhancement network involves mixing clean and noise samples to create noisy samples, performing noise reduction, framing enhanced speech frames, classifying their effectiveness, and training the network based on accuracy determinations to improve noise reduction and speech classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If speech enhancement network performs noise reduction on noisy speech signals, then speech quality is improved, but residual noises are generated in non-speech segments

Engineering Contradiction:
Improvespeech enhancement qualityVSAvoidresidual noises
Core Design Contradiction:
Manufacturing precisionVSObject-generated harmful factors

Solution Approach 1:

The enhanced speech sample is divided into multiple enhanced speech frames, and each frame is independently classified for speech effectiveness. This segmentation allows the system to identify and separately handle non-speech segments, preventing residual noise generation in those regions while maintaining enhancement quality in actual speech portions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing strategies to different segments of the speech signal based on their classification results. Speech frames are enhanced with noise reduction, while non-speech frames are handled differently to avoid generating residual noises. This local differentiation resolves the contradiction by optimizing quality where needed while preventing harm where not needed.

Inventive Principle:
Principle #3Local quality

2Ease of operation

If speech enhancement network processes all speech frames uniformly, then processing simplicity is maintained, but speech classification accuracy deteriorates

Engineering Contradiction:
Improveprocessing simplicityVSAvoidspeech classification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The speech signal is divided into frames that are independently classified and processed. This segmentation enables accurate speech classification by evaluating each frame separately, improving precision without significantly complicating the overall processing pipeline through modular operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system generates effectiveness distribution information from frame classification results and uses this feedback to guide the speech enhancement process. This feedback mechanism improves classification accuracy by continuously adjusting processing based on detected speech effectiveness, while maintaining operational simplicity through automated closed-loop control.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If effectiveness distribution is generated for each enhanced speech frame, then speech classification accuracy is improved, but computational load increases

Engineering Contradiction:
Improvespeech classification accuracyVSAvoidcomputational load
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

By dividing the speech signal into frames and processing them independently, the system achieves accurate speech classification through localized effectiveness distribution generation. This segmentation enables parallel processing of frames, improving accuracy while managing computational load through efficient resource utilization across multiple independent units.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system generates effectiveness distribution information selectively for frames that require classification, rather than uniformly processing all frames with the same computational intensity. This partial action approach improves speech classification accuracy where needed while reducing unnecessary computational expenditure in already-processed or clearly-non-speech segments.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250391419A1Method for training speech enhancement network, method for enhancing speech, and electronic device
Publication Date: 2025.12.25 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US20250391419A1 patent drawing
  • US20250391419A1 patent drawing
  • US20250391419A1 patent drawing

AI summary

A method for training a speech enhancement network, performed by an electronic device, includes: acquiring a first clean speech sample and a noise sample, and mixing them to generate a noisy speech sample; performing noise reduction on the noisy speech sample based on the speech enhancement network to obtain an enhanced speech sample; framing the enhanced speech sample into a plurality of enhanced speech frames, classifying speech effectiveness of the enhanced speech frames, and generating a first effectiveness distribution based on classification results of the enhanced speech frames; and determining a noise reduction accuracy based on the enhanced speech sample and the first clean speech sample, determining a speech classification accuracy based on the first effectiveness distribution, determining a speech enhancement accuracy based on the noise reduction accuracy and the speech classification accuracy, and training the speech enhancement network based on the speech enhancement accuracy.