Speech Enhancement Network Training for Non-Speech Noise Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech enhancement networks generate residual noises during processing of non-speech segments, leading to reduced speech enhancement quality.

Innovation Solution

A method for training a speech enhancement network that classifies speech effectiveness in enhanced speech frames, determines noise reduction and speech classification accuracy, and trains the network based on these accuracies to improve its capability to suppress noises in non-speech segments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If speech enhancement processing is performed on noisy speech encompassing non-speech segments, then speech enhancement quality is improved, but residual noises are generated in non-speech segments

Engineering Contradiction:
Improvespeech enhancement qualityVSAvoidresidual noises
Core Design Contradiction:
Manufacturing precisionVSObject-generated harmful factors

Solution Approach 1:

The patent divides the enhanced speech into multiple frames and classifies each frame as either speech or non-speech segment. This segmentation allows different processing strategies to be applied to different types of segments, thereby improving overall speech enhancement quality while specifically addressing residual noises in non-speech segments through targeted suppression techniques

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different quality standards and processing approaches to different parts of the speech signal. Speech segments are enhanced with one set of criteria while non-speech segments are processed with another set of criteria focused on noise suppression. This local differentiation allows the system to maintain high speech enhancement quality without generating residual noises in non-speech portions

Inventive Principle:
Principle #3Local quality

2Measurement precision

If speech effectiveness classification is added to the training process, then speech enhancement accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvespeech enhancement accuracyVSAvoidtraining process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines speech enhancement training with speech effectiveness classification training into a unified training framework. By merging these two tasks, the system simultaneously learns to enhance speech while classifying speech effectiveness, improving overall accuracy without requiring separate complex training processes. The multi-task learning approach integrates both objectives into a single coherent training pipeline

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The speech enhancement network is designed to perform multiple functions: enhancing speech quality and classifying speech effectiveness. This multi-functional design allows the same network structure to achieve both goals, reducing the need for separate specialized systems and thereby managing complexity while improving accuracy

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP4730329A1Training method for speech enhancement network, speech enhancement method, and electronic device
Publication Date: 2026.04.22 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • EP4730329A1 patent drawingFigure 1
  • EP4730329A1 patent drawingFigure 2~3
  • EP4730329A1 patent drawingFigure 4

AI summary

A training method for a speech enhancement network, a speech enhancement method, and an electronic device. The training method comprises: performing speech effectiveness classification on respective enhanced speech frames, and on the basis of the classification results of the enhanced speech frames, generating an effectiveness distribution for sample enhanced speech; by means of the effectiveness distribution, determining the speech classification accuracy of the speech enhancement network; and measuring the degree of change in speech effectiveness of each enhanced speech frame compared to that before noise reduction, and on this basis, determining speech enhancement accuracy according to noise reduction accuracy and the speech classification accuracy. In this way, the method can focus on improving the ability of the speech enhancement network to suppress noise in non-speech segments. When the trained speech enhancement network is used to perform noise reduction on speech to be processed, if the speech comprises non-speech segments, the trained speech enhancement network can effectively reduce the phenomenon of residual noise and improve the quality of speech enhancement. The invention can be widely applied in various scenarios such as cloud technology, artificial intelligence, smart transportation, and assisted driving.