Self-supervised voice noise reduction method and device

By introducing TCM module and ONT strategy into the DCUnet network, a self-supervised speech noise reduction model is built, which solves the adaptability and computing efficiency of the existing model in complex noise environments, and realizes efficient speech noise reduction and multi-sound source separation.

CN120496488AActive Publication Date: 2025-08-15HAINACORD (HUBEI) TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510499950.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-15
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

When dealing with complex noise environments and interlaced scenes of multiple sound sources, the existing speech noise reduction model is insufficient in adaptability and is difficult to adapt to dynamically changing noise scenarios. The computing resource requirements are high, which affects real-time and generalization capabilities.

Method used

The DCUnet network with a deep complex domain combined with the time-channel modeling (TCM) module is adopted, and through the self-supervised learning strategy ONT, the noise speech training model is used to build a speech noise reduction model, including an encoder, TCM module, Complex-TSTM module and decoder, optimize the time-frequency domain signal representation to reduce the dependence on clear speech targets.

Benefits of technology

It improves the reconstruction quality and noise reduction performance of speech signals, enhances the adaptability of the model in complex environments, reduces the cost of training data collection, and improves the generalization ability and computing efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496488A_ABST
    Figure CN120496488A_ABST
Patent Text Reader

Abstract

The invention provides a self-supervised voice noise reduction method and device, and relates to the field of acoustic signal processing, and the method comprises the steps: S1, obtaining a DCUnet network, and building a voice noise reduction model through adding a TCM module in the DCUnet network; s2, obtaining a noise voice set, training the voice noise reduction model through the noise voice set, and obtaining a trained voice noise reduction model; and S3, performing voice noise reduction through the trained voice noise reduction model. According to the method, the voice noise reduction model is constructed based on the DCUnet of the deep complex field, and the reconstruction quality and the noise reduction performance of the voice signal are improved in combination with a complex field overall processing strategy; change characteristics of a noise scene are dynamically captured through a TCM module in the voice noise reduction model, and the adaptability of the model to dynamic signals in a complex environment is enhanced; the ONT strategy is adopted to train the voice noise reduction model, the noise voice serves as training data, clear target voice data is not needed, the training data collection cost is reduced, and meanwhile the generalization ability of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of acoustic signal processing, and in particular to a self-supervised speech noise reduction method and device. Background Art

[0002] In complex urban environments, noise interference poses a significant challenge to the separation and analysis of sound sources. Especially in scenarios where multiple sound sources are intertwined, noise not only significantly degrades speech signal separation but can also mask critical speech or ambient sound information, thereby impacting the performance of downstream tasks such as smart city monitoring, noise control assessment, and the development of multimodal perception systems. Therefore, the research and application of speech noise reduction technologies in complex environments is extremely important.

[0003] Currently, most speech noise reduction tasks rely on "noisy-clean training" (NCT) strategies. These methods achieve noise reduction by training a network to map noisy speech to clean speech. However, obtaining completely clean speech signals requires expensive recording equipment and a strictly controlled recording environment. Data collection is time-consuming and expensive, and often limited in scale and diversity. To address this issue, self-supervised learning strategies have emerged, such as "noisy-noisy training" (NNT) and "noisier-noisy training" (NerNT). These methods achieve noise reduction by constructing a mapping between noisy speech and clean speech, but they still suffer from insufficient performance when dealing with multiple noise sources in complex environments. Furthermore, these methods have a strong dependence on the characteristics of the noise distribution, making them difficult to adapt to dynamically changing noise scenarios. Furthermore, the models lack generalization capabilities in high-noise scenarios, which can lead to oversmoothing of the speech signal or loss of important details.

[0004] Furthermore, from the perspective of model optimization, existing speech noise reduction networks are mostly based on real-valued domain computations, typically focusing solely on estimating spectral amplitudes while neglecting phase modeling, which limits signal reconstruction quality. Although deep complex networks (such as DCUnet) optimize signal-to-noise ratio loss by incorporating complex time-frequency masks, significantly improving signal reconstruction capabilities, existing methods struggle with handling long-range dependencies and global contextual information. Furthermore, the high complexity and computational resource requirements of these methods limit their practical application.

[0005] Therefore, developing a speech denoising method that does not rely on clean speech targets and can effectively model complex noise environments has become an important research direction in the field of speech signal processing. Denoising not only significantly improves the signal-to-noise ratio of the target speech but also provides a clearer signal input for the separation of mixed sounds, thereby reducing mutual interference between multiple sound sources and improving the overall performance of the separation algorithm. In practical applications, denoising technology lays the foundation for feature extraction and classification of mixed signals and is an indispensable step in achieving multi-source separation. One of the key paths in the development of speech denoising technology is to optimize the joint representation of time-frequency domain signals by combining complex domain processing and context modeling capabilities.

[0006] While self-supervised learning methods such as NNT and NerNT have, to some extent, eliminated the need for clear speech target data, they still have limitations when dealing with complex noisy environments and dynamic multi-source scenarios. These methods typically assume that the noise distribution has zero mean or specific statistical properties. However, the noise distribution in real environments is complex and variable, often inconsistent with these assumptions, resulting in a significant decrease in noise reduction performance. Furthermore, these methods place high demands on the initial model settings and the number of noise samples. When noise characteristics are not fully covered, the model may overfit to specific noise patterns and lack sufficient generalization capabilities.

[0007] In addition to the limitations of training strategies, existing noise reduction models also face challenges in improving their performance. Denoising networks, such as the Deep Complex U-Net (DCUnet), significantly improve signal amplitude and phase modeling capabilities by introducing complex time-frequency masks. However, their architecture still suffers from the following deficiencies: First, DCUnet's limited ability to model long-range dependencies and global contextual information leads to an inability to fully capture the dynamic interactions between multiple sound sources in complex signal separation tasks. Second, the model's high processing complexity for complex signals and its high computational resource requirements limit its application in resource-constrained environments. Finally, DCUnet's poor ability to reconstruct spectral details in high-noise environments can lead to loss of key features in speech signals or excessive smoothing.

[0008] On the one hand, existing noise reduction models generally lack adaptability to dynamic signal changes. Traditional self-supervised learning methods have poor adaptability to dynamic noise scenarios and struggle to adjust model parameters in real time to handle rapidly changing noise characteristics. This is primarily due to these methods' lack of comprehensive utilization of signal timing information and their neglect of the temporal correlation between noise and target signals, which limits the model's noise reduction performance and stability.

[0009] On the other hand, existing noise reduction methods still need to be strengthened in terms of real-time performance and computational efficiency. Especially when faced with multi-source interweaving, dynamic changes, and high-complexity noise scenarios, models generally require more computing resources to accurately separate and reconstruct signals. This not only affects the real-time performance of noise reduction algorithms, but also limits their deployment capabilities in resource-constrained devices such as smart devices and portable voice assistants. Furthermore, existing models often require separate processing of the real and imaginary parts when optimizing complex domain calculations. This design increases computational complexity and may affect the overall signal modeling effect. Summary of the Invention

[0010] In view of this, the purpose of the present invention is to provide a self-supervised speech noise reduction method and device to solve the technical problems of insufficient performance of existing noise reduction models in terms of adaptability to dynamic signal changes, real-time performance and computational efficiency.

[0011] The present invention provides a self-supervised speech noise reduction method, comprising the steps of:

[0012] S1: Obtain the DCUnet network and build a speech noise reduction model by adding the TCM module to the DCUnet network;

[0013] S2: Obtain a noisy speech set, train a speech noise reduction model using the noisy speech set, and obtain a trained speech noise reduction model;

[0014] S3: Perform speech noise reduction using the trained speech noise reduction model.

[0015] Preferred:

[0016] The speech noise reduction model includes an encoder, a TCM module, a Complex-TSTM module and a decoder connected in sequence.

[0017] Preferred:

[0018] The encoder consists of N Conv2D modules connected in sequence, and the Conv2D modules are numbered from 1 to N;

[0019] The decoder includes N Conv2D modules connected in sequence, and the Conv2D modules are numbered from N to 1;

[0020] The Conv2D modules with the same number in the encoder and decoder are connected to each other.

[0021] Preferred:

[0022] The TCM module includes a head token generation module, a multi-head self-attention mechanism module, and a classification token enhancement module connected in sequence;

[0023] The encoder is connected to the head Token generation module, and the classification Token enhancement module is connected to the Complex-TSTM module.

[0024] Preferred:

[0025] The Complex-TSTM module includes: RealTSTM part and Imag TSTM part;

[0026] The classification token enhancement module is connected with the RealTSTM part and the Imag TSTM part;

[0027] The Real TSTM part and the Imag TSTM part are connected to the decoder.

[0028] Preferably, step S2 is specifically as follows:

[0029] S21: converting the noisy speech into a noise image, inputting the noise image into an encoder in the speech noise reduction model for encoding, and obtaining a first feature image;

[0030] S22: Inputting the first feature image into the TCM module to enhance feature expression capability to obtain a second feature image;

[0031] S23: Input the second feature image into the Complex-TSTM module to calculate the comprehensive loss, and then input the second feature image into the decoder for decoding to obtain the initial noise-reduced speech;

[0032] S24: superimposing the initial noise-reduced speech and the noisy speech to obtain a final noise-reduced speech, and adjusting the parameters of the speech noise reduction model based on the final noise-reduced speech;

[0033] S25: Repeat steps S21-S24 until the comprehensive loss is less than a preset value, and obtain a trained speech noise reduction model.

[0034] Preferred:

[0035] The calculation formula of comprehensive loss L is:

[0036]

[0037] in, is the time domain loss, is the frequency loss, is the weighted signal-to-noise ratio loss, is the regularization loss, and α and β are hyperparameters.

[0038] A storage medium stores instructions and data for implementing the self-supervised speech noise reduction method.

[0039] A self-supervised speech noise reduction device comprises: a processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement the self-supervised speech noise reduction method.

[0040] The present invention has the following beneficial effects:

[0041] A speech denoising model is constructed based on the deep complex domain DCUnet network. Combined with the overall complex domain processing strategy, it improves the reconstruction quality and denoising performance of speech signals. The TCM module in the speech denoising model dynamically captures the changing characteristics of noise scenes, enhancing the model's adaptability to dynamic signals in complex environments. The speech denoising model is trained using the ONT strategy, using noisy speech as training data, eliminating the need for clear target speech data, reducing the cost of training data collection and improving the model's generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a flow chart of a method according to an embodiment of the present invention;

[0043] Figure 2 This is the structural diagram of the DCUnet network;

[0044] Figure 3 This is the structural diagram of the speech noise reduction model;

[0045] Figure 4 It is the structural diagram of the TCM module;

[0046] Figure 5 This is a structural diagram of the device according to an embodiment of the present invention;

[0047] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0048] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0049] Reference Figure 1 To address the problems of existing speech noise reduction methods in dealing with complex noise environments and multi-source interweaving scenarios, such as insufficient adaptability, limited modeling capabilities, and poor real-time performance, the present invention proposes a self-supervised speech noise reduction method that combines temporal-channel modeling (TCM module) with complex domain optimization, including the following steps:

[0050] S1: Obtain the DCUnet network and build a speech noise reduction model by adding the TCM module to the DCUnet network;

[0051] As an embodiment, the deep complex U-Net (DCUnet network) is used as the basic model in the network architecture, and is optimized and expanded. The structure of the DCUnet network is as follows: Figure 2 As shown in Figure 2. In the encoder and decoder modules, complex two-dimensional convolution and complex value deconvolution operations are combined to improve the adaptability and robustness of the model when processing complex noise signals. At the same time, high-resolution feature information is transmitted between the encoder and decoder through jump connections, effectively preserving the semantics and details of the signal. The structure of the speech denoising model is shown in Figure 2. Figure 3 shown.

[0052] The speech noise reduction model includes an encoder, a TCM module, a Complex-TSTM module and a decoder connected in sequence.

[0053] As an example:

[0054] The encoder consists of N Conv2D modules connected in sequence, and the Conv2D modules are numbered from 1 to N;

[0055] The decoder includes N Conv2D modules connected in sequence, and the Conv2D modules are numbered from N to 1;

[0056] The Conv2D modules with the same number in the encoder and decoder are connected to each other.

[0057] As an embodiment, the TCM module designed in the present invention significantly enhances the model's ability to capture the dependency between the time dimension and the channel dimension.

[0058] The TCM module includes a head token generation module, a multi-head self-attention mechanism module, and a classification token enhancement module connected in sequence;

[0059] The encoder is connected to the head Token generation module, and the classification Token enhancement module is connected to the Complex-TSTM module.

[0060] Specifically, the structure of the TCM module is as follows Figure 4 As shown, Figure 4 Here, Head Token Generation represents the head token generation module, Multi-Head Self-Attention represents the multi-head self-attention mechanism module, and ClassificationToken Enrichment represents the classification token enhancement module.

[0061] (1) Header Token generation module:

[0062] The TCM module first extracts the channel information of the input signal through the header Token generation component. The input sequence consists of classification Token CLS and time Token The time tokens in the sequence are divided into H segments along the channel dimension, and the dimension of each segment is d = D / H, where H represents the number of attention heads. Next, each segment generates channel features through time average pooling, and is projected back to D dimensions through a fully connected layer and GeLU activation function to form a head token. These tokens represent different parts of the channel information and are then concatenated with the input sequence to form a time-channel token sequence with a length of T+H+1.

[0063] (2) Multi-head self-attention mechanism module:

[0064] In the TCM module, the multi-head self-attention mechanism (MHSA) works similarly to the traditional MHSA, but its input sequence contains not only time tokens but also channel tokens. In order to learn the interaction between time and channel, the multi-head self-attention mechanism converts the time-channel token into Query (Q), Key (K) and Value (V). This process is achieved through the corresponding linear projection matrix The time-channel token is projected H times, where i represents the index of the attention head. Each projection generates a d-dimensional channel representation. Then, through the scaled dot product calculation, the self-attention operation calculates the appropriate weight along the time axis based on the correlation between each token. This process is performed in parallel in the H attention heads. Then, the outputs of all attention heads are spliced and the final linear projection matrix W is used. O Converted to output embedding. The overall formula of multi-head self-attention is as follows:

[0065] MultiHead(X)=Concat(head1,…,head H )W O

[0066]

[0067] (3) Classification Token Enhancement Module:

[0068] Although the classification token CLS in MHSA already extracts information from both time and channel tokens, to further enhance the information expressed in the classification token, the TCM module separates the time and header tokens from the MHSA output and performs average pooling on each. The pooled time and header mean tokens are then used to enrich the classification token, providing more comprehensive information support for the final prediction.

[0069] As an example:

[0070] The Complex-TSTM module includes: RealTSTM part and Imag TSTM part;

[0071] The classification token enhancement module is connected with the RealTSTM part and the Imag TSTM part;

[0072] The RealTSTM part and the Imag TSTM part are connected to the decoder.

[0073] Specifically, in the TSTM module, the complex features of the input are divided into real and imaginary parts, which are processed by the Real TSTM and Imag TSTM, respectively. Specifically, the Real TSTM processes the real part of the complex features, while the Imag TSTM processes the imaginary part. The two operate independently during the calculation process, extracting speech features along different dimensions. After processing, the results of the real and imaginary parts are recombined through complex arithmetic operations to generate a complete complex output, thus supporting the subsequent decoder.

[0074] S2: Obtain a noisy speech set, train a speech noise reduction model using the noisy speech set, and obtain a trained speech noise reduction model;

[0075] As an example, to achieve efficient training, the present invention uses an ONT strategy to generate training pairs from a single noisy speech sample, without relying on clear speech target data. The ONT strategy generates conditionally independent audio pairs through subsampling and combines it with a regularization loss term to optimize network performance. This significantly reduces the training's dependence on noise distribution assumptions, thereby improving the model's generalization ability in complex dynamic noise environments.

[0076] Step S2 is specifically as follows:

[0077] S21: converting the noisy speech into a noise image, inputting the noise image into an encoder in the speech noise reduction model for encoding, and obtaining a first feature image;

[0078] S22: Inputting the first feature image into the TCM module to enhance feature expression capability to obtain a second feature image;

[0079] S23: Input the second feature image into the Complex-TSTM module to calculate the comprehensive loss, and then input the second feature image into the decoder for decoding to obtain the initial noise-reduced speech;

[0080] Specifically, a comprehensive loss is constructed by combining time domain loss, frequency loss, weighted signal-to-noise ratio loss and regularization loss.

[0081] The calculation formula of comprehensive loss L is:

[0082]

[0083] in, is the time domain loss, is the frequency loss, is the weighted signal-to-noise ratio loss, is the regularization loss, and α and β are hyperparameters.

[0084] Specifically, time domain loss It is calculated by the mean square error (MSE) between the enhanced waveform and the clear waveform, which is defined as:

[0085]

[0086] Among them, s i and They represent the speech samples of the i-th clear speech sample and the denoised sample respectively, and N is the total number of audio samples.

[0087] Frequency domain loss It is used to monitor the model to learn more information, thereby improving the intelligibility and perceptual quality of speech. It is defined as:

[0088]

[0089] Among them, S and denote the clear spectrum and enhanced spectrum respectively, r and i denote the real and imaginary parts of the complex number respectively, T and F denote the number of frames and frequency bins respectively.

[0090] Weighted SNR loss It is used to directly optimize the commonly used evaluation indicators in the time domain, which are defined as:

[0091]

[0092] Among them, x represents the noisy sample, y represents the target sample, represents the estimated output, and α represents the energy ratio between the target speech and the noise.

[0093] For an audio pair s1(x) and s2(x) sampled from a noisy speech x, the present invention uses a regularization loss As an additional constraint, it is defined as:

[0094]

[0095] Among them, f θ represents the denoising network. In order to stabilize the training process, the update of s1(fθ (x)) and s2(f θ (x)) and gradually increase the hyperparameter γ to achieve the best training effect.

[0096] S24: superimposing the initial noise-reduced speech and the noisy speech to obtain a final noise-reduced speech, and adjusting the parameters of the speech noise reduction model based on the final noise-reduced speech;

[0097] S25: Repeat steps S21-S24 until the comprehensive loss is less than a preset value, and obtain a trained speech noise reduction model.

[0098] S3: Perform speech noise reduction using the trained speech noise reduction model.

[0099] Specifically, a trained speech noise reduction model enables comprehensive modeling of complex speech signals, significantly improving noise reduction performance, computational efficiency, and adaptability in practical applications. This provides an efficient, robust, and innovative solution for speech signal processing. Furthermore, this model is applicable to a variety of application scenarios, including multi-source separation, speech enhancement, and noise suppression in complex environments. It provides technical support for fields such as intelligent surveillance, speech recognition, and human-computer interaction.

[0100] See Figure 5 , Figure 5 4 is a schematic diagram of the working of the hardware device of an embodiment of the present invention, wherein the hardware device specifically includes: a self-supervised speech noise reduction device 401, a processor 402 and a storage medium 403.

[0101] A self-supervised speech noise reduction device 401: The self-supervised speech noise reduction device 401 implements the self-supervised speech noise reduction method.

[0102] Processor 402: The processor 402 loads and executes the instructions and data in the storage medium 403 to implement the self-supervised speech noise reduction method.

[0103] Storage medium 403: The storage medium 403 stores instructions and data; the storage medium 403 is used to implement the self-supervised speech noise reduction method.

[0104] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.

[0105] The serial numbers of the embodiments of the present invention are for descriptive purposes only and do not represent superiority or inferiority of the embodiments. In a unit claim that lists several means, several of these means may be embodied by the same item of hardware. The use of the terms first, second, and third, etc., does not denote any order and should be construed as identifiers.

[0106] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A self-supervised speech noise reduction method, characterized in that: Including steps: S1: Obtain the DCUnet network and build a speech noise reduction model by adding the TCM module to the DCUnet network; S2: Obtain a noisy speech set, train a speech noise reduction model using the noisy speech set, and obtain a trained speech noise reduction model; S3: Perform speech noise reduction using the trained speech noise reduction model.

2. The self-supervised speech noise reduction method according to claim 1, wherein: The speech noise reduction model includes an encoder, a TCM module, a Complex-TSTM module and a decoder connected in sequence.

3. The self-supervised speech noise reduction method according to claim 2, wherein: The encoder consists of N Conv2D modules connected in sequence, and the Conv2D modules are numbered from 1 to N; The decoder includes N Conv2D modules connected in sequence, and the Conv2D modules are numbered from N to 1; The Conv2D modules with the same number in the encoder and decoder are connected to each other.

4. The self-supervised speech noise reduction method according to claim 2, wherein: The TCM module includes a head token generation module, a multi-head self-attention mechanism module, and a classification token enhancement module connected in sequence; The encoder is connected to the head Token generation module, and the classification Token enhancement module is connected to the Complex-TSTM module.

5. The self-supervised speech noise reduction method according to claim 4, characterized in that: The Complex-TSTM module includes: RealTSTM part and Imag TSTM part; The classification token enhancement module is connected with the Real TSTM part and the Imag TSTM part; The RealTSTM part and the Imag TSTM part are connected to the decoder.

6. The self-supervised speech noise reduction method according to claim 2, characterized in that Step S2 is specifically as follows: S21: converting the noisy speech into a noise image, inputting the noise image into an encoder in the speech noise reduction model for encoding, and obtaining a first feature image; S22: Inputting the first feature image into the TCM module to enhance feature expression capability to obtain a second feature image; S23: Input the second feature image into the Complex-TSTM module to calculate the comprehensive loss, and then input the second feature image into the decoder for decoding to obtain the initial noise-reduced speech; S24: superimposing the initial noise-reduced speech and the noisy speech to obtain a final noise-reduced speech, and adjusting the parameters of the speech noise reduction model based on the final noise-reduced speech; S25: Repeat steps S21-S24 until the comprehensive loss is less than a preset value, and obtain a trained speech noise reduction model.

7. The self-supervised speech noise reduction method according to claim 6, characterized in that: The calculation formula of comprehensive loss L is: in, is the time domain loss, is the frequency loss, is the weighted signal-to-noise ratio loss, is the regularization loss, and α and β are hyperparameters.

8. A storage medium, characterized in that: The storage medium stores instructions and data for implementing the self-supervised speech noise reduction method according to any one of claims 1 to 7.

9. A self-supervised speech noise reduction device, characterized by: include: A processor and a storage medium; the processor loads and executes instructions and data in the storage medium to implement the self-supervised speech noise reduction method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Single-channel speech enhancement method based on deep reconvolution network

    CN114360567A

  • KAN convolution improvement-based DCU-Net intelligent sound box speech recognition and noise reduction method

    CN119132325A

  • Noise reduction system for dynamic noise reduction

    EP4531042A1