Speech enhancement method, device and equipment

Through the self-supervised module learning the robust characterization of speech and using the vocoder module for high-fidelity reconstruction, the problem of pattern collapse and reliance on a large amount of labeled data in the prior art is solved, and the naturalness and clarity improvement of speech enhancement is achieved.

CN120199265AActive Publication Date: 2025-06-24UNIV OF SCI & TECH OF CHINA
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510550499.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-06-24
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing speech enhancement technology has pattern collapse problems during training, resulting in the lack of naturalness of the generated speech signals. The discriminant model relies on a large amount of labeled data, which has poor generalization and is difficult to effectively improve the clarity of speech.

Method used

The self-supervised module is used to learn the robust representation of the speech to be enhanced through the feature encoder and the Transformer module, obtain the noisy context representation, and high-fidelity reconstruction is carried out through the multi-layer one-dimensional transposed convolution and multi-resolution fusion module of the vocoder module to generate enhanced speech.

Benefits of technology

It effectively suppresses noise and reverb in the speech to be enhanced, significantly improves the naturalness and clarity of the speech, and does not need to rely on paired data, greatly reducing training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199265A_ABST
    Figure CN120199265A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a speech enhancement method, device and equipment, and relates to the technical field of data processing. The method comprises the following steps: acquiring to-be-enhanced voice, wherein the to-be-enhanced voice is noisy voice; the method comprises the following steps: inputting to-be-enhanced speech into a self-supervision module of a speech enhancement model, and converting the to-be-enhanced speech into noisy context representation; the noisy context representation is converted into enhanced speech by inputting the noisy context representation into a vocoder module of a speech enhancement model, the enhanced speech being clean speech. Therefore, the voice enhancement method provided by the embodiment of the invention can perform voice enhancement on the to-be-enhanced voice (noisy voice), thereby effectively suppressing noise and reverberation in the to-be-enhanced voice, and remarkably improving the naturalness and definition of the enhanced voice (clean voice).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and particularly to a voice enhancement method, apparatus, and device. Background Art

[0002] Voice enhancement technology refers to the technology of improving the clarity and naturalness of speech by suppressing background noise and reverberation in noisy speech. In application scenarios such as voice interaction in smart homes and voice calls on mobile devices in noisy environments, the performance of voice enhancement technology has a direct impact on the user experience.

[0003] In related technologies, a voice enhancement model based on a deep neural network (DNN) is usually used for voice enhancement processing. Voice enhancement models mainly include two categories: generative models and discriminative models. Among them, the generative model reconstructs the noisy speech into a speech signal consistent with the probability distribution by learning the probability distribution of clean speech. However, the generative model has the problem of mode collapse during the training process, that is, the generative model may overly focus on certain specific speech features and ignore other speech features, resulting in the lack of naturalness of the generated speech signal. The discriminative model refers to minimizing the difference between the training speech output by the voice enhancement model and the clean speech corresponding to the noisy speech to suppress noise after inputting the noisy speech into the voice enhancement model. However, the discriminative model relies on a large number of "noisy speech-clean speech" data pairs for training. When facing unseen noise types, due to the lack of learning ability of the discriminative model for the unseen noise types, it is difficult to effectively improve the clarity of speech.

[0004] Therefore, how to perform voice enhancement on noisy speech to obtain cleaner speech with higher naturalness and clarity has become a technical problem to be solved urgently. Summary of the Invention

[0005] Based on the above problems, the present application provides a voice enhancement method, apparatus, and device, which can perform voice enhancement on noisy speech and obtain cleaner speech with higher naturalness and clarity.

[0006] The embodiments of the present application disclose the following technical solutions:

[0007] In a first aspect, the present application discloses a voice enhancement method, and the method includes:

[0008] Obtain the speech to be enhanced, where the speech to be enhanced is noisy speech;

[0009] By inputting the speech to be enhanced into the self-supervised module of the speech enhancement model, converting the speech to be enhanced into a noisy context representation;

[0010] By inputting the noisy context representation into the vocoder module of the speech enhancement model, converting the noisy context representation into enhanced speech, and the enhanced speech is clean speech.

[0011] Optionally, the self-supervised module includes a feature encoder and a Transformer module; the converting the speech to be enhanced into a noisy context representation includes:

[0012] Extracting features from the speech to be enhanced through the feature encoder to obtain noisy features;

[0013] Randomly masking the noisy features through the Transformer module, and obtaining a noisy context representation according to the masked noisy features.

[0014] Optionally, the vocoder module includes a generator, and the generator includes a multi-layer one-dimensional transposed convolution and a multi-resolution fusion module; the converting the noisy context representation into enhanced speech includes:

[0015] Upsampling the noisy context through the multi-layer one-dimensional transposed convolution to obtain intermediate features at multiple time scales, and the time resolution of each intermediate feature is aligned with the time resolution of the speech to be enhanced;

[0016] Fusing the intermediate features at the multiple time scales through the multi-resolution fusion module to obtain enhanced speech.

[0017] Optionally, the vocoder module further includes a discriminator, and the discriminator includes a multi-scale discriminator and a multi-period discriminator; the fusing the intermediate features at the multiple time scales through the multi-resolution fusion module to obtain enhanced speech includes:

[0018] Fusing the intermediate features at the multiple time scales through the multi-resolution fusion module to obtain fused speech;

[0019] Evaluating the fused speech at multiple resolution levels through the multi-scale discriminator, and evaluating the periodic components of the fused speech through the multi-period discriminator;

[0020] When both the level evaluation and the component evaluation meet the corresponding evaluation conditions, determining the fused speech as enhanced speech.

[0021] Optionally, the training method of the self-supervised module is specifically as follows:

[0022] Obtain a noisy sample speech and a clean sample speech;

[0023] Respectively perform feature extraction on the noisy sample speech and the clean sample speech through the feature encoder to obtain a noisy sample feature and a clean sample feature;

[0024] Randomly mask the noisy sample feature and the clean sample feature respectively through the Transformer module, and respectively obtain a noisy sample context representation and a clean sample context representation according to the masked noisy sample feature and the masked clean sample feature;

[0025] Train the self-supervised module according to the noisy sample context representation, the clean sample context representation, the noisy sample feature, the clean sample feature and the total loss function, wherein the total loss function includes a contrastive loss function and a consistency loss function.

[0026] Optionally, the training method of the vocoder module is specifically as follows:

[0027] Upsample the noisy sample context through the multi-layer one-dimensional transposed convolution to obtain sample intermediate features at multiple time scales, and the time resolution of each sample intermediate feature is aligned with the time resolution of the noisy sample speech;

[0028] Fuse the sample intermediate features at the multiple time scales through the multi-resolution fusion module to obtain an enhanced sample speech;

[0029] Train the generator of the vocoder module according to the clean sample speech, the enhanced sample speech and the loss function of the generator, wherein the loss function of the generator includes a first adversarial loss function, a feature matching loss function, a Mel spectrum loss function, an amplitude loss function and a complex number loss function;

[0030] Train the discriminator of the vocoder module according to the clean sample speech, the enhanced sample speech and a second adversarial loss function.

[0031] In a second aspect, the present application discloses a speech enhancement device, and the device includes: a speech acquisition module, a first conversion module and a second conversion module;

[0032] The speech acquisition module is used to acquire the speech to be enhanced, and the speech to be enhanced is a noisy speech;

[0033] The first conversion module is used to convert the speech to be enhanced into a noisy context representation by inputting the speech to be enhanced into the self-supervised module of the speech enhancement model;

[0034] The second conversion module is configured to convert the noisy context representation into enhanced speech, which is clean speech, by inputting the noisy context representation into the vocoder module of the speech enhancement model.

[0035] Optionally, the self-supervised module includes a feature encoder and a Transformer module; the first conversion module includes: a first conversion sub-module and a second conversion sub-module:

[0036] The first conversion sub-module is configured to extract features from the speech to be enhanced through the feature encoder to obtain noisy features;

[0037] The second conversion sub-module is configured to randomly mask the noisy features through the Transformer module and obtain a noisy context representation according to the masked noisy features.

[0038] Optionally, the vocoder module includes a generator, and the generator includes a multi-layer one-dimensional transposed convolution and a multi-resolution fusion module; the second conversion module includes: a third conversion sub-module and a fourth conversion sub-module;

[0039] The third conversion sub-module is configured to upsample the noisy context through the multi-layer one-dimensional transposed convolution to obtain intermediate features at multiple time scales, and the time resolution of each intermediate feature is aligned with the time resolution of the speech to be enhanced;

[0040] The fourth conversion sub-module is configured to fuse the intermediate features at the multiple time scales through the multi-resolution fusion module to obtain enhanced speech.

[0041] Optionally, the vocoder module further includes a discriminator, and the discriminator includes a multi-scale discriminator and a multi-period discriminator; the fourth conversion sub-module is specifically configured to:

[0042] Fuse the intermediate features at the multiple time scales through the multi-resolution fusion module to obtain fused speech;

[0043] Evaluate the fused speech at multiple resolution levels through the multi-scale discriminator, and evaluate the periodic components of the fused speech through the multi-period discriminator;

[0044] When both the level evaluation and the component evaluation meet the corresponding evaluation conditions, determine that the fused speech is enhanced speech.

[0045] Optionally, the training unit of the self-supervised module is specifically as follows:

[0046] The first training unit is configured to obtain sample noisy speech and sample clean speech;

[0047] A second training unit, configured to respectively perform feature extraction on the sample noisy speech and the sample clean speech through the feature encoder to obtain a sample noisy feature and a sample clean feature;

[0048] A third training unit, configured to respectively perform random masking on the sample noisy feature and the sample clean feature through the Transformer module, and respectively obtain a sample noisy context representation and a sample clean context representation according to the masked sample noisy feature and the masked sample clean feature;

[0049] A fourth training unit, configured to train the self-supervised module according to the sample noisy context representation, the sample clean context representation, the sample noisy feature, the sample clean feature, and a total loss function, where the total loss function includes a contrastive loss function and a consistency loss function.

[0050] Optionally, the training unit of the vocoder module is specifically as follows:

[0051] A fifth training unit, configured to perform upsampling on the sample noisy context through the multi-layer one-dimensional transposed convolution to obtain sample intermediate features at multiple time scales, and the time resolution of each sample intermediate feature is aligned with the time resolution of the sample noisy speech;

[0052] A sixth training unit, configured to fuse the sample intermediate features at multiple time scales through the multi-resolution fusion module to obtain sample enhanced speech;

[0053] A seventh training unit, configured to train the generator of the vocoder module according to the sample clean speech, the sample enhanced speech, and the loss function of the generator, where the loss function of the generator includes a first adversarial loss function, a feature matching loss function, a Mel spectrum loss function, an amplitude loss function, and a complex number loss function;

[0054] An eighth training unit, configured to train the discriminator of the vocoder module according to the sample clean speech, the sample enhanced speech, and a second adversarial loss function.

[0055] In a third aspect, the present application discloses a speech enhancement device, where the device includes: a memory and a processor;

[0056] The memory is configured to store a program;

[0057] The processor is configured to execute the program to implement each step of the speech enhancement method as described in the first aspect.

[0058] Compared with the prior art, the present application has the following beneficial effects:

[0059] The embodiments of the present application provide a voice enhancement method, apparatus and device. The method includes: obtaining the voice to be enhanced, where the voice to be enhanced is a noisy voice; converting the voice to be enhanced into a noisy context representation by inputting the voice to be enhanced into the self-supervised module of the voice enhancement model; and converting the noisy context representation into an enhanced voice by inputting the noisy context representation into the vocoder module of the voice enhancement model, where the enhanced voice is a clean voice. Thus, the voice enhancement method provided by the embodiments of the present application can perform voice enhancement on the voice to be enhanced (noisy voice), effectively suppressing the noise and reverberation in the voice to be enhanced, and significantly improving the naturalness and clarity of the voice to be enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0061] Figure 1 It is a flowchart of a voice enhancement method provided by an embodiment of the present application;

[0062] Figure 2A It is a schematic diagram of a voice enhancement model provided by an embodiment of the present application;

[0063] Figure 2B It is a training flowchart of a voice enhancement model provided by an embodiment of the present application;

[0064] Figure 3 It is a schematic diagram of a voice enhancement apparatus provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] As described above, speech enhancement processing is usually performed using a speech enhancement model based on a Deep Neural Network (DNN). Speech enhancement models mainly include two categories: Generative Models and Discriminative Models. Among them, the generative model reconstructs the noisy speech into a speech signal consistent with the probability distribution by learning the probability distribution of clean speech. However, the generative model has the problem of mode collapse during the training process, that is, the generative model may overly focus on certain specific speech features and ignore other speech features, resulting in the generated speech signal lacking naturalness. The discriminative model refers to minimizing the difference between the training speech output by the speech enhancement model and the clean speech corresponding to the noisy speech to suppress noise after inputting the noisy speech into the speech enhancement model. However, the discriminative model relies on a large number of "noisy speech - clean speech" data pairs for training. When faced with unseen noise types, due to the lack of learning ability of the discriminative model for the unseen noise types, it is difficult to effectively improve the clarity of speech enhancement.

[0066] After research, the inventors proposed a speech enhancement method, device, and equipment. In this application, a self-supervised module learns the robust representation of the speech to be enhanced (i.e., obtains the noisy context representation), and through the high-fidelity reconstruction ability of the vocoder module, converts the noisy context representation into enhanced speech, and can obtain the speech to be enhanced (clean speech) with higher naturalness and clarity, effectively solving the problem of speech distortion caused by mode collapse in traditional generative models, as well as the problems of the discriminative model relying on a large amount of labeled data and poor generalization. Further, the self-supervised module uses contrastive loss and consistency loss to learn anti-noise representations from unlabeled data, significantly improving the adaptability of the speech enhancement model to unseen noise; the vocoder module ensures the naturalness and sound quality of the enhanced speech through multi-scale adversarial training and spectral constraints. Therefore, the speech enhancement method provided in the embodiments of this application can still generate speech with high clarity and naturalness in a complex noise environment, and does not require relying on paired data, greatly reducing the training cost.

[0067] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0068] See Figure 1 , which is a flowchart of a speech enhancement method provided in an embodiment of this application. The method includes:

[0069] S101: Obtain the speech to be enhanced, where the speech to be enhanced is noisy speech.

[0070] Noisy speech refers to the original speech that includes background noise, reverberation, or other interferences, and needs to be enhanced through subsequent steps to restore clarity and naturalness.

[0071] S102: Convert the speech to be enhanced into a noisy context representation by inputting the speech to be enhanced into the self-supervised module of the speech enhancement model.

[0072] The self-supervised module includes a Feature Encoder and a Contextual Encoder. Specifically, when inputting the speech to be enhanced into the self-supervised module of the speech enhancement model, first perform feature extraction on the speech to be enhanced through the Feature Encoder to obtain noisy features; then randomly mask the noisy features through the Contextual Encoder, and obtain a noisy context representation (i.e., noise-robust features) based on the masked noisy features.

[0073] It can be understood that the above is an explanatory description of how to convert the speech to be enhanced into a noisy context representation in practical applications by inputting the speech to be enhanced into the self-supervised module of the speech enhancement model. Next, the training method of the self-supervised module will be explained:

[0074] See Figure 2A , which is a schematic diagram of a speech enhancement model provided by an embodiment of the present application. See Figure 2B , which is a training flowchart of a speech enhancement model provided by an embodiment of the present application.

[0075] As can be seen from Figure 2A , the speech enhancement model includes a self-supervised module and a vocoder module connected in series. Among them, the self-supervised module adopts a dual-branch structure (processing sample noisy speech and sample clean speech respectively), and is used to convert the speech waveforms of the sample noisy speech and the sample clean speech into corresponding sample noisy context representations and sample clean context representations respectively.

[0076] Exemplarily, the sample clean speech can come from datasets such as Librispeech, LibriVox, LibriTTS, and VCTK. The sample noisy speech can come from datasets such as Audioset, FreeSound, and FSD50K. For the specific data sources, the present application does not make any limitations.

[0077] Specifically, the feature encoder consists of M one-dimensional convolutional neural networks (e.g., 7 layers), and the stride and kernel size of each layer of convolution gradually decrease (e.g., the stride of the first layer is 5 and the kernel size is 10; the stride of the last layer is 2 and the kernel size is 2). The feature encoder extracts features from the noisy speech sample x noisy and the clean speech sample x clean respectively, to obtain the noisy sample feature z noisy and the clean sample feature z clean , that is, z noisy = f(x noisy ), z clean = f(x clean ). The Transformer module includes N Transformer blocks (e.g., 12 layers, 24 layers). The Transformer module performs random masking on the noisy sample feature z noisy and the clean sample feature z clean respectively (the purpose is to force the speech enhancement model to learn noise robustness), and then inputs them into the Transformer blocks respectively to obtain the noisy sample context representation c noisy and the clean sample context representation c clean , that is, c noisy = g(z noisy ), c clean = g(z clean ).

[0078] Subsequently, based on the noisy sample context representation c noisy , the clean sample context representation c clean , the noisy sample feature z noisy , the clean sample feature z clean and the total loss function L total , the self-supervised module is trained. Among them, the total loss function L total of the self-supervised module consists of a contrastive loss function L con and a consistency loss function L consis . In a specific implementation, the total loss function can be shown as the following formula (1):

[0079] L total = L con + α1·L consis (1)

[0080] where α1 is the first proportional value.

[0081] Specifically, the formula of the contrastive loss function L con can be shown as the following formula (2):

[0082]

[0083] Among them, sim is the cosine similarity, t is the time index, k is the temperature parameter, c noisy is the noisy context representation of the sample, z clean is the clean feature of the sample, z noisy is the noise feature of the sample. It can be understood that the contrast loss function L con The role is to pull in the sample noisy context representation c noisy and sample clean features z clean , and push the sample noisy context representation c noisy and sample noise feature z noisy The similarity enables the speech enhancement model to learn to extract features from the sample noisy speech that are consistent with the sample clean speech.

[0084] Consistency loss function L consis The formula can be shown as follows:

[0085] L consis =‖c noisy -c clean ||2 (3)

[0086] Among them, c noisy is the noisy context representation of the sample, c clean is the clean context representation of the sample. It can be understood that the consistency loss function L consis The role of is to constrain the distribution consistency of sample noisy speech and sample clean speech in the representation space, thereby improving the denoising effect.

[0087] S103: The noisy context representation is input into a vocoder module of a speech enhancement model to convert the noisy context representation into enhanced speech, and the enhanced speech is clean speech.

[0088] The vocoder module includes a generator, which includes multiple layers of one-dimensional transposed convolution and a multi-resolution fusion (MRF) module. Specifically, the noisy context representation is input into the vocoder module of the speech enhancement model, and the noisy context representation is first upsampled through multiple layers of one-dimensional transposed convolution to obtain intermediate features of multiple time scales, where the time resolution of each intermediate feature is aligned with the time resolution of the speech to be enhanced; then, the intermediate features of multiple time scales are fused through the multi-resolution fusion module to obtain enhanced speech.

[0089] The vocoder module further includes a discriminator, which includes a Multi-Scale Discriminator (MSD) and a Multi-Period Discriminator (MPD). Then, after the intermediate features of multiple time scales are fused through the multi-resolution fusion module (the fused speech is obtained at this time), the fused speech can be evaluated at multiple resolution levels through the multi-scale discriminator, and the periodic components of the fused speech can be evaluated through the multi-period discriminator. When both the level evaluation and the component evaluation meet the corresponding evaluation conditions, the fused speech is determined to be enhanced speech.

[0090] It can be understood that the above is an explanation of how to convert the noisy context representation into enhanced speech by inputting the noisy context representation into the vocoder module of the speech enhancement model in practical applications. Next, the training method of the vocoder module will be explained:

[0091] From Figure 2A it can be seen that the vocoder module includes a generator and two discriminators, which are used to reconstruct the sample noisy context representation c noisy to obtain the sample enhanced speech

[0092] Specifically, the generator includes multiple layers of one-dimensional transposed convolutions, which are used to upsample the sample noisy context representation c noisy by k u times (the upsample factor of each layer increases gradually) so that the time resolution of the sample noisy context representation c noisy is aligned with the time resolution of the original sample noisy speech. It can be understood that the output of the multiple layers of one-dimensional transposed convolutions is the preliminarily upsampled sample intermediate features (the waveform has not been formed yet). The generator also includes a multi-resolution fusion module, which is used to fuse the sample intermediate features of different time scales, thereby outputting the sample enhanced speech

[0093] The discriminator includes a multi-scale discriminator and a multi-period discriminator. The input of the discriminator is the sample clean speech x clean and the sample enhanced speech The output of the discriminator is a probability value (0-1), indicating the confidence of the sample enhanced speech being the sample clean speech x clean Among them, the multi-scale discriminator is used to perform multi-level downsampling (such as 2×, 4×, 8×) on the sample enhanced speech so as to evaluate the sample enhanced speech at multiple resolution levels The authenticity is ensured to prevent the generator from generating "locally reasonable but globally incoherent" speech (such as discontinuous waveforms). The multi-period discriminator is used to segment the sample-enhanced speech (such as cycle lengths = 2, 3, 5, 7), so as to evaluate the authenticity of the sample-enhanced speech from multiple cycle levels and avoid the generator from generating speech with "reasonable spectrum but chaotic phase" (such as mechanical voice or phase distortion). It can be seen that the multi-scale discriminator and the multi-period discriminator respectively constrain the generator from two dimensions of "time-domain multi-scale" and "frequency-domain periodicity", and can cover all physical characteristics of speech.

[0094] In a specific implementation, the loss function L of the generator G consists of the first adversarial loss function L adv (G; D), the feature matching loss function L fm , the Mel spectrum loss function L mel , the amplitude loss function L mag and the complex number loss function L com . The loss function L of the generator G can be shown as the following formula (4):

[0095] L G = L adv (G; D) + β1L fm + β2L mel + β3L mag + β4L com (4)

[0096] where β1, β2, β3, and β4 are the first hyperparameter, the second hyperparameter, the third hyperparameter, and the fourth hyperparameter respectively.

[0097] Specifically, the first adversarial loss function is used to solve the over-smoothing problem, enhance the global authenticity of the sample-enhanced speech , and retain the high-frequency details (such as plosives) of the sample-enhanced speech . The first adversarial loss function L adv (G; D) can be shown as the following formula (5):

[0098] L adv (G; D) = E[(D(G(c noisy )) - 1) 2 (5)

[0099] The feature matching loss function L fm is used to constrain the sample clean speech x clean and the sample-enhanced speech The distance in the discriminator feature space (i.e., minimizing feature differences) to prevent the generator from only optimizing the adversarial loss and ignoring the rationality of the speech structure (preventing mode collapse). The feature matching loss function L fm can be shown as in Equation (6) below:

[0100]

[0101] where T represents the number of layers of the discriminator, D i and N i represent the features and the number of features of the i-th layer of the discriminator, respectively.

[0102] The Mel-spectrum loss function L mel is used to minimize the L1 distance between the clean speech sample x clean and the enhanced speech sample in the Mel frequency domain, which can ensure that the auditory perception quality (such as timbre, clarity) of the enhanced speech sample is consistent with that of the clean speech sample x clean The Mel-spectrum loss function L mel can be shown as in Equation (7) below:

[0103]

[0104] where is the mapping function from the speech signal to the Mel spectrum.

[0105] The magnitude loss function L mag is used to minimize the difference in the short-time Fourier transform (STFT) magnitude spectra between the clean speech sample x clean and the enhanced speech sample , thereby avoiding spectral distortion by constraining the frequency-domain energy distribution. The magnitude loss function L mag can be shown as in Equation (8) below:

[0106]

[0107] where STFT(·) is the magnitude spectrum of the short-time Fourier transform.

[0108] The complex loss function L com is used to simultaneously minimize the differences in the real part (magnitude) and the imaginary part (phase) of the STFT, solve the problem of inaccurate phase estimation in traditional methods, and improve the waveform reconstruction accuracy. The complex loss function L com can be shown as in Equation (9) below:

[0109]

[0110] where Re(·) and Im(·) are the real part and the imaginary part of the STFT (i.e., phase information), respectively.

[0111] In a specific implementation, the loss function L of the discriminator D is the second adversarial loss function and can be as shown in the following formula (10):

[0112]

[0113] Therefore, after obtaining the sample clean speech and the sample enhanced speech, the generator of the vocoder module can be trained according to the sample clean speech, the sample enhanced speech, and the loss function of the generator; the discriminator of the vocoder module can be trained according to the sample clean speech, the sample enhanced speech, and the second adversarial loss function.

[0114] In summary, the present application discloses a speech enhancement method. The present application learns the robust representation of the speech to be enhanced (i.e., obtains the noisy context representation) through the self-supervised module, and converts the noisy context representation into enhanced speech through the high-fidelity reconstruction ability of the vocoder module, and can obtain the speech to be enhanced (clean speech) with higher naturalness and clarity, effectively solving the problem of speech distortion caused by mode collapse in traditional generative models, as well as the problems of discriminative models relying on a large amount of labeled data and poor generalization. Further, the self-supervised module uses contrastive loss and consistency loss to learn noise-resistant representations from unlabeled data, significantly improving the adaptability of the speech enhancement model to unseen noise; the vocoder module ensures the naturalness and sound quality of the enhanced speech through multi-scale adversarial training and spectral constraints. Therefore, the speech enhancement method provided by the embodiments of the present application can still generate speech with high clarity and high naturalness in a complex noise environment, and does not rely on paired data, greatly reducing the training cost.

[0115] See Figure 3 , which is a schematic diagram of a speech enhancement device provided by an embodiment of the present application. The speech enhancement device 300 includes: a speech acquisition module 301, a first conversion module 302, and a second conversion module 303;

[0116] The speech acquisition module 301 is configured to acquire the speech to be enhanced, and the speech to be enhanced is noisy speech;

[0117] The first conversion module 302 is configured to convert the speech to be enhanced into a noisy context representation by inputting the speech to be enhanced into the self-supervised module of the speech enhancement model;

[0118] The second conversion module 303 is configured to convert the noisy context representation into enhanced speech by inputting the noisy context representation into the vocoder module of the speech enhancement model, and the enhanced speech is clean speech.

[0119] Optionally, the self-supervised module includes a feature encoder and a Transformer module; the first conversion module includes: a first conversion sub-module and a second conversion sub-module:

[0120] The first conversion sub-module is used to extract features from the speech to be enhanced through the feature encoder to obtain noisy features;

[0121] The second conversion sub-module is used to randomly mask the noisy features through the Transformer module and obtain a noisy context representation according to the masked noisy features.

[0122] Optionally, the vocoder module includes a generator, and the generator includes multi-layer 1D transposed convolution and a multi-resolution fusion module; the second conversion module 303 includes: a third conversion sub-module and a fourth conversion sub-module;

[0123] The third conversion sub-module is used to upsample the noisy context through multi-layer 1D transposed convolution to obtain intermediate features at multiple time scales, and the time resolution of each intermediate feature is aligned with the time resolution of the speech to be enhanced;

[0124] The fourth conversion sub-module is used to fuse the intermediate features at multiple time scales through the multi-resolution fusion module to obtain enhanced speech.

[0125] Optionally, the vocoder module further includes a discriminator, and the discriminator includes a multi-scale discriminator and a multi-period discriminator; the fourth conversion sub-module is specifically used for:

[0126] Fusing the intermediate features at multiple time scales through the multi-resolution fusion module to obtain fused speech;

[0127] Evaluating the fused speech at multiple resolution levels through the multi-scale discriminator, and evaluating the periodic components of the fused speech through the multi-period discriminator;

[0128] When both the level evaluation and the component evaluation meet the corresponding evaluation conditions, determining the fused speech as enhanced speech.

[0129] Optionally, the training unit of the self-supervised module is specifically as follows:

[0130] The first training unit is used to obtain sample noisy speech and sample clean speech;

[0131] The second training unit is used to extract features from the sample noisy speech and the sample clean speech respectively through the feature encoder to obtain sample noisy features and sample clean features;

[0132] The third training unit is used to randomly mask the sample noisy features and the sample clean features through a Transformer module, and respectively obtain a sample noisy context representation and a sample clean context representation according to the masked sample noisy features and the masked sample clean features;

[0133] The fourth training unit is used to train the self-supervised module according to the sample noisy context representation, the sample clean context representation, the sample noisy features, the sample clean features, and the total loss function, where the total loss function includes a contrastive loss function and a consistency loss function.

[0134] Optionally, the training unit of the vocoder module is specifically as follows:

[0135] The fifth training unit is used to upsample the sample noisy context through a multi-layer one-dimensional transposed convolution to obtain sample intermediate features at multiple time scales, and the time resolution of each sample intermediate feature is aligned with the time resolution of the sample noisy speech;

[0136] The sixth training unit is used to fuse the sample intermediate features at multiple time scales through a multi-resolution fusion module to obtain sample enhanced speech;

[0137] The seventh training unit is used to train the generator of the vocoder module according to the sample clean speech, the sample enhanced speech, and the loss function of the generator, where the loss function of the generator includes a first adversarial loss function, a feature matching loss function, a Mel spectrum loss function, an amplitude loss function, and a complex number loss function;

[0138] The eighth training unit is used to train the discriminator of the vocoder module according to the sample clean speech, the sample enhanced speech, and a second adversarial loss function.

[0139] In summary, the present application discloses a speech enhancement device. The present application learns a robust representation of the speech to be enhanced (i.e., obtains a noisy context representation) through a self-supervised module, and converts the noisy context representation into enhanced speech through the high-fidelity reconstruction ability of the vocoder module, and can obtain the speech to be enhanced (clean speech) with higher naturalness and clarity, effectively solving the problem of speech distortion caused by mode collapse in traditional generative models, as well as the problems of discriminative models relying on a large amount of labeled data and poor generalization. Further, the self-supervised module uses contrastive loss and consistency loss to learn noise-resistant representations from unlabeled data, significantly improving the adaptability of the speech enhancement model to unseen noise; the vocoder module ensures the naturalness and sound quality of the enhanced speech through multi-scale adversarial training and spectral constraints. Therefore, the speech enhancement device provided by the embodiments of the present application can still generate speech with high clarity and high naturalness in a complex noise environment, and does not need to rely on paired data, greatly reducing the training cost.

[0140] The embodiments of the present application also provide corresponding voice enhancement devices for implementing the voice enhancement method provided by the embodiments of the present application.

[0141] Among them, the voice enhancement device includes a memory and a processor. The memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device executes a voice enhancement method according to any embodiment of the present application.

[0142] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.

[0143] It should be noted that the embodiments in this specification are all described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components referred to as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0144] As described above, this is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A speech enhancement method, characterized in that: The method comprises: Acquire a speech to be enhanced, where the speech to be enhanced is a noisy speech; The speech to be enhanced is input into a self-supervision module of a speech enhancement model, thereby converting the speech to be enhanced into a noisy context representation; The noisy context representation is input into a vocoder module of the speech enhancement model to convert the noisy context representation into enhanced speech, where the enhanced speech is clean speech.

2. The method according to claim 1, characterized in that: The self-supervision module includes a feature encoder and a Transformer module; the step of converting the speech to be enhanced into a noisy context representation includes: Extracting features of the speech to be enhanced by the feature encoder to obtain noisy features; The noisy features are randomly masked through the Transformer module, and a noisy context representation is obtained according to the masked noisy features.

3. The method according to claim 2, characterized in that The vocoder module includes a generator, wherein the generator includes a multi-layer one-dimensional transposed convolution and a multi-resolution fusion module; the converting the noisy context representation into enhanced speech includes: Upsampling the noisy context by the multi-layer one-dimensional transposed convolution to obtain intermediate features of multiple time scales, wherein the time resolution of each intermediate feature is aligned with the time resolution of the speech to be enhanced; The intermediate features of the multiple time scales are fused through the multi-resolution fusion module to obtain enhanced speech.

4. The method according to claim 3, characterized in that The vocoder module further includes a discriminator, and the discriminator includes a multi-scale discriminator and a multi-period discriminator; the multi-resolution fusion module is used to fuse the intermediate features of the multiple time scales to obtain enhanced speech, including: The intermediate features of the multiple time scales are fused through the multi-resolution fusion module to obtain fused speech; Performing level evaluation on the fused speech from multiple resolution levels by the multi-scale discriminator, and performing component evaluation on the periodic components of the fused speech by the multi-periodic discriminator; When both the level evaluation and the component evaluation satisfy corresponding evaluation conditions, the fused speech is determined to be enhanced speech.

5. The method according to claim 4, characterized in that The training method of the self-supervision module is as follows: Get sample noisy speech and sample clean speech; The feature encoder is used to extract features of the sample noisy speech and the sample clean speech respectively, so as to obtain sample noisy features and sample clean features; The noisy features of the samples and the clean features of the samples are randomly masked by the Transformer module, and a noisy context representation of the samples and a clean context representation of the samples are obtained according to the masked noisy features of the samples and the masked clean features of the samples; The self-supervision module is trained according to the sample noisy context representation, the sample clean context representation, the sample noisy features, the sample clean features and the total loss function, wherein the total loss function includes a contrast loss function and a consistency loss function.

6. The method according to claim 5, characterized in that The training method of the vocoder module is as follows: Upsampling the sample noisy context by the multi-layer one-dimensional transposed convolution to obtain sample intermediate features at multiple time scales, wherein the time resolution of each sample intermediate feature is aligned with the time resolution of the sample noisy speech; By means of the multi-resolution fusion module, the intermediate features of the samples at the multiple time scales are fused to obtain sample enhanced speech; Training the generator of the vocoder module according to the sample clean speech, the sample enhanced speech and the loss function of the generator, wherein the loss function of the generator includes a first adversarial loss function, a feature matching loss function, a Mel spectrum loss function, an amplitude loss function and a complex loss function; The discriminator of the vocoder module is trained according to the sample clean speech, the sample enhanced speech and the second adversarial loss function.

7. A speech enhancement device, characterized in that: The device comprises: a speech acquisition module, a first conversion module and a second conversion module; The speech acquisition module is used to acquire the speech to be enhanced, where the speech to be enhanced is noisy speech; The first conversion module is used to convert the speech to be enhanced into a noisy context representation by inputting the speech to be enhanced into a self-supervision module of a speech enhancement model; The second conversion module is used to convert the noisy context representation into enhanced speech by inputting the noisy context representation into the vocoder module of the speech enhancement model, and the enhanced speech is clean speech.

8. The device according to claim 7, characterized in that The self-supervision module includes a feature encoder and a Transformer module; The first conversion module includes: a first conversion submodule and a second conversion submodule: The first conversion submodule is used to extract features of the speech to be enhanced through the feature encoder to obtain noisy features; The second conversion submodule is used to randomly mask the noisy features through the Transformer module, and obtain the noisy context representation according to the masked noisy features.

9. The device according to claim 8, characterized in that The vocoder module includes a generator, and the generator includes a multi-layer one-dimensional transposed convolution and a multi-resolution fusion module; the second conversion module includes: a third conversion submodule and a fourth conversion submodule; The third conversion submodule is used to upsample the noisy context through the multi-layer one-dimensional transposed convolution to obtain intermediate features of multiple time scales, and the time resolution of each of the intermediate features is aligned with the time resolution of the speech to be enhanced; The fourth conversion submodule is used to fuse the intermediate features of the multiple time scales through the multi-resolution fusion module to obtain enhanced speech.

10. A speech enhancement device, characterized in that: The device comprises: a memory and a processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the speech enhancement method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Speech enhancement method of improved multi-resolution residual U-shaped network

    CN113707164A

  • Speech enhancement method and device, equipment and storage medium

    CN113707168A

  • Speech enhancement method fusing Transform and U-net network

    CN114141238A

  • Speech recognition model training method, speech recognition method and electronic equipment

    CN114582330A

  • Improved speech enhancement method and system, medium, equipment and terminal

    CN116343807A