Iterative self-supervised training methods, systems, and electronic devices for speech enhancement models
By using an iterative self-supervised training method to construct training data pairs using pure noise samples, the problem of poor speech enhancement effect under noise interference in existing technologies is solved, and efficient speech enhancement effect in real-world scenarios is achieved.
Patent Information
- Application Number
- CN202211696112.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2042-12-28
AI Technical Summary
Existing speech enhancement techniques perform poorly in real-world scenarios with noise interference. Methods based on statistical signal processing are inaccurate in noise estimation, while methods based on deep neural networks rely on large amounts of clean speech training data, which are difficult to obtain. Self-supervised training has poor generalization ability under certain noise types.
An iterative self-supervised training method for speech enhancement models is adopted. The noisy speech output by the speech enhancement model is preprocessed using pure noise samples to construct training data pairs. The neural network is then iteratively trained through a loss function until the loss function converges, thus avoiding reliance on clean speech data.
Without relying on clean speech data, iterative training improved speech enhancement, thereby enhancing the accuracy and robustness of speech recognition.
Smart Images

Figure CN115985335B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent speech, and in particular to an iterative self-supervised training method, system, and electronic device for speech enhancement models. Background Technology
[0002] With the development of intelligent voice technology, automatic speech recognition, speaker recognition, and other voice technologies are being used more and more in daily life. However, the performance of these technologies in real-world scenarios is often not as good as that in the ideal environment of a laboratory. A major factor contributing to this gap is environmental noise in the real world, which can significantly reduce the accuracy of speech recognition. Speech enhancement techniques can be used to remove interfering noise from speech, thereby improving speech recognition performance.
[0003] In speech enhancement, traditional methods based on statistical signal processing, processing methods based on deep neural networks (model training, neural network combination), or self-supervised training methods are commonly used.
[0004] In the process of realizing this invention, the inventors discovered at least the following problems in the related technology:
[0005] Speech enhancement methods based on statistical signal processing rely heavily on noise estimation algorithms. While these algorithms can accurately estimate steady-state noise, overestimation of the noise energy spectrum can lead to speech distortion, and underestimation of noise energy can result in residual noise. They also perform poorly in estimating non-steady-state noise. Methods based on deep neural networks all suffer from the problem of relying on large amounts of clean speech training data, which are difficult to obtain in practical applications. Directly using self-supervised training to train neural networks results in poor generalization ability for certain noise types, failing to achieve satisfactory results. Summary of the Invention
[0006] To address at least the problems in existing technologies where speech enhancement requires a large amount of clean speech training data and self-supervised training struggles to achieve satisfactory results, this invention provides, in a first aspect, an iterative self-supervised training method for speech enhancement models, comprising:
[0007] The noise-reduced noisy speech output by the speech enhancement model in the (k-1)th stage is preprocessed using pure noise samples to obtain training data pairs in the kth stage, wherein k>1, and the training data pairs include: noisy speech and noisy speech.
[0008] The training data pair of the k-th stage is input into the speech enhancement model, and the noisy speech in the training data pair of the k-th stage is denoised to obtain the noisy speech output by the k-th stage.
[0009] The speech enhancement model is trained using the loss function of the noisy speech output in the k-th stage and the noisy speech in the training data pair in the k-th stage. If the loss function does not converge, the speech enhancement model is iteratively trained in the next stage using the pure noise and the noisy speech output in the k-th stage until the loss function converges.
[0010] Secondly, embodiments of the present invention provide an iterative self-supervised training system for a speech enhancement model, comprising:
[0011] The preprocessing module is used to preprocess the noisy speech output by the speech enhancement model in the (k-1)th stage using pure noise samples to obtain training data pairs in the kth stage, wherein k>1, and the training data pairs include: noisy speech and noisy speech.
[0012] The speech enhancement module is used to input the training data pair of the k-th stage into the speech enhancement model, perform speech denoising on the noisy speech in the training data pair of the k-th stage, and obtain the denoised noisy speech output by the k-th stage.
[0013] The iterative training module is used to perform self-supervised learning training of the speech enhancement model in the k-th stage based on the loss function of the noisy speech output in the k-th stage and the noisy speech in the training data pair in the k-th stage. If the loss function does not converge, the speech enhancement model is iteratively trained in the next stage using the pure noise and the noisy speech output in the k-th stage until the loss function converges.
[0014] Thirdly, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the iterative self-supervised training method for a speech enhancement model according to any embodiment of the present invention.
[0015] Fourthly, embodiments of the present invention provide a storage medium storing a computer program thereon, characterized in that, when the program is executed by a processor, it implements the steps of an iterative self-supervised training method for a speech enhancement model according to any embodiment of the present invention.
[0016] The beneficial effects of this invention are as follows: by using noise unrelated to noisy speech as the noise source, entirely new training data pairs are constructed, and an iterative method is applied to self-supervised training. This allows for effective integration with the neural network structure, rather than simply using impure speech as the target for training. Superior speech enhancement results can be achieved without using pure training. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of an iterative self-supervised training method for a speech enhancement model provided in an embodiment of the present invention;
[0019] Figure 2 This is a framework diagram of an iterative training system for an iterative self-supervised training method for a speech enhancement model, provided in an embodiment of the present invention.
[0020] Figure 3 This is a model architecture diagram of an iterative self-supervised training method for a speech enhancement model provided in an embodiment of the present invention;
[0021] Figure 4 This is a diagram of a self-supervised speech enhancement architecture, which is an iterative self-supervised training method for a speech enhancement model provided in an embodiment of the present invention.
[0022] Figure 5 This is a schematic diagram of the training target data for an iterative self-supervised training method for a speech enhancement model provided in an embodiment of the present invention;
[0023] Figure 6 This is a schematic diagram of the sampling parameters p of an iterative self-supervised training method for a speech enhancement model provided in an embodiment of the present invention;
[0024] Figure 7 This is a schematic diagram illustrating the evaluation results of an iterative self-supervised training method for a speech enhancement model provided in an embodiment of the present invention on the VoiceBank+DEMAND dataset;
[0025] Figure 8 This is a schematic diagram of the structure of an iterative self-supervised training system for a speech enhancement model according to an embodiment of the present invention;
[0026] Figure 9This is a schematic diagram of an embodiment of an electronic device for iterative self-supervised training of a speech enhancement model, as provided in one embodiment of the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] like Figure 1 The diagram shows a flowchart of an iterative self-supervised training method for a speech enhancement model according to an embodiment of the present invention, which includes the following steps:
[0029] S11: Preprocess the noisy speech output by the speech enhancement model in stage k-1 using pure noise samples to obtain training data pairs in stage k, wherein k>1, and the training data pairs include: noisy speech and noisy speech.
[0030] S12: Input the training data pair of the k-th stage into the speech enhancement model, perform speech denoising on the noisy speech in the training data pair of the k-th stage, and obtain the denoised noisy speech output by the k-th stage.
[0031] S13: Based on the loss function of the noisy speech output in the k-th stage and the noisy speech in the training data pair in the k-th stage, perform self-supervised learning training for the speech enhancement model in the k-th stage. If the loss function does not converge, use the pure noise and the noisy speech output in the k-th stage to iteratively perform self-supervised learning training for the speech enhancement model in the next stage until the loss function converges.
[0032] In this embodiment, the most basic model for mono speech enhancement can be simply represented as follows:
[0033] y = X + n
[0034] in and Let represent noisy speech with added noise, clean speech, and pure noise used for adding noise, respectively. Consider using the STFT (short-time Fourier transform) domain, and transforming y, x, and n through Fourier transforms to obtain... and W is the frame window size, T is the number of frames, and F is the interval between samples in the frequency domain. In the most basic model, the goal is to find a denoising model F. in, This represents the estimated STFT of clean speech. Considering the difficulty in obtaining clean speech in existing techniques, this method utilizes noisy speech to train a speech enhancement model for denoising.
[0035] For step S11, the iterative self-supervised training structure of the speech enhancement model in this method is as follows: Figure 2 As shown, y represents the noisy speech input in this round, which is obtained from the output of the speech enhancement model in the previous round; n' represents other noise in the pure noise sample that is unrelated to the noise of the noisy speech y; F k-1 This represents the speech enhancement model at stage k-1; This means inputting y into F. k-1 The resulting noisy speech, F k Let represent the speech enhancement model at stage k.
[0036] As one implementation method, when k=1, the original noisy speech is preprocessed using pure noise samples to obtain training data pairs in the first stage.
[0037] Select a portion of pure noise that is not related to the noisy speech noise from the pure noise sample, and add the pure noise to the noisy speech according to a preset signal-to-noise ratio to obtain noisy speech.
[0038] In this embodiment, when k=1, which is the initial stage of this method, the original noisy speech y and various pure noise samples are prepared. The preprocessing process is divided into two parts. The first part is dataset preparation, where noise unrelated to the original noisy speech y is selected from the pure noise samples and added to the original noisy speech y to generate noisy speech. For example, if the noise in the original noisy speech is traffic noise, other types of noise that are not traffic noise are selected from the pure noise samples. The noise in the noisy speech is stronger than the noise in the noisy speech y. The noisy speech and the noisy speech are used as training data pairs {(y)}. i +n′ i y i The second part involves windowing and framing the speech signals in the training data pairs. For example, for speech with a sampling rate of 16kHz, Hamming windows are used for framing, with a frame length of 400 sampling points (i.e., 25ms) and a frame shift of 100 sampling points (i.e., 6.25ms). In the first training phase, the constructed training data pairs are used to train the speech enhancement model.
[0039] When k>1, the noisy speech y output by the speech enhancement model in the k-1 stage is preprocessed using pure noise samples. The specific preprocessing process is the same as when k=1, and will not be repeated here.
[0040] For step S12, the training data pair for the k-th stage determined in step S11 is input into the speech enhancement model. In the k-th stage, when k=0, the training data pair {(y i +n′ i y i The input is fed into the speech enhancement model for training. For stages >1, the target is selected (or randomly sampled) from the noisy speech output in stage k-1 after denoising. and noisy speech y. Where F k-1 This is the model trained in the (k-1)th stage. Specifically, we define two sets of data pairs:
[0041]
[0042] Then the training data pairs for the k-th training iteration are:
[0043]
[0044] Where p∈[0,1] is the probability Ω of sampling from the set. k-1 (1-p) is the probability of sampling from set y. That is, given model F... k-1 During training in the (k-1)th stage, the dataset Ω k-1 Based on F k-1 Then model F k It uses Ω k This is an iterative training process. When p = 1.0, the process degenerates into the existing iterative refinement method for image denoising, which only uses the estimated signal as the training target, thus obtaining the denoised noisy speech output in the k-th stage.
[0045] As one implementation, the speech enhancement model is an encoder-decoder structure, wherein:
[0046] The encoder includes multiple convolutional blocks and deconvolutional blocks corresponding to the multiple convolutional blocks, wherein the multiple convolutional blocks and the multiple deconvolutional blocks are connected by convolutional block attention modules, and each convolutional block attention module is composed of a channel attention block and a frequency attention block;
[0047] The encoder and the decoder are connected by multiple dual-path blocks, wherein the dual-path blocks include a first long short-term memory layer modeled in the frequency domain and a second long short-term memory layer modeled in the time domain.
[0048] In this embodiment, the method adjusts the speech enhancement model, such as... Figure 3 The diagram shows the structure of the speech enhancement model proposed in this method, which includes M Conv2D convolutional blocks, M corresponding ConvTrans2D deconvolutional blocks, and N dual-path blocks. Each Conv2D / ConvTrans2D block consists of a Conv2D / ConvTrans2D layer, a batch normalization layer, and a PReLu (Parametric ReLU, activation function). Each dual-path block comprises two LSTM (Long Short-Term Memory) layers. In the two LSTMs, the frequency dimension is modeled in the first LSTM, and the time dimension is modeled in the second LSTM. Each encoder and decoder block is connected to a CBAM (convolutional block attention module), which consists of a channel attention block and a frequency attention block.
[0049] For step S13, s and This represents the speech signal corresponding to the noisy speech in the training data pair of stage k and the noisy speech output in stage k.
[0050] Based on s and Let's evaluate the following loss functions used to train a neural network. For example, MSE (mean square error):
[0051]
[0052] Speech signal evaluation standards using SNR (SIGNAL-NOISE RATIO) and SI-SNR (Scale-invariant-SNR).
[0053]
[0054] in It is s and The projection, This is the estimated noise. When there is no projection, SI-SNR is equivalent to SNR.
[0055] MS-STFT (multi-scale-STFT): MS-STFT loss in the field of neural vocoders is very effective for speech enhancement. The i-th level STFT loss is defined as:
[0056]
[0057] Among them ||·|||F And ||·||1 are the Frobenius and L1 norms. The final MS-STFT loss is a combination of the STFT loss and L1 loss at each scale of the time-domain signal:
[0058]
[0059] By determining L MS-STFT The loss is backpropagated to update the neural network model parameters. Training is iteratively performed on the training dataset until the loss function converges and training stops. In summary, the overall block diagram of this method is as follows: Figure 4 As shown.
[0060] This implementation demonstrates that using noise unrelated to the noisy speech as the noise source to construct entirely new training data pairs, and applying iterative methods to self-supervised training, effectively combines with the neural network structure rather than simply using impure speech as the target for training. This achieves superior speech enhancement results without using pure training data.
[0061] The speech enhancement method described in this paper is illustrated through specific experiments using the VoiceBank+DEMAND dataset. This dataset has two training subsets; our method used a subset of 28 speakers, totaling 11,572 dialogues. The test set consists of 2 invisible speakers, containing 824 voice recordings. Furthermore, 1000 dialogues were randomly selected from the training subset as the validation set, with the remaining dialogues used as the training set. The DEMAND2 corpus contains 18 types of noise, each with a fixed length of 5 minutes.
[0062] The dataset was sampled and mixed at a sampling rate of 48 kHz, reducing the sampling of all speech to 16 kHz. A Hamming window of 400 samples (25 ms) was used to segment the speech signal, with a frame shift length of 100 samples (6.25 ms) and an FFT length of 512. The number of frames input to the neural network was 157 (1 s). Therefore, the window size W, frame number T, and frequency F were 400, 157, and 257, respectively.
[0063] To obtain a noisy version of the estimated speech during training, noise types are randomly selected from a DEMAND with random starting positions. Noisier speech is mixed by adding sampled speech and sampled noise with SNR levels ranging from -5dB to 20dB.
[0064] For iterative self-supervised training of the speech enhancement model, the encoder block size is M=3. The channel numbers of the convolutional layers in the encoder are 32, 64, and 128. The kernel size, stride, and padding are (5, 2), (2, 1), and (1, 1), respectively. Therefore, the channel and frequency dimensions of the encoder output are 128 and 32, respectively. The block size of the dual-path module is N=2. The decoder block size is the same as the encoder, and the other configurations of each deconvolutional layer are set as mirror images of the convolutional layers. The neural network has approximately 1.5 million model parameters.
[0065] The neural network is trained using Xavier initialization and the AdamW optimizer. The initial learning rate is 0.001, and a scheduler named "ReduceLROnPlateau" is used. The training process stops when the stopping policy is triggered or when the maximum of 300 training epochs is reached.
[0066] This method is compared with the following representative speech enhancement methods. Traditional MMSE (minimum mean-square error) based speech enhancement methods (e.g., OMLSA) and DNN (Deep Neural Networks) based supervised methods (e.g., DCCRN) are selected for comparison. In addition, two related self-supervised learning methods, Noise2Noise and Noisy2Target, are also compared.
[0067] To evaluate the quality of denoised speech, this method uses SISNR, PESQ-WB (Perceptual Evaluation of Speech Quality), and eSTOI (extended Short-Time Objective Intelligence) as evaluation metrics.
[0068] This method first conducts testing and research on the training objective and sampling parameters. For example... Figure 5 The evaluation results for four training objectives are shown, with the sampling parameter p = 0.5. MagNorm indicates whether the amplitude of the estimated signal is normalized to the input signal. Only the SI-SNR training objective converges successfully when MagNorm is not applied. When MagNorm is applied, SI-SNR still achieves the best performance. This is because training data pairs are generated online during model training, resulting in a random root mean square (RMS) level for the mixed speech. For denoising models, it is impossible to learn random RMS values. Of the four training objectives, only SI-SNR, which projects the estimated signal onto the reference signal, is insensitive to the RMS level of the signal. Therefore, SI-SNR is used as the training objective in the following experiments.
[0069] like Figure 6 The results are shown for different values of the sampling parameter p ranging from 0.0 to 1.0, with an interval of 0.1. When p = 0.0, the training strategy is equivalent to Noisy2Target, which uses only the original noisy speech as the training target. When p = 1.0, the training strategy is similar to the iterative thinning method for image denoising, which uses only the estimated speech as the training target. As can be seen from the figure, performance improves with increasing p. When p = 0.9, the PESQ-WB, eSTOI, and SI-SNR scores reach their maximum values of 2.512, 0.839, and 19.22 dB, respectively. However, when p increases to 1.0, the scores decrease to 2.262, 0.825, and 18.831 dB, respectively. Therefore, p = 0.9 was set in the subsequent comparative experiments.
[0070] Then, comparative experiments were conducted. Audio samples and supplements are available online. Figure 7 The evaluation results of the comparative methods on the VoiceBank+DEMAND test set are presented, with the highest score for each evaluation metric highlighted in bold. OMLSA does not require training and denoises the test set. The DCCRN and Noise2Noise denoising models are trained using the original settings. This method uses only the VoiceBank+DEMAND dataset for the Noisy2Target training strategy. The PESQ-WB, eSTOI, and SI-SNR scores on the test set are shown in the figure, p=0.0 (2.033, 0.791, and 12.98 dB, respectively).
[0071] like Figure 7 Experimental results show that: 1. The supervised learning method DCCRN achieves the highest PESQ-WB and eSTOI scores; 2. The Noise2Noise method has a better SI-SNR score than DCCRN; 3. The Noisy2Target method has the worst performance across all metrics, except compared to the traditional MMSE-based method OMLSA; 4. Our proposed method achieves the best performance among self-supervised methods across all metrics and obtains the highest SI-SNR score among all compared methods. From these results, it can be concluded that the proposed self-supervised method achieves competitive performance compared to state-of-the-art supervised speech enhancement methods.
[0072] In summary, this method is a self-supervised learning approach for speech enhancement that does not require real, clean speech. Specifically, it iteratively trains the model by progressively constructing noisy training data. Extensive experiments were conducted using four commonly used loss functions in speech enhancement, and only the SI-SNR loss function successfully converged. Experimental results demonstrate that the proposed self-supervised method achieves competitive performance compared to state-of-the-art supervised speech enhancement methods.
[0073] like Figure 8 The diagram shown is a schematic diagram of an iterative self-supervised training system for a speech enhancement model according to an embodiment of the present invention. The system can execute the iterative self-supervised training method for the speech enhancement model described in any of the above embodiments and is configured in a terminal.
[0074] This embodiment provides an iterative self-supervised training system 10 for a speech enhancement model, which includes a preprocessing module 11, a speech enhancement module 12, and an iterative training module 13.
[0075] The preprocessing module 11 is used to preprocess the noisy speech output by the speech enhancement model in the (k-1)th stage using pure noise samples to obtain the training data pair in the kth stage, where k > 1, and the training data pair includes noisy speech and noisy speech. The speech enhancement module 12 is used to input the training data pair in the kth stage into the speech enhancement model and perform speech denoising on the noisy speech in the training data pair in the kth stage to obtain the noisy speech output in the kth stage. The iterative training module 13 is used to perform self-supervised learning training of the speech enhancement model in the kth stage based on the loss function of the noisy speech output in the kth stage and the noisy speech in the training data pair in the kth stage. If the loss function does not converge, the speech enhancement model is iteratively trained in the next stage using pure noise and the noisy speech output in the kth stage until the loss function converges.
[0076] This invention also provides a non-volatile computer storage medium storing computer-executable instructions that can execute the iterative self-supervised training method for the speech enhancement model in any of the above method embodiments.
[0077] In one embodiment, the non-volatile computer storage medium of the present invention stores computer-executable instructions, which are configured as follows:
[0078] The noise-reduced noisy speech output by the speech enhancement model in the (k-1)th stage is preprocessed using pure noise samples to obtain training data pairs in the kth stage, wherein k>1, and the training data pairs include: noisy speech and noisy speech.
[0079] The training data pair of the k-th stage is input into the speech enhancement model, and the noisy speech in the training data pair of the k-th stage is denoised to obtain the noisy speech output by the k-th stage.
[0080] The speech enhancement model is trained using the loss function of the noisy speech output in the k-th stage and the noisy speech in the training data pair in the k-th stage. If the loss function does not converge, the speech enhancement model is iteratively trained in the next stage using the pure noise and the noisy speech output in the k-th stage until the loss function converges.
[0081] As a non-volatile computer-readable storage medium, it can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the embodiments of this invention. One or more program instructions are stored in the non-volatile computer-readable storage medium, and when executed by a processor, they perform the iterative self-supervised training method for the speech enhancement model in any of the above method embodiments.
[0082] Figure 9 This is a schematic diagram of the hardware structure of an electronic device for an iterative self-supervised training method for a speech enhancement model, as provided in another embodiment of this application. Figure 9 As shown, the device includes:
[0083] One or more processors 910 and memory 920, Figure 9 Taking a processor 910 as an example, the device for the iterative self-supervised training method of the speech enhancement model may also include an input device 930 and an output device 940.
[0084] The processor 910, memory 920, input device 930, and output device 940 can be connected via a bus or other means. Figure 9 Taking the example of a connection between China and Israel via a bus.
[0085] The memory 920, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the iterative self-supervised training method for the speech enhancement model in the embodiments of this application. The processor 910 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 920, thereby implementing the iterative self-supervised training method for the speech enhancement model in the above-described method embodiments.
[0086] The memory 920 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store data, etc. Furthermore, the memory 920 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 920 may optionally include memory remotely located relative to the processor 910, and these remote memories can be connected to the mobile device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0087] Input device 930 can receive input numerical or character information. Output device 940 may include display devices such as a display screen.
[0088] The one or more modules are stored in the memory 920, and when executed by the one or more processors 910, they perform the iterative self-supervised training method for the speech enhancement model in any of the above method embodiments.
[0089] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.
[0090] Non-volatile computer-readable storage media may include a stored program area and a stored data area, wherein the stored program area may store an operating system and an application program required for at least one function; the stored data area may store data created based on the use of the device, etc. Furthermore, the non-volatile computer-readable storage medium may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the non-volatile computer-readable storage medium may optionally include memory remotely located relative to the processor, and these remote memories may be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0091] This invention also provides an electronic device comprising: at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the iterative self-supervised training method for speech enhancement models according to any embodiment of this invention.
[0092] The electronic devices described in this application exist in various forms, including but not limited to:
[0093] (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and primarily aim to provide voice and data communication. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.
[0094] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, possessing computing and processing capabilities, and generally also have mobile internet access features. These terminals include PDAs, MIDs, and UMPCs, such as tablet computers.
[0095] (3) Portable entertainment devices: These devices can display and play multimedia content. This category includes audio and video players, handheld game consoles, e-book readers, as well as smart toys and portable car navigation devices.
[0096] (4) Other electronic devices with data processing functions.
[0097] In this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, without necessarily requiring or implying any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising" or "including" include not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0098] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0099] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An iterative self-supervised training method for a speech enhancement model, comprising: The noisy speech output by the speech enhancement model in stage k-1 is preprocessed using pure noise samples to obtain training data pairs in stage k, where k > 1. The training data pairs include noisy speech and noisy speech. The speech enhancement model is an encoder-decoder structure, wherein the encoder includes multiple convolutional blocks and deconvolutional blocks corresponding to the multiple convolutional blocks. The multiple convolutional blocks and the multiple deconvolutional blocks are connected by attention modules of each convolutional block, and each convolutional block attention module consists of channel attention blocks and frequency attention blocks. The encoder and the decoder are connected by multiple dual-path blocks, wherein the dual-path blocks include a first long short-term memory layer modeled in the frequency domain and a second long short-term memory layer modeled in the time domain. The training data pair of the k-th stage is input into the speech enhancement model, and the noisy speech in the training data pair of the k-th stage is denoised to obtain the noisy speech output by the k-th stage. The speech enhancement model is trained using the loss function of the noisy speech output in the k-th stage and the noisy speech in the training data pair in the k-th stage. If the loss function does not converge, the speech enhancement model is iteratively trained in the next stage using the pure noise and the noisy speech output in the k-th stage until the loss function converges.
2. The method according to claim 1, wherein, When k = 1, the method further includes: The original noisy speech was preprocessed using pure noise samples to obtain training data pairs for the first stage.
3. The method according to claim 1, wherein, The preprocessing of the noisy speech output by the speech enhancement model in the (k-1)th stage using pure noise samples includes: Select a portion of pure noise that is not related to the noisy speech noise from the pure noise sample, and add the pure noise to the noisy speech according to a preset signal-to-noise ratio to obtain noisy speech.
4. An iterative self-supervised training system for a speech enhancement model, comprising: A preprocessing module is used to preprocess the noisy speech output by the speech enhancement model in stage k-1 using pure noise samples to obtain training data pairs in stage k, where k > 1. The training data pairs include noisy speech and noisy speech. The speech enhancement model is an encoder-decoder structure, wherein the encoder includes multiple convolutional blocks and deconvolutional blocks corresponding to the multiple convolutional blocks. The multiple convolutional blocks and the multiple deconvolutional blocks are connected by attention modules of each convolutional block. Each attention module of the convolutional block consists of channel attention blocks and frequency attention blocks. The encoder and the decoder are connected by multiple dual-path blocks, wherein the dual-path blocks include a first long short-term memory layer modeled in the frequency domain and a second long short-term memory layer modeled in the time domain. The speech enhancement module is used to input the training data pair of the k-th stage into the speech enhancement model, perform speech denoising on the noisy speech in the training data pair of the k-th stage, and obtain the denoised noisy speech output by the k-th stage. The iterative training module is used to perform self-supervised learning training of the speech enhancement model in the k-th stage based on the loss function of the noisy speech output in the k-th stage and the noisy speech in the training data pair in the k-th stage. If the loss function does not converge, the speech enhancement model is iteratively trained in the next stage using the pure noise and the noisy speech output in the k-th stage until the loss function converges.
5. The system according to claim 4, wherein, The preprocessing module is used for: When k=1, the original noisy speech is preprocessed using pure noise samples to obtain training data pairs in the first stage.
6. The system according to claim 4, wherein, The preprocessing module is also used for: Select a portion of pure noise that is not related to the noisy speech noise from the pure noise sample, and add the pure noise to the noisy speech according to a preset signal-to-noise ratio to obtain noisy speech.
7. An electronic device comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the steps of the method according to any one of claims 1-3.
8. A storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-3.
Citation Information
Patent Citations
Speech recognition method for comparative predictive coding self-supervised structure joint training
CN112767922A
Speech enhancement model training method and device and speech enhancement method and device
CN113593594A