Training method of speech signal enhancement model, speech signal enhancement method and device
By training a speech signal enhancement model and utilizing predictive echo cancellation and noise reduction masking information separation processing, the problem of poor speech signal enhancement effect in existing technologies is solved, and more efficient speech signal enhancement is achieved.
Patent Information
- Application Number
- CN202211182522.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2042-09-27
AI Technical Summary
Existing speech signal enhancement models process speech signals with different enhancement needs in the same way, resulting in speech distortion and poor enhancement effect.
By acquiring the spectral information of sample speech signals and reference signals, a speech signal enhancement model is trained. Echo cancellation masking information and noise reduction masking information are used to separate echo cancellation and noise reduction functions, and the masking information is output separately for processing to meet different enhancement requirements.
It improves the enhancement effect of voice signals, avoids voice distortion, and meets the enhancement needs of different scenarios.
Smart Images

Figure CN115641862B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of speech processing, and particularly relates to a speech signal enhancement model training method, a speech signal enhancement method, a device, an electronic device, a storage medium and a computer program product. BACKGROUND
[0002] With the development of speech processing technology, a technology of processing a speech signal to perform noise suppression and echo cancellation on the speech signal has appeared. The noise suppression and echo cancellation on the speech signal are mainly to enhance the speech signal.
[0003] In the related art, the current speech signal enhancement method mainly performs enhancement processing on a speech signal to be processed by using a trained speech signal enhancement model (such as a neural network model). However, the trained speech signal enhancement model is one-step, and the same enhancement processing is performed on speech signals with different enhancement requirements, which is easy to cause speech distortion, thereby causing poor enhancement effect of the speech signal. SUMMARY
[0004] The present disclosure provides a speech signal enhancement model training method, a speech signal enhancement method, a device, an electronic device, a storage medium and a computer program product to at least solve the problem of poor enhancement effect of the speech signal in the related art. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, a speech signal enhancement model training method is provided, comprising:
[0006] obtaining a sample speech signal and a sample reference signal of the sample speech signal; the sample speech signal comprises a clean speech signal, a noise signal and an echo signal;
[0007] inputting spectrum information of the sample speech signal and spectrum information of the sample reference signal into a speech signal enhancement model to be trained to obtain predicted echo cancellation masking information and predicted noise reduction masking information;
[0008] obtaining initial enhancement spectrum information of the sample speech signal according to the spectrum information of the sample speech signal and the predicted echo cancellation masking information, and obtaining target enhancement spectrum information of the sample speech signal according to the initial enhancement spectrum information and the predicted noise reduction masking information;
[0009] According to differences between the initial enhanced spectrum information and first target spectrum information, and differences between the target enhanced spectrum information and second target spectrum information, the speech signal enhancement model to be trained is trained to obtain a trained speech signal enhancement model; the first target spectrum information includes spectrum information of the clean speech signal and spectrum information of the noise signal, and the second target spectrum information includes spectrum information of the clean speech signal.
[0010] In an example embodiment, the inputting of the spectrum information of the sample speech signal and the spectrum information of the sample reference signal into the speech signal enhancement model to be trained to obtain the predicted echo cancellation masking information and the predicted noise reduction masking information comprises:
[0011] The spectrum information of the sample speech signal and the spectrum information of the sample reference signal are input into the speech signal enhancement model to be trained for splicing processing to obtain spliced spectrum information of the sample speech signal.
[0012] The spliced spectrum information of the sample speech signal is subjected to feature extraction processing to obtain audio features of the sample speech signal.
[0013] The audio features of the sample speech signal are respectively subjected to first classification processing and second classification processing to obtain the predicted echo cancellation masking information and the predicted noise reduction masking information; the first classification processing is used for classifying the predicted echo cancellation masking information, and the second classification processing is used for classifying the predicted noise reduction masking information.
[0014] In an example embodiment, the training of the speech signal enhancement model to be trained according to differences between the initial enhanced spectrum information and first target spectrum information, and differences between the target enhanced spectrum information and second target spectrum information to obtain a trained speech signal enhancement model comprises:
[0015] A first loss value is obtained according to the differences between the initial enhanced spectrum information and the first target spectrum information, and a second loss value is obtained according to the differences between the target enhanced spectrum information and the second target spectrum information;
[0016] The first loss value and the second loss value are subjected to fusion processing to obtain a target loss value.
[0017] According to the target loss value, the speech signal enhancement model to be trained is trained until a training end condition is reached, and the speech signal enhancement model after training that reaches the training end condition is taken as the trained speech signal enhancement model.
[0018] In an example embodiment, the obtaining of the initial enhanced spectral information of the sample voice signal according to the spectral information of the sample voice signal and the predicted echo cancellation masking information, and the obtaining of the target enhanced spectral information of the sample voice signal according to the initial enhanced spectral information and the predicted noise reduction masking information, comprises:
[0019] fusing the spectral information of the sample voice signal and the predicted echo cancellation masking information to obtain the initial enhanced spectral information of the sample voice signal;
[0020] fusing the initial enhanced spectral information and the predicted noise reduction masking information to obtain the target enhanced spectral information of the sample voice signal.
[0021] In an example embodiment, the spectral information of the sample voice signal comprises the spectral information of the clean voice signal, the spectral information of the noise signal and the spectral information of the echo signal.
[0022] Before the training of the voice signal enhancement model according to the difference between the initial enhanced spectral information and the first target spectral information, and the difference between the target enhanced spectral information and the second target spectral information, to obtain the trained voice signal enhancement model, further comprising:
[0023] extracting the spectral information of the clean voice signal and the spectral information of the noise signal from the spectral information of the sample voice signal;
[0024] generating the first target spectral information according to the spectral information of the clean voice signal and the spectral information of the noise signal, and generating the second target spectral information according to the spectral information of the clean voice signal.
[0025] According to a second aspect of the embodiments of the present disclosure, a voice signal enhancement method is provided, comprising:
[0026] obtaining a voice signal to be enhanced and a reference signal of the voice signal to be enhanced;
[0027] inputting the spectral information of the voice signal to be enhanced and the spectral information of the reference signal into the trained voice signal enhancement model to obtain echo cancellation masking information and noise reduction masking information; the trained voice signal enhancement model is trained according to the voice signal enhancement model training method in any one of the embodiments of the first aspect;
[0028] updating the echo cancellation masking information and the noise reduction masking information according to a signal enhancement request of the voice signal to be enhanced to obtain updated echo cancellation masking information and updated noise reduction masking information;
[0029] enhancing the spectral information of the speech signal to be enhanced according to the updated echo cancellation masking information and the updated noise reduction masking information, to obtain target enhanced spectral information of the speech signal to be enhanced;
[0030] transforming the target enhanced spectral information to obtain an enhanced speech signal of the speech signal to be enhanced.
[0031] In an example embodiment, the inputting the spectral information of the speech signal to be enhanced and the spectral information of the reference signal into the trained speech signal enhancement model to obtain echo cancellation masking information and noise reduction masking information comprises:
[0032] inputting the spectral information of the speech signal to be enhanced and the spectral information of the reference signal into the trained speech signal enhancement model for splicing processing to obtain spliced spectral information of the speech signal to be enhanced;
[0033] performing feature extraction processing on the spliced spectral information of the speech signal to be enhanced to obtain audio features of the speech signal to be enhanced;
[0034] respectively performing first classification processing and second classification processing on the audio features of the speech signal to be enhanced to obtain the echo cancellation masking information and the noise reduction masking information; the first classification processing is used for classifying the echo cancellation masking information, and the second classification processing is used for classifying the noise reduction masking information.
[0035] In an example embodiment, the updating the echo cancellation masking information and the noise reduction masking information according to the signal enhancement request of the speech signal to be enhanced to obtain updated echo cancellation masking information and updated noise reduction masking information comprises:
[0036] confirming target echo cancellation amplitude and target noise reduction amplitude of the speech signal to be enhanced according to the signal enhancement request of the speech signal to be enhanced;
[0037] updating the echo cancellation masking information according to the target echo cancellation amplitude to obtain updated echo cancellation masking information, and updating the noise reduction masking information according to the target noise reduction amplitude to obtain updated noise reduction masking information.
[0038] In an example embodiment, the enhancing the spectral information of the speech signal to be enhanced according to the updated echo cancellation masking information and the updated noise reduction masking information, to obtain target enhanced spectral information of the speech signal to be enhanced comprises:
[0039] Fusing the spectrum information of the to-be-enhanced voice signal and the updated echo cancellation masking information to obtain initial enhanced spectrum information of the to-be-enhanced voice signal;
[0040] Fusing the initial enhanced spectrum information and the updated noise reduction masking information to obtain target enhanced spectrum information of the to-be-enhanced voice signal.
[0041] According to a third aspect of the embodiments of the present disclosure, a training device of a voice signal enhancement model is provided, comprising:
[0042] A sample obtaining unit configured to perform obtaining a sample voice signal and a sample reference signal of the sample voice signal; the sample voice signal comprises a clean voice signal, a noise signal and an echo signal;
[0043] An information predicting unit configured to perform inputting spectrum information of the sample voice signal and spectrum information of the sample reference signal into a to-be-trained voice signal enhancement model to obtain predicted echo cancellation masking information and predicted noise reduction masking information;
[0044] A spectrum processing unit configured to perform obtaining initial enhanced spectrum information of the sample voice signal according to the spectrum information of the sample voice signal and the predicted echo cancellation masking information, and obtaining target enhanced spectrum information of the sample voice signal according to the initial enhanced spectrum information and the predicted noise reduction masking information;
[0045] A model training unit configured to perform training the to-be-trained voice signal enhancement model according to differences between the initial enhanced spectrum information and first target spectrum information, and differences between the target enhanced spectrum information and second target spectrum information to obtain a trained voice signal enhancement model; the first target spectrum information comprises spectrum information of the clean voice signal and spectrum information of the noise signal, and the second target spectrum information comprises spectrum information of the clean voice signal.
[0046] In an exemplary embodiment, the information predicting unit is further configured to perform inputting the spectrum information of the sample voice signal and the spectrum information of the sample reference signal into the to-be-trained voice signal enhancement model for splicing processing to obtain spliced spectrum information of the sample voice signal; performing feature extraction processing on the spliced spectrum information of the sample voice signal to obtain audio features of the sample voice signal; respectively performing first classification processing and second classification processing on the audio features of the sample voice signal to obtain the predicted echo cancellation masking information and the predicted noise reduction masking information; the first classification processing is used for classifying the predicted echo cancellation masking information, and the second classification processing is used for classifying the predicted noise reduction masking information.
[0047] In an example embodiment, the model training unit is further configured to perform obtaining a first loss value according to a difference between the initial enhanced spectrum information and the first target spectrum information, and obtaining a second loss value according to a difference between the target enhanced spectrum information and the second target spectrum information; performing fusion processing on the first loss value and the second loss value to obtain a target loss value; training the to-be-trained speech signal enhancement model according to the target loss value until a training end condition is reached, and taking the trained speech signal enhancement model that reaches the training end condition as the trained speech signal enhancement model.
[0048] In an example embodiment, the spectrum processing unit is further configured to perform fusion processing on the spectrum information of the sample speech signal and the predicted echo cancellation masking information to obtain initial enhanced spectrum information of the sample speech signal; and fusion processing on the initial enhanced spectrum information and the predicted noise reduction masking information to obtain target enhanced spectrum information of the sample speech signal.
[0049] In an example embodiment, the spectrum information of the sample speech signal includes spectrum information of the clean speech signal, spectrum information of the noise signal, and spectrum information of the echo signal.
[0050] The apparatus further includes an information generation unit configured to perform extracting the spectrum information of the clean speech signal and the spectrum information of the noise signal from the spectrum information of the sample speech signal; generating the first target spectrum information according to the spectrum information of the clean speech signal and the spectrum information of the noise signal, and generating the second target spectrum information according to the spectrum information of the clean speech signal.
[0051] According to a fourth aspect of the embodiments of the present disclosure, a speech signal enhancement apparatus is provided, including:
[0052] A signal acquisition unit configured to perform acquiring a to-be-enhanced speech signal and a reference signal of the to-be-enhanced speech signal;
[0053] An information processing unit configured to perform inputting spectrum information of the to-be-enhanced speech signal and spectrum information of the reference signal into a trained speech signal enhancement model to obtain echo cancellation masking information and noise reduction masking information; the trained speech signal enhancement model being trained according to the training method of the speech signal enhancement model in any one of the embodiments of the first aspect.
[0054] The information update unit is configured to perform an update of the echo cancellation masking information and the noise reduction masking information based on the signal enhancement request of the speech signal to be enhanced, so as to obtain the updated echo cancellation masking information and the updated noise reduction masking information.
[0055] The information enhancement unit is configured to perform enhancement processing on the spectral information of the speech signal to be enhanced based on the updated echo cancellation masking information and the updated noise reduction masking information, to obtain the target enhanced spectral information of the speech signal to be enhanced;
[0056] The information transformation unit is configured to transform the target enhancement spectrum information to obtain the enhanced speech signal of the speech signal to be enhanced.
[0057] In an exemplary embodiment, the information processing unit is further configured to perform concatenation processing on a trained speech signal enhancement model by inputting the spectral information of the speech signal to be enhanced and the spectral information of the reference signal into the model to obtain concatenated spectral information of the speech signal to be enhanced; perform feature extraction processing on the concatenated spectral information of the speech signal to be enhanced to obtain audio features of the speech signal to be enhanced; and perform a first classification processing and a second classification processing on the audio features of the speech signal to be enhanced to obtain echo cancellation masking information and noise reduction masking information, respectively; the first classification processing is used to classify the echo cancellation masking information, and the second classification processing is used to classify the noise reduction masking information.
[0058] In an exemplary embodiment, the information update unit is further configured to perform the following actions based on the signal enhancement request of the speech signal to be enhanced: confirming the target echo cancellation amplitude and the target noise reduction amplitude of the speech signal to be enhanced; updating the echo cancellation masking information according to the target echo cancellation amplitude to obtain updated echo cancellation masking information; and updating the noise reduction masking information according to the target noise reduction amplitude to obtain updated noise reduction masking information.
[0059] In an exemplary embodiment, the information enhancement unit is further configured to perform a fusion process of the spectral information of the speech signal to be enhanced and the updated echo cancellation masking information to obtain initial enhanced spectral information of the speech signal to be enhanced; and to perform a fusion process of the initial enhanced spectral information and the updated noise reduction masking information to obtain target enhanced spectral information of the speech signal to be enhanced.
[0060] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising:
[0061] processor;
[0062] Memory used to store the processor's executable instructions;
[0063] The processor is configured to execute the instructions to implement the training method for the speech signal enhancement model as described in any embodiment of the first aspect, or the speech signal enhancement method as described in any embodiment of the second aspect.
[0064] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform a training method for a speech signal enhancement model as described in any embodiment of the first aspect, or a speech signal enhancement method as described in any embodiment of the second aspect.
[0065] According to a seventh aspect of the present disclosure, a computer program product is provided, the computer program product including instructions that, when executed by a processor of an electronic device, enable the electronic device to perform a training method for a speech signal enhancement model as described in any embodiment of the first aspect, or a speech signal enhancement method as described in any embodiment of the second aspect.
[0066] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:
[0067] The process involves acquiring a sample speech signal and a sample reference signal; the sample speech signal includes a clean speech signal, a noise signal, and an echo signal; then, the spectral information of the sample speech signal and the spectral information of the sample reference signal are input into the speech signal enhancement model to be trained to obtain predicted echo cancellation masking information and predicted noise reduction masking information; next, based on the spectral information of the sample speech signal and the predicted echo cancellation masking information, the initial enhanced spectral information of the sample speech signal is obtained, and based on the initial enhanced spectral information and the predicted noise reduction masking information, the target enhanced spectral information of the sample speech signal is obtained; finally, based on the difference between the initial enhanced spectral information and the first target spectral information, as well as the difference between the target enhanced spectral information and the second target spectral information, the speech signal enhancement model to be trained is trained to obtain the trained speech signal enhancement model; the first target spectral information includes the spectral information of the clean speech signal and the spectral information of the noise signal, and the second target spectral information includes the spectral information of the clean speech signal. In this way, when training the speech signal enhancement model, the differences between the initial enhanced spectrum information and the first target spectrum information obtained based on the predicted echo cancellation masking information, as well as the differences between the target enhanced spectrum information and the second target spectrum information obtained based on the predicted noise reduction masking information, are fully utilized to assist in the training of the speech signal enhancement model. Since the echo cancellation function and the noise reduction function are separate, the trained speech signal enhancement model can output the echo cancellation masking information and the noise reduction masking information separately. This facilitates the subsequent processing of the separately output echo cancellation masking information and noise reduction masking information to meet the enhancement requirements of the speech signal, and it is less likely to cause speech distortion. This improves the enhancement effect of the speech signal and avoids the defect of performing echo cancellation and noise reduction processing on the speech signal simultaneously in any scenario, which would lead to speech distortion and poor speech signal enhancement effect.
[0068] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0069] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0070] Figure 1 This is a flowchart illustrating a training method for a speech signal enhancement model according to an exemplary embodiment.
[0071] Figure 2 This is a flowchart illustrating the steps of obtaining predicted echo cancellation masking information and predicted noise reduction masking information according to an exemplary embodiment.
[0072] Figure 3This is a block diagram illustrating a speech signal enhancement model according to an exemplary embodiment.
[0073] Figure 4 The calculation of initial enhanced spectral information S is illustrated according to an exemplary embodiment. aec Flowcharts of (t, f) and target-enhanced spectral information S(t, f).
[0074] Figure 5 This is a flowchart illustrating a speech signal enhancement method according to an exemplary embodiment.
[0075] Figure 6 This is a flowchart illustrating another method for training a speech signal enhancement model according to an exemplary embodiment.
[0076] Figure 7 This is a block diagram illustrating a training apparatus for a speech signal enhancement model according to an exemplary embodiment.
[0077] Figure 8 This is a block diagram illustrating a speech signal enhancement device according to an exemplary embodiment.
[0078] Figure 9 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation
[0079] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0080] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0081] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0082] Figure 1 This is a flowchart illustrating a training method for a speech signal enhancement model according to an exemplary embodiment, such as... Figure 1As shown, the training method for this speech signal enhancement model is used in a terminal; it is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this exemplary embodiment, the method includes the following steps:
[0083] In step S110, the sample speech signal and the sample reference signal of the sample speech signal are acquired; the sample speech signal includes a clean speech signal, a noise signal and an echo signal.
[0084] Among them, the sample speech signal refers to the speech signal used to train the speech signal enhancement model, such as the near-end microphone signal.
[0085] The sample reference signal refers to the far-end reference signal, specifically the speech signal received from a distant location, such as a far-end microphone signal. This far-end reference signal is played back by the near-end speaker and then captured by the near-end microphone to form an echo signal. It should be noted that the sample reference signal can also be called the near-end speaker signal.
[0086] It should be noted that when voice communication is conducted between the near end (also known as the local end) and the far end (also known as the other end), the microphone at the near end can collect the voice signal at the near end and send the voice signal at the near end to the far end, where it is played out by the speaker at the far end; at the same time, the speaker at the near end can play the voice signal sent by the far end, such as the voice signal at the far end collected by the microphone at the far end.
[0087] The sample speech signal includes three types of speech signals: clean speech signal, noise signal, and echo signal. The clean speech signal refers to the pure speech in the sample speech signal, such as the voice of a broadcaster. The noise signal refers to the environmental noise in the sample speech signal, such as keyboard sounds. The echo signal refers to the echo in the sample speech signal, specifically the echo signal corresponding to the far-end reference signal, such as the echo signal of the audio signal played by the near-end speaker.
[0088] Specifically, the terminal obtains sample speech signals and sample reference signals from the local database, which facilitates the subsequent training of a speech signal enhancement model using the sample speech signals and sample reference signals.
[0089] In step S120, the spectral information of the sample speech signal and the spectral information of the sample reference signal are input into the speech signal enhancement model to be trained to obtain the predicted echo cancellation masking information and the predicted noise reduction masking information.
[0090] The spectral information of the sample speech signal refers to the spectrum of the sample speech signal, which is obtained by performing STFT (Short-Time Fourier Transform) processing on the sample speech signal.
[0091] The spectral information of the sample reference signal refers to the spectrum of the sample reference signal, which is obtained by performing STFT processing on the sample reference signal.
[0092] The speech signal enhancement model to be trained refers to a DNN (Deep Neural Networks) model, which specifically includes connected layers, convolutional layers, gated recurrent units, and two fully connected layers. It should be noted that traditional DNN models only include one fully connected layer.
[0093] It should be noted that traditional speech signal enhancement models (such as traditional DNN models) are one-step processes, which can only output an enhanced signal that has undergone both noise suppression and echo cancellation, meaning that the resulting enhanced signal does not contain noise or echo. However, different speech signals have different enhancement requirements. If the above one-step speech signal enhancement model is used to process each speech signal, it is easy to cause speech distortion, resulting in poor speech signal enhancement effect.
[0094] Among them, predictive echo cancellation masking information refers to an echo cancellation masking matrix that can achieve echo cancellation function, specifically used to eliminate echo signals in sample speech signals. Predictive denoising masking information refers to a denoising masking matrix that can achieve denoising function, specifically used to reduce noise signals in sample speech signals.
[0095] Specifically, the terminal performs STFT processing on the sample speech signal to obtain the spectral information of the sample speech signal; it also performs STFT processing on the sample reference signal to obtain the spectral information of the sample reference signal; the spectral information of the sample speech signal and the spectral information of the sample reference signal are input into the speech signal enhancement model to be trained; the speech signal enhancement model to be trained concatenates the spectral information of the sample speech signal and the spectral information of the sample reference signal, and extracts the audio features from the concatenated spectral information; the audio features are then processed twice with fully connected layers to obtain the predicted echo cancellation masking information and the predicted noise reduction masking information.
[0096] For example, the terminal inputs the spectral information of the sample speech signal and the spectral information of the sample reference signal into the DNN model to be trained, and outputs the predicted echo cancellation masking matrix and the predicted noise reduction masking matrix through the DNN model to be trained.
[0097] In step S130, the initial enhanced spectrum information of the sample speech signal is obtained based on the spectrum information of the sample speech signal and the predicted echo cancellation masking information, and the target enhanced spectrum information of the sample speech signal is obtained based on the initial enhanced spectrum information and the predicted noise reduction masking information.
[0098] The initial enhanced spectral information refers to the spectral information obtained after echo cancellation of the sample speech signal. The target enhanced spectral information refers to the spectral information obtained after echo cancellation and noise reduction of the sample speech signal.
[0099] Specifically, the terminal uses predictive echo cancellation masking information to perform echo cancellation on the spectral information of the sample speech signal, obtaining echo-cancelled spectral information as the initial enhanced spectral information of the sample speech signal; and uses predictive noise reduction masking information to perform noise reduction on the initial enhanced spectral information of the sample speech signal, obtaining echo-cancelled and noise-reduced spectral information as the target enhanced spectral information of the sample speech signal.
[0100] In step S140, the speech signal enhancement model to be trained is trained based on the difference between the initial enhanced spectrum information and the first target spectrum information, as well as the difference between the target enhanced spectrum information and the second target spectrum information, to obtain the trained speech signal enhancement model; the first target spectrum information includes the spectrum information of the clean speech signal and the spectrum information of the noise signal, and the second target spectrum information includes the spectrum information of the clean speech signal.
[0101] Specifically, the terminal acquires the spectral information of the clean speech signal and the spectral information of the noise signal, and combines the spectral information of the clean speech signal and the noise signal to obtain the first target spectral information; the spectral information of the clean speech signal is confirmed as the second target spectral information; based on the difference between the initial enhanced spectral information and the first target spectral information, and the difference between the target enhanced spectral information and the second target spectral information, a target loss value is obtained; based on the target loss value, the speech signal enhancement model to be trained is trained until the training termination condition is met; the trained speech signal enhancement model that meets the training termination condition is taken as the trained speech signal enhancement model.
[0102] It should be noted that the training termination condition refers to the current training iterations reaching the preset number of training iterations, the current loss value being less than the preset threshold, etc.
[0103] In the training method of the above-mentioned speech signal enhancement model, a sample speech signal and a sample reference signal of the sample speech signal are acquired. The sample speech signal includes a clean speech signal, a noise signal, and an echo signal. Then, the spectral information of the sample speech signal and the spectral information of the sample reference signal are input into the speech signal enhancement model to be trained to obtain predicted echo cancellation masking information and predicted noise reduction masking information. Next, based on the spectral information of the sample speech signal and the predicted echo cancellation masking information, the initial enhanced spectral information of the sample speech signal is obtained, and based on the initial enhanced spectral information and the predicted noise reduction masking information, the target enhanced spectral information of the sample speech signal is obtained. Finally, based on the difference between the initial enhanced spectral information and the first target spectral information, and the difference between the target enhanced spectral information and the second target spectral information, the speech signal enhancement model to be trained is trained to obtain the trained speech signal enhancement model. The first target spectral information includes the spectral information of the clean speech signal and the spectral information of the noise signal, and the second target spectral information includes the spectral information of the clean speech signal. In this way, when training the speech signal enhancement model, the differences between the initial enhanced spectrum information and the first target spectrum information obtained based on the predicted echo cancellation masking information, as well as the differences between the target enhanced spectrum information and the second target spectrum information obtained based on the predicted noise reduction masking information, are fully utilized to assist in the training of the speech signal enhancement model. Since the echo cancellation function and the noise reduction function are separate, the trained speech signal enhancement model can output the echo cancellation masking information and the noise reduction masking information separately. This facilitates the subsequent processing of the separately output echo cancellation masking information and noise reduction masking information to meet the enhancement requirements of the speech signal, and it is less likely to cause speech distortion. This improves the enhancement effect of the speech signal and avoids the defect of performing echo cancellation and noise reduction processing on the speech signal simultaneously in any scenario, which would lead to speech distortion and poor speech signal enhancement effect.
[0104] In one exemplary embodiment, such as Figure 2 As shown, in step S120, the spectral information of the sample speech signal and the spectral information of the sample reference signal are input into the speech signal enhancement model to be trained to obtain the predicted echo cancellation masking information and the predicted noise reduction masking information. This can be achieved through the following steps:
[0105] In step S210, the spectral information of the sample speech signal and the spectral information of the sample reference signal are input into the speech signal enhancement model to be trained for splicing processing to obtain the spliced spectral information of the sample speech signal.
[0106] In step S220, feature extraction processing is performed on the spliced spectrum information of the sample speech signal to obtain the audio features of the sample speech signal.
[0107] In step S230, the audio features of the sample speech signal are subjected to first classification processing and second classification processing respectively to obtain predicted echo cancellation masking information and predicted noise reduction masking information.
[0108] Among them, spliced spectrum information refers to the spectrum information obtained after splicing the spectrum information of the sample speech signal and the spectrum information of the sample reference signal.
[0109] Among them, the audio features of the sample speech signal are used to represent the speech features of the sample speech signal.
[0110] In this context, the first classification process refers to the first fully connected processing, and the second classification process refers to the second fully connected processing. Furthermore, the first classification process is used to classify the predicted echo cancellation masking information; that is, the result of the first classification process is the predicted echo cancellation masking information. The second classification process is used to classify the predicted denoising masking information; that is, the result of the second classification process is the predicted denoising masking information.
[0111] Specifically, the terminal inputs the spectral information of the sample speech signal and the spectral information of the sample reference signal into the speech signal enhancement model to be trained. The speech signal enhancement model then concatenates the spectral information of the sample speech signal and the spectral information of the sample reference signal to obtain the concatenated spectral information of the sample speech signal. Feature extraction is performed on the concatenated spectral information of the sample speech signal to obtain the extracted audio features, which serve as the audio features of the sample speech signal. The audio features of the sample speech signal undergo a first classification process to obtain the corresponding processing result, which serves as the predicted echo cancellation masking information. Simultaneously, the audio features of the sample speech signal undergo a second classification process to obtain the corresponding processing result, which serves as the predicted noise reduction masking information.
[0112] For example, see reference. Figure 3 The terminal will display the spectral information S of the sample speech signal. d (t, f) and the spectral information S of the sample reference signal r (t, f) is the input to the speech signal enhancement model to be trained. The spectral information S of the sample speech signal is processed by the splicing layer in the speech signal enhancement model to be trained. d (t, f) and the spectral information S of the sample reference signal r (t, f) is concatenated to obtain the concatenated spectrum information of the sample speech signal; multiple convolutional layers are used to extract features from the concatenated spectrum information of the sample speech signal to obtain the initial audio features of the sample speech signal; a gated recurrent unit is used to extract features from the initial audio features of the sample speech signal again to obtain the audio features of the sample speech signal; the first fully connected layer is used to fully connect the audio features of the sample speech signal to obtain the predicted echo cancellation masking matrix M. aecThe audio features of the sample speech signal are processed by a second fully connected layer to obtain the predicted noise reduction masking matrix M. ns Where t refers to time and f refers to frequency.
[0113] The technical solution provided in this disclosure, by inputting the spectral information of the sample speech signal and the spectral information of the sample reference signal into the speech signal enhancement model to be trained, obtains predicted echo cancellation masking information and predicted noise reduction masking information. This helps to distinguish between the echo cancellation function and the noise reduction function, enabling the trained speech signal enhancement model to output echo cancellation masking information and noise reduction masking information separately. This facilitates subsequent processing of the separately output echo cancellation masking information and noise reduction masking information to meet the speech signal enhancement requirements, reduces the likelihood of speech distortion, and thus improves the speech signal enhancement effect.
[0114] In an exemplary embodiment, step S140, which trains the speech signal enhancement model to be trained based on the difference between the initial enhanced spectrum information and the first target spectrum information, and the difference between the target enhanced spectrum information and the second target spectrum information, to obtain a trained speech signal enhancement model, specifically includes the following: obtaining a first loss value based on the difference between the initial enhanced spectrum information and the first target spectrum information, and obtaining a second loss value based on the difference between the target enhanced spectrum information and the second target spectrum information; fusing the first loss value and the second loss value to obtain a target loss value; training the speech signal enhancement model to be trained based on the target loss value until the training termination condition is met, and taking the trained speech signal enhancement model that has met the training termination condition as the trained speech signal enhancement model.
[0115] Specifically, the first loss value refers to the echo cancellation loss. Furthermore, a first target spectral information, including the spectral information of both the clean speech signal and the noise signal, is used as the target to monitor the effectiveness of echo cancellation.
[0116] The second loss value refers to the noise suppression loss. Furthermore, a second target spectral information, including the spectral information of the clean speech signal, is used as the target to monitor the effectiveness of noise suppression.
[0117] Specifically, the terminal obtains a first loss value based on the difference between the initial enhanced spectral information and the first target spectral information, which includes the spectral information of the clean speech signal and the spectral information of the noise signal, combined with a first loss function; it obtains a second loss value based on the difference between the target enhanced spectral information and the second target spectral information, which includes the spectral information of the clean speech signal, combined with a second loss function; the first loss value and the second loss value are added together to obtain the target loss value; the model parameters of the speech signal enhancement model to be trained are adjusted according to the target loss value to obtain the speech signal enhancement model with adjusted model parameters; the speech signal enhancement model with adjusted model parameters is retrained until the training termination condition is met, and the trained speech signal enhancement model that has met the training termination condition is taken as the trained speech signal enhancement model.
[0118] It should be noted that the first and second loss functions can be any loss function, such as the MSE (mean-square error) loss function.
[0119] For example, see reference. Figure 4 The terminal will display the spectral information S of the sample speech signal. d (t, f) and the spectral information S of the sample reference signal r (t, f) is the input to the speech signal enhancement model to be trained, and the predicted echo cancellation masking matrix M is obtained. aec and the prediction denoising masking matrix M ns Among them, the spectral information S of the sample speech signal d (t, f) includes the spectral information Ss(t, f) of the clean speech signal and the spectral information S of the noise signal. n (t, f) and the spectral information S of the echo signal e (t, f), i.e., S d (t, f) = Ss(t, f) + S n (t,f)+S e (t, f); then, the spectrum information of the first target is Ss(t, f) + S n The second target spectral information is Ss(t, f). Then, the terminal uses the spectral information Ss(t, f) of the sample speech signal as a basis. d (t, f) and the predicted echo cancellation masking matrix M aec The initial enhanced spectral information S of the sample speech signal is obtained. aec (t, f); based on the initial enhanced spectral information S of the sample speech signal aec (t, f) and the predicted noise reduction masking matrix M ns The target augmented spectrum information S(t, f) of the sample speech signal is obtained. Then, the terminal uses the initial augmented spectrum information S... aec(t, f) and the first target spectrum information Ss(t, f) + S n The difference between (t, f) is used to calculate the first loss value, loss(S). aec (t, f), Ss(t, f) + S n (t, f)); Based on the difference between the target enhanced spectral information S(t, f) and the second target spectral information Ss(t, f), the second loss value loss(S(t, f), Ss(t, f)) is calculated; The first loss value loss(S aec (t, f), Ss(t, f) + S n The target loss value Loss is obtained by adding the second loss value (S(t,f)) and the second loss value (S(t,f), Ss(t,f)).
[0120] Loss = loss(S) aec (t, f), Ss(t, f) + S n (t, f)) + loss (S (t, f), Ss (t, f));
[0121] Furthermore, if the target loss value Loss is less than a preset threshold, the terminal adjusts the model parameters of the speech signal enhancement model to be trained according to the target loss value, and retrains the speech signal enhancement model with the adjusted model parameters until the target loss value Loss obtained from the trained speech signal enhancement model is less than the preset threshold. Then, the trained speech signal enhancement model is regarded as the trained speech signal enhancement model.
[0122] The technical solution provided in this disclosure, when training a speech signal enhancement model, fully utilizes the difference between the initial enhanced spectrum information and the first target spectrum information obtained based on predicted echo cancellation masking information, as well as the difference between the target enhanced spectrum information and the second target spectrum information obtained based on predicted noise reduction masking information, to assist in the training of the speech signal enhancement model. This ensures that the echo cancellation function and noise reduction function of the trained speech signal enhancement model are distinguished, avoiding the defect that simultaneous echo cancellation and noise reduction processing of the speech signal in any scenario would lead to speech distortion and poor speech signal enhancement effect.
[0123] In an exemplary embodiment, step S140, which obtains initial enhanced spectrum information of the sample speech signal based on the spectrum information of the sample speech signal and the predicted echo cancellation masking information, and obtains target enhanced spectrum information of the sample speech signal based on the initial enhanced spectrum information and the predicted noise reduction masking information, specifically includes the following: fusing the spectrum information of the sample speech signal and the predicted echo cancellation masking information to obtain the initial enhanced spectrum information of the sample speech signal; and fusing the initial enhanced spectrum information and the predicted noise reduction masking information to obtain the target enhanced spectrum information of the sample speech signal.
[0124] Among them, fusion processing refers to multiplication.
[0125] Specifically, the terminal multiplies the spectral information of the sample speech signal with the predicted echo cancellation masking information to obtain the corresponding spectral information, which serves as the initial enhanced spectral information of the sample speech signal; the terminal then multiplies the initial enhanced spectral information with the predicted noise reduction masking information to obtain the corresponding spectral information, which serves as the target enhanced spectral information of the sample speech signal.
[0126] For example, see reference. Figure 4 The terminal will display the spectral information S of the sample speech signal. d (t, f) and the predicted echo cancellation masking matrix M aec Multiplying the samples yields the initial enhanced spectral information S of the sample speech signal. aec (t, f); the initial enhanced spectral information of the sample speech signal S aec (t, f) and the predicted noise reduction masking matrix M ns Multiply the samples to obtain the target augmented spectral information S(t, f) of the sample speech signal.
[0127] The technical solution provided in this disclosure obtains the initial enhanced spectrum information of the sample speech signal based on the spectrum information of the sample speech signal and the predicted echo cancellation masking information, and obtains the target enhanced spectrum information of the sample speech signal based on the initial enhanced spectrum information and the predicted noise reduction masking information. This is beneficial for subsequently calculating different loss values based on the initial enhanced spectrum information and the target enhanced spectrum information, and then using different loss values to supervise the echo cancellation effect and noise reduction effect of the speech signal enhancement model, so that the echo cancellation masking information and noise reduction masking information output by the trained speech signal enhancement model are more accurate.
[0128] In an exemplary embodiment, the spectral information of the sample speech signal includes the spectral information of the clean speech signal, the spectral information of the noise signal, and the spectral information of the echo signal. Therefore, before training the speech signal enhancement model to be trained based on the difference between the initial enhanced spectral information and the first target spectral information, and the difference between the target enhanced spectral information and the second target spectral information, step S140 further includes the following: extracting the spectral information of the clean speech signal and the spectral information of the noise signal from the spectral information of the sample speech signal; generating the first target spectral information based on the spectral information of the clean speech signal and the noise signal; and generating the second target spectral information based on the spectral information of the clean speech signal.
[0129] Since the sample speech signal includes clean speech signal, noise signal and echo signal, the spectral information of the sample speech signal also includes the spectral information of clean speech signal, noise signal and echo signal.
[0130] For example, the spectral information S of the sample speech signal d (t, f) includes the spectral information Ss(t, f) of the clean speech signal and the spectral information S of the noise signal. n (t, f) and the spectral information S of the echo signal e (t, f), i.e., S d (t, f) = Ss(t, f) + S n (t,f)+S e (t, f); The terminal obtains the spectral information S from the sample speech signal. d In (t, f), the spectral information Ss(t, f) of the clean speech signal and the spectral information S of the noise signal are identified. n (t, f), and from the spectral information S of the sample speech signal d From (t, f), extract the spectral information Ss(t, f) of the identified clean speech signal and the spectral information S of the noise signal. n (t, f); The spectral information Ss(t, f) of the clean speech signal and the spectral information S of the noise signal are... n Combining (t, f) yields the first target spectrum information, which is Ss(t, f) + S n (t, f); Based on the spectral information Ss(t, f) of the clean speech signal, the second target spectral information is generated, that is, the second target spectral information is Ss(t, f).
[0131] The technical solution provided in this disclosure generates first target spectrum information based on the spectrum information of the clean speech signal and the spectrum information of the noise signal, and generates second target spectrum information based on the spectrum information of the clean speech signal. This is beneficial for subsequently using the first target spectrum information to supervise the echo cancellation effect of the speech signal enhancement model, and using the second target spectrum information to supervise the noise reduction effect of the speech signal enhancement model. This makes the echo cancellation masking information and noise reduction masking information output by the trained speech signal enhancement model more accurate, and further improves the accuracy of determining the echo cancellation masking information and noise reduction masking information.
[0132] Figure 5 This is a flowchart illustrating a speech signal enhancement method according to an exemplary embodiment, such as... Figure 5 As shown, this voice signal enhancement method is used in a terminal and includes the following steps:
[0133] In step S510, the speech signal to be enhanced and the reference signal of the speech signal to be enhanced are obtained.
[0134] The speech signal to be enhanced refers to the speech signal that needs to be enhanced, such as the signal from a near-end microphone.
[0135] In this context, the reference signal for the speech signal to be enhanced refers to a far-end reference signal, such as a far-end microphone signal. This far-end reference signal is played back by the near-end speaker and then captured by the near-end microphone to form an echo signal. In other words, the speech signal to be enhanced includes the echo signal of the reference signal.
[0136] In step S520, the spectral information of the speech signal to be enhanced and the spectral information of the reference signal are input into the trained speech signal enhancement model to obtain echo cancellation masking information and noise reduction masking information; the trained speech signal enhancement model is trained according to the training method of speech signal enhancement model.
[0137] The spectral information of the speech signal to be enhanced refers to the spectrum of the speech signal to be enhanced, which is specifically obtained by performing STFT (Short-Time Fourier Transform) processing on the speech signal to be enhanced.
[0138] The spectral information of the reference signal refers to the spectrum of the reference signal, which is obtained by performing STFT processing on the reference signal.
[0139] Among them, the trained speech signal enhancement model refers to a speech signal enhancement model that can output echo cancellation masking information and noise reduction masking information independently. Specifically, it includes connection layers, convolutional layers, gated recurrent units and two fully connected layers, such as the trained DNN model.
[0140] Echo cancellation masking information refers to the echo cancellation masking matrix, which enables echo cancellation and is specifically used to eliminate echo signals in the speech signal to be enhanced. The specific value of the echo cancellation masking matrix represents the degree of echo cancellation.
[0141] Among them, noise reduction masking information refers to a noise reduction masking matrix that can achieve noise reduction function, specifically used to reduce noise signals in the speech signal to be enhanced. The specific value of the noise reduction masking matrix is used to represent the degree of noise suppression.
[0142] In step S530, the echo cancellation masking information and noise reduction masking information are updated according to the signal enhancement request of the speech signal to be enhanced, so as to obtain the updated echo cancellation masking information and the updated noise reduction masking information.
[0143] Among them, the signal enhancement request is used to indicate the specific enhancement requirements of the speech signal to be enhanced; for example, in some scenarios, only echo cancellation or noise suppression is needed; in other scenarios, strong noise suppression is not required and it is even desirable to retain some non-speech sounds, such as keyboard sounds, but at the same time, it is necessary to completely eliminate echo.
[0144] The noise reduction masking matrix, represented by the noise reduction masking information, has parameters ranging from 0 to 1, representing the degree of suppression of the amplitude corresponding to each time-frequency point. For example, if a time-frequency point contains only speech components, ideally the value of the noise reduction masking matrix at that time-frequency point should be 1, meaning no suppression is performed; if a time-frequency point contains only noise components, ideally the value of the noise reduction masking matrix at that time-frequency point should be 0, meaning complete suppression; if a time-frequency point contains both noise and speech components, ideally the value of the noise reduction masking matrix at that time-frequency point should be the proportion of the speech component at that time-frequency point, meaning only speech is preserved while noise is completely suppressed; if a more conservative noise reduction capability is required to avoid damage to speech or other components, the noise reduction amount needs to be controlled. For example, all values of the noise reduction masking matrix should not be lower than -20dB, i.e., the value of the noise reduction masking matrix should not be lower than 0.1, thus achieving controllable noise reduction. The lower the value of -20dB, the higher the degree of noise reduction.
[0145] It should be noted that the control of echo cancellation amount is similar to the control of noise reduction amount, and will not be repeated here.
[0146] Among them, the updated echo cancellation masking information can control the amount of echo cancellation of the speech signal to be enhanced; the updated noise reduction masking information can control the amount of noise reduction of the speech signal to be enhanced, thereby meeting the speech enhancement requirements of the speech signal to be enhanced.
[0147] In step S540, the spectral information of the speech signal to be enhanced is enhanced based on the updated echo cancellation masking information and the updated noise reduction masking information to obtain the target enhanced spectral information of the speech signal to be enhanced.
[0148] Among them, enhancing the spectral information of the speech signal to be enhanced refers to using the updated echo cancellation masking information and the updated noise reduction masking information to perform echo cancellation and noise reduction processing on the spectral information of the speech signal to be enhanced.
[0149] In step S550, the target enhancement spectrum information is transformed to obtain the enhanced speech signal of the speech signal to be enhanced.
[0150] The enhanced speech signal is obtained by performing ISTFT (Inverse Short-Time Fourier Transform) processing on the target enhancement spectrum information of the speech signal to be enhanced.
[0151] Specifically, in response to a signal enhancement request for the speech signal to be enhanced, the terminal obtains the speech signal to be enhanced and a reference signal; it performs STFT processing on the speech signal to be enhanced to obtain its spectral information; it also performs STFT processing on the reference signal to obtain its spectral information; the spectral information of the speech signal to be enhanced and the spectral information of the reference signal are input into a trained speech signal enhancement model. The trained speech signal enhancement model concatenates the spectral information of the speech signal to be enhanced and the spectral information of the reference signal, extracts audio features from the concatenated spectral information, and performs two fully connected processing steps on the audio features to obtain echo cancellation masking information and noise reduction masking information; based on the speech signal to be enhanced... The signal enhancement request is processed by updating the echo cancellation amount represented by the echo cancellation masking information and the noise reduction amount represented by the noise reduction masking information to obtain updated echo cancellation masking information and updated noise reduction masking information. Using the updated echo cancellation masking information, echo cancellation is performed on the spectral information of the speech signal to be enhanced, resulting in echo-cancelled spectral information, which serves as the initial enhanced spectral information of the speech signal to be enhanced. Using the updated noise reduction masking information, noise reduction is performed on the initial enhanced spectral information of the speech signal to be enhanced, resulting in echo-cancelled and noise-reduced spectral information, which serves as the target enhanced spectral information of the speech signal to be enhanced. Finally, ISTFT processing is performed on the target enhanced spectral information of the speech signal to be enhanced to obtain the enhanced speech signal.
[0152] For example, in a live-streaming scenario, the host wants to demonstrate the crisp sound of his newly purchased mechanical keyboard. In this case, it is necessary to preserve the keyboard noise while continuously suppressing the echo. Based on this requirement, the terminal corresponding to the host performs corresponding post-processing on the echo cancellation masking matrix and noise reduction masking matrix output by the trained speech signal enhancement model. Then, using the processed echo cancellation masking matrix and noise reduction masking matrix, the speech signal to be enhanced is enhanced to obtain an enhanced speech signal that includes the host's voice and keyboard noise, but does not contain the echo.
[0153] In the aforementioned speech signal enhancement method, the following steps are taken: First, the speech signal to be enhanced and its reference signal are acquired. Then, the spectral information of the speech signal to be enhanced and the spectral information of the reference signal are input into a trained speech signal enhancement model to obtain echo cancellation masking information and noise reduction masking information. Next, based on the signal enhancement request of the speech signal to be enhanced, the echo cancellation masking information and noise reduction masking information are updated to obtain updated echo cancellation masking information and updated noise reduction masking information. Finally, based on the updated echo cancellation masking information and updated noise reduction masking information, the spectral information of the speech signal to be enhanced is enhanced to obtain the target enhanced spectral information of the speech signal to be enhanced. Finally, the target enhanced spectral information is transformed to obtain the enhanced speech signal of the speech signal to be enhanced. In this way, the trained speech signal enhancement model outputs echo cancellation masking information and noise reduction masking information separately, and performs corresponding post-processing on the echo cancellation masking information and noise reduction masking information to meet the speech signal enhancement requirements. This makes it less likely to cause speech distortion, thereby improving the speech signal enhancement effect. It avoids the defect of performing echo cancellation and noise reduction processing on the speech signal simultaneously in any scenario, which would lead to speech distortion and poor speech signal enhancement effect.
[0154] In an exemplary embodiment, step S520, which involves inputting the spectral information of the speech signal to be enhanced and the spectral information of the reference signal into the trained speech signal enhancement model to obtain echo cancellation masking information and noise reduction masking information, specifically includes the following: inputting the spectral information of the speech signal to be enhanced and the spectral information of the reference signal into the trained speech signal enhancement model for splicing processing to obtain spliced spectral information of the speech signal to be enhanced; performing feature extraction processing on the spliced spectral information of the speech signal to be enhanced to obtain audio features of the speech signal to be enhanced; and performing first classification processing and second classification processing on the audio features of the speech signal to be enhanced to obtain echo cancellation masking information and noise reduction masking information.
[0155] The first classification process is used to classify echo cancellation masking information; that is, the result of the first classification process is echo cancellation masking information. The second classification process is used to classify noise reduction masking information; that is, the result of the second classification process is noise reduction masking information.
[0156] Specifically, the terminal inputs the spectral information of the speech signal to be enhanced and the spectral information of the reference signal into the trained speech signal enhancement model. The trained speech signal enhancement model concatenates the spectral information of the speech signal to be enhanced and the spectral information of the reference signal to obtain the concatenated spectral information of the speech signal to be enhanced. Feature extraction is performed on the concatenated spectral information of the speech signal to be enhanced to obtain the extracted audio features, which are used as the audio features of the sample speech signal. The audio features of the speech signal to be enhanced are subjected to a first classification process to obtain the corresponding processing result, which is used as echo cancellation masking information. At the same time, the audio features of the speech signal to be enhanced are subjected to a second classification process to obtain the corresponding processing result, which is used as noise reduction masking information.
[0157] For example, the terminal inputs the spectral information of the speech signal to be enhanced and the spectral information of the reference signal into the trained speech signal enhancement model. The splicing layer in the trained speech signal enhancement model splices the spectral information of the speech signal to be enhanced and the spectral information of the reference signal to obtain the spliced spectral information of the speech signal to be enhanced. Multiple convolutional layers perform feature extraction processing on the spliced spectral information of the speech signal to be enhanced to obtain the initial audio features of the speech signal to be enhanced. A gated recurrent unit performs further feature extraction processing on the initial audio features of the speech signal to be enhanced to obtain the audio features of the speech signal to be enhanced. A first fully connected layer performs fully connected processing on the audio features of the speech signal to be enhanced to obtain the echo cancellation masking matrix. A second fully connected layer performs fully connected processing on the audio features of the speech signal to be enhanced to obtain the noise reduction masking matrix.
[0158] The technical solution provided in this disclosure inputs the spectral information of the speech signal to be enhanced and the spectral information of the reference signal into a trained speech signal enhancement model to obtain echo cancellation masking information and noise reduction masking information. In this way, by separating the echo cancellation function and the noise reduction function of the model, it is beneficial to process the echo cancellation masking information and noise reduction masking information output by the model separately in a corresponding manner to meet the enhancement requirements of the speech signal, and it is not easy for speech distortion to occur, thereby improving the enhancement effect of the speech signal.
[0159] In an exemplary embodiment, step S530, which updates the echo cancellation masking information and noise reduction masking information according to the signal enhancement request of the speech signal to be enhanced, to obtain updated echo cancellation masking information and updated noise reduction masking information, specifically includes the following: confirming the target echo cancellation amplitude and target noise reduction amplitude of the speech signal to be enhanced according to the signal enhancement request of the speech signal to be enhanced; updating the echo cancellation masking information according to the target echo cancellation amplitude to obtain updated echo cancellation masking information; and updating the noise reduction masking information according to the target noise reduction amplitude to obtain updated noise reduction masking information.
[0160] Among them, the target echo cancellation magnitude refers to the final echo cancellation magnitude of the speech signal to be enhanced, such as 100%; the target noise reduction magnitude refers to the final noise reduction magnitude of the speech signal to be enhanced, such as 50%.
[0161] Specifically, the terminal parses the signal enhancement request of the speech signal to be enhanced to obtain the target echo cancellation amplitude and the target noise reduction amplitude of the speech signal to be enhanced; updates the echo cancellation amplitude represented by the echo cancellation masking information according to the target echo cancellation amplitude, so that the echo cancellation amplitude represented by the echo cancellation masking information is the target echo cancellation amplitude, thereby obtaining the updated echo cancellation masking information; updates the noise reduction amplitude represented by the noise reduction masking information according to the target noise reduction amplitude, so that the noise reduction amplitude represented by the noise reduction masking information is the target noise reduction amplitude, thereby obtaining the updated noise reduction masking information.
[0162] The technical solution provided in this disclosure updates the echo cancellation masking information and noise reduction masking information according to the signal enhancement request of the speech signal to be enhanced, thereby obtaining updated echo cancellation masking information and updated noise reduction masking information. This achieves the goal of making the echo cancellation amount and noise reduction amount controllable, meeting the specific enhancement requirements of the speech signal, and avoiding the defect that simultaneous echo cancellation and noise reduction processing of the speech signal is performed every time the specific enhancement requirements of the speech signal cannot be considered, which leads to speech distortion and poor speech signal enhancement effect.
[0163] In an exemplary embodiment, step S540 above, which enhances the spectral information of the speech signal to be enhanced based on the updated echo cancellation masking information and the updated noise reduction masking information to obtain the target enhanced spectral information of the speech signal to be enhanced, specifically includes the following: fusing the spectral information of the speech signal to be enhanced and the updated echo cancellation masking information to obtain the initial enhanced spectral information of the speech signal to be enhanced; and fusing the initial enhanced spectral information and the updated noise reduction masking information to obtain the target enhanced spectral information of the speech signal to be enhanced.
[0164] Specifically, the terminal multiplies the spectral information of the speech signal to be enhanced with the echo cancellation masking information to obtain the corresponding spectral information, which is used as the initial enhanced spectral information of the speech signal to be enhanced; and multiplies the initial enhanced spectral information of the speech signal to be enhanced with the noise reduction masking information to obtain the corresponding spectral information, which is used as the target enhanced spectral information of the speech signal to be enhanced.
[0165] The technical solution provided in this disclosure enhances the spectral information of the speech signal to be enhanced based on the updated echo cancellation masking information and the updated noise reduction masking information, thereby obtaining the target enhanced spectral information of the speech signal to be enhanced. This is beneficial for meeting the specific enhancement requirements of the speech signal, and it is less likely to cause speech distortion, thereby improving the enhancement effect of the speech signal.
[0166] Figure 6 This is a flowchart illustrating another training method for a speech signal enhancement model according to an exemplary embodiment, such as... Figure 6 As shown, the training method for this speech signal enhancement model, used in a terminal, includes the following steps:
[0167] In step S610, the sample speech signal and the sample reference signal of the sample speech signal are acquired; the sample speech signal includes a clean speech signal, a noise signal and an echo signal.
[0168] In step S620, the spectral information of the sample speech signal and the spectral information of the sample reference signal are input into the speech signal enhancement model to be trained for splicing processing to obtain the spliced spectral information of the sample speech signal.
[0169] In step S630, feature extraction processing is performed on the spliced spectrum information of the sample speech signal to obtain the audio features of the sample speech signal.
[0170] In step S640, the audio features of the sample speech signal are subjected to first classification processing and second classification processing respectively to obtain predicted echo cancellation masking information and predicted noise reduction masking information.
[0171] In step S650, the spectral information of the sample speech signal and the predicted echo cancellation masking information are fused to obtain the initial enhanced spectral information of the sample speech signal.
[0172] In step S660, the initial enhanced spectrum information and the predicted noise reduction masking information are fused to obtain the target enhanced spectrum information of the sample speech signal.
[0173] In step S670, the spectral information of the clean speech signal and the spectral information of the noise signal are extracted from the spectral information of the sample speech signal; a first target spectral information is generated based on the spectral information of the clean speech signal and the spectral information of the noise signal, and a second target spectral information is generated based on the spectral information of the clean speech signal.
[0174] In step S680, a first loss value is obtained based on the difference between the initial enhanced spectrum information and the first target spectrum information, and a second loss value is obtained based on the difference between the target enhanced spectrum information and the second target spectrum information; the first loss value and the second loss value are fused to obtain the target loss value.
[0175] In step S690, the speech signal enhancement model to be trained is trained according to the target loss value until the training termination condition is met. The trained speech signal enhancement model that has reached the training termination condition is taken as the trained speech signal enhancement model.
[0176] In the training method of the above-mentioned speech signal enhancement model, the difference between the initial enhanced spectrum information and the first target spectrum information obtained based on the predicted echo cancellation masking information, and the difference between the target enhanced spectrum information and the second target spectrum information obtained based on the predicted noise reduction masking information, are fully utilized to assist in the training of the speech signal enhancement model. Since the echo cancellation function and the noise reduction function are separate, the trained speech signal enhancement model can output the echo cancellation masking information and the noise reduction masking information separately. This facilitates the subsequent processing of the separately output echo cancellation masking information and noise reduction masking information to meet the enhancement requirements of the speech signal, and it is not easy to cause speech distortion. This improves the enhancement effect of the speech signal and avoids the defect of performing echo cancellation and noise reduction processing on the speech signal simultaneously in any scenario, which would lead to speech distortion and poor enhancement effect of the speech signal.
[0177] To more clearly illustrate the training method of the speech signal enhancement model provided in this disclosure, a specific embodiment is used to describe the training method of the speech signal enhancement model. In one embodiment, this disclosure also provides a neural network training method that separates noise suppression and echo cancellation functions, which can support the use of post-processing to control the noise suppression amount and echo cancellation amount separately in certain complex scenarios. Specifically, it includes the following:
[0178] The first step is model training.
[0179] Figure 4 For the training framework used in this disclosure, the speech signal enhancement model to be trained (such as a deep neural network) needs to predict two masking matrices, namely the echo cancellation masking matrix M. aec and noise reduction masking matrix M ns First, the echo cancellation masking matrix M aec It will affect the near-end microphone signal spectrum S d On (t, f), a signal spectrum S with only echo cancellation was generated. aec (t, f), the signal spectrum S that undergoes echo cancellation only next.aec (t, f) will interact with the noise reduction masking matrix M ns The noise is suppressed by multiplication, thus obtaining the final enhanced signal spectrum S(t,f).
[0180] Furthermore, in the field of speech enhancement, DNN models often use convolutional layers and gated recurrent units to extract audio features, and finally use a fully connected layer for feature integration. In this training scheme, however, this disclosure uses... Figure 3 The two parallel fully connected layers shown are used to process the extracted audio features, and finally output two masking matrices, namely the echo cancellation masking matrix M. aec and noise reduction masking matrix M ns The loss function used during training comprises two main parts: echo cancellation loss and noise suppression loss. The known spectrum S of the near-end microphone signal... d (t, f) is derived from the speech spectrum S s (t, f), noise spectrum S n (t, f) and echo spectrum S e (t, f) constitutes S d (t, f) = S s (t,f)+S n (t,f)+S e If (t, f), then the total loss is:
[0181] Loss = loss(S) aec (t, f), Ss(t, f) + S n (t, f)) + loss (S (t, f), Ss (t, f));
[0182] The first part of the loss uses the noisy signal without echo as the target to supervise the effect of echo cancellation, while the second part of the loss uses clean speech as the target to supervise the effect of noise suppression.
[0183] The second step is model application.
[0184] During the inference process, the two masking matrices output by the trained speech signal enhancement model are constrained through post-processing, thereby achieving controllable echo cancellation and noise reduction. For example, the minimum value of the noise reduction masking matrix is set to be no lower than -20dB, i.e., the minimum value of the noise reduction masking matrix is no lower than 0.1, thus achieving a noise suppression effect of no more than -20dB.
[0185] In the aforementioned neural network training method that separates noise suppression and echo cancellation functions, the trained speech signal enhancement model can simultaneously predict a noise reduction masking matrix and an echo cancellation masking matrix, and obtain the enhanced speech signal using a concatenated approach. During inference, post-processing is used to separately control the noise reduction and echo cancellation amounts.
[0186] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0187] It is understood that the same / similar parts between the various embodiments of the methods described above in this specification can be referred to each other. Each embodiment focuses on the differences from other embodiments, and relevant parts can be referred to the description of other method embodiments.
[0188] Based on the same inventive concept, this disclosure also provides a training apparatus for a speech signal enhancement model for implementing the training method of the speech signal enhancement model mentioned above.
[0189] Figure 7 This is a block diagram illustrating a training apparatus for a speech signal enhancement model according to an exemplary embodiment. (Refer to...) Figure 7 The device includes a sample acquisition unit 710, an information prediction unit 720, a spectrum processing unit 730, and a model training unit 740.
[0190] The sample acquisition unit 710 is configured to acquire a sample speech signal and a sample reference signal of the sample speech signal; the sample speech signal includes a clean speech signal, a noise signal and an echo signal.
[0191] The information prediction unit 720 is configured to input the spectral information of the sample speech signal and the spectral information of the sample reference signal into the speech signal enhancement model to be trained, and obtain the predicted echo cancellation masking information and the predicted noise reduction masking information.
[0192] The spectrum processing unit 730 is configured to perform the following operations: obtaining initial enhanced spectrum information of the sample speech signal based on the spectrum information of the sample speech signal and the predicted echo cancellation masking information; and obtaining target enhanced spectrum information of the sample speech signal based on the initial enhanced spectrum information and the predicted noise reduction masking information.
[0193] The model training unit 740 is configured to train the speech signal enhancement model to be trained based on the difference between the initial enhanced spectral information and the first target spectral information, as well as the difference between the target enhanced spectral information and the second target spectral information, to obtain the trained speech signal enhancement model; the first target spectral information includes the spectral information of the clean speech signal and the spectral information of the noise signal, and the second target spectral information includes the spectral information of the clean speech signal.
[0194] In an exemplary embodiment, the information prediction unit 720 is further configured to perform concatenation processing on the spectral information of the sample speech signal and the spectral information of the sample reference signal, which are input into the speech signal enhancement model to be trained, to obtain the concatenated spectral information of the sample speech signal; to perform feature extraction processing on the concatenated spectral information of the sample speech signal to obtain the audio features of the sample speech signal; and to perform a first classification processing and a second classification processing on the audio features of the sample speech signal to obtain predicted echo cancellation masking information and predicted noise reduction masking information, respectively; the first classification processing is used to classify the predicted echo cancellation masking information, and the second classification processing is used to classify the predicted noise reduction masking information.
[0195] In an exemplary embodiment, the model training unit 740 is further configured to perform the following operations: obtaining a first loss value based on the difference between the initial enhanced spectral information and the first target spectral information; obtaining a second loss value based on the difference between the target enhanced spectral information and the second target spectral information; fusing the first loss value and the second loss value to obtain a target loss value; training the speech signal enhancement model to be trained based on the target loss value until the training termination condition is met; and using the trained speech signal enhancement model that has met the training termination condition as the trained speech signal enhancement model.
[0196] In an exemplary embodiment, the spectrum processing unit 730 is further configured to perform a fusion process of the spectrum information of the sample speech signal and the predicted echo cancellation masking information to obtain the initial enhanced spectrum information of the sample speech signal; and to perform a fusion process of the initial enhanced spectrum information and the predicted noise reduction masking information to obtain the target enhanced spectrum information of the sample speech signal.
[0197] In an exemplary embodiment, the spectral information of the sample speech signal includes the spectral information of the clean speech signal, the spectral information of the noise signal, and the spectral information of the echo signal;
[0198] The training apparatus for the speech signal enhancement model provided in this disclosure further includes an information generation unit configured to perform the following actions: extracting the spectral information of a clean speech signal and the spectral information of a noise signal from the spectral information of a sample speech signal; generating a first target spectral information based on the spectral information of the clean speech signal and the spectral information of the noise signal; and generating a second target spectral information based on the spectral information of the clean speech signal.
[0199] Based on the same inventive concept, this disclosure also provides a speech signal enhancement apparatus for implementing the speech signal enhancement method described above.
[0200] Figure 8 This is a block diagram illustrating a speech signal enhancement device according to an exemplary embodiment. (Refer to...) Figure 8 The device includes a signal acquisition unit 810, an information processing unit 820, an information update unit 830, an information enhancement unit 840, and an information transformation unit 850.
[0201] The signal acquisition unit 810 is configured to acquire the speech signal to be enhanced and a reference signal for the speech signal to be enhanced.
[0202] The information processing unit 820 is configured to input the spectral information of the speech signal to be enhanced and the spectral information of the reference signal into the trained speech signal enhancement model to obtain echo cancellation masking information and noise reduction masking information; the trained speech signal enhancement model is trained according to the training method of the speech signal enhancement model in any embodiment of the first aspect.
[0203] The information update unit 830 is configured to update the echo cancellation masking information and noise reduction masking information according to the signal enhancement request of the speech signal to be enhanced, so as to obtain the updated echo cancellation masking information and the updated noise reduction masking information.
[0204] The information enhancement unit 840 is configured to perform enhancement processing on the spectral information of the speech signal to be enhanced based on the updated echo cancellation masking information and the updated noise reduction masking information, so as to obtain the target enhanced spectral information of the speech signal to be enhanced.
[0205] The information transformation unit 850 is configured to transform the target enhancement spectrum information to obtain the enhanced speech signal of the speech signal to be enhanced.
[0206] In an exemplary embodiment, the information processing unit 820 is further configured to perform splicing processing on the spectral information of the speech signal to be enhanced and the spectral information of the reference signal, which are input into a trained speech signal enhancement model, to obtain spliced spectral information of the speech signal to be enhanced; to perform feature extraction processing on the spliced spectral information of the speech signal to be enhanced to obtain audio features of the speech signal to be enhanced; and to perform a first classification processing and a second classification processing on the audio features of the speech signal to be enhanced to obtain echo cancellation masking information and noise reduction masking information, respectively; the first classification processing is used to classify the echo cancellation masking information, and the second classification processing is used to classify the noise reduction masking information.
[0207] In one exemplary embodiment, the information update unit 830 is further configured to perform the following actions based on the signal enhancement request of the speech signal to be enhanced: confirming the target echo cancellation amplitude and the target noise reduction amplitude of the speech signal to be enhanced; updating the echo cancellation masking information according to the target echo cancellation amplitude to obtain updated echo cancellation masking information; and updating the noise reduction masking information according to the target noise reduction amplitude to obtain updated noise reduction masking information.
[0208] In an exemplary embodiment, the information enhancement unit 840 is further configured to perform a fusion process of the spectral information of the speech signal to be enhanced and the updated echo cancellation masking information to obtain initial enhanced spectral information of the speech signal to be enhanced; and to perform a fusion process of the initial enhanced spectral information and the updated noise reduction masking information to obtain target enhanced spectral information of the speech signal to be enhanced.
[0209] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0210] The training device for the aforementioned speech signal enhancement model, or each module in the speech signal enhancement device, can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0211] Figure 9 This is a block diagram illustrating an electronic device 900 for implementing a training method or a speech signal enhancement method for a speech signal enhancement model, according to an exemplary embodiment. For example, the electronic device 900 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0212] Reference Figure 9The electronic device 900 may include one or more of the following components: processing component 902, memory 904, power supply component 906, multimedia component 908, audio component 910, input / output (I / O) interface 912, sensor component 914, and communication component 916.
[0213] Processing component 902 typically controls the overall operation of electronic device 900, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 902 may include one or more processors 920 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 902 may include one or more modules to facilitate interaction between processing component 902 and other components. For example, processing component 902 may include a multimedia module to facilitate interaction between multimedia component 908 and processing component 902.
[0214] Memory 904 is configured to store various types of data to support the operation of electronic device 900. Examples of this data include instructions for any application or method operating on electronic device 900, contact data, phonebook data, messages, pictures, videos, etc. Memory 904 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, optical disk, or graphene memory.
[0215] Power supply component 906 provides power to various components of electronic device 900. Power supply component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 900.
[0216] Multimedia component 908 includes a screen that provides an output interface between the electronic device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 908 includes a front-facing camera and / or a rear-facing camera. When the electronic device 900 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0217] Audio component 910 is configured to output and / or input audio signals. For example, audio component 910 includes a microphone (MIC) configured to receive external audio signals when electronic device 900 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 904 or transmitted via communication component 916. In some embodiments, audio component 910 also includes a speaker for outputting audio signals.
[0218] I / O interface 912 provides an interface between processing component 902 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0219] Sensor assembly 914 includes one or more sensors for providing state assessments of various aspects of electronic device 900. For example, sensor assembly 914 can detect the on / off state of electronic device 900, the relative positioning of components such as the display and keypad of electronic device 900, changes in position of electronic device 900 or its components, the presence or absence of user contact with electronic device 900, orientation or acceleration / deceleration of device 900, and temperature changes of electronic device 900. Sensor assembly 914 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 914 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 914 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0220] Communication component 916 is configured to facilitate wired or wireless communication between electronic device 900 and other devices. Electronic device 900 can access wireless networks based on communication standards, such as WiFi, carrier networks (such as 2G, 3G, 4G, or 5G), or combinations thereof. In one exemplary embodiment, communication component 916 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 916 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0221] In an exemplary embodiment, the electronic device 900 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0222] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 904 including instructions, which can be executed by a processor 920 of an electronic device 900 to perform the above-described method. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0223] In an exemplary embodiment, a computer program product is also provided, the computer program product including instructions that can be executed by a processor 920 of an electronic device 900 to perform the above-described method.
[0224] It should be noted that the above-mentioned apparatus, electronic equipment, computer-readable storage medium, computer program product, etc., may also include other implementation methods according to the description of the method embodiments. For specific implementation methods, please refer to the description of the relevant method embodiments, which will not be elaborated here.
[0225] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.
[0226] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A training method for a speech signal enhancement model, characterized in that, include: Acquire a sample speech signal and a sample reference signal of the sample speech signal; the sample speech signal includes a clean speech signal, a noise signal, and an echo signal; the sample speech signal refers to the speech signal used to train the speech signal enhancement model; the sample reference signal refers to a far-end reference signal; The spectral information of the sample speech signal and the spectral information of the sample reference signal are input into the speech signal enhancement model to be trained to obtain the predicted echo cancellation masking information and the predicted noise reduction masking information. Based on the spectral information of the sample speech signal and the predicted echo cancellation masking information, the initial enhanced spectral information of the sample speech signal is obtained, and based on the initial enhanced spectral information and the predicted noise reduction masking information, the target enhanced spectral information of the sample speech signal is obtained. Based on the difference between the initial enhanced spectrum information and the first target spectrum information, and the difference between the target enhanced spectrum information and the second target spectrum information, the speech signal enhancement model to be trained is trained to obtain a trained speech signal enhancement model; the first target spectrum information includes the spectrum information of the clean speech signal and the spectrum information of the noise signal, and the second target spectrum information includes the spectrum information of the clean speech signal.
2. The method according to claim 1, characterized in that, The step of inputting the spectral information of the sample speech signal and the spectral information of the sample reference signal into the speech signal enhancement model to be trained to obtain predicted echo cancellation masking information and predicted noise reduction masking information includes: The spectral information of the sample speech signal and the spectral information of the sample reference signal are input into the speech signal enhancement model to be trained for splicing processing to obtain the spliced spectral information of the sample speech signal. The spliced spectral information of the sample speech signal is processed by feature extraction to obtain the audio features of the sample speech signal; The audio features of the sample speech signal are subjected to a first classification process and a second classification process respectively to obtain the predicted echo cancellation masking information and the predicted noise reduction masking information; the first classification process is used to classify the predicted echo cancellation masking information, and the second classification process is used to classify the predicted noise reduction masking information.
3. The method according to claim 1, characterized in that, The step of training the speech signal enhancement model to be trained based on the difference between the initial enhanced spectral information and the first target spectral information, and the difference between the target enhanced spectral information and the second target spectral information, to obtain the trained speech signal enhancement model, includes: A first loss value is obtained based on the difference between the initial enhanced spectrum information and the first target spectrum information, and a second loss value is obtained based on the difference between the target enhanced spectrum information and the second target spectrum information; The first loss value and the second loss value are fused together to obtain the target loss value; The speech signal enhancement model to be trained is trained according to the target loss value until the training termination condition is met. The trained speech signal enhancement model that has met the training termination condition is taken as the trained speech signal enhancement model.
4. The method according to claim 1, characterized in that, The step of obtaining initial enhanced spectral information of the sample speech signal based on the spectral information of the sample speech signal and the predicted echo cancellation masking information, and obtaining target enhanced spectral information of the sample speech signal based on the initial enhanced spectral information and the predicted noise reduction masking information, includes: The spectral information of the sample speech signal and the predicted echo cancellation masking information are fused to obtain the initial enhanced spectral information of the sample speech signal. The initial enhanced spectrum information and the predicted noise reduction masking information are fused together to obtain the target enhanced spectrum information of the sample speech signal.
5. The method according to any one of claims 1 to 4, characterized in that, The spectral information of the sample speech signal includes the spectral information of the clean speech signal, the spectral information of the noise signal, and the spectral information of the echo signal; Before training the speech signal enhancement model to be trained based on the difference between the initial enhanced spectral information and the first target spectral information, and the difference between the target enhanced spectral information and the second target spectral information, to obtain the trained speech signal enhancement model, the method further includes: From the spectral information of the sample speech signal, extract the spectral information of the clean speech signal and the spectral information of the noise signal; The first target spectrum information is generated based on the spectrum information of the clean speech signal and the spectrum information of the noise signal, and the second target spectrum information is generated based on the spectrum information of the clean speech signal.
6. A method for enhancing speech signals, characterized in that, include: Acquire the speech signal to be enhanced and a reference signal for the speech signal to be enhanced; The reference signal for the speech signal to be enhanced refers to the far-end reference signal; The spectral information of the speech signal to be enhanced and the spectral information of the reference signal are input into the trained speech signal enhancement model to obtain echo cancellation masking information and noise reduction masking information. The trained speech signal enhancement model is obtained by training according to any one of claims 1 to 5; Based on the signal enhancement request of the speech signal to be enhanced, the echo cancellation masking information and the noise reduction masking information are updated to obtain the updated echo cancellation masking information and the updated noise reduction masking information. Based on the updated echo cancellation masking information and the updated noise reduction masking information, the spectral information of the speech signal to be enhanced is enhanced to obtain the target enhanced spectral information of the speech signal to be enhanced. The target enhancement spectrum information is transformed to obtain the enhanced speech signal of the speech signal to be enhanced.
7. The method according to claim 6, characterized in that, The step of inputting the spectral information of the speech signal to be enhanced and the spectral information of the reference signal into the trained speech signal enhancement model to obtain echo cancellation masking information and noise reduction masking information includes: The spectral information of the speech signal to be enhanced and the spectral information of the reference signal are input into the trained speech signal enhancement model and spliced together to obtain the spliced spectral information of the speech signal to be enhanced. Feature extraction processing is performed on the spliced spectral information of the speech signal to be enhanced to obtain the audio features of the speech signal to be enhanced; The audio features of the speech signal to be enhanced are subjected to a first classification process and a second classification process, respectively, to obtain the echo cancellation masking information and the noise reduction masking information; the first classification process is used to classify the echo cancellation masking information, and the second classification process is used to classify the noise reduction masking information.
8. The method according to claim 6, characterized in that, The step of updating the echo cancellation masking information and the noise reduction masking information according to the signal enhancement request of the speech signal to be enhanced, to obtain updated echo cancellation masking information and updated noise reduction masking information, includes: Based on the signal enhancement request of the speech signal to be enhanced, the target echo cancellation amplitude and target noise reduction amplitude of the speech signal to be enhanced are confirmed. The echo cancellation masking information is updated according to the target echo cancellation amplitude to obtain updated echo cancellation masking information, and the noise reduction masking information is updated according to the target noise reduction amplitude to obtain updated noise reduction masking information.
9. The method according to any one of claims 6 to 8, characterized in that, The step of enhancing the spectral information of the speech signal to be enhanced based on the updated echo cancellation masking information and the updated noise reduction masking information to obtain the target enhanced spectral information of the speech signal to be enhanced includes: The spectral information of the speech signal to be enhanced and the updated echo cancellation masking information are fused together to obtain the initial enhanced spectral information of the speech signal to be enhanced. The initial enhanced spectrum information and the updated noise reduction masking information are fused together to obtain the target enhanced spectrum information of the speech signal to be enhanced.
10. A training device for a speech signal enhancement model, characterized in that, include: The sample acquisition unit is configured to acquire a sample speech signal and a sample reference signal of the sample speech signal; the sample speech signal includes a clean speech signal, a noise signal, and an echo signal; the sample speech signal refers to the speech signal used to train the speech signal enhancement model; the sample reference signal refers to a far-end reference signal. The information prediction unit is configured to input the spectral information of the sample speech signal and the spectral information of the sample reference signal into the speech signal enhancement model to be trained, and obtain predicted echo cancellation masking information and predicted noise reduction masking information. The spectrum processing unit is configured to perform the following operations: obtaining initial enhanced spectrum information of the sample speech signal based on the spectrum information of the sample speech signal and the predicted echo cancellation masking information; and obtaining target enhanced spectrum information of the sample speech signal based on the initial enhanced spectrum information and the predicted noise reduction masking information. The model training unit is configured to train the speech signal enhancement model to be trained based on the difference between the initial enhanced spectral information and the first target spectral information, and the difference between the target enhanced spectral information and the second target spectral information, to obtain a trained speech signal enhancement model; the first target spectral information includes the spectral information of the clean speech signal and the spectral information of the noise signal, and the second target spectral information includes the spectral information of the clean speech signal.
11. A voice signal enhancement device, characterized in that, include: The signal acquisition unit is configured to acquire the speech signal to be enhanced and a reference signal of the speech signal to be enhanced; The reference signal for the speech signal to be enhanced refers to the far-end reference signal; The information processing unit is configured to input the spectral information of the speech signal to be enhanced and the spectral information of the reference signal into the trained speech signal enhancement model to obtain echo cancellation masking information and noise reduction masking information. The trained speech signal enhancement model is obtained by training according to any one of claims 1 to 5; The information update unit is configured to perform an update of the echo cancellation masking information and the noise reduction masking information based on the signal enhancement request of the speech signal to be enhanced, so as to obtain the updated echo cancellation masking information and the updated noise reduction masking information. The information enhancement unit is configured to perform enhancement processing on the spectral information of the speech signal to be enhanced based on the updated echo cancellation masking information and the updated noise reduction masking information, to obtain the target enhanced spectral information of the speech signal to be enhanced; The information transformation unit is configured to transform the target enhancement spectrum information to obtain the enhanced speech signal of the speech signal to be enhanced.
12. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the training method of the speech signal enhancement model as described in any one of claims 1 to 5, or the speech signal enhancement method as described in any one of claims 6 to 9.
13. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the training method of the speech signal enhancement model as described in any one of claims 1 to 5, or the speech signal enhancement method as described in any one of claims 6 to 9.