Training Method of Speech Enhancement Model and Speech Enhancement Method

By training the combined model of speech extractor and classifier, the problem that the speech enhancement system cannot filter and interfere with human voice is solved, and the speech signal of the specified speech object is accurately extracted from multiple speech object signals, improving the listening quality and intelligibility.

CN114758668BActive Publication Date: 2025-07-04BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210435868.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-24
Publication Date
2025-07-04
Estimated Expiration
2042-04-24

AI Technical Summary

Technical Problem

The existing voice enhancement system cannot effectively filter and interfere with the voice signal when multiple speaking objects speak at the same time, resulting in a decrease in the listening quality and intelligibility of the speech signal of the specified speaking object.

Method used

Using a combined model of speech extractor, speech characterization extractor and classifier, the pure speech signal and overlapping speech band noise signals in the sample are trained, and the cross entropy loss and scale-invariant signal to distortion ratio loss function is used to optimize the model to accurately extract the speech signal of the specified speech object.

Benefits of technology

It improves the listening quality and intelligibility of the voice signal of the designated speaking object, effectively filters interfering with the voice signal, and ensures the clarity of the voice signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114758668B_ABST
    Figure CN114758668B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a training method and a voice enhancement method for a voice enhancement model, including: obtaining training samples of multiple speakers; inputting the first clean voice signal samples of each speaker into a voice feature extractor; inputting the voice features of each speaker into a classifier; inputting the voice features of each speaker and the magnitude spectrum of the overlapping voice noisy signal samples into a voice extractor, and determining the predicted enhanced voice signal of the speaker according to the predicted magnitude spectrum mask of the enhanced voice signal of the speaker; calculating a loss according to the enhanced voice signal corresponding to each speaker, the second clean voice signal sample, the identification prediction result, and the identification label; adjusting the parameters of the voice extractor, the voice feature extractor, and the classifier through the loss to train the voice enhancement model. In this way, the trained voice enhancement model can accurately extract the voice signal of a specified speaker from the voice signals of multiple speakers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and more particularly, to a method for training a voice enhancement model, a voice enhancement method, an apparatus, an electronic device, a speaker, and a storage medium. Background Art

[0002] With the development of deep learning technologies, voice enhancement systems in related technologies can remove noise signals from noisy audio signals, thereby improving the auditory quality and intelligibility of non-noise signals.

[0003] However, voice enhancement systems in related technologies can only filter out non-interfering human voice signals such as environmental noise in noisy audio signals, and cannot filter out interfering human voice signals in noisy audio signals. For example, if there are multiple speakers speaking simultaneously in the current environment, the voice enhancement system in related technologies will retain the voice signals of all speakers. At this time, the unfiltered interfering human voice signals will interfere with the voice signals of the designated speaker, reducing the auditory quality and intelligibility of the voice signals of the designated speaker. Summary of the Invention

[0004] The present disclosure provides a method for training a voice enhancement model, a voice enhancement method, an apparatus, an electronic device, a speaker, and a storage medium, so as to at least solve the problem in the above related technologies that the voice enhancement system will retain the voice signals of all speakers, and at this time, the unfiltered interfering human voice signals will interfere with the voice signals of the designated speaker, reducing the auditory quality and intelligibility of the voice signals of the designated speaker.

[0005] According to a first aspect of the embodiments of the present disclosure, there is provided a method for training a voice enhancement model, where the voice enhancement model includes a voice extractor, a voice feature extractor, and a classifier. The training method includes: obtaining training samples of multiple speakers, where the training samples of each speaker include: a first clean voice signal sample of the speaker, a second clean voice signal sample of the speaker, an overlapping voice noisy signal sample of the speaker, and an identification label, where the overlapping voice noisy signal sample of the speaker is obtained by superimposing the voice signals of at least one other speaker on the second clean voice signal sample of the speaker; inputting the first clean voice signal sample of each speaker into the voice feature extractor to obtain the voice feature of the speaker; inputting the voice feature of each speaker into the classifier to obtain the identification prediction result of the speaker; inputting the voice feature of each speaker and the amplitude spectrum of the overlapping voice noisy signal sample into the voice extractor to obtain the amplitude spectrum mask of the predicted enhanced voice signal of the speaker, and determining the predicted enhanced voice signal of the speaker according to the amplitude spectrum mask of the predicted enhanced voice signal of the speaker, where the enhanced voice signal is the voice signal of the speaker extracted from the overlapping voice noisy signal sample of the speaker; calculating a loss according to the enhanced voice signal corresponding to each speaker, the second clean voice signal sample, the identification prediction result, and the identification label; adjusting the parameters of the voice extractor, the voice feature extractor, and the classifier through the loss to train the voice enhancement model.

[0006] Optionally, calculating the loss according to the enhanced voice signal corresponding to each speaker, the second clean voice signal sample, the identification prediction result, and the identification label includes: calculating a first loss according to the identification prediction result and the identification label corresponding to each speaker; calculating a second loss according to the enhanced voice signal and the second clean voice signal sample corresponding to each speaker; and performing a weighted sum of the first loss and the second loss to obtain the loss.

[0007] Optionally, the first loss is calculated by a cross-entropy loss function.

[0008] Optionally, the second loss is calculated by a scale-invariant signal-to-distortion ratio loss function.

[0009] Optionally, the voice feature extractor and the classifier are obtained through a first pre-training; the voice extractor is obtained through a second pre-training, where the first pre-training and the second pre-training are independent of each other.

[0010] Optionally, determining the predicted enhanced voice signal of the speaker according to the amplitude spectrum mask of the predicted enhanced voice signal of the speaker includes:

[0011] Multiply the amplitude spectrum of the noisy overlapping speech signal sample of the speaking object by the amplitude spectrum mask of the predicted enhanced speech signal of the speaking object to obtain the amplitude spectrum of the predicted enhanced speech signal of the speaking object;

[0012] Combine the amplitude spectrum of the predicted enhanced speech signal of the speaking object with the phase spectrum of the noisy overlapping speech signal sample and perform inverse short-time Fourier transform to obtain the predicted enhanced speech signal of the speaking object.

[0013] Optionally, the inputting the first clean speech signal sample of each speaking object into the speech feature extractor includes:

[0014] Perform short-time Fourier transform on the first clean speech signal sample of each speaking object to obtain the amplitude spectrum of the first clean speech signal sample of the speaking object;

[0015] Obtain the Mel cepstral features corresponding to the amplitude spectrum of the first clean speech signal sample of the speaking object;

[0016] Input the Mel cepstral features corresponding to the amplitude spectrum of the first clean speech signal sample of the speaking object into the speech feature extractor.

[0017] Optionally, the first clean speech signal sample is different from the second clean speech signal sample.

[0018] According to a second aspect of the embodiments of the present disclosure, a speech enhancement method is provided. The speech enhancement method includes: obtaining a measured speech signal; inputting the speech signal of a pre-registered target speaking object into a speech feature extractor trained according to the training method of the present disclosure to obtain the speech feature of the target speaking object; inputting the speech feature of the target speaking object and the measured speech signal into a speech extractor trained according to the training method of the present disclosure to obtain the speech signal of the target speaking object extracted from the measured speech signal.

[0019] According to a third aspect of the embodiments of the present disclosure, there is provided a training device for a voice enhancement model, where the voice enhancement model includes a voice extractor, a voice feature extractor, and a classifier. The training device includes: an acquisition module configured to acquire training samples of multiple speakers. Each training sample of a speaker includes: a first clean voice signal sample of the speaker, a second clean voice signal sample of the speaker, an overlapping voice noisy signal sample of the speaker, and an identification label, where the overlapping voice noisy signal sample of the speaker is obtained by superimposing the voice signals of at least one other speaker on the second clean voice signal sample of the speaker; a first input module configured to input the first clean voice signal sample of each speaker into the voice feature extractor to obtain the voice feature of the speaker; a second input module configured to input the voice feature of each speaker into the classifier to obtain the identification prediction result of the speaker; a third input module configured to input the voice feature of each speaker and the amplitude spectrum of the overlapping voice noisy signal sample into the voice extractor to obtain the amplitude spectrum mask of the predicted enhanced voice signal of the speaker, and determine the predicted enhanced voice signal of the speaker according to the amplitude spectrum mask of the predicted enhanced voice signal of the speaker, where the enhanced voice signal is the voice signal of the speaker extracted from the overlapping voice noisy signal sample of the speaker; and adjust the parameters of the voice extractor, the voice feature extractor, and the classifier through the loss to train the voice enhancement model.

[0020] Optionally, the calculation module is configured to: calculate a first loss according to the identification prediction result and the identification label corresponding to each speaker; calculate a second loss according to the enhanced voice signal corresponding to each speaker and the second clean voice signal sample; and perform a weighted sum of the first loss and the second loss to obtain the loss.

[0021] Optionally, the first loss is calculated by a cross-entropy loss function.

[0022] Optionally, the second loss is calculated by a scale-invariant signal-to-distortion ratio loss function.

[0023] Optionally, the voice feature extractor and the classifier are obtained through a first pre-training; the voice extractor is obtained through a second pre-training, where the first pre-training and the second pre-training are independent of each other.

[0024] Optionally, the third input module is configured to:

[0025] multiply the amplitude spectrum of the overlapping voice noisy signal sample of the speaker by the amplitude spectrum mask of the predicted enhanced voice signal of the speaker to obtain the amplitude spectrum of the predicted enhanced voice signal of the speaker;

[0026] Combine the amplitude spectrum of the enhanced speech signal of the predicted speaker with the phase spectrum of the overlapping speech band-noisy signal sample and perform an inverse short-time Fourier transform to obtain the enhanced speech signal of the predicted speaker.

[0027] Optionally, the first input module is configured to:

[0028] Perform a short-time Fourier transform on the first clean speech signal sample of each speaker to obtain the amplitude spectrum of the first clean speech signal sample of the speaker;

[0029] Obtain the Mel cepstral features corresponding to the amplitude spectrum of the first clean speech signal sample of the speaker;

[0030] Input the Mel cepstral features corresponding to the amplitude spectrum of the first clean speech signal sample of the speaker into the speech feature extractor.

[0031] Optionally, the first clean speech signal sample is different from the second clean speech signal sample.

[0032] According to a fourth aspect of the embodiments of the present disclosure, there is provided a speech enhancement device, including: an acquisition module configured to acquire a measured speech signal; a first input module configured to input the speech signal of a pre-registered target speaker into a speech feature extractor trained by a training device according to the present disclosure to obtain the speech feature of the target speaker; a second input module configured to input the speech feature of the target speaker and the measured speech signal into a speech extractor trained by the training device according to the present disclosure to obtain the speech signal of the target speaker extracted from the measured speech signal.

[0033] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the training method or the speech enhancement method of the speech enhancement model according to the present disclosure.

[0034] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the training method or the speech enhancement method of the speech enhancement model according to the present disclosure.

[0035] According to a seventh aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the training method or the speech enhancement method of the speech enhancement model according to the present disclosure.

[0036] According to an eighth aspect of the embodiments of the present disclosure, there is provided a speaker, including a voice enhancement device according to the present disclosure.

[0037] According to a ninth aspect of the embodiments of the present disclosure, there is provided a speaker, including: at least one processor; at least one memory storing computer-executable instructions, wherein when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the voice enhancement method according to the present disclosure.

[0038] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0039] In this way, in the training method of the voice enhancement model of the present disclosure, since a large number of pure voice signal samples of the speaking object and overlapping voice noisy signal samples are used to train the voice enhancement model, the trained voice enhancement model can accurately extract the voice signal of the specified speaking object from the voice signals of multiple speaking objects, avoiding the interference of interfering human voice signals on the voice signal of the specified speaking object, and improving the listening quality and intelligibility of the voice signal of the specified speaking object. Moreover, the voice enhancement method of the present disclosure can, with the help of the voice signal of the target speaking object pre-registered, use the trained voice enhancement model to extract the potential voice signal of the target speaking object from the measured voice signal, avoiding the interference of interfering human voice signals on the voice signal of the specified speaking object, achieving the effect of voice enhancement for the specified speaking object, and improving the listening quality and intelligibility of the voice signal of the specified speaking object.

[0040] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0042] Figure 1 is a flowchart showing a method for training a voice enhancement model according to an exemplary embodiment of the present disclosure;

[0043] Figure 2 is a schematic diagram showing a training stage of a voice enhancement model according to an exemplary embodiment of the present disclosure;

[0044] Figure 3 is a flowchart showing a voice enhancement method according to an exemplary embodiment of the present disclosure;

[0045] Figure 4It is a schematic diagram showing the test phase of a voice enhancement model according to an exemplary embodiment of the present disclosure;

[0046] Figure 5 It is a block diagram showing a training device of a voice enhancement model according to an exemplary embodiment of the present disclosure;

[0047] Figure 6 It is a block diagram showing a voice enhancement device according to an exemplary embodiment of the present disclosure;

[0048] Figure 7 It is a block diagram showing an electronic device according to an exemplary embodiment of the present disclosure;

[0049] Figure 8 It is a block diagram showing the structure of a speaker according to an exemplary embodiment of the present disclosure;

[0050] Figure 9 It is a block diagram showing the structure of a speaker according to another exemplary embodiment of the present disclosure. Detailed implementation manners

[0051] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0052] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0053] It should be noted here that "at least one of several items" in the present disclosure all represents the three types of parallel situations including "any one of the several items", "any combination of several items of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "performing at least one of step one and step two" means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing step one and step two.

[0054] Figure 1It is a flowchart showing a method for training a voice enhancement model according to an exemplary embodiment of the present disclosure. The voice enhancement model may include a voice extractor, a voice feature extractor, and a classifier.

[0055] Referring to Figure 1 , in step 101, training samples of multiple speaking objects can be obtained. Among them, the training samples of each speaking object may include: the first clean speech signal sample x(t), the second clean speech signal sample s(t), the overlapping speech noisy signal sample y(t), and the identification label (label) of the speaking object. The overlapping speech noisy signal sample y(t) of the speaking object is obtained by superimposing the speech signals of at least one other speaking object on the second clean speech signal sample s(t) of the speaking object. Further, the overlapping speech noisy signal sample y(t) of the speaking object may also contain different types of noise. For example, noise data may be added to the speech signal of the speaking object according to different signal-to-noise ratios (SNRs), and the noise may be the room impulse response (RIR) of different rooms, etc. The overlapping speech noisy signal sample y(t) of the speaking object may also be a speech noisy signal sample of a single speaking object. It should be noted that the speaking object may be an object with voice output ability. For example, it may be a real speaker or a virtual speaker, etc.

[0056] According to an exemplary embodiment of the present disclosure, the first clean speech signal sample x(t) is different from the second clean speech signal sample s(t), that is, the first clean speech signal sample x(t) and the second clean speech signal sample s(t) may be speech signal samples with different contents of the same speaking object, and both the first clean speech signal sample x(t) and the second clean speech signal sample s(t) are single-speaking-object speech signal samples with a higher SNR.

[0057] Further, the reason for requiring the first clean speech signal sample x(t) and the second clean speech signal sample s(t) to be speech signal samples with different contents of the same speaking object is as follows: Subsequently, it is necessary to use the pre-registered speech signal of the specified speaking object to extract the potential speech signal of the specified speaking object from the measured speech signal by using the trained voice enhancement model. However, in an actual test scenario, the pre-registered speech signal of each speaking object is often different from the potential speech signal of the specified speaking object in the measured speech signal.

[0058] For example, all employees of a company can pre-register their respective voice signals. During an online meeting among the company's employees, the content of each participant's speech during the meeting must be different from the content of their pre-registered voice signal. Therefore, during the training phase, in order to simulate the real test scenario as much as possible, the first clean voice signal sample x(t) and the second clean voice signal sample s(t) used for training can be set as voice signal samples with different contents, so as to ensure that the voice enhancement model can extract the potential voice signal of the specified speaker from the measured voice signal as accurately as possible in the actual test scenario.

[0059] It should be noted that during the training phase of the voice enhancement model, voice data of thousands of speakers are used to train a voice extractor, a voice feature extractor, and a classifier. The voice extractor, voice feature extractor, and classifier can all be constructed based on neural networks.

[0060] In step 102, the first clean voice signal sample x(t) of each speaker can be input into the voice feature extractor to obtain the voice feature of the speaker, where the voice feature of the speaker can be the voiceprint feature of the speaker.

[0061] According to an exemplary embodiment of the present disclosure, the first clean voice signal sample x(t) of each speaker can be subjected to a Short Time Fourier Transform (STFT) to obtain the magnitude spectrum |X(n,k)| of the first clean voice signal sample x(t) of the speaker.

[0062] It should be noted that if the original audio signal with a length of T is represented as x(t) in the time domain, where t represents time and 0 < t ≤ T, then after the short-time Fourier transform, x(t) can be represented in the time-frequency domain as:

[0063] X(n,k) = STFT(x(t))

[0064] where n is the frame sequence, 0 < n ≤ N, and N is the total number of frames; k is the center frequency sequence, 0 < k ≤ K, and K is the total number of frequency points. The complex-domain X(n,k) includes the magnitude spectrum |X(n,k)| and the phase value ∠X(n,k) in the time-frequency domain.

[0065] Then, the Mel cepstral features M(n,l) corresponding to the magnitude spectrum |X(n,k)| of the first clean voice signal sample x(t) of the speaker can be obtained, where l is the feature dimension.

[0066] Next, the Mel cepstral features M(n, l) corresponding to the magnitude spectrum |X(n, k)| of the first clean speech signal sample x(t) of the speaking object can be input into a speech feature extractor to obtain the speech feature of the speaking object. For example, a multi-layer attention mechanism-based delay network f(·) can be used to process the frame-based Mel cepstral features M(n, l) to obtain frame-level features R(n, i), where i is the number of nodes in this layer of neural network.

[0067] R(n, i) = f(Mel(|STFT(x(t))|))

[0068] After obtaining the frame-level features R(n, i), statistic pooling can be performed on the frame-level features R(n, i) to obtain the speech feature e of the speaking object at the sentence level. The speech feature e of the speaking object at the sentence level can be a vector:

[0069] e = statistic pooing(R(n, i))

[0070] In step 103, the speech feature e of each speaking object can be input into a classifier to obtain the identification prediction result of the speaking object. For example, a two-layer forward network can be used to transform the speech feature e corresponding to the speaking object and predict the identification prediction result of the speech feature e belonging to each speaking object among all speaking objects in the training data. Among them, the identification prediction result can be the identity prediction probability.

[0071] It should be noted that short-time Fourier transform can also be performed on the overlapping speech noisy signal sample y(t):

[0072] Y(n, k) = STFT(y(t))

[0073] Among them, the complex-domain Y(n, k) includes the magnitude spectrum |Y(n, k)| and the phase spectrum ∠Y(n, k).

[0074] In step 104, the speech feature e of each speaking object and the magnitude spectrum |Y(n, k)| of the overlapping speech noisy signal sample y(t) can be input into a speech extractor to obtain the magnitude spectrum mask M of the predicted enhanced speech signal of the speaking object, and the predicted enhanced speech signal s′(t) of the speaking object can be determined according to the magnitude spectrum mask M of the predicted enhanced speech signal of the speaking object. Among them, the enhanced speech signal s′(t) is the speech signal of the speaking object extracted from the overlapping speech noisy signal sample y(t) of the speaking object. It should be noted that the speech extractor of the present disclosure can include a forward neural network layer, a Gate Recurrent Unit (GRU) layer, and a dilated convolutional network layer.

[0075] According to an exemplary embodiment of the present disclosure, the magnitude spectrum |Y(n,k)| of the overlapping speech noisy signal sample y(t) of the speaking object can be multiplied by the magnitude spectrum mask M of the predicted enhanced speech signal of the speaking object to obtain the magnitude spectrum |Y′(n,k)| of the predicted enhanced speech signal of the speaking object. This process realizes the filtering of the magnitude spectrum mask M of the predicted enhanced speech signal of the speaking object. Next, the magnitude spectrum |Y′(n,k)| of the predicted enhanced speech signal of the speaking object can be combined with the phase spectrum ∠Y(n,k) of the overlapping speech noisy signal sample y(t) and perform an inverse short-time Fourier transform (iSTFT) to obtain the predicted enhanced speech signal s′(t) of the speaking object:

[0076] s′(t) = iSTFT(g(STFT(x(t)), STFT(y(t))))

[0077] Refer to Figure 2 , Figure 2 is a schematic diagram showing the training stage of a speech enhancement model according to an exemplary embodiment of the present disclosure. In Figure 2 , feature extraction can be performed on the first clean speech signal sample x(t) of each speaking object to obtain the Mel cepstral features M(n,l) corresponding to the first clean speech signal sample x(t) of the speaking object. Then, the Mel cepstral features M(n,l) corresponding to the first clean speech signal sample x(t) of the speaking object can be used as the input of the speech feature extractor to obtain the speech feature e of the speaking object.

[0078] For the overlapping speech noisy signal sample y(t) of the speaking object, a short-time Fourier transform can be performed on it to obtain Y(n,k) in the complex domain, and the Y(n,k) in the complex domain includes a magnitude spectrum |Y(n,k)| and a phase spectrum ∠Y(n,k). The speech feature e of the speaking object and the magnitude spectrum |Y(n,k)| of the overlapping speech noisy signal sample y(t) of the speaking object can be used as the input of the speech extractor. With the help of the speech feature e of the speaking object, the magnitude spectrum mask M of the predicted enhanced speech signal of the speaking object can be obtained.

[0079] Next, the amplitude spectrum mask M of the enhanced speech signal of the predicted speaker can be filtered. For example, the amplitude spectrum |Y(n,k)| of the overlapped speech noisy signal sample y(t) of the speaker can be multiplied point by point with the amplitude spectrum mask M of the predicted enhanced speech signal of the speaker to obtain the amplitude spectrum |Y′(n,k)| of the predicted enhanced speech signal of the speaker. Finally, the amplitude spectrum |Y′(n,k)| of the predicted enhanced speech signal of the speaker can be combined with the phase spectrum ∠Y(n,k) of the overlapped speech noisy signal sample y(t) and iSTFT can be performed to obtain the predicted enhanced speech signal s′(t) of the speaker.

[0080] Moreover, the speech representation e of the speaker can also be used as the input of the classifier to obtain the identification prediction result of the speaker, that is, the identity prediction probability of predicting the speaker as each of the speaker 1, speaker 2, ……, speaker m can be obtained.

[0081] In step 105, the loss can be calculated according to the enhanced speech signal s′(t), the second clean speech signal sample s(t), the identification prediction result, and the identification label corresponding to each speaker.

[0082] According to the exemplary embodiment of the present disclosure, the speech extractor needs to improve the SNR of the speech signal of the potential speaker extracted as much as possible, and the speech representation extractor needs to improve the classification accuracy of the speaker as much as possible. Therefore, a joint training method can be adopted, and the loss function corresponding to the speech extractor and the loss function corresponding to the speech representation extractor can be weighted and summed as the loss function used in actual training.

[0083] For example, the first loss J1 can be calculated according to the identification prediction result and the identification label corresponding to each speaker; the second loss J2 can be calculated according to the enhanced speech signal s′(t) and the second clean speech signal sample s(t) corresponding to each speaker. Next, the first loss J1 and the second loss J2 can be weighted and summed to obtain the loss J during actual training:

[0084] J = α * J1 + J2

[0085] where α is the weight.

[0086] According to the exemplary embodiment of the present disclosure, the first loss J1 can be calculated by the cross-entropy loss function:

[0087]

[0088] Where C is the total number of speakers in the training data. When the speech representation e of the speech segment is output by speaker c, P c is equal to 1; otherwise, P c = 0. P(c|e) is the identification prediction result that the classifier predicts the speech representation e of the speech segment as being output by speaker c.

[0089] According to an exemplary embodiment of the present disclosure, the second loss J2 can be calculated by a Scale-Invariant Signal To Distortion Ratio (SISDR) loss function:

[0090]

[0091] In step 106, the parameters of the speech extractor, the speech representation extractor, and the classifier can be adjusted by the loss J, so as to train the speech enhancement model.

[0092] According to an exemplary embodiment of the present disclosure, the speech representation extractor and the classifier can be obtained through a first pre-training; the speech extractor can be obtained through a second pre-training. Wherein, the first pre-training and the second pre-training are independent of each other.

[0093] At this time, the process of training the speech enhancement model according to the present disclosure can include three stages:

[0094] In the first training stage, a large number of clean speech signal samples of a single speaker can be used to pre-train the speech representation extractor and the classifier. The classifier can be used to obtain the prediction probability that the clean speech signal sample of a single speaker belongs to each speaker in the training dataset. Then, according to the true identity label of each speaker, the loss function of cross-entropy can be calculated, and the parameters of the speech representation extractor and the classifier can be adjusted by this cross-entropy loss function, so as to achieve the purpose of pre-training the speech representation extractor and the classifier.

[0095] In the second training stage, the pre-trained speech feature extractor can be used to extract the speech feature e corresponding to the first clean speech signal sample of the speaker. With the help of this speech feature e, the potential speech signal of the speaker can be extracted from the overlapping speech noisy signal sample of the speaker. Based on the second clean speech signal sample of the speaker, which is different from the content of the first clean speech signal sample of the speaker and is included in the overlapping speech noisy signal sample of the speaker, and the extracted speech signal of the speaker, the SISDR loss function can be calculated. Next, the parameters of the speech extractor can be adjusted through the SISDR loss function, thereby realizing the pre-training of the speech extractor. It should be noted that in the second training stage, the parameters of the pre-trained speech feature extractor need to be fixed, that is, the pre-trained speech feature extractor will not be retrained in the second training stage.

[0096] In the third training stage, the pre-trained speech feature extractor, the pre-trained classifier, and the pre-trained speech extractor can be jointly trained using a smaller learning rate. Among them, the loss function of the cross-entropy corresponding to the speech feature extractor and the classifier and the SISDR loss function corresponding to the speech extractor can be weighted and summed as the loss function in the third training stage. Next, the parameters of the pre-trained speech feature extractor, the pre-trained classifier, and the pre-trained speech extractor can be adjusted using the loss function obtained by weighted summation, thereby realizing the joint training of the speech enhancement model.

[0097] It should be noted that compared with the method of directly training the speech feature extractor, the classifier, and the speech extractor, this method of first pre-training each module separately and then jointly training each module has two advantages:

[0098] 1) The speech feature extractor and the classifier can be trained using more training data to improve their performance;

[0099] 2) By adopting the pre-training method, the speech enhancement model can be more easily trained and can converge faster.

[0100] Refer to Figure 3 , Figure 3 which is a flowchart showing a speech enhancement method according to an exemplary embodiment of the present disclosure.

[0101] Refer to Figure 3 , in step 301, the measured speech signal a(t) can be obtained. Among them, the measured speech signal a(t) may be the speech signal of one speaker or the speech signals of multiple speakers, and the measured speech signal a(t) may also contain environmental noise, for example, the noise of the air conditioner running, the noise of the user typing on the keyboard, and so on.

[0102] It should be noted that the voice enhancement model of the present disclosure can be deployed on the user's terminal. The user needs to pre-register a voice signal and store it on their own terminal for future use.

[0103] In step 302, the pre-registered voice signal of the target speaker can be input into the voice feature extractor trained according to the training method of the present disclosure to obtain the voice feature e of the target speaker.

[0104] In step 303, the voice feature e of the target speaker and the measured voice signal a(t) can be input into the voice extractor trained according to the training method of the present disclosure to obtain the voice signal s′(t) of the target speaker extracted from the measured voice signal a(t).

[0105] Refer to Figure 4 , Figure 4 is a schematic diagram showing the test stage of a voice enhancement model according to an exemplary embodiment of the present disclosure. In Figure 4 , the test process on terminal 1 of Zhang San and terminal 2 of Li Si is shown. In the test stage, Figure 2 the classifier and the calculation module of the loss function in are discarded.

[0106] When Zhang San uses the voice enhancement model of the present disclosure, he can perform feature extraction on a pre-registered voice signal z1(t) of his own to obtain the Mel cepstrum feature M(n, l)1 corresponding to the pre-registered voice signal z1(t). Then, the Mel cepstrum feature M(n, l)1 corresponding to Zhang San's pre-registered voice signal z1(t) can be used as the input of the voice feature extractor to obtain Zhang San's voice feature e1. When the voice enhancement model receives the measured voice a1(t), it can perform a short-time Fourier transform on it to obtain A1(n, k) in the complex domain. The A1(n, k) in the complex domain includes the amplitude spectrum |A1(n, k)| and the phase spectrum ∠A1(n, k). It should be noted that the measured voice a1(t) may contain at least one interfering human voice. The voice feature e1 of Zhang San and the amplitude spectrum |A1(n, k)| of the measured voice a1(t) can be used as the input of the voice extractor. With the help of Zhang San's voice feature e1, the amplitude spectrum mask M1 of the enhanced voice signal of Zhang San extracted from the measured voice a1(t) can be obtained.

[0107] Next, the amplitude spectrum mask M1 of the enhanced speech signal of Zhang San predicted to be extracted from the test speech a1(t) can be filtered. For example, the amplitude spectrum |A1(n,k)| of the test speech a1(t) can be multiplied point-by-point with the amplitude spectrum mask M1 of the enhanced speech signal of Zhang San predicted to be extracted from the test speech a1(t) to obtain the amplitude spectrum |A1′(n,k)| of the enhanced speech signal of Zhang San predicted. Finally, the amplitude spectrum |A1′(n,k)| of the enhanced speech signal of Zhang San predicted can be combined with the phase spectrum ∠A1(n,k) of the test speech a1(t) and iSTFT can be performed to obtain the enhanced speech signal s′1(t) of Zhang San predicted. It should be noted that the enhanced speech signal s′1(t) of Zhang San predicted only contains the voice of Zhang San, and the interfering human voices, environmental noises, etc. in the test speech are all filtered out.

[0108] When Li Si uses the speech enhancement model of the present disclosure, the process of extracting the potential speech signal of Li Si from the test speech is similar to the process of extracting the potential speech signal of Zhang San from the test speech described above, and will not be elaborated here. In Figure 4 , a pre-registered speech signal of Li Si is z2(t), the Mel cepstral features corresponding to Li Si are M(n,l)2, the speech representation of Li Si is e2, the test speech received by the speech enhancement model is a2(t), the amplitude spectrum obtained by performing short-time Fourier transform on it is |A2(n,k)|, the phase spectrum is ∠A2(n,k), the amplitude spectrum mask of the enhanced speech signal of Li Si predicted to be extracted from the test speech a2(t) is M2, the amplitude spectrum of the enhanced speech signal of Li Si predicted is |A2′(n,k)|, and the enhanced speech signal of Li Si predicted is s′2(t).

[0109] It should be noted that the speech enhancement system in the related art will retain the speech signals of all speakers. At this time, the unfiltered interfering human voice signals will interfere with the speech signals of the specified speaker, reducing the listening quality and intelligibility of the speech signals of the specified speaker. For example, when Wang Er participates in a remote video conference and gives a speech at his workstation, his colleague Xiao Zhou next to him is discussing problems with other colleagues. At this time, Wang Er's microphone will pick up Wang Er's voice, Xiao Zhou's voice, and environmental noises at the same time. For example, the environmental noise can be the noise of the air conditioner running or the noise of typing on the keyboard. The speech enhancement system in the related art can only filter out environmental noises and cannot filter out interfering human voices. At this time, the listeners at the other end of the remote video conference will hear Wang Er's voice and Xiao Zhou's voice at the same time. However, what the remote listeners really want to hear is only Wang Er's voice. Since Xiao Zhou's voice is also transmitted to the remote end, the listening quality and intelligibility of Wang Er's speech for the remote listeners will both decrease.

[0110] The voice enhancement model of the present disclosure can filter out all the picked-up ambient noise and interfering human voices during a video conference. Subsequently, the terminals of the participants can distribute the pure voice with the ambient noise and interfering human voices filtered out to the terminals of other participants through the server and play it. It should be noted that during a video conference, when multiple participants all enable the voice enhancement models on their respective terminals, the voice enhancement models on each participant's terminal will execute the same voice extraction process. In this way, the voice enhancement model of the present disclosure can extract the potential voice signal of a specific speaker from the measured voice signal with the help of the voice signal of the pre-registered speaker, avoiding the interference of the interfering human voice signal on the voice signal of the specific speaker, achieving the effect of voice enhancement for the specified speaker, and improving the listening quality and intelligibility of the voice signal of the specified speaker.

[0111] Figure 5 FIG. is a block diagram showing a training apparatus for a voice enhancement model according to an exemplary embodiment of the present disclosure. The voice enhancement model includes a voice extractor, a voice feature extractor, and a classifier.

[0112] Referring to Figure 5 , the training apparatus 500 for the voice enhancement model may include an acquisition module 501, a first input module 502, a second input module 503, a third input module 504, a calculation module 505, and an adjustment module 506.

[0113] The acquisition module 501 is configured to acquire training samples of multiple speakers. Among them, the training sample of each speaker includes: the first pure voice signal sample of the speaker, the second pure voice signal sample of the speaker, the overlapping voice noisy signal sample, and an identification label. The overlapping voice noisy signal sample of the speaker is obtained by superimposing the voice signals of at least one other speaker on the second pure voice signal sample of the speaker.

[0114] The first input module 502 is configured to input the first pure voice signal sample of each speaker into the voice feature extractor to obtain the voice feature of the speaker.

[0115] The second input module 503 is configured to input the voice feature of each speaker into the classifier to obtain the identification prediction result of the speaker.

[0116] A third input module 504, configured to input the speech representation of each speaker and the amplitude spectrum of the overlapping speech noisy signal samples into the speech extractor, obtain the amplitude spectrum mask of the predicted enhanced speech signal of the speaker, and determine the predicted enhanced speech signal of the speaker according to the amplitude spectrum mask of the predicted enhanced speech signal of the speaker, wherein the enhanced speech signal is the speech signal of the speaker extracted from the overlapping speech noisy signal samples of the speaker;

[0117] A calculation module 505, configured to calculate a loss according to the enhanced speech signal corresponding to each speaker, the second clean speech signal sample, the identification prediction result, and the identification label;

[0118] An adjustment module 506, configured to adjust the parameters of the speech extractor, the speech representation extractor, and the classifier through the loss to train the speech enhancement model.

[0119] According to an exemplary embodiment of the present disclosure, the calculation module 505 is configured to:

[0120] Calculate a first loss according to the identification prediction result and the identification label corresponding to each speaker;

[0121] Calculate a second loss according to the enhanced speech signal corresponding to each speaker and the second clean speech signal sample;

[0122] Perform a weighted sum of the first loss and the second loss to obtain the loss.

[0123] According to an exemplary embodiment of the present disclosure, the first loss is calculated by a cross-entropy loss function.

[0124] According to an exemplary embodiment of the present disclosure, the second loss is calculated by a scale-invariant signal-to-distortion ratio loss function.

[0125] According to an exemplary embodiment of the present disclosure, the speech representation extractor and the classifier are obtained through a first pre-training; the speech extractor is obtained through a second pre-training, wherein the first pre-training and the second pre-training are independent of each other.

[0126] According to an exemplary embodiment of the present disclosure, the third input module 504 is configured to:

[0127] Multiply the amplitude spectrum of the overlapping speech noisy signal sample of the speaker by the amplitude spectrum mask of the predicted enhanced speech signal of the speaker to obtain the amplitude spectrum of the predicted enhanced speech signal of the speaker;

[0128] Combine the amplitude spectrum of the predicted enhanced speech signal of the speaker object with the phase spectrum of the overlapping speech noisy signal sample and perform an inverse short-time Fourier transform to obtain the predicted enhanced speech signal of the speaker object.

[0129] According to an exemplary embodiment of the present disclosure, the first input module 502 is configured to:

[0130] Perform a Fourier transform on the first clean speech signal sample of each speaker object to obtain the amplitude spectrum of the first clean speech signal sample of the speaker object;

[0131] Obtain the Mel cepstral features corresponding to the amplitude spectrum of the first clean speech signal sample of the speaker object;

[0132] Input the Mel cepstral features corresponding to the amplitude spectrum of the first clean speech signal sample of the speaker object into the speech feature extractor.

[0133] According to an exemplary embodiment of the present disclosure, the first clean speech signal sample is different from the second clean speech signal sample.

[0134] Figure 6 It is a block diagram showing a speech enhancement device according to an exemplary embodiment of the present disclosure.

[0135] Refer to Figure 6 As shown in, the speech enhancement device 600 may include an acquisition module 601, a first input module 602, and a second input module 603.

[0136] The acquisition module 601 is configured to acquire the measured speech signal;

[0137] The first input module 602 is configured to input the speech signal of the pre-registered target speaker object into the speech feature extractor trained by the training device according to the present disclosure to obtain the speech feature of the target speaker object;

[0138] The second input module 603 is configured to input the speech feature of the target speaker object and the measured speech signal into the speech extractor trained by the training device according to the present disclosure to obtain the speech signal of the target speaker object extracted from the measured speech signal.

[0139] Figure 7 It is a block diagram showing an electronic device 700 according to an exemplary embodiment of the present disclosure.

[0140] Refer to Figure 7, the electronic device 700 includes at least one memory 701 and at least one processor 702. Instructions are stored in the at least one memory 701, and when the instructions are executed by the at least one processor 702, a method for training a voice enhancement model or a voice enhancement method according to an exemplary embodiment of the present disclosure is executed.

[0141] As an example, the electronic device 700 may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instructions. Here, the electronic device 700 does not have to be a single electronic device, and may also be a collection of any devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The electronic device 700 may also be a part of an integrated control system or a system manager, or may be configured as a portable electronic device that can be interconnected locally or remotely (e.g., via wireless transmission).

[0142] In the electronic device 700, the processor 702 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0143] The processor 702 may run instructions or code stored in the memory 701, where the memory 701 may also store data. The instructions and data may also be sent and received via a network interface device through a network, where the network interface device may adopt any known transmission protocol.

[0144] The memory 701 may be integrated with the processor 702. For example, RAM or flash memory may be arranged within an integrated circuit microprocessor, etc. In addition, the memory 701 may include a separate device, such as an external disk drive, a storage array, or other storage devices that can be used by any database system. The memory 701 and the processor 702 may be operatively coupled, or may communicate with each other, for example, through an I / O port, a network connection, etc., such that the processor 702 can read files stored in the memory.

[0145] In addition, the electronic device 700 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 700 may be connected to each other via a bus and / or a network.

[0146] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device can execute a method for training a voice enhancement model or a voice enhancement method according to the present disclosure.

[0147] Examples of the computer-readable storage medium herein include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device being configured to store a computer program and any associated data, data files and data structures in a non-transitory manner and provide the computer program and any associated data, data files and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system such that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.

[0148] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including a computer program, which when executed by a processor implements the training method or the voice enhancement method of the voice enhancement model according to the present disclosure.

[0149] Figure 8 is a block diagram showing the structure of a speaker according to an exemplary embodiment of the present disclosure. As Figure 8 shown, the speaker 800 according to an exemplary embodiment of the present disclosure includes: a voice enhancement device 600.

[0150] Figure 9 is a block diagram showing the structure of a speaker according to another exemplary embodiment of the present disclosure. As Figure 9As shown, the speaker 900 according to an exemplary embodiment of the present disclosure includes: at least one memory 901 and at least one processor 902. A set of computer-executable instructions is stored in the at least one memory 901. When the set of computer-executable instructions is executed by the at least one processor 902, the voice enhancement method described in the above exemplary embodiment is executed.

[0151] As an example, the speakers 800 and 900 in the above exemplary embodiments can be understood as devices integrated with speakers and / or microphones. For example, they can be smart speakers, home speakers, video conferencing devices, or teleconferencing devices. In addition, they can also be integrated on other devices. That is, it should be clear that as long as a speaker uses the voice enhancement method shown in the present disclosure for voice enhancement, it falls within the scope of protection of the present disclosure.

[0152] As an example, the speakers 800 and 900 may also include other components for the speakers to perform their own functions. For example, they may include at least one of the following items: a signal acquisition unit and a signal processing unit. The signal acquisition unit can collect sounds in the environment to form an audio signal, and the signal processing unit can process the audio signal collected by the signal acquisition unit (for example, amplification processing, etc.).

[0153] As an example, the speakers 800 and 900 can be applied to but not limited to at least one of the following scenarios: video conferencing scenarios, home environment scenarios, and online teaching scenarios. It should be understood that they can also be applied to other appropriate scenarios, and the present disclosure does not limit this. In different usage scenarios, the component structures of the speakers 800 and 900 may be different. It should be clear that as long as a speaker uses the voice enhancement method shown in the present disclosure for voice enhancement, it falls within the scope of protection of the present disclosure.

[0154] According to the training method, voice enhancement method, device, electronic device, speaker, and storage medium of the voice enhancement model of the present disclosure, since a large number of pure voice signal samples of speaking objects and overlapping voice noisy signal samples are used to train the voice enhancement model, the trained voice enhancement model can accurately extract the voice signal of a specified speaking object from the voice signals of multiple speaking objects, avoiding interference from interfering voice signals to the voice signal of the specified speaking object, and improving the listening quality and intelligibility of the voice signal of the specified speaking object. Moreover, the voice enhancement method of the present disclosure can, with the help of the voice signal of a pre-registered target speaking object, use the trained voice enhancement model to extract the potential voice signal of the target speaking object from the measured voice signal, avoiding interference from interfering voice signals to the voice signal of the specified speaking object, achieving the effect of voice enhancement for the specified speaking object, and improving the listening quality and intelligibility of the voice signal of the specified speaking object. Further, the first pure voice signal sample and the second pure voice signal sample for training can be set as voice signal samples with different contents, so as to ensure that in the actual test scenario, the voice enhancement model can extract the potential voice signal of the specified speaking object from the measured voice signal as accurately as possible. Further, the training method of the voice enhancement model of the present disclosure can first pre-train each module separately and then jointly train each module, which can ensure that the voice feature extractor and the classifier use more training data for training, making their performance more perfect; moreover, adopting the pre-training method can make the voice enhancement model easier to train and converge faster.

[0155] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only illustrative, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0156] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A training method for a voice enhancement model, characterized in that The speech enhancement model includes a speech extractor, a speech feature extractor, and a classifier. The training method includes: Obtaining training samples of multiple speakers. Each speaker's training sample includes: the first clean speech signal sample of the speaker, the second clean speech signal sample of the speaker, the overlapping speech noisy signal sample of the speaker, and an identification label. The overlapping speech noisy signal sample of the speaker is obtained by superimposing the speech signals of at least one other speaker on the second clean speech signal sample of the speaker; Inputting the first clean speech signal sample of each speaker into the speech feature extractor to obtain the speech feature of the speaker; Inputting the speech feature of each speaker into the classifier to obtain the identification prediction result of the speaker; Inputting the speech feature of each speaker and the magnitude spectrum of the overlapping speech noisy signal sample into the speech extractor to obtain the magnitude spectrum mask of the predicted enhanced speech signal of the speaker, and determining the predicted enhanced speech signal of the speaker according to the magnitude spectrum mask of the predicted enhanced speech signal of the speaker. The enhanced speech signal is the speech signal of the speaker extracted from the overlapping speech noisy signal sample of the speaker; Calculating a loss according to the enhanced speech signal, the second clean speech signal sample, the identification prediction result, and the identification label corresponding to each speaker; Adjusting the parameters of the speech extractor, the speech feature extractor, and the classifier through the loss to train the speech enhancement model; Wherein, calculating the loss according to the enhanced speech signal, the second clean speech signal sample, the identification prediction result, and the identification label corresponding to each speaker includes: Calculating a first loss according to the identification prediction result and the identification label corresponding to each speaker; Calculating a second loss according to the enhanced speech signal and the second clean speech signal sample corresponding to each speaker; Performing weighted summation on the first loss and the second loss to obtain the loss; Wherein, the second loss is calculated by a scale-invariant signal-to-distortion ratio loss function.

2. The training method according to claim 1, wherein The first loss is calculated by a cross-entropy loss function.

3. The training method according to claim 1, wherein, The speech feature extractor and the classifier are obtained through a first pre-training; the speech extractor is obtained through a second pre-training, where the first pre-training and the second pre-training are independent of each other.

4. The training method according to claim 1, wherein Determining the predicted enhanced speech signal of the speaker according to the magnitude spectrum mask of the predicted enhanced speech signal of the speaker includes: Multiplying the magnitude spectrum of the overlapping speech noisy signal sample of the speaker by the magnitude spectrum mask of the predicted enhanced speech signal of the speaker to obtain the magnitude spectrum of the predicted enhanced speech signal of the speaker; Combining the magnitude spectrum of the predicted enhanced speech signal of the speaker with the phase spectrum of the overlapping speech noisy signal sample and performing inverse short-time Fourier transform to obtain the predicted enhanced speech signal of the speaker.

5. The training method according to claim 1, wherein Inputting the first clean speech signal sample of each speaker into the speech feature extractor includes: Perform short-time Fourier transform on the first clean speech signal sample of each speaking object to obtain the amplitude spectrum of the first clean speech signal sample of the speaking object; Obtain the Mel cepstrum features corresponding to the amplitude spectrum of the first clean speech signal sample of the speaking object; Input the Mel cepstrum features corresponding to the amplitude spectrum of the first clean speech signal sample of the speaking object into the speech feature extractor.

6. The training method according to claim 1, wherein The first clean speech signal sample is different from the second clean speech signal sample.

7. A voice enhancement method, characterized in that, The speech enhancement method includes: Obtain the measured speech signal; Input the speech signal of a pre-registered target speaking object into the speech feature extractor trained by any one of the training methods in claims 1 to 6 to obtain the speech feature of the target speaking object; Input the speech feature of the target speaking object and the measured speech signal into the speech extractor trained by any one of the training methods in claims 1 to 6 to obtain the speech signal of the target speaking object extracted from the measured speech signal.

8. A training device for a voice enhancement model, characterized in that, The speech enhancement model includes a speech extractor, a speech feature extractor, and a classifier. The training device includes: An acquisition module configured to acquire training samples of multiple speaking objects. Each training sample of a speaking object includes: the first clean speech signal sample of the speaking object, the second clean speech signal sample, the overlapping speech noisy signal sample, and an identification label. The overlapping speech noisy signal sample of the speaking object is obtained by superimposing the speech signals of at least one other speaking object on the second clean speech signal sample of the speaking object; A first input module configured to input the first clean speech signal sample of each speaking object into the speech feature extractor to obtain the speech feature of the speaking object; A second input module configured to input the speech feature of each speaking object into the classifier to obtain the identification prediction result of the speaking object; A third input module configured to input the speech feature of each speaking object and the amplitude spectrum of the overlapping speech noisy signal sample into the speech extractor to obtain the amplitude spectrum mask of the predicted enhanced speech signal of the speaking object, and determine the predicted enhanced speech signal of the speaking object according to the amplitude spectrum mask of the predicted enhanced speech signal of the speaking object. The enhanced speech signal is the speech signal of the speaking object extracted from the overlapping speech noisy signal sample of the speaking object; A calculation module configured to calculate a loss according to the enhanced speech signal, the second clean speech signal sample, the identification prediction result, and the identification label corresponding to each speaking object; An adjustment module configured to adjust the parameters of the speech extractor, the speech feature extractor, and the classifier through the loss to train the speech enhancement model; Wherein, the calculation module is configured to: Calculate a first loss according to the identification prediction result and the identification label corresponding to each speaking object; Calculate a second loss according to the enhanced speech signal and the second clean speech signal sample corresponding to each speaking object; Perform weighted summation on the first loss and the second loss to obtain the loss; Among them, the second loss is calculated by a scale-invariant signal-to-distortion ratio loss function.

9. The training device according to claim 8, wherein, The first loss is calculated by a cross-entropy loss function.

10. The training device according to claim 8, wherein The speech feature extractor and the classifier are obtained through a first pre-training; the speech extractor is obtained through a second pre-training, where the first pre-training and the second pre-training are independent of each other.

11. The training device according to claim 8, wherein The third input module is configured to: Multiply the magnitude spectrum of the overlapping speech noisy signal sample of the speaker by the magnitude spectrum mask of the predicted enhanced speech signal of the speaker to obtain the magnitude spectrum of the predicted enhanced speech signal of the speaker; Combine the magnitude spectrum of the predicted enhanced speech signal of the speaker with the phase spectrum of the overlapping speech noisy signal sample and perform an inverse short-time Fourier transform to obtain the predicted enhanced speech signal of the speaker.

12. The training device according to claim 8, characterized in that, The first input module is configured to: Perform a short-time Fourier transform on the first clean speech signal sample of each speaker to obtain the magnitude spectrum of the first clean speech signal sample of the speaker; Obtain the Mel cepstral features corresponding to the magnitude spectrum of the first clean speech signal sample of the speaker; Input the Mel cepstral features corresponding to the magnitude spectrum of the first clean speech signal sample of the speaker into the speech feature extractor.

13. The training device according to claim 8, wherein, The first clean speech signal sample is different from the second clean speech signal sample.

14. A voice enhancement device, characterized in that, The speech enhancement device includes: An acquisition module, configured to acquire a measured speech signal; A first input module, configured to input the speech signal of a pre-registered target speaker into a speech feature extractor trained by a training device according to any one of claims 8 to 13 to obtain the speech feature of the target speaker; A second input module, configured to input the speech feature of the target speaker and the measured speech signal into a speech extractor trained by a training device according to any one of claims 8 to 13 to obtain the speech signal of the target speaker extracted from the measured speech signal.

15. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Among them, the processor is configured to execute the instructions to implement the training method of the speech enhancement model according to any one of claims 1 to 6 or the speech enhancement method according to claim 7.

16. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the training method of the speech enhancement model according to any one of claims 1 to 6 or the speech enhancement method according to claim 7.

17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the training method of the speech enhancement model according to any one of claims 1 to 6 or the speech enhancement method according to claim 7.

18. A speaker, characterized in that, Comprising: The speech enhancement device according to claim 14.

19. A speaker, characterized in that, Comprising: At least one processor; At least one memory for storing computer-executable instructions; Among them, when the computer-executable instructions are run by the at least one processor, the at least one processor is prompted to execute the speech enhancement method according to claim 7.

Citation Information

Patent Citations

  • Voice enhancement model training method and device and voice enhancement method and device

    CN112289333A

  • Voice filtering method and filtering system

    CN112687275A

  • Speaking object representation extraction model training method and speaking object identity recognition method

    CN113990327A

  • Speech enhancement model training method and device and speech enhancement method and device

    CN114121029A