Voice separation method, device, electronic device and storage medium
By constructing teacher models and multiple student models, and using pseudo-labels to guide student model training, the problem of lack of semi-supervised model training methods in the existing technology is solved, and the generalization of student models and the optimization of voice separation effect is achieved.
Patent Information
- Application Number
- CN202210301773.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-03-24
AI Technical Summary
There is a lack of semi-supervised model training method for speech separation in the prior art, and it is difficult to effectively utilize a large amount of unsupervised data and a small amount of supervised data.
By constructing teacher models and multiple student models, using teacher models to train supervised data, determine the pseudo-labels of unsupervised data based on the output results of teacher models and student models, and guide the student models to train, thereby realizing the training of semi-supervised models.
The student model has achieved a breakthrough in the limitations of the teacher model and obtained a student model with a separation effect better than the teacher model. At the same time, it improved the generalization of the student model, obtained a purer target pronunciation, and obtained a better pronunciation separation effect.
Smart Images

Figure CN114613387B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech signal processing, and in particular to a speech separation method, device, electronic equipment and storage medium. Background Art
[0002] Speech separation refers to separating clean speech from noisy mixed speech of different speakers. Neural networks rely on a large amount of labeled data to provide generalization of the model and prevent overfitting. High-quality supervised data is often expensive and difficult to obtain; while unsupervised data is usually large and easy to obtain, but is ignored due to the lack of effective utilization methods. For speech separation tasks based on deep learning, supervised data is even more difficult to obtain because noise, human voice interference, etc. are everywhere.
[0003] Therefore, how to use a large amount of unsupervised data combined with a small amount of supervised data to train a semi-supervised model for speech separation is still a problem that needs to be solved urgently in the field of speech separation. Summary of the invention
[0004] The present invention provides a speech separation method, device, electronic device and storage medium, which are used to solve the problem that there is a lack of a semi-supervised model training method for speech separation in the prior art.
[0005] The present invention provides a speech separation method, comprising:
[0006] Determine the speech to be separated;
[0007] Inputting the speech to be separated into a speech separation model to obtain a target speech of the speech to be separated output by the speech separation model;
[0008] The speech separation model is one of multiple student models, and the multiple student models are obtained by training multiple initial student models based on a first sample speech and a pseudo target speech of the first sample speech. The pseudo target speech of the first sample speech is determined based on a first speech separation result output by a teacher model and the multiple initial student models for the first sample speech, and the teacher model is obtained through supervised training.
[0009] According to a speech separation method provided by the present invention, the multiple student models are trained based on the following steps:
[0010] Inputting the first sample speech into the teacher model and the multiple initial student models respectively, to obtain first speech separation results outputted by the teacher model and the multiple initial student models respectively;
[0011] Determine a first speech separation result with the best separation effect among the first speech separation results, and determine the target speech in the first speech separation result with the best separation effect as the pseudo target speech of the first sample speech;
[0012] Based on the pseudo target speech of the first sample speech and the first speech separation results respectively output by the multiple initial student models, the multiple initial student models are iteratively updated with parameters to obtain the multiple student models.
[0013] According to a speech separation method provided by the present invention, determining the first speech separation result with the best separation effect among the first speech separation results includes:
[0014] Extracting spectral features of the target speech and the interference audio in each of the first speech separation results, and determining the similarity between the spectral features of the target speech and the spectral features of the interference audio in each of the first speech separation results;
[0015] Based on the similarities corresponding to the first speech separation results, determining the first speech separation result corresponding to the lowest similarity, and determining the first speech separation result corresponding to the lowest similarity as the first speech separation result with the best separation effect;
[0016] Alternatively, based on the similarities corresponding to the first speech separation results, the separation effects of the first speech separation results are determined, and based on the separation effects of the first speech separation results, the first speech separation result with the best separation effect is determined.
[0017] According to a speech separation method provided by the present invention, the step of extracting spectral features of the target speech and the interference audio in each of the first speech separation results comprises:
[0018] Inputting the target speech and the interference audio in each of the first speech separation results into a spectral feature extractor respectively to obtain the spectral features of the target speech and the spectral features of the interference audio;
[0019] The spectral feature extractor is a feature extractor in a speaker recognition model, and the speaker recognition model is trained based on a speaker's speech and speaker information of the speaker's speech.
[0020] According to a speech separation method provided by the present invention, the pseudo target speech based on the first sample speech and the first speech separation results respectively output by the multiple initial student models, performing iterative parameter updating on the multiple initial student models, comprises:
[0021] When the first speech separation result with the best separation effect is the first speech separation result output by the teacher model, iteratively updating the parameters of the multiple initial student models based on the pseudo target speech of the first sample speech and the first speech separation results respectively output by the multiple initial student models;
[0022] When the first speech separation result with the best separation effect is the first speech separation result output by any initial student model, the parameters of the other initial student models are iteratively updated based on the pseudo target speech of the first sample speech and the first speech separation results respectively output by other initial student models. The other initial student models are the initial student models among the multiple initial student models except any initial student model.
[0023] According to a speech separation method provided by the present invention, the multiple student models are also trained based on the following steps:
[0024] Inputting the second sample speech into the multiple initial student models respectively, and obtaining the second speech separation results respectively output by the multiple initial student models;
[0025] Based on the real target speech of the second sample speech and the second speech separation results respectively outputted by the multiple initial student models, the multiple initial student models are iteratively updated with parameters to obtain the multiple student models.
[0026] According to a speech separation method provided by the present invention, the speech separation model is determined based on the following steps:
[0027] Inputting the third sample speech into the plurality of student models respectively, and obtaining the third speech separation results respectively output by the plurality of student models;
[0028] Based on the real target speech of the third sample speech and the third speech separation results respectively output by the multiple student models, determining the performance evaluation results respectively corresponding to the multiple student models;
[0029] Based on the performance evaluation results respectively corresponding to the multiple student models, the speech separation model is determined from the multiple student models.
[0030] The present invention also provides a speech separation device, comprising:
[0031] A speech determination unit, used for determining the speech to be separated;
[0032] A speech separation unit, used for inputting the speech to be separated into a speech separation model to obtain a target speech of the speech to be separated output by the speech separation model;
[0033] The speech separation model is one of multiple student models, and the multiple student models are obtained by training multiple initial student models based on a first sample speech and a pseudo target speech of the first sample speech. The pseudo target speech of the first sample speech is determined based on a first speech separation result output by a teacher model and the multiple initial student models for the first sample speech, and the teacher model is obtained through supervised training.
[0034] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-mentioned speech separation methods is implemented.
[0035] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the speech separation method described in any one of the above is implemented.
[0036] The speech separation method, device, electronic device and storage medium provided by the present invention construct a teacher model and multiple student models to jointly participate in semi-supervised model training. The teacher model is obtained by training based on supervised speech data, and the pseudo-labels of the unsupervised speech data are determined according to the output results of the teacher model and the multiple student models for the unsupervised speech data respectively, and the multiple student models are guided to be trained, so that the student model can break through the limitations of the teacher model and obtain a student model with better separation effect than the teacher model, while improving the generalization of the student model. On this basis, the student model is applied to the speech separation task of the speech to be separated, so as to obtain a purer target speech and a better speech separation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0038] Figure 1 This is one of the flow charts of the speech separation method provided by the present invention;
[0039] Figure 2 It is one of the training process diagrams of multiple student models provided by the present invention;
[0040] Figure 3 It is a flowchart of a method for determining a first speech separation result with the best separation effect provided by the present invention;
[0041] Figure 4This is the second schematic diagram of the training process of multiple student models provided by the present invention;
[0042] Figure 5 It is a flow chart of a method for determining a speech separation model provided by the present invention;
[0043] Figure 6 It is a schematic diagram of the training process of the teacher model provided by the present invention;
[0044] Figure 7 This is the third schematic diagram of the training process of multiple student models provided by the present invention;
[0045] Figure 8 This is the second flow chart of the speech separation method provided by the present invention;
[0046] Fig. 9 It is a structural schematic diagram of the speech separation device provided by the present invention;
[0047] Fig.10 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0049] In recent years, some semi-supervised learning methods based on deep learning have been proposed. Mean-teacher is the most widely used and most basic semi-supervised training method. Its general idea is that the model acts as both a student and a teacher. As a teacher, it is used to generate goals for students to learn; as a student, it uses the goals generated by the teacher model to learn. This semi-supervised dual-model strategy has achieved a certain accuracy improvement in image classification tasks compared to the supervised model strategy. However, the biggest drawback of Mean-teacher is that the direct use of consistency constraints and model parameter smoothing will cause serious coupling between the teacher model and the student model, resulting in performance bottlenecks and model collapse; in addition, the teacher model does not learn more meaningful knowledge than the student model. If the student model makes incorrect predictions for certain samples, the teacher model will most likely retain such incorrect predictions and cannot be corrected.
[0050] In order to overcome the performance defects of Mean-teacher, the Dual-student method was proposed. Dual-student uses two student models for independent training, and no longer has a fixed teacher model, in order to achieve decoupling of the two models and prevent model collapse. Specifically, during the training process, the student model with more stable performance is used as a label to guide the training of another student model. The calculation criterion for stability depends on the classification posterior probability of the model. Similarly, there are methods such as Mixmatch, Remixmatch, and Fixmatch that use the classification posterior probability to guide the training of semi-supervised models. However, the biggest limitation of these methods is that they can only be applied to classification tasks in deep learning, such as image recognition, intent classification, and speech recognition, because they need to use classification posterior probabilities; they cannot be applied to regression tasks in deep learning, such as speech separation, speech denoising, and echo cancellation.
[0051] In fact, there are not many semi-supervised learning methods for speech separation so far, and Mixup-Breakdown is one of them. The core idea of Mixup-Breakdown is to decompose the unsupervised data into the first target speech and the second target speech (or noise) through the teacher model, and then reconstruct the mixed signal according to a certain signal-to-interference ratio (or signal-to-noise ratio) and continue to train the student model; at the same time, the output of the teacher model is directly used as the label for the student model training.
[0052] However, the biggest risk of this dual-model cascade approach is that it places very high demands on the teacher model’s speech separation performance. Because the output of the teacher model is directly used as the learning target of the student model and the input for building the student model, if the teacher model’s separation effect is not good enough, it will lead to problems such as label contamination, which will lead to worse training of the student model.
[0053] To this end, the present invention provides a speech separation method. Figure 1 It is one of the flow charts of the speech separation method provided by the present invention, such as Figure 1 As shown, the method includes:
[0054] Step 110: determine the speech to be separated.
[0055] Here, the speech to be separated is the mixed speech that needs to be separated, which can be collected in advance by a sound receiving device, or recorded in real time. The real-time recording can be a voice recording or a video recording. Speech separation can be used to separate the speech signal of a speaker from other audio signals such as background music and noise, and can also be used to separate speech signals corresponding to different speakers, which is not specifically limited in the embodiments of the present invention.
[0056] Step 120, inputting the speech to be separated into the speech separation model, and obtaining the target speech of the speech to be separated output by the speech separation model;
[0057] The speech separation model is one of multiple student models, and the multiple student models are obtained by training multiple initial student models based on a first sample speech and a pseudo target speech of the first sample speech. The pseudo target speech of the first sample speech is determined based on the first speech separation results output by the teacher model and multiple initial student models for the first sample speech, and the teacher model is obtained through supervised training.
[0058] Specifically, after determining the speech to be separated, the speech to be separated can be input into the speech separation model, and the speech separation model will perform speech separation on the speech to be separated, thereby obtaining the speech separation result of the speech to be separated output by the model, which includes the target speech and interference audio. The target speech here is the clean speech of the target speaker, and the interference audio is the audio that interferes with the target speech, such as the speech of other speakers, background music, noise, etc.
[0059] Before executing step 120, considering that supervised data is often difficult to obtain in speech separation tasks, while unsupervised data is easy to obtain, the embodiment of the present invention applies a semi-supervised model training method that combines a large amount of unsupervised data with a small amount of supervised data to obtain a speech separation model, so as to reduce the cost of label data production while increasing the generalization of the model. Based on this, considering that the existing semi-supervised model training method for speech separation tasks will over-trust the output of the supervised model, resulting in the upper limit of the training of the semi-supervised model being the supervised model, and the generalization of the model has not been substantially improved, the embodiment of the present invention simultaneously applies a teacher model and multiple initial student models, and determines the pseudo-label of the unsupervised data, that is, the pseudo-target speech of the first sample speech, through the results of the output of each model for the unsupervised data, using this as a learning goal to guide multiple initial student models to train the unsupervised data, thereby obtaining multiple trained student models, and determining the speech separation model from the multiple student models by random selection or performance testing.
[0060] The specific training steps may be: first, collect a large amount of unsupervised speech data as the first sample speech, and construct a teacher model. Here, the pre-trained model obtained by supervised training can be directly applied as the teacher model, or supervised speech data and its corresponding real labels can be first collected, and then the initial teacher model is supervised trained according to the supervised speech data and its corresponding real labels to obtain the teacher model; then, determine the pseudo-label corresponding to the first sample speech, that is, the pseudo-target speech, according to the first speech separation result output by the teacher model and multiple initial student models for the first sample speech, and use the first sample speech and the pseudo-target speech of the first sample speech to train multiple initial student models, thereby obtaining multiple student models.
[0061] Here, the multiple student models may be two student models or more than two student models; the model structures of the teacher model and each student model may be the same or different. The model structure may adopt a time-frequency domain network structure, such as DCCRN (Deep Complex Convolution Recurrent Network), FSMN (Feedforward Sequential Memory Networks) and CLDNN (CNN-LSTM-DNN), etc., or a time domain network structure may be adopted, such as TASnet and Conv-TASnet, etc. The embodiments of the present invention do not specifically limit this.
[0062] The pseudo target speech of the first sample speech may be determined by evaluating the speech separation effect of the first speech separation results output by each model, and then determining the pseudo target speech according to the target speech in the first speech separation result with the best speech separation effect. Alternatively, the speech purity of the target speech in the first speech separation results output by each model may be evaluated, and then determining the pseudo target speech according to the target speech with the highest speech purity. The embodiment of the present invention does not make any specific limitation on this. The training method of multiple student models may be to train only with the first sample speech, or to jointly train with the first sample speech and supervised speech data. The embodiment of the present invention does not make any specific limitation on this.
[0063] It should be noted that, unlike the prior art in which the output results of the teacher model are fixed as pseudo labels, the embodiments of the present invention continuously evaluate the first speech separation results output by each model, and determine the pseudo target speech of the first sample speech according to the evaluation results corresponding to each model, thereby urging the initial student model to continuously learn and gradually improve the model performance by continuously updating the learning objectives, so that the trained student model can break through the limitations of the teacher model and achieve a better separation effect than the teacher model. In addition, in order to avoid the extreme situation that the first speech separation results output by each initial student model are always the same, resulting in the training upper limit of each initial student model still being the teacher model, each initial student model in the embodiment of the present invention is a model with different initial model parameters, and the initial model parameters here can be randomly generated or pre-set, and each initial student model learns independently during the training process without sharing parameters.
[0064] The method provided by the embodiment of the present invention constructs a teacher model and multiple student models to jointly participate in the semi-supervised model training. The teacher model is obtained according to the training of supervised speech data, and the pseudo-labels of the unsupervised speech data are determined according to the output results of the teacher model and the multiple student models for the unsupervised speech data respectively, and the multiple student models are guided to be trained, so that the student model can break through the limitations of the teacher model and obtain a student model with better separation effect than the teacher model, and at the same time improve the generalization of the student model. On this basis, the student model is applied to the speech separation task of the speech to be separated, so as to obtain a purer target speech and a better speech separation effect.
[0065] Based on the above embodiments, Figure 2 is one of the training flow diagrams of multiple student models provided by the present invention, such as Figure 2 As shown, multiple student models are trained based on the following steps:
[0066] Step 210, inputting the first sample speech into the teacher model and the multiple initial student models respectively, and obtaining the first speech separation results outputted by the teacher model and the multiple initial student models respectively;
[0067] Step 220, determining the first speech separation result with the best separation effect among the first speech separation results, and determining the target speech in the first speech separation result with the best separation effect as the pseudo target speech of the first sample speech;
[0068] Step 230, based on the pseudo target speech of the first sample speech and the first speech separation results respectively output by the multiple initial student models, iteratively update the parameters of the multiple initial student models to obtain multiple student models.
[0069] Specifically, in order to more directly improve the speech separation effect of the student model, after obtaining the trained teacher model, the embodiment of the present invention first inputs the first sample speech into the teacher model to obtain the first speech separation result output by the teacher model, and inputs the first sample speech into multiple initial student models respectively to obtain the first speech separation results output by multiple initial student models. Then, the separation effect of the first speech separation results output by each model is evaluated, so as to determine the first speech separation result with the best separation effect among the first speech separation results according to the evaluation results. The separation effect here is used to characterize the degree of differentiation between the target speech and the interference audio in the corresponding first speech separation result.
[0070] On this basis, the target speech in the first speech separation result with the best separation effect can be determined as the pseudo target speech of the first sample speech. Finally, the loss value between the first speech separation results output by multiple initial student models and the pseudo target speech can be determined according to the loss function, and then the multiple initial student models can be guided to iteratively update their parameters according to the loss values corresponding to the multiple initial student models, thereby obtaining multiple student models.
[0071] The method provided by the embodiment of the present invention selects the output result with the best separation effect from the output results of the teacher model and each student model, uses the target speech therein as the pseudo-label of the unsupervised data, and uses this as the learning goal to guide multiple initial student models to train the unsupervised data, thereby breaking the traditional limitation of fixing the output results of the teacher model as pseudo-labels for model training, and it is easier to obtain a speech separation effect that is better than that of the supervised model.
[0072] Based on any of the above embodiments, Figure 3 is a flow chart of a method for determining a first speech separation result with the best separation effect provided by the present invention, such as Figure 3 As shown, determining the first speech separation result with the best separation effect among the first speech separation results includes:
[0073] Step 310, extracting spectral features of the target speech and the interfering audio in each first speech separation result, and determining the similarity between the spectral features of the target speech and the spectral features of the interfering audio in each first speech separation result;
[0074] Step 320, based on the similarities corresponding to the first speech separation results, determine the first speech separation result corresponding to the lowest similarity, and determine the first speech separation result corresponding to the lowest similarity as the first speech separation result with the best separation effect;
[0075] Alternatively, based on the similarities corresponding to the first speech separation results, the separation effects of the first speech separation results are determined, and based on the separation effects of the first speech separation results, the first speech separation result with the best separation effect is determined.
[0076] Specifically, the first speech separation result with the best separation effect can be determined in the following manner: first, spectral features are extracted for the target speech and the interfering audio in each first speech separation result respectively to obtain the spectral features of the target speech and the spectral features of the interfering audio in each first speech separation result; then, the similarity between the spectral features of the target speech and the spectral features of the interfering audio in each first speech separation result is calculated, and then the first speech separation result with the best separation effect is determined according to the similarities corresponding to each first speech separation result. The determination method here can be to compare the similarities corresponding to each first speech separation result, determine the first speech separation result corresponding to the lowest similarity, and then determine the first speech separation result corresponding to the lowest similarity as the first speech separation result with the best separation effect; or the separation effect of each first speech separation result can be calculated according to the similarities corresponding to each first speech separation result, and then the first speech separation result with the best separation effect is determined according to the separation effect of each first speech separation result.
[0077] Here, the spectral feature is used to characterize the characteristics of the corresponding audio on the spectrum, for example, it can be x-vector, i-vector, etc. The similarity can be calculated by using methods such as cosine similarity and Pearson correlation coefficient. The separation effect is used to characterize the distinction between the target speech and the interference audio, which can be obtained by subtracting the similarity from 1 or by dividing the similarity by a certain constant, and the embodiment of the present invention does not specifically limit this.
[0078] It can be understood that in the speech separation task, the two separated audios should belong to different sound sources, and their spectral features should be distinguishable. If the similarity between the two is higher, the separation effect of the corresponding first speech separation result is worse; if the similarity between the two is lower, the separation effect of the corresponding first speech separation result is better.
[0079] Based on any of the above embodiments, in step 310, spectral features are extracted from the target speech and the interference audio in each first speech separation result, including:
[0080] Inputting the target speech and the interference audio in each first speech separation result into a spectral feature extractor respectively to obtain the spectral features of the target speech and the spectral features of the interference audio;
[0081] The spectral feature extractor is a feature extractor in the speaker recognition model. The speaker recognition model is trained based on the speaker's speech and the speaker information of the speaker's speech.
[0082] Specifically, considering that the speaker recognition task can realize the speaker information corresponding to the recognized speech, it can be used to evaluate the discrimination between the target speech and the interference audio in each first speech separation result, and then obtain the separation effect of each first speech separation result. Therefore, the embodiment of the present invention applies the feature extractor in the speaker recognition model, that is, the spectral feature extractor, and inputs the target speech and the interference audio in each first speech separation result into the spectral feature extractor respectively, so as to obtain the spectral features of the target speech that can characterize the target speaker information, and the spectral features of the interference audio of other sound source information, which are used for the subsequent calculation of the separation effect of each first speech separation result.
[0083] Here, the speaker recognition model can be pre-trained based on the speaker's voice and the speaker information of the speaker's voice, where the speaker information is information used to characterize the speaker's identity. The model structure of the speaker recognition model can be, for example, TDNN (Time-Delay Neural Network), Transformer, etc., which is not specifically limited in the embodiment of the present invention.
[0084] Based on any of the above embodiments, in step 230, based on the pseudo target speech of the first sample speech and the first speech separation results respectively output by the multiple initial student models, iteratively updating the parameters of the multiple initial student models includes:
[0085] When the first speech separation result with the best separation effect is the first speech separation result output by the teacher model, iteratively updating the parameters of the multiple initial student models based on the pseudo target speech of the first sample speech and the first speech separation results respectively output by the multiple initial student models;
[0086] When the first speech separation result with the best separation effect is the first speech separation result output by any initial student model, the parameters of other initial student models are iteratively updated based on the pseudo-target speech of the first sample speech and the first speech separation results output by other initial student models respectively. The other initial student models are initial student models other than the initial student model among multiple initial student models.
[0087] Specifically, after obtaining the first speech separation results outputted by the teacher model and the multiple initial student models, the first speech separation result with the best separation effect among the first speech separation results can be determined first, and then corresponding processing can be performed according to different determination results:
[0088] When the first speech separation result with the best separation effect is the first speech separation result output by the teacher model, the target speech in the first speech separation result output by the teacher model is used as the pseudo target speech of the first sample speech, and the parameters of the multiple initial student models are iteratively updated according to the loss values between the first speech separation results respectively output by the multiple initial student models and the pseudo target speech;
[0089] When the first speech separation result with the best separation effect is the first speech separation result output by any initial student model, it means that the parameters of the initial student model do not need to be updated at present. The target speech in the first speech separation result output by the initial student model is used as the pseudo target speech of the first sample speech. According to the loss values between the first speech separation results output by other initial student models and the pseudo target speech, the parameters of other initial student models are iteratively updated. The other initial student models here are the initial student models other than the initial student model among the multiple initial student models.
[0090] It should be noted that if the first speech separation result with the best separation effect is the first speech separation result output by any initial student model, then there is no need to train the initial student model, and the loss value corresponding to the initial student model can be 0, and only other initial student models need to be trained. By continuously comparing the separation effects of the first speech separation results output by each model, the target speech in the first speech separation result with the best separation effect is selected as the learning target, and the initial student model is guided to be trained, so that the student model can break through the limitations of the teacher model and achieve a better separation effect than the teacher model.
[0091] Based on any of the above embodiments, Figure 4 This is the second schematic diagram of the training process of multiple student models provided by the present invention, such as Figure 4 As shown, multiple student models are also trained based on the following steps:
[0092] Step 410, inputting the second sample speech into a plurality of initial student models respectively, and obtaining second speech separation results respectively output by the plurality of initial student models;
[0093] Step 420, based on the real target speech of the second sample speech and the second speech separation results respectively output by the multiple initial student models, iteratively update the parameters of the multiple initial student models to obtain multiple student models.
[0094] Specifically, in order to give full play to the role of supervised speech data, maintain the advantages of supervised models, and further improve the speech separation effect of student models, in the training stage of multiple student models, an embodiment of the present invention also applies a small amount of supervised speech data as the second sample speech to train multiple initial student models. The specific process can be that the second sample speech is first input into multiple initial student models respectively to obtain the second speech separation results output by multiple initial student models respectively, and then according to the real label corresponding to the second sample speech, that is, the real target speech, the second speech separation results output by multiple initial student models respectively, combined with the unsupervised speech data, that is, the pseudo target speech of the first sample speech and the first speech separation results output by multiple initial student models respectively, the parameters of multiple initial student models are iteratively updated, thereby obtaining multiple student models.
[0095] It should be noted that in the process of iteratively updating the parameters of multiple initial student models by combining unsupervised speech data and supervised speech data, the unsupervised speech data can be first applied to iteratively update the parameters to obtain multiple updated initial student models, and then the supervised speech data can be applied to continue to iteratively update the parameters of the updated multiple initial student models, and finally multiple student models are obtained. It can also be that the supervised speech data is first applied to iteratively update the parameters to obtain multiple updated initial student models, and then the unsupervised speech data is applied to continue to iteratively update the parameters of the updated multiple initial student models, and finally multiple student models are obtained. It is also possible to combine the unsupervised speech data and the supervised speech data into a mixed training sample, and then apply the mixed training sample to iteratively update the parameters of multiple initial student models, and finally obtain multiple student models. The embodiments of the present invention do not make specific limitations on this.
[0096] Based on any of the above embodiments, Figure 5 It is a flow chart of the method for determining the speech separation model provided by the present invention, such as Figure 5 As shown, the speech separation model is determined based on the following steps:
[0097] Step 510, inputting the third sample speech into a plurality of student models respectively, and obtaining third speech separation results respectively output by the plurality of student models;
[0098] Step 520, based on the real target speech of the third sample speech and the third speech separation results respectively output by the plurality of student models, determining the performance evaluation results respectively corresponding to the plurality of student models;
[0099] Step 530, based on the performance evaluation results corresponding to the multiple student models, determine the speech separation model from the multiple student models.
[0100] Specifically, after obtaining multiple trained student models, in order to test the performance of multiple student models and select the student model with the best performance as the speech separation model, the embodiment of the present invention uses the supervised speech data that did not participate in the training as the third sample speech, and obtains the true label corresponding to the third sample speech, that is, the true target speech of the third sample speech, and inputs the third sample speech into multiple student models respectively to obtain the third speech separation results output by the multiple student models respectively. Then, according to the true target speech of the third sample speech and the third speech separation results output by the multiple student models respectively, the performance of the multiple student models is evaluated, so as to obtain the performance evaluation results corresponding to the multiple student models respectively, and then the speech separation model is determined from the multiple student models according to the performance evaluation results corresponding to the multiple student models respectively.
[0101] Here, the performance evaluation method can be to directly use performance indicators such as SDR (signal-to-distortion ratio) and PESQ (Perceptual Evaluation of Speech Quality) for evaluation, or to apply a loss function for evaluation. For example, a loss function is applied to calculate the loss value between the target speech in the third speech separation results output by multiple student models and the true target speech of the third sample speech, and then the student model with the smallest loss value is selected from the multiple student models as the speech separation model.
[0102] Furthermore, in the training stage of multiple student models, in addition to inputting the first sample speech and other training sample sets, other supervised speech data can also be input as a development set. After each round of training, the development set is used to evaluate whether the loss function of the model converges. If not, the next round of training continues. If converged, the training ends. At this time, the speech separation model can be determined based on the loss values corresponding to each student model obtained by applying the development set data in the last round of training, or new supervised speech data can be re-input, and the speech separation model can be determined based on the loss values corresponding to each student model obtained thereby. The embodiment of the present invention does not specifically limit this.
[0103] Based on any of the above embodiments, since the speech separation task is a regression task and there is no concept of classification posterior probability or confidence, the dual-student training method cannot be used. Existing semi-supervised model training methods for speech separation tasks usually over-trust the output of the supervised model, resulting in the upper limit of the semi-supervised model training being the supervised model (excluding additional measures such as data enhancement). Due to the constraints of the supervised model, the contribution of unsupervised data is small, and the generalization of the model has not been substantially improved.
[0104] In this regard, the present invention provides a semi-supervised model training method for speech separation tasks. Taking the multiple student models as two student models as an example, the specific implementation process of the method is as follows:
[0105] Step 1: Use a supervised dataset to pre-train a supervised model (as a teacher model).
[0106] Figure 6 It is a schematic diagram of the training process of the teacher model provided by the present invention, such as Figure 6 As shown, specifically:
[0107] Step 1-1: prepare the speaker's high-fidelity speech as the real target speech, and at least one of the speech of other speakers, actual collected noise and simulated scattered noise as the real interference audio.
[0108] Step 1-2: Superimpose the real target speech and the real interference audio in the time domain according to a certain range (for example, -10db-10db) of signal-to-interference ratio or signal-to-noise ratio, thereby constructing a mixed speech as a training sample of the teacher model, that is, the second sample speech, to simulate mixed speech in various scenarios.
[0109] Step 1-3: Train the teacher model and use SI-SNR (Scale Invariant Signal-to-noise Ratio) as the loss function (i.e. Figure 6 The supervised loss in , constrains the output target speech to approach the true target speech label of the second sample speech, that is:
[0110] Loss = SISNR (Output target , Label target )
[0111] Among them, Out target is the target speech output by the teacher model for the second sample speech, Label target is the true target speech of the second sample speech.
[0112] Step 2: Use the supervised dataset and the unsupervised dataset to train a semi-supervised model (dual student model).
[0113] The purpose of conducting semi-supervised model training instead of unsupervised model training is to make full use of the resources of supervised data, guide model learning and maintain the advantages of supervised models. For the supervised training data in the training set, that is, the second sample speech, the same approach as in Step 1-3 is adopted, using the real target speech label as the learning target to train the two student models; for the unsupervised training data, that is, the first sample speech, the specific training process is as follows:
[0114] Figure 7 This is the third schematic diagram of the training process of multiple student models provided by the present invention, such as Figure 7 As shown, specifically:
[0115] Step 2-1: Construct three models: a teacher model (T model) and two student models (S1 model and S2 model). Since conv-tasnet has the advantages of excellent separation effect, small model size and small delay, the embodiment of the present invention adopts conv-tasnet as the model structure of the teacher model and the two student models; the teacher model is the supervised model trained in Step 1-3, the model parameters are initialized using the model parameters trained in Step 1-3, and the parameters are kept unchanged during the semi-supervised training process, and only the target speech and interference audio are forward output for reference; the two student models learn independently, do not share parameters, model parameters such as weights are randomly initialized, and training parameters such as step size, learning rate, and optimizer are kept consistent.
[0116] Step 2-2: The first sample speech is used as the input of the teacher model and each student model.
[0117] Step 2-3: In the speaker recognition task, DNN (Deep Neural Networks) projects the speech segment of variable length into a fixed-dimensional speaker embedding, called x-vector. The embodiment of the present invention uses a spectral feature extractor pre-trained in the speaker recognition task, namely, an x-vector extractor (TDNN), to generate a fixed-dimensional (e.g., 400-dimensional) spectral feature embedding of the target speech and interference audio output by the teacher model and the student model. The embedding here can be used for similarity calculation to evaluate the separation effect corresponding to the unsupervised data.
[0118] Step 2-4: Calculate the cosine distance between the embedding of the target speech and the embedding of the interference audio output by each model to evaluate the separation effect of the first speech separation result output by each model. The calculation formula is:
[0119]
[0120] Among them, Emb target That is, the embedding of the target speech output by each model, Emb interf That is, the embedding of the interference audio output by each model, That is the cosine distance between the two.
[0121] The cosine similarity cos(A, B) is defined as:
[0122]
[0123] Among them, ||A|| 2 =1,||B|| 2 = 1. The value range of cosine similarity is [-1, 1], and the similarity between two identical vectors is 1. Therefore, the value range of cosine distance is [0, 2], and the cosine distance between two identical vectors is 0. In the speech separation task, the spectral features of the two speakers separated, i.e., embedding, should be discriminative, i.e. The larger the value, the better the effect of speech separation between the two speakers.
[0124] Step 2-5: According to the cosine distances corresponding to the teacher model and the two student models calculated in step 2-4, and the cosine distance judgment condition, the loss feedback and gradient update are determined. The specific cosine distance judgment condition is:
[0125] 1) If and At this time, the separation effects of the two student models are worse than those of the teacher model. For this piece of unsupervised data, the pseudo-label output by the T model is used as the learning target, and the S1 and S2 model parameters are updated at the same time. The loss function is:
[0126] Loss_S1=SISNR(out S1 ,out T )
[0127] Loss_S2 = SISNR (out S2 ,out t )
[0128] 2) If At this time, the separation effect of one student model is worse than that of the teacher model, and the separation effect of the other student model is better than that of the teacher model. For this piece of unsupervised data, the output of the S2 model is used as the learning target, and only the parameters of the S1 model are updated. The purpose of this is to break through the limitations of the teacher model and achieve a better separation effect than the teacher model. The loss function is:
[0129] Loss_S1=SISNR(out S1 ,out S2 )
[0130] Loss_S2=0
[0131] like Same reason.
[0132] 3) If At this time, the separation effect of the two student models is better than that of the teacher model. Similar to 2), for this unsupervised data, the output of the S2 model is used as the learning target, and only the parameters of the S1 model are updated. The loss function is:
[0133] Loss_S1=SISNR(out S1 ,out S2 )
[0134] Loss_S2=0
[0135] like Same reason.
[0136] Step 3: Model testing
[0137] According to the set supervised third sample speech, evaluation and comparison are performed, and the student model with better performance is selected from the S1 model or the S2 model as the speech separation model.
[0138] Figure 8 FIG. 2 is a flow chart of the speech separation method provided by the present invention. Figure 8 As shown, when the model is applied, the speech to be separated in the time domain is input into the speech separation model. After the three modules of encoding, separation and decoding, the separated time domain speech signal output by the speech separation model is obtained, which contains the clean target speech and the interference audio.
[0139] It should be noted that during the training phase of the two student models, if the parameters of the teacher model continue to be updated, due to the large amount of unsupervised data, if the student model makes incorrect predictions about certain samples, the teacher model will most likely be affected by the errors and retain such incorrect predictions, which can easily lead to the same output results of the three models and a loss value of 0, causing the three models to stop learning and ultimately obtaining a model with poor performance.
[0140] The present invention relates to the field of speech separation and semi-supervised model training based on deep learning, which integrates the ideas of mean-teacher and dual-student, and innovatively proposes to construct a teacher model and multiple student models, and uses the cosine distance of the embedding of each model output result as an indicator for evaluating the separation effect. The output result of the model with good separation effect is used as a pseudo-label of unsupervised data to guide the model with weak separation effect to train unsupervised data, thereby breaking the limitation of the traditional fixed teacher model output result as a pseudo-label for model training, and it is easier to obtain a speech separation effect that is better than that of a supervised model. Among them, when training the semi-supervised model, the cosine distance of the spectral features (embedding) of the speech output of each model will be referenced to determine whether the learning objectives, loss functions and model parameters are updated.
[0141] The speech separation device provided by the present invention is described below. The speech separation device described below and the speech separation method described above can be referred to each other.
[0142] Based on any of the above embodiments, an embodiment of the present invention provides a speech separation device. Fig. 9 Schematic diagram of the structure of the speech separation device provided by the present invention. Fig. 9 As shown, the device comprises:
[0143] A speech determination unit 910, configured to determine the speech to be separated;
[0144] A speech separation unit 920 is used to input the speech to be separated into the speech separation model to obtain the target speech of the speech to be separated output by the speech separation model;
[0145] The speech separation model is one of multiple student models, and the multiple student models are obtained by training multiple initial student models based on a first sample speech and a pseudo target speech of the first sample speech. The pseudo target speech of the first sample speech is determined based on the first speech separation results output by the teacher model and multiple initial student models for the first sample speech, and the teacher model is obtained through supervised training.
[0146] The device provided by the embodiment of the present invention constructs a teacher model and multiple student models to jointly participate in the semi-supervised model training. The teacher model is obtained according to the supervised speech data training, and the pseudo-labels of the unsupervised speech data are determined according to the output results of the teacher model and the multiple student models for the unsupervised speech data respectively, and the multiple student models are guided to be trained, so that the student model can break through the limitations of the teacher model and obtain a student model with better separation effect than the teacher model, and at the same time improve the generalization of the student model. On this basis, the student model is applied to the speech separation task of the speech to be separated, so as to obtain a purer target speech and a better speech separation effect.
[0147] Based on any of the above embodiments, multiple student models are trained based on the following steps:
[0148] Inputting the first sample speech into the teacher model and the multiple initial student models respectively, and obtaining the first speech separation results outputted by the teacher model and the multiple initial student models respectively;
[0149] Determine a first speech separation result with the best separation effect among the first speech separation results, and determine the target speech in the first speech separation result with the best separation effect as the pseudo target speech of the first sample speech;
[0150] Based on the pseudo target speech of the first sample speech and the first speech separation results respectively output by the multiple initial student models, the parameters of the multiple initial student models are iteratively updated to obtain multiple student models.
[0151] Based on any of the above embodiments, determining a first speech separation result having the best separation effect among the first speech separation results includes:
[0152] Extracting spectral features of the target speech and the interfering audio in each first speech separation result, and determining the similarity between the spectral features of the target speech and the spectral features of the interfering audio in each first speech separation result;
[0153] Based on the similarities corresponding to the first speech separation results, determining the first speech separation result corresponding to the lowest similarity, and determining the first speech separation result corresponding to the lowest similarity as the first speech separation result with the best separation effect;
[0154] Alternatively, based on the similarities corresponding to the first speech separation results, the separation effects of the first speech separation results are determined, and based on the separation effects of the first speech separation results, the first speech separation result with the best separation effect is determined.
[0155] Based on any of the above embodiments, performing spectral feature extraction on the target speech and the interference audio in each first speech separation result respectively includes:
[0156] Inputting the target speech and the interference audio in each first speech separation result into a spectral feature extractor respectively to obtain the spectral features of the target speech and the spectral features of the interference audio;
[0157] The spectral feature extractor is a feature extractor in the speaker recognition model. The speaker recognition model is trained based on the speaker's speech and the speaker information of the speaker's speech.
[0158] Based on any of the above embodiments, based on the pseudo target speech of the first sample speech and the first speech separation results respectively output by the multiple initial student models, iteratively updating the parameters of the multiple initial student models includes:
[0159] When the first speech separation result with the best separation effect is the first speech separation result output by the teacher model, iteratively updating the parameters of the multiple initial student models based on the pseudo target speech of the first sample speech and the first speech separation results respectively output by the multiple initial student models;
[0160] When the first speech separation result with the best separation effect is the first speech separation result output by any initial student model, the parameters of other initial student models are iteratively updated based on the pseudo-target speech of the first sample speech and the first speech separation results output by other initial student models respectively. The other initial student models are initial student models other than any initial student model among multiple initial student models.
[0161] Based on any of the above embodiments, multiple student models are also trained based on the following steps:
[0162] Inputting the second sample speech into the multiple initial student models respectively, and obtaining the second speech separation results outputted by the multiple initial student models respectively;
[0163] Based on the real target speech of the second sample speech and the second speech separation results respectively output by the multiple initial student models, the parameters of the multiple initial student models are iteratively updated to obtain the multiple student models.
[0164] Based on any of the above embodiments, the speech separation model is determined based on the following steps:
[0165] Inputting the third sample speech into the plurality of student models respectively, and obtaining the third speech separation results outputted by the plurality of student models respectively;
[0166] Based on the real target speech of the third sample speech and the third speech separation results respectively output by the plurality of student models, determining the performance evaluation results respectively corresponding to the plurality of student models;
[0167] Based on the performance evaluation results respectively corresponding to the multiple student models, a speech separation model is determined from the multiple student models.
[0168] Fig.10 An example of a physical structure diagram of an electronic device is shown in FIG. Fig.10As shown, the electronic device may include: a processor 1010, a communication interface 1020, a memory 1030 and a communication bus 1040, wherein the processor 1010, the communication interface 1020 and the memory 1030 communicate with each other through the communication bus 1040. The processor 1010 may call the logic instructions in the memory 1030 to execute the speech separation method, which includes: determining the speech to be separated; inputting the speech to be separated into a speech separation model to obtain the target speech of the speech to be separated output by the speech separation model; the speech separation model is one of a plurality of student models, the plurality of student models are obtained by training a plurality of initial student models based on a first sample speech and a pseudo target speech of the first sample speech, the pseudo target speech of the first sample speech is determined based on the first speech separation results output by the teacher model and the plurality of initial student models for the first sample speech, respectively, and the teacher model is obtained by supervised training.
[0169] In addition, the logic instructions in the above-mentioned memory 1030 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.
[0170] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the speech separation method provided by the above-mentioned methods, which includes: determining the speech to be separated; inputting the speech to be separated into a speech separation model to obtain the target speech of the speech to be separated output by the speech separation model; the speech separation model is one of a plurality of student models, and the plurality of student models are obtained by training a plurality of initial student models based on a first sample speech and a pseudo target speech of the first sample speech, and the pseudo target speech of the first sample speech is determined based on a first speech separation result output by a teacher model and the plurality of initial student models for the first sample speech, and the teacher model is obtained by supervised training.
[0171] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the speech separation method provided by the above-mentioned methods, the method comprising: determining a speech to be separated; inputting the speech to be separated into a speech separation model to obtain a target speech of the speech to be separated output by the speech separation model; the speech separation model is one of a plurality of student models, the plurality of student models are obtained by training a plurality of initial student models based on a first sample speech and a pseudo target speech of the first sample speech, the pseudo target speech of the first sample speech is determined based on a first speech separation result output by a teacher model and the plurality of initial student models for the first sample speech, and the teacher model is obtained by supervised training.
[0172] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0173] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A speech separation method, It is characterized in that include: Determine the speech to be separated; Inputting the speech to be separated into a speech separation model to obtain a target speech of the speech to be separated output by the speech separation model; The speech separation model is one of a plurality of student models, the plurality of student models are obtained by training a plurality of initial student models based on a first sample speech and a pseudo target speech of the first sample speech, the pseudo target speech of the first sample speech is determined based on a first speech separation result outputted by a teacher model and the plurality of initial student models for the first sample speech, and the teacher model is obtained by supervised training; The speech separation model is determined based on the following steps: Inputting the third sample speech into the plurality of student models respectively, and obtaining the third speech separation results respectively output by the plurality of student models; Based on the real target speech of the third sample speech and the third speech separation results respectively output by the multiple student models, determining the performance evaluation results respectively corresponding to the multiple student models; Based on the performance evaluation results respectively corresponding to the multiple student models, the speech separation model is determined from the multiple student models.
2. The speech separation method according to claim 1, It is characterized in that The multiple student models are trained based on the following steps: Inputting the first sample speech into the teacher model and the multiple initial student models respectively, to obtain first speech separation results outputted by the teacher model and the multiple initial student models respectively; Determine a first speech separation result with the best separation effect among the first speech separation results, and determine the target speech in the first speech separation result with the best separation effect as the pseudo target speech of the first sample speech; Based on the pseudo target speech of the first sample speech and the first speech separation results respectively output by the multiple initial student models, the multiple initial student models are iteratively updated with parameters to obtain the multiple student models.
3. The speech separation method according to claim 2, It is characterized in that The determining of the first speech separation result having the best separation effect among the first speech separation results comprises: Extracting spectral features of the target speech and the interference audio in each of the first speech separation results, and determining the similarity between the spectral features of the target speech and the spectral features of the interference audio in each of the first speech separation results; Based on the similarities corresponding to the first speech separation results, determining the first speech separation result corresponding to the lowest similarity, and determining the first speech separation result corresponding to the lowest similarity as the first speech separation result with the best separation effect; Alternatively, based on the similarities corresponding to the first speech separation results, the separation effects of the first speech separation results are determined, and based on the separation effects of the first speech separation results, the first speech separation result with the best separation effect is determined.
4. The speech separation method according to claim 3, It is characterized in that The extracting spectral features of the target speech and the interference audio in each of the first speech separation results respectively includes: Inputting the target speech and the interference audio in each of the first speech separation results into a spectral feature extractor respectively to obtain the spectral features of the target speech and the spectral features of the interference audio; The spectral feature extractor is a feature extractor in a speaker recognition model, and the speaker recognition model is trained based on a speaker's speech and speaker information of the speaker's speech.
5. The speech separation method according to claim 2, It is characterized in that The pseudo target speech based on the first sample speech and the first speech separation results respectively output by the multiple initial student models, iteratively updating the parameters of the multiple initial student models, comprises: When the first speech separation result with the best separation effect is the first speech separation result output by the teacher model, iteratively updating the parameters of the multiple initial student models based on the pseudo target speech of the first sample speech and the first speech separation results respectively output by the multiple initial student models; When the first speech separation result with the best separation effect is the first speech separation result output by any initial student model, the parameters of the other initial student models are iteratively updated based on the pseudo target speech of the first sample speech and the first speech separation results respectively output by other initial student models. The other initial student models are the initial student models among the multiple initial student models except any initial student model.
6. The speech separation method according to claim 2, It is characterized in that The multiple student models are also trained based on the following steps: Inputting the second sample speech into the multiple initial student models respectively, and obtaining the second speech separation results respectively output by the multiple initial student models; Based on the real target speech of the second sample speech and the second speech separation results respectively outputted by the multiple initial student models, the multiple initial student models are iteratively updated with parameters to obtain the multiple student models.
7. A speech separation device, It is characterized in that include: A speech determination unit, used for determining the speech to be separated; A speech separation unit, used for inputting the speech to be separated into a speech separation model to obtain a target speech of the speech to be separated output by the speech separation model; The speech separation model is one of a plurality of student models, the plurality of student models are obtained by training a plurality of initial student models based on a first sample speech and a pseudo target speech of the first sample speech, the pseudo target speech of the first sample speech is determined based on a first speech separation result outputted by a teacher model and the plurality of initial student models for the first sample speech, and the teacher model is obtained by supervised training; The speech separation model is determined based on the following steps: Inputting the third sample speech into the plurality of student models respectively, and obtaining the third speech separation results respectively output by the plurality of student models; Based on the real target speech of the third sample speech and the third speech separation results respectively output by the multiple student models, determining the performance evaluation results respectively corresponding to the multiple student models; Based on the performance evaluation results respectively corresponding to the multiple student models, the speech separation model is determined from the multiple student models.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the speech separation method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the speech separation method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Model training method and device and voice signal processing method and device
CN113380268A