Method and device for training human voice accompaniment separation model, and electronic equipment
By training sample audio through a hybrid training method using a vocal accompaniment separation network and a discriminator, the problems of limited supervised audio quantity and high cost are solved, improving the generalization ability and separation accuracy of the vocal accompaniment separation model and enhancing the singing experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-02
- Publication Date
- 2026-03-31
AI Technical Summary
In existing technologies, the amount of supervised audio is small and the cost of obtaining it is high, resulting in high training costs for vocal accompaniment separation models and affecting the singing experience.
The sample audio is separated by the human voice and accompaniment separation network in the human voice and accompaniment separation model. The prediction results are judged by combining the human voice discriminator and the accompaniment discriminator. Supervised and unsupervised sample audio are used for mixed training to reduce the dependence on supervised sample audio.
It improves the generalization ability and separation accuracy of the vocal accompaniment separation model, reduces training costs, and enhances the singing experience.
Smart Images

Figure CN116259329B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of Internet technology, and in particular to a training method, apparatus and electronic device for a vocal accompaniment separation model. Background Technology
[0002] With the development of internet technology, singing through karaoke apps has become an increasingly popular form of entertainment. However, the audio provided to these apps by record companies is a mixed track, meaning vocals and accompaniment are superimposed. When a singer uses these apps, the accompaniment includes vocals, preventing them from enjoying a clean performance and resulting in a poor singing experience. Therefore, separating the vocals and accompaniment from the audio is a problem that needs to be solved.
[0003] Currently, the common approach is to train a vocal-accompaniment separation model using supervised audio to separate the vocals and accompaniment. In this model, the supervised audio serves as the input, while the individual vocals and accompaniment are used as references.
[0004] The problem with the above technical solutions is that the amount of supervised audio is small and the acquisition cost is high, resulting in high training costs for the vocal accompaniment separation model. Summary of the Invention
[0005] This disclosure provides a training method, apparatus, and electronic device for a vocal-accompaniment separation model, which improves the generalization ability of the model and enhances the accuracy of the trained model in separating vocals and accompaniment. The technical solution of this disclosure is as follows:
[0006] According to one aspect of the embodiments of this disclosure, a method for training a vocal accompaniment separation model is provided, comprising:
[0007] For any one of the multiple sample audios, the human voice and accompaniment are separated based on the human voice and accompaniment separation network in the human voice and accompaniment separation model to obtain the predicted human voice and the predicted accompaniment. The multiple sample audios include supervised sample audios with supervised information and unsupervised sample audios without supervised information. The human voice and accompaniment separation model is used to separate the human voice and accompaniment in the input audio. The human voice and accompaniment separation network is pre-trained based on the supervised sample audios in the multiple sample audios.
[0008] Based on the human voice discriminator in the human voice accompaniment separation model, the predicted human voice is discriminated to obtain a first discrimination result, which is used to indicate whether it is a predicted human voice;
[0009] Based on the accompaniment discriminator in the vocal accompaniment separation model, the predicted accompaniment is discriminated to obtain a second discrimination result, which is used to indicate whether it is a predicted accompaniment;
[0010] The vocal-accompaniment separation model is trained based on the sample audio, the predicted human voice, the predicted accompaniment, the first discrimination result, and the second discrimination result.
[0011] According to another aspect of the present disclosure, a training apparatus for a vocal accompaniment separation model is provided, comprising:
[0012] The separation unit is configured to perform vocal-accompaniment separation on any sample audio from a plurality of sample audios, based on the vocal-accompaniment separation network in the vocal-accompaniment separation model, to obtain predicted vocals and predicted accompaniment. The plurality of sample audios include supervised sample audios with supervised information and unsupervised sample audios without supervised information. The vocal-accompaniment separation model is used to separate vocals and accompaniment in the input audio. The vocal-accompaniment separation network is pre-trained based on supervised sample audios from the plurality of sample audios.
[0013] The first discrimination unit is configured to discriminate the predicted human voice based on the human voice discriminator in the human voice accompaniment separation model, and obtain a first discrimination result, wherein the first discrimination result is used to indicate whether it is a predicted human voice;
[0014] The second discrimination unit is configured to discriminate the predicted accompaniment based on the accompaniment discriminator in the vocal accompaniment separation model, and obtain a second discrimination result, which is used to indicate whether it is a predicted accompaniment.
[0015] The training unit is configured to train the vocal-accompaniment separation model based on the sample audio, the predicted human voice, the predicted accompaniment, the first discrimination result, and the second discrimination result.
[0016] In some embodiments, the training unit includes:
[0017] The first determining subunit is configured to determine a first generator loss based on the sample audio, the predicted human voice, the predicted accompaniment, the first discrimination result, and the second discrimination result when the sample audio is supervised sample audio. The first generator loss is used to indicate the loss of the human voice and accompaniment separation network.
[0018] The second determining subunit is configured to determine a first discriminator loss based on the sample audio, the first discrimination result, the second discrimination result, the third discrimination result, and the fourth discrimination result. The third discrimination result is the discrimination result of the human voice discriminator on the reference human voice in the sample audio, and the fourth discrimination result is the discrimination result of the accompaniment discriminator on the reference accompaniment in the sample audio. The first discriminator loss is used to indicate the loss of the human voice discriminator and the accompaniment discriminator.
[0019] The first training subunit is configured to train the vocal accompaniment separation model based on the first generator loss and the first discriminator loss.
[0020] In some embodiments, the first determining subunit is configured to, when the sample audio is supervised sample audio, determine a first training loss based on a reference human voice in the sample audio, the predicted human voice, and the first discrimination result; determine a second training loss based on a reference accompaniment in the sample audio, the predicted accompaniment, and the second discrimination result; and determine the sum of the first training loss and the second training loss as the first generator loss.
[0021] In some embodiments, the second determining subunit is configured to determine a third training loss based on a reference human voice of the sample audio, the first discrimination result, and the third discrimination result; determine a fourth training loss based on a reference accompaniment of the sample audio, the second discrimination result, and the fourth discrimination result; and determine the sum of the third training loss and the fourth training loss as the first discriminator loss.
[0022] In some embodiments, the training unit includes:
[0023] The third determining subunit is configured to determine a second discriminator loss based on the sample audio, the first discrimination result, and the second discrimination result when the sample audio is unsupervised sample audio. The second discriminator loss is used to indicate the loss of the human voice discriminator and the accompaniment discriminator.
[0024] The fourth determining subunit is configured to determine a second generator loss based on the sample audio, the predicted human voice, the predicted accompaniment, the first discrimination result, and the second discrimination result. The second generator loss is used to indicate the loss of the human voice and accompaniment separation network.
[0025] The second training subunit is configured to train the vocal accompaniment separation model based on the second generator loss and the second discriminator loss.
[0026] In some embodiments, the third determining subunit is configured to, when the sample audio is unsupervised sample audio, determine a fifth training loss based on the sample audio and the first discrimination result; determine a sixth training loss based on the sample audio and the second discrimination result; and determine the sum of the fifth training loss and the sixth training loss as the second discriminator loss.
[0027] In some embodiments, the separation unit is configured to extract features from the sample audio based on the vocal-accompaniment separation model to obtain multiple audio feature information; cluster the multiple audio feature information based on the vocal-accompaniment separation model to obtain multiple clustering results, the clustering results being used to indicate the vocals and accompaniment in the sample audio; and process the multiple clustering results based on the vocal-accompaniment separation model to obtain the predicted vocals and the predicted accompaniment.
[0028] In some embodiments, the apparatus further includes:
[0029] The transform unit is configured to perform a short-time Fourier transform on any supervised audio sample to obtain the amplitude spectrum and phase spectrum of the supervised audio sample in the frequency domain.
[0030] The processing unit is configured to process the amplitude spectrum of the supervised sample audio based on the human voice accompaniment separation network to obtain the amplitude spectrum of the pre-trained human voice and the amplitude spectrum of the pre-trained accompaniment.
[0031] The acquisition unit is configured to obtain the pre-trained human voice and the pre-trained accompaniment based on the phase spectrum of the supervised sample audio, the amplitude spectrum of the pre-trained human voice, and the amplitude spectrum of the pre-trained accompaniment.
[0032] The determining unit is configured to determine the pre-training loss based on the difference between the pre-trained human voice and the reference human voice of the supervised sample audio, and the difference between the pre-trained accompaniment and the reference accompaniment of the supervised sample audio.
[0033] The pre-training unit is configured to pre-train the vocal accompaniment separation network based on the pre-training loss.
[0034] In some embodiments, the determining unit is configured to determine a first signal distortion ratio based on the difference between the pre-trained human voice and a reference human voice in the supervised sample audio; determine a second signal distortion ratio based on the difference between the pre-trained accompaniment and a reference accompaniment in the supervised sample audio; and determine the pre-training loss by weighted summation of the first signal distortion ratio and the second signal distortion ratio.
[0035] According to another aspect of the embodiments of this disclosure, an electronic device is provided, comprising:
[0036] One or more processors;
[0037] Memory used to store the executable program code of the processor;
[0038] The processor is configured to execute the program code to implement the training method of the above-mentioned vocal accompaniment separation model.
[0039] According to another aspect of the present disclosure, a computer-readable storage medium is provided, which, when the program code in the computer-readable storage medium is executed by the processor of an electronic device, enables the electronic device to perform the training method of the above-described vocal accompaniment separation model.
[0040] According to another aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the training method of the above-described vocal accompaniment separation model.
[0041] This disclosure provides a training method for a vocal-accompaniment separation model. By using a vocal-accompaniment separation network within the model, predicted vocals and accompaniment are separated from any sample audio, achieving separation of vocals and accompaniment in the sample audio. Then, based on the vocal discriminator and accompaniment discriminator in the model, the predicted vocals and accompaniment are separately discriminated, yielding a first discrimination result and a second discrimination result. The training loss is then determined using the sample audio, the predicted vocals and accompaniment, and the discriminator's discrimination results, thereby updating the parameters in the vocal-accompaniment separation model and training it. Since the sample audio includes a small number of supervised sample audios and a large number of unsupervised sample audios, the high cost of acquiring supervised sample audios is avoided, improving the generalization ability of the vocal-accompaniment separation model and increasing the accuracy of the separated vocals and accompaniment by the trained model.
[0042] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0043] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0044] Figure 1 This is a schematic diagram illustrating the implementation environment of a training method for a vocal accompaniment separation model according to an exemplary embodiment.
[0045] Figure 2 This is a flowchart illustrating a training method for a vocal accompaniment separation model according to an exemplary embodiment.
[0046] Figure 3 This is a flowchart illustrating another training method for a vocal accompaniment separation model according to an exemplary embodiment.
[0047] Figure 4 This is a flowchart illustrating a supervised sample audio pre-trained vocal accompaniment separation network according to an exemplary embodiment.
[0048] Figure 5 This is a flowchart illustrating a vocal accompaniment separation model trained based on supervised sample audio, according to an exemplary embodiment.
[0049] Figure 6 This is a flowchart illustrating a vocal accompaniment separation model trained based on unsupervised sample audio, according to an exemplary embodiment.
[0050] Figure 7 This is a schematic diagram illustrating the application of a vocal accompaniment separation model according to an exemplary embodiment.
[0051] Figure 8 This is a block diagram illustrating a training apparatus for a vocal accompaniment separation model according to an exemplary embodiment.
[0052] Figure 9 This is a block diagram of a training apparatus for another vocal accompaniment separation model, according to an exemplary embodiment.
[0053] Figure 10 This is a block diagram illustrating an electronic device according to an exemplary embodiment.
[0054] Figure 11 This is a schematic diagram of the structure of a server according to an exemplary embodiment. Detailed Implementation
[0055] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0056] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0057] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this disclosure are authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the sample audio involved in this disclosure was obtained with full authorization.
[0058] Figure 1 This is a schematic diagram illustrating the implementation environment of a training method for a vocal accompaniment separation model according to an exemplary embodiment. See also Figure 1 The implementation environment specifically includes: terminal 101 and server 102.
[0059] Terminal 101 can be at least one of the following devices: smartphone, smartwatch, desktop computer, laptop, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), and laptop computer. Terminal 101 can have an application for audio playback installed and running. Terminal 101 can connect to server 102 via a wireless network or a wired network.
[0060] Terminal 101 can refer to one of a plurality of terminals; this embodiment uses terminal 101 as an example only. Those skilled in the art will understand that the number of terminals can be more or less. For example, there may be only a few terminals, or dozens or hundreds, or even more. This disclosure does not limit the number of terminals or the type of device.
[0061] Server 102 can be at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 102 can be connected to terminal 101 and other terminals via a wireless network or a wired network. Optionally, the number of servers can be more or less, and this disclosure does not limit this. Of course, server 102 may also include other functional servers to provide more comprehensive and diversified services.
[0062] Figure 2 This is a flowchart illustrating a training method for a vocal accompaniment separation model according to an exemplary embodiment, such as... Figure 2 As shown, this method is executed by the server and includes the following steps.
[0063] In step S201, for any sample audio from multiple sample audios, the server performs voice and accompaniment separation on the sample audio based on the voice and accompaniment separation network in the voice and accompaniment separation model to obtain predicted voice and predicted accompaniment. The multiple sample audios include supervised sample audios with supervised information and unsupervised sample audios without supervised information. The voice and accompaniment separation model is used to separate the voice and accompaniment in the input audio. The voice and accompaniment classification network is pre-trained based on supervised sample audios from multiple sample audios.
[0064] In this embodiment, the multiple sample audio files are used to train a vocal-accompaniment separation model. These sample audio files include a limited number of supervised sample audio files and a large number of unsupervised sample audio files. The server can train the vocal-accompaniment separation model based on the supervised and unsupervised sample audio files. Accordingly, based on the pre-trained vocal-accompaniment separation network, the server performs vocal-accompaniment separation on any input sample audio file to obtain the predicted vocals and predicted accompaniment. This approach enables hybrid training of the vocal-accompaniment separation model using both supervised and unsupervised sample audio files, reducing the need for supervised sample audio files and thus reducing the cost of acquiring them. This reduces the cost of training the vocal-accompaniment separation model and improves its generalization ability.
[0065] In step S202, the server uses the voice discriminator in the voice accompaniment separation model to discriminate the predicted voice and obtain a first discrimination result. The first discrimination result is used to indicate whether it is the predicted voice.
[0066] In this embodiment, the voice discriminator can discriminate the input audio, and the output discrimination result varies depending on the discrimination. For example, if the voice discriminator discriminates the input audio as a human voice predicted by the voice accompaniment separation network, then the voice discriminator outputs 0; if the voice discriminator discriminates the input audio as a reference human voice used as supervision information, then the voice discriminator outputs 1. Typically, the discriminator cannot discriminate perfectly accurately, therefore the discriminator output is a prediction result, i.e., the first discrimination result includes the probability of predicting 0 and the probability of predicting 1. Therefore, the server can use this voice discriminator to discriminate the predicted human voice separated by the voice accompaniment separation network to obtain the first discrimination result. By discriminating the predicted human voice with the voice discriminator, the obtained first discrimination result can be used to determine the training loss of the voice discriminator, thereby training the voice discriminator.
[0067] In step S203, the server uses the accompaniment discriminator in the vocal accompaniment separation model to discriminate the predicted accompaniment and obtain a second discrimination result. The second discrimination result is used to indicate whether it is the predicted accompaniment.
[0068] In this embodiment, the accompaniment discriminator can discriminate the input audio, and the output discrimination result varies depending on the discrimination. For example, if the accompaniment discriminator discriminates the input audio as the accompaniment predicted by the vocal accompaniment separation network, then the accompaniment discriminator outputs 0; if the accompaniment discriminator discriminates the input audio as a reference accompaniment used as supervision information, then the accompaniment discriminator outputs 1. Typically, the discriminator cannot discriminate perfectly accurately, therefore the discriminator output is a prediction result, i.e., the second discrimination result includes the probability of a prediction of 0 and the probability of a prediction of 1. Therefore, the server can use this accompaniment discriminator to discriminate the predicted accompaniment separated by the vocal accompaniment separation network to obtain the second discrimination result. By discriminating the predicted accompaniment using the accompaniment discriminator, the obtained second discrimination result can be used to determine the training loss of the accompaniment discriminator, thereby training the accompaniment discriminator.
[0069] In step S204, the server trains the human voice and accompaniment separation model based on the sample audio, predicted human voice, predicted accompaniment, first discrimination result, and second discrimination result.
[0070] In this embodiment, since the predicted human voice and predicted accompaniment can reflect the separation effect of the human voice and accompaniment separation network, and the first and second discrimination results can reflect the discrimination status of the human voice discriminator and the accompaniment discriminator, the server can obtain the training losses of the corresponding human voice and accompaniment separation network, the human voice discriminator, and the accompaniment discriminator based on the sample audio, the predicted human voice, the predicted accompaniment, the first discrimination result, and the second discrimination result, thereby training the human voice and accompaniment separation model and improving the accuracy of the human voice and accompaniment obtained by the human voice and accompaniment separation model.
[0071] This disclosure provides a training method for a vocal-accompaniment separation model. By using a vocal-accompaniment separation network within the model, predicted vocals and accompaniment are separated from any sample audio, achieving separation of vocals and accompaniment in the sample audio. Then, based on the vocal discriminator and accompaniment discriminator in the model, the predicted vocals and accompaniment are separately discriminated, yielding a first discrimination result and a second discrimination result. The training loss is then determined using the sample audio, the predicted vocals and accompaniment, and the discriminator's discrimination results, thereby updating the parameters in the vocal-accompaniment separation model and training it. Since the sample audio includes a small number of supervised sample audios and a large number of unsupervised sample audios, the high cost of acquiring supervised sample audios is avoided, improving the generalization ability of the vocal-accompaniment separation model and increasing the accuracy of the separated vocals and accompaniment by the trained model.
[0072] In some embodiments, the vocal-accompaniment separation model is trained based on sample audio, predicted human voice, predicted accompaniment, a first discrimination result, and a second discrimination result, including:
[0073] When the sample audio is supervised sample audio, the first generator loss is determined based on the sample audio, predicted human voice, predicted accompaniment, first discrimination result and second discrimination result. The first generator loss is used to indicate the loss of the human voice and accompaniment separation network.
[0074] Based on the sample audio, the first discrimination result, the second discrimination result, the third discrimination result, and the fourth discrimination result, the first discriminator loss is determined. The third discrimination result is the discrimination result of the voice discriminator on the reference human voice in the sample audio, and the fourth discrimination result is the discrimination result of the accompaniment discriminator on the reference accompaniment in the sample audio. The first discriminator loss is used to indicate the loss of the voice discriminator and the accompaniment discriminator.
[0075] The vocal accompaniment separation model is trained based on the first generator loss and the first discriminator loss.
[0076] By using supervised audio samples, the first generator loss and the first discriminator loss are determined based on the supervision information of the sample audio, the prediction and discrimination of the vocal accompaniment separation model, and thus the training of the vocal accompaniment separation model can be achieved.
[0077] In some embodiments, when the sample audio is supervised sample audio, a first generator loss is determined based on the sample audio, predicted human voice, predicted accompaniment, a first discrimination result, and a second discrimination result, including:
[0078] When the sample audio is supervised sample audio, the first training loss is determined based on the reference human voice, the predicted human voice and the first discrimination result in the sample audio.
[0079] The second training loss is determined based on the reference accompaniment, the predicted accompaniment, and the second discrimination result in the sample audio.
[0080] The sum of the first training loss and the second training loss is determined as the first generator loss.
[0081] By determining the first generator loss of the vocal accompaniment separation network, the vocal accompaniment separation network can be trained based on the first generator loss.
[0082] In some embodiments, the first discriminator loss is determined based on the sample audio, the first discrimination result, the second discrimination result, the third discrimination result, and the fourth discrimination result, including:
[0083] Based on the reference human voice in the sample audio, the first discrimination result, and the third discrimination result, the third training loss is determined;
[0084] Based on the reference accompaniment of the sample audio, the second discrimination result, and the fourth discrimination result, the fourth training loss is determined;
[0085] The sum of the third and fourth training losses is determined as the first discriminator loss.
[0086] By determining the first discriminator loss for the voice discriminator and the accompaniment discriminator, the voice discriminator and the accompaniment discriminator can be trained based on the first discriminator loss.
[0087] In some embodiments, the vocal-accompaniment separation model is trained based on sample audio, predicted human voice, predicted accompaniment, a first discrimination result, and a second discrimination result, including:
[0088] When the sample audio is unsupervised sample audio, the second discriminator loss is determined based on the sample audio, the first discrimination result, and the second discrimination result. The second discriminator loss is used to indicate the loss of the voice discriminator and the accompaniment discriminator.
[0089] Based on the sample audio, predicted human voice, predicted accompaniment, first discrimination result, and second discrimination result, the second generator loss is determined. The second generator loss is used to indicate the loss of the human voice and accompaniment separation network.
[0090] The vocal accompaniment separation model is trained based on the second generator loss and the second discriminator loss.
[0091] By determining the second generator loss and the second discriminator loss based on the sample audio, the prediction and discrimination results of the vocal accompaniment separation model when the sample audio is unsupervised, the training of the vocal accompaniment separation model can be achieved.
[0092] In some embodiments, when the sample audio is unsupervised sample audio, the second discriminator loss is determined based on the sample audio, the first discrimination result, and the second discrimination result, including:
[0093] When the sample audio is unsupervised sample audio, the fifth training loss is determined based on the sample audio and the first discrimination result;
[0094] Based on the sample audio and the second discrimination result, the sixth training loss is determined;
[0095] The sum of the fifth and sixth training losses is determined as the second discriminator loss.
[0096] By determining the second discriminator loss of the vocal and accompaniment separation network, the vocal discriminator and the accompaniment discriminator can be trained based on this second discriminator loss.
[0097] In some embodiments, for any sample audio from a plurality of sample audios, the sample audio is separated into a predicted human voice and a predicted accompaniment based on the human voice and accompaniment separation network in the human voice and accompaniment separation model, including:
[0098] Based on the vocal accompaniment separation model, feature extraction is performed on the sample audio to obtain multiple audio feature information;
[0099] Based on the human voice and accompaniment separation model, multiple audio feature information is clustered to obtain multiple clustering results. The clustering results are used to indicate the human voice and accompaniment in the sample audio.
[0100] Based on the vocal and accompaniment separation model, multiple clustering results are processed to obtain predicted vocals and predicted accompaniment.
[0101] By clustering the features of sample audio, predicted vocals and predicted accompaniment are obtained, thereby achieving the separation of vocals and accompaniment in audio.
[0102] In some embodiments, the method further includes:
[0103] For any supervised audio sample, perform a short-time Fourier transform on the supervised audio sample to obtain the amplitude spectrum and phase spectrum of the supervised audio sample in the frequency domain;
[0104] Based on the human voice accompaniment separation network, the amplitude spectrum of supervised sample audio is processed to obtain the amplitude spectrum of pre-trained human voice and the amplitude spectrum of pre-trained accompaniment.
[0105] Based on the phase spectrum of supervised sample audio, the amplitude spectrum of pre-trained human voice, and the amplitude spectrum of pre-trained accompaniment, the pre-trained human voice and pre-trained accompaniment are obtained.
[0106] The pre-training loss is determined based on the differences between the pre-trained human voice and the reference human voice in the supervised sample audio, as well as the differences between the pre-trained accompaniment and the reference accompaniment in the supervised sample audio.
[0107] The vocal accompaniment separation network is pre-trained based on the pre-training loss.
[0108] By obtaining the amplitude spectrum and phase spectrum of supervised sample audio in the frequency domain, it is possible to obtain pre-trained vocals and pre-trained accompaniment based on the amplitude spectrum and phase spectrum, thereby realizing the pre-training of the vocal and accompaniment separation network.
[0109] In some embodiments, the pre-training loss is determined based on the differences between the pre-trained human voice and the reference human voice in the supervised sample audio, and the differences between the pre-trained accompaniment and the reference accompaniment in the supervised sample audio, including:
[0110] The first signal distortion ratio is determined based on the difference between the pre-trained human voice and the reference human voice in the supervised sample audio.
[0111] The second signal distortion ratio is determined based on the difference between the pre-trained accompaniment and the reference accompaniment of the supervised sample audio.
[0112] The pre-training loss is determined by weighted summation of the first signal distortion ratio and the second signal distortion ratio.
[0113] By comparing the signal distortion ratio between the actual vocals and accompaniment and the predicted vocals and accompaniment, the error between the vocal and accompaniment separation network and the actual situation during the pre-training process can be obtained, thus enabling the pre-training of the vocal and accompaniment separation network.
[0114] Figure 3 This is a flowchart illustrating another training method for a vocal accompaniment separation model according to an exemplary embodiment, such as... Figure 3 As shown, this method is executed by the server and includes the following steps.
[0115] In step S301, for any supervised audio sample, the server performs a short-time Fourier transform on the supervised audio sample to obtain the amplitude spectrum and phase spectrum of the supervised audio sample in the frequency domain.
[0116] In this embodiment, the server can pre-train a vocal-accompaniment separation network based on supervised sample audio from the sample audio, preparing for subsequent training of the vocal-accompaniment separation model. The supervised sample audio includes original sample audio and mixed sample audio. Original sample audio refers to unprocessed audio, such as an unprocessed song. The mixed sample audio is obtained by mixing reference vocals and reference accompaniment from the original sample audio, with the mixing method being random mixing of reference vocals and reference accompaniment from different original sample audios. For any supervised sample audio, the server can perform a short-time Fourier transform on the supervised sample audio to obtain its amplitude spectrum and phase spectrum. By obtaining the amplitude and phase spectra of the supervised sample audio in the frequency domain, pre-trained vocals and pre-trained accompaniment can be obtained based on the amplitude and phase spectra, thereby achieving pre-training of the vocal-accompaniment separation network.
[0117] For example, any original sample audio can be represented in the time domain as a pair of data, i.e., (x(t), s1(t), s2(t)). Here, x(t) represents the original sample audio; s1(t) represents the reference vocal in the original sample audio; and s2(t) represents the reference accompaniment in the original sample audio. Since the number of original sample audios is finite, the server can mix the reference vocal and reference accompaniment in the original sample audio to obtain a mixed sample audio. This mixed sample audio can also be represented as a pair of data, i.e., (x1(t), s3(t), s6(t)). Here, x1(t) represents the mixed sample audio, obtained by mixing the reference vocal s3(t) and the reference accompaniment s6(t); s3(t) represents the reference vocal in any original sample audio; and s6(t) represents the reference accompaniment in any original sample audio other than the original sample audio to which s3(t) belongs.
[0118] In some embodiments, for any supervised audio sample, the server can perform a short-time Fourier transform on the supervised audio sample, converting it from the time domain to the frequency domain. The server can then obtain the frequency domain features of the supervised audio sample, namely the amplitude spectrum features and phase spectrum features, as the features of the supervised audio sample. By converting the supervised audio sample in the time domain to the frequency domain, the amplitude spectrum features and phase spectrum features in the frequency domain are obtained, thereby enabling pre-training of the vocal accompaniment separation network based on the frequency domain features.
[0119] In some embodiments, the server can perform a short-time Fourier transform on any supervised audio sample in the time domain using the following formula (1).
[0120] X(n,k)=STFT(x(t)) (1)
[0121] Where X(n, k) represents the frequency domain representation of the supervised sample audio x(t) after short-time Fourier transform; t represents time, with a value range of 0 < t ≤ T, and T represents the signal length of the supervised sample audio; n represents the frame sequence, with a value range of 0 < n ≤ N, and N represents the total number of frames; k represents the center frequency sequence, with a value range of 0 < k ≤ K, and K represents the total number of frequency points; STFT() represents the short-time Fourier transform function.
[0122] In some embodiments, the server can determine the amplitude spectrum |X(n,k)| and phase spectrum ∠X(n,k) of the supervised sample audio using the above formula (1).
[0123] It should be noted that this disclosure uses frequency domain feature extraction technology as an example for illustration. The server can also extract features from supervised audio samples using time domain or complex domain feature extraction technology. For example, for time domain feature extraction technology, a one-dimensional convolutional neural network can be used to convert the supervised audio sample signal in the time domain into a feature representation of the signal. For complex domain feature extraction technology, a short-time Fourier transform can be performed on the supervised audio sample to obtain a complex spectrum. This disclosure does not impose any limitations on this.
[0124] In step S302, the server processes the amplitude spectrum of the supervised sample audio based on the human voice accompaniment separation network to obtain the amplitude spectrum of the pre-trained human voice and the amplitude spectrum of the pre-trained accompaniment.
[0125] In this embodiment, the server uses the amplitude spectrum of the supervised sample audio obtained in the above steps as input to the vocal accompaniment separation network. This vocal accompaniment separation network can be composed of at least one of a feedforward neural network, a convolutional neural network, or a recurrent neural network; this embodiment does not impose any limitation on this. The vocal accompaniment separation network can estimate the masks of the pre-trained vocals and pre-trained accompaniments separately, thereby obtaining the amplitude spectra of the pre-trained vocals and pre-trained accompaniments, providing data preparation for obtaining the pre-trained vocals and pre-trained accompaniments.
[0126] In some embodiments, the server can obtain the amplitude spectrum of the pre-trained vocals and the amplitude spectrum of the pre-trained accompaniment by the following formula (2).
[0127]
[0128] Where i represents pre-trained vocals or pre-trained accompaniment; Si This represents the amplitude spectrum of the pre-trained vocals or the pre-trained accompaniment; M i |X| represents the mask for pre-trained vocals or pre-trained accompaniment; |X| represents the amplitude spectrum of supervised sample audio; G represents the vocal / accompaniment separation network.
[0129] In step S303, the server obtains the pre-trained human voice and the pre-trained accompaniment based on the phase spectrum of the supervised sample audio, the amplitude spectrum of the pre-trained human voice, and the amplitude spectrum of the pre-trained accompaniment.
[0130] In this embodiment of the disclosure, since the amplitude spectra of the pre-trained human voice and the pre-trained accompaniment in the supervised sample audio are obtained in the frequency domain, the server can use inverse short-time Fourier transform and overlap-and-add (OLA) algorithm to transform the amplitude spectra of the pre-trained human voice and the pre-trained accompaniment from the frequency domain to the time domain, thereby obtaining the pre-trained human voice and the pre-trained accompaniment, and thus achieving the separation of the human voice and the accompaniment in the audio.
[0131] In some embodiments, the server can obtain the pre-trained vocals and pre-trained accompaniment with supervised sample audio in the time domain by the following formula (3).
[0132]
[0133] in, Represents pre-trained vocals or pre-trained accompaniment in the time domain; OLA() represents the overlap-add algorithm; iSTFT() represents the inverse short-time Fourier transform function; S i ∠X represents the amplitude spectrum of the pre-trained vocal or the pre-trained accompaniment; ∠X represents the phase spectrum of the supervised sample audio, including the phase spectrum of the pre-trained vocal and the phase spectrum of the pre-trained accompaniment.
[0134] In step S304, the server determines the pre-training loss based on the differences between the pre-trained human voice and the reference human voice in the supervised sample audio, as well as the differences between the pre-trained accompaniment and the reference accompaniment in the supervised sample audio.
[0135] In this embodiment, the reference vocals and accompaniment in the supervised sample audio are known supervisory information. The pre-trained vocals and accompaniment are obtained by a vocal-accompaniment separation network. Therefore, the server can determine the pre-training loss of the vocal-accompaniment separation network based on the difference between the actual vocals and accompaniment in the supervised sample audio and the vocals and accompaniment predicted by the vocal-accompaniment separation network, thereby pre-training the vocal-accompaniment separation network.
[0136] In some embodiments, the server can determine the signal distortion ratios between the pre-trained vocal and reference vocal, and between the pre-trained accompaniment and reference accompaniment, respectively, to obtain the pre-training loss of the vocal-accompaniment separation network. Accordingly, the server determines a first signal distortion ratio based on the difference between the pre-trained vocal and the reference vocal in the supervised sample audio. Then, the server determines a second signal distortion ratio based on the difference between the pre-trained accompaniment and the reference accompaniment in the supervised sample audio. The server then performs a weighted sum of the first and second signal distortion ratios to determine the pre-training loss. The weights corresponding to the first and second signal distortion ratios are not limited in this embodiment. By determining the error between the actual vocal and accompaniment and the predicted vocal and accompaniment during pre-training and the actual situation based on the signal distortion ratios between the actual vocal and accompaniment, the pre-training of the vocal-accompaniment separation network can be obtained, thereby enabling pre-training of the vocal-accompaniment separation network.
[0137] It should be noted that the embodiments of this disclosure determine the pre-training loss of the vocal accompaniment separation network through the scale-invariant signal-to-distortion ratio (SISDR) in the time domain. For example, the first signal-to-distortion ratio is the distortion ratio of the signal pair between the pre-trained vocal and the reference vocal in the supervised sample audio. The server can also determine the pre-training loss of the vocal accompaniment separation network through the average mean square error between the spectra of the pre-trained vocal and the reference vocal, and the average mean square error between the spectra of the pre-trained accompaniment and the reference accompaniment. The embodiments of this disclosure do not impose any limitations on this.
[0138] In some embodiments, the server can determine the signal distortion ratio of the pre-trained human voice and the signal distortion ratio of the pre-trained accompaniment using the following formula (4).
[0139]
[0140] Wherein, SISDR() represents the scale-invariant signal-to-distortion ratio function; 's' indicates a pre-trained vocal or pre-trained accompaniment; 's' indicates a reference vocal or reference accompaniment; '<>' indicates inner product calculation.
[0141] In some embodiments, the server can determine the pre-training loss of the vocal accompaniment separation network using the following formula (5).
[0142]
[0143] Where J represents the pre-training loss of the vocal accompaniment separation network; J1 represents the pre-training loss between the pre-trained vocal and the reference vocal; and J2 represents the pre-training loss between the pre-trained accompaniment and the reference accompaniment. s1 represents the pre-trained human voice; s1 represents the reference human voice; SISDR() represents the scale-invariant signal-to-distortion ratio function. s1 indicates pre-training accompaniment; s2 indicates reference accompaniment.
[0144] In step S305, the server pre-trains the vocal accompaniment separation network based on the pre-training loss.
[0145] In this embodiment, the server obtains the pre-training loss of the vocal accompaniment separation network using the above formula (5). Then, the server can pre-train the vocal accompaniment separation network based on the pre-training loss, so that the parameters of the vocal accompaniment separation network can meet the needs of separating the vocals and accompaniment in the audio, thereby preparing for training the vocal accompaniment separation model.
[0146] For example, Figure 4 This is a flowchart illustrating a supervised sample audio pre-trained vocal accompaniment separation network according to an exemplary embodiment. See also Figure 4 As shown, the server pre-trains the vocal-accompaniment separation network based on any supervised audio sample x. The vocal-accompaniment separation network separates the vocals and accompaniment from the input supervised audio sample x, obtaining the pre-trained vocal... and pre-training accompaniment Then, the server uses the reference human voice s1 and the pre-trained human voice as a basis. The difference between the reference vocals and the pre-trained vocals is used to determine the pre-training loss J1 between them. Simultaneously, the server uses the reference accompaniment s2 and the pre-trained accompaniment... The difference between the reference accompaniment and the pre-training accompaniment is used to determine the pre-training loss J2 between them. Then, the server performs a weighted sum of the two pre-training losses to obtain the pre-training loss J of the vocal accompaniment separation network, which can then be used to pre-train the vocal accompaniment separation network.
[0147] In step S306, for any sample audio from multiple sample audios, the server performs voice and accompaniment separation on the sample audio based on the voice and accompaniment separation network in the voice and accompaniment separation model to obtain the predicted voice and the predicted accompaniment. The multiple sample audios include supervised sample audios with supervised information and unsupervised sample audios without supervised information. The voice and accompaniment separation model is used to separate the voice and accompaniment in the input audio. The voice and accompaniment separation network is obtained by pre-training based on the above steps S301 to S305.
[0148] In this embodiment, the sample audio includes a limited number of supervised sample audios and a large number of unsupervised sample audios. The supervised sample audios only include the mixed sample audio obtained by mixing the reference vocals and reference accompaniment from the original sample audio in step 301, along with the supervision information of the mixed sample audios. The unsupervised sample audios only include the mixed sample audios and do not include the reference vocals and reference accompaniment. The server can separate the vocals and accompaniment from any sample audio based on the vocal-accompaniment separation model to obtain predicted vocals and predicted accompaniment, thereby training the vocal-accompaniment separation model. Since the sample audios include supervised sample audios with supervised information and unsupervised sample audios without supervised information, the vocal-accompaniment separation model trained by the server based on the sample audios improves the model's generalization ability compared to a model trained only on supervised sample audios.
[0149] In some embodiments, the server performs voice-accompaniment separation on any sample audio based on the voice-accompaniment separation model to obtain the predicted voice and predicted accompaniment. This process is consistent with steps 301 to 303 above and will not be repeated here.
[0150] In some embodiments, the server can also cluster the features of the sample audio using a vocal-accompaniment separation model to obtain predicted vocals and predicted accompaniment. Accordingly, the server extracts features from the sample audio based on the vocal-accompaniment separation model, obtaining multiple audio feature information. Then, the server clusters these multiple audio feature information based on the vocal-accompaniment separation model, obtaining multiple clustering results. The server then processes these clustering results based on the vocal-accompaniment separation model to obtain predicted vocals and predicted accompaniment. These clustering results are used to indicate the vocals and accompaniment in the sample audio. By clustering the features of the sample audio to obtain predicted vocals and predicted accompaniment, the separation of vocals and accompaniment in the audio is achieved.
[0151] In step S307, the server uses the voice discriminator in the voice accompaniment separation model to discriminate the predicted voice and obtain a first discrimination result. The first discrimination result is used to indicate whether it is the predicted voice.
[0152] In this embodiment, the voice discriminator is based on a deep neural network and is used to discriminate input audio. The training objective of the voice discriminator is: if the input audio is a human voice predicted by a voice accompaniment separation network, then the first discrimination result obtained by the voice discriminator is 0. If the input audio is a real human voice used as supervision information, then the first discrimination result obtained by the voice discriminator is 1. The voice discriminator typically outputs probabilities of discrimination as 0 and 1. By discriminating the input audio based on the voice discriminator, the voice discriminator is trained.
[0153] In step S308, the server uses the accompaniment discriminator in the vocal accompaniment separation model to discriminate the predicted accompaniment and obtain a second discrimination result. The second discrimination result is used to indicate whether it is the predicted accompaniment.
[0154] In this embodiment, the accompaniment discriminator is based on a deep neural network and is used to discriminate input audio. The goal of the accompaniment discriminator is: if the input audio is an accompaniment predicted by a vocal accompaniment separation network, then the second discrimination result obtained by the accompaniment discriminator is 0. If the input audio is the actual accompaniment used as supervision information, then the second discrimination result obtained by the vocal discriminator is 1. The accompaniment discriminator typically outputs probabilities of discrimination as 0 and 1. By discriminating the input audio based on the accompaniment discriminator, the accompaniment discriminator is trained.
[0155] It should be noted that the voice and accompaniment separation model in this embodiment of the present disclosure includes both a voice discriminator and an accompaniment discriminator, which distinguish between voices and accompaniment in the sample audio using a time-domain processing method. The server can also distinguish between voices and accompaniment using a frequency-domain processing method. This embodiment of the present disclosure does not impose any limitations on this.
[0156] In step S309, the server trains the human voice and accompaniment separation model based on the sample audio, predicted human voice, predicted accompaniment, first discrimination result, and second discrimination result.
[0157] In this embodiment, since the predicted vocals and accompaniment are obtained by the vocal-accompaniment separation network, and the first and second discrimination results are obtained by the vocal discriminator and the accompaniment discriminator, the server can obtain the training loss of the vocal-accompaniment separation model based on the difference between the prediction results obtained by the vocal-accompaniment separation model and the actual sample audio, thereby training the vocal-accompaniment separation model.
[0158] In some embodiments, the sample audio includes both supervised and unsupervised sample audio. The server obtains different training losses when training the vocal accompaniment separation model based on these two types of sample audio. Therefore, the server can train the vocal accompaniment separation model in the following two ways.
[0159] In Method 1, when the sample audio is supervised audio, the server can obtain a first generator loss and a first discriminator loss based on the predicted human voice and predicted accompaniment obtained from the human voice and accompaniment separation network, the discrimination results of the human voice discriminator and the accompaniment discriminator, and the supervision information of the supervised sample audio. The first generator loss indicates the loss of the human voice and accompaniment separation network. The first discriminator loss indicates the losses of the human voice discriminator and the accompaniment discriminator. Correspondingly, when the sample audio is supervised audio, the server determines the first generator loss based on the sample audio, the predicted human voice, the predicted accompaniment, the first discrimination result, and the second discrimination result. Then, the server determines the first discriminator loss based on the sample audio, the first discrimination result, the second discrimination result, the third discrimination result, and the fourth discrimination result. Then, the server trains the human voice and accompaniment separation model based on the first generator loss and the first discriminator loss. The third discrimination result is the discrimination result of the human voice discriminator on the reference human voice in the sample audio. The fourth discrimination result is the discrimination result of the accompaniment discriminator on the reference accompaniment in the sample audio. By determining the first generator loss and the first discriminator loss based on the supervision information of the sample audio, the prediction and discrimination results of the vocal accompaniment separation model, when the sample audio is supervised, the training of the vocal accompaniment separation model can be achieved.
[0160] In some embodiments, when the sample audio is supervised audio, the first generator loss corresponding to the vocal accompaniment separation network can be obtained by adding the first training loss and the second training loss. Accordingly, when the sample audio is supervised audio, the server determines the first training loss based on the reference vocal, the predicted vocal, and the first discrimination result in the sample audio. Then, the server determines the second training loss based on the reference accompaniment, the predicted accompaniment, and the second discrimination result in the sample audio. The server then determines the sum of the first and second training losses as the first generator loss. By determining the first generator loss of the vocal accompaniment separation network, the network can be trained based on this first generator loss.
[0161] In some embodiments, the server may determine the first training loss using the following formula (6).
[0162]
[0163] Where J1 represents the first training loss; SISDR() represents the scale-invariant signal-to-distortion ratio function; s1 represents the predicted human voice; s1 represents the reference human voice; x represents the supervised sample audio. D1 represents the minimum mean square loss; D1 represents the human voice discriminator.
[0164] The server can determine the second training loss using the following formula (7).
[0165]
[0166] Where J2 represents the second training loss; SISDDR() represents the scale-invariant signal-to-distortion ratio function; s1 represents the predicted accompaniment; s2 represents the reference accompaniment; x represents the supervised sample audio. D1 represents the minimum mean square loss; D2 represents the accompaniment discriminator.
[0167] The server can determine the first generator loss of the vocal accompaniment separation network using the following formula (8).
[0168] J G1 =J1+J2 (8)
[0169] Where G represents the vocal accompaniment separation network; J G1 J1 represents the first generator loss of the vocal accompaniment separation network when the sample audio is supervised sample audio; J2 represents the first training loss; J2 represents the second training loss.
[0170] In some embodiments, when the sample audio is supervised sample audio, the first discriminator loss corresponding to the voice discriminator and the accompaniment discriminator can be obtained by adding the third training loss and the fourth training loss. Accordingly, the server determines the third training loss based on the reference voice of the sample audio, the first discrimination result, and the third discrimination result. Then, the server determines the fourth training loss based on the reference accompaniment of the sample audio, the second discrimination result, and the fourth discrimination result. Then, the server determines the sum of the third training loss and the fourth training loss as the first discriminator loss. By determining the first discriminator loss of the voice discriminator and the accompaniment discriminator, the voice discriminator and the accompaniment discriminator can be trained based on the first discriminator loss.
[0171] In some embodiments, the server may determine the third training loss using the following formula (9).
[0172]
[0173] Where J3 represents the third training loss; x represents the supervised sample audio; and s1 represents the reference human voice. D1 represents the minimum mean square loss; D1 represents the human voice discriminator. This indicates a prediction of human voices.
[0174] The server can determine the fourth training loss using the following formula (10).
[0175]
[0176] Where J4 represents the fourth training loss; x represents the supervised sample audio; and s2 represents the reference accompaniment. D1 represents the minimum mean square loss; D2 represents the accompaniment discriminator. This indicates a predicted accompaniment.
[0177] The server can determine the first discriminator loss of the voice discriminator and the accompaniment discriminator using the following formula (11).
[0178] J D1 =J3+J4 (11)
[0179] Where D represents the vocal discriminator and the accompaniment discriminator; J D1 J1 represents the first discriminator loss of the human voice discriminator and the accompaniment discriminator when the sample audio is supervised sample audio; J2 represents the third training loss; J4 represents the fourth training loss.
[0180] For example, Figure 5 This is a flowchart illustrating a supervised training model for separating vocal accompaniment based on supervised sample audio, according to an exemplary embodiment. See also... Figure 5 As shown, when the sample audio is supervised audio, the server trains the vocal-accompaniment separation model based on this supervised audio sample. The vocal-accompaniment separation network separates the vocals and accompaniment from the input supervised audio sample x, obtaining the predicted vocals. And predict accompaniment Then, predict human voices The reference human voice s1 and the predicted human voice s1 are input into the human voice discriminator to obtain the first and third discrimination results, respectively. Then, based on the reference human voice s1 and the predicted human voice s1, the human voice is further processed. Based on the first discrimination result, determine the first training loss J1. Then, based on the reference voice s1 and the predicted voice... The first and third discrimination results determine the third training loss J3. Then, the predicted accompaniment... The reference accompaniment s2 and the accompaniment s2 are input into the accompaniment discriminator to obtain the second and fourth discrimination results, respectively. Then, based on the reference accompaniment s2, the predicted accompaniment is... Based on the second discrimination result, the second training loss J2 is determined. Then, based on the reference accompaniment s2 and the predicted accompaniment... The second and fourth discrimination results determine the fourth training loss J4. Then, the server can determine the first generator loss of the vocal accompaniment separation network through the first and second training losses, and determine the first discriminator loss of the vocal discriminator and the accompaniment discriminator through the third and fourth training losses, thereby enabling the vocal accompaniment separation model to be trained based on the first generator loss and the first discriminator loss.
[0181] In the second method, when the sample audio is unsupervised, the server can derive a second generator loss and a second discriminator loss based on the predicted human voice and accompaniment obtained from the human voice and accompaniment separation network, the discrimination results of the human voice discriminator and the accompaniment discriminator, and the unsupervised sample audio. The second generator loss indicates the loss of the human voice and accompaniment separation network. The second discriminator loss indicates the losses of the human voice discriminator and the accompaniment discriminator. Accordingly, when the sample audio is unsupervised, the server determines the second discriminator loss based on the sample audio, the first discrimination result, and the second discrimination result. Then, the server determines the second generator loss based on the sample audio, the predicted human voice, the predicted accompaniment, the first discrimination result, and the second discrimination result. Finally, the server trains the human voice and accompaniment separation model based on the second generator loss and the second discriminator loss. By determining the second generator loss and the second discriminator loss based on the sample audio, the prediction and discrimination results of the vocal accompaniment separation model when the sample audio is unsupervised, the training of the vocal accompaniment separation model can be achieved.
[0182] In some embodiments, when the sample audio is unsupervised, the server can determine the fifth training loss and the sixth training loss respectively using the first discrimination result of the voice discriminator and the second discrimination result of the accompaniment discriminator, thereby obtaining the second discriminator loss. Correspondingly, when the sample audio is unsupervised, the server determines the fifth training loss based on the sample audio and the first discrimination result. Then, the server determines the sixth training loss based on the sample audio and the second discrimination result. The server then determines the sum of the fifth and sixth training losses as the second discriminator loss. By determining the second discriminator loss of the voice-accompaniment separation network, the voice discriminator and the accompaniment discriminator can be trained based on this second discriminator loss.
[0183] In some embodiments, the server can determine the second generator loss of the vocal accompaniment separation network using the following formula (12).
[0184]
[0185] Where G represents the vocal accompaniment separation network; J G2 This represents the second generator loss of the vocal accompaniment separation network when the sample audio is unsupervised; SISDR() represents the scale-invariant signal-to-distortion ratio function. Indicates a prediction of human voice; y represents the predicted accompaniment; y represents the unsupervised sample audio. D1 represents the minimum mean square loss; D2 represents the vocal discriminator; D3 represents the accompaniment discriminator.
[0186] The server can determine the fifth training loss using the following formula (13).
[0187]
[0188] Where J5 represents the fifth training loss; y represents the unsupervised sample audio; D1 represents the minimum mean square loss; D1 represents the human voice discriminator. This indicates a prediction of human voices.
[0189] The server can determine the sixth training loss using the following formula (14).
[0190]
[0191] Where J6 represents the sixth training loss; y represents the unsupervised sample audio; D1 represents the minimum mean square loss; D2 represents the accompaniment discriminator. This indicates a predicted accompaniment.
[0192] The server can determine the second discriminator loss of the vocal discriminator and the accompaniment discriminator using the following formula (15).
[0193] J D2 =J5+J6 (15)
[0194] Where D represents the vocal discriminator and the accompaniment discriminator; J D2 J5 represents the second discriminator loss of the human voice discriminator and the accompaniment discriminator when the sample audio is unsupervised; J6 represents the fifth training loss; J5 represents the sixth training loss.
[0195] For example, Figure 6 This is a flowchart illustrating an unsupervised sample audio training model for separating vocal accompaniment, according to an exemplary embodiment. See also... Figure 6 As shown, when the sample audio is unsupervised, the server trains the vocal-accompaniment separation model based on this unsupervised sample audio. The vocal-accompaniment separation network separates the vocals and accompaniment from the input unsupervised sample audio y, obtaining the predicted vocals. And predict accompaniment Then, predict human voices The input is fed into the voice discriminator to obtain the first discrimination result. Then, based on the predicted human voice... Based on the first discrimination result, the fifth training loss J5 is determined. Then, the predicted accompaniment is... The input is fed into the accompaniment discriminator to obtain the second discrimination result. Then, based on the predicted accompaniment... Based on the second discrimination result, the sixth training loss J6 is determined. Then, based on the unsupervised sample audio y and the predicted human voice... And predict accompaniment The second generator loss is determined. Then, the server can determine the second discriminator loss for the vocal discriminator and the accompaniment discriminator using the fifth training loss and the sixth training loss, thereby enabling the vocal-accompaniment separation model to be trained based on the second generator loss and the second discriminator loss.
[0196] It should be noted that the supervised sample audio used for pre-training the vocal accompaniment separation network and the sample audio used for training the vocal accompaniment separation model in this embodiment are both single-channel speech. The server can also implement the training method of the vocal accompaniment separation model proposed in this disclosure based on multi-channel speech or a mixture of single-channel and multi-channel speech, and this embodiment does not limit this.
[0197] For example, Figure 7 This is a schematic diagram illustrating the application of a vocal accompaniment separation model according to an exemplary embodiment. See also... Figure 7 As shown, taking a song as an example, the server inputs any song into the vocal-accompaniment separation model to obtain the vocals and accompaniment of the song. The vocal-accompaniment separation model separates the vocals and accompaniment of a song, allowing the singer to perform with a clean accompaniment without vocals, improving the singer's performance experience and enhancing the accuracy of the vocals and accompaniment.
[0198] This disclosure provides a training method for a vocal-accompaniment separation model. By using a vocal-accompaniment separation network within the model, predicted vocals and accompaniment are separated from any sample audio, achieving separation of vocals and accompaniment in the sample audio. Then, based on the vocal discriminator and accompaniment discriminator in the model, the predicted vocals and accompaniment are separately discriminated, yielding a first discrimination result and a second discrimination result. The training loss is then determined using the sample audio, the predicted vocals and accompaniment, and the discriminator's discrimination results, thereby updating the parameters in the vocal-accompaniment separation model and training it. Since the sample audio includes a small number of supervised sample audios and a large number of unsupervised sample audios, the high cost of acquiring supervised sample audios is avoided, improving the generalization ability of the vocal-accompaniment separation model and increasing the accuracy of the separated vocals and accompaniment by the trained model.
[0199] Figure 8 This is a block diagram illustrating a training apparatus for a vocal accompaniment separation model according to an exemplary embodiment. (Refer to...) Figure 8 As shown, the device includes a separation unit 801, a first discrimination unit 802, a second discrimination unit 803, and a training unit 804.
[0200] Separation unit 801 is configured to separate the human voice and accompaniment in any sample audio from multiple sample audios based on the human voice and accompaniment separation network in the human voice and accompaniment separation model, to obtain the predicted human voice and the predicted accompaniment. The multiple sample audios include supervised sample audios with supervised information and unsupervised sample audios without supervised information. The human voice and accompaniment separation model is used to separate the human voice and accompaniment in the input audio. The human voice and accompaniment separation network is pre-trained based on supervised sample audios from multiple sample audios.
[0201] The first discrimination unit 802 is configured to discriminate the predicted human voice based on the human voice accompaniment separation model and obtain a first discrimination result. The first discrimination result is used to indicate whether it is the predicted human voice.
[0202] The second discrimination unit 803 is configured to discriminate the predicted accompaniment based on the accompaniment separation model and obtain a second discrimination result. The second discrimination result is used to indicate whether it is the predicted accompaniment.
[0203] Training unit 804 is configured to train the vocal-accompaniment separation model based on sample audio, predicted vocals, predicted accompaniment, first discrimination result, and second discrimination result.
[0204] In some embodiments, Figure 9 This is a block diagram illustrating a training apparatus for another vocal accompaniment separation model according to an exemplary embodiment. (Refer to...) Figure 9 As shown, training unit 804 includes:
[0205] The first determining subunit 901 is configured to determine a first generator loss based on the sample audio, predicted human voice, predicted accompaniment, first discrimination result, and second discrimination result when the sample audio is supervised sample audio. The first generator loss is used to indicate the loss of the human voice and accompaniment separation network.
[0206] The second determining subunit 902 is configured to determine a first discriminator loss based on the sample audio, a first discrimination result, a second discrimination result, a third discrimination result, and a fourth discrimination result. The third discrimination result is the discrimination result of the voice discriminator on the reference voice in the sample audio, and the fourth discrimination result is the discrimination result of the accompaniment discriminator on the reference accompaniment in the sample audio. The first discriminator loss is used to indicate the loss of the voice discriminator and the accompaniment discriminator.
[0207] The first training subunit 903 is configured to train the vocal accompaniment separation model based on the first generator loss and the first discriminator loss.
[0208] In some embodiments, the first determining subunit 901 is configured to, when the sample audio is supervised sample audio, determine a first training loss based on a reference human voice, a predicted human voice, and a first discrimination result in the sample audio; determine a second training loss based on a reference accompaniment, a predicted accompaniment, and a second discrimination result in the sample audio; and determine the sum of the first training loss and the second training loss as the first generator loss.
[0209] In some embodiments, the second determining subunit 902 is configured to determine a third training loss based on a reference human voice in the sample audio, a first discrimination result, and a third discrimination result; determine a fourth training loss based on a reference accompaniment in the sample audio, a second discrimination result, and a fourth discrimination result; and determine the sum of the third training loss and the fourth training loss as the first discriminator loss.
[0210] In some embodiments, the training unit 804 includes:
[0211] The third determining subunit 904 is configured to determine the second discriminator loss based on the sample audio, the first discrimination result, and the second discrimination result when the sample audio is unsupervised sample audio. The second discriminator loss is used to indicate the loss of the voice discriminator and the accompaniment discriminator.
[0212] The fourth determining subunit 905 is configured to determine the second generator loss based on the sample audio, predicted human voice, predicted accompaniment, first discrimination result, and second discrimination result. The second generator loss is used to indicate the loss of the human voice and accompaniment separation network.
[0213] The second training subunit 906 is configured to train the vocal accompaniment separation model based on the second generator loss and the second discriminator loss.
[0214] In some embodiments, the third determining subunit 904 is configured to, when the sample audio is unsupervised sample audio, determine a fifth training loss based on the sample audio and a first discrimination result; determine a sixth training loss based on the sample audio and a second discrimination result; and determine the sum of the fifth training loss and the sixth training loss as the second discriminator loss.
[0215] In some embodiments, the separation unit 801 is configured to extract features from the sample audio based on the vocal-accompaniment separation model to obtain multiple audio feature information; cluster the multiple audio feature information based on the vocal-accompaniment separation model to obtain multiple clustering results, the clustering results being used to indicate the vocals and accompaniment in the sample audio; and process the multiple clustering results based on the vocal-accompaniment separation model to obtain predicted vocals and predicted accompaniment.
[0216] In some embodiments, the apparatus further includes:
[0217] Transform unit 805 is configured to perform a short-time Fourier transform on any supervised sample audio to obtain the amplitude spectrum and phase spectrum of the supervised sample audio in the frequency domain.
[0218] The processing unit 806 is configured to process the amplitude spectrum of supervised sample audio based on the human voice accompaniment separation network to obtain the amplitude spectrum of pre-trained human voice and the amplitude spectrum of pre-trained accompaniment.
[0219] The acquisition unit 807 is configured to obtain the pre-trained human voice and the pre-trained accompaniment based on the phase spectrum of supervised sample audio, the amplitude spectrum of the pre-trained human voice, and the amplitude spectrum of the pre-trained accompaniment.
[0220] The determination unit 808 is configured to determine the pre-training loss based on the difference between the pre-trained human voice and the reference human voice of the supervised sample audio, and the difference between the pre-trained accompaniment and the reference accompaniment of the supervised sample audio.
[0221] Pre-training unit 809 is configured to pre-train the vocal accompaniment separation network based on pre-training loss.
[0222] In some embodiments, determining unit 808 is configured to determine a first signal distortion ratio based on the difference between a pre-trained human voice and a reference human voice in supervised sample audio; determine a second signal distortion ratio based on the difference between a pre-trained accompaniment and a reference accompaniment in supervised sample audio; and determine a pre-training loss by weighted summation of the first signal distortion ratio and the second signal distortion ratio.
[0223] This disclosure provides a training device for a vocal-accompaniment separation model. By using a vocal-accompaniment separation network within the model, predicted vocals and accompaniment are separated from any sample audio, achieving separation of vocals and accompaniment in the sample audio. Then, based on the vocal and accompaniment discriminators within the model, the predicted vocals and accompaniment are separately discriminated, yielding a first discrimination result and a second discrimination result. The training loss can then be determined using the sample audio, the predicted vocals and accompaniment, and the discriminator's discrimination results, thereby updating the parameters in the vocal-accompaniment separation model and training it. Since the sample audio includes a small number of supervised sample audios and a large number of unsupervised sample audios, the high cost of acquiring supervised sample audios is avoided, improving the generalization ability of the vocal-accompaniment separation model and increasing the accuracy of the separated vocals and accompaniment by the trained model.
[0224] Figure 10This is a block diagram illustrating an electronic device 1000 according to an exemplary embodiment. Typically, the electronic device 1000 includes a processor 1001 and a memory 1002.
[0225] Processor 1001 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1001 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1001 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1001 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0226] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 are used to store at least one program code, which is executed by the processor 1001 to implement the training method for the vocal accompaniment separation model provided in the method embodiments of this disclosure.
[0227] In some embodiments, the electronic device 1000 may optionally include a peripheral device interface 1003 and at least one peripheral device. The processor 1001, memory 1002, and peripheral device interface 1003 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1003 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 1004, a display screen 1005, a camera assembly 1006, an audio circuit 1007, and a power supply 1008.
[0228] Peripheral device interface 1003 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1001 and memory 1002. In some embodiments, processor 1001, memory 1002 and peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1001, memory 1002 and peripheral device interface 1003 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0229] The radio frequency (RF) circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1004 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1004 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1004 can communicate with other electronic devices via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1004 may also include circuitry related to NFC (Near Field Communication), which is not limited in this disclosure.
[0230] Display screen 1005 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1005 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1001 for processing. In this case, display screen 1005 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1005, which serves as the front panel of electronic device 1000; in other embodiments, there may be at least two display screens, respectively disposed on different surfaces of electronic device 1000 or in a folded design; in some embodiments, display screen 1005 may be a flexible display screen, disposed on a curved or folded surface of electronic device 1000. Furthermore, display screen 1005 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 1005 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).
[0231] The camera assembly 1006 is used to acquire images or videos. Optionally, the camera assembly 1006 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the electronic device, and the rear-facing camera is located on the back of the electronic device. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1006 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0232] The audio circuit 1007 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1001 for processing, or input to the radio frequency circuit 1004 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the electronic device 1000. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1001 or the radio frequency circuit 1004 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1007 may also include a headphone jack.
[0233] The power supply 1008 is used to supply power to the various components in the electronic device 1000. The power supply 1008 can be AC power, DC power, a disposable battery, or a rechargeable battery. When the power supply 1008 includes a rechargeable battery, the rechargeable battery can support wired or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0234] Those skilled in the art will understand that Figure 10 The structure shown does not constitute a limitation on the electronic device 1000, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0235] When a computer device is configured as a server, Figure 11 This is a schematic diagram of a server structure according to an exemplary embodiment. The server 1100 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 1101 and one or more memories 1102. The memories 1102 store at least one computer program, which is loaded and executed by the processor 1101 to implement the training method for the vocal accompaniment separation model provided in the various method embodiments described above. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated here.
[0236] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 1002 including instructions, which can be executed by a processor 1001 of an electronic device 1000 to complete the training method for the vocal accompaniment separation model described above. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.
[0237] In an exemplary embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the training method for the above-described vocal accompaniment separation model.
[0238] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0239] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A training method for a vocal accompaniment separation model, characterized in that, The method comprises: For any sample audio in a plurality of sample audios, performing vocal accompaniment separation on the sample audio based on a vocal accompaniment separation network in a vocal accompaniment separation model to obtain predicted vocals and predicted accompaniment, the plurality of sample audios comprising supervised sample audios with supervision information and unsupervised sample audios without supervision information, the vocal accompaniment separation model being used for separating vocals and accompaniment in input audio, and the vocal accompaniment separation network being pre-trained based on the supervised sample audios in the plurality of sample audios; Discriminating the predicted vocals based on a vocal discriminator in the vocal accompaniment separation model to obtain a first discrimination result, the first discrimination result being used to indicate whether the predicted vocals are correct; Discriminating the predicted accompaniment based on an accompaniment discriminator in the vocal accompaniment separation model to obtain a second discrimination result, the second discrimination result being used to indicate whether the predicted accompaniment is correct; In a case where the sample audio is a supervised sample audio, determining a first generator loss based on the sample audio, the predicted vocals, the predicted accompaniment, the first discrimination result, and the second discrimination result, the first generator loss being used to indicate a loss of the vocal accompaniment separation network; Determining a first discriminator loss based on the sample audio, the first discrimination result, the second discrimination result, a third discrimination result, and a fourth discrimination result, the third discrimination result being a discrimination result of the vocal discriminator on reference vocals in the sample audio, the fourth discrimination result being a discrimination result of the accompaniment discriminator on reference accompaniment in the sample audio, and the first discriminator loss being used to indicate a loss of the vocal discriminator and the accompaniment discriminator; Training the vocal accompaniment separation model based on the first generator loss and the first discriminator loss; In a case where the sample audio is an unsupervised sample audio, determining a second discriminator loss based on the sample audio, the first discrimination result, and the second discrimination result, the second discriminator loss being used to indicate a loss of the vocal discriminator and the accompaniment discriminator; Determining a second generator loss based on the sample audio, the predicted vocals, the predicted accompaniment, the first discrimination result, and the second discrimination result, the second generator loss being used to indicate a loss of the vocal accompaniment separation network; Training the vocal accompaniment separation model based on the second generator loss and the second discriminator loss.
2. The method of claim 1, wherein, The determining, in a case where the sample audio is a supervised sample audio, of a first generator loss based on the sample audio, the predicted vocals, the predicted accompaniment, the first discrimination result, and the second discrimination result comprises: In a case where the sample audio is a supervised sample audio, determining a first training loss based on reference vocals in the sample audio, the predicted vocals, and the first discrimination result; Determining a second training loss based on reference accompaniment in the sample audio, the predicted accompaniment, and the second discrimination result; A sum of the first training loss and the second training loss is determined as the first generator loss. 3.The method of training a vocal accompaniment separation model of claim 1, wherein, The first discriminator loss is determined based on the sample audio, the first discrimination result, the second discrimination result, a third discrimination result, and a fourth discrimination result. A third training loss is determined based on a reference vocal of the sample audio, the first discrimination result, and the third discrimination result. A fourth training loss is determined based on a reference accompaniment of the sample audio, the second discrimination result, and the fourth discrimination result. A sum of the third training loss and the fourth training loss is determined as the first discriminator loss.
4. The method of claim 1, wherein, The second discriminator loss is determined based on the sample audio, the first discrimination result, and the second discrimination result in a case where the sample audio is an unsupervised sample audio. A fifth training loss is determined based on the sample audio and the first discrimination result in a case where the sample audio is an unsupervised sample audio. A sixth training loss is determined based on the sample audio and the second discrimination result. A sum of the fifth training loss and the sixth training loss is determined as the second discriminator loss.
5. The method of claim 1, wherein, The vocal-accompaniment separation of the sample audio by the vocal-accompaniment separation network in the vocal-accompaniment separation model includes: Feature extraction of the sample audio by the vocal-accompaniment separation model to obtain a plurality of audio feature information. Clustering of the plurality of audio feature information by the vocal-accompaniment separation model to obtain a plurality of clustering results, the clustering results being used to indicate vocals and accompaniments in the sample audio. Processing of the plurality of clustering results by the vocal-accompaniment separation model to obtain the predicted vocals and the predicted accompaniments.
6. The method of claim 1, wherein, The method further includes: Short-time Fourier transform of the supervised sample audio to obtain an amplitude spectrum and a phase spectrum of the supervised sample audio in a frequency domain. Processing of the amplitude spectrum of the supervised sample audio by the vocal-accompaniment separation network to obtain an amplitude spectrum of a pre-trained vocal and an amplitude spectrum of a pre-trained accompaniment. Obtaining of the pre-trained vocal and the pre-trained accompaniment based on the phase spectrum of the supervised sample audio, the amplitude spectrum of the pre-trained vocal, and the amplitude spectrum of the pre-trained accompaniment. Determination of a pre-training loss based on a difference between the pre-trained vocal and a reference vocal of the supervised sample audio and a difference between the pre-trained accompaniment and a reference accompaniment of the supervised sample audio. Pre-training of the vocal-accompaniment separation network based on the pre-training loss.
7. The method of claim 6, wherein, The pre-training loss is determined based on a difference between the pre-trained vocal and a reference vocal of the supervised sample audio and a difference between the pre-trained accompaniment and a reference accompaniment of the supervised sample audio, including: Determination of a first signal distortion ratio based on the difference between the pre-trained vocal and the reference vocal of the supervised sample audio. determine a second signal distortion ratio based on a difference between the pre-training accompaniment and a reference accompaniment of the supervised sample audio; determine the pre-training loss by weighted sum of the first signal distortion ratio and the second signal distortion ratio.
8. A training device for a vocal accompaniment separation model, characterized in that, Comprising: a separation unit configured to, for any sample audio in a plurality of sample audios, perform vocal accompaniment separation on the sample audio based on a vocal accompaniment separation network in a vocal accompaniment separation model to obtain predicted vocals and predicted accompaniment, the plurality of sample audios including supervised sample audios with supervision information and unsupervised sample audios without supervision information, the vocal accompaniment separation model being used for separating vocals and accompaniment in an input audio, the vocal accompaniment separation network being pre-trained based on the supervised sample audios in the plurality of sample audios; a first discrimination unit configured to, based on a vocal discriminator in the vocal accompaniment separation model, discriminate the predicted vocals to obtain a first discrimination result, the first discrimination result being used to indicate whether it is predicted vocals; a second discrimination unit configured to, based on an accompaniment discriminator in the vocal accompaniment separation model, discriminate the predicted accompaniment to obtain a second discrimination result, the second discrimination result being used to indicate whether it is predicted accompaniment; a training unit configured to, in a case that the sample audio is a supervised sample audio, determine a first generator loss based on the sample audio, the predicted vocals, the predicted accompaniment, the first discrimination result and the second discrimination result, the first generator loss being used to indicate a loss of the vocal accompaniment separation network; determine a first discriminator loss based on the sample audio, the first discrimination result, the second discrimination result, a third discrimination result and a fourth discrimination result, the third discrimination result being a discrimination result of the vocal discriminator on reference vocals in the sample audio, the fourth discrimination result being a discrimination result of the accompaniment discriminator on reference accompaniment in the sample audio, the first discriminator loss being used to indicate a loss of the vocal discriminator and the accompaniment discriminator; train the vocal accompaniment separation model based on the first generator loss and the first discriminator loss; in a case that the sample audio is an unsupervised sample audio, determine a second discriminator loss based on the sample audio, the first discrimination result and the second discrimination result, the second discriminator loss being used to indicate a loss of the vocal discriminator and the accompaniment discriminator; determine a second generator loss based on the sample audio, the predicted vocals, the predicted accompaniment, the first discrimination result and the second discrimination result, the second generator loss being used to indicate a loss of the vocal accompaniment separation network; train the vocal accompaniment separation model based on the second generator loss and the second discriminator loss.
9. An electronic device, comprising: Comprising: one or more processors; a memory for storing program codes executable by the processors; wherein the processors are configured to execute the program codes to implement the training method of the vocal accompaniment separation model according to any one of claims 1 to 7.
10. A computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enable the electronic device to perform the method of training a vocal accompaniment separation model according to any one of claims 1 to 7.
Citation Information
Patent Citations
Voice separation model training method and device, and voice separation method and device
CN111179962A
Main melody extraction model training method and component and singing detection method and component
CN115114993A