Training of Speech Conversion Model and Speech Conversion Method, Device and Related Equipment

By adopting cyclic adversarial generation network and gradient inversion technology in the speech conversion model, the problem of instability and poor training of speech conversion models in the existing technology is solved, and a more accurate and stable speech conversion effect is achieved.

CN114882897BActive Publication Date: 2025-06-20PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210517643.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-13
Publication Date
2025-06-20
Estimated Expiration
2042-05-13

AI Technical Summary

Technical Problem

In the prior art, the speech conversion model training is unstable, the effect is poor, and it is difficult to fully learn the speech distribution.

Method used

The speech conversion model training method based on the loop adversarial generation network is adopted, and the voice data is trained cyclically through the loop adversarial network, and the generator generation ability is optimized using gradient inversion technology.

Benefits of technology

A more robust voice conversion model is realized, improving the accuracy and stability of voice conversion, and ensuring the maintenance of voice content during the conversion process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114882897B_ABST
    Figure CN114882897B_ABST
Patent Text Reader

Abstract

This application relates to artificial intelligence technology, and provides a method and device for training a voice conversion model, as well as related equipment. The method includes: inputting first timbre domain voice data and second timbre domain voice data into a cyclic generative adversarial network to train the cyclic generative adversarial network, and obtaining the discrimination result of the discriminator in the cyclic generative adversarial network; if it is necessary to optimize the generator in the cyclic generative adversarial network and it is determined to perform gradient reversal according to the discrimination result, then reverse the gradient calculated according to the loss function corresponding to the generator, and update the model parameters of the generator according to the reversed gradient; if it is determined not to perform gradient reversal according to the discrimination result, then calculate the gradient according to the loss function corresponding to the generator, and update the model parameters of the generator according to the gradient, and iterate the training until the model converges. Through this application, a more robust voice conversion model is obtained, and at the same time, the accuracy of voice conversion is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular, to a method and device for training a voice conversion model, a voice conversion method, and related equipment. Background Art

[0002] Voice conversion (VC) is to convert the voice of a source speaker into the voice timbre of a target speaker by timbre extraction and content decoupling while keeping the speech content the same. Its application scenarios include dubbing in film and television dramas and voice timbre conversion in e-book reading to automatically match different story characters.

[0003] The current voice conversion methods mainly include voice conversion based on a Generative Adversarial Network (GAN) and voice conversion based on conditional VAE. The voice conversion method based on GAN can synthesize target voices with relatively high similarity, but the model training based on GAN is unstable. The method based on conditional VAE is relatively simple to train, but it is difficult to comprehensively learn the latent variables with the same distribution as the target voice. Summary of the Invention

[0004] In order to solve the technical problems in the prior art that the model training of voice conversion is unstable, the effect is not good, or it is difficult to comprehensively learn the voice distribution. The present application provides a method and device for training a voice conversion model, a voice conversion method, and related equipment. Its main purpose is to obtain a more robust voice conversion model and enhance the accuracy of voice conversion at the same time.

[0005] To achieve the above object, the present application provides a method for training a voice conversion model based on a cyclic adversarial network. The method for training the voice conversion model includes:

[0006] Obtaining a first voice dataset in a first timbre domain and a second voice dataset in a second timbre domain as training samples;

[0007] Inputting the first voice data in the first timbre domain selected from the first voice dataset in the first timbre domain and the second voice data in the second timbre domain selected from the second voice dataset in the second timbre domain into the constructed cyclic generative adversarial network for cyclic adversarial training of the cyclic generative adversarial network, and obtaining the discrimination result of the discriminator in the cyclic generative adversarial network;

[0008] If it is necessary to optimize the generator in the cyclic generative adversarial network, determine whether to perform gradient reversal according to the discrimination result;

[0009] If it is determined to perform gradient reversal according to the discrimination result, calculate the first loss function corresponding to the generator, and reverse the first gradient calculated according to the first loss function.

[0010] Update the model parameters of the generator according to the reversed first gradient;

[0011] If it is determined not to perform gradient reversal according to the discrimination result, calculate the first loss function corresponding to the generator, calculate the first gradient according to the first loss function, and update the model parameters of the generator according to the first gradient;

[0012] If the cyclic generative adversarial network has not converged, re - execute the above steps of inputting the first - timbre - domain speech data selected from the first - timbre - domain speech dataset and the second - timbre - domain speech data selected from the second - timbre - domain speech dataset into the constructed cyclic generative adversarial network for cyclic adversarial training of the cyclic generative adversarial network until the cyclic generative adversarial network converges.

[0013] To achieve the above - mentioned purpose, the present application also provides a voice conversion method based on a cyclic adversarial generative network. The voice conversion method includes:

[0014] Use the trained voice conversion model based on the cyclic adversarial generative network to perform timbre conversion on the input first - timbre - domain speech to be converted, and obtain the corresponding target second - timbre - domain speech data, where the trained voice conversion model based on the cyclic adversarial generative network is obtained according to the training method of the voice conversion model based on the cyclic adversarial generative network in any of the previous items.

[0015] To achieve the above - mentioned purpose, the present application also provides a training device for a voice conversion model based on a cyclic adversarial generative network. The training device for the voice conversion model includes:

[0016] A sample acquisition module, configured to acquire a first - timbre - domain speech dataset and a second - timbre - domain speech dataset as training samples;

[0017] A training module, configured to input the first - timbre - domain speech data selected from the first - timbre - domain speech dataset and the second - timbre - domain speech data selected from the second - timbre - domain speech dataset into the constructed cyclic generative adversarial network to perform cyclic adversarial training on the cyclic generative adversarial network, and obtain the discrimination result of the discriminator in the cyclic generative adversarial network;

[0018] A judgment module, configured to determine whether to perform gradient reversal according to the discrimination result if it is necessary to optimize the generator in the cyclic generative adversarial network;

[0019] A reversal module, configured to calculate the first loss function corresponding to the generator and reverse the first gradient calculated according to the first loss function if it is determined to perform gradient reversal according to the discrimination result,

[0020] A first parameter update module, configured to update the model parameters of the generator according to the reversed first gradient;

[0021] A second parameter update module, configured to calculate a first loss function corresponding to a generator if it is determined according to the discrimination result that gradient reversal is not performed, calculate a first gradient according to the first loss function, and update the model parameters of the generator according to the first gradient;

[0022] An iteration module, configured to jump to the training module if the cyclic generative adversarial network has not converged until the cyclic generative adversarial network converges.

[0023] To achieve the above object, the present application further provides a computer device, including a memory, a processor, and computer-readable instructions stored on the memory and executable on the processor. When the processor executes the computer-readable instructions, it executes the steps of the training method of the voice conversion model based on the cyclic generative adversarial network as described in any one of the preceding items, or when the processor executes the computer-readable instructions, it executes the steps of the voice conversion method based on the cyclic generative adversarial network as described in any one of the preceding items.

[0024] To achieve the above object, the present application further provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the processor is caused to execute the steps of the training method of the voice conversion model based on the cyclic generative adversarial network as described in any one of the preceding items, or the processor is caused to execute the steps of the voice conversion method based on the cyclic generative adversarial network as described in any one of the preceding items.

[0025] The training of the voice conversion model, the voice conversion method, device and related equipment proposed by the present application utilize the advantage of the cyclic generative adversarial network in fitting the real data distribution, establish a voice conversion model based on the cyclic generative adversarial network, support unsupervised training, and can use a large amount of unsupervised data to perform the conversion of the source voice to the target voice timbre, without the need for a large amount of paired first timbre domain voice data and second timbre domain voice data; it realizes the conversion of the voice of the source speaker to the voice timbre of the target speaker through timbre extraction and content decoupling while keeping the voice content the same. In addition, the present application uses the gradient reversal technology for the cyclic generative adversarial network to optimize the generation ability of the generator, enhances the discrimination ability between the target timbre and the source timbre, thereby reversely improving the generation ability of the generator for the target timbre, and enhancing the adversarial training between the generator and the discriminator. The present application utilizes the cycle consistency of the cyclic generative adversarial network to make the reconstruction of voice conversion more stable, the model training more accurate, and the obtained voice conversion model more robust, thereby improving the accuracy of voice conversion. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a schematic flowchart of the training method of the voice conversion model based on the cyclic generative adversarial network in an embodiment of the present application;

[0027] Figure 2It is a schematic structural diagram of a cyclic generative adversarial network in an embodiment of the present application;

[0028] Figure 3 It is a structural block diagram of a training device for a voice conversion model based on a cyclic adversarial generative network in an embodiment of the present application;

[0029] Figure 4 It is an internal structural block diagram of a computer device in an embodiment of the present application.

[0030] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Specific Embodiments

[0031] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application. It should be understood that the specific embodiments described herein are only for explaining the present application and are not used to limit the present application.

[0032] Figure 1 It is a schematic flowchart of a training method for a voice conversion model based on a cyclic adversarial generative network in an embodiment of the present application. Refer to Figure 1 The training method for the voice conversion model based on the cyclic adversarial generative network includes the following steps S100 - S700.

[0033] S100: Obtain a first timbre domain voice data set and a second timbre domain voice data set as training samples.

[0034] Specifically, the first timbre domain voice data set includes multiple first timbre domain voice data, and the second timbre domain voice data set includes multiple second timbre domain voice data. The voice data is a kind of spectrum, specifically, it can be the spectrum obtained by performing short-time Fourier transform on audio data. This embodiment does not require a large number of paired first timbre domain voice data and second timbre domain voice data. The first timbre domain voice data selected from the first timbre domain voice data set and the second timbre domain voice data selected from the second timbre domain voice data set form a group of training samples.

[0035] S200: Input the first timbre domain voice data selected from the first timbre domain voice data set and the second timbre domain voice data selected from the second timbre domain voice data set into the constructed cyclic generative adversarial network to perform cyclic adversarial training on the cyclic generative adversarial network, and obtain the discrimination result of the discriminator in the cyclic generative adversarial network.

[0036] Specifically, a Cycle Generative Adversarial Network (Cycle GAN for short) includes a generator (generative model or generator) and a discriminator (discriminator or discriminative model or discriminator). The purpose of this application is to achieve the conversion from the source tone domain to the target tone domain. Therefore, the Cycle GAN of this application learns the tone distribution of the speech data in the dataset through adversarial training to complete the tone transfer from the source tone domain to the target tone domain. The Cycle GAN should not only fit the tone distribution of the speech data in the target tone domain but also maintain the content features of the speech data in the source domain.

[0037] The generator is used to capture or fit the distribution of the training data, generate a similar data distribution similar to the real training data, and the pursuit is that the more similar to the real training data, the better.

[0038] After the speech data is input into the Cycle GAN, it will be converted into a hidden layer vector representation through feature extraction to represent the original speech data.

[0039] In this embodiment, the generator of the Cycle GAN is specifically used to fit the distribution of the speech data, generate generated speech data imitating the speech data in the second tone domain according to the speech data in the first tone domain, or generate generated speech data imitating the speech data in the first tone domain according to the speech data in the second tone domain.

[0040] The discriminator is used to learn the mapping relationship between the input speech data and the output class label, that is, estimate the probability that the generated speech data generated by the generator comes from the training data, judge whether the generated speech data is real or fake, and feedback the discrimination result to the generator. If it is estimated that the generated speech data comes from the training data, the probability output by the discriminator is large; otherwise, the probability output by the discriminator is small.

[0041] The purpose of the generator is to deceive the discriminator, while the purpose of the discriminator is not to be deceived by the generator. The generator learns the mapping relationship through adversarial training with the discriminator. The two networks are alternately trained until the data generated by the generator can pass off as real and reach a certain balance with the discrimination ability of the discriminator.

[0042] The discriminator in this embodiment is specifically used to discriminate whether the generated speech data imitating the speech data in the second tone domain is the speech data in the second tone domain or conforms to the distribution of the speech data in the second tone domain, and output the corresponding probability; or discriminate whether the generated speech data imitating the speech data in the first tone domain is the speech data in the first tone domain or conforms to the distribution of the speech data in the first tone domain, and output the corresponding probability. The discrimination result is the corresponding probability output.

[0043] S300: If it is necessary to optimize the generator in the cyclic generative adversarial network, it is determined whether to perform gradient reversal according to the discrimination result.

[0044] Specifically, the cyclic generative adversarial network is trained, that is, the model parameters or network parameters of the generator and the discriminator are optimized. It is possible to fix the network parameters of the generator and optimize the network parameters of the discriminator using the cross-entropy loss function; fix the network parameters of the discriminator and optimize the network parameters of the generator using the cross-entropy loss function. Of course, it is also possible to optimize the network parameters of the discriminator and the generator simultaneously.

[0045] Regardless of which optimization method is used, if it is currently necessary to optimize the generator or the current training node is the generator optimization node, it is necessary to determine whether to perform gradient reversal (Gradient Reversal) on the gradient corresponding to the generator according to the discrimination result of the discriminator. That is, it is determined whether to perform gradient reversal (flipping) according to whether the discriminator determines that the generated speech data comes from the speech data set.

[0046] If it is not currently necessary to optimize the model parameters of the generator, it is not necessary to determine whether to perform gradient reversal according to the discrimination result, nor is it necessary to perform gradient reversal.

[0047] After step S300 is executed, step S400 or step S600 is entered.

[0048] S400: If it is determined to perform gradient reversal according to the discrimination result, calculate the first loss function corresponding to the generator, and reverse the first gradient calculated according to the first loss function.

[0049] Specifically, if it is determined according to the discrimination result that the gradient of the generator needs to be reversed, calculate the first loss function corresponding to the generator of the current training node, calculate the corresponding first gradient according to the first loss function, and reverse the first gradient.

[0050] By calculating the gap between the distribution of the generated speech data and the distribution of the real speech data, the first loss function corresponding to the generator can be obtained.

[0051] S500: Update the model parameters of the generator according to the reversed first gradient.

[0052] Specifically, update the model parameters of the generator according to the reversed first gradient to obtain a new pre-trained cyclic generative adversarial network. After step S500 is executed, step S700 is entered, and the new pre-trained cyclic generative adversarial network is iteratively trained again using the data set.

[0053] S600: If it is determined not to perform gradient reversal according to the discrimination result, calculate the first loss function corresponding to the generator, calculate the first gradient according to the first loss function, and update the model parameters of the generator according to the first gradient.

[0054] Specifically, if it is determined not to perform gradient reversal according to the discrimination result, calculate the first loss function corresponding to the generator of the current training node, calculate the corresponding first gradient according to the first loss function, and update the model parameters of the generator according to the non-reversed first gradient to obtain a new pre-trained cyclic generative adversarial network. After step S600 is executed, step S700 is entered, and the new pre-trained cyclic generative adversarial network is iteratively trained again using the data set.

[0055] In this embodiment, it is determined whether to perform reversal according to the discrimination result of the discriminator. When the discrimination result is that the generated speech data comes from the training sample, reversal is performed. When the discrimination result is that the generated speech data does not belong to the training sample, the gradient is normally backpropagated and the gradient is not changed.

[0056] In this embodiment, the gradient can be calculated through the backpropagation algorithm (taking the gap between the predicted value and the true value as the loss and passing it backward layer by layer, and each layer of the network will calculate the gradient according to the passed-back loss), and the stochastic gradient descent algorithm and the calculated gradient are used to update the model parameters of the cyclic generative adversarial network.

[0057] S700: If the cyclic generative adversarial network has not converged, re-execute step S200 until the cyclic generative adversarial network converges.

[0058] Specifically, during the training of the cyclic generative adversarial network, the generator and the discriminator are alternately trained (alternately optimizing the parameters). When one party is fixed and the network parameters (model parameters) of the other party are updated, and alternately iterated. During this process, both the generator and the discriminator strive to optimize their own networks, thus forming a competitive confrontation until both parties reach a dynamic balance (Nash equilibrium).

[0059] If the cyclic generative adversarial network does not meet the convergence condition, jump to execute step S200 and the subsequent steps, and iteratively train the new pre-trained cyclic generative adversarial network using the data set until the cyclic generative adversarial network converges. The convergence condition can be that the number of training times reaches the preset number of iterations or the loss function of the cyclic generative adversarial network is less than the loss threshold.

[0060] In addition, the first tone domain speech data set can be divided into a first training set and a first test set, and the second tone domain speech data set can be divided into a second training set and a second test set. The cyclic generative adversarial network is trained using the first training set and the second training set, and the trained cyclic generative adversarial network is verified using the first test set and the second test set.

[0061] In this embodiment, by taking advantage of the superiority of the cycle generative adversarial network in fitting the real data distribution, a voice conversion model based on the cycle generative adversarial network is established, which supports unsupervised training. It can perform the conversion of the source voice to the target voice timbre between the source domain and the target domain without the need for one-to-one correspondence of the voice data in the two data sets, that is, a large amount of unsupervised data is used for the conversion of the source voice to the target voice timbre, and a large number of paired voice data in the first timbre domain and the second timbre domain are not required to achieve the timbre migration of the voice.

[0062] In this embodiment, the cycle generative adversarial network is also used to constrain and limit the generator to retain the voice content features of the source domain voice. It realizes the conversion of the voice of the source speaker to the voice timbre of the target speaker through timbre extraction and content decoupling while keeping the voice content the same.

[0063] In addition, this embodiment uses the gradient reversal mechanism for the cycle generative adversarial network, which enhances the discrimination ability between the target timbre and the source timbre, thereby improving the generation ability of the generator for the target timbre in the reverse direction. The cycle consistency of the cycle generative adversarial network makes the reconstruction of the voice conversion more stable and the model training more accurate, thus improving the accuracy of the voice conversion.

[0064] Figure 2 It is a schematic structural diagram of the cycle generative adversarial network in an embodiment of the present application. Refer to Figure 2 , the cycle generative adversarial network includes a first generator, a second generator, a first discriminator, and a second discriminator;

[0065] Step S200 specifically includes:

[0066] The first generator generates first generated voice data imitating the second timbre domain from the voice data in the first timbre domain, and the second generator reconstructs the first generated voice data to obtain second generated voice data imitating the first timbre domain;

[0067] The second generator generates third generated voice data imitating the first timbre domain from the voice data in the second timbre domain, and the first generator reconstructs the third generated voice data to obtain fourth generated voice data imitating the second timbre domain;

[0068] The second discriminator determines whether the first generated voice data is voice data in the second timbre domain to obtain a first discrimination result;

[0069] The first discriminator determines whether the third generated voice data is voice data in the first timbre domain to obtain a second discrimination result.

[0070] Specifically, the Cycle Generative Adversarial Network (Cycle GAN model) consists of two sets of GAN models in a dual form, and each set of GAN models includes a generator and an associated discriminator. The input of the first generator is the speech data in the first timbre domain or the speech data imitating the first timbre domain. The first generator is used to generate the speech data imitating the second timbre domain from the speech data in the first timbre domain, or to reconstruct the speech data imitating the second timbre domain from the speech data imitating the first timbre domain.

[0071] The input of the second generator is the speech data in the second timbre domain or the speech data imitating the second timbre domain. The second generator is used to generate the speech data imitating the first timbre domain from the speech data in the second timbre domain, or to reconstruct the speech data imitating the first timbre domain from the speech data imitating the second timbre domain.

[0072] The first discriminator is used to determine whether the generated speech data imitating the first timbre domain is the speech data in the first timbre domain or conforms to the distribution of the speech data in the first timbre domain, that is, to output the probability that the generated speech data imitating the first timbre domain is the speech data in the first timbre domain. More specifically, the speech data in the first timbre domain and the third generated speech data are used as the input of the first discriminator, and the first discriminator is used to determine whether the third generated speech data is the speech data in the first timbre domain, that is, to determine the probability that the third generated speech data comes from the speech data in the first timbre domain, and a second discrimination result is obtained.

[0073] The second discriminator is used to determine whether the generated speech data imitating the second timbre domain is the speech data in the second timbre domain or conforms to the distribution of the speech data in the second timbre domain, that is, to output the probability that the generated speech data imitating the second timbre domain is the speech data in the second timbre domain. More specifically, the speech data in the second timbre domain and the first generated speech data are used as the input of the second discriminator, and the second discriminator is used to determine whether the first generated speech data is the speech data in the second timbre domain, that is, to determine the probability that the first generated speech data comes from the speech data in the second timbre domain, and a first discrimination result is obtained.

[0074] In addition, step S200 specifically further includes:

[0075] Input the speech data in the first timbre domain into the first discriminator for discrimination to obtain a third discrimination result;

[0076] Input the speech data in the second timbre domain into the second discriminator for discrimination to obtain a fourth discrimination result.

[0077] Specifically, the first discriminator is further configured to determine whether the speech data in the first timbre domain is the speech data in the first timbre domain, so as to obtain a third determination result, and further prompt the first discriminator to learn the speech distribution of the speech data in the first timbre domain. This is conducive to the first discriminator better distinguishing the speech data imitating the first timbre domain from the real speech data in the first timbre domain, and strengthening the ability to distinguish true and false.

[0078] The second discriminator is further configured to determine whether the speech data in the second timbre domain is the speech data in the second timbre domain, so as to obtain a fourth determination result, and further prompt the second discriminator to learn the speech distribution of the speech data in the second timbre domain. This is conducive to the second discriminator better distinguishing the speech data imitating the second timbre domain from the real speech data in the second timbre domain, and strengthening the ability to distinguish true and false.

[0079] Generally, the discriminator is a binary classifier. If the discriminator determines that the generated speech data comes from the training data or dataset, the output probability is 1; otherwise, the output probability is 0. When the true and false discrimination ability of the discriminator is weak, it may not be able to distinguish the true and false of the generated speech data. Therefore, the true and false discrimination ability of the discriminator can be trained through continuous iterative training. When the generation ability of the generator is weak, the generated speech data may be easily recognized by the discriminator. Through continuous iterative training, the generator can generate speech data that can pass for real to deceive the discriminator. In this way, through the game between the generator and the discriminator, the ability of the generator to generate imitation speech data approaching the target domain speech signal is improved. At the same time, the discrimination ability of the discriminator for the imitation speech data is improved.

[0080] In addition, the first loss function corresponding to the generator is calculated according to the speech data in the first timbre domain, the speech data in the second timbre domain, and the generated speech data, where the generated speech data includes at least one of the first generated speech data, the second generated speech data, the third generated speech data, and the fourth generated speech data.

[0081] In one embodiment, the cyclic generative adversarial network further includes a first gradient reversal layer and a second gradient reversal layer;

[0082] Steps S400 and S500 specifically include:

[0083] If the first discrimination result is that the first generated speech data is the speech data in the second timbre domain, calculate the first sub-loss function corresponding to the first generator, calculate the first sub-gradient corresponding to the first generator according to the first sub-loss function, reverse the first sub-gradient through the first gradient reversal layer, and update the model parameters of the first generator according to the reversed first sub-gradient. Among them, the first sub-loss function is calculated according to the speech data in the first timbre domain, the speech data in the second timbre domain, and the first converted speech data. The first converted speech data includes at least one of the first generated speech data, the second generated speech data, the third generated speech data, and the fourth generated speech data;

[0084] If the second discrimination result is that the third generated speech data is the speech data in the first timbre domain, calculate the second sub-loss function corresponding to the second generator, calculate the second sub-gradient corresponding to the second generator according to the second sub-loss function, reverse the second sub-gradient through the second gradient reversal layer, and update the model parameters of the second generator according to the reversed second sub-gradient. Among them, the second sub-loss function is calculated according to the speech data in the first timbre domain, the speech data in the second timbre domain, and the second converted speech data. The second converted speech data includes at least one of the first generated speech data, the second generated speech data, the third generated speech data, and the fourth generated speech data.

[0085] Specifically, the model parameters of the first generator and the second generator can be updated at the same training node or at different training nodes. However, whether the first generator and the second generator perform gradient reversal is not synchronized. The first generator determines whether to perform gradient reversal according to the first discrimination result, and the second generator determines whether to perform gradient reversal according to the second discrimination result. The first discrimination result and the second discrimination result are not necessarily the same discrimination result. Therefore, it is possible that the gradient of the first generator needs to be reversed while the gradient of the second generator does not need to be reversed, and it is also possible that the gradient of the first generator does not need to be reversed while the gradient of the second generator needs to be reversed. There is also a possibility that the gradients of both the first generator and the second generator need to be reversed, and there is also a possibility that the gradients of both the first generator and the second generator do not need to be reversed.

[0086] The first loss function includes the first sub-loss function corresponding to the first generator and the second sub-loss function corresponding to the second generator. If gradient reversal is required, the first generator updates its model parameters according to the reversed first sub-gradient. The first sub-gradient is calculated according to the first sub-loss function. The second generator updates its model parameters according to the reversed second sub-gradient. The second sub-gradient is calculated according to the second sub-loss function. The first gradient includes the first sub-gradient and the second sub-gradient.

[0087] In addition, the first sub-loss function and the second sub-loss function are losses obtained by calculating the difference between the generated speech data and the source speech data or the target speech data. The greater the difference, the higher the penalty the generator will receive.

[0088] In one embodiment, the method further includes:

[0089] Before the cyclic generative adversarial network converges, if the discriminator needs to be optimized, calculate the second loss function corresponding to the discriminator, calculate the second gradient corresponding to the discriminator according to the second loss function, and update the model parameters of the discriminator according to the second gradient.

[0090] Specifically, the second loss function includes a third sub-loss function corresponding to the first discriminator and a fourth sub-loss function corresponding to the second discriminator. The second gradient includes a third sub-gradient corresponding to the first discriminator and a fourth sub-gradient corresponding to the second discriminator.

[0091] The second loss function corresponding to the discriminator can be obtained by calculating the classification or discrimination result of the discriminator on the generated speech data.

[0092] Calculate the third sub-loss function corresponding to the first discriminator, calculate the third sub-gradient corresponding to the first discriminator according to the third sub-loss function, and update the model parameters of the first discriminator according to the third sub-gradient, where the third sub-loss function is calculated based on at least two of the first tone domain speech data, the second tone domain speech data, the first generated speech data, the second generated speech data, the third generated speech data, and the fourth generated speech data.

[0093] Calculate the fourth sub-loss function corresponding to the second discriminator, calculate the fourth sub-gradient corresponding to the second discriminator according to the fourth sub-loss function, and update the model parameters of the second discriminator according to the fourth sub-gradient, where the fourth sub-loss function is calculated based on at least two of the first tone domain speech data, the second tone domain speech data, the first generated speech data, the second generated speech data, the third generated speech data, and the fourth generated speech data.

[0094] When training the discriminator, there is no need to train the generator, so there is no need for the gradient of the generator nor to determine whether to perform gradient reversal.

[0095] In one embodiment, the method further includes:

[0096] Calculate the total loss function of the cyclic generative adversarial network, where the total loss function is calculated based on the cycle consistency loss function and the adversarial loss function, or, the cycle consistency loss function, the adversarial loss function, and the identity mapping loss function;

[0097] If the total loss function is less than the preset convergence value, it is determined that the Cycle Generative Adversarial Network converges.

[0098] Specifically, in the Cycle Generative Adversarial Network, the closer the distribution of the speech data generated by the generator is to the real speech data, the better. The stronger the discriminative ability of the discriminator to distinguish between true and false speech, the better. The optimization process of the Cycle Generative Adversarial Network is a Minimax game problem.

[0099] The total loss function of the Cycle Generative Adversarial Network (Cycle GAN) is specifically:

[0100] L = L1 + λ1L2 + λ2L3 or, L = L1 + λ1L2

[0101] Where, L is the total loss function, L1 is the cycle consistency loss function, L2 is the adversarial loss function, and L3 is the identity mapping loss function. λ1 and λ2 are hyperparameters set to represent the weights between the cycle consistency loss function, the adversarial loss function, and the identity mapping function.

[0102] The cycle consistency loss is specifically the difference or distance between the real speech data and the corresponding reconstructed generated speech data. In this application, the cycle consistency loss is the distance between the speech data in the first timbre domain and the second generated speech data, and the distance between the speech data in the second timbre domain and the fourth generated speech data.

[0103] The calculation formula of L1 is shown in formula (1):

[0104] L1 = L cyc (G, F) = E A~Pdata(A) [||F(G(A)) - A||1] + E B~Pdata(B) [||G(F(B)) - B||1] Formula (1)

[0105] Where, the mapping function G corresponding to the first generator is used for A -> B, and G(A) represents generating speech data imitating the second timbre domain according to the input of the speech data in the first timbre domain. F(G(A)) represents reconstructing the speech data imitating the second timbre domain into speech data imitating the first timbre domain.

[0106] The mapping function F corresponding to the second generator is used for B -> A, and F(B) represents generating speech data imitating the first timbre domain according to the input of the speech data in the second timbre domain. G(F(B)) represents reconstructing the speech data imitating the first timbre domain into speech data imitating the second timbre domain.

[0107] Where, A is the real speech data in the first timbre domain, and B is the real speech data in the second timbre domain. E A~Pdata(A)Represents the expected value of the voice data distribution in domain A (the first timbre domain), E B~Pdata(B) Represents the expected value of the voice data distribution in domain B (the second timbre domain).

[0108] The adversarial loss function, i.e., the discriminative loss function (GAN loss or Adversarial loss), is determined by the discrimination result of the discriminator, as shown in the following formulas (2) - (4):

[0109] L2 = L GAN (G, D B , A, B) + L GAN (F, D A , A, B) Formula (2)

[0110] Among them, L GAN (G, D B , A, B) = E B~Pdata(B) [log D B (B)] +

[0111] E A~Pdata(A) [log(1 - D B (G(A))]) Formula (3)

[0112] L GAN (F, D A , A, B) = E A~Pdata(A) [log D A (A)] +

[0113] E B~Pdata(B) [log(1 - D A (F(B))]) Formula (4)

[0114] The ontology mapping function is specifically shown in the following formula (5):

[0115] L3 = L id (G, F) = E B~Pdata(B) [||G(B) - B||1] +

[0116] E A~Pdata(A) [||F(A) - A||1] Formula (5)

[0117] Among them, D B (B) is the probability that the second discriminator determines that the real voice data in the second timbre domain comes from the voice data set in the second timbre domain (the probability of real voice data), D B (G(A)) is the probability that the second discriminator determines that the first generated voice data G(A) generated by the first generator G comes from the voice data set in the second timbre domain (the probability of real voice data).

[0118] Among them, DA (A) is the probability that the first discriminator determines that the real first timbre domain voice data comes from the first timbre domain voice data set (the probability of real voice data), D A (F(B)) is the probability that the first discriminator determines that the third generated voice data F(B) generated by the second generator F comes from the first timbre domain voice data set (the probability of real voice data).

[0119] During the model training process of the cyclic generative adversarial network, the model parameters in the network structure are optimized by continuously reducing the total loss function L until the model converges. The loss functions calculated during the gradient backpropagation process need to be backpropagated.

[0120] In the cyclic generative adversarial network, it is trained according to stochastic gradient descent. When the model parameters of the generator are fixed, the model parameters of the discriminator are updated. When the model parameters of the discriminator are fixed, the model parameters of the generator are updated. Preferably, the discriminator is trained first, and then the generator is trained.

[0121] The first sub-loss function of the first generator is the total loss function L. Gradient reversal can reverse the gradient calculated according to the total loss function L, or it can only reverse the gradient calculated by the loss function L GAN (G, D B , A, B) and the gradients calculated by other loss functions in the total loss function L are not reversed.

[0122] The second sub-loss function of the second generator is the total loss function L. Gradient reversal can reverse the gradient calculated according to the total loss function L, or it can only reverse the gradient calculated by the loss function L GAN (F, D A , A, B) and the gradients calculated by other loss functions in the total loss function L are not reversed.

[0123] In one embodiment, reversing the first gradient includes: taking the reciprocal of the first gradient, or multiplying the first gradient by a preset negative number.

[0124] Specifically, for example, if the first gradient is g, the reversed first gradient obtained after taking the reciprocal of the first gradient is 1 / g.

[0125] The preset negative number can specifically be -1, and the value of the preset negative number can be specifically set according to the actual situation. This application is not limited thereto.

[0126] In one embodiment, the generator includes a first convolutional layer, a max pooling layer, a second convolutional layer, and an LSTM layer arranged in sequence, and the discriminator includes a third convolutional layer, an unfolding layer, a fully connected layer, and a normalization layer;

[0127] The LSTM layer is connected to the third convolutional layer of the corresponding discriminator through the corresponding gradient reversal layer.

[0128] Specifically, each generator includes at least two first convolutional layers arranged in sequence, and the size of the first convolutional layer is larger than that of the second convolutional layer. Each discriminator includes at least two third convolutional layers arranged in sequence. The size of the third convolutional layer can be the same as that of the second convolutional layer. The expansion layer is used to reduce the dimension of the output of the third convolutional layer, and the normalization layer is used to output probability values. The normalization layer specifically uses the softmax function.

[0129] The LSTM layer of the first generator is connected to the third convolutional layer of the second discriminator through the first reversal layer, and the LSTM layer of the second generator is connected to the third convolutional layer of the first discriminator through the second reversal layer.

[0130] Among them, the corresponding English for the reversal layer is the Gradient Reversal Layer. The reversal layer does not affect the parameter change of the discriminator. The standard reversal layer will change the gradient transmission from positive to negative, or from negative to positive, or it can also become the reciprocal, so as to change the parameters of the generator according to the reversed gradient to optimize in order to minimize the loss.

[0131] This application also provides a voice conversion method based on a cyclic adversarial generation network. The voice conversion method includes:

[0132] Using the trained voice conversion model based on the cyclic adversarial generation network to perform timbre conversion on the input first timbre domain voice to be converted, and obtaining the corresponding target second timbre domain voice data, where the trained voice conversion model based on the cyclic adversarial generation network is obtained according to the training method of the voice conversion model based on the cyclic adversarial generation network in any of the previous items.

[0133] This application uses the cyclic consistency loss of the cyclic generative adversarial network to enable unsupervised training in the model training of voice conversion, and can also make the reconstruction of the codec more stable during training. The gradient flipping technology is introduced to optimize the generation ability of the generator, enhance the adversarial training between the generator and the discriminator, enhance the discrimination ability of the discriminator, and thus improve the generation ability of the generator for the target timbre in reverse, making the trained voice conversion model more robust and the voice conversion ability stronger.

[0134] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0135] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0136] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0137] Figure 3 It is a structural block diagram of a training device for a voice conversion model based on a cyclic adversarial generation network in an embodiment of the present application. Refer to Figure 3 As shown in the figure, the training device for the voice conversion model based on the cyclic adversarial generation network includes:

[0138] A sample acquisition module 100, configured to acquire a first voice dataset in a first timbre domain and a second voice dataset in a second timbre domain as training samples;

[0139] A training module 200, configured to input the first voice data selected from the first voice dataset in the first timbre domain and the second voice data selected from the second voice dataset in the second timbre domain into a constructed cyclic generative adversarial network to perform cyclic adversarial training on the cyclic generative adversarial network, and obtain the discrimination result of the discriminator in the cyclic generative adversarial network;

[0140] A judgment module 300, configured to determine whether to perform gradient reversal according to the discrimination result if it is necessary to optimize the generator in the cyclic generative adversarial network;

[0141] A reversal module 400, configured to calculate the first loss function corresponding to the generator and reverse the first gradient calculated according to the first loss function if it is determined to perform gradient reversal according to the discrimination result,

[0142] A first parameter update module 500, configured to update the model parameters of the generator according to the reversed first gradient;

[0143] A second parameter update module 600, configured to calculate the first loss function corresponding to the generator, calculate the first gradient according to the first loss function, and update the model parameters of the generator according to the first gradient if it is determined not to perform gradient reversal according to the discrimination result;

[0144] An iterative module 700, which is used to jump to the training module 200 if the cyclic generative adversarial network does not converge until the cyclic generative adversarial network converges.

[0145] In one embodiment, the cyclic generative adversarial network includes a first generator, a second generator, a first discriminator, and a second discriminator;

[0146] The training module 200 specifically includes:

[0147] A first generation module, which is used to generate first generated speech data imitating the second timbre domain from the first timbre domain speech data through the first generator,

[0148] A second generation module, which is used to reconstruct the first generated speech data through the second generator to obtain second generated speech data imitating the first timbre domain;

[0149] The second generation module is further used to generate third generated speech data imitating the first timbre domain from the second timbre domain speech data through the second generator,

[0150] The first generation module is further used to reconstruct the third generated speech data through the first generator to obtain fourth generated speech data imitating the second timbre domain;

[0151] A first discrimination module, which is used to judge whether the first generated speech data is second timbre domain speech data through the second discriminator to obtain a first discrimination result;

[0152] A second discrimination module, which is used to judge whether the third generated speech data is first timbre domain speech data through the first discriminator to obtain a second discrimination result.

[0153] In one embodiment, the cyclic generative adversarial network further includes a first gradient reversal layer and a second gradient reversal layer;

[0154] The reversal module 400 includes:

[0155] A first reversal module, which is used to calculate a first sub-loss function corresponding to the first generator if the first discrimination result is that the first generated speech data is second timbre domain speech data, calculate a first sub-gradient corresponding to the first generator according to the first sub-loss function, and reverse the first sub-gradient through the first gradient reversal layer, where the first sub-loss function is calculated according to the first timbre domain speech data, the second timbre domain speech data, and first converted speech data, and the first converted speech data includes at least one of the first generated speech data, the second generated speech data, the third generated speech data, and the fourth generated speech data;

[0156] A second inversion module, configured to calculate a second sub-loss function corresponding to the second generator if the second discrimination result indicates that the third generated voice data is voice data in the first timbre domain, calculate a second sub-gradient corresponding to the second generator according to the second sub-loss function, and invert the second sub-gradient through a second gradient inversion layer, where the second sub-loss function is calculated according to the voice data in the first timbre domain, the voice data in the second timbre domain, and the second converted voice data, and the second converted voice data includes at least one of the first generated voice data, the second generated voice data, the third generated voice data, and the fourth generated voice data.

[0157] The first parameter update module 500 specifically includes:

[0158] A first sub-update module, configured to update the model parameters of the first generator according to the inverted first sub-gradient;

[0159] A second sub-update module, configured to update the model parameters of the second generator according to the inverted second sub-gradient.

[0160] In one embodiment, the apparatus further includes:

[0161] A third parameter update module, configured to calculate a second loss function corresponding to the discriminator if the discriminator needs to be optimized before the cyclic generative adversarial network converges, calculate a second gradient corresponding to the discriminator according to the second loss function, and update the model parameters of the discriminator according to the second gradient.

[0162] In one embodiment, the apparatus further includes:

[0163] A total loss calculation module, configured to calculate a total loss function of the cyclic generative adversarial network, where the total loss function is calculated according to a cyclic consistency loss function and an adversarial loss function, or, a cyclic consistency loss function, an adversarial loss function, and an identity mapping loss function;

[0164] A convergence determination module, configured to determine that the cyclic generative adversarial network converges if the total loss function is less than a preset convergence value.

[0165] In one embodiment, the inversion module 400 includes:

[0166] An inversion unit, configured to take the reciprocal of the first gradient, or multiply the first gradient by a preset negative number.

[0167] In one embodiment, the generator includes a first convolutional layer, a max pooling layer, a second convolutional layer, and an LSTM layer arranged in sequence, and the discriminator includes a third convolutional layer, an unfolding layer, a fully connected layer, and a normalization layer arranged in sequence;

[0168] The LSTM layer is connected to the third convolutional layer of the corresponding discriminator through the corresponding gradient reversal layer.

[0169] The meanings of "first" and "second" in the above modules / units are only used to distinguish different modules / units, and are not used to limit which module / unit has a higher priority or other limiting meanings. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or modules does not necessarily have to be limited to those steps or modules clearly listed, but may include other steps or modules not clearly listed or inherent to these processes, methods, products or devices. The division of modules in this application is only a logical division, and there may be other division methods in actual implementation.

[0170] For the specific limitations of the training device of the voice conversion model based on the cyclic adversarial generation network, reference can be made to the limitations of the training method of the voice conversion model based on the cyclic adversarial generation network in the above text, which will not be elaborated here. Each module in the above training device of the voice conversion model based on the cyclic adversarial generation network can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0171] Figure 4 It is the internal structure block diagram of the computer device in an embodiment of this application. As Figure 4As shown in the figure, the computer device includes a processor, a memory, a network interface, an input device, and a display screen connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory includes a storage medium and an internal memory. The storage medium can be a non-volatile storage medium or a volatile storage medium. The storage medium stores an operating system and can also store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can implement the training method of the voice conversion model based on the cyclic adversarial generation network or the voice conversion method based on the cyclic adversarial generation network. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the storage medium. The internal memory can also store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the training method of the voice conversion model based on the cyclic adversarial generation network or the voice conversion method based on the cyclic adversarial generation network. The network interface of the computer device is used to communicate with an external server via a network connection. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0172] In one embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions (such as a computer program) stored on the memory and executable on the processor. When the processor executes the computer-readable instructions, it implements the steps of the training method of the voice conversion model based on the cyclic adversarial generation network or the voice conversion method based on the cyclic adversarial generation network in the above embodiments, such as Figure 1 The steps S100 to S700 shown and the extensions and related steps of the method. Or, when the processor executes the computer-readable instructions, it implements the functions of each module / unit of the training device of the voice conversion model based on the cyclic adversarial generation network or the voice conversion device based on the cyclic adversarial generation network in the above embodiments, such as Figure 3 The functions of the modules 100 to 700 shown. To avoid repetition, it will not be elaborated here.

[0173] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The processor is the control center of the computer device and connects all parts of the entire computer device through various interfaces and circuits.

[0174] The memory can be used to store computer-readable instructions and / or modules. The processor realizes various functions of the computer device by running or executing the computer-readable instructions and / or modules stored in the memory, and by invoking the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, video data, etc.).

[0175] The memory can be integrated in the processor or can be separately provided from the processor.

[0176] Those skilled in the art can understand that Figure 4 the structure shown in

[0177] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. Figure 1 In one embodiment, a computer-readable storage medium is provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, the steps of the training method of the voice conversion model based on the cyclic adversarial generation network or the voice conversion method based on the cyclic adversarial generation network in the above embodiments are realized, such as Figure 3The functions of modules 100 to 700 shown are not described again here to avoid repetition.

[0178] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium. When the computer-readable instructions are executed, they can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0179] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, article, or method including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, device, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, device, article, or method including that element.

[0180] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments. Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0181] The above are only the preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A training method for a voice conversion model based on a cyclic adversarial generation network, characterized in that, The training method of the voice conversion model includes: Obtaining a first timbre domain voice dataset and a second timbre domain voice dataset as training samples; Inputting the first timbre domain voice data selected from the first timbre domain voice dataset and the second timbre domain voice data selected from the second timbre domain voice dataset into the constructed cyclic generative adversarial network to perform cyclic adversarial training on the cyclic generative adversarial network, and obtaining the discrimination result of the discriminator in the cyclic generative adversarial network; the cyclic generative adversarial network includes a first generator, a second generator, a first discriminator, and a second discriminator; the step of inputting the first timbre domain voice data selected from the first timbre domain voice dataset and the second timbre domain voice data selected from the second timbre domain voice dataset into the constructed cyclic generative adversarial network to perform cyclic adversarial training on the cyclic generative adversarial network and obtaining the discrimination result of the discriminator in the cyclic generative adversarial network includes generating, by the first generator, the first generated voice data imitating the second timbre domain from the first timbre domain voice data, and reconstructing, by the second generator, the first generated voice data to obtain the second generated voice data imitating the first timbre domain; generating, by the second generator, the third generated voice data imitating the first timbre domain from the second timbre domain voice data, and reconstructing, by the first generator, the third generated voice data to obtain the fourth generated voice data imitating the second timbre domain; judging, by the second discriminator, whether the first generated voice data is the second timbre domain voice data to obtain a first discrimination result; judging, by the first discriminator, whether the third generated voice data is the first timbre domain voice data to obtain a second discrimination result; If it is necessary to optimize the generator in the cyclic generative adversarial network, determine whether to perform gradient reversal according to the discrimination result; If it is determined to perform gradient reversal according to the discrimination result, calculate the first loss function corresponding to the generator, and reverse the first gradient calculated according to the first loss function, Update the model parameters of the generator according to the reversed first gradient; If it is determined not to perform gradient reversal according to the discrimination result, calculate the first loss function corresponding to the generator, calculate the first gradient according to the first loss function, and update the model parameters of the generator according to the first gradient; If the cyclic generative adversarial network does not converge, re-execute the above step of inputting the first timbre domain voice data selected from the first timbre domain voice dataset and the second timbre domain voice data selected from the second timbre domain voice dataset into the constructed cyclic generative adversarial network to perform cyclic adversarial training on the cyclic generative adversarial network until the cyclic generative adversarial network converges.

2. The method according to claim 1, characterized in that, The cyclic generative adversarial network further includes a first gradient reversal layer and a second gradient reversal layer; If it is determined to perform gradient reversal according to the discrimination result, calculate a first loss function corresponding to the generator, reverse a first gradient calculated according to the first loss function, and update model parameters of the generator according to the reversed first gradient, including: If the first discrimination result is that the first generated speech data is speech data in the second timbre domain, calculate a first sub-loss function corresponding to the first generator, calculate a first sub-gradient corresponding to the first generator according to the first sub-loss function, reverse the first sub-gradient through the first gradient reversal layer, and update model parameters of the first generator according to the reversed first sub-gradient, where the first sub-loss function is calculated according to the first timbre domain speech data, the second timbre domain speech data, and first converted speech data, and the first converted speech data includes at least one of the first generated speech data, the second generated speech data, the third generated speech data, and the fourth generated speech data; If the second discrimination result is that the third generated speech data is speech data in the first timbre domain, calculate a second sub-loss function corresponding to the second generator, calculate a second sub-gradient corresponding to the second generator according to the second sub-loss function, reverse the second sub-gradient through the second gradient reversal layer, and update model parameters of the second generator according to the reversed second sub-gradient, where the second sub-loss function is calculated according to the first timbre domain speech data, the second timbre domain speech data, and second converted speech data, and the second converted speech data includes at least one of the first generated speech data, the second generated speech data, the third generated speech data, and the fourth generated speech data.

3. The method according to claim 2, characterized in that, The method further includes: Before the cyclic generative adversarial network converges, if it is necessary to optimize the discriminator, calculate a second loss function corresponding to the discriminator, calculate a second gradient corresponding to the discriminator according to the second loss function, and update model parameters of the discriminator according to the second gradient.

4. The method according to claim 1, characterized in that, The method further includes: Calculate a total loss function of the cyclic generative adversarial network, where the total loss function is calculated according to a cycle consistency loss function and an adversarial loss function, or, a cycle consistency loss function, an adversarial loss function, and an identity mapping loss function; If the total loss function is less than a preset convergence value, determine that the cyclic generative adversarial network converges.

5. The method according to claim 1, characterized in that, Reversing the first gradient includes: taking the reciprocal of the first gradient, or multiplying the first gradient by a preset negative number.

6. The method according to claim 1, characterized in that, The generator includes a first convolutional layer, a max pooling layer, a second convolutional layer, and an LSTM layer arranged in sequence, and the discriminator includes a third convolutional layer, an unfolding layer, a fully connected layer, and a normalization layer arranged in sequence; The LSTM layer is connected to the third convolutional layer of the corresponding discriminator through a corresponding gradient reversal layer.

7. A voice conversion method based on a cyclic adversarial generation network, characterized in that, The speech conversion method includes: Use the trained voice conversion model based on the cyclic adversarial network to perform voice conversion on the input first timbre domain voice to be converted, and obtain the corresponding target second timbre domain voice data, where the trained voice conversion model based on the cyclic adversarial network is obtained according to the training method of the voice conversion model based on the cyclic adversarial network according to any one of claims 1-6.

8. A training device for a voice conversion model based on a cyclic adversarial generation network, characterized in that, The training device of the voice conversion model includes: A sample acquisition module, configured to acquire a first timbre domain voice data set and a second timbre domain voice data set as training samples; A training module, configured to input the first timbre domain voice data selected from the first timbre domain voice data set and the second timbre domain voice data selected from the second timbre domain voice data set into the constructed cyclic generative adversarial network to perform cyclic adversarial training on the cyclic generative adversarial network, and obtain the discrimination result of the discriminator in the cyclic generative adversarial network; the cyclic generative adversarial network includes a first generator, a second generator, a first discriminator, and a second discriminator; the training module includes: a first generation module, configured to generate first generated voice data imitating the second timbre domain by the first generator; a second generation module, configured to reconstruct the first generated voice data through the second generator to obtain second generated voice data imitating the first timbre domain; the second generation module is further configured to generate third generated voice data imitating the first timbre domain by the second generator; the first generation module is further configured to reconstruct the third generated voice data through the first generator to obtain fourth generated voice data imitating the second timbre domain; a first discrimination module, configured to determine whether the first generated voice data is the second timbre domain voice data through the second discriminator to obtain a first discrimination result; a second discrimination module, configured to determine whether the third generated voice data is the first timbre domain voice data through the first discriminator to obtain a second discrimination result; A judgment module, configured to determine whether to perform gradient reversal according to the discrimination result if it is necessary to optimize the generator in the cyclic generative adversarial network; A reversal module, configured to calculate a first loss function corresponding to the generator and reverse the first gradient calculated according to the first loss function if it is determined to perform gradient reversal according to the discrimination result, A first parameter update module, configured to update the model parameters of the generator according to the reversed first gradient; A second parameter update module, configured to calculate a first loss function corresponding to the generator, calculate a first gradient according to the first loss function, and update the model parameters of the generator according to the first gradient if it is determined not to perform gradient reversal according to the discrimination result; An iteration module, configured to jump to the training module if the cyclic generative adversarial network does not converge until the cyclic generative adversarial network converges.

9. A computer device, comprising a memory, a processor, and computer-readable instructions stored on the memory and executable on the processor, characterized in that, When the processor executes the computer-readable instructions, it performs the steps of the training method of the voice conversion model based on the cyclic adversarial generation network according to any one of claims 1-6, or when the processor executes the computer-readable instructions, it performs the steps of the voice conversion method based on the cyclic adversarial generation network according to claim 7.

10. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that, When the computer-readable instructions are executed by the processor, the processor is caused to perform the steps of the training method of the voice conversion model based on the cyclic adversarial generation network according to any one of claims 1-6, or the processor is caused to perform the steps of the voice conversion method based on the cyclic adversarial generation network according to claim 7.

Citation Information

Patent Citations

  • Method, apparatus and equipment for establishing voice enhancement network and computer storage medium

    CN109147810A

  • Text information classification method and device based on GAN and storage medium

    CN113010675A