Audio generation method, training method for related model, and related device

By introducing a conversion model and a tone discriminator in the audio generation method, the problem of tone and accent coupling in the prior art is solved, and the audio of specified tone and accent is generated is realized, and the accuracy of generation is improved.

CN114420083BActive Publication Date: 2025-06-13XIAN XUNFEI SUPER BRAIN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111491439.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-08
Publication Date
2025-06-13
Estimated Expiration
2041-12-08

AI Technical Summary

Technical Problem

After the existing audio generation method, after the specified tone, the synthesized audio is difficult to convert to other accents, and after the specified accent, the synthesized audio is difficult to convert to other accents, resulting in the coupling of tone and accents, making it difficult to generate audio of any tone and accent.

Method used

It provides a training method and audio generation method for converting models. By inputting the sample text, the identification of sample accents and the identification of sample tones, the sample synthesized spectrum corresponding to the sample accents and sample tones is generated, and the model parameters are adjusted using the tone discriminator to generate audio for the specified tones and the specified accent.

Benefits of technology

The audio generated with specified tone and specified accent is realized, avoiding the problem of the generated audio carrying the accent or tone of the original audio, and improving the accuracy of audio generation for specific accents and specific tone.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114420083B_ABST
    Figure CN114420083B_ABST
Patent Text Reader

Abstract

The present application discloses an audio generation method, a training method for a related model, and related devices and equipment. Among them, the audio generation method includes: performing text encoding of a target accent on the target text to obtain a target text encoding vector of the target accent; performing decoding processing on the target text encoding vector and a target timbre vector corresponding to the target timbre to obtain target sub-spectra at a plurality of moments corresponding to the target timbre; performing synthesis processing on the target sub-spectra at the plurality of moments to obtain a target synthesis spectrum corresponding to the target text. Through the above method, it is possible to generate an audio with a specified timbre and a specified accent by using the text. In addition, a timbre discriminator can be used to train a conversion model for the spectrum of the audio that generates the target accent and the target timbre based on the target text, which can make the timbre of the synthesis spectrum generated by the trained conversion model tend to be consistent with the specified timbre and improve the accuracy of model conversion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of audio processing, and in particular to an audio generation method, a training method for a related model, and a related device. Background Art

[0002] Human speech can reflect the timbre of the speaker. Timbre is a characteristic of the speaker itself and can be used to distinguish different speakers. Speech can also reflect the accent of the speaker. An accent is the general pronunciation feature of different groups of people using the same language. For example, the American accent and British accent of English, Mandarin and Sichuan dialect of Chinese, etc.

[0003] Existing audio generation methods usually perform specific conversions on the original audio to generate new audio. However, during the long-term research and development process, the applicant of this application found that when generating audio using existing audio generation methods, since the original audio has the accent and phonemes of the speaker, after specifying the timbre, the synthesized audio has the accent of the speaker with that timbre and cannot be changed to other accents. Correspondingly, after specifying the accent, the synthesized audio has the timbre of the speaker with that accent and cannot be changed to other timbres. The coupling of timbre and accent makes it difficult to generate audio with arbitrary timbres and accents. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide an audio generation method, a training method for a related model, and a related device that can generate audio corresponding to a specified timbre and a specified accent.

[0005] To solve the above technical problem, a technical solution adopted by this application is: to provide a training method for a conversion model, the method includes the following steps to train the conversion model: input the sample text, the identifier of the sample accent, and the identifier of the sample timbre into the conversion model to obtain a sample synthesis spectrum corresponding to the sample accent and the sample timbre; use a timbre discriminator to perform timbre discrimination on the sample synthesis spectrum to obtain a first predicted timbre; based on the sample timbre and the first predicted timbre, adjust the parameters of the conversion model.

[0006] To solve the above technical problem, another technical solution adopted by this application is: to provide an audio generation method, the method includes: performing text encoding of the target accent on the target text to obtain a target text encoding vector of the target accent; performing decoding processing on the target text encoding vector and a target timbre vector corresponding to the target timbre to obtain target sub-spectrums at several moments corresponding to the target timbre; performing synthesis processing on the target sub-spectrums at several moments to obtain a target synthesis spectrum corresponding to the target text.

[0007] To solve the above technical problems, another technical solution adopted by this application is: to provide a training device for a conversion model, which includes an input module, a timbre discrimination module, and a model adjustment module. Among them, the input module is used to input a sample text, an identifier of a sample accent, and an identifier of a sample timbre into the conversion model to obtain a sample synthesis spectrum corresponding to the sample accent and the sample timbre; the timbre discrimination module is used to use a timbre discriminator to perform timbre discrimination on the sample synthesis spectrum to obtain a first predicted timbre; the model adjustment module is used to adjust the parameters of the conversion model based on the sample timbre and the first predicted timbre.

[0008] To solve the above technical problems, another technical solution adopted by this application is: to provide an audio generation device, which includes a text encoding module, a decoding module, and a synthesis module. Among them, the text encoding module is used to perform text encoding of a target accent on a target text to obtain a target text encoding vector of the target accent; the decoding module is used to perform decoding processing on the target text encoding vector and a target timbre vector corresponding to the target timbre to obtain target sub-spectrums at several moments corresponding to the target timbre; the synthesis module is used to perform synthesis processing on the target sub-spectrums at several moments to obtain a target synthesis spectrum corresponding to the target text.

[0009] To solve the above technical problems, another technical solution adopted by this application is: to provide an electronic device, which includes a processor and a memory. The memory is used to store program data, and the processor is used to execute the program data to implement the above conversion model training method or audio generation method.

[0010] In the above solution, by performing text encoding on the target text to obtain the target text encoding vector of the target accent, and performing decoding processing on the target text encoding vector and the target timbre, the spectrum corresponding to the target text is obtained. Therefore, without relying on the original audio, directly relying on the text to obtain the spectrum of the audio with the content of the target text, the target accent, and the target timbre, so it is possible to use the text to generate an audio with a specified timbre and a specified accent. Moreover, since there is no need to rely on the original audio, the problem that the generated audio carries the accent or timbre of the original audio can be avoided, and the accuracy of generating an audio with a specific accent and a specific timbre can be improved.

[0011] In addition, in the above solution, the sample text, the identifier of the sample accent, and the identifier of the sample timbre can be input into the conversion model to obtain a sample synthesis spectrum corresponding to the sample accent and the sample timbre. The timbre discriminator is used to train and discriminate the timbre of the sample synthesis spectrum. Based on the sample timbre and the first predicted timbre, the parameters of the conversion model are adjusted to obtain a conversion model that can generate a conversion model corresponding to the specified timbre and the specified accent. The conversion model is used to perform text encoding of the target accent on the target text, and then the obtained target text encoding vector and the target timbre vector corresponding to the target timbre are decoded. The obtained target sub-spectrums at several moments are synthesized to obtain a target synthesis spectrum corresponding to the target text, the target timbre, and the target accent. Through the above method, using the timbre discriminator to assist in the training of the model can make the timbre of the synthesis spectrum generated by the conversion model tend to be consistent with the specified timbre, improve the accuracy of model conversion, and finally the conversion model can generate a target synthesis spectrum with any specified timbre and any specified accent. Description of the Drawings

[0012] Figure 1 is a schematic flowchart of an embodiment of the audio generation method of the present application;

[0013] Figure 2 is a schematic flowchart of another embodiment of the audio generation method of the present application;

[0014] Figure 3 is a schematic flowchart of another embodiment of step S220 of the present application;

[0015] Figure 4 is a schematic flowchart of an embodiment of the layer normalization processing of the present application;

[0016] Figure 5 is a schematic flowchart of another embodiment of step S230 of the present application;

[0017] Figure 6 is a schematic flowchart of another embodiment of the audio generation method of the present application;

[0018] Figure 7 is a schematic flowchart of an embodiment of the training method of the conversion model of the present application;

[0019] Figure 8 is a schematic flowchart of another embodiment of step S730 of the present application;

[0020] Figure 9 is a schematic flowchart of another embodiment of step S833 of the present application;

[0021] Figure 10 is a schematic flowchart of another embodiment of step S740 of the present application;

[0022] Figure 11 It is a schematic framework diagram of an embodiment of the training device for the conversion model of the present application;

[0023] Figure 12 It is a schematic framework diagram of an embodiment of the audio generation device of the present application;

[0024] Figure 13 It is a schematic framework diagram of an embodiment of the electronic device of the present application;

[0025] Figure 14 It is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners

[0026] To make the objectives, technical solutions and effects of the present application clearer and more definite, the following further describes the present application in detail with reference to the accompanying drawings and by way of examples.

[0027] In this article, the term "and / or" merely describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two.

[0028] It can be understood that the method of the present application can include any one of the following method embodiments and the combination of any non-conflicting following method embodiments.

[0029] It can be understood that the audio generation method in the present application can be executed by an electronic device, which can also be simply referred to as a device. The device can be any device with execution ability, such as a mobile phone, a computer, a tablet computer, etc.

[0030] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the audio generation method of the present application. The method includes:

[0031] Step S110: Perform text encoding on the target text with the target accent to obtain the target text encoding vector with the target accent.

[0032] It should be noted that a person's voice can reflect the text content to be expressed by the speaker, and can also reflect the speaker's timbre and accent. Among them, the timbre is the characteristic of the speaker himself and can be used to distinguish different speakers, while the accent is the general pronunciation characteristic of different groups using the same language. For example, the British accent and the American accent in English, and the Sichuan dialect and Mandarin in Chinese, etc.

[0033] This application provides an audio generation method that can specify text content, accent, and voice. After the user determines the target text, target accent, and target voice, the audio corresponding to the target synthesis spectrum generated by the audio generation method in this application is the audio expressing the target text in the target accent and target voice.

[0034] It should be noted that the process of generating the target synthesis spectrum can be, but is not limited to, implemented using a pre-trained conversion model. Specifically, the process of generating the target synthesis spectrum can be divided into two stages. First, a target text encoding vector is obtained based on the target text and target accent. The target text encoding vector obtained in this stage does not contain voice information. Then, the target synthesis spectrum is obtained based on the target voice and the target text encoding vector.

[0035] The target text can be the text input by the user, and the length and language of the target text are not restricted. Generally speaking, the target text can be a sentence, such as "I love China", "Hello, world", etc.

[0036] Step S120: Decode the target text encoding vector and the target voice vector corresponding to the target voice to obtain the target sub-spectra at several moments corresponding to the target voice.

[0037] In this embodiment, the target synthesis spectrum can be divided into target sub-spectra at several moments. Specifically, for example, if the target text contains 6 phonemes, the target synthesis spectrum corresponding to these 6 phonemes can be divided into 60 target sub-spectra at moments. The number obtained here can be adjusted according to the actual needs of the user. By performing relevant decoding processing depending on the target text encoding vector and the target voice vector corresponding to the target voice, the 60 target sub-spectra at moments can be obtained. It can be understood that this decoding processing can be looped several times. Each time it loops, a target sub-spectrum at one moment can be obtained. Through several decoding operations, the target sub-spectra at several moments corresponding to the target voice can be obtained.

[0038] Step S130: Synthesize the target sub-spectra at several moments to obtain the target synthesis spectrum corresponding to the target text.

[0039] After obtaining the target sub-spectra at several moments, all the target sub-spectra are synthesized. For example, the target sub-spectra at several moments can be concatenated in sequence, and thus the target synthesis spectrum corresponding to the target text can be obtained.

[0040] In the above solution, the conversion model is used to perform text encoding of the target accent on the target text, and then decoding processing is performed on the obtained target text encoding vector and the target timbre vector corresponding to the target timbre. The obtained target sub-spectra at several moments are synthesized to obtain a target synthesis spectrum corresponding to the target text, target timbre, and target accent. In this way, the conversion model can generate a target synthesis spectrum with any specified timbre and any specified accent.

[0041] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another embodiment of the audio generation method of this application. The method includes:

[0042] Step S210: Convert the target text into several phonemes.

[0043] Among them, a phoneme is divided according to the natural attributes of speech and is the smallest speech unit or the smallest speech segment that constitutes a syllable. Analyzing according to the pronunciation actions in the syllable, one pronunciation action constitutes one phoneme. Specifically, for example, the Chinese syllable "ā" has only one phoneme, "ài" has two phonemes, and "dài" has three phonemes. It should be noted that different language systems use different phoneme systems, and the phonemes corresponding to different accents in the same language system may also be different. For example, the phoneme systems used in English and Chinese are different, and the phonemes corresponding to Cantonese and Mandarin in Chinese are also different.

[0044] Therefore, converting the target text into several phonemes is carried out according to the language and accent of the target text. For the same target text, if the target accent is different, the obtained several phonemes may also be different. Specifically, the device can convert the target text into several phonemes through the corresponding dictionary.

[0045] Step S220: Perform text encoding of the target accent on the target text to obtain a target text encoding vector of the target accent.

[0046] In this embodiment, the device can pre-store a trained conversion model, which can be used to implement the audio generation method in this embodiment. Specifically, it can include a decoder and several text encoders, and one text encoder corresponds to one accent. After determining the target accent, the text encoder to be used in the audio generation process can be determined, without involving the use of other text encoders.

[0047] It can be understood that the training process of several text encoders pre-stored in the device can include classifying the corpus (that is, the training samples) according to the accent, and respectively using each class of samples to train a text encoder, thus obtaining several text encoders corresponding to different accents respectively.

[0048] Specifically, the device can select the text encoder corresponding to the target accent from several text encoders based on the accent identifier of the target accent selected by the user. For example, it selects the text encoder associated with the accent identifier of the target accent. Since only the text encoder corresponding to the target accent is involved in the audio generation process and other text encoders are not used, for the sake of simplicity of description and easy understanding, the text encoder corresponding to the target accent is simply referred to as the text encoder in the subsequent audio generation steps.

[0049] Specifically, the device inputs the several phonemes obtained by converting the target text into the text encoder, and uses the text encoder to perform text encoding of the target accent on the several phonemes to obtain the target text encoding vectors corresponding to the several phonemes, and each phoneme corresponds to a target text encoding vector.

[0050] Please refer to Figure 2 and Figure 3 , Figure 3 which is a schematic flowchart of another embodiment of step S220 of this application. Step S220 includes:

[0051] Step S321: Obtain the first original phoneme vectors of several phonemes corresponding to the target text.

[0052] Among them, each phoneme corresponds to a first original phoneme vector, and the first original phoneme vector can characterize the characteristics of the corresponding phoneme.

[0053] In some embodiments, the text encoder includes a word embedding layer, a fully connected layer, and a first recurrent network layer. Among them, the word embedding layer can be used to convert phonemes into phoneme vectors.

[0054] Then step S321 can be implemented through the word embedding layer of the text encoder, and specifically can include: using the word embedding layer to obtain the second original phoneme vectors of several phonemes, and using the fully connected layer of the text encoder to process the second original phoneme vectors to obtain the first original phoneme vectors of each phoneme.

[0055] Step S322: Perform timbre removal processing on the first original phoneme vectors of each phoneme to obtain the target phoneme vectors of each phoneme.

[0056] It should be noted that there are differences in the timbre information carried by the audio of different speakers. The timbre information carried by the audio of the same speaker has relatively stable characteristics. The speaker of the audio can be identified through the characteristics of the timbre information carried by the audio. Removing the timbre means removing the characteristics of the timbre information belonging to the speaker, so that the timbre information is in an average state. This average state can be a standard normal distribution, rather than removing all timbre information. The timbre information of the audio after removing the timbre does not contain the characteristics of the timbre information belonging to any speaker, and thus the speaker cannot be identified. Then it can be said that this piece of audio does not carry characteristic timbre information. The purpose of removing the timbre in step S220 is to make the characteristics of the timbre information of the finally obtained target sub-spectrum be controlled only by the target timbre vector in step S230, which is beneficial to making the timbre of the target sub-spectrum close to the target timbre.

[0057] Specifically, step S322 can be to use the first original phoneme vector as the vector to be processed, and perform layer normalization processing on the vector to be processed to obtain the target phoneme vector, that is, the target vector. Among them, removing the timbre is achieved through layer normalization processing. It should be noted that each phoneme corresponds to a first original phoneme vector output by the fully connected layer. Taking the first original phoneme vector of each phoneme as the vector to be processed, layer normalization processing is respectively performed on each of them. Hereinafter, taking the processing of the first original phoneme vector of one phoneme as an example for illustration.

[0058] Please refer to Figure 2 、 Figure 3 and Figure 4 , Figure 4 is a schematic flowchart of an embodiment of layer normalization processing in the present application. The layer normalization processing includes:

[0059] Step S410: Obtain the first statistical value and the second statistical value of the vector to be processed.

[0060] Among them, the first statistical value is used to reflect the central tendency of the elements in the vector to be processed, for example, the average value. The second statistical value is used to reflect the degree of dispersion of the elements in the vector to be processed, for example, the variance.

[0061] Specifically, step S410 can be specifically implemented by the following formula:

[0062]

[0063]

[0064] where a is the vector to be processed, a iis an element of the vector to be processed, H is the number of elements in the vector to be processed, 1 ≤ i ≤ H, μ is the first statistical value, calculated by formula 1, which is the average value of the elements in the vector to be processed, and σ is the second statistical value, calculated by formula 2, which is the variance of the elements in the vector to be processed.

[0065] The vector to be processed here is the first original phoneme vector corresponding to a phoneme. The first original phoneme vector is output after being processed by a fully connected layer. Then the number of elements in the first original phoneme vector is determined by the number of nodes H in the fully connected layer, that is, the number of elements in the first original phoneme vector is the same as the number of nodes H in the fully connected layer. For a simple example, if the first original phoneme vector a is (1, 2, 3), then H = 3, a 1 = 1, a 2 = 2, a 3 = 3, and μ is the average value of a i and σ is the variance of a i

[0066] Step S420: Obtain an intermediate vector using the vector to be processed, the first statistical value, and the second statistical value, and perform activation processing on the intermediate vector using an activation function to obtain a target vector.

[0067] Specifically, step S420 can be specifically implemented through the following formula:

[0068]

[0069] where h is the target vector, f is the activation function, a is the vector to be processed, μ is the first statistical value, σ is the second statistical value, g and b are model parameters, determined during the model training process, φ is a minimum value, and the purpose of φ is to prevent a division-by-zero error when the variance is equal to 0.

[0070] For a vector to be processed, after steps S410 and S420, a target vector is obtained. The target vector is the result obtained after removing the timbre from the vector to be processed. Encoding the target vector obtained after removing the timbre can make the encoded result, that is, the timbre information of the audio corresponding to the target text encoding vector, not have the characteristic timbre information belonging to any speaker.

[0071] Step S323: Encode the target phoneme vectors of several phonemes to obtain a target text encoding vector of the target accent.

[0072] ​It should be noted that even for the same text content, the corresponding synthesized spectrum may be different. This is because the spectrum synthesis is also affected by the previous text, that is, the historical synthesis information. For a simple example, the target texts are "good thing" and "thing" respectively, both of which contain the text content "thing". However, in the target synthesized spectra corresponding to these two target texts, the spectra corresponding to the text content "thing" are also different. This is caused by the influence of the previous text on the audio. In step S323, each target phoneme vector is encoded in sequence to generate a target text encoding vector. The final target text encoding vector of the target accent contains all the target text encoding vectors corresponding to several phonemes of the target text. When generating each target text encoding vector, the influence of the already generated historical target text encoding vectors needs to be considered.

[0073] Since the recurrent network layer can retain certain historical processing information in the subsequent processing during the processing, it is suitable for processing audio. Step S323 can be implemented by the first recurrent network layer of the text encoder. Among them, the first recurrent network layer can encode the target phoneme vectors of several phonemes in sequence and output the target text encoding vectors corresponding to several phonemes in sequence, and finally obtain the target text encoding vector of the target accent.

[0074] Step S230: Decode the target text encoding vector and the target timbre vector corresponding to the target timbre to obtain the target sub-spectra at several moments corresponding to the target timbre.

[0075] It should be noted that step S230 can be implemented by a decoder. By continuously performing the decoding operation, the target sub-spectra at each moment can be obtained, so as to obtain the target sub-spectra at several moments corresponding to the target text.

[0076] During the process of generating the target sub-spectrum at the current moment, the influence of the target sub-spectrum at the previous moment also needs to be considered. Therefore, the work of the decoder is based on the target sub-spectrum, the target text encoding vector and the target timbre at the previous moment.

[0077] In a specific implementation scenario, the decoder is an autoregressive decoder. The decoder specifically includes a one-dimensional convolutional neural network layer, a timbre embedding layer and a second recurrent network layer. Among them, the second recurrent network layer is a position-sensitive attention and decoder recurrent network layer.

[0078] Please refer to Figure 2 and Figure 5 , Figure 5 which is a schematic flowchart of another embodiment of step S230 of the present application. Step S230 includes:

[0079] Step S531: Extract features from the target sub-spectrum at the previous moment to obtain a second spectral feature vector.

[0080] Specifically, step S531 can be implemented by a one-dimensional convolutional neural network layer of a decoder. The target sub-spectrum at the previous moment is input into the one-dimensional convolutional neural network layer, and a second spectral feature vector is output. The second spectral feature can be used to characterize the features of the target sub-spectrum.

[0081] Step S532: Perform timbre removal processing on the second spectral feature vector to obtain a first spectral feature vector.

[0082] The processing process of step S532 is similar to that of step S322. The relevant description of step S532 can refer to the relevant content of the foregoing step S322.

[0083] Specifically, step S532 can be: taking the second spectral feature vector as the vector to be processed, performing layer normalization processing on the vector to be processed, and obtaining the first spectral feature vector, that is, the target vector. Among them, the timbre removal processing is implemented through layer normalization processing.

[0084] Please refer to Figure 2 、 Figure 4 and Figure 5 , Figure 4 which is a schematic flowchart of an embodiment of layer normalization processing in this application. The layer normalization processing includes:

[0085] Step S410: Obtain a first statistical value and a second statistical value of the vector to be processed.

[0086] Specifically, step S410 can be specifically implemented by the following formula:

[0087]

[0088]

[0089] where a is the vector to be processed, a i is an element of the vector to be processed, H is the number of elements in the vector to be processed, 1≤i≤H, μ is the first statistical value, if calculated by formula 1, that is, the average value of the elements in the vector to be processed, and σ is the second statistical value, if calculated by formula 2, that is, the variance of the elements in the vector to be processed.

[0090] The vector to be processed here is the second spectral feature vector of the target sub-spectrum at the previous moment, and the second spectral feature vector is output after being processed by the one-dimensional convolutional neural network layer.

[0091] Step S420: Obtain an intermediate vector by using the vector to be processed, the first statistical value and the second statistical value, and perform activation processing on the intermediate vector by using an activation function to obtain the target vector.

[0092] Specifically, step S420 can be specifically implemented through the following formula:

[0093]

[0094] Among them, h is the target vector, f is the activation function, a is the vector to be processed, μ is the first statistical value, calculated through formula 1, σ is the second statistical value, calculated through formula 2, g and b are model parameters, φ is the minimum value, and the purpose of φ is to prevent division by zero error when the variance is equal to 0.

[0095] For a vector to be processed, after steps S410 and S420, a target vector is obtained. The target vector is the result obtained after removing the timbre information from the vector to be processed. Here, the target vector is the first spectral feature vector, and the first spectral feature vector is the feature vector of the target sub-spectrum at the previous moment without any speaker's timbre information. The timbre removal process here can ensure that when generating the target sub-spectrum at the current moment, it is not affected by the timbre information in the target sub-spectrum at the previous moment. That is to say, the timbre information feature of the target sub-spectrum at the current moment is only controlled by the target timbre vector, so that the timbre of the target sub-spectrum can be closer to the target timbre.

[0096] After being processed by steps S531 and S352, the feature vector of the target sub-spectrum at the previous moment without any speaker's characteristic timbre information is obtained, which can be used in the subsequent decoding process of the target sub-spectrum at the current moment.

[0097] Step S533: Perform fusion processing on the first spectral feature vector corresponding to the target sub-spectrum at the previous moment and the target timbre vector of the target timbre to obtain a first fusion vector.

[0098] If the current moment is the first moment after starting decoding, then all frequencies in the spectrum at the previous moment are defaulted to 0, and the first spectral feature vector and the second spectral feature vector obtained are both zero vectors.

[0099] It can be understood that after the user specifies the target timbre, the device can convert the target timbre vector according to the target timbre identifier, and the above conversion process can be implemented using the timbre embedding layer of the decoder.

[0100] In some embodiments, the above fusion processing can be element-wise addition processing.

[0101] Step S534: Use the first fusion vector and the target text encoding vector to obtain the target sub-spectrum at the current moment.

[0102] Step S534 can be implemented by the second recurrent network layer of the decoder. Specifically, the second recurrent network layer can be a position-sensitive attention and decoder recurrent network layer. Among them, the target text encoding vectors input to the position-sensitive attention and decoder recurrent network layer of the decoder are all the target text encoding vectors corresponding to several phonemes of the target text. Through position-sensitive attention processing, all the target text encoding vectors are weighted and summed to determine the weights of all the target text encoding vectors in the current decoding operation, and then decoding is performed according to the weights.

[0103] If the current time is t 0 , after steps S531 - S534 are executed and the target sub-spectrum at the current time is obtained, then the generation of the target sub-spectrum at the next time, that is, time t 1 , can be carried out. Then, the target sub-spectrum corresponding to time t 0 is used as the target sub-spectrum of the previous time of time t 1 . The device can then execute steps S531 - S534 again to generate the target sub-spectrum at time t 1 . This process is repeated continuously until all the target text encoding vectors are processed, and thus the target sub-spectra at several times are obtained.

[0104] Step S240: Synthesize the target sub-spectra at several times to obtain the target synthesized spectrum corresponding to the target text.

[0105] Specifically, synthesizing the target sub-spectra at several times can be to splice the target sub-spectra at several times in sequence to obtain the target synthesized spectrum.

[0106] Step S250: Generate a target audio with a target accent and a target timbre based on the target synthesized spectrum.

[0107] The device can pre-store a vocoder. Through the vocoder, a target audio can be generated based on the target synthesized spectrum. The target audio is an audio that expresses the target text with the target accent and the target timbre.

[0108] In the above solution, the conversion model is used to perform text encoding of the target accent on the target text, and then the obtained target text encoding vectors and the target timbre vectors corresponding to the target timbre are decoded. The target sub-spectra at several times obtained are synthesized to obtain the target synthesized spectrum corresponding to the target text, the target timbre, and the target accent. Using the target synthesized spectrum to obtain the target audio. In this way, the conversion model can generate the target synthesized spectrum with any specified timbre and any specified accent, and finally obtain the target audio.

[0109] Please refer to Figure 6 , Figure 6It is a schematic flowchart of another embodiment of the audio generation method of the present application. The method includes:

[0110] Step S610: Perform text encoding of the target accent on the target text to obtain a target text encoding vector of the target accent.

[0111] For the relevant description of step S610, reference can be made to the relevant content of step S220 above, and details will not be elaborated here.

[0112] Step S620: Perform variational auto-encoding processing on the target text encoding vector to obtain a target sentence-level encoding vector.

[0113] The conversion model in this embodiment further includes a variational auto-encoder. Step S620 can be implemented by the variational auto-encoder of the conversion model. The target text may include several phonemes, so there are several target text encoding vectors corresponding to the several phonemes. Inputting all the target text encoding vectors into the variational auto-encoder of the conversion model can obtain a sentence-level encoding vector corresponding to the target text.

[0114] It should be noted that the corpus used in the model training process may carry some other characteristic information of the speaker, such as emotion, speech rate, rhythm, stress, etc. In order to make the synthesized audio not carry other characteristics of the speaker, it is necessary to consider excluding the disturbance of the above other characteristic information during the process of synthesizing the audio. Step S620 is to achieve the exclusion of disturbance, where the target sentence-level encoding vector is the characteristic representing the disturbance, and is used to exclude the disturbance in the subsequent process of generating the spectrum, so that the target audio does not contain disturbance information. Since the target sentence-level encoding vector is obtained based on the target text encoding vector, the target sentence-level encoding vector can be understood as the synthesized disturbance characteristic.

[0115] In some embodiments, step S620 may specifically include performing dot-product self-attention processing on the target text encoding vector to obtain a target hidden vector, and then using the target hidden vector to obtain a target sentence-level encoding vector. For example, passing the target hidden vector through a mean, variance network layer and resampling to generate the target sentence-level encoding vector. Among them, the number of target text encoding vectors corresponding to the target text is indefinite, and only one target sentence-level encoding vector is generated for the target text. The above conversion can be achieved through dot-product self-attention processing.

[0116] Step S630: Perform fusion processing on the first spectral feature vector corresponding to the target sub-spectrum at the previous moment and the target timbre vector of the target timbre to obtain a first fusion vector.

[0117] For the relevant description of step S630, reference can be made to the relevant content of step S533 above, and details will not be elaborated here.

[0118] Step S640: Using the first fusion vector, the target sentence-level encoding vector, and the target text encoding vector, obtain the target sub-spectrum at the current moment.

[0119] For the relevant description of step S640, reference can be made to the relevant content of the aforementioned step S534. Specifically, step S640 may include fusing the first fusion vector and the target sentence-level encoding vector to obtain a second fusion vector, and using the second recurrent network layer of the decoder to process the second fusion vector and the target text encoding vector to obtain the target sub-spectrum at the current moment. Here, the target text encoding vector is all the target text encoding vectors corresponding to the target samples.

[0120] In the above solution, the target text is encoded with the target accent using the conversion model, and then the obtained target text encoding vector and the target timbre vector corresponding to the target timbre are decoded. The target sub-spectrums at several moments obtained are synthesized to obtain the target synthesized spectrum corresponding to the target text, target timbre, and target accent. Among them, a variational autoencoder is also used to eliminate perturbations, thereby improving the accuracy of audio generation.

[0121] Please refer to Figure 7 , Figure 7 which is a schematic flowchart of an embodiment of the training method of the conversion model of the present application.

[0122] It should be noted that the present application provides a conversion model and a timbre discriminator. Among them, the conversion model is used to generate a synthesized spectrum, and the timbre discriminator is used to discriminate the timbre of the input spectrum. The timbre discriminator can be used to discriminate the timbre of the synthesized spectrum. The conversion model and the timbre discriminator are trained adversarially. The timbre discriminator is only used to assist the conversion model in training, and the timbre discriminator is not required when the conversion model is applied.

[0123] The steps of the training method of the conversion model can be divided into the steps of training the timbre discriminator and the steps of training the conversion model. It should be noted that in the embodiments of the present application, an example of one training step is used for illustration. The device can perform multiple trainings to finally obtain a trained conversion model and / or timbre discriminator. The timbre discriminator and the conversion model can be trained in a sequential training or cross-training manner. Regardless of whether a sequential training or cross-training manner is adopted, the steps of one training step for the timbre discriminator and the conversion model will not change. Different training methods change the order of multiple trainings.

[0124] If the sequential training method is adopted, the entire training process includes multiple trainings, which can be divided into two stages. In the first stage, the training object is the timbre discriminator, and this training stage can be the step of performing multiple iterative trainings on the timbre discriminator. After the timbre discriminator is trained, the second stage begins. In the second stage, the training object is the conversion model, and this training process can be the step of performing multiple iterative trainings on the conversion model. Then, when the conversion model starts training, the timbre discriminator can be used to accurately judge the timbre of the synthesized spectrum.

[0125] If the cross-training method of the timbre discriminator and the conversion model is adopted, the entire training process includes multiple trainings. Each training trains one of the timbre discriminator and the conversion model. The difference from the sequential training method is that the samples and objects for each training can be randomly selected, and the objects for two adjacent trainings can be the same or different. Only the parameters of the training object for that training will be adjusted in one training. Then the device can select the sample text and / or the sample real spectrum as the training sample for this time. If the training sample for this time includes the sample text, it can be determined that the training object for this time is the conversion model, and the conversion model is trained using the sample text. If the training sample for this time does not include the sample text, it can be determined that the training object for this time is the timbre discriminator, and the timbre discriminator is trained using the sample real spectrum. It should be noted that when the conversion model is trained once using the sample text, the training steps can also include using the timbre discriminator to judge the timbre of the sample synthesized spectrum, so as to adjust the parameters of the conversion model using the discrimination result and the sample timbre. Since the training object for this time is the conversion model, the parameters of the timbre discriminator are not adjusted. In addition, before cross-training the timbre discriminator and the conversion model, the timbre discriminator can be pre-trained using the sample real spectrum to obtain a timbre discriminator with a certain timbre discrimination accuracy.

[0126] The following takes the case where the timbre discriminator has been trained before training the conversion model as an example for illustration. Among them, step S710 and step S720 are the steps of training the timbre discriminator once, and step S730-step S750 are the steps of training the conversion model once. The method includes:

[0127] Step S710: Use the timbre discriminator to perform timbre discrimination on the sample real spectrum to obtain the second predicted timbre.

[0128] The device can pre-store the real spectrum and the text corresponding to the real spectrum as the samples for model training. Among them, the real spectrum is collected in reality and converted from the audio emitted by the speaker, and the actual timbre and accent of the sample real spectrum are known.

[0129] Specifically, input the sample true spectrum into the timbre discriminator, and the timbre discriminator outputs the second predicted timbre as the discrimination result.

[0130] Step S720: Based on the actual timbre of the sample true spectrum and the second predicted timbre, adjust the parameters of the timbre discriminator.

[0131] The actual timbre of the sample true spectrum is known. After obtaining the second predicted timbre output by the timbre discriminator, if there is a difference between the second predicted timbre and the actual timbre, the parameters of the timbre discriminator can be adjusted using a loss function. Specifically, the loss function of the timbre discriminator can be calculated by the following formula:

[0132] L D =-E x (y x logD(x)) (Formula 4)

[0133] where L D is the loss function of the timbre discriminator, D represents the timbre discriminator, x is the sample true spectrum, and y x is the timbre category of the actual timbre corresponding to the sample true spectrum.

[0134] By adjusting the parameters in the timbre discriminator, the timbre discriminator can more accurately judge the timbre of the spectrum input into it. The device can execute Step S710 and Step S720 multiple times. After the discrimination accuracy rate of the timbre discriminator is greater than the preset value, it can be considered that the training of the timbre discriminator is completed. Then, the timbre discriminator can be used to judge whether the sample synthetic spectrum generated by the conversion model is consistent with the sample timbre.

[0135] Step S730: Input the sample text, the identifier of the sample accent, and the identifier of the sample timbre into the conversion model to obtain the sample synthetic spectrum corresponding to the sample accent and the sample timbre.

[0136] Among them, the sample text, the sample accent, and the sample timbre can be the text, accent, and timbre corresponding to the same true spectrum. The conversion model includes a text encoder and a decoder. Step S730 can be implemented through the text encoder and decoder of the conversion model. The relevant content of Step S730 can refer to the relevant description of the foregoing embodiment of the audio generation method.

[0137] Please refer to Figure 7 and Figure 8 , Figure 8 which is a schematic flowchart of another embodiment of Step S730 of this application. Step S730 includes:

[0138] Step S831: Use the text encoder of the conversion model to perform text encoding of the sample accent on the sample text to obtain the sample text encoding vector of the sample accent.

[0139] For the relevant description of step S831, reference can be made to the foregoing relevant content about step S220, which will not be elaborated here.

[0140] It should be noted that the samples used in the process of model training may carry some other characteristic information of the speaker, such as emotion, speech rate, rhythm, stress, etc. Different samples may contain contradictory other characteristic information. For example, some samples carry an angry emotion, and some samples carry an excited emotion. Then, training the conversion model with samples having contradictory characteristics will cause the conversion model to be unable to determine which characteristic among the contradictory characteristics the output spectrum should carry, which will lead to errors or crashes of the conversion model. To avoid the above problems, a variational autoencoder is provided in this application. Through the variational autoencoder, other characteristic information in the input spectrum / text encoding vector can be extracted, that is, the real perturbation feature / synthetic perturbation feature. In the subsequent decoding process, the extracted perturbation features can be removed, so that the output of the model during training is not affected by the possibly contradictory other characteristic information, reducing the possibility of the model making errors or crashing and making the model more stable.

[0141] It should be noted that steps S832 - S833 are optional. If the conversion model also includes a reference variational autoencoder and a variational autoencoder of the conversion model, then steps S832 - S834 can be executed to solve the above problem of model instability.

[0142] Step S832: Perform variational autoencoding on the sample text encoding vector using the variational autoencoder of the conversion model to obtain a sample sentence-level encoding vector.

[0143] It should be noted that the generation process of the sample sentence-level encoding vector in step S832 is the same as that of the target sentence-level encoding vector in the audio generation process.

[0144] Specifically, step S832 may include performing dot-product self-attention processing on the sample text encoding vector using the variational autoencoder of the conversion model to obtain a first hidden vector, and then obtaining a sample sentence-level encoding vector based on the first hidden vector.

[0145] The order of steps S832 and S833 can be swapped, and steps S835 and S835 can be executed after step S833 is completed.

[0146] Step S833: Perform variational autoencoding on the real spectrum corresponding to the sample text using the reference variational autoencoder to obtain a reference sentence-level encoding vector.

[0147] The reference sentence-level encoding vector obtained by inputting the true spectrum corresponding to the sample text, sample accent, and sample timbre into the reference variational autoencoder can reflect the characteristics of the perturbations in the true spectrum. Thus, in the subsequent generation process, the characteristics of the perturbations in the true spectrum are removed, so that the output of the conversion model is not affected by the characteristics of the perturbations and the output is more stable.

[0148] Please refer to Figure 7 , Figure 8 and Figure 9 , Figure 9 which is a schematic flowchart of another embodiment of step S833 of this application. Step S833 includes:

[0149] Step S9331: Use several second convolutional groups of the reference variational autoencoder to respectively extract features from the true spectrum to obtain a sequence of feature vectors corresponding to each second convolutional group.

[0150] It should be noted that the reference variational autoencoder includes several second convolutional groups, a fully connected layer, and mean and variance network layers. Among them, the second convolutional group includes a one-dimensional convolutional layer, a ReLU activation function, and a batch normalization layer connected in sequence. Among them, the ReLU activation function and the batch normalization layer are used to improve the stability of the model.

[0151] Among them, several second convolutional groups are in a parallel relationship and are all used to extract features from the true spectrum entering the reference variational autoencoder. Taking one second convolutional group as an example, the true spectrum includes sub-spectrums at several moments. Use this group of second convolutional groups to sequentially extract features from several sub-spectrums to obtain the feature vectors corresponding to each moment of the true spectrum, that is, the sequence of feature vectors corresponding to this second convolutional group. The processing process of each second convolutional group is similar, so that a sequence of feature vectors corresponding to each second convolutional group can be obtained.

[0152] Taking the example that the true spectrum includes sub-spectrums at two moments for a simple illustration, a second convolutional group sequentially extracts features from the sub-spectrums at these two moments, and respectively obtains the feature vector a (a 1 , a 2 , a 3 ) of the sub-spectrum at the first moment and the feature vector b (b 1 , b 2 , b 3 ) of the sub-spectrum at the second moment. The vector a and the vector b form the sequence of feature vectors corresponding to this second convolutional group.

[0153] Step S9332: Statistically analyze each feature vector in the sequence of feature vectors to obtain a first statistical vector and a second statistical vector.

[0154] It can be understood that the operation in step S9332 is to perform statistics on all the feature vectors corresponding to a second convolutional group at once. The channels of all the feature vectors in a feature vector sequence are correspondingly the same. By performing statistics on the element values of the same channel of all the feature vectors in this sequence, a first statistical vector and a second statistical vector are obtained. The elements included in the first statistical vector are the first statistical values corresponding to each channel, and the elements included in the second statistical vector are the second statistical values corresponding to each channel. Among them, the first statistical value can reflect the central tendency of the element values of each channel, and the second statistical value can reflect the degree of dispersion of the element values of each channel.

[0155] Continuing with the above example where the true spectrum includes sub-spectrums at two moments, the first statistical value is the average value, and the second statistical value is the variance. Among them, a 1 and b 1 correspond to the same channel, a 2 and b 2 correspond to the same channel, a 3 and b 3 correspond to the same channel. Respectively, the first statistical values of a 1 and b 1 , a 2 and b 2 , a 3 and b 3 are calculated, and thus the three average values corresponding to the three channels are obtained. These three average values form the first statistical vector. Respectively, the second statistical values of a 1 and b 1 , a 2 and b 2 , a 3 and b 3 are calculated, which are the variances. Thus, the three variances corresponding to the three channels are obtained. These three variances form the second statistical vector.

[0156] Step S9333: Concatenate the first statistical vector and the second statistical vector to obtain a concatenated vector.

[0157] It should be noted that since the durations corresponding to different true spectra are different, the number of sub-spectra included in different true spectra is also different. By processing the features of the true spectra in the above manner, it can be ensured that no matter what duration of true spectrum is input into the reference variational autoencoder, the output result is a sentence-level encoded vector, and thus the features of spectra with different time lengths are unified in size.

[0158] Step S9334: Perform dimensionality reduction processing on the concatenated vector to obtain a second hidden vector.

[0159] It should be noted that step S9334 can be implemented by referring to the fully connected layer of the variational autoencoder.

[0160] Step S9335: Based on the second latent vector, obtain the reference sentence-level encoding vector.

[0161] It should be noted that the operations of the reference variational autoencoder and the variational autoencoder of the conversion model to process the latent vector to obtain the reference sentence-level encoding vector can be the same, both implemented through the mean and variance network layers.

[0162] Step S834: Use the decoder of the conversion model to perform decoding processing on the sample text encoding vector and the sample timbre vector corresponding to the sample timbre to obtain the sample sub-spectrums at several moments corresponding to the sample timbre. It should be noted that since the characteristics of the perturbation are brought by the real spectrum, during the training process, in order to make the model generate a sample synthesis spectrum without perturbation more accurately, the sample sentence-level encoding vector output by the reference variational autoencoder is used to replace the sample sentence-level encoding vector output by the variational autoencoder of the conversion model, that is, the real perturbation characteristics are removed during decoding instead of the synthesized perturbation characteristics.

[0163] That is, when the conversion model includes a variational autoencoder, step S834 can specifically be to use the decoder of the conversion model to perform decoding processing on the sample text encoding vector, the reference sentence-level encoding vector, and the sample timbre vector to obtain the sample sub-spectrums at several moments corresponding to the sample timbre.

[0164] When using the conversion model to generate the target synthesis spectrum, there is no real spectrum that can be used as a reference at this time, but it is still necessary to remove the perturbation characteristics from the spectrum output by the conversion model. Then, only the target sentence-level encoding vector output by the variational autoencoding vector of the conversion model can be used, that is, the synthesized perturbation characteristics are used to replace the real perturbation characteristics, so as to achieve the purpose of removing the perturbation characteristics.

[0165] It should be noted that the process of decoding to obtain the sub-spectrum at the current moment also includes processing the sub-spectrum at the previous moment. The relevant descriptions of this part can refer to the relevant content in the foregoing embodiments of audio generation. In the foregoing embodiments of using the conversion model to generate the target sub-spectrum, the target sub-spectrum at the previous moment is input to the decoder. When training the conversion model, since the sample spectrum corresponds to a known real spectrum, during the process of decoding to obtain the sample sub-spectrum at the current moment, the real spectrum at the previous moment or the sample sub-spectrum at the previous moment can be input to the decoder.

[0166] Step S835: Use the conversion model to perform synthesis processing on the sample sub-spectrums at several moments to obtain the sample synthesis spectrum corresponding to the sample text.

[0167] Among them, the synthesis process may be to splice the sample sub-spectrums at several moments in sequence.

[0168] Step S740: Use a timbre discriminator to perform timbre discrimination on the sample synthesized spectrum to obtain a first predicted timbre.

[0169] It should be noted that the process of using the timbre discriminator to perform timbre discrimination on the sample true spectrum is the same as the process of performing timbre discrimination on the sample synthesized spectrum. Here, the process of performing timbre discrimination on the sample synthesized spectrum is taken as an example for illustration.

[0170] Please refer to Figure 7 and Figure 10 , Figure 10 which is a schematic flowchart of another embodiment of step S740 of this application. Step S740 includes:

[0171] Step S1041: Use several first convolutional groups of the timbre discriminator to perform feature extraction on the sample synthesized spectrum respectively to obtain spectrum features corresponding to each first convolutional group.

[0172] It should be noted that the timbre discriminator includes several first convolutional groups, a recurrent network layer, and a Softmax layer. Among them, each first convolutional group includes a two-dimensional convolutional layer, a ReLU activation function, and a batch normalization layer connected in sequence. The several first convolutional groups are in a parallel relationship and all perform feature extraction on the spectrum input to the timbre discriminator, and each first convolutional group correspondingly obtains spectrum features.

[0173] Step S1042: Use the recurrent network layer of the timbre discriminator to process the spectrum features to obtain a prediction result.

[0174] It should be noted that the recurrent network layer of the timbre discriminator will process the sample sub-spectrums at each moment in sequence to obtain a prediction result. The prediction result includes the first probability that the sample sub-spectrums at each moment in the sample synthesized spectrum belong to each preset timbre.

[0175] Step S1043: Statistically analyze the first probabilities that the sample sub-spectrums at each moment belong to each preset timbre to determine the second probability that the sample synthesized spectrum belongs to each preset timbre.

[0176] Due to the characteristic of the recurrent network layer to retain historical processing information, it can be considered that the prediction result corresponding to the last moment output by the recurrent network layer also represents the first probability that the sample sub-spectrums at each moment belong to each preset timbre. Therefore, the Softmax layer can be used to determine the second probability that the sample synthesized spectrum belongs to each preset timbre. Specifically, the Softmax layer can map the real number space of the recurrent network layer to a probability space, thereby determining the second probability that the sample synthesized spectrum belongs to each preset timbre.

[0177] Step S1044: Select a preset timbre whose second probability meets the preset condition as the first predicted timbre.

[0178] The preset condition can be that the second probability is the highest, which means that the spectrum input to the timbre discriminator is most likely to be the preset timbre. Then, this preset timbre can be used as the result of this timbre discrimination, that is, the first predicted timbre. Specifically, the timbre with the highest probability value is output by the Softmax layer.

[0179] Step S750: Based on the sample timbre and the first predicted timbre, adjust the parameters of the conversion model.

[0180] It can be understood that if there is a difference between the sample timbre and the first predicted timbre, the parameters of the conversion model can be adjusted using the loss function. The parameters of the conversion model adjusted here can include the parameters of the text encoder and decoder, and can also include the parameters of the reference variational autoencoder and the variational autoencoder of the conversion model.

[0181] Among them, the loss function in step S750 can be calculated by the following formula:

[0182]

[0183] Among them, represents the additional timbre error loss of the conversion model, D represents the timbre discriminator, x' is the sample synthetic spectrum generated according to the sample timbre, and y x' is the timbre category corresponding to the sample timbre.

[0184] In some embodiments, during the training of the conversion model, the parameters of the conversion model can also be adjusted based on the difference between the sample synthetic spectrum and the real spectrum corresponding to the sample text, sample accent, and sample timbre, that is, the reconstruction loss.

[0185] Step S760: Use the difference between the sample sentence-level encoding vector and the reference sentence-level encoding vector to adjust the parameters of the variational autoencoder.

[0186] It should be noted that step S760 is an optional step. If the conversion model includes a variational autoencoder, then step S760 is executed. If the conversion model does not include a variational autoencoder, then step S760 does not need to be executed.

[0187] In some embodiments, the variational autoencoder can be a part of the decoder, so that steps S750 and S760 can be executed as one step.

[0188] When training the conversion model, during the process of using the conversion model to generate the sample synthetic frequency spectrum, the sample sentence-level encoding vector output by the variational autoencoder of the conversion model is not used. Instead, during the process of using the conversion model to generate the target synthetic frequency spectrum, only the target sentence-level encoding vector output by the variational autoencoder of the conversion model can be used. To make the target sentence-level encoding vector output by the variational autoencoder of the conversion model close to the true perturbation features during the subsequent process of using the conversion model to generate the target synthetic frequency spectrum, during the training process, the differences between the sample sentence-level encoding vector and the reference sentence-level encoding vector can be utilized to adjust the parameters of the reference variational autoencoder and the variational autoencoder of the conversion model. This makes the sample sentence-level encoding vector and the reference sentence-level encoding vector tend to be consistent for the same sample, that is, the synthetic perturbation features obtained through the variational autoencoder of the conversion model tend to be consistent with the true perturbation features of the true frequency spectrum obtained through the reference variational autoencoder.

[0189] The device can execute step S730 - step S760 multiple times to complete the training of the conversion model.

[0190] In some embodiments, the timbre discriminator and the conversion model are cross-trained. It can be understood that the training process can be divided into multiple rounds. In the first several rounds of training, the conversion model is trained using the sample text, and the timbre discriminator is trained using the sample true frequency spectrum and the actual timbre of the sample true frequency spectrum. If the timbre discriminator makes a wrong judgment on the timbre of the sample true frequency spectrum, the loss function is used to correct it. After several rounds of training, the timbre discriminator is cross-trained using the sample true frequency spectrum and the sample synthetic frequency spectrum obtained using the conversion model. During this process, if the timbre judgment on the sample synthetic frequency spectrum is wrong, the loss is applied to the timbre conversion model; if the timbre judgment on the sample true frequency spectrum is wrong, the loss is applied to the timbre discriminator. That is, for the first several rounds of training, each round of training executes the steps of training the timbre discriminator once, that is, step S710 and step S720. After several rounds of training are completed, each round of training executes step S710 and step S720 or step S740 - step S760 until the training of the conversion model is completed.

[0191] After completing the training of the conversion model, the target synthetic audio can be generated using the conversion model. The relevant steps for generating the target synthetic audio can refer to the relevant embodiments of the foregoing audio generation method and will not be elaborated here.

[0192] In the above solution, the sample text, the identifier of the sample accent, and the identifier of the sample timbre are input into the conversion model to obtain a sample synthesis spectrum corresponding to the sample accent and the sample timbre. The timbre discriminator is used to train and discriminate the timbre of the sample synthesis spectrum. Based on the sample timbre and the first predicted timbre, the parameters of the conversion model are adjusted to obtain a conversion model capable of generating a conversion model corresponding to the specified timbre and the specified accent. The conversion model is used to perform text encoding of the target accent on the target text, and then decoding processing is performed on the obtained target text encoding vector and the target timbre vector corresponding to the target timbre. The obtained target sub-spectra at several moments are synthesized to obtain a target synthesis spectrum corresponding to the target text, the target timbre, and the target accent. Among them, a variational autoencoder is also used to eliminate perturbations. In the above manner, by using the timbre discriminator to assist in training the model, it can be ensured that the timbre of the synthesis spectrum generated by the conversion model is consistent with the specified timbre, improving the accuracy of the model. And finally, the conversion model can generate a target synthesis spectrum with any specified timbre and any specified accent, and using the variational autoencoder can exclude the influence of other feature information of the sample on the model, improving the stability of the model.

[0193] Please refer to Figure 11 , Figure 11 which is a schematic framework diagram of an embodiment of the training device of the conversion model of the present application.

[0194] In this embodiment, the training device 110 of the conversion model includes an input module 111, a timbre discrimination module 112, and a model adjustment module 113. Among them, the input module 111 is used to input the sample text, the identifier of the sample accent, and the identifier of the sample timbre into the conversion model to obtain a sample synthesis spectrum corresponding to the sample accent and the sample timbre. The timbre discrimination module 112 is used to use the timbre discriminator to discriminate the timbre of the sample synthesis spectrum to obtain the first predicted timbre. The model adjustment module 113 is used to adjust the parameters of the conversion model based on the sample timbre and the first predicted timbre.

[0195] In the above solution, the sample text, the identifier of the sample accent, and the identifier of the sample timbre are input into the conversion model to obtain a sample synthesis spectrum corresponding to the sample accent and the sample timbre. The timbre discriminator is used to train and discriminate the timbre of the sample synthesis spectrum. Based on the sample timbre and the first predicted timbre, the parameters of the conversion model are adjusted to obtain a conversion model capable of generating a conversion model corresponding to the specified timbre and the specified accent. In the above manner, by using the timbre discriminator to assist in training the model, it can be ensured that the timbre of the synthesis spectrum generated by the conversion model is consistent with the specified timbre, improving the accuracy of the model.

[0196] Please refer to Figure 12 , Figure 12 which is a schematic framework diagram of an embodiment of the audio generation device of the present application.

[0197] In this embodiment, the audio generation device 120 includes a text encoding module 121, a decoding module 122, and a synthesis module 123. Among them, the text encoding module 121 is configured to perform text encoding of a target accent on the target text to obtain a target text encoding vector of the target accent. The decoding module 122 is configured to perform decoding processing on the target text encoding vector and a target timbre vector corresponding to the target timbre to obtain target sub-spectra at several moments corresponding to the target timbre. The synthesis module 123 is configured to perform synthesis processing on the target sub-spectra at several moments to obtain a target synthesis spectrum corresponding to the target text.

[0198] In the above solution, the sample text, the identifier of the sample accent, and the identifier of the sample timbre are input into the conversion model to obtain a sample synthesis spectrum corresponding to the sample accent and the sample timbre. The timbre discriminator is used to train and discriminate the timbre of the sample synthesis spectrum. Based on the sample timbre and the first predicted timbre, the parameters of the conversion model are adjusted to obtain a conversion model that can generate a conversion model corresponding to the specified timbre and the specified accent. The conversion model is used to perform text encoding of the target accent on the target text, and then decoding processing is performed on the obtained target text encoding vector and the target timbre vector corresponding to the target timbre. The obtained target sub-spectra at several moments are synthesized to obtain a target synthesis spectrum corresponding to the target text, the target timbre, and the target accent. Through the above method, the timbre discriminator is used to assist the model training, which can make the timbre of the synthesis spectrum generated by the conversion model consistent with the specified timbre, improve the accuracy of the model, and finally the conversion model can generate a target synthesis spectrum with any specified timbre and any specified accent.

[0199] Please refer to Figure 13 , Figure 13 which is a schematic framework diagram of an embodiment of the electronic device of the present application.

[0200] In this embodiment, the electronic device 130 includes a memory 131 and a processor 132, where the memory 131 is coupled to the processor 132. Specifically, the various components of the electronic device 130 can be coupled together through a bus, or the processor 132 of the electronic device 130 is respectively connected to other components one by one. The electronic device 130 can be any device with processing capabilities, such as a computer, a tablet computer, a mobile phone, etc.

[0201] The memory 131 is used to store program data executed by the processor 132 and data during the processing of the processor 132. For example, the conversion model, the target text, the target timbre, etc. Among them, the memory 131 includes a non-volatile storage part for storing the above program data.

[0202] The processor 132 controls the operation of the electronic device 130. The processor 132 may also be referred to as a CPU (Central Processing Unit). The processor 132 may be an integrated circuit chip with the ability to process signals. The processor 132 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 132 may be implemented jointly by multiple integrated circuit chips.

[0203] The processor 132 is used to execute instructions to implement the training method of any of the above conversion models or any of the audio generation methods by calling the program data stored in the memory 131.

[0204] For example, the processor 132 may perform text encoding of the target text with the target accent to obtain a target text encoding vector, and perform decoding processing on the target text encoding vector and the target timbre vector to obtain the target sub-spectrum at several moments corresponding to the target timbre.

[0205] In the above solution, the sample text, the identifier of the sample accent, and the identifier of the sample timbre are input into the conversion model to obtain a sample synthesis spectrum corresponding to the sample accent and the sample timbre. The timbre discriminator is used to train and discriminate the timbre of the sample synthesis spectrum. Based on the sample timbre and the first predicted timbre, the parameters of the conversion model are adjusted to obtain a conversion model capable of generating a conversion model corresponding to the specified timbre and the specified accent. The conversion model is used to perform text encoding of the target text with the target accent, and then decoding processing is performed on the obtained target text encoding vector and the target timbre vector corresponding to the target timbre. The obtained target sub-spectrum at several moments is synthesized to obtain a target synthesis spectrum corresponding to the target text, the target timbre, and the target accent. Through the above method, using the timbre discriminator to assist in training the model can make the timbre of the synthesis spectrum generated by the conversion model consistent with the specified timbre, improve the accuracy of the model, and finally the conversion model can generate a target synthesis spectrum with any specified timbre and any specified accent.

[0206] Please refer to Figure 14 , Figure 14 which is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application.

[0207] In this embodiment, the computer-readable storage medium 140 stores program data 141 that can be run by the processor. The program data can be executed to implement the training method of any of the above conversion models or any of the audio generation methods.

[0208] The computer-readable storage medium 140 may specifically be a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc., which are media that can store program data. Or it may also be a server storing the program data. The server can send the stored program data to other devices for running, or it can also run the stored program data itself.

[0209] In some embodiments, the computer-readable storage medium 140 may also be a memory as Figure 13 shown.

[0210] In the above solution, the target text encoding vector of the target accent is obtained by text-encoding the target text, and the target text encoding vector and the target timbre are decoded to obtain the spectrum corresponding to the target text. Therefore, without relying on the original audio, the spectrum of the audio with the content of the target text, the target accent, and the target timbre can be directly obtained depending on the text. Thus, it is possible to generate an audio with a specified timbre and a specified accent by using the text. Moreover, since there is no need to rely on the original audio, the problem that the generated audio carries the accent or timbre of the original audio can be avoided, and the accuracy of generating an audio with a specific accent and a specific timbre can be improved.

[0211] In addition, in the above solution, the sample text, the identifier of the sample accent, and the identifier of the sample timbre can also be input into the conversion model to obtain the sample synthesis spectrum corresponding to the sample accent and the sample timbre. The timbre discriminator is used to train and discriminate the timbre of the sample synthesis spectrum. Based on the sample timbre and the first predicted timbre, the parameters of the conversion model are adjusted to obtain a conversion model that can generate a conversion model corresponding to the specified timbre and the specified accent. The target text is text-encoded with the target accent by using the conversion model, and then the obtained target text encoding vector and the target timbre vector corresponding to the target timbre are decoded, and the obtained target sub-spectra at several moments are synthesized to obtain the target synthesis spectrum corresponding to the target text, the target timbre, and the target accent. By the above method, using the timbre discriminator to assist in training the model can make the timbre of the synthesis spectrum generated by the conversion model tend to be consistent with the specified timbre, improve the accuracy of model conversion, and finally the conversion model can generate the target synthesis spectrum with any specified timbre and any specified accent.

[0212] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present application by the same token.

Claims

1. A method for training a conversion model, characterized in that, it includes the following steps to train the conversion model: Input the sample text, the identifier of the sample accent, and the identifier of the sample timbre into the conversion model to obtain a sample synthesis spectrum corresponding to the sample accent and sample timbre; Use a timbre discriminator to perform timbre discrimination on the sample synthesis spectrum to obtain a first predicted timbre; Based on the sample timbre and the first predicted timbre, adjust the parameters of the conversion model; The trained conversion model is used to generate a target synthesis spectrum based on a target text, a target accent, and a target timbre vector corresponding to the target timbre, and the target synthesis spectrum is used to generate a target audio with the target accent and the target timbre.

2. The method according to claim 1, characterized in that, the timbre discriminator is trained and completed before training the conversion model; or, the timbre discriminator and the conversion model are cross-trained.

3. The method according to claim 2, characterized in that, when the timbre discriminator and the conversion model are cross-trained, the method includes: Select the sample text and / or the sample real spectrum as the current training sample; If the current training sample includes the sample text, use the sample text to perform the training on the conversion model; If the current training sample does not include the sample text, use the sample real spectrum to train the timbre discriminator.

4. The method according to claim 2 or 3, characterized in that, the training steps of the timbre discriminator include: Use the timbre discriminator to perform timbre discrimination on the sample real spectrum to obtain a second predicted timbre; Based on the actual timbre of the sample real spectrum and the second predicted timbre, adjust the parameters of the timbre discriminator.

5. The method according to claim 1, characterized in that, the using the timbre discriminator to perform timbre discrimination on the sample synthesis spectrum to obtain a first predicted timbre includes: Use several first convolutional groups of the timbre discriminator to perform feature extraction on the sample synthesis spectrum respectively to obtain spectrum features corresponding to each of the first convolutional groups; Process the spectrum features through the recurrent network layer of the timbre discriminator to obtain a prediction result, where the prediction result includes the first probability that each sample sub-spectrum at each moment in the sample synthesis spectrum belongs to each preset timbre; Statistically analyze the first probabilities that each sample sub-spectrum at each moment belongs to each preset timbre to determine the second probability that the sample synthesis spectrum belongs to each of the preset timbres; Select the preset timbre whose second probability meets the preset condition as the first predicted timbre.

6. The method according to claim 1, characterized in that, the inputting the sample text, the identifier of the sample accent, and the identifier of the sample timbre into the conversion model to obtain a sample synthesis spectrum corresponding to the sample accent and sample timbre includes: Use the text encoder of the conversion model to perform text encoding of the sample accent on the sample text to obtain a sample text encoding vector of the sample accent; The decoder of the conversion model performs decoding processing based at least on the sample text encoding vector and the sample timbre vector corresponding to the sample timbre, so as to obtain sample sub-spectrums at a plurality of moments corresponding to the sample timbre; Performing synthesis processing on the sample sub-spectrums at the plurality of moments by using the conversion model to obtain a sample synthesis spectrum corresponding to the sample text.

7. The method according to claim 6, wherein, The inputting the sample text, the identifier of the sample accent, and the identifier of the sample timbre into the conversion model to obtain a sample synthesis spectrum corresponding to the sample accent and the sample timbre further includes: Performing variational auto-encoding on the sample text encoding vector by using the variational auto-encoder of the conversion model to obtain a sample sentence-level encoding vector; Performing variational auto-encoding on the sample true spectrum corresponding to the sample text by using a reference variational auto-encoder to obtain a reference sentence-level encoding vector; The using the decoder of the conversion model to perform decoding processing based at least on the sample text encoding vector and the sample timbre vector corresponding to the sample timbre so as to obtain sample sub-spectrums at a plurality of moments corresponding to the sample timbre includes: Using the decoder of the conversion model to perform decoding processing on the sample text encoding vector, the reference sentence-level encoding vector, and the sample timbre vector to obtain sample sub-spectrums at a plurality of moments corresponding to the sample timbre; The method further includes: Adjusting the parameters of the variational auto-encoder of the conversion model and the reference variational auto-encoder by using the difference between the sample sentence-level encoding vector and the reference sentence-level encoding vector.

8. The method according to claim 7, wherein, The using the variational auto-encoder of the conversion model to perform variational auto-encoding on the sample text encoding vector to obtain a sample sentence-level encoding vector includes: Performing dot-product self-attention processing on the sample text encoding vector by using the variational auto-encoder of the conversion model to obtain a first hidden vector; Based on the first hidden vector, obtaining the sample sentence-level encoding vector; And / or, the using the reference variational auto-encoder to perform variational auto-encoding on the true spectrum corresponding to the sample text to obtain a reference sentence-level encoding vector includes: Using a plurality of second convolutional groups of the reference variational auto-encoder to respectively perform feature extraction on the true spectrum to obtain a sequence of feature vectors corresponding to each of the second convolutional groups, and each sequence of feature vectors includes feature vectors corresponding to each moment of the true spectrum; Statistically analyzing each feature vector in the sequence of feature vectors to obtain a first statistical vector and a second statistical vector; Concatenating the first statistical vector and the second statistical vector to obtain a concatenated vector; Performing dimensionality reduction processing on the concatenated vector to obtain a second hidden vector; Based on the second hidden vector, obtaining the reference sentence-level encoding vector.

9. An audio generation method, wherein, includes: Performing text encoding of a target accent on a target text to obtain a target text encoding vector of the target accent; Decode the target text encoding vector and the target timbre vector corresponding to the target timbre to obtain target sub-spectrums at several moments corresponding to the target timbre; Perform synthesis processing on the target sub-spectrums at several moments to obtain a target synthesis spectrum corresponding to the target text; Generate a target audio with the target accent and target timbre based on the target synthesis spectrum.

10. The method according to claim 9, wherein, The encoding the target text with the target accent to obtain the target text encoding vector with the target accent includes: Based on the accent identifier of the target accent, select the text encoder corresponding to the target accent; Use the text encoder corresponding to the target accent to perform text encoding of the target accent on the target text to obtain the target text encoding vector with the target accent.

11. The method according to claim 9 or 10, wherein, The encoding the target text with the target accent to obtain the target text encoding vector with the target accent includes: Obtain first original phoneme vectors of several phonemes corresponding to the target text; Perform timbre removal processing on the first original phoneme vectors of each phoneme to obtain target phoneme vectors of each phoneme; Encode the target phoneme vectors of the several phonemes to obtain the target text encoding vector with the target accent.

12. The method according to claim 11, wherein, Before the obtaining the first original phoneme vectors of several phonemes corresponding to the target text, the method further includes: Convert the target text into several phonemes; The obtaining the first original phoneme vectors of several phonemes corresponding to the target text includes: Use the word embedding layer of the text encoder to obtain second original phoneme vectors of the several phonemes; Use the fully connected layer of the text encoder to process the second original phoneme vectors of each phoneme to obtain the first original phoneme vectors of each phoneme; The encoding the target phoneme vectors of the several phonemes to obtain the target text encoding vector with the target accent includes: Use the first recurrent network layer of the text encoder to perform encoding processing on the target phoneme vectors of the several phonemes to obtain the target text encoding vector with the target accent.

13. The method according to claim 9, wherein, The decoding the target text encoding vector and the target timbre vector corresponding to the target timbre to obtain target sub-spectrums at several moments corresponding to the target timbre includes: Perform fusion processing on the first spectrum feature vector corresponding to the target sub-spectrum at the previous moment and the target timbre vector to obtain a first fusion vector; Use the first fusion vector and the target text encoding vector to obtain the target sub-spectrum at the current moment.

14. The method according to claim 13, wherein, The fusion processing is element-wise addition; and / or, The decoding the target text encoding vector and the target timbre vector corresponding to the target timbre to obtain target sub-spectrums at several moments corresponding to the target timbre further includes: Extract features from the target sub-spectrum at the previous moment to obtain a second spectral feature vector; Perform timbre removal processing on the second spectral feature vector to obtain the first spectral feature vector.

15. The method according to claim 11 or 14, wherein, Performing timbre removal processing on the first original phoneme vector of each phoneme to obtain the target phoneme vector of each phoneme, or performing timbre removal processing on the second spectral feature vector to obtain the first spectral feature vector, includes: Taking the first original phoneme vector as the vector to be processed and the target phoneme vector as the target vector, or taking the second spectral feature vector as the vector to be processed and the first spectral feature vector as the target vector; Performing layer normalization processing on the vector to be processed to obtain the target vector.

16. The method according to claim 15, wherein, The performing layer normalization processing on the vector to be processed to obtain the target vector includes: Obtaining a first statistical value and a second statistical value of the vector to be processed, wherein the first statistical value is used to reflect the central tendency of each element in the vector to be processed, and the second statistical value is used to reflect the degree of dispersion of each element in the vector to be processed; Using the vector to be processed, the first statistical value and the second statistical value to obtain an intermediate vector, and using an activation function to perform activation processing on the intermediate vector to obtain the target vector.

17. The method according to claim 13, wherein, After performing text encoding of the target accent on the target text to obtain the target text encoding vector of the target accent, the method further includes: Performing variational auto-encoding processing on the target text encoding vector to obtain a target sentence-level encoding vector; The using the first fusion vector and the target text encoding vector to obtain the target sub-spectrum at the current moment includes: Using the first fusion vector, the target sentence-level encoding vector and the target text encoding vector to obtain the target sub-spectrum at the current moment.

18. The method according to claim 17, wherein, The performing variational auto-encoding processing on the target text encoding vector to obtain a target sentence-level encoding vector includes: Performing dot-product self-attention processing on the target text encoding vector to obtain a target hidden vector; Using the target hidden vector to obtain the target sentence-level encoding vector; And / or, the using the first fusion vector, the target sentence-level encoding vector and the target text encoding vector to obtain the target spectrum at the current moment includes: Fusing the first fusion vector and the target sentence-level encoding vector to obtain a second fusion vector; Using the second recurrent network layer of the decoder to process the second fusion vector and the target text encoding vector to obtain the target sub-spectrum at the current moment.

19. The method according to claim 9, wherein, The performing synthesis processing on the target sub-spectra at the several moments to obtain the target synthesis spectrum corresponding to the target text includes: Stitching the target sub-spectra at the several moments to obtain the target synthesis spectrum.

20. The method according to claim 9, wherein, the steps of performing text encoding of the target text with the target accent to obtain the target text encoding vector with the target accent to performing synthesis processing on the target spectrograms at the several moments to obtain the target synthesis spectrogram corresponding to the target text are executed by a conversion model; the method further includes: training the conversion model by using the method according to any one of claims 1 to 8.

21. A training device for a conversion model, wherein, it includes: an input module, configured to input a sample text, an identifier of a sample accent, and an identifier of a sample timbre into the conversion model to obtain a sample synthesis spectrogram corresponding to the sample accent and the sample timbre; a timbre discrimination module, configured to perform timbre discrimination on the sample synthesis spectrogram by using a timbre discriminator to obtain a first predicted timbre; a model adjustment module, configured to adjust parameters of the conversion model based on the sample timbre and the first predicted timbre; the trained conversion model is configured to generate a target synthesis spectrogram based on a target text, a target accent, and a target timbre vector corresponding to the target timbre, and the target synthesis spectrogram is used to generate a target audio having the target accent and the target timbre.

22. An audio generation device, wherein, it includes: a text encoding module, configured to perform text encoding of the target text with the target accent to obtain the target text encoding vector with the target accent; a decoding module, configured to perform decoding processing on the target text encoding vector and the target timbre vector corresponding to the target timbre to obtain target sub-spectrograms at several moments corresponding to the target timbre; a synthesis module, configured to perform synthesis processing on the target sub-spectrograms at the several moments to obtain the target synthesis spectrogram corresponding to the target text; and generate a target audio having the target accent and the target timbre based on the target synthesis spectrogram.

23. An electronic device, wherein, the device includes a processor and a memory, the memory is configured to store program data, and the processor is configured to execute the program data to implement the method according to any one of claims 1-8 or claims 9-20.

24. A computer-readable storage medium storing program instructions that can be run by a processor, wherein, when the program instructions are executed by the processor, the method according to any one of claims 1-8 or claims 9-20 is implemented.

Citation Information

Patent Citations

  • System and method for voice-to-voice conversion

    CN111201565A

  • Speech synthesis model training method and speech synthesis method

    CN113450756A