Speech synthesis model training and speech synthesis method, device and related equipment

By classifying audio data in the speech synthesis model and performing multi-stage learning, the problem of the inability to utilize paired audio data in traditional technology is solved, and more accurate and efficient speech synthesis is achieved.

CN114882869BActive Publication Date: 2025-05-06PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210517644.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-13
Publication Date
2025-05-06
Estimated Expiration
2042-05-13

AI Technical Summary

Technical Problem

Traditional speech synthesis models cannot effectively utilize unpaired audio data during training, resulting in insufficient utilization of resources.

Method used

By classifying the audio data of the same speaker in the training sample into paired audio and unpaired audio, and using a combination of encoder, vector quantization layer and decoder to perform supervised, unsupervised and semi-supervised learning, the parameters of the vector quantization layer are optimized to utilize unpaired data.

Benefits of technology

It realizes the effective utilization of unpaired data, enriches the code book content in the speech synthesis model, and thus improves the accuracy and training efficiency of synthetic speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114882869B_ABST
    Figure CN114882869B_ABST
Patent Text Reader

Abstract

The present invention discloses a training method for a speech synthesis model, which is applied to the field of artificial intelligence. The speech synthesis model provided by the present invention includes a vector quantization layer, the vector quantization layer includes a code book, and the method provided by the present invention includes: classifying the audio data to be trained into paired audio and unpaired audio, and converting the paired audio and unpaired audio into a first continuous variable and a second continuous variable respectively through an encoder; training the first continuous variable in a supervised learning manner, and optimizing the parameters of the vector quantization layer according to the obtained first loss; training the second continuous variable in an unsupervised learning manner, and improving the code book; sending the first continuous variable and the second continuous variable to the speech synthesis model for training in a semi-supervised learning manner, optimizing the parameters of the vector quantization layer according to the obtained second loss, until the second loss is minimized, and obtaining a trained speech synthesis model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a speech synthesis model training and speech synthesis method, device and related equipment. Background Art

[0002] Speech synthesis is the process of generating corresponding speech content from text content, that is, the input is a piece of text, and the output is a playable audio file. Traditional speech synthesis technologies include: concatenative synthesis, parametric synthesis, and end-to-end synthesis. Among them, concatenative synthesis technology and end-to-end synthesis technology only use paired speech and text data sets for speech synthesis, and cannot use unpaired speech and text data sets to obtain audio data. Summary of the invention

[0003] Embodiments of the present invention provide a speech synthesis model training and speech synthesis method, apparatus, computer equipment and storage medium to solve the problem that unpaired audio data cannot be used in the training process of a traditional speech synthesis model.

[0004] A method for training a speech synthesis model, the speech synthesis model comprising an encoder, a vector quantization layer, and a decoder, the vector quantization layer comprising a codebook, the method comprising:

[0005] Classify audio data of the same speaker included in the training sample into paired audio and unpaired audio;

[0006] Convert the paired audio into a first continuous variable corresponding to each audio data respectively through an encoder;

[0007] Extracting the first continuous variable and inputting the extracted first continuous variable into the vector quantization layer for vectorization processing to obtain a first discrete variable and a first phoneme vector, adding the first phoneme vector to the code book, converting the first discrete variable into a first reconstructed audio through the decoder, and calculating a first loss between the first reconstructed audio and the paired audio;

[0008] Determine whether the first loss is minimized, and if not, optimize the parameters of the vector quantization layer according to the first loss, and loop the steps from extracting the first continuous variable to determining whether the first loss is minimized, until the first loss is minimized, thereby completing the first stage of training;

[0009] Convert the unpaired audio into a second continuous variable corresponding to each audio data respectively through an encoder;

[0010] Inputting the second continuous variable into the vector quantization layer for vectorization processing, obtaining a second phoneme vector including a pseudo phoneme label after passing through the code book, and adding the second phoneme vector to the code book to obtain an updated vector quantization layer;

[0011] Input the first continuous variable and the second continuous variable into the updated vector quantization layer for processing to obtain a second discrete variable, convert the second discrete variable into a second reconstructed audio through the decoder, and calculate a second loss between the second reconstructed audio and the paired audio or the unpaired audio;

[0012] Determine whether the second loss reaches a minimum. If not, optimize the parameters of the updated vector quantization layer according to the second loss, and loop the steps of inputting the first continuous variable and the second continuous variable into the updated vector quantization layer for processing to determining whether the second loss reaches a minimum, until the second loss reaches a minimum, thereby obtaining a trained speech synthesis model.

[0013] A method for performing speech synthesis according to the speech synthesis model trained by the above method, the method comprising:

[0014] Recognize the input text through the phoneme recognition tool to obtain the phoneme label to be synthesized;

[0015] The to-be-synthesized phoneme label is input into the speech synthesis model, and the speech synthesis model queries the to-be-synthesized latent vector corresponding to the to-be-synthesized phoneme label according to a pre-selected code book;

[0016] The latent vector to be synthesized is input into a decoder in the speech synthesis model to obtain a target synthesized speech.

[0017] A training device for a speech synthesis model, the speech synthesis model comprising an encoder, a vector quantization layer, and a decoder, the vector quantization layer comprising a codebook, the device comprising:

[0018] An audio data classification module, used for classifying audio data of the same speaker included in the training sample into paired audio and unpaired audio;

[0019] A first data conversion module, configured to convert the paired audio into a first continuous variable corresponding to each audio data through an encoder;

[0020] A first loss calculation module is used to extract the first continuous variable and input the extracted first continuous variable to the vector quantization layer for vectorization processing to obtain a first discrete variable and a first phoneme vector, add the first phoneme vector to the code book, convert the first discrete variable into a first reconstructed audio through the decoder, and calculate a first loss between the first reconstructed audio and the paired audio;

[0021] A first loss loop module, used to determine whether the first loss is minimized, and if not, optimize the parameters of the vector quantization layer according to the first loss, and loop the steps from extracting the first continuous variable to determining whether the first loss is minimized, until the first loss is minimized, completing the first stage of training;

[0022] A second data conversion module, configured to convert the unpaired audio into a second continuous variable corresponding to each audio data through an encoder;

[0023] A pseudo phoneme label module, used for inputting the second continuous variable into the vector quantization layer for vectorization processing, obtaining a second phoneme vector containing a pseudo phoneme label after passing through the code book, and adding the second phoneme vector to the code book to obtain an updated vector quantization layer;

[0024] A second loss calculation module, configured to input the first continuous variable and the second continuous variable into the updated vector quantization layer for processing to obtain a second discrete variable, convert the second discrete variable into a second reconstructed audio through the decoder, and calculate a second loss between the second reconstructed audio and the paired audio or the unpaired audio;

[0025] The second loss loop module is used to determine whether the second loss is minimized. If not, the parameters of the updated vector quantization layer are optimized according to the second loss, and the steps of inputting the first continuous variable and the second continuous variable into the updated vector quantization layer for processing to determining whether the second loss is minimized are looped until the second loss is minimized to obtain a trained speech synthesis model.

[0026] A device for performing speech synthesis according to the speech synthesis model provided by the training device for the speech synthesis model, the device comprising:

[0027] The phoneme recognition module is used to recognize the input text through the phoneme recognition tool to obtain the phoneme label to be synthesized;

[0028] A phoneme query module, used for inputting the to-be-synthesized phoneme label into the speech synthesis model, and the speech synthesis model queries the to-be-synthesized latent vector corresponding to the to-be-synthesized phoneme label according to a pre-selected code book;

[0029] The target speech synthesis module is used to input the latent vector to be synthesized into the decoder in the speech synthesis model to obtain the target synthesized speech.

[0030] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the processor implements the steps of the method for training the above-mentioned speech synthesis model or the method for performing speech synthesis according to the speech synthesis model.

[0031] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for training the above-mentioned speech synthesis model or the method for performing speech synthesis based on the speech synthesis model.

[0032] The training of the speech synthesis model and the speech synthesis method, device, computer equipment and storage medium first send paired data to the speech synthesis model for supervised learning training, then send unpaired data to the speech synthesis model for unsupervised learning training, and finally send the paired data and unpaired data together to the speech synthesis model for semi-supervised learning training. A large amount of unpaired data that cannot be used by traditional technology is utilized, and the content of the code book in the speech synthesis model is further enriched through the unpaired data, so that the content of the synthesized speech obtained in the end is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0034] Figure 1 It is a schematic diagram of an application environment for training a speech synthesis model and a speech synthesis method in one embodiment of the present invention;

[0035] Figure 2 is a flow chart of a method for training a speech synthesis model in one embodiment of the present invention;

[0036] Figure 3 is a schematic diagram of the architecture of a speech synthesis model before training in one embodiment of the present invention;

[0037] Figure 4 is a flow chart of a speech synthesis method according to an embodiment of the present invention;

[0038] Figure 5is a structural schematic diagram of a training device for a speech synthesis model in one embodiment of the present invention;

[0039] Figure 6 is a structural schematic diagram of a speech synthesis device in one embodiment of the present invention;

[0040] Figure 7 is a schematic diagram of a type of computer device in one embodiment of the present invention;

[0041] Figure 8 is a schematic diagram of another type of computer device in an embodiment of the present invention. DETAILED DESCRIPTION

[0042] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0043] The training of the speech synthesis model and the speech synthesis method provided in this application can be applied in Figure 1 In the application environment, the computer device can communicate with an external device through a network, and the external device is, for example, a server. The computer device can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices. The server can be implemented as an independent server or a server cluster consisting of multiple servers.

[0044] In one embodiment, if Figure 2 As shown, a training method for a speech synthesis model is provided, wherein the speech synthesis model includes a vector quantization layer, and the vector quantization layer includes a code book. Figure 1 The server in the example is used as an example to illustrate, including the following steps S101 to S108:

[0045] S101. Classify audio data of the same speaker included in a training sample into paired audio and unpaired audio.

[0046] Specifically, the paired audio contains audio data and text data corresponding to the audio data. The unpaired audio only contains audio data, but does not contain corresponding text data. In traditional speech synthesis technology, the unpaired audio cannot be used to train the model. In this embodiment, first, the paired data is used to train the speech synthesis model in a supervised learning manner to complete the first stage of training. Then, based on the results of the first stage of training, the unpaired data is used to train the speech synthesis model in an unsupervised learning manner to complete the second stage of training. Finally, based on the results of the first stage of training and the second stage of training, the paired data and the unpaired data are used to train the speech synthesis model in a semi-supervised learning manner to complete the third stage of training. After the third stage of training is completed, a trained speech synthesis model is obtained.

[0047] S102: Convert the paired audio into a first continuous variable corresponding to each audio data through an encoder.

[0048] S103, extracting the first continuous variable and inputting the extracted first continuous variable into the vector quantization layer for vectorization processing to obtain a first discrete variable and a first phoneme vector, adding the first phoneme vector to the code book, converting the first discrete variable into a first reconstructed audio through the decoder, and calculating a first loss between the first reconstructed audio and the paired audio.

[0049] Among them, the code book contains the phoneme labels of the same speaker and the audio vectors corresponding to the phoneme labels, but the code book is limited to the use of the same speaker, that is, the audio vectors in the code book are all derived from the speech data of the same speaker. After the training of the speech synthesis model is completed, the code book is used in the process of speech synthesis using the speech synthesis model, and the synthesized speech obtained only contains the physical properties of the speaker's speech, and the physical properties of speech include the pitch, sound intensity, sound duration and timbre of the sound.

[0050] Specifically, firstly, the first continuous variable is vector quantized by the preset distance formula in the vector quantization layer to obtain the first discrete variable. Then, the code book is queried by the preset search function in the vector quantization layer to obtain the first phoneme label set corresponding to each variable in the first discrete variable. Then, the first phoneme label set is filtered by the normalized exponential function preset in the vector quantization layer, so that each variable in the first discrete variable corresponds to a unique phoneme label. Finally, each variable in the first discrete variable and the unique phoneme label corresponding to each variable are converted into the first phoneme vector one by one, and the first phoneme vector is added to the code book. It should be particularly noted that before adding the first phoneme vector to the code book, it is also possible to query whether the same first phoneme vector already exists in the code book. If so, the adding step is not performed. The first discrete variable is also sent to a preset decoder, and the decoder converts the first discrete variable into a first reconstructed audio. The first loss of the first reconstructed audio and the paired audio is calculated by a preset loss function.

[0051] Before the first discrete variable is converted into the first reconstructed audio through the decoder, the method further includes: firstly merging the variables corresponding to the same phoneme label in the first discrete variable to obtain a merged discrete variable; then replacing the first discrete variable with the merged discrete variable, and then sending the first discrete variable to the decoder. The merging operation further compresses the amount of data contained in the first discrete variable, so that the amount of tasks for the subsequent decoder to process the first discrete variable is further reduced, so that the efficiency of the entire training process of the speech synthesis model is further improved.

[0052] In a specific embodiment, Figure 3 As shown, the preset distance formula of the vector quantization layer is L2 distance (Euclidean Distance), and the distance formula can also be replaced by L1 distance (Manhattan). The L2 distance formula and the L1 distance formula can obtain discrete variables after processing continuous variables, and the specific processing methods and principles are not repeated here. The query function adopts the argmin function, and the query function can also be replaced by the argmax function. The mathematical characteristics and usage of the argmin function and the argmax function are not repeated here. The normalized exponential function adopts the softmin function, and the normalized exponential function can also be replaced by the softmax function. The mathematical characteristics and usage of the softmin function and the softmax function are not repeated here. The loss function is a mean absolute error loss function or a mean square error loss function, or an improved version of the mean absolute error loss function and the mean square error loss function.

[0053] S104, determining whether the first loss is minimized; if not, optimizing the parameters of the vector quantization layer according to the first loss, and looping the steps from extracting the first continuous variable to determining whether the first loss is minimized, until the first loss is minimized, thereby completing the first stage of training.

[0054] Specifically, the parameters in the distance formula in the vector quantization layer are optimized according to the first loss. Because vector quantization is a lossy data compression technology, optimizing the parameters in the distance formula can retain the characteristics of the audio data to the greatest extent in the process of lossy compression of the first continuous variable into the first discrete variable. At the same time, because the lossy data compression technology of vector quantization is adopted, the subsequent data processing task volume and task difficulty are reduced during the training process of the speech synthesis model, and the overall training process efficiency of the speech synthesis model is improved. Among them, the query function is optimized according to the first loss, so that when the code book is queried again, a phoneme label with a higher matching degree can be obtained. Among them, the normalized exponential function is optimized according to the first loss, so that when the normalized exponential function is used to filter the first phoneme label set, a timbre label with a higher matching degree can be retained.

[0055] Among them, in the first stage of using the paired data to loop the speech synthesis model, the audio data contained in the paired data and the text data corresponding to the audio data are utilized to continuously improve the codebook of the vector quantization layer in the speech synthesis model, and adjust the various parameters of the vector quantization layer, thereby establishing a technical foundation for training the unpaired data in an unsupervised manner.

[0056] S105. Convert the unpaired audio into a second continuous variable corresponding to each audio data through an encoder.

[0057] S106. Input the second continuous variable into the vector quantization layer for vectorization processing, obtain a second phoneme vector containing a pseudo phoneme label after passing through the code book, add the second phoneme vector to the code book, and obtain an updated vector quantization layer.

[0058] Specifically, first, the second continuous variable is vectorized by the preset distance formula in the vector quantization layer to obtain the third discrete variable. Then, the code book that has completed the first stage of training is queried by the preset search function to obtain the pseudo phoneme label corresponding to each variable in the third discrete variable. It should be noted that because the third discrete variable is converted from unpaired data, the accurate label cannot be determined, but a relatively close pseudo phoneme label can be obtained based on the enriched code book after the first stage of training. Finally, the variable corresponding to each pseudo phoneme label is further sharpened by the preset sharpening function to obtain the corresponding relationship between the sharpened variable and the pseudo phoneme label, and the corresponding relationship is converted into a second phoneme vector. The second phoneme vector is added to the code book to further enrich the content of the code book and complete the second stage of training.

[0059] Wherein, the sharpening function can be a softmax function or a softmin function, and the mathematical characteristics and usage of the softmin function and the softmax function are not described here. It should be noted that due to the uncertainty of the unpaired data, in the second stage of training, there will be a situation where some variables in the third discrete variable cannot match the pseudo phoneme label in the code book, because the search function searches for the phoneme label as the pseudo phoneme label within the preset first range. When the situation that the search function cannot match occurs, the phoneme label within the preset second range is output by the search function. It can be known that the second range must be greater than the first range. Finally, it is necessary to determine whether a new phoneme label needs to be generated according to the number of phoneme labels within the second range, that is, when the phoneme labels within the second range are less than the first preset number, a new phoneme label is generated according to the preset phoneme label generation method, and the new phoneme label and the variable corresponding to the new phoneme label are converted into a second phoneme vector, and the second phoneme vector is added to the code book. When the number of timbre labels within the second range is greater than or equal to the first preset number, the current variable is treated as noise data in training and discarded.

[0060] S107. Input the first continuous variable and the second continuous variable into the updated vector quantization layer for processing to obtain a second discrete variable, convert the second discrete variable into a second reconstructed audio through the decoder, and calculate the second loss of the second reconstructed audio and the paired audio or the unpaired audio.

[0061] Among them, after completing the first stage training and the second stage training of the speech synthesis model using the paired data and the unpaired data respectively, the code book of the speech synthesis model has been sufficiently enriched. At this time, the paired data and the unpaired data need to be resent to the speech synthesis model for training in a semi-supervised learning manner, so as to further optimize the various parameters of the vector quantization layer in the speech synthesis model.

[0062] Specifically, compared with the first training stage and the second training stage, in the training process of the third training stage at this time, the paired data and the unpaired data are processed in a series to obtain a second discrete variable, the second discrete variable directly queries the corresponding phoneme label set from the code book, and then performs the aforementioned phoneme merging operation, and finally converts the second discrete variable after the phoneme merging operation into a second reconstructed audio through a decoder. The second loss of the second reconstructed audio and the audio data in the paired data or the audio data in the unpaired data is calculated by the preset damage function.

[0063] Among them, at this time, the third stage training process does not add operations to the code book in the speech synthesis model, so according to the second loss, only the parameters of the preset distance formula and the parameters of the preset search function in the vector quantization layer are optimized.

[0064] S108. Determine whether the second loss reaches a minimum. If not, optimize the parameters of the updated vector quantization layer according to the second loss, and loop the steps of inputting the first continuous variable and the second continuous variable into the updated vector quantization layer for processing to determining whether the second loss reaches a minimum, until the second loss reaches a minimum, thereby obtaining a trained speech synthesis model.

[0065] The training method of the speech synthesis model proposed in this embodiment first sends paired data to the speech synthesis model for supervised learning training, then sends unpaired data to the speech synthesis model for unsupervised learning training, and finally sends the paired data and unpaired data together to the speech synthesis model for semi-supervised learning training. This method not only utilizes a large amount of unpaired data that cannot be used by traditional technologies, but also enriches the content of the code book in the speech synthesis model through the unpaired data, making the final synthesized speech content more accurate. It also further compresses the data in the training process through vector quantization and phoneme synchronization methods, thereby greatly improving the training efficiency of the speech synthesis model.

[0066] Figure 4is a flow chart of a method for performing speech synthesis using a speech synthesis model trained by the speech synthesis model training method according to one embodiment of the present invention. According to another embodiment of the present invention, a method for performing speech synthesis using a speech synthesis model trained by the speech synthesis model training method according to the above-mentioned speech synthesis model training method is proposed, such as Figure 4 As shown, the method includes the following steps S201 to S203.

[0067] S201 . Recognize the input text by a phoneme recognition tool to obtain a phoneme label to be synthesized.

[0068] The phoneme recognition tool is a recognition tool that has been trained with a large amount of data in advance and can accurately convert the input text into a corresponding phoneme tag set. It should be noted that the phoneme tags to be synthesized can all be found in the code book of the speech synthesis model.

[0069] S202: input the to-be-synthesized phoneme label into the speech synthesis model, and the speech synthesis model queries the to-be-synthesized latent vector corresponding to the to-be-synthesized phoneme label according to a pre-selected code book.

[0070] Before performing speech synthesis processing, a corresponding codebook needs to be pre-selected, and the pre-selected codebook is obtained by training the aforementioned speech synthesis model based on the speech data of the target speaker. Selecting different codebooks will eventually result in different physical properties of the synthesized speech, that is, different target speakers are speaking from the auditory sense.

[0071] S203: input the latent vector to be synthesized into a decoder in the speech synthesis model to obtain target synthesized speech.

[0072] It should be noted that the latent vector to be synthesized is not converted from a continuous variable, so in the process of using the speech synthesis model to perform speech synthesis, it is not necessary to perform phoneme merging operation on the latent vector to be synthesized, which further shortens the speech synthesis steps and improves the efficiency of the speech synthesis process.

[0073] In one embodiment, a training device 100 for a speech synthesis model is provided, wherein the speech synthesis model includes a vector quantization layer, and the vector quantization layer includes a codebook. The training device 100 for the speech synthesis model corresponds to the training method for the speech synthesis model in the above embodiment. Figure 5 As shown, the training device 100 of the speech synthesis model includes an audio data classification module 11, a first data conversion module 12, a first loss calculation module 13, a first loss circulation module 14, a second data conversion module 15, a pseudo phoneme label module 16, a second loss calculation module 17 and a second loss circulation module 18. The functional modules are described in detail as follows:

[0074] An audio data classification module 11, used to classify audio data of the same speaker included in the training sample into paired audio and unpaired audio;

[0075] A first data conversion module 12, configured to convert the paired audio into a first continuous variable corresponding to each audio data through an encoder;

[0076] A first loss calculation module 13 is used to extract the first continuous variable and input the extracted first continuous variable to the vector quantization layer for vectorization processing to obtain a first discrete variable and a first phoneme vector, add the first phoneme vector to the code book, convert the first discrete variable into a first reconstructed audio through the decoder, and calculate a first loss between the first reconstructed audio and the paired audio;

[0077] A first loss loop module 14 is used to determine whether the first loss is minimized. If not, the parameters of the vector quantization layer are optimized according to the first loss, and the steps from extracting the first continuous variable to determining whether the first loss is minimized are looped until the first loss is minimized, thereby completing the first stage of training;

[0078] A second data conversion module 15, configured to convert the unpaired audio into a second continuous variable corresponding to each audio data through an encoder;

[0079] A pseudo phoneme label module 16 is used to input the second continuous variable into the vector quantization layer for vectorization processing, obtain a second phoneme vector containing a pseudo phoneme label after passing through the code book, and add the second phoneme vector to the code book to obtain an updated vector quantization layer;

[0080] A second loss calculation module 17 is used to input the first continuous variable and the second continuous variable into the updated vector quantization layer for processing to obtain a second discrete variable, convert the second discrete variable into a second reconstructed audio through the decoder, and calculate a second loss between the second reconstructed audio and the paired audio or the unpaired audio;

[0081] The second loss loop module 18 is used to determine whether the second loss has reached a minimum. If not, the parameters of the updated vector quantization layer are optimized according to the second loss, and the steps of inputting the first continuous variable and the second continuous variable into the updated vector quantization layer for processing to determining whether the second loss has reached a minimum are looped until the second loss reaches a minimum, thereby obtaining a trained speech synthesis model.

[0082] Furthermore, the first loss calculation module 13 also includes:

[0083] A first discrete variable calculation submodule, configured to perform vector quantization processing on the first continuous variable using a preset distance formula to obtain a first discrete variable;

[0084] A first phoneme label set submodule, configured to query the code book through a preset search function to obtain a first phoneme label set corresponding to each variable in the first discrete variable;

[0085] A normalization processing submodule, used for filtering the first phoneme label set by a preset normalization exponential function so that each variable in the first discrete variable corresponds to a unique phoneme label;

[0086] A first phoneme vector submodule, configured to convert each variable in the first discrete variables and the unique phoneme label corresponding to each variable into the first phoneme vector one by one, and add the first phoneme vector to the code book;

[0087] A phoneme label merging submodule, used for merging the variables corresponding to the same phoneme label in the first discrete variables to obtain a merged discrete variable;

[0088] The discrete variable sending submodule is used to replace the first discrete variable with the combined discrete variable, and then send the first discrete variable to a decoder.

[0089] Furthermore, the pseudo-phoneme label module 16 further includes:

[0090] A third discrete variable calculation submodule, configured to vectorize the second continuous variable using a preset distance formula to obtain a third discrete variable;

[0091] A pseudo-phoneme label submodule, used to query the codebook that has completed the first stage training through a preset search function to obtain a pseudo-phoneme label corresponding to each variable in the third discrete variable;

[0092] The second phoneme vector submodule is used to further sharpen the variables corresponding to each pseudo phoneme label through a preset sharpening function, obtain the corresponding relationship between the sharpened variables and the pseudo phoneme labels, and convert the corresponding relationship into a second phoneme vector.

[0093] For the specific definition of the training device of the speech synthesis model, please refer to the definition of the training method of the speech synthesis model in the above text, which will not be repeated here. The various modules in the above-mentioned training device of the speech synthesis model can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0094] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data involved in the training method of the speech synthesis model. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a training method for a speech synthesis model is implemented.

[0095] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, a training method for a speech synthesis model is implemented.

[0096] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for training a speech synthesis model in the above embodiment are implemented, for example: Figure 2 Alternatively, when the processor executes the computer program, the functions of each module / unit of the training device for the speech synthesis model in the above embodiment are realized, for example Figure 5 The functions of modules 11 to 18 are shown in Figure 1. To avoid repetition, they will not be described here.

[0097] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the training method of the speech synthesis model in the above embodiment are implemented, such as Figure 2Alternatively, when the computer program is executed by a processor, the functions of each module / unit of the training device for the speech synthesis model in the above embodiment are realized, for example Figure 5 The functions of modules 11 to 18 are shown in Figure 1. To avoid repetition, they will not be described here.

[0098] Figure 6 FIG. 2 is a schematic diagram of the structure of a speech synthesis device 200 according to an embodiment of the present invention. Figure 6 As shown, the device 200 for performing speech synthesis based on the speech synthesis model provided by the speech synthesis model training device 100 includes a phoneme recognition module 21, a phoneme query module 22, and a target speech synthesis module 23. The functional modules are described in detail as follows:

[0099] The phoneme recognition module 21 is used to recognize the input text through a phoneme recognition tool to obtain a phoneme label to be synthesized;

[0100] A phoneme query module 22, configured to input the to-be-synthesized phoneme label into the speech synthesis model, and the speech synthesis model queries the to-be-synthesized latent vector corresponding to the to-be-synthesized phoneme label according to a pre-selected code book;

[0101] The target speech synthesis module 23 is used to input the latent vector to be synthesized into the decoder in the speech synthesis model to obtain the target synthesized speech.

[0102] The meaning of "first" and "second" in the above modules / units is only to distinguish different modules / units, and is not used to define which module / unit has a higher priority or other limiting meanings. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices. The division of modules in this application is only a logical division, and there may be other division methods when implemented in actual applications.

[0103] For the specific definition of the speech synthesis device, please refer to the definition of the speech synthesis method above, which will not be repeated here. Each module in the above-mentioned speech synthesis device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0104] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data involved in the speech synthesis method. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a speech synthesis method is implemented.

[0105] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, a speech synthesis method is implemented.

[0106] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the speech synthesis method in the above embodiment are implemented, such as Figure 4 Alternatively, when the processor executes the computer program, the functions of each module / unit of the speech synthesis device in the above embodiment are realized, for example, Figure 6 The functions of modules 21 to 23 are shown in Figure 2. To avoid repetition, they will not be described here.

[0107] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the computer device, and various interfaces and lines are used to connect various parts of the entire computer device.

[0108] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the computer device by running or executing the computer program and / or module stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, video data, etc.), etc.

[0109] The memory may be integrated into the processor or may be arranged separately from the processor.

[0110] In one embodiment, a computer-readable storage medium is provided on which a computer program is stored. When the computer program is executed by a processor, the steps of the speech synthesis method in the above embodiment are implemented, such as Figure 4 Alternatively, when the computer program is executed by the processor, the functions of each module / unit of the speech synthesis device in the above embodiment are realized, for example, Figure 6 The functions of modules 21 to 23 are shown in Figure 2. To avoid repetition, they will not be described here.

[0111] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0112] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0113] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A method for training a speech synthesis model, characterized in that: The speech synthesis model includes an encoder, a vector quantization layer, and a decoder, the vector quantization layer includes a codebook, and the method includes: Classify audio data of the same speaker included in the training sample into paired audio and unpaired audio; Convert the paired audio into a first continuous variable corresponding to each audio data respectively through an encoder; Extracting the first continuous variable and inputting the extracted first continuous variable into the vector quantization layer for vectorization processing to obtain a first discrete variable and a first phoneme vector, adding the first phoneme vector to the code book, converting the first discrete variable into a first reconstructed audio through the decoder, and calculating a first loss between the first reconstructed audio and the paired audio; Determine whether the first loss is minimized, and if not, optimize the parameters of the vector quantization layer according to the first loss, and loop the steps from extracting the first continuous variable to determining whether the first loss is minimized, until the first loss is minimized, thereby completing the first stage of training; Convert the unpaired audio into a second continuous variable corresponding to each audio data respectively through an encoder; Inputting the second continuous variable into the vector quantization layer for vectorization processing, obtaining a second phoneme vector including a pseudo phoneme label after passing through the code book, and adding the second phoneme vector to the code book to obtain an updated vector quantization layer; Input the first continuous variable and the second continuous variable into the updated vector quantization layer for processing to obtain a second discrete variable, convert the second discrete variable into a second reconstructed audio through the decoder, and calculate a second loss between the second reconstructed audio and the paired audio or the unpaired audio; Determine whether the second loss reaches a minimum. If not, optimize the parameters of the updated vector quantization layer according to the second loss, and loop the steps of inputting the first continuous variable and the second continuous variable into the updated vector quantization layer for processing to determining whether the second loss reaches a minimum, until the second loss reaches a minimum, thereby obtaining a trained speech synthesis model.

2. The method for training a speech synthesis model according to claim 1, characterized in that: The step of inputting the first continuous variable into the vector quantization layer for vectorization processing to obtain a first discrete variable and a first phoneme vector comprises: Performing vector quantization processing on the first continuous variable using a preset distance formula to obtain a first discrete variable; Searching the codebook by a preset search function to obtain a first phoneme label set corresponding to each variable in the first discrete variable; Filtering the first phoneme label set by a preset normalized exponential function so that each variable in the first discrete variable corresponds to a unique phoneme label; Each variable in the first discrete variables and the unique phoneme label corresponding to each variable are converted into the first phoneme vector one by one, and the first phoneme vector is added to the code book.

3. The method for training a speech synthesis model according to claim 2, characterized in that: The step of inputting the second continuous variable into the vector quantization layer for vectorization processing, and obtaining a second phoneme vector including a pseudo phoneme label through the code book further comprises: The second continuous variable is vectorized by the preset distance formula to obtain a third discrete variable; By searching the codebook that has completed the first stage of training through the preset search function, a pseudo phoneme label corresponding to each variable in the third discrete variable is obtained; The variable corresponding to each of the pseudo-phoneme labels is further sharpened by a preset sharpening function to obtain a correspondence between the sharpened variable and the pseudo-phoneme label, and the correspondence is converted into the second phoneme vector.

4. The method for training a speech synthesis model according to claim 1, characterized in that: Before converting the first discrete variable into a first reconstructed audio through the decoder, the method further includes: Merging the variables corresponding to the same phoneme label in the first discrete variables to obtain a merged discrete variable; The combined discrete variable replaces the first discrete variable, and then the replaced first discrete variable is sent to a decoder.

5. The method for training a speech synthesis model according to claim 1, characterized in that: The loss function for calculating the first loss is a mean absolute error loss function or a mean square error loss function.

6. A method for performing speech synthesis using a speech synthesis model obtained by the training method according to any one of claims 1 to 5, characterized in that: include: Recognize the input text through the phoneme recognition tool to obtain the phoneme label to be synthesized; The to-be-synthesized phoneme label is input into the speech synthesis model, and the speech synthesis model queries the to-be-synthesized latent vector corresponding to the to-be-synthesized phoneme label according to a pre-selected code book; The latent vector to be synthesized is input into a decoder in the speech synthesis model to obtain a target synthesized speech.

7. A training device for a speech synthesis model, characterized in that: The speech synthesis model includes an encoder, a vector quantization layer, and a decoder, the vector quantization layer includes a codebook, and the device includes: An audio data classification module, used for classifying audio data of the same speaker included in the training sample into paired audio and unpaired audio; A first data conversion module, configured to convert the paired audio into a first continuous variable corresponding to each audio data through an encoder; A first loss calculation module is used to extract the first continuous variable and input the extracted first continuous variable to the vector quantization layer for vectorization processing to obtain a first discrete variable and a first phoneme vector, add the first phoneme vector to the code book, convert the first discrete variable into a first reconstructed audio through the decoder, and calculate a first loss between the first reconstructed audio and the paired audio; A first loss loop module, used to determine whether the first loss is minimized, and if not, optimize the parameters of the vector quantization layer according to the first loss, and loop the steps from extracting the first continuous variable to determining whether the first loss is minimized, until the first loss is minimized, completing the first stage of training; A second data conversion module, configured to convert the unpaired audio into a second continuous variable corresponding to each audio data through an encoder; A pseudo phoneme label module, used for inputting the second continuous variable into the vector quantization layer for vectorization processing, obtaining a second phoneme vector containing a pseudo phoneme label after passing through the code book, and adding the second phoneme vector to the code book to obtain an updated vector quantization layer; A second loss calculation module, configured to input the first continuous variable and the second continuous variable into the updated vector quantization layer for processing to obtain a second discrete variable, convert the second discrete variable into a second reconstructed audio through the decoder, and calculate a second loss between the second reconstructed audio and the paired audio or the unpaired audio; The second loss loop module is used to determine whether the second loss is minimized. If not, the parameters of the updated vector quantization layer are optimized according to the second loss, and the steps of inputting the first continuous variable and the second continuous variable into the updated vector quantization layer for processing to determining whether the second loss is minimized are looped until the second loss is minimized to obtain a trained speech synthesis model.

8. A device for performing speech synthesis according to the speech synthesis model of the speech synthesis model training device provided in claim 7, characterized in that: include: The phoneme recognition module is used to recognize the input text through the phoneme recognition tool to obtain the phoneme label to be synthesized; A phoneme query module, used for inputting the to-be-synthesized phoneme label into the speech synthesis model, and the speech synthesis model queries the to-be-synthesized latent vector corresponding to the to-be-synthesized phoneme label according to a pre-selected code book; The target speech synthesis module is used to input the latent vector to be synthesized into the decoder in the speech synthesis model to obtain the target synthesized speech.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Voice conversion model training method and device, voice conversion model application method and device, equipment and storage medium

    CN113345454A

  • Voice processing method and device

    CN113823258A