Training of a speech synthesis model, speech synthesis method, apparatus, device, and medium
Patent Information
- Application Number
- CN202211690583.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-27
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2042-12-27
AI Technical Summary
[0004]本发明提供一种语音合成模型的训练、语音合成方法、装置、设备及介质,用以解决现有技术中在可用数据量较低,录音质量不佳的业务场景下,语音合成的质量差的缺陷
[0035] The speech synthesis model training, speech synthesis method, device, equipment and medium provided by the present invention, based on phoneme classification results and phoneme labels, iterates the parameters of the initial synthesis model, thereby obtaining a speech synthesis model with better speech synthesis effect, improving the accuracy and reliability of the speech synthesis model. At the same time, the phoneme classification result is obtained by classifying the predicted acoustic features into phonemes, which further improves the phoneme classification effect of the speech synthesis model.
Smart Images

Figure CN116013245B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech synthesis technology, and in particular to a speech synthesis model training, speech synthesis method, apparatus, device and medium. Background Technology
[0002] In recent years, various speech synthesis systems based on generative models such as VAE (Variational Autoencoder) and Normalizing-Flow models have emerged in the industry. The effectiveness of these speech synthesis systems is highly dependent on high-quality and accurately labeled speech data, and they also have strict requirements on the amount of data.
[0003] In existing technologies, speech synthesis systems based on the Normalizing-Flow model generally suffer from problems such as unclear pronunciation and imperfect sound quality, resulting in poor speech synthesis quality. However, in business scenarios with limited available data and poor recording quality, the quality of speech synthesis is even worse. Summary of the Invention
[0004] This invention provides a training method, apparatus, device, and medium for speech synthesis, which addresses the shortcomings of existing technologies in business scenarios with low available data and poor recording quality, resulting in poor speech synthesis quality.
[0005] This invention provides a training method for a speech synthesis model, comprising:
[0006] Obtain the sample text and the phoneme tags of the sample speech corresponding to the sample text;
[0007] Based on the initial synthesis model, speech synthesis is performed on the sample text to obtain the predicted acoustic features of the sample text;
[0008] Phoneme classification is performed on the predicted acoustic features to obtain the phoneme classification results of the predicted acoustic features;
[0009] Based on the phoneme classification results and the phoneme labels, the initial synthesis model is iterated to obtain the speech synthesis model.
[0010] According to a training method for a speech synthesis model provided by the present invention, the step of iterating the parameters of the initial synthesis model based on the phoneme classification results and the phoneme labels to obtain the speech synthesis model includes:
[0011] Based on the phoneme classification results and phoneme labels, as well as the text encoding of the sample text and the speech encoding of the sample speech, the initial synthesis model is iterated to obtain the speech synthesis model.
[0012] The initial synthesis model includes a cascaded initial encoder and an initial decoder. The text encoding is obtained by encoding the sample text based on the initial encoder, and the speech encoding is obtained by inverse encoding the sample speech based on the initial decoder.
[0013] According to a training method for a speech synthesis model provided by the present invention, the initial synthesis model is iterated to obtain the speech synthesis model by performing parameter iteration based on the phoneme classification results and the phoneme labels, as well as the text encoding of the sample text and the speech encoding of the sample speech, including:
[0014] Based on the difference between the phoneme classification results and the phoneme labels, the classification loss is determined;
[0015] The distribution loss is determined based on the difference between the text encoding and the speech encoding;
[0016] Based on the classification loss and the distribution loss, the initial synthesis model is iterated to obtain the speech synthesis model.
[0017] According to a training method for a speech synthesis model provided by the present invention, the step of iterating the parameters of the initial synthesis model based on the classification loss and the distribution loss to obtain the speech synthesis model includes:
[0018] The classification loss and the distribution loss are weighted and fused to obtain the fusion loss;
[0019] Based on the fusion loss, the parameters of the initial synthesis model are iterated to obtain the speech synthesis model.
[0020] According to a training method for a speech synthesis model provided by the present invention, the step of performing phoneme classification on the predicted acoustic features to obtain the phoneme classification result of the predicted acoustic features includes:
[0021] Based on the phoneme classifier, the predicted acoustic features are classified into phonemes to obtain the phoneme classification results of the predicted acoustic features;
[0022] The phoneme classifier is trained based on the acoustic features of samples carrying phoneme labels, and the configuration parameters of the sample acoustic features are consistent with those of the predicted acoustic features.
[0023] The present invention also provides a speech synthesis method, comprising:
[0024] Obtain the text to be synthesized;
[0025] Based on the speech synthesis model, speech synthesis is performed on the text to be synthesized.
[0026] The speech synthesis model is obtained by iterating the parameters of an initial synthesis model based on the phoneme classification results of the sample text and the phoneme tags carried by the corresponding sample speech. The phoneme classification results are obtained by classifying the predicted acoustic features of the sample text synthesized based on the initial synthesis model.
[0027] The present invention also provides a training device for a speech synthesis model, comprising:
[0028] The acquisition unit is used to acquire sample text and sample speech corresponding to the sample text, wherein the sample speech carries phoneme tags.
[0029] A speech synthesis unit is used to perform speech synthesis on the sample text based on an initial synthesis model to obtain the predicted acoustic features of the sample text.
[0030] A phoneme classification unit is used to classify the predicted acoustic features into phonemes to obtain the phoneme classification results of the predicted acoustic features.
[0031] The parameter iteration unit is used to perform parameter iteration on the initial synthesis model based on the phoneme classification results and the phoneme labels to obtain the speech synthesis model.
[0032] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a training method for any of the above-described speech synthesis models.
[0033] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the training method of the speech synthesis model as described above.
[0034] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements a training method for any of the above-described speech synthesis models.
[0035] The speech synthesis model training, speech synthesis method, device, equipment and medium provided by the present invention, based on phoneme classification results and phoneme labels, iterates the parameters of the initial synthesis model, thereby obtaining a speech synthesis model with better speech synthesis effect, improving the accuracy and reliability of the speech synthesis model. At the same time, the phoneme classification result is obtained by classifying the predicted acoustic features into phonemes, which further improves the phoneme classification effect of the speech synthesis model. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0037] Figure 1 This is a flowchart illustrating the training method for the speech synthesis model provided by the present invention;
[0038] Figure 2 This is a flowchart illustrating step 141 in the training method of the speech synthesis model provided by the present invention;
[0039] Figure 3 This is a flowchart illustrating step 230 in the training method of the speech synthesis model provided by the present invention;
[0040] Figure 4 This is a flowchart illustrating step 130 in the training method of the speech synthesis model provided by the present invention;
[0041] Figure 5 This is a flowchart illustrating the speech synthesis method provided by the present invention;
[0042] Figure 6 This is a schematic diagram of the structure of the training device for the speech synthesis model provided by the present invention;
[0043] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0045] In existing technologies, speech synthesis models can be mainly divided into two categories: autoregressive models (such as the Tacotron model, Transformer-TTS (Transformer-Text to Speech, a speech synthesis model based on Transformer) and non-autoregressive models (such as the Glow-TTS model, FastSpeech model, etc.).
[0046] In existing technologies, autoregressive models can achieve state-of-the-art results in speech acoustic feature modeling tasks, which is determined by the inherent properties of speech synthesis tasks. However, autoregressive models have very limited application scenarios due to their low inference efficiency. Non-autoregressive models (such as the FastSpeech model) are used in many business scenarios because of their parallel computing advantages.
[0047] Speech synthesis systems based on the Normalizing-Flow model generally suffer from unclear pronunciation and imperfect sound quality, resulting in poor speech synthesis quality. However, in business scenarios with limited available data and poor recording quality, the quality of speech synthesis is even worse.
[0048] To address the aforementioned problems, embodiments of the present invention provide a training method for a speech synthesis model. Figure 1 This is a flowchart illustrating the training method for the speech synthesis model provided by the present invention, as shown below. Figure 1 As shown, the method includes:
[0049] Step 110: Obtain the sample text and the phoneme tags of the sample speech corresponding to the sample text.
[0050] Specifically, sample text and the phoneme tags of the corresponding sample speech can be obtained. The sample text can be directly input by the user, or it can be obtained by transcribing the collected audio, or it can be obtained by acquiring images through image acquisition devices such as scanners, mobile phones, cameras, and tablets, and then performing OCR (Optical Character Recognition) on the images. This embodiment of the invention does not specifically limit this.
[0051] The sample speech corresponding to the sample text here refers to the speech at the speech level that is consistent with the text in the sample text, and the sample speech here carries phoneme tags.
[0052] Step 120: Based on the initial synthesis model, perform speech synthesis on the sample text to obtain the predicted acoustic features of the sample text.
[0053] Specifically, after obtaining the sample text, speech synthesis can be performed on the sample text based on the initial synthesis model to obtain the predicted acoustic features of the sample text. The initial synthesis model here refers to the initial speech synthesis model, which can include a cascaded initial encoder and initial decoder. For example, the initial synthesis model can be a Transformer model.
[0054] Here, speech synthesis of sample text refers to converting sample text into human-like speech output. Speech synthesis of sample text can be performed using the Normalizing-Flow model or VAE (Variational Autoencoder), etc. This embodiment of the invention does not specifically limit the specific methods used.
[0055] The predicted acoustic features here reflect acoustic-level features. These features can be Mel Frequency Cepstrum Coefficient (MFCC) features or Perceptual Linear Predictive (PLP) features, etc. This embodiment of the invention does not impose any specific limitations on them.
[0056] Step 130: Perform phoneme classification on the predicted acoustic features to obtain the phoneme classification results of the predicted acoustic features.
[0057] Specifically, after obtaining the predicted acoustic features of the sample text, phoneme classification can be performed on the predicted acoustic features to obtain the phoneme classification results of the predicted acoustic features.
[0058] Here, a phoneme classifier can be used to classify the predicted acoustic features into phonemes, and the phoneme classification results of the predicted acoustic features can be obtained.
[0059] The phoneme classifier here may include a convolutional neural network (CNN) and a Transformer model, or it may include a CNN network, a Transformer model and a fully connected layer (FC). This embodiment of the invention does not specifically limit this.
[0060] Step 140: Based on the phoneme classification results and the phoneme labels, perform parameter iteration on the initial synthesis model to obtain the speech synthesis model.
[0061] Specifically, after obtaining the phoneme classification results, the initial synthesis model can be iterated based on the phoneme classification results and phoneme labels to obtain the speech synthesis model. The parameters of the initial synthesis model can be randomly generated or pre-set.
[0062] That is, the classification loss can be determined based on the phoneme classification results and phoneme labels. The classification loss is used to reflect the difference between the phoneme classification results and phoneme labels. Then, the initial synthesis model can be iterated based on the classification loss, and the initial synthesis model after parameter iteration can be determined as a speech synthesis model.
[0063] It is understandable that the greater the difference between the phoneme classification result and the phoneme label, the greater the classification loss; the smaller the difference between the phoneme classification result and the phoneme label, the smaller the classification loss.
[0064] Based on the phoneme classification results and phoneme labels, the parameters of the initial synthesis model are iterated. As a result, the obtained speech synthesis model has better speech synthesis effect, and at the same time, the phoneme classification effect of the speech synthesis model is further improved.
[0065] The method provided in this embodiment of the invention iterates the parameters of the initial synthesis model based on the phoneme classification results and phoneme labels. As a result, the obtained speech synthesis model has better speech synthesis effect and improves the accuracy and reliability of the speech synthesis model. At the same time, the phoneme classification results are obtained by classifying the predicted acoustic features into phonemes, which further improves the phoneme classification effect of the speech synthesis model.
[0066] Based on the above embodiments, step 140 includes:
[0067] Step 141: Based on the phoneme classification results and the phoneme labels, as well as the text encoding of the sample text and the speech encoding of the sample speech, perform parameter iteration on the initial synthesis model to obtain the speech synthesis model;
[0068] The initial synthesis model includes a cascaded initial encoder and an initial decoder. The text encoding is obtained by encoding the sample text based on the initial encoder, and the speech encoding is obtained by inverse encoding the sample speech based on the initial decoder.
[0069] Specifically, the initial synthesis model may include a cascaded initial encoder and an initial decoder. The initial synthesis model can be a Transformer model, where the initial encoder can be an Encoder and the initial decoder can be a Decoder. The parameters of the initial synthesis model can be randomly generated or pre-set.
[0070] Sample text, corresponding sample speech, and phoneme tags of the corresponding sample speech can be collected in advance.
[0071] During the training of a speech synthesis model, sample text can be input into the initial encoder, which then encodes the sample text to obtain and output the text encoding. The formula for the log-likelihood of the text encoding distribution is as follows:
[0072]
[0073] Where c represents the text encoding.
[0074] Based on the text encoding distribution, the mean μ of the prior Gaussian distribution can be obtained.
[0075] In addition, the sample speech can be input into the initial decoder, which will perform speech inverse encoding on the sample speech to obtain and output the speech code of the sample speech.
[0076] Here, the speech coding distribution follows a Gaussian distribution with a mean of μ.
[0077] Furthermore, based on the initial synthesis model, speech synthesis can be performed on the sample text to obtain the predicted acoustic features of the sample text, and phoneme classification can be performed on the predicted acoustic features to obtain the phoneme classification results of the predicted acoustic features.
[0078] After obtaining the phoneme classification results, as well as the text encoding of the sample text and the speech encoding of the sample speech, the classification loss can be determined based on the phoneme classification results and phoneme labels. The classification loss is used to reflect the difference between the phoneme classification results and the phoneme labels. The distribution loss can also be determined based on the text encoding of the sample text and the speech encoding of the sample speech. The distribution loss is used to reflect the difference between the text encoding of the sample text and the speech encoding of the sample speech.
[0079] After obtaining the classification loss and distribution loss, the initial synthesis model can be iterated based on the classification loss and distribution loss, and the initial synthesis model after parameter iteration can be determined as the speech synthesis model.
[0080] The method provided in this embodiment of the invention iterates the parameters of the initial synthesis model based on phoneme classification results and phoneme labels, as well as the text encoding of the sample text and the speech encoding of the sample speech, to obtain a speech synthesis model, thereby further improving the accuracy and reliability of the speech synthesis model.
[0081] Based on the above embodiments, Figure 2 This is a flowchart illustrating step 141 of the training method for the speech synthesis model provided by the present invention, as shown below. Figure 2 As shown, step 141 includes:
[0082] Step 210: Determine the classification loss based on the difference between the phoneme classification result and the phoneme label;
[0083] Step 220: Determine the distribution loss based on the difference between the text encoding and the speech encoding;
[0084] Step 230: Based on the classification loss and the distribution loss, perform parameter iteration on the initial synthesis model to obtain the speech synthesis model.
[0085] Specifically, after obtaining the phoneme classification results, as well as the text encoding of the sample text and the speech encoding of the sample speech, a classification loss can be determined based on the phoneme classification results and phoneme labels. The classification loss is used to reflect the difference between the phoneme classification results and the phoneme labels. A distribution loss can also be determined based on the text encoding of the sample text and the speech encoding of the sample speech. The distribution loss is used to reflect the difference between the text encoding of the sample text and the speech encoding of the sample speech.
[0086] It is understandable that the greater the difference between the text encoding of the sample text and the speech encoding of the sample speech, the greater the distribution loss; the smaller the difference between the text encoding of the sample text and the speech encoding of the sample speech, the smaller the distribution loss.
[0087] After obtaining the classification loss and distribution loss, the initial synthesis model can be iterated based on the classification loss and distribution loss, or based on the weighted sum of the classification loss and distribution loss, and the initial synthesis model after parameter iteration is determined as the speech synthesis model.
[0088] The method provided in this invention iterates the parameters of the initial synthesis model based on classification loss and distribution loss to obtain a speech synthesis model, thereby further improving the accuracy and reliability of the speech synthesis model.
[0089] Based on the above embodiments, Figure 3 This is a flowchart illustrating step 230 in the training method of the speech synthesis model provided by the present invention, as shown below. Figure 3 As shown, step 230 includes:
[0090] Step 231: Perform a weighted fusion of the classification loss and the distribution loss to obtain the fusion loss;
[0091] Step 232: Based on the fusion loss, perform parameter iteration on the initial synthesis model to obtain the speech synthesis model.
[0092] Specifically, after obtaining the classification loss and the distribution loss, the classification loss and the distribution loss can be weighted and fused to obtain the fusion loss, as shown in the following formula:
[0093] L = L nll +α·L ce
[0094] Among them, L nll L represents the distributed loss. ce Let L represent the classification loss, L represent the fusion loss, and α represent the weighting parameter.
[0095] After obtaining the fusion loss, the parameters of the initial synthesis model can be iterated based on the fusion loss, and the initial synthesis model after parameter iteration can be determined as the speech synthesis model.
[0096] The method provided in this embodiment of the invention iterates the parameters of the initial synthesis model based on fusion loss to obtain a speech synthesis model. The fusion loss is obtained by weighted fusion of classification loss and distribution loss, thereby improving the accuracy and reliability of the speech synthesis model.
[0097] Based on the above embodiments, Figure 4 This is a flowchart illustrating step 130 in the training method of the speech synthesis model provided by the present invention, as shown below. Figure 4 As shown, step 130 includes:
[0098] Based on the phoneme classifier, the predicted acoustic features are classified into phonemes to obtain the phoneme classification results of the predicted acoustic features;
[0099] The phoneme classifier is trained based on the acoustic features of samples carrying phoneme labels, and the configuration parameters of the sample acoustic features are consistent with those of the predicted acoustic features.
[0100] Specifically, after obtaining the predicted acoustic features, the predicted acoustic features can be classified into phonemes based on a phoneme classifier to obtain the phoneme classification results of the predicted acoustic features.
[0101] The phoneme classifier here may include a CNN network and a Transformer model, or it may include a CNN network, a Transformer model and a fully connected layer. This embodiment of the invention does not specifically limit this.
[0102] The Transformer model here may include 12 Transformer modules, but this embodiment of the invention does not specifically limit this. The parameters of the phoneme classifier here can be randomly generated or preset.
[0103] Before training the phoneme classifier, sample acoustic features carrying phoneme labels can be collected in advance, and an initial phoneme classifier can be built in advance. The configuration parameters of the sample acoustic features and the predicted acoustic features are consistent. That is, the sample acoustic features and the predicted acoustic features can be configured with the same window length parameter, the same window shift parameter, or the same window length parameter and window shift parameter, etc. The embodiments of the present invention do not make specific limitations on this.
[0104] During the training process of the phoneme classifier, the acoustic features of the samples carrying phoneme labels can be input into the initial phoneme classifier, which then performs phoneme classification on the acoustic features of the samples to obtain the phoneme classification results of the acoustic features of the samples.
[0105] After obtaining the phoneme classification results of the sample acoustic features, a loss function value can be determined based on the phoneme classification results and phoneme labels. This loss function value reflects the difference between the phoneme classification results and phoneme labels of the sample acoustic features. The loss function here can be either the cross-entropy loss function or the mean squared error loss function; this embodiment of the invention does not specifically limit its application.
[0106] It is understandable that the greater the difference between the phoneme classification result and the phoneme label of the sample acoustic features, the larger the loss function value; the smaller the difference between the phoneme classification result and the phoneme label of the sample acoustic features, the smaller the loss function value.
[0107] After obtaining the loss function value, the initial phoneme classifier can be iterated based on the loss function value, and the initial phoneme classifier after parameter iteration can be determined as the phoneme classifier.
[0108] Therefore, based on the phoneme classifier, the predicted acoustic features are classified into phonemes, and the phoneme classification results of the predicted acoustic features are obtained, which improves the phoneme classification effect of the subsequent speech synthesis model.
[0109] The method provided in this embodiment of the invention classifies predicted acoustic features based on a phoneme classifier, thereby obtaining the phoneme classification results of the predicted acoustic features. This improves the phoneme classification effect of the subsequent speech synthesis model. Furthermore, the configuration parameters of the sample acoustic features and the predicted acoustic features are consistent, which improves the training efficiency of the phoneme classifier.
[0110] Considering that after training the speech synthesis model, it can be applied.
[0111] Based on the above embodiments, the present invention provides a speech synthesis method. Figure 5 This is a flowchart illustrating the speech synthesis method provided by the present invention, as shown below. Figure 5 As shown, the method includes:
[0112] Step 510: Obtain the text to be synthesized.
[0113] Specifically, the text to be synthesized can be obtained. The text to be synthesized here is the text that needs to be synthesized into speech later. The text to be synthesized can be directly input by the user, or it can be obtained by transcribing the collected audio into speech, or it can be obtained by acquiring images through image acquisition devices such as scanners, mobile phones, cameras, tablets, etc., and performing OCR (Optical Character Recognition) on the images. This embodiment of the invention does not make specific limitations on this.
[0114] Step 520: Based on the speech synthesis model, perform speech synthesis on the text to be synthesized;
[0115] The speech synthesis model is obtained by iterating the parameters of an initial synthesis model based on the phoneme classification results of the sample text and the phoneme tags carried by the corresponding sample speech. The phoneme classification results are obtained by classifying the predicted acoustic features of the sample text synthesized based on the initial synthesis model.
[0116] Specifically, after obtaining the text to be synthesized, speech synthesis can be performed on the text based on the speech synthesis model.
[0117] Before training the speech synthesis model, sample texts and phoneme tags carried by the corresponding sample speech can be collected in advance, and an initial synthesis model can be built in advance. The parameters of the initial synthesis model can be randomly generated or pre-set.
[0118] Specifically, during the training process of the speech synthesis model, sample text can be input into the initial synthesis model, which then performs speech synthesis on the sample text to obtain the predicted acoustic features of the sample text. Here, the initial synthesis model can use a Normalizing-Flow model or a VAE (Variational Autoencoder) to perform speech synthesis on the sample text; this embodiment of the invention does not impose any specific limitations on this.
[0119] The predicted acoustic features here reflect acoustic-level features. These features can be Mel Frequency Cepstrum Coefficient (MFCC) features or Perceptual Linear Predictive (PLP) features, etc. This embodiment of the invention does not impose any specific limitations on them.
[0120] After obtaining the predicted acoustic features of the sample text, phoneme classification can be performed on the predicted acoustic features to obtain the phoneme classification results of the predicted acoustic features.
[0121] After obtaining the phoneme classification results, the classification loss can be determined based on the phoneme classification results of the sample text and the phoneme labels carried by the corresponding sample speech. The classification loss reflects the difference between the phoneme classification results of the sample text and the phoneme labels carried by the corresponding sample speech.
[0122] It is understandable that the greater the difference between the phoneme classification result of the sample text and the phoneme label carried by the corresponding sample speech, the greater the classification loss; the smaller the difference between the phoneme classification result of the sample text and the phoneme label carried by the corresponding sample speech, the smaller the classification loss.
[0123] After obtaining the classification loss, the parameters of the initial synthesis model can be iterated based on the classification loss, and the initial synthesis model after parameter iteration can be determined as the speech synthesis model.
[0124] Therefore, the determined speech synthesis model is a model with speech synthesis capabilities, which improves the accuracy and reliability of speech synthesis.
[0125] The method provided in this invention is based on a speech synthesis model to synthesize speech from text. The speech synthesis model is obtained by iterating the parameters of an initial synthesis model based on the phoneme classification results of the sample text and the phoneme tags carried by the corresponding sample speech. The phoneme classification results are obtained by classifying the predicted acoustic features of the sample text synthesized based on the initial synthesis model. This improves the accuracy and reliability of the speech synthesis model.
[0126] Based on any of the above embodiments, a training method for a speech synthesis model includes the following steps:
[0127] The first step is to obtain the sample text and the phoneme labels of the corresponding sample speech.
[0128] The second step involves synthesizing speech from the sample text based on the initial synthesis model to obtain the predicted acoustic features of the sample text.
[0129] The third step is to classify the predicted acoustic features based on the phoneme classifier to obtain the phoneme classification results of the predicted acoustic features. The phoneme classifier here is trained based on the sample acoustic features carrying phoneme labels, and the configuration parameters of the sample acoustic features are consistent with those of the predicted acoustic features.
[0130] The fourth step is to determine the classification loss based on the difference between the phoneme classification results and the phoneme labels.
[0131] The fifth step is to determine the distribution loss based on the differences between text encoding and speech encoding.
[0132] The sixth step is to perform a weighted fusion of the classification loss and the distribution loss to obtain the fusion loss.
[0133] Step 7: Based on the fusion loss, iterate the parameters of the initial synthesis model to obtain the speech synthesis model.
[0134] The training apparatus for the speech synthesis model provided by the present invention will be described below. The training apparatus for the speech synthesis model described below can be referred to in correspondence with the training method for the speech synthesis model described above.
[0135] Based on any of the above embodiments, the present invention provides a training device for a speech synthesis model. Figure 6 This is a schematic diagram of the structure of the training device for the speech synthesis model provided by the present invention, as shown below. Figure 6 As shown, the device includes:
[0136] The acquisition unit 610 is used to acquire sample text and sample speech corresponding to the sample text, wherein the sample speech carries phoneme tags.
[0137] The speech synthesis unit 620 is used to perform speech synthesis on the sample text based on an initial synthesis model to obtain the predicted acoustic features of the sample text.
[0138] Phoneme classification unit 630 is used to classify the predicted acoustic features into phonemes to obtain the phoneme classification results of the predicted acoustic features;
[0139] The parameter iteration unit 640 is used to perform parameter iteration on the initial synthesis model based on the phoneme classification results and the phoneme labels to obtain the speech synthesis model.
[0140] The apparatus provided in this embodiment of the invention iterates the parameters of the initial synthesis model based on the phoneme classification results and phoneme labels. As a result, the obtained speech synthesis model has better speech synthesis effect, which improves the accuracy and reliability of the speech synthesis model. At the same time, the phoneme classification results are obtained by classifying the predicted acoustic features into phonemes, which further improves the phoneme classification effect of the speech synthesis model.
[0141] Based on any of the above embodiments, the parameter iteration unit is specifically used for:
[0142] The parameter iteration subunit is used to perform parameter iteration on the initial synthesis model based on the phoneme classification results and the phoneme labels, as well as the text encoding of the sample text and the speech encoding of the sample speech, to obtain the speech synthesis model;
[0143] The initial synthesis model includes a cascaded initial encoder and an initial decoder. The text encoding is obtained by encoding the sample text based on the initial encoder, and the speech encoding is obtained by inverse encoding the sample speech based on the initial decoder.
[0144] Based on any of the above embodiments, the parameter iteration subunit is specifically used for:
[0145] A classification loss unit is defined to determine the classification loss based on the difference between the phoneme classification result and the phoneme label;
[0146] A distribution loss unit is determined to determine the distribution loss based on the difference between the text encoding and the speech encoding;
[0147] An iterative unit is used to iterate the parameters of the initial synthesis model based on the classification loss and the distribution loss to obtain the speech synthesis model.
[0148] Based on any of the above embodiments, the iteration unit is specifically used for:
[0149] A fusion loss unit is determined to perform a weighted fusion of the classification loss and the distribution loss to obtain the fusion loss;
[0150] An iterative subunit is used to iterate the parameters of the initial synthesis model based on the fusion loss to obtain the speech synthesis model.
[0151] Based on any of the above embodiments, the phoneme classification unit is specifically used for:
[0152] Based on the phoneme classifier, the predicted acoustic features are classified into phonemes to obtain the phoneme classification results of the predicted acoustic features;
[0153] The phoneme classifier is trained based on the acoustic features of samples carrying phoneme labels, and the configuration parameters of the sample acoustic features are consistent with those of the predicted acoustic features.
[0154] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute a training method for a speech synthesis model. This method includes: acquiring sample text and phoneme labels of sample speech corresponding to the sample text; performing speech synthesis on the sample text based on an initial synthesis model to obtain predicted acoustic features of the sample text; classifying the predicted acoustic features into phonemes to obtain phoneme classification results; and iterating the parameters of the initial synthesis model based on the phoneme classification results and the phoneme labels to obtain the speech synthesis model.
[0155] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0156] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the training method of the speech synthesis model provided by the above methods. The method includes: acquiring sample text and phoneme labels of sample speech corresponding to the sample text; performing speech synthesis on the sample text based on an initial synthesis model to obtain predicted acoustic features of the sample text; classifying the predicted acoustic features by phonemes to obtain phoneme classification results of the predicted acoustic features; and iterating the parameters of the initial synthesis model based on the phoneme classification results and the phoneme labels to obtain the speech synthesis model.
[0157] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a training method for the speech synthesis model provided by the above methods. The method includes: acquiring sample text and phoneme labels of sample speech corresponding to the sample text; performing speech synthesis on the sample text based on an initial synthesis model to obtain predicted acoustic features of the sample text; classifying the predicted acoustic features by phonemes to obtain phoneme classification results of the predicted acoustic features; and iterating the parameters of the initial synthesis model based on the phoneme classification results and the phoneme labels to obtain the speech synthesis model.
[0158] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0159] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A training method for a speech synthesis model, characterized in that, include: Obtain the sample text and the phoneme tags of the sample speech corresponding to the sample text; Based on the initial synthesis model, speech synthesis is performed on the sample text to obtain the predicted acoustic features of the sample text; The predicted acoustic features are classified based on a phoneme classifier to obtain the phoneme classification result of the predicted acoustic features. The phoneme classifier is trained based on the sample acoustic features carrying phoneme labels. The sample acoustic features and the predicted acoustic features are configured with the same window length parameter and window shift parameter. Based on the phoneme classification results and phoneme labels, as well as the text encoding of the sample text and the speech encoding of the sample speech, the initial synthesis model is iterated to obtain a speech synthesis model. The initial synthesis model includes a cascaded initial encoder and an initial decoder. The text encoding is obtained by encoding the sample text based on the initial encoder, and the speech encoding is obtained by inverse encoding the sample speech based on the initial decoder. The mean of the prior Gaussian distribution is obtained based on the text encoding, and the speech encoding distribution follows the Gaussian distribution of the mean. The step of iterating the parameters of the initial synthesis model based on the phoneme classification results and phoneme labels, as well as the text encoding of the sample text and the speech encoding of the sample speech, to obtain a speech synthesis model includes: Based on the difference between the phoneme classification results and the phoneme labels, the classification loss is determined; The distribution loss is determined based on the difference between the text encoding and the speech encoding; Based on the classification loss and the distribution loss, the initial synthesis model is iterated to obtain the speech synthesis model.
2. The training method for the speech synthesis model according to claim 1, characterized in that, The process of iterating the parameters of the initial synthesis model based on the classification loss and the distribution loss to obtain the speech synthesis model includes: The classification loss and the distribution loss are weighted and fused to obtain the fusion loss; Based on the fusion loss, the parameters of the initial synthesis model are iterated to obtain the speech synthesis model.
3. A speech synthesis method, characterized in that, include: Obtain the text to be synthesized; Based on the speech synthesis model, speech synthesis is performed on the text to be synthesized. The speech synthesis model is obtained by iteratively processing parameters of an initial synthesis model based on phoneme classification results and phoneme labels, as well as text encoding of sample text and speech encoding of sample speech. The initial synthesis model includes a cascaded initial encoder and an initial decoder. The text encoding is obtained by encoding the sample text based on the initial encoder, and the speech encoding is obtained by inverse encoding the sample speech based on the initial decoder. The mean of a prior Gaussian distribution is obtained based on the text encoding, and the speech encoding distribution follows a Gaussian distribution of the mean. The phoneme classification result is obtained by the phoneme classifier classifying the predicted acoustic features of the sample text synthesized based on the initial synthesis model. The phoneme classifier is trained based on the sample acoustic features carrying phoneme labels. The sample acoustic features and the predicted acoustic features are configured with the same window length parameter and window shift parameter. The method of iterating the parameters of the initial synthesis model based on the phoneme classification results and the phoneme labels, as well as the text encoding of the sample text and the speech encoding of the sample speech, includes: determining the classification loss based on the difference between the phoneme classification results and the phoneme labels; determining the distribution loss based on the difference between the text encoding and the speech encoding; and iterating the parameters of the initial synthesis model based on the classification loss and the distribution loss to obtain the speech synthesis model.
4. A training device for a speech synthesis model, characterized in that, include: The acquisition unit is used to acquire sample text and sample speech corresponding to the sample text, wherein the sample speech carries phoneme tags. A speech synthesis unit is used to perform speech synthesis on the sample text based on an initial synthesis model to obtain the predicted acoustic features of the sample text. A phoneme classification unit is used to classify the predicted acoustic features based on a phoneme classifier to obtain the phoneme classification result of the predicted acoustic features. The phoneme classifier is trained based on sample acoustic features carrying phoneme labels. The sample acoustic features and the predicted acoustic features are configured with the same window length parameter and window shift parameter. A parameter iteration unit is used to perform parameter iteration on the initial synthesis model based on the phoneme classification result and the phoneme label, as well as the text encoding of the sample text and the speech encoding of the sample speech, to obtain a speech synthesis model. This includes: determining a classification loss based on the difference between the phoneme classification result and the phoneme label; determining a distribution loss based on the difference between the text encoding and the speech encoding; and performing parameter iteration on the initial synthesis model based on the classification loss and the distribution loss to obtain the speech synthesis model. The initial synthesis model includes a cascaded initial encoder and an initial decoder. The text encoding is obtained by encoding the sample text based on the initial encoder, and the speech encoding is obtained by inverse encoding the sample speech based on the initial decoder. The mean of a prior Gaussian distribution is obtained based on the text encoding, and the speech encoding distribution follows a Gaussian distribution of the mean.
5. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the training method of the speech synthesis model as described in any one of claims 1 to 2 or the speech synthesis method as described in claim 3.
6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the training method of the speech synthesis model as described in any one of claims 1 to 2 or the speech synthesis method as described in claim 3.
Citation Information
Patent Citations
Information synthesis method and device, electronic equipment and computer readable storage medium
CN112786005A
Training speech synthesis to generate different speech sounds
CN114787913A