Training Method, Device, Electronic Device and Storage Medium for Speech Synthesis Model
By introducing an adversarial network of generators and discriminators into the speech synthesis model, voice information is directly generated based on text information, and the problems of cumulative error and high training costs in traditional technologies are solved, achieving efficient, high-speed and high-quality speech synthesis effects.
Patent Information
- Application Number
- CN202210095968.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-01-26
AI Technical Summary
In traditional speech synthesis technology, the acoustic model and the vocoder need to be trained separately, resulting in cumulative errors, low speech quality, and a large number of training samples are required to make the model converge, which increases the training cost and time.
A training method of speech synthesis model is adopted, without introducing intermediate conversion links such as the Mel spectrum, and directly generate speech information based on text information. The overall training is carried out through the adversarial network of the generator and discriminator to eliminate cumulative errors, and only a certain amount of training samples can be used to train the model to convergence.
It improves the speech synthesis effect of the speech synthesis model, reduces training costs and time, and achieves high-speed, high-quality and high-efficiency training.
Smart Images

Figure CN114512112B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of natural language processing, and in particular, to a method, apparatus, electronic device, and storage medium for training a speech synthesis model. Background Art
[0002] Speech synthesis technology can convert text information into corresponding speech information. Traditional speech synthesis technologies generally include speech synthesis technology based on statistical parameter modeling (also known as parametric synthesis), speech synthesis technology based on unit selection and waveform splicing (also known as concatenative synthesis), and speech synthesis technology based on neural network models. Among them, the speech synthesis technology based on neural network models first converts text information into specific acoustic features, such as Mel spectrogram, and then uses a vocoder to convert acoustic features such as Mel spectrogram into audio waveform data, so as to realize the conversion from text information to speech information.
[0003] However, the inventors of the present application found that the speech synthesis technology that first converts text information into specific acoustic features and then converts the acoustic features into audio waveform data is a two-stage technology. Converting text information into specific acoustic features is achieved by an acoustic model, and converting acoustic features into audio waveform data is achieved by a vocoder. The acoustic model and the vocoder need to be trained separately. When the trained acoustic model and vocoder are connected in series to form a speech synthesis system, it will bring cumulative errors, resulting in low quality of the generated speech information. If the acoustic model and the vocoder are first spliced together and then trained, a large amount of training samples are required to make the model converge, greatly increasing the training cost and training time. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a method, apparatus, electronic device, and storage medium for training a speech synthesis model, which does not introduce intermediate conversion links such as Mel spectrogram, eliminates cumulative errors, improves the speech synthesis effect of the model, and only requires a certain amount of training samples to train the model to converge, reducing the training cost.
[0005] To solve the above technical problems, an embodiment of the present application provides a method for training a speech synthesis model. The speech synthesis model includes a generator and a discriminator. The training method includes the following steps: obtaining a plurality of first audio data labeled with text information to generate a training sample set; wherein, the text information is used to characterize the content of the first audio data; inputting the text information into the generator to obtain second audio data output by the generator; inputting the first audio data and the second audio data into the discriminator to obtain a discrimination result output by the discriminator; wherein, the discrimination result is used to characterize the similarity degree between the first audio data and the second audio data; and iteratively training the speech synthesis model according to the discrimination result and a preset loss function.
[0006] An embodiment of the present application further provides a training device for a speech synthesis model. The device includes a construction module for constructing a speech synthesis model, the speech synthesis model including a generator and a discriminator; an acquisition module for obtaining a plurality of first audio data labeled with text information to generate a training sample set, wherein the text information is used to characterize the content of the first audio data; and a training module for inputting the text information into the generator to obtain second audio data output by the generator, inputting the first audio data and the second audio data into the discriminator to obtain a discrimination result output by the discriminator, and iteratively training the speech synthesis model according to the discrimination result and a preset loss function, wherein the discrimination result is used to characterize the similarity degree between the first audio data and the second audio data.
[0007] An embodiment of the present application further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method for training a speech synthesis model.
[0008] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, and the computer program, when executed by a processor, implements the above method for training a speech synthesis model.
[0009] The training method, device, electronic device, and storage medium for a voice synthesis model provided by an embodiment of the present application. The voice synthesis model includes a generator and a discriminator. When training the voice synthesis model, first obtain a number of first audio data labeled with text information to generate a training sample set. The text information labeled on the first audio data can represent the content of the first audio data. The server inputs the text information into the generator to obtain the second audio data output by the generator, and then inputs both the first audio data and the second audio data into the discriminator to obtain the discrimination result output by the discriminator, which is used to represent the similarity degree between the first audio data and the second audio data. Finally, according to the discrimination result and a preset loss function, the entire voice synthesis model is iteratively trained. Considering the two-stage voice synthesis technology, the acoustic model and the vocoder need to be trained separately. When the acoustic model and the vocoder are cascaded into a voice synthesis system later, cumulative errors will occur. However, if the acoustic model and the vocoder are first spliced together and then trained, a large amount of training samples are required to make the model converge. In the embodiment of the present application, without introducing intermediate conversion links such as Mel spectrograms, the generator can directly generate voice information based on text information, and can train the entire voice synthesis model, naturally eliminating cumulative errors and improving the voice synthesis effect of the model. At the same time, based on the principle of the adversarial network, the discriminator can effectively judge the data generated by the generator, and only a certain amount of training samples are required to train the voice synthesis model to convergence, reducing the training cost.
[0010] In addition, the training sample set includes K training samples, where K is an integer greater than 1. The iterative training of the voice synthesis model according to the discrimination result and the preset loss function includes: fixing the discriminator, selecting N training samples from the training sample set, and performing N training sessions on the generator according to the discrimination results of the N training samples and the preset loss function; where N is an integer greater than 0 and less than K; obtaining the remaining K - N training samples in the training sample set, and performing alternating training on the generator and the discriminator according to the discrimination results of the remaining K - N training samples and the loss function. In the embodiment of the present application, when performing iterative training on the voice synthesis model, first fix the parameters of the discriminator and perform N training sessions on the generator alone, so that the generator can quickly obtain the basic ability to generate audio data according to text information, and then perform alternating training on the generator and the discriminator to gradually stabilize and improve the voice synthesis ability of the voice synthesis model, realizing high-speed, high-quality, and high-efficiency training of the voice synthesis model.
[0011] In addition, iteratively training the speech synthesis model according to the discrimination result and a preset loss function includes: obtaining a verification metric of the speech synthesis model according to the discrimination result and the preset loss function; wherein, the verification metric includes the accuracy rate, recall rate, and F1 value of the speech synthesis model; optimizing the parameters of the speech synthesis model according to the verification metric; wherein, the optimized parameters of the speech synthesis model at least include the number of convolution kernels, the window value of the convolution kernel, the L2 value, and the learning rate. The accuracy rate, recall rate, and F1 value can well measure the training effect of the model. Updating parameters such as the number of convolution kernels, the window value of the convolution kernel, the L2 value, and the learning rate of the speech synthesis model according to these verification metrics can well improve the speech synthesis effect of the speech synthesis model.
[0012] In addition, the preset loss function includes a first loss term and a second loss term. The first loss term is a feature matching loss, and the second loss term is a multi-scale short-time Fourier transform loss. The multi-scale short-time Fourier transform loss includes a full-band multi-scale short-time Fourier transform loss and a sub-band multi-scale short-time Fourier transform loss. Considering that although the feature matching loss can stabilize the training of the entire speech synthesis model, the feature matching loss alone cannot effectively measure the similarity and difference between the first audio data (real audio) and the second audio data (predicted audio). Therefore, in addition to using the conventional feature matching loss, the embodiments of the present application also add a multi-scale short-time Fourier transform loss that can scientifically and accurately measure the similarity and difference between the real audio and the predicted audio, further improving the training effect of the speech synthesis model.
[0013] In addition, obtaining a training sample set by obtaining a plurality of first audio data labeled with text information includes: obtaining a plurality of first audio data labeled with text information; filtering the text information according to a preset filtering rule to remove symbol information and / or redundant information in the text information; converting the filtered text information into a vector sequence according to a preset pronunciation rule and the content of the filtered text information; generating a training sample set according to the vector sequence and the first audio data corresponding to the vector sequence. Filtering the obtained first audio data to remove meaningless symbol information and duplicate content therein, and converting it into a vector sequence according to the pronunciation rule. Subsequently, directly inputting the vector sequence into the generator can facilitate the generator to generate the second audio data, and at the same time can improve the training effect on the generator and the entire speech synthesis model.
[0014] In addition, after obtaining a plurality of first audio data marked with text information to generate a training sample set, the following steps are included: selecting a plurality of training samples from the training sample set as verification samples and generating a verification sample set; after iteratively training the speech synthesis model, the following steps are included: inputting the text information marked on the verification samples into the generator of the trained speech synthesis model to obtain third speech data output by the generator; performing a Mean Opinion Score (MOS) on the third speech data to obtain a MOS value; if the MOS value is lower than a preset threshold, re-train the speech synthesis model. MOS scoring is a scientific and reasonable speech quality scoring method. In this embodiment, after the model training is completed, the verification samples are used to verify the model to obtain the MOS value. If the MOS value is too low, it indicates that the model training effect is not good. Re-training the speech synthesis model can effectively improve the model training effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings, and these exemplary illustrations do not constitute a limitation on the embodiments.
[0016] Figure 1 is the flowchart of the training method of the speech synthesis model according to an embodiment of the present application Figure 1 ;
[0017] Figure 2 is the schematic diagram of a generator of a speech synthesis model provided in an embodiment of the present application;
[0018] Figure 3 is the schematic diagram of a discriminator of a speech synthesis model provided in an embodiment of the present application;
[0019] Figure 4 is the flowchart of iteratively training the speech synthesis model according to the discriminant result and a preset loss function in an embodiment of the present application;
[0020] Figure 5 is the flowchart of alternately training the generator and the discriminator according to the discriminant results and loss functions of the remaining K - N training samples in an embodiment of the present application;
[0021] Figure 6 is the flowchart of obtaining a plurality of first audio data marked with text information to generate a training sample set in an embodiment of the present application;
[0022] Figure 7 is the flowchart of the training method of the speech synthesis model according to another embodiment of the present application Figure 2 ;
[0023] Figure 8 It is a schematic structural diagram of a training device for a speech synthesis model according to another embodiment of the present application;
[0024] Figure 9 It is a schematic structural diagram of an electronic device according to another embodiment of the present application. Detailed implementation manners
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in the embodiments of the present application, many technical details are provided for readers to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions required to be protected by the present application can be implemented. The following division of each embodiment is for convenience of description and should not constitute any limitation on the specific implementation manner of the present application. Each embodiment can be combined and cross-referenced with each other on the premise of not conflicting with each other.
[0026] An embodiment of the present application relates to a method for training a speech synthesis model, which is applied to an electronic device. Among them, the electronic device can be a terminal or a server. In this embodiment and the following embodiments, the electronic device is described by taking the server as an example. The implementation details of the method for training the speech synthesis model in this embodiment will be specifically described below. The following content is only implementation details provided for convenient understanding and is not necessary for implementing the solution.
[0027] The specific process of the method for training the speech synthesis model in this embodiment can be as Figure 1 shown and includes:
[0028] Step 101, obtaining a plurality of first audio data marked with text information to generate a training sample set.
[0029] In a specific implementation, when the server trains the speech synthesis model, it first obtains a plurality of first audio data marked with text information as training samples, thereby generating a training sample set. Among them, each training sample, that is, the text information marked on each first audio data, is used to represent the content of the first audio data.
[0030] In an example, the content of a piece of first audio data is "What's the weather like tomorrow", and the text information marked on the first audio data is the text form of "What's the weather like tomorrow".
[0031] In an example, the server can obtain 100,000 pieces of first audio data marked with text information, randomly select 90,000 pieces of the first audio data as training samples to generate a training sample set, and use the remaining 10,000 pieces of first audio data as verification samples to generate a verification sample set.
[0032] Step 102: Input the text information into the generator to obtain the second audio data output by the generator.
[0033] Specifically, after generating the training sample set, the server can input the training samples, that is, the text information annotated on the first audio data, into the generator of the speech synthesis model to obtain the second audio data generated and output by the generator according to the text information. The first audio data can be regarded as the real audio data, so the second audio data can be called the predicted audio data.
[0034] In one example, the schematic structural diagram of the generator of the speech synthesis model can be as Figure 2 shown. The generator includes units such as a phoneme inserter, an encoder, a variance predictor, a decoder, several convolutional layers, an upsampling layer, and a residual network. The input of the generator is the text information annotated on the first audio data, and the output of the generator is the second audio data.
[0035] Step 103: Input the first audio data and the second audio data into the discriminator to obtain the discrimination result output by the discriminator.
[0036] In a specific implementation, after the server obtains the second audio data generated by the generator according to the text information annotated on the training samples from the generator, it can input both the second audio data and the first audio data into the discriminator of the speech synthesis model to obtain the discrimination result output by the discriminator. Among them, the discrimination result output by the discriminator is used to represent the similarity degree between the first audio data and the second audio data, that is, the similarity degree between the real audio data and the predicted audio data, that is, the discriminator can identify the first audio data and the second audio data.
[0037] In one example, the schematic structural diagram of the discriminator of the speech synthesis model can be as Figure 3 shown. The discriminator includes a downsampling layer and several convolutional layers. The input of the discriminator is the first audio data and the second audio data, and the output of the discriminator is the discrimination result of the discriminator on the first audio data and the second audio data.
[0038] In a specific implementation, the speech synthesis model is built based on a generative adversarial network (GAN), including a generator and a discriminator. The generator is connected to the discriminator. The input of the generator is text information, and the output is second audio data. The input of the discriminator is the first audio data and the second audio data, and the output is the discrimination result of the similarity degree between the first audio data and the second audio data. The generator creates realistic data to deceive the discriminator, and the discriminator continuously discriminates between the realistic data and the real data, and performs iterative training in this way until the realistic data created by the generator can deceive the discriminator, that is, the discriminator determines that the data generated by the generator is real data, so as to obtain a generator that meets the requirements.
[0039] Step 104, perform iterative training on the speech synthesis model according to the discrimination result and a preset loss function.
[0040] In a specific implementation, after the server obtains the discrimination result from the discriminator, it can perform iterative training on the entire speech synthesis model according to the discrimination result and a preset loss function. Among them, the preset loss function can be set by those skilled in the art according to actual needs.
[0041] In an example, the server can obtain the verification metrics of the speech synthesis model according to the discrimination result output by the discriminator and a preset loss function. Among them, the verification metrics of the speech synthesis model include the accuracy rate, recall rate, and F1 value of the speech synthesis model. Among them, F1 value = (accuracy rate * recall rate * 2) / (accuracy rate + recall rate). After the server obtains the verification metrics of the speech synthesis model, it can optimize the parameters of the speech synthesis model according to these verification metrics. The parameters of the speech synthesis model at least include the number of convolutional kernels, the window value of the convolutional kernel, the L2 value of the weight decay norm, and the learning rate, etc. The accuracy rate, recall rate, and F1 value can well measure the training effect of the model. Updating the parameters such as the number of convolutional kernels, the window value of the convolutional kernel, the L2 value, and the learning rate of the speech synthesis model according to these verification metrics can well improve the speech synthesis effect of the speech synthesis model.
[0042] In this embodiment, the speech synthesis model includes a generator and a discriminator. When training the speech synthesis model, first, a number of first audio data marked with text information are obtained to generate a training sample set. The text information marked on the first audio data can represent the content of the first audio data. The server inputs the text information into the generator to obtain the second audio data output by the generator. Then, both the first audio data and the second audio data are input into the discriminator to obtain the discriminant result output by the discriminator, which is used to represent the similarity between the first audio data and the second audio data. Finally, according to the discriminant result and a preset loss function, the entire speech synthesis model is iteratively trained. Considering the two-stage speech synthesis technology, the acoustic model and the vocoder need to be trained separately. When the acoustic model and the vocoder are cascaded into a speech synthesis system later, cumulative errors will occur. However, if the acoustic model and the vocoder are first spliced together and then trained, a large amount of training samples are required to make the model converge. In the embodiment of this application, without introducing intermediate conversion links such as Mel spectrograms, the generator can directly generate speech information based on the text information, and the entire speech synthesis model can be trained, naturally eliminating the cumulative errors and improving the speech synthesis effect of the model. At the same time, based on the principle of the adversarial network, the discriminator can effectively judge the data generated by the generator, and only a certain amount of training samples are required to train the speech synthesis model to convergence, reducing the training cost.
[0043] In one embodiment, the training sample set generated by the server includes K training samples, where K is an integer greater than 1. The server iteratively trains the speech synthesis model according to the discriminant result and a preset loss function, which can be implemented through the following steps: Figure 4 as shown, specifically including:
[0044] Step 201, fix the discriminator, select N training samples from the training sample set, and train the generator N times according to the discriminant results of the N training samples and a preset loss function.
[0045] In a specific implementation, when the server iteratively trains the speech synthesis model, it first fixes the parameters of the discriminator and only trains the generator N times. That is, the server first selects N training samples from the training sample set, where N is an integer greater than 0 and less than K. The server sequentially inputs these N training samples into the generator and the discriminator, and trains the generator N times according to the discriminant results of these N training samples and a preset loss function, so that the generator can quickly obtain the basic ability to generate audio data based on the text information.
[0046] In an example, the training sample set contains 90,000 training samples. The server randomly selects 10,000 training samples from them, fixes the parameters of the discriminator, and first trains the generator 10,000 times.
[0047] Step 202: Obtain the remaining K - N training samples in the training sample set, and alternately train the generator and the discriminator according to the discrimination results and loss function of the remaining K - N training samples.
[0048] In a specific implementation, after the server trains the generator N times, it can obtain the remaining K - N training samples in the training sample set, input these remaining K - N training samples into the generator and the discriminator, and alternately train the generator and the discriminator according to the discrimination results and loss function of these remaining K - N training samples.
[0049] In an example, the training sample set contains 90,000 training samples. The server randomly selects 10,000 training samples from them, fixes the parameters of the discriminator, first trains the generator 10,000 times, and then alternately trains the generator and the discriminator according to the remaining 80,000 training samples.
[0050] In an example, the server alternately trains the generator and the discriminator according to the discrimination results and loss function of the remaining K - N training samples, which can be implemented through the following steps: Figure 5 as shown below, specifically including:
[0051] Step 301: Fix the discriminator, obtain the i-th training sample among the K - N training samples, and train the generator according to the discrimination result and loss function of the i-th training sample.
[0052] Step 302: Fix the generator, obtain the (i + 1)-th training sample among the K - N training samples, and train the discriminator according to the discrimination result and loss function of the (i + 1)-th training sample.
[0053] In a specific implementation, when the server alternately trains the generator and the discriminator, it can first fix the parameters of the discriminator, obtain the i-th training sample among the K - N training samples, input this i-th training sample into the generator and the discriminator. The server trains the generator once according to the discrimination result and loss function of the i-th training sample. Then the server fixes the parameters of the generator, obtains the (i + 1)-th training sample among the K - N training samples. The server inputs the (i + 1)-th training sample into the generator and the discriminator. The server trains the discriminator once according to the discrimination result and loss function of the (i + 1)-th training sample. Repeat this process until all K - N training samples are involved in the training. Alternately training the generator and the discriminator can gradually stabilize and improve the speech synthesis ability of the speech synthesis model, and achieve high - speed, high - quality, and high - efficiency training of the speech synthesis model.
[0054] In this embodiment, when iteratively training the speech synthesis model, first fix the parameters of the discriminator and train the generator alone for N times, so that the generator can quickly obtain the basic ability to generate audio data according to text information. Then, alternately train the generator and the discriminator to gradually stabilize and improve the speech synthesis ability of the speech synthesis model, realizing high-speed, high-quality, and high-efficiency training of the speech synthesis model.
[0055] In one embodiment, the preset loss function used for training the speech synthesis model includes a first loss term and a second loss term. The first loss term is the feature matching loss, and the second loss term is the multi-scale short-time Fourier transform loss. The multi-scale short-time Fourier transform loss includes the multi-scale short-time Fourier transform loss of the full frequency band and the multi-scale short-time Fourier transform loss of the sub-frequency band. Considering that although the feature matching loss can stabilize the training of the entire speech synthesis model, the similarity and difference between the first audio data (real audio data) and the second audio data (predicted audio data) cannot be effectively measured only based on the feature matching loss. Therefore, in the embodiments of this application, in addition to using the conventional feature matching loss, the multi-scale short-time Fourier transform loss that can scientifically and accurately measure the similarity and difference between the real audio and the predicted audio is added, further improving the training effect of the speech synthesis model.
[0056] In one example, the second loss term can be expressed by the following formula:
[0057]
[0058] In the formula, is the multi-scale short-time Fourier transform loss of the full frequency band, is the multi-scale short-time Fourier transform loss of the sub-frequency band, and L 2 is the second loss term.
[0059] In one embodiment, the server obtains a training sample set by generating a number of first audio data labeled with text information, which can be realized through the steps shown in Figure 6 as follows, specifically including:
[0060] Step 401: Obtain a number of first audio data labeled with text information.
[0061] Step 402: Filter the text information according to the preset filtering rules to remove the symbol information and / or redundant information in the text information.
[0062] In a specific implementation, the text information annotated on the first audio data is basically machine-annotated. For example, the first audio data is input into a mature speech recognition model to obtain the text content of the first audio data. Machine annotation may identify symbol information and / or redundant information, which actually has no practical significance for speech synthesis. Therefore, after the server obtains a number of first audio data annotated with text information, it can filter the annotated text information according to a preset filtering rule to remove the symbol information and / or redundant information in the text information.
[0063] Step 403: Convert the filtered text information into a vector sequence according to the preset pronunciation rule and the content of the filtered text information.
[0064] Step 404: Generate a training sample set according to the vector sequence and the first audio data corresponding to the vector sequence.
[0065] In a specific implementation,
[0066] In this embodiment, after the server removes the symbol information and duplicate content without practical significance in the text information, it can convert the text information into a vector sequence according to the preset pronunciation rule. The server generates a training sample set according to the vector sequence and the first audio data corresponding to the vector sequence. Subsequently, directly inputting the vector sequence into the generator can facilitate the generator to generate the second audio data, and at the same time can improve the training effect on the generator and the entire speech synthesis model.
[0067] Another embodiment of the present application relates to a method for training a speech synthesis model. The implementation details of the method for training the speech synthesis model in this embodiment will be specifically described below. The following content is only the implementation details provided for convenience of understanding and is not necessary for implementing the solution. The specific process of the method for training the speech synthesis model in this embodiment can be as Figure 7 shown, including:
[0068] Step 501: Obtain a number of first audio data annotated with text information to generate a training sample set.
[0069] Among them, step 501 is substantially the same as step 101 and will not be elaborated here.
[0070] Step 502: Select a number of training samples from the training sample set as verification samples and generate a verification sample set.
[0071] Specifically, after the server generates the training sample set, it can extract a number of training samples from the training sample set as verification samples for testing the speech synthesis effect of the speech synthesis model, and generate a verification sample set.
[0072] Step 503: Input the text information into the generator to obtain the second audio data output by the generator.
[0073] Step 504: Input the first audio data and the second audio data into a discriminator to obtain the discrimination result output by the discriminator.
[0074] Step 505: Iteratively train the speech synthesis model according to the discrimination result and a preset loss function.
[0075] Among them, steps 503 to 505 are substantially the same as steps 102 to 104, and will not be elaborated here.
[0076] Step 506: Input the text information annotated on the verification sample into the generator of the trained speech synthesis model to obtain the third audio data output by the generator.
[0077] Step 507: Perform a Mean Opinion Score (MOS) rating on the third audio data to obtain a MOS value.
[0078] Step 508: If the MOS value is lower than a preset threshold, retrain the speech synthesis model.
[0079] In a specific implementation, after the server completes the iterative training of the speech synthesis model, it can use a verification sample set to verify the speech synthesis effect of the speech synthesis type. The server inputs the text information annotated on the verification sample into the generator of the trained speech synthesis model to obtain the third audio data output by the generator, and uses the Mean Opinion Score (MOS) rating method to rate the third audio data to obtain a MOS value. If the MOS value is lower than a preset threshold, retrain the speech synthesis model. MOS rating is a scientific and reasonable speech quality rating method. In this embodiment, after the model training is completed, the verification sample is used to verify the model to obtain a MOS value. If the MOS value is too low, it indicates that the model training effect is not good. Retraining the speech synthesis model can effectively improve the model training effect.
[0080] In an example, the server uses a two-stage speech synthesis model composed of an "acoustic model + vocoder" as the control group, and a speech synthesis model composed of a "generator + discriminator" as the experimental group, and uses three verification sample sets, namely verification sample set A, verification sample set B, and verification sample set C, to rate the experimental group and the control group. Among them, the total audio data duration of verification sample set A is 4.5 hours, the total audio data duration of verification sample set B is 10 hours, and the total audio data duration of verification sample set C is 50 hours. The MOS rating results of the control group and the experimental group are shown in Table 1 (the full score of MOS rating is 5 points):
[0081] Table 1: MOS rating results of the control group and the experimental group
[0082] Verification sample set Control group (acoustic model + vocoder) Experimental group (generator + discriminator) A 3.7 3.9 B 3.9 4.0 C 3.9 4.2
[0083] It can be seen that the speech synthesis effect of the end-to-end speech synthesis model composed of a "generator + discriminator" is significantly better than that of the two-stage speech synthesis model composed of an "acoustic model + vocoder".
[0084] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are within the protection scope of this patent; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process are within the protection scope of this patent.
[0085] Another embodiment of this application relates to a training device for a speech synthesis model. The implementation details of the training device for the speech synthesis model in this embodiment will be specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this solution. The schematic diagram of the training device for the speech synthesis model in this embodiment can be as Figure 8 shown. The device includes: a construction module 601, an acquisition module 602, and a training module 603, where the training module 603 is respectively connected to the construction module 601 and the acquisition module 602.
[0086] The construction module 601 is used to construct a speech synthesis model, and the speech synthesis model includes a generator and a discriminator.
[0087] Specifically, when using the training device for the speech synthesis model to train the speech synthesis model, the construction module 601 first constructs a speech synthesis model including a generator and a discriminator, and provides the constructed speech synthesis model to the training module 603. Among them, the generator of the speech synthesis model is connected to the discriminator. The input of the generator is text information, and the output is the second audio data. The input of the discriminator is the first audio data and the second audio data, and the output is the discrimination result of the similarity degree between the first audio data and the second audio data.
[0088] The acquisition module 602 is used to acquire a number of first audio data marked with text information to generate a training sample set, where the text information is used to characterize the content of the first audio data.
[0089] Specifically, after the construction module 601 constructs the speech synthesis model, the acquisition module 602 can acquire a number of first audio data marked with text information as training samples, so as to generate a training sample set, and send the training sample set to the training module 603. Among them, each training sample, that is, the text information marked on each first audio data, is used to characterize the content of the first audio data.
[0090] The training module 603 is configured to input text information into a generator, obtain second audio data output by the generator, input the first audio data and the second audio data into a discriminator, obtain a discrimination result output by the discriminator, and perform iterative training on the speech synthesis model according to the discrimination result and a preset loss function, where the discrimination result is used to characterize the similarity between the first audio data and the second audio data.
[0091] Specifically, after obtaining the speech synthesis model and the training sample set, the training module 603 may input the training samples, that is, the text information annotated on the first audio data, into the generator of the speech synthesis model, obtain the second audio data generated and output by the generator according to the text information, then input both the second audio data and the first audio data into the discriminator of the speech synthesis model, obtain the discrimination result output by the discriminator, and finally perform iterative training on the entire speech synthesis model according to the discrimination result and the preset loss function.
[0092] It is worth mentioning that each module involved in this embodiment is a logical module. In practical applications, a logical unit may be a physical unit, a part of a physical unit, or may be implemented as a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, units that are not closely related to solving the technical problems proposed in this application are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.
[0093] Another embodiment of this application relates to an electronic device, such as Figure 9 shown, including: at least one processor 701; and a memory 702 communicatively connected to the at least one processor 701; wherein, the memory 702 stores instructions executable by the at least one processor 701, and the instructions are executed by the at least one processor 701 to enable the at least one processor 701 to execute the training method of the speech synthesis model in the above embodiments.
[0094] Wherein, the memory and the processor are connected by a bus. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and the memory together. The bus may also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and therefore will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted over the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.
[0095] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor when executing operations.
[0096] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the method embodiments described above are implemented.
[0097] That is, those skilled in the art can understand that all or part of the steps of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium, including several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs, etc., which can store program codes.
[0098] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present application.
Claims
1. A training method for a speech synthesis model, characterized in that, the model includes a generator and a discriminator, and the training method includes: obtaining a training sample set by generating a number of first audio data labeled with text information; wherein, the text information is used to characterize the content of the first audio data; inputting the text information into the generator to obtain second audio data output by the generator; inputting the first audio data and the second audio data into the discriminator to obtain a discrimination result output by the discriminator; wherein, the discrimination result is used to characterize the similarity degree between the first audio data and the second audio data; iteratively training the speech synthesis model according to the discrimination result and a preset loss function; wherein, the training sample set includes K training samples, K is an integer greater than 1, and the iteratively training the speech synthesis model according to the discrimination result and the preset loss function includes: fixing the discriminator, selecting N training samples from the training sample set, and training the generator N times according to the discrimination results of the N training samples and the preset loss function; wherein, N is an integer greater than 0 and less than K; obtaining the remaining K-N training samples in the training sample set, and alternately training the generator and the discriminator according to the discrimination results of the remaining K-N training samples and the loss function.
2. The training method for a speech synthesis model according to claim 1, characterized in that, the alternately training the generator and the discriminator according to the discrimination results of the remaining K-N training samples and the loss function includes: fixing the discriminator, obtaining the i-th training sample among the K-N training samples, and training the generator according to the discrimination result of the i-th training sample and the loss function; fixing the generator, obtaining the (i + 1)-th training sample among the K-N training samples, and training the discriminator according to the discrimination result of the (i + 1)-th training sample and the loss function.
3. The training method for a speech synthesis model according to claim 1, characterized in that, the iteratively training the speech synthesis model according to the discrimination result and the preset loss function includes: obtaining a verification index of the speech synthesis model according to the discrimination result and the preset loss function; wherein, the verification index includes the accuracy rate, recall rate and F1 value of the speech synthesis model; optimizing the parameters of the speech synthesis model according to the verification index; wherein, the optimized parameters of the speech synthesis model at least include the number of convolution kernels, the window value of the convolution kernel, the weight decay norm L2 value and the learning rate.
4. The training method for a speech synthesis model according to any one of claims 1-3, characterized in that, The preset loss function includes a first loss term and a second loss term. The first loss term is a feature matching loss, and the second loss term is a multi-scale short-time Fourier transform loss. The multi-scale short-time Fourier transform loss includes a multi-scale short-time Fourier transform loss in the full frequency band and a multi-scale short-time Fourier transform loss in the sub-frequency band.
5. The training method of the speech synthesis model according to claim 4, wherein, the second loss term is represented by the following formula: Among them, is the multi-scale short-time Fourier transform loss of the full frequency band, is the multi-scale short-time Fourier transform loss of the sub-frequency band, and L 2 is the second loss term.
6. The training method of the speech synthesis model according to any one of claims 1-3, wherein, the obtaining of a plurality of first audio data labeled with text information to generate a training sample set includes: obtaining a plurality of first audio data labeled with text information; filtering the text information according to a preset filtering rule to remove symbol information and / or redundant information in the text information; converting the filtered text information into a vector sequence according to a preset pronunciation rule and the content of the filtered text information; generating a training sample set according to the vector sequence and the first audio data corresponding to the vector sequence.
7. The training method of the speech synthesis model according to claim 1, wherein, after obtaining a plurality of first audio data labeled with text information to generate a training sample set, it includes: selecting a plurality of training samples from the training sample set as validation samples and generating a validation sample set; after iteratively training the speech synthesis model, it includes: inputting the text information labeled on the validation samples into the generator of the trained speech synthesis model to obtain third speech data output by the generator; performing a mean opinion score (MOS) on the third speech data to obtain a MOS value; if the MOS value is lower than a preset threshold, retraining the speech synthesis model.
8. The training method of the speech synthesis model according to any one of claims 1-3, wherein, the speech synthesis model is built based on a generative adversarial network (GAN). The input of the generator is the text information, the output of the generator is the second audio data, the input of the discriminator is the first audio data and the second audio data, and the output of the discriminator is a discrimination result on the similarity between the first audio data and the second audio data.
9. A training device for a speech synthesis model, wherein, it includes: a construction module for constructing a speech synthesis model, which includes a generator and a discriminator; an acquisition module for obtaining a plurality of first audio data labeled with text information to generate a training sample set, where the text information is used to characterize the content of the first audio data; A training module, configured to input the text information into the generator, obtain second audio data output by the generator, input the first audio data and the second audio data into the discriminator, obtain a discrimination result output by the discriminator, and iteratively train the speech synthesis model according to the discrimination result and a preset loss function, wherein the discrimination result is used to represent the similarity degree between the first audio data and the second audio data; there are K training samples in the training sample set, K being an integer greater than 1, and the iteratively training the speech synthesis model according to the discrimination result and the preset loss function includes: fixing the discriminator, selecting N training samples from the training sample set, and training the generator N times according to the discrimination results of the N training samples and the preset loss function; wherein, N is an integer greater than 0 and less than K; obtaining the remaining K-N training samples in the training sample set, and alternately training the generator and the discriminator according to the discrimination results of the remaining K-N training samples and the loss function.
10. An electronic device, characterized in that it includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the training method of the speech synthesis model according to any one of claims 1 to 8.
11. A computer-readable storage medium storing a computer program, characterized in that when the computer program is executed by a processor, it implements the training method of the speech synthesis model according to any one of claims 1 to 8.
Citation Information
Patent Citations
Translation model training method, medium and device and computing equipment
CN110110337A
Speech synthesis method and device with mood, computing equipment and storage medium
CN111161703A
Speech synthesis model training method and device, terminal equipment and storage medium
CN112786003A
Speech synthesis method and device, and electronic equipment
CN112837670A