Speech synthesis method, model training method, device and storage medium

By using pre-trained sets and target duration prediction networks in the speech synthesis model and using noise-added speech samples in the fine-tuning stage, the problem of the model requiring a large amount of recording data and easy fitting noise in the prior art is solved, and a more natural and smooth speech synthesis effect is achieved.

CN114283783BActive Publication Date: 2025-06-13UNIV OF SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111674186.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-06-13
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

When building a model of a specific speaker, existing speech synthesis technology requires a large amount of recording data, and the model is easy to fit noise, and the synthesized speech has a serious mechanical feeling and is not smooth enough to fully match the speaker's tone and rhythm.

Method used

The preset speech synthesis model is pre-trained through the pre-training set, and the target duration prediction network corresponding to the target application scenario is used to replace the original duration prediction network to obtain the target speech synthesis model, and fine-tune it in the fine-tuning stage using a small number of speech samples from the target speaker and the noise-plused speech samples from the same type of speakers for fine-tuning.

Benefits of technology

It improves the generalization ability and robustness of the speech synthesis model, making the synthesized speech more natural and smooth, fits the speaking style of a specific speaker, and reduces the risk of overfitting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283783B_ABST
    Figure CN114283783B_ABST
Patent Text Reader

Abstract

The present application provides a speech synthesis method, a model training method, a device and a storage medium. The speech synthesis method includes: obtaining a text to be synthesized, inputting it into a trained target speech synthesis model to obtain a target speech corresponding to the text to be synthesized; wherein, the speech synthesis model is pre-trained according to a pre-training set to obtain the speech synthesis model; replacing the duration prediction network of the speech synthesis model with a target duration prediction network corresponding to a target application scenario to obtain a target speech synthesis model; obtaining a target training set, where the target training set includes speech samples of a target speaker; selecting speech samples of the same type of speaker as the target speaker from the pre-training set for mask noise addition to obtain noise-added speech samples; training the target speech synthesis model according to the target training set and the noise-added speech samples to obtain a trained target speech synthesis model. The present application can synthesize high-quality natural and fluent speech that is more in line with the speaking style of a specific speaker.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of speech synthesis, and in particular, to a speech synthesis method, a model training method, a device, and a storage medium. Background Art

[0002] Speech synthesis, also known as text-to-speech (TTS), aims to convert input text into smooth and natural output speech, and is a key technology for realizing intelligent human-machine speech interaction.

[0003] In traditional speech synthesis technology, to build a speech synthesis model for a specific speaker, 10-20 hours of recording data of this specific speaker is required, and the better the quality of the recording data, the better the effect of the synthesized speech. With the diversification of application scenarios, such as mobile phone assistants, in-vehicle navigation, and replication of the voices of relatives, each application scenario requires a large amount of recording data of its representative speaker, which is difficult and costly. However, existing speech synthesis models built based on a small amount of training data are prone to fitting noise, and the synthesized speech has a serious mechanical feeling, is not smooth enough, and is also very different from the speaking styles of the speaker, such as timbre and rhythm. Summary of the Invention

[0004] The present application provides a speech synthesis method, a model training method, a device, and a storage medium, aiming to improve the generalization ability and robustness of the speech synthesis model, so that the speech synthesis model can synthesize high-quality natural and smooth speech that is more in line with the speaking style of a specific speaker.

[0005] In a first aspect, the present application provides a speech synthesis method, and the method includes:

[0006] Obtain the text to be synthesized, input it into the trained target speech synthesis model, and obtain the target speech corresponding to the text to be synthesized, where the trained target speech synthesis model is obtained through the following method:

[0007] Pre-train a preset speech synthesis model according to a pre-training set to obtain a speech synthesis model, where the pre-training set includes speech samples of multiple speakers, and the speech synthesis model includes a duration prediction network;

[0008] Determine the target duration prediction network corresponding to the target application scenario, and replace the duration prediction network of the speech synthesis model with the target duration prediction network to obtain a target speech synthesis model;

[0009] Obtain a target training set, where the target training set includes speech samples of a target speaker;

[0010] Obtain the speech samples of the same - type speakers as the target speaker from the pre - training set, perform mask noise - adding processing on the speech samples of the same - type speakers to obtain noise - added speech samples;

[0011] Train the target speech synthesis model according to the target training set and the noise - added speech samples to obtain the trained target speech synthesis model.

[0012] In a second aspect, the present application provides a method for training a speech synthesis model. The method includes:

[0013] Pre - train a preset speech synthesis model according to a pre - training set to obtain a speech synthesis model, where the pre - training set includes speech samples of multiple speakers, and the speech synthesis model includes a duration prediction network;

[0014] Determine the target duration prediction network corresponding to the target application scenario, and replace the duration prediction network of the speech synthesis model with the target duration prediction network to obtain a target speech synthesis model;

[0015] Obtain a target training set, where the target training set includes speech samples of a target speaker;

[0016] Obtain the speech samples of the same - type speakers as the target speaker from the pre - training set, perform mask noise - adding processing on the speech samples of the same - type speakers to obtain noise - added speech samples;

[0017] Train the target speech synthesis model according to the target training set and the noise - added speech samples to obtain the trained target speech synthesis model.

[0018] In a third aspect, the present application further provides a computer device. The computer device includes:

[0019] A memory and a processor;

[0020] Wherein, the memory is connected to the processor and is used to store programs;

[0021] The processor is used to implement the steps of any speech synthesis method provided in the embodiments of the present application, or implement the steps of any method for training a speech synthesis model provided in the embodiments of the present application by running the programs stored in the memory.

[0022] Fourthly, the present application also provides a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to implement the steps of any of the speech synthesis methods provided in the embodiments of the present application, or to implement the steps of any of the training methods of the speech synthesis model provided in the embodiments of the present application.

[0023] The speech synthesis method, model training method, device and storage medium disclosed in the present application. For the speech synthesis method, the text to be synthesized is obtained and input into a trained target speech synthesis model to obtain the target speech corresponding to the text to be synthesized. Among them, the trained target speech synthesis model is obtained through the following method: in the pre-training stage, a speech synthesis model is pre-trained using speech samples of multiple speakers, and then the target duration prediction network corresponding to the target application scenario is used to replace the duration prediction network in the speech synthesis model to obtain the target speech synthesis model. Thus, in the fine-tuning stage, a small number of speech samples of the target speaker and the noise-added speech samples of the same-kind speakers of the target speaker are used to fine-tune the target speech synthesis model to obtain the trained target speech synthesis model. On the one hand, compared with the duration prediction network pre-trained using speech samples of multiple speakers, the target duration prediction network corresponding to the target application scenario has higher stability. Fine-tuning on the target duration prediction network corresponding to the target application scenario enables the trained target speech synthesis model to predict the phoneme duration more in line with the pronunciation style of the target speaker, so that the rhythm of the speech synthesized by the trained target speech synthesis model is closer to the real speech of the target speaker, improving the naturalness and fluency of the synthesized speech. On the other hand, the noise-added speech samples of the same-kind speakers of the target speaker are obtained by performing mask noise addition on the speech samples of the same-kind speakers of the target speaker used in the pre-training stage. When fine-tuning the target speech synthesis model, it can prevent the target speech synthesis model from "carrying over" historical information to predict the present, effectively reducing overfitting of the trained target speech synthesis model, thereby enhancing the generalization ability of the trained target speech synthesis model, making the trained target speech synthesis model highly robust to noise scenarios, and further improving the quality of the synthesized speech.

[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1a It is a schematic diagram of constructing a personalized speech synthesis system by an existing few-shot technology provided by an embodiment of the present application;

[0027] Figure 1b It is a schematic diagram of constructing a personalized speech synthesis system by an existing one-shot technology provided by an embodiment of the present application;

[0028] Figure 2 It is a schematic flowchart of the steps of a speech synthesis method provided by an embodiment of the present application;

[0029] Figure 3 It is a schematic structural diagram of a preset speech synthesis model provided by an embodiment of the present application;

[0030] Figure 4 It is an example diagram of performing mask and noise addition processing on speech samples of the same type of target speaker provided by an embodiment of the present application;

[0031] Figure 5 It is an example diagram of characterizing acoustic features as images provided by an embodiment of the present application;

[0032] Figure 6 It is a schematic flowchart of the steps of a training method of a speech synthesis model provided by an embodiment of the present application;

[0033] Figure 7 It is a schematic block diagram of a computer device provided by an embodiment of the present application.

[0034] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Detailed implementation manners

[0035] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0036] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may be changed according to the actual situation.

[0037] It should be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0038] It should be understood that, for the convenience of clearly describing the technical solutions of the embodiments of this application, in the embodiments of this application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. For example, the first image set and the second image set are only used to distinguish different image sets and do not limit their sequence. Those skilled in the art can understand that the terms such as "first" and "second" do not limit the quantity and execution order, and the terms such as "first" and "second" do not necessarily limit differences.

[0039] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0040] Currently, the mainstream speech synthesis models obtained based on a small amount of training data are all based on sequence-to-sequence modeling: 1) Among them, models with an attention structure such as TACO, etc., use a content-based attention mechanism to achieve the alignment between the input text feature sequence and the output acoustic feature sequence. At this time, the attention weights are calculated based on the query (Query) and the key value (Key), without any constraints on monotonicity and local information. This leads to stability problems such as repeated pronunciation, missed reading, mispronunciation, and inability to end during speech synthesis due to alignment errors, especially when the amount of synthetic training set data is too small. 2) For models without an attention structure, since the duration prediction model and the acoustic model are modeled separately, the fluency of the synthesis is insufficient.

[0041] In related technologies, few-shot techniques and one-shot (also known as zero-shot) techniques are also used to construct personalized speech synthesis models. Such as Figure 1aAs shown in the figure, the few-shot technology uses multi-person data to train a multi-speaker model, and then uses a small amount of single-person data to fine-tune the model. The purpose is to use a large amount of multi-person data to guide the model learning, and then use knowledge transfer for single-person fine-tuning to obtain a personalized model. However, due to various application scenarios, there are different noises in the recorded data. When fine-tuning the model, due to the small amount of data, the model is easily overfitted to the noise, resulting in different noises in the personalized synthesis model. At the same time, the duration prediction model is a statistical model, and it is difficult to statistically obtain the duration from 5-10 minutes of training data, resulting in a serious mechanical feeling in the duration of the personalized model. As Figure 1b shown, oneshot extracts features from the input audio through a timbre extraction network (such as voiceprint features i-vector, d-vector, etc.) as the condition and inputs them into the multi-speaker model. During testing, only one sentence of the speaker is needed to extract the voiceprint and input it into the multi-speaker model to obtain a personalized model. However, due to the fact that the oneshot scheme does not fine-tune the model and the model parameters remain unchanged, there are two problems: 1) The timbre of the personalized model is not very similar; 2) The personalized synthesis model does not have the prosody information of the speaker.

[0042] Therefore, the present application provides a speech synthesis method, a model training method, a device, and a storage medium. Among them, the speech synthesis method pre-trains a speech synthesis model based on a pre-training set, replaces the duration prediction network in the speech synthesis model with a target duration prediction network corresponding to the target application scenario to obtain a target speech synthesis model, and then fine-tunes the target speech synthesis model based on the speech samples of the target speaker and the noisy speech samples of the same type of speakers of the target speaker to obtain a trained target speech synthesis model. Finally, the trained target speech synthesis model is used to synthesize the target speech corresponding to the text to be synthesized of the target speaker. The generalization ability and robustness of the trained target speech synthesis model are improved, and when used to synthesize the exclusive speech of a specific speaker, it can synthesize high-quality natural and fluent speech that is more in line with the speaking style of the specific speaker.

[0043] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0044] Please refer to Figure 2 , Figure 2 which is a speech synthesis method provided by an embodiment of the present application. This method can be applied to a computer device, specifically to a computer device for speech synthesis, such as including a GPU, etc.

[0045] As Figure 2 shown, the speech synthesis method includes steps S101 to S106.

[0046] S101. Obtain the text to be synthesized and input it into the trained target speech synthesis model to obtain the target speech corresponding to the text to be synthesized.

[0047] Among them, the trained target speech synthesis model is a personalized speech synthesis model exclusive to the target speaker.

[0048] Obtain the text to be synthesized and input the text to be synthesized into the trained target speech synthesis model, and the target speech corresponding to the text to be synthesized output by the target speech synthesis model can be obtained. It can be understood that the target speech is very similar to the real speech of the target speaker.

[0049] The trained target speech synthesis model is obtained through the following method:

[0050] S102. Pre-train the preset speech synthesis model according to the pre-training set to obtain a speech synthesis model. Among them, the pre-training set includes speech samples of multiple speakers, and the speech synthesis model includes a duration prediction network.

[0051] Training the trained target speech synthesis model mainly includes two processes: one is to pre-train the preset speech synthesis model based on the speech samples of multiple speakers in the pre-training set to obtain a pre-trained speech synthesis model; the other is to perform personalized fine-tuning training on the pre-trained speech synthesis model again based on a small number of speech samples of the target speaker in the personalized training set to obtain the trained target speech synthesis model.

[0052] It should be noted that the preset speech synthesis model is a multi-speaker speech synthesis model. The speech samples of multiple speakers in the pre-training set do not include the speech samples of the target speaker.

[0053] It can be understood that before pre-training the preset speech synthesis model according to the pre-training set to obtain a speech synthesis model, a pre-training set is first established. Specifically: obtain the first audio data corresponding to multiple speakers and the text of the first audio data, and the second audio data and the text of the second audio data from the preset speech library, where the quality of the first audio data is higher than that of the second audio data; establish a pre-training set according to the first audio data and the text of the first audio data, and the second audio data and the text of the second audio data.

[0054] To quickly establish a pre-training set, a limited amount of original pre-training data can be collected from a preset speech database, and the original pre-training data can be processed to establish the pre-training set. Specifically, first collect the first audio data and the second audio data corresponding to multiple speakers from the preset speech database, as well as the text corresponding to the first audio data and the text corresponding to the second audio data. Among them, the quality of the first audio data is higher than that of the second audio data. The first audio data can be characterized as high-quality audio data with low audio distortion, high intensity, high frequency, and / or high signal-to-noise ratio, and the second audio data can be characterized as low-quality audio data with high audio distortion, low intensity, low frequency, and / or low signal-to-noise ratio. In addition, the first audio data and the second audio data corresponding to each speaker can be audio files of a sentence respectively, meeting a certain speech duration, such as 5 min - 10 min.

[0055] After that, extract the acoustic features corresponding to the first audio data, label the clean label for the first audio data, and determine the phoneme duration corresponding to the text of the first audio data. Also, extract the acoustic features corresponding to the second audio data, label the noise label for the second audio data, and determine the phoneme duration corresponding to the text of the second audio data.

[0056] Considering that the original pre-training data collected from the preset speech database is limited and data augmentation is required, extended audio data is obtained based on the collected first audio data. Exemplarily, noise is added to the first audio data, for example, randomly adding 5 - 25 db of noise, and the noise scenarios are diverse. The extended audio data is labeled with the noise label.

[0057] Finally, based on the acoustic features and clean labels corresponding to the first audio data, the phoneme duration corresponding to the text of the first audio data, the acoustic features and noise labels corresponding to the second audio data, the phoneme duration corresponding to the text of the second audio data, and combined with the noise label corresponding to the extended audio data, a pre-training set is established. Exemplarily, the way to establish the pre-training set is, for example, the pre-training sample set is a set, and the elements in the set are the speech samples of each speaker, that is:

[0058] Pre-training set = {Speech sample 1 of speaker 1, Speech sample 2 of speaker 2,..., Speech sample n of speaker n, Extended speech sample x}

[0059] = {(Acoustic feature 1 of the audio data of speaker 1, Label 1 marked, Phoneme duration 1), (Acoustic feature 2 of the audio data of speaker 2, Label 2 marked, Phoneme duration 2),..., (Acoustic feature n of the audio data of speaker n, Label n marked, Phoneme duration n), (Acoustic feature x of the extended audio data, Label x marked, Phoneme duration x)}

[0060] = {[(acoustic feature 1' of the first audio data of speaker 1, label clean1', phoneme duration 1'), (acoustic feature 1" of the second audio data of speaker 1, label noise1", phoneme duration 1")], [(acoustic feature 2' of the first audio data of speaker 2, label clean2', phoneme duration 2'), (acoustic feature 2" of the second audio data of speaker 2, label noise2", phoneme duration 2")],..., [(acoustic feature n' of the first audio data of speaker n, label cleann', phoneme duration n'), (acoustic feature n" of the second audio data of speaker n, label noisen", phoneme duration n")], (acoustic feature x of the extended audio data, label noisex, phoneme duration x)}.

[0061] That is to say, the speech samples of multiple speakers in the pre-training set include the acoustic features and labels of the audio data of multiple speakers, as well as the phoneme durations corresponding to the texts of the audio data of multiple speakers. Among them, the acoustic features and labels of the audio data of multiple speakers include the acoustic features and clean labels of the first audio data of multiple speakers, and the acoustic features and noise labels of the second audio data. The phoneme durations corresponding to the texts of the audio data of multiple speakers include the phoneme durations corresponding to the texts of the first audio data of multiple speakers, and the phoneme durations corresponding to the texts of the second audio data.

[0062] It can be understood that since the extended audio data is obtained by adding noise to the first audio data, the acoustic features of the extended audio data are the same as those of the first audio data, the text of the extended audio data is the same as the text of the first audio data, and thus the phoneme duration corresponding to the text of the extended audio data is the same as the phoneme duration corresponding to the text of the first audio data.

[0063] In some embodiments, to determine the phoneme duration corresponding to the text of the first audio data, specifically, a phoneme sequence corresponding to the text of the first audio data is obtained, and the phoneme sequence is forced to align with the text of the first audio data to obtain the phoneme duration of each phoneme in the phoneme sequence.

[0064] To determine the phoneme duration corresponding to the text of the first audio data, front-end processing is performed on the text of the first audio data to obtain a phoneme sequence (also referred to as a pronunciation sequence) corresponding to the text of the first audio data; then, the phoneme sequence corresponding to the text of the first audio data is forced to align with the text of the first audio data by using a pre-trained ASR (Automatic Speech Recognition) model to obtain the phoneme duration of each phoneme in the phoneme sequence corresponding to the text of the first audio data.

[0065] Similarly, in order to determine the phoneme duration corresponding to the text of the second audio data, the text of the second audio data is preprocessed to obtain the phoneme sequence corresponding to the text of the second audio data; then, the pre-trained ASR model is used to force-align the phoneme sequence corresponding to the text of the second audio data with the text of the second audio data to obtain the phoneme duration of each phoneme in the phoneme sequence corresponding to the text of the second audio data.

[0066] Among them, the ASR model is pre-trained in the open-source speech recognition tool kaldi by using the text of the first audio data and the phoneme sequence corresponding to the text of the first audio data, as well as the text of the second audio data and the phoneme sequence corresponding to the text of the second audio data, to obtain the trained ASR model.

[0067] After obtaining the pre-training set, the preset speech synthesis model is pre-trained according to the pre-training set to generate a pre-trained speech synthesis model.

[0068] Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of the preset speech synthesis model. The preset speech synthesis model includes a sequence-to-sequence network without an attention mechanism in the autoregressive model, a duration prediction network, and a variance prediction network. Among them, the sequence-to-sequence network without an attention mechanism in the autoregressive model serves as an acoustic model and includes an encoder end and a decoder end. The introduction of autoregression is to make the audio synthesized by the speech synthesis system smoother and more natural; the duration prediction network includes two layers of bidirectional LSTM (Long Short-Term Memory) networks. There are two objectives when pre-training the preset speech synthesis model: one is for the sequence-to-sequence network without an attention mechanism in the autoregressive model to learn to predict acoustic features, and the other is for the duration prediction network to learn to predict the duration of each phoneme.

[0069] In some embodiments, the pre-training of the preset speech synthesis model according to the pre-training set to obtain the speech synthesis model is specifically: pre-training the sequence-to-sequence network and the duration prediction network of the preset speech synthesis model according to the pre-training set, and saving the parameters of the sequence-to-sequence network and the duration prediction network to obtain the speech synthesis model.

[0070] Pre-train the sequence-to-sequence network without an attention mechanism in the autoregressive model and the duration prediction network simultaneously according to the pre-training set, and save the parameters of the sequence-to-sequence network without an attention mechanism in the autoregressive model and the duration prediction network to obtain the speech synthesis model.

[0071] In some embodiments, the pre-training of the sequence-to-sequence network and the duration prediction network of the preset speech synthesis model according to the pre-training set is specifically as follows: Input the speech samples of multiple speakers in the pre-training set into the preset speech synthesis model, encode the acoustic features and the phoneme durations at the encoding end of the sequence-to-sequence network to obtain an acoustic feature encoding vector and a phoneme duration encoding vector; add an embedding operation for noise at the decoding end of the sequence-to-sequence network to obtain a noise embedding vector; use the phoneme duration encoding vector as the input and the phoneme duration as the prediction target to train the duration prediction network; based on the acoustic feature encoding vector, the phoneme duration encoding vector, and the noise embedding vector, use the acoustic features as the prediction target to train the sequence-to-sequence network.

[0072] Input the speech samples of multiple speakers in the pre-training set into the preset speech synthesis model, deeply encode (encoding) the acoustic features and the phoneme durations in the speech samples of multiple speakers at the encoding end to obtain an acoustic feature encoding vector and a phoneme duration encoding vector output by the encoding end. Use the phoneme duration encoding vector output by the encoding end as the input and the phoneme durations in the speech samples of multiple speakers as the prediction target to train the duration prediction network.

[0073] It should be noted that an embedding of noise is added at the decoding end to absorb the noise information of the channel, so that the target of the autoregressive sequence-to-sequence network without an attention mechanism is clean acoustic features when the input label is a clean label; and the target is noisy or low-quality acoustic features when the input label is a noise label, so that the speech synthesized by the speech synthesis model can be controlled to be clean or not through the noise label and the clean label. That is, an embedding operation for noise is added at the decoding end to obtain a noise embedding vector.

[0074] Based on the acoustic feature encoding vector and the phoneme duration encoding vector output by the encoding end, as well as the noise embedding vector, use the acoustic features in the speech samples of multiple speakers as the prediction target to train the autoregressive sequence-to-sequence network without an attention mechanism.

[0075] In some embodiments, the training method of the speech synthesis model further includes: calculating a first loss function of the duration prediction network during the training of the duration prediction network; calculating a second loss function of the sequence-to-sequence network during the training of the sequence-to-sequence network; calculating a loss function of the preset speech synthesis model according to the first loss function and the second loss function until the loss function converges to obtain the speech synthesis model.

[0076] When training the duration prediction network, obtain the predicted phoneme durations output by the duration prediction network, and construct a loss function (defined as the first loss function) for the duration prediction network according to the predicted phoneme durations and the actual phoneme durations in the pre-training set.

[0077] When training the sequence2sequence network without the attention mechanism of autoregression, obtain the predicted acoustic features output by the sequence2sequence network without the attention mechanism of autoregression, and construct a loss function (defined as the second loss function) for the sequence2sequence network without the attention mechanism of autoregression according to the predicted acoustic features and the actual acoustic features in the pre-training set.

[0078] In some embodiments, when training the sequence-to-sequence network, calculating the second loss function of the sequence-to-sequence network specifically includes: obtaining a fused prediction vector according to the acoustic feature encoding vector, the phoneme duration encoding vector, and the noise embedding vector; performing autoregressive decoding on the fused prediction vector at the decoding end of the sequence-to-sequence network, so that the sequence-to-sequence network learns to predict clean acoustic features when the label is the clean label and learns to predict noise acoustic features when the label is the noise label, and calculating the second loss function of the sequence-to-sequence network.

[0079] The phoneme duration encoding vector output by the encoding end is also used as the input of the variance prediction module. After normalizing the mean and variance of the predicted phoneme duration output by the duration prediction network and the variance output by the variance prediction module, perform Gaussian upsampling together with the acoustic feature encoding vector output by the encoding end to obtain a Gaussian upsampling vector; also perform positional encoding on the predicted phoneme duration output by the duration prediction network to obtain a positional encoding vector.

[0080] Fuse the positional encoding vector, the Gaussian upsampling vector, and the noise embedding vector to obtain a fused prediction vector, input the fused prediction vector into the decoding end for autoregressive decoding. The sequence2sequence network without the attention mechanism of autoregression uses the acoustic features in the speech samples of multiple speakers as the prediction target, learns to predict clean acoustic features when the label in the speech samples of multiple speakers is the clean label, and learns to predict noise acoustic features when the label is the noise label, so as to obtain the predicted acoustic features, and then construct the second loss function according to the predicted acoustic features and the actual acoustic features in the pre-training set.

[0081] Further, according to the first loss function and the second loss function, calculate the loss function of a preset speech synthesis model until the loss function converges, and obtain a pre-trained speech synthesis model. Specifically, the loss of the speech synthesis model includes two parts. One part is the loss of the sequence2sequence network without the attention mechanism in autoregressive mode, and the other part is the loss of the duration prediction network. The loss function L of the speech synthesis model is as follows:

[0082] L = L spec+ λL dur

[0083]

[0084]

[0085] where L spec represents the loss function of the sequence2sequence network without the attention mechanism in autoregressive mode, T represents the number of frames of acoustic features, K represents the dimension of acoustic features (the shape of acoustic features is T*K), y and y' represent two predicted acoustic features, and y * represents the actual acoustic feature;

[0086] L dur represents the loss function of the duration prediction network, N represents the number of phonemes, d represents the predicted phoneme duration, and d * represents the actual phoneme duration;

[0087] λ represents a hyperparameter, and its value can be set flexibly according to the actual situation. For example, λ = 1.0.

[0088] That is, by adding the loss of the sequence2sequence network without the attention mechanism in autoregressive mode and the loss of the duration prediction network, the loss function of the speech synthesis model can be obtained. Through iterative training until the loss function of the speech synthesis model converges, the pre-trained speech synthesis model can be obtained. The sequence2sequence network without the attention mechanism in autoregressive mode in the pre-trained speech synthesis model is used to predict the acoustic features of the text, and the duration prediction network is used to predict the duration of each phoneme of the text.

[0089] Further, in some embodiments, the type or degree of added noise can also be controlled, and different degrees of noise labels are used for control, so that the pre-trained speech synthesis model can control the cleanliness of the synthesized speech.

[0090] After the pre-training is completed to obtain the pre-trained speech synthesis model, fine-tune the pre-trained speech synthesis model, that is, perform personalized training on the pre-trained speech synthesis model.

[0091] Step S103: Determine the target duration prediction network corresponding to the target application scenario, and replace the duration prediction network of the speech synthesis module with the target duration prediction network to obtain a target speech synthesis model.

[0092] By pre-statistically analyzing the phoneme durations of different speakers, we found that there are commonalities in the phoneme durations of different speakers. For example, when everyone says "the People's Republic of China", the pauses and the pronunciation times of each character are basically the same. The duration prediction network is a statistical model. In low-resource scenarios (with less training data), it is difficult for the duration prediction network to statistically obtain duration information from just a few sentences. That is to say, due to the small number of speech samples in the pre-training set, when the preset speech synthesis model is pre-trained using the pre-training set, the stability of the obtained duration prediction network is not high. At the same time, since various application scenarios (such as the customer service scenario, in-vehicle scenario, etc.) have a lot of training data from which a good duration prediction model can be statistically obtained, therefore, it is proposed that in the fine-tuning stage, instead of fine-tuning on the duration prediction network obtained in the pre-training stage, fine-tuning is performed on the duration prediction network corresponding to the speaker at the phonetic database level in the same domain (abnormal phoneme durations are removed during the fine-tuning process, for example, removing phoneme times less than 30 ms or greater than 500 ms to ensure the stability of the duration prediction model).

[0093] Specifically, before fine-tuning the pre-trained speech synthesis model, first obtain the duration prediction network corresponding to the personalized application scenario (i.e., the target application scenario of the target speaker), which is defined as the target duration prediction network. Then, replace the duration prediction network obtained during pre-training with the target duration prediction network corresponding to the target application scenario, and define the pre-trained speech synthesis model with the replaced duration prediction network as the target speech synthesis model.

[0094] It can be understood that the duration prediction network corresponding to each application scenario is pre-trained in advance based on a large number of speech samples of representative speakers in each application scenario. Exemplarily, for example, if the target application scenario is the customer service scenario, the duration prediction model corresponding to the speaker at the customer service scenario phonetic database level is used to replace the duration prediction network in the pre-trained speech synthesis model.

[0095] In some embodiments, the determining of the target duration prediction network corresponding to the target application scenario is specifically: selecting the duration prediction network with the closest cosine distance to the duration prediction network from the duration prediction networks corresponding to a preset variety of application scenarios as the target duration prediction network corresponding to the target application scenario; or determining the target duration prediction network corresponding to the target application scenario from the duration prediction networks corresponding to a preset variety of application scenarios.

[0096] One way to determine the target duration prediction network corresponding to the target application scenario is to select the duration prediction network with the closest cosine distance to the duration prediction network from the pre-trained duration prediction networks corresponding to the various application scenarios as the target duration prediction network corresponding to the target application scenario. Specifically, extract the voiceprint features of the audio data of multiple speakers in the pre-training set, and also extract the voiceprint features from the speech samples of the representative speakers of each application scenario, respectively calculate the cosine distances between the voiceprint features corresponding to the pre-training set and the voiceprint features corresponding to the speech samples of the representative speakers of each application scenario, and sort the calculated cosine distances from large to small. The smaller the cosine distance, the smaller the cosine distance between the duration prediction network obtained during pre-training and the duration prediction network corresponding to the corresponding application scenario. Select the duration prediction network with the closest cosine distance to the duration prediction network as the target duration prediction network corresponding to the target application scenario.

[0097] In some embodiments, determining the target duration prediction network corresponding to the target application scenario is specifically: matching the duration prediction network corresponding to the target application scenario from the duration prediction networks corresponding to the preset multiple application scenarios as the target duration prediction network.

[0098] The second way to determine the target duration prediction network corresponding to the target application scenario is to clarify the target application scenario, match the personalized scenario with multiple application scenarios, and directly match the duration prediction network corresponding to the target application scenario from the pre-trained duration prediction networks corresponding to multiple application scenarios as the target duration prediction network.

[0099] Step S104: establishing a target training set, wherein the target training set includes speech samples of a target speaker.

[0100] In order to fine-tune the target speech synthesis model to obtain a personalized speech synthesis model, it is necessary to establish a personalized training set corresponding to the target speaker (defined as the target training set). The target training set includes the speech sample of the target speaker (user). The speech sample of the target speaker includes the acoustic features and labels of the audio data of the target speaker, and the phoneme duration of the text corresponding to the audio data of the target speaker. It is understandable that the speech sample of the target speaker is small, wherein the audio data of the target speaker can be an audio file of a sentence of the target speaker, which meets a certain speech duration, such as 5-10 minutes.

[0101] Step S105 : selecting speech samples of speakers similar to the target speaker from the pre-training set, and performing mask noise processing on the speech samples of the speakers similar to the target speaker to obtain noisy speech samples.

[0102] Among them, when the sequence2sequence network without the attention mechanism in the autoregressive model is trained, in addition to receiving the input from the encoder at the decoder end, there is also the input of historical information. However, the acoustic feature is a very smooth feature with a large correlation before and after. When the training data is small, the sequence2sequence network without the attention mechanism in the autoregressive model is very likely to learn to "carry" the historical information to predict the present, resulting in the sequence2sequence network without the attention mechanism in the autoregressive model not learning information from the encoder, and the sequence2sequence network without the attention mechanism in the autoregressive model is prone to overfitting (overfitting means that the model has a high accuracy on the training set but a low accuracy on the validation set). Therefore, the tf-mask noise addition method is creatively applied in fine-tuning to effectively reduce overfitting. Among them, the tf-mask noise addition method refers to the method of masking and adding noise in the time domain and frequency domain.

[0103] Specifically, select the speech samples of the same-category speakers of the target speaker from the pre-training set. The same-category speakers of the target speaker include speakers with the same gender, age, speaking style, and / or speaking environment as the target speaker. The speech samples of the same-category speakers of the target speaker are the historical information that is easily "carried" during the fine-tuning of the autoregressive model. Perform mask noise addition processing on the speech samples of the same-category speakers, and define the speech samples of the same-category speakers after mask noise addition processing as the noise-added speech samples.

[0104] In some embodiments, the performing mask noise addition processing on the speech samples of the same-category speakers is specifically: adding noise to the acoustic features in the speech samples of the same-category speakers in the time domain and frequency domain.

[0105] Performing mask noise addition processing on the speech samples of the same-category speakers, that is, adding noise to the acoustic features of the speech samples of the same-category speakers of the target speaker in the time domain and frequency domain, as Figure 4 shown.

[0106] Step S106: Train the target speech synthesis model according to the target training set and the noise-added speech samples to obtain the trained target speech synthesis model.

[0107] Thus, fine-tune the target speech synthesis model according to the target training set and the noise-added speech samples to obtain the trained target speech synthesis model. The trained target speech synthesis model is the personalized speech synthesis model corresponding to the target speaker.

[0108] In some embodiments, step S105 is specifically as follows: Input the speech samples of the target speaker in the target training set and the noise-added speech samples into the target speech synthesis model for training, redefine the loss function until the redefined loss function converges, and obtain the trained target speech synthesis model.

[0109] In some embodiments, redefining the loss function specifically means: Redefine the second loss function of the sequence-to-sequence network as the structural similarity SSIM loss function; redefine the loss function of the target speech synthesis model according to the SSIM loss function.

[0110] Input the speech samples of the target speaker in the target training set and the noise-added speech samples into the target speech synthesis model for training. Among them, considering that the loss function of the sequence-to-sequence network without the attention mechanism in autoregressive is essentially the mean squared error (MSE) loss function, and the MSE loss function tends to make the model biased towards the average level. Therefore, redefine the loss function of the sequence-to-sequence network without the attention mechanism in autoregressive as the SSIM loss function. The SSIM loss function is a measure of structural similarity, and the formula is as follows:

[0111]

[0112] where c 1 and c 2 represent constants;

[0113] The acoustic features are characterized as an image of T*D, denoted as x, and u x is the mean of x;

[0114] The predicted acoustic features are characterized as another image of T*D, denoted as y, and u y is the mean of y;

[0115] σ xy is the covariance matrix of x and y, is the variance of x, is the variance of y.

[0116] SSIM can better reflect the judgment of the human visual system on the similarity of two images than MSE. As Figure 5 shown, the MSE values of the middle and right images are the same, but obviously the middle image has more noise. The SSIM value of the right image is significantly higher than that of the middle (the higher the SSIM, the better). Therefore, we add the SSIM loss during personalized fine-tuning. In fact, in actual experiments, it is also verified that adding this loss will make the spectrum brighter and cleaner.

[0117] As the loss function of the sequence2sequence network without attention mechanism in the autoregressive model is redefined as the SSIM loss function, the loss function of the target speech synthesis model will also be defined as the sum of the SSIM loss function and the loss function of the duration prediction network. By iteratively training until the loss function of the redefined target speech synthesis model converges, a personalized speech synthesis system can be generated.

[0118] Thus, by fine-tuning the target speech synthesis model with a small number of speech samples of the target speaker, a personalized speech synthesis model exclusive to the target speaker can be obtained.

[0119] Then, the text to be synthesized is input into the trained target speech synthesis model. The target phoneme duration is predicted through the target duration prediction network, and then the target acoustic features are predicted through the sequence2sequence network without attention mechanism in the autoregressive model. Finally, the target speech corresponding to the text to be synthesized is synthesized according to the target phoneme duration and the target acoustic features. Thus, through the personalized speech synthesis model exclusive to the target speaker, personalized speech corresponding to any text to be synthesized by the target speaker can be synthesized.

[0120] The voice synthesis method provided by each of the above embodiments obtains the text to be synthesized, inputs it into the trained target voice synthesis model, and obtains the target voice corresponding to the text to be synthesized. Among them, the trained target voice synthesis model is obtained through the following method: in the pre-training stage, a voice synthesis model is pre-trained using a small number of voice samples of multiple speakers, and then the target duration prediction network corresponding to the target application scenario is used to replace the duration prediction network in the voice synthesis model to obtain the target voice synthesis model. Thus, in the fine-tuning stage, a small number of voice samples of the target speaker and the noisy voice samples of the same-type speakers of the target speaker are used to fine-tune the target voice synthesis model to obtain the trained target voice synthesis model. On the one hand, compared with the duration prediction network pre-trained using the voice samples of multiple speakers, the target duration prediction network corresponding to the target application scenario has higher stability. Fine-tuning on the target duration prediction network corresponding to the target application scenario enables the trained target voice synthesis model to predict the phoneme durations that are more in line with the pronunciation style of the target speaker, so that the rhythm of the voice synthesized by the trained target voice synthesis model is closer to the real voice of the target speaker, improving the naturalness and fluency of the synthesized voice. On the other hand, the noisy voice samples of the same-type speakers of the target speaker are obtained by performing mask noise addition processing on the voice samples of the same-type speakers of the target speaker used in the pre-training stage. When fine-tuning the target voice synthesis model, it can prevent the target voice synthesis model from "carrying" historical information over to predict the present, effectively reducing overfitting of the trained target voice synthesis model, thereby improving the generalization ability of the trained target voice synthesis model, making the trained target voice synthesis model highly robust to noise scenarios, and further improving the quality of the synthesized voice.

[0121] Please participate Figure 6 , Figure 6 This is a training method for a voice synthesis model provided by an embodiment of the present application. This method can be applied to a computer device, specifically to a computer device for dedicated model training, such as including a GPU, etc.

[0122] As Figure 6 shown, the training method for this voice synthesis model includes steps S201 to S205.

[0123] S201. Pre-train a preset voice synthesis model according to a pre-training set to obtain a voice synthesis model, where the pre-training set includes voice samples of multiple speakers, and the voice synthesis model includes a duration prediction network;

[0124] S202. Determine the target duration prediction network corresponding to the target application scenario, and use the target duration prediction network to replace the duration prediction network of the voice synthesis model to obtain a target voice synthesis model;

[0125] S203. Obtain a target training set, where the target training set includes speech samples of a target speaker.

[0126] S204. Obtain speech samples of speakers of the same type as the target speaker from the pre-training set, perform mask noise addition processing on the speech samples of the speakers of the same type to obtain noise-added speech samples.

[0127] S205. Train the target speech synthesis model according to the target training set and the noise-added speech samples to obtain a trained target speech synthesis model.

[0128] The training method of this speech synthesis model mainly includes two processes: one is to pre-train a preset speech synthesis model based on speech samples of multiple speakers in the pre-training set to obtain a pre-trained speech synthesis model; the other is to perform personalized fine-tuning training on the pre-trained speech synthesis model again based on a small number of speech samples of the target speaker in the personalized training set to generate a personalized speech synthesis model exclusive to the target speaker.

[0129] Among them, for the specific pre-training process and personalized fine-tuning training process, please refer to the above embodiments and will not be elaborated here.

[0130] The training method of this speech synthesis model improves the generalization ability and robustness of the speech synthesis model, so that the trained target speech synthesis model can synthesize high-quality natural and fluent speech that is more in line with the speaking style of a specific speaker.

[0131] Please refer to Figure 7 , Figure 7 which is a schematic block diagram of a computer device provided by an embodiment of the present application. As Figure 7 shown, the computer device 700 includes one or more processors 701 and a memory 702. The processor 701 and the memory 702 are connected through a bus, and this bus is, for example, an I2C (Inter-integrated Circuit) bus.

[0132] Among them, one or more processors 701 work alone or jointly to execute the steps of the training method of the speech synthesis model or the action image generation method provided by the above embodiments.

[0133] Specifically, the processor 701 can be a micro-control unit (MCU), a central processing unit (CPU), a digital signal processor (DSP), etc.

[0134] Specifically, the memory 702 can be a Flash chip, a read-only memory (ROM), a magnetic disk, an optical disc, a USB flash drive, a mobile hard disk, etc.

[0135] The processor 701 is configured to run a computer program stored in the memory 702, and when executing the computer program, implement the steps of the speech synthesis method or the training method of the speech synthesis model provided in the foregoing embodiments.

[0136] Exemplarily, the processor 701 is configured to run a computer program stored in the memory 702, and when executing the computer program, implement the following steps:

[0137] Obtain the text to be synthesized, input it into the trained target speech synthesis model, and obtain the target speech corresponding to the text to be synthesized. The trained target speech synthesis model is obtained through the following method: pre-train a preset speech synthesis model according to a pre-training set to obtain a speech synthesis model, where the pre-training set includes speech samples of multiple speakers, and the speech synthesis model includes a duration prediction network; determine a target duration prediction network corresponding to a target application scenario, and replace the duration prediction network of the speech synthesis model with the target duration prediction network to obtain a target speech synthesis model; obtain a target training set, where the target training set includes speech samples of a target speaker; obtain speech samples of speakers of the same type as the target speaker from the pre-training set, perform mask noise addition processing on the speech samples of the speakers of the same type to obtain noise-added speech samples; train the target speech synthesis model according to the target training set and the noise-added speech samples to obtain the trained target speech synthesis model.

[0138] In some embodiments, the preset speech synthesis model includes a sequence-to-sequence network and a duration prediction network; when the processor implements pre-training the preset speech synthesis model according to a pre-training set to obtain a speech synthesis model, it specifically further implements the following steps:

[0139] Pre-train the sequence-to-sequence network and the duration prediction network of the preset speech synthesis model according to the pre-training set, save the parameters of the sequence-to-sequence network and the duration prediction network, and obtain a speech synthesis model.

[0140] In some embodiments, the speech samples of the multiple speakers include the acoustic features of the audio data of the multiple speakers and the phoneme durations corresponding to the texts of the audio data of the multiple speakers; when the processor implements pre-training the sequence-to-sequence network and the duration prediction network of the preset speech synthesis model according to the pre-training set, it specifically further implements the following steps:

[0141] Input the speech samples of multiple speakers in the pre-training set into a preset speech synthesis model. Encode the acoustic features and the phoneme durations at the encoding end of the sequence-to-sequence network to obtain an acoustic feature encoding vector and a phoneme duration encoding vector. Add an embedding operation for noise at the decoding end of the sequence-to-sequence network to obtain a noise embedding vector. Use the phoneme duration encoding vector as the input and the phoneme duration as the prediction target to train the duration prediction network. Based on the acoustic feature encoding vector, the phoneme duration encoding vector, and the noise embedding vector, use the acoustic features as the prediction target to train the sequence-to-sequence network.

[0142] In some embodiments, when the processor executes the computer program, the following steps are further specifically implemented:

[0143] When training the duration prediction network, calculate the first loss function of the duration prediction network. When training the sequence-to-sequence network, calculate the second loss function of the sequence-to-sequence network. According to the first loss function and the second loss function, calculate the loss function of the preset speech synthesis model until the loss function converges to obtain the speech synthesis model.

[0144] In some embodiments, the speech samples of the multiple speakers include labels corresponding to the audio data of the multiple speakers, and the labels include clean labels and noise labels. When the processor implements calculating the second loss function of the sequence-to-sequence network when training the sequence-to-sequence network, the following steps are further specifically implemented:

[0145] Obtain a fusion prediction vector according to the acoustic feature encoding vector, the phoneme duration encoding vector, and the noise embedding vector. Perform autoregressive decoding on the fusion prediction vector at the decoding end of the sequence-to-sequence network, so that the sequence-to-sequence network learns to predict clean acoustic features when the label is a clean label and learns to predict noise acoustic features when the label is a noise label, and calculate the second loss function of the sequence-to-sequence network.

[0146] In some embodiments, when the processor implements training the target speech synthesis model according to the target training set and the noise-added speech samples to obtain the trained target speech synthesis model, the following steps are further specifically implemented:

[0147] Input the speech samples of the target speaker in the target training set and the noise-added speech samples into the target speech synthesis model for training, redefine the loss function until the redefined loss function converges to obtain the trained target speech synthesis model.

[0148] In some embodiments, when the processor implements the redefined loss function, the following steps are specifically further implemented:

[0149] Redefine the second loss function of the sequence-to-sequence network as a structural similarity SSIM loss function; redefine the loss function of the target speech synthesis model according to the SSIM loss function.

[0150] In some embodiments, before the processor implements pre-training a preset speech synthesis model according to a pre-training set to obtain a speech synthesis model, the following steps are specifically further implemented:

[0151] Obtain first audio data corresponding to multiple speakers and the text of the first audio data, as well as the second audio data and the text of the second audio data from a preset speech library, where the quality of the first audio data is higher than that of the second audio data;

[0152] Establish a pre-training set according to the first audio data and the text of the first audio data, as well as the second audio data and the text of the second audio data.

[0153] In some embodiments, when the processor implements determining a target duration prediction network corresponding to a target application scenario and replacing the duration prediction network of the speech synthesis model with the target duration prediction network to obtain a target speech synthesis model, the following steps are specifically further implemented:

[0154] Select the duration prediction network with the closest cosine distance to the duration prediction network from the duration prediction networks corresponding to a preset variety of application scenarios as the target duration prediction network corresponding to the target application scenario; or match the duration prediction network corresponding to the target application scenario from the duration prediction networks corresponding to a preset variety of application scenarios as the target duration prediction network.

[0155] In some embodiments, when the processor implements mask noise addition processing on the speech samples of the same type of speakers, the following steps are specifically further implemented:

[0156] Add noise to the acoustic features of the speech samples of the same type of speakers in the time domain and the frequency domain.

[0157] Exemplarily, the processor 701 is configured to run a computer program stored in the memory 702, and when executing the computer program, implement the following steps:

[0158] Pre-train a preset speech synthesis model according to a pre-training set to obtain a speech synthesis model, where the pre-training set includes speech samples of multiple speakers, and the speech synthesis model includes a duration prediction network; determine a target duration prediction network corresponding to a target application scenario, and replace the duration prediction network of the speech synthesis model with the target duration prediction network to obtain a target speech synthesis model; obtain a target training set, where the target training set includes speech samples of a target speaker; obtain speech samples of speakers of the same type as the target speaker from the pre-training set, perform mask noise addition processing on the speech samples of the speakers of the same type to obtain noise-added speech samples; train the target speech synthesis model according to the target training set and the noise-added speech samples to obtain a trained target speech synthesis model.

[0159] An embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the steps of the speech synthesis method or the training method of the speech synthesis model provided in the above embodiment.

[0160] Wherein, the computer-readable storage medium may be an internal storage unit of the computer device described in any of the foregoing embodiments, such as the hard disk or memory of the terminal device. The computer-readable storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal device.

[0161] As mentioned above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A voice synthesis method, characterized in that, the method includes: obtaining the text to be synthesized, inputting it into a trained target voice synthesis model, and obtaining the target voice corresponding to the text to be synthesized, wherein the trained target voice synthesis model is obtained by the following method: pre-training a preset voice synthesis model according to a pre-training set to obtain a voice synthesis model, wherein the pre-training set includes voice samples of multiple speakers, and the voice synthesis model includes a duration prediction network; determining a target duration prediction network corresponding to a target application scenario, and replacing the duration prediction network of the voice synthesis model with the target duration prediction network to obtain a target voice synthesis model; obtaining a target training set, wherein the target training set includes voice samples of a target speaker; obtaining voice samples of speakers of the same type as the target speaker from the pre-training set, performing mask noise addition processing on the voice samples of the speakers of the same type to obtain noise-added voice samples; training the target voice synthesis model according to the target training set and the noise-added voice samples to obtain the trained target voice synthesis model.

2. The method according to claim 1, characterized in that, the preset voice synthesis model includes a sequence-to-sequence network and a duration prediction network; the pre-training the preset voice synthesis model according to a pre-training set to obtain a voice synthesis model includes: pre-training the sequence-to-sequence network and the duration prediction network of the preset voice synthesis model according to the pre-training set, saving the parameters of the sequence-to-sequence network and the duration prediction network, and obtaining a voice synthesis model.

3. The method according to claim 2, characterized in that, the voice samples of the multiple speakers include the acoustic features of the audio data of the multiple speakers and the phoneme durations corresponding to the text of the audio data of the multiple speakers; the pre-training the sequence-to-sequence network and the duration prediction network of the preset voice synthesis model according to the pre-training set includes: inputting the voice samples of multiple speakers in the pre-training set into the preset voice synthesis model, encoding the acoustic features and the phoneme durations at the encoding end of the sequence-to-sequence network to obtain an acoustic feature encoding vector and a phoneme duration encoding vector; adding an embedding operation for noise at the decoding end of the sequence-to-sequence network to obtain a noise embedding vector; using the phoneme duration encoding vector as an input and the phoneme duration as a prediction target to train the duration prediction network; based on the acoustic feature encoding vector, the phoneme duration encoding vector, and the noise embedding vector, using the acoustic features as a prediction target to train the sequence-to-sequence network.

4. The method according to claim 3, characterized in that, the method further includes: when training the duration prediction network, calculating a first loss function of the duration prediction network; when training the sequence-to-sequence network, calculating a second loss function of the sequence-to-sequence network; Calculate the loss function of the preset speech synthesis model according to the first loss function and the second loss function until the loss function converges, and obtain the speech synthesis model.

5. The method according to claim 4, wherein, the speech samples of the multiple speakers include labels corresponding to the audio data of the multiple speakers, and the labels include clean labels and noise labels; when training the sequence-to-sequence network, calculating the second loss function of the sequence-to-sequence network includes: obtaining a fused prediction vector according to the acoustic feature encoding vector, the phoneme duration encoding vector, and the noise embedding vector; performing autoregressive decoding on the fused prediction vector at the decoding end of the sequence-to-sequence network, so that the sequence-to-sequence network learns to predict clean acoustic features when the label is a clean label and learns to predict noise acoustic features when the label is a noise label, and calculating the second loss function of the sequence-to-sequence network.

6. The method according to any one of claims 2-5, wherein, training the target speech synthesis model according to the target training set and the noise-added speech samples to obtain the trained target speech synthesis model includes: inputting the speech samples of the target speaker in the target training set and the noise-added speech samples into the target speech synthesis model for training, redefining the loss function until the redefined loss function converges, and obtaining the trained target speech synthesis model.

7. The method according to claim 6, wherein, redefining the loss function includes: redefining the second loss function of the sequence-to-sequence network as a structural similarity SSIM loss function; redefining the loss function of the target speech synthesis model according to the SSIM loss function.

8. The method according to claim 1, wherein, before training the preset speech synthesis model according to the pre-training set to obtain the speech synthesis model, it includes: obtaining the first audio data corresponding to multiple speakers and the text of the first audio data, as well as the second audio data and the text of the second audio data from a preset speech library, wherein the quality of the first audio data is higher than that of the second audio data; establishing a pre-training set according to the first audio data and the text of the first audio data, and the second audio data and the text of the second audio data.

9. The method according to claim 1, wherein, determining the target duration prediction network corresponding to the target application scenario and replacing the duration prediction network of the speech synthesis model with the target duration prediction network to obtain the target speech synthesis model includes: selecting the duration prediction network with the closest cosine distance to the duration prediction network from the duration prediction networks corresponding to the preset multiple application scenarios as the target duration prediction network corresponding to the target application scenario; or matching the duration prediction network corresponding to the target application scenario from the duration prediction networks corresponding to the preset multiple application scenarios as the target duration prediction network.

10. The method according to claim 1, wherein, the masking and noise adding process for the speech samples of the same type of speakers includes: adding noise to the acoustic features of the speech samples of the same type of speakers in the time domain and frequency domain.

11. A method for training a speech synthesis model, wherein, the method includes: pre-training a preset speech synthesis model according to a pre-training set to obtain a speech synthesis model, wherein the pre-training set includes speech samples of multiple speakers, and the speech synthesis model includes a duration prediction network; determining a target duration prediction network corresponding to a target application scenario, and replacing the duration prediction network of the speech synthesis model with the target duration prediction network to obtain a target speech synthesis model; obtaining a target training set, wherein the target training set includes speech samples of a target speaker; obtaining speech samples of the same type of speakers as the target speaker from the pre-training set, performing a masking and noise adding process on the speech samples of the same type of speakers to obtain noise-added speech samples; training the target speech synthesis model according to the target training set and the noise-added speech samples to obtain a trained target speech synthesis model.

12. A computer device, wherein, the computer device includes: a memory and a processor; wherein, the memory is connected to the processor and is used for storing programs; the processor is configured to implement the steps of the speech synthesis method according to any one of claims 1-10, or implement the steps of the method for training a speech synthesis model according to claim 11 by running the programs stored in the memory.

13. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the steps of the speech synthesis method according to any one of claims 1-10, or implement the steps of the method for training a speech synthesis model according to claim 11.

Citation Information

Patent Citations

  • Lightweight multi-speaker voice synthesis system and electronic equipment

    CN112133282A

  • Speech conversion method and system based on semi-parallel corpus

    CN112530403A