Speech synthesis method, network training method, device, equipment and storage medium
By using an end-to-end trained front-end network, acoustic model network, and vocoder network, the problem of insufficient intonation in Chinese speech synthesis was solved, achieving high-quality speech synthesis, reducing manpower and resource consumption, and improving user experience and cross-language adaptability.
Patent Information
- Application Number
- CN202310124566.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-16
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2043-02-16
AI Technical Summary
Existing technologies struggle to effectively address the issue of intonation and rhythm in Chinese speech synthesis, resulting in synthesized speech lacking naturalness, limited cross-language applicability, and high human and resource consumption.
End-to-end training is performed using a serially connected front-end network, acoustic model network, and vocoder network. High-quality audio waveforms are generated by acquiring prosodic features and phoneme sequences, reducing reliance on specialized linguistic knowledge and improving the naturalness and cross-linguistic adaptability of speech synthesis.
It improves the intonation and rhythm of Chinese speech synthesis, reduces the consumption of manpower and resources, enhances the quality of synthesized speech and user experience, and strengthens the ability to transfer models across languages.
Smart Images

Figure BDA0004081511150000101 
Figure HDA0004081511160000011 
Figure HDA0004081511160000021
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of speech processing, in particular to the technical field of artificial intelligence and speech synthesis, and more particularly to a speech synthesis method, a network training method, an apparatus, a device and a storage medium. BACKGROUND
[0002] Speech synthesis is a technology of converting text information into understandable and fluent spoken language output. The speech synthesis process is a process of converting text information into linguistic features or phonemes, and then converting the linguistic features or phonemes into audio waveforms.
[0003] Compared with other foreign languages, Chinese is a tonal language, and the feeling of ups and downs of the synthesized speech is an important indicator of naturalness. SUMMARY
[0004] The present disclosure provides a speech synthesis method, a network training method, an apparatus, a device and a storage medium to solve at least one of the above-mentioned defects.
[0005] According to a first aspect of the present disclosure, a speech synthesis method is provided, which comprises:
[0006] In response to receiving a text to be synthesized, obtaining prosodic features of the text to be synthesized and a phoneme sequence corresponding to the text to be synthesized by using a front-end network;
[0007] Obtaining acoustic features corresponding to the text to be synthesized according to the prosodic features and the phoneme sequence by using an acoustic model network;
[0008] Obtaining an audio waveform of the synthesized speech corresponding to the text to be synthesized according to the acoustic features by using a vocoder network;
[0009] The front-end network is a neural network for generating prosodic features and phoneme sequences, the acoustic model network is a neural network for generating acoustic features, and the vocoder network is a neural network for generating audio waveforms; the front-end network, the acoustic model network and the vocoder network are serially connected to form a speech synthesis network; the front-end network, the acoustic model network and the vocoder network are obtained by pre-training the speech synthesis network in an end-to-end manner.
[0010] According to a second aspect of the present disclosure, a network training method is provided for training a speech synthesis network, wherein the speech synthesis network comprises a front-end network, an acoustic model network and a vocoder network connected in series, the front-end network is a neural network for generating prosodic features and phoneme sequences, the acoustic model network is a neural network for generating acoustic features, and the vocoder network is a neural network for generating audio waveforms, and the method comprises:
[0011] inputting the to-be-trained text into the front-end network, obtaining prosody features and a phoneme sequence output by the front-end network;
[0012] inputting prosody labels and phoneme labels corresponding to the to-be-trained text into the acoustic model network, obtaining acoustic features output by the acoustic model network; and inputting the acoustic features output by the acoustic model network into the vocoder network to obtain audio waveforms output by the vocoder network;
[0013] constructing a first loss function according to the prosody features output by the front-end network and the prosody labels corresponding to the to-be-trained text, and the phoneme sequence output by the front-end network and the phoneme labels corresponding to the to-be-trained text;
[0014] constructing a second loss function according to the audio waveforms output by the vocoder network and audio waveforms of synthesized speech corresponding to the to-be-trained text;
[0015] adjusting network parameters of the speech synthesis network based on the first loss function and the second loss function.
[0016] According to a third aspect of the present disclosure, a speech synthesis apparatus is provided, which comprises:
[0017] a text front-end module configured to, in response to receiving to-be-synthesized text, obtain prosody features of the to-be-synthesized text and a phoneme sequence corresponding to the to-be-synthesized text by using a front-end network;
[0018] an acoustic model module configured to obtain acoustic features corresponding to the to-be-synthesized text according to the prosody features and the phoneme sequence by using an acoustic model network;
[0019] a vocoder module configured to obtain audio waveforms of synthesized speech corresponding to the to-be-synthesized text according to the acoustic features by using a vocoder network;
[0020] The front-end network, the acoustic model network and the vocoder network are obtained by performing end-to-end training on the speech synthesis network in advance.
[0021] According to a fourth aspect of the present disclosure, a network training apparatus is provided for training a speech synthesis network, the speech synthesis network comprising a front-end network, an acoustic model network, and a vocoder network connected in series, the front-end network being a neural network for generating prosodic features and a phoneme sequence, the acoustic model network being a neural network for generating acoustic features, and the vocoder network being a neural network for generating an audio waveform, the apparatus comprising:
[0022] a front-end training module configured to input a to-be-trained text into the front-end network, and obtain prosodic features and a phoneme sequence output by the front-end network;
[0023] an acoustic training module configured to input prosodic labels and phoneme labels corresponding to the to-be-trained text into the acoustic model network, and obtain acoustic features output by the acoustic model network; and input the acoustic features output by the acoustic model network into the vocoder network to obtain an audio waveform output by the vocoder network;
[0024] a first loss module configured to construct a first loss function according to the prosodic features output by the front-end network and prosodic labels corresponding to the to-be-trained text, and the phoneme sequence output by the front-end network and phoneme labels corresponding to the to-be-trained text;
[0025] a second loss module configured to construct a second loss function according to the audio waveform output by the vocoder network and an audio waveform of synthesized speech corresponding to the to-be-trained text;
[0026] a back propagation module configured to adjust network parameters of the speech synthesis network based on the first loss function and the second loss function.
[0027] According to a fifth aspect of the present disclosure, an electronic device is provided, the electronic device comprising:
[0028] at least one processor; and
[0029] a memory communicatively connected to the at least one processor; wherein
[0030] the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the speech synthesis method and the network training method.
[0031] According to a sixth aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to perform the speech synthesis method and the network training method.
[0032] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the above voice synthesis method and network training method.
[0033] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0034] The accompanying drawings are used to better understand the present scheme and do not constitute a limitation on the present disclosure. Among them:
[0035] Figure 1 is a flowchart of a voice synthesis method provided by an embodiment of the present disclosure;
[0036] Figure 2 is a flowchart of part of steps of another voice synthesis method provided by an embodiment of the present disclosure;
[0037] Figure 3 is a flowchart of part of steps of another voice synthesis method provided by an embodiment of the present disclosure;
[0038] Figure 4 is a flowchart of a network training method provided by an embodiment of the present disclosure;
[0039] Figure 5 is a structural schematic diagram of a voice synthesis device provided by an embodiment of the present disclosure;
[0040] Figure 6 is a structural schematic diagram of a network training device provided by an embodiment of the present disclosure;
[0041] Figure 7 is a block diagram of an electronic device for implementing the voice synthesis method and network training method of the embodiments of the present disclosure. DETAILED DESCRIPTION
[0042] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, the description below omits the description of well-known functions and structures.
[0043] In some related technologies, text information is converted into linguistic features or phonemes by setting rules.
[0044] This conversion mode requires the rule setting developer to have strong professional linguistic knowledge, is prone to errors, difficult to maintain, and extremely high in manpower cost. Moreover, due to the great changes in grammar rules of different voices, the rules cannot be borrowed from each other, so that linguistic experts of corresponding languages need to be involved in the development of different languages, resulting in weak cross-voice application.
[0045] In some related technologies, after solving the text regularization problem by setting rules, the regularized text is processed by using a neural network to perform word segmentation, tone conversion, and multi-phonetic character prediction, to obtain linguistic features or phonemes corresponding to the text information.
[0046] In some related technologies, the linguistic features or phonemes corresponding to the text information are converted into acoustic features, such as mel-spectrum features, by using an acoustic model.
[0047] In some related technologies, the acoustic features corresponding to the text information are converted into audio waveforms by using a vocoder.
[0048] By introducing a neural network, the demand for professional linguistic knowledge is reduced, the manpower cost is reduced, and the model can be migrated between different languages by replacing a data set. However, after the linguistic features or phonemes corresponding to the text information obtained by the neural network are input into the acoustic model and the vocoder, the synthesized voice has a strong mechanical feeling.
[0049] The speech synthesis method and the network training method provided in the embodiments of the present disclosure aim to solve at least one of the above technical problems in the prior art.
[0050] The speech synthesis method and the network training method provided in the embodiments of the present disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be a vehicle-mounted device, a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The method can be implemented by a processor invoking computer-readable program instructions stored in a memory. Alternatively, the method can be executed by a server.
[0051] Figure 1 A flowchart of a speech synthesis method provided in the embodiments of the present disclosure is shown, as shown in Figure 1 The method can mainly include the following steps.
[0052] In step S110, in response to receiving a text to be synthesized, a front-end network is used to obtain prosodic features of the text to be synthesized and a phoneme sequence corresponding to the text to be synthesized.
[0053] In step S120, an acoustic model network is adopted to obtain acoustic features corresponding to the text to be synthesized according to the prosodic features and the phoneme sequence;
[0054] In step S130, a vocoder network is adopted to obtain an audio waveform of the synthesized speech corresponding to the text to be synthesized according to the acoustic features;
[0055] The front-end network is a neural network for generating prosodic features and phoneme sequences, the acoustic model network is a neural network for generating acoustic features, and the vocoder network is a neural network for generating an audio waveform. The front-end network, the acoustic model network, and the vocoder network are serially connected to form a speech synthesis network. The front-end network, the acoustic model network, and the vocoder network are obtained by performing end-to-end training on the speech synthesis network in advance.
[0056] For example, in step S110, the text to be synthesized can be a text input by a user through a human-computer interaction device.
[0057] In some possible implementation manners, after the human-computer interaction device obtains the text to be synthesized input by the user, the text to be synthesized can be obtained from the human-computer interaction device in a wired or wireless manner.
[0058] In some possible implementation manners, the text to be synthesized can be a Chinese text, and specifically can be a text composed of one or more Chinese sentences.
[0059] In some possible implementation manners, the text to be synthesized can contain other types of input, such as codes, cardinal numbers, ordinal numbers, numerical ranges, dates, temperatures, fractions, decimals, percentages, and telephone numbers.
[0060] In some possible implementation manners, the text to be synthesized is subjected to text normalization processing, and other types of input in the text to be synthesized are converted into pure text input.
[0061] In some possible implementation manners, text normalization is a process of converting numbers, symbols, abbreviations, and the like in a text into language characters.
[0062] In some possible implementation manners, the prosodic features corresponding to the text to be synthesized can be prosodic levels to which text words and text sentences included in the text to be synthesized belong.
[0063] In some possible implementation manners, the prosodic levels are one of prosodic words, prosodic phrases, intonation phrases, and ends of sentences.
[0064] In some possible implementation manners, the phoneme sequence corresponding to the text to be synthesized can be a set of phonemes corresponding to each character of the text to be synthesized arranged in sequence.
[0065] In some possible implementation manners, the front-end network is used to perform sentence segmentation, word segmentation, part-of-speech tagging, prosody prediction, and multi-phonetic character prediction on the text to be synthesized, and obtain the prosody features and the phoneme sequence corresponding to the text to be synthesized.
[0066] In some possible implementation manners, the front-end model can be composed of a feature extraction subnetwork and a plurality of subnetworks corresponding to a plurality of subtasks. The feature extraction subnetwork is used to extract semantic features of the text to be synthesized, and each subnetwork is used to complete one of the tasks of sentence segmentation, word segmentation, part-of-speech tagging, prosody prediction, phonetic conversion, and multi-phonetic character prediction according to the extracted semantic features. The prosody prediction subnetwork is a neural network used to generate prosody features, and the phonetic conversion subnetwork is a neural network used to generate a phoneme sequence.
[0067] That is, the sentence segmentation, word segmentation, part-of-speech tagging, prosody prediction, multi-phonetic character prediction, and generation of the phoneme sequence are completed in parallel.
[0068] In some possible implementation manners, the front-end model can be composed of a feature extraction subnetwork, a sentence segmentation subnetwork, a word segmentation subnetwork, a part-of-speech tagging subnetwork, a prosody prediction subnetwork, a phonetic conversion subnetwork, and a multi-phonetic character prediction subnetwork connected in series.
[0069] That is, the feature extraction subnetwork extracts semantic features of the text to be synthesized, the sentence segmentation subnetwork is used to obtain at least one text sentence corresponding to the text to be synthesized according to the extracted semantic features, the word segmentation subnetwork is used to obtain a plurality of text words corresponding to the text to be synthesized according to the output of the sentence segmentation subnetwork, the part-of-speech tagging subnetwork is used to obtain the parts of speech of the plurality of text words according to the output of the sentence segmentation subnetwork, the prosody prediction subnetwork is used to obtain the prosody features of the text to be synthesized according to the output of the part-of-speech tagging subnetwork, the phonetic conversion subnetwork is used to obtain the phoneme sequence of the text to be synthesized according to the output of the prosody prediction subnetwork, and the multi-phonetic character prediction subnetwork is used to determine the phonemes corresponding to the multi-phonetic characters in the text to be synthesized.
[0070] In some possible implementation manners, a rule-based and front-end model combined manner can be used to obtain the prosody features of the text to be synthesized and the phoneme features corresponding to the text to be synthesized.
[0071] In some possible implementation manners, the text to be synthesized can be subjected to sentence segmentation, word segmentation, and part-of-speech tagging processing based on preset rules, a plurality of text words corresponding to the text to be synthesized are obtained, the prosody features and the phoneme sequence corresponding to the text to be synthesized are obtained based on the front-end model and the plurality of text words corresponding to the text to be synthesized, and the phonemes in the obtained phoneme sequence are subjected to tonal conversion processing based on the preset rules, so as to finally determine the phoneme sequence corresponding to the text to be synthesized.
[0072] Among them, tone change refers to the phenomenon that the tone values of some syllables are affected by the following tone and change when the syllables are continuously emitted.
[0073] In some possible implementations, in step S120, the acoustic model network can be any neural network that can generate acoustic features.
[0074] In some possible implementations, the acoustic model network can be an autoregressive deep learning network, or a non-autoregressive deep learning network.
[0075] In some specific implementations, the acoustic model network can be a non-autoregressive FastSpeech2 network.
[0076] In some possible implementations, the acoustic features corresponding to the text to be synthesized can be the Mel spectrum corresponding to the text to be synthesized.
[0077] In some possible implementations, the LN (Layer Normalization) layer of the acoustic model network (such as FastSpeech2) can be replaced by a learnable adaptive Tensor, which also changes during the training of the acoustic model network. During the training of the acoustic model network, the Tensor will automatically change towards the optimal angle of the loss function.
[0078] The Tensor is used for adaptive normalization processing of the features output by the neural network layer connected thereto, so that the acoustic network model can perform better in the speech synthesis task.
[0079] In some possible implementations, the vocoder network can be any neural network that can generate audio waveforms.
[0080] In some specific implementations, the vocoder network can be a pre-trained Multi-band MelGAN.
[0081] In some possible implementations, the audio waveform output by the vocoder network can be used to synthesize corresponding speech and provided to a playback device to play the speech, so that the user can obtain the synthesized speech corresponding to the text to be synthesized.
[0082] In some specific implementations, the playback device can be a device with specific playback functions, such as a smart speaker.
[0083] In some possible implementations, the front-end network, the acoustic model network, and the vocoder network are connected in series and trained together.
[0084] That is, the front-end network, the acoustic model network, and the vocoder network form an end-to-end speech synthesis network, and the trained front-end network, acoustic model network, and vocoder network are obtained by training the speech synthesis network end to end.
[0085] In some possible manners, the text to be trained is input into the front-end network to obtain prosodic features and a phoneme sequence output by the front-end network; the prosodic label and the phoneme label corresponding to the text to be trained are input into the acoustic model network to obtain acoustic features output by the acoustic model network; the acoustic features output by the acoustic model network are input into the vocoder network to obtain an audio waveform output by the vocoder network; a first loss function is constructed according to the prosodic features output by the front-end network and the prosodic label corresponding to the text to be trained, and the phoneme sequence output by the front-end network and the phoneme label corresponding to the text to be trained; a second loss function is constructed according to the audio waveform output by the vocoder network and the audio waveform of the synthesized speech corresponding to the text to be trained; and the network parameters of the speech synthesis network are adjusted based on the first loss function and the second loss function.
[0086] In the speech synthesis method provided in the embodiments of the present disclosure, the front-end network, the acoustic model network, and the vocoder network are serially connected to form an end-to-end trained and end-to-end inferred speech synthesis network, and the text is provided to the speech synthesis network, so that the audio waveform of the synthesized speech corresponding to the text can be directly obtained without the need for other processing by the user or the engineer, thereby reducing the waste of manpower and material resources and being more friendly and efficient for the user and the engineer.
[0087] Meanwhile, the prosodic features of the text to be synthesized are obtained by the front-end network, and the acoustic model network learns the prosodic features of the text to be synthesized, thereby enhancing the prosodic sense of the synthesized speech, reducing the mechanical sense of the synthesized speech, improving the quality of the obtained synthesized speech, and further improving the user experience.
[0088] The speech synthesis method provided in the embodiments of the present disclosure is described below.
[0089] As described above, in some possible implementation manners, the text to be synthesized can be segmented, tokenized, and morphologically annotated based on a preset rule to obtain a plurality of text words corresponding to the text to be synthesized, the prosodic features and the phoneme sequence corresponding to the text to be synthesized can be obtained based on the front-end model and the plurality of text words corresponding to the text to be synthesized, and the phonemes in the obtained phoneme sequence can be processed for tonal variation based on a preset rule to finally determine the phoneme sequence corresponding to the text to be synthesized.
[0090] Figure 2A flowchart of a step of obtaining prosodic features and a phoneme sequence of the text to be synthesized based on a preset rule and a front-end model in a speech synthesis method provided by an embodiment of the present disclosure is shown in FIG. 2. As shown in FIG. 2, the step mainly can include: Figure 2
[0091] In step S210, text normalization processing is performed on the text to be synthesized to obtain a normalized text corresponding to the text to be synthesized.
[0092] In step S220, a plurality of text words corresponding to the text to be synthesized are determined according to the normalized text.
[0093] In step S230, prosodic features and a phoneme sequence are obtained based on the front-end network and the plurality of text words corresponding to the text to be synthesized.
[0094] In step S240, the phoneme sequence is subjected to tonal conversion processing.
[0095] In some possible implementation manners, in step S210, the text normalization processing on the text to be synthesized can be converting various non-standard types of text in the text to be synthesized into pure text input.
[0096] In some possible implementation manners, the obtained pure text input is determined as the normalized text corresponding to the text to be synthesized.
[0097] In some specific implementation manners, the text normalization is completed using a pure rule manner, and an original input and a normalized output of the text normalization unit are shown in the following table.
[0098]
[0099] In some possible implementation manners, in step S220, the plurality of text words corresponding to the text to be synthesized are determined according to the normalized text corresponding to the text to be synthesized, which can be dividing the normalized text into at least one text sentence based on a preset rule; and dividing the text sentence into a plurality of text words based on a preset rule.
[0100] In some possible implementation manners, dividing the normalized text into at least one text sentence based on a preset rule can be performing sentence division on the normalized text according to punctuation marks or other rules to obtain the text sentence corresponding to the text to be synthesized.
[0101] In some possible implementation manners, dividing the text sentence into a plurality of text words based on a preset rule can be performing word segmentation processing on the text sentence to divide the text sentence into a plurality of text words.
[0102] In some possible implementation ways, the front-end network comprises a feature extraction sub-network, a prosody prediction sub-network, and a grapheme-to-phoneme conversion sub-network.
[0103] The feature extraction sub-network is used for extracting semantic features, the prosody prediction sub-network is a neural network used for generating prosody features, and the grapheme-to-phoneme conversion sub-network is a neural network used for generating a phoneme sequence.
[0104] In some possible implementation ways, the prosody features comprise a prosody hierarchy sequence corresponding to the text to be synthesized, and in some possible implementation ways, the prosody hierarchy is one of a prosodic word, a prosodic phrase, a tonal phrase, and a sentence end.
[0105] In some possible implementation ways, the phoneme sequence corresponding to the text to be synthesized can be a set of phonemes corresponding to each character in the text to be synthesized arranged in sequence.
[0106] In some possible implementation ways, in step S230, based on the front-end network, the prosody features and the phoneme sequence are obtained according to the multiple text words corresponding to the text to be synthesized, which can comprise: inputting the multiple text words corresponding to the text to be synthesized into the feature extraction sub-network to obtain semantic features corresponding to the text to be synthesized; inputting the semantic features into the prosody prediction sub-network to obtain the prosody features; and inputting the semantic features into the grapheme-to-phoneme conversion sub-network to obtain the phoneme sequence.
[0107] In some possible implementation ways, the feature extraction sub-network can be an ERNIE (Enhanced Representation from kNowledge IntEgration) network.
[0108] In some possible implementation ways, the ERNIE network can be connected in parallel with multiple sub-neural networks corresponding to multiple sub-tasks, such as the prosody prediction sub-network and the grapheme-to-phoneme conversion sub-network.
[0109] Compared with a serial connection mode, parallel execution of multiple sub-tasks (such as prosody prediction and grapheme-to-phoneme conversion) can not only improve the processing speed, but also avoid mutual dependence and mutual influence between the sub-tasks, such as mutual influence between the prosody prediction task and the grapheme-to-phoneme conversion task, thereby improving the accuracy and efficiency of the prosody prediction task and the grapheme-to-phoneme conversion task, and further improving the accuracy of the obtained prosody features and phoneme sequence and the accuracy of the speech synthesis.
[0110] Meanwhile, the prosody prediction task, the grapheme-to-phoneme conversion task, and the tasks of sentence segmentation, word segmentation, and tone conversion are processed in series, which is equivalent to pre-processing the input of the feature extraction sub-network, thereby improving the quality of the input of the feature extraction sub-network, further improving the quality of the extracted semantic features, and improving the accuracy of the obtained prosody features and phoneme sequence and the accuracy of the speech synthesis.
[0111] It should be emphasized that in the related art, there is no consistent conclusion on the relationship between the prosody prediction task and the grapheme-to-phoneme conversion task, and the relationship with the sentence segmentation, word segmentation, and tone change tasks. After a plurality of control tests, when the relationship between the prosody prediction task and the grapheme-to-phoneme conversion task, and the relationship with the sentence segmentation, word segmentation, and tone change tasks are as shown in the speech synthesis method provided in the embodiments of the present disclosure, the performance and effect of the speech synthesis task can be optimal.
[0112] In some possible implementation manners, the text corpus of Zhabei and AISHLL3 can be used as a training set to pre-train the feature extraction sub-network and the prosody prediction sub-network.
[0113] In some possible implementation manners, the grapheme-to-phoneme conversion sub-network can be a G2P (Grapheme-to-Phoneme) conversion network.
[0114] In some possible implementation manners, the grapheme-to-phoneme conversion sub-network can include a multi-pronunciation word prediction sub-network. The multi-pronunciation word prediction sub-network is configured to acquire multi-pronunciation words in the text to be synthesized according to the semantic features of the text to be synthesized, and determine phonemes corresponding to the multi-pronunciation words.
[0115] In some possible implementation manners, inputting the semantic features into the grapheme-to-phoneme conversion sub-network to acquire the phoneme sequence can include inputting the semantic features into the grapheme-to-phoneme conversion sub-network, determining multi-pronunciation words in the text to be synthesized, and determining phonemes corresponding to the multi-pronunciation words.
[0116] The multi-pronunciation word prediction sub-network can be used to predict phonemes corresponding to multi-pronunciation words in the text information, improve the accuracy of the phonemes corresponding to the multi-pronunciation words, and further improve the accuracy of the synthesized speech corresponding to the multi-pronunciation words.
[0117] In some possible implementation manners, in step S240, the tone change processing on the phoneme sequence can be tone change processing on the phoneme sequence based on a preset rule.
[0118] Tone change refers to a phenomenon that when Chinese syllables are continuously pronounced, the tone values of some syllables are affected by the following tone and change.
[0119] Tone change processing is a processing of using a tone change method on syllables of words and sentences.
[0120] Text normalization, sentence processing, word segmentation processing, and tone change processing are all simple and clear text processing. Compared with using a neural network to perform text normalization, sentence processing, word segmentation processing, and tone change processing on the synthesized text, the accuracy of the results obtained by performing text normalization, sentence processing, word segmentation processing, and tone change processing on the synthesized text based on rules is not low, but the implementation is simpler, the resources occupied are less, and the processing efficiency is higher.
[0121] The accuracy of the results obtained by performing prosody prediction and multi-phonetic character prediction based on rules is much lower than the accuracy of the results obtained by performing prosody prediction and multi-phonetic character prediction using a neural network. Therefore, in summary, based on rules, text normalization, sentence processing, word segmentation processing, and tone change processing are performed on the synthesized text, and prosody prediction and multi-phonetic character prediction are performed using a neural network, which can reduce the occupation of resources and improve the processing efficiency on the basis of ensuring accuracy.
[0122] The following is a specific example to show how to obtain the prosodic features of the synthesized text and the phoneme sequence corresponding to the synthesized text in the speech synthesis method provided by the embodiments of the present disclosure.
[0123] In the case of the original input, that is, the synthesized text is "There are 112 211 colleges in the country", the normalized text obtained by performing normalization processing on the synthesized text is "There are 112 211 colleges in the country"; the text sentence obtained by performing sentence processing on the normalized text is "There are 112 211 colleges in the country" (since the original input is a sentence, the text sentence obtained is the same as the normalized text); the text words obtained by performing word segmentation on the text sentence are "There are 112 211 colleges in the country", that is, the input is divided into "There are 112 211 colleges in the country" multiple text words; based on the front-end network, the phoneme sequence "quan2 guo2 yi2 gong4 you3 yi4 bai3 yi1 shi2er4 suo3 er4 yao1 yao1 gao1 xiao4" is obtained, and the prosodic level "There are #2 112 211 colleges #2 in the country #4" is obtained, wherein #1 is a prosodic word, #2 is a prosodic phrase, #3 is a tone phrase, and #4 is the end of a sentence.
[0124] In some possible implementation manners, the acoustic model network can be any neural network that can obtain the acoustic features of the synthesized text according to the prosodic features and the phoneme sequence.
[0125] In some specific implementation manners, the acoustic model network can be a non-autoregressive FastSpeech2 (fast speech) network.
[0126] The FastSpeech2 model is composed of a Phoneme Embedding, an Encoder, a Variance adaptor, a Mel-spectrogram, and the like. The Phoneme Embedding and the Encoder are used to extract corresponding phoneme features. The Variance adaptor is used to predict and process the duration, pitch, and energy of the phoneme features, and then send the features to a Decoder to synthesize a Mel-spectrogram.
[0127] Any module capable of achieving a function can be used in the embodiments of the present disclosure, and therefore, the embodiments of the present disclosure do not limit the structure of the Phoneme Embedding, the Encoder, the Variance adaptor, and the Mel-spectrogram.
[0128] In some possible implementations, the acoustic features corresponding to the text to be synthesized can be Mel (Mell) spectrums corresponding to the text to be synthesized.
[0129] In some possible implementations, the MFA (Montreal-Forced-Aligner, speech forced aligner) can be used to align the prosodic labels with the duration of the TIMIT corpus, and then a version of the labels without duration is designed using #1-4 of the original corpus. The FastSpeech2 is trained using the training data with the labels with duration and the training data with the labels without duration, respectively.
[0130] In some specific implementations, since the acoustic model network obtained by training the FastSpeech2 using the training data with the labels with duration has better effect, the FastSpeech2 is pre-trained using the training data with the labels with duration to obtain the acoustic model network.
[0131] In some possible implementations, the LN (Layer Normalization, layer normalization) layer of the acoustic model network (such as FastSpeech2) can be replaced by a learnable adaptive Tensor (tensor), i.e., an adaptive parameter. The Tensor also changes in the training process of the acoustic model network. In the training process of the acoustic model network, the Tensor will automatically change towards the angle that optimizes the loss function.
[0132] Figure 3This diagram illustrates the process of replacing the LN layers of an acoustic model network (such as FastSpeech2) with learnable adaptive Tensors, and then using the acoustic model network to obtain the acoustic features corresponding to the text to be synthesized based on prosodic features and phoneme sequences. Figure 3 As shown, this step specifically includes:
[0133] In step S310, prosodic features and phoneme sequences are input into the acoustic model network to obtain the features output by the neural network layer of the acoustic model network;
[0134] In step S320, adaptive parameters are used to normalize the features output by the neural network layer of the acoustic model network, and the acoustic features of the text to be synthesized are determined based on the normalized features.
[0135] In some possible implementations, in step S310, obtaining the features of the neural network layer output of the acoustic model network can be obtaining the features of the neural network layer output connected to the adaptive parameters, that is, the features of the neural network layer output connected to the LN layer in the original FastSpeech2.
[0136] In some possible implementations, the adaptive parameter can be the mean-variance parameter, which can be used to normalize the features output by the neural network layers.
[0137] In some possible implementations, in step S320, determining the acoustic features of the text to be synthesized based on the normalized features can be achieved by obtaining the output of an acoustic model network in which the LN layer is replaced by a learnable adaptive Tensor as the acoustic features of the text to be synthesized.
[0138] By using adaptive parameters to adaptively normalize the features output by the connected neural network layers, the learning of acoustic network models can be made more flexible, enabling them to perform better in speech synthesis tasks.
[0139] Meanwhile, the acoustic model network, which replaces the LN layer with a learnable adaptive Tensor, can flexibly process prosodic labels with and without duration, making the acoustic model network more widely applicable and leveraging its applications.
[0140] In some possible implementations, the vocoder network can be any neural network capable of acquiring audio waveforms based on acoustic features.
[0141] In some specific implementations, the vocoder network can be a pre-trained Multi-band MelGAN (Multi-band Mel Adversarial Network).
[0142] The main structure of the multi-band MelGAN model includes two parts, a Generator and a Discriminator. The input of the Generator is a Mel spectrum, and the output is an audio waveform. The Discriminator is used to determine whether the audio waveform generated by the Generator is true. The Generator is trained according to the determination result of the Discriminator.
[0143] Before the audio waveform generated by the Generator is input into the Discriminator, it needs to be processed by an Avg Pool (average pooling) to obtain audio waveforms of different bandwidths. The audio waveforms of different bandwidths are input into an Analysis FilterBank (analysis filter bank) for processing, such as suppressing high-frequency parts and processing overlapping parts, to obtain processed audio waveforms of different bandwidths for input into the Discriminator.
[0144] The main part of the Generator includes a Conv1D (one-dimensional convolution layer), an Upsample (up-sampling layer), a ResidualBlock (residual block), a Conv1D Tanh (one-dimensional convolution plus hyperbolic tangent layer), and other modules. These modules are commonly used in deep learning and will not be explained in detail.
[0145] In the inference process of the Multi-band MelGAN model, the trained Generator is used, and the Discriminator is not used.
[0146] In some possible implementations, the audio waveform output by the vocoder network can be used to synthesize corresponding speech and provided to a playback device to play the speech for a user to obtain synthesized speech corresponding to the text to be synthesized.
[0147] In some specific implementations, the playback device can be a device with specific playback functions, such as a smart speaker.
[0148] Figure 4 A flowchart of a network training method provided by an embodiment of the present disclosure is shown, as shown in Figure 4 The method includes the following steps.
[0149] In step S410, the text to be trained is input into the front-end network to obtain prosodic features and phoneme sequences output by the front-end network.
[0150] In step S420, the prosodic label and the phoneme label corresponding to the text to be trained are input into the acoustic model network to obtain acoustic features output by the acoustic model network; and the acoustic features output by the acoustic model network are input into the vocoder network to obtain audio waveforms output by the vocoder network.
[0151] In step S430, a first loss function is constructed according to the prosodic features output by the front-end network and the prosodic label corresponding to the text to be trained, and the phoneme sequence output by the front-end network and the phoneme label corresponding to the text to be trained.
[0152] In step S440, a second loss function is constructed according to the audio waveforms output by the vocoder network and the audio waveforms of the synthesized speech corresponding to the text to be trained.
[0153] In step S450, the network parameters of the speech synthesis network are adjusted based on the first loss function and the second loss function.
[0154] The front-end network, the acoustic model network, and the vocoder network are serially connected to form the speech synthesis network, the front-end network is a neural network for generating prosodic features and phoneme sequences, the acoustic model network is a neural network for generating acoustic features, and the vocoder network is a neural network for generating audio waveforms.
[0155] For example, in step S410, the text to be trained can be the text corpus of Pinyin and AISHELL3.
[0156] In some possible implementation manners, the prosodic label corresponding to the text to be trained can be a label of a prosodic level corresponding to the text to be trained.
[0157] The following table is the prosodic level and corresponding label of the text corpus of Pinyin (csmsc) and AISHELL3.
[0158] ryh_token csmsc aishll3 % #1 % ` #2 Figure 1 #3 $ #4 $
[0159] The Pinyin has four prosodic levels, namely #1 prosodic word, #2 prosodic phrase, #3 intonation phrase, and #4 end of sentence, and the AISHELL3 has two prosodic levels, namely #1 prosodic word and #4 end of sentence, rhy_token is a special placeholder corresponding to different prosodic levels, and rhy_token is a label for real training and prediction, and then the corresponding prosodic level is restored by using the mapping.
[0160] That is, in the training process, the label % represents a prosodic word, the label ` represents a prosodic phrase, the label ˉ represents an intonation phrase, and the label $ represents an end of sentence.
[0161] The phoneme label corresponding to the text to be trained can be a set of labels of actual phonemes of each character in the text to be trained arranged in sequence.
[0162] In some possible implementation manners, the front-end network comprises a feature extraction sub-network, a prosody prediction sub-network, and a grapheme-to-phoneme conversion sub-network.
[0163] The feature extraction sub-network is a neural network for extracting semantic features; the prosody prediction sub-network is a neural network for generating prosody features; and the grapheme-to-phoneme conversion sub-network is a neural network for generating a phoneme sequence.
[0164] In some possible implementation manners, the inputting of the text to be trained into the front-end network and the obtaining of the prosody features and the phoneme sequence output by the front-end network can comprise: obtaining a plurality of text vocabularies corresponding to the text to be trained, inputting the plurality of text vocabularies corresponding to the text to be trained into the feature extraction sub-network, and obtaining semantic features corresponding to the text to be trained; inputting the semantic features into the prosody prediction sub-network, and obtaining the prosody features; and inputting the semantic features into the grapheme-to-phoneme conversion sub-network, and obtaining the phoneme sequence.
[0165] In some possible implementation manners, the feature extraction sub-network can be an ERNIE (Enhanced Representation from kNowledge IntEgration) network.
[0166] In some possible implementation manners, the ERNIE network can be connected to a plurality of sub-neural networks corresponding to a plurality of sub-tasks, such as the prosody prediction sub-network and the grapheme-to-phoneme conversion sub-network, in parallel.
[0167] In some possible implementation manners, in step S420, the MFA can be used to align the prosody labels with time lengths of the text to be trained, and the prosody labels with time lengths are used to train the speech synthesis model.
[0168] In some specific implementation manners, the MFA is used to determine time length information corresponding to the prosody labels of the text to be trained, the prosody labels corresponding to the text to be trained, the phoneme labels corresponding to the text to be trained, and the time length information corresponding to the prosody labels of the text to be trained are input into the acoustic model network, and acoustic features output by the acoustic model network are obtained.
[0169] In some possible implementation manners, the prosody labels corresponding to the text to be trained and the phoneme labels are fused together to train the speech synthesis model.
[0170] In some specific implementation manners, the prosody labels corresponding to the text to be trained and the phoneme labels corresponding to the text to be trained are input into the acoustic model network to obtain training fusion features, and the acoustic features output by the acoustic model network are obtained based on the training fusion features.
[0171] In some possible implementation manners, features are extracted from the prosody label corresponding to the to-be-trained text and the phoneme label corresponding to the to-be-trained text respectively, and then the extracted features are fused to train the speech synthesis model.
[0172] In some specific implementation manners, the prosody label corresponding to the to-be-trained text is input into the acoustic model network to obtain training prosody features; the phoneme label corresponding to the to-be-trained text is input into the acoustic model network to obtain training phoneme features; the training prosody features and the training phoneme features are fused to obtain acoustic features output by the acoustic model network based on the fused features.
[0173] Through multiple sets of comparative tests, among the three different training manners, the speech synthesis model obtained by using the MFA to align the prosody label with a time length of the to-be-trained text and training the speech synthesis model by using the prosody label with a time length is better.
[0174] In some possible implementation manners, in step S450, adjusting the network parameters of the speech synthesis network based on the first loss function and the second loss function can include: adjusting the network parameters of the front-end network based on the first loss function; and adjusting the network parameters of the acoustic model network and the vocoder network based on the second loss function.
[0175] That is, the end-to-end training process of the speech synthesis model can be that the to-be-trained text is input into the front-end network to obtain prosody features and phoneme sequences predicted by the front-end network, then loss alignment is performed on the prosody features and the phoneme sequences predicted by the front-end network and the real prosody label and the phoneme sequence of the to-be-trained text, and a first loss function is constructed; meanwhile, the real prosody label and the phoneme sequence of the to-be-trained text are input into the acoustic model network and the vocoder network in series to obtain audio waveforms predicted by the vocoder network, then loss alignment is performed on the audio waveforms predicted by the vocoder network and the audio waveforms of the real synthesized speech, and a second loss function is constructed; the network parameters of the front-end network are modified through back propagation according to the first loss function; and the network parameters of the acoustic model network and the vocoder network are modified through back propagation according to the second loss function.
[0176] In some possible implementation manners, the sum of the first loss function and the second loss function can be used as a total loss function, and all the network parameters of the speech synthesis network are adjusted through the total loss function.
[0177] In some possible implementation manners, the acoustic model network can be any neural network that can generate acoustic features.
[0178] In some specific implementation manners, the acoustic model network can be a non-autoregressive FastSpeech2 (fast speech) network.
[0179] The FastSpeech2 model is composed of a phoneme embedding, an encoder, a variance adaptor, a mel-spectrogram, and the like. Among them, the phoneme embedding and the encoder are used to extract corresponding phoneme features, the variance adaptor is used to predict and process the duration, pitch, and energy of the phoneme features, and then the features are sent to the decoder to synthesize the mel spectrum.
[0180] Any module that can realize a function can be used in the embodiments of the present disclosure, and therefore, the embodiments of the present disclosure do not limit the structure of the phoneme embedding, the encoder, the variance adaptor, and the mel-spectrogram.
[0181] In some possible implementation manners, the LN (Layer Normalization) layer of an acoustic model network (such as FastSpeech2) can be replaced by a learnable adaptive tensor, that is, an adaptive parameter, which is also changed during the training of the acoustic model network. During the training of the acoustic model network, the adaptive parameter will automatically change towards an angle that optimizes the loss function.
[0182] In some possible implementation manners, inputting the prosody label and the phoneme label corresponding to the text to be trained into the acoustic model network to obtain the acoustic feature output by the acoustic model network can include: inputting the prosody label and the phoneme label corresponding to the text to be trained into the acoustic model network to obtain the feature output by the neural network layer of the acoustic model network; performing normalization processing on the feature output by the neural network layer of the acoustic model network using the adaptive parameter, and obtaining the acoustic feature output by the acoustic model network according to the normalized feature.
[0183] In some possible implementation manners, obtaining the feature output by the neural network layer of the acoustic model network can be obtaining the feature output by the neural network layer connected to the adaptive parameter, that is, the feature output by the neural network layer connected to the LN layer in the original FastSpeech2.
[0184] In some possible implementation manners, the adaptive parameter can be a mean-variance parameter, which can be used to perform normalization processing on the feature output by the neural network layer.
[0185] In some possible implementation manners, determining the acoustic feature of the text to be synthesized according to the normalized feature can be obtaining the output of the acoustic model network in which the LN layer is replaced by the learnable adaptive tensor as the acoustic feature of the text to be synthesized.
[0186] By using the adaptive parameter to adaptively normalize the features output by the neural network layer connected thereto, the learning of the acoustic network model can be more flexible, so that the acoustic network model can perform better in the speech synthesis task.
[0187] At the same time, the acoustic model network in which the LN layer is replaced by a learnable adaptive Tensor can flexibly process the prosodic label with time length and the prosodic label without time length, so that the acoustic model network has a wider range of use and more applications of the acoustic model network.
[0188] In some possible implementation manners, the vocoder network can be any neural network that can generate an audio waveform.
[0189] In some specific implementation manners, the vocoder network can be a pre-trained Multi-band MelGAN.
[0190] The main structure of the Multi-band MelGAN model includes two parts, a Generator and a Discriminator. The input of the Generator is a Mel spectrum, and the output is an audio waveform. The Discriminator is used to determine whether the audio waveform generated by the Generator is true, and the Generator is trained according to the determination result of the Discriminator.
[0191] Before the audio waveform generated by the Generator is input into the Discriminator, it needs to be processed by an Avg Pool to obtain audio waveforms of different bandwidths, and the audio waveforms of different bandwidths are input into an Analysis FilterBank for processing, such as suppressing the high-frequency part and processing the overlapping part, to obtain the processed audio waveforms of different bandwidths for input into the Discriminator.
[0192] The main part of the Generator includes a Conv1D, an Upsample, a Residual Block, a Conv1DTanh and the like. These modules are commonly used in deep learning and will not be explained in detail here.
[0193] In the network training method provided in the embodiments of the present disclosure, the front-end network, the acoustic model network and the vocoder network are serially connected to form a speech synthesis network, and end-to-end training of the speech synthesis network is implemented, that is, only the text to be trained and the corresponding label are required in the training process of the speech synthesis network, and no other processing is required, so that the training of the speech synthesis network is more convenient, and resource waste caused by information transmission in different processing processes is reduced.
[0194] Meanwhile, the prosody features of the text to be synthesized are obtained through the front-end network, and the prosody features of the text to be synthesized are learned by the acoustic model network, so that the prosody of the synthesized speech is enhanced, the mechanical feeling of the speech generated by the trained model is reduced, and the quality of the generated speech is improved.
[0195] Based on the same principle as the method shown in Figure 5 Figure 5 The structure of the speech synthesis device provided by the embodiments of the present disclosure is shown in the schematic diagram, as shown in the schematic diagram, the speech synthesis device 50 can include: Figure 1
[0196] The text front-end module 501 is configured to, in response to receiving the text to be synthesized, obtain the prosody features of the text to be synthesized and the phoneme sequence corresponding to the text to be synthesized by using the front-end network;
[0197] The acoustic model module 502 is configured to obtain the acoustic features corresponding to the text to be synthesized according to the prosody features and the phoneme sequence by using the acoustic model network;
[0198] The vocoder module 503 is configured to obtain the audio waveform of the synthesized speech corresponding to the text to be synthesized according to the acoustic features by using the vocoder network.
[0199] The front-end network is a neural network for generating prosody features and phoneme sequences, the acoustic model network is a neural network for generating acoustic features, and the vocoder network is a neural network for generating audio waveforms. The front-end network, the acoustic model network, and the vocoder network are serially connected to form a speech synthesis network. The front-end network, the acoustic model network, and the vocoder network are obtained by performing end-to-end training on the speech synthesis network in advance.
[0200] In the speech synthesis device provided by the embodiments of the present disclosure, the front-end network, the acoustic model network, and the vocoder network are serially connected to form an end-to-end trained and end-to-end inferred speech synthesis network. When the text is provided to the speech synthesis network, the audio waveform of the synthesized speech corresponding to the text can be directly obtained without the need for other processing by the user or the engineer, thereby reducing the waste of manpower and resources, and being more friendly and efficient for the user and the engineer.
[0201] Meanwhile, the prosody features of the text to be synthesized are obtained through the front-end network, and the prosody features of the text to be synthesized are learned by the acoustic model network, so that the prosody of the synthesized speech is enhanced, the mechanical feeling of the synthesized speech is reduced, and the quality of the synthesized speech is improved, thereby further improving the user experience.
[0202] It can be understood that the above-mentioned modules of the speech synthesis device in the embodiments of the present disclosure have the functions of realizing Figure 1 The embodiments shown illustrate the functions of corresponding steps in the speech synthesis method. These functions can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functions. These modules can be software and / or hardware, and each module can be implemented individually or multiple modules can be integrated. For a detailed description of the functions of each module in the aforementioned speech synthesis device, please refer to [link to relevant documentation]. Figure 4 The corresponding descriptions of the speech synthesis methods in the embodiments shown are not repeated here.
[0203] Based on and Figure 6 The method shown follows the same principle. Figure 6 A schematic diagram of the structure of a network training device provided in an embodiment of this disclosure is shown, such as... Figure 4 As shown, the network training device 60 is used to train a speech synthesis network. The speech synthesis network includes a front-end network, an acoustic model network, and a vocoder network connected in series. The front-end network is a neural network that generates prosodic features and phoneme sequences, the acoustic model network is a neural network that generates acoustic features, and the vocoder network is a neural network that generates audio waveforms. The network training device 60 may include:
[0204] The front-end training module 601 is used to input the text to be trained into the front-end network and obtain the prosodic features and phoneme sequences output by the front-end network.
[0205] The acoustic training module 602 is used to input the prosody tags and phoneme tags corresponding to the text to be trained into the acoustic model network to obtain the acoustic features output by the acoustic model network; and to input the acoustic features output by the acoustic model network into the vocoder network to obtain the audio waveform output by the vocoder network.
[0206] The first loss module 603 is used to construct a first loss function based on the prosodic features output by the front-end network and the prosodic labels corresponding to the text to be trained, as well as the phoneme sequence output by the front-end network and the phoneme labels corresponding to the text to be trained.
[0207] The second loss module 604 is used to construct a second loss function based on the audio waveform output by the vocoder network and the audio waveform of the synthesized speech corresponding to the text to be trained.
[0208] The backpropagation module 605 is used to adjust the network parameters of the speech synthesis network based on the first loss function and the second loss function.
[0209] In the network training apparatus provided in the embodiments of the present disclosure, the front-end network, the acoustic model network and the vocoder network are serially composed into the speech synthesis network, end-to-end training of the speech synthesis network is realized, that is, only the text to be trained and the corresponding label are needed in the training process of the speech synthesis network, and no other processing is needed, so that the training of the speech synthesis network is more convenient, and resource waste caused by information transmission in different processing processes is reduced.
[0210] Meanwhile, the prosody features of the text to be synthesized are obtained through the front-end network, and the acoustic model network learns the prosody features of the text to be synthesized, so that the prosody of the synthesized speech is enhanced, the mechanical feeling of the speech generated by the trained model is reduced, and the quality of the generated speech is improved.
[0211] It can be understood that the above modules of the network training apparatus in the embodiments of the present disclosure have the functions of the corresponding steps of the network training method in the embodiments shown in the above. Figure 4 The functions can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. The above modules can be software and / or hardware, and the above modules can be implemented separately or integrated. For the function description of the modules of the network training apparatus, reference can be made to the corresponding description of the network training method in the embodiments shown in the above, which will not be repeated here. Figure 7
[0212] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations, and do not violate public order and good customs.
[0213] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0214] The electronic device includes at least one processor, and a memory connected with the at least one processor in communication; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the speech synthesis method and the network training method provided in the embodiments of the present disclosure.
[0215] Compared with the prior art, the front-end network, the acoustic model network and the vocoder network are serially composed into an end-to-end trained and end-to-end inferred speech synthesis network, the text is provided to the speech synthesis network, and the audio waveform of the synthesized speech corresponding to the text can be directly obtained without other processing by the user or the engineer, so that the waste of manpower and material resources is reduced, and the user and the engineer are more friendly and efficient.
[0216] Meanwhile, the prosody features of the text to be synthesized are obtained through the front-end network, and the prosody features of the text to be synthesized are learned by the acoustic model network, so that the prosody of the synthesized speech is enhanced, the mechanical feeling of the synthesized speech is reduced, the quality of the obtained synthesized speech is improved, and the user experience is further improved.
[0217] The readable storage medium is a non-transitory computer readable storage medium storing computer instructions, where the computer instructions are used to make a computer execute the speech synthesis method and the network training method provided in the embodiments of the present disclosure.
[0218] Compared with the prior art, the front-end network, the acoustic model network, and the vocoder network are serially composed into an end-to-end training and end-to-end inference speech synthesis network, and the text is provided to the speech synthesis network, so that the audio waveform of the synthesized speech corresponding to the text can be directly obtained without other processing by the user or the engineer, the waste of manpower and material resources is reduced, and the user and the engineer are more friendly and efficient.
[0219] Meanwhile, the prosody features of the text to be synthesized are obtained through the front-end network, and the prosody features of the text to be synthesized are learned by the acoustic model network, so that the prosody of the synthesized speech is enhanced, the mechanical feeling of the synthesized speech is reduced, the quality of the obtained synthesized speech is improved, and the user experience is further improved.
[0220] The computer program product includes a computer program, and the computer program realizes the speech synthesis method and the network training method provided in the embodiments of the present disclosure when executed by a processor.
[0221] Compared with the prior art, the front-end network, the acoustic model network, and the vocoder network are serially composed into an end-to-end training and end-to-end inference speech synthesis network, and the text is provided to the speech synthesis network, so that the audio waveform of the synthesized speech corresponding to the text can be directly obtained without other processing by the user or the engineer, the waste of manpower and material resources is reduced, and the user and the engineer are more friendly and efficient.
[0222] Meanwhile, the prosody features of the text to be synthesized are obtained through the front-end network, and the prosody features of the text to be synthesized are learned by the acoustic model network, so that the prosody of the synthesized speech is enhanced, the mechanical feeling of the synthesized speech is reduced, the quality of the obtained synthesized speech is improved, and the user experience is further improved.
[0223] Figure 7A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0224] like As shown, the electronic device 700 includes a computing unit 710, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 720 or a computer program loaded from a storage unit 780 into a random access memory (RAM) 730. The RAM 730 may also store various programs and data required for the operation of the device 700. The computing unit 710, ROM 720, and RAM 730 are interconnected via a bus 740. An input / output (I / O) interface 750 is also connected to the bus 740.
[0225] Multiple components in device 700 are connected to I / O interface 750, including: input unit 760, such as keyboard, mouse, etc.; output unit 770, such as various types of monitors, speakers, etc.; storage unit 780, such as disk, optical disk, etc.; and communication unit 790, such as network card, modem, wireless transceiver, etc. Communication unit 790 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0226] The computing unit 710 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 710 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The computing unit 710 performs the speech synthesis method and the network training method provided in the embodiments of the present disclosure. For example, in some embodiments, the speech synthesis method and the network training method provided in the embodiments of the present disclosure can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 780. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 720 and / or the communication unit 790. When the computer program is loaded onto the RAM 730 and executed by the computing unit 710, one or more steps of the speech synthesis method and the network training method provided in the embodiments of the present disclosure can be performed. Alternatively, in other embodiments, the computing unit 710 can be configured to perform the speech synthesis method and the network training method provided in the embodiments of the present disclosure by any other appropriate means, such as by means of firmware.
[0227] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0228] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0229] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0230] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0231] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0232] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0233] It should be understood that the various forms of flow shown above can be re-ordered, added to, or have steps deleted, using the steps described above. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, which are not limited herein.
[0234] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other phonemes. Any modifications, equivalent replacements, and improvements within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A speech synthesis method, comprising: In response to receiving the text to be synthesized, the text to be synthesized is subjected to text regularization processing to obtain the regularized text corresponding to the text to be synthesized; Based on preset rules, the regularized text is divided into at least one text statement, and the text statement is divided into multiple text words. Multiple text words corresponding to the text to be synthesized are input into the feature extraction subnetwork of the front-end network to obtain semantic features corresponding to the text to be synthesized; the semantic features are input into the prosodic prediction subnetwork of the front-end network to obtain prosodic features; the prosodic features include the prosodic hierarchical sequence corresponding to the text to be synthesized; the semantic features are input into the phoneme conversion subnetwork of the front-end network to obtain phoneme sequences, and the phoneme sequences are subjected to tone shifting based on preset rules; An acoustic model network is used to obtain the acoustic features corresponding to the text to be synthesized based on the prosodic features and the phoneme sequence; wherein, the acoustic model network uses adaptive parameters to normalize the features output by the neural network layer of the acoustic model network; A vocoder network is used to obtain the audio waveform of the synthesized speech corresponding to the text to be synthesized based on the acoustic features; The front-end network is a neural network that generates prosodic features and phoneme sequences, the acoustic model network is a neural network that generates acoustic features, and the vocoder network is a neural network that generates audio waveforms. The front-end network, the acoustic model network, and the vocoder network are sequentially combined to form a speech synthesis network. The front-end network, the acoustic model network, and the vocoder network are obtained by pre-training the speech synthesis network end-to-end.
2. The method according to claim 1, wherein, The front-end network includes a feature extraction subnetwork, a prosody prediction subnetwork, and a word-sound conversion subnetwork.
3. The method according to claim 1, wherein, The step of inputting the semantic features into the phoneme conversion subnetwork to obtain the phoneme sequence includes: The semantic features are input into the phonetic conversion subnetwork to determine the polyphonic characters in the text to be synthesized, and to determine the phonemes corresponding to the polyphonic characters.
4. The method according to claim 1, wherein, The network parameters of the acoustic model network include adaptive parameters; The method of using an acoustic model network to obtain the acoustic features corresponding to the text to be synthesized based on the prosodic features and the phoneme sequence includes: The prosodic features and the phoneme sequence are input into the acoustic model network to obtain the features output by the neural network layer of the acoustic model network; The features output by the neural network layer of the acoustic model network are normalized using adaptive parameters, and the acoustic features of the text to be synthesized are determined based on the normalized features.
5. A network training method for training a speech synthesis network, the speech synthesis network comprising a front-end network, an acoustic model network, and a vocoder network connected in series, wherein the front-end network is a neural network for generating prosodic features and phoneme sequences, the acoustic model network is a neural network for generating acoustic features, and the vocoder network is a neural network for generating audio waveforms, the method comprising: Perform text regularization on the text to be trained to obtain the regularized text corresponding to the text to be trained; Based on preset rules, the regularized text is divided into at least one text statement, and the text statement is divided into multiple text words. Multiple text words corresponding to the text to be trained are input into the feature extraction subnetwork of the front-end network to obtain semantic features corresponding to the text to be trained; the semantic features are input into the prosodic prediction subnetwork of the front-end network to obtain prosodic features; the prosodic features include the prosodic hierarchical sequence corresponding to the text to be trained; the semantic features are input into the phoneme conversion subnetwork of the front-end network to obtain phoneme sequences, and the phoneme sequences are subjected to tone sandhi processing based on preset rules; The prosodic labels and phoneme labels corresponding to the text to be trained are input into the acoustic model network to obtain the acoustic features output by the acoustic model network; wherein, the acoustic model network uses adaptive parameters to normalize the features output by the neural network layer of the acoustic model network; the acoustic features output by the acoustic model network are input into the vocoder network to obtain the audio waveform output by the vocoder network; A first loss function is constructed based on the prosodic features output by the front-end network and the prosodic labels corresponding to the text to be trained, as well as the phoneme sequence output by the front-end network and the phoneme labels corresponding to the text to be trained. A second loss function is constructed based on the audio waveform output by the vocoder network and the audio waveform of the synthesized speech corresponding to the text to be trained; The network parameters of the speech synthesis network are adjusted based on the first loss function and the second loss function.
6. The method according to claim 5, wherein, The step of inputting the prosodic labels and phoneme labels corresponding to the text to be trained into the acoustic model network and obtaining the acoustic features output by the acoustic model network includes: The prosodic labels of the text to be trained are determined using a speech forced aligner. The prosodic label, phoneme label, and duration information corresponding to the prosodic label of the text to be trained are input into the acoustic model network to obtain the acoustic features output by the acoustic model network.
7. The method according to claim 5, wherein, The step of inputting the prosodic labels and phoneme labels corresponding to the text to be trained into the acoustic model network and obtaining the acoustic features output by the acoustic model network includes: Input the prosodic labels corresponding to the text to be trained into the acoustic model network to obtain training prosodic features; Input the phoneme labels corresponding to the text to be trained into the acoustic model network to obtain training phoneme features; The training prosodic features and the training phoneme features are fused together, and the acoustic features output by the acoustic model network are obtained based on the fused features.
8. The method according to claim 5, wherein, The step of inputting the prosodic labels and phoneme labels corresponding to the text to be trained into the acoustic model network and obtaining the acoustic features output by the acoustic model network includes: The prosodic labels and phoneme labels corresponding to the text to be trained are input into the acoustic model network to obtain training fusion features; Based on the training fusion features, the acoustic features output by the acoustic model network are obtained.
9. The method according to claim 5, wherein, The network parameters of the acoustic model network include adaptive parameters; The step of inputting the prosodic labels and phoneme labels corresponding to the text to be trained into the acoustic model network and obtaining the acoustic features output by the acoustic model network includes: The prosodic labels and phoneme labels corresponding to the text to be trained are input into the acoustic model network to obtain the features output by the neural network layer of the acoustic model network; The features output by the neural network layer of the acoustic model network are normalized using adaptive parameters, and the acoustic features output by the acoustic model network are obtained based on the normalized features.
10. The method according to claim 5, wherein, The step of adjusting the network parameters of the speech synthesis network based on the first loss function and the second loss function includes: Based on the first loss function, the network parameters of the front-end network are adjusted; Based on the second loss function, the network parameters of the acoustic model network and the vocoder network are adjusted.
11. A speech synthesis device, comprising: A text front-end module is used to respond to received text to be synthesized by performing text regularization on the text to be synthesized to obtain regularized text corresponding to the text to be synthesized; based on preset rules, dividing the regularized text into at least one text sentence, and dividing the text sentence into multiple text words; inputting the multiple text words corresponding to the text to be synthesized into the feature extraction subnetwork of the front-end network to obtain semantic features corresponding to the text to be synthesized; inputting the semantic features into the prosody prediction subnetwork of the front-end network to obtain prosodic features; the prosodic features include the prosodic hierarchical sequence corresponding to the text to be synthesized; inputting the semantic features into the phoneme conversion subnetwork of the front-end network to obtain a phoneme sequence, and performing tone shifting on the phoneme sequence based on preset rules; An acoustic model module is used to obtain the acoustic features corresponding to the text to be synthesized based on the prosodic features and the phoneme sequence using an acoustic model network; wherein, the acoustic model network uses adaptive parameters to normalize the features output by the neural network layer of the acoustic model network; A vocoder module is used to obtain the audio waveform of the synthesized speech corresponding to the text to be synthesized based on the acoustic features using a vocoder network; The front-end network is a neural network that generates prosodic features and phoneme sequences, the acoustic model network is a neural network that generates acoustic features, and the vocoder network is a neural network that generates audio waveforms. The front-end network, the acoustic model network, and the vocoder network are sequentially combined to form a speech synthesis network. The front-end network, the acoustic model network, and the vocoder network are obtained by pre-training the speech synthesis network end-to-end.
12. A network training apparatus for training a speech synthesis network, the speech synthesis network comprising a front-end network, an acoustic model network, and a vocoder network connected in series, wherein the front-end network is a neural network for generating prosodic features and phoneme sequences, the acoustic model network is a neural network for generating acoustic features, and the vocoder network is a neural network for generating audio waveforms, the apparatus comprising: The front-end training module is used to perform text regularization processing on the text to be trained and obtain the regularized text corresponding to the text to be trained. Based on preset rules, the regularized text is divided into at least one text statement, and the text statement is divided into multiple text words; the multiple text words corresponding to the text to be trained are input into the feature extraction subnetwork of the front-end network to obtain the semantic features corresponding to the text to be trained; the semantic features are input into the prosodic prediction subnetwork of the front-end network to obtain prosodic features; the prosodic features include the prosodic hierarchical sequence corresponding to the text to be trained; the semantic features are input into the phoneme conversion subnetwork of the front-end network to obtain the phoneme sequence, and the phoneme sequence is subjected to tone sandhi processing based on preset rules; An acoustic training module is used to input the prosodic tags and phoneme tags corresponding to the text to be trained into the acoustic model network to obtain the acoustic features output by the acoustic model network; wherein, the acoustic model network uses adaptive parameters to normalize the features output by the neural network layer of the acoustic model network; and inputs the acoustic features output by the acoustic model network into the vocoder network to obtain the audio waveform output by the vocoder network; The first loss module is used to construct a first loss function based on the prosodic features output by the front-end network and the prosodic labels corresponding to the text to be trained, as well as the phoneme sequence output by the front-end network and the phoneme labels corresponding to the text to be trained. The second loss module is used to construct a second loss function based on the audio waveform output by the vocoder network and the audio waveform of the synthesized speech corresponding to the text to be trained. The backpropagation module is used to adjust the network parameters of the speech synthesis network based on the first loss function and the second loss function.
13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the speech synthesis method of any one of claims 1-4 and the network training method of any one of claims 5-10.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the speech synthesis method according to any one of claims 1-4 and the network training method according to any one of claims 5-10.
15. A computer program product comprising a computer program that, when executed by a processor, implements the speech synthesis method according to any one of claims 1-4 and the network training method according to any one of claims 5-10.
Citation Information
Patent Citations
Speech synthesis device and method, electronic equipment and storage medium
CN113096636A
Text processing method and device, text processing model training method and device, equipment and storage medium
CN114550692A