Speech synthesis method and related devices, equipment and storage media
By using a combination of feature prediction model and acoustic model in speech synthesis technology to predict and generate pronunciation features and waveforms, the stability and naturalness problems in existing speech synthesis technologies are solved, and higher quality speech synthesis is achieved.
Patent Information
- Application Number
- CN202510224088.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The existing speech synthesis technology has stability problems such as multiple readings, missed readings, and misreading, as well as natural problems such as unnatural and rereading errors.
A speech synthesis method is adopted to predict the pronunciation characteristics of the character sequence to be synthesized based on the feature prediction model, combine the acoustic model to predict the pronunciation characteristics and the hidden layer characteristics of the sequence, and finally the waveform recovery is performed by the vocoder to generate the synthetic speech.
This method improves the stability and naturalness of speech synthesis by referring to the semanticity and rhythmicity of the character sequence to be synthesized, and avoids the problem of acoustic characteristics deviation.
Smart Images

Figure CN119724148B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of speech processing, and in particular, to a speech synthesis method, related devices, equipment, and storage media. Background Art
[0002] Speech synthesis technology, also known as Text to Speech (TTS) technology, is an important branch in the field of artificial intelligence. It involves converting text information into natural and fluent speech output, enabling a computer to "speak", and being applied to, including but not limited to: assisting visually impaired people in reading, voice navigation systems, automatic customer service systems, audiobook production, etc.
[0003] However, the speech synthesized by existing speech synthesis technologies still has stability problems such as redundant reading, missed reading, misreading, etc., and naturalness problems such as unnaturalness and incorrect stress. In view of this, how to improve the stability and naturalness of speech synthesis has become an urgent problem to be solved. Summary of the Invention
[0004] The main technical problem to be solved by this application is to provide a speech synthesis method, related devices, equipment, and storage media, which can improve the stability and naturalness of speech synthesis.
[0005] To solve the above technical problem, in the first aspect of this application, a speech synthesis method is provided, including: predicting the pronunciation features of a character sequence to be synthesized based on a feature prediction model; wherein, the character sequence to be synthesized is a text sequence or a phoneme sequence, and the pronunciation features at least include feature information related to the pronunciation content and pronunciation prosody of the character sequence to be synthesized; predicting acoustic features based on an acoustic model for the pronunciation features and the sequence hidden layer features of the character sequence to be synthesized; wherein, the sequence hidden layer features are the hidden layer features obtained during the process of the feature prediction model processing the character sequence to be synthesized; and restoring the waveform of the acoustic features based on a vocoder to obtain the synthesized speech.
[0006] To solve the above technical problem, in the second aspect of this application, a speech synthesis device is provided, including: a pronunciation prediction module, an acoustic prediction module, and a waveform restoration module. The pronunciation prediction module is used to predict the pronunciation features of a character sequence to be synthesized based on a feature prediction model; wherein, the character sequence to be synthesized is a text sequence or a phoneme sequence, and the pronunciation features at least include feature information related to the pronunciation content and pronunciation prosody of the character sequence to be synthesized; the acoustic prediction module is used to predict acoustic features based on an acoustic model for the pronunciation features and the sequence hidden layer features of the character sequence to be synthesized; wherein, the sequence hidden layer features are the hidden layer features obtained during the process of the feature prediction model processing the character sequence to be synthesized; and the waveform restoration module is used to restore the waveform of the acoustic features based on a vocoder to obtain the synthesized speech.
[0007] To solve the above technical problems, a third aspect of the present application provides an electronic device, which at least includes a memory and a processor coupled to each other. At least program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the speech synthesis method in the first aspect above.
[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are used to implement the speech synthesis method in the first aspect above.
[0009] In the above solution, the pronunciation features of the character sequence to be synthesized are predicted based on the feature prediction model, and the character sequence to be synthesized is a text sequence or a phoneme sequence. The pronunciation features at least include the feature information related to the pronunciation content and pronunciation prosody of the character sequence to be synthesized, and the acoustic features are predicted based on the acoustic model for the pronunciation features and the sequence hidden layer features of the character sequence to be synthesized. The sequence hidden layer features are the hidden layer features obtained during the process of the feature prediction model processing the character sequence to be synthesized. Then, the acoustic features are restored to waveforms based on the vocoder to obtain the synthesized speech. On the one hand, the pronunciation features predicted based on the character sequence to be synthesized at least include the feature information related to the pronunciation content and pronunciation prosody of the character sequence to be synthesized. Therefore, using these pronunciation features for subsequent acoustic feature prediction and speech waveform restoration can ensure that the feature information related to the pronunciation content, that is, the semanticity of the character sequence to be synthesized, is referred to during the speech synthesis process, which helps to improve the stability of speech synthesis as much as possible. It can also ensure that the feature information related to the pronunciation prosody, that is, the prosody of the character sequence to be synthesized, is referred to during the speech synthesis process, which helps to improve the naturalness of speech synthesis as much as possible. On the other hand, the acoustic model not only makes predictions based on the pronunciation features but also further supplements with the sequence hidden layer features. Therefore, it can further perform constraints by combining the sequence hidden layer features during the acoustic feature prediction process, which helps to avoid the predicted acoustic features deviating from the character sequence to be synthesized as much as possible. Therefore, the stability and naturalness of speech synthesis can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a schematic flowchart of an embodiment of the speech synthesis method of the present application;
[0011] Figure 2a is a schematic diagram of an embodiment of the feature extraction model of the present application;
[0012] Figure 2b is a schematic diagram of an embodiment of the speaker enhancement model of the present application;
[0013] Figure 2c is a schematic diagram of an embodiment of the feature prediction model of the present application;
[0014] Figure 2d is a schematic diagram of an embodiment of the acoustic model of the present application;
[0015] Figure 2e is a schematic diagram of an embodiment of the joint training of the present application;
[0016] Figure 2f is a process schematic diagram of an embodiment of the speech synthesis method of the present application;
[0017] Figure 3 is a framework schematic diagram of an embodiment of the speech synthesis device of the present application;
[0018] Figure 4 is a framework schematic diagram of an embodiment of the electronic device of the present application;
[0019] Figure 5 is a framework schematic diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners
[0020] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.
[0021] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.
[0022] The terms "system" and "network" are often used interchangeably in this article. The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the fragment " / " in this article generally represents an "or" relationship between the associated objects before and after. In addition, "multiple" in this article means two or more than two.
[0023] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the speech synthesis method of the present application. Specifically, it may include the following steps:
[0024] Step S11: Predict the pronunciation features of the character sequence to be synthesized based on the feature prediction model.
[0025] In the embodiments of the present disclosure, the character sequence to be synthesized is a text sequence or a phoneme sequence, and the pronunciation features at least include the feature information related to the pronunciation content and pronunciation rhythm of the character sequence to be synthesized. It should be noted that the text sequence is a sequence composed of text characters as basic units, and the phoneme sequence is a sequence composed of phoneme characters as basic units.
[0026] In an implementation scenario, the feature prediction model may include, but is not limited to, network architectures such as Transformer, and specifically may adopt network structures including, but not limited to, Llama. Here, the network architectures included in the feature prediction model and the network structures adopted are not limited.
[0027] In an implementation scenario, the feature prediction model can be trained based on the first character sequence and the first pronunciation feature of the first speech to which it belongs, with the first pronunciation feature as the optimization target. The first pronunciation feature of the first speech can be extracted by a feature extraction model. The feature extraction model is trained with the second speech as the model input, with the first attribute annotated by the second speech as the optimization target and the second attribute annotated by the second speech as the degradation target. The first attribute may include at least one of the second character sequence of the second speech and the true fundamental frequency. The second attribute may at least include the true speaker of the second speech. That is to say, the training objective of the feature extraction model is to extract pronunciation features that are content-related, prosody-related but speaker-independent from the input speech as much as possible. Thus, the first pronunciation feature extracted by the feature extraction model from the first speech can also ensure that it is as much as possible content-related, prosody-related but speaker-independent from the first speech. Furthermore, training the feature prediction model with the first pronunciation feature as the optimization target can ensure that the feature prediction model predicts pronunciation features that are content-related, prosody-related but speaker-independent as much as possible.
[0028] In a specific implementation scenario, please refer to Figure 2a , Figure 2a which is a schematic diagram of an embodiment of the feature extraction model of this application. As Figure 2a shown, before using the feature extraction model to extract features from the first speech, the feature extraction model can be trained first. Specifically, the second speech hidden layer feature of the second speech can be extracted first, and encoded based on the second speech hidden layer feature to obtain the second speech encoded feature. Then, based on the second speech encoded feature, predictions can be made to obtain the prediction information of the second speech about the first attribute and the prediction information about the second attribute. On this basis, at least based on the first loss measured by the difference between the annotation information and the prediction information of the second speech about the first attribute, and the second loss measured by the difference between the annotation information and the prediction information of the second speech about the second attribute, the network parameters of the feature extraction model can be adjusted, and the gradient of the second loss is flipped during the parameter adjustment process. It should be noted that the feature extraction model may specifically include Figure 2aIt includes three parts: open-source models such as Hubert, WavLM, and AE for extracting speech hidden-layer features, an encoder, and vector quantization (VQ). In the above method, at least based on the first loss measured by the difference between the annotation information and the prediction information of the second speech regarding the first attribute, and the second loss measured by the difference between the annotation information and the prediction information of the second speech regarding the second attribute, the network parameters of the feature extraction model are adjusted, and the gradient of the second loss is flipped during the parameter adjustment process, which can force the feature extraction model to extract as much feature information as possible that is independent of the speaker but related to the pronunciation content and pronunciation rhythm to obtain pronunciation features.
[0029] As a possible example, open-source models such as Hubert, WavLM, and AE can be used to extract features from the second speech to obtain the second speech hidden-layer features of the second speech. For the sake of convenience in description, the second speech hidden-layer features can be denoted as .
[0030] As a possible example, the second speech hidden-layer features can be encoded by an encoder (Encoder) to obtain the second speech encoded features. For the sake of convenience in description, the second speech encoded features can be denoted as .
[0031] As a possible example, when the first attribute annotated in the second speech includes the second character sequence of the second speech, the second speech encoded features can pass through a text enhancement model to obtain the prediction information of the second speech regarding the first attribute (i.e., the predicted character sequence of the second speech). For example, when the second character sequence is the text sequence of the second speech, the prediction information of the second speech regarding the first attribute can include the text sequence predicted by the text enhancement model from the second speech encoded features ; or, when the second character sequence is the phoneme sequence of the second speech, the prediction information of the second speech regarding the first attribute can include the phoneme sequence predicted by the text enhancement model from the second speech encoded features . It should be noted that the text enhancement model can include, but is not limited to, convolutional layers, long short-term memory networks, etc. The specific structure of the text enhancement model is not limited here.
[0032] As a possible example, when the first attribute annotated in the second speech includes the true fundamental frequency of the second speech, the second speech encoded features can pass through a fundamental frequency enhancement model to obtain the prediction information of the second speech regarding the first attribute (i.e., the predicted fundamental frequency of the second speech). It should be noted that the fundamental frequency enhancement model can include, but is not limited to, convolutional layers, long short-term memory networks, etc. The specific structure of the fundamental frequency enhancement model is not limited here.
[0033] As a possible example, the second speech coding feature The speaker enhancement model can also be used to obtain prediction information about the second attribute of the second speech (i.e., the predicted speaker of the second speech). It should be noted that the speaker enhancement model may include but is not limited to: a convolutional layer, a long short-term memory network, etc., and the specific structure of the speaker enhancement model is not limited here.
[0034] As a possible example, see Figure 2a , after obtaining the second speech coding feature Afterwards, the second speech coding feature On the other hand, vector quantization can be performed to obtain a second pronunciation feature, and decoding can be performed based on the second pronunciation feature to obtain a predicted acoustic feature of the second speech, so that a third loss can be obtained based on the second pronunciation feature and the vector quantization codebook, and a fourth loss can be obtained based on the real acoustic feature and the predicted acoustic feature of the second speech, and then the network parameters of the feature extraction model can be adjusted based on the first loss, the second loss, the third loss and the fourth loss. It should be noted that in the embodiments of the present disclosure, the "acoustic features" such as acoustic features, predicted acoustic features, and real acoustic features can specifically be any acoustic features such as MFCC and FBank, and the specific type of "acoustic features" is not limited here. For the sake of ease of description, the training loss of the feature extraction model It can be expressed as:
[0035] …… (1)
[0036] In the above formula (1), represents the first loss of the character sequence in the first attribute (i.e., the difference between the second character sequence and the predicted character sequence is measured by a loss function such as cross entropy), represents the first loss about the fundamental frequency in the first attribute (i.e., the difference between the true fundamental frequency and the predicted fundamental frequency is measured by a loss function such as L1), represents the second loss about the second attribute (i.e., the difference between the real speaker and the predicted speaker is measured by a loss function such as cross entropy), represents the fourth loss (e.g., the difference between the real acoustic features and the predicted acoustic features can be measured by a loss function such as L1), represents the third loss. As a possible example, It can be expressed as:
[0037] …… (2)
[0038] In the above formula (2), sg represents the stop-gradient operation. , is a hyperparameter, is the codebook, where n represents the number of codebooks, d represents the dimension of a codebook, and has the same dimension as It should be noted that when the third loss is measured by the above formula (2), when adjusting the network parameters of the feature extraction model, the first term in formula (2) (i.e., ) only acts on the codebook, and the second term (i.e., ) only acts on the Encoder. Please refer to Figure 2b . Figure 2b is a schematic diagram of an embodiment of the speaker enhancement model of the present application. As Figure 2b shown, the speaker enhancement model can be composed of a Gradient Reversal Layer and a speaker prediction network. Thus, when adjusting the network parameters of the feature extraction model based on the foregoing training loss, it can force the feature extraction model to extract as much as possible feature information that is irrelevant to the speaker but related to the pronunciation content and pronunciation rhythm. In addition, during the process of adjusting the parameters based on the training loss, the foregoing open-source models such as Hubert, WavLM, and AE used to extract speech hidden layer features do not participate in the parameter update. After the training converges, the network parameters of the feature extraction model can be kept fixed.
[0039] In a specific implementation scenario, after the feature extraction model training converges, the feature extraction model can be used to extract features from the first speech. Specifically, the first speech hidden layer features of the first speech can be extracted, and based on the first speech hidden layer features, encoding is performed to obtain the first speech encoding features, and then based on the first speech encoding features, vector quantization is performed to obtain the first pronunciation features. As a possible example, please continue to refer to Figure 2a , the first speech can be processed by open-source models such as Hubert, WavLM, and AE in the feature extraction model used to extract speech hidden layer features to extract the first speech hidden layer features. The first hidden layer features can be processed by the encoder in the feature extraction model to extract the first speech encoding features. After the first speech encoding features undergo vector quantization, the first pronunciation features can be obtained, and thus it can be made that the first pronunciation features contain as much as possible feature information related to the pronunciation content and pronunciation rhythm of the first speech and are irrelevant to the speaker of the first speech.
[0040] In a specific implementation scenario, after extracting the first pronunciation feature of the first speech, the feature prediction model can be trained based on the first character sequence of the first speech (such as the text sequence, phoneme sequence, etc. of the first speech) and the first pronunciation feature of the first speech, and the first pronunciation feature is used as the optimization target during the training process, that is, forcing the feature prediction model to predict the first pronunciation feature as accurately as possible according to the first character sequence.
[0041] As a possible example, the first character sequence can be predicted based on the feature prediction model to obtain the predicted pronunciation feature of the first speech, and then the network parameters of the feature prediction model can be adjusted based on the difference between the first pronunciation feature and the predicted pronunciation feature of the first speech.
[0042] As another possible example, different from the aforementioned training method, please refer to Figure 2c , Figure 2c is a schematic diagram of an embodiment of the feature prediction model of the present application. As described above, the character sequence of the first speech can include the true text sequence, true phoneme sequence, etc. of the first speech, and then the true text sequence or true phoneme sequence of the first speech can be selected as the first character sequence with a preset probability. For example, the true text sequence can be selected as the first character sequence with a probability of 50%, and the true phoneme sequence can be selected as the first character sequence with a probability of 50%. Then, the first character sequence is predicted based on the feature prediction model to obtain the predicted pronunciation feature of the first character sequence, and thus, based on the predicted pronunciation feature, a predicted character sequence aligned with the predicted pronunciation feature is obtained. For example, a CTC (Connectionist Temporal Classification) module can be used to predict the predicted pronunciation feature to obtain a predicted character sequence aligned with the predicted pronunciation feature. Furthermore, the network parameters of the feature prediction model can be adjusted based on the difference between the predicted pronunciation feature and the first pronunciation feature of the first speech, and the difference between the labeled character sequence and the predicted character sequence of the first speech. When the first character sequence is the true text sequence, the predicted character sequence is the predicted phoneme sequence, and the labeled character sequence is the true phoneme sequence. When the first character sequence is the true phoneme sequence, the predicted character sequence is the predicted text sequence, and the labeled character sequence is the true text sequence. For ease of description, the training loss of the feature prediction model can be expressed as:
[0043] …… (3)
[0044] In the above formula (3), represents the loss value measured based on a loss function such as cross-entropy for the difference between the predicted pronunciation feature and the first pronunciation feature of the first speech, It represents the loss value measured based on the difference between the labeled character sequence and the predicted character sequence of the first speech using CTC loss. In the above manner, by further adding constraints on the alignment between the text and phonemes based on the feature prediction model, on the one hand, it can assist the feature prediction model to better learn to align the predicted pronunciation features with the input character sequence, and as much as possible ensure that the features predicted by the feature prediction model are related to the pronunciation content and pronunciation rhythm but independent of the speaker. On the other hand, it can assist the feature prediction model to better learn the conversion relationship from phonemes to text and from text to phonemes, and assist the feature prediction model to better learn pronunciation. In addition, as the input of the synthesis system, the text can enable the feature prediction model to better understand the semantic information conveyed in the text, making the synthesis effect more natural and the stress more reasonable.
[0045] It should be noted that through the above process steps, the feature prediction model can be obtained, and then the pronunciation features can be predicted based on the feature prediction model for the character sequence to be synthesized. The pronunciation features include feature information related to the character sequence to be synthesized, pronunciation content, and pronunciation rhythm, but are independent of the speaker.
[0046] Step S12: Based on the acoustic model, predict the sequence hidden layer features of the pronunciation features and the character sequence to be synthesized to obtain acoustic features.
[0047] In the embodiments of the present disclosure, the sequence hidden layer features are the hidden layer features obtained during the process of the feature prediction model processing the character sequence to be synthesized. In addition, the acoustic model may include, but is not limited to, network architectures such as Transformer, Conformer, and convolutional neural networks. The network structure of the acoustic model is not limited here. That is, the pronunciation features and the sequence hidden layer features of the character sequence to be synthesized can be input into the acoustic model to obtain the acoustic features predicted by the acoustic model, such as types like MFCC. The specific type of the acoustic features is not limited here. It should be noted that the acoustic features can be trained based on the third speech, its third pronunciation features, and the third character sequence after the feature extraction model converges. The third pronunciation features can be obtained by the feature extraction model extracting features from the third speech. For details, reference can be made to the relevant description of the aforementioned feature extraction model, which will not be elaborated here. In addition, the third character sequence can be a phoneme sequence, or the third character sequence can also be a text sequence, which is not limited here.
[0048] In an implementation scenario, please refer to Figure 2d , Figure 2d which is a schematic diagram of an embodiment of the acoustic model of this application. As Figure 2dAs shown, as a possible example, the third pronunciation feature of the third speech can be first extracted based on the feature extraction model, and the predicted acoustic feature of the third speech can be obtained by predicting the third pronunciation feature and the third character sequence of the third speech based on the acoustic model. Then, based on the acoustic reconstruction loss measured by the difference between the true acoustic feature and the predicted acoustic feature of the third speech, the network parameters of the acoustic model can be adjusted.
[0049] In another implementation scenario, different from the foregoing implementation manner, as another possible example, the third pronunciation feature of the third speech can be first extracted based on the feature extraction model, and the predicted acoustic feature of the third speech can be obtained by predicting the third pronunciation feature and the third character sequence of the third speech based on the acoustic model. Then, based on the acoustic reconstruction loss measured by the difference between the true acoustic feature and the predicted acoustic feature of the third speech, and the generative adversarial loss of the generative adversarial network, the network parameters of the acoustic model can be adjusted, and the discriminator in the generative adversarial network is at least used to discriminate whether the predicted acoustic feature is a true acoustic feature. For ease of description, the training loss of the acoustic model can be expressed as:
[0050] ……(4)
[0051] In the above formula (4), represents the acoustic reconstruction loss measured by the L1 loss function based on the difference between the true acoustic feature and the predicted acoustic feature of the third speech, represents the generation loss of the generator in the generative adversarial network, represents the discrimination loss of the discriminator in the generative adversarial network. The generator is used to generate acoustic features that are as difficult as possible for the discriminator to distinguish between true and false, while the discriminator is used to accurately distinguish the true and false attributes of the acoustic features. For specific details, refer to the technical details of the generative adversarial network, which will not be elaborated here. In the above manner, based on the acoustic reconstruction loss measured by the difference between the true acoustic feature and the predicted acoustic feature of the third speech, and the generative adversarial loss of the generative adversarial network, the network parameters of the acoustic model are adjusted, which can combine the generative adversarial network to ensure the authenticity of the generated predicted acoustic features as much as possible, and help improve the audio quality of subsequent speech synthesis.
[0052] In another implementation scenario, after the feature prediction model and the acoustic model are trained separately, they can be jointly trained based on the fourth voice data. The fourth voice data can include a fourth voice and a fourth pronunciation feature of the fourth voice, a fourth character sequence, a labeled character sequence and a real acoustic feature. The fourth pronunciation feature can be extracted by the feature extraction model, and when the fourth character sequence is a real phoneme sequence of the fourth voice, the labeled character sequence of the fourth voice is a real text sequence of the fourth voice, and when the fourth character sequence is a real text sequence of the fourth voice, the labeled character sequence of the fourth voice is a real phoneme sequence of the fourth voice. That is to say, after the joint training is completed, the feature prediction model, the acoustic model and the following vocoder can be used to implement speech synthesis through the process steps of the embodiment of the present disclosure. Please refer to Figure 2e , Figure 2e Schematic diagram of an embodiment of joint training of the present application. Figure 2e As shown, the fourth character sequence can be predicted based on the feature prediction model to obtain the predicted pronunciation features of the fourth character sequence, and then prediction can be made based on the predicted pronunciation features of the fourth character sequence to obtain a predicted character sequence aligned with the predicted pronunciation features of the fourth character sequence, and the predicted pronunciation features of the fourth character sequence and the sample sequence hidden features of the fourth character sequence can be predicted based on the acoustic model to obtain the predicted acoustic features of the fourth speech, and the hidden features of the sample sequence are the hidden features obtained in the process of processing the fourth character sequence with the feature prediction model, and when the fourth character sequence is the real text sequence of the fourth speech, the predicted character sequence is the predicted phoneme sequence. When the fourth character sequence is the real phoneme sequence of the fourth voice, the predicted character sequence is the predicted text sequence, and then the network parameters of the feature prediction model and the acoustic model can be adjusted based on the difference between the predicted pronunciation feature and the fourth pronunciation feature, the difference between the predicted character sequence and the annotated character sequence of the fourth voice, the difference between the predicted acoustic feature of the fourth voice and the real acoustic feature, and the generative adversarial loss of the generative adversarial network, and the discriminator in the generative adversarial network is at least used to discriminate whether the predicted acoustic feature is a real acoustic feature, which can further improve the model performance of the feature prediction model and the acoustic model, so as to help improve the audio quality of the subsequent synthesized speech. It should be noted that the detailed details of the above-mentioned joint training can be combined with the training process of the aforementioned feature prediction model and the prediction process of the aforementioned acoustic model. The main difference is that the character sequence can be directly input when the acoustic model is trained alone, while the acoustic model needs to input the hidden features obtained by the character sequence processed by the feature prediction model during the joint training process. Other details can be found in the above-mentioned related descriptions, which will not be repeated here.
[0053] Step S13: Perform waveform restoration on the acoustic features based on the vocoder to obtain synthesized speech.
[0054] Specifically, the vocoder can be a parallel vocoder, an autoregressive vocoder, etc. The specific type of the vocoder is not limited herein. Please refer to Figure 2f , Figure 2f which is a schematic process diagram of an embodiment of the voice synthesis method of the present application. As Figure 2f shown, the character sequence to be synthesized can be predicted to obtain pronunciation features through a feature prediction model. The pronunciation features and the sequence hidden layer features of the character sequence to be synthesized can be predicted through an acoustic model to obtain acoustic features. Finally, the acoustic features can be processed by a vocoder to restore the voice waveform, that is, the synthesized voice.
[0055] In the above solution, the pronunciation features of the character sequence to be synthesized are predicted based on the feature prediction model, and the character sequence to be synthesized is a text sequence or a phoneme sequence. The pronunciation features at least include the feature information related to the pronunciation content and pronunciation prosody of the character sequence to be synthesized. The acoustic features are predicted based on the pronunciation features and the sequence hidden layer features of the character sequence to be synthesized, and the sequence hidden layer features are the hidden layer features obtained during the process of the feature prediction model processing the character sequence to be synthesized. Then, the acoustic features are restored to waveforms based on the vocoder to obtain the synthesized voice. On the one hand, the pronunciation features predicted based on the character sequence to be synthesized at least include the feature information related to the pronunciation content and pronunciation prosody of the character sequence to be synthesized. Therefore, using this pronunciation feature for subsequent acoustic feature prediction and voice waveform restoration can ensure that the feature information related to the pronunciation content, that is, the semanticity of the character sequence to be synthesized, is referred to during the voice synthesis process, which helps to improve the stability of the voice synthesis as much as possible. It can also ensure that the feature information related to the pronunciation prosody, that is, the prosody of the character sequence to be synthesized, is referred to during the voice synthesis process, which helps to improve the naturalness of the voice synthesis as much as possible. On the other hand, the acoustic model not only predicts based on the pronunciation features but also further supplements with the sequence hidden layer features. Therefore, it can further constrain by combining the sequence hidden layer features during the acoustic feature prediction process, which helps to avoid the predicted acoustic features deviating from the character sequence to be synthesized as much as possible. Therefore, the stability and naturalness of the voice synthesis can be improved.
[0056] Please refer to Figure 3 , Figure 3It is a schematic framework diagram of an embodiment of the voice synthesis device of the present application. The voice synthesis device 30 includes: a pronunciation prediction module 31, an acoustic prediction module 32, and a waveform restoration module 33. The pronunciation prediction module 31 is used to predict the pronunciation features of the character sequence to be synthesized based on a feature prediction model. Among them, the character sequence to be synthesized is a text sequence or a phoneme sequence, and the pronunciation features at least include the feature information related to the pronunciation content and pronunciation prosody of the character sequence to be synthesized. The acoustic prediction module 32 is used to predict acoustic features based on an acoustic model for the pronunciation features and the sequence hidden layer features of the character sequence to be synthesized. Among them, the sequence hidden layer features are the hidden layer features obtained during the process of the feature prediction model processing the character sequence to be synthesized. The waveform restoration module 33 is used to restore the waveform of the acoustic features based on a vocoder to obtain the synthesized speech.
[0057] In the above solution, the voice synthesis device 30 predicts the pronunciation features of the character sequence to be synthesized based on a feature prediction model, and the character sequence to be synthesized is a text sequence or a phoneme sequence. The pronunciation features at least include the feature information related to the pronunciation content and pronunciation prosody of the character sequence to be synthesized, and predicts acoustic features based on an acoustic model for the pronunciation features and the sequence hidden layer features of the character sequence to be synthesized. The sequence hidden layer features are the hidden layer features obtained during the process of the feature prediction model processing the character sequence to be synthesized. Then, the waveform of the acoustic features is restored based on a vocoder to obtain the synthesized speech. On the one hand, the pronunciation features predicted based on the character sequence to be synthesized at least include the feature information related to the pronunciation content and pronunciation prosody of the character sequence to be synthesized. Therefore, using this pronunciation feature for subsequent acoustic feature prediction and speech waveform restoration can ensure that the feature information related to the pronunciation content, that is, the semanticity of the character sequence to be synthesized, is referred to during the voice synthesis process, which helps to improve the stability of voice synthesis as much as possible. It can also ensure that the feature information related to the pronunciation prosody, that is, the prosody of the character sequence to be synthesized, is referred to during the voice synthesis process, which helps to improve the naturalness of voice synthesis as much as possible. On the other hand, the acoustic model not only predicts based on the pronunciation features but also further supplements with the sequence hidden layer features. Therefore, it can further perform constraints by combining the sequence hidden layer features during the acoustic feature prediction process, which helps to avoid the predicted acoustic features deviating from the character sequence to be synthesized as much as possible. Therefore, the stability and naturalness of voice synthesis can be improved.
[0058] In some disclosed embodiments, the feature prediction model is trained based on a first character sequence and the first pronunciation features of its affiliated first voice with the first pronunciation features as the optimization target. The first pronunciation features of the first voice are extracted by a feature extraction model. The feature extraction model is trained with a second voice as the model input and with the first attribute annotated by the second voice as the optimization target and the second attribute annotated by the second voice as the degradation target. The first attribute includes at least one of the second character sequence and the true fundamental frequency of the second voice, and the second attribute at least includes the true speaker of the second voice.
[0059] In some disclosed embodiments, the speech synthesis device 30 includes a first hidden layer extraction module for extracting the first speech hidden layer features of a first speech; the speech synthesis device 30 includes a first speech encoding module for encoding based on the first speech hidden layer features to obtain first speech encoding features; the speech synthesis device 30 includes a first vector quantization module for performing vector quantization based on the first speech encoding features to obtain first pronunciation features.
[0060] In some disclosed embodiments, the speech synthesis device 30 includes a second hidden layer extraction module for extracting the second speech hidden layer features of a second speech; the speech synthesis device 30 includes a second speech encoding module for encoding based on the second speech hidden layer features to obtain second speech encoding features; the speech synthesis device 30 includes a speech attribute prediction module for predicting based on the second speech encoding features to obtain prediction information about a first attribute and prediction information about a second attribute of the second speech; the speech synthesis device 30 includes a first model optimization module for adjusting the network parameters of the feature extraction model based at least on a first loss measured by the difference between the labeled information and the prediction information about the first attribute of the second speech, and a second loss measured by the difference between the labeled information and the prediction information about the second attribute of the second speech; wherein, the second loss has its gradient flipped during the parameter adjustment process.
[0061] In some disclosed embodiments, the speech synthesis device 30 includes a first acoustic prediction module for performing vector quantization based on the second speech encoding features to obtain second pronunciation features, and decoding based on the second pronunciation features to obtain predicted acoustic features of the second speech; the speech synthesis device 30 includes a loss function measurement module for obtaining a third loss based on the second pronunciation features and the vector quantization codebook, and obtaining a fourth loss based on the true acoustic features and the predicted acoustic features of the second speech; the first model optimization module is specifically configured to adjust the network parameters of the feature extraction model based on the first loss, the second loss, the third loss, and the fourth loss.
[0062] In some disclosed embodiments, the speech synthesis device 30 includes a character sequence selection module for selecting, with a preset probability, the true text sequence or the true phoneme sequence of the first speech as the first character sequence; the speech synthesis device 30 includes a first pronunciation prediction module for predicting the first character sequence based on a feature prediction model to obtain the predicted pronunciation features of the first character sequence; the speech synthesis device 30 includes a first sequence prediction module for predicting based on the predicted pronunciation features to obtain a predicted character sequence aligned with the predicted pronunciation features; the speech synthesis device 30 includes a second model optimization module for adjusting the network parameters of the feature prediction model based on the difference between the predicted pronunciation features and the first pronunciation features of the first speech, and the difference between the labeled character sequence and the predicted character sequence of the first speech; wherein, when the first character sequence is the true text sequence, the predicted character sequence is the predicted phoneme sequence, and the labeled character sequence is the true phoneme sequence, and when the first character sequence is the true phoneme sequence, the predicted character sequence is the predicted text sequence, and the labeled character sequence is the true text sequence.
[0063] In some disclosed embodiments, the acoustic model is trained after the feature extraction model converges. The speech synthesis device 30 includes a pronunciation feature extraction module for extracting the third pronunciation features of the third speech based on the feature extraction model; the speech synthesis device 30 includes a second acoustic prediction module for predicting the predicted acoustic features of the third speech based on the acoustic model and the third character sequence of the third speech; wherein, the third character sequence is the true text sequence or the true phoneme sequence of the third speech; the speech synthesis device 30 includes a third model optimization module for adjusting the network parameters of the acoustic model based on the acoustic reconstruction loss measured by the difference between the true acoustic features and the predicted acoustic features of the third speech, and the generative adversarial loss of the generative adversarial network; wherein, the discriminator in the generative adversarial network is at least used to discriminate whether the predicted acoustic features are true acoustic features.
[0064] In some disclosed embodiments, the feature prediction model and the acoustic model are separately trained and then jointly trained based on the fourth speech data. The fourth speech data includes the fourth speech and the fourth pronunciation features, the fourth character sequence, the labeled character sequence, and the true acoustic features of the fourth speech. The fourth pronunciation features are extracted by the feature extraction model, and when the fourth character sequence is the true phoneme sequence of the fourth speech, the labeled character sequence of the fourth speech is the true text sequence of the fourth speech, and when the fourth character sequence is the true text sequence of the fourth speech, the labeled character sequence of the fourth speech is the true phoneme sequence of the fourth speech.
[0065] In some disclosed embodiments, the speech synthesis device 30 includes a second pronunciation prediction module configured to predict a fourth character sequence based on a feature prediction model to obtain predicted pronunciation features of the fourth character sequence; the speech synthesis device 30 includes a second sequence prediction module configured to predict based on the predicted pronunciation features of the fourth character sequence to obtain a predicted character sequence aligned with the predicted pronunciation features of the fourth character sequence; the speech synthesis device 30 includes a third acoustic prediction module configured to predict based on an acoustic model the predicted pronunciation features of the fourth character sequence and the hidden layer features of the sample sequence of the fourth character sequence to obtain predicted acoustic features of the fourth speech; wherein, the hidden layer features of the sample sequence are the hidden layer features obtained during the process of the feature prediction model processing the fourth character sequence, and when the fourth character sequence is the true text sequence of the fourth speech, the predicted character sequence is the predicted phoneme sequence, and when the fourth character sequence is the true phoneme sequence of the fourth speech, the predicted character sequence is the predicted text sequence; the speech synthesis device 30 includes a fourth model optimization module configured to adjust the network parameters of the feature prediction model and the acoustic model based on the difference between the predicted pronunciation features and the fourth pronunciation features, the difference between the predicted character sequence and the labeled character sequence of the fourth speech, the difference between the predicted acoustic features and the true acoustic features of the fourth speech, and the generative adversarial loss of the generative adversarial network; wherein, the discriminator in the generative adversarial network is at least configured to discriminate whether the predicted acoustic features are true acoustic features.
[0066] Please refer to Figure 4 , Figure 4 is a schematic framework diagram of an embodiment of the electronic device of the present application. The electronic device 40 at least includes a memory 41 and a processor 42 that are coupled to each other. The memory 41 stores at least program instructions, and the processor 42 is configured to execute the program instructions to implement the steps in any of the above speech synthesis method embodiments. Specifically, reference can be made to the foregoing disclosed embodiments, which will not be elaborated herein. As a possible example, the electronic device 40 may include, but is not limited to, an office notebook, a smart phone, a tablet computer, a learning machine, a server, etc. The specific type of the electronic device 40 is not limited herein.
[0067] Specifically, the processor 42 is used to control itself and the memory 41 to implement the steps in any of the above speech synthesis method embodiments. The processor 42 can also be referred to as a CPU (Central Processing Unit). The processor 42 may be an integrated circuit chip with signal processing capabilities. The processor 42 can also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Additionally, the processor 42 can be implemented jointly by integrated circuit chips.
[0068] In the above solution, the electronic device 40 predicts the pronunciation features of the character sequence to be synthesized based on the feature prediction model, and the character sequence to be synthesized is a text sequence or a phoneme sequence. The pronunciation features at least include the feature information related to the pronunciation content and pronunciation prosody of the character sequence to be synthesized, and predicts the acoustic features based on the acoustic model for the pronunciation features and the sequence hidden layer features of the character sequence to be synthesized. The sequence hidden layer features are the hidden layer features obtained during the process of the feature prediction model processing the character sequence to be synthesized. Then, based on the vocoder, the acoustic features are restored to waveforms to obtain the synthesized speech. On the one hand, the pronunciation features predicted based on the character sequence to be synthesized at least include the feature information related to the pronunciation content and pronunciation prosody of the character sequence to be synthesized. Therefore, using this pronunciation feature for subsequent acoustic feature prediction and speech waveform restoration can ensure that the feature information related to the pronunciation content, that is, the semanticity of the character sequence to be synthesized, is referred to in the speech synthesis process, which helps to improve the stability of speech synthesis as much as possible. It can also ensure that the feature information related to the pronunciation prosody, that is, the prosody of the character sequence to be synthesized, is referred to in the speech synthesis process, which helps to improve the naturalness of speech synthesis as much as possible. On the other hand, the acoustic model not only predicts based on the pronunciation features but also further supplements with the sequence hidden layer features. Therefore, it can further constrain by combining the sequence hidden layer features during the acoustic feature prediction process, which helps to avoid the predicted acoustic features deviating from the character sequence to be synthesized as much as possible. Therefore, the stability and naturalness of speech synthesis can be improved.
[0069] Please refer to Figure 5 , Figure 5It is a schematic framework diagram of an embodiment of the computer-readable storage medium 50 of the present application. The computer-readable storage medium 50 stores program instructions 51 that can be run by a processor, and the program instructions 51 are used to implement the steps in any of the above-described embodiments of the speech synthesis method.
[0070] In the above solution, the computer-readable storage medium 50 predicts the pronunciation features of the character sequence to be synthesized based on a feature prediction model, and the character sequence to be synthesized is a text sequence or a phoneme sequence. The pronunciation features at least include the feature information related to the pronunciation content and pronunciation prosody of the character sequence to be synthesized. Then, based on an acoustic model, the pronunciation features and the sequence hidden layer features of the character sequence to be synthesized are predicted to obtain acoustic features. The sequence hidden layer features are the hidden layer features obtained during the process of the feature prediction model processing the character sequence to be synthesized. Then, based on a vocoder, the acoustic features are restored to waveforms to obtain synthesized speech. On the one hand, the pronunciation features predicted based on the character sequence to be synthesized at least include the feature information related to the pronunciation content and pronunciation prosody of the character sequence to be synthesized. Therefore, using these pronunciation features for subsequent acoustic feature prediction and speech waveform restoration can ensure that the feature information related to the pronunciation content, that is, the semanticity of the character sequence to be synthesized, is referred to during the speech synthesis process, which helps to improve the stability of speech synthesis as much as possible. It can also ensure that the feature information related to the pronunciation prosody, that is, the prosody of the character sequence to be synthesized, is referred to during the speech synthesis process, which helps to improve the naturalness of speech synthesis as much as possible. On the other hand, the acoustic model not only predicts based on the pronunciation features but also further supplements with the sequence hidden layer features. Therefore, it can further perform constraints by combining the sequence hidden layer features during the acoustic feature prediction process, which helps to avoid the predicted acoustic features deviating from the character sequence to be synthesized as much as possible. Therefore, the stability and naturalness of speech synthesis can be improved.
[0071] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0072] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.
[0073] In several embodiments provided by this application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the apparatuses or units can be in electrical, mechanical or other forms.
[0074] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0075] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0076] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in each embodiment of this application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical disks and other various media that can store program codes.
[0077] If the technical solution of this application involves personal information, before the product applying the technical solution of this application processes personal information, it has clearly informed the personal information processing rules and obtained the independent consent of the individual. If the technical solution of this application involves sensitive personal information, before the product applying the technical solution of this application processes sensitive personal information, it has obtained the individual's separate consent and at the same time meets the requirements of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform that the personal information collection scope has been entered and personal information will be collected. If an individual voluntarily enters the collection scope, it is regarded as consenting to the collection of their personal information; or on the device for personal information processing, when the personal information processing rules are informed by obvious signs / information, personal authorization is obtained through pop-up messages or by asking the individual to upload their personal information by themselves, etc.; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
Claims
1. A speech synthesis method, characterized in that: include: Predicting pronunciation features of a character sequence to be synthesized based on a feature prediction model; wherein the character sequence to be synthesized is a text sequence or a phoneme sequence, and the pronunciation features at least include feature information related to the pronunciation content and pronunciation rhythm of the character sequence to be synthesized; Predicting the pronunciation features and the sequence hidden features of the character sequence to be synthesized based on the acoustic model to obtain acoustic features; wherein the sequence hidden features are the hidden features obtained in the process of the feature prediction model processing the character sequence to be synthesized; Performing waveform restoration on the acoustic features based on a vocoder to obtain synthesized speech; The feature prediction model and the acoustic model are trained separately and then jointly trained based on fourth speech data, wherein the fourth speech data includes a fourth speech and a fourth pronunciation feature of the fourth speech, a fourth character sequence, a marked character sequence and a real acoustic feature, and the step of the joint training includes: Predicting the fourth character sequence based on the feature prediction model to obtain a predicted pronunciation feature of the fourth character sequence; Predicting based on the predicted pronunciation features of the fourth character sequence to obtain a predicted character sequence aligned with the predicted pronunciation features of the fourth character sequence, and predicting the predicted pronunciation features of the fourth character sequence and the sample sequence hidden features of the fourth character sequence based on the acoustic model to obtain predicted acoustic features of the fourth speech; wherein the sample sequence hidden features are hidden features obtained in the process of processing the fourth character sequence by the feature prediction model; Based on the difference between the predicted pronunciation feature and the fourth pronunciation feature, the difference between the predicted character sequence and the annotated character sequence of the fourth voice, the difference between the predicted acoustic feature of the fourth voice and the real acoustic feature, and the generative adversarial loss of the generative adversarial network, the network parameters of the feature prediction model and the acoustic model are adjusted; wherein the discriminator in the generative adversarial network is at least used to discriminate whether the predicted acoustic feature is the real acoustic feature.
2. The method according to claim 1, characterized in that The feature prediction model is trained based on the first pronunciation feature of the first character sequence and the first speech to which it belongs, with the first pronunciation feature as the optimization target. The first pronunciation feature of the first speech is extracted by a feature extraction model. The feature extraction model is trained with the second speech as the model input and with the first attribute marked by the second speech as the optimization target and the second attribute marked by the second speech as the degradation target. The first attribute includes at least one of the second character sequence and the real fundamental frequency of the second speech, and the second attribute includes at least the real speaker of the second speech.
3. The method according to claim 2, characterized in that The step of extracting the first pronunciation feature comprises: Extracting first speech hidden layer features of the first speech; Encoding based on the first speech hidden layer feature to obtain a first speech coding feature; Vector quantization is performed based on the first speech coding feature to obtain the first pronunciation feature.
4. The method according to claim 2 or 3, characterized in that: The training steps of the feature extraction model include: Extracting second speech hidden layer features of the second speech; Encoding based on the second speech hidden layer feature to obtain a second speech coding feature; Predicting based on the second speech coding feature to obtain prediction information of the second speech with respect to the first attribute and prediction information with respect to the second attribute; Adjust the network parameters of the feature extraction model based at least on a first loss measured by the difference between the labeling information and the prediction information of the second speech regarding the first attribute, and a second loss measured by the difference between the labeling information and the prediction information of the second speech regarding the second attribute; wherein the gradient of the second loss is flipped during the parameter adjustment process.
5. The method according to claim 4, characterized in that Before adjusting the network parameters of the feature extraction model based at least on the first loss measured by the difference between the labeling information and the prediction information of the second speech with respect to the first attribute, and the second loss measured by the difference between the labeling information and the prediction information of the second speech with respect to the second attribute, the method further includes: Performing vector quantization based on the second speech coding feature to obtain a second pronunciation feature, and performing decoding based on the second pronunciation feature to obtain a predicted acoustic feature of the second speech; Based on the second pronunciation feature and the vector quantized codebook, a third loss is obtained, and based on the real acoustic feature and the predicted acoustic feature of the second speech, a fourth loss is obtained; The adjusting of the network parameters of the feature extraction model based on at least a first loss measured by a difference between the labeling information and the prediction information of the second speech with respect to the first attribute, and a second loss measured by a difference between the labeling information and the prediction information of the second speech with respect to the second attribute, comprises: Based on the first loss, the second loss, the third loss and the fourth loss, a network parameter of the feature extraction model is adjusted.
6. The method according to claim 2, characterized in that The training steps of the feature prediction model include: Selecting a real text sequence or a real phoneme sequence of the first speech as the first character sequence with a preset probability; Predicting the first character sequence based on the feature prediction model to obtain predicted pronunciation features of the first character sequence; Perform prediction based on the predicted pronunciation feature to obtain a predicted character sequence aligned with the predicted pronunciation feature; Based on the difference between the predicted pronunciation feature of the first speech and the first pronunciation feature, and the difference between the marked character sequence of the first speech and the predicted character sequence, the network parameters of the feature prediction model are adjusted; wherein, when the first character sequence is the real text sequence, the predicted character sequence is a predicted phoneme sequence, and the marked character sequence is the real phoneme sequence, and when the first character sequence is the real phoneme sequence, the predicted character sequence is a predicted text sequence, and the marked character sequence is the real text sequence.
7. The method according to claim 2, characterized in that The acoustic model is trained after the feature extraction model training converges, and the training steps of the acoustic model include: Performing feature extraction on the third speech based on the feature extraction model to obtain a third pronunciation feature of the third speech; Predicting the third pronunciation feature and the third character sequence of the third speech based on the acoustic model to obtain the predicted acoustic feature of the third speech; wherein the third character sequence is a real text sequence or a real phoneme sequence of the third speech; Based on the acoustic reconstruction loss measured by the difference between the real acoustic features and the predicted acoustic features of the third speech, and the generative adversarial loss of the generative adversarial network, the network parameters of the acoustic model are adjusted; wherein the discriminator in the generative adversarial network is at least used to determine whether the predicted acoustic features are the real acoustic features.
8. The method according to claim 2, characterized in that: The fourth pronunciation feature is extracted by the feature extraction model, and when the fourth character sequence is a true phoneme sequence of the fourth voice, the annotated character sequence of the fourth voice is a true text sequence of the fourth voice, and when the fourth character sequence is a true text sequence of the fourth voice, the annotated character sequence of the fourth voice is a true phoneme sequence of the fourth voice.
9. The method according to claim 8, characterized in that When the fourth character sequence is a true text sequence of the fourth voice, the predicted character sequence is a predicted phoneme sequence; when the fourth character sequence is a true phoneme sequence of the fourth voice, the predicted character sequence is a predicted text sequence.
10. A speech synthesis device, characterized in that: include: A pronunciation prediction module, used to predict the pronunciation features of the character sequence to be synthesized based on the feature prediction model; wherein the character sequence to be synthesized is a text sequence or a phoneme sequence, and the pronunciation features at least include feature information related to the pronunciation content and pronunciation rhythm of the character sequence to be synthesized; An acoustic prediction module, used for predicting the pronunciation features and the sequence hidden features of the character sequence to be synthesized based on the acoustic model to obtain acoustic features; wherein the sequence hidden features are the hidden features obtained in the process of the feature prediction model processing the character sequence to be synthesized; A waveform recovery module, used for performing waveform recovery on the acoustic features based on a vocoder to obtain synthesized speech; The feature prediction model and the acoustic model are trained separately and then jointly trained based on fourth speech data, the fourth speech data includes a fourth speech and a fourth pronunciation feature of the fourth speech, a fourth character sequence, a marked character sequence and a real acoustic feature, and the speech synthesis device also includes: A second pronunciation prediction module, used for predicting the fourth character sequence based on the feature prediction model to obtain a predicted pronunciation feature of the fourth character sequence; A second sequence prediction module, configured to make a prediction based on the predicted pronunciation feature of the fourth character sequence to obtain a predicted character sequence aligned with the predicted pronunciation feature of the fourth character sequence; a third acoustic prediction module, configured to predict the predicted pronunciation features of the fourth character sequence and the sample sequence hidden features of the fourth character sequence based on the acoustic model to obtain the predicted acoustic features of the fourth speech; wherein the sample sequence hidden features are the hidden features obtained in the process of processing the fourth character sequence by the feature prediction model; The fourth module optimization module is used to adjust the network parameters of the feature prediction model and the acoustic model based on the difference between the predicted pronunciation feature and the fourth pronunciation feature, the difference between the predicted character sequence and the marked character sequence of the fourth voice, the difference between the predicted acoustic feature of the fourth voice and the real acoustic feature, and the generative adversarial loss of the generative adversarial network; wherein the discriminator in the generative adversarial network is at least used to discriminate whether the predicted acoustic feature is the real acoustic feature.
11. An electronic device, characterized in that: The invention at least comprises a memory and a processor coupled to each other, wherein the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the speech synthesis method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: Program instructions that can be executed by a processor are stored, and the program instructions are used to implement the speech synthesis method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Speech synthesis method and device thereof
CN113488022A
Construction method of voice synthesizer, and voice synthesis method and device
CN113823257A
Model training method and device, server and medium
CN115206284A
Speech processing method, feature set learning method, computing device and computer readable storage medium
CN118571228A