Speech synthesis method, apparatus, device, and storage medium

CN115762468BActive Publication Date: 2026-09-18IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211397831.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-09
Publication Date
2026-09-18
Estimated Expiration
2042-11-09

AI Technical Summary

Technical Problem

本案申请人发现,在某些语音合成场景中,用户希望合成后的语音能够符合用户所要表达的情感,而现有技术无法解决这一问题

Benefits of technology

[0018] By employing the above technical solution, this application pre-configures an acoustic information generation module. This module, based on phonemes extracted from the text to be synthesized, generates acoustic information matching the phonemes, with the goal of generating acoustic information that can be used to predict the emotional type of the text to be synthesized. Then, based on the generated acoustic information, synthesized speech is obtained. Therefore, this application, when generating acoustic information based on the phonemes of the text to be synthesized, specifies the direction of acoustic information generation. That is, it ensures that the generated acoustic information can be used as a basis to predict the emotional type of the text to be synthesized, thereby guaranteeing that the generated acoustic information contains the emotional information expressed by the text to be synthesized. Furthermore, when performing speech synthesis based on this acoustic information containing the emotional information expressed by the text to be synthesized, the synthesized speech can conform to the emotion to be expressed by the text to be synthesized, thus improving the emotional expression capability of the synthesized speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115762468B_ABST
    Figure CN115762468B_ABST
Patent Text Reader

Abstract

This application discloses a speech synthesis method, apparatus, device, and storage medium. The application includes a pre-configured acoustic information generation module. This module generates acoustic information matching the phonemes extracted from the text to be synthesized, with the goal of generating acoustic information that can be used to predict the emotional type of the text. Based on the generated acoustic information, synthesized speech is obtained. Therefore, this application specifies the direction of acoustic information generation, enabling the generated acoustic information to predict the emotional type of the text to be synthesized. This ensures that the generated acoustic information contains the emotional information expressed by the text to be synthesized. Furthermore, when speech synthesis is performed based on this acoustic information containing the emotional information expressed by the text to be synthesized, the synthesized speech conforms to the emotion expressed by the text to be synthesized, thus improving the emotional expression capability of the synthesized speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech synthesis technology, and more specifically, to a speech synthesis method, apparatus, device, and storage medium. Background Technology

[0002] Speech synthesis technology is an intelligent voice technology that converts text into speech, and it is one of the core technologies for realizing human-computer interaction. With the continuous development and improvement of speech synthesis technology, it has now been widely applied to all aspects of social life, including public services (information broadcasting, intelligent customer service, etc.), smart hardware (smart speakers, smart robots, etc.), smart transportation (voice navigation, intelligent in-vehicle devices, etc.), education (smart classrooms, foreign language learning, etc.), and entertainment (audio reading, film dubbing, virtual IP, etc.), creating extensive economic and social value.

[0003] Current research in speech synthesis technology generally focuses on how to synthesize speech with different genders, accents, speech rates, and timbres. For example, in the speech synthesis process, speaker representation vectors are used to control the prosody and timbre of the synthesized speech. The applicant in this case discovered that in certain speech synthesis scenarios, users want the synthesized speech to match the emotions they wish to express, a problem that existing technologies cannot solve. Summary of the Invention

[0004] In view of the above problems, this application is made to provide a speech synthesis method, apparatus, device, and storage medium to improve the emotional expression ability of synthesized speech. The specific solution is as follows:

[0005] Firstly, a speech synthesis method is provided, including:

[0006] Obtain the text to be synthesized;

[0007] The phonemes of the text to be synthesized are extracted, and the pre-configured acoustic information generation module is used to generate matching acoustic information based on the phonemes. The acoustic information is acoustic information that can be used to predict the emotion type of the text to be synthesized.

[0008] Based on the acoustic information, synthesized speech is obtained.

[0009] Secondly, a speech synthesis device is provided, comprising:

[0010] The text to be synthesized acquisition unit is used to acquire the text to be synthesized;

[0011] A phoneme extraction unit is used to extract the phonemes of the text to be synthesized;

[0012] An acoustic information generation unit is used to generate matching acoustic information based on the phonemes using a pre-configured acoustic information generation module, wherein the acoustic information is acoustic information that can be used to predict the emotion type of the text to be synthesized.

[0013] A speech synthesis unit is used to obtain synthesized speech based on the acoustic information.

[0014] Thirdly, a speech synthesis device is provided, including: a memory and a processor;

[0015] The memory is used to store programs;

[0016] The processor is used to execute the program to implement the various steps of the speech synthesis method described above.

[0017] Fourthly, a storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the various steps of the speech synthesis method described above.

[0018] By employing the above technical solution, this application pre-configures an acoustic information generation module. This module, based on phonemes extracted from the text to be synthesized, generates acoustic information matching the phonemes, with the goal of generating acoustic information that can be used to predict the emotional type of the text to be synthesized. Then, based on the generated acoustic information, synthesized speech is obtained. Therefore, this application, when generating acoustic information based on the phonemes of the text to be synthesized, specifies the direction of acoustic information generation. That is, it ensures that the generated acoustic information can be used as a basis to predict the emotional type of the text to be synthesized, thereby guaranteeing that the generated acoustic information contains the emotional information expressed by the text to be synthesized. Furthermore, when performing speech synthesis based on this acoustic information containing the emotional information expressed by the text to be synthesized, the synthesized speech can conform to the emotion to be expressed by the text to be synthesized, thus improving the emotional expression capability of the synthesized speech. Attached Figure Description

[0019] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0020] Figure 1 This is a schematic flowchart of a speech synthesis method provided in an embodiment of this application;

[0021] Figure 2 An example is shown in the schematic diagram of an optional structure for an acoustic information generation model;

[0022] Figure 3An example is shown in the schematic diagram of an optional training process for an acoustic information generation model;

[0023] Figure 4 An example of an alternative structural diagram for another acoustic information generation model is provided;

[0024] Figure 5 An example is provided, illustrating the spatial distribution of emotional intensity.

[0025] Figure 6 An illustration of an alternative training process for another acoustic information generation model is provided.

[0026] Figure 7 This is a schematic diagram of a speech synthesis device provided in an embodiment of this application;

[0027] Figure 8 This is a schematic diagram of the structure of the speech synthesis device provided in the embodiments of this application. Detailed Implementation

[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0029] This application provides a speech synthesis solution applicable to various scenarios requiring speech synthesis, such as public service scenarios (information broadcasting, intelligent customer service, etc.), smart hardware scenarios (smart speakers, smart robots, etc.), smart transportation scenarios (voice navigation, smart in-vehicle devices, etc.), education scenarios (smart classrooms, foreign language learning, etc.), and entertainment scenarios (audio reading, film dubbing, virtual IP, etc.). Using this speech synthesis solution, synthesized speech can convey emotions consistent with the intended message of the text to be synthesized, thereby enhancing the emotional expressiveness of the synthesized speech.

[0030] The proposed solution can be implemented using a terminal with data processing capabilities, such as a mobile phone, computer, server, or cloud platform.

[0031] Next, combined Figure 1 As shown, the speech synthesis method of this application may include the following steps:

[0032] Step S100: Obtain the text to be synthesized.

[0033] The text to be synthesized is the text that needs to be processed through speech synthesis. This text can be user-inputted text or a response generated by the machine based on human-computer dialogue information.

[0034] Step S110: Extract the phonemes of the text to be synthesized.

[0035] Step S120: Using a pre-configured acoustic information generation module, generate matching acoustic information based on the phonemes, wherein the acoustic information is acoustic information that can be used to predict the emotion type of the text to be synthesized.

[0036] Specifically, the acoustic information generation module configured in this application has the ability to generate acoustic information based on the input phonemes, which can be used as a basis to predict the sentiment type of the text to be synthesized. That is, the acoustic information generation module is configured to generate acoustic information matching the phonemes with the goal of generating acoustic information that can be used to predict the sentiment type of the text to be synthesized.

[0037] Based on this, this application inputs the phonemes extracted from the text to be synthesized into the acoustic information generation module to obtain the acoustic information output by the acoustic information generation module.

[0038] The acoustic information contains the emotional information expressed by the text to be synthesized.

[0039] Step S130: Based on the acoustic information, synthesized speech is obtained.

[0040] Specifically, acoustic information can take many forms, such as spectral features (e.g., Mel spectral features, linear spectral features, etc.) and speech waveforms. Based on this acoustic information, synthesized speech can be obtained.

[0041] Taking acoustic information as a spectral feature as an example, the spectral feature can be fed into a vocoder to obtain the synthesized speech output by the vocoder.

[0042] The speech synthesis method provided in this application includes a pre-configured acoustic information generation module. This module generates acoustic information matching the phonemes extracted from the text to be synthesized, with the goal of generating acoustic information that can be used to predict the emotional type of the text. Based on the generated acoustic information, synthesized speech is obtained. Therefore, this application specifies the direction of acoustic information generation when generating acoustic information based on the phonemes of the text to be synthesized. This ensures that the generated acoustic information can be used to predict the emotional type of the text, thus guaranteeing that the generated acoustic information contains the emotional information expressed by the text. Furthermore, when speech synthesis is performed based on this acoustic information containing the emotional information expressed by the text, the synthesized speech conforms to the emotion expressed by the text, improving the emotional expression capability of the synthesized speech.

[0043] In some embodiments of this application, an optional implementation of the acoustic information generation module described above is introduced. The acoustic information generation module can adopt a neural network model structure, such as an acoustic information generation model. This acoustic information generation model, through pre-training, can be configured to acquire acoustic information matching the phonemes based on the phonemes of the input text to be synthesized, with the acquisition direction being to obtain acoustic information that can be used to predict the emotion type of the text to be synthesized corresponding to the phonemes.

[0044] Based on this, the extracted phonemes can be input into the acoustic information generation model to obtain the acoustic information output by the acoustic information generation model.

[0045] Combination Figure 2 As shown in the embodiments of this application, an optional structural composition of the acoustic information generation model is provided, which may include:

[0046] Phoneme encoding module and acoustic decoding module.

[0047] The phoneme encoding module is used to encode the input phonemes to obtain phoneme encoding features.

[0048] The phoneme encoding module can use a TTS (Text To Speech) Encoder to encode the input phonemes and obtain phoneme encoding features.

[0049] The acoustic decoding module is used to decode based on the phoneme encoding features to obtain decoded acoustic information, which is acoustic information that can be used to predict the emotion type of the text to be synthesized.

[0050] The acoustic decoding module can use a TTS (Text To Speech) Encoder to decode the phoneme encoding features and obtain the decoded acoustic information.

[0051] In order to enable the acoustic information generated by the acoustic information generation model to be used as a basis for predicting the sentiment type of the text to be synthesized, based on the above... Figure 2 The example acoustic information generation model structure, in this application embodiment, provides an optional training method for the acoustic information generation model, which uses supervised sentiment constraints and unsupervised bidirectional sentiment constraints for training.

[0052] Specifically, the training process of the acoustic information generation model can be referred to Figure 3 As shown, it may include the following steps:

[0053] S11. Obtain the training phonemes and training acoustic information of the training speech corresponding to the training text.

[0054] Specifically, this application can pre-collect training text and corresponding training speech. Phonemes are extracted from the training text to obtain training phonemes. Acoustic information is extracted from the training speech to obtain training acoustic information.

[0055] The training speech corresponding to the training text can be the collected training speech that matches the training text, such as performing text recognition on the collected audio to obtain the training text and training speech, or extracting subtitle text from multimedia data carrying subtitles as training text, and extracting audio that matches the subtitles from multimedia data as training speech.

[0056] In addition, if only training text is available, it can be synthesized into acoustic information using TTS (Text To Speech) and used as training acoustic information.

[0057] Furthermore, the training texts obtained in this step are labeled with their respective sentiment type tags. This application can predefine several sentiment type tags, such as happiness, anger, sadness, and neutrality, and then label the training texts with their respective sentiment type tags.

[0058] S12. Using the acoustic information generation model, generate matching target acoustic information based on the training phonemes.

[0059] The acoustic information generation model can generate matching acoustic information based on the input training phonemes, as described in the foregoing embodiments, to serve as the target acoustic information.

[0060] S13. Using a preset first emotion classification model, emotion classification is performed based on the target acoustic information to obtain the first emotion classification result.

[0061] Specifically, this application can pre-set a first sentiment classification model, which is used to predict the corresponding sentiment type label based on the input target acoustic information. Based on this, the target acoustic information output by the aforementioned acoustic information generation model is input into the first sentiment classification model to obtain the first sentiment classification result output by the first sentiment classification model.

[0062] S14. Using a preset second emotion classification model, emotion classification is performed based on the training acoustic information to obtain the second emotion classification result.

[0063] Specifically, in order to achieve supervised emotion constraints and unsupervised bidirectional emotion constraints, this application embodiment also pre-configures a second emotion classification model for performing emotion classification based on the input training acoustic information to obtain the second emotion classification result.

[0064] S15. Based on the second sentiment classification result and the sentiment type label of the training text, calculate the sentiment classification loss; based on the first sentiment classification result and the second sentiment classification result, calculate the classification result alignment loss.

[0065] Specifically, based on the second sentiment classification result output by the second sentiment classification model and the sentiment type labels of the training text, supervised sentiment constraints can be performed, that is, sentiment classification loss can be calculated.

[0066] Furthermore, unsupervised bidirectional sentiment constraints can be applied based on the first and second sentiment classification results, that is, the alignment loss between the first and second sentiment classification results can be calculated.

[0067] S16. Based on the sentiment classification loss and the classification result alignment loss, calculate the total loss, and train the network parameters of the model with minimizing the total loss as the training objective.

[0068] Specifically, by combining the two losses mentioned above, the total loss can be obtained. The network parameters of the model are trained with the goal of minimizing the total loss. Here, the network parameters of the model can include the network parameters of the acoustic information generation model and the first and second sentiment classification models.

[0069] After the set training termination conditions are met, the trained acoustic information generation model can be obtained.

[0070] The training method for the acoustic information generation model provided in this embodiment improves the overall emotional encoding expressiveness of the acoustic information generation model by setting supervised emotional constraints and bidirectional emotional constraints.

[0071] In some embodiments of this application, combined with Figure 4 As shown in the embodiments of this application, another optional structural composition of the acoustic information generation model is provided, which may include:

[0072] Phoneme encoding module, emotion intensity adjustment module, and acoustic decoding module.

[0073] The phoneme encoding module is used to encode the input phonemes to obtain phoneme encoding features.

[0074] The phoneme encoding module can use a TTS (Text To Speech) Encoder to encode the input phonemes and obtain phoneme encoding features.

[0075] The emotional intensity adjustment module is used to locate the emotional intensity direction of the phoneme encoding feature in the emotional intensity space, and to adjust the direction vector of the phoneme encoding feature according to the emotional intensity direction to obtain the adjusted phoneme encoding feature.

[0076] It should be noted that different texts to be synthesized may express the same type of emotion, but the intensity of the emotion may differ. This application allows for pre-defining emotion intensity levels, and each emotion type can be further subdivided into multiple emotion intensity levels. (See reference...) Figure 5 This example illustrates a spatial distribution diagram of emotional intensity. Emotional intensity can be categorized into three levels: neutral, medium, and high.

[0077] For example, the text to be synthesized is "My parents took me to the zoo today.", which corresponds to the sentiment type "pleasure" and the sentiment intensity "neutral".

[0078] The text to be synthesized, “I had so much fun going to the zoo with my parents today,” corresponds to the “pleasure” sentiment type and the “medium” sentiment intensity.

[0079] The text to be synthesized, “Today my parents took me to the zoo, and I was so happy!”, corresponds to the “pleasure” emotion type and the “high” emotion intensity.

[0080] In this embodiment, in order to achieve continuous adjustment of emotional intensity and further improve the emotional intensity expression ability of synthesized speech, an emotional intensity adjustment module was added to the acoustic information generation model.

[0081] By training the acoustic information generation model, the emotion intensity adjustment module can locate the emotion intensity direction of the input phoneme coding features in the emotion intensity space, and adjust the direction vector of the input phoneme coding features according to the emotion intensity direction to obtain the adjusted phoneme coding features.

[0082] An acoustic decoding module is used to decode based on the adjusted phoneme encoding features to obtain decoded acoustic information, which is acoustic information that can be used to predict the emotion type of the text to be synthesized.

[0083] The acoustic decoding module can use a TTS (Text To Speech) Encoder to decode the phoneme encoding features and obtain the decoded acoustic information.

[0084] Compared to Figure 2The corresponding acoustic information generation model, as described in this embodiment, adds an emotion intensity adjustment module between the phoneme encoding module and the acoustic decoding module. This module can locate the emotion intensity direction of the input phoneme encoding features in the emotion intensity space, adjust the direction vector of the input phoneme encoding features according to the emotion intensity direction, and send the adjusted phoneme encoding features to the acoustic decoding module for decoding. This makes the acoustic information obtained by decoding further contain emotion intensity information, and the synthesized speech has a stronger ability to express the emotion intensity.

[0085] In the above Figure 4 Based on the structure of the example acoustic information generation model, this application embodiment further provides an optional training method for the above-mentioned acoustic information generation model, which is still trained with supervised emotional constraints and unsupervised bidirectional emotional constraints.

[0086] Specifically, the training process of the acoustic information generation model can be referred to Figure 6 As shown, it may include the following steps:

[0087] S21. Obtain the training phonemes corresponding to the training text, and obtain the training acoustic information of the training speech corresponding to the training text, wherein the training text is labeled with its corresponding emotion type label and emotion intensity label.

[0088] This step is similar to step S11 in the previous embodiment, except that the training text obtained in this embodiment is further labeled with an emotion intensity label in addition to the emotion type label. The emotion intensity label can be predefined, for example, three emotion intensity labels: neutral, medium, and high.

[0089] S22. Using the acoustic information generation model, generate matching target acoustic information based on the trained phonemes.

[0090] In this embodiment, the acoustic information generation model can be based on... Figure 4 The example structure generates target acoustic information.

[0091] S23. Using a preset first emotion classification model, emotion classification and emotion intensity classification are performed based on the target acoustic information to obtain the first emotion classification result.

[0092] Specifically, this step corresponds to step S13 in the aforementioned embodiments, as detailed above.

[0093] S24. Using a preset emotion coding module, the training acoustic information is encoded to obtain emotion coding features. The emotion intensity adjustment module is used to adjust the direction vector of the emotion coding features to obtain the adjusted emotion coding features.

[0094] Specifically, in this embodiment, an emotion encoding module can be pre-configured to encode the input training acoustic information to obtain emotion encoding features. Further, an emotion intensity adjustment module in the acoustic information generation model can be used to locate the emotion intensity direction of the emotion encoding features in the emotion intensity space. According to the emotion intensity direction, the direction vector of the emotion encoding features is adjusted to obtain the adjusted emotion encoding features, which are then fed into the second emotion classification model.

[0095] S25. Using the preset second emotion classification model, emotion classification and emotion intensity classification are performed based on the adjusted emotion coding features to obtain the second emotion classification result.

[0096] Specifically, the second sentiment classification model can classify sentiment types and sentiment intensities based on the adjusted sentiment encoding features of the input, thus obtaining the second sentiment classification result.

[0097] S26. Based on the second sentiment classification result and the sentiment type label and sentiment intensity label of the training text, calculate the sentiment and intensity classification loss; based on the first sentiment classification result and the second sentiment classification result, calculate the classification result alignment loss.

[0098] Specifically, the second sentiment classification result includes sentiment type classification result and sentiment intensity classification result. Sentiment type classification loss is calculated based on sentiment type classification result and labeled sentiment type label, and sentiment intensity classification loss is calculated based on sentiment intensity classification result and labeled sentiment intensity label. The sentiment and intensity classification loss are composed of sentiment type classification loss and sentiment intensity classification loss.

[0099] Furthermore, based on the first and second sentiment classification results, the alignment loss of the two classification results is calculated. Specifically, the alignment loss of the sentiment type classification result and the alignment loss of the sentiment intensity classification result in the first and second sentiment classification results can be calculated separately, and the final classification result alignment loss is composed of the two alignment losses.

[0100] S27. Based on the emotion and intensity classification loss and the classification result alignment loss, calculate the total loss, and train the network parameters of the model with minimizing the total loss as the training objective.

[0101] Specifically, by combining the two losses mentioned above, the total loss can be obtained. The network parameters of the model are trained with the goal of minimizing the total loss. Here, the network parameters of the model can include the network parameters of the acoustic information generation model, the first and second sentiment classification models, and the sentiment encoding module.

[0102] After the set training termination conditions are met, the trained acoustic information generation model can be obtained.

[0103] The training method for the acoustic information generation model provided in this embodiment improves the overall emotional encoding expressiveness of the acoustic information generation model by setting supervised emotional constraints and bidirectional emotional constraints.

[0104] The speech synthesis apparatus provided in the embodiments of this application is described below. The speech synthesis apparatus described below can be referred to in correspondence with the speech synthesis method described above.

[0105] See Figure 7 , Figure 7 This is a schematic diagram of the structure of a speech synthesis device disclosed in an embodiment of this application.

[0106] like Figure 7 As shown, the device may include:

[0107] The text to be synthesized acquisition unit 11 is used to acquire the text to be synthesized;

[0108] Phoneme extraction unit 12 is used to extract the phonemes of the text to be synthesized;

[0109] The acoustic information generation unit 13 is used to generate matching acoustic information based on the phonemes using a pre-configured acoustic information generation module, wherein the acoustic information is acoustic information that can be used to predict the emotion type of the text to be synthesized.

[0110] The speech synthesis unit 14 is used to obtain synthesized speech based on the acoustic information.

[0111] Optionally, the aforementioned acoustic information may include spectral characteristics, acoustic waveforms, etc.

[0112] Optionally, the aforementioned acoustic information generation module can be an acoustic information generation model. The process by which the acoustic information generation unit uses the acoustic information generation module to generate matching acoustic information based on the phonemes can include:

[0113] The phonemes are input into the acoustic information generation model to obtain the acoustic information output by the model;

[0114] The acoustic information generation model is configured to acquire acoustic information matching the phonemes based on the input phonemes, with the acquisition direction being to obtain acoustic information that can be used to predict the emotion type of the text corresponding to the phonemes.

[0115] Optionally, embodiments of this application disclose two different structural compositions of the aforementioned acoustic information generation model: the first,

[0116] The acoustic information generation model may include: a phoneme encoding module and an acoustic decoding module;

[0117] The phoneme encoding module is used to encode the input phonemes to obtain phoneme encoding features;

[0118] The acoustic decoding module is used to decode based on the phoneme encoding features to obtain decoded acoustic information, which is acoustic information that can be used to predict the emotion type of the text to be synthesized.

[0119] Based on the structure of the first acoustic information generation model described above, the apparatus of this application may further include:

[0120] The first model training unit is used to train the acoustic information generation model of the first structure described above. The training process may include:

[0121] Obtain the training phonemes corresponding to the training text, and obtain the training acoustic information of the training speech corresponding to the training text, wherein the training text is labeled with its respective emotion type tag.

[0122] The acoustic information generation model is used to generate matching target acoustic information based on the trained phonemes;

[0123] Using a preset first emotion classification model, emotion classification is performed based on the target acoustic information to obtain the first emotion classification result;

[0124] Using a pre-defined second emotion classification model, emotion classification is performed based on the training acoustic information to obtain the second emotion classification result;

[0125] Based on the second sentiment classification result and the sentiment type labels of the training text, calculate the sentiment classification loss; based on the first sentiment classification result and the second sentiment classification result, calculate the classification result alignment loss.

[0126] Based on the sentiment classification loss and the classification result alignment loss, the total loss is calculated, and the network parameters of the model are trained with minimizing the total loss as the training objective.

[0127] The second type

[0128] Acoustic information generation models may include:

[0129] Phoneme encoding module, emotion intensity adjustment module, and acoustic decoding module;

[0130] The phoneme encoding module is used to encode the input phonemes to obtain phoneme encoding features;

[0131] The emotional intensity adjustment module is used to locate the emotional intensity direction of the phoneme coding feature in the emotional intensity space, and adjust the direction vector of the phoneme coding feature according to the emotional intensity direction to obtain the adjusted phoneme coding feature.

[0132] The acoustic decoding module is used to decode based on the adjusted phoneme encoding features to obtain decoded acoustic information, which is acoustic information that can be used to predict the emotion type of the text to be synthesized.

[0133] Based on the structure of the second acoustic information generation model described above, the apparatus of this application may further include:

[0134] The second model training unit is used to train the acoustic information generation model of the second structure described above. The training process may include:

[0135] Obtain the training phonemes corresponding to the training text, and obtain the training acoustic information of the training speech corresponding to the training text, wherein the training text is labeled with its corresponding emotion type label and emotion intensity label.

[0136] The acoustic information generation model is used to generate matching target acoustic information based on the trained phonemes;

[0137] Using a preset first emotion classification model, emotion classification and emotion intensity classification are performed based on the target acoustic information to obtain the first emotion classification result;

[0138] The training acoustic information is encoded using a preset emotion coding module to obtain emotion coding features. The direction vector of the emotion coding features is adjusted using the emotion intensity adjustment module to obtain adjusted emotion coding features.

[0139] Using a pre-defined second emotion classification model, emotion classification and emotion intensity classification are performed based on the adjusted emotion coding features to obtain the second emotion classification result;

[0140] Based on the second sentiment classification result and the sentiment type label and sentiment intensity label of the training text, calculate the sentiment and intensity classification loss; based on the first sentiment classification result and the second sentiment classification result, calculate the classification result alignment loss;

[0141] Based on the emotion and intensity classification loss and the classification result alignment loss, the total loss is calculated, and the network parameters of the model are trained with minimizing the total loss as the training objective.

[0142] The speech synthesis device provided in this application embodiment can be applied to speech synthesis devices such as terminals, servers, and cloud computing. Optionally, Figure 8 The hardware structure block diagram of the speech synthesis device is shown below, with reference to... Figure 8 The hardware structure of a speech synthesis device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;

[0143] In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4;

[0144] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0145] Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;

[0146] The memory stores a program, which the processor can call. The program is used for:

[0147] Obtain the text to be synthesized;

[0148] The phonemes of the text to be synthesized are extracted, and the pre-configured acoustic information generation module is used to generate matching acoustic information based on the phonemes. The acoustic information is acoustic information that can be used to predict the emotion type of the text to be synthesized.

[0149] Based on the acoustic information, synthesized speech is obtained.

[0150] Optionally, the refined and extended functions of the program can be found in the description above.

[0151] This application embodiment also provides a storage medium that can store a program suitable for execution by a processor, the program being used for:

[0152] Obtain the text to be synthesized;

[0153] The phonemes of the text to be synthesized are extracted, and the pre-configured acoustic information generation module is used to generate matching acoustic information based on the phonemes. The acoustic information is acoustic information that can be used to predict the emotion type of the text to be synthesized.

[0154] Based on the acoustic information, synthesized speech is obtained.

[0155] Optionally, the refined and extended functions of the program can be found in the description above.

[0156] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0157] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0158] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A speech synthesis method, characterized in that, include: Obtain the text to be synthesized; The phonemes of the text to be synthesized are extracted, and a pre-configured acoustic information generation model is used to generate matching acoustic information based on the phonemes. The acoustic information is acoustic information that can be used to predict the emotion type of the text to be synthesized. Based on the acoustic information, synthesized speech is obtained; The training process of the acoustic information generation model includes: Obtain the training phonemes corresponding to the training text, and obtain the training acoustic information of the training speech corresponding to the training text, wherein the training text is labeled with its respective emotion type tag. The acoustic information generation model is used to generate matching target acoustic information based on the trained phonemes; Using a preset first emotion classification model, emotion classification is performed based on the target acoustic information to obtain the first emotion classification result; Using a pre-defined second emotion classification model, emotion classification is performed based on the training acoustic information to obtain the second emotion classification result; Based on the second sentiment classification result and the sentiment type labels of the training text, calculate the sentiment classification loss; based on the first sentiment classification result and the second sentiment classification result, calculate the classification result alignment loss. Based on the sentiment classification loss and the classification result alignment loss, the total loss is calculated, and the network parameters of the model are trained with minimizing the total loss as the training objective.

2. The method according to claim 1, characterized in that, The process of generating matching acoustic information based on the phonemes using a pre-configured acoustic information generation model includes: The phonemes are input into the acoustic information generation model to obtain the acoustic information output by the model; The acoustic information generation model is configured to acquire acoustic information matching the phonemes based on the input phonemes, with the acquisition direction being to obtain acoustic information that can be used to predict the emotion type of the text corresponding to the phonemes.

3. The method according to claim 2, characterized in that, The acoustic information generation model includes: a phoneme encoding module and an acoustic decoding module; The phoneme encoding module is used to encode the input phonemes to obtain phoneme encoding features; The acoustic decoding module is used to decode based on the phoneme encoding features to obtain decoded acoustic information, which is acoustic information that can be used to predict the emotion type of the text to be synthesized.

4. The method according to any one of claims 1-3, characterized in that, The acoustic information is spectral characteristics.

5. A speech synthesis method, characterized in that, include: Obtain the text to be synthesized; The phonemes of the text to be synthesized are extracted, and a pre-configured acoustic information generation model is used to generate matching acoustic information based on the phonemes. The acoustic information is acoustic information that can be used to predict the emotion type of the text to be synthesized. Based on the acoustic information, synthesized speech is obtained; The training process of the acoustic information generation model includes: Obtain the training phonemes corresponding to the training text, and obtain the training acoustic information of the training speech corresponding to the training text, wherein the training text is labeled with its corresponding emotion type label and emotion intensity label. The acoustic information generation model is used to generate matching target acoustic information based on the trained phonemes; Using a preset first emotion classification model, emotion classification and emotion intensity classification are performed based on the target acoustic information to obtain the first emotion classification result; The training acoustic information is encoded using a preset emotion coding module to obtain emotion coding features. The direction vector of the emotion coding features is adjusted using an emotion intensity adjustment module to obtain adjusted emotion coding features. Using a pre-defined second emotion classification model, emotion classification and emotion intensity classification are performed based on the adjusted emotion coding features to obtain the second emotion classification result; Based on the second sentiment classification result and the sentiment type label and sentiment intensity label of the training text, calculate the sentiment and intensity classification loss; based on the first sentiment classification result and the second sentiment classification result, calculate the classification result alignment loss; Based on the emotion and intensity classification loss and the classification result alignment loss, the total loss is calculated, and the network parameters of the model are trained with minimizing the total loss as the training objective.

6. The method according to claim 5, characterized in that, The process of generating matching acoustic information based on the phonemes using a pre-configured acoustic information generation model includes: The phonemes are input into the acoustic information generation model to obtain the acoustic information output by the model; The acoustic information generation model is configured to acquire acoustic information matching the phonemes based on the input phonemes, with the acquisition direction being to obtain acoustic information that can be used to predict the emotion type of the text corresponding to the phonemes.

7. The method according to claim 6, characterized in that, The acoustic information generation model includes: a phoneme encoding module, an emotion intensity adjustment module, and an acoustic decoding module; The phoneme encoding module is used to encode the input phonemes to obtain phoneme encoding features; The emotional intensity adjustment module is used to locate the emotional intensity direction of the phoneme coding feature in the emotional intensity space, and adjust the direction vector of the phoneme coding feature according to the emotional intensity direction to obtain the adjusted phoneme coding feature. The acoustic decoding module is used to decode based on the adjusted phoneme encoding features to obtain decoded acoustic information, which is acoustic information that can be used to predict the emotion type of the text to be synthesized.

8. The method according to any one of claims 5-7, characterized in that, The acoustic information is spectral characteristics.

9. A speech synthesis device, characterized in that, include: The text to be synthesized acquisition unit is used to acquire the text to be synthesized; A phoneme extraction unit is used to extract the phonemes of the text to be synthesized; An acoustic information generation unit is used to generate matching acoustic information based on the phonemes using a pre-configured acoustic information generation model, wherein the acoustic information is acoustic information that can be used to predict the emotion type of the text to be synthesized. A speech synthesis unit is used to obtain synthesized speech based on the acoustic information; The training process of the acoustic information generation model includes: Obtain the training phonemes corresponding to the training text, and obtain the training acoustic information of the training speech corresponding to the training text, wherein the training text is labeled with its respective emotion type tag. The acoustic information generation model is used to generate matching target acoustic information based on the trained phonemes; Using a preset first emotion classification model, emotion classification is performed based on the target acoustic information to obtain the first emotion classification result; Using a pre-defined second emotion classification model, emotion classification is performed based on the training acoustic information to obtain the second emotion classification result; Based on the second sentiment classification result and the sentiment type labels of the training text, calculate the sentiment classification loss; based on the first sentiment classification result and the second sentiment classification result, calculate the classification result alignment loss. Based on the sentiment classification loss and the classification result alignment loss, the total loss is calculated, and the network parameters of the model are trained with minimizing the total loss as the training objective.

10. A speech synthesis device, characterized in that, include: The text to be synthesized acquisition unit is used to acquire the text to be synthesized; A phoneme extraction unit is used to extract the phonemes of the text to be synthesized; An acoustic information generation unit is used to generate matching acoustic information based on the phonemes using a pre-configured acoustic information generation model, wherein the acoustic information is acoustic information that can be used to predict the emotion type of the text to be synthesized. A speech synthesis unit is used to obtain synthesized speech based on the acoustic information; The training process of the acoustic information generation model includes: Obtain the training phonemes corresponding to the training text, and obtain the training acoustic information of the training speech corresponding to the training text, wherein the training text is labeled with its corresponding emotion type label and emotion intensity label. The acoustic information generation model is used to generate matching target acoustic information based on the trained phonemes; Using a preset first emotion classification model, emotion classification and emotion intensity classification are performed based on the target acoustic information to obtain the first emotion classification result; The training acoustic information is encoded using a preset emotion coding module to obtain emotion coding features. The direction vector of the emotion coding features is adjusted using an emotion intensity adjustment module to obtain adjusted emotion coding features. Using a pre-defined second emotion classification model, emotion classification and emotion intensity classification are performed based on the adjusted emotion coding features to obtain the second emotion classification result; Based on the second sentiment classification result and the sentiment type label and sentiment intensity label of the training text, calculate the sentiment and intensity classification loss; based on the first sentiment classification result and the second sentiment classification result, calculate the classification result alignment loss; Based on the emotion and intensity classification loss and the classification result alignment loss, the total loss is calculated, and the network parameters of the model are trained with minimizing the total loss as the training objective.

11. A speech synthesis device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is configured to execute the program to implement the various steps of the speech synthesis method as described in any one of claims 1 to 4 or 5 to 8.

12. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the various steps of the speech synthesis method as described in any one of claims 1 to 4 or 5 to 8.

Citation Information

Patent Citations

  • Speech synthesis model, model training method and speech synthesis method

    CN113920977A

  • Speech synthesis method and system

    CN114627851A