Speech synthesis method, related device, electronic device and storage medium

By extracting and utilizing global and local pronunciation features, the problem of difficulty in adapting to different scenarios in the existing technology is solved, and the pronunciation diversity and adaptability of pronunciation synthesis are improved.

CN114283781BActive Publication Date: 2025-05-30UNIV OF SCI & TECH OF CHINA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111650035.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-05-30
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

Existing speech synthesis technologies are difficult to adapt to different scenarios, such as intelligent interaction and novel reading, and cannot freely synthesize speech with different rhythms.

Method used

By obtaining the text to be synthesized, the pronunciation attributes and the speaker logo, the global and local pronunciation characteristics are extracted, and the pronunciation synthesis is performed based on these characteristics, the synthesis of different pronunciations is achieved.

Benefits of technology

It has achieved improved adaptability to different scenarios, and can accurately and freely synthesize pronunciations with different rhythms, enhancing the scope of application of pronunciation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283781B_ABST
    Figure CN114283781B_ABST
Patent Text Reader

Abstract

The present application discloses a speech synthesis method, related devices, electronic devices, and storage media. The speech synthesis method includes: obtaining a text to be synthesized, a first speech attribute, and a second speech attribute; wherein the first speech attribute includes at least one of an emotion category and a style category, and the second speech attribute includes a speaker identifier; obtaining a global prosody feature having the first speech attribute, and making a prediction based on the text to be synthesized, the first speech attribute, and the second speech attribute to obtain a local prosody feature; wherein the global prosody feature contains sentence-level prosody feature information, and the local prosody feature contains word-level prosody feature information; performing synthesis based on the text to be synthesized, the global prosody feature, and the local prosody feature to obtain a synthesized speech. The above solution can freely synthesize voices with different prosodies and improve the adaptability to different scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of speech synthesis, and in particular, to a speech synthesis method, related devices, electronic devices, and storage media. Background Art

[0002] Speech synthesis technology is one of the branches in the field of artificial intelligence research. It mainly converts text information into audible voice information, that is, enables machines to talk like humans.

[0003] Currently, although speech synthesis technology can generate synthetic speech approaching natural speech, the synthetic speech often represents the average prosody of the database. Therefore, it is difficult to adapt to different scenarios such as novels, news, customer service, and live broadcasts. For example, in the intelligent interaction scenario, existing intelligent customer services cannot empathize with people's speaking emotions. Even when people are very angry, the synthetic speech still gives answers in the same tone; or, in the novel reading scenario, the identities of the narrator and dialogue characters in the novel are unclear, and for the different emotions of different characters, the synthetic speech is also in the same tone. In view of this, how to freely synthesize speech with different prosodies has become an urgent problem to be solved. Summary of the Invention

[0004] The main technical problem to be solved by this application is to provide a speech synthesis method, related devices, electronic devices, and storage media, which can freely synthesize speech with different prosodies and improve the adaptability to different scenarios.

[0005] To solve the above technical problem, in the first aspect of this application, a speech synthesis method is provided, including: obtaining a text to be synthesized, a first speech attribute, and a second speech attribute; wherein, the first speech attribute includes at least one of an emotion category and a style category, and the second speech attribute includes a speaker identifier; obtaining a global prosody feature with the first speech attribute, and making a prediction based on the text to be synthesized, the first speech attribute, and the second speech attribute to obtain a local prosody feature; wherein, the global prosody feature contains sentence-level prosody feature information, and the local prosody feature contains word-level prosody feature information; synthesizing based on the text to be synthesized, the global prosody feature, and the local prosody feature to obtain a synthetic speech.

[0006] To solve the above technical problems, a second aspect of the present application provides a speech synthesis device, including: an acquisition module, a global prosody module, a local prosody module, and a synthesis module. The acquisition module is used to acquire the text to be synthesized, a first speech attribute, and a second speech attribute; wherein, the first speech attribute includes at least one of an emotion category and a style category, and the second speech attribute includes a speaker identifier; the global prosody module is used to acquire global prosody features with the first speech attribute, and the local prosody module is used to make predictions based on the text to be synthesized, the first speech attribute, and the second speech attribute to obtain local prosody features; wherein, the global prosody features contain sentence-level prosody feature information, and the local prosody features contain word-level prosody feature information; the synthesis module is used to perform synthesis based on the text to be synthesized, the global prosody features, and the local prosody features to obtain synthesized speech.

[0007] To solve the above technical problems, a third aspect of the present application provides an electronic device, including a memory and a processor coupled to each other. Program instructions are stored in the memory, and the processor is used to execute the program instructions to implement the speech synthesis method in the first aspect above.

[0008] To solve the above technical problems, a fourth aspect of the present application provides a computer-readable storage medium storing program instructions that can be run by a processor, and the program instructions are used to implement the speech synthesis method in the first aspect above.

[0009] In the above solution, the text to be synthesized, the first speech attribute, and the second speech attribute are acquired, and the first speech attribute includes at least one of an emotion category and a style category, and the second speech attribute includes a speaker identifier. Based on this, global prosody features with the first speech attribute are acquired, and predictions are made based on the text to be synthesized, the first speech attribute, and the second speech attribute to obtain local prosody features. The global prosody features contain sentence-level prosody feature information, and the local prosody features contain word-level prosody feature information. In addition, synthesis is performed based on the text to be synthesized, the global prosody features, and the local prosody features to obtain synthesized speech. On the one hand, referring to the sentence-level prosody feature information during the synthesis process can ensure that the overall synthesized speech conforms to the first speech attribute. On the other hand, further referring to the word-level prosody features during the synthesis process can control the subtle changes in the local prosody of the synthesized speech. Therefore, speech with different prosodies can be accurately and freely synthesized, improving the adaptability to different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a schematic flowchart of an embodiment of the speech synthesis method of the present application;

[0011] Figure 2 is a schematic diagram of the process of an embodiment of training a global prosody extraction network;

[0012] Figure 3It is a schematic process diagram of an embodiment of prosody conversion;

[0013] Figure 4 It is a schematic process diagram of an embodiment of training a local prosody extraction network;

[0014] Figure 5 It is a schematic process diagram of an embodiment of training a local prosody prediction network;

[0015] Figure 6 It is a schematic process diagram of an embodiment of transfer learning;

[0016] Figure 7 It is a schematic process diagram of an embodiment of training a prosody encoding network and a text encoding network;

[0017] Figure 8 It is a schematic process diagram of an embodiment of training a synthesis network;

[0018] Figure 9 It is a schematic process diagram of an embodiment of emotional style transfer;

[0019] Figure 10 It is a schematic framework diagram of an embodiment of the voice synthesis device of the present application;

[0020] Figure 11 It is a schematic framework diagram of an embodiment of the electronic device of the present application;

[0021] Figure 12 It is a schematic framework diagram of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners

[0022] The solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings of the specification.

[0023] In the following description, specific details such as specific system architectures, interfaces, and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the present application.

[0024] The terms "system" and "network" are often used interchangeably herein. The term "and / or" herein merely describes an association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after. In addition, "plurality" herein means two or more than two.

[0025] Please refer to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the voice synthesis method of the present application.

[0026] Specifically, the following steps may be included:

[0027] Step S11: Obtain the text to be synthesized, the first voice attribute, and the second voice attribute.

[0028] In the embodiments of the present disclosure, the first voice attribute includes at least one of an emotion category and a style category. For example, the first voice attribute may only include the emotion category; or, the first voice attribute may only include the style category; or, in order to improve the applicability of voice synthesis, the first voice attribute may include the emotion category and the style category, which is not limited herein.

[0029] In an implementation scenario, the emotion category can be selected from several preset emotion categories. Taking an interactive scenario as an example, the emotion of the user's voice can be recognized to obtain the target emotion category of the user's voice. On this basis, the emotion category can be selected from several preset emotion categories as the first voice attribute, and the synthesized voice can be obtained through subsequent steps to respond to the user's voice. Exemplarily, when the emotion of the user's voice is recognized and the target emotion category is "angry", the preset emotion category "sorry" can be selected. Other situations can be inferred by analogy and will not be exemplified one by one here. In addition, the several preset emotion categories may include but are not limited to: comfort, cute, doting, naughty, encouraging, sorry, etc., that is, the fluctuations of the preset emotion categories are relatively gentle, so that in various scenarios, even if the emotion category is selected incorrectly among the several preset emotion categories, the impact on the user experience can be reduced as much as possible. For example, for the aforementioned text to be synthesized "Searching, please wait a moment", its correct emotion category is "neutral", and the selected emotion category is "naughty". In this case, not only will it not affect the user experience, but it will also give the user a better interactive experience, which is conducive to greatly improving the fault tolerance of emotion category selection. Of course, in addition, the several preset emotion categories may also include but are not limited to: sad, happy, angry, surprised, questioning, etc., which is not limited herein.

[0030] In an implementation scenario, the style category can be selected from several preset style categories. Specifically, the style category can be selected according to the usage scenario. In addition, the style category may include but is not limited to: novel, news, customer service, interaction, colloquial, etc., which is not limited herein.

[0031] In the embodiments of the present disclosure, the second voice attribute includes a speaker identifier. It should be noted that different speakers have different speaker identifiers, and the speaker identifier is used to distinguish the timbres of different speakers, so that in the process of voice synthesis, the synthesized voice with the expected timbre can be obtained according to the speaker identifier. For specific details, please refer to the subsequent related descriptions and will not be elaborated herein.

[0032] In an implementation scenario, the text to be synthesized can be pre-input. Taking the novel reading scenario as an example, texts such as narration and dialogue can be pre-input as the text to be synthesized; or, taking the news scenario as an example, news texts can be pre-input as the text to be synthesized. Other cases can be inferred by analogy and will not be exemplified one by one here.

[0033] In an implementation scenario, the text to be synthesized can also be predicted. Taking the customer service scenario as an example, corresponding response interaction texts can be generated based on user interaction texts as the text to be synthesized. It should be noted that the user interaction texts can be input by users or obtained by recognizing user voices. In addition, a question-and-answer model such as based on the Encoder-Decoder (i.e., encoder-decoder) structure can be used to predict the user interaction texts to obtain the corresponding response interaction texts. Other cases can be inferred by analogy and will not be exemplified one by one here.

[0034] Step S12: Obtain global prosody features with the first voice attribute, and make predictions based on the text to be synthesized, the first voice attribute, and the second voice attribute to obtain local prosody features.

[0035] In the embodiments of the present disclosure, the global prosody features include sentence-level prosody feature information. That is to say, the global prosody features can reflect the overall prosody information of sentences, so that the speech synthesis can be overall controlled through the global prosody features.

[0036] In an implementation scenario, the synthesized speech can be synthesized based on a speech synthesis model. The speech synthesis model can be trained based on sample data. The sample data includes sample speeches marked with sample voice attributes and sample texts corresponding to the sample speeches. And the global prosody features can be obtained based on the sample global prosody features of the reference sample speech. The reference sample speech can specifically be a sample speech with the first voice attribute. In the above manner, the synthesized speech is synthesized through the speech synthesis model, and the speech synthesis model is trained based on sample data. The sample data includes sample speeches marked with sample voice attributes and sample texts corresponding to the sample speeches, while the global prosody features are obtained based on the sample global prosody features of the reference sample speech. The reference sample speech is a sample speech with the first voice attribute, that is, the global prosody features can be obtained from the sample data used to train the synthesis model, so it is beneficial to improve the convenience of extracting the global prosody features.

[0037] In a specific implementation scenario, the foregoing sample voice attributes may include a sample first voice attribute and a sample second voice attribute. The sample first voice attribute may include at least one of the emotional category of the sample voice and the style category of the sample voice. In addition, the sample second voice attribute may include the speaker identifier of the sample voice (i.e., the unique identifier of the speaker who emits the sample voice). For the specific meanings of the emotional category and the style category, reference may be made to the foregoing related descriptions, which will not be elaborated herein. It should be noted that, different from the foregoing first voice attribute and second voice attribute, the sample first voice attribute and the sample second voice attribute can be marked according to the actual situation of the sample voice. For example, if the emotional category actually reflected by the sample voice is "neutral" and the actual generated scenario is "news", then the sample first voice attributes "neutral" and "news" can be marked for it. And if the speaker who actually emits the sample voice is Speaker No. 1, then the second voice attribute "001" can be marked for it. Other situations can be deduced by analogy, and no further examples will be given here.

[0038] In a specific implementation scenario, the foregoing voice synthesis model may include a global prosody extraction network, and the global prosody extraction network is used to assist in obtaining global prosody features. The specific process of the assisted extraction will not be elaborated herein for the time being. It should be noted that after the global prosody extraction network is trained, it can decouple the global prosody features it extracts from the speaker and the speech content. Specifically, please refer to Figure 2 , Figure 2 which is a schematic diagram of the process of an embodiment of training the global prosody extraction network. As Figure 2 shown, the first sample global prosody feature of the sample voice can be extracted. Exemplarily, the global prosody extraction network can perform global prosody extraction on the acoustic features of the sample voice to obtain the first sample global prosody feature. It should be noted that in the disclosed embodiments of the present application, for the convenience of distinction, the global prosody features extracted at different stages and for different data are distinguished by different naming methods. In the real scenario, these global prosody features can all be represented by vectors with a preset dimension. On this basis, a first sample synthesized voice can be synthesized based on the first sample global prosody feature and the sample text corresponding to the sample voice, and a predicted first voice attribute of the sample voice can be obtained based on the first sample global prosody feature, and the second sample global prosody feature of the first sample synthesized voice can be extracted. It should be noted that as Figure 2As shown, the speech synthesis model may further include a synthesis network, and the synthesis network includes a phoneme encoder and a decoder. After the phoneme sequence of the sample text is encoded by the phoneme encoder, phoneme features can be obtained. After the first sample global prosody feature is weighted by a weighting factor, it is input into the decoder together with the phoneme features for decoding to synthesize acoustic features. Based on the synthesized acoustic features, the first sample synthesized speech can be further synthesized. The specific synthesis process will not be elaborated here. In addition, similar to the sample first speech attribute, predicting the first speech attribute may include at least one of the predicted emotion category and the predicted style category of the sample speech. The specific meaning can refer to the relevant descriptions of the emotion category and the style category mentioned above and will not be elaborated here. On this basis, based on the differences between the predicted first speech attribute and the sample first speech attribute, and the differences between the first sample global prosody feature and the second sample global prosody feature, the network parameters of the global prosody extraction network can be adjusted at least. Specifically, during the training process, the classification difference between the predicted first speech attribute and the sample first speech attribute can be minimized, and the distribution difference between the first sample global prosody feature and the second sample global prosody feature can be minimized. Exemplarily, the classification difference between the predicted first speech attribute and the sample first speech attribute can be measured by a loss function such as cross entropy, and the distribution difference between the first sample global prosody feature and the second sample global prosody feature can be measured by a loss function such as mean square error. The specific process of difference measurement can refer to the technical details of loss functions such as cross entropy and mean square error and will not be elaborated here. In addition, the global prosody extraction network may include, but is not limited to, AE (AutoEncoder), VAE (Variational AutoEncoder), CVAE (Conditional Variational AutoEncoder), GST (i.e., Global Style Tokens), etc., which are not limited here. It should be noted that during the process of adjusting the network parameters, the network parameters of the synthesis network can also be adjusted together. Of course, the network parameters of the global prosody extraction network can also be adjusted first based on the differences between the predicted first speech attribute and the sample first speech attribute until the global prosody extraction network converges. On this basis, the network parameters of the synthesis network can be adjusted based on the differences between the first sample global prosody feature and the second sample global prosody feature to train the synthesis network until it converges. The training method is not limited here.In the above method, the speech synthesis model includes a global prosody extraction network, and the global prosody extraction network is used to extract the global prosody features of the samples. The sample speech attributes at least include the first sample speech attribute. Based on this, the first sample global prosody feature of the sample speech is extracted, and then based on the first sample global prosody feature and the sample text corresponding to the sample speech, the first sample synthesized speech is synthesized, and a prediction is made based on the first sample global prosody feature to obtain the predicted first speech attribute of the sample speech, and the second sample global prosody feature of the first sample synthesized speech is extracted. On this basis, at least the network parameters of the global prosody extraction network are adjusted based on the difference between the predicted first speech attribute and the first sample speech attribute, and the difference between the first sample global prosody feature and the second sample global prosody feature. Therefore, through the classification loss constraint of the speech attributes and the distribution loss constraint of the prosody features, the global prosody extraction network can be decoupled from information such as the speaker and the text during the feature extraction process, so that the extracted global prosody features only represent feature information such as emotion and style, improving the accuracy of the global prosody features.

[0039] In a specific implementation scenario, based on the convergence of the training of the global prosody extraction network, global prosody features with the first speech attribute can be obtained. Specifically, the sample global prosody features of each reference sample speech can be fused to obtain the global prosody features. Exemplarily, sample speeches with the same sample first speech attribute as the labeled sample can be obtained, and the trained and converged global prosody extraction network can be used to extract features from these sample speeches respectively to obtain the corresponding sample global prosody features. On this basis, these sample global prosody features can be processed by means of calculating the mean value, sampling, etc. to obtain the global prosody features with the first speech attribute. Alternatively, the sample global prosody features of the first target sample speech can also be used as the global prosody features. It should be noted that the first target sample speech is selected from the reference sample speeches, and the text similarity between the sample text corresponding to the first target sample speech and the text to be synthesized meets the first condition. Exemplarily, the text similarity can be calculated between the text to be synthesized and the sample texts corresponding to each reference sample speech respectively, and the reference sample speech with the highest text similarity can be taken as the first target sample speech. That is to say, the first condition can be specifically set as the highest text similarity. On this basis, the trained and converged global prosody extraction network can be used to extract features from the first target sample speech to obtain the global prosody features with the first speech attribute. It should be noted that the text similarity can be obtained by processing the text to be synthesized and the sample text using a text matching model, and the text matching model can be obtained by adjusting the downstream task using a big data unsupervised network, or the similarity can be calculated by means of rules such as template matching, which is not limited here. Alternatively, the sample global prosody features of the second target sample speech can also be used as the global prosody features. Similar to the first target sample speech, the second target sample speech is also selected from the reference sample speeches, and the presentation intensity of the second target sample speech with respect to the first speech attribute meets the second condition. Exemplarily, the trained and converged global prosody extraction network can be used to extract features from each reference sample speech respectively to obtain the sample global prosody features of each reference sample speech. On this basis, classification prediction can be performed based on the sample global prosody features of each reference sample speech to obtain the prediction probability values of each reference sample speech for the emotional category (and / or style category) in the first speech attribute. Thus, the reference sample speech corresponding to the highest prediction probability value can be taken as the second target sample speech. That is to say, the presentation intensity with respect to the first speech attribute can be measured by the prediction probability values of each reference sample speech for the emotional category (and / or style category) in the first speech attribute, and the second condition can be specifically set as the highest presentation intensity with respect to the first speech attribute (i.e., the largest prediction probability value).In the above method, the global prosody features are obtained by integrating the sample global prosody features of each reference sample speech, or the sample global prosody features of the first target sample speech are used as the global prosody features, or the sample global prosody features of the second target sample speech are used as the global prosody features. The first target sample speech and the second target sample speech are both selected from the reference sample speech. The text similarity between the sample text corresponding to the first target sample speech and the text to be synthesized satisfies the first condition, and the presentation intensity of the second target sample speech with respect to the first speech attribute satisfies the second condition. The global prosody features with the first speech attribute can be obtained in multiple ways, which is beneficial to further improving the adaptability to different scenarios and expanding the application scope of speech synthesis.

[0040] In the embodiments of the present disclosure, the local prosody features include word-level prosody feature information. That is to say, the local prosody features can reflect the prosody information of a single word, so that the local detailed prosody changes can be improved through the local prosody features. It should be noted that in the embodiments of the present disclosure, the local prosody features can be obtained based on the text to be synthesized, the first speech attribute, and the second speech attribute by selecting the prosody conversion method as needed, or the local prosody features can be obtained based on the text to be synthesized, the first speech attribute, and the second speech attribute by selecting the transfer learning method. The two methods of prosody conversion and transfer learning are specifically introduced below.

[0041] In an implementation scenario, please refer to Figure 3 , Figure 3 which is a schematic diagram of the process of an embodiment of prosody conversion. As Figure 3 shown, taking sound library A as the source, the emotions and styles of sound library A can be transferred to the timbre of sound library B. Specifically, the speaker ID (i.e., speaker identifier) of the source and the corresponding emotion category and style category can be used for prosody conversion to generate local prosody features. Combining the global prosody features for speech synthesis and referring to the speaker ID (i.e., speaker identifier) of sound library B during the synthesis process, a synthesized speech with the target timbre (i.e., the speaker timbre of sound library B) and conforming to the prosody of the source (i.e., the emotions and styles of sound library A) can be generated. Therefore, during the training process, only the local prosody prediction network of the target speaker needs to be trained. Of course, in order to further expand the application scope of speech synthesis, the local prosody prediction networks of each speaker can also be trained, so that when the target speaker changes, only the local prosody prediction network of the corresponding speaker needs to be selected, without having to retrain.

[0042] Specifically, as described above, the synthetic speech can be synthesized based on a speech synthesis model, and the speech synthesis model may further include a local prosody prediction network, which is used to directly predict local prosody features based on the text to be synthesized, the first speech attribute, and the second speech attribute. In the above manner, the local prosody features are directly predicted by the local prosody prediction network for the text to be synthesized, the first speech attribute, and the second speech attribute, which is beneficial to improving the convenience of obtaining local prosody features.

[0043] In a specific implementation scenario, in addition to the global prosody extraction network, the speech synthesis model may further include a local prosody extraction network. As described above, the speech synthesis model can be trained based on sample data, and the sample data may include sample speech annotated with sample speech attributes and the sample text corresponding to the sample speech. Specifically, the local prosody prediction network can be trained based on the sample data after the local prosody extraction network converges. It should be noted that, similar to the global prosody extraction network, the local prosody extraction network and the local prosody prediction network may include, but are not limited to: AE, VAE, CVAE, GST, etc., which are not limited herein. In the above manner, the speech synthesis model further includes a local prosody extraction network, the speech synthesis model is trained based on sample data, the sample data includes sample speech annotated with sample speech attributes and the sample text corresponding to the sample speech, and the local prosody prediction network is trained based on the sample data after the local prosody extraction network converges, which can further train the local prosody prediction network after the local prosody extraction network is trained to assist the training of the local prosody prediction network, and is beneficial to improving the accuracy of the local prosody prediction network.

[0044] In a specific implementation scenario, during the process of training the local prosody extraction network, please refer to Figure 4 , Figure 4 is a schematic diagram of the process of an embodiment of training the local prosody extraction network. As Figure 4As shown, the first sample local prosodic feature of the sample speech can be extracted first, and the first text feature of the sample text can be extracted. It should be noted that a local prosody extraction network can be used to extract the first sample local prosodic feature of the sample speech. In addition, a pre-trained language model such as BERT (Bidirectional Encoder Representation from Transformers, that is, the Encoder of the bidirectional Transformer) can be used to extract the first text feature of the sample text. For specific details, reference can be made to the technical details of pre-trained language models such as BERT, which will not be elaborated here. On this basis, a prediction can be made based on the first sample local prosodic feature to obtain the predicted second speech attribute of the sample speech, and based on the first sample local prosodic feature and the sample text corresponding to the sample speech, a second sample synthesized speech can be synthesized, and the second sample local prosodic feature of the second sample synthesized speech can be extracted. It should be noted that as mentioned above, the speech synthesis model can include a synthesis network, so the first sample local prosodic feature and the sample text corresponding to the sample speech can be input into the synthesis network to obtain the second sample synthesized speech. For the specific process, reference can be made to the relevant description of the first sample synthesized speech above, which will not be elaborated here. At the same time, a local prosody extraction network can be used to perform local prosody extraction on the second sample synthesized speech to obtain the second sample local prosody extraction feature. In addition, in the embodiments of the present disclosure, for the convenience of distinction, at different stages and for different data, different names are used for the local prosodic features respectively extracted by the local prosody extraction network. In the real scenario, these local prosodic features can all be represented by vectors of a preset dimension. On this basis, at least the network parameters of the local prosody extraction network can be adjusted based on the difference between the predicted second speech attribute and the sample second speech attribute, the difference between the first sample local prosodic feature and the first text feature, and the difference between the first sample local prosodic feature and the second sample local prosodic feature. It should be noted that the classification difference between the predicted second speech attribute and the sample second speech attribute can be maximized, the distribution difference between the first sample local prosodic feature and the second sample local prosodic feature can be minimized, and the mutual information between the first sample local prosodic feature and the first text feature can be minimized, so that the local prosodic features extracted by the local prosody extraction network can be decoupled from the speaker information and text information, and then only contain feature information related to prosody information such as fundamental frequency fluctuations and pronunciation durations. In addition, in the real scenario, a loss function such as cross entropy can be used to measure the classification difference between the predicted second speech attribute and the sample second speech attribute, and a loss function such as mean square error can be used to measure the distribution difference between the first sample local prosodic feature and the second sample local prosodic feature.Further, during the training process, the local prosody extraction network can be adjusted based on the differences between the predicted second speech attribute and the sample second speech attribute, as well as the differences between the first sample local prosody feature and the first text feature, so as to train the local prosody extraction network until convergence. On this basis, the network parameters of the synthesis network can be adjusted based on the differences between the first sample local prosody feature and the second sample local prosody feature. Therefore, during the training process, it can be gradually trained in two stages, which is beneficial to reducing the training difficulty as much as possible and improving the training efficiency. In the above method, the sample speech attribute at least includes the sample second speech attribute. During the training process, the first sample local prosody feature of the sample speech is extracted, and the first text feature of the sample text is extracted. Then, based on the first sample local prosody feature, the predicted second speech attribute of the sample speech is obtained. Based on the first sample local prosody feature and the sample text corresponding to the sample speech, the second sample synthesized speech is synthesized, and the second sample local prosody feature of the second sample synthesized speech is extracted. On this basis, based on the differences between the predicted second speech attribute and the sample second speech attribute, the differences between the first sample local prosody feature and the first text feature, and the differences between the first sample local prosody feature and the second sample local prosody feature, at least the network parameters of the local prosody extraction network are adjusted, which can decouple the local prosody features extracted by the local prosody extraction network from the speaker information and text information, and further make it only contain feature information related to prosody information such as fundamental frequency fluctuations and pronunciation duration, which is beneficial to improving the accuracy of the local prosody features.

[0045] In a specific implementation scenario, based on the convergence of the training of the local prosody extraction network, the local prosody prediction network can be further trained. Please refer to Figure 5 , Figure 5It is a schematic process diagram of an embodiment for training a local prosody prediction network. During the training process, the third sample local prosody feature of the sample speech can be extracted based on the local prosody extraction network, and the sample text corresponding to the sample speech and the sample speech attributes annotated for the sample speech can be predicted based on the local prosody prediction network to obtain the predicted local prosody feature. On this basis, the network parameters of the local prosody prediction network can be adjusted based on the difference between the third sample local prosody feature and the predicted local prosody feature. Specifically, during the processing of the local prosody prediction network, the phoneme sequence and text information of the sample text can be used as inputs, and the semantic information in the text information can be extracted through a pre-trained language model. The semantic information is fused with the phoneme sequence through alignment encoding (including but not limited to attention mechanisms, autoregressive models, etc.) to obtain the predicted local prosody feature. During this process, the third sample local prosody feature extracted after the sample speech passes through the locally trained and converged local prosody extraction network can be used as the target to constrain the local prosody prediction network. By minimizing the difference between the third sample local prosody feature and the predicted local prosody feature, the two can be made as close as possible, thereby improving the accuracy of the local prosody prediction network as much as possible. Specifically, the mutual information between the third sample local prosody feature and the predicted local prosody feature can be calculated, and during the training process, the network parameters of the local prosody prediction network can be adjusted by maximizing the mutual information. In the above manner, the third sample local prosody feature of the sample speech is extracted based on the local prosody extraction network, and the sample text corresponding to the sample speech and the sample speech attributes annotated for the sample speech are predicted based on the local prosody prediction network to obtain the predicted local prosody feature. On this basis, the network parameters of the local prosody prediction network are adjusted based on the difference between the third sample local prosody feature and the predicted local prosody feature, which is beneficial to improving the accuracy of the local prosody prediction network.

[0046] In one implementation scenario, please refer to Figure 6 , Figure 6 which is a schematic process diagram of an embodiment of transfer learning. As Figure 6 shown, different from the aforementioned prosody conversion, transfer learning uses mixed data of multiple styles, multiple emotions, and multiple speakers, and controls the generation of local prosody features through speaker ID (i.e., speaker identifier), emotion ID (i.e., emotion category), and style ID (i.e., style category). By using the target speaker ID (i.e., target speaker identifier) and changing different emotion IDs (i.e., emotion categories) and style IDs (i.e., style categories), local prosody features are generated, enabling the target speaker to achieve speech synthesis with different styles and emotions by only changing the emotion and style controls, which is more convenient for engine development, system integration, and flexible control.

[0047] Specifically, the synthetic speech can be synthesized based on a speech synthesis model, and the speech synthesis model includes a prosody encoding network and a text encoding network. The text encoding network is used to predict speech attribute features based on the text to be synthesized and the first speech attribute, and the speech attribute features are independent of the speaker. The prosody encoding network is used to predict local prosody features based on the speech attribute features, the text to be synthesized, and the second speech attribute. It should be noted that the prosody encoding network can include, but is not limited to, CVAE, etc., and is not limited herein. In the above manner, by using the text encoding network to predict speech attribute features based on the text to be synthesized and the first speech attribute, and by using the prosody encoding network to predict local prosody features based on the speech attribute features, the text to be synthesized, and the second speech attribute, and the speech attribute features are independent of the speaker, it is possible to further decouple speech attributes such as emotion, style, and speaker through the prosody encoding network, which is beneficial to improving the control over emotion and style and enhancing the distinguishability of emotion and style.

[0048] In a specific implementation scenario, as described above, the speech synthesis model can further include a local prosody extraction network. Regarding the network structure and training process of the local prosody extraction network, reference can be made to the foregoing related descriptions and will not be elaborated herein. In addition, the speech synthesis model can be trained based on sample data. The sample data includes sample speech labeled with sample speech attributes and sample texts corresponding to the sample speech, and the prosody encoding network and the text encoding network are trained based on the sample data after the local prosody extraction network converges. After the local prosody extraction network is trained, the prosody encoding network and the text encoding network can be further trained to assist the training of the prosody encoding network and the text encoding network through the trained and converged local prosody extraction network, which is beneficial to improving the accuracy of the prosody encoding network and the text encoding network.

[0049] In a specific implementation scenario, please refer to Figure 7 , Figure 7 is a schematic diagram of the process of an embodiment of training the prosody encoding network and the text encoding network. As Figure 7As shown, during the training process of the prosody encoding network and the text encoding network, the fourth sample local prosody feature of the sample speech can be extracted first. It should be noted that the fourth sample local prosody feature can be extracted by the aforementioned local prosody extraction network. On this basis, the sample speech attribute feature of the fourth sample local prosody feature can be extracted, and based on the sample text and the sample first speech attribute, the predicted speech attribute feature can be obtained. It should be noted that the prosody encoding network can be used to extract the feature of the fourth sample local prosody feature to obtain the sample speech attribute feature. In addition, the sample text and the sample first speech attribute can be input into the text encoding network to obtain the predicted speech attribute feature. Based on this, the fifth sample local prosody feature can be predicted based on the sample text, the sample second speech attribute, and the predicted speech attribute feature. Specifically, the above sample text, the sample second speech attribute, and the predicted speech attribute feature can be input into the prosody encoding network to obtain the fifth sample local prosody feature. Thus, based on the difference between the sample speech attribute feature and the predicted speech attribute feature, and the difference between the fourth sample local prosody feature and the fifth sample local prosody feature, the network parameters of the prosody encoding network and the text encoding network can be adjusted. Specifically, the distribution difference between the sample speech attribute feature and the predicted speech attribute feature can be minimized, and the distribution difference between the fourth sample local prosody feature and the fifth sample local prosody feature can be minimized. It should be noted that in the real scenario, the distribution difference between the sample speech attribute feature and the predicted speech attribute feature can be measured by a loss function such as the mean square error, and the distribution difference between the fourth sample local prosody feature and the fifth sample local prosody feature can be measured by a loss function such as the mean square error. For the specific measurement method, the technical details of loss functions such as the mean square error can be referred to, which will not be elaborated here. By the above method, the fourth sample local prosody feature of the sample speech is extracted. Based on this, the sample speech attribute feature of the fourth sample local prosody feature is further extracted, and the predicted speech attribute feature is obtained based on the sample text and the sample first speech attribute. Thus, the fifth sample local prosody feature is predicted based on the sample text, the sample second speech attribute, and the predicted speech attribute feature. On this basis, based on the difference between the sample speech attribute feature and the predicted speech attribute feature, and the difference between the fourth sample local prosody feature and the fifth sample local prosody feature, the network parameters of the prosody encoding network and the text encoding network are adjusted, which can further decouple speech attributes such as emotion, style, and speaker, thereby improving the control ability of emotion and style, and being conducive to improving the distinguishability of emotion and style.

[0050] It should be noted that, up to this point, the global prosody extraction network, local prosody extraction network, local prosody prediction network, prosody encoding network, text encoding network, etc. can be trained, so that in the process of speech synthesis, the above various networks can be selectively used for speech synthesis according to actual needs. For specific details, please refer to the following relevant descriptions and will not be elaborated here.

[0051] Step S13: Synthesize based on the text to be synthesized, global prosody features, and local prosody features to obtain synthesized speech.

[0052] In an implementation scenario, as described above, the synthesized speech can be obtained by synthesizing based on a speech synthesis model. The speech synthesis model includes a global prosody extraction network, a local prosody extraction network, and a synthesis network. The speech synthesis model is trained based on sample data. The sample data includes sample speech with labeled sample speech attributes and sample texts corresponding to the sample speech. And the synthesis network is trained based on the sample data after the global prosody extraction network and the local prosody extraction network are respectively trained and converge. In the above manner, after the global prosody extraction network and the local prosody extraction network are respectively trained and converge, further training of the synthesis network is beneficial to improving the speech synthesis effect.

[0053] In a specific implementation scenario, please refer to Figure 8 , Figure 8 which is a schematic diagram of the process of an embodiment of training the synthesis network. As Figure 8 described, the third sample global prosody feature and the sixth sample local prosody feature of the sample speech can be extracted. It should be noted that specifically, the global prosody extraction network can be used to extract the prosody of the sample speech to obtain the third sample global prosody feature, and the local prosody extraction network can be used to extract the prosody of the sample speech to obtain the sixth sample layout prosody feature. On this basis, synthesis can be performed based on the sample text corresponding to the sample speech, the third sample global prosody feature, and the sixth sample local prosody feature to obtain the sample synthesized speech, and at least the network parameters of the synthesis network can be adjusted based on the difference between the sample speech and the sample synthesized speech. It should be noted that the sample text corresponding to the sample speech, the third sample global prosody feature, and the sixth sample local prosody feature can be input into the synthesis network to obtain the sample synthesized speech. In addition, as Figure 8As shown, during the synthesis process, the global prosody features of the third sample can also be scaled by a weighting factor. Further, during the training process, the network parameters of the local prosody extraction network can also be adjusted simultaneously. It should be noted that since high-quality style data and emotion data require the interpretation of professionals and the data recording cost is relatively high, in order to enable different speakers to generate synthetic speech with rich prosody variations. Through the above method, by decoupling the global prosody features, local prosody features from the speaker, the prosody of a certain speaker can be transferred to a speaker with a different prosody, so that even a speaker without emotion data can have the expressions of different emotions of an emotional speaker, greatly reducing the demand for high-quality emotion and style data.

[0054] In a specific implementation scenario, after the synthesis network training is completed, the text to be synthesized, global prosody features, and local prosody features can be input into the synthesis network to obtain synthetic speech. Please refer to Figure 8 , the global prosody features can be scaled by a weighting factor and then input into the synthesis network together with the local prosody features and the text to be synthesized, and finally the synthetic speech is obtained, and the synthetic speech has the emotion category and style category defined by the first speech attribute and the timbre corresponding to the speaker identifier defined by the second speech attribute.

[0055] In an implementation scenario, different from the aforementioned method of speech synthesis through the text to be synthesized, in the case of obtaining the speech to be transferred, the timbre of the speech to be transferred can be migrated to obtain synthetic speech with the timbre of the target speaker, and the synthetic speech maintains the speech attributes such as emotion and style of the original speech to be transferred and the speech content. Please refer to Figure 9 , Figure 9 is a schematic diagram of the process of an embodiment of emotion style transfer. As Figure 9 shown, after the speech to be transferred is subjected to prosody extraction by the global prosody extraction network and the local prosody extraction network respectively, the global prosody features and local prosody features can be obtained. At the same time, the text to be synthesized, global prosody features, local prosody features corresponding to the speech to be transferred, and the speaker identifier of the target speaker can be input into the synthesis network to obtain synthetic speech, and the synthetic speech has the original prosody and content of the speech to be transferred, but its timbre is the timbre of the target speaker.

[0056] In the above solution, the text to be synthesized, the first speech attribute, and the second speech attribute are obtained. The first speech attribute includes at least one of the emotion category and the style category, and the second speech attribute includes the speaker identifier. Based on this, the global prosody features with the first speech attribute are obtained, and predictions are made based on the text to be synthesized, the first speech attribute, and the second speech attribute to obtain the local prosody features. The global prosody features contain sentence-level prosody feature information, and the local prosody features contain word-level prosody feature information. And synthesis is performed based on the text to be synthesized, the global prosody features, and the local prosody features to obtain the synthesized speech. On the one hand, referring to the sentence-level prosody feature information during the synthesis process can ensure that the overall synthesized speech conforms to the first speech attribute. On the other hand, further referring to the word-level prosody features during the synthesis process can control the subtle changes in the local prosody of the synthesized speech. Therefore, it is possible to accurately and freely synthesize voices with different prosodies, improving the adaptability to different scenarios.

[0057] Please refer to Figure 10 , Figure 10 FIG. is a schematic framework diagram of an embodiment of the voice synthesis device 100 of the present application. The voice synthesis device 100 includes: an acquisition module 101, a global prosody module 102, a local prosody module 103, and a synthesis module 104. The acquisition module 101 is used to acquire the text to be synthesized, the first speech attribute, and the second speech attribute; wherein, the first speech attribute includes at least one of the emotion category and the style category, and the second speech attribute includes the speaker identifier; the global prosody module 102 is used to acquire the global prosody features with the first speech attribute, and the local prosody module 103 is used to make predictions based on the text to be synthesized, the first speech attribute, and the second speech attribute to obtain the local prosody features; wherein, the global prosody features contain sentence-level prosody feature information, and the local prosody features contain word-level prosody feature information; the synthesis module 104 is used to perform synthesis based on the text to be synthesized, the global prosody features, and the local prosody features to obtain the synthesized speech.

[0058] In the above solution, on the one hand, referring to the sentence-level prosody feature information during the synthesis process can ensure that the overall synthesized speech conforms to the first speech attribute. On the other hand, further referring to the word-level prosody features during the synthesis process can control the subtle changes in the local prosody of the synthesized speech. Therefore, it is possible to accurately and freely synthesize voices with different prosodies, improving the adaptability to different scenarios.

[0059] In some publicly disclosed embodiments, the synthesized speech is obtained by synthesizing based on a voice synthesis model, and the voice synthesis model is trained based on sample data. The sample data includes sample voices annotated with sample speech attributes and sample texts corresponding to the sample voices, and the global prosody features are obtained based on the sample global prosody features of the reference sample voices, and the reference sample voices are sample voices with the first speech attribute.

[0060] Therefore, a synthesized speech is obtained through a speech synthesis model, and the speech synthesis model is trained based on sample data. The sample data includes sample speeches marked with sample speech attributes and sample texts corresponding to the sample speeches. The global prosody feature is obtained based on the sample global prosody feature of a reference sample speech, and the reference sample speech is a sample speech with a first speech attribute. That is, the global prosody feature can be obtained from the sample data used to train the synthesis model, so it is beneficial to improve the convenience of extracting the global prosody feature.

[0061] In some disclosed embodiments, the global prosody module 102 includes a first acquisition sub-module for fusing the sample global prosody features of each reference sample speech to obtain a global prosody feature; the global prosody module 102 includes a second acquisition sub-module for using the sample global prosody feature of a first target sample speech as the global prosody feature; the global prosody module 102 includes a third acquisition sub-module for using the sample global prosody feature of a second target sample speech as the global prosody feature; wherein, the first target sample speech and the second target sample speech are both selected from the reference sample speeches, and the text similarity between the sample text corresponding to the first target sample speech and the text to be synthesized satisfies a first condition, and the presentation intensity of the second target sample speech with respect to the first speech attribute satisfies a second condition.

[0062] Therefore, by fusing the sample global prosody features of each reference sample speech to obtain a global prosody feature, or using the sample global prosody feature of a first target sample speech as the global prosody feature, or using the sample global prosody feature of a second target sample speech as the global prosody feature, and the first target sample speech and the second target sample speech are both selected from the reference sample speeches, and the text similarity between the sample text corresponding to the first target sample speech and the text to be synthesized satisfies a first condition, and the presentation intensity of the second target sample speech with respect to the first speech attribute satisfies a second condition, the global prosody feature with the first speech attribute can be obtained in multiple ways, which is beneficial to further improving the adaptability to different scenarios and expanding the application scope of speech synthesis.

[0063] In some disclosed embodiments, the speech synthesis model includes a global prosody extraction network, and the global prosody extraction network is used to extract sample global prosody features. The sample speech attributes at least include a sample first speech attribute. The speech synthesis device 100 includes a global extraction network training module. The global extraction network training module includes a first sample global prosody extraction sub-module for extracting the first sample global prosody feature of the sample speech; the global extraction network training module includes a first sample synthesis sub-module for synthesizing a first sample synthesized speech based on the first sample global prosody feature and the sample text corresponding to the sample speech; the global extraction network training module includes a first prediction sub-module for predicting, based on the first sample global prosody feature, a predicted first speech attribute of the sample speech; the global extraction network training module includes a second sample global prosody extraction sub-module for extracting the second sample global prosody feature of the first sample synthesized speech; the global extraction network training module includes a first adjustment sub-module for at least adjusting the network parameters of the global prosody extraction network based on the difference between the predicted first speech attribute and the sample first speech attribute, and the difference between the first sample global prosody feature and the second sample global prosody feature.

[0064] Therefore, the speech synthesis model includes a global prosody extraction network, and the global prosody extraction network is used to extract sample global prosody features. The sample speech attributes at least include a sample first speech attribute. Based on this, the first sample global prosody feature of the sample speech is extracted, and then a first sample synthesized speech is synthesized based on the first sample global prosody feature and the sample text corresponding to the sample speech. And a predicted first speech attribute of the sample speech is obtained by prediction based on the first sample global prosody feature, and the second sample global prosody feature of the first sample synthesized speech is extracted. On this basis, at least the network parameters of the global prosody extraction network are adjusted based on the difference between the predicted first speech attribute and the sample first speech attribute, and the difference between the first sample global prosody feature and the second sample global prosody feature. Therefore, it is possible to decouple the global prosody extraction network from information such as the speaker and text during the feature extraction process through the classification loss constraint of the speech attribute and the distribution loss constraint of the prosody feature, so that the extracted global prosody features only represent feature information such as emotion and style, and improve the accuracy of the global prosody features.

[0065] In some disclosed embodiments, the synthesized speech is synthesized based on a speech synthesis model, and the speech synthesis model includes a local prosody prediction network. The local prosody prediction network is used to directly predict local prosody features based on the text to be synthesized, the first speech attribute, and the second speech attribute.

[0066] Therefore, directly predicting the text to be synthesized, the first speech attribute, and the second speech attribute through the local prosody prediction network to obtain local prosody features is beneficial to improving the convenience of obtaining local prosody features.

[0067] In some disclosed embodiments, the speech synthesis model further includes a local prosody extraction network. The speech synthesis model is trained based on sample data, where the sample data includes sample speech labeled with sample speech attributes and sample text corresponding to the sample speech, and the local prosody prediction network is trained based on the sample data after the local prosody extraction network converges.

[0068] Therefore, the speech synthesis model further includes a local prosody extraction network. The speech synthesis model is trained based on sample data, where the sample data includes sample speech labeled with sample speech attributes and sample text corresponding to the sample speech, and the local prosody prediction network is trained based on the sample data after the local prosody extraction network converges. It can further train the local prosody prediction network after the local prosody extraction network is trained to assist in the training of the local prosody prediction network, which is beneficial to improving the accuracy of the local prosody prediction network.

[0069] In some disclosed embodiments, the sample speech attributes at least include sample second speech attributes. The speech synthesis device 100 includes a local extraction network training module. The local extraction network training module includes a first sample local prosody extraction sub-module for extracting the first sample local prosody features of the sample speech. The local extraction network training module includes a first text feature extraction sub-module for extracting the first text features of the sample text. The local extraction network training module includes a second prediction sub-module for predicting based on the first sample local prosody features to obtain the predicted second speech attributes of the sample speech. The local extraction network training module includes a second sample synthesis sub-module for synthesizing a second sample synthesis speech based on the first sample local prosody features and the sample text corresponding to the sample speech. The local extraction network training module includes a second sample local prosody extraction sub-module for extracting the second sample local prosody features of the second sample synthesis speech. The local extraction network training module includes a second adjustment sub-module for at least adjusting the network parameters of the local prosody extraction network based on the differences between the predicted second speech attributes and the sample second speech attributes, the differences between the first sample local prosody features and the first text features, and the differences between the first sample local prosody features and the second sample local prosody features.

[0070] Therefore, the sample voice attributes at least include the sample second voice attributes. During the training process, the first sample local prosody features of the sample voice are extracted, and the first text features of the sample text are extracted. Then, based on the first sample local prosody features, predictions are made to obtain the predicted second voice attributes of the sample voice. Based on the first sample local prosody features and the sample text corresponding to the sample voice, a second sample synthesized voice is synthesized, and the second sample local prosody features of the second sample synthesized voice are extracted. On this basis, based on the differences between the predicted second voice attributes and the sample second voice attributes, the differences between the first sample local prosody features and the first text features, and the differences between the first sample local prosody features and the second sample local prosody features, at least the network parameters of the local prosody extraction network are adjusted, which can decouple the local prosody features extracted by the local prosody extraction network from the speaker information and text information, and then make it only contain feature information related to prosody information such as fundamental frequency fluctuations and pronunciation durations, which is beneficial to improving the accuracy of the local prosody features.

[0071] In some disclosed embodiments, the voice synthesis device 100 includes a local prediction network training module. The local prediction network training module includes a third sample local prosody extraction sub-module for extracting the third sample local prosody features of the sample voice based on the local prosody extraction network. The local prediction network training module includes a local prosody prediction sub-module for predicting the predicted local prosody features based on the local prosody prediction network for the sample text corresponding to the sample voice and the sample voice attributes annotated for the sample voice. The local prediction network training module includes a third adjustment sub-module for adjusting the network parameters of the local prosody prediction network based on the differences between the third sample local prosody features and the predicted local prosody features.

[0072] Therefore, based on the local prosody extraction network, the third sample local prosody features of the sample voice are extracted, and based on the local prosody prediction network, predictions are made for the sample text corresponding to the sample voice and the sample voice attributes annotated for the sample voice to obtain the predicted local prosody features. On this basis, based on the differences between the third sample local prosody features and the predicted local prosody features, the network parameters of the local prosody prediction network are adjusted, which is beneficial to improving the accuracy of the local prosody prediction network.

[0073] In some disclosed embodiments, the synthesized voice is synthesized based on a voice synthesis model, and the voice synthesis model includes a prosody encoding network and a text encoding network. The text encoding network is used to predict voice attribute features based on the text to be synthesized and the first voice attributes, and the voice attribute features are independent of the speaker. The prosody encoding network is used to predict local prosody features based on the voice attribute features, the text to be synthesized, and the second voice attributes.

[0074] Therefore, by using the text encoding network to predict the speech attribute features for the text to be synthesized and the first speech attribute, and by using the prosody encoding network to predict the local prosody features for the speech attribute features, the text to be synthesized, and the second speech attribute, and the speech attribute features are independent of the speaker, so that the prosody encoding network can further decouple speech attributes such as emotion, style, and speaker, which is beneficial to improving the control over emotion and style and enhancing the distinguishability of emotion and style.

[0075] In some disclosed embodiments, the speech synthesis model further includes a local prosody extraction network. The speech synthesis model is trained based on sample data, where the sample data includes sample speech annotated with sample speech attributes and sample text corresponding to the sample speech, and the prosody encoding network and the text encoding network are trained based on the sample data after the local prosody extraction network converges.

[0076] Therefore, after the local prosody extraction network is trained, the prosody encoding network and the text encoding network can be further trained to assist the training of the prosody encoding network and the text encoding network through the trained local prosody extraction network, which is beneficial to improving the accuracy of the prosody encoding network and the text encoding network.

[0077] In some disclosed embodiments, the sample speech attributes include sample first speech attributes and sample second speech attributes. The speech synthesis device 100 includes an encoding network training module, and the encoding network training module includes a fourth sample local prosody extraction sub-module for extracting the fourth sample local prosody features of the sample speech; the encoding network training module includes a sample speech attribute extraction sub-module for extracting the sample speech attribute features of the fourth sample local prosody features; the encoding network training module includes a speech attribute prediction sub-module for predicting the predicted speech attribute features based on the sample text and the sample first speech attributes; the encoding network training module includes a fifth sample local prosody prediction sub-module for predicting the fifth sample local prosody features based on the sample text, the sample second speech attributes, and the predicted speech attribute features; the encoding network training module includes a fourth adjustment sub-module for adjusting the network parameters of the prosody encoding network and the text encoding network based on the difference between the sample speech attribute features and the predicted speech attribute features, and the difference between the fourth sample local prosody features and the fifth sample local prosody features.

[0078] Therefore, the fourth sample local prosody feature of the sample speech is extracted. Based on this, the sample speech attribute feature of the fourth sample local prosody feature is further extracted. And based on the sample text and the sample first speech attribute, the predicted speech attribute feature is obtained. Thus, based on the sample text, the sample second speech attribute and the predicted speech attribute feature, the fifth sample local prosody feature is predicted. On this basis, further based on the differences between the sample speech attribute feature and the predicted speech attribute feature, and between the fourth sample local prosody feature and the fifth sample local prosody feature, the network parameters of the prosody encoding network and the text encoding network are adjusted, which can further decouple speech attributes such as emotion, style, and speaker, thereby improving the control ability of emotion and style and being beneficial to enhancing the distinguishability of emotion and style.

[0079] In some disclosed embodiments, the synthesized speech is synthesized based on a speech synthesis model. The speech synthesis model includes a global prosody extraction network, a local prosody extraction network, and a synthesis network. The speech synthesis model is trained based on sample data. The sample data includes sample speech with labeled sample speech attributes and the sample text corresponding to the sample speech. And the synthesis network is trained based on the sample data after the global prosody extraction network and the local prosody extraction network are respectively trained to convergence.

[0080] Therefore, after the global prosody extraction network and the local prosody extraction network are respectively trained to convergence, further training the synthesis network is beneficial to improving the speech synthesis effect.

[0081] In some disclosed embodiments, the speech synthesis device 100 includes a synthesis network training module. The synthesis network training module includes a sample prosody extraction sub-module for extracting the third sample global prosody feature and the sixth sample local prosody feature of the sample speech; the network training module includes a sample synthesis sub-module for synthesizing based on the sample text corresponding to the sample speech, the third sample global prosody feature and the sixth sample local prosody feature to obtain the sample synthesized speech; the synthesis network training module includes a fifth adjustment sub-module for at least adjusting the network parameters of the synthesis network based on the difference between the sample speech and the sample synthesized speech.

[0082] Therefore, by decoupling the global prosody feature, the local prosody feature and the speaker, the prosody of a certain speaker can be transferred to a speaker with a different prosody, so that even a speaker without emotional data can have different emotional expressions of an emotional speaker, greatly reducing the demand for high-quality emotion and style data.

[0083] Please refer to Figure 11 , Figure 11It is a schematic diagram of the framework of an embodiment of the electronic device 110 of the present application. The electronic device 110 includes a memory 111 and a processor 112 that are coupled to each other. Program instructions are stored in the memory 111, and the processor 112 is configured to execute the program instructions to implement the steps in any of the above-described embodiments of the speech synthesis method. Specifically, the electronic device 110 may include, but is not limited to: a desktop computer, a laptop computer, a server, a mobile phone, a tablet computer, etc., which are not limited herein.

[0084] Specifically, the processor 112 is configured to control itself and the memory 111 to implement the steps in any of the above-described embodiments of the speech synthesis method. The processor 112 may also be referred to as a CPU (Central Processing Unit). The processor 112 may be an integrated circuit chip with signal processing capabilities. The processor 112 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Additionally, the processor 112 may be implemented jointly by integrated circuit chips.

[0085] In the above solution, on the one hand, by referring to the sentence-level prosody feature information during the synthesis process, the overall synthesized speech can be controlled to conform to the first speech attribute. On the other hand, by further referring to the word-level prosody features during the synthesis process, the subtle changes in the local prosody of the synthesized speech can be controlled. Therefore, it is possible to accurately and freely synthesize speeches with different prosodies, improving the adaptability to different scenarios.

[0086] Please refer to Figure 12 , Figure 12 It is a schematic diagram of the framework of an embodiment of the computer-readable storage medium 120 of the present application. The computer-readable storage medium 120 stores program instructions 121 that can be run by a processor, and the program instructions 121 are used to implement the steps in any of the above-described embodiments of the speech synthesis method.

[0087] In the above solution, on the one hand, by referring to the sentence-level prosody feature information during the synthesis process, the overall synthesized speech can be controlled to conform to the first speech attribute. On the other hand, by further referring to the word-level prosody features during the synthesis process, the subtle changes in the local prosody of the synthesized speech can be controlled. Therefore, it is possible to accurately and freely synthesize speeches with different prosodies, improving the adaptability to different scenarios.

[0088] In some embodiments, the functions or modules included in the apparatus provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0089] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.

[0090] In several embodiments provided in the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0091] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0092] In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0093] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

Claims

1. A speech synthesis method, characterized in that, it includes: obtaining the text to be synthesized, a first speech attribute, and a second speech attribute; wherein, the first speech attribute includes at least one of an emotion category and a style category, and the second speech attribute includes a speaker identifier; obtaining a global prosody feature with the first speech attribute, and making a prediction based on the text to be synthesized, the first speech attribute, and the second speech attribute to obtain a local prosody feature; wherein, the global prosody feature contains sentence-level prosody feature information, and the local prosody feature contains word-level prosody feature information; performing synthesis based on the text to be synthesized, the global prosody feature, and the local prosody feature to obtain a synthesized speech.

2. The method according to claim 1, characterized in that, the synthesized speech is obtained by synthesizing based on a speech synthesis model, the speech synthesis model is trained based on sample data, the sample data includes sample speeches labeled with sample speech attributes and the sample texts corresponding to the sample speeches, and the global prosody feature is obtained based on the sample global prosody feature of a reference sample speech, and the reference sample speech is a sample speech with the first speech attribute.

3. The method according to claim 2, characterized in that, the obtaining of the global prosody feature with the first speech attribute, includes any one of the following: fusing the sample global prosody features of each of the reference sample speeches to obtain the global prosody feature; using the sample global prosody feature of a first target sample speech as the global prosody feature; using the sample global prosody feature of a second target sample speech as the global prosody feature; wherein, the first target sample speech and the second target sample speech are both selected from the reference sample speeches, and the text similarity between the sample text corresponding to the first target sample speech and the text to be synthesized satisfies a first condition, and the presentation intensity of the second target sample speech with respect to the first speech attribute satisfies a second condition.

4. The method according to claim 2, characterized in that, the speech synthesis model includes a global prosody extraction network, and the global prosody extraction network is used to extract sample global prosody features, the sample speech attributes at least include sample first speech attributes, and the training steps of the global prosody extraction network include: extracting the first sample global prosody feature of the sample speech; synthesizing a first sample synthesized speech based on the first sample global prosody feature and the sample text corresponding to the sample speech, making a prediction based on the first sample global prosody feature to obtain the predicted first speech attribute of the sample speech, and extracting the second sample global prosody feature of the first sample synthesized speech; at least adjusting the network parameters of the global prosody extraction network based on the difference between the predicted first speech attribute and the sample first speech attribute, and the difference between the first sample global prosody feature and the second sample global prosody feature.

5. The method according to claim 1, characterized in that, The synthesized speech is synthesized based on a speech synthesis model, and the speech synthesis model includes a local prosody prediction network, which is used to directly predict the local prosody features based on the text to be synthesized, the first speech attribute, and the second speech attribute.

6. The method according to claim 5, wherein, the speech synthesis model further includes a local prosody extraction network, the speech synthesis model is trained based on sample data, the sample data includes sample speech labeled with sample speech attributes and the sample text corresponding to the sample speech, and the local prosody prediction network is trained based on the sample data after the local prosody extraction network converges.

7. The method according to claim 6, wherein, the sample speech attributes at least include sample second speech attributes, and the training steps of the local prosody extraction network include: extracting the first sample local prosody features of the sample speech, and extracting the first text features of the sample text; predicting based on the first sample local prosody features to obtain the predicted second speech attributes of the sample speech, synthesizing a second sample synthesized speech based on the first sample local prosody features and the sample text corresponding to the sample speech, and extracting the second sample local prosody features of the second sample synthesized speech; adjusting at least the network parameters of the local prosody extraction network based on the difference between the predicted second speech attributes and the sample second speech attributes, the difference between the first sample local prosody features and the first text features, and the difference between the first sample local prosody features and the second sample local prosody features.

8. The method according to claim 6, wherein, the training steps of the local prosody prediction network include: extracting the third sample local prosody features of the sample speech based on the local prosody extraction network, and predicting the sample text corresponding to the sample speech and the sample speech attributes labeled on the sample speech based on the local prosody prediction network to obtain predicted local prosody features; adjusting the network parameters of the local prosody prediction network based on the difference between the third sample local prosody features and the predicted local prosody features.

9. The method according to claim 1, wherein, the synthesized speech is synthesized based on a speech synthesis model, and the speech synthesis model includes a prosody encoding network and a text encoding network, the text encoding network is used to predict speech attribute features based on the text to be synthesized and the first speech attribute, the speech attribute features are independent of the speaker, and the prosody encoding network is used to predict the local prosody features based on the speech attribute features, the text to be synthesized, and the second speech attribute.

10. The method according to claim 9, wherein, The speech synthesis model further includes a local prosody extraction network. The speech synthesis model is trained based on sample data, where the sample data includes sample speeches labeled with sample speech attributes and sample texts corresponding to the sample speeches, and the prosody encoding network and the text encoding network are trained based on the sample data after the local prosody extraction network converges.

11. The method according to claim 10, wherein, the sample speech attributes include sample first speech attributes and sample second speech attributes, and the training steps of the prosody encoding network and the text encoding network include: extracting a fourth sample local prosody feature of the sample speech; extracting sample speech attribute features of the fourth sample local prosody feature, and predicting predicted speech attribute features based on the sample text and the sample first speech attributes; predicting a fifth sample local prosody feature based on the sample text, the sample second speech attributes, and the predicted speech attribute features; adjusting the network parameters of the prosody encoding network and the text encoding network based on the difference between the sample speech attribute features and the predicted speech attribute features, and the difference between the fourth sample local prosody feature and the fifth sample local prosody feature.

12. The method according to claim 1, wherein, the synthesized speech is synthesized based on a speech synthesis model, and the speech synthesis model includes a global prosody extraction network, a local prosody extraction network, and a synthesis network. The speech synthesis model is trained based on sample data, where the sample data includes sample speeches labeled with sample speech attributes and sample texts corresponding to the sample speeches, and the synthesis network is trained based on the sample data after the global prosody extraction network and the local prosody extraction network converge respectively.

13. The method according to claim 12, wherein, the training steps of the synthesis network include: extracting a third sample global prosody feature and a sixth sample local prosody feature of the sample speech; performing synthesis based on the sample text corresponding to the sample speech, the third sample global prosody feature, and the sixth sample local prosody feature to obtain a sample synthesized speech; at least adjusting the network parameters of the synthesis network based on the difference between the sample speech and the sample synthesized speech.

14. A speech synthesis device, wherein, it includes: an acquisition module, configured to acquire a text to be synthesized, a first speech attribute, and a second speech attribute; wherein, the first speech attribute includes at least one of an emotion category and a style category, and the second speech attribute includes a speaker identifier; a global prosody module and a local prosody module, where the global prosody module is configured to acquire global prosody features with the first speech attribute, and the local prosody module is configured to perform prediction based on the text to be synthesized, the first speech attribute, and the second speech attribute to obtain local prosody features; wherein, the global prosody features include sentence-level prosody feature information, and the local prosody features include word-level prosody feature information; A synthesis module, configured to perform synthesis based on the text to be synthesized, the global prosody features, and the local prosody features, to obtain synthesized speech.

15. An electronic device, characterized in that it includes a memory and a processor that are coupled to each other, program instructions are stored in the memory, and the processor is configured to execute the program instructions to implement the speech synthesis method according to any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that it stores program instructions that can be run by a processor, and the program instructions are used to implement the speech synthesis method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Text information processing method and device, computer equipment and readable storage medium

    CN111274807A

  • Speech synthesis method, electronic equipment and storage device

    CN112786004A