Speech synthesis method, device, computer equipment and storage medium

By obtaining the target emotional categories and intensity, and using emotional tokens and text features to synthesize speech, the problem of insufficient emotional expression in the prior art is solved, emotional speech synthesis is achieved, and the flexibility of auditory effects and speech synthesis is improved.

CN115376486BActive Publication Date: 2025-08-22TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211016347.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-23
Publication Date
2025-08-22
Estimated Expiration
2042-08-23

AI Technical Summary

Technical Problem

The existing speech synthesis technology lacks emotional expression, resulting in poor auditory effects of synthetic audio.

Method used

By obtaining the target emotional categories and intensity of the text to be synthesized, using the emotional token and text features, determine the emotional characteristics, and synthesize emotional voice.

Benefits of technology

It realizes emotional speech synthesis, improves user auditory effects, and enhances the flexibility and controllability of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115376486B_ABST
    Figure CN115376486B_ABST
Patent Text Reader

Abstract

The present application discloses a speech synthesis method, apparatus, computer device, and storage medium. The method includes the following steps: obtaining a target emotion category and target emotion intensity corresponding to a text to be synthesized; determining the emotion characteristics of the text to be synthesized based on emotion tokens, target emotion categories, and target emotion intensities corresponding to a plurality of preset emotion categories; using any emotion token to represent the characteristics of the corresponding preset emotion category; determining the emotion text characteristics of the text to be synthesized based on the text characteristics and emotion characteristics of the text to be synthesized; and synthesizing an emotional speech that conforms to the target emotion category and target emotion intensity of the text to be synthesized based on the emotion text characteristics. This method can achieve the synthesis of emotionally rich speech, thereby improving the user's auditory experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech technology, and in particular to a speech synthesis method, apparatus, computer equipment, and storage medium. Background Art

[0002] Current speech synthesis technologies typically focus solely on synchronizing text and speech content, which can result in a stilted and poorly perceived audio quality. However, emotion, in addition to the text itself, reflects a person's attitude toward something (such as inner feelings and intentions). Therefore, synthesizing emotionally charged speech to enhance auditory quality has become a hot topic in current speech synthesis research. Summary of the Invention

[0003] The embodiments of the present application provide a speech synthesis method, apparatus, computer device and storage medium, which can realize the synthesis of emotional speech to improve the user's auditory effect.

[0004] The first aspect of the embodiments of the present application discloses a speech synthesis method, the method comprising:

[0005] Obtain the target sentiment category and target sentiment intensity corresponding to the text to be synthesized;

[0006] Determining the emotional features of the text to be synthesized based on the emotional tokens corresponding to the plurality of preset emotional categories, the target emotional category, and the target emotional intensity; any emotional token is used to represent the features of the corresponding preset emotional category;

[0007] Determining the emotional text features of the text to be synthesized based on the text features of the text to be synthesized and the emotional features;

[0008] The emotional speech of the text to be synthesized conforming to the target emotional category and the target emotional intensity is synthesized according to the emotional text features.

[0009] A second aspect of an embodiment of the present application discloses a speech synthesis device, comprising:

[0010] An acquisition unit, used to acquire a target emotion category and a target emotion intensity corresponding to the text to be synthesized;

[0011] A first determining unit is configured to determine the emotional features of the text to be synthesized based on the emotional tokens corresponding to a plurality of preset emotional categories, the target emotional category, and the target emotional intensity; any emotional token is used to represent the features of the corresponding preset emotional category;

[0012] A second determining unit is configured to determine the emotional text feature of the text to be synthesized based on the text feature of the text to be synthesized and the emotional feature;

[0013] The synthesis unit is used to synthesize the emotional speech of the text to be synthesized in accordance with the target emotional category and the target emotional intensity according to the emotional text features.

[0014] A third aspect of an embodiment of the present application discloses a computer device, comprising a processor, a memory, and a network interface, wherein the processor, the memory, and the network interface are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method of the first aspect above.

[0015] A fourth aspect of an embodiment of the present application discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method of the first aspect.

[0016] A fifth aspect of the present application discloses a computer program product or computer program, comprising program instructions stored in a computer-readable storage medium. A processor of a computer device reads the program instructions from the computer-readable storage medium and executes the program instructions, causing the computer device to perform the method of the first aspect.

[0017] In an embodiment of the present application, a computer device can obtain a target emotion category and a target emotion intensity corresponding to a text to be synthesized; and determine the emotion features of the text to be synthesized based on the emotion tokens, target emotion categories, and target emotion intensity corresponding to a plurality of preset emotion categories; any emotion token is used to characterize the features of the corresponding preset emotion category, and further, the emotion text features of the text to be synthesized can be determined based on the text features and emotion features of the text to be synthesized, so that the text to be synthesized can be synthesized according to the emotion text features to conform to the target emotion category and the target emotion intensity. By implementing the above method, the emotion features corresponding to the text to be synthesized can be obtained by specifying the emotion category and emotion intensity, and the emotion text features rich in text information and emotion information can be further obtained by combining the text features of the text to be synthesized, so as to synthesize a synthesized speech of a specified emotion category and emotion intensity through the emotion text features, thereby realizing the synthesis of emotion-rich speech, so as to improve the user's auditory effect, and also improve the flexibility and controllability of the synthesized speech. In addition, when determining the emotion features, the emotion tokens corresponding to the various emotion categories can also be used to characterize the emotion, so as to enhance the controllability and interpretability of the emotion representation. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1a This is a structural diagram of a speech synthesis scenario provided by an embodiment of the present application;

[0020] Figure 1b This is a schematic diagram of the architecture of a speech synthesis system provided in an embodiment of the present application;

[0021] Figure 2 This is a flow chart of a speech synthesis method provided in an embodiment of the present application;

[0022] Figure 3a This is a schematic diagram of a structure for obtaining a target emotion category and a target emotion intensity provided by an embodiment of the present application;

[0023] Figure 3b This is a schematic diagram of the structure of an emotional speech synthesis model provided in an embodiment of the present application;

[0024] Figure 3c This is a schematic diagram of the structure of an emotion control network provided by an embodiment of the present application;

[0025] Figure 3d This is a schematic diagram of the structure of another emotional speech synthesis model provided in an embodiment of the present application;

[0026] Figure 4 This is a flowchart of another speech synthesis method provided in an embodiment of the present application;

[0027] Figure 5a This is a schematic diagram of the structure of a reference model provided in an embodiment of the present application;

[0028] Figure 5b This is a schematic diagram of the structure of an emotion extraction network provided in an embodiment of the present application;

[0029] Figure 5c This is a schematic diagram of the structure of an emotion representation network provided in an embodiment of the present application;

[0030] Figure 5d is a structural diagram of another reference model provided in an embodiment of the present application;

[0031] Figure 6 This is a structural diagram of a speech synthesis device provided in an embodiment of the present application;

[0032] Figure 7It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application.

[0034] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0035] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0036] Key technologies in speech technology include automatic speech recognition (ASR), text-to-speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech becoming one of the most promising methods of human-computer interaction.

[0037] Natural language processing (NLP) is a key area of ​​research in computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. Natural language processing (NLP) integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language we use in everyday life—and is closely linked to the study of linguistics. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robotic question answering, and knowledge graphs.

[0038] Based on the speech technology and machine learning technologies mentioned in the above-mentioned artificial intelligence technologies, this application proposes an emotional speech synthesis model (i.e., a model for emotionally controlling text and obtaining corresponding emotional speech). This emotional speech synthesis model can emotionally control the text to be synthesized based on the emotion category and emotion intensity to obtain the corresponding emotional speech. Furthermore, an embodiment of this application also proposes a speech synthesis solution based on this emotional speech synthesis model. Specifically, the general principle of this solution is as follows: for the text to be synthesized, the target emotion category and target emotion intensity corresponding to the text to be synthesized can be first obtained; then, based on the target emotion category and target emotion intensity, the text to be synthesized can be emotionally controlled to obtain the corresponding emotional speech. For example, the emotional features of the text to be synthesized can be determined based on the emotion tokens corresponding to multiple preset emotion categories, the target emotion category, and the target emotion intensity. Any emotion token is used to represent the characteristics of the preset emotion category. The text features of the text to be synthesized can also be further obtained to obtain the emotional speech corresponding to the text to be synthesized under the target emotion category and target emotion intensity based on the emotion features and text features. For example, the emotional text features of the text to be synthesized can be determined, and the corresponding emotional speech can be determined based on the emotional text features.

[0039] In one implementation, the above-mentioned speech synthesis scheme can be implemented by calling an end-to-end emotional speech synthesis model, such as directly accepting the text to be synthesized, the target emotional category and the target emotional intensity as input to obtain the corresponding emotional speech through the emotional speech synthesis model.

[0040] In summary, the speech synthesis scheme proposed in the embodiment of the present application may have the following beneficial effects: the emotional features of the text to be synthesized can be obtained through the specified emotional category and emotional intensity, and the emotional text features rich in text information and emotional information can be further obtained in combination with the text features of the text to be synthesized, so as to synthesize a synthetic speech of the specified emotional category and emotional intensity through the emotional text features, thereby realizing the synthesis of emotional speech to improve the user's auditory effect, and also to improve the flexibility and controllability of the synthesized speech. In addition, when determining the emotional features, the emotion tokens corresponding to various emotion categories can also be used to represent the emotion to enhance the controllability and interpretability of the emotion representation. In addition, by calling the emotional speech synthesis model to realize the synthesis of emotional speech, the automation and intelligence of speech synthesis can also be improved.

[0041] The embodiments of the present application can be applied to various application scenarios such as human-computer interaction, speech synthesis, voice interaction, voice assistant, and audio novel generation, so as to utilize the above-mentioned speech synthesis scheme to realize the simultaneous control of emotion category and emotion intensity, obtain emotion-rich speech, and thus improve user experience. For example, in the audio novel generation scenario, the user can enter the desired emotion category and emotion intensity in the audio novel application. After the audio novel application receives the emotion category and emotion intensity entered by the user, the novel can be read aloud with a specific emotion category and emotion intensity based on the novel text and the input emotion category and emotion intensity. For another example, in application scenarios such as human-computer interaction, voice interaction, and voice assistant, the user can enter the desired emotion category and emotion intensity on the relevant interface in advance. In subsequent interactions, the user can hear the emotion-rich speech that he is interested in, thereby increasing the user's auditory experience.

[0042] In the application scenarios mentioned above, users only need to provide the text to be synthesized, the emotion category and the emotion intensity as input. The end-to-end emotion-controllable speech synthesis model can directly accept the text to be synthesized, the emotion category and the emotion intensity as input and output the synthesized speech with the specified emotion category and intensity. Figure 1a As shown, if the user wants to synthesize a voice with the text "Happy New Year!", the emotion category "happy", and the emotion intensity "0.9", the user can directly enter the above three items respectively in the relevant interface of the application corresponding to the application scenario. When the application receives the three items, the application can automatically return the corresponding emotional voice. Specifically, the application can obtain the corresponding emotional voice according to the voice synthesis scheme proposed in this application, and broadcast the emotional voice.

[0043] In a specific implementation, the execution subject of the above-mentioned speech synthesis solution can be a computer device, which includes but is not limited to a terminal or a server. In other words, the computer device can be a server or a terminal, or a system composed of a server and a terminal. Among them, the terminal mentioned above can be an electronic device, including but not limited to mobile phones, tablet computers, desktop computers, laptop computers, PDAs, vehicle-mounted devices, intelligent voice interaction devices, augmented reality / virtual reality (AR / VR) devices, helmet displays, wearable devices, smart speakers, smart home appliances, aircraft, digital cameras, cameras and other mobile Internet devices (Mobile Internet Device, MID) with network access capabilities. Among them, the server mentioned above can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, vehicle-road collaboration, content delivery networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms.

[0044] In one implementation, when the computer device is a server, the embodiment of the present application provides a speech synthesis system, such as Figure 1b As shown, the speech synthesis system may include at least one terminal and at least one server; taking the terminal as an example, the user can input the target emotion category and target emotion intensity corresponding to the text to be synthesized on the terminal interface, and the terminal can obtain the target emotion category and target emotion intensity. After the terminal obtains the target emotion category and target emotion intensity, the target emotion category and target emotion intensity can be uploaded to the server, so that the server synthesizes the emotion speech of the text to be synthesized according to the obtained target emotion category and target emotion intensity to obtain the corresponding emotion speech. After the emotion speech is synthesized, the server can also return the emotion speech to the terminal so that the emotion speech can be voice broadcast on the terminal interface.

[0045] Based on the speech synthesis solution provided above, the embodiment of the present application provides a speech synthesis method, which can be executed by the computer device mentioned above. Figure 2 The speech synthesis method includes but is not limited to the following steps:

[0046] S201: Obtain the target emotion category and target emotion intensity corresponding to the text to be synthesized.

[0047] The text to be synthesized can be any text, for example, a novel, news, or user conversation text. The target emotion category can be one of one or more preset emotion categories, for example, happiness, anger, sadness, surprise, fear, disgust, etc., which are not described here one by one. The target emotion category can be any of these preset emotion categories, such as happiness or fear. The target emotion intensity can be a representation of the emotion intensity corresponding to the target emotion category. The emotion intensity can be represented by a numerical value. For example, the emotion intensity can range from 0 to 1, where a larger value corresponds to a stronger emotion intensity, and a smaller value corresponds to a weaker emotion intensity. That is, the target emotion intensity can be any value between 0 and 1, such as 0.4 or 0.9. For example, assuming the target emotion category is happiness, a target emotion intensity of 0.9 indicates a high intensity of happiness, while a target emotion intensity of 0.1 indicates a low intensity of happiness.

[0048] In one implementation, when there is a demand for speech synthesis, the target emotion category and target emotion intensity corresponding to the text to be synthesized can be obtained.

[0049] Optionally, the presence of a speech synthesis requirement may be determined when the computer device receives a speech synthesis request. For example, a user may send a speech synthesis request for a speech to be synthesized to the computer device, causing the computer device to receive the speech synthesis request. After the computer device receives the speech synthesis request, it is determined that there is a speech synthesis requirement for the speech to be synthesized. In one possible implementation, when a user needs to perform speech synthesis on a text, the user may perform relevant operations on a user operation interface output by a terminal to send a speech synthesis request for the text to be synthesized to the computer device.

[0050] See for example Figure 3a As shown: The terminal used by the user can display a user operation interface on the terminal screen, and the user operation interface can include at least a data setting area marked by 301 and a confirmation control marked by 302, wherein the data setting area can include a text setting area, an emotion category setting area, and an emotion intensity setting area. The text setting area can be used to enter text information. For example, a text can be directly entered in the text setting area, or the storage area address where the text is located can be entered. The storage area address is used to link to the corresponding text; the emotion category setting area can be used to enter a specific emotion category; the emotion intensity setting area can be used to enter a specific emotion intensity.

[0051] If the user wants to perform speech synthesis on a certain text, the user can enter relevant information in the data setting area 301, such as Figure 3a The text displayed in the text is "Happy New Year", the emotion category is "Happy", and the emotion intensity is "0.9"; then, the user can perform a trigger operation (such as a click operation, a press operation, etc.) on the confirmation control 302, thereby triggering the terminal used by the user to obtain input data for speech synthesis from the data setting area 301. The input data may include the text to be synthesized, the target emotion category and the target emotion intensity. These data are the data set in the data setting area 301, such as the text to be synthesized is "Happy New Year", the target emotion category is "Happy", and the target emotion intensity is "0.9". After the terminal obtains these input data, it can send a speech synthesis request for the input data to the computer device, and the speech synthesis request may include the input data.

[0052] Optionally, the demand for speech synthesis may also be generated by triggering a speech synthesis timer task. For example, a speech synthesis timer task may be set, which indicates the triggering conditions for speech synthesis of the text to be synthesized. The text to be synthesized may be pre-stored in a certain storage area. When the speech synthesis timer task is triggered, the text to be synthesized may be directly obtained from the storage area, and the speech synthesis operation may be performed. The target emotion category and target emotion intensity of the text to be synthesized may be associated with the corresponding text to be synthesized and stored. When the text to be synthesized is obtained, the corresponding target emotion category and target emotion intensity may also be obtained. The triggering condition may be that the current time reaches the preset speech synthesis time; or the remaining storage space of the storage area reaches the preset remaining storage space; or a new text to be synthesized is added to the storage area, and so on.

[0053] S202: Determine the emotional features of the text to be synthesized according to the emotional tokens, target emotional categories, and target emotional intensities corresponding to a plurality of preset emotional categories.

[0054] Each preset emotion category corresponds to an emotion token, and any emotion token can be used to characterize the characteristics of the corresponding preset emotion category. The emotion tokens corresponding to the multiple preset emotion categories are obtained during the training of the emotion speech synthesis model. The method for obtaining the emotion tokens can be found in the relevant description in step S402.

[0055] The number of target emotion categories can be one or more, and correspondingly, the number of target emotion intensities can also be one or more, and one target emotion category can correspond to one target emotion intensity. In the embodiment of the present application, a target emotion category and a target emotion intensity are used as an example for description.

[0056] In a broad sense, when representing the emotional features in the entire emotional space, the emotional tokens corresponding to each preset emotional category in the entire emotional space can be weighted and summed. The weighted summation result is the emotional feature in a broad sense, as shown in the following formula (1).

[0057]

[0058] Where E represents the emotional feature, N represents the total number of preset emotional categories (or the total number of emotional tokens), and T i Represents the emotion token corresponding to the i-th preset emotion category (or directly understood as the i-th emotion token), ω i It represents the token weight of the i-th emotion token, and the token weight of the i-th emotion token is the emotion intensity corresponding to the i-th preset emotion category.

[0059] Based on this, it can be seen that the representation of the above-mentioned emotional features includes both broad emotional categories and broad emotional intensity. In the embodiment of the present application, the purpose is to control the emotional intensity of the text to be synthesized specifically to a specific emotional category. In this case, a baseline emotional category can be set first. For example, the baseline emotional category can be a neutral emotional category or other emotional category. Considering that the neutral emotional category is a category that does not contain any emotion relative to other emotional categories, the embodiment of the present application takes the neutral emotional category as an example of the baseline emotional category for relevant explanation. Assume that the neutral emotional category is the emotional representation when the emotional intensity of all emotional categories is the weakest, that is, the emotional representation with an emotional intensity of 0, and other emotional categories except the neutral emotional category are the emotional representation when the emotional intensity of the emotional category is the strongest, that is, the emotional representation with an emotional intensity of 1. And in the high-dimensional emotional space, the closer to the emotional token of the neutral emotional category, the smaller the emotional intensity; conversely, the closer to the specific emotional token, the stronger the emotional intensity of that category. Here, the specific emotion token is the emotion token corresponding to the target emotion category, that is, the emotion intensity of the emotion token corresponding to the neutral emotion category is 0, and the emotion intensity of the emotion token corresponding to the target emotion category is 1. The embodiment of the present application needs to obtain the emotion representation of the target emotion category under a specific emotion intensity (that is, the target emotion intensity). Then, it can be considered to adopt an interpolation operation between the emotion intensity of a certain specific emotion token (that is, the emotion token corresponding to the target emotion category) and the emotion intensity of the neutral emotion token (that is, the emotion token corresponding to the neutral emotion category), that is, to interpolate between 0 and 1, to determine the emotion characteristics of the text to be synthesized. For example, the obtained target emotion intensity can be used as the token weight of the specific emotion token, and the difference between 1 and the target emotion intensity can be used as the token weight of the neutral emotion token. In this case, the token weights of the emotion tokens corresponding to other preset emotion categories are defaulted to 0. For example, this implementation method of determining the emotion characteristics can be expressed using the following formula (2).

[0060] E=αT hap +(1-α)T neu (2)

[0061] Among them, T hap Indicates the emotion token corresponding to the target emotion category, T neu represents the emotional token corresponding to the neutral emotional category, α represents the token weight of the emotional token corresponding to the target emotional category, and the token weight is the target emotional intensity corresponding to the target emotional category. (1-α) represents the token weight of the emotional token corresponding to the neutral emotional category. For example, taking the target emotional category as happy, then T in formula (2) hapIt can be understood as the emotional token corresponding to the happy emotional category, and α represents the token weight corresponding to the happy emotional category, that is, the target emotional intensity.

[0062] As can be seen from the above, we can use a weighted combination of emotion tokens from multiple preset emotion categories to enhance the controllability and interpretability of emotion representation. In addition, the emotion representation space can be expanded by using neutral emotion as the representation when the emotion intensity is 0 and extreme specific emotion as the representation when the emotion intensity is 1.

[0063] In one implementation, step S202 may be implemented by an emotion control network, which may be a network in an emotion speech synthesis model, such as Figure 3b As shown. That is, in a specific implementation, the emotional features can be obtained by calling the emotional speech synthesis model, which can be obtained by performing emotional control on the emotional control network in the emotional speech synthesis model. The emotional control network can have the weighted logic required in the embodiment of the present application (the weighted logic is also the specific implementation logic for determining the emotional features of the text to be synthesized as shown in the above formula (2)), and the emotional control network is pre-configured with multiple emotional tokens corresponding to preset emotional categories. As shown Figure 3c The diagram below shows the network structure of the emotion control network. Figure 3c It can be seen that the main idea of ​​the emotion control network is to perform weighted combination of multiple different emotion tokens.

[0064] The specific implementation of obtaining the emotional features through the emotional control network can be: inputting the target emotional category and target emotional intensity into the emotional control network in the emotional speech synthesis model, and weighting the emotional tokens corresponding to multiple preset emotional categories based on the target emotional category and target emotional intensity in the emotional control network, thereby obtaining the emotional features of the text to be synthesized. For example, the token weights of the emotional tokens corresponding to multiple preset emotional categories can be determined first, and then the token weights can be used to perform weighted processing with the corresponding emotional tokens to obtain multiple weighted processing results, and the multiple weighted processing results can be summed up, and the summation result is the emotional feature. Among them, the token weights of the emotional tokens corresponding to multiple preset emotional categories can be determined in the following way: matching the target emotional category with multiple preset emotional categories, setting the token weights of the emotional tokens corresponding to the matched preset emotional categories to: target emotional intensity; setting the token weights of the emotional tokens corresponding to the neutral emotional category in the multiple preset emotional categories to: the difference between 1 and the target emotional intensity; and setting the token weights of the emotional tokens corresponding to other preset emotional categories to: 0.

[0065] It can be seen that the emotion control network can artificially adjust the emotion category and emotion intensity of the synthesized speech, and then synthesize speech of any different emotion category and emotion intensity to improve the controllability and flexibility of speech synthesis.

[0066] In the embodiment of the present application, it is possible to control the emotion category and emotion intensity at the sentence level, and it is also possible to control the emotion category and emotion intensity at the phoneme level. For sentence-level control, it is only necessary to regulate the emotion at the sentence level and then expand it to be equal to the length of the phoneme sequence; for phoneme-level control, it is necessary to regulate the emotion corresponding to each phoneme separately. For example, an emotion intensity can be added to each phoneme in the phoneme sequence, and the phonemes are weighted by the emotion intensity to achieve the regulation of the phonemes, wherein the emotion intensity corresponding to each phoneme can be set to gradually change from strong to weak (such as the emotion intensity corresponding to each phoneme is 0.5, 0.3, 0.1, ...), or from weak to strong (such as the emotion intensity corresponding to each phoneme is 0.1, 0.3, 0.5, ...), or a smooth gradient between different emotions (i.e., using different emotion categories as baseline emotions, such as the neutral emotion token in the above formula (2) can be a sad emotion token or a fear emotion token, etc.) and other fine-grained emotion change effects.

[0067] It should be noted that for the emotion control network, the key is to use a variety of emotion tokens to represent the emotion category and emotion intensity. The specific representation method can be linear (such as the weighted processing mentioned above) or nonlinear, such as taking the logarithm or mapping to other spaces, etc. The representation method is not specifically limited.

[0068] S203: Acquire text features of the text to be synthesized, and determine the emotional text features of the text to be synthesized based on the emotional features and the text features.

[0069] In one implementation, the text features are obtained by encoding the phoneme sequence corresponding to the text to be synthesized. Optionally, the text features can be obtained by calling the acoustic network in the emotional speech synthesis model, and the acoustic network can include an encoding network and a decoding network; first, the text to be synthesized can be converted into a phoneme sequence, and after obtaining the phoneme sequence, the phoneme sequence is input into the encoding network for encoding processing to obtain the text features of the text to be synthesized. Among them, the conversion of the text to be synthesized into a phoneme sequence can also be implemented by a network, such as the network can be a phoneme conversion network, that is, the acoustic network can also include a phoneme conversion network; in a specific implementation, the text to be synthesized can be input into the phoneme conversion network for text-to-phoneme processing to obtain a phoneme sequence, such as Figure 3d shown.

[0070] After obtaining the text features, the emotional text features of the text to be synthesized can be determined based on the text features and the emotional features. The emotional text features of the text to be synthesized can be determined in the following three ways.

[0071] Case (1): The text features and sentiment features can be directly fused to obtain sentiment text features. The fusion process can refer to the addition process. It can be understood that the text features can be composed of one or more text feature values, and the sentiment features can also be composed of one or more sentiment feature values. The feature length of the text features and the feature length of the sentiment features are the same. Then, the text feature values ​​in the text features and the sentiment feature values ​​in the sentiment features can be added together, and the addition result can be used as the sentiment text feature. For example, assuming that the text features can be represented as (A1, A2, A3, A4) and the sentiment features can be represented as (B1, B2, B3, B4), then the sentiment text features can be represented as (A1+B1, A2+B2, A3+B3, A4+B4).

[0072] Case (2): First, a reference weight sequence corresponding to the text features of the text to be synthesized can be obtained, and a vector embedding process can be performed on the reference weight sequence to obtain a reference weight embedding vector for the reference weight sequence; then, based on the emotional features, text features and the reference weight embedding vector, the emotional text features of the text to be synthesized can be determined. For example, the emotional features, text features and the reference weight embedding vector can be fused to obtain the corresponding emotional text features, such as Figure 3d As shown. For the understanding of the fusion process, please refer to the above description and will not be repeated here. The reference weight sequence may include one or more reference weights, and the sequence length of the reference weight sequence is the same as the sequence length of the phoneme sequence. The reference weight in the reference weight sequence may represent the contribution of each phoneme in the phoneme sequence to the current emotional representation, or the contribution of each phoneme to the emotional representation under the target emotional category; and the numerical value corresponding to the reference weight is positively correlated with the contribution, that is, the larger the numerical value corresponding to the reference weight, the greater the corresponding contribution, and the smaller the numerical value corresponding to the reference weight, the smaller the corresponding contribution. In one possible implementation, all reference weights in the reference weight sequence can be set to the highest weight value. For example, if the range of any reference weight is 0-1, the highest weight value is 1, which means that all reference weights in the reference weight sequence are set to 1. Then, the reference weight embedding vector corresponding to the reference weight sequence is added to the output (text feature) of the encoding network, that is, the fusion process of the reference weight embedding vector and the text feature described above.

[0073] As mentioned above, the larger the numerical value corresponding to the reference weight, the greater the corresponding contribution. Setting all the reference weights in the reference weight sequence to 1 can maximize the contribution of each phoneme to the emotional representation under the target emotional category. Then, when the reference weight embedding vector corresponding to this reference weight sequence is subsequently fused to obtain the emotional text feature, the emotional representation of the target emotional category can be made stronger and more extreme. As mentioned above, when determining the emotional features (emotional features are used to fuse with text features and reference weight embedding vectors to obtain emotional text features), the target emotional category and the neutral emotional category are used to determine them. Then, by setting 1, while making the emotional representation of the target emotional category stronger and more extreme, the gap in emotional representation between the target emotional category and the neutral emotional category can also be widened. That is to say, the space for emotional representation can be widened and the averaging effect of emotional representation can be reduced, thereby improving the speech synthesis effect.

[0074] Case (3): After obtaining the emotional text features through case (1) or case (2), the emotional text features can also be subjected to variable information adaptation processing. The variable information adaptation processing can be understood as integrating the features of the variable information into the emotional text features to obtain the emotional text features required in the end. For example, variable information can be duration, pitch, energy, etc., that is, the features of these dimensions can also be integrated into the emotional text features to strengthen the emotional representation of the emotional text features, thereby improving the emotional representation effect. Among them, the variable information adaptation processing can be implemented by a variable information adaptation network, such as the variable information adaptation network can be a network in the acoustic network, that is, the acoustic network can also include a variable information adaptation network; in a specific implementation, the emotional text features can be input into the variable information adaptation network to perform variable information adaptation processing to obtain the emotional text features required in the end, such as Figure 3d shown.

[0075] It should be noted that, for the end-to-end acoustic network in the embodiment of the present application, its specific implementation can be various network structures as described above (such as Figure 3d As shown, the acoustic network can be FastSpeech 2) or based on other acoustic networks, such as Tacotron 2; this application does not specifically limit the acoustic network.

[0076] S204: Synthesizing the emotional speech corresponding to the target emotional category and target emotional intensity of the text to be synthesized according to the emotional text features.

[0077] In one implementation, the emotional text features can be decoded to obtain a mel-spectrogram of the text to be synthesized, which is a spectrum that converts the frequency into a mel scale. The decoding process can be implemented by a decoding network in the acoustic network, that is, the emotional text features can be input into the decoding network for decoding to obtain a mel-spectrogram, such as Figure 3d As shown. Then, the mel-spectrogram is vocoded to synthesize (i.e., obtain) the emotional speech corresponding to the target emotional category and target emotional intensity of the text to be synthesized. Among them, a neural vocoder (such as WaveRNN) can be used to vocode the mel-spectrogram to generate a waveform, thereby obtaining the emotional speech corresponding to the target emotional category and target emotional intensity of the text to be synthesized. In specific application scenarios (such as voice interaction, voice assistant, etc.), after obtaining the emotional speech, the emotional speech can also be broadcast.

[0078] From the above description, it can be seen that the embodiment of the present application can realize the synthesis of emotional speech by calling the emotional speech synthesis model; wherein, steps S201-S204 can be understood as the processing process (or reasoning process) of the emotional speech synthesis model in the actual application scenario. In short, the embodiment of the present application provides an end-to-end emotional speech synthesis model with controllable emotion category and emotion degree. The emotional speech synthesis model can directly accept the text to be synthesized, the emotion category and the emotion intensity as input to obtain the corresponding emotional speech.

[0079] In an embodiment of the present application, the emotional features for the text to be synthesized can be obtained by specifying the emotional category and emotional intensity, and the emotional text features rich in text information and emotional information can be further obtained in combination with the text features of the text to be synthesized, so as to synthesize a synthetic speech of the specified emotional category and emotional intensity through the emotional text features, thereby realizing the synthesis of emotional speech, so as to improve the user's auditory effect, and also to improve the flexibility and controllability of the synthesized speech. In addition, when determining the emotional features, the emotion tokens corresponding to various emotional categories can also be used to represent the emotion, so as to enhance the controllability and interpretability of the emotion representation. In addition, an end-to-end emotional speech synthesis model with controllable emotional categories and emotional degrees can be established, and the emotional speech synthesis model can be directly used to realize the synthesis of emotional speech, so as to improve the automation and intelligence of speech synthesis.

[0080] See also Figure 4 , Figure 4 This is a flow chart of another speech synthesis method provided by an embodiment of the present application. The method is applied to a computer device and can be executed by the computer device; the embodiment of the present application mainly describes the training process of obtaining an emotional speech synthesis model by training a reference model, wherein the reference model includes an acoustic network and an emotional network. Figure 4 As shown, the speech synthesis method may include:

[0081] S401: Obtain training sample pairs.

[0082] The training sample pair may include sample data and label data corresponding to the sample data. The sample data may include sample text and sample audio. The label data may include an emotion category label corresponding to the sample data. The emotion category label may be used to indicate the emotion category corresponding to the sample data. For example, the emotion category label corresponding to a certain sample data is used to indicate the happy emotion category. It should be understood that the number of training sample pairs may include one or more. The embodiment of the present application mainly describes model training using one training sample pair as an example.

[0083] S402: Obtain a sample mel-spectrogram of the sample audio, and input the sample mel-spectrogram into the emotion network in the reference model to obtain sample emotion features and weights of emotion tokens corresponding to multiple preset emotion categories.

[0084] In one implementation, a mel-spectrogram corresponding to the sample audio can be first obtained. For example, the mel-spectrogram here can be referred to as a sample mel-spectrogram; the sample mel-spectrogram can be used as the input of the emotion network in the reference model, so that the emotion network can be used to process the sample mel-spectrogram to obtain the sample emotion features and the weights of the emotion tokens corresponding to multiple preset emotion categories. The weights of the emotion tokens corresponding to the multiple preset emotion categories can be used to calculate the subsequent model loss value; the sample mel-spectrogram here can also be used to calculate the subsequent model loss value, as described in step S404.

[0085] Among them, the model structure diagram of the reference model can be as follows Figure 5a As shown, Figure 5aAs shown, the reference model may include an acoustic network and an emotional network. The acoustic network may be an end-to-end network, wherein the emotional network may include an emotional extraction network and an emotional representation network. Optionally, the specific implementation method of using the emotional network to implement step S402 may be: the sample Mel spectrum map may be first input into the emotional extraction network for feature extraction to obtain the global emotional features of the sample audio, and the global emotional features may be directly the output of the emotional extraction network; further, the global emotional features may be input into the emotional representation network for feature extraction again to obtain the sample emotional features of the sample audio under multiple preset emotional categories, and the weights of the emotional tokens corresponding to the multiple preset emotional categories. Among them, the sample emotional features may be directly the output of the emotional representation network; the weights of multiple emotional tokens may be obtained in the process of processing the emotional representation network, and the specific acquisition process may refer to the description of the emotional representation network below. The following is an introduction to the emotional extraction network and the emotional representation network.

[0086] (1) Sentiment Extraction Network

[0087] The network structure diagram of the emotion extraction network can be shown as follows Figure 5b As shown in , the emotion extraction network can also be called a sample audio coding network. Figure 5bAs shown, the input of the emotion extraction network can be a sample Mel-spectrogram. As described above, the sample Mel-spectrogram can be a Mel-spectrogram corresponding to a sample audio; the output of the emotion extraction network can include a global emotion feature and can also include a sample weight sequence. The sequence length of the sample weight sequence can be the same as the sequence length of the sample phoneme sequence corresponding to the sample Mel-spectrogram. The network structure of the emotion extraction network (sample audio coding network) can mainly include L convolutional networks, a gate-controlled recurrent network, a length adjustment network, and an attention mechanism network. Among them, the L convolutional networks can be used to extract emotion features from the sample Mel-spectrogram, and each convolutional network can include a one-dimensional convolution, batch normalization, and an activation function. The gate-controlled recurrent network can further extract emotion features from the sample Mel-spectrogram. The length adjustment network can convert the sequence length. For example, in this application, a sequence of the same length as the frame sequence can be converted into a sequence of the same length as the sample phoneme sequence. The sequence length of the emotional features extracted from the sample Mel-spectrogram through the gate-controlled recurrent network is the same as the sequence length of the frame sequence. The frame sequence here can refer to the original sequence corresponding to the sample Mel-spectrogram. Usually, the original sequence corresponding to the sample Mel-spectrogram is longer than the sample phoneme sequence corresponding to the sample text (the understanding of the sample phoneme sequence can refer to the description in the following step S403). In order to ensure the same length of the sequence, the sequence length can be adjusted by the length adjuster. Specifically, the length adjustment network can adopt a frame merging operation to take the average or maximum value of all adjacent frames belonging to the same phoneme, thereby converting a sequence of the same length as the frame sequence into a sequence of the same length as the sample phoneme sequence.

[0088] The attention mechanism network can be used to obtain global sentiment features and a sequence of sample weights. The attention mechanism network can be a single-headed attention mechanism network. For the attention mechanism network, the output of the length adjustment network can serve as the key and value of the attention mechanism network, and a randomly initialized learnable vector can be used as the query. Through the attention mechanism network, a global sentiment feature and a sequence of sample weights can be obtained. The global sentiment feature can represent the global sentiment information of the sample audio. Each sample weight in the sample weight sequence can represent the contribution of each phoneme in the sample audio to the current sentiment representation. Generally, the phonemes included in the sample audio are consistent with the phonemes included in the sample text. In this case, each sample weight in the sample weight sequence can also be understood as representing: the contribution of the corresponding phoneme in the sample phoneme sequence corresponding to the sample text to the sentiment representation, or the contribution of the corresponding phoneme to the sentiment representation under the sentiment category corresponding to the sentiment category label.

[0089] It should be noted that for the emotion extraction network, the key to the emotion extraction network is to use the attention mechanism to extract global emotion features and sample weight sequences. Figure 5b In addition to the attention mechanism network structure shown, in other implementations, the emotion extraction network can also be based on a convolutional neural network, or an LSTM, or a fully connected network structure, etc., which is not specifically limited in this application.

[0090] (2) Emotion Representation Network

[0091] The network structure diagram of the emotion representation network can be shown as follows Figure 5c As shown in , the key to the emotion representation network is to limit different emotion categories to different emotion tokens, so the emotion representation network can also be called the emotion token layer. Figure 5c As shown, the input of the emotion representation network can be the global emotion feature from the output of the sample audio encoder (emotion extraction network), and the output can include a sentence-level emotion vector, which can be the sample emotion feature mentioned above. The output can also include emotion tokens corresponding to multiple preset emotion categories. For example, if the number of multiple preset emotion categories is N, then the output is N emotion tokens, that is, emotion tokens corresponding to N preset emotion categories, such as the multiple preset emotion categories can include neutral, happy, angry, sad, surprised, afraid, and disgust. The network structure of the emotion representation network (emotion token layer) can mainly include an attention mechanism network, which can be used to use global emotion features for further emotion feature extraction to obtain sample emotion features. The attention mechanism in the attention mechanism network can be a single-headed attention mechanism. Simply put, the attention mechanism network can be understood as a single-headed attention mechanism network specifically for emotion representation.

[0092] The attention mechanism network can use N (N is the total number corresponding to multiple preset emotion categories) randomly initialized emotion tokens as keys and values, and the input of the emotion representation network (i.e., global emotion features) as queries to implement the processing mechanism of the attention mechanism network. It is understandable that the core of the attention mechanism is to let the network pay attention to where it needs more attention, which is generally reflected in the form of attention weights. Usually, the essence of the attention mechanism can be understood as weighted summation. Correspondingly, applied in the embodiment of the present application, the attention mechanism can be understood as the weighted summation of N emotion tokens. Among them, the sample emotion feature is the weighted summation result of the N emotion tokens. It can be seen that compared with the global emotion feature obtained by the emotion extraction network, the sample emotion feature is more interpretable. It can also be seen that in the weighted summation process, the weights corresponding to each emotion token in the N emotion tokens are involved (also called attention weights).

[0093] The weight of an emotion token can be obtained by calculating the similarity between the global emotion feature and the emotion token. For example, the weight of each emotion token can be determined by calculating the similarity between the global emotion feature and each emotion token to obtain the similarity between the global emotion feature and each emotion token. Each similarity is then normalized (such as softmax), and each normalized result is the weight of the corresponding emotion token. In this way, the weight corresponding to each emotion token can be obtained, and then the sample emotion feature can be obtained by weighting N emotion tokens based on the weight.

[0094] When training the reference model, each emotion token can be randomly initialized first. Then, through continuous iterative training, the N emotion tokens and the weights corresponding to the N emotion tokens are also continuously updated. After the reference model training is completed, the obtained N emotion tokens and the weights corresponding to the N emotion tokens (that is, the N emotion tokens in the emotion representation network in the last iterative training and the weights corresponding to the N emotion tokens) are the data required by the embodiment of the present application. Among them, the N emotion tokens here are the emotion tokens involved in the above-mentioned step S202; the weights corresponding to the N emotion tokens are what are needed to calculate the model loss value in the following step S404. One emotion token corresponds to a preset emotion category, and the weights corresponding to the N emotion tokens here are the weights of the emotion tokens corresponding to multiple preset emotion categories, and N is the total number of multiple preset emotion categories.

[0095] In an embodiment of the present application, for the attention mechanism network of the emotion representation network, a cross-entropy loss can be added to the weights in the attention mechanism network (that is, the weights of each emotion token), wherein the cross-entropy loss is a multi-classification loss. Then, a cross-entropy loss is added to the weights of multiple emotion tokens, that is, each emotion token is classified (that is, each emotion token corresponds to an emotion category), and in the training of the reference model, each sample data corresponds to an emotion category label for indicating the emotion category. In the training of the reference model, when the cross entropy loss is calculated based on the weights of multiple emotion tokens obtained from the sample data and the emotion category label corresponding to the sample data, if the classification is wrong (that is, the predicted emotion category and the emotion category label are different), a large loss value will be generated when calculating the cross entropy loss. Then, when training through the loss value, this loss value will gradually decrease. When this loss value is updated to a smaller value, it can be indicated that the classification is correct, that is, the emotion category label corresponding to the sample data can be restricted to the corresponding emotion token. For example, assuming that the emotion category label corresponding to the sample data is happy, then in the calculation of the cross entropy loss, the emotion representation of the happy emotion category can be restricted to the happy emotion token. Then, through training with a large amount of sample data, the emotion representation corresponding to each emotion category can be restricted to a specified corresponding emotion token, so that one emotion token can correspond to the emotion representation of one emotion category. In this way, the weighted combination of multiple emotion tokens can well represent the entire emotion space.

[0096] It should be noted that for the emotion representation network, the key is to limit the emotion representation of different emotion categories to different emotion tokens; the specific number and type of emotion tokens depend on the data used. In addition to the emotion categories mentioned above, other categories of emotions can also be used, such as relief, admiration, etc. This application does not make specific limitations on the preset emotion categories.

[0097] It is understandable that the sample audio is usually an audio signal in the time domain. In signal processing, data in the frequency domain is usually used for processing, so the sample audio in the time domain can be converted into sample audio in the frequency domain. Based on this, the method for obtaining the sample Mel spectrum of the sample audio can be: first use the fast Fourier transform to convert the sample audio from the time domain to the frequency domain to obtain the corresponding spectrum; wherein, the spectrum can represent the distribution of the sample audio at different frequencies, and the spectrum is characterized by the frequency as a logarithmic scale. Compared with the sample audio in the time domain, using the frequency domain to represent the sample audio can more easily highlight the characteristics of the sample audio, and the amount of data can also be reduced, so as to increase the data processing speed. After obtaining the spectrum corresponding to the sample audio, the frequency in the spectrum can be further converted to obtain a Mel spectrum, such as converting the frequency from a logarithmic scale to a Mel scale to form a Mel spectrum.

[0098] S403: Input the sample text and the sample sentiment features into the acoustic network in the reference model to obtain a predicted mel-spectrogram of the sample text.

[0099] In one implementation, the acoustic network may include an encoding network and a decoding network; wherein the encoding network may be used to encode the sample phoneme sequence corresponding to the sample text to obtain the sample text features of the sample text; based on this, it can be seen that the encoding network processes phonemes rather than directly processing the text, and the sample text may be first subjected to text-to-phoneme processing to obtain the sample phoneme sequence corresponding to the sample text. The text-to-phoneme processing may be implemented by a phoneme conversion network, which may be included in the acoustic network, such as Figure 5d Then, after obtaining the sample phoneme sequence, the sample phoneme sequence can be input into the encoding network in the acoustic network for encoding processing to obtain the sample text features of the sample text.

[0100] The sample emotion features can also be input into the acoustic network in the reference model to obtain the sample emotion text features in the acoustic network based on the sample emotion features and the sample text features. The sample emotion text features can include features of multiple dimensions (i.e., text and emotion); further, the sample emotion text features can be input into the decoding network for decoding processing to obtain the predicted Mel spectrum of the sample text.

[0101] Optionally, the specific implementation methods for obtaining the sample emotional text features based on the sample emotional features and the sample text features may include the following three situations:

[0102] Case (1): directly fusing the sample sentiment feature and the sample text feature to obtain the corresponding sample sentiment text feature. The fusing process can be referred to the description in step S203 above, which will not be repeated here.

[0103] Case (2): A sample weight sequence is obtained from the output of the sentiment extraction network; the sequence length of the sample weight sequence is the same as the sequence length of the sample phoneme sequence, and each sample weight in the sample weight sequence can be used to represent: the contribution of the corresponding phoneme in the sample phoneme sequence to the sentiment representation, wherein the relevant understanding of the sample weight sequence can refer to the description in step S402. After obtaining the sample weight sequence, the sample weight sequence can be vector embedded to obtain a sample weight embedding vector for the sample weight sequence; through the vector embedding process, the dimension of the sample weight sequence can be converted to be consistent with the dimension of the sample sentiment feature or the sample text feature, so that fusion processing can be performed. Furthermore, the sample sentiment text feature can be obtained based on the sample sentiment feature, the sample text feature and the sample weight embedding vector, such as the sample sentiment feature, the sample text feature and the sample weight embedding vector can be fused to obtain the sample sentiment text feature. The fusion processing can refer to the description in the above step S203 and will not be repeated here.

[0104] Optionally, each sample weight in the sample weight sequence can be represented by a numerical value. Considering that the sample weight sequence is composed of a large number of discrete numerical values, the numerical values ​​corresponding to each sample weight may be different. Then, when the sample weight sequence is subsequently vector embedded, a large number of different numerical values ​​need to be vector embedded, which requires a large amount of calculation and affects the model training speed. The sample weights in the sample weight sequence can be binned. The binning operation can be understood as quantization, that is, the numerical values ​​in a certain range can be classified into one numerical value to reduce the amount of calculation of subsequent vector embedding, thereby improving the model training speed.

[0105] The specific implementation of the binning operation can be: determining H binning categories corresponding to the sample weight sequence and the bin value corresponding to each binning category; wherein each binning category represents a numerical range, such as one binning category in the H binning categories is 0-0.1, another binning category is 0.1-0.2, and so on; the bin value corresponding to the binning category is to set all values ​​in the binning category to a single value, such as all values ​​in the binning category 0-0.1 can be set to 0.05. When determining the binning category, the H binning categories can be divided based on the maximum and minimum values ​​in the sample weight sequence. For example, the difference between the maximum and minimum values ​​can be calculated, and the ratio between the difference and the reference value can be calculated. The result of rounding down or rounding up the ratio is then used as the number of binning categories (i.e., H). The numerical range between the maximum and minimum values ​​is then divided into H numerical ranges (the range size of each numerical range can be the same or different). The H numerical ranges are then H binning categories. For example, if H is 4, the maximum and minimum values ​​are 0.1-0.9 respectively, then the four binning categories can be 0.1-0.3, 0.3-0.5, 0.5-0.7, and 0.7-0.9. H can also be preset, such as 10, 8, and other values. After determining the H binning categories, the binning category corresponding to each sample weight in the sample weight sequence is determined based on the H binning categories, and then the value of each sample weight after the binning operation is determined based on the binning value corresponding to the binning category. For example, for a sample weight, the sample weight can be reset to the binning value corresponding to the binning category in which it is located. For example, assuming a sample weight is 0.01, its corresponding binning category is 0-0.1, and the binning value of the binning category is 0.05, then after the binning operation, the sample weight can be updated to 0.05.

[0106] Case (3): After obtaining the sample emotion text features through case (1) or case (2), the sample emotion text features can also be subjected to variable information adaptation processing. The variable information adaptation processing can be understood as fusing one or more variable information (such as duration, pitch, energy, etc.) features into the sample emotion text features to obtain the final required sample emotion text features; wherein, the variable information adaptation processing can be implemented by a variable information adaptation network, such as the variable information adaptation network can be a network in the acoustic network, that is, the acoustic network can also include a variable information adaptation network; in a specific implementation, the sample emotion text features can be input into the variable information adaptation network for variable information adaptation processing to obtain the final required sample emotion text features, such as Figure 5d shown.

[0107] S404: Train the reference model according to the predicted mel-spectrogram, the sample mel-spectrogram, the emotion category label, and the weights of the emotion tokens corresponding to the plurality of preset emotion categories to obtain a trained reference model, and determine the emotion speech synthesis model according to the trained reference model.

[0108] In one implementation, a model loss value of a reference model can be determined based on the predicted mel-spectrogram, the sample mel-spectrogram, the predicted emotion category label, and the emotion category label. The reference model can then be trained using the model loss value to obtain a trained reference model. For example, model parameters in the reference model can be updated in a direction that reduces the model loss value to achieve the desired model training effect.

[0109] Optionally, a specific implementation method for determining the model loss value of the reference model may be: using the predicted mel-spectrogram and the sample mel-spectrogram to determine the first loss value of the reference model, and using the weights of the emotion tokens corresponding to multiple preset emotion categories and the emotion category labels to determine the second loss value of the reference model; wherein the loss function used to determine the first loss value may be a mean square error loss function, and the loss function used to determine the second loss value may be a cross-entropy loss function. After obtaining the first loss value and the second loss value, the model loss value may be determined based on these two loss values. For example, the sum of the first loss value and the second loss value may be used as the model loss value; for another example, the first loss value and the second loss value may be weighted, and the sum of the two weighted processing results may be used as the model loss value. The weighted value corresponding to the weighted processing of the first loss value and the weighted value corresponding to the weighted processing of the second loss value may be pre-set, and their values ​​are not specifically limited in this application. For example, the weighted value corresponding to the first loss value may be 0.3, and the weighted value corresponding to the second loss value may be 0.7. Based on this, it can be seen that using the above training method, the reference model can be trained in combination with losses in multiple dimensions to improve the robustness of the model and make the speech synthesis effect of the reference model better.

[0110] In one implementation, after obtaining the trained reference model, the emotional speech synthesis model can be determined based on the trained reference model. Optionally, emotional tokens corresponding to multiple preset emotional categories can be obtained from the output of the emotional representation network, and an emotional control network can be constructed by the emotional tokens corresponding to multiple preset emotional categories; the output here can refer to the output of the emotional representation network during the last training in the iterative training process of the reference model. Then the acoustic network in the trained reference model is extracted to determine the emotional speech synthesis model based on the acoustic network and the emotional control network in the trained reference model, that is, the emotional speech synthesis model can be composed of the acoustic network and the emotional control network in the trained reference model, as can be seen in the emotional speech synthesis model. Figure 3bor Figure 3d shown.

[0111] The above method embodiments are all examples of the method of the present application. The description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments. For example, after training the emotional speech synthesis model, the target emotional category and target emotional intensity corresponding to the text to be synthesized can be obtained, so as to control the emotional category and emotional intensity of the text to be synthesized based on the emotional speech synthesis model, and obtain the emotional speech of the text to be synthesized under the target emotional category and target emotional intensity. It will not be repeated here.

[0112] In an embodiment of the present application, data without emotion intensity labels can be used directly for training, and there is no need to perform additional preprocessing operations such as emotion intensity sorting or extraction on the emotion data set, so as to simplify the training process of the reference model; using the trained reference model, an end-to-end emotion category and emotion degree controllable emotion speech synthesis model can be established to directly accept the text to be synthesized, emotion category and emotion intensity as input, and obtain the corresponding emotion speech, so as to improve the automation and intelligence of speech synthesis.

[0113] To better understand the speech synthesis method in this application, the following combination Figure 3d and Figure 5d The model structure diagram shown is further explained. In order to achieve speech synthesis under specific emotion categories and emotion intensities, the embodiment of the present application may involve an acoustic network. Taking FastSpeech 2 as an example, an emotion extraction network (also called a sample audio encoder), an emotion representation network (also called an emotion token layer) and an emotion control network are added on this basis, which can be used to extract, represent and control emotions respectively. The embodiment of the present application involves a training phase (such as step S401-step S404) and an inference phase (such as step S201-step S204). Among them, there are differences in the model frameworks used in the training phase and the inference phase. The model used in the training phase can be the reference model mentioned above (the reference model includes an acoustic network, an emotion extraction network and an emotion extraction network), such as Figure 5d The model used in the inference stage can be the emotional speech synthesis model mentioned above (the emotional speech synthesis model includes the acoustic network and the emotional control network), as shown in Figure 3d The following describes these two stages.

[0114] During the training phase, the reference model can receive <sample text, sample audio> pairs with emotion category labels as input. The emotion extraction network can extract global emotion features from the sample mel-spectrogram corresponding to the sample audio, and then input the global emotion features into the emotion representation network to obtain sentence-level emotion features of a specific emotion category and expand them into sample emotion features of the same length as the sample phoneme sequence corresponding to the sample audio. Furthermore, the sample emotion features can be added to the output of the encoding network in the acoustic network (i.e., the sample text features). In addition, the sample weight embedding vector corresponding to the sample weight sequence in the output of the emotion extraction network can also be added to the output of the encoding network to obtain the final required sample emotion text features. Optionally, the sample weight sequence can be binned first, and then each category can be vector embedded separately to obtain the sample weight embedding vector. After obtaining the sample emotion text features, the predicted mel-spectrogram is further obtained based on the sample emotion text features, thereby realizing the training of the reference model, such as Figure 5d shown.

[0115] In the inference stage, speech of a specific emotion category and emotion intensity can be synthesized by artificially given text to be synthesized, target emotion category and target emotion intensity. Among them, the emotion control network receives the target emotion category and target emotion intensity as input, performs specific weighted representation on the multiple emotion tokens obtained in the training stage, outputs the emotion features of the specified emotion category and emotion intensity, and then adds them to the output of the encoding network in the acoustic network (i.e., text features). In addition, the reference weight embedding vector corresponding to the reference weight sequence can also be added to the output of the encoding network to obtain the emotion text feature; further, a mel spectrum graph can be obtained based on the emotion text feature, and the emotion speech can be obtained based on the mel spectrum graph, such as Figure 3d shown.

[0116] It can be seen that the embodiment of the present application provides a speech synthesis method that realizes controllable emotion categories and intensities. In the training stage, the method can first obtain the global emotion features at the sentence level and the sample weight sequence of the same length as the sample phoneme sequence through the emotion extraction network; then, different categories of emotions are respectively limited to different emotion tokens through a specific emotion representation network; thereby obtaining the sentence-level emotion representation vector (i.e., sample emotion features) and multiple emotion tokens representing preset emotion categories, and these multiple emotion tokens can well represent the entire emotion space. In the reasoning stage, different weighted combinations of emotion tokens can be artificially specified through the emotion control network to obtain emotion representations of different emotion categories and emotion intensities, and further synthesize the corresponding audio through the end-to-end acoustic network. Among them, the emotion representation control can be done at the sentence level or at a more fine-grained phoneme level. In addition, the embodiment of the present application also introduces weight embedding (such as the sample weight embedding vector in the training stage and the reference weight embedding vector in the reasoning stage) to further expand the space of emotion representation and reduce the impact of emotion representation averaging, thereby improving the quality of speech synthesis.

[0117] See also Figure 6 , is a schematic diagram of the structure of a speech synthesis device provided in an embodiment of the present application. The speech synthesis device described in this embodiment includes:

[0118] An acquisition unit 601 is used to acquire a target emotion category and a target emotion intensity corresponding to the text to be synthesized;

[0119] The first determining unit 602 is configured to determine the emotional features of the text to be synthesized based on the emotional tokens corresponding to the plurality of preset emotional categories, the target emotional category, and the target emotional intensity; any emotional token is used to represent the features of the corresponding preset emotional category;

[0120] A second determining unit 603 is configured to determine the emotional text feature of the text to be synthesized based on the text feature of the text to be synthesized and the emotional feature;

[0121] The synthesis unit 604 is configured to synthesize the emotional speech of the text to be synthesized according to the emotional text features, so that the text complies with the target emotional category and the target emotional intensity.

[0122] In one implementation, the first determining unit 602 is specifically configured to:

[0123] The target emotion category and the target emotion intensity are input into the emotion control network in the emotion speech synthesis model, and the emotion control network weights the emotion tokens corresponding to the multiple preset emotion categories based on the target emotion category and the target emotion intensity to obtain the emotion characteristics of the text to be synthesized; the emotion tokens corresponding to the preset emotion categories are obtained during the training process of the emotion speech synthesis model.

[0124] In one implementation, the emotional speech synthesis model is obtained by training a reference model, where the reference model includes an acoustic network and an emotional network; the apparatus further includes a training unit 605, specifically configured to:

[0125] Acquire a training sample pair, the training sample pair including sample data and label data corresponding to the sample data, the sample data including sample text and sample audio, and the label data including an emotion category label corresponding to the sample data;

[0126] Obtaining a sample Mel-spectrogram of the sample audio, and inputting the sample Mel-spectrogram into the emotion network to obtain sample emotion features and weights of emotion tokens corresponding to the plurality of preset emotion categories;

[0127] Inputting the sample text and the sample emotion feature into the acoustic network to obtain a predicted Mel-spectrogram of the sample text;

[0128] The reference model is trained according to the predicted mel-spectrogram, the sample mel-spectrogram, the emotion category label, and the weights of the emotion tokens corresponding to the multiple preset emotion categories to obtain a trained reference model, and the emotional speech synthesis model is determined based on the trained reference model.

[0129] In one implementation, the emotion network includes an emotion extraction network and an emotion representation network; the training unit 605 is specifically configured to:

[0130] Inputting the sample Mel-spectrogram into the emotion extraction network for feature extraction to obtain the global emotion feature of the sample audio;

[0131] The global emotion feature is input into the emotion representation network for feature extraction to obtain the sample emotion features of the sample audio under the multiple preset emotion categories, and the weights of the emotion tokens corresponding to the multiple preset emotion categories.

[0132] In one implementation, the acoustic network includes an encoding network and a decoding network; the training unit 605 is specifically configured to:

[0133] Performing text-to-phoneme processing on the sample text to obtain a sample phoneme sequence corresponding to the sample text;

[0134] Inputting the sample phoneme sequence into the encoding network for encoding processing to obtain sample text features of the sample text;

[0135] Obtaining a sample emotional text feature according to the sample emotional feature and the sample text feature;

[0136] The sample emotional text features are input into the decoding network for decoding processing to obtain a predicted Mel-spectrogram of the sample text.

[0137] In one implementation, the training unit 605 is specifically configured to:

[0138] Obtaining a sample weight sequence from the output of the emotion extraction network; the sequence length of the sample weight sequence is the same as the sequence length of the sample phoneme sequence, and each sample weight in the sample weight sequence is used to represent the contribution of the corresponding phoneme in the sample phoneme sequence to the emotion representation;

[0139] Performing vector embedding on the sample weight sequence to obtain a sample weight embedding vector for the sample weight sequence;

[0140] A sample emotional text feature is obtained according to the sample emotional feature, the sample text feature and the sample weight embedding vector.

[0141] In one implementation, the training unit 605 is specifically configured to:

[0142] Obtaining emotion tokens corresponding to a plurality of preset emotion categories from the output of the emotion representation network;

[0143] Constructing an emotion control network using emotion tokens corresponding to the plurality of preset emotion categories;

[0144] The emotional speech synthesis model is determined based on the acoustic network and the emotional control network in the trained reference model.

[0145] In one implementation, the second determining unit 603 is specifically configured to:

[0146] Obtaining a reference weight sequence corresponding to the text features of the text to be synthesized, wherein the sequence length of the reference weight sequence is the same as the sequence length of the phoneme sequence of the text to be synthesized;

[0147] Performing vector embedding processing on the reference weight sequence to obtain a reference weight embedding vector for the reference weight sequence;

[0148] The emotional text feature of the text to be synthesized is determined according to the emotional feature, the text feature and the reference weight embedding vector.

[0149] In one implementation, the synthesis unit 604 is specifically configured to:

[0150] Decoding the emotional text features to obtain a mel-spectrogram of the text to be synthesized;

[0151] The mel-spectrogram is subjected to voice coding conversion to obtain emotional speech of the text to be synthesized that meets the target emotion category and the target emotion intensity.

[0152] It is understood that the division of units in the embodiments of the present application is schematic and is merely a logical functional division. In actual implementation, other division methods may be used. The functional units in the embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0153] See also Figure 7 , is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device described in this embodiment includes: a processor 701, a memory 702, and a network interface 703. The processor 701, the memory 702, and the network interface 703 can exchange data.

[0154] The processor 701 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.

[0155] The memory 702 may include a read-only memory and a random access memory, and provides program instructions and data to the processor 701. A portion of the memory 702 may also include a non-volatile random access memory. When the processor 701 calls the program instructions, it is used to execute:

[0156] Obtain the target sentiment category and target sentiment intensity corresponding to the text to be synthesized;

[0157] Determining the emotional features of the text to be synthesized based on the emotional tokens corresponding to the plurality of preset emotional categories, the target emotional category, and the target emotional intensity; any emotional token is used to represent the features of the corresponding preset emotional category;

[0158] Determining the emotional text features of the text to be synthesized based on the text features of the text to be synthesized and the emotional features;

[0159] The emotional speech of the text to be synthesized conforming to the target emotional category and the target emotional intensity is synthesized according to the emotional text features.

[0160] In one implementation, the processor 701 is specifically configured to:

[0161] The target emotion category and the target emotion intensity are input into the emotion control network in the emotion speech synthesis model, and the emotion control network weights the emotion tokens corresponding to the multiple preset emotion categories based on the target emotion category and the target emotion intensity to obtain the emotion characteristics of the text to be synthesized; the emotion tokens corresponding to the preset emotion categories are obtained during the training process of the emotion speech synthesis model.

[0162] In one implementation, the emotional speech synthesis model is obtained by training a reference model, where the reference model includes an acoustic network and an emotional network; the processor 701 is further configured to:

[0163] Acquire a training sample pair, the training sample pair including sample data and label data corresponding to the sample data, the sample data including sample text and sample audio, and the label data including an emotion category label corresponding to the sample data;

[0164] Obtaining a sample Mel-spectrogram of the sample audio, and inputting the sample Mel-spectrogram into the emotion network to obtain sample emotion features and weights of emotion tokens corresponding to the plurality of preset emotion categories;

[0165] Inputting the sample text and the sample emotion feature into the acoustic network to obtain a predicted Mel-spectrogram of the sample text;

[0166] The reference model is trained according to the predicted mel-spectrogram, the sample mel-spectrogram, the emotion category label, and the weights of the emotion tokens corresponding to the multiple preset emotion categories to obtain a trained reference model, and the emotional speech synthesis model is determined based on the trained reference model.

[0167] In one implementation, the emotion network includes an emotion extraction network and an emotion representation network; the processor 701 is specifically configured to:

[0168] Inputting the sample Mel-spectrogram into the emotion extraction network for feature extraction to obtain the global emotion feature of the sample audio;

[0169] The global emotion feature is input into the emotion representation network for feature extraction to obtain the sample emotion features of the sample audio under the multiple preset emotion categories, and the weights of the emotion tokens corresponding to the multiple preset emotion categories.

[0170] In one implementation, the acoustic network includes an encoding network and a decoding network; the processor 701 is specifically configured to:

[0171] Performing text-to-phoneme processing on the sample text to obtain a sample phoneme sequence corresponding to the sample text;

[0172] Inputting the sample phoneme sequence into the encoding network for encoding processing to obtain sample text features of the sample text;

[0173] Obtaining a sample emotional text feature according to the sample emotional feature and the sample text feature;

[0174] The sample emotional text features are input into the decoding network for decoding processing to obtain a predicted Mel-spectrogram of the sample text.

[0175] In one implementation, the processor 701 is specifically configured to:

[0176] Obtaining a sample weight sequence from the output of the emotion extraction network; the sequence length of the sample weight sequence is the same as the sequence length of the sample phoneme sequence, and each sample weight in the sample weight sequence is used to represent the contribution of the corresponding phoneme in the sample phoneme sequence to the emotion representation;

[0177] Performing vector embedding on the sample weight sequence to obtain a sample weight embedding vector for the sample weight sequence;

[0178] A sample emotional text feature is obtained according to the sample emotional feature, the sample text feature and the sample weight embedding vector.

[0179] In one implementation, the processor 701 is specifically configured to:

[0180] Obtaining emotion tokens corresponding to a plurality of preset emotion categories from the output of the emotion representation network;

[0181] Constructing an emotion control network using emotion tokens corresponding to the plurality of preset emotion categories;

[0182] The emotional speech synthesis model is determined based on the acoustic network and the emotional control network in the trained reference model.

[0183] In one implementation, the processor 701 is specifically configured to:

[0184] Obtaining a reference weight sequence corresponding to the text features of the text to be synthesized, wherein the sequence length of the reference weight sequence is the same as the sequence length of the phoneme sequence of the text to be synthesized;

[0185] Performing vector embedding processing on the reference weight sequence to obtain a reference weight embedding vector for the reference weight sequence;

[0186] The emotional text feature of the text to be synthesized is determined according to the emotional feature, the text feature and the reference weight embedding vector.

[0187] In one implementation, the processor 701 is specifically configured to:

[0188] Decoding the emotional text features to obtain a Mel-spectrogram of the text to be synthesized;

[0189] The mel-spectrogram is subjected to voice coding conversion to obtain emotional speech of the text to be synthesized that meets the target emotion category and the target emotion intensity.

[0190] The embodiment of the present application further provides a computer storage medium in which program instructions are stored. When the program is executed, the program may include: Figure 2 or Figure 4 Part or all of the steps of the speech synthesis method in the corresponding embodiment.

[0191] It should be noted that for the aforementioned various method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0192] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0193] The present application also provides a computer program product or computer program, which includes program instructions stored in a computer-readable storage medium. A processor of a computer device reads the program instructions from the computer-readable storage medium and executes the program instructions, causing the computer device to perform the steps performed in the above-described method embodiments.

[0194] The above is a detailed introduction to a speech synthesis method, device, computer equipment and storage medium provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the contents of this specification should not be understood as limiting the present application.

Claims

1. A speech synthesis method, characterized in that: The method comprises: Obtain the target sentiment category and target sentiment intensity corresponding to the text to be synthesized; Determining the emotional features of the text to be synthesized based on the emotional tokens corresponding to the plurality of preset emotional categories, the target emotional category, and the target emotional intensity; any emotional token is used to represent the features of the corresponding preset emotional category; Determining the emotional text features of the text to be synthesized based on the text features of the text to be synthesized, a reference weight sequence corresponding to the text features, and the emotional features, wherein the reference weight sequence includes one or more reference weights, and the reference weights in the reference weight sequence are used to represent the contribution of each phoneme in the phoneme sequence of the text to be synthesized to the emotional representation under the target emotional category, and the sequence length of the reference weight sequence is the same as the sequence length of the phoneme sequence of the text to be synthesized; The emotional speech of the text to be synthesized conforming to the target emotional category and the target emotional intensity is synthesized according to the emotional text features.

2. The method according to claim 1, characterized in that The step of determining the emotional features of the text to be synthesized based on the emotional tokens corresponding to the plurality of preset emotional categories, the target emotional category, and the target emotional intensity includes: The target emotion category and the target emotion intensity are input into the emotion control network in the emotion speech synthesis model, and the emotion control network weights the emotion tokens corresponding to the multiple preset emotion categories based on the target emotion category and the target emotion intensity to obtain the emotion characteristics of the text to be synthesized; the emotion tokens corresponding to the preset emotion categories are obtained during the training process of the emotion speech synthesis model.

3. The method according to claim 2, characterized in that The emotional speech synthesis model is obtained by training a reference model, wherein the reference model includes an acoustic network and an emotional network; the method further includes: Acquire a training sample pair, the training sample pair including sample data and label data corresponding to the sample data, the sample data including sample text and sample audio, and the label data including an emotion category label corresponding to the sample data; Obtaining a sample Mel-spectrogram of the sample audio, and inputting the sample Mel-spectrogram into the emotion network to obtain sample emotion features and weights of emotion tokens corresponding to the plurality of preset emotion categories; Inputting the sample text and the sample emotion feature into the acoustic network to obtain a predicted Mel-spectrogram of the sample text; The reference model is trained according to the predicted mel-spectrogram, the sample mel-spectrogram, the emotion category label, and the weights of the emotion tokens corresponding to the multiple preset emotion categories to obtain a trained reference model, and the emotional speech synthesis model is determined based on the trained reference model.

4. The method according to claim 3, characterized in that The emotion network includes an emotion extraction network and an emotion representation network; the step of inputting the sample Mel-spectrogram into the emotion network to obtain the sample emotion features and the weights of the emotion tokens corresponding to the plurality of preset emotion categories includes: Inputting the sample Mel-spectrogram into the emotion extraction network for feature extraction to obtain the global emotion feature of the sample audio; The global emotion feature is input into the emotion representation network for feature extraction to obtain the sample emotion features of the sample audio under the multiple preset emotion categories, and the weights of the emotion tokens corresponding to the multiple preset emotion categories.

5. The method according to claim 4, characterized in that The acoustic network includes an encoding network and a decoding network; inputting the sample text and the sample emotion feature into the acoustic network to obtain a predicted Mel-spectrogram of the sample text includes: Performing text-to-phoneme processing on the sample text to obtain a sample phoneme sequence corresponding to the sample text; Inputting the sample phoneme sequence into the encoding network for encoding processing to obtain sample text features of the sample text; Obtaining a sample emotional text feature according to the sample emotional feature and the sample text feature; The sample emotional text features are input into the decoding network for decoding processing to obtain a predicted Mel-spectrogram of the sample text.

6. The method according to claim 5, characterized in that The step of obtaining the sample emotional text feature according to the sample emotional feature and the sample text feature includes: Obtaining a sample weight sequence from the output of the emotion extraction network; the sequence length of the sample weight sequence is the same as the sequence length of the sample phoneme sequence, and each sample weight in the sample weight sequence is used to represent the contribution of the corresponding phoneme in the sample phoneme sequence to the emotion representation; Performing vector embedding on the sample weight sequence to obtain a sample weight embedding vector for the sample weight sequence; A sample emotional text feature is obtained according to the sample emotional feature, the sample text feature and the sample weight embedding vector.

7. The method according to claim 4, characterized in that Determining the emotional speech synthesis model according to the trained reference model includes: Obtaining emotion tokens corresponding to a plurality of preset emotion categories from the output of the emotion representation network; Constructing an emotion control network using emotion tokens corresponding to the plurality of preset emotion categories; The emotional speech synthesis model is determined based on the acoustic network and the emotional control network in the trained reference model.

8. The method according to claim 1 or 2, characterized in that The step of determining the emotional text features of the text to be synthesized based on the text features of the text to be synthesized, the reference weight sequence corresponding to the text features, and the emotional features includes: Obtaining a reference weight sequence corresponding to the text features of the text to be synthesized; Performing vector embedding processing on the reference weight sequence to obtain a reference weight embedding vector for the reference weight sequence; The emotional text feature of the text to be synthesized is determined according to the emotional feature, the text feature and the reference weight embedding vector.

9. The method according to claim 1 or 2, characterized in that The step of synthesizing the emotional speech of the text to be synthesized according to the emotional text feature so as to conform to the target emotional category and the target emotional intensity includes: Decoding the emotional text features to obtain a Mel-spectrogram of the text to be synthesized; The mel-spectrogram is subjected to voice coding conversion to obtain emotional speech of the text to be synthesized that meets the target emotion category and the target emotion intensity.

10. A computer device, characterized in that: The method comprises a processor, a memory and a network interface, wherein the processor, the memory and the network interface are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1 to 9.

11. A computer storage medium, characterized in that The computer storage medium stores a computer program, which includes program instructions. When the program instructions are executed by a processor, a computer device having the processor executes the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Emotional speech generating method and apparatus for controlling emotional intensity

    US20210090551A1