Speech synthesis method and related device

By obtaining and adjusting the emotional distribution in the target text, and generating speech synthesis audio with multiple emotional and emotional ups and downs, the problem of insufficient emotional expression in the prior art is solved and the auditory experience of speech synthesis is improved.

CN115440185BActive Publication Date: 2025-05-30SHANGHAI ZHENGDA XIMALAYA NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211083730.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-06
Publication Date
2025-05-30
Estimated Expiration
2042-09-06

AI Technical Summary

Technical Problem

The existing speech synthesis technology has shortcomings in emotional expression, especially in audio reading scenarios, where the lack of long-term TTS audio with emotional ups and downs leads to auditory fatigue.

Method used

By obtaining the emotional type distribution and emotional intensity distribution in the target text, and adjusting the emotional type distribution according to the emotional intensity distribution, the target emotional distribution is obtained, and finally a synthetic speech is generated based on the target emotional distribution.

Benefits of technology

The synthetic pronunciation has achieved a variety of emotional and emotional ups and downs, which enhances the emotional expression of speech synthesis and reduces auditory fatigue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115440185B_ABST
    Figure CN115440185B_ABST
Patent Text Reader

Abstract

In the speech synthesis method and related devices provided by the present application, the speech synthesis device obtains the emotional type distribution and emotional intensity distribution in the target text, adjusts the emotional type distribution according to the emotional intensity distribution to obtain the target emotional distribution; finally, according to the target emotional distribution, the synthetic speech of the target text is generated. Since the emotional type distribution represents the distribution ratio of multiple preset emotions in the target text, and the emotional intensity distribution represents the distribution ratio of multiple preset emotional intensities in the target text, therefore, after the emotional intensity distribution acts on the emotional type distribution, the synthetic speech in the audiobook scenario of a large number of dialogues has multiple emotions and emotional fluctuations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech synthesis, and more particularly, to a speech synthesis method and related devices. Background Art

[0002] Text-to-Speech (TTS) technology is a technology that automatically generates speech according to text and has currently been widely applied to scenarios such as voice assistants and news broadcasts. Although current TTS technology can already rival human voices in terms of timbre and sound quality, there are still certain deficiencies in emotional expression. Insufficient emotional expressiveness is one of the common problems in current TTS technology. Especially in the scenario of audiobook reading with a large amount of dialogue between characters, long-duration TTS audio lacking emotional fluctuations will cause auditory fatigue and make listeners lose the desire to continue listening. Summary of the Invention

[0003] To overcome at least one deficiency in the prior art, this application provides a speech synthesis method and related devices for synthesizing synthetic speech with multiple emotions and emotional fluctuations, specifically including:

[0004] In a first aspect, this application provides a speech synthesis method applied to a speech synthesis device, the method including:

[0005] Obtain the emotional type distribution and emotional intensity distribution in the target text, where the emotional type distribution represents the distribution ratio of multiple preset emotions in the target text, and the emotional intensity distribution represents the distribution ratio of multiple preset emotional intensities in the target text;

[0006] Adjust the emotional type distribution according to the emotional intensity distribution to obtain a target emotional distribution;

[0007] Generate synthetic speech of the target text according to the target emotional distribution.

[0008] In a second aspect, this application provides a speech synthesis device applied to a speech synthesis device, the speech synthesis device including:

[0009] An emotion analysis module for obtaining the emotional type distribution and emotional intensity distribution in the target text, where the emotional type distribution represents the distribution ratio of multiple preset emotions in the target text, and the emotional intensity distribution represents the distribution ratio of multiple preset emotional intensities in the target text;

[0010] An emotion adjustment module for adjusting the emotional type distribution according to the emotional intensity distribution to obtain a target emotional distribution;

[0011] A speech synthesis module, configured to generate a synthesized speech of the target text according to the target emotion distribution.

[0012] In a third aspect, the present application provides a computer-readable storage medium storing a computing program, which when executed by a processor, implements the speech synthesis method described above.

[0013] In a fifth aspect, the present application provides a speech synthesis device, which includes a processor and a memory. The memory stores a computer program, which when executed by the processor, implements the speech synthesis method described above.

[0014] Compared with the prior art, the present application has the following beneficial effects:

[0015] In the speech synthesis method and related devices provided by the present application, the speech synthesis device obtains the emotion type distribution and emotion intensity distribution in the target text, adjusts the emotion type distribution according to the emotion intensity distribution, and obtains the target emotion distribution; finally, according to the target emotion distribution, a synthesized speech of the target text is generated. Since the emotion type distribution represents the distribution ratio of multiple preset emotions in the target text, and the emotion intensity distribution represents the distribution ratio of multiple preset emotion intensities in the target text, therefore, after the emotion intensity distribution acts on the emotion type distribution, the synthesized speech in a large number of audiobook scenarios has multiple emotions and emotional fluctuations. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a schematic flowchart of the method provided by the embodiment of the present application;

[0018] Figure 2 It is a schematic overall framework diagram provided by the embodiment of the present application;

[0019] Figure 3 It is a schematic diagram of the device structure provided by the embodiment of the present application;

[0020] Figure 4 It is a schematic diagram of the device structure provided by the embodiment of the present application.

[0021] Icons: 101 - Sentiment Analysis Module; 102 - Sentiment Adjustment Module; 103 - Speech Synthesis Module; 201 - Memory; 202 - Processor; 203 - Communication Unit; 204 - System Bus. Detailed Implementation Manner

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. The components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations.

[0023] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0024] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0025] In the description of the present application, it should be noted that the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance. In addition, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0026] Research has found that although current speech synthesis technology can already rival that of real humans in terms of timbre and sound quality, there are still certain deficiencies in emotional expression. Insufficient emotional expressiveness is one of the common problems in current speech synthesis technology. Especially in the scenario of audiobook reading with a large amount of dialogue between characters, long-duration TTS audio lacking emotional fluctuations will cause auditory fatigue and make listeners lose the desire to continue listening.

[0027] For example, real human emotions are usually complex emotions composed of a combination of multiple emotions. The emotion labels obtained by existing speech synthesis technologies based on text prediction are usually single-category, such as "happy" or "surprised". This results in the inability to reflect complex emotions that contain two emotions simultaneously, such as "surprised and delighted". Therefore, there are certain deficiencies in the emotional expressiveness of current speech synthesis technologies.

[0028] For another example, there is a lack of control over emotional intensity. The emotional intensity of all synthesized audio with the emotion label "happy" is the same, lacking the ups and downs of emotions. When users judge the emotional information contained in the text, they will consider at least two aspects: certainty and intensity:

[0029] 1. Certainty: Certainty indicates how certain one is that the text expresses a certain emotion.

[0030] 2. Intensity: Intensity indicates the strength of the emotion expressed by the text.

[0031] Generally speaking, the higher the intensity, the higher the certainty; however, when the certainty is high, the intensity is not necessarily high. For example, in the sentence "He smiled faintly and said to the young man, 'Please show me the way.'", it can be very certain that this sentence contains a happy emotion from the phrase "smiled faintly", but the degree of happiness is weak. Existing technologies usually ignore this dimension of emotional intensity, resulting in a sense of incongruity sometimes when synthesizing audio due to overly high or low emotional intensity.

[0032] It should be noted that the defects existing in the above solutions of the prior art are all the results obtained by the inventors through practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by the embodiments of the present application below for the above problems should be the contributions made by the inventors to the present application during the invention and creation process, and should not be understood as the technical content known to those skilled in the art.

[0033] In view of this, this embodiment provides a speech synthesis method applied to a speech synthesis device. By obtaining the emotional type distribution and emotional intensity distribution in the target text, and using the emotional intensity distribution to make certain adjustments to the emotional type distribution to obtain the target emotional distribution, and obtaining the synthesized speech of the target text based on the target emotional distribution. Since the emotional type distribution represents the distribution ratio of multiple preset emotions in the target text, and the emotional intensity distribution represents the distribution ratio of multiple preset emotional intensities in the target text, therefore, after the emotional intensity distribution acts on the emotional type distribution, the synthesized speech in the scenario of audiobook reading of a large number of dialogues has multiple emotions and emotional ups and downs.

[0034] Among them, in some embodiments, the speech synthesis device may be a server. For example, the server may be a single server or a server group. The server group may be centralized or distributed (for example, the server may be a distributed system). In some embodiments, the server may be local or remote relative to the user terminal. In some embodiments, the server may be implemented on a cloud platform; by way of example only, the cloud platform may include private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, inter-cloud, multi-cloud, etc., or any combination thereof. In some embodiments, the server may be implemented on an electronic device having one or more components.

[0035] Exemplarily, when the speech synthesis device is a server, the user can access the server through the user terminal, upload articles, novels, stories to the server, so that the server converts them into synthesized speech.

[0036] Of course, in some other embodiments, the speech synthesis device may also be a user terminal. For example, by way of example only, it includes mobile terminals, tablet computers, laptop computers, desktop computers, etc. In some embodiments, the mobile terminal may include smart home devices, smart mobile devices, virtual reality devices, or augmented reality devices, etc., or any combination thereof. In some embodiments, the smart home device may include a smart TV, a smart speaker, etc. In some embodiments, the smart mobile device may include a smart phone, a personal digital assistant (PDA), a navigation device, etc., or any combination thereof.

[0037] Exemplarily, when the speech synthesis device is a smart speaker, the user can send articles, novels, stories, etc. to the smart speaker in text form, so that the smart speaker generates corresponding synthesized speech and plays it.

[0038] Based on the above related introduction, the following will be combined with Figure 1 elaborate on each step of the speech synthesis method in detail. However, it should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application. As Figure 1 shown, the method includes:

[0039] S101, obtaining the emotional type distribution and emotional intensity distribution in the target text.

[0040] Among them, the emotional type distribution represents the distribution ratio of multiple preset emotions in the target text, and the emotional intensity distribution represents the distribution ratio of multiple preset emotional intensities in the target text.

[0041] In related technologies, a variety of methods have been proposed to obtain the emotional type distribution and emotional intensity distribution possessed by the target text. In this embodiment, the emotional information contained in the target text, that is, the emotional type distribution and emotional intensity distribution, is obtained through a pre-trained text sentiment analysis model.

[0042] In order to obtain this text sentiment analysis model, this embodiment also provides a corresponding model training method. The model training method includes two links: corpus annotation and model training. The following will introduce these two links in detail.

[0043] Corpus annotation:

[0044] Book texts in different fields can be collected as training samples. When annotating, the annotator can see the complete context information and annotate the preset emotional types and preset emotional intensities expressed by all the dialogues in the text according to the context. Among them, the emotional category can be single-choice or multiple-choice. In this embodiment, 9 types of emotional category labels are used, namely happy, angry, sad, afraid, surprised, disgusted, vigilant, admired, and neutral.

[0045] For the emotional intensity, this embodiment uses 3 labels of preset emotional intensities, namely strong, medium, and weak. Based on these three labels of preset emotional intensities, each dialogue is simultaneously annotated by multiple annotators (for example, 5 annotators are selected). Finally, the annotation results of all annotators are accumulated to obtain the annotation of the emotional type distribution and emotional intensity distribution corresponding to each dialogue.

[0046] Model training:

[0047] The tasks in the training process include two subtasks: emotional category prediction and emotional intensity prediction. The dialogue text with the annotation of emotional type distribution and emotional intensity distribution and its context are used as the model input at the same time. Based on a large-scale pre-trained language model (any pre-trained model can be used, for example, BERT), fine-tuning is performed. The pre-trained language model is used as a semantic feature extractor. The output layer of the subtask predicts the corresponding distribution according to the semantic features extracted from the dialogue text and its context. The KL divergence loss function is used to measure the difference degree between the predicted distribution and the annotated distribution to guide the training process.

[0048] In this way, a trained text sentiment analysis model is obtained. Based on this text sentiment analysis model, the target text and its context are input into the text sentiment analysis model to obtain the sentiment type distribution and sentiment intensity distribution of the target text. It should be noted that for non-dialogue texts, in this embodiment, a neutral sentiment type and a neutral intensity are uniformly used.

[0049] Based on the above introduction to the sentiment type distribution and sentiment intensity distribution, continue to refer to Figure 1 , the speech synthesis method further includes:

[0050] S102. Adjust the sentiment type distribution according to the sentiment intensity distribution to obtain a target sentiment distribution.

[0051] In this embodiment, in order to obtain a synthesized speech with multiple emotions and emotional fluctuations, it is necessary to apply the sentiment intensity distribution to the sentiment type distribution. It can be understood that if the overall sentiment intensity of the target text is low, the proportion of non-neutral emotions in the sentiment type distribution needs to be weakened; conversely, the proportion of non-neutral emotions needs to be increased. Therefore, step S102 includes the following specific implementation manners:

[0052] S102-1. Obtain a sentiment intensity coefficient according to the sentiment intensity distribution.

[0053] Among them, the sentiment intensity coefficient is used to represent the overall sentiment intensity of the target text. In this regard, in this embodiment, the speech synthesis device can weight the proportion corresponding to each preset sentiment intensity in the sentiment intensity distribution according to the weights of multiple preset sentiment intensities to obtain the sentiment intensity coefficient.

[0054] Exemplarily, assuming that the sentiment intensity distribution of the target text is P = {strong: 0.3, medium: 0.6, weak: 0.1}, then according to the weights V = {strong: 1.0, medium: 0.5, weak: 0.25} of multiple preset sentiment intensities, the sentiment intensity coefficient I can be obtained according to the following calculation method:

[0055] I = P 强 * V 强 + P 中 * V 中 + P 弱 * V 弱 = 0.3 * 1.0 + 0.6 * 0.5 + 0.1 * 0.25 = 0.625

[0056] Among them, the weights of multiple preset emotional intensities are empirical values, which can be set according to the matching degree between the training samples labeled during the training of the text sentiment analysis model and the emotional intensities of the emotional recordings in the emotional audio library. The specific setting principle is that if the overall emotional intensity of the training samples is weak, while the emotional intensities of the emotional recordings in the emotional audio library are strong, then the weights of the preset emotional intensities "strong", "medium", and "weak" are reduced; otherwise, the weights of the preset emotional intensities "strong", "medium", and "weak" are increased. In this way, the influence of the emotional intensity of the emotional recordings in the emotional audio library on the emotional intensity of the synthesized speech is reduced.

[0057] It should be understood here that the emotional recordings in the emotional audio library are obtained by recording the single-category emotional audio of at least 1 recorder. For example, the emotional audio library includes emotional recordings with emotional information such as happiness, anger, sadness, fear, surprise, disgust, vigilance, admiration, and neutrality. The emotional recordings in the emotional audio library are used to obtain the feature vectors of the preset emotional types for subsequent use in synthesizing speech.

[0058] S102-2. Adjust the proportion of non-neutral emotions in the emotion type distribution through the emotion intensity coefficient, and adjust the proportion of neutral emotions in the emotion type distribution according to the adjusted proportion of non-neutral emotions to obtain the target emotion distribution.

[0059] As introduced in the above embodiments, if the overall emotional intensity of the target text is low, then it is necessary to weaken the proportion of non-neutral emotions in the emotion type distribution. Therefore, for step S102-2, this embodiment may include the following specific implementation manners for reducing the proportion of non-neutral emotions in the emotion type distribution and increasing the proportion of neutral emotions in the emotion type distribution:

[0060] S102-2-1. Obtain the total proportion of non-neutral emotions in the emotion type distribution.

[0061] S102-2-2. If the total proportion is greater than the emotion intensity coefficient, then reduce the proportion of non-neutral emotions in the emotion type distribution through the emotion intensity coefficient and increase the proportion of neutral emotions in the emotion type distribution to obtain the target emotion distribution.

[0062] S102-2-3. If the total proportion is less than the emotion intensity coefficient, then increase the proportion of non-neutral emotions in the emotion type distribution through the emotion intensity coefficient and reduce the proportion of neutral emotions in the emotion type distribution to obtain the target emotion distribution.

[0063] As an alternative implementation, the speech synthesis device may obtain the ratio between the total proportion and the emotional intensity coefficient as the emotional adjustment coefficient; then, divide the proportion of each non-neutral emotion by the emotional adjustment coefficient to obtain the adjusted proportion of each non-neutral emotion; finally, according to the adjusted proportion of each non-neutral emotion, take the remaining proportion as the adjusted proportion of the neutral emotion, so as to obtain the target emotion distribution.

[0064] Exemplarily, continuing to assume that the emotional adjustment coefficient I = 0.625 and the emotion type distribution is: {happy: 0.3, angry: 0.0, sad: 0.0, afraid: 0.0, surprised: 0.5, disgusted: 0.0, vigilant: 0.0, admired: 0.15, neutral: 0.05}, then the total proportion of non-neutral emotions is:

[0065] 0.3 + 0.5 + 0.15 = 0.95 > 0.625

[0066] Therefore, take the ratio between the total proportion and the emotional intensity coefficient as the emotional adjustment coefficient:

[0067] 0.95 / 0.625 = 1.52

[0068] Since the emotional adjustment coefficient at this time is greater than 1, therefore, after dividing the proportion of each non-neutral emotion by the emotional adjustment coefficient, it can play a role in reducing the proportion of non-neutral emotions:

[0069] {happy: (0.3 / 1.52) = 0.1974, angry: 0.0, sad: 0.0, afraid: 0.0, surprised: (0.5 / 1.52) = 0.3289, disgusted: 0.0, vigilant: 0.0, admired: (0.15 / 1.52) = 0.0987, neutral: (1 - 0.1974 - 0.3289 - 0.0987) = 0.375}.

[0070] Of course, if the emotional adjustment coefficient at this time is less than 1, then after dividing the proportion of each non-neutral emotion by the emotional adjustment coefficient, it can play a role in increasing the proportion of non-neutral emotions; and once the proportion of non-neutral emotions is increased, it means that the proportion of neutral emotions needs to be reduced. This embodiment will not elaborate on this.

[0071] Based on the above introduction of the target emotion distribution, continue to refer to Figure 1 , the speech synthesis method further includes:

[0072] S103, generate the synthetic speech of the target text according to the target emotion distribution.

[0073] In this regard, in this embodiment, the speech synthesis device integrates the speech features of multiple preset emotion types into the target text according to the target emotion distribution, so as to obtain a synthesized speech with multiple emotions and emotional fluctuations. Therefore, as an optional implementation manner, step S103 includes:

[0074] S103-1, taking the proportion of each preset emotion type in the target emotion distribution as a weight, weighting the feature vectors of each preset emotion type to obtain a mixed emotion feature.

[0075] Among them, the feature vector of each preset emotion type is obtained by using an embedding network layer to extract features from the emotion recordings in the emotion audio library. The calculation expression of the mixed emotion feature Emix is:

[0076] Emix = P 高兴 *E 高兴 +P 愤怒 *E 愤怒 +P 悲伤 *E 悲伤 +P 害怕 *E 害怕 +P 吃惊 *E 吃惊 +P 厌恶 *E 厌恶 +P 警惕 *E 警惕 +P 钦佩 *E 钦佩 +P 中性 *E 中性

[0077] In the formula, E 高兴 represents the feature vector obtained by extracting features from the emotion audio containing "happy" emotion information, and P 高兴 represents the proportion of the preset emotion type of "happy" in the target emotion distribution; E 愤怒 represents the feature vector obtained by extracting features from the emotion audio containing "angry" emotion information, and P 愤怒 represents the proportion of the preset emotion type of "angry" in the target emotion distribution; and so on, the physical meanings of other symbols in the expression can be obtained. This embodiment will not be elaborated further.

[0078] S103-2, inputting the mixed emotion feature and the linguistic feature of the target text into the speech synthesis model together to obtain the synthesized speech of the target text.

[0079] That is, in this embodiment, the mixed emotion feature is used as the emotion representation, and the linguistic feature is used as the phoneme representation and timbre representation, and they are input into the speech synthesis model together to obtain the synthesized speech.

[0080] As an optional implementation, the speech synthesis model includes an acoustic model and a vocoder model. Among them, the emotional representation obtained by adding an emotional embedding layer to any mainstream speech synthesis acoustic model (for example, the DurIAN model) can be used as the input of the decoder together with the phoneme representation and the timbre representation to guide the training process, so as to obtain the acoustic model in this embodiment. Similarly, the vocoder model can be trained using any mainstream vocoder model (for example, the HiFi-GAN model). Therefore, based on the acoustic model and the vocoder model in this embodiment, step S103-2 includes the following specific implementation manners:

[0081] S103-2-1, input the mixed emotional features and the linguistic features into the acoustic model together to obtain the emotion-rich acoustic features.

[0082] S103-2-2, input the emotion-rich acoustic features into the vocoder model to obtain the synthetic speech of the target text.

[0083] To enable those skilled in the art to use the content of this application, a process framework as shown is provided to outline the entire above implementation manner. As shown, the entire framework includes an emotion branch, a linguistic branch, and a speech synthesis branch. Figure 2 shown Figure 2 As shown, the entire framework includes an emotion branch, a linguistic branch, and a speech synthesis branch.

[0084] Emotion branch:

[0085] For the target text, the speech synthesis device obtains the emotion type distribution and the emotion intensity distribution in the target text through the emotion analysis model; then, in the fusion calculation link, the emotion intensity distribution acts on the emotion type distribution to obtain the target emotion distribution; finally, interpolation calculation (weighting) is performed on the feature vectors of multiple preset emotions and the target emotion distribution to obtain the mixed emotional features of multiple preset emotions.

[0086] Linguistic branch:

[0087] For the target text, the speech synthesis device extracts the linguistic features of it to obtain the linguistic features in the target text.

[0088] Speech synthesis branch:

[0089] The speech synthesis device inputs the mixed emotional features of the target text and the linguistic features in the target text into the acoustic model together to obtain the acoustic features of the target text, and then inputs the acoustic features of the target text into the vocoder model, so as to obtain the synthetic language of the target text. In this way, the synthetic speech in the audio reading scenario with a large number of dialogues has multiple emotions and emotional fluctuations.

[0090] In summary, the speech synthesis device obtains the emotion type distribution and emotion intensity distribution in the target text, and adjusts the emotion type distribution according to the emotion intensity distribution to obtain the target emotion distribution; finally, according to the target emotion distribution, the synthesized speech of the target text is generated. Since the emotion type distribution represents the distribution ratio of multiple preset emotions in the target text, and the emotion intensity distribution represents the distribution ratio of multiple preset emotion intensities in the target text, after the emotion intensity distribution is applied to the emotion type distribution, the synthesized speech in the audio reading scene with a large amount of dialogue has multiple emotions and changes in emotion fluctuations.

[0091] Based on the same inventive concept as the above speech synthesis method, this embodiment also provides a speech synthesis device applied to a speech synthesis device. The speech synthesis device includes at least one software function module that can be stored in the memory 201 in software form or solidified in the operating system (OS) of the speech synthesis device. The processor 202 in the speech synthesis device is used to execute the executable module stored in the memory 201. For example, the software function modules and computer programs included in the speech synthesis device. Please refer to Figure 3 , from the functional point of view, the speech synthesis device may include:

[0092] The sentiment analysis module 101 is used to obtain the sentiment type distribution and sentiment intensity distribution in the target text, wherein the sentiment type distribution represents the distribution ratio of multiple preset sentiments in the target text, and the sentiment intensity distribution represents the distribution ratio of multiple preset sentiment intensities in the target text.

[0093] In this embodiment, the sentiment analysis module 101 is used to implement Figure 1 For a detailed description of the sentiment analysis model, please refer to the detailed description of step S101.

[0094] The emotion adjustment module 102 is used to adjust the emotion type distribution according to the emotion intensity distribution to obtain the target emotion distribution.

[0095] In this embodiment, the emotion adjustment module 102 is used to implement Figure 1 For the detailed description of step S102 in the emotion adjustment module 102, please refer to the detailed description of step S102.

[0096] The speech synthesis module 103 is used to generate synthesized speech of the target text according to the target emotion distribution.

[0097] In this embodiment, the speech synthesis module 103 is used to implement Figure 1 For the detailed description of the speech synthesis module 103, please refer to the detailed description of step S103.

[0098] It is also worth noting that since the speech synthesis device and the speech synthesis method have the same inventive concept, the above emotion analysis module 101, emotion adjustment module 102 and speech synthesis module 103 can also be used to implement other steps or sub-steps of the speech synthesis method of this embodiment, which will not be repeated in this embodiment.

[0099] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0100] It should also be understood that if the above implementation is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application.

[0101] Therefore, this embodiment further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the speech synthesis method provided in this embodiment is implemented. The computer-readable storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0102] This embodiment further provides a speech synthesis device, which may include a processor 202 and a memory 201. The processor 202 and the memory 201 may communicate via a system bus 204. In addition, the memory 201 stores a computer program, and the processor implements the speech synthesis method provided in this embodiment by reading and executing the computer program corresponding to the above implementation in the memory 201.

[0103] like Figure 4 As shown, the speech synthesis device may further include a communication unit 203, and the memory 201, the processor 202, and the communication unit 203 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines.

[0104] Among them, the memory 201 can be an information recording device based on any electronic, magnetic, optical or other physical principles, and is used to record execution instructions, data, etc. In some embodiments, the memory 201 can be, but is not limited to, a volatile memory, a non-volatile memory, a storage drive, etc.

[0105] Among them, by way of example only, the volatile memory can be a Random Access Memory (RAM). The non-volatile memory can be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electric Erasable Programmable Read-Only Memory (EEPROM), a flash memory, etc.; the storage drive can be a disk drive, a solid state drive, any type of storage disk (such as an optical disk, a DVD, etc.), or a similar storage medium, or a combination thereof, etc.

[0106] The communication unit 203 is used to send and receive data through a network. In some embodiments, the network can include a wired network, a wireless network, an optical fiber network, a telecommunication network, an intranet, the Internet, a Local Area Network (LAN), a Wide Area Network (WAN), a Wireless Local Area Networks (WLAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), a Public Switched Telephone Network (PSTN), a Bluetooth network, a ZigBee network, or a Near Field Communication (NFC) network, etc., or any combination thereof. In some embodiments, the network can include one or more network access points. For example, the network can include a wired or wireless network access point, such as a base station and / or a network switching node, and one or more components of the service request processing system can be connected to the network through the access point to exchange data and / or information.

[0107] The processor 202 may be an integrated circuit chip with the ability to process signals, and the processor may include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, the above-mentioned processor may include a Central Processing Unit (CPU), an Application Specific Integrated Circuit (ASIC), an Application Specific Instruction-set Processor (ASIP), a Graphics Processing Unit (GPU), a Physics Processing Unit (PPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), a controller, a microcontroller unit, a Reduced Instruction Set Computing (RISC), or a microprocessor, etc., or any combination thereof.

[0108] It should be understood that the devices and methods disclosed in the above embodiments may also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

[0109] As described above, these are only various embodiments of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A speech synthesis method, characterized in that, applied to a speech synthesis device, the method includes: Obtain the emotional type distribution and emotional intensity distribution in the target text, where the emotional type distribution represents the distribution ratio of multiple preset emotions in the target text, and the emotional intensity distribution represents the distribution ratio of multiple preset emotional intensities in the target text; Adjust the emotional type distribution according to the emotional intensity distribution to obtain a target emotional distribution; Generate the synthesized speech of the target text according to the target emotional distribution.

2. The speech synthesis method according to claim 1, characterized in that, The adjusting the emotional type distribution according to the emotional intensity distribution to obtain a target emotional distribution includes: Obtain an emotional intensity coefficient according to the emotional intensity distribution, where the emotional intensity coefficient is used to represent the overall emotional intensity of the target text; Adjust the proportion of non-neutral emotions in the emotional type distribution through the emotional intensity coefficient, and adjust the proportion of neutral emotions in the emotional type distribution according to the adjusted proportion of non-neutral emotions to obtain the target emotional distribution.

3. The speech synthesis method according to claim 2, characterized in that, The adjusting the proportion of non-neutral emotions in the emotional type distribution through the emotional intensity coefficient, and adjusting the proportion of neutral emotions in the emotional type distribution according to the adjusted proportion of non-neutral emotions to obtain the target emotional distribution includes: Obtain the total proportion of non-neutral emotions in the emotional type distribution; If the total proportion is greater than the emotional intensity coefficient, reduce the proportion of non-neutral emotions in the emotional type distribution through the emotional intensity coefficient, and increase the proportion of neutral emotions in the emotional type distribution to obtain the target emotional distribution; If the total proportion is less than the emotional intensity coefficient, increase the proportion of non-neutral emotions in the emotional type distribution through the emotional intensity coefficient, and reduce the proportion of neutral emotions in the emotional type distribution to obtain the target emotional distribution.

4. The speech synthesis method according to claim 3, characterized in that, The reducing the proportion of non-neutral emotions in the emotional type distribution through the emotional intensity coefficient, and increasing the proportion of neutral emotions in the emotional type distribution to obtain the target emotional distribution includes: Obtain the ratio between the total proportion and the emotional intensity coefficient as an emotional adjustment coefficient; Divide the proportion of each non-neutral emotion by the emotional adjustment coefficient to obtain the adjusted proportion of each non-neutral emotion; According to the adjusted proportion of each non-neutral emotion, use the remaining proportion as the adjusted proportion of the neutral emotion, thereby obtaining the target emotional distribution.

5. The speech synthesis method according to claim 2, characterized in that, The obtaining an emotional intensity coefficient according to the emotional intensity distribution includes: Weight the proportion corresponding to each preset emotional intensity in the emotional intensity distribution according to the respective weights of the multiple preset emotional intensities to obtain the emotional intensity coefficient.

6. The speech synthesis method according to claim 1, It is characterized in that generating a synthetic speech of the target text according to the target emotion distribution includes: using the proportion of each preset emotion type in the target emotion distribution as a weight to weight the feature vectors of each preset emotion type to obtain a mixed emotion feature; inputting the mixed emotion feature and the linguistic feature of the target text into a speech synthesis model to obtain a synthetic speech of the target text.

7. The speech synthesis method according to claim 6, It is characterized in that the speech synthesis model includes an acoustic model and a vocoder model, and inputting the mixed emotion feature and the linguistic feature of the target text into the speech synthesis model to obtain a synthetic speech of the target text includes: inputting the mixed emotion feature and the linguistic feature into the acoustic model to obtain an emotion-rich acoustic feature; inputting the emotion-rich acoustic feature into the vocoder model to obtain a synthetic speech of the target text.

8. A speech synthesis device, It is characterized in that applied to a speech synthesis device, the speech synthesis device includes: an emotion analysis module, configured to obtain an emotion type distribution and an emotion intensity distribution in a target text, where the emotion type distribution represents the distribution ratio of multiple preset emotions in the target text, and the emotion intensity distribution represents the distribution ratio of multiple preset emotion intensities in the target text; an emotion adjustment module, configured to adjust the emotion type distribution according to the emotion intensity distribution to obtain a target emotion distribution; a speech synthesis module, configured to generate a synthetic speech of the target text according to the target emotion distribution.

9. A computer-readable storage medium, It is characterized in that the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the speech synthesis method according to any one of claims 1-7 is implemented.

10. A speech synthesis device, It is characterized in that the speech synthesis device includes a processor and a memory, the memory stores a computer program, and when the computer program is executed by the processor, the speech synthesis method according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Speech synthesis model training method, speech synthesis method and related device

    CN115312027A