Model processing method and device, emotional speech synthesis method and device

By acquiring and adjusting the voice data of the target vocalization object, the speech synthesis model is trained to predict the spectral map and generate time domain waveforms, the problem of insufficient emotional expression in the prior art is solved, and the synthesis and application of emotional speech data is realized.

CN114724540BActive Publication Date: 2025-09-05ALIBABA GROUP HOLDING LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202011543098.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-21
Publication Date
2025-09-05
Estimated Expiration
2040-12-21

AI Technical Summary

Technical Problem

Existing speech synthesis techniques are difficult to synthesize speech with emotional expressiveness and cannot clearly perceive the emotions of the vocal object.

Method used

By obtaining multiple emotional voice data of the target vocal object, adjusting the sound elements and combining them into an emotional voice data set, the speech synthesis model is trained to predict the spectral map and generate time domain waveforms, and combining emotional intensity adjustment to achieve emotional voice synthesis.

Benefits of technology

The synthesis of emotional voice data is realized, and the output of emotional expressive voice data can be based on text information and emotional markers. It is suitable for live broadcasts, e-books, and video dubbing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114724540B_ABST
    Figure CN114724540B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a method and device for processing speech data, a method and device for processing models, and a method and device for synthesizing emotional speech. Among them, by obtaining multiple first emotional speech data of a target sound-making object and adjusting the target sound elements of at least one first emotional speech data, the second emotional speech data is obtained, so that the multiple first emotional speech data and the second emotional speech data are merged into an emotional speech data set of the target sound-making object. Afterwards, the target identity information of the target sound-making object, as well as the lines and emotional tags corresponding to the emotional speech data samples in the emotional speech data set can be used as input, and the emotional speech data samples can be used as training labels to train the speech synthesis model to be trained to obtain an emotional speech synthesis model. Thereafter, in the application stage, the emotional speech synthesis model can synthesize speech data with emotional expressiveness based on the input text information and emotional tags.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of speech synthesis technology, specifically, to methods and devices for processing speech data, methods and devices for model processing, methods and devices for emotional speech synthesis, methods and devices for emotional speech synthesis based on live broadcast, methods and devices for emotional speech synthesis based on e-books, and methods and devices for emotional speech synthesis based on videos. Background Art

[0002] Speech synthesis technology has evolved over decades, progressing through stages of understanding, naturalness, and expressiveness. However, current speech synthesis technology often struggles to synthesize emotionally expressive speech. Emotional expressiveness typically refers to the ability to clearly perceive the speaker's emotions, such as excitement, sadness, or neutrality, after hearing a sound.

[0003] Therefore, there is an urgent need for a reasonable and reliable solution to synthesize emotionally expressive speech. Summary of the Invention

[0004] The embodiments of this specification provide methods and devices for processing voice data, methods and devices for model processing, methods and devices for emotional speech synthesis, methods and devices for emotional speech synthesis based on live broadcast, methods and devices for emotional speech synthesis based on e-books, and methods and devices for emotional speech synthesis based on videos.

[0005] In a first aspect, an embodiment of the present specification provides a method for processing voice data, including: obtaining multiple first emotion voice data of a target voice-producing object, the multiple first emotion voice data corresponding to multiple lines, the multiple lines corresponding to at least one emotion tag, wherein the first emotion voice data is obtained by recording the sound emitted by the target voice-producing object when reading the corresponding lines; adjusting the target sound element of at least one first emotion voice data to obtain second emotion voice data; merging the multiple first emotion voice data and the second emotion voice data into an emotion voice data set of the target voice-producing object.

[0006] In some embodiments, the target sound elements include speaking rate and / or intonation.

[0007] In some embodiments, the dialogue sentences include dialogues from any of the following works: literary works, dramatic works, and film and television works.

[0008] In some embodiments, the at least one sentiment tag includes at least one of: neutral, positive sentiment, negative sentiment.

[0009] In some embodiments, the positive emotion includes at least one of the following: excitement, relief, happiness, and admiration; the negative emotion includes at least one of the following: sadness, anger, disgust, and fear.

[0010] In some embodiments, before obtaining multiple first emotional voice data of the target sound-making object, the method also includes: obtaining at least one text; for an emotional tag in the at least one emotional tag, extracting multiple lines of dialogue with the emotion indicated by the emotional tag from the at least one text; providing the extracted lines to the target sound-making object so that the target sound-making object reads the extracted lines, thereby obtaining the multiple first emotional voice data.

[0011] On the second aspect, an embodiment of this specification provides a model processing method, including: obtaining target identity information and an emotional speech data set of a target sound-making object, as well as lines and emotional tags corresponding to the emotional speech data samples in the emotional speech data set; using the target identity information, the lines and emotional tags as input, and the emotional speech data samples as training labels, to train a speech synthesis model to be trained, and obtain an emotional speech synthesis model.

[0012] In some embodiments, the speech synthesis model to be trained is pre-trained in the following manner: taking the sample identity information and text information of at least one sample sound-making object as input, taking the speech data of the sample sound-making object reading the text information as a training label, and training the initial speech synthesis model, wherein the sample sound-making object is different from the target sound-making object.

[0013] In some embodiments, the speech synthesis model to be trained includes a spectrogram prediction network and a vocoder, and the first processing process of the speech synthesis model to be trained includes: using the spectrogram prediction network to predict a spectrogram based on the input target identity information, lines and emotional tags; using the vocoder to generate a time domain waveform based on the spectrogram predicted by the spectrogram prediction network.

[0014] In some embodiments, the training of the speech synthesis model to be trained includes: determining the prediction loss based on the time domain waveform and the emotional speech data sample, and adjusting the network parameters in the spectrum prediction network with the goal of reducing the prediction loss.

[0015] In some embodiments, the spectrum prediction network is associated with an emotion intensity coefficient corresponding to at least one emotion marker, and the emotion intensity coefficient is used for emotion intensity adjustment; and in the application stage of the emotion speech synthesis model, the second processing process of the emotion speech synthesis model includes: using the spectrum prediction network to adjust the emotion intensity according to the emotion intensity coefficient corresponding to the input emotion marker.

[0016] In some embodiments, the spectrogram prediction network includes an encoder and a decoder; and using the spectrogram prediction network to predict a spectrogram based on input target identity information, lines and emotional tags includes: using the encoder to convert the input target identity information, lines and emotional tags into vectors respectively, and splicing the converted vectors to obtain a spliced ​​vector; using the decoder to predict a spectrogram based on the spliced ​​vector.

[0017] In some embodiments, the encoder includes an emotion tag embedding module, an identity embedding module and a character encoding module; and using the encoder to convert the input target identity information, lines and emotion tags into vectors respectively, including: using the emotion tag embedding module to map the input emotion tag into an emotion embedding vector; using the identity embedding module to map the input target identity information into an identity embedding vector; using the character encoding module to map the input lines into a character embedding vector, and encoding the character embedding vector to obtain a character encoding vector.

[0018] In some embodiments, the emotion tag embedding module associates an emotion intensity coefficient corresponding to at least one emotion tag, and the emotion intensity coefficient is used for emotion intensity adjustment; and in the application stage of the emotion speech synthesis model, the second processing process of the emotion speech synthesis model includes: using the emotion tag embedding module, after mapping the input emotion tag into an emotion embedding vector, the product of the emotion embedding vector and the emotion intensity coefficient corresponding to the emotion tag is determined as the emotion embedding vector after emotion intensity adjustment.

[0019] In some embodiments, the spectrogram comprises a mel-frequency spectrogram.

[0020] In the third aspect, an embodiment of this specification provides an emotional speech synthesis method, including: obtaining text information of the speech to be synthesized and its corresponding emotional tag; inputting the text information and the emotional tag into an emotional speech synthesis model trained using the method described in any implementation method in the second aspect, so that the emotional speech synthesis model outputs synthesized emotional speech data.

[0021] In the fourth aspect, an embodiment of this specification provides an emotional speech synthesis method, which is applied to a client, including: obtaining text information of the speech to be synthesized and its corresponding emotional tag; sending the text information and the emotional tag to the speech synthesis end, so that the speech synthesis end inputs the text information and the emotional tag into an emotional speech synthesis model trained using the method described in any implementation method in the second aspect, so that the emotional speech synthesis model outputs synthesized emotional speech data.

[0022] In the fifth aspect, an embodiment of this specification provides an emotional speech synthesis method based on live broadcast, which is applied to the anchor client, including: obtaining the dubbing text of the virtual anchor of the live broadcast, and the emotional tag corresponding to the dubbing text; sending the dubbing text and the emotional tag to the server, so that the server inputs the dubbing text and the emotional tag into the emotional speech synthesis model trained by the method described in any implementation method in the second aspect, so that the emotional speech synthesis model outputs synthesized emotional speech data; and providing the emotional speech data to the corresponding audience client via the server.

[0023] In the sixth aspect, an embodiment of this specification provides an emotional speech synthesis method based on an e-book, comprising: obtaining a target text in the e-book, and an emotional tag corresponding to the target text; inputting the target text and the emotional tag into an emotional speech synthesis model trained using the method described in any implementation method in the second aspect, so that the emotional speech synthesis model outputs synthesized emotional speech data; and providing the emotional speech data based on an e-book client.

[0024] In the seventh aspect, an embodiment of this specification provides a video-based emotional speech synthesis method, including: obtaining the dubbing text of the video to be dubbed, and the emotional tags corresponding to the dubbing text; inputting the dubbing text and the emotional tags into an emotional speech synthesis model trained using the method described in any implementation method in the second aspect, so that the emotional speech synthesis model outputs synthesized emotional speech data; and providing the emotional speech data based on a video client.

[0025] In an eighth aspect, an embodiment of the present specification provides a speech synthesis model, comprising: a spectrogram prediction network for predicting a spectrogram based on the target identity information of the input target sound-making object, and the lines and emotional tags corresponding to the emotional speech data samples of the target sound-making object; and a vocoder for generating a time domain waveform based on the spectrogram predicted by the spectrogram prediction network.

[0026] In some embodiments, the spectrum prediction network is associated with an emotion intensity coefficient corresponding to at least one emotion tag, and the emotion intensity coefficient is used for emotion intensity adjustment; and in the model application stage, the spectrum prediction network is also used to: adjust the emotion intensity according to the emotion intensity coefficient corresponding to the input emotion tag.

[0027] In some embodiments, the spectrogram prediction network includes: an encoder for converting input target identity information, dialogue sentences, and emotional tags into vectors respectively, and splicing the converted vectors to obtain a spliced ​​vector; a decoder for predicting a spectrogram based on the spliced ​​vector.

[0028] In some embodiments, the encoder includes: an emotion tag embedding module for mapping the input emotion tag into an emotion embedding vector; an identity embedding module for mapping the input target identity information into an identity embedding vector; a character encoding module for mapping the input dialogue sentence into a character embedding vector, and encoding the character embedding vector to obtain a character encoding vector.

[0029] In some embodiments, the emotion tag embedding module associates at least one emotion tag with an emotion intensity coefficient corresponding to each emotion tag, and the emotion intensity coefficient is used for emotion intensity adjustment; and in the model application stage, the emotion tag embedding module is also used to: after mapping the input emotion tag into an emotion embedding vector, the product of the emotion embedding vector and the emotion intensity coefficient corresponding to the emotion tag is determined as the emotion embedding vector after emotion intensity adjustment.

[0030] In the ninth aspect, an embodiment of the present specification provides a voice data processing device, comprising: an acquisition unit, configured to acquire multiple first emotion voice data of a target sound-making object, the multiple first emotion voice data corresponding to multiple lines, the multiple lines corresponding to at least one emotion tag, wherein the first emotion voice data is obtained by recording the sound emitted when the target sound-making object reads the corresponding lines; an adjustment unit, configured to adjust the target sound element of at least one first emotion voice data to obtain second emotion voice data; a generation unit, configured to merge the multiple first emotion voice data and the second emotion voice data into an emotion voice data set of the target sound-making object.

[0031] In the tenth aspect, an embodiment of the present specification provides a model processing device, comprising: an acquisition unit, configured to acquire target identity information and an emotional speech data set of a target sound-making object, as well as lines and emotional tags corresponding to the emotional speech data samples in the emotional speech data set; a model training unit, configured to take the target identity information, the lines and emotional tags as input, and the emotional speech data samples as training labels, to train a speech synthesis model to be trained, and obtain an emotional speech synthesis model.

[0032] In the eleventh aspect, an embodiment of the present specification provides an emotional speech synthesis device, comprising: an acquisition unit, configured to acquire text information of the speech to be synthesized and its corresponding emotional tag; a speech synthesis unit, configured to input the text information and the emotional tag into an emotional speech synthesis model trained using the method described in any implementation method in the second aspect, so that the emotional speech synthesis model outputs synthesized emotional speech data.

[0033] In the twelfth aspect, an embodiment of the present specification provides an emotional speech synthesis device, which is applied to a client, comprising: an acquisition unit, configured to acquire text information of the speech to be synthesized and its corresponding emotional tag; a sending unit, configured to send the text information and the emotional tag to the speech synthesis end, so that the speech synthesis end inputs the text information and the emotional tag into an emotional speech synthesis model trained using the method described in any implementation method in the second aspect, so that the emotional speech synthesis model outputs synthesized emotional speech data.

[0034] In the thirteenth aspect, an embodiment of the present specification provides an emotional speech synthesis device based on live broadcast, which is applied to an anchor client, including: an acquisition unit, configured to acquire the dubbing text of the virtual anchor of the live broadcast, and the emotional tag corresponding to the dubbing text; a sending unit, configured to send the dubbing text and the emotional tag to the server, so that the server inputs the dubbing text and the emotional tag into an emotional speech synthesis model trained by the method described in any implementation method in the second aspect, so that the emotional speech synthesis model outputs synthesized emotional speech data; a processing unit, configured to provide the emotional speech data to the corresponding audience client via the server.

[0035] In the fourteenth aspect, an embodiment of the present specification provides an emotional speech synthesis device based on an e-book, comprising: an acquisition unit, configured to acquire a target text in the e-book, and an emotional tag corresponding to the target text; a speech synthesis unit, configured to input the target text and the emotional tag into an emotional speech synthesis model trained using the method described in any implementation method in the second aspect, so that the emotional speech synthesis model outputs synthesized emotional speech data; a processing unit, configured to provide the emotional speech data based on an e-book client.

[0036] In the fifteenth aspect, an embodiment of the present specification provides a video-based emotional speech synthesis device, comprising: an acquisition unit, configured to acquire the dubbing text of the video to be dubbed, and the emotional tag corresponding to the dubbing text; a speech synthesis unit, configured to input the dubbing text and the emotional tag into an emotional speech synthesis model trained using the method described in any implementation method in the second aspect, so that the emotional speech synthesis model outputs synthesized emotional speech data; a processing unit, configured to provide the emotional speech data based on a video client.

[0037] In the sixteenth aspect, an embodiment of this specification provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, the computer is caused to execute the method described in any one of the implementation methods in the first to seventh aspects.

[0038] In the seventeenth aspect, an embodiment of this specification provides a computing device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the method described in any one of the implementation methods in the first to seventh aspects.

[0039] In an eighteenth aspect, an embodiment of this specification provides a computer program, which, when executed in a computer, causes the computer to execute the method described in any one of the implementations in the first to seventh aspects.

[0040] The method and apparatus provided by the above-mentioned embodiments of this specification obtain multiple first emotional speech data of the target sound-making object, and then adjust the target sound elements of at least one first emotional speech data to obtain second emotional speech data, so as to merge the multiple first emotional speech data and the second emotional speech data into a larger emotional speech data set of the target sound-making object. Afterwards, by taking the target identity information of the target sound-making object and the lines and emotional tags corresponding to the emotional speech data samples in the emotional speech data set as input, and using the emotional speech data samples as training labels, the speech synthesis model to be trained is trained, and an emotional speech synthesis model with better emotional speech synthesis effect can be obtained. Thereafter, in the application stage, the emotional speech synthesis model can synthesize emotionally expressive speech data based on the input text information and emotional tags. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0042] Figure 1 is an exemplary system architecture diagram to which some embodiments of this specification may be applied;

[0043] Figure 2 is a flowchart of an embodiment of a method for processing voice data according to the present specification;

[0044] Figure 3 is a flow chart of an embodiment of a model processing method according to the present specification;

[0045] Figure 4a is a schematic diagram of the first processing step of the speech synthesis model to be trained;

[0046] Figure 4b This is a schematic diagram of the processing of the spectrum prediction network;

[0047] Figure 4c It is a schematic diagram of the encoder's processing;

[0048] Figure 5 is a flowchart of an embodiment of an emotional speech synthesis method according to the present specification;

[0049] Figure 6 is a schematic diagram of an embodiment of an emotional speech synthesis method according to the present specification;

[0050] Figure 7 This is a schematic diagram of the emotional speech synthesis method in a live broadcast scenario;

[0051] Figure 8 This is a schematic diagram of the emotional speech synthesis method in the audio reading scenario;

[0052] Figure 9 This is a schematic diagram of the emotional speech synthesis method in the video dubbing scenario;

[0053] Figure 10 is a structural diagram of a voice data processing device according to this specification;

[0054] Figure 11 is a structural diagram of a model processing device according to this specification;

[0055] Figure 12 is a structural diagram of an emotional speech synthesis device according to this specification;

[0056] Figure 13 is a structural diagram of an emotional speech synthesis device according to this specification;

[0057] Figure 14 This is a structural diagram of a live broadcast-based emotional speech synthesis device according to this specification;

[0058] Figure 15 This is a structural diagram of an e-book-based emotional speech synthesis device according to this specification;

[0059] Figure 16 This is a structural diagram of a video-based emotional speech synthesis device according to this specification. DETAILED DESCRIPTION

[0060] This specification is further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to explain the relevant invention and are not intended to limit the invention. The embodiments described herein are merely a portion of the embodiments of this specification, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments in this specification without creative effort are intended to fall within the scope of protection of this application.

[0061] It should be noted that, for ease of description, only the portions relevant to the invention are shown in the accompanying drawings. The embodiments and features within these embodiments may be combined unless there is a conflict. Furthermore, terms such as "first" and "second" in this specification are used solely for informational purposes and do not constitute any limitation.

[0062] Some embodiments of this specification provide a method for processing speech data, a method for processing a model, and a method for synthesizing emotional speech, which can realize the synthesis of emotionally expressive speech data. Specifically, Figure 1 An exemplary system architecture diagram suitable for use with these embodiments is shown.

[0063] like Figure 1 As shown, it shows a sample management system, a model training system, a speech synthesis system and a client. Among them, the sample management system and the model training system can be the same system or different systems, which is not specifically limited here.

[0064] The sample management system can obtain multiple pieces of first emotional speech data of the target sounding object and establish an emotional speech data set of the target sounding object based on the multiple pieces of first emotional speech data. The emotional speech data in the emotional speech data set can be used as emotional speech data samples.

[0065] The target speaker is typically a natural person. The plurality of first emotional speech data corresponds to a plurality of lines, each of which corresponds to at least one emotional tag. It should be noted that the plurality of first emotional speech data and the plurality of lines may have a one-to-one correspondence.

[0066] In practice, the plurality of first emotional speech data and the at least one emotional tag also have a corresponding relationship. Specifically, a line sentence has an emotion indicated by a corresponding emotional tag, and accordingly, the first emotional speech data corresponding to the line sentence also has the emotion indicated by the emotional tag.

[0067] Dialogue can be any line spoken by any character. Furthermore, dialogue can include lines from any of the following works: literary works, dramatic works, film and television dramas, etc. It should be understood that the character can be a character from any of these works. Furthermore, the character can be a human or animal character, etc., without specific limitations here. Literary works can include novels and / or screenplays. Dramatic works can include plays, operas, local dramas, and / or radio dramas. Film and television dramas can include films and / or television dramas.

[0068] Furthermore, lines can include monologues, asides, or dialogues. Monologues are typically words spoken by any character to express their feelings or personal wishes. A aside is typically a character speaking to the audience without the knowledge of other characters. Dialogues are typically conversations between characters. Dialogues typically have strong emotions, and therefore, lines can specifically include dialogues.

[0069] The emotional markers in this specification can be markers used to represent any emotion. Specifically, the at least one emotional marker can include neutral, positive emotions and / or negative emotions. The positive emotions can include excitement, relief, happiness and / or admiration, etc. The negative emotions can include sadness, anger, disgust and / or fear, etc. Optionally, the neutral emotions can include surprise, boredom and / or fatigue, etc.

[0070] The first emotional voice data among the above-mentioned multiple first emotional voice data are emotionally expressive voice data of the target voice-producing object, which are obtained by recording the sound emitted by the target voice-producing object when reading the corresponding lines.

[0071] During the model training phase, the model training system may take as input the target identity information of the target speaker, as well as the lines and emotion tags corresponding to the emotional speech data samples in the emotional speech dataset, and use the emotional speech data samples as training labels to train the speech synthesis model to obtain the emotional speech synthesis model. The target identity information may include any information used to indicate the identity of the target speaker, such as, but not limited to, the user ID, ID number, employee number, and / or telephone number of the target speaker, etc., which are not specifically limited here.

[0072] After obtaining the emotional speech synthesis model, it can be applied to the speech synthesis system. Specifically, during the model application phase, the speech synthesis system can, for example, obtain the text information of the speech to be synthesized, as well as the emotional tag corresponding to the text information, from the client, and input the text information and the emotional tag into the emotional speech synthesis model, so that the emotional speech synthesis model outputs synthesized emotional speech data. The speech synthesis system can then provide the emotional speech data to the client, causing the client to play the emotional speech data to the user, and / or provide the emotional speech data to other clients other than the client, causing the other clients to play the emotional speech data to the user.

[0073] Among them, the speech synthesis system can be applied to different scenarios, such as live broadcast scenarios, audio reading scenarios and / or video dubbing scenarios, etc. In the live broadcast scenario, the text information of the speech to be synthesized may include the dubbing text of the virtual anchor, the source client of the dubbing text may include the anchor client, and the above-mentioned other clients may include the audience client. In the audio reading scenario, the text information of the speech to be synthesized may include the target text in the e-book, and the target text may be any text in the e-book, which is not specifically limited here. In addition, the source client of the target text may include an e-book client. In the video dubbing scenario, the text information of the speech to be synthesized may include the dubbing text of the video to be dubbed, and the source client of the dubbing text may include the video client.

[0074] The specific implementation steps of the above method are described below in conjunction with specific embodiments.

[0075] See Figure 2 , which shows a process 200 of an embodiment of a method for processing speech data. The execution subject of the method may be Figure 1 The sample management system shown. The method comprises the following steps:

[0076] Step 201: Acquire multiple pieces of first emotional speech data of a target sounding subject, wherein the multiple pieces of first emotional speech data correspond to multiple lines, and the multiple lines correspond to at least one emotional tag, wherein the first emotional speech data are obtained by recording the sound emitted by the target sounding subject when reading the corresponding lines;

[0077] Step 202: adjusting the target sound element of at least one first emotional speech data to obtain second emotional speech data;

[0078] Step 203: Merge the plurality of first emotional speech data and the second emotional speech data into an emotional speech data set of the target vocal object.

[0079] The above steps are further explained below.

[0080] In step 201, the plurality of first emotional speech data may be uploaded to the sample management system by a person responsible for voice recording, wherein the person and the target voice object may be the same person or different persons, which is not specifically limited here.

[0081] In addition, the above-mentioned multiple lines of dialogue may be manually selected or not manually selected, and there is no specific limitation here.

[0082] Optionally, before step 201, the execution subject may obtain at least one text, wherein the text contains dialogue sentences. Then, for the emotion tag in the at least one emotion tag, a plurality of dialogue sentences having the emotion indicated by the emotion tag may be extracted from the at least one text. Then, the extracted dialogue sentences may be provided to the target sound-producing object so that the target sound-producing object reads the extracted dialogue sentences, thereby obtaining the plurality of first emotion speech data. Among them, the emotion tag may correspond to the dialogue extraction rule in advance, and according to the dialogue extraction rule, a plurality of dialogue sentences having the emotion indicated by the emotion tag may be extracted from the at least one text. It should be understood that the dialogue extraction rule may be set according to actual needs and is not specifically limited here.

[0083] It should be noted that by adopting this non-manual selection method, the lines corresponding to the at least one emotional tag can be quickly obtained. Compared with the manual selection method, it can effectively save labor costs and time costs.

[0084] It should be noted that the text in the at least one text mentioned above can come from any of the works listed in the previous article.

[0085] In practice, for speech synthesis models, the larger the data size, the better the overall synthesis effect. However, due to strict requirements for emotional expressiveness and emotional intensity control, only speech data with different emotions from the same person can be used, and the data size is limited. Speech lines, especially dialogue lines, are both colloquial and emotional. For each of the at least one emotion tag mentioned above, multiple (e.g., 500-1000) dialogue lines are selected for that emotion tag, allowing for complete recording in a relatively short time, thereby effectively controlling costs.

[0086] After obtaining the plurality of first emotional speech data of the target sounding object through recording, in order to expand the emotional speech data sample of the target sounding object, step 202 may be performed to implement sample expansion.

[0087] Specifically, in step 202, the target sound elements of at least one of the plurality of first emotional speech data items can be adjusted to obtain second emotional speech data. It should be understood that the second emotional speech data is the adjusted first emotional speech data. The target sound elements are elements related to the characteristics of the sound. Furthermore, the target sound elements may include, for example, speech rate and / or intonation.

[0088] In step 203, the plurality of first emotional speech data and the second emotional speech data may be merged into an emotional speech data set.

[0089] In addition, the execution entity may also store the target speaker's emotional speech dataset in a designated database, and may also store corresponding relationship information related to the emotional speech dataset in the database. The corresponding relationship information is used to represent the corresponding relationship between the emotional speech data in the emotional speech dataset and the lines and emotional tags.

[0090] The voice data processing method provided in this embodiment can expand the emotional voice data sample of the target sound-producing object by obtaining multiple first emotional voice data obtained through recording of the target sound-producing object, and then adjusting the target sound elements of at least one first emotional voice data to obtain second emotional voice data. Then, the multiple first emotional voice data and second emotional voice data can be merged into an emotional voice data set with a larger data scale for the target sound-producing object. The emotional voice data in the emotional voice data set, as well as the lines and emotional tags corresponding to the emotional voice data, can be used to train an emotional voice synthesis model with better emotional voice synthesis effect.

[0091] Next, we will further introduce the application of the emotional speech dataset in the model training stage.

[0092] See Figure 3 , which shows a process 300 of an embodiment of a model processing method. The execution subject of the method can be Figure 1 The model training system shown. The method includes the following steps:

[0093] Step 301: obtaining target identity information and an emotional speech dataset of a target sounding subject, as well as lines and emotional tags corresponding to emotional speech data samples in the emotional speech dataset;

[0094] Step 302 : Taking the target identity information, dialogue sentences and emotion tags as input and the emotion speech data samples as training labels, the speech synthesis model to be trained is trained to obtain an emotion speech synthesis model.

[0095] The above steps are further explained below.

[0096] In step 301, the target identity information and emotional speech data set of the target sound-making object, as well as the lines and emotional tags corresponding to the emotional speech data samples in the emotional speech data set, can be received from the sample management system or obtained from the database as described above, without specific limitation here.

[0097] In step 302, the target identity information, as well as the lines and emotional tags corresponding to the emotional speech data samples in the emotional speech dataset can be used as input, and the emotional speech data samples are used as training labels to train the speech synthesis model to be trained to obtain an emotional speech synthesis model.

[0098] In practice, the speech synthesis model to be trained can be a pre-trained model. Specifically, the speech synthesis model to be trained can be pre-trained in the following manner: taking the sample identity information and text information of at least one sample sound-making object as input, and using the speech data of the sample sound-making object reading the text information as a training label to train the initial speech synthesis model, wherein the sample sound-making object is usually a natural person and is different from the target sound-making object. The information items included in the sample identity information are similar to those in the target identity information and will not be repeated here. Based on this, by training the pre-trained speech synthesis model to obtain an emotional speech synthesis model, the amount of emotional speech data of the target sound-making object can be greatly reduced.

[0099] Typically, the initial speech synthesis model can be an untrained speech synthesis model. During pre-training of the initial speech synthesis model, no emotion tag is input into the model. Therefore, the speech data of the at least one sample utterance can be considered as emotionless speech data.

[0100] It should be noted that although no emotion tags are input to the initial speech synthesis model during pre-training, the model can be pre-associated with at least one emotion tag as described above and randomly assign an emotion tag to the input text information from the at least one emotion tag. The speech synthesis model to be trained using this pre-training method can ensure speech intelligibility.

[0101] Optionally, the speech synthesis model to be trained may include, but is not limited to, a spectrogram prediction network and a vocoder. The spectrogram prediction network may be a neural network for predicting a spectrogram, and the vocoder may be a neural network for converting a spectrogram into a time-domain waveform. Typically, the spectrogram prediction network may introduce an attention mechanism, which may include, for example, a position-sensitive attention mechanism. By introducing this attention mechanism, the accumulated attention weights of previous decoding processes may be used as an additional feature, thereby enabling the model to remain consistent as it moves forward along the input sequence, reducing potential subsequence duplication or omission during the decoding process.

[0102] Furthermore, the spectrum prediction network is used to predict the spectrogram based on the input target identity information, lines and emotional tags. The vocoder is used to generate a time domain waveform based on the spectrogram. Based on this, in the model training phase, the first processing step of the speech synthesis model to be trained may include: using the spectrum prediction network to predict the spectrogram based on the input target identity information, lines and emotional tags; using the vocoder to generate a time domain waveform based on the spectrogram predicted by the spectrum prediction network. Figure 4a As shown, it is a schematic diagram of the first processing process of the above-mentioned speech synthesis model to be trained.

[0103] Specifically, during the model training phase, the spectrogram prediction network can convert the input target identity information, dialogue sentences, and emotional tags into vectors respectively, and concatenate the converted vectors to obtain a concatenated vector, and predict the spectrogram based on the concatenated vector.

[0104] It should be noted that the spectrogram in this specification is a time-varying spectrum. This spectrogram may include, but is not limited to, a Mel-frequency spectrogram. Typically, a Mel-frequency spectrogram is referred to as a Mel-spectrum and can be obtained by transforming the corresponding original spectrogram using a Mel-scaled filter bank.

[0105] Optionally, the spectrum prediction network can associate at least one emotion marker with an emotion intensity coefficient corresponding to each emotion marker. During the application phase of the emotional speech synthesis model, the spectrum prediction network can adjust the emotion intensity based on the emotion intensity coefficient corresponding to the input emotion marker. Based on this, during the application phase of the emotional speech synthesis model, the second processing step of the emotional speech synthesis model can include: utilizing the spectrum prediction network to adjust the emotion intensity based on the emotion intensity coefficient corresponding to the input emotion marker.

[0106] Specifically, for an input emotion tag, the spectrogram prediction network first maps the emotion tag into an emotion embedding vector. It then multiplies the emotion embedding vector by the emotion intensity coefficient corresponding to the emotion tag to determine the emotion embedding vector adjusted for emotion intensity. This allows for effective control of emotion intensity during the application phase of the emotional speech synthesis model.

[0107] The above-mentioned emotion intensity coefficient can be, for example, within the range of [0.01, 2]. In addition, the default value of the above-mentioned emotion intensity coefficient can be 1. For any emotion tag, when the value of the emotion intensity coefficient corresponding to the emotion tag is 0.01, the emotion indicated by the emotion tag is slightly inclined. When the value of the emotion intensity coefficient is 2, the default emotion intensity is doubled.

[0108] Optionally, the spectrogram prediction network may include but is not limited to an encoder and a decoder. The decoder may introduce the attention mechanism as described above. The encoder is used to convert the input target identity information, lines and emotional markers into vectors respectively, and splice the converted vectors, and input the obtained spliced ​​vectors into the decoder. The decoder is used to predict the spectrogram based on the spliced ​​vector. Based on this, the above-mentioned first processing process may further include: using the encoder to convert the input target identity information, lines and emotional markers into vectors respectively, and splicing the converted vectors to obtain a spliced ​​vector; using the decoder to predict the spectrogram based on the spliced ​​vector. As Figure 4bAs shown, it is a schematic diagram of the processing process of the spectrum prediction network in the above-mentioned first processing process.

[0109] Furthermore, the encoder may include but is not limited to an emotion tag embedding module, an identity embedding module and a character encoding module. Among them, the emotion tag embedding module is used to map the input emotion tag into an emotion embedding vector. The identity embedding module is used to map the input target identity information into an identity embedding vector. The character encoding module is used to map the input lines into character embedding vectors, and encode the character embedding vectors to obtain character encoding vectors. Based on this, the above-mentioned first processing process may further include: using the emotion tag embedding module to map the input emotion tag into an emotion embedding vector; using the identity embedding module to map the input target identity information into an identity embedding vector; using the character encoding module to map the input lines into character embedding vectors. Figure 4c As shown, it is a schematic diagram of the processing process of the encoder in the above-mentioned first processing process.

[0110] It should be understood that during the model training phase, the vectors output by the sentiment tag embedding module, the identity embedding module, and the character encoding module are respectively used to splice into the spliced ​​vector as described above.

[0111] Furthermore, the emotion marker embedding module can associate at least one emotion marker with an emotion intensity coefficient respectively corresponding to the emotion marker, and the emotion intensity coefficient is used for emotion intensity adjustment. In the application stage of the emotion speech synthesis model, the emotion marker embedding module can also be used to: after mapping the input emotion marker into an emotion embedding vector, determine the product of the emotion embedding vector and the emotion intensity coefficient corresponding to the emotion marker as the emotion embedding vector after emotion intensity adjustment. Based on this, the above-mentioned second processing process can further include: using the emotion marker embedding module, after mapping the input emotion marker into an emotion embedding vector, determine the product of the emotion embedding vector and the emotion intensity coefficient corresponding to the emotion marker as the emotion embedding vector after emotion intensity adjustment. Thus, in the application stage of the emotion speech synthesis model, effective control of emotion intensity can be achieved.

[0112] Optionally, training the speech synthesis model to be trained may include training a spectrogram prediction network. It should be understood that if the vocoder in the speech synthesis model to be trained has high accuracy, only the spectrogram prediction network in the speech synthesis model to be trained may be trained.

[0113] As an implementation method, training the aforementioned speech synthesis model to be trained specifically includes determining a prediction loss based on an emotional speech data sample serving as a training label and a spectrogram predicted by a spectrogram prediction network, and adjusting network parameters in the spectrogram prediction network with the goal of reducing the prediction loss. The prediction loss may be the degree of inconsistency between the spectrogram of the emotional speech data sample and the predicted spectrogram.

[0114] As another implementation, training the speech synthesis model to be trained specifically includes determining a prediction loss based on an emotional speech data sample serving as a training label and a time-domain waveform generated by a vocoder, and adjusting network parameters in the spectrogram prediction network with the goal of reducing the prediction loss. The prediction loss may be the degree of inconsistency between the time-domain waveform of the emotional speech data sample and the time-domain waveform generated by the vocoder.

[0115] Optionally, in addition to training the spectrogram prediction network, the vocoder can also be trained. For example, the spectrogram of an emotional speech data sample in an emotional speech dataset can be used as input, and the time domain waveform of the emotional speech data sample can be used as training labels to train the vocoder.

[0116] Optionally, the above-mentioned speech synthesis model to be trained can adopt an improved architecture of the Tacotron2 architecture. Among them, Tacotron2 is an end-to-end speech synthesis model based on deep learning. In practice, the Tacotron2 architecture includes a spectrogram prediction network, a vocoder, and an intermediate connection layer. The spectrogram prediction network is a feature prediction network based on a cyclic Seq2seq that introduces an attention mechanism, which is used to predict a Mel spectrum frame sequence from an input character sequence. The vocoder is a revised version of WaveNet, which is used to generate time domain waveform samples based on the predicted Mel spectrum frame sequence. The intermediate connection layer uses a low-level acoustic representation-Mel frequency spectrogram to connect the spectrogram prediction network and the vocoder.

[0117] Seq2seq is a variant of a recurrent neural network that includes an encoder and a decoder. WaveNet is a deep neural network used to generate raw audio.

[0118] In the Tacotron2 architecture, the spectrogram prediction network consists of an encoder and a decoder. The encoder consists solely of a character encoding module, which typically includes a character embedding layer, three convolutional layers, and a bidirectional LSTM (Long Short-Term Memory) network.

[0119] In some embodiments, the Tacotron2 architecture can be improved by adding the emotion tag embedding module and identity embedding module as described above to the encoder in the Tacotron2 architecture. The improved Tacotron2 architecture with the addition of the emotion tag embedding module and identity embedding module can be used as the architecture of the speech synthesis model to be trained.

[0120] The model processing method provided in this embodiment obtains the target identity information and emotional speech data set of the target sound-making object, as well as the lines and emotional tags corresponding to the emotional speech data samples in the emotional speech data set, and then uses the target identity information, the lines and emotional tags as input, and the emotional speech data samples as training labels to train the speech synthesis model to be trained, so as to obtain an emotional speech synthesis model with better emotional speech synthesis effect.

[0121] Next, we will introduce the relevant content of the emotional speech synthesis model in the application stage.

[0122] See Figure 5 , which shows a process 500 of an embodiment of an emotional speech synthesis method. The execution subject of the method can be Figure 1 The speech synthesis system shown. The method comprises the following steps:

[0123] Step 501: obtaining text information of the speech to be synthesized and its corresponding emotion tag;

[0124] Step 502: Input the text information and the emotion tag into the emotion speech synthesis model, so that the emotion speech synthesis model outputs synthesized emotion speech data.

[0125] Among them, the emotional speech synthesis model in this embodiment adopts Figure 3 The corresponding embodiment is described by the method for training.

[0126] It should be noted that, in this embodiment, the text information of the speech to be synthesized can be any type of text information, such as the dubbing text mentioned above, or the target text in an e-book, etc., and is not specifically limited here.

[0127] It should be noted that, as described above, the emotional speech synthesis model can include a spectrogram prediction network and a vocoder. The spectrogram prediction network can include an encoder and a decoder. The encoder can include an emotion tag embedding module, an identity embedding module, and a character encoding module.

[0128] During the application phase, the textual information of the speech to be synthesized and its corresponding emotion tag serve as input to the emotional speech synthesis model. Specifically, the emotion tag serves as input to the emotion tag embedding module, which outputs an emotion embedding vector based on the input emotion tag. The textual information of the speech to be synthesized serves as input to the character encoding module, which outputs a character encoding vector based on the input textual information. It should be understood that the concatenated vector, resulting from the concatenation of the emotion embedding vector and the character encoding vector, serves as input to the decoder.

[0129] There are different ways to implement the sentiment tag embedding module.

[0130] As an implementation method, the emotion tag embedding module can map the input emotion tag into an emotion embedding vector and output the emotion embedding vector.

[0131] As another implementation method, the emotion tag embedding module can associate at least one emotion tag with an emotion intensity coefficient respectively corresponding to the emotion tag, and the emotion intensity coefficient is used for emotion intensity adjustment. The emotion tag embedding module can be further used to: after mapping the input emotion tag into an emotion embedding vector, determine the product of the emotion embedding vector and the emotion intensity coefficient corresponding to the emotion tag as the emotion embedding vector after emotion intensity adjustment, and output the emotion embedding vector after emotion intensity adjustment. It should be understood that the emotion embedding vector after emotion intensity adjustment is used for splicing with the corresponding character encoding vector. By adopting this implementation method, effective control of emotion intensity can be achieved.

[0132] The emotional speech synthesis method provided in this embodiment obtains the text information of the speech to be synthesized and its corresponding emotional tag, and then inputs the text information and emotional tag into the emotional speech synthesis model, enabling the emotional speech synthesis model to synthesize emotionally expressive speech data. Furthermore, without the need to input additional information, such as reference audio, the synthesis effect and emotional intensity can be effectively controlled.

[0133] Further reading Figure 6 , which is a schematic diagram of an embodiment of an emotional speech synthesis method. Figure 1 The client shown in Figure 1 The interaction process between the speech synthesis system shown in FIG.

[0134] like Figure 6 As shown, the emotional speech synthesis method may include the following steps:

[0135] Step 601: The client obtains text information of the speech to be synthesized and its corresponding emotion tag;

[0136] Step 602: The client sends the text information and the emotion tag to the speech synthesis terminal;

[0137] In step 603, the speech synthesis end inputs the text information and the emotion tag into the emotion speech synthesis model, so that the emotion speech synthesis model outputs synthesized emotion speech data.

[0138] In step 601, the client may obtain the text information and its corresponding emotion tag in response to a user's speech synthesis instruction for the text information to be synthesized. The speech synthesis instruction may include the text information or a text tag of the text information, and the text tag may be pre-assigned to the emotion tag.

[0139] Optionally, the speech synthesis instruction may include an emotion tag and any one of the following: text information of the speech to be synthesized, and a text identifier of the text information. The emotion tag may be an emotion tag selected by the user for the text information.

[0140] In step 603, the speech synthesis end uses the emotional speech synthesis model to synthesize emotional speech data based on the text information and the emotional tag. Figure 3 The corresponding embodiment is described by the method for training.

[0141] Optionally, after step 603 , the speech synthesis end may provide the emotional speech data to the client, and / or provide the emotional speech data to other clients except the client.

[0142] Figure 6 The speech synthesis method described in the corresponding embodiment obtains the text information of the speech to be synthesized and its corresponding emotional tags through the client, and then sends the text information and emotional tags to the speech synthesis end, so that the speech synthesis end inputs the text information and emotional tags into the emotional speech synthesis model, which can enable the emotional speech synthesis model to output personalized emotional speech data, and the emotional speech data has strong emotional expressiveness.

[0143] Figure 6 The speech synthesis method described in the corresponding embodiment can be applied to different scenarios, such as live broadcast scenarios, audio reading scenarios and / or video dubbing scenarios.

[0144] As an example, in a live broadcast scenario, the text information of the speech to be synthesized may include the dubbing text of the live virtual host. Figure 7As shown, it shows a schematic diagram of the emotional speech synthesis method in a live broadcast scenario. Specifically, in a live broadcast scenario, the emotional speech synthesis method may include: step 701, the anchor client obtains the dubbing text of the virtual anchor of the live broadcast, and the emotional tag corresponding to the dubbing text; step 702, the anchor client sends the dubbing text and the emotional tag to the server; step 703, the server inputs the dubbing text and the emotional tag into the emotional speech synthesis model as described above, so that the emotional speech synthesis model outputs the synthesized emotional speech data; step 704, the server provides the emotional speech data to the corresponding audience client. Among them, the emotional speech data serves as the dubbing speech data of the virtual anchor. The server includes the speech synthesis terminal as described above. The audience client can be a client of the audience user who watches the live broadcast corresponding to the dubbing text. During the live broadcast, the audience client can play the emotional speech data to the audience users to which it belongs.

[0145] Optionally, in step 704, the server may provide the emotional voice data to a corresponding viewer client in response to receiving a play request related to the emotional voice data from the anchor client.

[0146] Optionally, the server may also provide the emotional voice data to the anchor client, so that the anchor client plays the emotional voice data to the user.

[0147] In the audio reading scenario, the text information to be synthesized into speech may include, for example, any text in an e-book. Figure 8 As shown, it shows a schematic diagram of the emotional speech synthesis method in the audio reading scenario. Specifically, in the audio reading scenario, the emotional speech synthesis method may include: step 801, the e-book client obtains the target text in the e-book, and the emotional tag corresponding to the target text; step 802, the e-book client sends the target text and the emotional tag to the speech synthesis end; step 803, the speech synthesis end inputs the target text and the emotional tag into the emotional speech synthesis model as described above, so that the emotional speech synthesis model outputs the synthesized emotional speech data; step 804, the speech synthesis end provides the emotional speech data to the e-book client, so that the e-book client provides the emotional speech data to the user. Among them, the emotional speech data serves as the speech data corresponding to the target text. The target text can be the text to be read aloud selected by the user in the e-book. In addition, the text category of the target text can include, for example, novels, essays or poems.

[0148] In the video dubbing scenario, the text information of the speech to be synthesized may include, for example, the dubbing text of the video to be dubbed. Figure 9As shown, it shows a schematic diagram of the emotional speech synthesis method in the video dubbing scenario. Specifically, in the video dubbing scenario, the emotional speech synthesis method may include: step 901, the video client obtains the dubbing text of the video to be dubbed, and the emotional tag corresponding to the dubbing text; step 902, the video client sends the dubbing text and the emotional tag to the speech synthesis end; step 903, the speech synthesis end inputs the dubbing text and the emotional tag into the emotional speech synthesis model as described above, so that the emotional speech synthesis model outputs the synthesized emotional speech data; step 904, the speech synthesis end provides the emotional speech data to the video client, so that the video client provides the emotional speech data to the user. Among them, the emotional speech data serves as the dubbing speech data of the video to be dubbed.

[0149] The above only provides examples of the application of the emotional speech synthesis method in live broadcast scenarios, audio reading scenarios, and video dubbing scenarios. The application of the emotional speech synthesis method in other scenarios can be inferred based on the examples described above, and no further examples will be given here.

[0150] Further references Figure 10 This specification provides an embodiment of a device for processing speech data. Figure 2 Corresponding to the method embodiment shown, the device can be applied to Figure 1 The sample management system shown.

[0151] like Figure 10 As shown, the speech data processing device 1000 of this embodiment includes: an acquisition unit 1001, an adjustment unit 1002 and a generation unit 1003. The acquisition unit 1001 is configured to acquire multiple first emotion speech data of the target sound-producing object, the multiple first emotion speech data corresponding to multiple lines, the multiple lines corresponding to at least one emotion tag, wherein the first emotion speech data is obtained by recording the sound emitted by the target sound-producing object when reading the corresponding lines; the adjustment unit 1002 is configured to adjust the target sound element of at least one emotion speech data to obtain second emotion speech data; the generation unit 1003 is configured to merge the multiple first emotion speech data and the second emotion speech data into an emotion speech data set of the target sound-producing object.

[0152] Optionally, the target sound elements may include speech speed and / or intonation, etc.

[0153] Optionally, the dialogue sentences may include dialogues from any of the following works: literary works, dramatic works, and film and television works.

[0154] Optionally, the at least one emotion tag may include at least one of the following: neutral, positive emotion, negative emotion, etc. Positive emotion may include at least one of the following: excitement, relief, happiness, admiration, etc. Negative emotion may include at least one of the following: sadness, anger, disgust, fear, etc.

[0155] Optionally, the acquisition unit 1001 can also be configured to: acquire at least one text; and the above-mentioned device 1000 can also include: an extraction unit (not shown in the figure), configured to extract, for an emotion tag in the above-mentioned at least one emotion tag, a plurality of lines of dialogue having the emotion indicated by the emotion tag from the at least one text; a sending unit (not shown in the figure), configured to provide the extracted lines of dialogue to the target sound-making object so that the target sound-making object reads the extracted lines of dialogue, thereby obtaining the above-mentioned plurality of first emotion voice data.

[0156] Further references Figure 11 This specification provides an embodiment of a model processing device. Figure 3 Corresponding to the method embodiment shown, the device can be applied to Figure 1 The model training system shown.

[0157] like Figure 11 As shown, the model processing device 1100 of this embodiment includes: an acquisition unit 1101 and a model training unit 1102. The acquisition unit 1101 is configured to acquire the target identity information and the emotional speech data set of the target sounding object, as well as the lines and emotional tags corresponding to the emotional speech data samples in the emotional speech data set; the model training unit 1102 is configured to use the target identity information, the lines and emotional tags as input, and the emotional speech data samples as training labels to train the speech synthesis model to be trained, thereby obtaining an emotional speech synthesis model.

[0158] Optionally, the speech synthesis model to be trained is pre-trained in the following manner: taking the sample identity information and text information of at least one sample sound-making object as input, using the speech data of the sample sound-making object reading the text information as a training label, and training the initial speech synthesis model, wherein the sample sound-making object is different from the target sound-making object.

[0159] Optionally, the speech synthesis model to be trained may include: a spectrogram prediction network and a vocoder. The first processing process of the speech synthesis model to be trained includes: using the spectrogram prediction network to predict the spectrogram based on the input target identity information, lines and emotional tags; using the vocoder to generate a time domain waveform based on the spectrogram predicted by the spectrogram prediction network.

[0160] Optionally, the model training unit 1102 may be further configured to: train a spectrum prediction network.

[0161] Optionally, the model training unit 1102 may be further configured to: determine the prediction loss based on the time domain waveform and the emotional speech data sample, and adjust the network parameters in the spectrum prediction network with the goal of reducing the prediction loss.

[0162] Optionally, the spectrum prediction network can associate at least one emotion marker with an emotion intensity coefficient corresponding to each emotion marker, and the emotion intensity coefficient is used for emotion intensity adjustment; and in the application stage of the emotion speech synthesis model, the second processing process of the emotion speech synthesis model may include: using the spectrum prediction network to adjust the emotion intensity according to the emotion intensity coefficient corresponding to the input emotion marker.

[0163] Optionally, the spectrogram prediction network may include an encoder and a decoder; and the above-mentioned first processing process may further include: using the encoder to convert the input target identity information, dialogue sentences and emotional tags into vectors respectively, and splicing the converted vectors to obtain a spliced ​​vector; using the decoder to predict the spectrogram based on the spliced ​​vector.

[0164] Optionally, the encoder may include an emotion tag embedding module, an identity embedding module and a character encoding module; and the above-mentioned first processing process may specifically include: using the emotion tag embedding module to map the input emotion tag into an emotion embedding vector; using the identity embedding module to map the input target identity information into an identity embedding vector; using the character encoding module to map the input dialogue sentence into a character embedding vector, and encoding the character embedding vector to obtain a character encoding vector.

[0165] Optionally, the emotion marker embedding module can associate at least one emotion marker with an emotion intensity coefficient corresponding to each emotion marker, and the emotion intensity coefficient is used for emotion intensity adjustment; and in the application stage of the emotion speech synthesis model, the second processing process of the emotion speech synthesis model can further include: using the emotion marker embedding module, after mapping the input emotion marker into an emotion embedding vector, the product of the emotion embedding vector and the emotion intensity coefficient corresponding to the emotion marker is determined as the emotion embedding vector after emotion intensity adjustment.

[0166] Optionally, the spectrogram may include a mel-frequency spectrogram.

[0167] Further references Figure 12 This specification provides an embodiment of an emotional speech synthesis device. Figure 5 Corresponding to the method embodiment shown, the device can be applied to Figure 1 The speech synthesis system shown.

[0168] like Figure 12 As shown, the emotional speech synthesis device 1200 of this embodiment includes: an acquisition unit 1201 and a speech synthesis unit 1202. The acquisition unit 1201 is configured to acquire the text information of the speech to be synthesized and its corresponding emotional tag; the speech synthesis unit 1202 is configured to input the text information and the emotional tag into the speech synthesis unit 1202. Figure 3 The emotional speech synthesis model obtained by training the method described in the corresponding embodiment enables the emotional speech synthesis model to output synthesized emotional speech data.

[0169] Further references Figure 13 This specification provides an embodiment of an emotional speech synthesis device. Figure 6 Corresponding to the method embodiment shown, the device can be applied to Figure 1 The client shown.

[0170] like Figure 13 As shown, the emotional speech synthesis device 1300 of this embodiment includes: an acquisition unit 1301 and a sending unit 1302. The acquisition unit 1301 is configured to acquire the text information of the speech to be synthesized and its corresponding emotional tag; the sending unit 1302 is configured to send the text information and emotional tag to the speech synthesis end, so that the speech synthesis end uses the text information and emotional tag as input. Figure 3 The emotional speech synthesis model obtained by training the method described in the corresponding embodiment enables the emotional speech synthesis model to output synthesized emotional speech data.

[0171] Further references Figure 14 This specification provides an embodiment of an emotional speech synthesis device based on live broadcast, and the embodiment of the device is similar to Figure 7 Corresponding to the method embodiment shown, the device can be applied to the anchor client in the live broadcast scene.

[0172] like Figure 14 As shown, the emotional speech synthesis device 1400 of this embodiment includes: an acquisition unit 1401, a sending unit 1402 and a processing unit 1403. Among them, the acquisition unit 1401 is configured to acquire the dubbing text of the live virtual anchor and the emotional tag corresponding to the dubbing text; the sending unit 1402 is configured to send the dubbing text and the emotional tag to the server, so that the server inputs the dubbing text and the emotional tag using the following method: Figure 3 The emotional speech synthesis model obtained by training the method described in the corresponding embodiment enables the emotional speech synthesis model to output synthesized emotional speech data; the processing unit 1403 is configured to provide the emotional speech data to the corresponding audience client via the server.

[0173] Further references Figure 15 This specification provides an embodiment of an emotional speech synthesis device based on an e-book. Figure 8 Corresponding to the method embodiment shown, the device can be applied to a speech synthesis terminal (such as Figure 1 The speech synthesis system shown).

[0174] like Figure 15 As shown, the emotional speech synthesis device 1500 of this embodiment includes: an acquisition unit 1501, a speech synthesis unit 1502 and a processing unit 1503. The acquisition unit 1501 is configured to acquire the target text in the e-book and the emotional tag corresponding to the target text; the speech synthesis unit 1502 is configured to input the target text and the emotional tag into the e-book. Figure 3 The emotional speech synthesis model obtained by training the method described in the corresponding embodiment enables the emotional speech synthesis model to output synthesized emotional speech data; the processing unit 1503 is configured to provide the emotional speech data based on the e-book client.

[0175] Further references Figure 16 This specification provides an embodiment of a video-based emotional speech synthesis device. Figure 9 Corresponding to the method embodiment shown, the device can be applied to a speech synthesis terminal (such as Figure 1 The speech synthesis system shown).

[0176] like Figure 16 As shown, the emotional speech synthesis device 1600 of this embodiment includes: an acquisition unit 1601, a speech synthesis unit 1602 and a processing unit 1603. The acquisition unit 1601 is configured to acquire the dubbing text of the video to be dubbed and the emotional tag corresponding to the dubbing text; the speech synthesis unit 1602 is configured to input the dubbing text and the emotional tag into the video. Figure 3 The emotional speech synthesis model obtained by training the method described in the corresponding embodiment enables the emotional speech synthesis model to output synthesized emotional speech data; the processing unit 1603 is configured to provide emotional speech data based on the video client.

[0177] exist Figure 10-16 In the corresponding device embodiments, the specific processing of each unit and the technical effects brought about by it can be referred to the relevant description in the method embodiment above, and will not be repeated here.

[0178] An embodiment of this specification further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, the computer is caused to execute the method described in any of the above method embodiments.

[0179] An embodiment of this specification also provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in any of the above method embodiments is implemented.

[0180] The embodiments of this specification also provide a computer program that, when executed on a computer, causes the computer to perform the method described in any of the above method embodiments. The computer program may include, for example, an APP (Application) or a mini-program.

[0181] Those skilled in the art will appreciate that, in one or more of the above examples, the functions described in the various embodiments disclosed in this specification may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0182] In some cases, the actions or steps recited in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the accompanying drawings do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0183] The specific implementation methods described above further illustrate in detail the purposes, technical solutions and beneficial effects of the multiple embodiments disclosed in this specification. It should be understood that the above description is only the specific implementation methods of the multiple embodiments disclosed in this specification, and is not intended to limit the protection scope of the multiple embodiments disclosed in this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the multiple embodiments disclosed in this specification should be included in the protection scope of the multiple embodiments disclosed in this specification.

Claims

1. A method for processing voice data, comprising: Acquiring multiple pieces of first emotional speech data of a target sounding subject, the multiple pieces of first emotional speech data corresponding to multiple lines, the multiple lines corresponding to at least one emotional tag, wherein the first emotional speech data is obtained by recording the sound emitted by the target sounding subject when reading the corresponding lines; Adjusting a target sound element of at least one piece of first emotional speech data to obtain second emotional speech data; Merging the plurality of first emotional speech data and the second emotional speech data into an emotional speech data set of the target vocal object; Obtaining target identity information and an emotional speech data set of a target sounding object, as well as lines and emotional tags corresponding to emotional speech data samples in the emotional speech data set, wherein the emotional speech data in the emotional speech data set is used as the emotional speech data sample; The target identity information, the lines and the emotion tag are used as input, and the emotional speech data sample is used as a training label to train the speech synthesis model to obtain an emotional speech synthesis model; In the application stage of the emotional speech synthesis model, the second processing process of the emotional speech synthesis model includes: The spectrum prediction network in the emotional speech synthesis model is used to adjust the emotional intensity according to the emotional intensity coefficient corresponding to the input emotional tag, wherein the spectrum prediction network associates the emotional intensity coefficients corresponding to multiple emotional tags respectively.

2. The method according to claim 1, wherein The target voice elements include speech rate and / or intonation.

3. The method according to claim 1, wherein The lines include lines from any of the following works: literary works, dramatic works, and film and television works.

4. The method according to claim 1, wherein The at least one emotion tag includes at least one of the following: neutral, positive emotion, and negative emotion.

5. The method according to claim 4, wherein The positive emotions include at least one of the following: excitement, relief, happiness, and admiration; The negative emotion includes at least one of the following: sadness, anger, disgust, and fear.

6. The method according to any one of claims 1 to 5, wherein: Before acquiring the plurality of first emotional speech data of the target sound utterance object, the method further includes: Get at least one text; For an emotion tag in the at least one emotion tag, extracting a plurality of lines having the emotion indicated by the emotion tag from the at least one text; The extracted dialogue sentences are provided to the target sounding object so that the target sounding object reads the extracted dialogue sentences, thereby obtaining the plurality of first emotion voice data.

7. A model processing method comprising: Obtain target identity information and an emotional speech data set of a target sound-producing object, as well as lines and emotional tags corresponding to emotional speech data samples in the emotional speech data set, wherein the emotional speech data in the emotional speech data set is used as the emotional speech data sample, and the emotional speech data set includes first emotional speech data and second emotional speech data, where the second emotional speech data is obtained by adjusting a target sound element of at least one piece of the first emotional speech data; The target identity information, the lines and the emotion tag are used as input, and the emotional speech data sample is used as a training label to train the speech synthesis model to obtain an emotional speech synthesis model; In the application stage of the emotional speech synthesis model, the second processing process of the emotional speech synthesis model includes: The spectrum prediction network in the emotional speech synthesis model is used to adjust the emotional intensity according to the emotional intensity coefficient corresponding to the input emotional tag, wherein the spectrum prediction network associates the emotional intensity coefficients corresponding to multiple emotional tags respectively.

8. The method according to claim 7, wherein: The speech synthesis model to be trained is pre-trained in the following manner: The sample identity information and text information of at least one sample sound-making object are used as input, and the speech data of the sample sound-making object reading the text information is used as a training label to train the initial speech synthesis model, wherein the sample sound-making object is different from the target sound-making object.

9. The method according to claim 7, wherein: The speech synthesis model to be trained includes a spectrum prediction network and a vocoder. The first processing process of the speech synthesis model to be trained includes: Utilizing the spectrogram prediction network, predicting a spectrogram based on input target identity information, lines, and emotional tags; The vocoder is used to generate a time domain waveform based on the spectrogram predicted by the spectrogram prediction network.

10. The method according to claim 9, wherein: The training of the speech synthesis model to be trained includes: Based on the time domain waveform and the emotional speech data sample, a prediction loss is determined, and with the goal of reducing the prediction loss, network parameters in the spectrum prediction network are adjusted.

11. The method according to claim 9, wherein: The spectrum prediction network is associated with an emotion intensity coefficient corresponding to each of the at least one emotion marker, and the emotion intensity coefficient is used for emotion intensity adjustment.

12. The method according to claim 9, wherein The spectrogram prediction network includes an encoder and a decoder; and The method of using the spectrogram prediction network to predict a spectrogram based on input target identity information, lines and emotion tags includes: Using the encoder, the input target identity information, lines and emotion tags are converted into vectors respectively, and the converted vectors are spliced ​​to obtain a spliced ​​vector; A spectrogram is predicted using the decoder based on the concatenated vector.

13. The method according to claim 12, wherein: The encoder includes an emotion tag embedding module, an identity embedding module and a character encoding module; as well as The encoder is used to convert the input target identity information, lines and emotion tags into vectors, including: Mapping the input emotion tag into an emotion embedding vector using the emotion tag embedding module; Mapping the input target identity information into an identity embedding vector using the identity embedding module; The character encoding module is used to map the input dialogue sentences into character embedding vectors, and the character embedding vectors are encoded to obtain character encoding vectors.

14. The method according to claim 13, wherein The emotion marker embedding module associates the emotion intensity coefficient corresponding to at least one emotion marker, and the emotion intensity coefficient is used for emotion intensity adjustment; as well as The emotion tag embedding module maps the input emotion tag into an emotion embedding vector, and then multiplies the emotion embedding vector by the emotion intensity coefficient corresponding to the emotion tag to determine the emotion embedding vector adjusted by emotion intensity.

15. The method according to claim 9 or 12, wherein: Spectrograms include Mel-frequency spectrograms.

16. A method for emotional speech synthesis, comprising: Obtain the text information of the speech to be synthesized and its corresponding emotion tag; The text information and the emotion tag are input into an emotion speech synthesis model trained using the method according to claim 7, so that the emotion speech synthesis model outputs synthesized emotion speech data.

17. An emotional speech synthesis method, applied to a client, comprising: The text information and the emotional tag are sent to the speech synthesis end, so that the speech synthesis end inputs the text information and the emotional tag into the emotional speech synthesis model trained using the method as claimed in claim 7, so that the emotional speech synthesis model outputs synthesized emotional speech data.

18. A method for emotional speech synthesis based on live broadcast, applied to a host client, comprising: Obtain the dubbing text of the virtual anchor of the live broadcast and the emotional tag corresponding to the dubbing text; Sending the dubbing text and the emotion tag to a server, so that the server inputs the dubbing text and the emotion tag into an emotion speech synthesis model trained using the method according to claim 7, so that the emotion speech synthesis model outputs synthesized emotion speech data; The emotional voice data is provided to the corresponding viewer client via the server.

19. An emotional speech synthesis method based on an e-book, comprising: Obtaining a target text in an e-book and a sentiment tag corresponding to the target text; Inputting the target text and the emotion tag into an emotional speech synthesis model trained by the method according to claim 7, so that the emotional speech synthesis model outputs synthesized emotional speech data; The emotional voice data is provided based on an e-book client.

20. A video-based emotional speech synthesis method, comprising: Obtaining the dubbing text of the video to be dubbed and the emotion tag corresponding to the dubbing text; Inputting the dubbing text and the emotion tag into the emotion speech synthesis model trained by the method according to claim 7, so that the emotion speech synthesis model outputs synthesized emotion speech data; The emotional voice data is provided based on a video client.

21. A speech synthesis model, comprising: A spectrogram prediction network is used to predict a spectrogram based on the input target identity information of the target sounding object and the lines and emotion tags corresponding to the emotional speech data sample of the target sounding object; a vocoder for generating a time domain waveform based on the spectrogram predicted by the spectrogram prediction network; The speech synthesis model is obtained in the following way: Obtaining target identity information and an emotional speech dataset of a target sounding object, as well as lines and emotional tags corresponding to emotional speech data samples in the emotional speech dataset; The target identity information, the lines and the emotion tag are used as input, and the emotion speech data sample is used as a training label to train the speech synthesis model to obtain a speech synthesis model, wherein the emotion speech data in the emotion speech data set is used as the emotion speech data sample, and the emotion speech data set includes first emotion speech data and second emotion speech data, and the second emotion speech data is obtained by adjusting the target sound element of at least one piece of the first emotion speech data; During the application phase of the emotional speech synthesis model, the second processing step of the speech synthesis model includes: The spectrum prediction network in the speech synthesis model is used to adjust the emotion intensity according to the emotion intensity coefficient corresponding to the input emotion tag, wherein the spectrum prediction network associates the emotion intensity coefficients corresponding to multiple emotion tags respectively.

22. The speech synthesis model according to claim 21, wherein: The spectrum prediction network is associated with an emotion intensity coefficient corresponding to each of the at least one emotion marker, and the emotion intensity coefficient is used for emotion intensity adjustment.

23. The speech synthesis model according to claim 21, wherein: The spectrum prediction network includes: The encoder is used to convert the input target identity information, lines and emotion tags into vectors respectively, and concatenate the converted vectors to obtain a concatenated vector; A decoder is configured to predict a spectrogram based on the concatenated vector.

24. The speech synthesis model according to claim 23, wherein: The encoder comprises: Sentiment tag embedding module, used to map the input sentiment tag into a sentiment embedding vector; The identity embedding module is used to map the input target identity information into an identity embedding vector; The character encoding module is used to map the input dialogue sentences into character embedding vectors and encode the character embedding vectors to obtain character encoding vectors.

25. The speech synthesis model according to claim 24, wherein: The emotion marker embedding module associates the emotion intensity coefficient corresponding to at least one emotion marker, and the emotion intensity coefficient is used for emotion intensity adjustment; as well as During the model application phase, the sentiment tag embedding module is also used to: After the input emotion tag is mapped into an emotion embedding vector, the product of the emotion embedding vector and the emotion intensity coefficient corresponding to the emotion tag is determined as the emotion embedding vector adjusted by emotion intensity.

26. A device for processing speech data, comprising: an acquisition unit configured to acquire a plurality of first emotional speech data of a target sounding subject, the plurality of first emotional speech data corresponding to a plurality of lines, the plurality of lines corresponding to at least one emotional tag, wherein the first emotional speech data is obtained by recording the sound emitted by the target sounding subject when reading the corresponding lines; an adjusting unit configured to adjust a target sound element of at least one first emotional speech data to obtain adjusted second emotional speech data; a generating unit configured to merge the plurality of first emotion speech data and the second emotion speech data into an emotion speech data set of the target vocal object; The voice data processing device is further configured to: Obtaining target identity information and an emotional speech data set of a target sounding object, as well as lines and emotional tags corresponding to emotional speech data samples in the emotional speech data set, wherein the emotional speech data in the emotional speech data set is used as the emotional speech data sample; The target identity information, the lines and the emotion tag are used as input, and the emotional speech data sample is used as a training label to train the speech synthesis model to obtain an emotional speech synthesis model; In the application stage of the emotional speech synthesis model, the second processing process of the emotional speech synthesis model includes: The spectrum prediction network in the emotional speech synthesis model is used to adjust the emotional intensity according to the emotional intensity coefficient corresponding to the input emotional tag, wherein the spectrum prediction network associates the emotional intensity coefficients corresponding to multiple emotional tags respectively.

27. A model processing device comprising: an acquisition unit configured to acquire target identity information and an emotional speech data set of a target sound-producing object, as well as lines and emotional tags corresponding to emotional speech data samples in the emotional speech data set, wherein the emotional speech data in the emotional speech data set is used as the emotional speech data sample, the emotional speech data set includes first emotional speech data and second emotional speech data, and the second emotional speech data is obtained by adjusting a target sound element of at least one piece of the first emotional speech data; A model training unit is configured to take the target identity information, the lines and the emotion tag as input, and the emotional speech data sample as a training label, to train the speech synthesis model to be trained, and obtain an emotional speech synthesis model; The model processing device is further configured to: In the application stage of the emotional speech synthesis model, the second processing process of the emotional speech synthesis model includes: The spectrum prediction network in the emotional speech synthesis model is used to adjust the emotional intensity according to the emotional intensity coefficient corresponding to the input emotional tag, wherein the spectrum prediction network associates the emotional intensity coefficients corresponding to multiple emotional tags respectively.

28. An emotional speech synthesis device, comprising: An acquisition unit configured to acquire text information of the speech to be synthesized and its corresponding emotion tag; The speech synthesis unit is configured to input the text information and the emotion tag into an emotion speech synthesis model trained using the method according to claim 7, so that the emotion speech synthesis model outputs synthesized emotion speech data.

29. An emotional speech synthesis device, applied to a client, comprising: The sending unit is configured to send the text information and the emotional tag to the speech synthesis end, so that the speech synthesis end inputs the text information and the emotional tag into the emotional speech synthesis model trained using the method as claimed in claim 7, so that the emotional speech synthesis model outputs synthesized emotional speech data.

30. A live broadcast-based emotional speech synthesis device, applied to a host client, comprising: An acquisition unit is configured to acquire a dubbing text of a live virtual anchor and an emotion tag corresponding to the dubbing text; a sending unit configured to send the dubbing text and the emotion tag to a server, so that the server inputs the dubbing text and the emotion tag into an emotional speech synthesis model trained using the method according to claim 7, so that the emotional speech synthesis model outputs synthesized emotional speech data; The processing unit is configured to provide the emotional voice data to the corresponding audience client via the server.

31. An emotional speech synthesis device based on an electronic book, comprising: An acquisition unit is configured to acquire a target text in an e-book and a sentiment tag corresponding to the target text; a speech synthesis unit configured to input the target text and the emotion tag into an emotional speech synthesis model trained using the method according to claim 7, so that the emotional speech synthesis model outputs synthesized emotional speech data; The processing unit is configured to provide the emotional voice data based on the e-book client.

32. A video-based emotional speech synthesis device, comprising: An acquisition unit is configured to acquire a dubbing text of a video to be dubbed and an emotion tag corresponding to the dubbing text; a speech synthesis unit configured to input the dubbing text and the emotion tag into an emotional speech synthesis model trained using the method according to claim 7, so that the emotional speech synthesis model outputs synthesized emotional speech data; The processing unit is configured to provide the emotional voice data based on the video client.

33. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed in a computer, the computer is caused to execute the method according to any one of claims 1 to 20.

34. A computing device comprising a memory and a processor, wherein: The memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 20 is implemented.

35. A computer program, which, when executed in a computer, causes the computer to perform the method according to any one of claims 1 to 20.

Citation Information

Patent Citations

  • Emotional speech synthesizing method and device

    CN102005205A

  • Training method of emotion recognition model, emotion recognition method, device, equipment, and storage medium

    CN109817246A

  • Speech synthesis method and device and computer readable storage medium

    CN110136690A

  • Speech synthesis method, speech synthesis system, terminal device and readable storage medium

    CN110379409A

  • Voice synthesis method and voice synthesis device

    CN111192568A