Emotional voice generation method and device, equipment and medium
By obtaining the target emotional prompt text described in natural language and combining the target text, using the emotion encoder and text encoder to generate embedded vectors and input them into the joint speech generation model, it solves the problem that the existing technology is difficult to achieve delicate or complex emotional expression, and realizes natural and emotionally vivid speech generation.
Patent Information
- Application Number
- CN202510347786.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-24
AI Technical Summary
The existing emotional control speech synthesis technology is difficult to achieve delicate or complex emotional expression, and the interaction method is single, resulting in unnatural speech generated, reducing the flexibility and synthesis effect of emotional control speech synthesis.
By obtaining the target emotion prompt text of the target text and natural language description, the pre-trained emotion encoder and text encoder generate emotion embedding vectors and semantic embedding vectors, and input them into the joint speech generation model for fusion, generating speech features with expected emotions, and finally generating target emotion speech through decoding processing.
It improves the flexibility of emotional speech generation and control interaction, and the generated speech is natural and vivid, effectively improving the synthesis effect of emotional speech.
Smart Images

Figure CN120199282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an emotional speech generation method, device, equipment and medium. Background Art
[0002] Text-To-Speech (TTS) technology refers to the process of generating speech from text. With the development of artificial intelligence technology, speech synthesis plays an increasingly important role in the field of human-computer dialogue. As an important supplement to speech synthesis, emotional speech generation greatly expands the application scenarios of speech synthesis. For example, in the field of medical and health, emotional speech generation can be used to help those who have lost their voices express emotions through a computer, and can also provide more realistic emotional expressions in intelligent diagnosis and treatment systems; or in the field of fintech business, emotional speech generation can be used in financial intelligent customer service to make the speech of financial intelligent customer service more natural and emotional, so as to reduce users' resistance to machine customer service.
[0003] However, there are still some problems in existing emotion-controlled speech synthesis. Traditional TTS emotion control methods usually rely on predefined discrete emotion labels (such as "happy" and "sad"), which limits the richness and flexibility of emotion expression and is difficult to meet more delicate or complex emotion needs. Moreover, most of the existing technologies achieve emotion control by directly adjusting emotion embedding vectors or input parameters, with a single interaction method, making it difficult for users to intuitively express the desired emotion effect, resulting in the synthesized emotional speech sounding rigid and unnatural, and reducing the flexibility and synthesis effect of emotion-controlled speech synthesis. Summary of the Invention
[0004] In view of the above deficiencies of the prior art, the purpose of the present invention is to provide an emotional speech generation method, device, equipment and medium that can be applied to the medical field, fintech or other related fields, and its main purpose is to improve the flexibility of emotional speech generation control interaction and the synthesis effect of emotional speech.
[0005] The technical solution of the present invention is as follows:
[0006] The first aspect of the present invention provides an emotional speech generation method, including:
[0007] Obtaining a target text of the speech to be generated and a target emotion prompt text described in natural language;
[0008] Inputting the target emotion prompt text into a pre-trained emotion encoder to obtain a corresponding emotion embedding vector;
[0009] Inputting the target text into a pre-trained text encoder to obtain a corresponding semantic embedding vector;
[0010] Input the emotional embedding vector and the semantic embedding vector into a pre-trained joint speech generation model, perform speech feature prediction by fusing the emotional embedding vector and the semantic embedding vector, and generate speech features with the expected emotion;
[0011] Perform decoding processing on the speech features to generate the target emotional speech corresponding to the target text.
[0012] The second aspect of the present invention provides an emotional speech generation device, including:
[0013] An acquisition module, configured to acquire the target text of the speech to be generated and the target emotional prompt text described based on natural language;
[0014] An emotion encoding module, configured to input the target emotional prompt text into a pre-trained emotion encoder to obtain the corresponding emotional embedding vector;
[0015] A text encoding module, configured to input the target text into a pre-trained text encoder to obtain the corresponding semantic embedding vector;
[0016] A speech generation module, configured to input the emotional embedding vector and the semantic embedding vector into a pre-trained joint speech generation model, perform speech feature prediction by fusing the emotional embedding vector and the semantic embedding vector, and generate speech features with the expected emotion;
[0017] A speech decoding module, configured to perform decoding processing on the speech features to generate the target emotional speech corresponding to the target text.
[0018] The third aspect of the present invention provides a computer device, including at least one processor; and,
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above-mentioned emotional speech generation method.
[0021] The fourth aspect of the present invention provides a non-volatile computer-readable storage medium, which stores computer-executable instructions, and when the computer-executable instructions are executed by one or more processors, the one or more processors can be enabled to execute the above-mentioned emotional speech generation method.
[0022] Beneficial effects: The present invention discloses an emotional speech generation method, device, equipment and medium. Compared with the prior art, in the embodiments of the present invention, a target text of the speech to be generated and a target emotional prompt text described in natural language are obtained; the target emotional prompt text is input into a pre-trained emotional encoder to obtain a corresponding emotional embedding vector; the target text is input into a pre-trained text encoder to obtain a corresponding semantic embedding vector; the emotional embedding vector and the semantic embedding vector are input into a pre-trained joint speech generation model to perform a fused speech feature prediction on the emotional embedding vector and the semantic embedding vector, and generate a speech feature with an expected emotion; the speech feature is decoded to generate a target emotional speech corresponding to the target text. By using the emotional prompt text described in natural language to control the emotion of the generated speech, the flexibility of the emotional speech generation control interaction is improved, and through the fused speech generation of emotional embedding and semantic embedding, the generated speech is natural and vivid in emotion, effectively improving the synthesis effect of emotional speech. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the solutions in the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 It is a schematic diagram of an application environment for the emotional speech generation method provided by the embodiments of the present invention;
[0025] Figure 2 It is a flowchart of the emotional speech generation method provided by the embodiments of the present invention;
[0026] Figure 3 It is a flowchart of step S202 of the emotional speech generation method provided by the embodiments of the present invention;
[0027] Figure 4 It is a flowchart of step S204 of the emotional speech generation method provided by the embodiments of the present invention;
[0028] Figure 5 It is a flowchart of step S205 of the emotional speech generation method provided by the embodiments of the present invention;
[0029] Figure 6 It is another flowchart of the emotional speech generation method provided by the embodiments of the present invention;
[0030] Figure 7 It is another flowchart of the emotional speech generation method provided by the embodiments of the present invention;
[0031] Figure 8 Schematic diagram of the functional modules of the emotional speech generation device provided by the embodiment of the present invention;
[0032] Figure 9 Schematic diagram of the hardware structure of the computer device provided by the embodiment of the present invention. Detailed implementation manners
[0033] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the present invention is further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The embodiments of the present invention are introduced below with reference to the accompanying drawings.
[0034] The emotional speech generation method provided by the embodiment of the present invention can be applied in an application environment such as Figure 1 which includes a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 can include various connection types, such as wired and / or wireless communication links, etc.
[0035] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only for example).
[0036] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0037] The server 105 may be a server that provides various services, such as a background server (only for example) that supports the content browsed by the user using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device. The server 105 may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of large management difficulty and weak business scalability existing in traditional physical hosts and VPS services (″Virtual Private Server″, or simply ″VPS″). The server 105 may also be a server of a distributed system, or a server combined with a blockchain.
[0038] It should be noted that the emotional speech generation method provided by the embodiments of the present application can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the emotional speech generation device provided by the embodiments of the present invention can also be set in the first terminal device 101, the second terminal device 102, or the third terminal device 103. Or, the emotional speech generation method provided by the embodiments of the present invention can generally also be executed by the server 105. Correspondingly, the emotional speech generation device provided by the embodiments of the present invention can generally be set in the server 105.
[0039] It should be understood that the numbers of the above terminal devices, networks, and servers are only illustrative. According to actual needs, any number of terminal devices, networks, and servers can be provided.
[0040] As Figure 2 shown, the emotional speech generation method provided by the embodiments of the present invention specifically includes the following steps:
[0041] S201. Obtain a target text of the speech to be generated and a target emotion prompt text described in natural language.
[0042] In this embodiment, the target text refers to the text content for which the user hopes to generate speech. Specifically, it can be the text data manually input by the user, or the text data automatically obtained according to a preset reply template, etc. The target emotion prompt text is the type of emotion that the user hopes to generate speech to express described in natural language. Here, natural language description means using everyday language to express emotions, rather than simple labels, so as to more precisely convey the nuances of emotions. That is, the target emotion prompt text in this embodiment does not rely on predefined discrete emotion labels (such as "happy", "sad", etc.), but controls the speech emotion through natural language description. For example, "Read with a gentle and slightly sad tone", "A happy but slightly nervous tone", etc. This not only reduces the interaction threshold of emotion description but also supports more complex and diverse emotion prompts.
[0043] Exemplarily, in the scenario of emotion speech generation in the medical field, for example, in the voice guidance system of a hospital, the target text obtained can be "Your test results will take some time. Please wait patiently", and the target emotion prompt text can be "Read with a gentle and slightly reassuring tone" to relieve the patient's anxiety.
[0044] Or, in the scenario of emotion speech generation in the financial field, for example, in the intelligent customer service system of a bank, the target text can be "I'm sorry, your loan application has not passed the review", and the target emotion prompt text can be "Read with a gentle and slightly regretful tone" to alleviate the customer's disappointment.
[0045] S202. Input the target emotion prompt text into a pre-trained emotion encoder to obtain a corresponding emotion embedding vector.
[0046] In this embodiment, the obtained target emotion prompt text is input into a pre-trained emotion encoder. The emotion encoder is specifically a neural network model, whose function is to convert the emotion prompt text described in natural language into a numerical vector that can be understood and processed by a computer, namely an emotion embedding vector; the emotion embedding vector is a high-dimensional vector that can capture the complex features of emotions, thereby converting the abstract emotion description into specific numerical features and providing an emotion basis for subsequent speech generation.
[0047] For example, in the scenario of the medical field, the pre-trained emotion encoder converts "read aloud in a gentle and slightly reassuring tone" into an emotion embedding vector as the emotion guidance for speech generation, making the generated speech sound more gentle and reassuring when informing the patient of the waiting for the examination results; for another example, in the scenario of the financial field, the pre-trained emotion encoder converts "read aloud in a gentle and slightly regretful tone" into an emotion embedding vector as the emotion guidance for speech generation, making the generated speech sound more gentle and regretful when informing the customer of the loan application result, providing accurate generation guidance information for generating accurate emotion speech and improving the generation effect of emotion speech.
[0048] S203. Input the target text into the pre-trained text encoder to obtain the corresponding semantic embedding vector.
[0049] In this embodiment, to achieve accurate text-to-speech generation conversion, the target text of the speech to be generated is input into the pre-trained text encoder. The text encoder can specifically be a large-scale pre-trained language model, such as BERT, RoBERTa, or GPT, etc. The text encoder performs semantic parsing on the target text, extracts the context information and semantic features of the input target text, thereby generating a semantic embedding vector, converting the semantic information of the text into numerical features, providing an accurate semantic basis for subsequent speech generation, and achieving accurate and reliable text-to-emotion speech generation conversion.
[0050] S204. Input the emotion embedding vector and the semantic embedding vector into the pre-trained joint speech generation model, perform fusion speech feature prediction on the emotion embedding vector and the semantic embedding vector, and generate speech features with the expected emotion.
[0051] In this embodiment, after obtaining the emotion embedding vector and the semantic embedding vector, the pre-trained joint semantic generation model is used to co-model the text content and the emotion information to generate speech features with the expected emotion. Among them, the joint speech generation model is a pre-trained deep learning model, such as a Transformer model, a Conformer model, etc. These models can process sequence data, fuse the emotion embedding vector and the semantic embedding vector. The fusion method can be splicing or an attention mechanism, etc. Splicing is to directly splice the two vectors into a longer vector, while the attention mechanism dynamically adjusts the contribution degrees of the two vectors in a weighted manner, and both can achieve the fusion input of the emotion embedding vector and the semantic embedding vector.
[0052] The joint speech generation model processes the emotion embedding vector and the semantic embedding vector through a multi-layer neural network (such as the encoder-decoder structure of Transformer), and makes a speech generation prediction for the semantic embedding vector under the guidance of the emotion of the emotion embedding vector, so as to dynamically fuse the emotion embedding into the semantic embedding, realize the fusion prediction of emotion and semantic information, and map the text and emotion information to speech features. The speech features specifically refer to Mel spectrogram features, which contain speech content and emotion information. Through the joint speech generation prediction, it is ensured that the emotion expression of the output speech features is consistent with the requirements of the target emotion prompt text, and natural, vivid and emotional speech features are generated.
[0053] S205. Perform decoding processing on the speech features to generate the target emotion speech corresponding to the target text.
[0054] In this embodiment, decoding processing is performed based on the generated speech features with expected emotion. For example, for the generated Mel spectrogram features, which can capture the frequency and time information of the speech, the key features carried therein can be decoded and converted into high-quality audio waveforms by decoding the Mel spectrogram features, so as to generate the actual target emotion speech corresponding to the target text. At this time, the generated target emotion speech not only conveys the semantic information of the text, but also expresses the emotion consistent with the target emotion prompt text, improving the naturalness and expressiveness of speech synthesis.
[0055] For example, in the scenario of the medical field, convert "Read in a gentle and slightly comforting tone" into an emotion embedding vector, and convert "Your test results will take some time. Please wait patiently" into a semantic embedding vector. The joint speech generation model fuses the emotion embedding vector of "Read in a gentle and slightly comforting tone" and the semantic embedding vector of "Your test results will take some time. Please wait patiently", decodes the generated speech features into the speech of "Your test results will take some time. Please wait patiently", and the emotion of "Read in a gentle and slightly comforting tone" is contained in the speech, relieving the patient's anxiety.
[0056] In the scenario of the financial field, convert "Read in a gentle and slightly regretful tone" into an emotion embedding vector, and convert "I'm sorry, your loan application has not been approved" into a semantic embedding vector. The joint speech generation model fuses the emotion embedding vector of "Read in a gentle and slightly regretful tone" and the semantic embedding vector of "I'm sorry, your loan application has not been approved", decodes the generated speech features into the speech of "I'm sorry, your loan application has not been approved", and the emotion of "Read in a gentle and slightly regretful tone" is contained in the speech, reducing the customer's disappointment.
[0057] In the above embodiments, the present invention discloses an emotional speech generation method, which includes obtaining a target text of the speech to be generated and a target emotional prompt text described in natural language; inputting the target emotional prompt text into a pre-trained emotional encoder to obtain a corresponding emotional embedding vector; inputting the target text into a pre-trained text encoder to obtain a corresponding semantic embedding vector; inputting the emotional embedding vector and the semantic embedding vector into a pre-trained joint speech generation model to perform fusion speech feature prediction on the emotional embedding vector and the semantic embedding vector, generating speech features with an expected emotion; and performing decoding processing on the speech features to generate the target emotional speech corresponding to the target text. By using the emotional prompt text described in natural language to control the emotion of the generated speech, the flexibility of the emotional speech generation control interaction is improved, and through the fusion speech generation of emotional embedding and semantic embedding, the generated speech is natural and vivid in emotion, effectively improving the synthesis effect of emotional speech.
[0058] In one embodiment, the emotional encoder includes a pre-trained language model, an emotional feature extractor, and an emotional mapper. As Figure 3 shown, step S202 includes:
[0059] S301. Input the target emotional prompt text into the pre-trained language model, perform semantic parsing on the target emotional prompt text, and obtain emotional prompt semantic information;
[0060] S302. Input the emotional prompt semantic information into the emotional feature extractor, perform feature extraction on the emotional prompt semantic information, and obtain corresponding emotional prompt features;
[0061] S303. Perform embedding mapping on the emotional prompt features through the emotional mapper, map the emotional prompt features to a preset embedding space for speech generation, and obtain the emotional embedding vector.
[0062] In this embodiment, a pre-trained language model (such as BERT, RoBERTa, etc.) is pre-trained with a large amount of text data, capable of understanding the semantic information of natural language. The target emotion prompt text is input into the pre-trained language model, and the large-scale pre-trained language model is used to perform semantic parsing on the target emotion prompt text, capture the context information of the target emotion prompt text, and obtain the corresponding emotion prompt semantic information. For example, in the scenario of the medical field, the target emotion prompt text "Read in a gentle and slightly comforting tone" is input into the pre-trained language model, and the model outputs a high-dimensional semantic representation, capturing the semantic information of "gentle" and "comforting"; while in the scenario of the financial field, the target emotion prompt text "Read in a gentle and slightly regretful tone" is input into the pre-trained language model, and the model outputs a high-dimensional semantic representation, capturing the semantic information of "gentle" and "regretful".
[0063] The parsed and obtained emotion prompt semantic information is input into the emotion feature extractor to extract features from the emotion prompt semantic information, and extract emotion-related information therefrom, such as emotion prompt features like intonation, intensity, rhythm, etc., to reflect the subtle differences of different emotions. Then, the emotion prompt features are input into the emotion mapper, and through linear or non-linear transformation, the features are mapped into a preset embedding space to generate an emotion embedding vector, specifically an embedding space that can be used for speech generation, so that the generated emotion embedding vector can be better compatible with the speech generation model, improving the naturalness and expressiveness of speech synthesis.
[0064] For example, in the voice guidance system of a hospital, the extracted emotion prompt features (related to "gentle" and "comforting", especially features such as intonation, intensity, rhythm, etc.) are mapped into a preset embedding space through the emotion mapper to generate corresponding emotion embedding vectors, which can accurately capture emotion information while ensuring that the generated emotion embedding vectors are compatible with the speech generation model, improving the naturalness and expressiveness of speech synthesis, and enhancing the emotion speech generation effect in the hospital voice guidance system.
[0065] In one embodiment, the joint speech generation model includes a Transformer model, such as Figure 4 shown, step S204 includes:
[0066] S401. Fuse and input the emotion embedding vector and the semantic embedding vector into the encoder of the Transformer model, and perform encoding processing on the fused input embedding vectors based on the multi-head attention mechanism to obtain a fused encoded representation;
[0067] S402. Input the fused encoding representation and the emotion embedding vector into the decoder of the Transformer model, and decode the fused encoding representation under the guidance of the emotion embedding vector to generate speech features with the expected emotion.
[0068] In this embodiment, when co-modeling text content and emotion information to generate speech features with the expected emotion, the joint speech generation model adopts a generation framework based on Transformer, such as the Transformer model or other variant models also based on the self-attention mechanism, such as the Conformer model, etc. The Transformer model is an end-to-end sequence generation model based on the self-attention mechanism and is widely used in natural language processing tasks. It consists of an encoder and a decoder. The encoder is responsible for encoding the input sequence into a high-dimensional context representation, and the decoder uses these representations to generate the output sequence. Moreover, the Transformer model adopts the multi-head attention mechanism. By parallelizing multiple attention heads, it captures various different patterns and relationships in the data. Each attention head focuses on different parts of the input sequence, thereby improving the model's representation ability.
[0069] When generating speech by jointly using text and emotion cues, the emotion embedding vector and the semantic embedding vector are combined into a high-dimensional vector through fusion methods such as concatenation or weighted summation, and then input into the encoder of the Transformer model. Through the multi-head attention mechanism and the feed-forward neural network, the fused input embedding vector is encoded to capture the context information and emotion information in the input sequence, generating a fused encoding representation and realizing the joint modeling of emotion and text. The encoding process through the multi-head attention mechanism enables the model to better capture the complex relationships and emotion information in the input sequence, generate a richer context representation, and provide more accurate input for subsequent decoding.
[0070] The decoder in the Transformer model utilizes the fused encoding representation generated by the encoder and the sentiment embedding vector to gradually predict and generate the output sequence, i.e., the speech features with the expected sentiment. Specifically, the decoder of the Transformer model gradually generates speech features through a multi-layer neural network. These features represent the spectral information of the speech. At each step of prediction and generation, the decoder will refer to the output of the encoder and the previously generated output, and generate the next output through the self-attention mechanism and the encoder-decoder attention mechanism. At the same time, at each step of prediction and generation, it will also refer to the sentiment embedding vector to adjust the sentiment expression in the generated speech features, such as prosodic features like intonation, intensity, and rhythm. Through the sentiment embedding vector for sentiment guidance during the decoding process, it ensures that the generated speech features conform to the expected sentiment. The generated speech features not only convey the semantic information of the text but also express the corresponding sentiment, making the entire speech segment sound more natural and coherent, and enhancing the naturalness and expressiveness of emotional speech synthesis.
[0071] In one embodiment, as Figure 5 shown, step S205 includes:
[0072] S501. Decode the speech features through a vocoder to generate an initial speech waveform corresponding to the target text and having the expected sentiment;
[0073] S502. Perform smoothing filtering on the initial speech waveform through a dynamic filter to generate the target emotional speech.
[0074] In this embodiment, a vocoder is a model that converts speech features (such as Mel spectrogram) into an actual speech waveform. Specifically, vocoders such as WaveNet, WaveRNN, and MelGAN can be used, or high-fidelity vocoders such as HiFi-GAN or WaveGlow. The speech features generated by the joint speech generation model, i.e., the Mel spectrogram, are input into the vocoder for conversion from frequency-domain features to time-domain waveforms, thereby decoding the Mel spectrogram into a high-quality speech waveform. The generated speech not only conveys the semantic information of the text but also expresses the corresponding sentiment. Further, to improve the overall quality of emotional speech synthesis, a dynamic filter is also used to perform smoothing filtering on the initial speech waveform output by the vocoder. The dynamic filter can dynamically adjust the filter parameters according to the changes in the speech signal, thereby smoothing and optimizing the speech waveform. Specifically, it smooths the emotional transitions in the initial speech waveform, and then generates the target emotional speech, making the target emotional speech sound smoother and more natural, and improving the overall effect of emotional speech synthesis.
[0075] In specific implementation, the position of emotional transition can be determined in different ways. For example, during the decoding process, the emotional embedding vector is not only used to generate speech features but also can be used to adjust the subsequent guiding dynamic filter. According to the emotional embedding vector, the intensity and change information of emotion are provided to identify the position of emotional transition in the initial speech waveform; or perform spectral analysis on the initial speech waveform to identify discontinuous or mutated parts in the spectrum, which may correspond to emotional transitions and need to be smoothed, etc. By smoothing the speech waveform at the identified position of emotional transition, the speech expression becomes more natural and vivid.
[0076] For example, in a voice-guided diagnosis system in the medical field, the initial speech waveform of "Your test results will take some time. Please wait patiently" is generated, and the emotion of "read in a gentle and slightly reassuring tone" is contained in the speech. The initial speech waveform is smoothed and filtered through a dynamic filter to identify and smooth the position of emotional transition. There is a slight change in emotion between "will take some time" and "Please wait patiently", so the filter parameters are adjusted to perform smoothing filtering on this position to generate the final target emotional speech, making the entire speech segment sound more natural and coherent.
[0077] Another example is in the intelligent customer service system of a bank. The initial speech waveform of "I'm sorry, your loan application has not passed the review" is generated, and the emotion of "read in a gentle and slightly regretful tone" is contained in the speech. The initial speech waveform is smoothed and filtered through a dynamic filter to identify and smooth the position of emotional transition. There is an emotional transition between "I'm sorry" and "has not passed the review", so the filter parameters are adjusted to perform smoothing filtering on this position to generate the final target emotional speech, improving the naturalness and quality of the speech.
[0078] In one embodiment, as Figure 6 shown, before step S202, the method further includes:
[0079] S601. Collect emotional cue training texts based on natural language descriptions;
[0080] S602. Construct multiple groups of positive and negative sample pairs according to the emotional cue training texts;
[0081] S603. Input the multiple groups of positive and negative sample pairs into the emotional encoder to be trained for contrastive learning training to obtain a trained emotional encoder.
[0082] In this embodiment, during the training phase, a large number of emotion prompt training texts based on natural language descriptions are collected to perform contrastive learning training on the emotion encoder, thereby optimizing the separability of emotion embeddings. Specifically, the emotion prompt training texts are the input data for training the emotion encoder, usually containing natural language descriptions of emotion prompts, such as "read in a gentle and slightly comforting tone" or "read in a professional and confident tone", etc. These texts can be collected in various ways, such as extracting from existing emotion annotation datasets, collecting user-annotated data through public platforms, or extracting from actual application scenarios in professional fields (such as medical and financial), etc. Abundant training texts can improve the generalization ability of the emotion encoder, enabling it to accurately understand and generate various emotion prompts.
[0083] Construct multiple pairs of positive and negative samples based on the collected emotion prompt training texts. Among them, the positive sample refers to the sample that matches the reference prompt sample, and the negative sample refers to the sample that does not match the reference prompt sample. By constructing positive and negative sample pairs, the emotion encoder can be trained to recognize the differences between different emotions. Specifically, the reference prompt sample, positive sample, and negative sample can be randomly selected from the collected emotion prompt training texts to ensure that one in each pair of samples matches the target emotion and the other does not. Of course, corresponding positive and negative samples can also be constructed based on the selected reference prompt sample through data augmentation and other means. For example, if the reference prompt sample is "gentle and slightly comforting", the positive sample can be "read in a gentle and slightly comforting tone", and the negative sample can be "read in a serious and direct tone", etc. And if the reference prompt sample is "gentle and slightly regretful", the positive sample can be "gentle and regretful", and the negative sample can be "professional and confident", etc.
[0084] Input the constructed multiple pairs of positive and negative samples into the emotion encoder to be trained for contrastive learning training. Contrastive learning is an unsupervised or semi-supervised learning method. By comparing positive and negative sample pairs, the model is trained to learn the differences between different samples. By optimizing the loss function (such as the contrastive loss function), the model can distinguish positive samples and negative samples. Through contrastive learning training, the emotion encoder can more accurately identify and generate embedding vectors of different emotions, improving the quality and naturalness of emotion speech synthesis.
[0085] In one embodiment, as Figure 7 shown, before step S204, the method further includes:
[0086] S701. Collect emotion prompt samples, sample texts, and sample voices corresponding to the sample texts;
[0087] S702. Input the emotion prompt samples into a pre-trained emotion encoder to obtain corresponding emotion sample embedding vectors;
[0088] S703. Input the sample text into a pre-trained text encoder to obtain a corresponding semantic sample embedding vector;
[0089] S704. Input the emotion sample embedding vector and the semantic sample embedding vector into a joint speech generation module to be trained for speech generation processing;
[0090] S705. Perform multi-task learning training on the joint speech generation model to be trained according to the speech generation result, the emotion prompt sample, and the sample speech to obtain a trained joint speech generation model.
[0091] In this embodiment, by collecting emotion prompt samples, sample texts, and sample speeches corresponding to the sample texts as training data, model training for multi-task learning of the joint speech generation model is performed. Among them, the emotion prompt samples can directly use the training texts when training the emotion encoder, which contain emotion prompts described in natural language, such as "read in a gentle and slightly comforting tone" or "read in a professional and confident tone", etc.; the sample texts are the text contents corresponding to the emotion prompt samples, such as "Your test results will take some time. Please wait patiently", and the sample language is the actual speech data corresponding to the sample text. By enriching the training data, the generalization ability of the joint speech generation model is improved, enabling it to accurately understand and generate various emotion prompts.
[0092] Furthermore, after collecting various samples, data augmentation processing can also be performed on the collected emotion prompt samples, sample texts, and sample speeches to generate corresponding extended samples with different emotion combinations, namely extended emotion prompt samples, extended sample texts, and extended sample speeches. Specifically, data augmentation processing can be carried out through methods such as synonym replacement and emotion dictionary expansion. For example, replace the keywords in the emotion prompt with synonyms to generate new emotion prompt samples. For example, "read in a gentle and slightly comforting tone" can be replaced with "read in a gentle and soothing tone with comfort", etc.; expand the emotion words in the emotion prompt into richer emotion expressions through the emotion dictionary, such as "gentle" can be expanded to "gentle, soft, soothing", etc. Through data augmentation, more diverse extended samples are generated, and the extended samples are jointly involved in the training of the joint speech generation model, increasing the diversity and quantity of the training data and improving the generalization ability of the model.
[0093] Similar to the above reasoning process, during the training phase, emotional prompt samples are input into a pre-trained emotion encoder to obtain corresponding emotional sample embedding vectors, and the sample texts are input into a pre-trained text encoder to obtain corresponding semantic sample embedding vectors. The emotional sample embedding vectors and semantic sample embedding vectors are input into a joint speech generation module to be trained, and speech generation processing is performed through a model with an encoder-decoder structure to generate corresponding predicted speech features and decode them to obtain predicted speech. According to the speech generation results, i.e., the predicted speech, emotional prompt samples, and sample speech, multi-task learning training is performed on the joint speech generation model to be trained to obtain a trained joint speech generation model. Specifically, according to training objectives such as the emotional consistency difference between the emotional prompt samples and the predicted speech, the emotional classification difference between the predicted speech and the sample speech, and the speech naturalness of the predicted speech, multi-task learning training is performed on the joint speech generation model. By optimizing loss functions (such as contrastive loss functions, cross-entropy loss functions, etc.), the model can simultaneously learn emotion recognition and speech generation tasks, improving the quality and naturalness of speech synthesis.
[0094] It should be noted that there is not necessarily a certain order among the above steps. Those of ordinary skill in the art can understand from the description of the embodiments of the present invention that in different embodiments, the above steps can have different execution orders, that is, they can be executed in parallel or exchanged, etc.
[0095] Further referring to Figure 8 , as an implementation of the above Figure 2 shown method, an embodiment of an emotional speech generation device is provided by the present invention. This device embodiment corresponds to the Figure 2 shown method embodiment, and this device can be specifically applied to various electronic devices.
[0096] As Figure 8 shown, the emotional speech generation device 80 described in this embodiment includes:
[0097] An acquisition module 801, configured to acquire a target text of the speech to be generated and a target emotional prompt text described in natural language;
[0098] An emotion encoding module 802, configured to input the target emotional prompt text into a pre-trained emotion encoder to obtain a corresponding emotional embedding vector;
[0099] A text encoding module 803, configured to input the target text into a pre-trained text encoder to obtain a corresponding semantic embedding vector;
[0100] A speech generation module 804, configured to input the emotion embedding vector and the semantic embedding vector into a pre-trained joint speech generation model, perform speech feature prediction by fusing the emotion embedding vector and the semantic embedding vector, and generate speech features with an expected emotion;
[0101] A speech decoding module 805, configured to perform decoding processing on the speech features to generate a target emotion speech corresponding to the target text.
[0102] The module referred to in the present invention refers to a series of computer program instruction segments capable of completing specific functions, which is more suitable for describing the execution process of emotion speech generation than a program. For the specific implementation manners of each module, please refer to the corresponding method embodiments above, and will not be elaborated herein.
[0103] In one embodiment, the emotion encoder includes a pre-trained language model, an emotion feature extractor, and an emotion mapper. The emotion encoding module 802 includes:
[0104] A semantic parsing unit, configured to input the target emotion prompt text into the pre-trained language model, perform semantic parsing on the target emotion prompt text, and obtain emotion prompt semantic information;
[0105] An emotion feature extraction unit, configured to input the emotion prompt semantic information into the emotion feature extractor, perform feature extraction on the emotion prompt semantic information, and obtain corresponding emotion prompt features;
[0106] An embedding unit, configured to perform embedding mapping on the emotion prompt features through the emotion mapper, map the emotion prompt features to a preset embedding space for speech generation, and obtain the emotion embedding vector.
[0107] In one embodiment, the joint speech generation model includes a Transformer model. The speech generation module 804 includes:
[0108] A fusion encoding unit, configured to fuse and input the emotion embedding vector and the semantic embedding vector into the encoder of the Transformer model, perform encoding processing on the fused input embedding vectors based on the multi-head attention mechanism, and obtain a fusion encoding representation;
[0109] A joint decoding unit, configured to input the fusion encoding representation and the emotion embedding vector into the decoder of the Transformer model, and perform decoding on the fusion encoding representation under the guidance of the emotion embedding vector to generate speech features with an expected emotion.
[0110] In one embodiment, the speech decoding module 805 includes:
[0111] A voice decoding unit, configured to decode the voice features through a vocoder to generate an initial voice waveform corresponding to the target text and having an expected emotion;
[0112] A filtering unit, configured to perform smoothing filtering on the initial voice waveform through a dynamic filter to generate the target emotion voice.
[0113] In one embodiment, the apparatus 80 further includes:
[0114] A prompt text acquisition module, configured to acquire emotion prompt training texts described in natural language;
[0115] A sample construction module, configured to construct multiple groups of positive and negative sample pairs according to the emotion prompt training texts;
[0116] A contrast training module, configured to input the multiple groups of positive and negative sample pairs into a to-be-trained emotion encoder for contrast learning training to obtain a trained emotion encoder.
[0117] In one embodiment, the apparatus 80 further includes:
[0118] A training sample acquisition module, configured to acquire emotion prompt samples, sample texts, and sample voices corresponding to the sample texts;
[0119] A sample emotion encoding module, configured to input the emotion prompt samples into a pre-trained emotion encoder to obtain corresponding emotion sample embedding vectors;
[0120] A sample text encoding module, configured to input the sample texts into a pre-trained text encoder to obtain corresponding semantic sample embedding vectors;
[0121] A sample voice generation module, configured to input the emotion sample embedding vectors and semantic sample embedding vectors into a to-be-trained joint voice generation module for voice generation processing;
[0122] A multi-task training module, configured to perform multi-task learning training on a to-be-trained joint voice generation model according to the voice generation result, the emotion prompt samples, and the sample voices to obtain a trained joint voice generation model.
[0123] In one embodiment, the apparatus 80 further includes:
[0124] A data augmentation module, configured to perform data augmentation processing on the emotion prompt samples, sample texts, and sample voices to generate corresponding extended samples with different emotion combinations.
[0125] In the above embodiments, the present invention discloses an emotional speech generation device, which obtains a target text of the speech to be generated and a target emotional prompt text described in natural language; inputs the target emotional prompt text into a pre-trained emotional encoder to obtain a corresponding emotional embedding vector; inputs the target text into a pre-trained text encoder to obtain a corresponding semantic embedding vector; inputs the emotional embedding vector and the semantic embedding vector into a pre-trained joint speech generation model to perform speech feature prediction by fusing the emotional embedding vector and the semantic embedding vector, and generates a speech feature with an expected emotion; performs decoding processing on the speech feature to generate a target emotional speech corresponding to the target text. By using an emotional prompt text described in natural language to control the emotion of the generated speech, the flexibility of the emotional speech generation control interaction is improved, and through the speech generation by fusing emotional embedding and semantic embedding, the generated speech is natural and vivid in emotion, effectively improving the synthesis effect of emotional speech.
[0126] Another embodiment of the present invention provides a computer device, as Figure 9 shown, the computer device 90 includes:
[0127] One or more processors 901 and a memory 902, Figure 9 Taking one processor 901 as an example for introduction, the processor 901 and the memory 902 can be connected through a bus or other means, Figure 9 Taking the connection through the bus as an example.
[0128] The processor 901 is used to complete various control logics of the computer device 90. It can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components. Moreover, the processor 901 can also be any conventional processor, microprocessor, or state machine. The processor 901 can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP and / or any other such configuration.
[0129] The memory 902, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as program instructions corresponding to the emotional speech generation method in the embodiments of the present invention. The processor 901 executes various functional applications and data processing of the computer device 90 by running the non-volatile software programs, instructions, and units stored in the memory 902, that is, implements the emotional speech generation method in the above method embodiments.
[0130] The memory 902 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device 90, etc. In addition, the memory 902 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 902 optionally includes a memory remotely provided relative to the processor 901, and these remote memories can be connected to the computer device 90 through a network. Examples of the above networks include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. One or more units are stored in the memory 902 and, when executed by one or more processors 901, perform the steps of the emotional speech generation method in any of the above method embodiments.
[0131] In the above embodiments, the present invention discloses a computer device. By obtaining a target text of the speech to be generated and a target emotion prompt text described in natural language; inputting the target emotion prompt text into a pre-trained emotion encoder to obtain a corresponding emotion embedding vector; inputting the target text into a pre-trained text encoder to obtain a corresponding semantic embedding vector; inputting the emotion embedding vector and the semantic embedding vector into a pre-trained joint speech generation model to perform a fused speech feature prediction on the emotion embedding vector and the semantic embedding vector to generate a speech feature with an expected emotion; performing a decoding process on the speech feature to generate a target emotion speech corresponding to the target text. The emotion of the generated speech is controlled by the emotion prompt text described in natural language, which improves the flexibility of the emotional speech generation control interaction. Moreover, through the fused speech generation of emotion embedding and semantic embedding, the generated speech is natural and vivid in emotion, effectively improving the synthesis effect of the emotional speech.
[0132] The embodiments of the present invention provide a non-volatile computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by one or more processors, they perform the steps of the emotional speech generation method in any of the above method embodiments.
[0133] In the above embodiments, the present invention discloses a non-volatile computer-readable storage medium. By obtaining a target text of the speech to be generated and a target emotion prompt text described in natural language; inputting the target emotion prompt text into a pre-trained emotion encoder to obtain a corresponding emotion embedding vector; inputting the target text into a pre-trained text encoder to obtain a corresponding semantic embedding vector; inputting the emotion embedding vector and the semantic embedding vector into a pre-trained joint speech generation model to perform fused speech feature prediction on the emotion embedding vector and the semantic embedding vector, generating speech features with an expected emotion; and performing decoding processing on the speech features to generate a target emotion speech corresponding to the target text. By using an emotion prompt text described in natural language to control the emotion of the generated speech, the flexibility of the emotion speech generation control interaction is improved, and through the fused speech generation of emotion embedding and semantic embedding, the generated speech is natural and the emotion is vivid, effectively improving the synthesis effect of the emotion speech.
[0134] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0135] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0136] In summary, in a method, apparatus, device, and medium for generating emotional speech disclosed by the present invention, the method includes: obtaining a target text of the speech to be generated and a target emotional prompt text described in natural language; inputting the target emotional prompt text into a pre-trained emotional encoder to obtain a corresponding emotional embedding vector; inputting the target text into a pre-trained text encoder to obtain a corresponding semantic embedding vector; inputting the emotional embedding vector and the semantic embedding vector into a pre-trained joint speech generation model to perform a speech feature prediction of fusing the emotional embedding vector and the semantic embedding vector, and generating a speech feature with an expected emotion; and performing a decoding process on the speech feature to generate a target emotional speech corresponding to the target text. By using an emotional prompt text described in natural language to control the emotion of the generated speech, the flexibility of the emotional speech generation control interaction is improved, and through the speech generation of fusing emotional embedding and semantic embedding, the generated speech is natural and vivid in emotion, effectively improving the synthesis effect of the emotional speech.
[0137] Certainly, those of ordinary skill in the art can understand that all or part of the processes in implementing the methods of the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it can include the processes of the above method embodiments. The storage medium can be a memory, a magnetic disk, a floppy disk, a flash memory, an optical memory, etc.
[0138] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for illustrative introduction and do not represent actual use. It should be understood that the application of the present invention is not limited to the above examples, and those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A method for generating emotional speech, characterized in that: include: Obtaining a target text of speech to be generated and a target emotional prompt text based on a natural language description; Inputting the target emotion prompt text into a pre-trained emotion encoder to obtain a corresponding emotion embedding vector; Inputting the target text into a pre-trained text encoder to obtain a corresponding semantic embedding vector; Inputting the emotion embedding vector and the semantic embedding vector into a pre-trained joint speech generation model, performing speech feature prediction on the emotion embedding vector and the semantic embedding vector to generate speech features with expected emotions; The speech features are decoded to generate target emotional speech corresponding to the target text.
2. The emotional speech generation method according to claim 1, characterized in that: The emotion encoder includes a pre-trained language model, an emotion feature extractor and an emotion mapper, and the target emotion prompt text is input into the pre-trained emotion encoder to obtain a corresponding emotion embedding vector, including: Inputting the target emotion prompt text into the pre-trained language model, performing semantic analysis on the target emotion prompt text, and obtaining emotion prompt semantic information; Inputting the emotion cue semantic information into the emotion feature extractor, performing feature extraction on the emotion cue semantic information, and obtaining corresponding emotion cue features; The emotion prompt feature is embedded and mapped by the emotion mapper, and the emotion prompt feature is mapped to a preset embedding space for speech generation to obtain the emotion embedding vector.
3. The emotional speech generation method according to claim 1, characterized in that: The joint speech generation model includes a Transformer model, and the emotion embedding vector and the semantic embedding vector are input into the pre-trained joint speech generation model, and the fused speech feature prediction of the emotion embedding vector and the semantic embedding vector is performed to generate speech features with expected emotions, including: The sentiment embedding vector and the semantic embedding vector are fused and input into the encoder of the Transformer model, and the fused input embedding vector is encoded based on a multi-head attention mechanism to obtain a fused encoding representation; The fused coded representation and the emotion embedding vector are input into the decoder of the Transformer model, and the fused coded representation is decoded under the guidance of the emotion embedding vector to generate speech features with expected emotions.
4. The emotional speech generation method according to claim 1, characterized in that: The decoding process of the speech feature to generate the target emotional speech corresponding to the target text includes: Decoding the speech features through a vocoder to generate an initial speech waveform corresponding to the target text and having the expected emotion; The target emotional speech is generated by performing smoothing filtering on the initial speech waveform through a dynamic filter.
5. The emotional speech generation method according to claim 1, characterized in that: Before inputting the target emotion prompt text into a pre-trained emotion encoder to obtain a corresponding emotion embedding vector, the method further includes: Collect emotional prompt training text based on natural language description; Constructing multiple groups of positive and negative sample pairs according to the emotional cue training text; The multiple groups of positive and negative sample pairs are input into the emotion encoder to be trained for contrast learning training to obtain a trained emotion encoder.
6. The emotional speech generation method according to claim 1, characterized in that: Before inputting the emotion embedding vector and the semantic embedding vector into the pre-trained joint speech generation model, the method further includes: Collecting emotional prompt samples, sample texts, and sample speech corresponding to the sample texts; Inputting the emotion prompt sample into a pre-trained emotion encoder to obtain a corresponding emotion sample embedding vector; Inputting the sample text into a pre-trained text encoder to obtain a corresponding semantic sample embedding vector; Inputting the emotion sample embedding vector and the semantic sample embedding vector into the joint speech generation module to be trained for speech generation processing; Multi-task learning training is performed on the joint speech generation model to be trained according to the speech generation result, the emotion prompt sample and the sample speech to obtain a trained joint speech generation model.
7. The emotional speech generation method according to claim 6, characterized in that: After collecting the emotion prompt sample, the sample text and the sample speech corresponding to the sample text, the method further includes: Data enhancement processing is performed on the emotion prompt samples, sample texts and sample voices to generate corresponding extended samples with different emotion combinations.
8. An emotional speech generating device, characterized in that: include: An acquisition module is used to acquire a target text of the speech to be generated and a target emotional prompt text based on a natural language description; The emotion encoding module is used to input the target emotion prompt text into a pre-trained emotion encoder to obtain a corresponding emotion embedding vector; A text encoding module, used for inputting the target text into a pre-trained text encoder to obtain a corresponding semantic embedding vector; A speech generation module, used for inputting the emotion embedding vector and the semantic embedding vector into a pre-trained joint speech generation model, performing speech feature prediction on the fused emotion embedding vector and the semantic embedding vector, and generating speech features with expected emotions; The speech decoding module is used to decode the speech features to generate target emotional speech corresponding to the target text.
9. A computer device, characterized in that: comprising at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the emotional speech generation method described in any one of claims 1-7.
10. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores computer-executable instructions, which, when executed by one or more processors, enable the one or more processors to execute the emotional speech generation method described in any one of claims 1-7.
Citation Information
Cited By
Speech synthesis method and device, electronic equipment and storage medium
CN120895024A
Speech synthesis method and device, electronic equipment and storage medium
CN120895024B
Speech synthesis method and device, electronic equipment and storage medium
CN121034283A
Speech synthesis method and device, electronic equipment and storage medium
CN121034283B