Emotion speech synthesis method based on emotion analysis
By combining the BERT text sentiment analysis model and end-to-end speech synthesis module, a multi-emotional speech synthesis system is created, which solves the problem that traditional speech synthesis is difficult to generate multi-emotional and natural speech, and achieves a higher sense of reality and user experience.
Patent Information
- Application Number
- CN202510135855.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional speech synthesis methods are difficult to generate multi-emotional and natural voice, and cannot meet the rich and natural emotional expression needs in application scenarios such as smart customer service and virtual assistants.
The BERT-based text sentiment analysis model is used to create an emotional speech synthesis model and end-to-end Char2Wav or Tacotron module, and multi-emotional speech synthesis is achieved through the emotion speech library and emotion control model.
It improves the authenticity and user experience of speech synthesis, and the generated voice is more natural and smooth, and can flexibly adjust emotional colors and adaptive adjustments to different emotional texts.
Smart Images

Figure CN120108377A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an emotional speech synthesis method based on emotional analysis. Background Art
[0002] Speech synthesis, also known as text-to-speech technology, can convert any text information into standard and fluent speech in real time, which is equivalent to installing an artificial mouth on the machine. Traditionally, there are the following methods for speech synthesis:
[0003] Method 1: Multi-emotion speech synthesis method based on speech feature adjustment. This method realizes emotion synthesis by adjusting the basic features of speech (such as pitch, speed, loudness, duration, etc.); for example, parameters such as pitch and speech speed can be adjusted according to emotional labels (such as "happy", "angry", "sad") to imitate the expression of different emotions. Method 2: Emotional speech synthesis method based on splicing. This method synthesizes emotional speech by splicing pre-recorded speech fragments (such as words, phrases, etc.) on demand; during synthesis, the corresponding speech samples can be selected for splicing according to the input text and its emotional label. Method 3: Emotional speech synthesis method based on HMM (Hidden Markov Model). This method uses HMM to model the dynamic changes of speech and synthesizes by analyzing the relationship between the pattern of emotional changes and the audio signal; HMM can be modeled by considering the information such as phonemes and rhythm of speech and combining emotional labels to realize emotion synthesis.
[0004] However, traditional methods can usually only generate speech with a single emotion, lack flexibility, and the generated speech expression is not natural enough. When faced with application scenarios that require emotional interaction, such as intelligent customer service, virtual assistants, emotional voice navigation, etc., they cannot provide rich and practical emotional expressions, resulting in poor user experience.
[0005] Therefore, how to provide an emotional speech synthesis method based on sentiment analysis to improve the realism and user experience of emotional speech synthesis has become a technical problem that needs to be solved urgently. Summary of the invention
[0006] The technical problem to be solved by the present invention is to provide an emotional speech synthesis method based on emotion analysis, so as to improve the realism and user experience of emotional speech synthesis.
[0007] The present invention is implemented as follows: an emotional speech synthesis method based on emotional analysis comprises the following steps:
[0008] Step S1: creating a text sentiment analysis model for sentiment classification based on BERT, and training the text sentiment analysis model;
[0009] Step S2, creating an emotional speech database, obtaining a large amount of speech data containing different emotions, extracting speech features of each of the speech data, annotating each of the speech data based on the speech features and the emotion category, and storing the annotated speech data in the emotional speech database;
[0010] Step S3, obtaining a large amount of historical emotional texts, classifying each historical emotional text through the text emotion analysis model, obtaining the emotion category corresponding to each of the historical emotional texts, matching the corresponding voice data from the emotional voice library based on the emotion category, and constructing a data set based on the matched historical emotional texts and voice data;
[0011] Step S4, creating an emotional speech synthesis model based on the end-to-end Char2Wav module or the Tacotron module, and training the emotional speech synthesis model through the data set;
[0012] Step S5: creating an emotion control model for regulating emotion, and training the emotion control model using the data set;
[0013] Step S6: creating a multi-emotion speech synthesis system based on the emotion speech synthesis model and the emotion control model, and verifying the multi-emotion speech synthesis system;
[0014] Step S7: performing speech synthesis by using the verified multi-emotion speech synthesis system.
[0015] Furthermore, in step S1, the text sentiment analysis model consists of an input embedding layer, several Transformer encoders, and a sentiment classification output layer;
[0016] The input embedding layer is used to convert each word of the sentiment text into an embedding vector:
[0017] X=[E(w 1 ),E(w 2 ),...,E(w n )]+P;
[0018] Where X represents the embedding vector; w n Represents the nth word in the sentiment text; E() represents the embedding function; P represents the position encoding of the word;
[0019] The Transformer encoder consists of a self-attention network and several layers of feedforward neural networks, which are used to encode the embedding vector to obtain the feature vector h. The formula is:
[0020] H l=Attention(H l-1 )+H l-1 ;
[0021] Among them, H l represents the lth layer of feedforward neural network; H l-1 represents the l-1th layer of feedforward neural network; Attention() represents the self-attention network; the input of the first layer of the feedforward neural network is the embedding vector X, and the output of the last layer of the feedforward neural network is the feature vector h;
[0022] The sentiment classification output layer is used to perform sentiment prediction based on the feature vector h and output the sentiment category:
[0023]
[0024] in, represents the predicted probability of the i-th emotion category; exp() represents the exponential function; W represents the weight matrix; b represents the bias term; C represents the total number of emotion categories; (Wh+b) i represents the i-th eigenvector adjusted based on the weight matrix and bias term;
[0025] Furthermore, in step S1, the text sentiment analysis model is trained based on the cross entropy loss function:
[0026]
[0027] Among them, L represents the loss value of the cross entropy loss function; y i represents the true sentiment category corresponding to the i-th word; p i represents the probability distribution of the predicted emotion category; C represents the total number of emotion categories.
[0028] Furthermore, in step S2, the voice feature extraction formula is:
[0029] MFCC(t)=DCT(log|S(t)|);
[0030] Among them, MFCC(t) represents the speech features of the speech data at time t; DCT() represents the discrete cosine transform function; S(t) represents the Mel spectrum obtained by performing short-time Fourier transform on the speech data at time t.
[0031] Furthermore, in step S4, the Char2Wav module is used to encode the emotional text through a bidirectional recurrent neural network to obtain a hidden state h t , for the hidden state h t Decode to get acoustic features Then use SampleRNN to transform the acoustic features Convert to raw audio waveform.
[0032] Furthermore, the hidden state h t The calculation formula is:
[0033]
[0034] Among them, h t represents the hidden state at time t; represents the forward hidden state at time t; represents the backward hidden state at time t; [·; ·] represents the concatenation operation; RNN forward () represents the forward recurrent neural network in the bidirectional recurrent neural network; RNN backward () represents the backward recurrent neural network in the bidirectional recurrent neural network; x t A character vector representing the sentiment text input at time t; represents the forward hidden state at time t-1; Represents the backward hidden state at time t+1;
[0035] The acoustic characteristics The calculation formula is:
[0036]
[0037] in, represents the acoustic characteristics at time t; Represents the acoustic features at time t-1; RNN decoder () represents the RNN decoder.
[0038] Furthermore, in step S4, the Tacotron module is used to extract text features h from the sentiment text through the CBHG network. t' , based on the attention mechanism for the text feature h t' Align to get the context vector c t' , based on the context vector c t' Compute Mel-spectrogram subgraph Based on each Mel spectrum subgraph Calculate the Mel spectrum The Mel-spectrogram is transformed into Convert to raw audio waveform.
[0039] Furthermore, in step S5, the emotion control model is used to input the emotion label E z Adjust the hidden state h of the emotional speech synthesis model z , get the optimized hidden state h' z, based on the optimized hidden state h' z The original audio waveform output by the emotional speech synthesis model is decoded to obtain synthesized speech.
[0040] Furthermore, the optimized hidden state h' z The calculation formula is:
[0041] h' z =h z *W E (E z );
[0042] Among them, W E (E z ) represents the sentiment label E z The associated weight matrix.
[0043] Furthermore, in step S6, the verification of the multi-emotion speech synthesis system is specifically as follows:
[0044] The multi-emotion speech synthesis system is verified based on MOS scores and emotion recognition accuracy.
[0045] The advantages of the present invention are:
[0046] Create a text sentiment analysis model for sentiment classification through BERT, and train the text sentiment analysis model; then obtain a large amount of speech data containing different emotions, extract the speech features of each speech data, annotate each speech data based on the speech features and sentiment categories, and store the annotated speech data in the created sentiment speech library; then obtain a large amount of historical sentiment texts, classify each historical sentiment text through the text sentiment analysis model, obtain the sentiment category corresponding to each historical sentiment text, match the corresponding speech data from the sentiment speech library based on the sentiment category, and build a data set based on the matched historical sentiment texts and speech data; then create an emotion speech synthesis model based on the end-to-end Char2Wav module or Tacotron module, and train the emotion speech synthesis model through the data set; create an emotion control model for regulating emotions, and train the emotion control model through the data set; create a multi-emotion speech synthesis system based on the emotion speech synthesis model and the emotion control model, and train the multi-emotion speech synthesis system The verification was carried out, and finally the speech synthesis was performed through the verified multi-emotional speech synthesis system; that is, the text sentiment analysis model created by BERT was used for sentiment classification, which can accurately identify the emotion category from the emotional text, not only based on simple speech features (such as pitch, speaking speed, etc.) to adjust the emotion, but also can identify more complex and delicate emotional changes (such as the emotional expression of a text in different contexts); through the end-to-end emotional speech synthesis model, no manually designed intermediate steps are required from the emotional text input to the synthesized speech output, and the generated synthesized speech is more natural and fluent compared with the splicing-based synthesis method; through the emotion control model to adjust the emotion, the emotional color of the synthesized speech can be flexibly adjusted during the synthesis process, and the combination with the text sentiment analysis model can further improve the control accuracy of the emotion; through the combination of the embedding vector and the emotional speech synthesis model, the generated speech features can be adaptively adjusted for different emotional texts, and the naturalness and diversity of emotional expression can be improved, which ultimately greatly improves the realism and user experience of emotional speech synthesis. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The present invention will be further described below in conjunction with embodiments with reference to the accompanying drawings.
[0048] Figure 1 It is a circuit principle block diagram of an emotional speech synthesis method based on emotional analysis of the present invention. DETAILED DESCRIPTION
[0049] The technical solution in the embodiments of the present application has the following overall idea: by performing sentiment classification through the text sentiment analysis model created by BERT, sentiment categories can be accurately identified from sentiment texts. It not only adjusts sentiment based on simple speech features, but can also identify more complex and delicate sentiment changes; through an end-to-end emotional speech synthesis model, no manually designed intermediate steps are required from emotional text input to synthesized speech output, and the generated synthesized speech is more natural and fluent compared to the splicing-based synthesis method; by adjusting emotions through the emotion control model, the emotional color of the synthesized speech can be flexibly adjusted during the synthesis process, and the combination with the text sentiment analysis model can further improve the control accuracy of emotions; through the combination of embedding vectors and the emotional speech synthesis model, the generated speech features can be adaptively adjusted for different emotional texts, thereby improving the naturalness and diversity of emotional expression, thereby improving the realism of emotional speech synthesis and user experience.
[0050] Please refer to Figure 1 As shown, a preferred embodiment of the emotional speech synthesis method based on emotional analysis of the present invention comprises the following steps:
[0051] Step S1, creating a text sentiment analysis model for sentiment classification based on BERT (Bidirectional Encoder Representations from Transformers), and training the text sentiment analysis model; that is, classifying sentiment texts based on deep learning, identifying sentiment categories (such as happiness, sadness, anger, etc.) in sentiment texts, optimizing sentiment classification accuracy, and ensuring efficient recognition of multi-sentiment texts; BERT has powerful context modeling capabilities;
[0052] Step S2, creating an emotional speech library, obtaining a large amount of speech data containing different emotions, extracting speech features of each of the speech data, annotating each of the speech data based on the speech features and the emotion category, storing each of the annotated speech data in the emotional speech library, providing materials for the training of the emotional speech synthesis model and the emotion control model, and exploring speech feature extraction and representation methods (such as emotion embedding, emotion labeling) to support flexible regulation of emotions in speech synthesis;
[0053] Step S3, obtaining a large amount of historical emotional texts, classifying each historical emotional text through the text emotion analysis model, obtaining the emotion category corresponding to each of the historical emotional texts, matching the corresponding voice data from the emotional voice library based on the emotion category, and constructing a data set based on the matched historical emotional texts and voice data;
[0054] Step S4, creating an emotional speech synthesis model based on the end-to-end Char2Wav module or Tacotron module, training the emotional speech synthesis model through the data set, and during the training process, adjusting the hyperparameters (attention mechanism, decoder structure) to achieve flexible regulation of speech emotion, and using a multi-task learning method to optimize speech quality and emotional expression at the same time; the emotional speech synthesis model supports dynamic adjustment of parameters such as intonation, speech speed, and sound intensity during the speech synthesis process to ensure the natural expression of different emotions;
[0055] Step S5: creating an emotion control model for regulating emotion, and training the emotion control model using the data set;
[0056] Step S6: creating a multi-emotion speech synthesis system based on the emotion speech synthesis model and the emotion control model, and verifying the multi-emotion speech synthesis system to improve the naturalness and accuracy of speech emotion expression;
[0057] Step S7: performing speech synthesis by using the verified multi-emotion speech synthesis system.
[0058] In step S1, the text sentiment analysis model consists of an input embedding layer, several Transformer encoders, and a sentiment classification output layer;
[0059] The input embedding layer is used to convert each word of the sentiment text into an embedding vector:
[0060] X=[E(w 1 ),E(w 2 ),...,E(w n )]+P;
[0061] Where X represents the embedding vector; w n Represents the nth word in the sentiment text; E() represents the embedding function; P represents the position encoding of the word;
[0062] The Transformer encoder consists of a self-attention network and several layers of feed-forward neural networks, which are used to encode the embedding vector to obtain the feature vector h. The formula is:
[0063] H l =Attention(H l-1 )+H l-1 ;
[0064] Among them, H l represents the lth layer of feedforward neural network; H l-1represents the l-1th layer of feedforward neural network; Attention() represents the self-attention network; the input of the first layer of the feedforward neural network is the embedding vector X, and the output of the last layer of the feedforward neural network is the feature vector h;
[0065] The sentiment classification output layer is used to perform sentiment prediction based on the feature vector h and output the sentiment category:
[0066]
[0067] in, represents the predicted probability of the i-th emotion category; exp() represents the exponential function; W represents the weight matrix; b represents the bias term; C represents the total number of emotion categories; (Wh+b) i represents the i-th eigenvector adjusted based on the weight matrix and bias term;
[0068] In step S1, the text sentiment analysis model is trained based on the cross entropy loss function:
[0069]
[0070] Among them, L represents the loss value of the cross entropy loss function; y i represents the true sentiment category corresponding to the i-th word; p i represents the probability distribution of the predicted emotion category; C represents the total number of emotion categories.
[0071] In step S2, the voice feature extraction formula is:
[0072] MFCC(t)=DCT(logS(t));
[0073] Among them, MFCC(t) represents the speech features of the speech data at time t; DCT() represents the discrete cosine transform function; S(t) represents the Mel spectrum obtained by performing short-time Fourier transform on the speech data at time t.
[0074] The speech feature is the MFCC (Mel Frequency Cepstral Coefficient) feature. The MFCC feature is a common speech feature that can effectively represent the frequency characteristics of speech data. It is obtained by performing pre-emphasis, window function processing, FFT (Fast Fourier Transform), Mel filter bank processing and other steps on the speech data, as follows:
[0075] Enhance the high-frequency components of speech data through a pre-emphasis filter:
[0076] x[n]=s[n]-as[n-1];
[0077] Wherein, s[n] represents the original speech data at time point n, n represents the time index, which is used to represent the position in the speech data, t is a continuous time variable, which represents the time point of the speech data, and n is the discretized version of t; s[n-1] represents the original speech data at time point n-1; a represents the pre-emphasis coefficient, and the value is preferably 0.95; x[n] represents the speech data after pre-emphasis;
[0078] The speech data is divided into frames at intervals of 20-40 milliseconds, and a window function (such as a Hamming window) is added to each frame to reduce edge effects. Then, each frame is subjected to a fast Fourier transform (FFT) to obtain a frequency domain representation X[n]. The frequency domain representation X[n] is then input into a series of Mel filter banks to obtain a Mel spectrum:
[0079]
[0080] Among them, S k represents the Mel spectrum; N represents the length of the frequency domain representation X[n], that is, the number of frequency domain samples after each frame is fast Fourier transformed; H k (n) represents the response of the kth Mel filter;
[0081] Then, the Mel spectrum S k Perform a logarithmic compression operation:
[0082] logS k =log(S k +ε);
[0083] Among them, ε represents an infinitesimally small constant, which is used to prevent numerical instability in logarithmic operations;
[0084] Then perform discrete cosine transform to obtain MFCC features:
[0085]
[0086] Among them, K represents the number of Mel filters; m represents the index of MFCC features.
[0087] In step S4, the Char2Wav module is used to encode the emotional text through a bidirectional recurrent neural network (Bi-RNN) to obtain a hidden state h t , for the hidden state h t Decode to get acoustic features (Mel spectrum), and then use SampleRNN to transform the acoustic features Convert to raw audio waveform. SampleRNN is a RNN-based generative model that can generate audio waveforms frame by frame.
[0088] The hidden state ht The calculation formula is:
[0089]
[0090] Among them, h t represents the hidden state at time t; represents the forward hidden state at time t; represents the backward hidden state at time t; [·; ·] represents the concatenation operation; RNN forward () represents the forward recurrent neural network in the bidirectional recurrent neural network; RNN backward () represents the backward recurrent neural network in the bidirectional recurrent neural network; x t A character vector representing the sentiment text input at time t; represents the forward hidden state at time t-1; Represents the backward hidden state at time t+1;
[0091] The acoustic characteristics The calculation formula is:
[0092]
[0093] in, represents the acoustic characteristics at time t; Represents the acoustic features at time t-1; RNN decoder () represents the RNN decoder.
[0094] In step S4, the Tacotron module is used to extract text features h from the sentiment text through the CBHG network (Convolution Bank, Highway Network, and BiRNN). t' , based on the attention mechanism for the text feature h t' Align to get the context vector c t' , based on the context vector c t' Compute Mel-spectrogram subgraph Based on each Mel spectrum subgraph Calculate the Mel spectrum The Mel-spectrogram is transformed into Convert to raw audio waveform.
[0095] The text feature h t' The formula for alignment is:
[0096] c t' =Attention(h t' ,c t'-1 );
[0097] Among them, c t' represents the context vector at time t'; c t'-1 Represents the context vector at time t'-1; Attention() represents the attention mechanism;
[0098] The Mel-spectrogram subgraph The calculation formula is:
[0099]
[0100] Among them, Decoder() represents the Mel spectrum decoding function;
[0101] The formula of the Griffin-Lim algorithm is:
[0102]
[0103] in, Represents the initial estimated waveform; ISTFT() represents the inverse short-time Fourier transform; represents the estimated waveform of the k+1th iteration; represents element-by-element multiplication; i represents an imaginary unit, which performs complex exponential operation on the phase information to recover the phase information; ∠ represents the phase information; represents the estimated waveform of the kth iteration.
[0104] In step S5, the emotion control model is used to input the emotion label E z Adjust the hidden state h of the emotional speech synthesis model z , get the optimized hidden state h' z , based on the optimized hidden state h' z The original audio waveform output by the emotional speech synthesis model is decoded to obtain synthesized speech.
[0105] The optimized hidden state h' z The calculation formula is:
[0106] h' z =h z *W E (E z );
[0107] Among them, W E (E z ) represents the sentiment label E z The associated weight matrix.
[0108] In speech synthesis, in addition to accurately converting emotional text into synthesized speech, the generated synthesized speech also needs to have a certain degree of naturalness and expressiveness. This means that it must not only be pronounced correctly but also be able to convey appropriate emotions. The emotion control model is designed to adjust the output according to the given emotion label so that the generated synthesized speech can reflect the specified emotional state. Multi-task learning is used during the training process of the emotion control model to balance speech quality and emotional expression.
[0109] Through multi-task learning, voice quality and emotional expression can be optimized simultaneously, avoiding the trade-off problem between voice quality and emotional expression in traditional methods; through joint training, emotion recognition and speech synthesis can promote each other and further improve the overall performance.
[0110] Emotion regulation is mainly achieved by modifying the hidden state, which can be regarded as an internal representation of the model for the input data, which contains most of the information of the input data. In order to introduce emotional factors, a weight matrix related to the emotional label is introduced, which is dynamically generated according to the given emotional label.
[0111] First, the emotional label needs to be encoded into a form that can be understood by the model, usually a vector representing one or more emotional states, such as happiness, sadness, etc.; then a weight matrix is generated based on the emotional label. This weight matrix is used to weight the original hidden state in order to adjust the information related to specific emotions in the hidden state; the original hidden state is adjusted by dot multiplication with the weight matrix, so that the original hidden state is endowed with more emotional information, which will be used to generate speech with specific emotional colors in the subsequent decoding process; finally, the optimized hidden state after emotional regulation is sent to the decoder, and the decoder generates the final synthesized speech based on the optimized hidden state; since the optimized hidden state already contains emotional information, the generated synthesized speech will also have corresponding emotional colors.
[0112] In step S6, the verification of the multi-emotion speech synthesis system is specifically as follows:
[0113] The multi-emotion speech synthesis system is verified based on MOS (Mean Opinion Score) scoring and emotion recognition accuracy.
[0114] MOS score is used to evaluate the naturalness and intelligibility of speech synthesis. The formula is:
[0115]
[0116] Among them, N z Indicates the number of test samples; Score i Represents the score of the th sample.
[0117] The formula for the emotion recognition accuracy is:
[0118]
[0119] In summary, the advantages of the present invention are:
[0120] Create a text sentiment analysis model for sentiment classification through BERT, and train the text sentiment analysis model; then obtain a large amount of speech data containing different emotions, extract the speech features of each speech data, annotate each speech data based on the speech features and sentiment categories, and store the annotated speech data in the created sentiment speech library; then obtain a large amount of historical sentiment texts, classify each historical sentiment text through the text sentiment analysis model, obtain the sentiment category corresponding to each historical sentiment text, match the corresponding speech data from the sentiment speech library based on the sentiment category, and build a data set based on the matched historical sentiment texts and speech data; then create an emotion speech synthesis model based on the end-to-end Char2Wav module or Tacotron module, and train the emotion speech synthesis model through the data set; create an emotion control model for regulating emotions, and train the emotion control model through the data set; create a multi-emotion speech synthesis system based on the emotion speech synthesis model and the emotion control model, and train the multi-emotion speech synthesis system The verification was carried out, and finally the speech synthesis was performed through the verified multi-emotional speech synthesis system; that is, the text sentiment analysis model created by BERT was used for sentiment classification, which can accurately identify the emotion category from the emotional text, not only based on simple speech features (such as pitch, speaking speed, etc.) to adjust the emotion, but also can identify more complex and delicate emotional changes (such as the emotional expression of a text in different contexts); through the end-to-end emotional speech synthesis model, no manually designed intermediate steps are required from the emotional text input to the synthesized speech output, and the generated synthesized speech is more natural and fluent compared with the splicing-based synthesis method; through the emotion control model to adjust the emotion, the emotional color of the synthesized speech can be flexibly adjusted during the synthesis process, and the combination with the text sentiment analysis model can further improve the control accuracy of the emotion; through the combination of the embedding vector and the emotional speech synthesis model, the generated speech features can be adaptively adjusted for different emotional texts, and the naturalness and diversity of emotional expression can be improved, which ultimately greatly improves the realism and user experience of emotional speech synthesis.
[0121] Although the specific implementation modes of the present invention are described above, those skilled in the art should understand that the specific implementation modes described are only illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included in the scope of protection of the claims of the present invention.
Claims
1. An emotional speech synthesis method based on sentiment analysis, characterized in that: The steps include: Step S1: creating a text sentiment analysis model for sentiment classification based on BERT, and training the text sentiment analysis model; Step S2, creating an emotional speech database, obtaining a large amount of speech data containing different emotions, extracting speech features of each of the speech data, annotating each of the speech data based on the speech features and the emotion category, and storing the annotated speech data in the emotional speech database; Step S3, obtaining a large amount of historical emotional texts, classifying each historical emotional text through the text emotion analysis model, obtaining the emotion category corresponding to each of the historical emotional texts, matching the corresponding voice data from the emotional voice library based on the emotion category, and constructing a data set based on the matched historical emotional texts and voice data; Step S4, creating an emotional speech synthesis model based on the end-to-end Char2Wav module or the Tacotron module, and training the emotional speech synthesis model through the data set; Step S5: creating an emotion control model for regulating emotion, and training the emotion control model using the data set; Step S6: creating a multi-emotion speech synthesis system based on the emotion speech synthesis model and the emotion control model, and verifying the multi-emotion speech synthesis system; Step S7: performing speech synthesis by using the verified multi-emotion speech synthesis system.
2. The emotional speech synthesis method based on emotional analysis as claimed in claim 1, characterized in that: In step S1, the text sentiment analysis model consists of an input embedding layer, several Transformer encoders, and a sentiment classification output layer; The input embedding layer is used to convert each word of the sentiment text into an embedding vector: X=[E(w1),E(w2),...,E(w n )]+P; Where X represents the embedding vector; w n Represents the nth word in the sentiment text; E() represents the embedding function; P represents the position encoding of the word; The Transformer encoder consists of a self-attention network and several layers of feedforward neural networks, which are used to encode the embedding vector to obtain the feature vector h. The formula is: H l =Attention(H l-1 )+H l-1 ; Among them, H l represents the lth layer of feedforward neural network; H l-1 represents the l-1th layer of feedforward neural network; Attention() represents the self-attention network; the input of the first layer of the feedforward neural network is the embedding vector X, and the output of the last layer of the feedforward neural network is the feature vector h; The sentiment classification output layer is used to perform sentiment prediction based on the feature vector h and output the sentiment category: in, represents the predicted probability of the i-th emotion category; exp() represents the exponential function; W represents the weight matrix; b represents the bias term; C represents the total number of emotion categories; (Wh+b) i Represents the i-th eigenvector adjusted based on the weight matrix and bias term.
3. The emotional speech synthesis method based on emotional analysis as claimed in claim 1, characterized in that: In step S1, the text sentiment analysis model is trained based on the cross entropy loss function: Among them, L represents the loss value of the cross entropy loss function; y i represents the true sentiment category corresponding to the i-th word; p i represents the probability distribution of the predicted emotion category; C represents the total number of emotion categories.
4. The emotional speech synthesis method based on emotional analysis as claimed in claim 1, characterized in that: In step S2, the voice feature extraction formula is: MFCC(t)=DCT(log|S(t)|); Among them, MFCC(t) represents the speech features of the speech data at time t; DCT() represents the discrete cosine transform function; S(t) represents the Mel spectrum obtained by performing short-time Fourier transform on the speech data at time t.
5. The emotional speech synthesis method based on emotional analysis as claimed in claim 1, characterized in that: In step S4, the Char2Wav module is used to encode the emotional text through a bidirectional recurrent neural network to obtain a hidden state h t , for the hidden state h t Decode to get acoustic features Then use SampleRNN to transform the acoustic features Convert to raw audio waveform.
6. The emotional speech synthesis method based on emotional analysis as claimed in claim 5, characterized in that: The hidden state h t The calculation formula is: Among them, h t represents the hidden state at time t; represents the forward hidden state at time t; represents the backward hidden state at time t; [·; ·] represents the concatenation operation; RNN forward () represents the forward recurrent neural network in the bidirectional recurrent neural network; RNN backward () represents the backward recurrent neural network in the bidirectional recurrent neural network; x t A character vector representing the sentiment text input at time t; represents the forward hidden state at time t-1; Represents the backward hidden state at time t+1; The acoustic characteristics The calculation formula is: in, represents the acoustic characteristics at time t; Represents the acoustic features at time t-1; RNN decoder () represents the RNN decoder.
7. The emotional speech synthesis method based on emotional analysis as claimed in claim 1, characterized in that: In step S4, the Tacotron module is used to extract text features h from the sentiment text through the CBHG network. t' , based on the attention mechanism for the text feature h t' Align to get the context vector c t' , based on the context vector c t' Compute Mel-spectrogram subgraph Based on each Mel spectrum subgraph Calculate the Mel spectrum The Mel-spectrogram is transformed into Convert to raw audio waveform.
8. The emotional speech synthesis method based on emotional analysis as claimed in claim 1, characterized in that: In step S5, the emotion control model is used to input the emotion label E z Adjust the hidden state h of the emotional speech synthesis model z , get the optimized hidden state h' z , based on the optimized hidden state h' z The original audio waveform output by the emotional speech synthesis model is decoded to obtain synthesized speech.
9. The emotional speech synthesis method based on emotional analysis as claimed in claim 8, characterized in that: The optimized hidden state h' z The calculation formula is: h' z =h z *W E (HAVE BEEN z ); Among them, W E (E z ) represents the sentiment label E z The associated weight matrix.
10. The emotional speech synthesis method based on emotional analysis according to claim 1, characterized in that: In the step S6, the verification of the multi-emotion speech synthesis system specifically includes: verifying the multi-emotion speech synthesis system based on MOS score and emotion recognition accuracy.