Voice conversion method, device, equipment and storage medium based on sentiment analysis
By combining sentiment analysis and acoustic structure, the problem of lack of natural emotion and personalization in TTS technology is solved, and more natural, fluent and emotionally rich speech generation is achieved to adapt to the needs of different scenarios and speakers.
Patent Information
- Application Number
- CN202510022497.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-01-07
AI Technical Summary
Existing TTS technology lacks natural emotional expression when generating speech, cannot adapt to real-time input text and changing emotional expressions, and cannot generate personalized speech based on the speaker's characteristics, affecting its application in medicine and finance.
By obtaining a sample data set for preprocessing, the sentiment analysis structure and acoustic structure are used to perform sentiment classification and acoustic feature extraction on the text data, combined with the speech generator to synthesize speech, and the preset loss function is used to iteratively train the emotional speech synthesis model to generate personalized and emotionally rich speech.
It improves the naturalness and fluency of speech, enhances the ability to express emotions, improves adaptability to real-time input and the diversity of personalized speech, and enhances user experience and satisfaction.
Smart Images

Figure CN119864014B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of artificial intelligence, financial technology, and digital medical technology, and in particular to a speech conversion method, apparatus, device, and storage medium based on sentiment analysis. Background Art
[0002] In the current medical and financial fields, text-to-speech (TTS) technology, as an important means of information transmission, has been widely used in scenarios such as remote medical consultation, intelligent customer service, and financial information broadcasting. Existing TTS products and technologies have made some progress in converting text to speech and can provide basic speech output capabilities. However, current TTS products and technologies still face many shortcomings and defects in practical applications, limiting their in-depth application in fields such as medicine and finance.
[0003] First, most existing TTS systems lack natural emotional expression when generating speech, resulting in monotonous and robotic output. In medicine, emotional connection between doctors and patients is crucial. Speech lacking emotional expression may fail to accurately convey the doctor's care and instructions, compromising treatment effectiveness. Similarly, in finance, personalized, emotionally charged voice services can boost customer satisfaction and loyalty, but the lack of emotion in existing TTS technology makes it difficult to meet this need.
[0004] Secondly, to improve processing speed and efficiency, some TTS systems simplify data during processing, often resulting in a loss of emotional and semantic information in the speech. In fields such as medicine and finance, the accuracy and completeness of information are crucial; any loss of information can lead to misunderstandings or risks. Therefore, the information loss caused by data simplification needs to be addressed urgently.
[0005] Furthermore, existing TTS technologies often rely on pre-recorded voice clips and are less adaptable to real-time text input and changing emotional expressions. In scenarios with high real-time requirements, such as medicine and finance, this lack of adaptability can lead to information lag and distortion, impacting user experience and decision-making effectiveness.
[0006] Finally, many TTS systems fail to generate personalized voices based on the speaker's characteristics and emotional needs, resulting in a lack of diversity. In medicine, different roles—doctors, patients, and family members—may require different voice styles to meet their specific communication needs. In finance, personalized voice services are also a key means of enhancing brand image and customer retention. However, the lack of personalization in existing TTS technology limits its potential for application in these areas. Summary of the Invention
[0007] The purpose of the embodiments of the present application is to propose a speech conversion method, device, equipment and storage medium based on sentiment analysis to solve the technical problems that the speech generated by existing TTS technology lacks natural emotional expression, has poor adaptability to real-time input text and changing emotional expressions, and cannot generate personalized speech based on the speaker's characteristics and emotional needs.
[0008] In order to solve the above technical problems, the present application provides a speech conversion method based on sentiment analysis, which adopts the following technical solutions:
[0009] Acquire a sample data set, wherein the sample data set includes a large number of paired data of text data samples and corresponding acoustic feature samples;
[0010] Preprocessing the text data sample to obtain a standard text data set;
[0011] Inputting the standard text dataset into a pre-built conversion-synthesis model, wherein the conversion-synthesis model includes a sentiment analysis structure, an acoustic structure, and a speech generator;
[0012] Performing sentiment classification and sentiment embedding on the standard text dataset using the sentiment analysis structure to obtain a text sentiment embedding vector;
[0013] Extracting acoustic features by fusing the standard text dataset and the text sentiment embedding vector through the acoustic structure to obtain predicted acoustic features;
[0014] Inputting the predicted acoustic features and the acoustic feature samples into the speech generator to synthesize speech waveforms, respectively, to obtain predicted speech and real speech;
[0015] Calculating a loss based on the predicted acoustic features, the acoustic feature samples, the predicted speech, and the real speech according to a preset loss function;
[0016] Adjusting the model parameters of the conversion synthesis model based on the loss, continuing the iterative training until an iteration stop condition is met, and obtaining a final emotional speech synthesis model;
[0017] The text to be converted is obtained, and the text to be converted is input into the emotional speech synthesis model to synthesize the target speech.
[0018] In order to solve the above technical problems, the embodiment of the present application further provides a speech conversion device based on sentiment analysis, which adopts the following technical solution:
[0019] An acquisition module is used to acquire a sample data set, wherein the sample data set includes a large number of paired data of text data samples and corresponding acoustic feature samples;
[0020] A preprocessing module, configured to preprocess the text data sample to obtain a standard text data set;
[0021] An input module, configured to input the standard text dataset into a pre-built conversion-synthesis model, wherein the conversion-synthesis model includes a sentiment analysis structure, an acoustic structure, and a speech generator;
[0022] A sentiment analysis module is used to perform sentiment classification and sentiment embedding on the standard text dataset using the sentiment analysis structure to obtain a text sentiment embedding vector;
[0023] An acoustic feature extraction module, configured to extract acoustic features by fusing the standard text dataset and the text sentiment embedding vector with the acoustic structure to obtain predicted acoustic features;
[0024] A speech prediction module, configured to input the predicted acoustic features and the acoustic feature samples into the speech generator respectively to synthesize speech waveforms, thereby obtaining predicted speech and real speech;
[0025] A loss calculation module, configured to calculate the loss based on the predicted acoustic features, the acoustic feature samples, the predicted speech, and the real speech according to a preset loss function;
[0026] An iterative module, configured to adjust the model parameters of the conversion synthesis model based on the loss, continue iterative training until an iteration stop condition is met, and obtain a final emotional speech synthesis model;
[0027] The speech generation module is used to obtain the text to be converted, input the text to be converted into the emotional speech synthesis model, and synthesize the target speech.
[0028] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:
[0029] The computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the speech conversion method based on sentiment analysis as described above are implemented.
[0030] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:
[0031] The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the above-mentioned speech conversion method based on sentiment analysis.
[0032] Compared with the prior art, this application has the following beneficial effects:
[0033] The present application provides a speech conversion method based on sentiment analysis, which can improve the text quality by preprocessing text samples in a sample data set, help to more accurately understand the text content, and accurately identify the emotional tendency in the text; through the sentiment analysis structure, sentiment classification and sentiment embedding are performed on the standard text data set, thereby realizing automatic recognition and expression of text emotions and enhancing the ability of emotional expression; through the acoustic structure, the standard text data set and the text sentiment embedding vector are integrated to extract acoustic features, and the acoustic features are input into the speech generator for speech waveform synthesis, which can give the speech more emotional color, make the synthesized speech more natural, reduce monotony and mechanical feeling, improve the accuracy of speech recognition, enhance the coherence and fluency of speech, and improve user experience and satisfaction. By converting text to speech through the emotional speech synthesis model, it can improve the adaptability to real-time input, and can also generate personalized speech according to the speaker's characteristics and emotional needs, thereby improving the diversity of speech synthesis. In addition, the emotional speech synthesis model can handle multiple speaker characteristics and multiple emotional scenarios, providing possibilities for a wider range of applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0035] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;
[0036] Figure 2 is a flowchart of an embodiment of a speech conversion method based on sentiment analysis according to the present application;
[0037] Figure 3 1 is a schematic structural diagram of an embodiment of a speech conversion device based on sentiment analysis according to the present application;
[0038] Figure 4 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION
[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.
[0040] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0041] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0042] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0043] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0044] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.
[0045] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .
[0046] It should be noted that the speech conversion method based on sentiment analysis provided in the embodiment of the present application is generally executed by a server / terminal device, and accordingly, the speech conversion device based on sentiment analysis is generally set in the server / terminal device.
[0047] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0048] Continue to refer Figure 2 , shows a flow chart of an embodiment of a speech conversion method based on sentiment analysis according to the present application, comprising the following steps:
[0049] Step S201 : obtaining a sample data set, wherein the sample data set includes a large number of paired data of text data samples and corresponding acoustic feature samples.
[0050] Collect a speech dataset. This can be collected using professional recording studio equipment or obtained from public speech databases such as LJSpeech and Emo-DB. Preprocess the raw speech dataset, including noise removal and voice quality adjustment, to obtain a preprocessed speech dataset. Obtain the corresponding text dataset, extract the acoustic features of the speech data in the speech dataset, and annotate the text dataset with the extracted acoustic features. This matches the text data with the corresponding acoustic features, establishing a mapping between text and speech and providing supervised learning data for subsequent model training.
[0051] In this embodiment, the electronic device (e.g. Figure 1The server / terminal device shown in the figure can receive the acquired sample data set via a wired connection or a wireless connection. It should be noted that the wireless connection method may include but is not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future.
[0052] The text dataset with annotated acoustic features is used as a sample dataset. The sample dataset contains a large amount of paired data, and each paired data includes a text data feature and its corresponding acoustic feature sample.
[0053] Among them, acoustic features include Mel spectrogram, duration information, pitch (F0) and energy. The Mel spectrogram is a visual representation of the frequency content of the speech signal; duration information represents the duration of each phoneme or syllable; pitch (F0) represents the fundamental frequency of the speech signal, which is related to human pitch perception; energy represents the intensity or loudness of the speech signal.
[0054] Step S202: pre-process the text data sample to obtain a standard text data set.
[0055] In this embodiment, the text data samples are preprocessed to standardize the text data samples into standard text data. Specifically, the preprocessing includes cleaning, normalization, word segmentation, etc., wherein cleaning includes removing irrelevant characters and special symbols, removing stop words, spelling checking, grammar correction, removing duplicate content, ASCII conversion, lowercase conversion, abbreviation expansion, number expansion and space compression, for example, the abbreviation "Dr." is expanded to "doctor" and the number "2" is converted to "two". Normalization is to scale or standardize the text data to make it conform to a specific data distribution or range. For example, the data value is scaled to the range of [0,1] using a normalization method.
[0056] Step S203: input the standard text dataset into a pre-built conversion-synthesis model, where the conversion-synthesis model includes a sentiment analysis structure, an acoustic structure, and a speech generator.
[0057] In this embodiment, the pre-built conversion synthesis model includes a sentiment analysis structure, an acoustic structure and a speech generator. The sentiment analysis structure is used to perform sentiment analysis on the text, and the acoustic structure is used to convert the graphics elements (written characters) in the text into phonemes (sound units). The corresponding acoustic features of the text data sample are predicted based on the converted phonemes combined with the sentiment analysis results of the sentiment analysis structure. The speech generator converts the acoustic features into audible speech waveforms, and the synthesized speech also has natural emotional expression.
[0058] Step S204: sentiment classification and sentiment embedding are performed on the standard text dataset using a sentiment analysis structure to obtain a text sentiment embedding vector.
[0059] The sentiment analysis structure performs sentiment analysis on standard text data in the standard text dataset. Sentiment analysis includes sentiment classification and sentiment embedding generation.
[0060] In this embodiment, the sentiment analysis structure includes a pre-trained sentiment classification model and an embedding generation layer. The standard text dataset is input into the sentiment classification model to perform text sentiment classification to obtain text sentiment categories; the text sentiment categories are sentimentally embedded through the embedding generation layer to generate text sentiment embedding vectors.
[0061] The sentiment classification model is a pre-trained model that has learned a large number of language patterns and emotional expressions. The sentiment classification model can perform sentiment classification on standard text data in a standard text dataset. The embedding generation layer generates sentiment embeddings based on the results of sentiment classification.
[0062] In some optional implementations of this embodiment, the sentiment classification model includes an input layer, an embedding layer, a spatial expansion layer, an encoding layer, and a fully connected layer. The step of inputting the standard text dataset into the sentiment classification model to perform text sentiment classification and obtain the text sentiment category includes:
[0063] Preprocess the standard text dataset through the input layer to obtain input text data;
[0064] The input text data is converted into a vector through the embedding layer to obtain the text embedding vector;
[0065] The text embedding vector is input into the spatial expansion layer, and the text expansion vector is obtained by concatenating the text embedding vector and the introduced hint vector;
[0066] The emotional features in the text expansion vector are extracted through the self-attention mechanism of the encoding layer to obtain the emotional semantic features;
[0067] The emotional semantic features are input into the fully connected layer for classification to obtain the text sentiment category.
[0068] The sentiment classification model uses an improved RoBERTa model. The input layer preprocesses a standard text dataset to obtain input text data in a format suitable for model processing. This input text data is then fed into the embedding layer for vector conversion. This conversion involves character embedding (token embedding), segmentation embedding, and position embedding.
[0069] Character embeddings represent the primary semantic information of a character; segment embeddings represent sentence structure, distinguishing between sentences or paragraphs; and position embeddings provide information about the position of a word within a sentence, adding temporal information to the attention mechanism. The text embedding vector is obtained by adding the vectors corresponding to the character embedding, segment embedding, and position embedding.
[0070] The spatial expansion layer introduces a learnable prompt vector pool to spatially expand the text embedding vector to enrich the feature representation. Specifically, the text embedding vector is input into the spatial expansion layer, which then obtains a pre-trained prompt vector pool containing S vectors, where S is a positive integer. K pre-trained prompt vectors are obtained from the prompt vector pool, where k is a positive integer less than or equal to S. The text embedding vector is concatenated with the k prompt vectors to obtain a text expansion vector. By introducing the prompt vector, additional information or constraints are introduced to the original feature space, thereby expanding the feature space. This allows the model to capture more features related to sentiment information, thereby improving sentiment classification accuracy.
[0071] The encoding layer is a plurality of stacked connected Transformer modules, including a plurality of stacked connected encoders and a plurality of stacked connected decoders. The semantic and contextual features of the text extension vector are extracted through the encoder's multi-head attention mechanism to obtain the text encoding features. The text encoding features are input into the decoder. The global dependency between the text encoding features is captured through the multi-head attention mechanism and the cross attention mechanism to obtain the emotional semantic features.
[0072] The emotional semantic features are input into the fully connected layer, and the classification probability distribution of the emotional category is output through the Softmax activation function. The text emotional category is obtained according to the classification probability distribution.
[0073] Text sentiment classification through sentiment classification models can efficiently and accurately identify the emotional expression of the text and improve the accuracy of sentiment classification.
[0074] The embedding generation layer can use embedding layers from deep learning models, such as the BERT embedding layer and the Transformer embedding layer, or word embedding models such as Word2Vec, GloVe, and FastText. The embedding generation layer performs sentiment embedding based on the text's sentiment category and generates the corresponding text sentiment embedding vector.
[0075] By performing sentiment classification and sentiment embedding on standard text datasets, the accuracy of text sentiment recognition can be improved, which helps voice convey the emotional color in the text more accurately. Sentiment embedding can generate personalized voice based on the emotional needs of the input text, which helps to flexibly adjust according to the emotional needs of different scenarios and improve adaptability to the input text and emotional changes.
[0076] Step S205 , extracting acoustic features by fusing the standard text dataset and the text sentiment embedding vector through acoustic structure to obtain predicted acoustic features.
[0077] The core goal of acoustic structure training is to build a model that can convert text information into acoustic features. The acoustic structure includes a phoneme conversion layer, an acoustic encoder, a variance adapter, and an acoustic decoder. A standard text dataset and a text sentiment embedding vector are input into the acoustic structure. These are processed sequentially through the phoneme conversion layer, the acoustic encoder, the variance adapter, and the acoustic decoder to produce predicted acoustic features that incorporate sentiment information.
[0078] In this embodiment, the step of extracting acoustic features by fusing a standard text dataset and a text sentiment embedding vector using an acoustic structure to obtain predicted acoustic features includes:
[0079] The standard text dataset and text sentiment embedding vector are mapped to phonemes through the phoneme conversion layer to obtain a phoneme embedding sequence;
[0080] The phoneme embedding sequence is subjected to feature extraction through an acoustic encoder to obtain a phoneme feature sequence;
[0081] Input the phoneme feature sequence into the variance adapter for acoustic feature prediction to obtain the acoustic feature hidden sequence;
[0082] The acoustic feature hidden sequence is decoded by the acoustic decoder to obtain the predicted acoustic features.
[0083] In some optional implementations of this embodiment, the phoneme conversion layer adopts a grapheme-to-phoneme conversion (G2P) model, and the G2P model includes an input layer, a feature extraction layer, a sequence modeling layer, and an output layer. Specifically, the input text data set is preprocessed by the input layer, including operations such as word segmentation and part-of-speech tagging, to form a text sequence suitable for model processing; the input text sequence and text sentiment embedding vector are feature extracted by the feature extraction layer to capture the phoneme features of the text sequence and the sentiment features of the text sentiment embedding vector; the sentiment features are embedded into the phoneme features and input into the sequence modeling layer for sequence modeling to obtain a phoneme representation vector that integrates each phoneme feature and sentiment feature; the phoneme representation vector is decoded by the output layer using a beam search-based decoding algorithm to convert it into a phoneme symbol sequence to obtain a complete phoneme embedding sequence.
[0084] In some optional implementations, the acoustic encoder includes multiple layers of self-attention layers and 1D convolution layers; position information is added to the phoneme embedding sequence through position encoding to obtain a phoneme embedding sequence containing position information; the phoneme embedding sequence containing position information is input into the multiple layers of self-attention layers, and the phoneme embedding sequence is processed using the multiple layers of self-attention mechanism to capture the dependency between phonemes to obtain a phoneme semantic sequence; and the phoneme semantic sequence is feature extracted through the 1D convolution layer to obtain a phoneme feature sequence to enhance the model's understanding of the phoneme sequence.
[0085] In some optional implementations, the variance adapter includes a duration predictor, a pitch predictor, and an energy predictor, which are respectively used to predict the duration, pitch, and energy of a phoneme. Specifically, the phoneme feature sequence is input into the duration predictor, and the logarithmic duration value of each phoneme is predicted to obtain a predicted duration value; the pitch predictor predicts the logarithmic pitch value of each phoneme in the phoneme feature sequence to obtain a predicted pitch value; the energy predictor predicts the energy level of each phoneme in the phoneme feature sequence to obtain predicted energy information; the predicted duration value, predicted pitch value, and predicted energy information are embedded and converted to obtain an acoustic feature vector; and the acoustic feature vector and the phoneme feature sequence are concatenated to obtain an acoustic feature hidden sequence.
[0086] In this embodiment, the duration predictor uses the first fully connected layer to predict the logarithmic duration value of each phoneme based on the phoneme feature sequence output by the acoustic encoder, and converts the logarithmic duration value into the actual predicted duration value through an exponential function. The pitch predictor uses the second fully connected layer to predict the logarithmic pitch value of each phoneme based on the phoneme feature sequence, and converts the logarithmic pitch value into the actual predicted pitch value (Hz) through an exponential function. The energy predictor is used to predict the energy level of each phoneme, usually scaled by a sigmoid function to ensure that the energy value is within a reasonable range. The predicted duration value, predicted pitch value and predicted energy information are feature integrated and embedded in a high-dimensional space to obtain an acoustic feature vector; the acoustic feature vector is added to the phoneme feature sequence to form a new adaptive hidden sequence containing acoustic feature information, namely, an acoustic feature hidden sequence.
[0087] By using the variance adapter to predict duration, pitch, and energy, the naturalness, intelligibility, expressiveness, and clarity of speech can be significantly improved, making the synthesized speech closer to natural human pronunciation, thereby meeting the needs of various application scenarios and having better adaptability to multiple scenarios.
[0088] In some optional implementations, the acoustic decoder includes a self-attention layer, a 1D convolution layer, and an upsampling and residual connection layer, which uses the self-attention layer and the 1D convolution layer to convert the adaptive acoustic feature hidden sequence into an acoustic feature spectrum sequence, and gradually increases the size of the feature map through upsampling and residual connection until the target speech length is reached, and outputs the predicted acoustic features.
[0089] In a specific example, the acoustic decoder is a Mel spectrogram decoder. Mel spectrogram is a type of acoustic feature and is commonly used in speech synthesis. Therefore, the acoustic feature hidden sequence is directly decoded into a Mel spectrogram.
[0090] By fusing text phonemes and text emotions into acoustic structure for acoustic feature prediction, it is possible to learn the correspondence between input text and acoustic features, and capture the corresponding acoustic features and conversion rules of the text, which helps to improve the naturalness, fluency, intelligibility and accuracy of synthesized speech, and optimize the efficiency and quality of speech synthesis.
[0091] In some optional implementations, during the training of the transductive synthesis model, the acoustic structure is first trained. The loss function during the training process is used to measure the difference between the acoustic features predicted by the model and the actual acoustic features, and the model parameters are optimized to minimize this difference. The loss function may include:
[0092] 1) Mel-spectrogram loss: measures the difference between the predicted Mel-spectrogram and the true Mel-spectrogram, usually calculated using L1 loss or L2 loss.
[0093] 2) Duration loss: measures the difference between the predicted phoneme duration and the true phoneme duration.
[0094] 3) Pitch loss: measures the difference between the predicted pitch and the true pitch.
[0095] 4) Energy loss: measures the difference between the predicted energy level and the actual energy level.
[0096] 5) Total loss: The weighted sum of all the above losses is obtained to obtain the total loss, which is used for end-to-end training of the model.
[0097] A loss function is used to calculate the loss between the predicted acoustic features and the true acoustic features. Based on this loss, the model parameters of the acoustic structure are updated using the backpropagation algorithm. An optimization algorithm (such as Adam or SGD) is used to adjust the model parameters to minimize the loss function. After adjusting the model parameters, the acoustic structure is iteratively trained until satisfactory performance on the validation set is achieved or a predetermined number of iterations is reached. To prevent overfitting, dropout or L2 regularization can be introduced during training.
[0098] Step S206: input the predicted acoustic features and the acoustic feature samples into a speech generator to synthesize speech waveforms, thereby obtaining predicted speech and real speech.
[0099] In the speech synthesis stage, the goal is to convert acoustic features into audible speech waveforms while ensuring that the synthesized speech has natural emotional expression.
[0100] In this embodiment, the speech generator includes multiple convolutional layers and residual blocks. The predicted acoustic features and acoustic feature samples are respectively input into the speech generator. The predicted acoustic features and acoustic feature samples are respectively convolved through multiple convolutional layers to obtain corresponding predicted spectral features and sample spectral features. The predicted spectral features and sample spectral features are converted into waveforms through the residual block to obtain corresponding predicted speech and real speech.
[0101] In a specific example, when the predicted acoustic feature output by the acoustic structure is a predicted Mel spectrogram, the predicted Mel spectrogram is converted into a high-resolution audio waveform through multiple convolutional layers and residual blocks, that is, predicted speech is generated.
[0102] By using a speech generator to synthesize speech waveforms based on acoustic features, the naturalness and fluency of speech synthesis can be improved, multi-language and multi-style speech synthesis can be supported, the diversity of speech synthesis can be increased, and personalized customization of speech synthesis can be achieved, thereby enhancing the flexibility and scalability of speech synthesis.
[0103] Step S207 , calculating the loss based on the predicted acoustic features, the acoustic feature samples, the predicted speech, and the real speech according to a preset loss function.
[0104] During model training, the speech synthesis architecture consists of a speech generator and a speech discriminator. The speech generator attempts to generate realistic audio waveforms, while the speech discriminator attempts to distinguish between real and generated waveforms. The speech discriminator evaluates the difference between the generated audio waveform and the real waveform—that is, the difference between the predicted speech and the real speech—to guide the training of the speech generator.
[0105] In this embodiment, a speech discriminator corresponding to the speech generator is obtained, and the predicted speech and real speech are input into the speech discriminator to calculate the adversarial loss. The spectral loss between the predicted acoustic features and the acoustic feature samples is calculated. The L1 loss is calculated based on the predicted and real speech. The adversarial loss, spectral loss, and L1 loss are weighted and summed to obtain the final total loss. The spectral loss is the Mel-spectrogram loss.
[0106] Specifically, the differential privacy sensitivity is calculated based on the predicted speech and the real speech; based on the differential privacy sensitivity, the privacy loss value, namely the adversarial loss, is calculated using the Markov formula.
[0107] By combining multiple loss functions to calculate losses, we can enhance the robustness of the model, maintain the accuracy of speech features, optimize the training process, improve model performance, balance the quality and diversity of generated speech, and flexibly adapt to different application scenarios and needs.
[0108] Step S208: Adjust the model parameters of the conversion synthesis model based on the loss, continue iterative training until the iteration stop condition is met, and obtain the final emotional speech synthesis model.
[0109] Specifically, the Adam or SGD optimizer is used to adjust the model parameters of the conversion synthesis model according to the loss, and the adjusted model is iteratively trained until the iteration stop condition is met, that is, the number of iterations reaches the preset number or the loss does not change significantly, and the final emotional speech synthesis model is output.
[0110] Step S209: obtaining the text to be converted, inputting the text to be converted into an emotional speech synthesis model, and synthesizing the target speech.
[0111] In this embodiment, the acquired text to be converted is fed into an emotional speech synthesis model. By combining sentiment analysis and speech synthesis technology, a personalized and emotionally rich target speech is generated. The emotional speech synthesis model generates emotionally charged speech based on the context of the text, automatically identifying and expressing the text's emotions, and improving the naturalness and expressiveness of the speech.
[0112] It should be emphasized that in order to further ensure the privacy and security of the text to be converted, the above text to be converted can also be stored in a node of a blockchain.
[0113] The blockchain referred to in this application is a new application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of this information (to prevent counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product service layer, and the application service layer.
[0114] This application can improve the text quality by preprocessing the text samples in the sample data set, help to more accurately understand the text content, and accurately identify the emotional tendencies in the text; through the sentiment analysis structure, sentiment classification and sentiment embedding are performed on the standard text data set, thereby realizing the automatic recognition and expression of text emotions and enhancing the emotional expression ability; through the acoustic structure, the standard text data set and the text emotion embedding vector are integrated to extract acoustic features, and the acoustic features are input into the speech generator for speech waveform synthesis, which can give the speech more emotional color, make the synthesized speech more natural, reduce monotony and mechanical feeling, improve the accuracy of speech recognition, enhance the coherence and fluency of speech, and improve user experience and satisfaction. By converting text to speech through the emotional speech synthesis model, the adaptability to real-time input can be improved, and personalized speech can be generated according to the speaker's characteristics and emotional needs, thereby improving the diversity of speech synthesis. In addition, the emotional speech synthesis model can handle multiple speaker characteristics and multiple emotional scenarios, providing possibilities for a wider range of applications.
[0115] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0116] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0117] The sentiment analysis-based speech conversion method of this application gives the speech synthesis system a deeper level of naturalness and emotional expression capabilities. For Ping An Group, the application of the sentiment analysis-based speech conversion method will greatly enhance the interactivity and affinity of its financial services. In the insurance business, this method can provide customers with a more personalized and empathetic service experience through emotionally rich voice interaction. Especially in scenarios such as claims consultation and customer care, it can more effectively convey warmth and understanding. In the banking business, the application of this method makes automated voice services no longer cold and impersonal machine language, but can adjust the tone and expression according to the customer's emotions and needs, enhancing customer trust and satisfaction. Furthermore, in the asset management and investment consulting fields, this method can provide a more vivid and persuasive way to convey information, helping customers better understand investment products and market dynamics. Overall, the application of this application will make Ping An Group's services more humane, enhance customer loyalty, and improve its brand image, thereby gaining an advantage in the highly competitive financial market.
[0118] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0119] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0120] Further references Figure 3 , as a response to the above Figure 2 The present application provides an embodiment of a speech conversion device based on sentiment analysis, which is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0121] like Figure 3 As shown, the speech conversion device 300 based on sentiment analysis in this embodiment includes: an acquisition module 301, a preprocessing module 302, an input module 303, a sentiment analysis module 304, an acoustic feature extraction module 305, a speech prediction module 306, a loss calculation module 307, an iteration module 308, and a speech generation module 309. Among them:
[0122] The acquisition module 301 is used for acquiring a sample data set, wherein the sample data set includes a large number of paired data of text data samples and corresponding acoustic feature samples;
[0123] The preprocessing module 302 is used to preprocess the text data sample to obtain a standard text data set;
[0124] The input module 303 is used to input the standard text dataset into a pre-built conversion synthesis model, wherein the conversion synthesis model includes a sentiment analysis structure, an acoustic structure and a speech generator;
[0125] The sentiment analysis module 304 is used to perform sentiment classification and sentiment embedding on the standard text dataset using the sentiment analysis structure to obtain a text sentiment embedding vector;
[0126] The acoustic feature extraction module 305 is used to extract acoustic features by fusing the standard text dataset and the text sentiment embedding vector with the acoustic structure to obtain predicted acoustic features;
[0127] The speech prediction module 306 is used to input the predicted acoustic features and the acoustic feature samples into the speech generator respectively to synthesize speech waveforms to obtain predicted speech and real speech;
[0128] The loss calculation module 307 is used to calculate the loss according to the predicted acoustic features, the acoustic feature samples, the predicted speech and the real speech according to a preset loss function;
[0129] The iterative module 308 is used to adjust the model parameters of the conversion synthesis model based on the loss, continue iterative training until the iteration stop condition is met, and obtain the final emotional speech synthesis model;
[0130] The speech generation module 309 is used to obtain the text to be converted, input the text to be converted into the emotional speech synthesis model, and synthesize the target speech.
[0131] It should be emphasized that in order to further ensure the privacy and security of the text to be converted, the above text to be converted can also be stored in a node of a blockchain.
[0132] The above-mentioned speech conversion device 300 based on sentiment analysis can improve the text quality by preprocessing the text samples in the sample data set, help to more accurately understand the text content, and accurately identify the emotional tendency in the text; through the sentiment analysis structure, sentiment classification and sentiment embedding are performed on the standard text data set, thereby realizing automatic recognition and expression of text emotions and enhancing the ability of emotional expression; through the acoustic structure, the standard text data set and the text emotion embedding vector are integrated to extract acoustic features, and the acoustic features are input into the speech generator for speech waveform synthesis, which can give the speech more emotional color, make the synthesized speech more natural, reduce monotony and mechanical feeling, improve the accuracy of speech recognition, enhance the coherence and fluency of speech, and improve user experience and satisfaction. By converting text to speech through the emotional speech synthesis model, it can improve adaptability to real-time input, and can also generate personalized speech according to the characteristics and emotional needs of the speaker, thereby improving the diversity of speech synthesis.
[0133] In some optional implementations, the sentiment analysis structure includes a pre-trained sentiment classification model and an embedding generation layer, and the sentiment analysis module 304 includes:
[0134] A classification submodule is used to input the standard text data set into the sentiment classification model to perform text sentiment classification and obtain text sentiment categories;
[0135] The emotion embedding submodule is used to perform emotion embedding on the text emotion category through the embedding generation layer to generate a text emotion embedding vector.
[0136] By performing sentiment classification and sentiment embedding on standard text datasets, the accuracy of text sentiment recognition can be improved, which helps speech convey the emotional color in the text more accurately and improves the adaptability to input text and emotional changes.
[0137] In some optional implementations of this embodiment, the sentiment classification model includes an input layer, an embedding layer, a spatial expansion layer, and a 、 The encoding layer and the fully connected layer, the classification submodule includes:
[0138] An input unit, configured to preprocess the standard text data set through the input layer to obtain input text data;
[0139] An embedding unit, configured to perform vector conversion on the input text data through the embedding layer to obtain a text embedding vector;
[0140] an expansion unit, configured to input the text embedding vector into the spatial expansion layer, and obtain a text expansion vector by concatenating the text embedding vector and the introduced hint vector;
[0141] An encoding unit, configured to extract sentiment features from the text extension vector through a self-attention mechanism of the encoding layer to obtain sentiment semantic features;
[0142] The classification unit is used to input the emotional semantic features into the fully connected layer for classification to obtain the text emotion category.
[0143] Text sentiment classification through sentiment classification models can efficiently and accurately identify the emotional expression of the text and improve the accuracy of sentiment classification.
[0144] In some optional implementations, the acoustic structure includes a phoneme conversion layer, an acoustic encoder, a variance adapter, and an acoustic decoder, and the acoustic feature extraction module 305 includes:
[0145] A phoneme conversion submodule, configured to map the standard text dataset and the text sentiment embedding vector into phonemes through the phoneme conversion layer to obtain a phoneme embedding sequence;
[0146] an acoustic coding submodule, configured to extract features from the phoneme embedding sequence through the acoustic encoder to obtain a phoneme feature sequence;
[0147] An adapter module, configured to input the phoneme feature sequence into the variance adapter to perform acoustic feature prediction to obtain an acoustic feature hidden sequence;
[0148] The acoustic decoding submodule is used to decode the acoustic feature hidden sequence through the acoustic decoder to obtain the predicted acoustic feature.
[0149] By fusing text phonemes and text emotions into acoustic structure for acoustic feature prediction, it is possible to learn the correspondence between input text and acoustic features, and capture the corresponding acoustic features and conversion rules of the text, which helps to improve the naturalness, fluency, intelligibility and accuracy of synthesized speech, and optimize the efficiency and quality of speech synthesis.
[0150] In some optional implementations of this embodiment, the variance adapter includes a duration predictor, a pitch predictor, and an energy predictor, and the adapter submodule is further configured to:
[0151] Inputting the phoneme feature sequence into the duration predictor, predicting the logarithmic duration value of each phoneme to obtain a predicted duration value;
[0152] Predicting the pitch logarithm of each phoneme in the phoneme feature sequence by the pitch predictor to obtain a predicted pitch value;
[0153] Predicting the energy level of each phoneme in the phoneme feature sequence by the energy predictor to obtain predicted energy information;
[0154] Embedding and converting the predicted duration value, the predicted pitch value, and the predicted energy information to obtain an acoustic feature vector;
[0155] The acoustic feature vector and the phoneme feature sequence are concatenated to obtain an acoustic feature hidden sequence.
[0156] By using the variance adapter to predict duration, pitch, and energy, the naturalness, intelligibility, expressiveness, and clarity of speech can be significantly improved, making the synthesized speech closer to natural human pronunciation, thereby meeting the needs of various application scenarios and having better adaptability to multiple scenarios.
[0157] In some optional implementations, the speech generator includes multiple convolutional layers and residual blocks, and the speech prediction module 306 is further configured to:
[0158] Performing convolution operations on the predicted acoustic features and the acoustic feature samples respectively through the multiple convolution layers to obtain predicted spectral features and sample spectral features respectively;
[0159] The predicted spectral features and the sample spectral features are waveform-converted by the residual block to obtain corresponding predicted speech and real speech.
[0160] By using a speech generator to synthesize speech waveforms based on acoustic features, the naturalness and fluency of speech synthesis can be improved, multi-language and multi-style speech synthesis can be supported, the diversity of speech synthesis can be increased, and personalized customization of speech synthesis can be achieved, thereby enhancing the flexibility and scalability of speech synthesis.
[0161] In some optional implementations of this embodiment, the loss calculation module 307 is further configured to:
[0162] Obtaining a speech discriminator corresponding to the speech generator, and inputting the predicted speech and the real speech into the speech discriminator to calculate the adversarial loss;
[0163] Calculating a spectral loss between the predicted acoustic feature and the acoustic feature sample;
[0164] Calculating L1 loss based on the predicted speech and the real speech;
[0165] The adversarial loss, the spectrum loss and the L1 loss are weightedly summed to obtain the final loss.
[0166] By combining multiple loss functions to calculate losses, we can enhance the robustness of the model, maintain the accuracy of speech features, optimize the training process, and improve model performance.
[0167] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0168] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 with a memory 41, a processor 42, and a network interface 43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0169] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.
[0170] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk equipped on the computer device 4, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for the speech conversion method based on sentiment analysis. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.
[0171] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions or process data stored in the memory 41, such as computer-readable instructions for executing the speech conversion method based on sentiment analysis.
[0172] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.
[0173] By preprocessing the text samples in the sample data set, the text quality can be improved, which helps to more accurately understand the text content and accurately identify the emotional tendencies in the text; by using the sentiment analysis structure to perform sentiment classification and sentiment embedding on the standard text data set, the automatic recognition and expression of text emotions is realized, and the emotional expression ability is enhanced; by fusing the standard text data set and the text emotion embedding vector through the acoustic structure to extract acoustic features, and inputting the acoustic features into the speech generator for speech waveform synthesis, the speech can be given more emotional color, the synthesized speech is more natural, the monotony and mechanical feeling are reduced, the accuracy of speech recognition is improved, the coherence and fluency of the speech are enhanced, and the user experience and satisfaction are improved. By converting text to speech through the emotional speech synthesis model, the adaptability to real-time input can be improved, and personalized speech can be generated according to the speaker's characteristics and emotional needs, thereby improving the diversity of speech synthesis.
[0174] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the speech conversion method based on sentiment analysis as described above.
[0175] By preprocessing the text samples in the sample data set, the text quality can be improved, which helps to more accurately understand the text content and accurately identify the emotional tendencies in the text; by using the sentiment analysis structure to perform sentiment classification and sentiment embedding on the standard text data set, the automatic recognition and expression of text emotions is realized, and the emotional expression ability is enhanced; by fusing the standard text data set and the text emotion embedding vector through the acoustic structure to extract acoustic features, and inputting the acoustic features into the speech generator for speech waveform synthesis, the speech can be given more emotional color, the synthesized speech is more natural, the monotony and mechanical feeling are reduced, the accuracy of speech recognition is improved, the coherence and fluency of the speech are enhanced, and the user experience and satisfaction are improved. By converting text to speech through the emotional speech synthesis model, the adaptability to real-time input can be improved, and personalized speech can be generated according to the speaker's characteristics and emotional needs, thereby improving the diversity of speech synthesis.
[0176] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0177] It should be noted that the non-Company's software tools or components that appear in the embodiments of this application are merely examples and do not represent actual use.
[0178] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.
Claims
1. A speech conversion method based on sentiment analysis, characterized in that: The steps include: Acquire a sample data set, wherein the sample data set includes a large number of paired data of text data samples and corresponding acoustic feature samples; Preprocessing the text data sample to obtain a standard text data set; Inputting the standard text dataset into a pre-built conversion-synthesis model, wherein the conversion-synthesis model includes a sentiment analysis structure, an acoustic structure, and a speech generator; Performing sentiment classification and sentiment embedding on the standard text dataset using the sentiment analysis structure to obtain a text sentiment embedding vector; Extracting acoustic features by fusing the standard text dataset and the text sentiment embedding vector through the acoustic structure to obtain predicted acoustic features; Inputting the predicted acoustic features and the acoustic feature samples into the speech generator to synthesize speech waveforms, respectively, to obtain predicted speech and real speech; Calculating a loss based on the predicted acoustic features, the acoustic feature samples, the predicted speech, and the real speech according to a preset loss function; Adjusting the model parameters of the conversion synthesis model based on the loss, continuing the iterative training until an iteration stop condition is met, and obtaining a final emotional speech synthesis model; The text to be converted is obtained, and the text to be converted is input into the emotional speech synthesis model to synthesize the target speech.
2. The method for voice conversion based on sentiment analysis according to claim 1, characterized in that The sentiment analysis structure includes a pre-trained sentiment classification model and an embedding generation layer. The step of performing sentiment classification and sentiment embedding on the standard text dataset using the sentiment analysis structure to obtain a text sentiment embedding vector includes: Inputting the standard text data set into the sentiment classification model to perform text sentiment classification to obtain text sentiment categories; The text emotion category is emotionally embedded through the embedding generation layer to generate a text emotion embedding vector.
3. The speech conversion method based on sentiment analysis according to claim 2, characterized in that The sentiment classification model includes an input layer, an embedding layer, and a spatial expansion layer. 、 The coding layer and the fully connected layer, the step of inputting the standard text data set into the sentiment classification model to perform text sentiment classification and obtain the text sentiment category includes: Preprocessing the standard text data set through the input layer to obtain input text data; Performing vector conversion on the input text data through the embedding layer to obtain a text embedding vector; Inputting the text embedding vector into the spatial expansion layer, and obtaining a text expansion vector by concatenating the text embedding vector and the introduced hint vector; Extracting sentiment features from the text extension vector through the self-attention mechanism of the encoding layer to obtain sentiment semantic features; The emotional semantic features are input into the fully connected layer for classification to obtain the text emotion category.
4. The method for voice conversion based on sentiment analysis according to claim 1, wherein: The acoustic structure includes a phoneme conversion layer, an acoustic encoder, a variance adapter, and an acoustic decoder. The step of extracting acoustic features by fusing the standard text dataset and the text sentiment embedding vector through the acoustic structure to obtain predicted acoustic features includes: Mapping the standard text dataset and the text sentiment embedding vector into phonemes through the phoneme conversion layer to obtain a phoneme embedding sequence; Performing feature extraction on the phoneme embedding sequence by the acoustic encoder to obtain a phoneme feature sequence; Inputting the phoneme feature sequence into the variance adapter to perform acoustic feature prediction to obtain an acoustic feature hidden sequence; The acoustic feature hidden sequence is decoded by the acoustic decoder to obtain the predicted acoustic feature.
5. The method for voice conversion based on sentiment analysis according to claim 4, characterized in that: The variance adapter includes a duration predictor, a pitch predictor, and an energy predictor. The step of inputting the phoneme feature sequence into the variance adapter for acoustic feature prediction to obtain an acoustic feature hidden sequence includes: Inputting the phoneme feature sequence into the duration predictor, predicting the logarithmic duration value of each phoneme to obtain a predicted duration value; Predicting the pitch logarithm of each phoneme in the phoneme feature sequence by the pitch predictor to obtain a predicted pitch value; Predicting the energy level of each phoneme in the phoneme feature sequence by the energy predictor to obtain predicted energy information; Embedding and converting the predicted duration value, the predicted pitch value, and the predicted energy information to obtain an acoustic feature vector; The acoustic feature vector and the phoneme feature sequence are concatenated to obtain an acoustic feature hidden sequence.
6. The method for voice conversion based on sentiment analysis according to claim 1, characterized in that The speech generator includes a plurality of convolutional layers and residual blocks. The steps of inputting the predicted acoustic features and the acoustic feature samples into the speech generator for speech waveform synthesis to obtain the predicted speech and the real speech include: Performing convolution operations on the predicted acoustic features and the acoustic feature samples respectively through the multiple convolution layers to obtain predicted spectral features and sample spectral features respectively; The predicted spectral features and the sample spectral features are waveform-converted by the residual block to obtain corresponding predicted speech and real speech.
7. The method for voice conversion based on sentiment analysis according to claim 1, characterized in that: The step of calculating the loss according to the predicted acoustic features, the acoustic feature samples, the predicted speech, and the real speech according to a preset loss function includes: Obtaining a speech discriminator corresponding to the speech generator, and inputting the predicted speech and the real speech into the speech discriminator to calculate the adversarial loss; Calculating a spectral loss between the predicted acoustic feature and the acoustic feature sample; Calculating L1 loss based on the predicted speech and the real speech; The adversarial loss, the spectrum loss and the L1 loss are weightedly summed to obtain the final loss.
8. A speech conversion device based on sentiment analysis, characterized in that: include: An acquisition module is used to acquire a sample data set, wherein the sample data set includes a large number of paired data of text data samples and corresponding acoustic feature samples; A preprocessing module, configured to preprocess the text data sample to obtain a standard text data set; An input module, configured to input the standard text dataset into a pre-built conversion-synthesis model, wherein the conversion-synthesis model includes a sentiment analysis structure, an acoustic structure, and a speech generator; A sentiment analysis module is used to perform sentiment classification and sentiment embedding on the standard text dataset using the sentiment analysis structure to obtain a text sentiment embedding vector; An acoustic feature extraction module, configured to extract acoustic features by fusing the standard text dataset and the text sentiment embedding vector with the acoustic structure to obtain predicted acoustic features; A speech prediction module, configured to input the predicted acoustic features and the acoustic feature samples into the speech generator respectively to synthesize speech waveforms, thereby obtaining predicted speech and real speech; A loss calculation module, configured to calculate the loss based on the predicted acoustic features, the acoustic feature samples, the predicted speech, and the real speech according to a preset loss function; An iterative module, configured to adjust the model parameters of the conversion synthesis model based on the loss, continue iterative training until an iteration stop condition is met, and obtain a final emotional speech synthesis model; The speech generation module is used to obtain the text to be converted, input the text to be converted into the emotional speech synthesis model, and synthesize the target speech.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the speech conversion method based on emotion analysis as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the speech conversion method based on sentiment analysis according to any one of claims 1 to 7.