Method, system and equipment for converting characters into voice through multi-mode emotion driving

Through a multimodal emotion-driven deep learning model, combining text emotions and user voice characteristics, personalized and emotionally rich voice is generated, which solves the problems of speech naturalness and personalized customization in the existing technology, and is suitable for virtual assistants, navigation and smart homes and other fields.

CN120496496APending Publication Date: 2025-08-15SHENZHEN LIUFENG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510140395.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing text-to-voice technology has shortcomings in the naturalness of speech, emotional expression and personalized customization, and it is difficult to generate voices that conform to the user's voice style and emotional color.

Method used

Through multi-modal emotion analysis and user emotional state, the deep learning model is used to extract and fuse the emotional characteristics of the text and the user's personalized speech characteristics, generate joint feature vectors, perform speech synthesis, and adjust the audio waveform through multi-task learning and joint optimization methods to improve the naturalness and emotional performance of speech.

Benefits of technology

The generated voice not only conveys the text content, but can also be customized according to the emotional characteristics of the text and the user's voice style, improving the naturalness and emotional level of the voice, and is suitable for high-real-time scenarios such as voice assistants and voice interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496496A_ABST
    Figure CN120496496A_ABST
Patent Text Reader

Abstract

The invention provides a method, a system and equipment for converting characters into voice through multi-modal emotion driving, and the method comprises the following steps: S1, inputting a to-be-processed text, carrying out emotion analysis, and recognizing the emotion features of the to-be-processed text; s2, inputting voice data provided by a user, and extracting personalized voice features of the voice data; s3, fusing the emotion features and the personalized speech features to generate a joint feature vector, and embedding the joint feature vector into a deep learning model to perform speech synthesis; s4, inputting a to-be-processed text and the joint feature vector, and generating an audio waveform through a deep learning model; s5, analyzing the context of the to-be-processed text, and adjusting and optimizing the audio waveform to obtain a final voice result; multi-modal sentiment analysis is combined with a user emotional state, personalized voice customization is realized by using a deep learning model, and a context understanding module can intelligently adjust voice features according to context information, so that the naturalness and adaptability of voice are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing and speech extraction technology, and in particular to a multimodal emotion-driven text-to-speech method, system, and device. Background Art

[0002] Currently, text-to-speech technology is widely used in a variety of fields, including virtual assistants, navigation, e-book reading, and smart homes. Traditional text-to-speech technology relies primarily on two core approaches: 1. Concatenative speech synthesis, which generates synthesized speech by concatenating pre-recorded speech segments; and 2. Generative speech synthesis based on deep learning, which generates speech directly from text using neural networks.

[0003] Despite significant progress in these technologies, challenges remain in speech naturalness, emotional expression, and personalized customization. Existing technologies mostly focus on improving speech quality and generation speed, while innovation in multi-dimensional customization and personalized expression is relatively weak.

[0004] Therefore, the current relevant technologies and application research still need to be further improved. Summary of the Invention

[0005] In view of this, the present invention proposes a multimodal emotion-driven text-to-speech method, system and device, which solves the technical problems of insufficient speech naturalness, emotional expression and personalized customization in the prior art.

[0006] The technical solution of the present invention is achieved as follows:

[0007] In one aspect, the present invention provides a multimodal emotion-driven text-to-speech method, comprising the following steps:

[0008] S1, inputting a text to be processed, performing sentiment analysis, and identifying the sentiment features of the text to be processed;

[0009] S2, inputting voice data provided by the user and extracting personalized voice features of the voice data;

[0010] S3, fusing the emotional features and the personalized speech features to generate a joint feature vector, and embedding the joint feature vector into a deep learning model for speech synthesis;

[0011] S4, inputs the text to be processed and the joint feature vector, and generates an audio waveform through a deep learning model;

[0012] S5, analyzes the context of the text to be processed, adjusts and optimizes the audio waveform, and obtains the final speech result.

[0013] On the basis of this technical solution, it is further preferred that the personalized speech features are extracted through self-supervised learning and transfer learning techniques.

[0014] On the basis of this technical solution, further preferably, the personalized voice features include pitch, speaking speed or stress pattern.

[0015] On the basis of this technical solution, it is further preferred that the adjustment and optimization adopt a multi-task learning method and a joint optimization method to optimize the emotional features and personalized voice features.

[0016] On the basis of this technical solution, it is further preferred to jointly optimize personalized speech features, specifically including using Tacotron2 to generate Mel spectrum, then converting the Mel spectrum into audio waveform through WaveNet, and combining emotional and personalized features for optimization to generate high-quality speech.

[0017] In a second aspect, the present invention further provides a multimodal emotion-driven text-to-speech system, comprising:

[0018] The sentiment analysis module is used to input the text to be processed, perform sentiment analysis, and identify the sentiment characteristics of the text to be processed;

[0019] A feature extraction module, configured to input voice data provided by a user and extract personalized voice features from the voice data;

[0020] A feature fusion module is used to fuse the emotional features and the personalized speech features to generate a joint feature vector, and the joint feature vector is embedded in a deep learning model for speech synthesis;

[0021] The speech synthesis module is used to input the text to be processed and the joint feature vector, and generate the audio waveform through the deep learning model;

[0022] The output result module is used to analyze the context of the text to be processed, adjust and optimize the audio waveform, and obtain the final speech result.

[0023] On the basis of this technical solution, more preferably, the output result module also includes a context analysis unit and a speech feature adjustment unit.

[0024] The context analysis unit is used to analyze the context type of the text to be processed;

[0025] The speech feature adjustment unit is used to adjust the speech feature according to the context type so that the generated speech matches the context.

[0026] On the basis of this technical solution, it is further preferred that the sentiment analysis module includes a sentiment text analysis unit, an emotional state monitoring unit and a sentiment feature generation unit.

[0027] Among them, the sentiment text analysis unit is used to input the text to be processed and output the sentiment category and sentiment intensity;

[0028] An emotional state monitoring unit, used to obtain the user's emotional state information;

[0029] The emotion feature generation unit is used to combine the results of the emotion text analysis unit and the emotion state monitoring unit to generate an emotion feature vector.

[0030] On the basis of this technical solution, it is further preferred that the feature extraction module includes a speech data acquisition unit, a speech feature extraction unit and a speech model training unit.

[0031] A voice data collection unit, used to collect voice samples provided by users;

[0032] A speech feature extraction unit, used to extract personalized speech features from speech samples;

[0033] The speech model training unit is used to train a personalized speech model using personalized speech features.

[0034] In a third aspect, the present invention further provides a device comprising a processor and a memory for storing processor-executable instructions, wherein the processor executes the instructions to implement the multimodal emotion-driven text-to-speech method as described in the first aspect.

[0035] The multimodal emotion-driven text-to-speech method, system, and device described in the present invention have the following advantages over the prior art:

[0036] Emotional speech synthesis not only conveys the content of the text, but also generates emotionally charged speech based on the emotional characteristics of the text. Dynamic adjustment of emotional characteristics makes speech synthesis more vivid and natural. At the same time, personalized speech customization is achieved. By learning the user's voice characteristics, the system can generate speech that matches the user's voice style, improving the user's interactive experience and enhancing user affinity. Through multi-task learning and joint optimization methods, the present invention can optimize emotional expression while ensuring the naturalness of the speech, making the synthesized speech both natural and rich in emotional layers.

[0037] Due to the end-to-end training of the deep learning model, the present invention can generate high-quality speech in a shorter time and support real-time speech generation. It is suitable for scenarios with high real-time requirements such as voice assistants and voice interactions. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0039] Figure 1 This is a flow chart of the multimodal emotion-driven text-to-speech method described in Example 1 of the present invention;

[0040] Figure 2 This is a framework diagram of the multimodal emotion-driven text-to-speech system described in Example 3 of the present invention. DETAILED DESCRIPTION

[0041] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0042] The present invention provides a multimodal emotion-driven text-to-speech method, system, and device. By combining multimodal emotion analysis with the user's emotional state and utilizing a deep learning model to achieve personalized voice customization, the context understanding module can intelligently adjust voice features based on contextual information, thereby improving the naturalness and adaptability of the voice.

[0043] Example 1

[0044] like Figure 1 As shown, this embodiment provides a multimodal emotion-driven text-to-speech method, including the following steps:

[0045] S1, inputting a text to be processed, performing sentiment analysis, and identifying the sentiment features of the text to be processed;

[0046] S2, inputting voice data provided by the user and extracting personalized voice features of the voice data;

[0047] S3, fusing the emotional features and the personalized speech features to generate a joint feature vector, and embedding the joint feature vector into a deep learning model for speech synthesis;

[0048] S4, inputs the text to be processed and the joint feature vector, and generates an audio waveform through a deep learning model;

[0049] S5, analyzes the context of the text to be processed, adjusts and optimizes the audio waveform, and obtains the final speech result.

[0050] In this embodiment, specifically, the personalized speech features are extracted through self-supervised learning and transfer learning techniques.

[0051] In this embodiment, specifically, the personalized voice features include pitch, speaking speed or stress pattern.

[0052] In this embodiment, specifically, the adjustment and optimization adopt a multi-task learning method and a joint optimization method to optimize the emotional features and personalized voice features.

[0053] In this embodiment, specifically, personalized speech features are jointly optimized, specifically including using Tacotron2 to generate Mel spectrum, then converting the Mel spectrum into an audio waveform through WaveNet, and optimizing it in combination with emotion and personalized features to generate high-quality speech.

[0054] Example 2

[0055] A multimodal emotion-driven text-to-speech method comprises the following steps:

[0056] S1: Input the text to be processed T, perform sentiment analysis, and identify the sentiment features e of the text to be processed f , including emotion category e and emotion intensity s;

[0057] S2, input the voice data D provided by the user u , extract the personalized voice feature u of the voice data f ; Specifically, the personalized speech features are extracted through self-supervised learning and transfer learning techniques; In this embodiment, specifically, the personalized speech features include pitch, speaking speed or stress pattern;

[0058] S3, integrating the emotional features e f and the personalized voice feature u f , generate the joint feature vector f fusion ;

[0059] Specifically, f fusion =f(e f ,u f ), the joint feature vector f fusion Embed deep learning models for speech synthesis;

[0060] S4, input the text to be processed T and the joint feature vector f fusion , through the improved Tacotron2 and WaveNet network, generate audio waveform y with emotional expression and personalized style;

[0061] S5, analyzes the context of the text to be processed, adjusts and optimizes the audio waveform y, and obtains the final speech result.

[0062] Specifically, the adjustment and optimization adopts multi-task learning (MTL) and joint optimization methods to simultaneously optimize the emotional features e f and personalized voice features u f of learning.

[0063] Among them, the emotional feature e f The optimization goal is to maximize the accuracy of the sentiment classification task, and the loss function is defined as L emotion , whose optimization goal is to minimize the sentiment classification error:

[0064]

[0065] Among them, y i is the actual emotion label, is the sentiment label predicted by the system model based on this method.

[0066] Personalized voice optimization: personalized voice features f The optimization goal is to maximize the quality of personalized speech generation, and the loss function is defined as L personalized ,The optimization goal is to minimize the variance of speech generation.

[0067]

[0068] Among them, d i Personalized voice data provided to users, The generated speech waveform.

[0069] More specifically, the ultimate optimization goal is to simultaneously optimize the sentiment features e f and personalized voice features u f , the loss function L for joint optimization joint Defined as:

[0070] L joint =λ1L emotion +λ2L personalized

[0071] Where λ1 and λ2 are weighted coefficients that control the relative importance of sentiment optimization and personalized optimization.

[0072] In this embodiment, specifically, personalized speech features are jointly optimized, specifically including using Tacotron2 to generate Mel spectrum, then converting the Mel spectrum into an audio waveform through WaveNet, and optimizing it in combination with emotion and personalized features to generate high-quality speech.

[0073] Example 3

[0074] like Figure 2 As shown, a multimodal emotion-driven text-to-speech system includes:

[0075] The sentiment analysis module is used to input the text to be processed, perform sentiment analysis, and identify the sentiment characteristics of the text to be processed;

[0076] Specifically, by incorporating sentiment analysis of text, user emotional state, and contextual understanding into the text-to-speech (TTS) process, emotionally rich speech is generated. Traditional TTS systems rely solely on text content, ignoring the emotional changes in speech. Through the emotion-driven mechanism, the system described in this embodiment can adjust the intonation, speed, and emotional color of speech based on the emotional characteristics of the text, context, and user emotional state, making the speech more humane and expressive.

[0077] A feature extraction module, configured to input voice data provided by a user and extract personalized voice features from the voice data;

[0078] Specifically, through deep learning technology, a small amount of user voice samples is used to train a personalized voice feature model. The system described in this embodiment can generate voice that conforms to the user's personalized characteristics such as tone, voice style, and speaking speed. Compared with traditional TTS technology, a large amount of training data is required to simulate the voice characteristics of different individuals. This system can achieve personalized voice synthesis with a small amount of user data, greatly improving the convenience and effectiveness of voice customization.

[0079] A feature fusion module is used to fuse the emotional features and the personalized speech features to generate a joint feature vector, and the joint feature vector is embedded in a deep learning model for speech synthesis;

[0080] The speech synthesis module is used to input the text to be processed and the joint feature vector, and generate the audio waveform through the deep learning model;

[0081] The output result module is used to analyze the context of the text to be processed, adjust and optimize the audio waveform, and obtain the final speech result.

[0082] Specifically, by combining deep learning models with natural language processing technology, the present invention can automatically understand the context of input text (such as imperative, interrogative, and exclamatory sentences) and dynamically adjust speech features such as speech rate, intonation, and stress during the speech synthesis process, making the generated speech more natural, fluent, and contextually appropriate. This technology overcomes the limitations of traditional TTS technology, which outputs fixed speech features, and makes speech synthesis more flexible and diverse.

[0083] In this embodiment, specifically, the output result module further includes a context analysis unit and a speech feature adjustment unit.

[0084] The context analysis unit is used to analyze the context type of the text to be processed;

[0085] The speech feature adjustment unit is used to adjust the speech feature according to the context type so that the generated speech matches the context.

[0086] In this embodiment, specifically, the sentiment analysis module includes a sentiment text analysis unit, an emotional state monitoring unit and a sentiment feature generation unit.

[0087] Among them, the sentiment text analysis unit is used to input the text to be processed and output the sentiment category and sentiment intensity;

[0088] An emotional state monitoring unit, used to obtain the user's emotional state information;

[0089] The emotion feature generation unit is used to combine the results of the emotion text analysis unit and the emotion state monitoring unit to generate an emotion feature vector.

[0090] In this embodiment, specifically, the feature extraction module includes a speech data acquisition unit, a speech feature extraction unit and a speech model training unit.

[0091] A voice data collection unit, used to collect voice samples provided by users;

[0092] A speech feature extraction unit, used to extract personalized speech features from speech samples;

[0093] The speech model training unit is used to train a personalized speech model using personalized speech features.

[0094] Specifically, when this system is applied to a usage scenario, it is implemented in the following ways:

[0095] The sentiment analysis module is used to input "Hello, nice to meet you", perform sentiment analysis, identify the sentiment characteristics of the text to be processed, and output a result of high sentiment;

[0096] A feature extraction module, configured to input voice data provided by a user and extract personalized voice features from the voice data;

[0097] Specifically, the main task of the feature extraction module is to extract personalized speech features from the speech data provided by the user. These features may include the pitch, speaking speed, tone, intonation and pronunciation characteristics of the speech.

[0098] For example, when a user says “Hello, nice to meet you”, the feature extraction module extracts the following:

[0099] Pitch feature: Identify the pitch changes of each word and determine whether there is a rising or falling tone.

[0100] Speech rate features: Speech rate features are extracted based on the user's pronunciation speed, which may be a faster or slower speaking rhythm.

[0101] Emotional features: Determine the user's emotions when saying this sentence, such as whether it contains emotions such as pleasure and friendliness.

[0102] Personalized voice features: Extract personal features such as accent and voice style based on the user's timbre and unique pronunciation patterns.

[0103] These features are converted into a set of digital vectors, representing the user's personalized voice features. Suppose in the sentence "Hello, nice to meet you", the system recognizes:

[0104] Pitch changes: rising tone at the beginning of a sentence and falling tone at the end of a sentence.

[0105] Speaking speed: relatively steady and moderate.

[0106] Emotion: Friendly and warm.

[0107] These feature data are then passed to the next module.

[0108] A feature fusion module is used to fuse the emotional features and the personalized speech features to generate a joint feature vector, and the joint feature vector is embedded in a deep learning model for speech synthesis;

[0109] The feature fusion module primarily combines the extracted emotional features with the personalized speech features to form a joint feature vector. This allows for speech synthesis to consider not only the text content but also the influence of emotion and personalized speech, resulting in more natural and personalized speech.

[0110] For the sentence "Hello, nice to meet you", assume that the extracted features include:

[0111] Emotional characteristics: friendly and cheerful.

[0112] Personalized voice features: gentle voice style and slower speaking speed.

[0113] These features are weighted and fused in this module, possibly using some algorithms, such as weighted average or deep learning network, to fuse different features (such as emotion and timbre features) into a joint feature vector.

[0114] This joint feature vector represents a comprehensive speech feature, including emotion, personalized timbre, pitch, speaking speed, etc. For example, for the sentence "Hello, nice to meet you", it may contain the following fusion information:

[0115] Pitch characteristics: Gentle rising and falling tones.

[0116] Emotional characteristics: Pleasant emotion, with a slight smile in the voice.

[0117] Speaking speed characteristics: slow, clear pronunciation.

[0118] This joint feature vector will be fed into the next speech synthesis module.

[0119] The speech synthesis module is used to input the text to be processed and the joint feature vector, and generate the audio waveform through the deep learning model;

[0120] In the speech synthesis module, the input is the text to be processed, that is, "Hello, nice to meet you" and the joint feature vector, and then the deep learning model will generate an audio waveform based on these inputs.

[0121] The working process of the speech synthesis module includes:

[0122] Text-to-speech conversion: First, the system converts the text "Hello, nice to meet you" into voice content and generates the corresponding audio waveform.

[0123] Audio Adjustment: Based on the input joint feature vector, the system adjusts the audio pitch, speaking speed, intonation, and other aspects to make the voice reflect the user's personalized characteristics and emotions. For example, the system adjusts the audio to a gentle, pleasant voice with a moderate speaking speed based on the feature vector.

[0124] Therefore, the generated audio waveform not only reflects the text content, but also incorporates the user's personalized voice characteristics and emotions.

[0125] The output result module is used to analyze the context of the text to be processed, adjust and optimize the audio waveform, and obtain the final speech result;

[0126] The task of the output result module is to optimize and adjust the generated audio waveform to ensure that the final output speech is more natural, fluent, and consistent with the context.

[0127] In the "Hello, nice to meet you" example:

[0128] Contextual analysis: The output module analyzes the context of this sentence and understands that it is a friendly greeting, usually used during a first meeting, with a light-hearted and cheerful mood.

[0129] Audio optimization: Depending on the context, the system may optimize the audio waveform and adjust the pitch and rhythm to make the voice sound more appropriate to the actual situation (such as adjusting the speaking speed to make the voice clearer or fine-tuning the pitch to make the tone sound more friendly).

[0130] Finally, after context analysis and audio optimization, the output voice will be a natural and friendly "Hello, nice to meet you" with appropriate emotional color and user-personalized voice characteristics.

[0131] At the same time, a traditional speech synthesis model was evaluated against the system model of this solution. MOS scores, a standard method for evaluating the naturalness of synthesized speech, were used. MOS scores typically range from 1 to 5, with 5 representing the most natural and closest to human speech, and 1 representing the least natural. The scores were based on clarity, emotional expression, and audio quality, with higher scores indicating better model performance. The results are shown in the following table:

[0132] Table 1 Evaluation results of each model

[0133]

[0134] As can be seen from the table, Tacotron2+WaveNet in Example 2 is significantly superior to traditional speech synthesis models in terms of clarity, emotional expression, and audio quality.

[0135] In a preferred embodiment, a device is also provided, comprising a processor and a memory for storing processor-executable instructions, wherein the processor executes the instructions to implement the multimodal emotion-driven text-to-speech method as described in one aspect.

[0136] In summary, the present invention provides a multimodal, emotion-driven text-to-speech method, system, and device based on deep learning. By combining deep learning models with natural language processing technology, the present invention can automatically understand the context of the input text (such as imperative, interrogative, and exclamatory sentences) and dynamically adjust speech features such as speech rate, intonation, and stress during the speech synthesis process, making the generated speech more natural, fluent, and contextually appropriate. This technology breaks through the limitations of fixed speech feature output in traditional TTS technology, making speech synthesis more flexible and diverse.

[0137] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A multimodal emotion-driven text-to-speech method, characterized in that: The steps include: S1, inputting a text to be processed, performing sentiment analysis, and identifying the sentiment features of the text to be processed; S2, inputting voice data provided by the user and extracting personalized voice features of the voice data; S3, fusing the emotional features and the personalized speech features to generate a joint feature vector, and embedding the joint feature vector into a deep learning model for speech synthesis; S4, inputs the text to be processed and the joint feature vector, and generates an audio waveform through a deep learning model; S5, analyzes the context of the text to be processed, adjusts and optimizes the audio waveform, and obtains the final speech result.

2. The multimodal emotion-driven text-to-speech method according to claim 1, wherein: The personalized speech features are extracted through self-supervised learning and transfer learning techniques.

3. The multimodal emotion-driven text-to-speech method according to claim 2, wherein: The personalized voice features include pitch, speaking speed or stress pattern.

4. The multimodal emotion-driven text-to-speech method according to claim 1, wherein: The adjustment and optimization adopts a multi-task learning method and a joint optimization method to optimize the emotional features and personalized voice features.

5. The multimodal emotion-driven text-to-speech method according to claim 4, wherein: Jointly optimize personalized speech features, specifically using Tacotron2 to generate Mel spectrograms, then converting the Mel spectrograms into audio waveforms through WaveNet, and optimizing them in combination with emotional and personalized features to generate high-quality speech.

6. A multimodal emotion-driven text-to-speech system, characterized by: include: The sentiment analysis module is used to input the text to be processed, perform sentiment analysis, and identify the sentiment characteristics of the text to be processed; A feature extraction module, configured to input voice data provided by a user and extract personalized voice features from the voice data; A feature fusion module is used to fuse the emotional features and the personalized speech features to generate a joint feature vector, and the joint feature vector is embedded in a deep learning model for speech synthesis; The speech synthesis module is used to input the text to be processed and the joint feature vector, and generate the audio waveform through the deep learning model; The output result module is used to analyze the context of the text to be processed, adjust and optimize the audio waveform, and obtain the final speech result.

7. The multimodal emotion-driven text-to-speech system according to claim 6, wherein: The output result module also includes a context analysis unit and a speech feature adjustment unit. The context analysis unit is used to analyze the context type of the text to be processed; The speech feature adjustment unit is used to adjust the speech feature according to the context type so that the generated speech matches the context.

8. The multimodal emotion-driven text-to-speech system according to claim 6, wherein: The sentiment analysis module includes a sentiment text analysis unit, a sentiment state monitoring unit and a sentiment feature generation unit. Among them, the sentiment text analysis unit is used to input the text to be processed and output the sentiment category and sentiment intensity; An emotional state monitoring unit, used to obtain the user's emotional state information; The emotion feature generation unit is used to combine the results of the emotion text analysis unit and the emotion state monitoring unit to generate an emotion feature vector.

9. The multimodal emotion-driven text-to-speech system according to claim 6, wherein: The feature extraction module includes a speech data acquisition unit, a speech feature extraction unit and a speech model training unit. A voice data collection unit, used to collect voice samples provided by users; A speech feature extraction unit, used to extract personalized speech features from speech samples; The speech model training unit is used to train a personalized speech model using personalized speech features.

10. A device comprising a processor and a memory for storing instructions executable by the processor, characterized in that: The processor executes instructions to implement the multimodal emotion-driven text-to-speech method as described in any one of claims 1-5.