Fine emotion control TTS method and device based on large model, equipment and medium
By combining a large language model with a neural network, the generation and dynamic mapping of high-dimensional continuous emotional feature vectors are achieved, solving the problem of rough emotional expression in speech synthesis in existing technologies and improving the sophistication of speech synthesis and user experience.
Patent Information
- Application Number
- CN202511058827.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-09-26
AI Technical Summary
Existing speech synthesis technology has difficulty accurately understanding complex emotional states and their dynamic changes in highly emotional interaction scenarios, resulting in insufficient subtlety in the emotional expression of synthesized speech and an inability to convey delicate emotional information, affecting user experience and trust.
A method based on a large language model and pre-trained neural network is adopted to obtain the emotional feature vector of the input text, perform hierarchical coding and emotional curve analysis, generate a high-dimensional continuous emotional feature vector, and combine it with visual emotional features to achieve dynamic mapping between speech coding vectors and emotional curve vectors to generate the target speech.
It significantly improves the emotional expression delicacy and dynamic continuity of speech synthesis, can accurately express the emotional levels and temporal changes in complex texts, and enhances the immersive and realistic experience of user interaction.
Smart Images

Figure CN120708660A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a TTS method, device, equipment and medium for fine emotion control based on a large model. Background Art
[0002] In highly emotionally sensitive fields like banking, financial customer service, insurance consulting, and medical consultation, human-computer voice interaction has become a crucial service model. Text-to-speech (TTS) technology, as the core output vehicle for this interaction model, directly determines the user's perceived quality and trust in the service. In these scenarios, users not only expect accurate information but also a personalized, contextually relevant interactive experience. For example, they desire understanding and empathy during psychological counseling or personalized emotional adaptation in high-end customer service.
[0003] However, current mainstream speech synthesis technology has significant bottlenecks in meeting the above-mentioned high emotional interaction needs. A prominent common problem is that the emotional expression of its synthesized speech is not delicate enough. Specifically, it is difficult for the system to accurately understand and model the complex emotional state and its dynamic changes (such as emotional categories and intensity levels) contained in the input text. This directly leads to the system's inability to effectively perceive the emotional intensity gradient of the text at key interaction nodes that require intelligent emotional feedback, and then it is difficult to coordinate and generate multi-dimensional acoustic features that accurately match it (such as the ups and downs of the tone, the changes in the speed of speech, the warmth or calmness of the timbre, the rhythm of pauses, etc.). As a result, the synthesized speech often appears to be emotionally monotonous and mechanical, lacking the rich emotional levels and natural and smooth changes in human speech, and unable to convey delicate emotional information.
[0004] This lack of emotional expressiveness severely restricts the effective application and expansion of speech synthesis technology in scenarios involving deep emotional interaction. Users are unable to obtain the expected emotional resonance or adaptive feedback, resulting in a stilted interaction experience, increased alienation, and reduced trust. This can even lead to user dissatisfaction or resistance, hindering the development of more natural and humanized human-computer interaction. Summary of the Invention
[0005] The embodiments of the present invention provide a TTS method, apparatus, device, and medium for fine emotion control based on a large model, aiming to solve the problem of insufficient subtlety in emotion expression in existing TTS technologies.
[0006] In a first aspect, an embodiment of the present invention provides a TTS method for fine emotion control based on a large model, which includes:
[0007] Obtaining input text, and collecting sentiment feature vectors of the input text based on a preset large language model;
[0008] Encoding the emotion feature vector to obtain a speech coding vector;
[0009] Determining a sentiment curve vector of the input text based on a pre-trained neural network model;
[0010] A target speech corresponding to the input text is generated based on the speech coding vector and the emotion curve vector.
[0011] A further technical solution is that the emotional feature vector of the input text is collected based on a preset large language model, including:
[0012] Get prompt word configuration information;
[0013] The prompt word configuration information and the input text are input into the large language model, so that the large language model collects the emotional feature vector of the input text according to the prompt word configuration information.
[0014] A further technical solution is that encoding the emotion feature vector to obtain a speech coding vector includes:
[0015] The emotion feature vector is encoded by a preset hierarchical encoder to obtain a speech coding vector including a mutually independent global emotion category vector, a local emotion intensity vector and an NVs trigger vector.
[0016] A further technical solution is that the emotion curve vector of the input text is determined based on the pre-trained neural network model, including:
[0017] Inputting the input text into the neural network model, so as to obtain an initial sentiment curve vector of the input text through prediction by the neural network model;
[0018] The preset Transformer model is used to adjust the weight information of the initial emotion curve vector based on the attention mechanism to obtain the emotion curve vector.
[0019] A further technical solution is that generating a target speech corresponding to the input text based on the speech coding vector and the emotion curve vector includes:
[0020] generating an NVs segment vector based on the NVs trigger vector using a pre-trained diffusion model;
[0021] Through a preset TTS model, a target speech corresponding to the input text is generated based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector and the NVs segment vector.
[0022] A further technical solution is that the target speech corresponding to the input text is generated based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector and the NVs segment vector by using a preset TTS model, including:
[0023] Identifying the emotional focus words of the input text and generating a prosody reinforcement vector for the emotional focus words;
[0024] Acquire an input image and collect a visual emotion feature vector of the input image;
[0025] Generate a fusion vector based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector, the NVs segment vector, the rhythm reinforcement vector and the visual emotion feature vector;
[0026] The fusion vector is input into the TTS model so that the TTS model generates a target speech corresponding to the input text.
[0027] A further technical solution is that the identifying of the emotional focus words in the input text includes:
[0028] Identifying sentiment focus words of the input text based on a pre-trained fine-grained sentiment analysis model;
[0029] The collecting of the visual emotion feature vector of the input image includes:
[0030] The visual emotion feature vector of the input image is collected based on a pre-trained image feature collection model.
[0031] In a second aspect, an embodiment of the present invention further provides a large-model-based fine-grained emotion control TTS device, which includes a unit for executing the above method.
[0032] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.
[0033] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program can implement the above method when executed by a processor.
[0034] Embodiments of the present invention provide a large-scale model-based fine-grained emotion control TTS method, apparatus, device, and medium. The method comprises: obtaining input text, collecting an emotional feature vector of the input text based on a preset large language model; encoding the emotional feature vector to obtain a speech encoding vector; determining an emotional curve vector of the input text based on a pre-trained neural network model; and generating a target speech corresponding to the input text based on the speech encoding vector and the emotional curve vector. By constructing a large-scale language model-based fine-grained emotion control TTS method, the present invention significantly improves the emotional expression sophistication and dynamic continuity of speech synthesis. Traditional speech synthesis technology is limited by the coarse-grained control of discrete emotion tags, making it difficult to capture the implicit emotional levels and temporal changes in the text, resulting in mechanical and monotonous emotional expression in the synthesized speech. For example, when faced with complex text such as "She forced a smile and said she wasn't tired," traditional systems can only generate a monotonous "sad" tone, failing to convey the repressed and contradictory emotions behind the "forced smile." This solution innovatively incorporates a large language model to perform deep semantic analysis of input text, generating a high-dimensional, continuous emotional feature vector. This vector accurately represents emotional intensity, ambivalence, and implicit differences, laying the foundation for precise control. A pre-trained neural network then analyzes the text's grammatical structure and emotional turning points, dynamically generating a time-evolving emotional curve vector. This allows the speech prosody to naturally transition from calm to passionate. Finally, the speech encoding vector and the emotional curve vector are used in synergy to drive the acoustic model, achieving a dynamic mapping of emotional parameters to speech parameters. This multi-level processing mechanism enables the system to distinguish subtle differences within the same emotional category, such as generating distinct fundamental frequency fluctuation patterns when expressing "relief" and "ecstasy." It also supports the coherent evolution of emotion within long sentences, such as the anthropomorphic transition from a low-pitched buildup to an explosive climax in a suspenseful narrative. The technical breakthrough lies in upgrading static emotional labels to a continuous, dynamic emotional parameter space, fundamentally resolving the industry challenge of crude and fragmented emotional expression. This significantly enhances the immersive and authentic user experience in scenarios such as virtual assistants and audiobooks. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0036] Figure 1 A schematic diagram of a flow chart of a TTS method for fine emotion control based on a large model provided by an embodiment of the present invention;
[0037] Figure 2A schematic block diagram of a TTS device for fine emotion control based on a large model provided by an embodiment of the present invention;
[0038] Figure 3 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0040] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0041] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0042] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0043] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0044] See also Figure 1 The embodiment of the present invention provides a TTS method for fine emotion control based on a large model, which includes the following steps:
[0045] S1, obtaining input text, and collecting the sentiment feature vector of the input text based on a preset large language model.
[0046] In a specific implementation, the input text can be input by the user or generated based on a previous process (for example, the answer of the question-answering model), which is not specifically limited by the present invention. The large language model can be ChatGPT or Qwen2.5, which is not specifically limited by the present invention.
[0047] Traditional TTS systems typically rely on predefined emotion labels (such as "neutral" and "joy") to directly modulate acoustic parameters, resulting in coarse and homogenized emotional expression. This solution, however, extracts emotional feature vectors from text using a large language model. Essentially, this encodes the text's implicit emotional semantics (such as "restrained anger" or "mixed feelings of sadness and joy") into a mathematical representation in a high-dimensional continuous space. This process simulates humans' ability to deeply interpret textual emotion. For example, when inputting "He trembled as he accepted the trophy," the large language model can capture the excitement and tension implied by "trembling," generating an emotion vector distinct from the more common "joy."
[0048] In some preferred embodiments, the above step of "collecting the emotional feature vector of the input text based on a preset large language model" specifically includes the following steps: obtaining prompt word configuration information; inputting the prompt word configuration information and the input text into the large language model, so that the large language model collects the emotional feature vector of the input text according to the prompt word configuration information.
[0049] In specific implementations, prompt word configuration information can be user-entered, a practice not specifically limited by this invention. This prompt word configuration information guides the large language model's emotion extraction process, empowering the system with scenario-specific emotion adaptation capabilities and accurate analysis of user intent. Traditional emotional TTS lacks contextual awareness. The sentence "Prices have increased by 50%" conveys vastly different emotions in a financial report versus a consumer complaint (objective and calm in the former, anxious and dissatisfied in the latter). In this solution, prompt words are input into the large language model as meta-instructions (e.g., "serious news broadcast" or "children's story interpretation"), essentially modifying the model's internal attention weight distribution. Technically, a large language model (such as the GPT architecture) adds a learnable contextual embedding to the input layer via prompt words, enabling the model to emphasize semantic features related to the prompt word when calculating emotional features. For example, when the prompt word is "emergency alert," the model assigns higher emotional weights to verbs like "explosion" and "collapse," outputting a high-arousal feature vector. Conversely, the prompt word "bedtime story" suppresses intense emotions and amplifies the weights of adjectives like "warm" and "gentle." This dynamic intervention mechanism solves three key problems: first, it avoids emotional ambiguity within a single text (e.g., "really interesting" can be both sarcastic and admiring); second, it improves cross-scenario adaptability (educational TTS requires exaggerated emotions, while medical TTS requires restrained emotions); and third, it reduces manual annotation costs—users can adjust the emotional style simply by using prompts, without having to modify the text. Experiments have shown that the addition of prompts significantly improves sentiment classification accuracy and enables rapid emotional switching between multiple characters in the virtual anchor system (e.g., a news anchor and a cartoon character share the same model).
[0050] S2, encode the emotional feature vector to obtain a speech coding vector.
[0051] In a specific implementation, encoding the emotion feature vector can be accomplished by an encoder. The present invention does not specifically limit the type of encoder. After encoding by the encoder, a speech encoding vector can be obtained. The speech encoding vector refers to a structured representation obtained by encoding the emotion feature vector, and includes core acoustic parameters for speech synthesis.
[0052] For example, in some preferred embodiments, the above step of "encoding the emotional feature vector to obtain a speech coding vector" specifically includes the following steps: encoding and decoupling the emotional feature vector into independent global emotional category vectors, local emotional intensity vectors and NVs trigger vectors through a preset hierarchical encoder, which together constitute the speech coding vector. The global emotional category vector, the local emotional intensity vector and the NVs trigger vector can be understood as three parts of the speech coding vector. The global emotional category vector is a discrete vector that describes the overall emotional tone of the text (such as "sadness" and "joy"), and its function is to control the macro-emotional direction of the speech and ensure the emotional consistency of the synthesized speech. The local emotional intensity vector is a time series vector that encodes word-level emotional intensity fluctuations, and its function is to fine-tune the emotional expression of key words (such as strengthening the "very" in "very painful"). The NVs trigger vector is a control vector that triggers non-speech sound effects (such as laughter and sobbing), and its function is to control the type, triggering timing and duration of NVs.
[0053] In a specific implementation, the layered encoder may be a layered VAE encoder, which is not specifically limited in the present invention. Below, the VAE encoder is used for illustration.
[0054] A hierarchical VAE encoder is used to decouple the emotion feature vector into a global emotion category vector, a local emotion intensity vector, and a NVs (Non-Verbal Sounds) trigger vector, enabling modularization and fine-tuning of emotion parameters. Technically, the VAE's latent variable separation feature forces different vectors to focus on independent features: the global vector identifies the primary emotion category (e.g., "sad"), the local vector controls word-level intensity (e.g., "very" is 30% stronger than "a little"), and the NVs trigger vector manages non-verbal elements (e.g., sobs, laughter). For example, when generating the sentence "She smiled and said: It's okay," the global vector selects the "forced smile" category, the local vector injects an intensity peak at the word "no," and the NVs trigger vector adds a trembling breath sound at the end of the sentence. This decoupling allows developers to adjust certain parameters independently (e.g., enhancing intensity without changing the emotion category), resolving the problem of adjustment distortion caused by the coupling of emotion elements in end-to-end models.
[0055] Specifically, this scheme uses the latent space decoupling characteristics of variational autoencoder (VAE) to split the emotional features into three vectors:
[0056] Global emotion category vector: represents the overall emotional tone (such as the "sad" category) and ensures its discrete distinguishability through the VAE's category constraint loss (Classification Loss);
[0057] Local sentiment intensity vector: encodes word-level sentiment intensity fluctuations (such as the intensity peak of “very” in “very painful”), and the temporal convolutional layer extracts the dependency relationship between adjacent words;
[0058] NVs trigger vector: dedicated to controlling the triggering timing and type of non-verbal sounds, such as sobs and laughter. Its training relies on adversarial learning to distinguish between speech and NVs.
[0059] For example, when generating the sentence "She choked and said, 'I'm fine,'" the global vector selects the "fighting back tears" category, the local vector injects an intensity pulse at the "choking" part, and the NVs trigger vector generates a 0.5-second sob at the beginning of the sentence. This decoupling offers two major advantages: first, it allows independent parameter adjustment—users can increase intensity independently without changing the emotion type; second, it improves synthesis efficiency. After layered encoding, NVs generation can be offloaded to the diffusion model (see Section 5), avoiding overloading the main TTS model.
[0060] S3: Determine the sentiment curve vector of the input text based on a pre-trained neural network model.
[0061] In a specific implementation, a neural network model is trained using pre-labeled training data to enable it to recognize the sentiment curve vector of the input text. Furthermore, the sentiment curve vector of the input text is determined using the pre-trained neural network model. The neural network model may be, for example, a Bi-LSTM, which is not specifically limited in the present invention. A sentiment curve vector refers to a time series vector that represents the dynamic evolution of text sentiment over time. Sentiment curve vectors are used to model (or represent) the gradual change of emotion in long sentences (e.g., from calm to passionate), addressing the fragmented nature of traditional TTS emotional expression.
[0062] The introduction of sentiment curve vectors further addresses the issue of emotional dynamics: a pre-trained neural network analyzes the temporal structure of text (such as clause transitions and interjection placement) to generate a trajectory of emotional intensity over time. For example, in the sentence "We won the game...but my teammate was injured," the curve vector constructs a parabola that first rises and then falls, causing the synthesized speech to rise in pitch at the beginning of the sentence "We won" and then shift to a low, vibrato at the end of the sentence "Injured."
[0063] In some preferred embodiments, the above step of "determining the sentiment curve vector of the input text based on a pre-trained neural network model" specifically includes the following steps: inputting the input text into the neural network model so that the neural network model predicts the initial sentiment curve vector of the input text; through a preset Transformer model, adjusting the weight information of the initial sentiment curve vector based on the attention mechanism to obtain the sentiment curve vector.
[0064] In practice, a dual-stage sentiment curve generation approach using a neural network and a Transformer solves the problems of temporal drift and weight imbalance in sentiment expression in long texts. Traditional RNN models suffer from sentiment decay when processing long sequences (e.g., a drop in sentiment intensity at the end of a paragraph), while a single Transformer can over-focus on local features. In this approach, a pre-trained neural network (e.g., a Bi-LSTM) first generates an initial sentiment curve vector. This vector establishes a basic temporal profile by analyzing the text's grammatical structure (e.g., rising intonation in interrogative sentences) and sentiment word density (e.g., cumulative intensity for three consecutive negative words). However, this approach can result in underweighting key sentiment words (e.g., the negation "not" is ignored). Therefore, the Transformer attention mechanism undergoes a secondary calibration: its Query-Key-Value mechanism calculates the sentiment contribution of each word, assigning higher attention weights to transition words (e.g., "but") and adverbs of degree (e.g., "extremely"). Taking the text "The weather is nice, but I heard a typhoon is coming" as an example, the initial curve may form a peak at "nice"; the Transformer will recognize the transitional semantics of "but", shift the peak weight back to "typhoon", and generate a steep drop curve there. In terms of technical principles, the attention weight matrix is dynamically adjusted through a trainable scaled dot product (Scaled Dot-Product) to ensure that the sentiment curve is strictly aligned with the semantic center of gravity. In addition, positional encoding (Positional Encoding) retains word order information to avoid time sequence confusion. This design effectively reduces the consistency error of the sentiment curve in long texts (>50 words), and significantly improves the emotional coherence in the broadcast of multi-turn plots, especially.
[0065] S4: Generate a target speech corresponding to the input text based on the speech coding vector and the emotion curve vector.
[0066] In specific implementations, the fusion of speech coding vectors and emotion curves maps abstract emotions into specific parameters such as fundamental frequency, duration, and spectrum through a parameterized acoustic model (such as Tacotron), thereby accurately controlling the generation of the target speech corresponding to the input text. Its effect far exceeds simple emotion classification - it allows for intensity gradients within the same emotion category (such as a smooth transition from "smile" to "laugh"), and can handle complex and contradictory emotions (such as "bitter laughter"). This end-to-end fine control enables the synthesized speech to have anthropomorphic emotional expressiveness in scenarios such as film and television dubbing and psychotherapy robots.
[0067] In some preferred embodiments, the above step of "generating the target speech corresponding to the input text based on the speech coding vector and the emotion curve vector" specifically includes the following steps: generating an NVs segment vector based on the NVs trigger vector through a pre-trained diffusion model; generating the target speech corresponding to the input text based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector and the NVs segment vector through a preset TTS model.
[0068] In specific implementations, a diffusion model is introduced to generate NVs segment vectors, which work in conjunction with the main TTS model to achieve dynamic, high-fidelity synthesis of non-speech elements and seamless fusion of multiple signals. NVs segment vectors refer to the waveform features of non-speech sound effects generated by the diffusion model. Their function is to replace pre-made sound effect libraries and achieve high-fidelity, dynamically adaptive NVs synthesis (such as breath-dominated sobbing sounds).
[0069] Existing technologies usually use pre-made sound effect libraries to splice NVs, resulting in three major defects: mechanical sound quality, emotional disconnection (such as misalignment of laughter and voice rhythm), and inability to adapt to new scenarios.
[0070] In this approach, a diffusion model (such as WaveGrad) generates NVs through an iterative denoising process. The NVs trigger vector serves as a conditional input, guiding the gradual evolution of random noise into the target acoustic feature. For example, when the trigger vector points to "sob," the diffusion model generates a waveform dominated by breathy sounds, with increasing amplitude and vibrato through iterative denoising, which is far more natural than a fixed sample synthesized by regularity.
[0071] Subsequently, the TTS model integrates the NVs segment with the global emotion category vector, the local emotion intensity vector, and the emotion curve vector through a gated fusion unit. Specifically, the NVs segment is aligned with the speech on the time axis (e.g., laughter is inserted 500ms after the end of a sentence), and spectral mixing is used to avoid a sense of discontinuity. The technical advantages are threefold: first, the diffusion model can generate NVs outside the training data (e.g., new mechanical sound effects); second, high emotional consistency—the global emotion category vector constrains the NVs style (e.g., "joy" corresponds to short laughter, "sadness" corresponds to long breaths); and third, computational efficiency is optimized, with NVs generation and speech synthesis performed in parallel.
[0072] In some preferred embodiments, the above step of "generating the target speech corresponding to the input text based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector and the NVs segment vector through a preset TTS model" specifically includes the following steps: identifying the emotion focus words of the input text and generating a prosody reinforcement vector of the emotion focus words; acquiring an input image and collecting the visual emotion feature vector of the input image; generating a fusion vector based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector and the NVs segment vector, the prosody reinforcement vector and the visual emotion feature vector; inputting the fusion vector into the TTS model so that the TTS model generates the target speech corresponding to the input text.
[0073] In specific implementations, the combination of emotional focus word reinforcement and visual emotional feature fusion achieves multimodal emotional enhancement and highlights key information. Traditional multimodal TTS often simply splices text and image features, resulting in modal conflict (for example, the synthesized intonation is dissonant when the word "happy" is paired with a sad image).
[0074] This solution first locates the emotional focus words (such as "finally" in "finally waiting for you") through a fine-grained sentiment analysis model. Its technical principle is that object-level sentiment classification (such as the BERT+CRF model of ACL 2020) identifies objects / phrases that carry core emotions in the text. The prosody enhancement vector that is subsequently generated will apply acoustic modulation to the focus words: the fundamental frequency is increased by 15-30%, the syllable is extended by 50%, and amplitude tremors are added (simulating the vibration of the vocal cords when excited). For example, in "absolutely no failure is allowed", the word "absolutely" is enhanced to an explosive accent. The prosody enhancement vector refers to the acoustic modulation parameters for emotional focus words (such as "finally"), and its function is to highlight key emotional words and enhance expressiveness by increasing the fundamental frequency (15-30%), extending the syllable (50%), and amplitude tremors. The prosody enhancement vector can be pre-set by a person skilled in the art, or predicted by a pre-trained neural network model based on the emotional focus words, and the present invention is not specifically limited.
[0075] At the same time, image feature models (such as CLIP-ViT) extract visual emotion feature vectors. Their cross-modal alignment capability maps visual elements to acoustic parameters (e.g., a furrowed brow corresponds to increased glottalization). Visual emotion feature vectors are cross-modal emotional representations extracted from the input image. Their function is to fuse visual emotion (e.g., furrowed brows → glottalization) and resolve emotional conflicts between images and text.
[0076] Furthermore, the key innovation of the present invention lies in the generation mechanism of the fusion vector: Multi-Head Attention is used to calculate the correlation weights among the global emotion category vector, the local emotion intensity vector, the emotion curve vector, the NVs segment vector generation, the rhythm reinforcement vector and the visual emotion feature vector, and weighted fusion is performed to obtain the fusion vector. When the text conflicts with the visual emotion (such as the text "awesome" with an image of rolling eyes), the system automatically reduces the emotional intensity of the text and injects sarcastic breath. Furthermore, the fusion vector is input into the TTS model, and the TTS model generates the target speech corresponding to the input text based on the fusion vector.
[0077] In some preferred embodiments, the above step of "identifying the sentiment focus words of the input text" specifically includes the following steps: identifying the sentiment focus words of the input text based on a pre-trained fine-grained sentiment analysis model.
[0078] In specific implementation, the traditional method uses overall text sentiment classification (such as judging the whole sentence as "positive"), which is unable to locate the specific emotional object, resulting in the loss of focus of rhythm reinforcement. This solution uses a pre-trained fine-grained sentiment analysis model (such as the BERT-based fine-grained model of ACL 2020), and accurately labels the emotional polarity of each object in the text through joint training of entity recognition and sentiment attribution. For example, in "The restaurant environment is elegant, but the service is poor", the model labels "environment" as positive (requires a soothing tone) and "service" as negative (requires rhythm reinforcement as a sharp sound). Its technical principle relies on dependency syntactic analysis-the emotional word "bad" is associated with the object "service" through a dependency arc (DependencyArc), ensuring that the reinforcement vector is accurately applied to "service" rather than the entire sentence.
[0079] In some preferred embodiments, the above step of "collecting the visual emotion feature vector of the input image" specifically includes the following steps: collecting the visual emotion feature vector of the input image based on a pre-trained image feature collection model.
[0080] In specific implementations, the image feature acquisition model can use a visual emotion knowledge distillation architecture (such as the ResNet + attention module pre-trained by AffectNet) to locate the emotional regions of the image (for example, the tear region contributes 70% to "sadness") through the gradient weighted class activation map (Grad-CAM), avoiding background noise interference. The visual emotion feature vector output by the model is physically interpretable: high-frequency components correspond to tension (such as dilated pupils), and low-frequency components correspond to depression (such as drooping corners of the mouth). This high-precision recognition provides reliable input for subsequent fusion, effectively reducing the emotion mismatch rate in cross-modal TTS.
[0081] An embodiment of the present invention proposes a large-model-based, fine-grained emotion control text-to-speech (TTS) method, comprising: obtaining input text, collecting an emotion feature vector of the input text based on a preset large language model; encoding the emotion feature vector to obtain a speech encoding vector; determining an emotion curve vector of the input text based on a pretrained neural network model; and generating a target speech corresponding to the input text based on the speech encoding vector and the emotion curve vector. By constructing a large-language-based, fine-grained emotion control TTS method, the present invention significantly improves the emotional sophistication and dynamic continuity of speech synthesis. Traditional speech synthesis technology, limited by the coarse-grained control of discrete emotion labels, struggles to capture the underlying emotional layers and temporal variations in text, resulting in mechanical and monotonous emotional expression in synthesized speech. For example, when faced with complex text such as "She forced a smile and said she wasn't tired," traditional systems can only generate a monotonous "sad" tone, failing to convey the repressed and conflicting emotions behind the "forced smile." This solution innovatively introduces a large language model to perform deep semantic analysis of the input text, generating a high-dimensional, continuous emotion feature vector that accurately represents emotional intensity, conflict, and implicit feature differences, laying the foundation for fine-grained control. A pre-trained neural network then analyzes the text's grammatical structure and emotional turning points, dynamically generating a time-evolving emotion curve vector. This allows the speech prosody to naturally transition from calm to passionate. Finally, the speech encoding vector and emotion curve vector are used in conjunction to drive the acoustic model, dynamically mapping emotion parameters to speech parameters. This multi-level processing mechanism enables the system to distinguish subtle differences within the same emotion category, such as generating distinct fundamental frequency fluctuation patterns when expressing "relief" and "ecstasy." It also supports the coherent evolution of emotion within long sentences, such as the anthropomorphic transition from subdued foreshadowing to explosive climax in a suspenseful narrative. Its technological breakthrough lies in upgrading static emotion labels to a continuous and dynamic emotion parameter space. This fundamentally addresses the industry challenge of crude and fragmented emotional expression, significantly enhancing the immersive and authentic user experience in scenarios such as virtual assistants and audiobooks.
[0082] This embodiment of the present invention proposes a large-scale model-based refined emotion control TTS method that can generate voice services with delicate emotional expression in highly emotionally sensitive fields such as financial customer service, insurance consulting, and medical guidance, thereby improving the user experience. Specific application cases are as follows:
[0083] 1. Banking scenario: Enter the text: "Recent stock market volatility has intensified, and your stock portfolio may face a risk of a drawdown of more than 10%."
[0084] Emotional control is achieved:
[0085] Global sentiment category vector → "serious warning";
[0086] Local emotion intensity vector → triggering intensity peaks (sharp rise in pitch + syllable extension) at “above 10%”;
[0087] Emotional curve vector → a gradual transition from a smooth narrative to a sharp warning;
[0088] NVs segment vector → Add a short warning sound effect at the end of the sentence.
[0089] 2. Insurance Scenario
[0090] Enter text: "We understand you're feeling bad, but this accident isn't covered."
[0091] Emotional control is achieved:
[0092] Prompt word configuration → "Empathic communication";
[0093] Global sentiment category vector → “regretful but determined”;
[0094] Prosodic reinforcement vector → adding a warm tone to the word “understand” and a heavy tone to “not”;
[0095] Emotional curve vector → The first half of the sentence slows down (empathy), while the second half speeds up (principle).
[0096] Reduce customers’ negative emotions when they are rejected and maintain brand trust.
[0097] 3. Medical guidance
[0098] Enter the text: "Please bring your family with you. Your CT report requires further discussion."
[0099] Emotional control is achieved:
[0100] Fusion of visual and emotional features → Generate a repressive low-frequency spectrum based on the "stage IV" label on the report card;
[0101] Local emotion intensity vector → "further discussion" adds a slight tremolo (simulating the doctor's hesitation);
[0102] NVs trigger vector → Generate a 0.8 second silent pause at the beginning of the sentence.
[0103] The severity of the disease is conveyed through acoustic metaphors, providing patients with a psychological buffer period.
[0104] 4. Financial Customer Service
[0105] Enter the text: "Unusual overseas consumption has been detected. Please verify immediately whether it is my operation."
[0106] Emotional control is achieved:
[0107] Emotional curve vector → The first two seconds are a steady narration of facts, and the last three seconds gradually increase the sense of urgency;
[0108] Local sentiment intensity vector → The base frequency of the word “immediately” increases by 25%;
[0109] Rhythmic reinforcement vector → "abnormal" adds explosive accents;
[0110] Multimodal collaboration → When displaying the fraudulent transaction map simultaneously, voice will dynamically enhance regional keywords along with visual hotspots.
[0111] Balance the intensity of warnings and operational guidance to reduce the rate of erroneous operations.
[0112] See also Figure 2 , Figure 2 This is a schematic block diagram of a TTS device 20 for fine emotion control based on a large model provided by an embodiment of the present invention. Corresponding to the above-mentioned TTS method for fine emotion control based on a large model, the present invention also provides a TTS device 20 for fine emotion control based on a large model. The TTS device 20 for fine emotion control based on a large model includes a unit for executing the above-mentioned TTS method for fine emotion control based on a large model. The TTS device 20 for fine emotion control based on a large model can be configured in a desktop computer, tablet computer, laptop computer, or other terminal. Specifically, the TTS device 20 for fine emotion control based on a large model includes:
[0113] An acquisition unit 21 is configured to acquire an input text and collect a sentiment feature vector of the input text based on a preset large language model;
[0114] An encoding unit 22, configured to encode the emotion feature vector to obtain a speech encoding vector;
[0115] a determining unit 23, configured to determine a sentiment curve vector of the input text based on a pre-trained neural network model;
[0116] The generating unit 24 is configured to generate a target speech corresponding to the input text based on the speech coding vector and the emotion curve vector.
[0117] In some preferred embodiments, the collecting of the sentiment feature vector of the input text based on a preset large language model includes:
[0118] Get prompt word configuration information;
[0119] The prompt word configuration information and the input text are input into the large language model, so that the large language model collects the emotional feature vector of the input text according to the prompt word configuration information.
[0120] In some preferred embodiments, encoding the emotion feature vector to obtain a speech coding vector includes:
[0121] Through a preset hierarchical encoder, the emotion feature vector is encoded and decoupled into independent global emotion category vectors, local emotion intensity vectors and NVs trigger vectors, which together constitute the speech coding vector.
[0122] In some preferred embodiments, determining the sentiment curve vector of the input text based on a pre-trained neural network model includes:
[0123] Inputting the input text into the neural network model, so as to obtain an initial sentiment curve vector of the input text through prediction by the neural network model;
[0124] The preset Transformer model is used to adjust the weight information of the initial emotion curve vector based on the attention mechanism to obtain the emotion curve vector.
[0125] In some preferred embodiments, generating a target speech corresponding to the input text based on the speech coding vector and the emotion curve vector includes:
[0126] generating an NVs segment vector based on the NVs trigger vector using a pre-trained diffusion model;
[0127] Through a preset TTS model, a target speech corresponding to the input text is generated based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector and the NVs segment vector.
[0128] In some preferred embodiments, generating the target speech corresponding to the input text based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector, and the NVs segment vector by using a preset TTS model includes:
[0129] Identifying the emotional focus words of the input text and generating a prosody reinforcement vector for the emotional focus words;
[0130] Acquire an input image and collect a visual emotion feature vector of the input image;
[0131] Generate a fusion vector based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector, the NVs segment vector, the rhythm reinforcement vector and the visual emotion feature vector;
[0132] The fusion vector is input into the TTS model so that the TTS model generates a target speech corresponding to the input text.
[0133] In some preferred embodiments, identifying the emotional focus words of the input text includes:
[0134] Identifying sentiment focus words of the input text based on a pre-trained fine-grained sentiment analysis model;
[0135] The collecting of the visual emotion feature vector of the input image includes:
[0136] The visual emotion feature vector of the input image is collected based on a pre-trained image feature collection model.
[0137] It should be noted that technical personnel in the relevant field can clearly understand that the specific implementation process of the above-mentioned large-model-based fine emotion control TTS device 20 and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and conciseness of the description, it will not be repeated here.
[0138] The above-mentioned fine emotion control TTS device 20 based on a large model can be implemented in the form of a computer program. The computer program can be used in Figure 3 Runs on the computer equipment shown.
[0139] See also Figure 3 , Figure 3 This is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 can be a terminal or a server. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, personal digital assistant, wearable device, or other electronic device with communication capabilities. The server can be a standalone server or a server cluster consisting of multiple servers.
[0140] The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .
[0141] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 can execute a large model-based fine emotion control TTS method.
[0142] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0143] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a large model-based fine emotion control TTS method.
[0144] The network interface 505 is used to communicate with other devices over the network. Those skilled in the art will appreciate that the above structure is merely a block diagram of a portion of the structure related to the present invention and does not limit the computer device 500 to which the present invention is applied. A specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0145] The processor 502 is configured to execute a computer program 5032 stored in the memory to implement the following steps:
[0146] Obtaining input text, and collecting sentiment feature vectors of the input text based on a preset large language model;
[0147] Encoding the emotion feature vector to obtain a speech coding vector;
[0148] Determining a sentiment curve vector of the input text based on a pre-trained neural network model;
[0149] A target speech corresponding to the input text is generated based on the speech coding vector and the emotion curve vector.
[0150] In some preferred embodiments, the collecting of the sentiment feature vector of the input text based on a preset large language model includes:
[0151] Get prompt word configuration information;
[0152] The prompt word configuration information and the input text are input into the large language model, so that the large language model collects the emotional feature vector of the input text according to the prompt word configuration information.
[0153] In some preferred embodiments, encoding the emotion feature vector to obtain a speech coding vector includes:
[0154] Through a preset hierarchical encoder, the emotion feature vector is encoded and decoupled into independent global emotion category vectors, local emotion intensity vectors and NVs trigger vectors, which together constitute the speech coding vector.
[0155] In some preferred embodiments, determining the sentiment curve vector of the input text based on a pre-trained neural network model includes:
[0156] Inputting the input text into the neural network model, so as to obtain an initial sentiment curve vector of the input text through prediction by the neural network model;
[0157] The preset Transformer model is used to adjust the weight information of the initial emotion curve vector based on the attention mechanism to obtain the emotion curve vector.
[0158] In some preferred embodiments, generating a target speech corresponding to the input text based on the speech coding vector and the emotion curve vector includes:
[0159] generating an NVs segment vector based on the NVs trigger vector using a pre-trained diffusion model;
[0160] Through a preset TTS model, a target speech corresponding to the input text is generated based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector and the NVs segment vector.
[0161] In some preferred embodiments, generating the target speech corresponding to the input text based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector, and the NVs segment vector by using a preset TTS model includes:
[0162] Identifying the emotional focus words of the input text and generating a prosody reinforcement vector for the emotional focus words;
[0163] Acquire an input image and collect a visual emotion feature vector of the input image;
[0164] Generate a fusion vector based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector, the NVs segment vector, the rhythm reinforcement vector and the visual emotion feature vector;
[0165] The fusion vector is input into the TTS model so that the TTS model generates a target speech corresponding to the input text.
[0166] In some preferred embodiments, identifying the emotional focus words of the input text includes:
[0167] Identifying sentiment focus words of the input text based on a pre-trained fine-grained sentiment analysis model;
[0168] The collecting of the visual emotion feature vector of the input image includes:
[0169] The visual emotion feature vector of the input image is collected based on a pre-trained image feature collection model.
[0170] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0171] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.
[0172] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program. When the computer program is executed by a processor, the processor performs the following steps:
[0173] Obtaining input text, and collecting sentiment feature vectors of the input text based on a preset large language model;
[0174] Encoding the emotion feature vector to obtain a speech coding vector;
[0175] Determining a sentiment curve vector of the input text based on a pre-trained neural network model;
[0176] A target speech corresponding to the input text is generated based on the speech coding vector and the emotion curve vector.
[0177] In some preferred embodiments, the collecting of the sentiment feature vector of the input text based on a preset large language model includes:
[0178] Get prompt word configuration information;
[0179] The prompt word configuration information and the input text are input into the large language model, so that the large language model collects the emotional feature vector of the input text according to the prompt word configuration information.
[0180] In some preferred embodiments, encoding the emotion feature vector to obtain a speech coding vector includes:
[0181] Through a preset hierarchical encoder, the emotion feature vector is encoded and decoupled into independent global emotion category vectors, local emotion intensity vectors and NVs trigger vectors, which together constitute the speech coding vector.
[0182] In some preferred embodiments, determining the sentiment curve vector of the input text based on a pre-trained neural network model includes:
[0183] Inputting the input text into the neural network model, so as to obtain an initial sentiment curve vector of the input text through prediction by the neural network model;
[0184] The preset Transformer model is used to adjust the weight information of the initial emotion curve vector based on the attention mechanism to obtain the emotion curve vector.
[0185] In some preferred embodiments, generating a target speech corresponding to the input text based on the speech coding vector and the emotion curve vector includes:
[0186] generating an NVs segment vector based on the NVs trigger vector using a pre-trained diffusion model;
[0187] Through a preset TTS model, a target speech corresponding to the input text is generated based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector and the NVs segment vector.
[0188] In some preferred embodiments, generating the target speech corresponding to the input text based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector, and the NVs segment vector by using a preset TTS model includes:
[0189] Identifying the emotional focus words of the input text and generating a prosody reinforcement vector for the emotional focus words;
[0190] Acquire an input image and collect a visual emotion feature vector of the input image;
[0191] Generate a fusion vector based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector, the NVs segment vector, the rhythm reinforcement vector and the visual emotion feature vector;
[0192] The fusion vector is input into the TTS model so that the TTS model generates a target speech corresponding to the input text.
[0193] In some preferred embodiments, identifying the emotional focus words of the input text includes:
[0194] Identifying sentiment focus words of the input text based on a pre-trained fine-grained sentiment analysis model;
[0195] The collecting of the visual emotion feature vector of the input image includes:
[0196] The visual emotion feature vector of the input image is collected based on a pre-trained image feature collection model.
[0197] The storage medium is a physical, non-transient storage medium, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a magnetic disk, or an optical disk, etc. Any physical storage medium capable of storing program code can be non-volatile or volatile.
[0198] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0199] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.
[0200] The steps in the methods of the embodiments of the present invention may be adjusted in order, combined, or deleted as needed. The units in the devices of the embodiments of the present invention may be combined, divided, or deleted as needed. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0201] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.
[0202] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0203] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, to the extent such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to encompass such changes and modifications.
[0204] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A TTS method for fine emotion control based on a large model, characterized by: include: Obtaining input text, and collecting sentiment feature vectors of the input text based on a preset large language model; Encoding the emotion feature vector to obtain a speech coding vector; Determining a sentiment curve vector of the input text based on a pre-trained neural network model; A target speech corresponding to the input text is generated based on the speech coding vector and the emotion curve vector.
2. The TTS method for fine emotion control based on a large model according to claim 1 is characterized in that The collecting of the sentiment feature vector of the input text based on the preset large language model includes: Get prompt word configuration information; The prompt word configuration information and the input text are input into the large language model, so that the large language model collects the emotional feature vector of the input text according to the prompt word configuration information.
3. The TTS method for fine emotion control based on a large model according to claim 1 is characterized in that The step of encoding the emotion feature vector to obtain a speech coding vector includes: The emotion feature vector is encoded by a preset hierarchical encoder to obtain a speech coding vector including a mutually independent global emotion category vector, a local emotion intensity vector and an NVs trigger vector.
4. The TTS method for fine emotion control based on a large model according to claim 1 is characterized in that The determining of the sentiment curve vector of the input text based on the pre-trained neural network model includes: Inputting the input text into the neural network model, so as to obtain an initial sentiment curve vector of the input text through prediction by the neural network model; The preset Transformer model is used to adjust the weight information of the initial emotion curve vector based on the attention mechanism to obtain the emotion curve vector.
5. The TTS method for fine emotion control based on a large model according to claim 3 is characterized in that Generating a target speech corresponding to the input text based on the speech coding vector and the emotion curve vector includes: generating an NVs segment vector based on the NVs trigger vector using a pre-trained diffusion model; Through a preset TTS model, a target speech corresponding to the input text is generated based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector and the NVs segment vector.
6. The TTS method for fine emotion control based on a large model according to claim 5 is characterized in that: The method of generating a target speech corresponding to the input text based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector, and the NVs segment vector by using a preset TTS model includes: Identifying the emotional focus words of the input text and generating a prosody reinforcement vector for the emotional focus words; Acquire an input image and collect a visual emotion feature vector of the input image; Generate a fusion vector based on the global emotion category vector, the local emotion intensity vector, the emotion curve vector, the NVs segment vector, the rhythm reinforcement vector and the visual emotion feature vector; The fusion vector is input into the TTS model so that the TTS model generates a target speech corresponding to the input text.
7. The TTS method for fine emotion control based on a large model according to claim 6 is characterized in that: The identifying of the emotional focus words of the input text includes: Identifying sentiment focus words of the input text based on a pre-trained fine-grained sentiment analysis model; The collecting of the visual emotion feature vector of the input image includes: The visual emotion feature vector of the input image is collected based on a pre-trained image feature collection model.
8. A TTS device with fine emotion control based on a large model, characterized in that: The method comprises a unit for executing the method according to any one of claims 1 to 7.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the computer program can implement the method according to any one of claims 1 to 7.
Citation Information
Cited By
TTS speech emotion enhancement method, electronic equipment and storage medium
CN121096314A
Man-machine interaction voice perception method and system based on gradient intelligent dispatch subnet pool
CN121148370A
Speech synthesis method and device based on semantic sentiment analysis and rhythm regulation and control, equipment and storage medium
CN121862077A
A speech synthesis method and device based on semantic sentiment analysis and rhythm regulation, equipment and storage medium
CN121862077B