Emotional speech synthesis method and system based on dialogue context and personal experience
By constructing a multi-dimensional emotion enhancement condition set and an automatic annotation pipeline, the problem of lacking dialogue context and personal experience modeling in existing emotional speech synthesis technologies is solved, achieving emotional speech synthesis with high naturalness, high emotional accuracy and character consistency, and applicable to a variety of application scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI JIAOTONG UNIV
- Filing Date
- 2026-05-07
- Publication Date
- 2026-07-10
AI Technical Summary
Existing emotional speech synthesis technologies lack modeling of personal experiences and dialogue context, resulting in coarse-grained emotional expression. This makes it difficult to generate speech that fits the plot logic and is consistent with the characters. Furthermore, continuous acoustic features cannot be directly utilized by language models, affecting the accuracy of emotion judgment.
By acquiring or constructing a multi-dimensional set of emotion enhancement conditions, including personal experience descriptions, dialogue context descriptions, paralinguistic descriptions, and open-vocabulary emotion tags, and combining the semantic generation module and the acoustic reconstruction module, target emotional speech is generated. By utilizing an automatic annotation pipeline and acoustic feature quantization methods, high naturalness and accuracy of emotional speech are achieved.
It significantly improves the emotional naturalness, emotional accuracy, and person consistency of speech, reduces data construction costs, enhances the scale and consistency of annotation, and strengthens the naturalness of emotional speech and the reliability of emotional judgment.
Smart Images

Figure CN122369424A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of emotional speech synthesis technology, specifically to an emotional speech synthesis method and system based on dialogue context and personal experience. Background Technology
[0002] Emotional Text-to-Speech (Emotional TTS) aims to synthesize natural, expressive speech with a target emotional state based on given text. It is widely used in scenarios such as intelligent customer service, audio content production, educational companionship, digital humans, game interaction, and film dubbing. Current mainstream technologies typically control speech generation through discrete emotion tags, style tags, or short natural language commands, enabling the model to output preset or open-ended emotional lexical expressions such as "happy," "sad," and "serious."
[0003] However, most existing technologies still use single sentences as the basic modeling unit, primarily relying on the current text to be synthesized and a short prompt for emotion control. But genuine human emotional expression is not generated in isolation; it is simultaneously influenced by the speaker's personal experiences, the current dialogue context, the audience, plot development, and paralinguistic features. For example, the same phrase "I'm fine" may represent feigned composure, concealed disappointment, polite comfort, or genuine relief in different characters' backgrounds and dialogue situations. Solutions that rely solely on sentence-level tags or single-sentence prompts often only provide coarse-grained emotional results, failing to express nuanced and logically consistent tone, pauses, and rhythm.
[0004] Meanwhile, existing emotional speech datasets have significant limitations. Most datasets only provide text transcription and sentence-level sentiment labels, lacking higher-level semantic information such as character experiences, dialogue context, open-word sentiment labels, and paralinguistic descriptions. While some dialogue datasets provide historical context, they typically do not explicitly describe the character's past experiences and long-term cognitive baseline, thus still struggling to support models in understanding why a character speaks in a particular way at this particular moment. Under these data conditions, models are more likely to learn surface tone patterns rather than the causal logic behind emotions.
[0005] In recent years, open-vocabulary sentiment-based speech generation and speech generation technologies based on large language models have made progress. However, most existing methods simply concatenate text descriptions directly to the input side, and the model still needs to guess the appropriate emotional expression on its own. In scenarios involving complex character relationships, continuous plots, multi-turn dialogues, and fine-grained emotional changes, problems such as averaged expression, unstable characters, deviations in emotional direction, and inconsistencies between tone and plot can easily occur.
[0006] On the other hand, while large language models excel in text understanding and long-range inference, they struggle to directly interpret acoustic information such as fundamental frequency, energy, pauses, spectral distribution, and harmonic noise ratio in continuous speech. Large audio models, though capable of processing audio, are less adept at combining narrative text and character background to infer the causes of emotions. Existing open-vocabulary emotion control systems typically only specify "what emotion is desired," failing to reliably explain "why this emotion occurs" and "how it should be expressed." Consequently, in complex narrative scenarios, they are prone to exhibiting the correct emotional direction but incorrect expression.
[0007] In summary, existing technologies lack a complete solution that can both transform acoustic information into a structured representation usable by language models and integrate character experiences, dialogue context, and paralinguistic information for automatic annotation and speech generation. This makes it impossible to transform emotion modeling from "sentence-level label-driven" to "context and personal experience-driven," and makes it difficult to achieve highly natural, highly emotionally accurate, highly consistent with characters, and highly relevant to the plot in emotional speech synthesis. Summary of the Invention
[0008] In view of the above-mentioned deficiencies of the prior art, the present invention at least solves the following technical problems: 1. Existing emotional speech synthesis technologies rely solely on single sentences and emotional tags, lacking personal experience and dialogue context modeling, resulting in coarse-grained emotional expression and difficulty in generating speech that fits the plot logic and is consistent with the characters; 2. Existing sentiment data relies on manual annotation, making it difficult to obtain multi-dimensional information such as personal experiences, dialogue context, paralinguistic descriptions, and open-vocabulary sentiment tags at scale; 3. Continuous acoustic features cannot be directly utilized by language models, resulting in acoustic information not being effectively integrated into annotation and inference, thus affecting the accuracy of sentiment judgment.
[0009] To achieve the above objectives, this invention discloses an emotional speech synthesis method based on dialogue context and personal experience, the method comprising the following steps: S1: Obtain the text to be synthesized, and obtain or construct a multi-dimensional sentiment enhancement condition set corresponding to the text to be synthesized; wherein, the multi-dimensional sentiment enhancement condition set includes at least personal experience description and dialogue context description, and may also include at least one of paralinguistic description and open-word sentiment tags; the personal experience description is used to represent the long-term experience, identity background, relationship status or cognitive baseline of the target character or speaker, and the dialogue context description is used to represent the current dialogue scene, historical dialogue, speaking object or plot development information; S2: Input the text to be synthesized and the multi-dimensional emotion enhancement condition set into the semantic generation module according to a preset format, and the semantic generation module generates a semantic representation sequence related to emotion expression; S3: Input the semantic representation sequence into the acoustic reconstruction module, and combine it with at least one of speaker features, reference acoustic features, or role prior information to generate an intermediate acoustic representation; S4: Input the intermediate acoustic representation into the vocode module to generate the target emotional speech.
[0010] Furthermore, the multi-dimensional sentiment enhancement condition set includes personal experience description, dialogue context description, paralinguistic description, and open-word sentiment tags, and the personal experience description, dialogue context description, paralinguistic description, and open-word sentiment tags are all used as input condition fields for the semantic generation module; the multi-dimensional sentiment enhancement condition set is represented as follows:
[0011] in, This is a description of personal experience. For the context description of the dialogue, For secondary language description, For open vocabulary sentiment tags.
[0012] Furthermore, the multi-dimensional emotion enhancement condition set can be obtained through at least one of the following: automatic annotation results, role knowledge base, historical scripts, long dialogue summaries, reference speech analysis results, manual configuration results, user input information, or information provided by external business systems. When any field in the multi-dimensional emotion enhancement condition set is missing, the field can be generated or supplemented through at least one of the following methods: knowledge base retrieval, historical context retrieval, model inference, reference speech analysis, template completion, or default placeholder.
[0013] Furthermore, the multi-dimensional sentiment enhancement condition set can be generated through an automatic annotation process, which includes: acquiring speech data and corresponding text; automatically transcribing, forcing alignment, or matching dialogue between the speech data and the corresponding text to obtain the correspondence between speech segments and text segments; extracting acoustic paralinguistic features from the speech segments and converting them into structured acoustic representations; acquiring at least one text context information related to the text segment, including current text, historical dialogue, scene description, or character background information; and inputting the text context information and the structured acoustic representation into a structured inference model to generate at least one annotation information among personal experience description, dialogue context description, paralinguistic description, and open-vocabulary sentiment labels.
[0014] Furthermore, the acoustic sub-language features include at least one of fundamental frequency, speech rate, number of pauses, pause duration, spectral centroid, root mean square energy, MFCC variance, and harmonic signal-to-noise ratio; the structured acoustic representation includes a discrete token sequence, which is obtained by performing at least one of downsampling, normalization, intervalization, cluster quantization, vector quantization, or learning-based discretization on the continuous acoustic features.
[0015] Furthermore, for any continuous acoustic feature z, the discretized integer token is obtained using the following interval quantization method:
[0016] in, It is a continuous acoustic feature. This is the lower bound for acoustic feature statistics. To quantize the step size, Based on the interval number, Expand the interval number for outliers. It is a discretized integer token.
[0017] This invention also discloses an emotional speech synthesis system based on dialogue context and personal experience, which includes a condition acquisition and construction module, a semantic generation module, an acoustic reconstruction module, and a vocode module: The condition acquisition and construction module is used to acquire or construct a multi-dimensional sentiment enhancement condition set corresponding to the text to be synthesized. The multi-dimensional sentiment enhancement condition set includes at least personal experience description and dialogue context description, and may also include at least one of paralinguistic description and open-vocabulary sentiment labels. The semantic generation module is used to generate a semantic representation sequence based on the text to be synthesized and the multi-dimensional sentiment enhancement condition set; The acoustic reconstruction module is used to generate intermediate acoustic representations based on the semantic representation sequence; The vocode module is used to convert the intermediate acoustic representation into the target emotional speech.
[0018] Furthermore, the condition acquisition and construction module includes at least one of a data acquisition module, an alignment module, an acoustic feature extraction and quantization module, and an automatic annotation module; the data acquisition module is used to acquire speech data and corresponding text; the alignment module is used to automatically transcribe, force alignment, or perform dialogue matching on the speech data and corresponding text; the acoustic feature extraction and quantization module is used to extract acoustic paralinguistic features and generate structured acoustic representations; the automatic annotation module is used to generate annotation information based on text context information and structured acoustic representations.
[0019] Furthermore, the system also includes a data filtering module for filtering out labeled samples with alignment errors, context conflicts, or low quality.
[0020] This invention achieves the following beneficial technical effects by introducing a multi-dimensional emotion enhancement condition set, an automated annotation pipeline, and an acoustic feature quantization encoding method: 1. By integrating personal experiences, dialogue context, paralinguistic information, and open emotional tags into the generation process, the emotional naturalness, emotional accuracy, character consistency, and plot fit of the voice are greatly improved; 2. By using an automated annotation pipeline to generate multidimensional information in batches, the cost of data construction is significantly reduced, the scalability and consistency of annotation are improved, and large-scale high-quality datasets can be supported. 3. Quantizing continuous acoustic features into structured tokens effectively bridges the modal gap between acoustic information and language models, improving the quality of automatic annotation and the reliability of sentiment judgment.
[0021] Thanks to the overall design described above, this invention can be stably implemented in real products, possessing strong engineering practicality and industrial promotion value. It can be widely used in various application scenarios such as audiobooks, radio dramas, film and television dubbing, digital humans, educational companionship, intelligent customer service, in-vehicle interaction, and game character voices. Compared with the traditional method that relies solely on sentence-level sentiment tags, this invention simultaneously combines personal experience, dialogue context, paralinguistic descriptions, and open-vocabulary sentiment tags during the speech generation process, making the output speech more consistent with character settings, plot development, and interaction logic. It is especially suitable for the high-expressive speech needs of long dramas, multi-turn dialogues, and complex character relationships. From a technical implementation perspective, this invention simultaneously possesses automated data construction and end-to-end speech synthesis capabilities, and can be compatible with and integrated with existing speech synthesis systems. The front end continuously expands high-quality corpora through automatic speech recognition, forced alignment, dialogue extraction, acoustic feature extraction, and quantization encoding. The back end supports offline training and online inference through enhanced condition set construction, semantic generation, and acoustic reconstruction. It can be easily modularly integrated into existing speech platforms, content production platforms, and digital human systems. From the perspectives of performance and industrial value, this invention can significantly reduce the cost of acquiring high-quality emotional speech data, improve the consistency and scalability of multi-dimensional annotation, and further enhance the naturalness, emotional accuracy, and role consistency of synthesized speech. Related experimental results show that after introducing personal experience and contextual information, the subjective score for the continuation scene increased from 74.4 to 76.2, and the emotional naturalness score reached 63.9 with full configuration, which is superior to the 58.8 of the comparison system CosyVoice3. The overall performance is outstanding, demonstrating good engineering applicability and patent commercialization prospects. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of the emotional speech synthesis method based on context and personal experience according to the present invention; Figure 2 This is a schematic diagram of context-aware emotion-based speech samples from the present invention. Figure 3 This is a schematic diagram of the automatic labeling pipeline structure of the present invention; Figure 4 This is a schematic diagram comparing the inputs of the present invention with traditional emotional speech synthesis schemes; Figure 5 This is a schematic diagram of the framework of the emotional speech synthesis system based on context and personal experience according to the present invention. Detailed Implementation
[0023] The following description, with reference to the accompanying drawings, illustrates several preferred embodiments of the present invention to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0024] In the accompanying drawings, components with the same structure are indicated by the same numerical designation, and components with similar structures or functions are indicated by similar numerical designations. The dimensions and thicknesses of each component shown in the drawings are arbitrary, and the present invention does not limit the dimensions and thicknesses of each component. To make the illustrations clearer, the thickness of some components has been appropriately exaggerated in the drawings.
[0025] This invention provides a method for synthesizing emotional speech based on context and personal experience, such as... Figure 1 As shown, the method includes the following steps: S1: Obtain the text to be synthesized, and obtain or construct a multi-dimensional sentiment enhancement condition set corresponding to the text to be synthesized; The multi-dimensional emotion enhancement condition set includes at least personal experience description and dialogue context description, and may also include at least one of paralinguistic description and open-word emotion label; the personal experience description is used to represent the long-term experience, identity background, relationship status or cognitive baseline of the target character or speaker, and the dialogue context description is used to represent the current dialogue scene, historical dialogue, speaking object or plot development information; S2: Input the text to be synthesized and the multi-dimensional emotion enhancement condition set into the semantic generation module according to a preset format to generate a semantic representation sequence related to emotion expression; S3: Input the semantic representation sequence into the acoustic reconstruction module, and combine it with at least one of speaker features, reference acoustic features, or role prior information to generate an intermediate acoustic representation; S4: Input the intermediate acoustic representation into the vocode module to generate the target emotional speech.
[0026] In its overall implementation, this invention can either directly construct a multi-dimensional emotion enhancement condition set using information from a character knowledge base, historical scripts, long dialogue summaries, reference speech analysis results, manual configuration results, user input information, or external business systems, or generate the same multi-dimensional emotion enhancement condition set from speech data and corresponding text through an automatic annotation process. The automatic annotation process is a preferred data construction implementation method and is not a necessary step in performing emotional speech synthesis.
[0027] The implementation of this invention does not depend on specific automatic annotation models, forced alignment tools, or acoustic feature quantization algorithms; as long as a multi-dimensional emotion enhancement condition set containing personal experience descriptions and dialogue context descriptions can be obtained or constructed, emotional speech synthesis can be achieved according to the above speech generation process.
[0028] In one optional data construction implementation, the system collects audiobooks, radio dramas, and various audio data containing character dialogues from multiple speakers, and matches them with corresponding plot texts. Initial transcription can be obtained through automatic speech recognition, and then character-level or word-level timestamps are obtained using a forced alignment tool, thereby establishing a precise correspondence between audio segments and text. For dialogue content in the original text, it can be extracted based on quotation marks, dialogue markers, or character identifiers, and the robustness of speech-text alignment can be improved by combining sequential buffer matching, similarity matching, or pinyin-level matching methods.
[0029] Subsequently, the system inputs the current text, historical dialogue, scene description, character background information, and quantified acoustic tokens into the structured inference model, sequentially performing context analysis, multimodal fusion, annotation generation, and consistency verification operations, ultimately outputting a personal experience description, dialogue context description, paralinguistic description, and open-vocabulary sentiment tags. For example... Figure 2 As shown, the context-aware emotional speech samples generated by this invention contain complete information such as text transcription content, emotion tags, personal experiences, dialogue context, and paralinguistic features, providing sufficient training signals for the model. After removing alignment errors, contextual conflicts, and low-quality samples by the data filtering module, a standardized context-aware emotional speech dataset can be obtained. If necessary, the data quality can be further improved through manual sampling and correction. In this embodiment, the above scheme can construct a large-scale dataset of approximately 730,000 samples, 1335 hours, and covering 3359 open-vocabulary emotion tag entries. Manual verification results show that the average score of the automatically labeled paralinguistic descriptions is close to 4.0 / 5, demonstrating high accuracy and consistency.
[0030] During the automatic annotation stage, the system extracts multiple paralinguistic features from each speech segment, including fundamental frequency, speech rate, number and duration of pauses, spectral centroid, root mean square energy, MFCC variance, and harmonic noise ratio. For time-series features, downsampling is first performed, followed by quantization intervals based on statistical ranges. Continuous values are mapped to discrete integer tokens, and outliers exceeding the normal range are represented by extended intervals. The automatic annotation pipeline of this invention is as follows: Figure 3 As shown, the entire system works in sequence through the collaborative efforts of units such as automatic transcription, forced alignment, dialogue extraction, acoustic feature extraction, quantization encoding, and structured inference to automatically generate multidimensional annotation information.
[0031] In terms of sample organization, each training sample contains at least the target text, character background description, dialogue context description, quantized acoustic token, paralinguistic description, open-vocabulary sentiment tags, reference speech, and target semantic representation. The system can create indexes according to character, scene, sentiment category, or speech source to facilitate subsequent training sampling, data cleaning, and continuous expansion of the corpus.
[0032] In a preferred embodiment, the multi-dimensional sentiment enhancement condition set includes personal experience descriptions, dialogue context descriptions, paralinguistic descriptions, and open-vocabulary sentiment tags, and all four descriptions serve as input condition fields for the semantic generation module, which can be formally expressed as follows:
[0033] in, This indicates a description of personal experience. This indicates a description of the dialogue context. Indicates a secondary language description, This indicates an open-ended sentiment label. The above conditions and text... These factors together constitute the input to the semantic generation module, enabling the model to simultaneously combine the character's long-term state, the current dialogue environment, and the target's speaking style for generation. For example... Figure 4 As shown, compared with traditional solutions that rely solely on text or single emotion tags, this invention achieves more granular and plot-aligned emotion control through multi-dimensional conditional input, effectively avoiding the problems of homogenized emotion expression and logical deviation.
[0034] In the acoustic feature quantization stage, for any continuous acoustic feature z, interval quantization is used to obtain the discretized integer token:
[0035] in, It is a continuous acoustic feature. This is the lower bound for acoustic feature statistics. To quantize the step size, Based on the interval number, Expand the interval number for outliers. Discretized integer tokens. This method transforms continuous acoustic signals into structured input sequences that can be directly processed by language models, effectively bridging the modal gap between acoustic information and text models. This quantization scheme achieved significant performance improvements in ablation experiments on the IEMOCAP dataset.
[0036] In the emotional speech generation stage, the system first constructs a multi-dimensional emotion enhancement condition set based on the text to be synthesized. The condition set can be derived from automatic annotation results, role knowledge base, historical scripts, long dialogue summaries, reference speech analysis modules, or manual configuration. When some fields are missing, downgraded reasoning can be achieved through retrieval, completion, or default placeholders to ensure that the system can run stably in different application scenarios.
[0037] After constructing the condition set, the system concatenates the text and the enhanced condition set according to a preset format and inputs them into the semantic generation module, where the model predicts the discrete semantic representation. Unlike traditional solutions that rely solely on text or sentiment tags, this embodiment simultaneously integrates the character's historical state, the current dialogue environment, and the target's speaking style during the semantic generation stage, enabling the generation of prosodic patterns that better align with the character's character and the scene's logic. The semantic generation module employs an autoregressive generation method, and its training objective can be represented as:
[0038] in, For the target semantic representation sequence, The text to be synthesized. For a multi-dimensional emotion enhancement condition set, For parameters of the semantic generation module, For the semantic unit at time t, The sequence of semantic units generated before time t. During training, the model learns the constraints of character background, context, paralinguistics, and sentiment tags on emotional expression, making the output more consistent with the causes of emotions and the logic of the plot.
[0039] In the training implementation, the system uses the aforementioned automatically labeled dataset to supervise the training of the semantic generation module. The training samples include at least text, an augmented condition set, and a target semantic representation. The model learns the influence of multidimensional conditions on emotion and prosody under the constraints of supervised loss.
[0040] To further enhance character consistency, the system can incorporate speaker embeddings, reference audio features, or role-level prior information during the training or inference phases. During the inference phase, if only open-vocabulary sentiment labels are provided, the model can output basic sentiment control results; if personal experiences and context are input simultaneously, more nuanced and logically consistent emotional expressions can be generated; if reference speech exists but lacks paralinguistic descriptions, the system can first analyze the reference speech, generate candidate paralinguistic descriptions, and then perform speech synthesis.
[0041] After obtaining the semantic representation, the system inputs the semantic representation, speaker features, and reference acoustic features into a conditional acoustic decoder to generate intermediate acoustic representations such as Mel spectrum, which are then reconstructed into the final speech waveform by a vocoder. Through this two-stage processing approach, the system can achieve precise emotion control by relying on enhanced conditional sets while maintaining speaker timbre consistency and high speech fidelity. From the generation process perspective, the system first determines the target tone and prosodic trend based on the semantic representation, then gradually reconstructs acoustic details by combining reference acoustic features and speaker features, ultimately obtaining synthesized speech that matches both the character's identity and the current emotional context. In the implementation using a conditional flow matching acoustic decoder, the acoustic reconstruction process can be represented as:
[0042] in, for Acoustic state at any given moment For semantic representation sequence, For reference acoustic characteristics, Speaker characteristics This is a time-dependent vector field. This process, while maintaining the speaker's timbre stability, progressively renders the semantic representation into the target acoustic result. The acoustic decoder can employ a conditional stream matching decoder, a diffusion decoder, or a variational decoder, while the vocoder can use a generative adversarial network-based vocoder, a diffusion vocoder, a streaming vocoder, or other neural vocoders. From a generation perspective, the system first determines the target tone and prosodic trend based on the semantic representation, then combines reference acoustic features and speaker features to progressively reconstruct acoustic details, ultimately obtaining synthesized speech that matches both the character's identity and the current emotional context.
[0043] To implement the aforementioned emotional speech synthesis method, this embodiment provides a corresponding emotional speech synthesis system based on context and personal experience, such as... Figure 5 As shown, the system includes a condition acquisition and construction module, a semantic generation module, an acoustic reconstruction module, and a vocode module: The condition acquisition and construction module is used to acquire or construct a multi-dimensional sentiment enhancement condition set corresponding to the text to be synthesized, and can generate or complete condition fields based on at least one of the following: automatic annotation results, role knowledge base, historical scripts, long dialogue summaries, reference speech analysis results, manual configuration results, user input information, or information provided by external business systems; the semantic generation module is used to generate a semantic representation sequence based on the text to be synthesized and the multi-dimensional sentiment enhancement condition set; the acoustic reconstruction module is used to generate an intermediate acoustic representation based on the semantic representation sequence; and the vocode module is used to convert the intermediate acoustic representation into the target sentiment speech.
[0044] In an optional extended implementation, the condition acquisition and construction module includes a data acquisition module, an alignment module, an acoustic feature extraction and quantization module, and an automatic annotation module, used for offline construction of training data or online completion of missing fields; the data acquisition module is used to acquire speech data and corresponding text; the alignment module is used to automatically transcribe, force alignment, or perform dialogue matching on the speech data and corresponding text; the acoustic feature extraction and quantization module is used to extract acoustic paralinguistic features and generate structured acoustic representations; the automatic annotation module is used to generate annotation information based on text context information and structured acoustic representations. These extended modules can be integrated and deployed with the core emotion-based speech synthesis system, or they can be deployed as independent data construction services.
[0045] The system also includes a data filtering module to remove labeled samples with alignment errors, contextual conflicts, or low quality, thereby further improving the quality of the dataset.
[0046] The system can deploy and run the automatic annotation module, condition construction module, and semantic generation module as a series of services. The front end continuously extracts character background, context, and paralinguistic descriptions from new corpora through the automatic annotation process, while the back end completes speech synthesis through condition construction, semantic generation, and acoustic reconstruction. It supports both batch construction of offline training data and online inference and dynamic completion of missing fields, effectively expanding the system's applicable scenarios and deployment flexibility. The various modules of the system work collaboratively, enabling both batch construction of offline training data and dynamic completion of missing fields during online inference, further improving the system's applicability and operational stability.
[0047] In practical applications, the aforementioned methods and systems enable the same text to express differentiated emotions based on the character's personal experiences and the context of the dialogue. This effectively avoids problems such as uniform expression, character instability, emotional deviation, and mismatch between tone and plot found in traditional solutions. For example, if the text to be synthesized is "I'm really happy for you," and the character has a long-standing competitive relationship with the other party and is currently in a forced congratulatory scenario, the system will not output a straightforward expression of happiness. Instead, it will generate a restrained, suppressed, and feigned congratulatory expression, with pauses, speech rate, and emotional intensity that better align with the plot logic. When the text content is relatively bland, but the character has just experienced a setback and needs to maintain a calm demeanor, the system can use enhanced condition sets to infer an expression that is outwardly calm but inwardly suppressed, avoiding the generation of neutral speech that is out of sync with the plot.
[0048] Based on the above method and system, this embodiment can be deployed in an electronic device including a processor and a memory. The memory stores a computer program that can run on the processor. When the processor executes the program, it can realize all steps such as data construction, automatic annotation, condition construction, semantic generation, acoustic reconstruction, and speech output. Accordingly, the present invention also includes a computer-readable storage medium storing a computer program thereon. When the program is executed by the processor, it implements the above-described emotional speech synthesis method.
[0049] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0050] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A method for synthesizing emotional speech based on dialogue context and personal experience, characterized in that, The method includes the following steps: S1: Obtain the text to be synthesized, and obtain or construct a multi-dimensional sentiment enhancement condition set corresponding to the text to be synthesized; The multi-dimensional emotion enhancement condition set includes at least personal experience description and dialogue context description, and may also include at least one of paralinguistic description and open-word emotion label; the personal experience description is used to represent the long-term experience, identity background, relationship status or cognitive baseline of the target character or speaker, and the dialogue context description is used to represent the current dialogue scene, historical dialogue, speaking object or plot development information; S2: Input the text to be synthesized and the multi-dimensional emotion enhancement condition set into the semantic generation module according to a preset format, and the semantic generation module generates a semantic representation sequence related to emotion expression; S3: Input the semantic representation sequence into the acoustic reconstruction module, and combine it with at least one of speaker features, reference acoustic features, or role prior information to generate an intermediate acoustic representation; S4: Input the intermediate acoustic representation into the vocode module to generate the target emotional speech.
2. The emotional speech synthesis method based on dialogue context and personal experience according to claim 1, characterized in that, The multi-dimensional sentiment enhancement condition set includes personal experience description, dialogue context description, paralinguistic description, and open-vocabulary sentiment tags, and the personal experience description, dialogue context description, paralinguistic description, and open-vocabulary sentiment tags are all used as input condition fields for the semantic generation module; the multi-dimensional sentiment enhancement condition set is represented as follows: in, This is a description of personal experience. For the context description of the dialogue, For secondary language description, For open vocabulary sentiment tags.
3. The emotional speech synthesis method based on dialogue context and personal experience according to claim 1 or 2, characterized in that, The multi-dimensional emotion enhancement condition set is obtained through at least one of the following sources: automatic annotation results, role knowledge base, historical scripts, long dialogue summaries, reference speech analysis results, manual configuration results, user input information, or information provided by external business systems.
4. The emotional speech synthesis method based on dialogue context and personal experience according to any one of claims 1 to 3, characterized in that, When any field in the multi-dimensional emotion enhancement condition set is missing, the field is generated or completed by at least one of the following methods: knowledge base retrieval, historical context retrieval, model inference, reference speech analysis, template completion, or default placeholder.
5. The emotional speech synthesis method based on dialogue context and personal experience according to any one of claims 1 to 4, characterized in that, When the multi-dimensional emotion enhancement condition set is obtained through automatic annotation results, the automatic annotation results are generated through the following steps: Acquire speech data and corresponding text, and perform automatic transcription, forced alignment or dialogue matching on the speech data and corresponding text to obtain the correspondence between speech segments and text segments; Acoustic paralinguistic features are extracted from the speech segments and converted into structured acoustic representations; Obtain at least one text context information related to the text fragment, including current text, historical dialogue, scene narration, or character background information; The text context information and the structured acoustic representation are input into the structured inference model to generate at least one of the following annotation information: personal experience description, dialogue context description, paralinguistic description, and open-vocabulary sentiment label.
6. The emotional speech synthesis method based on dialogue context and personal experience according to claim 5, characterized in that, The acoustic sub-language features include at least one of fundamental frequency, speech rate, number of pauses, pause duration, spectral centroid, root mean square energy, MFCC variance, and harmonic signal-to-noise ratio; the structured acoustic representation includes a discrete token sequence, which is obtained by performing at least one of downsampling, normalization, intervalization, cluster quantization, vector quantization, or learning-based discretization on the continuous acoustic features.
7. The emotional speech synthesis method based on dialogue context and personal experience according to claim 6, characterized in that, For any continuous acoustic feature z, the discretized integer token is obtained using the following interval quantization method: in, It is a continuous acoustic feature. This is the lower bound for acoustic feature statistics. To quantize the step size, Based on the interval number, Expand the interval number for outliers. It is a discretized integer token.
8. An emotion-based speech synthesis system based on dialogue context and personal experience, characterized in that, It includes a condition acquisition and construction module, a semantic generation module, an acoustic reconstruction module, and a vocode module: The condition acquisition and construction module is used to acquire or construct a multi-dimensional sentiment enhancement condition set corresponding to the text to be synthesized. The multi-dimensional sentiment enhancement condition set includes at least personal experience description and dialogue context description, and may also include at least one of paralinguistic description and open-vocabulary sentiment labels. The semantic generation module is used to generate a semantic representation sequence based on the text to be synthesized and the multi-dimensional sentiment enhancement condition set; The acoustic reconstruction module is used to generate intermediate acoustic representations based on the semantic representation sequence; The vocode module is used to convert the intermediate acoustic representation into the target emotional speech.
9. The emotional speech synthesis system based on dialogue context and personal experience according to claim 8, characterized in that, The condition acquisition and construction module includes at least one of the following: a data acquisition module, an alignment module, an acoustic feature extraction and quantization module, and an automatic annotation module. The data acquisition module is used to acquire voice data and corresponding text; The alignment module is used to automatically transcribe, force alignment, or match dialogue between voice data and corresponding text. The acoustic feature extraction and quantization module is used to extract acoustic paralinguistic features and generate structured acoustic representations; The automatic annotation module is used to generate annotation information based on text context information and structured acoustic representation.
10. The emotional speech synthesis system based on dialogue context and personal experience according to claim 8, characterized in that, The system also includes a data filtering module, which is used to filter out samples with alignment errors, context conflicts, or low quality annotations.