Dialogue voice generation method and system based on emotion perception adapter and large model reasoning

Voice emotional features are extracted through the speech encoder and the temporal and hierarchical attention network, and combined with the emotion perception coding module to align with the large language model, and used some low-rank adaptive networks to generate emotionally consistent voice, solving the problem of mismatch between speech and emotional context and improving the quality of dialogue speech generation.

CN120260539AInactive Publication Date: 2025-07-04STATE GRID ZHEJIANG ELECTRIC POWER CO MARKETING SERVICE CENT

Patent Information

Application Number
CN202510731766.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing dialogue speech synthesis technology is difficult to effectively integrate the emotional information in the speech and text, resulting in the mismatch between the generated speech and the emotional context, lacking in-depth understanding of the evolution of multiple rounds of dialogue emotions, and the generation speech is inconsistent with the actual context.

Method used

Voice encoder is used to extract speech emotional characteristics with the time and hierarchical attention network, and aligned with the large language model through the emotion-perceptual coding module, and combined with the adapter of some low-rank adaptive networks to generate emotionally consistent target speech.

Benefits of technology

Effectively narrow the gap between emotions and text, generate target voice with emotional consistency, and improve the reasoning ability and the quality of speech generation in multiple rounds of dialogue emotional context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260539A_ABST
    Figure CN120260539A_ABST
Patent Text Reader

Abstract

The invention discloses a dialogue voice generation method and system based on an emotion perception adapter and large model reasoning. The dialogue voice generation method adopted by the invention comprises the following steps: extracting voice emotion characteristics in original dialogue voice data by using a voice encoder and a time and hierarchical attention network; aligning the speech emotion features with text features of the large language model through an emotion perception coding module based on a query converter network, and generating emotion embedding compatible with the large language model; generating text embedding from statements of the input dialogue text by using a word segmentation device of a large language model; performing emotion embedding and text embedding by adopting an emotion adapter and a text adapter based on a partial low-rank adaptive network, and reasoning a text reply and a reply emotion state of the dialogue; and in combination with the text reply and the reply emotional state, generating a target voice conforming to an emotional context by using a voice generation model. According to the invention, the difference between the emotion and the text is effectively reduced, and the target voice with emotion consistency is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular, to a dialogue speech generation method and system based on emotion-aware adapter and large model reasoning. Background Art

[0002] Emotion context-aware speech generation aims to generate speech output with natural emotional expressions according to the context and emotional information in the dialogue, and is widely used in scenarios such as intelligent voice assistants, virtual customer service, and voice dialogue systems. It is one of the important technologies to improve the human-computer interaction experience. In actual conversations, speech expression not only depends on the text content, but is also significantly affected by the emotional context. For the same conversation content, due to different emotional states of the speakers, there are also differences in the intonation, rhythm, and intensity of the speech.

[0003] However, traditional text-to-speech synthesis technology mainly generates speech based on text, ignoring the context and emotional information in the dialogue, and it is difficult to generate natural speech that conforms to the context. To solve this problem, dialogue speech synthesis technology infers the emotional context through dialogue history and generates matching speech. However, existing dialogue speech synthesis methods mainly rely on text information and ignore the emotional features in the speech, resulting in incomplete understanding of the emotional context and inaccurate emotional expressions of the generated speech. Some research has introduced methods such as multi-modal modeling and contrast learning, attempting to combine speech and text information to improve the emotional context reasoning ability, but there are still the following deficiencies: First, the emotion perception is insufficient, and it is difficult to accurately capture the complex emotional changes in the dialogue; second, the emotional context reasoning ability is limited, lacking a deep understanding of the emotional evolution of multi-round conversations, resulting in the generated speech not matching the actual context.

[0004] Therefore, how to effectively integrate the emotional information in speech and text, enhance the model's reasoning ability for the emotional context of multi-round conversations, and improve the quality of emotional speech generation in the case of insufficient data has become an urgent problem to be solved in the current field of emotion context-aware speech generation. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the problem of unnatural dialogue caused by the mismatch between the speech generated by the existing method and the emotional context, and to provide a dialogue speech generation method and system based on emotion-aware adapter and large model reasoning, so as to effectively narrow the gap between emotion and text and generate target speech with emotional consistency.

[0006] To this end, the present invention adopts the following technical solutions: A dialogue speech generation method based on emotion-aware adapter and large model reasoning, which includes the steps of: S1, using a speech encoder and a temporal and hierarchical attention network to extract the speech emotional features in the original dialogue speech data; S2. Align the speech emotion features with the text features of the large language model through an emotion-aware encoding module based on a query transformer network to generate emotion embeddings compatible with the large language model; S3. Use the tokenizer of the large language model to generate text embeddings for the sentences of the input dialogue text; S4. Adopt an emotion adapter and a text adapter based on a partial low-rank adaptation network for the emotion embeddings and text embeddings to infer the text response and response emotion state of the dialogue; S5. Combine the text response and the response emotion state and use a speech generation model to generate a target voice that conforms to the emotional context.

[0007] The present invention first uses a speech encoder and a temporal and hierarchical attention network to extract emotion features in the dialogue speech, aligns the emotion features with the text features through an emotion-aware encoding module to generate emotion embeddings compatible with the large language model; then, adopts an adapter based on a partial low-rank adaptation network to adapt the large language model, and fine-tunes the emotion embeddings and text embeddings respectively, effectively narrowing the gap between emotion and text, and generating a target voice with emotional consistency. The present invention introduces an emotion-aware encoder module to align speech emotion and text information, and combines an adapter based on a partial low-rank adaptation network to enhance the emotion context reasoning ability of the large language model under multi-modal input, and enhance the emotional consistency of speech generation.

[0008] Further, the specific content of step S1 includes: S11. Use a whisper speech encoder to extract the original dialogue speech data as a feature representation; S12. Use a temporal and hierarchical attention network to encode the feature representation into speech emotion features.

[0009] Further, the specific content of step S2 includes: S21. Initialize a set of learnable query vectors in the query transformer network, and each query vector is designed to capture a representation related to a specific emotion dimension; S22. Use the cross-attention network in the query transformer network to interact the query vectors with the extracted speech emotion features, so as to align the speech emotion features with the text features of the large language model; S23. To make the output speech emotion features acceptable to the large language model, project the cross-attention output through a linear layer to the input embedding dimension required by the large language model to generate emotion embeddings.

[0010] Further, the specific content of step S3 includes: Given the dialogue text of the kth sentence, use the tokenizer of the large language model to generate text embeddings.

[0011] Further, the specific content of step S4 includes: S41. Use a sentiment adapter based on a partial low-rank adaptation network to perform linear transformation and low-rank adaptation on the sentiment embedding to generate an adapted sentiment embedding. S42. Apply a text adapter based on a partial low-rank adaptation network to the text embedding for low-rank adaptation to finely control the sensitivity of the text semantic representation to the sentiment state, and generate an adapted text embedding. S43. Concatenate the adapted sentiment embedding and the text embedding as the input to the large language model. S44. The large language model infers the text response and response sentiment state of the conversation.

[0012] Further, the specific content of step S5 includes: Using FastSpeech2 as the backbone model for speech synthesis, and combining the text response and response sentiment state inferred by the large language model to generate a target speech consistent with the sentiment context.

[0013] The present invention also provides a dialogue speech generation system based on sentiment-aware adapter and large model inference for implementing the above dialogue speech generation method, which includes: Speech sentiment feature extraction unit: Use a speech encoder and a temporal and hierarchical attention network to extract speech sentiment features from the original dialogue speech data. Sentiment embedding generation unit: Align the speech sentiment features with the text features of the large language model through a sentiment-aware encoding module based on a query transformer network to generate a sentiment embedding compatible with the large language model. Dialogue text embedding generation unit: Use the tokenizer of the large language model to generate text embeddings for the sentences of the input dialogue text. Model inference unit: Use a sentiment adapter and a text adapter based on a partial low-rank adaptation network to infer the text response and response sentiment state of the conversation for the sentiment embedding and the text embedding. Target speech generation unit: Combine the text response and response sentiment state and use a speech generation model to generate a target speech that conforms to the sentiment context.

[0014] The beneficial effects of the present invention are as follows: The present invention introduces a sentiment-aware encoding module, a sentiment adapter, and a text adapter based on a partial low-rank adaptation network to align the sentiment information of the dialogue speech to the input of the large language model, and combines the powerful generation ability of the large language model to provide a dialogue speech generation method and system based on sentiment-aware adapter and large model inference, which overcomes the problems of insufficient sentiment perception and limited sentiment context inference ability in the existing methods, effectively reduces the gap between sentiment and text, and generates a target speech with sentiment consistency. Description of the Drawings

[0015] Figure 1It is a flowchart of a method for generating dialogue speech based on an emotion-aware adapter and large model reasoning according to the present invention; Figure 2 It is a schematic diagram of the process of a method for generating dialogue speech based on an emotion-aware adapter and large model reasoning according to the present invention; Figure 3 It is a composition diagram of a system for generating dialogue speech based on an emotion-aware adapter and large model reasoning according to the present invention; Figure 4 It is a composition diagram of the emotion embedding generation unit according to the present invention; Figure 5 It is a composition diagram of the model reasoning unit according to the present invention. Specific Embodiments

[0016] The present invention will be further described and explained below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0017] Embodiment 1 As shown in FIG. 1 and Figure 2 as shown, this embodiment provides a method for generating dialogue speech based on an emotion-aware adapter and large model reasoning, and the steps are as follows: S1. Use a voice encoder and a temporal and hierarchical attention network to extract voice emotion features from the original dialogue speech data; S2. Align the voice emotion features with the text features of the large language model through an emotion-aware encoding module based on a query transformer network to generate emotion embeddings compatible with the large language model; S3. Use the tokenizer of the large language model to generate text embeddings for the sentences of the input dialogue text; S4. Adopt an emotion adapter and a text adapter based on a partial low-rank adaptation network to infer the text response and response emotion state of the dialogue for the emotion embeddings and text embeddings; S5. Combine the text response and response emotion state, and use a voice generation model to generate target voices that conform to the emotional context.

[0018] Specifically, the content of step S1 is as follows: S11. Use a whisper voice encoder to extract the original dialogue speech data as a feature representation : ; S12. Use a temporal and hierarchical attention network to encode the feature representation into voice emotion features : , wherein, represents a temporal and hierarchical attention network.

[0019] Specifically, the content of step S2 is as follows: S21. Initialize a set of learnable query vectors in the query transformer network , each query vector is designed to capture representations related to specific emotional dimensions and is used to interact with speech emotional features, ultimately achieving alignment with text semantic representations.

[0020] S22. Use the cross-attention network in the query transformer network (Q-former) to enable the query vectors to interact with the extracted speech emotional features, thereby aligning the speech emotional features with the text features of the large language model: , wherein, respectively represent the projection matrices of query, key, and value in the attention mechanism, is the dimension of the key-value set; is the speech emotional feature, represents the cross-attention output, which is used to capture the alignment information between the speech emotional feature and the query vector; represents matrix inversion.

[0021] S23. To enable the output speech emotional feature to be accepted by the large language model, project the cross-attention output onto the input embedding dimension required by the large language model through a linear layer to generate an emotion embedding: , wherein, represents the generated emotion embedding, represents a trainable linear transformation matrix.

[0022] Specifically, the content of step S3 is as follows: Given the dialogue text of the k-th utterance , use the tokenizer of the large language model to generate a text embedding : , In the formula, represents the tokenizer.

[0023] Specifically, the content of step S4 is as follows: S41. Use the emotion adapter based on the partial low-rank adaptation network to perform linear transformation and low-rank adaptation on the emotion embedding to generate an adapted emotion embedding :

[0024] wherein, is the low-rank adaptation weight matrix of the sentiment embedding, is the update obtained by the low-rank adaptation of the sentiment embedding.

[0025] S42. Apply a text adapter based on a partial low-rank adaptation network to the text embedding for low-rank adaptation to finely control the sensitivity of the text semantic representation to the sentiment state, and generate the adapted text embedding :

[0026] Among them, represents the low-rank adaptation weight matrix of the text embedding; is the update obtained by the low-rank adaptation of the text embedding; By constraining the gradient update of the fully connected weights in the low-rank subspace, the low-rank adaptation significantly reduces the parameter complexity and improves the multi-modal fusion effect.

[0027] S43. Concatenate the adapted sentiment embedding and text embedding as the input of the large language model for joint inference:

[0028] Among them, represents the concatenation result; represents the concatenation operation in the token dimension, and this concatenation represents fusing the sentiment state and language content of the current round of conversation to form the joint inference input of the large language model; S44. The large language model infers the text reply and reply sentiment state of the conversation:

[0029] Among them, is the context content, is the sentiment category of the target statement, is the sentiment intensity of the target statement, is the generated k-th round text reply, represents the large language model, represents the probability function.

[0030] Specifically, the content of step S5 is as follows: Use FastSpeech2 as the backbone model for speech synthesis, and combine the text reply and reply sentiment state inferred by the large language model to generate the target speech consistent with the sentiment context: .

[0031] Next, the above dialogue speech generation method based on sentiment-aware adapters and large model inference is applied as follows.

[0032] Six speech datasets are used for model training and evaluation, namely DailyTalk, CREMA-D, EmoV-DB, IEMOCAP, MEAD, and TESS, with a total of 80.6 hours of speech, covering seven emotion categories (happy, sad, surprised, angry, fearful, disgusted, and neutral) and three intensity levels (weak, medium, and strong). In the present invention, three emotion intensity labels from ECSS are added to the DailyTalk dataset to make up for the lack of intensity information in the original. In the text-emotion alignment and emotion inference stages, all audio is resampled to 16 kHz. In the speech generation stage, the audio generates an 80-dimensional Mel spectrogram through short-time Fourier transform (STFT) and is downsampled to 22.05 kHz for speech synthesis. The data is divided into a training set, a validation set, and a test set in the ratio of 8:1:1.

[0033] In the present invention, Whisper Large v3 is used for speech recognition, and Vicuna-7B is used as the large language model, which is fine-tuned based on LLaMA. In Q-former, 25 trainable query tokens (with a dimension of 768) are used, and multiple adapters (PLoRA) are injected into the query and key-value projection layers of the LLaMA self-attention layer, with the rank set to 8 and the scaling factor set to 4.0. The AdamW optimizer is used in the training process, and the parameters are set as = 0.9, = 0.999, weight decay 0.05, and cosine learning rate decay is adopted (peak 3×10 -5 , minimum 1×10 -5 , with 3k steps of linear warm-up). The training is carried out for 180k steps, using 4 NVIDIA RTX A6000 GPUs. In the training stage, only the parameters of the Q-former and PLoRA modules are updated.

[0034] The present invention is compared with other emotion speech synthesis (CSS) frameworks, including ECSS and GRU-based methods. In addition, the present invention is also compared with the speech synthesis backbone model FastSpeech2 without context modeling, where GT represents the ground truth.

[0035] The present invention comprehensively evaluates the model performance using subjective and objective indicators. The subjective evaluation includes N-DMOS (speech naturalness in the dialogue context) and E-DMOS (consistency between emotional expression and the emotional context of the dialogue), which are scored by 20 volunteers. The objective indicators include the emotion classification accuracy (ECA), which evaluates the accuracy of emotional expression through emotion2vec plus large, Whisper and wav2vec 2.0 are respectively used to calculate the word error rate (WER), and the speech duration prediction is evaluated through the mean absolute duration difference (DDUR). The emotional context reasoning ability is evaluated by Macro-F1.

[0036] The quantitative evaluation results of the accuracy on the DailyTalk test set are shown in Table 1.

[0037] Table 1 Quantitative evaluation results of the accuracy on the DailyTalk test set

[0038] The results show that: the model proposed by the present invention can achieve better results than the existing state-of-the-art models on the dailytalk test set, and performs excellently in both subjective evaluation and various objective evaluations. It proves the superiority of the method of the present invention in terms of generality and accuracy, and has better application value.

[0039] Example 2 This embodiment provides a dialogue speech generation system based on emotion-aware adapter and large model reasoning, which is used to implement the dialogue speech generation method described in Embodiment 1, as Figure 3 shown, and it is composed of a speech emotion feature extraction unit, an emotion embedding generation unit, a dialogue text embedding generation unit, a model reasoning unit and a target speech generation unit.

[0040] The described speech emotion feature extraction unit: uses a speech encoder and a temporal and hierarchical attention network to extract speech emotion features from the original dialogue speech data.

[0041] The speech emotion feature extraction unit is composed of a feature representation extraction subunit and an encoding subunit.

[0042] The described feature representation extraction subunit: uses the whisper speech encoder to extract the original dialogue speech data into a feature representation : ; The described encoding subunit: uses a temporal and hierarchical attention network to encode the feature representation into speech emotion features : , Among them, represents the temporal and hierarchical attention network.

[0043] The described emotion embedding generation unit: Aligns the speech emotion features with the text features of the large language model through an emotion-aware encoding module based on the query transformer network, and generates an emotion embedding compatible with the large language model.

[0044] The described emotion embedding generation unit is as shown in Figure 4 and consists of an initialization subunit, a feature alignment subunit, and a projection subunit.

[0045] The described initialization subunit: Initializes a set of learnable query vectors in the query transformer network , and each query vector is designed to capture the representation related to a specific emotion dimension.

[0046] The described feature alignment subunit: Uses the cross-attention network in the query transformer network to interact the query vectors with the extracted speech emotion features, thereby aligning the speech emotion features with the text features of the large language model: , where respectively represent the projection matrices of query, key, and value in the attention mechanism, is the dimension of the key-value set; is the speech emotion feature, represents the cross-attention output, which is used to capture the alignment information between the speech emotion feature and the query vector; represents matrix inversion.

[0047] The described projection subunit: To make the output speech emotion feature acceptable to the large language model, projects the cross-attention output to the input embedding dimension required by the large language model through a linear layer to generate an emotion embedding: , where represents the generated emotion embedding, represents the trainable linear transformation matrix.

[0048] The described dialogue text embedding generation unit: Uses the tokenizer of the large language model to generate text embeddings for the sentences of the input dialogue text.

[0049] The described model inference unit: Is used to infer the text response and response emotion state of the dialogue by using an emotion adapter and a text adapter based on the partial low-rank adaptation network for the described emotion embedding and text embedding.

[0050] The described model inference unit is as shown inFigure 5 As shown, it consists of an emotion embedding adapter unit, a text embedding adapter unit, a splicing unit, and an inference unit.

[0051] The emotion embedding adapter unit: uses an emotion adapter based on a partial low-rank adaptation network to perform linear transformation and low-rank adaptation on the emotion embedding to generate the adapted emotion embedding :

[0052] where is the low-rank adaptation weight matrix of the emotion embedding; is the update obtained by the emotion embedding through low-rank adaptation.

[0053] The text embedding adapter unit: applies a text adapter based on a partial low-rank adaptation network to perform low-rank adaptation on the text embedding for fine-grained control of the sensitivity of the text semantic representation to the emotional state, and generates the adapted text embedding :

[0054] where represents the low-rank adaptation weight matrix of the text embedding; is the update obtained by the text embedding through low-rank adaptation.

[0055] The splicing unit: splices the adapted emotion embedding and text embedding, and uses it as the input of the large language model for joint inference:

[0056] where represents the splicing result; represents the splicing operation in the token dimension, and this splicing represents fusing the emotional state and language content of the current round of conversation to form the joint inference input of the large language model.

[0057] The inference unit: the large language model infers the text response and response emotional state of the conversation:

[0058] where is the context content, is the emotional category of the target statement, is the emotional intensity of the target statement, is the generated k-th round text response, represents the large language model, represents the probability function.

[0059] The target voice generation unit: It is used to combine the text response and the response emotional state, and use a voice generation model to generate a target voice that conforms to the emotional context.

[0060] It should be noted that each unit in the above-mentioned dialogue voice generation system based on emotion-aware adapter and large model reasoning can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned units can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above units. For the specific limitations of a dialogue voice generation system based on emotion-aware adapter and large model reasoning, refer to the limitations of a dialogue voice generation method based on emotion-aware adapter and large model reasoning (i.e., Embodiment 1) in the above text. The two have the same functions and effects, and will not be elaborated here.

[0061] Obviously, those skilled in the art can easily make various modifications to the above embodiments and apply the general principles described here to other embodiments without creative efforts. Therefore, the present invention is not limited to the above embodiments, and the improvements and modifications made by those skilled in the art to the present invention according to the disclosure of the present invention should be within the protection scope of the present invention.

Claims

1. A dialogue speech generation method based on emotion perception adapter and large model reasoning, characterized in that Including the steps: S1. Use a speech encoder and a temporal and hierarchical attention network to extract speech emotion features from the original dialogue speech data; S2. Align the speech emotion features with the text features of the large language model through an emotion-aware encoding module based on the query transformer network to generate emotion embeddings compatible with the large language model; S3. Use the tokenizer of the large language model to generate text embeddings for the sentences of the input dialogue text; S4. Adopt an emotion adapter and a text adapter based on the partial low-rank adaptation network to the emotion embeddings and text embeddings, and infer the text response and response emotion state of the dialogue; S5. Combine the text response and response emotion state, and use a speech generation model to generate target speech that conforms to the emotional context.

2. The method for generating dialogue voice according to claim 1, wherein The specific content of step S1 includes: S11. Use the whisper speech encoder to extract the original dialogue speech data as feature representations; S12. Use a temporal and hierarchical attention network to encode the feature representations into speech emotion features.

3. The method for generating dialogue voice according to claim 1, wherein The specific content of step S2 includes: S21. Initialize a set of learnable query vectors in the query transformer network, and each query vector is designed to capture the representations related to specific emotion dimensions; S22. Use the cross-attention network in the query transformer network to interact the query vectors with the extracted speech emotion features, so as to align the speech emotion features with the text features of the large language model; S23. To make the output speech emotion features acceptable to the large language model, project the cross-attention output through a linear layer to the input embedding dimension required by the large language model to generate emotion embeddings.

4. The method for generating dialogue voice according to claim 1, characterized in that, The specific content of step S3 includes: Given the dialogue text of the k-th sentence, use the tokenizer of the large language model to generate text embeddings.

5. The method for generating dialogue voice according to claim 1, characterized in that The specific content of step S4 includes: S41. Use an emotion adapter based on the partial low-rank adaptation network to perform linear transformation and low-rank adaptation on the emotion embeddings to generate adapted emotion embeddings; S42. Apply a text adapter based on the partial low-rank adaptation network to the text embeddings for low-rank adaptation, which is used to finely control the sensitivity of the text semantic representation to the emotion state, and generate adapted text embeddings; S43. Concatenate the adapted emotion embeddings and text embeddings as the input of the large language model; S44. The large language model infers the text response and response emotion state of the dialogue.

6. The method for generating dialogue voice according to claim 1, wherein The specific content of step S5 includes: Adopt FastSpeech2 as the backbone model for speech synthesis, and combine the text response and response emotion state inferred by the large language model to generate target speech consistent with the emotional context.

7. A dialogue speech generation system based on emotion perception adapter and large model reasoning, for implementing the dialogue speech generation method according to any one of claims 1-6, characterized in that, Including: Speech emotion feature extraction unit: Use a speech encoder and a temporal and hierarchical attention network to extract speech emotion features from the original dialogue speech data; Emotion embedding generation unit: Align the speech emotion features with the text features of the large language model through an emotion-aware encoding module based on the query transformer network to generate emotion embeddings compatible with the large language model; Dialogue text embedding generation unit: Use the tokenizer of the large language model to generate text embeddings for the sentences of the input dialogue text; Model Inference Unit: It is used to infer the text response and response emotion state of the dialogue by using the emotion adapter and text adapter based on the partial low-rank adaptation network for the emotion embedding and text embedding; Target Speech Generation Unit: It is used to combine the text response and response emotion state and use a speech generation model to generate the target speech that conforms to the emotional context.

8. The dialogue voice generation system according to claim 7, wherein The speech emotion feature extraction unit includes: Feature Representation Extraction Sub-unit: It uses a whisper speech encoder to extract the original dialogue speech data into feature representations; Encoding Sub-unit: It uses a temporal and hierarchical attention network to encode the feature representations into speech emotion features.

9. The dialogue voice generation system according to claim 7, wherein, The emotion embedding generation unit includes: Initialization Sub-unit: It initializes a set of learnable query vectors in the query transformer network, and each query vector is designed to capture the representation related to a specific emotion dimension; Feature Alignment Sub-unit: It uses the cross-attention network in the query transformer network to interact the query vectors with the extracted speech emotion features, so as to align the speech emotion features with the text features of the large language model; Projection Sub-unit: In order to make the output speech emotion features acceptable to the large language model, it projects the cross-attention output through a linear layer to the input embedding dimension required by the large language model to generate emotion embeddings.

10. The dialogue voice generation system according to claim 7, wherein, The model inference unit includes: Emotion Embedding Adaptation Sub-unit: It uses an emotion adapter based on the partial low-rank adaptation network to perform linear transformation and low-rank adaptation on the emotion embedding to generate the adapted emotion embedding; Text Embedding Adaptation Sub-unit: It applies a text adapter based on the partial low-rank adaptation network to the text embedding for low-rank adaptation, which is used to finely control the sensitivity of the text semantic representation to the emotion state, and generates the adapted text embedding; Concatenation Sub-unit: It concatenates the adapted emotion embedding and text embedding as the input of the large language model; Inference Sub-unit: The large language model infers the text response and response emotion state of the dialogue.

Citation Information

Patent Citations

  • Robot oriented multimodal emotion data interaction method and device

    CN106773923A

  • Emotion dialogue generation method and device based on self-attention mechanism

    CN110427490A

  • Emotional speech synthesis method and device based on AI large model

    CN117174073A

  • Multi-modal opinion expression recognition system and method driven by large language model

    CN118917407A

  • Generative AI-based emotion connection interaction method, system and equipment

    CN119166824A

Cited By

  • Voice interaction method and device, electronic equipment and storage medium

    CN121528210A