Speech interaction method, speech interaction system and storage medium

By combining interactive voice and text, and using multiple emotion feature extraction models and multilayer perceptrons to generate response voices with whole sentences and local prosodic features, the problem of intelligent dialogue systems being unable to recognize user emotions is solved, and emotionally rich and natural voice interaction is achieved.

JP2026027170AActive Publication Date: 2026-02-18NANJING SILICON INTELLIGENCE TECH CO LTD

Patent Information

Application Number
JP2025006547
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-05
Filing Date
2025-01-17
Publication Date
2026-02-18
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

Existing intelligent dialogue systems are unable to effectively recognize users' emotions and control the emotional changes in their response voices, resulting in insufficient emotional communication and overly simplistic responses.

Method used

By combining interactive voice and interactive text, using multiple emotion feature extraction models and multilayer perceptrons for emotion classification, responsive voice with whole sentence and local prosodic features is generated, achieving emotionally rich and natural voice interaction.

Benefits of technology

It improves the accuracy of emotion classification and the naturalness of generated response speech, enhancing the user's emotional communication experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026027170000001_ABST
    Figure 2026027170000001_ABST
Patent Text Reader

Abstract

To provide a voice interaction method, a voice interaction system and a storage medium capable of improving richness and naturalness of feeling.SOLUTION: Receiving an interaction speech input by a user, determining an emotion tag corresponding to the interaction speech according to the interaction speech and an interaction text corresponding to the interaction speech, determining a response text corresponding to the interaction text and a first prosodic feature and a second prosodic feature corresponding to the response text according to the emotion tag, and generating and outputting a response speech corresponding to the interaction speech according to the response text, the first prosodic feature, and the second prosodic feature; The first prosodic feature represents a prosodic feature of an entire sentence of the response text, and the second prosodic feature represents a local prosodic feature of each character in the response text.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application relates to the field of computer technology, and in particular to a voice interaction method, a voice interaction system, and a storage medium. [Background technology]

[0002] In recent years, with the advancement of artificial intelligence technology, the technology of man-machine interaction using an intelligent dialogue system has also been rapidly developing. The intelligent dialogue system can recognize an interaction text or an interaction voice input by a user and perform a series of processes, and then generate and output a response text to the interaction text or output a response voice to the interaction voice.

[0003] However, conventional intelligent dialogue systems generally cannot effectively recognize the user's emotions based on the interaction text or interaction voice, and are unable to control the degree of emotional change in the generated response voice, which is disadvantageous for emotional communication with the user. Summary of the Invention [Problem to be solved by the invention]

[0004] To solve the above problems, the embodiments of the present application provide a voice interaction method that can improve the accuracy of emotion classification of interaction voices and improve the emotional richness and naturalness of generated response voices. Specifically, the embodiments of the present application disclose the following technical configurations. [Means for solving the problem]

[0005] A voice interaction method according to a first aspect of an embodiment of the present application includes the steps of: receiving an interaction voice input by a user; determining an emotion tag corresponding to the interaction voice based on the interaction voice and an interaction text corresponding to the interaction voice; determining a response text corresponding to the interaction text and a first prosodic feature and a second prosodic feature corresponding to the response text based on the emotion tag; and generating and outputting a response voice corresponding to the interaction voice based on the response text, the first prosodic feature and the second prosodic feature, wherein the first prosodic feature represents the prosodic feature of an entire sentence of the response text, and the second prosodic feature represents the local prosodic feature of each character in the response text.

[0006] In some embodiments, determining an emotion tag corresponding to the interaction speech based on the interaction speech and an interaction text corresponding to the interaction speech includes: determining text emotion features based on the interaction speech and the interaction text, determining voice emotion features based on the interaction speech, and determining an emotion tag based on the text emotion features and the voice emotion features.

[0007] In some embodiments, determining text emotion features based on the interaction speech and the interaction text includes: processing the interaction speech with a first emotion feature extraction model to obtain whole-sentence emotion features; processing the interaction text with a second emotion feature extraction model to obtain character emotion features; and determining text emotion features based on the whole-sentence emotion features and the character emotion features.

[0008] In some embodiments, the step of determining voice emotion features based on the interaction speech includes: processing the interaction speech with a first emotion feature extraction model to obtain latent voice emotion features; processing the interaction speech with a third emotion feature extraction model to obtain explicit voice emotion features; and determining voice emotion features based on the latent voice emotion features and the explicit voice emotion features.

[0009] In some embodiments, determining a response text corresponding to the interaction text and a first prosodic feature and a second prosodic feature corresponding to the response text based on the emotion tag includes: generating a response text corresponding to the emotion tag based on the emotion tag; processing the response text and the emotion tag based on a first prosodic prediction model and a global prosodic variation parameter to obtain the first prosodic feature; and processing the response text and the first prosodic feature based on a second prosodic prediction model and a local prosodic variation parameter to obtain the second prosodic feature.

[0010] In some embodiments, the step of processing the response text and the emotion tag to obtain the first prosodic feature based on the first prosodic prediction model and the global prosodic variation parameter includes the steps of: performing a coding process on the emotion tag to obtain the prosodic feature of the entire sentence corresponding to the response text; and performing a noise reduction process on the prosodic feature of the entire sentence based on the first prosodic prediction model and the global prosodic variation parameter to obtain the first prosodic feature.

[0011] In some embodiments, the method further includes: performing a noise addition process on the prosodic features of the entire sample sentence based on a first prosody prediction model and the global prosody change parameter to obtain prosodic features of the entire sample sentence after noise addition; performing a noise removal process based on the global prosody change parameter and the prosodic features of the entire sample sentence after noise addition to obtain prosodic features of the entire sample sentence after noise removal; determining a whole-sentence prosodic loss function corresponding to the first prosody prediction model based on the prosodic features of the entire sample sentence and the prosodic features of the entire sample sentence after noise removal; and optimizing the whole-sentence prosodic loss function based on a first optimization parameter of the whole-sentence prosodic loss function.

[0012] In some embodiments, the step of processing the response text and the first prosodic features to obtain second prosodic features based on the second prosodic prediction model and the local prosodic variation parameter includes: performing a coding process on the response text by an encoder to obtain response text features; fusing the first prosodic features and the response text features to obtain fused text features; processing the fused text features based on a random prosodic predictor in the second prosodic prediction model and the local prosodic variation parameter to obtain random prosodic features corresponding to the fused text features; processing the fused text features based on a fixed prosodic predictor in the second prosodic prediction model to obtain fixed prosodic features corresponding to the fused text features; and determining the second prosodic features based on the random prosodic features, the fixed prosodic features and a control coefficient, wherein the fused text features are response text features including the first prosodic features, the random prosodic features include random fundamental frequency features, random energy features and random duration features, and the fixed prosodic features include fixed fundamental frequency features, fixed energy features and fixed duration features.

[0013] In some embodiments, the step of processing the fused text features to obtain random prosodic features corresponding to the fused text features based on the random prosodic predictor and the local prosodic variation parameter in the second prosodic prediction model includes the steps of: performing a noise addition process on the fused text features based on the random prosodic predictor and the local prosodic variation parameter to obtain the noise-added fused text features; and performing a noise removal process on the noise-added fused text features based on the random prosodic predictor and the local prosodic variation parameter to obtain the random prosodic features.

[0014] In some embodiments, the step of processing the fused text feature based on the random prosody predictor and the local prosody variation parameter in the second prosody prediction model to obtain a random prosodic feature corresponding to the fused text feature includes: processing the fused text feature based on the random fundamental frequency predictor and the local prosody variation parameter in the random prosody predictor to obtain a random fundamental frequency feature; processing the fused text feature and the random fundamental frequency feature based on the random energy predictor and the local prosody variation parameter in the random prosody predictor to obtain a random energy feature; processing the fused text feature, the random fundamental frequency feature and the random energy feature based on the random duration predictor and the local prosody variation parameter in the random prosody predictor to obtain a random duration feature; and determining the random prosodic feature based on the random fundamental frequency feature, the random energy feature and the random duration feature.

[0015] In some embodiments, the control coefficients include a first control coefficient for determining a weight of the random prosodic feature and a second control coefficient for determining a weight of the fixed prosodic feature, and determining the second prosodic feature based on the random prosodic feature, the fixed prosodic feature and the control coefficients includes determining the second prosodic feature based on the first control coefficient, the second control coefficient, the random prosodic feature and the fixed prosodic feature.

[0016] In some embodiments, the method further includes the steps of: obtaining sample response speech; obtaining sample fundamental frequency information, sample energy information, and sample duration information corresponding to the sample response speech; and coding the sample fundamental frequency information, sample energy information, and sample duration information to obtain sample local prosodic features; determining a random prosodic loss function corresponding to a random prosodic predictor based on the sample local prosodic features and the random prosodic features, and optimizing the random prosodic predictor based on second optimization parameters of the random prosodic loss function; determining a fixed prosodic loss function corresponding to the fixed prosodic predictor based on the sample local prosodic features and the fixed prosodic features, and optimizing the fixed prosodic loss function based on third optimization parameters of the fixed prosodic loss function, wherein the sample local prosodic features include the sample fundamental frequency feature, the sample energy feature, and the sample duration feature; the random prosodic loss function includes the random fundamental frequency loss function, the random energy loss function, and the random duration loss function; and the fixed prosodic loss function includes the fixed fundamental frequency loss function, the fixed energy loss function, and the fixed duration loss function.

[0017] In some embodiments, the step of performing a coding process on the response text by the encoder to obtain response text features includes: performing a coding process on the response text by a first encoder to obtain character-level features corresponding to the response text; performing a coding process on the response text by a second encoder to obtain phoneme-level features corresponding to the response text; and coding the sum of the character-level features and the phoneme-level features to obtain the response text features by a third encoder.

[0018] In some embodiments, the step of generating and outputting a response speech corresponding to the interaction speech based on the response text, the first prosodic feature, and the second prosodic feature includes the steps of determining speech features based on the first prosodic feature, the second prosodic feature, and the response text feature, and generating and outputting the response speech based on the speech features.

[0019] In some embodiments, the above method further includes, after the step of receiving an interaction speech input by a user, performing a speech recognition process on the interaction speech to obtain an interaction text corresponding to the interaction speech; and, after the step of determining an emotion tag corresponding to the interaction speech, further includes, based on a language model, processing the interaction text to obtain a response text corresponding to the interaction text.

[0020] A voice interaction system according to a second aspect of an embodiment of the present application includes: an input module configured to receive interaction speech input by a user; an emotion classification module configured to determine an emotion tag corresponding to the interaction speech based on the interaction speech and an interaction text corresponding to the interaction speech; a prosody prediction module configured to determine a response text corresponding to the interaction text and first and second prosody features corresponding to the response text based on the emotion tag; and an output module configured to generate and output a response speech corresponding to the interaction speech based on the response text, the first and second prosody features, wherein the first prosody feature represents the prosody feature of an entire sentence of the response text, and the second prosody feature represents the local prosody feature of each character in the response text.

[0021] A computer-readable storage medium according to a third aspect of an embodiment of the present application stores computer program instructions, which, when read by a computer, execute the voice interaction method according to the first aspect above.

[0022] A computer program product according to a fourth aspect of an embodiment of the present application includes a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions that, when executed by a computer, cause the computer to perform the voice interaction method described in the first aspect above. [Effects of the Invention]

[0023] In a speech interaction method according to an embodiment of the present application, an interaction speech input by a user is received; an emotion tag corresponding to the interaction speech is determined based on the interaction speech and an interaction text corresponding to the interaction speech; a response text corresponding to the interaction text and a first prosodic feature and a second prosodic feature corresponding to the response text are determined based on the emotion tag; and a response speech corresponding to the interaction speech can be generated and output based on the response text, the first prosodic feature and the second prosodic feature, wherein the first prosodic feature represents the prosodic feature of the entire sentence of the response text, and the second prosodic feature represents the local prosodic feature of each character in the response text.

[0024] According to this technical configuration, by determining an emotion tag corresponding to the interaction speech using two modalities, namely, the interaction speech and the interaction text, the accuracy of emotion classification for the interaction speech input by the user can be improved; and by determining the prosodic features of the entire sentence corresponding to the response text and the local prosodic features of each character based on the emotion tag, the emotional richness and naturalness of the generated response speech can be improved. [Brief explanation of the drawings]

[0025] In the following, in order to more clearly explain the technical configuration of the embodiments of the present invention, drawings that need to be used in the embodiments will be briefly introduced. The drawings in the following description are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort. [Figure 1] 1 is a flowchart of a voice interaction method according to some embodiments of the present application. [Figure 2] 1 is a schematic diagram of a voice interaction system according to some embodiments of the present application; [Figure 3] FIG. 2 is a schematic diagram of an emotion classification module according to some embodiments of the present application. [Figure 4] FIG. 1 is a schematic diagram of extracting MFCC features according to some embodiments of the present application. [Figure 5] 4 is a flowchart of another voice interaction method according to some embodiments of the present application. [Figure 6] FIG. 1 is a schematic diagram of a diffusion model according to some embodiments of the present application. [Figure 7] FIG. 2 is a schematic diagram of a second prosody prediction model according to some embodiments of the present application; [Figure 8] FIG. 1 is a schematic diagram of a coding process for text information according to some embodiments of the present application; [Figure 9] FIG. 1 is a schematic diagram of a VITS model according to some embodiments of the present application. [Figure 10] FIG. 2 is a schematic diagram of another voice interaction system according to some embodiments of the present application. [Figure 11] 1 is a schematic diagram of an electronic device according to some embodiments of the present application. DETAILED DESCRIPTION OF THE INVENTION

[0026] Hereinafter, in order to allow those skilled in the art to better understand the technical configuration of the embodiments of the present invention and to more easily comprehend the above-mentioned objectives, features and advantages of the embodiments of the present invention, the technical configuration of the embodiments of the present invention will be described in more detail with reference to the drawings.

[0027] Conventional intelligent dialogue systems generally have the following problems.

[0028] 1. Some existing intelligent dialogue systems only support text-based responses, do not effectively use voice, and have poor emotional intervention and comfort functions for users. In addition, when responding to user interaction text or interaction voice based on knowledge graph and database formats, the generated response text is too flat, making it difficult to realize the advantages of diversified responses and emotional responses of artificial intelligence.

[0029] 2. Some conventional intelligent dialogue systems judge and classify the user's feelings based solely on information of a single modality (i.e., text information or voice information), resulting in relatively large classification errors.

[0030] 3. When generating response speech, some conventional intelligent dialogue systems generally use a VITS (Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech) model to predict the prosodic features of the response text and generate the response speech based on the response text and prosodic features. However, the VITS model lacks fine-grained modeling of prosodic features, and the prosody of the response speech generated based on the same response text is completely identical. This is disadvantageous for emotional interaction with the user, as it cannot control the emotional changes in the response speech.

[0031] In view of the above problems, the present application provides a voice interaction method, a voice interaction system, and a storage medium that can perform voice interaction with a user, improve the accuracy of emotion classification for interaction voice input by the user, and control emotional changes in the output response voice to improve the emotional richness and naturalness of the response voice.

[0032] The voice interaction method according to the present application will be described in detail below in conjunction with the accompanying drawings.

[0033] Fig. 1 is a flowchart of a voice interaction method according to some embodiments of the present application, and Fig. 2 is a voice interaction system according to some embodiments of the present application. The voice interaction method shown in Fig. 1 can be realized by a voice interaction system 200 shown in Fig. 2. As shown in Fig. 1, the voice interaction method can include steps 110 to 140.

[0034] In step 110, interaction speech input by a user is received.

[0035] In some embodiments, as shown in FIG. 2, the voice interaction system 200 includes an input module 210, a processing module 220, an emotion classification module 230, a reasoning module 240, a voice synthesis module 250, and an output module 260.

[0036] In some embodiments, a user can input interaction speech into the voice interaction system 200, which can acquire the interaction speech via the input module 210 and send the interaction speech to the processing module 220.

[0037] In step 120, an emotion tag corresponding to the interaction voice is determined based on the interaction voice and the interaction text corresponding to the interaction voice.

[0038] In some embodiments, the processing module 220 may perform speech recognition processing on the interaction speech to convert the interaction speech into corresponding text information (i.e., interaction text). Illustratively, the processing module 220 may perform speech recognition processing on the interaction speech based on a speech recognition model, and the embodiments of the present application do not limit the type of the speech recognition model.

[0039] In some embodiments, after the processing module 220 obtains the interaction text, the emotion classification module 230 obtains the interaction text and the interaction speech, determines text emotion features based on the interaction speech and the interaction text, and determines voice emotion features based on the interaction speech, thereby classifying the user's emotion through two modalities: text emotion features and voice emotion features, i.e., determining an emotion tag corresponding to the interaction speech.

[0040] See the schematic diagram of the emotion classification module shown in Figure 3. As shown in Figure 3, the emotion classification module 230 acquires interaction speech and interaction text, and then extracts emotion features from the interaction speech using a first emotion feature extraction model to obtain 512-dimensional emotion features. These emotion features have a strong correlation with the emotion of the interaction text, and the interaction text can be treated as sentence-dimensional features, i.e., as the emotion features of the entire sentence of the interaction text. Then, the second emotion feature extraction model extracts emotion features from the interaction text to obtain character-dimensional features of the interaction text (i.e., character emotion features of the interaction text). Next, the acquired sentence-dimensional emotion features and character emotion features of the interaction text jointly constitute text modality features, and text emotion features can be obtained by performing emotion feature fusion on the sentence-dimensional emotion features and character emotion features in a text embedding layer (text_embedding).

[0041] The first emotional feature extraction model may be a Wav2Vec-2.0 model or the like, and the second emotional feature extraction model may be a MegatronBert model or the like. The embodiments of the present application do not limit the specific types of the first emotional feature extraction model and the second emotional feature extraction model.

[0042] The emotion classification module 230 can extract latent voice emotion features from the interaction speech using a first emotion feature extraction model, and extract explicit voice emotion features from the interaction speech using a third emotion feature extraction model. The third emotion feature extraction model extracts features from the Mel-Frequency Cepstral Coefficients (MFCCs) of the interaction speech to obtain MFCC features of the interaction speech, and can use the MFCC features of the interaction speech as explicit voice emotion features of the interaction speech.

[0043] See the schematic diagram for extracting MFCC features shown in Figure 4. As shown in Figure 4, when extracting MFCC features from interaction audio, first pre-emphasize the interaction audio, i.e., extend the high-frequency portion of the interaction audio using a high-pass filter, which serves to balance the spectrum and improve the signal-to-noise ratio. The formula for the transfer function of the high-pass filter can be seen in Equation (1) below. y(t)=x(t)-αx(t-1)······(1) Here, t is time, x(t) is the audio sampling value corresponding to time t, x(t-1) is the audio sampling value corresponding to time t-1, y(t) is the pre-emphasis result corresponding to time t, and α is the pre-emphasis coefficient, whose possible values ​​are generally within the range of [0.9, 1.0].

[0044] Next, after pre-emphasis on the interaction voice, the interaction voice is subjected to frame division (framing), i.e., the interaction voice is divided into multiple short-duration frames. After the interaction voice is divided into frames, a window function is used to perform windowing on the interaction voice for each frame to increase the continuity between frames and reduce spectral leakage. For example, a Hamming window can be used to perform windowing on the interaction voice for each frame. The formula for the Hamming window function can be seen in Equation (2) below.

number

[0045] After windowing, the windowed interaction audio for each frame is subjected to a Discrete Fourier Transform (DFT) to convert the time-domain signal into a frequency-domain signal. In the frequency domain, a Mel filterbank is used to extract spectral features from the interaction audio, so that the extracted features better match the perceptual characteristics of the human auditory system, providing strong support for emotion classification. The spectral features obtained by the Mel filterbank are then logarithmically transformed to obtain logarithmic spectral features. Finally, a Dual-Cosine Transform (DCT) is performed on the logarithmic spectral features to obtain the MFCC features of the interaction audio.

[0046] In some embodiments, after obtaining the latent and explicit speech emotion features of the interaction speech, speech modality features are constructed using both the latent and explicit speech emotion features, i.e., speech emotion features are obtained by performing emotion feature fusion on the latent and explicit speech emotion features in the speech embedding layer (Speech_embedding).

[0047] The emotion classification module 230 then uses a multilayer perceptron (MLP) to fuse the acquired text emotion features and audio emotion features, and performs emotion classification on the interaction voice based on the fused emotion features, thereby obtaining emotion tags corresponding to the emotion types of the interaction voice. For example, the emotion tags may include positive, negative, neutral, anger, sadness, joy, fear, surprise, disgust, etc.

[0048] According to the above technical configuration, by simultaneously using two modalities, namely, interaction text and interaction voice, it is possible to determine the emotional features corresponding to the two modalities, and determine the emotional features according to the correspondence between the interaction information of the two modalities and the latent space, thereby performing emotion classification, thereby more comprehensively capturing the user's emotional features and improving the accuracy of emotion classification.

[0049] In step 130, a response text corresponding to the interaction text and a first prosodic feature and a second prosodic feature corresponding to the response text are determined based on the emotion tag.

[0050] In some embodiments, the interaction speech and the emotional features of the interaction text are classified, an emotional tag is determined, and then a response content (i.e., a response text) to the interaction text is generated based on the emotional tag and the interaction text, a first prosodic feature corresponding to the response text is determined based on the emotional tag and the response text, and then a second prosodic feature corresponding to the response text is determined based on the response text and the first prosodic feature.

[0051] 5 is a flowchart of another voice interaction method according to some embodiments of the present application. As shown in FIG. 5, the above step 130 may include steps 510 to 530.

[0052] In step 510, a response text corresponding to the emotion tag is generated based on the emotion tag.

[0053] In some embodiments, after classifying the emotional features of the interaction voice and the interaction text and determining the emotional tag, the reasoning module 240 can generate a response content (i.e., a response text) to the interaction text based on the emotional tag and the interaction text.

[0054] For example, the reasoning module 240 can process the interaction text using a language model to generate interaction text that matches the user emotion corresponding to the emotion tag. For example, the language model can be composed of a main model (e.g., a ChatGLM model) and a fine-tuning model (e.g., a Low-Rank Adaptation of Large Language Models (LoRA) fine-tuning model). If the user's negative feelings need to be alleviated and soothed, the language model can be trained using the psyQA and / or efaqa psychological first aid datasets to generate a LoRA fine-tuning model for psychological aspects. The reasoning module 240 can load the ChatGLM model, then load the LoRA fine-tuning model, and replace the parameters of the ChatGLM model with those of the LoRA model so that the response text output from the language model better matches the language style of a psychological counselor. Finally, the interaction text is input to the ChatGLM model after the parameters have been replaced, and the ChatGLM model outputs a response text corresponding to the interaction text through reasoning.

[0055] In step 520, the response text and the emotion tag are processed based on the first prosody prediction model and the global prosody variation parameter to obtain first prosodic features.

[0056] In some embodiments, after determining the response text, the speech synthesis module 250 may generate a response voice that matches the user emotion corresponding to the emotion tag based on the emotion tag and the response text. For example, if the emotion tag corresponding to the interaction voice input by the user is negative (i.e., the user emotion is negative), the generated response voice may have a consoling tone, and if the emotion tag corresponding to the interaction voice input by the user is positive (i.e., the user emotion is positive), the generated response voice may have a positive and approving tone.

[0057] For example, when generating a response speech, not only can the prosody of the entire sentence of the response speech (i.e., the prosody at the sentence level) be adjusted, but also the local prosody of the response speech (i.e., the prosody of each character in the response speech) can be adjusted, so that the generated response speech can not only have a unified emotional style overall, but also have fine emotional changes in each local area.

[0058] In some embodiments, the speech synthesis module 250 includes a first prosody prediction model, which can be used to determine the first prosody feature of the response speech (i.e., the prosody feature of the entire sentence at the sentence level). The sentence-level prosody predictor can be constructed based on a diffusion model (Denoising Diffusion Probabilistic Models, DDPM). See the schematic diagram of the diffusion model shown in FIG. 6. As shown in FIG. 6, the diffusion model is a generative model that uses a diffusion process (a process of gradually adding noise) and a de-diffusion process (a process of gradually removing noise). In the diffusion process, a feature x is generated by a Markov chain. O Gradually change the white noise to x T and in the de-diffusion process, this white noise x T Gradually feature x O The process of training the first prosody prediction model and the process of determining the first prosody feature based on the first prosody prediction model will be described in detail below.

[0059] In some embodiments, the process of training the first prosody prediction model includes: performing a noise addition process on the prosodic features of the whole sentence of the sample based on the first prosody prediction model and the global prosody variation parameter to obtain the prosodic features of the whole sentence of the sample after noise addition; performing a noise removal process based on the global prosody variation parameter and the prosodic features of the whole sentence of the sample after noise addition to obtain the prosodic features of the whole sentence of the sample after noise removal; determining a whole-sentence prosodic loss function corresponding to the first prosody prediction model based on the prosodic features of the whole sentence of the sample and the prosodic features of the whole sentence of the sample after noise removal; and optimizing the whole-sentence prosodic loss function based on a first optimization parameter of the whole-sentence prosodic loss function.

[0060] In some examples, in the process of training the sentence-level prosody predictor, when noise is added to the prosodic features of the entire sample sentence, noise is added to the prosodic features of the entire sample sentence gradually (i.e., the prosodic features of the entire sample sentence x ) based on the global prosodic change parameter (i.e., the time step or the diffusion number). O ) and the prosodic features x t The time step can be set according to the actual situation. The realization of the diffusion process can be referred to Equation (3).

number

[0061] After noise addition is completed, the prosodic features of the entire sentence of the sample after noise addition x t In the de-diffusion process, the prosodic features x of the entire sample sentence after adding noise are t We sample at each time step and use the neural network μ θThe noise is gradually removed by

number

number

[0062] Then, the prosodic features of the whole sample sentence x O and prosodic features of the whole sentence of the sample after noise removal

number

number

number

number

[0063] In some embodiments, the step of processing the response text and the emotion tag to obtain the first prosodic feature based on the first prosodic prediction model and the global prosodic variation parameter includes the steps of: performing a coding process on the emotion tag to obtain the prosodic feature of the entire sentence corresponding to the response text; and performing a noise reduction process on the prosodic feature of the entire sentence based on the first prosodic prediction model and the global prosodic variation parameter to obtain the first prosodic feature.

[0064] In the process of reasoning based on the sentence-level prosody predictor, first, a coding process is performed on the emotion tag obtained by the emotion classification module 230 to obtain the prosody features of the entire sentence corresponding to the response text, and then, based on the sentence-level prosody predictor and the overall prosody change parameters, a noise removal process is performed on the prosody features of the entire sentence to obtain the predicted value (i.e., the first prosody feature) output by the sentence-level prosody predictor.

[0065] By increasing or decreasing the time step of the sentence-level prosody predictor, the emotional expression effect of the first prosodic features can be correspondingly increased or decreased. Increasing the time step can strengthen the sentence-level prosody predictor's ability to remove noise from the prosodic features of the entire input sentence, and the output first prosodic features will have higher accuracy and more pronounced emotional expression effect. Therefore, the time step of the sentence-level prosody predictor can be adjusted based on changes in the user's mood. For example, a relatively large time step is used when the user's mood is relatively negative, and a relatively small time step is used when the user's negative emotion has weakened, thereby adjusting the first prosodic features according to changes in the user's mood and improving the user experience.

[0066] In step 530, the response text and the first prosodic features are processed to obtain second prosodic features based on the second prosodic prediction model and the local prosodic variation parameters.

[0067] In some embodiments, the speech synthesis module 250 further includes a second prosodic prediction model, which includes a fundamental frequency predictor, an energy predictor, and a duration predictor, and can model fine-grained emotions (i.e., character-level prosody), thereby obtaining second prosodic features that allow emotions to be expressed more richly and accurately, and the second prosodic features are used to represent the local prosodic features of each character in the response text. The second prosodic prediction model can also be constructed based on the diffusion model shown in FIG. 6. The training process of the second prosodic prediction model and the process of determining the second prosodic features based on the second prosodic prediction model are described in detail below.

[0068] Please refer to the schematic diagram of the second prosody prediction model shown in Figure 7. As shown in Figure 7, the process of training the second prosody prediction model and the process of obtaining the second prosody feature based on the second prosody prediction model can be performed simultaneously, that is, the process of training the second prosody prediction model and the process of inferring can be performed simultaneously.

[0069] In some embodiments, the step of processing the response text and the first prosodic features to obtain the second prosodic features based on the second prosodic prediction model and the local prosodic variation parameter includes: performing a coding process on the response text by an encoder to obtain the response text features; fusing the first prosodic features and the response text features to obtain fused text features; processing the fused text features based on a random prosodic predictor in the second prosodic prediction model and the local prosodic variation parameter to obtain random prosodic features corresponding to the fused text features; processing the fused text features by a fixed prosodic predictor in the second prosodic prediction model to obtain fixed prosodic features corresponding to the fused text features; and determining the second prosodic features based on the random prosodic features, the fixed prosodic features and the control coefficients.

[0070] The fused text feature is a response text feature including a first prosodic feature, the second prosodic prediction model includes a random prosodic predictor and a fixed prosodic predictor, the random prosodic predictor includes a random fundamental frequency predictor, a random energy predictor, and a random duration predictor, the fixed prosodic predictor includes a fixed fundamental frequency predictor, a fixed energy predictor, and a fixed duration predictor, the random prosodic features include a random fundamental frequency feature, a random energy feature, and a random duration feature, and the fixed prosodic features include a fixed fundamental frequency feature, a fixed energy feature, and a fixed duration feature.

[0071] FIG. 8 is a schematic diagram illustrating a coding process for text information. As shown in FIG. 8, a first encoder (e.g., a Bert pre-trained model) can perform a coding process on the response text to obtain character-level features corresponding to the response text. A pre-trained Bert language model can obtain complex relationships between words and sentences in the response text to generate character-level features corresponding to the response text. These character-level features can represent complex prosodic information, thereby effectively optimizing the prosodic effect of the synthesized response speech. A second encoder (e.g., a deep learning phonetic model, a speech recognition model, etc.) can then perform a coding process on the response text to obtain phoneme-level features corresponding to the response text. The character-level features are then mapped using an embedding layer and transformed into the same dimensions as the phoneme-level features using a convolutional network. Finally, a third encoder (i.e., a text encoder) adds the two together to obtain the resulting features, which are then coded to obtain response text features. The embodiments of the present application do not limit the specific types of the first, second, and third encoders.

[0072] Next, the first prosodic feature and the response text feature are fused to obtain a fused text feature, where the fused text feature is a response text feature including the first prosodic feature. Then, the fused text feature is processed based on a random fundamental frequency predictor and a local prosodic variation parameter (i.e., a time step or a spreading number) to obtain a random fundamental frequency feature corresponding to the fused text feature, the fused text feature that has undergone fundamental frequency prediction is processed based on a random energy predictor and the local prosodic variation parameter to obtain a random energy feature corresponding to the fused text feature, and then the fused text feature that has undergone fundamental frequency and energy prediction is processed based on a random duration predictor and the local prosodic variation parameter to obtain a random duration feature corresponding to the fused text feature.

[0073] The process principle of processing the fused text feature using the random fundamental frequency predictor, random energy predictor, and random duration predictor to obtain the random fundamental frequency feature, random energy feature, and random duration feature respectively is the same: all of them perform noise addition processing on the fused text feature based on the local prosody variation parameter to obtain the noise-added fused text feature, and then perform noise removal processing on the noise-added fused text feature based on the local prosody variation parameter to obtain the fused text feature containing the random prosody feature. For specific processes and related formulas, please refer to the embodiment corresponding to Figure 6 above, and no further description will be given here.

[0074] In some embodiments, the process of training the second prosody prediction model includes the steps of: obtaining sample response speech; obtaining sample fundamental frequency information, sample energy information, and sample duration information corresponding to the sample response speech; and coding the sample fundamental frequency information, sample energy information, and sample duration information to obtain sample local prosody features; determining a random prosody loss function corresponding to the random prosody predictor based on the sample local prosody features and the random prosody features, and optimizing the random prosody predictor based on second optimization parameters of the random prosody loss function; and determining a fixed prosody loss function corresponding to the fixed prosody predictor based on the sample local prosody features and the fixed prosody features, and optimizing the fixed prosody predictor based on third optimization parameters of the fixed prosody loss function.

[0075] The sample local prosodic features include a sample fundamental frequency feature, a sample energy feature, and a sample duration feature; the random prosodic loss functions include a random fundamental frequency loss function, a random energy loss function, and a random duration loss function; and the fixed prosodic loss functions include a fixed fundamental frequency loss function, a fixed energy loss function, and a fixed duration loss function.

[0076] For example, at least one sample response voice is obtained from the sample voice database, the fundamental frequency of the sample response voice is extracted using the pysptk.sptk.rapt packet, and the fundamental frequency of the sample response voice is subjected to trimming and interpolation to obtain character-level sample fundamental frequency information. Furthermore, a short-time Fourier transform (STFT) operation is performed on the sample response voice using the librosa packet to obtain amplitude and phase information, and character-level sample energy information is obtained by calculating the sum of squares for each column and then taking the square root of the amplitude information. Furthermore, sample duration information corresponding to the sample response voice is obtained by forcibly aligning the sample response voice with the sample response text corresponding to the sample response voice using a monotonic alignment search (MAS). The obtained sample fundamental frequency information, sample energy information, and sample duration information are then coded to obtain sample fundamental frequency features, sample energy features, and sample duration features.

[0077] For example, a random prosody loss function corresponding to the random prosody predictor can be determined based on the sample local prosody feature and the random prosody feature output by each random prosody predictor, and the random prosody predictor can be optimized based on a second optimization parameter of the random prosody loss function. For example, a random fundamental frequency loss function corresponding to the random fundamental frequency predictor can be determined based on the random fundamental frequency feature and the sample fundamental frequency feature output by the random fundamental frequency predictor, a random energy loss function corresponding to the random energy predictor can be determined based on the random energy feature and the sample energy feature output by the random energy predictor, and a random duration loss function corresponding to the random duration predictor can be determined based on the random duration feature and the sample duration feature output by the random duration predictor.

[0078] The determination of the random fundamental frequency loss function, the random energy loss function, or the random time length loss function can be referred to Equation (5) in the embodiment corresponding to FIG. 6 , and the fused text feature x O Random prosodic features predicted by adding and removing noise to

number

number

number

[0079] In some embodiments, in addition to determining random prosodic features, fixed prosodic features corresponding to the fused text features can be determined based on a fixed prosodic predictor. The fixed fundamental frequency predictor / fixed energy predictor / fixed duration predictor can be a predictor including two layers of long short-term memory (LSTM) networks and one fully connected layer, so that the output of the predictor has a relatively stable structure. After the fused text features are input to the fixed fundamental frequency predictor / fixed energy predictor / fixed duration predictor, fixed fundamental frequency features, fixed energy features, and fixed duration features corresponding to the fused text features can be obtained, respectively. The formula for processing the fused text features through the fixed fundamental frequency predictor / fixed energy predictor / fixed duration predictor to obtain the fixed fundamental frequency features / fixed energy features / fixed duration features corresponding to the fused text features can be seen in Equation (7).

number

[0080] Then, a fixed prosody loss function corresponding to the fixed prosody predictor is determined based on the fixed prosody features predicted by the fixed prosody predictor and the sample local prosody features, and the fixed prosody predictor can be optimized based on the third optimization parameter in the fixed prosody loss function. Here, the formula for determining the fixed prosody loss function can refer to Equation (8).

number

[0081] Illustratively, an algorithm such as gradient descent can be used to optimize the fixed prosody predictor. When optimizing the fixed prosody predictor based on a gradient descent algorithm, the formula

number

number

[0082] After determining the random prosody loss function and the fixed prosody loss function, a loss function L corresponding to the second prosody prediction model is calculated. total =L2+L 固定予測器 It should be noted that optimizing the random prosody loss function and the fixed prosody loss function is optimizing the second prosody prediction model.

[0083] In some embodiments, the control coefficients include a first control coefficient and a second control coefficient, and determining the second prosodic feature based on the random prosodic feature, the fixed prosodic feature, and the control coefficients includes determining the second prosodic feature based on the first control coefficient, the second control coefficient, the random prosodic feature, and the fixed prosodic feature, wherein the first control coefficient is used to determine a weight of the random prosodic feature, and the second control coefficient is used to determine a weight of the fixed prosodic feature.

[0084] After determining the random prosodic features and the fixed prosodic features, the weights of the random prosodic features and the fixed prosodic features can be adjusted by a first control coefficient corresponding to the random prosodic features and a second control coefficient corresponding to the fixed prosodic features, respectively, thereby adjusting the prosodic diversity and prosodic stability, respectively, to determine the second prosodic features corresponding to the response text. For example, the first control coefficient and the second control coefficient can be adjusted to appropriate values, so that the local prosody of the response speech subsequently generated based on the random prosodic features and the fixed prosodic features has good variability and good stability.

[0085] According to the above technical configuration, the present application utilizes the distribution sampling generation features of the diffusion model itself, so that the predicted random prosodic features not only have the overall characteristics within the distribution but also have one-to-many randomness, thereby determining different prosody for the same response text. Furthermore, the time step corresponding to the noise removal process can be controlled to correspondingly control the degree of prosodic change, thereby improving the naturalness and controllability of the prosody. Furthermore, the present application predicts duration features for the response text features and also adds predictions of fundamental frequency and energy features to the response text features, thereby jointly determining character-level prosodic features of the response text based on the fundamental frequency, energy, and duration features, thereby improving the emotional richness and naturalness of the character-level prosodic features.

[0086] In step 140, a response speech corresponding to the interaction speech is generated and output based on the response text, the first prosodic feature, and the second prosodic feature.

[0087] In some embodiments, the speech synthesis module 250 further includes a speech synthesis model. After determining the first prosodic feature and the second prosodic feature, response text features (i.e., target text features) including the first prosodic feature and the second prosodic feature are input to the speech synthesis model. The speech synthesis model can determine corresponding speech features based on the target text features and generate a response speech corresponding to the interaction speech based on the speech features. Finally, the output module 260 outputs the response speech, thereby realizing speech interaction with the user.

[0088] For example, the speech synthesis model may be a Variational Inference with Adversarial Learning for End-to-End Text-to-Speech (VITS) model for end-to-end text-to-speech conversion. See the schematic diagram of the VITS model shown in FIG. 9. As shown in FIG. 9, the target text features determined by the second prosody prediction model are input into a flow-based model in the VITS model, whereby response speech features corresponding to the target text features can be obtained by the flow model. Then, a decoder in the VITS model processes the response speech features to convert the response speech features into corresponding audio waveforms, thereby obtaining response speech.

[0089] In some embodiments, referring again to Figure 9, the quality and naturalness of the generated response speech can be improved by training a VITS model. The training process for the VITS model can be seen in the following example.

[0090] 1. Audio Reconstruction Loss The audio reconstruction loss is used to measure the difference between the Mel spectrum generated by the VITS model and the real Mel spectrum. For example, a linear spectrum corresponding to a training sample speech is extracted, and a posterior encoder in the VITS model calculates the Mel spectrum y corresponding to the training sample speech based on the linear spectrum corresponding to the training sample speech. mel (This Mel spectrum is the real Mel spectrum of the training sample speech x), and then the post-encoder uses this Mel spectrum y mel Based on the latent feature z and the corresponding mean and variance, the posterior distribution q is calculated based on the mean and variance of the latent feature z. Φ(z|x), and finally, the response speech with latent feature z estimated by the decoder.

number

number

number

[0091] 2.KL Divergence Loss The KL discreteness loss is calculated based on the posterior distribution q Φ (z|x) and the conditional prior p θ It is used to measure the difference between (z|c,A). The loss function L corresponding to the KL discreteness kl can refer to equation (10). L kl =logq Φ (z|x)-logp θ (z|c,A) (10) Here, the conditional prior distribution p θ (z|c,A) is the target text feature output from the second prosody prediction model. After the expressive power of this target text feature is improved by the flow model, it is aligned with the above latent feature z output from the post-encoder to obtain the alignment matrix A.

[0092] 3. Adversarial Training Loss The adversarial training loss improves the quality of the generated response speech by distinguishing between real audio and synthetic audio generated by the decoder by introducing a discriminator D into the VITS model. The adversarial training loss includes the following three parts: a. Discriminator Loss The classifier loss is used to measure the discrimination ability of the classifier D between the training sample speech x and the synthetic audio G(z). The loss function L corresponding to the classifier loss is adv (D) can refer to equation (11). L adv (D)=Ε x,z [(D(x)-1) 2 +D(G(z)) 2 ]······(11) b.Generator Loss The generator loss is used to measure the ability of the synthetic audio G(z) to be classified as the training sample audio x by the classifier D. The loss function L corresponding to the generator loss is adv (G) can refer to equation (12). L adv (G)=Ε z [(D(G(z))-1) 2 ]······(12) c. Feature Matching Loss The feature matching loss is used to measure the difference between different hierarchical features of the classifier D between the synthetic audio G(z) and the training sample speech x. The loss function L corresponding to the feature matching loss is fm (G) can refer to equation (13).

number

[0093] According to the above technical configuration, the audio re-establishment quality, the alignment of the latent expression, and the authenticity of the generated audio in the VITS model are comprehensively considered, and these loss functions are optimized by the audio re-establishment loss, the KL discreteness loss, and the adversarial training loss optimization model, thereby enabling the VITS model to generate high-quality and natural response speech.

[0094] By applying the technical configuration of the present application, the emotion classification module in the voice interaction system performs emotion classification on the interaction speech input by the user based on multi-modality (text modality and speech modality), thereby making the emotion classification result more accurate. And the speech synthesis module in the voice interaction system determines sentence-level prosodic features (i.e., prosodic features of the entire sentence) and character-level prosodic features corresponding to the response text based on the response text corresponding to the interaction speech, so that the generated response speech not only has a uniform emotional style overall, but also has subtle local emotional changes according to the meaning of different sentences.

[0095] 10 is a schematic diagram of another voice interaction system according to some embodiments of the present application. As shown in FIG. 10, the voice interaction system 1000 includes an input module 1001, an emotion classification module 1002, a prosody prediction module 1003, and an output module 1004.

[0096] The input module 1001 is configured to receive an interaction voice input by a user.

[0097] The emotion classification module 1002 is configured to determine an emotion tag corresponding to the interaction speech based on the interaction speech and an interaction text corresponding to the interaction speech.

[0098] The prosody prediction module 1003 is configured to determine, based on the emotion tag, a response text corresponding to the interaction text and a first prosodic feature and a second prosodic feature corresponding to the response text, wherein the first prosodic feature is for representing the prosodic feature of the entire sentence of the response text, and the second prosodic feature is for representing the local prosodic feature of each character in the response text.

[0099] The output module 1004 is configured to generate and output a response speech corresponding to the interaction speech based on the response text, the first prosodic feature, and the second prosodic feature.

[0100] In some embodiments, the emotion classification module 1002 is specifically configured to determine text emotion features based on the interaction speech and the interaction text, determine voice emotion features based on the interaction speech, and determine emotion tags based on the text emotion features and the voice emotion features.

[0101] In some embodiments, the emotion classification module 1002 is specifically configured to process the interaction speech through a first emotion feature extraction model to obtain whole-sentence emotion features, process the interaction text through a second emotion feature extraction model to obtain character emotion features, and determine text emotion features based on the whole-sentence emotion features and the character emotion features.

[0102] In some embodiments, the emotion classification module 1002 is specifically configured to process the interaction speech through a first emotion feature extraction model to obtain latent voice emotion features, process the interaction speech through a third emotion feature extraction model to obtain overt voice emotion features, and determine voice emotion features based on the latent voice emotion features and the overt voice emotion features.

[0103] In some embodiments, the prosody prediction module 1003 is specifically configured to: generate, based on the emotion tag, a response text corresponding to the emotion tag; process the response text and the emotion tag to obtain a first prosodic feature based on a first prosody prediction model and a global prosodic variation parameter; and process the response text and the first prosodic feature to obtain a second prosodic feature based on a second prosodic prediction model and a local prosodic variation parameter.

[0104] In some embodiments, the prosody prediction module 1003 is specifically configured to perform a coding process on the emotion tag to obtain prosodic features of the entire sentence corresponding to the response text, and perform a noise reduction process on the prosodic features of the entire sentence based on the first prosody prediction model and the global prosodic variation parameter to obtain the first prosodic features.

[0105] As shown in FIG. 10, the voice interaction system 1000 further includes a first training module 1005 .

[0106] In some embodiments, the first training module 1005 is configured to: perform a noise addition process on the prosodic features of the whole sentence of the sample, based on the first prosody prediction model and the global prosody variation parameter, to obtain prosodic features of the whole sentence of the sample after noise addition; perform a noise removal process, based on the global prosody variation parameter and the prosodic features of the whole sentence of the sample after noise addition, to obtain prosodic features of the whole sentence of the sample after noise removal; determine a whole-sentence prosodic loss function corresponding to the first prosody prediction model, based on the prosodic features of the whole sentence of the sample and the prosodic features of the whole sentence of the sample after noise removal; and optimize the whole-sentence prosodic loss function based on a first optimization parameter of the whole-sentence prosodic loss function.

[0107] In some embodiments, the prosody prediction module 1003 is specifically configured to: perform a coding process on the response text by an encoder to obtain response text features; fuse the first prosody features and the response text features to obtain fused text features; process the fused text features based on a random prosody predictor and a local prosody variation parameter in a second prosody prediction model to obtain random prosody features corresponding to the fused text features; process the fused text features based on a fixed prosody predictor in the second prosody prediction model to obtain fixed prosody features corresponding to the fused text features; and determine the second prosody features based on the random prosody features, the fixed prosody features and the control coefficients, where the fused text features are response text features including the first prosody features, the random prosody features including a random fundamental frequency feature, a random energy feature, and a random duration feature, and the fixed prosody features including a fixed fundamental frequency feature, a fixed energy feature, and a fixed duration feature.

[0108] In some embodiments, the prosody prediction module 1003 is specifically configured to perform a noise addition process on the fused text feature based on the random prosody predictor and the local prosody variation parameter to obtain a noise-added fused text feature, and to perform a denoising process on the noise-added fused text feature based on the random prosody predictor and the local prosody variation parameter to obtain a random prosody feature.

[0109] In some embodiments, the prosody prediction module 1003 is specifically configured to: process the fused text features based on a random fundamental frequency predictor and a local prosody variation parameter in the random prosody predictor to obtain a random fundamental frequency feature; process the fused text features and the random fundamental frequency feature based on a random energy predictor and a local prosody variation parameter in the random prosody predictor to obtain a random energy feature; process the fused text features, the random fundamental frequency feature, and the random energy feature based on a random duration predictor and a local prosody variation parameter in the random prosody predictor to obtain a random duration feature; and determine the random prosody feature based on the random fundamental frequency feature, the random energy feature, and the random duration feature.

[0110] In some embodiments, the control coefficients include a first control coefficient for determining a weight of the random prosodic feature and a second control coefficient for determining a weight of the fixed prosodic feature, and the prosody prediction module 1003 is specifically configured to determine the second prosodic feature based on the first control coefficient, the second control coefficient, the random prosodic feature and the fixed prosodic feature.

[0111] As shown in FIG. 10, the voice interaction system 1000 further includes a second training module 1006 .

[0112] In some embodiments, the second training module 1006 is configured to: acquire a sample response speech; acquire sample fundamental frequency information, sample energy information, and sample duration information corresponding to the sample response speech; code the sample fundamental frequency information, sample energy information, and sample duration information to acquire sample local prosodic features; determine a random prosodic loss function corresponding to the random prosodic predictor based on the sample local prosodic features and the random prosodic features; optimize the random prosodic predictor based on second optimization parameters of the random prosodic loss function; determine a fixed prosodic loss function corresponding to the fixed prosodic predictor based on the sample local prosodic features and the fixed prosodic features; and optimize the fixed prosodic predictor based on third optimization parameters of the fixed prosodic loss function, wherein the sample local prosodic features include the sample fundamental frequency feature, the sample energy feature, and the sample duration feature; the random prosodic loss function includes a random fundamental frequency loss function, a random energy loss function, and a random duration loss function; and the fixed prosodic loss function includes a fixed fundamental frequency loss function, a fixed energy loss function, and a fixed duration loss function.

[0113] In some embodiments, the prosody prediction module 1003 is specifically configured to: perform a coding process on the response text using a first encoder to obtain character-level features corresponding to the response text; perform a coding process on the response text using a second encoder to obtain phoneme-level features corresponding to the response text; and code the sum of the character-level features and the phoneme-level features using a third encoder to obtain response text features.

[0114] In some embodiments, the output module 1004 is specifically configured to determine speech features based on the first prosodic feature, the second prosodic feature, and the response text feature, and generate and output a response speech based on the speech features.

[0115] As shown in FIG. 10, the voice interaction system 1000 further includes a processing module 1007 .

[0116] In some embodiments, the processing module 1007 is configured to perform speech recognition processing on the interaction speech to obtain an interaction text corresponding to the interaction speech, and process the interaction text through a language model to obtain a response text corresponding to the interaction text.

[0117] 11 is a schematic diagram of an electronic device according to some embodiments of the present application. In some embodiments, the electronic device includes one or more processors and a memory. The memory is configured to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors can implement the voice interaction method in the above embodiments.

[0118] 11, the electronic device 1100 includes a processor 1101 and a memory 1102. By way of example, the electronic device 1100 may further include a communications interface 1103 and a communications bus 1104.

[0119] The processor 1101, memory 1102, and communication interface 1103 communicate with each other via a communication bus 1104. The communication interface 1103 is for communicating with other devices (for example, clients and network elements such as other servers).

[0120] In some embodiments, the processor 1101 executes a program 1105, which may specifically execute the relevant steps in the above-described embodiments of the voice interaction method. Specifically, the program 1105 may include program code including computer-executable instructions.

[0121] Illustratively, processor 1101 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement some embodiments of the present application. Electronic device 1100 may include one or more processors, which may be the same type of processor (e.g., one or more CPUs) or different types of processors (e.g., one or more CPUs and one or more ASICs).

[0122] In some embodiments, memory 1102 is for storing programs 1105. Memory 1102 may include high-speed RAM memory and may further include non-volatile memory (NVM) (e.g., at least one magnetic disk memory).

[0123] Specifically, the program 1105 can be invoked by the processor 1101 to cause the electronic device 1100 to execute steps of a voice interaction method.

[0124] A computer-readable storage medium according to some embodiments of the present application stores at least one executable instruction, which, when executed by the electronic device 1100, causes the electronic device 1100 to perform steps of the voice interaction method in the above-described embodiments.

[0125] The executable instructions can be used, in particular, to cause the electronic device 1100 to perform the operations of a voice interaction method.

[0126] For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0127] The beneficial effects that can be achieved by the computer-readable storage medium according to some embodiments of the present application may refer to the beneficial effects in the corresponding voice interaction methods provided above, and will not be further described here.

[0128] It should be noted that in this application, terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not require or imply the existence of any actual relationship or order between those entities or operations. Furthermore, the terms "comprise," "comprises," or any other variations thereof are intended to cover a non-exclusive inclusion, whereby a process, method, article, or device that includes a set of elements not only includes those elements, but also other elements not explicitly stated, or may further include elements inherent in such process, method, article, or device. Absent further limitations, an element qualified by the phrase "comprises ..." does not exclude the presence of other identical elements in the process, method, article, or device that includes the element.

[0129] The embodiments in this specification are described in a related form, and the same or similar parts between the embodiments can be referred to each other, and each embodiment will be described focusing on the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so they will be described relatively briefly, and for related points, please refer to the description of the method embodiments.

[0130] The logic and / or steps depicted in flowcharts or otherwise described may be considered, for example, to be a sequential listing of executable instructions to implement logical functions, and may in particular be used by an instruction execution system, device, or apparatus (e.g., a computer-based system, a system including a processor, or other system capable of retrieving instructions from an instruction execution system, device, or apparatus and executing the instructions) on any computer-readable medium, or by a combination of these instruction execution systems, devices, or apparatus.

[0131] As used herein, a "computer-readable medium" may be any device that can contain, store, communicate, propagate, or transmit a program for use by an instruction execution system, device, or apparatus, or a combination of such instruction execution systems, devices, or apparatus.

[0132] More specific examples (non-exhaustive list) of computer-readable media include an electrical connection having one or more wires (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), fiber optic device, and portable disk read-only memory (CDROM).

[0133] The computer-readable medium may also be paper or other suitable medium on which the program is printed, and the program may be obtained in electronic form by, for example, optically scanning the paper or other medium and editing, interpreting, or processing it in any other suitable form as needed, and storing the program in computer memory. Note that parts of this application may be realized in hardware, software, firmware, or a combination thereof.

[0134] In the above-described embodiments, a plurality of steps or methods may be implemented by software or firmware stored in a memory and executed by an appropriate instruction execution system. For example, when implemented by hardware, as in the other embodiments, the steps or methods may be implemented by any one or combination of techniques well known in the art, such as a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application specific integrated circuit having appropriate combinational logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0135] The above-described embodiments of the present application do not limit the protection scope of the present application.

Claims

1. receiving an interaction voice input by a user; determining an emotion tag corresponding to the interaction speech based on the interaction speech and an interaction text corresponding to the interaction speech; determining a response text corresponding to the interaction text and a first prosodic feature and a second prosodic feature corresponding to the response text based on the emotion tag; generating and outputting a response speech corresponding to the interaction speech based on the response text, the first prosodic feature, and the second prosodic feature; the first prosodic feature represents the prosodic feature of an entire sentence of the response text; the second prosodic features represent local prosodic features of each character in the response text; A voice interaction method comprising:

2. determining an emotion tag corresponding to the interaction voice based on the interaction voice and an interaction text corresponding to the interaction voice, determining text emotion features based on the interaction speech and the interaction text, and determining voice emotion features based on the interaction speech; determining the emotion tag based on the text emotion features and the voice emotion features; 2. The voice interaction method according to claim 1, wherein:

3. determining a text emotion feature based on the interaction voice and the interaction text, Processing the interaction speech through a first emotion feature extraction model to obtain emotion features of the whole sentence; processing the interaction text through a second emotion feature extraction model to obtain character emotion features; determining the text sentiment feature based on the whole sentence sentiment feature and the character sentiment feature; 3. The voice interaction method according to claim 2, wherein:

4. The step of determining a voice emotion feature based on the interaction voice includes: processing the interaction speech through a first emotion feature extraction model to obtain latent speech emotion features; processing the interaction speech through a third emotion feature extraction model to obtain overt speech emotion features; determining the voice emotion features based on the latent voice emotion features and the overt voice emotion features; 3. The voice interaction method according to claim 2, wherein:

5. determining a response text corresponding to the interaction text and a first prosodic feature and a second prosodic feature corresponding to the response text based on the emotion tag, generating the response text corresponding to the emotion tag based on the emotion tag; processing the response text and the emotion tag based on a first prosody prediction model and a global prosody variation parameter to obtain the first prosody feature; processing the response text and the first prosodic features based on a second prosodic prediction model and local prosodic variation parameters to obtain the second prosodic features; 2. The voice interaction method according to claim 1, wherein:

6. processing the response text and the emotion tag to obtain the first prosodic features based on the first prosodic prediction model and the global prosodic variation parameter, performing a coding process on the emotion tag to obtain prosodic features of the entire sentence corresponding to the response text; and performing a noise reduction process on the prosodic features of the entire sentence based on the first prosodic prediction model and the overall prosodic variation parameter to obtain the first prosodic features.

6. A voice interaction method according to claim 5, wherein:

7. The voice interaction method includes: performing a noise addition process on the prosodic features of the entire sample sentence based on the first prosodic prediction model and the overall prosodic variation parameter to obtain the prosodic features of the entire sample sentence after noise addition; performing a noise removal process based on the overall prosodic variation parameter and the prosodic features of the entire sentence of the noise-added sample to obtain the prosodic features of the entire sentence of the noise-removed sample; determining a whole-sentence prosodic loss function corresponding to the first prosody prediction model based on the whole-sentence prosodic features of the sample and the whole-sentence prosodic features of the denoised sample; optimizing the prosodic loss function of the whole sentence based on a first optimization parameter of the prosodic loss function of the whole sentence; 7. A voice interaction method according to claim 6, wherein:

8. processing the response text and the first prosodic features to obtain the second prosodic features based on the second prosodic prediction model and local prosodic variation parameters, performing a coding process on the response text by an encoder to obtain response text features; fusing the first prosodic feature and the response text feature to obtain a fused text feature; processing the fused text features to obtain random prosodic features corresponding to the fused text features based on a random prosodic predictor in the second prosodic prediction model and the local prosodic variation parameters; processing the fused text features to obtain fixed prosodic features corresponding to the fused text features based on a fixed prosodic predictor in the second prosodic prediction model; determining the second prosodic feature based on the random prosodic feature, the fixed prosodic feature and a control coefficient; the fused text feature is a response text feature including the first prosodic feature; the random prosodic features include a random fundamental frequency feature, a random energy feature, and a random duration feature; the fixed prosodic features include a fixed fundamental frequency feature, a fixed energy feature, and a fixed duration feature; 6. A voice interaction method according to claim 5, wherein:

9. processing the fused text features to obtain random prosodic features corresponding to the fused text features based on a random prosodic predictor in the second prosodic prediction model and the local prosodic variation parameters, performing a noise addition process on the fused text feature based on the random prosody predictor and the local prosody variation parameter to obtain a noise-added fused text feature; and performing a noise removal process on the noise-added fused text features based on the random prosody predictor and the local prosody variation parameters to obtain the random prosody features.

9. A voice interaction method according to claim 8, wherein:

10. processing the fused text features to obtain random prosodic features corresponding to the fused text features based on a random prosodic predictor in the second prosodic prediction model and the local prosodic variation parameters, processing the fused text features to obtain the random fundamental frequency features based on a random fundamental frequency predictor in the random prosody predictor and the local prosody variation parameters; processing the fused text features and the random fundamental frequency features to obtain the random energy features based on a random energy predictor in the random prosody predictor and the local prosody variation parameters; processing the fused text feature, the random fundamental frequency feature and the random energy feature according to a random duration predictor in the random prosody predictor and the local prosody variation parameter to obtain the random duration feature; determining the random prosodic features based on the random fundamental frequency features, the random energy features and the random duration features.

9. A voice interaction method according to claim 8, wherein:

11. the control coefficients include a first control coefficient for determining weights of the random prosodic features and a second control coefficient for determining weights of the fixed prosodic features; determining the second prosodic feature based on the random prosodic feature, the fixed prosodic feature and a control coefficient, determining the second prosodic feature based on the first control coefficient, the second control coefficient, the random prosodic feature, and the fixed prosodic feature; A voice interaction method according to any one of claims 8 to 10.

12. The voice interaction method includes: obtaining a sample response speech; obtaining sample fundamental frequency information, sample energy information, and sample duration information corresponding to the sample response speech, and coding the sample fundamental frequency information, the sample energy information, and the sample duration information to obtain sample local prosodic features; determining a random prosody loss function corresponding to the random prosody predictor based on the sample local prosody features and the random prosody features, and optimizing the random prosody predictor based on a second optimization parameter of the random prosody loss function; determining a fixed prosody loss function corresponding to the fixed prosody predictor based on the sample local prosody features and the fixed prosody features; and optimizing the fixed prosody predictor based on a third optimization parameter of the fixed prosody loss function; the sample local prosodic features include a sample fundamental frequency feature, a sample energy feature, and a sample duration feature; the random prosody loss function includes a random fundamental frequency loss function, a random energy loss function, and a random duration loss function; The fixed prosody loss function includes a fixed fundamental frequency loss function, a fixed energy loss function, and a fixed duration loss function. A voice interaction method according to any one of claims 8 to 10.

13. The step of performing a coding process on the response text by an encoder to obtain response text features includes: performing a coding process on the response text by a first encoder to obtain character-level features corresponding to the response text; performing a coding process on the response text by a second encoder to obtain phoneme-level features corresponding to the response text; and coding the character-level features and the phoneme-level features by a third encoder to obtain the response text features. A voice interaction method according to any one of claims 8 to 10.

14. generating and outputting a response speech corresponding to the interaction speech based on the response text, the first prosodic feature, and the second prosodic feature, determining speech features based on the first prosodic features, the second prosodic features and the response text features; generating and outputting the response voice based on the voice characteristics. A voice interaction method according to any one of claims 1 to 10.

15. The voice interaction method includes: after receiving the interaction speech input by the user, further comprising: performing speech recognition processing on the interaction speech to obtain the interaction text corresponding to the interaction speech; After determining the emotion tag corresponding to the interaction voice, the method further includes processing the interaction text through a language model to obtain the response text corresponding to the interaction text. A voice interaction method according to any one of claims 1 to 10.

16. an input module configured to receive interaction speech input by a user; an emotion classification module configured to determine an emotion tag corresponding to the interaction speech based on the interaction speech and an interaction text corresponding to the interaction speech; a prosody prediction module configured to determine a response text corresponding to the interaction text and a first prosodic feature and a second prosodic feature corresponding to the response text based on the emotion tag; an output module configured to generate and output a response speech corresponding to the interaction speech based on the response text, the first prosodic feature, and the second prosodic feature; the first prosodic feature represents the prosodic feature of an entire sentence of the response text; the second prosodic features represent local prosodic features of each character in the response text; A voice interaction system characterized by:

Citation Information

Patent Citations

  • Speech synthesis method and related device, electronic equipment and storage medium

    CN114283781A

  • Speech synthesis method, speech synthesis system, electronic equipment and storage medium

    CN116682411A

  • Speech synthesis method and device and speech synthesis model training method and device

    CN117877460A

  • Method and system for generating sympathetic back-channel signal

    US20240221742A1

Cited By

  • Conversation monitoring program, conversation monitoring device, conversation monitoring system, and conversation monitoring method

    JP7893542B1