Voice interaction method and device, computer equipment and storage medium

By identifying emotional information in user voice conversations and generating emotional matching voice reply, the problem of user emotions not being paid attention to in the prior art is solved, and the authenticity and user experience of voice conversations are improved.

CN119993216APending Publication Date: 2025-05-13CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510249728.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing voice dialogue system cannot effectively judge the user's emotions and respond accordingly, resulting in poor user experience.

Method used

By obtaining the user's voice conversation data, the user's target emotional information is identified, and a voice reply that conforms to the user's emotions is generated based on the information, including feedback tone categories, speech speed categories and emotion categories.

Benefits of technology

It realizes providing voice replies that conform to current emotions to users, enhances the anthropomorphic authenticity of voice conversations, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993216A_ABST
    Figure CN119993216A_ABST
Patent Text Reader

Abstract

The invention relates to a voice interaction method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring voice dialogue data of a user; obtaining current target emotion information of the user according to the voice dialogue data; based on the target emotion information, feedback emotion information of reply is obtained, and the feedback emotion information comprises a feedback mood category, a feedback speed category and a feedback emotion category required by feedback; and obtaining reply voice data according to the feedback emotion information. By adopting the method, the problem that a dialogue system lacks emotion feedback in the prior art can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech generation technology, and in particular to a speech interaction method, apparatus, computer equipment and storage medium. Background Art

[0002] At present, the voice dialogue system is mainly composed of five modules: speech recognition (ASR), natural speech understanding (NLU), dialogue management (DM), natural language generation (NLG), and speech synthesis (TTS). The above methods can well complete speech recognition, semantic understanding and voice reply based on semantic information, and then complete the function of real-time dialogue with users.

[0003] However, in actual applications, the user's emotions are not taken into account in the conversation. It is not possible to judge the user's emotions well and respond accordingly. Instead, the user can only respond with dull text and voice, which affects the user experience. Summary of the invention

[0004] Based on this, a voice interaction method, apparatus, computer device and storage medium are provided to improve the problem of lack of emotional feedback in the dialogue system in the prior art.

[0005] On the one hand, a voice interaction method is provided, comprising:

[0006] Get the user's voice conversation data;

[0007] Obtaining the user's current target emotion information according to the voice conversation data;

[0008] Based on the target emotion information, feedback emotion information of the reply is obtained, wherein the feedback emotion information includes a feedback tone category, a feedback speech speed category, and a feedback emotion category required for the feedback;

[0009] According to the feedback emotion information, reply voice data is obtained.

[0010] In one embodiment, obtaining the user's current target emotion information according to the voice conversation data includes:

[0011] According to the voice conversation data, obtaining corresponding conversation text data;

[0012] Based on the conversation text data, obtaining a first emotion recognition result of text emotion recognition;

[0013] Obtaining a second emotion recognition result of speech emotion recognition according to the speech dialogue data;

[0014] The target emotion information is obtained according to the first emotion recognition result and the second emotion recognition result.

[0015] In one embodiment, obtaining a second emotion recognition result of speech emotion recognition according to the speech dialogue data includes:

[0016] Discretization processing is performed on the voice dialogue data to obtain a single feature encoding vector at the word level;

[0017] According to the single feature encoding vector at the word level, a hidden vector at the word level is obtained;

[0018] According to the hidden vector of the word level, local information representation is obtained based on the attention mechanism;

[0019] According to the local information representation, the hidden vector at the sentence level is obtained;

[0020] According to the hidden vector at the sentence level, a global information representation is obtained based on an attention mechanism;

[0021] The second emotion recognition result is obtained based on the global information representation.

[0022] In one embodiment, obtaining the local information representation includes determining the hidden vector h at the word level according to the following mathematical expression: it , word-level weight a it , local information representation g i :

[0023] h it =tanh(w w *e it +b w );

[0024]

[0025] g i =∑ t a it e it ;

[0026] Among them, w w is the weight matrix of the multilayer perceptron in the local information level, b w is the bias of the multilayer perceptron in the local information level, h w is a context vector that measures the importance of local features in a sentence, e it is the t-th single feature encoding vector in the ith sentence after discretization, and tanh(*) represents the hyperbolic tangent function;

[0027] The obtaining of the global information representation includes determining the hidden vector h at the sentence level according to the following mathematical expression: i , sentence-level weight a i, global information representation s:

[0028] h i =tanh(w s *e i +b s );

[0029]

[0030] s=∑ i a i e i ;

[0031] Among them, w s is the weight matrix of the multilayer perceptron in the global information level, b s is the bias of the multilayer perceptron in the global information level; e i is the input of the global information level; h s It is a context vector that measures the importance of local features in the entire input.

[0032] In one embodiment, obtaining the feedback emotion information of the reply based on the target emotion information includes:

[0033] Inputting the target emotion information into a pre-trained large model emotion descriptor;

[0034] and prompting the large model emotion descriptor according to the prompt word engineering to obtain a speech behavior description, and using the speech behavior description as the feedback emotion information;

[0035] The speech behavior description is a behavior strategy including the feedback tone category, feedback speech speed category, and feedback emotion category.

[0036] In one embodiment, obtaining reply voice data according to the feedback emotion information includes:

[0037] Taking the feedback emotion information and the context information in the voice dialogue data as input, obtaining feedback text according to a pre-trained text generation model, wherein the text generation model is obtained by fine-tuning a general large language model;

[0038] The reply voice data corresponding to the feedback text is obtained according to the feedback emotion information and the feedback text.

[0039] In one embodiment, after obtaining the user's current target emotion information, the method further includes:

[0040] Determining whether the user is in an abnormal emotion according to the target emotion information;

[0041] When the user is in an abnormal emotion, feedback emotion information in response to the abnormal emotion is obtained to obtain response voice data with the feedback emotion.

[0042] In another aspect, a voice interaction device is provided, the device comprising:

[0043] A voice acquisition module is used to acquire the user's voice conversation data;

[0044] An emotion recognition module, used to obtain the user's current target emotion information based on the voice conversation data;

[0045] A feedback emotion description module, used to obtain the feedback emotion information of the reply based on the target emotion information, wherein the feedback emotion information includes the feedback tone category, feedback speech speed category, and feedback emotion category required for the feedback;

[0046] The speech generation module is used to obtain reply speech data according to the feedback emotion information.

[0047] In another aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method is implemented when the processor executes the computer program.

[0048] A computer-readable storage medium is also provided, on which a computer program is stored, and when the computer program is executed by a processor, the method described above is implemented.

[0049] The above-mentioned voice interaction method, device, computer equipment and storage medium obtain the user's current target emotion information by performing emotion recognition on the voice dialogue data, and obtain the feedback emotion information that should be used for the voice reply based on the target emotion information, including the feedback tone category, feedback speech speed category, and feedback emotion category required for the feedback. According to the feedback tone category, feedback speech speed category, and feedback emotion category, voice generation is performed, and the user can be provided with a voice reply that matches the user's current emotion, thereby enhancing the authenticity of the anthropomorphism of the voice dialogue. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a flow chart of voice dialogue in the prior art;

[0051] Figure 2 A diagram of an application environment of a voice interaction method in an embodiment;

[0052] Figure 3 is a flowchart of a voice interaction method in one embodiment;

[0053] Figure 4 A schematic diagram of a process for obtaining target emotion information in one embodiment;

[0054] Figure 5 is a model structure diagram of speech recognition in one embodiment;

[0055] Figure 6 A flowchart of voice interaction in one embodiment;

[0056] Figure 7 is a structural block diagram of a voice interaction device in one embodiment;

[0057] Figure 8 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0059] The voice dialogue system uses technologies such as natural language processing (NLP) and machine learning to achieve natural and smooth interaction with users. It can understand the user's language input and provide corresponding answers or perform corresponding operations based on the user's intentions, thus playing an important role in many fields.

[0060] In the related technologies, the analysis method for voice dialogue system mainly consists of five modules: speech recognition (ASR), natural language understanding (NLU), dialogue management (DM), natural language generation (NLG), and text-to-speech synthesis (TTS). The above methods can well complete speech recognition, semantic understanding and voice reply based on semantic information, and then complete the function of real-time dialogue communication with users.

[0061] However, in actual applications, the user's emotions at the time are not taken into account in the conversation. When the user is in a sad or other emotional state and needs comfort, the existing methods cannot judge the user's emotions well and respond accordingly based on the emotions. They can only reply with dull text and voice, which will affect the user experience and reduce the user's intention to use the dialogue system.

[0062] Taking the conventional in-vehicle voice dialogue system as an example, its processing flow is as follows Figure 1 Shown:

[0063] Speech recognition: The dialogue system receives user speech input as the starting point of the dialogue. The speech input is converted into text through speech recognition technology.

[0064] Natural semantic understanding: Language understanding is an important component of a dialogue system, which is responsible for converting user input into a machine-understandable form to identify user intent, extract key information, and understand context.

[0065] Dialogue management: Dialogue management is the core component of the dialogue system, which is responsible for managing the flow and decision-making of the dialogue. The dialogue manager determines the actions that the system should take based on the user's intent and the state of the system, and generates the system's response.

[0066] Language generation: It is responsible for generating the system's responses according to the instructions of the dialogue manager. The language generator can use templates, rules, statistical models, or deep learning models to generate natural and fluent language responses.

[0067] Speech output: The dialogue system converts the generated responses into speech or text form for display to the user.

[0068] The above processing methods cannot well understand user emotions and respond to them accordingly, which affects the user experience.

[0069] The present application provides a voice interaction method that can provide corresponding emotional voice responses based on the user's emotions.

[0070] The voice interaction method provided in this application can be applied to Figure 2 In the application environment shown, the terminal 102 communicates with the server 104 through the network. The terminal 102 can be, but is not limited to, various vehicle terminals, personal computers, laptops, smart phones, tablet computers, and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0071] In one embodiment, the voice interaction method is applied to the terminal 102 as an example for description. Figure 3 As shown, the following steps are included:

[0072] Step 110, obtaining the user's voice conversation data.

[0073] Taking in-car voice conversation as an example, the in-car microphone can collect voice signals from the driver or other users.

[0074] Step 120, obtaining the user's current target emotion information based on the voice conversation data.

[0075] In the conversation scenario, the user's emotions behind the recognition can be the target emotion information, which can be the user's current emotion category. The emotion category is exemplarily divided into positive, neutral, and negative. Positive emotions are subdivided into like, happiness, gratitude, etc.; negative emotions are subdivided into complaints, anger, disgust, fear, sadness, etc., and neutral emotions are calm. In some possible implementations, with the help of emotion identification tools such as emotion wheels, emotion classification can be more refined.

[0076] Speech emotion recognition technology uses the acoustic features of a speech (such as pitch, speaking speed, volume, pauses, etc.) to identify the speaker's emotional state. Traditional methods usually include two steps: emotion feature extraction and statistical modeling. At present, emotion recognition models based on speech signals are mainly divided into two categories: discrete form emotion description models and continuous form emotion description models. Taking the discrete form emotion description model as an example, emotions are described in the form of discrete, adjective labels, such as angry, happy, surprised, disgusted, afraid, and sad. It can be implemented based on general classification models such as Gaussian mixture model (GMM), hidden Markov model (HMM), support vector machine (SVM), etc.

[0077] In this embodiment, after obtaining the voice dialogue data when the user speaks, the voice dialogue data is input into a pre-trained emotion recognition model to obtain the target emotion information when the user speaks. The pre-trained emotion recognition model is, for example, a BERT model, a Transformer model, and the like.

[0078] Step 130: obtaining reply feedback emotion information based on the target emotion information.

[0079] For example, in one implementation, the feedback emotion information includes feedback tone categories, such as the following:

[0080] Declarative mood: used to state facts or opinions and express definite information;

[0081] Interrogative tone: used to ask questions, express uncertainty or information that needs confirmation;

[0082] Imperative mood: used to express commands, requests or suggestions, etc.

[0083] Exclamatory mood: used to express strong emotions, such as surprise, joy, etc.

[0084] These mood categories convey the speaker's attitude and emotions through different grammatical forms and sound changes. For example, declarative mood usually uses declarative sentences, interrogative mood uses interrogative sentences, imperative mood uses imperative sentences, and exclamatory mood uses exclamatory sentences. In terms of sound, different mood categories have subtle changes in size, strength, speed, and timbre.

[0085] In more detailed classification, tone categories include gentle, shy, sad, happy, angry, serious, lazy, hesitant, sarcastic, etc.

[0086] Exemplarily, in one implementation, the feedback emotion information includes a feedback speech speed category.

[0087] Examples of speech speed categories include: slow, medium and fast. In different scenarios, choosing an appropriate speech speed can help users better understand and convey information. For example, slow speech is suitable for expressing heavy, melancholy, sad moods or describing solemn scenes; fast speech is suitable for expressing strong emotions such as fear, anger, excitement, etc.

[0088] Exemplarily, in one implementation, the feedback emotion information includes a feedback emotion category.

[0089] When responding to the user, the dialogue system generates voice responses with corresponding emotions, making the overall user experience better.

[0090] In the actual implementation process, according to the current target emotion information of the identified user, the pre-trained large model emotion descriptor is used to give the feedback emotion information of the feedback. For example, if the user emotion recognition result is sadness, the large model emotion descriptor should make a speech behavior description such as "soothing with a soothing speed, calm tone, and gentle voice". The speech behavior description includes the feedback tone category, feedback speed category, and feedback emotion category required for the feedback.

[0091] Step 140, obtaining reply voice data according to the feedback emotion information.

[0092] Based on the feedback tone category, feedback speed category, and feedback emotion category obtained in the above steps, the content, voice intonation, speed, tone, and other dimensions of the reply to the user are fine-tuned to obtain the required and more accurate voice reply. For example, the cosyvoice model is used for voice generation, the text to be replied and the feedback emotion information are input, and the prompt word engineering prompt is used to obtain the emotional voice output.

[0093] The above-mentioned voice interaction method, by identifying the user's current target emotion information, corresponds to the target emotion information, calculates the feedback tone category, feedback speech speed category, and feedback emotion category that should be fed back, and generates reply voice data under special emotions to the user based on the feedback tone category, feedback speech speed category, and feedback emotion category, thereby improving the anthropomorphic characteristics of voice broadcasting and improving the user's experience.

[0094] In one embodiment, for step 120, target emotion information is obtained by combining text emotion recognition and speech emotion recognition. Figure 4 As shown, the following steps are included:

[0095] Step 121, obtaining corresponding conversation text data according to the voice conversation data;

[0096] Step 122, obtaining a first emotion recognition result of text emotion recognition based on the conversation text data;

[0097] Step 123, obtaining a second emotion recognition result of voice emotion recognition according to the voice dialogue data;

[0098] Step 124: Obtain target emotion information according to the first emotion recognition result and the second emotion recognition result.

[0099] For step 121, the dialogue text data is the content of the user's voice. The dialogue system receives the user's voice input as the starting point of the dialogue, and the voice input is converted into text through voice recognition technology.

[0100] Speech recognition technology pre-processes the collected speech conversation data, including denoising, enhancement, and framing, extracts features from the audio signal, and converts it into a form that can be understood by the computer. These features can effectively represent the audio characteristics of the speech signal, facilitate subsequent pattern matching, and compare the extracted features with known speech models. Commonly used models include hidden Markov models (HMMs), deep neural networks (DNNs), etc. These models can learn and recognize different speech patterns, thereby converting speech conversation data into text.

[0101] Speech recognition technology also uses complex machine learning models such as recurrent neural networks (RNN) and long short-term memory networks (LSTM), which can capture subtle differences in speech and generate more accurate text output.

[0102] For step 122, text emotion recognition is performed based on the content of the voice conversation, and the emotional tendency of the text is judged by analyzing the vocabulary, grammar and semantics in the text. In the traditional way, it is achieved through predefined rules and vocabulary. These rules usually include regular expression matching, part-of-speech tagging, syntactic analysis, etc. for specific words or phrases. With the development of deep learning technology, deep learning models can automatically extract features from texts and perform emotion classification through multi-layer neural networks, thereby improving the accuracy and efficiency of emotion recognition.

[0103] For step 123, speech emotion recognition technology identifies the emotional state of the speaker through the acoustic features of a speech (such as pitch, speaking speed, volume, pauses, etc.).

[0104] In one embodiment, in the process of speech emotion recognition, based on a discretization tool such as wav2vec, the continuous feature z is converted into a discrete feature z' through a quantization module, thereby realizing the transformation of the feature space from infinite continuous to finite discrete. After the features are extracted using a convolutional neural network, a layer attention mechanism is used to capture contextual relationships and hierarchical emotional changes in speech emotion recognition.

[0105] The speech emotion recognition process exemplarily includes the following steps:

[0106] S1, discretizes the speech dialogue data to obtain a single feature encoding vector at the word level;

[0107] S2, obtain the hidden vector at the word level according to the single feature encoding vector at the word level;

[0108] S3, based on the hidden vector at the word level, obtains local information representation based on the attention mechanism;

[0109] S4, based on the local information representation, obtain the hidden vector at the sentence level;

[0110] S5, based on the hidden vector at the sentence level, obtains the global information representation based on the attention mechanism;

[0111] S6, obtaining a second emotion recognition result based on the global information representation.

[0112] Steps S1 to S6 are described as follows:

[0113] Use Figure 5 The model structure shown introduces an attention layer at the local information level of discrete features, that is, the attention layer at the word level; and an attention layer at the global information level, that is, the attention layer at the sentence level.

[0114] Based on wav2vec, we get the tth single data l in the ith sentence after discretization. it As the input of the model, it is taken as the t-th single feature encoding vector e in the discretized i-th sentence it , the attention layer at the local information level is the t-th single feature encoding vector e in the discretized i-th sentence it , use MLP (Multilayer Perceptron) to get a word-level hidden vector h it , after Softmax (normalization), the word-level weight a is obtained it .h w is a context vector that measures the importance of local features in a sentence and is randomly initialized and learned jointly during the training process. The final local information representation g i It is a it and e it The mathematical expression is as follows:

[0115] h it =tanh(w w *e it +b w );

[0116]

[0117] g i =∑ t a it e it .

[0118] Among them, w w is the weight matrix of the multilayer perceptron in the local information level, b w is the bias of the multilayer perceptron in the local information level, and tanh(*) represents the hyperbolic tangent function.

[0119] The attention layer at the global information level and the attention layer at the local information level use the same operation to represent the local information g i As input to the global information level i :

[0120] h i =tanh(w s *e i +b s );

[0121]

[0122] s=∑ i a i e i .

[0123] Among them, w s is the weight matrix of the multilayer perceptron in the global information level, b s is the bias of the multilayer perceptron in the global information level, h i is the hidden vector at the sentence level, h s is the context vector that measures the importance of local features in the entire input, s is the global information representation, and a i is the weight of the sentence level. Finally, the second emotion recognition result of speech emotion classification is obtained according to the category of the global information representation.

[0124] In the above process, the model is forced to learn the fine-grained features of each dimension of the attention layer of the local information level, i.e. the word level, and the attention layer of the global information level, i.e. the sentence context level. This improves the flexibility of the model and effectively identifies fine-grained emotional representations.

[0125] For step 120, based on the analysis of emotion information by speech and text, each emotion category can obtain a corresponding confidence level, which is normalized as a weight, and the target emotion information is obtained based on the weight. For example, the target emotion information is the emotion category with the largest current weight.

[0126] In some implementations, for a user's continuous voice input, different sentences may contain different emotional information. For example, the first sentence is sad and the second sentence is gentle. Then these two emotional descriptions will constitute an emotion pair. Different emotions have corresponding confidence levels. The confidence levels are normalized and used as input data. They are input into a large model emotion descriptor, and feedback emotional information is calculated based on the emotion pairs, that is, each emotion category after confidence level normalization is used as the target emotional information.

[0127] For step 130, in some embodiments, the target emotion information is input into a pre-trained large model emotion descriptor; and the large model emotion descriptor is prompted according to the prompt word engineering to obtain a speech behavior description, and the speech behavior description is used as the feedback emotion information; wherein the speech behavior description is a behavior strategy including feedback tone category, feedback speech speed category, and feedback emotion category.

[0128] The large-model emotion descriptor is based on large language models such as Qwen2.5-72B and chat-GPT, combined with prompt word engineering.

[0129] Prompt Engineering is a technique for pre-trained language models that guides the model to generate high-quality, accurate, and targeted outputs by designing, experimenting, and optimizing input prompts.

[0130] In the actual implementation process, according to the analysis of emotional information of speech and text, the confidence is normalized and used as the weight to input into the large model emotional descriptor, and the large language model is prompted with the designed prompt words. According to the target emotional information, it is determined how to give feedback and give the emotional description of the feedback. For example, if the overall result of the analysis of the user's speech and text information is sadness, the descriptor should make a voice behavior description of "soothing with a soothing, peaceful, and gentle voice". The voice behavior description includes the feedback tone category, feedback speed category, feedback emotion category, and the behavior strategy for soothing. After that, when generating a specific voice response, the voice behavior description can be expanded and fine-tuned.

[0131] For step 140, in one embodiment, when generating reply voice data, the feedback emotion information and the context information in the voice conversation data are used as input, and the feedback text is obtained according to a pre-trained text generation model, wherein the text generation model is obtained by fine-tuning according to a general large language model; and the reply voice data corresponding to the feedback text is obtained according to the feedback emotion information and the feedback text.

[0132] The context information includes the voice conversation data of the current round of conversation, and may also include the voice conversation data of the previous round.

[0133] Common large language models such as Qwen2.5 72B, chat-GPT model, Transformer model, etc.

[0134] The fine-tuning process is based on LoRA technology. LoRA fine-tunes the model by training low rank matrices and then injecting these parameters into the original model. This method does not require modifying the structure of the original model and can complete the training with only a small amount of data.

[0135] In the actual implementation process, the feedback emotional information and contextual information from the analysis of emotional information on speech and text are used as input. Through the large-scale pre-trained Transformer model, loRA is used for fine-tuning. The text emotion and speech emotion judged in the early stage can be combined to generate text based on the contextual information through autoregression.

[0136] For example, if the emotion judgment received earlier is "sadness" and the content is "I'm heartbroken", the reply content is: "I know you must be feeling very bad right now. Heartbreak is like a piece missing from your heart...". The reply will generate text content based on the emotion and the voice content itself, which can better comfort the person.

[0137] When generating emotional speech, we use the prompt word engineering to use information such as emotion, tone, and speech speed as descriptive prompt information for speech generation, and generate voice responses with corresponding emotions, making the overall user experience better.

[0138] In some embodiments, by identifying the user's abnormal emotions and providing emotional responses to the abnormal emotions, unnecessary emotional output can be reduced and the amount of calculation can be reduced.

[0139] Exemplarily, based on the target emotion information, it is determined whether the user is in an abnormal emotion; abnormal emotions include "sadness, joy" and the like, and normal emotions include calmness, etc. The specific classification rules are determined according to actual needs.

[0140] When the user is in an abnormal emotion, feedback emotional information is obtained in response to the abnormal emotion. For example, for the emotion of "sadness", feedback is given in a positive tone, moderate speed, and encouraging mood; for the emotion of "joy", feedback is given in a normal tone, moderate speed, and with a hint of fun.

[0141] like Figure 6 As shown, a flow chart of the speech interaction method of the present application is provided. First, speech recognition is performed, and text emotion recognition and speech emotion recognition are performed respectively. The recognition results are input into the large model emotion descriptor to obtain the speech behavior description required for feedback. Then, the emotion recognition results and context information are integrated to generate text, and finally speech generation is performed. A reply speech with rich emotional feedback is provided.

[0142] It should be understood that although Figure 3 , Figure 4 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 3 , Figure 4 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0143] In one embodiment, Figure 7 As shown, a voice interaction device is provided, including: a voice acquisition module 210, an emotion recognition module 220, a feedback emotion description module 230 and a voice generation module 240, wherein:

[0144] The voice acquisition module 210 is used to acquire the user's voice conversation data;

[0145] The emotion recognition module 220 is used to obtain the user's current target emotion information based on the voice dialogue data;

[0146] The feedback emotion description module 230 is used to obtain the feedback emotion information of the reply based on the target emotion information, and the feedback emotion information includes the feedback tone category, feedback speech speed category, and feedback emotion category required for the feedback;

[0147] The speech generation module 240 is used to obtain reply speech data according to the feedback emotion information.

[0148] The above-mentioned voice interaction device obtains the user's current target emotion information by performing emotion recognition on the voice dialogue data, and obtains the feedback emotion information that should be used for the voice reply based on the target emotion information, including the feedback tone category, feedback speech speed category, and feedback emotion category required for the feedback. According to the feedback tone category, feedback speech speed category, and feedback emotion category, voice generation is performed, and the user can be provided with a voice reply that matches the user's current emotion, thereby enhancing the authenticity of the voice dialogue.

[0149] In one embodiment, the emotion recognition module 220 obtains corresponding conversation text data based on the voice conversation data; obtains a first emotion recognition result of text emotion recognition based on the conversation text data; obtains a second emotion recognition result of voice emotion recognition based on the voice conversation data; obtains target emotion information based on the first emotion recognition result and the second emotion recognition result, and adopts a multi-modal combination of voice and text to improve emotion recognition capabilities.

[0150] In one embodiment, the emotion recognition module 220 discretizes the voice dialogue data to obtain a single feature coding vector at the word level; obtains a hidden vector at the word level based on the single feature coding vector at the word level; obtains a local information representation based on the hidden vector at the word level based on an attention mechanism; obtains a sentence-level hidden vector based on the local information representation; obtains a global information representation based on the sentence-level hidden vector based on an attention mechanism; and obtains a second emotion recognition result based on the global information representation.

[0151] The above process forces the model to learn fine-grained features in various dimensions, improving the flexibility of the model and effectively identifying fine-grained emotion representations.

[0152] In one embodiment, the feedback emotion description module 230 inputs the target emotion information into a pre-trained large model emotion descriptor; and prompts the large model emotion descriptor according to the prompt word engineering to obtain a speech behavior description, so as to obtain reply speech data based on the speech behavior description; wherein the speech behavior description is a behavior strategy including feedback tone category, feedback speech speed category, and feedback emotion category.

[0153] In one embodiment, the speech generation module 240 takes the feedback emotion information and the context information in the speech conversation data as input, and obtains the feedback text according to a pre-trained text generation model, wherein the text generation model is obtained by fine-tuning according to a general large language model; and obtains the reply speech data corresponding to the feedback text according to the feedback emotion information and the feedback text.

[0154] In one embodiment, the emotion recognition module 220 is also used to determine whether the user is in an abnormal emotion based on the target emotion information; if the user is in an abnormal emotion, obtain feedback emotion information in response to the abnormal emotion to obtain response voice data with feedback emotion.

[0155] For the specific definition of the voice interaction device, please refer to the definition of the voice interaction method above, which will not be repeated here. Each module in the above-mentioned voice interaction device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0156] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a voice interaction method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0157] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0158] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program:

[0159] Get the user's voice conversation data;

[0160] According to the voice conversation data, obtain the user's current target emotion information;

[0161] Based on the target emotion information, feedback emotion information of the reply is obtained, where the feedback emotion information includes the feedback tone category, feedback speech speed category, and feedback emotion category required for the feedback;

[0162] According to the feedback emotion information, the reply voice data is obtained.

[0163] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0164] According to the voice conversation data, obtaining corresponding conversation text data;

[0165] Based on the conversation text data, obtaining a first emotion recognition result of text emotion recognition;

[0166] Obtaining a second emotion recognition result of speech emotion recognition according to the speech dialogue data;

[0167] Target emotion information is obtained according to the first emotion recognition result and the second emotion recognition result.

[0168] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0169] Discretize the speech dialogue data to obtain a single feature encoding vector at the word level;

[0170] According to the single feature encoding vector at the word level, a hidden vector at the word level is obtained;

[0171] According to the hidden vector at the word level, local information representation is obtained based on the attention mechanism;

[0172] According to the local information representation, the hidden vector at the sentence level is obtained;

[0173] According to the hidden vector at the sentence level, the global information representation is obtained based on the attention mechanism;

[0174] A second emotion recognition result is obtained based on the global information representation.

[0175] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0176] Obtaining local information representation includes determining the hidden vector h at the word level according to the following mathematical expression: it , the word-level weight a it , local information representation g i :

[0177] h it =tanh(w w *e it )+b w );

[0178]

[0179] g i =∑ t a it e it ;

[0180] Among them, w w is the weight matrix of the multilayer perceptron in the local information level, b w is the bias of the multilayer perceptron in the local information level, h w is a context vector that measures the importance of local features in a sentence, e it is the t-th single feature encoding vector in the ith sentence after discretization, and tanh(*) represents the hyperbolic tangent function;

[0181] Obtaining global information representation includes determining the sentence-level hidden vector h according to the following mathematical expression: i , sentence-level weight q i , global information representation s:

[0182] h i =tanh(w s *e i +b s );

[0183]

[0184] s=∑ i a i e i ;

[0185] Among them, w s is the weight matrix of the multilayer perceptron in the global information level, b s is the bias of the multilayer perceptron in the global information level; e i is the input of the global information level; h s It is a context vector that measures the importance of local features in the entire input.

[0186] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0187] Input the target emotion information into the pre-trained large model emotion descriptor;

[0188] and prompting the large model emotion descriptor according to the prompt word engineering to obtain a speech behavior description, and using the speech behavior description as the feedback emotion information;

[0189] The speech behavior is described as a behavior strategy including feedback tone category, feedback speech speed category, and feedback emotion category.

[0190] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0191] Taking the feedback emotion information and the context information in the voice conversation data as input, the feedback text is obtained according to the pre-trained text generation model, wherein the text generation model is obtained by fine-tuning the general large language model;

[0192] According to the feedback emotion information and the feedback text, the reply voice data corresponding to the feedback text is obtained.

[0193] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:

[0194] According to the target emotion information, determine whether the user is in an abnormal emotion;

[0195] When the user is in an abnormal emotion, feedback emotion information in response to the abnormal emotion is obtained to obtain response voice data with the feedback emotion.

[0196] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0197] Get the user's voice conversation data;

[0198] According to the voice conversation data, obtain the user's current target emotion information;

[0199] Based on the target emotion information, feedback emotion information of the reply is obtained, where the feedback emotion information includes the feedback tone category, feedback speech speed category, and feedback emotion category required for the feedback;

[0200] According to the feedback emotion information, the reply voice data is obtained.

[0201] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0202] According to the voice conversation data, obtaining corresponding conversation text data;

[0203] Based on the conversation text data, obtaining a first emotion recognition result of text emotion recognition;

[0204] Obtaining a second emotion recognition result of speech emotion recognition according to the speech dialogue data;

[0205] Target emotion information is obtained according to the first emotion recognition result and the second emotion recognition result.

[0206] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0207] Discretize the speech dialogue data to obtain a single feature encoding vector at the word level;

[0208] According to the single feature encoding vector at the word level, a hidden vector at the word level is obtained;

[0209] According to the hidden vector at the word level, local information representation is obtained based on the attention mechanism;

[0210] According to the local information representation, the hidden vector at the sentence level is obtained;

[0211] According to the hidden vector at the sentence level, the global information representation is obtained based on the attention mechanism;

[0212] A second emotion recognition result is obtained based on the global information representation.

[0213] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0214] Input the target emotion information into the pre-trained large model emotion descriptor;

[0215] And according to the prompt word engineering, the large model emotion descriptor is prompted to obtain the voice behavior description, so as to obtain the reply voice data based on the voice behavior description;

[0216] The speech behavior is described as a behavior strategy including feedback tone category, feedback speech speed category, and feedback emotion category.

[0217] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:

[0218] Taking the feedback emotion information and the context information in the voice conversation data as input, the feedback text is obtained according to the pre-trained text generation model, wherein the text generation model is obtained by fine-tuning the general large language model;

[0219] According to the feedback emotion information and the feedback text, the reply voice data corresponding to the feedback text is obtained.

[0220] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0221] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0222] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A voice interaction method, characterized in that: include: Get the user's voice conversation data; Obtaining the user's current target emotion information according to the voice conversation data; Based on the target emotion information, feedback emotion information of the reply is obtained, wherein the feedback emotion information includes a feedback tone category, a feedback speech speed category, and a feedback emotion category required for the feedback; According to the feedback emotion information, reply voice data is obtained.

2. The voice interaction method according to claim 1, characterized in that: According to the voice conversation data, the user's current target emotion information is obtained, including: According to the voice conversation data, obtaining corresponding conversation text data; Based on the conversation text data, obtaining a first emotion recognition result of text emotion recognition; Obtaining a second emotion recognition result of speech emotion recognition according to the speech dialogue data; The target emotion information is obtained according to the first emotion recognition result and the second emotion recognition result.

3. The voice interaction method according to claim 2, characterized in that: The step of obtaining a second emotion recognition result of speech emotion recognition according to the speech dialogue data comprises: Discretization processing is performed on the voice dialogue data to obtain a single feature encoding vector at the word level; According to the single feature encoding vector at the word level, a hidden vector at the word level is obtained; According to the hidden vector of the word level, local information representation is obtained based on the attention mechanism; According to the local information representation, the hidden vector at the sentence level is obtained; According to the hidden vector at the sentence level, a global information representation is obtained based on an attention mechanism; The second emotion recognition result is obtained based on the global information representation.

4. The voice interaction method according to claim 3, characterized in that: The obtaining of the local information representation includes determining the hidden vector h at the word level according to the following mathematical expression: it , word-level weight a it , local information representation g i : h it =tanh(w w *e it )+b w ); g i =∑ t a it yes it ; Among them, w w is the weight matrix of the multilayer perceptron in the local information level, b w is the bias of the multilayer perceptron in the local information level, h w is a context vector that measures the importance of local features in a sentence, e it is the t-th single feature encoding vector in the ith sentence after discretization, and tanh(*) represents the hyperbolic tangent function; The obtaining of the global information representation includes determining the hidden vector h at the sentence level according to the following mathematical expression: i , sentence-level weight a i , global information representation s: h i =tanh(w s *e i +b s ); s=∑ i the i and i ; Among them, w s is the weight matrix of the multilayer perceptron in the global information level, b s is the bias of the multilayer perceptron in the global information level; e i is the input of the global information level; h s It is a context vector that measures the importance of local features in the entire input.

5. The voice interaction method according to claim 1, characterized in that: The step of obtaining the feedback emotion information of the reply based on the target emotion information includes: Inputting the target emotion information into a pre-trained large model emotion descriptor; and prompting the large model emotion descriptor according to the prompt word engineering to obtain a speech behavior description, and using the speech behavior description as the feedback emotion information; The speech behavior description is a behavior strategy including the feedback tone category, feedback speech speed category, and feedback emotion category.

6. The voice interaction method according to claim 1, characterized in that: The step of obtaining reply voice data according to the feedback emotion information includes: Taking the feedback emotion information and the context information in the voice dialogue data as input, obtaining feedback text according to a pre-trained text generation model, wherein the text generation model is obtained by fine-tuning a general large language model; The reply voice data corresponding to the feedback text is obtained according to the feedback emotion information and the feedback text.

7. The voice interaction method according to claim 1, characterized in that: After obtaining the current target emotion information of the user, the method further includes: Determining whether the user is in an abnormal emotion according to the target emotion information; When the user is in an abnormal emotion, feedback emotion information in response to the abnormal emotion is obtained to obtain response voice data with the feedback emotion.

8. A voice interaction device, characterized in that: The device comprises: A voice acquisition module is used to acquire the user's voice conversation data; An emotion recognition module, used to obtain the user's current target emotion information based on the voice conversation data; A feedback emotion description module, used to obtain the feedback emotion information of the reply based on the target emotion information, wherein the feedback emotion information includes the feedback tone category, feedback speech speed category, and feedback emotion category required for the feedback; The speech generation module is used to obtain reply speech data according to the feedback emotion information.

9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Data processing method and device, computer equipment and storage medium

    CN110688499A

  • User emotion recognition and reply method, system and device and storage medium

    CN113676600A

  • Speech emotion layered recognition method and system based on phoneme level

    CN114360584A

  • Voice emotion recognition method and device

    CN116259307A

  • Intelligent calling method based on large language model

    CN117336410A