Emotion-based voice interaction method and device, electronic equipment and storage medium

By recognizing users' emotions and personality types, and generating prompts with fine-grained emotion tags, this technology solves the problem that existing voice interaction methods struggle to achieve accurate emotion reporting under low latency, thus improving the real-time performance and accuracy of voice interaction.

CN121148383APending Publication Date: 2025-12-16IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511322777.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing voice interaction methods struggle to achieve accurate and coherent emotional voice delivery while maintaining low-latency interaction, thus impacting user experience.

Method used

By acquiring users' voice data, identifying users' emotional types and a set of preset speaker personality types, determining prompts with emotional tags, and inputting them into the first main model for streaming output, generating fine-grained emotionally tagged response text, and finally synthesizing the response speech.

Benefits of technology

It improves the accuracy of emotional voice broadcasting and the real-time nature of voice interaction, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121148383A_ABST
    Figure CN121148383A_ABST
Patent Text Reader

Abstract

The invention provides an emotion-based voice interaction method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining voice data of a user, recognizing an emotion type of the user, and determining a character type set of a preset speaker; based on the voice data, the emotion type and the character type set, determining a prompt instruction prompt with an emotion label; inputting a prompt instruction prompt into the first large model to obtain a streaming output reply text with a fine-grained emotion tag; and synthesizing reply voice based on the reply text, and broadcasting the reply voice. According to the emotion-based voice interaction method provided by the invention, the prompt instruction prompt with the emotion label is generated by integrating multi-dimensional information, the macroscopic emotion basic tone of the reply text is determined in advance, and on the basis, the reply text with the fine-grained emotion label is generated based on the prompt instruction prompt, so that the user experience is improved. Therefore, the accuracy of emotional voice broadcasting and the real-time performance of interaction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and in particular to a voice interaction method and device based on emotion, an electronic device and a storage medium. BACKGROUND

[0002] In modern human-computer interaction, giving rich and accurate emotion to speech synthesis can greatly improve user experience. At present, the existing voice interaction method based on emotion guarantees the accuracy of emotion prediction by accumulating long text, which affects the response speed; guarantees the fluency of interaction by real-time emotion prediction on streaming short text, which affects the accuracy of emotional speech broadcasting. Therefore, how to realize accurate and coherent emotional speech broadcasting under the premise of ensuring low-latency interaction is a technical problem to be solved in the field. SUMMARY

[0003] The present application provides a voice interaction method and device based on emotion, an electronic device and a storage medium to solve the defects that the existing voice interaction method is difficult to balance the accuracy and real-time performance of emotional speech broadcasting.

[0004] The present application provides a voice interaction method based on emotion, comprising: Obtaining voice data of a user, identifying the emotion type of the user, and determining a set of personality types of a preset voice actor; Based on the voice data, emotion type and set of personality types, determining a prompt instruction prompt with an emotional label; Inputting the prompt instruction prompt into a first large model to obtain a reply text with a fine-grained emotional label output by the first large model in a streaming manner; Synthesizing a reply voice based on the reply text and broadcasting the reply voice.

[0005] In some embodiments, the determination of the prompt instruction prompt with an emotional label based on the voice data, emotion type and set of personality types comprises: Identifying the voice data to obtain an identified text; Inputting the identified text, set of personality types and emotion type into a pre-constructed emotion recommendation model to obtain a prompt instruction prompt with an emotional label output by the emotion recommendation model; The emotion recommendation model is trained based on sample identified text, sample set of personality types and sample emotion type, and corresponding prompt instruction prompt label.

[0006] In some embodiments, after broadcasting the reply voice, the method further comprises: Obtaining feedback information of the user; identify a new emotion type of the user based on the feedback information; determine an emotion change of the user based on the emotion type and the new emotion type; adjust parameters of the emotion recommendation model based on the emotion change of the user, to obtain an adjusted emotion recommendation model.

[0007] In some embodiments, after obtaining the adjusted emotion recommendation model, the method further comprises: inputting the new recognized text and the new emotion type of the new voice data of the user, and the personality type set or the updated personality type set into the adjusted emotion recommendation model, to obtain a new prompt with a new emotion label output by the adjusted emotion recommendation model.

[0008] In some embodiments, obtaining the prompt with an emotion label output by the emotion recommendation model comprises: matching the emotion type with the personality type set based on the emotion recommendation model, determining a target personality type from the personality type set, generating the prompt with an emotion label according to the recognized text, the target personality type and the emotion type, and outputting the prompt.

[0009] In some embodiments, generating the prompt with an emotion label according to the recognized text, the target personality type and the emotion type comprises: obtaining a recommended emotion label based on the target personality type and the emotion type; generating the prompt based on the recognized text, the target personality type, the emotion type and the emotion label.

[0010] In some embodiments, identifying the emotion type of the user comprises: obtaining video data of the user; performing feature extraction on the video data to obtain facial expression features and body movement features of the user; performing feature extraction on the voice data to obtain voice features of the user; identifying the emotion type of the user based on the facial expression features, the body movement features and the voice features.

[0011] In some embodiments, the training process of the emotion recommendation model comprises: obtaining a plurality of sample voice data, performing recognition on the plurality of sample voice data to obtain a plurality of sample recognized texts, determining a plurality of sample personality type sets of sample speakers, and obtaining a plurality of sample emotion types; construct a training data set based on the plurality of sample recognized texts, the plurality of sample personality type sets and the plurality of sample emotion types, and determine a prompt instruction prompt label set with emotion labels; train an initial emotion recommendation model based on the training data set and the prompt instruction prompt label set, and obtain the emotion recommendation model after the training is completed.

[0012] In some embodiments, the constructing of the training data set based on the plurality of sample recognized texts, the plurality of sample personality type sets and the plurality of sample emotion types comprises: randomly matching the plurality of sample recognized texts, the plurality of sample personality type sets and the plurality of sample emotion types based on a second large model to obtain a plurality of training data, each piece of training data comprising a sample recognized text, a sample personality type set and a sample emotion type; constructing a training data set based on the plurality of training data.

[0013] The application further provides a voice interaction device based on emotion, comprising: an acquisition unit configured to acquire voice data of a user, recognize an emotion type of the user, and determine a personality type set of a preset voice person; a determination unit configured to determine a prompt instruction prompt with emotion labels based on the voice data, the emotion type and the personality type set; a reply unit configured to input the prompt instruction prompt into a first large model to obtain a reply text with fine-grained emotion labels that is output by the first large model in a streaming manner; a voice broadcast unit configured to synthesize a reply voice based on the reply text and broadcast the reply voice.

[0014] The application further provides an electronic device comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor implements the voice interaction method based on emotion as described above when executing the program.

[0015] The application further provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executable on a processor to implement the voice interaction method based on emotion as described above.

[0016] The emotion-based voice interaction method, device, electronic equipment and storage medium provided by the application, by acquiring voice data of a user, identifying the emotion type of the user, determining a set of personality types of a preset voice actor; based on the voice data, emotion type and set of personality types, determining a prompt instruction prompt with an emotion label, determining the macro emotional tone of the reply text in advance; inputting the prompt instruction prompt into a first large model, and outputting the reply text with fine-grained emotion labels from the first large model in a streaming manner, synthesizing the reply voice based on the reply text, and playing the reply voice, thereby improving the accuracy of emotion voice playing and the real-time performance of voice interaction. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0018] Figure 1 is a flowchart of the emotion-based voice interaction method provided by the prior art.

[0019] Figure 2 is one of the flowcharts of the emotion-based voice interaction method provided by the embodiments of the application.

[0020] Figure 3 is the second flowchart of the emotion-based voice interaction method provided by the embodiments of the application.

[0021] Figure 4 is the flowchart of the training process of the emotion recommendation model provided by the embodiments of the application.

[0022] Figure 5 is the structural schematic diagram of the emotion-based voice interaction device provided by the embodiments of the application.

[0023] Figure 6 is the structural schematic diagram of the electronic equipment provided by the embodiments of the application. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the application more clear, the technical solutions in the application will be described clearly and completely below in combination with the drawings in the application. Obviously, the described embodiments are part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.

[0025] The terms "first", "second", and the like in the present disclosure are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second" are generally a class, not limited to the number of objects, for example, the first object can be one or more.

[0026] Figure 1 A flowchart of the prior art emotion-based voice interaction method is provided. As shown in Figure 1 The prior art provides an emotion-based voice interaction method, which includes: obtaining voice data of a user, identifying the voice data to obtain identified text; inputting the identified text into a large model to stream the reply text from the large model; inputting the reply text into an emotion prediction model to obtain the reply text with emotion labels output by the emotion prediction model; synthesizing a reply voice based on the reply text and playing the reply voice.

[0027] Currently, the emotion prediction model needs to rely on a long text to obtain an accurate prediction result; accumulating the stream-delivered text to a certain amount of text, and then inputting the long text into the emotion prediction model can improve the accuracy of emotion prediction, but the delay is large, which affects the real-time performance of the broadcast response; directly inputting the stream-delivered text into the emotion prediction model can ensure the real-time performance of the broadcast response, but the emotion prediction effect is poor, which easily leads to inaccurate broadcast emotion and emotion jump between sentences.

[0028] Therefore, the present embodiment provides an emotion-based voice interaction method, device, electronic equipment and storage medium, which determines a prompt instruction prompt with an emotion label based on the voice data of the user, the emotion type and the personality type set of the preset voice actor; inputs the prompt instruction prompt into a first large model to stream the reply text with a fine-grained emotion label from the first large model, synthesizes a reply voice based on the reply text, and plays the reply voice. The present embodiment can improve the accuracy of emotion voice broadcast and the real-time performance of voice interaction.

[0029] Figure 2 A flowchart of the emotion-based voice interaction method provided by the present embodiment is shown in Figure 2 As shown in the figure, an emotion-based voice interaction method is provided, which includes the following steps: step 210, step 220, step 230 and step 240. The method flow steps are only used as one possible implementation of the present disclosure.

[0030] Step 210, obtaining voice data of a user, identifying the emotion type of the user, and determining a personality type set of a preset voice actor.

[0031] wherein the voice data includes language content, pitch, volume, speech rate, pause, tone, etc. information; the preset voice person refers to an AI virtual image or voice for voice interaction with the user, and the user can pre-set the role of the voice person, for example, an assistant named "Xiaoming"; the personality type set is a set of personality labels set for the preset voice person, such as {tender and considerate type, professional and rigorous type, lively and cheerful type}, and the user can pre-set the role of the voice person and the corresponding personality type set.

[0032] Optionally, according to the current configuration or the user's selection, the personality type set of the preset voice person is determined from the preset personality library.

[0033] Optionally, the voice data of the user is captured in real time through a microphone, a camera or the like, and the voice data is pre-processed to obtain pre-processed voice data; the pre-processed voice data is recognized through automatic speech recognition technology to obtain recognized text.

[0034] Optionally, the voice data and / or video data of the user are processed and analyzed through emotion recognition technology to identify the emotional type of the user in real time, such as happy, sad, angry, neutral, etc.

[0035] Step 220, determining a prompt instruction prompt with an emotional label based on the voice data, the emotional type and the personality type set; Optionally, the voice data is recognized to obtain recognized text; a target personality type matching the emotional type and the recognized text is determined from the personality type set; a prompt instruction prompt with an emotional label is determined based on the recognized text, the emotional type and the target personality type.

[0036] wherein the emotional label is an explicit emotional instruction indicating the macro emotional tone of the reply text, such as [calm down], [happy], [sorry], [confirm], etc.; the prompt instruction Prompt includes a specific task description for guiding the semantic direction of the content generated by the large model.

[0037] Optionally, the prompt instruction Prompt at least includes the recognized text of the voice data of the user, the target personality type of the preset voice person, the emotional type of the user and the emotional label.

[0038] For example, the prompt instruction Prompt is: the user inputs "How is the weather today?" in voice, the personality of the preset voice person is the cute type, and the current user is a little sad, please help me reply to the user, and the emotional label is carried when streaming.

[0039] Step 230, inputting the prompt instruction prompt into the first large model to obtain a reply text with a fine-grained emotional label output by the first large model.

[0040] The first large model is a large language model specially fine-tuned. Unlike general large models, the model can not only understand and execute complex natural language instructions (i.e., prompt instructions), but also be specially trained to output fine-grained emotion labels corresponding to the generated text.

[0041] The first large model is trained based on sample prompt instructions with sample emotion labels and reply text labels with fine-grained emotion labels. The sample prompt instructions at least include: sample recognized text of sample voice data of a sample user, sample personality type of a sample speaker, sample emotion type of the sample user, and sample emotion labels.

[0042] The fine-grained emotion label is an instruction label that acts on a small text unit (such as a word, phrase, or sentence) and is used to accurately guide the emotional performance of voice synthesis.

[0043] Optionally, the reply text includes multiple text units, and different text units have different fine-grained emotion labels. For example, the reply text is: [Softly] I hear a little cloud in your voice. [Cute] But that's okay, let me help you blow it away! [Happy] It's a big sunny day outside today, the sun is shining like a big smiley face, [Encouraging] just waiting for you to go out and say hello to it! It should be noted that the streaming output means outputting the reply text word by word or phrase by phrase; the subsequent voice synthesis module does not need to wait for the first large model to generate a complete sentence. As long as the first text segment with fine-grained emotion labels is received, the voice synthesis can start immediately, thereby improving the real-time performance of voice interaction.

[0044] Optionally, the prompt instruction prompt is input into the first large model, the prompt instruction prompt is parsed based on the first large model to obtain a parsed prompt instruction prompt, and the reply text with fine-grained emotion labels is generated based on the parsed prompt instruction prompt.

[0045] It can be understood that by inputting the prompt instruction prompt into the first large model, the reply text with fine-grained emotion labels output by the first large model in a streaming manner can greatly reduce the interaction delay, improve the expressiveness and personification of voice broadcasting, and improve the user experience.

[0046] Step 240, synthesizing a reply voice based on the reply text and broadcasting the reply voice.

[0047] Optionally, the reply voice is synthesized based on the reply text and the reply voice is broadcasted in a streaming manner.

[0048] Specifically, the voice synthesis module receives a data stream of the reply text output by the first large model in real time; the received data segment is quickly parsed to separate the fine-grained emotional label and the corresponding text content; a synthesis strategy matched with the fine-grained emotional label is called to synthesize the reply voice corresponding to the data segment.

[0049] For example, a data segment "[happy] today the weather is really good!" is received. It is identified that the emotional instruction is [happy] and the text to be synthesized is "today the weather is really good!". Higher average pitch, faster speech rate, and larger volume fluctuation are used to synthesize "today the weather is really good!".

[0050] In the embodiment of the application, the voice data of the user is obtained, the emotional type of the user is identified, and the personality type set of the preset voice is determined; based on the voice data, the emotional type, and the personality type set, a prompt instruction prompt with an emotional label is determined; the prompt instruction prompt is input to the first large model, the first large model is used to stream output a reply text with a fine-grained emotional label, a reply voice is synthesized based on the reply text, the reply voice is played, and the accuracy of emotional voice playing and the real-time performance of voice interaction are improved.

[0051] In some embodiments, determining the prompt instruction prompt with the emotional label based on the voice data, the emotional type, and the personality type set comprises: identifying the voice data to obtain an identified text; inputting the identified text, the personality type set, and the emotional type to a pre-constructed emotional recommendation model to obtain a prompt instruction prompt with an emotional label output by the emotional recommendation model; The emotional recommendation model is obtained by training based on sample identified texts, sample personality type sets, and sample emotional types, and corresponding prompt instruction prompts.

[0052] The emotional recommendation model integrates multiple input information (what the user said, how the user feels, and the role setting of the AI itself) to recommend or decide which emotion the AI should use to respond next to achieve the best interaction effect. The emotional recommendation model changes the traditional emotion prediction task into a decision-making task based on rich context. Through pre-training, the model can learn how to make the most appropriate emotional response in various complex situations.

[0053] In the embodiment of the application, the identified text, the set of personality types and the emotional type are input into a pre-constructed emotional recommendation model to obtain a prompt instruction prompt with an emotional label output by the emotional recommendation model, thereby improving the efficiency, accuracy and emotional richness of the prompt instruction generation, laying a foundation for subsequent rapid and accurate generation of reply texts with fine-grained emotional labels.

[0054] Figure 3 Fig. 2 is a flowchart of a method for emotional-based voice interaction according to an embodiment of the application. Figure 3 As shown in the figure, in some embodiments, a method for emotional-based voice interaction is provided, which comprises the following steps: Obtaining voice data of a user, identifying the voice data to obtain an identified text, determining a set of personality types of a preset voice person, and identifying an emotional type of the user; Inputting the identified text, the set of personality types and the emotional type into a pre-constructed emotional recommendation model to obtain a prompt instruction prompt with an emotional label output by the emotional recommendation model; Inputting the prompt instruction prompt into a first large model to obtain a reply text with fine-grained emotional labels output by the first large model in a streaming manner; Synthesizing a reply voice based on the reply text and playing the reply voice; Obtaining feedback information of the user; Identifying a new emotional type of the user based on the feedback information; Determining an emotional change of the user based on the emotional type and the new emotional type; Adjusting parameters of the emotional recommendation model based on the emotional change of the user to obtain an adjusted emotional recommendation model.

[0055] The preset voice person refers to an AI virtual image or voice for voice interaction with the user; the set of personality types is a set of personality labels pre-set for the preset voice person, such as {gentle and considerate type, professional and rigorous type, lively and cheerful type}.

[0056] Optionally, the set of personality types of the preset voice person is determined based on user input or according to system configuration.

[0057] The emotional recommendation model is trained based on sample identified texts, sample sets of personality types and sample emotional types, and corresponding prompt instruction prompt labels.

[0058] The emotional recommendation model is a decision model; the emotional label is an explicit emotional instruction, indicating the macro emotional tone of the reply text.

[0059] The fine-grained emotion label is a kind of instruction label that acts on a micro text unit and is used for accurately guiding the emotion performance of speech synthesis.

[0060] Optionally, the reply text includes a plurality of text units, and different text units are provided with different fine-grained emotion labels.

[0061] Optionally, the reply speech is synthesized based on the reply text, and the reply speech is played in a streaming manner.

[0062] Optionally, after the reply speech is played, feedback information of the user is obtained, the feedback information including new speech data and / or new video data of the user; and a new emotion type of the user is identified based on the new speech data and / or the new video data.

[0063] The emotion change of the user can be a positive change, a negative change or no change.

[0064] Optionally, a reward value is calculated based on the emotion change of the user, and parameters of the emotion recommendation model are adjusted according to the reward value.

[0065] If the emotion change is positive: this indicates that the decision made by the emotion recommendation model last time (for example, the Prompt strategy of [playful]+[encouragement] is recommended when the user is sad) is successful. The neural network weight that produces this decision is slightly enhanced, so that it is more likely to use this successful strategy again in the future when encountering similar situations.

[0066] If the emotion change is negative or no change: this indicates that the last strategy is a failure or ineffective. The neural network weight that produces the decision is slightly weakened to reduce the probability of repeating the strategy in a similar situation in the future.

[0067] It should be noted that, before the reply text is generated, the emotion recommendation model can make an emotion decision to obtain a prompt instruction prompt with an emotion label, the prompt instruction prompt is input to the first large model, and then the reply text with the fine-grained emotion label that is output in a streaming manner by the first large model can be directly obtained, without the need for emotion prediction.

[0068] In the embodiments of the present application, by identifying the new emotion type of the user, determining the emotion change of the user based on the emotion type and the new emotion type, and adjusting the parameters of the emotion recommendation model based on the emotion change of the user, the emotion recommendation model can be corrected and improved in actual application, the deficiencies of initial training are made up, and the emotion recommendation model can make more reasonable and appropriate responses when facing various complex, unexpected user emotions and situations.

[0069] In some embodiments, after obtaining the adjusted emotion recommendation model, the method further includes: The new recognized text and the new emotion type of the new voice data of the user, and the personality type set or the updated personality type set are input into the adjusted emotion recommendation model to obtain a new prompt instruction prompt with a new emotion label output by the adjusted emotion recommendation model.

[0070] Optionally, based on the user feedback information, the personality type set is updated, such as deleting, modifying, or adding at least one personality type, to obtain an updated personality type set. For example, the user continuously makes negative emotional feedback on the "playful" personality, and the "playful" personality is deleted from the personality type set.

[0071] Optionally, the input of the adjusted emotion recommendation model can be "new recognized text, new emotion type, and personality type set", or "new recognized text, new emotion type, and updated personality type set".

[0072] Optionally, the new prompt instruction prompt with the new emotion label is input into the first large model to obtain a new reply text with a new fine-grained emotion label output by the first large model in a streaming manner; and a new reply voice is synthesized based on the new reply text, and the new reply voice is played.

[0073] It can be understood that a new round of emotion recommendation is performed based on the adjusted emotion recommendation model, so that the accuracy and personalization level of emotional voice playing can be improved.

[0074] In some embodiments, the prompt instruction prompt with an emotion label output by the emotion recommendation model comprises: Based on the emotion recommendation model, the emotion type is matched with the personality type set, a target personality type is determined from the personality type set, and a prompt instruction prompt with an emotion label is generated according to the recognized text, the target personality type, and the emotion type, and the prompt instruction prompt is output.

[0075] Optionally, the emotion type is matched with the personality type set to obtain a matching degree of the emotion type and a plurality of personality types, and a target personality type is determined from the plurality of personality types. For example, the emotion type of the user is "sorrow", and the matching target personality type is "playful".

[0076] It can be understood that by matching the emotion type with the personality type set, a target personality type is determined from the personality type set, and a prompt instruction prompt with an emotion label is generated according to the recognized text, the target personality type, and the emotion type, the robustness and personalization level of voice interaction are improved, and the user experience is improved.

[0077] In some embodiments, the prompt instruction prompt with an emotion label is generated according to the recognized text, the target personality type, and the emotion type, comprising: Based on the target personality type and emotion type, recommended emotion tags are obtained; Based on the identified text, target personality type and emotion type, as well as emotion tags, a prompt instruction is generated.

[0078] Optionally, based on personality type and emotion type, multiple candidate emotion tags are recommended, and the final emotion tag is selected from the multiple candidate emotion tags.

[0079] In some embodiments, identifying the user's emotion type includes: Obtain user video data; Feature extraction is performed on video data to obtain the user's facial expression features and body movement features; Feature extraction is performed on the voice data to obtain the user's voice features; Based on facial expression features, body movement features, and voice features, the system identifies the user's emotion type.

[0080] Optionally, facial expression features, body movement features, and voice features can be input into the emotion recognition model to obtain the user's emotion type output by the emotion recognition model.

[0081] Understandably, identifying users' emotion types based on facial expression features, body movement features, and voice features improves the accuracy of emotion recognition.

[0082] Figure 4 This is a flowchart illustrating the training process of the sentiment recommendation model provided in an embodiment of the present invention, as shown below. Figure 4 As shown, in some embodiments, the training process of the sentiment recommendation model includes: Step 410: Acquire multiple sample speech data, recognize multiple sample speech data to obtain multiple sample recognized text, determine multiple sample personality type sets of sample speakers, and obtain multiple sample emotion types; Step 420: Based on multiple sample recognition texts, multiple sample personality type sets, and multiple sample emotion types, construct a training dataset and determine the prompt label set with emotion labels; Step 430: Based on the training dataset and the prompt label set, train the initial sentiment recommendation model. After training, the sentiment recommendation model is obtained.

[0083] Optionally, the sample recognition text, sample personality type set, and sample sentiment type are input into a pre-built initial sentiment recommendation model to obtain a prediction prompt with predicted sentiment label output by the initial sentiment recommendation model.

[0084] Optionally, based on the predicted prompt and the prompt label, a loss function value is calculated, and based on the loss function value, parameters of the initial sentiment recommendation model are iteratively optimized to obtain the sentiment recommendation model.

[0085] In the embodiment of the application, by identifying the plurality of sample recognition texts, the plurality of sample personality type sets and the plurality of sample sentiment types, a training data set is constructed, and a prompt label set with a sentiment label is determined; based on the training data set and the prompt label set, an initial sentiment recommendation model is trained, and after the training is completed, the sentiment recommendation model is obtained, thereby improving the robustness of the sentiment recommendation model.

[0086] In some embodiments, based on the plurality of sample recognition texts, the plurality of sample personality type sets and the plurality of sample sentiment types, a training data set is constructed, comprising: Based on the second large model, the plurality of sample recognition texts, the plurality of sample personality type sets and the plurality of sample sentiment types are randomly matched to obtain a plurality of training data, each piece of training data comprising a sample recognition text, a sample personality type set and a sample sentiment type; Based on the plurality of training data, a training data set is constructed.

[0087] The second large model is a general large model.

[0088] In the embodiment of the application, based on the second large model, the plurality of sample recognition texts, the plurality of sample personality type sets and the plurality of sample sentiment types are randomly matched to obtain a plurality of training data, which can quickly generate a large amount of training data, improve the efficiency of training the sentiment recommendation model, reduce the cost of training the sentiment recommendation model, and enhance the generalization ability and robustness of the sentiment recommendation model.

[0089] The following describes the emotion-based voice interaction device provided by the embodiment of the application, and the emotion-based voice interaction device described below can be correspondingly referred to the emotion-based voice interaction method described above.

[0090] Figure 5 The structure schematic diagram of the emotion-based voice interaction device provided by the embodiment of the application is shown in Figure 5 As shown in the figure, the emotion-based voice interaction device 500 comprises: An acquisition unit 510 is configured to acquire voice data of a user, identify a sentiment type of the user, and determine a personality type set of a preset voice actor. A determination unit 520 is configured to determine a prompt with a sentiment label based on the voice data, the sentiment type and the personality type set. The reply unit 530 is configured to input the prompt instruction prompt into the first large model to obtain reply text with a fine-grained emotional label output by the first large model in a streaming manner. The voice broadcast unit 540 is configured to synthesize reply voice based on the reply text and broadcast the reply voice.

[0091] Optionally, the prompt instruction prompt with the emotional label is determined based on the voice data, the set of emotional types, and the set of personality types, and includes the following steps. The voice data is recognized to obtain recognized text. The recognized text, the set of emotional types, and the set of personality types are input into a pre-constructed emotional recommendation model to obtain a prompt instruction prompt with an emotional label output by the emotional recommendation model. The emotional recommendation model is trained based on sample recognized text, a sample set of personality types, and a sample emotional type, and a corresponding prompt instruction prompt label.

[0092] Optionally, the voice interaction device based on emotions further includes a model adjustment unit configured to: Obtain feedback information of the user. Identify a new emotional type of the user based on the feedback information. Determine an emotional change of the user based on the emotional type and the new emotional type. Adjust parameters of the emotional recommendation model based on the emotional change of the user to obtain an adjusted emotional recommendation model.

[0093] In some embodiments, the voice interaction device based on emotions further includes: An emotional recommendation unit configured to input a new recognized text of new voice data of the user and a new emotional type, and a set of personality types or an updated set of personality types into the adjusted emotional recommendation model to obtain a new prompt instruction prompt with a new emotional label output by the adjusted emotional recommendation model.

[0094] Optionally, obtaining the prompt instruction prompt with the emotional label output by the emotional recommendation model includes: Based on the emotional recommendation model, the emotional type is matched with the set of personality types to determine a target personality type from the set of personality types, and a prompt instruction prompt with an emotional label is generated based on the recognized text, the target personality type, and the emotional type, and the prompt instruction prompt is output.

[0095] Optionally, the prompt instruction prompt with the emotional label is generated based on the recognized text, the target personality type, and the emotional type, and includes the following steps. A recommended emotional label is obtained based on the target personality type and the emotional type. Generate a prompt instruction based on the recognized text, the target personality type and the emotion type, and the emotion label.

[0096] Optionally, the emotion type of the user is identified, including: Obtain video data of the user; Feature extraction is performed on the video data to obtain facial expression features and body movement features of the user; Feature extraction is performed on the voice data to obtain voice features of the user; The emotion type of the user is identified based on the facial expression features, the body movement features and the voice features.

[0097] Optionally, the training process of the emotion recommendation model includes: Obtain a plurality of sample voice data, identify the plurality of sample voice data to obtain a plurality of sample recognized texts, determine a plurality of sample personality type sets of sample speakers, and obtain a plurality of sample emotion types; Based on the plurality of sample recognized texts, the plurality of sample personality type sets and the plurality of sample emotion types, a training data set is constructed, and a prompt instruction prompt label set with emotion labels is determined; Based on the training data set and the prompt instruction prompt label set, an initial emotion recommendation model is trained, and after the training is completed, the emotion recommendation model is obtained.

[0098] Optionally, based on the plurality of sample recognized texts, the plurality of sample personality type sets and the plurality of sample emotion types, the training data set is constructed, including: Based on the second large model, the plurality of sample recognized texts, the plurality of sample personality type sets and the plurality of sample emotion types are randomly matched to obtain a plurality of training data, each training data including a sample recognized text, a sample personality type set and a sample emotion type; Based on the plurality of training data, the training data set is constructed.

[0099] It should be noted that the emotion-based voice interaction device provided by the embodiment of the present application can realize all method steps realized by the above-mentioned emotion-based voice interaction method embodiment, and can achieve the same technical effects. The same parts and beneficial effects in this embodiment as the method embodiment will not be described in detail.

[0100] Figure 6 The structural schematic diagram of the electronic device provided by the embodiment of the present application is as follows: Figure 6As shown, the electronic device can include a processor 610, a communications interface 620, a memory 630, and a communications bus 640, wherein the processor 610, the communications interface 620, and the memory 630 complete mutual communication through the communications bus 640. The processor 610 can invoke a logic instruction in the memory 630 to execute an emotion-based voice interaction method, which includes: obtaining voice data of a user, identifying an emotion type of the user, determining a personality type set of a preset speaker; determining a prompt instruction prompt with an emotion label based on the voice data, the emotion type, and the personality type set; inputting the prompt instruction prompt to a first large model to obtain a reply text with a fine-grained emotion label that is streamed out by the first large model; synthesizing a reply voice based on the reply text, and playing the reply voice.

[0101] In addition, the logic instruction in the memory 630 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0102] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement an emotion-based voice interaction method provided by the above-mentioned methods, which includes: obtaining voice data of a user, identifying an emotion type of the user, determining a personality type set of a preset speaker; determining a prompt instruction prompt with an emotion label based on the voice data, the emotion type, and the personality type set; inputting the prompt instruction prompt to a first large model to obtain a reply text with a fine-grained emotion label that is streamed out by the first large model; synthesizing a reply voice based on the reply text, and playing the reply voice.

[0103] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0104] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and the necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0105] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An emotion-based voice interaction method, characterized in that, include: Acquire user's voice data, identify the user's emotional type, and determine a preset set of speaker personality types; Based on the voice data, emotion type, and personality type set, a prompt instruction with an emotion tag is determined; The prompt instruction is input into the first large model to obtain the reply text with fine-grained sentiment tags, which is streamed from the first large model. A reply voice is synthesized based on the reply text and then broadcast.

2. The emotion-based voice interaction method according to claim 1, characterized in that, The step of determining a prompt instruction with an emotion tag based on the voice data, emotion type, and personality type set includes: The voice data is recognized to obtain the recognized text; The identified text, personality type set, and emotion type are input into a pre-built emotion recommendation model to obtain a prompt instruction with emotion tag output by the emotion recommendation model; The emotion recommendation model is trained based on sample recognition text, sample personality type set and sample emotion type, as well as corresponding prompt labels.

3. The emotion-based voice interaction method according to claim 2, characterized in that, After broadcasting the reply voice, the following is also included: Obtain the user's feedback information; Based on the feedback information, the user's new emotional type is identified; Based on the stated emotion type and the new emotion type, the user's emotional changes are determined; Based on the user's emotional changes, the parameters of the emotional recommendation model are adjusted to obtain the adjusted emotional recommendation model.

4. The emotion-based voice interaction method according to claim 3, characterized in that, After obtaining the adjusted sentiment recommendation model, the following is also included: The newly recognized text and new emotion type of the user's new voice data, as well as the personality type set or the updated personality type set, are input into the adjusted emotion recommendation model to obtain a new prompt instruction with a new emotion tag output by the adjusted emotion recommendation model.

5. The emotion-based voice interaction method according to claim 2, characterized in that, The prompt instruction with sentiment tag output by the sentiment recommendation model includes: Based on the emotion recommendation model, the emotion type is matched with the set of personality types, the target personality type is determined from the set of personality types, and a prompt instruction with an emotion tag is generated and output based on the identified text, the target personality type and the emotion type.

6. The emotion-based voice interaction method according to claim 5, characterized in that, The step of generating a prompt instruction with an emotion tag based on the identified text, target personality type, and emotion type includes: Based on the target personality type and emotion type, recommended emotion tags are obtained; Based on the identified text, target personality type and emotion type, and the emotion tag, the prompt instruction is generated.

7. The emotion-based voice interaction method according to claim 1, characterized in that, The identification of the user's emotional type includes: Obtain the user's video data; Feature extraction is performed on the video data to obtain the user's facial expression features and body movement features; The user's voice features are obtained by extracting features from the voice data; Based on the facial expression features, body movement features, and voice features, the user's emotion type is identified.

8. The emotion-based voice interaction method according to claim 2, characterized in that, The training process of the sentiment recommendation model includes: Multiple sample speech data are acquired, the multiple sample speech data are recognized to obtain multiple sample recognized text, multiple sample personality type sets of sample speakers are determined, and multiple sample emotion types are obtained. Based on the multiple sample recognition texts, multiple sample personality type sets, and multiple sample emotion types, a training dataset is constructed, and a set of prompt labels with emotion tags is determined. Based on the training dataset and the prompt label set, an initial sentiment recommendation model is trained, and after training, the sentiment recommendation model is obtained.

9. The emotion-based voice interaction method according to claim 8, characterized in that, The training dataset is constructed based on the multiple sample text recognitions, multiple sample personality type sets, and multiple sample emotion types, including: Based on the second major model, the multiple sample recognition texts, multiple sample personality type sets, and multiple sample emotion types are randomly matched to obtain multiple training data. Each training data includes a sample recognition text, a sample personality type set, and a sample emotion type. Based on the aforementioned training data, a training dataset is constructed.

10. An emotion-based voice interaction device, characterized in that, include: The acquisition unit is used to acquire the user's voice data, identify the user's emotional type, and determine a preset set of speaker personality types. The determining unit is used to determine a prompt instruction with an emotion tag based on the voice data, the emotion type and the personality type set; The response unit is used to input the prompt instruction into the first large model to obtain the response text with fine-grained sentiment tags that is streamed from the first large model; A voice broadcasting unit is used to synthesize a reply voice based on the reply text and broadcast the reply voice.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the emotion-based voice interaction method as described in any one of claims 1 to 9.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the emotion-based voice interaction method as described in any one of claims 1 to 9.