A method and system for interactive emotional expressions in a voice robot

By recognizing user emotions and outputting corresponding facial expressions and actions based on the voice robot's persona, the problem of in-vehicle voice robots being unable to recognize user emotions has been solved, enabling emotional interaction between users and robots and improving the user experience.

CN116524921BActive Publication Date: 2026-05-26SHENZHEN LANYOU TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN LANYOU TECHNOLOGY CO LTD
Filing Date
2023-03-09
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing in-vehicle voice robots cannot recognize changes in a user's emotions, cannot adjust their emotions accordingly, and cannot provide preset responses based on the user's preferred persona, resulting in a rigid and unintelligent user experience.

Method used

By capturing audio information and identifying the emotional type of the audio information, the system outputs corresponding response information based on the preset human persona of the voice robot, and identifies the emotional type of the response information to output corresponding facial expression commands, executes facial expression actions, and realizes emotional interaction between the user and the robot.

Benefits of technology

It provides an intelligent emotional interaction experience, allowing users to feel that their emotions are synchronized with the robot's, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116524921B_ABST
    Figure CN116524921B_ABST
Patent Text Reader

Abstract

This invention discloses a method for interactive emotional expressions with a voice robot, comprising the following steps: Step S100: capturing audio information; Step S200: identifying the emotion type of the audio information; Step S300: outputting response information corresponding to the emotion type of the audio information according to a preset voice robot persona; Step S400: identifying the emotion type of the response information and outputting facial expression action instructions corresponding to the emotion type of the response information; Step S500: executing facial expression actions according to the facial expression action instructions. This invention achieves multi-dimensional interaction between the voice robot and the user's emotions, providing the user with a pleasant experience of having a close friend in the car.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle networking technology, and in particular to a method and system for interactive emotional expressions of a voice robot. Background Technology

[0002] Currently, vehicles equipped with vehicle-to-everything (V2X) technology can connect to more smart hardware devices via Bluetooth, such as aftermarket in-vehicle voice robots. However, current aftermarket in-vehicle voice robots cannot recognize changes in the user's emotions, cannot adjust their responses based on the user's emotional state, and cannot provide preset replies based on the user's preferred persona. Instead, they use uniform responses and do not use different expressions based on the persona type, resulting in a rigid, unintelligent, and emotionless user experience. Summary of the Invention

[0003] The main objective of this invention is to address the shortcomings of existing in-vehicle voice robots, which cannot recognize user emotions or coordinate with facial expression changes, by providing a method and system for voice robot emotional expression interaction.

[0004] To achieve the above objectives, the present invention provides a voice robot's emotional expression interaction method and system, comprising the following steps:

[0005] Step S100: Capture audio information;

[0006] Step S200: Identify the emotion type of the audio information;

[0007] Step S300: Based on the preset persona type of the voice robot, output the response information corresponding to the emotion type of the audio information;

[0008] Step S400: Identify the emotion type of the reply message and output the facial expression and action instruction information corresponding to the emotion type of the reply message;

[0009] Step S500: Execute the facial expression action according to the facial expression action instruction information.

[0010] Preferably, step S200 includes: converting audio information into text information, performing emotion recognition on the text information using an emotion recognition model, and outputting the emotion type of the audio information.

[0011] Preferred, step S300 includes: reading preset persona type information in the voice robot, selecting the corresponding sub-database in the response database according to the persona type for querying, and outputting response text information corresponding to the emotion type of the audio information.

[0012] Preferably, step S400 includes: performing emotion recognition on the reply text information through an emotion recognition model, outputting the emotion type of the reply text information, comparing the emotion type of the reply text information with the command titles in the facial expression and action command database, and outputting the facial expression and action command information stored in the block where the command title that matches the emotion type of the reply text information is located.

[0013] Preferably, step S300 further includes: converting the text information of the reply into audio information of the reply, and transmitting it to the voice robot to broadcast the audio information of the reply.

[0014] In addition, to achieve the above objectives, the present invention also provides an emotional expression interaction system for a voice robot, including a voice robot for capturing audio information, broadcasting reply audio information, and executing facial expression actions according to facial expression action instruction information; and a remote server for identifying the emotion type of the audio information and reply text information, and outputting reply text information and reply audio information corresponding to the emotion type of the audio information according to the preset human type of the voice robot, and outputting facial expression action instruction information corresponding to the emotion type of the reply text information.

[0015] Preferably, the remote server includes a cloud-based voice platform, which converts audio information into text information and converts reply text information into reply audio information.

[0016] Preferably, the remote server further includes an emotion recognition module, which identifies and outputs the emotion type of the response text information and the text information converted in the cloud voice platform through an internally stored emotion recognition model.

[0017] Preferably, the remote server further includes a response query module, which reads the preset persona type information of the voice robot, selects the corresponding sub-database in the response database according to the persona type information, and outputs the response text information corresponding to the emotion type of the audio information.

[0018] Preferably, the remote server further includes a comparison module, which compares the emotion type of the reply text information with the command titles in the facial expression and action command database, and outputs the facial expression and action command information stored in the block where the command title that matches the emotion type is located.

[0019] This invention provides a method for interactive voice robot based on emotional expressions, which has the following beneficial effects: It captures audio information by recording the emotional sounds emitted by the user, enabling interaction between the user and the robot; it identifies the emotional type of the audio information, analyzes the user's emotions using an emotion recognition model, and provides an algorithmic basis for the response; it outputs response information corresponding to the emotional type of the audio information based on the preset persona type of the voice robot; it intelligently provides response information for that type of persona and emotion by positioning the voice robot's persona and emotions; it identifies the emotional type of the response information and outputs facial expression action instructions corresponding to the emotional type of the response information; it outputs corresponding facial expression instructions based on the emotional basis of the response, and these facial expressions, combined with the response broadcast, provide the user with a multi-faceted experience. The user executes facial expression actions according to the facial expression action instructions, and the intelligent device presents the facial expression action instructions to the user in the form of machine facial expression actions. The user sees that the intelligent device knows their emotions or is synchronized with their own emotions, giving them a positive experience of finding a kindred spirit. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort:

[0021] Figure 1 The diagram shown is a flowchart illustrating an emotional expression interaction method for a voice robot according to an embodiment of the present invention.

[0022] Figure 2 The diagram shown is a schematic diagram of the system module structure of an emotional expression interaction system for a voice robot according to an embodiment of the present invention. Detailed Implementation

[0023] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Typical embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0025] The general idea of ​​this invention is to address the shortcomings of existing in-vehicle voice robots, which cannot recognize user emotions or synchronize with facial expressions, by providing a method and system for interactive emotional expressions in a voice robot. This invention records the emotional sounds emitted by the user to enable interaction between the user and the robot; identifies the emotion type of the audio information; analyzes the user's emotions using an emotion recognition model to provide an algorithmic basis for the response; outputs response information corresponding to the emotion type of the audio information based on a preset voice robot persona; intelligently provides responses tailored to the specific persona and emotion of the voice robot; identifies the emotion type of the response information and outputs corresponding facial expression action instructions; outputs corresponding facial expression instructions based on the emotion of the response, and these instructions, combined with the response, provide a multi-sensory experience for the user. The user executes the facial expression actions according to the instructions, and the intelligent device presents these actions to the user in the form of machine-generated facial expressions. Seeing that the intelligent device understands or synchronizes with their own emotions provides a positive experience of finding a kindred spirit.

[0026] To better understand the above technical solutions, the following detailed description, in conjunction with the accompanying drawings and specific embodiments, will be provided. It should be understood that the embodiments and specific features described herein are detailed explanations of the technical solutions of this application, and not limitations thereof. Where there is no conflict, the embodiments and technical features described herein can be combined with each other. (Refer to...) Figure 1 , Figure 1 The diagram shown is a flowchart illustrating a method for interactive emotional expressions in a voice robot according to an embodiment of the present invention. In this embodiment, the method for interactive emotional expressions in a voice robot includes:

[0027] Step S100: Capture audio information;

[0028] By recording and transmitting the emotional sounds emitted by users, interaction between users and robots can be achieved.

[0029] Specifically, in one embodiment of the present invention, the voice robot can display various facial expressions and actions according to instructions, and can also play sounds according to audio signals. It can also be equipped with a screen displaying various expressions, prompts, etc. In this invention, the voice robot has a sound recording module, which can record the audio waveforms of various sounds near the vehicle system, such as the voices of users inside the vehicle and the sounds of vehicle movement, and transmit the recorded audio signals to a cloud-based voice platform for processing.

[0030] Of course, in this invention, the voice robot has the function of communicating with a remote server, and it can connect to the remote server through 4G / 5G / WIFI and other means to transmit various data.

[0031] Step S200: Identify the emotion type of the audio information;

[0032] This invention uses an emotion recognition model to identify and analyze user emotions, providing an algorithmic basis for responses. Specifically, to identify the emotion type of audio information, the audio information is converted into text information, and the emotion recognition model is used to identify the emotion of the text information, outputting the emotion type of the audio information. In this invention, the voice robot transmits audio information to a cloud-based voice platform on a remote server via a wireless network. The cloud-based voice platform, acting as a voice processing platform, converts audio information into text information using ASR (Automatic Speech Recognition), processes it using NLP (Natural Language Processing), and outputs the text information. The cloud-based voice platform then transmits the text information to the emotion recognition module on the remote server. The emotion recognition module identifies the emotion type of the text information. In this invention, the emotion recognition module is a trained model obtained through extensive input text information training. It can intelligently determine the emotion type based on the text and output the result. Any text input can yield the corresponding emotion result, such as "Today we talked about..." Inputting the emotion type into the emotion recognition module yields a "positive and happy" emotion type, such as "What kind of leader is this? He assigns me a ton of work every day, but I haven't seen a raise in my salary." Inputting the emotion type into the emotion recognition module yields a "negative and angry" emotion type, such as "My girlfriend broke up with me today." Inputting the emotion type into the emotion recognition module yields a "sad" emotion type, such as "This stray dog ​​is so pitiful, waiting for food here every day." Inputting the emotion type into the emotion recognition module yields a "sympathetic" emotion type, such as "I have a competition to participate in today." Inputting the emotion type into the emotion recognition module yields a "neutral" emotion type. After obtaining the emotion type result for the current user, the emotion type result is transmitted to the response query module on the remote server.

[0033] Step S300: Based on the preset persona type of the voice robot, output the response information corresponding to the emotion type of the audio information;

[0034] By identifying the voice robot's persona and emotions, it can intelligently provide responses tailored to that type of persona and emotion.

[0035] Specifically, to understand how the response information is generated, the system reads the preset persona type information from the voice robot, selects the corresponding sub-database in the response database based on the persona type, and outputs the response text information corresponding to the emotion type of the audio information.

[0036] The response query module receives the emotion type recognition result of the user's voice and queries the response database based on the emotion type result to find the corresponding response text information. In this invention, the voice robot has multiple personas, such as gentle and considerate, cute, aloof, and cool. Each voice robot has its persona type pre-set at the factory, and the persona type information of each voice robot is also stored in the background. When the response query module receives the emotion type recognition result of the user's voice, it reads the corresponding persona type of the voice robot from the background, selects the response data sub-database corresponding to the persona type from the response database in the remote server, and then inputs the text information of the user's voice and the emotion type of the user's voice into the response data sub-database to intelligently retrieve the response text information. The response database contains a large number of question answers, and it can automatically reply with different answers based on the personality characteristics of the persona. It can realize that the same question input, different emotion types, and different persona types will have different responses. For example, for all gentle personas, if the text "I fell off my bicycle" is input, the emotion will be aggrieved. The output response text is "Where are you hurt?" with a sympathetic emoji. Conversely, with the same text, if the input emotion is anger, the output text might be "I probably can't help you, I suggest you go to the hospital," with a pitiful emoji. For example, for a cool and aloof persona, if the input text is "I fell off my bicycle," with a wronged emotion, the output text is "Watch where you're going next time, don't fall again," with a frowning emoji. Conversely, with the same text, if the input emotion is anger, the output text might be: "What's it to me, go see a doctor," with a cold emoji.

[0037] To enable better voice interaction between the voice robot and users, the text of the reply is converted into audio information and transmitted to the voice robot for voice playback.

[0038] After the reply text information is retrieved by the reply query module, it is transmitted to the cloud voice platform. The cloud voice platform converts the reply text information into reply audio information through TTS and transmits it to the voice robot, which then reads the reply audio information.

[0039] Step S400: Identify the emotion type of the reply message and output the facial expression and action instruction information corresponding to the emotion type of the reply message;

[0040] Based on the emotion of the reply, the corresponding emoji command is output. This emoji, in conjunction with the reply, provides users with a multi-faceted experience.

[0041] Specifically, the emotion recognition model is used to identify the emotion of the reply text information, output the emotion type of the reply text information, compare the emotion type of the reply text information with the command titles in the facial expression and action command database, and output the facial expression and action command information stored in the block where the command title that matches the emotion type of the reply text information is located.

[0042] In one embodiment of the present invention, after the reply query module obtains the reply text information through persona type and emotion type, in order to output the corresponding facial expression command, it is necessary to perform emotion recognition on the reply text information. At this time, the reply query module transmits the reply text information to the emotion recognition module, and performs emotion recognition on the reply text information through the emotion recognition model to obtain the emotion type of the reply text information. The facial expression command of the voice robot is retrieved based on the emotion type. The facial expression action commands of the voice robot are all stored in the facial expression action command database on the remote server. The emotion type information is compared with the command titles in the facial expression action command database, and the facial expression action command information stored in the block where the matching command title is located is output. The remote server stores various databases, including a dedicated facial expression action command database. Each facial expression action's facial expression action command database has its own command title. Here, the command title is preset and corresponds to the emotion type. That is, the emotion type information serves as the command title of the facial expression action command. The facial expression action command information for each command title is stored in the corresponding block according to the specified storage address. Once the facial expression action command information is found, it is directly transmitted and output to the intelligent robot.

[0043] For example, when a remote server receives the emotion type information of "happy", it scans the command title database for facial expression commands and finds the facial expression command block with "happy" as the command title. The action command block contains commands such as "left upper arm and right upper arm clapping action command, buttock wiggling left and right command, and double leg jumping command". The facial expression command information in the block is then transmitted to the intelligent robot.

[0044] Step S500: Execute the facial expression action according to the facial expression action instruction information.

[0045] The intelligent robot receives facial expression and gesture instructions, performs the corresponding facial expressions and gestures, and displays the expressions on its screen. For example, if it receives a facial expression and gesture instruction with "happy" as the title, the robot will clap its left and right upper arms, wiggle its hips from side to side, jump with both legs, and display a happy emoji on the screen.

[0046] Based on the above, the present invention provides a voice robot emotion expression interaction method that captures audio information, records the emotional sounds emitted by the user, and realizes the interaction between the user and the robot; identifies the emotion type of the audio information, analyzes the user's emotions through an emotion recognition model, and provides an algorithmic basis for the reply; outputs reply information corresponding to the emotion type of the audio information according to the preset persona type of the voice robot; intelligently provides reply information for that persona and emotion through persona and emotion positioning; identifies the emotion type of the reply information and outputs the facial expression action instruction information corresponding to the emotion type of the reply information; outputs the corresponding facial expression instruction according to the emotion basis of the reply, and this facial expression, in conjunction with the reply, brings multiple experiences to the user. The user executes the facial expression action according to the facial expression action instruction information, and the intelligent device presents the facial expression action instruction to the user in the form of machine facial expression actions. The user sees that the intelligent device knows their emotions or is synchronized with their emotions, thus having a good experience of finding a kindred spirit.

[0047] Accordingly, the present invention also provides an interactive system for emotional expressions in a voice robot, referring to... Figure 2 , Figure 2 The diagram shows a system module structure of a voice robot emotion expression interaction system according to an embodiment of the present invention. The system performs emotion expression interaction through the emotion expression interaction method described above, and includes: a voice robot, used to capture audio information, broadcast reply audio information, and execute expression actions according to expression action instruction information; and a remote server, used to identify the emotion type of the audio information and reply text information, and output reply text information and reply audio information corresponding to the emotion type of the audio information according to the preset voice robot persona type, and output expression action instruction information corresponding to the emotion type of the reply text information.

[0048] Preferably, the remote server includes a cloud-based voice platform, which converts audio information into text information and converts reply text information into reply audio information.

[0049] Preferably, the remote server further includes an emotion recognition module, which identifies and outputs the emotion type of the response text information and the text information converted in the cloud voice platform through an internally stored emotion recognition model.

[0050] Preferably, the remote server further includes a response query module, which reads the preset persona type information of the voice robot, selects the corresponding sub-database in the response database according to the persona type information, and outputs the response text information corresponding to the emotion type of the audio information.

[0051] Preferably, the remote server further includes a comparison module, which compares the emotion type of the reply text information with the command titles in the facial expression and action command database, and outputs the facial expression and action command information stored in the block where the command title that matches the emotion type is located.

[0052] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0053] Similarly, it should be understood that, in order to simplify this disclosure and aid in understanding one or more of the various aspects of the invention, in the above description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, this method of disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into this detailed description, wherein each claim itself is a separate embodiment of the invention.

[0054] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and further, they can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0055] Furthermore, those skilled in the art will understand that although some embodiments herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of the invention and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0056] The various component embodiments of the present invention can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such programs implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.

[0057] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.

Claims

1. An emotional expression interaction method for a voice robot, characterized in that It includes the following steps: Step S100: Capture audio information; Step S200: Identify the emotion type of the audio information; Step S300: According to the preset personality type of the voice robot, output reply language information corresponding to the emotion type of the audio information. Among them, the personality type information of each voice robot is stored in the background, and when the emotion type recognition result of the user's voice is received, the corresponding personality type of the voice robot will be read from the background, so that for the same question input, different personality types have different replies; Step S400: Identify the emotion type of the reply language information and output expression action instruction information corresponding to the emotion type of the reply language information; Step S500: Execute expression actions according to the expression action instruction information, where the expression action instruction is presented in the form of machine expression actions; The said Step S200 includes: Convert the audio information into text information, perform emotion recognition on the text information through an emotion recognition model, and output the emotion type of the audio information; The said Step S300 includes: Read the preset personality type information in the voice robot, select the corresponding sub-library in the reply language database according to the personality type for query, and output the reply language text information corresponding to the emotion type of the audio information; The said Step S400 includes: Perform emotion recognition on the reply language text information through an emotion recognition model, output the emotion type of the reply language text information, compare the emotion type of the reply language text information with the instruction titles in the expression action instruction database, and output the expression action instruction information stored in the block where the instruction title matching the emotion type of the reply language text information is located.

2. The emotional expression interaction method of the voice robot according to claim 1, wherein The said Step S300 further includes: Convert the reply language text information into reply voice frequency information and transmit it to the voice robot to perform voice broadcast on the reply voice frequency information.

3. An emotional expression interaction system for a voice robot, which is used to execute the method described in any one of claims 1 to 2, and is characterized in that, It includes: A voice robot, used to capture audio information, broadcast reply voice frequency information, and execute expression actions according to the expression action instruction information; A remote server, used to identify the emotion types of the audio information and the reply language text information, and according to the preset personality type of the voice robot, output the reply language text information and the reply voice frequency information corresponding to the emotion type of the audio information, and output the expression action instruction information corresponding to the emotion type of the reply language text information; Among them, the personality type information of each voice robot is stored in the background, and when the emotion type recognition result of the user's voice is received, the corresponding personality type of the voice robot will be read from the background, so that for the same question input, different personality types have different replies, and the expression action instruction is presented in the form of machine expression actions; The said remote server includes a cloud voice platform, and the cloud voice platform converts the audio information into text information and converts the reply language text information into reply voice frequency information.

4. The emotional expression interaction system of the voice robot according to claim 3, wherein The said remote server further includes an emotion recognition module, and the emotion recognition module performs emotion type recognition on the reply language text information and the text information converted in the cloud voice platform through the emotion recognition model stored inside and outputs the result.

5. The emotional expression interaction system of the voice robot according to claim 4, wherein The remote server further includes a reply query module, which reads the preset persona type information of the voice robot, selects the corresponding sub-database in the reply database according to the persona type information for query, and outputs the reply text information of the audio information corresponding to the emotion type.

6. The emotional expression interaction system of the voice robot according to claim 5, wherein, The remote server further includes a comparison module, which compares the emotion type of the reply text information with the instruction titles in the expression action instruction database, and outputs the expression action instruction information stored in the block where the instruction title matching the emotion type is located.