An emotional companion robot system, method, and emotional companion robot

By integrating a non-contact vital sign acquisition module and a cloud-based multimodal fusion module into the emotional companion robot system, efficient analysis of user voice, facial images, breathing and heart rate is achieved, solving the problem of inaccurate emotion recognition and improving the interactive performance and emotion regulation effect of the emotional companion robot.

CN118514093BActive Publication Date: 2026-05-26SHENYANG SIASUN ROBOT & AUTOMATION
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENYANG SIASUN ROBOT & AUTOMATION
Filing Date
2024-05-22
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing emotional companion robots suffer from inaccurate emotion recognition due to insufficient data collection, which affects human-computer interaction performance.

Method used

An emotional companion robot system was designed, which integrates a non-contact vital sign acquisition module, including a camera, a respiratory rate sensor, and a heart rate sensor. Facial tracking is achieved through drive motor control. Combined with a multimodal fusion module on a cloud server, voice, facial images, respiratory and heart rate are analyzed to generate accurate control commands.

Benefits of technology

It improves the accuracy of emotion recognition, enhances the applicability and adjustment methods of emotional companion robots, making them more naturally meet user expectations and improve the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118514093B_ABST
    Figure CN118514093B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of robot emotion recognition, specifically an emotion-companion robot system, method, and emotion-companion robot. It includes: a controller, used to receive user voice and user facial images, respiratory rate, and heart rate collected by a non-contact vital sign acquisition module; it converts the user's voice into a voice-text message and sends all the above messages to a cloud server for processing; it receives control commands returned by the cloud server after processing, and upon receiving the control commands, controls the emotion-companion robot to complete the corresponding actions; the cloud server receives information sent from the controller and judges the voice-text messages. If it is a command-type message, it generates commands to control the robot's actions; otherwise, it analyzes the above information and feeds back the resulting control commands to the controller. This invention fuses and analyzes the five types of information collected, improving the accuracy of emotion recognition and ensuring the effectiveness of the adjustment process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot emotion recognition, specifically an emotional companion robot system, method, and emotional companion robot. Background Technology

[0002] In recent years, with the development and application of various multimodal fusion technologies and adaptive learning algorithms, emotional companion robots have also made rapid progress. Through monitoring users, they can recognize users' emotions and provide feedback. However, in the current monitoring process, due to the limited amount of data collected, the emotion recognition process cannot be completed effectively, resulting in inaccurate robot feedback behavior. These issues all affect the human-computer interaction performance of emotional companion robots, thus requiring further research and improvement. Summary of the Invention

[0003] The purpose of this invention is to provide an emotional companion robot system. First, it proposes the design of a companion robot with good human-computer interaction performance. Then, it presents an emotion recognition and fusion strategy that can process multiple sensor data. The entire system can provide the user's emotional state according to different usage situations and provide specific adjustment schemes.

[0004] The technical solution adopted by the present invention to achieve the above objectives is: an emotional companion robot system, comprising: a controller, a cloud server, and a non-contact vital sign acquisition module installed on the emotional companion robot;

[0005] The controller receives user voice and facial images, respiratory rate, and heart rate collected by the non-contact vital sign acquisition module. It converts user voice into voice-text messages and sends all of these messages to the cloud server for processing. At the same time, it receives control commands returned by the cloud server after processing and controls the emotional companion robot to complete the corresponding actions.

[0006] The non-contact vital sign acquisition module is used to recognize the user's facial image and send the recognition result to the cloud server through the controller; at the same time, it monitors the user's breathing and heart rate and sends them to the cloud server through the controller.

[0007] The cloud server receives user audio content, converted voice-to-text messages, user facial images, breathing rate, and heart rate from the controller. It then judges the voice-to-text messages. If they are instruction-type messages, it generates instructions to control the robot's actions. Otherwise, it analyzes the voice-to-text messages, user voice, user facial images, breathing rate, and heart rate, and feeds back the resulting control instructions to the controller.

[0008] The cloud server includes: an instruction analysis module, an analysis and processing module, a multimodal fusion module, and an instruction message generation module;

[0009] The instruction analysis module receives user audio content, converted voice-text messages, user facial images, breathing frequency, and heart rate from the controller, and analyzes the voice-text messages. If the message is an instruction, it is sent to the instruction message generation module to generate instructions that can control the robot's actions; otherwise, the voice-text, user voice, user facial images, breathing frequency, and heart rate are sent to the analysis and processing module for analysis.

[0010] The analysis and processing module is used to receive voice text, user voice, user facial image, respiratory rate and heart rate sent from the instruction analysis module, analyze them, obtain analysis results X1, X2, X3, X4 and X5, and send them to the multimodal fusion module.

[0011] The multimodal fusion module, as the core of the cloud server, is used to send the results of weight assignment of X1, X2, and X3 to the instruction message generation module.

[0012] The instruction message generation module is used to process the content received from the instruction analysis module or the multimodal fusion module and generate control instructions that the robot can recognize.

[0013] The analysis and processing module includes: a text analysis module, an expression recognition module, a speech recognition module, a heartbeat analysis module, and a breathing analysis module connected to the multimodal fusion module;

[0014] The text analysis module is used to receive voice text, perform intent recognition and text-based emotion recognition, and generate result X1;

[0015] The facial expression recognition module is used to receive user videos sent by the video capture module, recognize the current facial expression, and generate result X2;

[0016] The speech recognition module is used to receive speech audio, recognize the tone and voiceprint of the current speech, and generate result X3;

[0017] The heartbeat analysis module analyzes the current heartbeat frequency and generates a result X4, where X4 represents the current real-time heartbeat frequency.

[0018] The respiratory analysis module analyzes the current respiratory rate and generates a result X5, where X5 represents the current real-time respiratory rate.

[0019] The sub-modules in the analysis and processing module send the corresponding generated results X1, X2, X3, X4, and X5 to the multimodal fusion module.

[0020] The non-contact vital sign acquisition module includes: camera A, respiratory rate sensor, heart rate sensor, and drive motor;

[0021] Camera A, respiratory rate sensor, and heart rate sensor are connected to the controller to send the collected user facial image, respiratory rate, and heart rate to the controller, respectively.

[0022] The respiratory rate sensor and heart rate sensor form an adjustable angle structure with camera A, which is used to ensure that the controller can control the drive motor so that camera A can track the face in real time, and during the tracking process, the respiratory rate sensor and heart rate sensor can be aligned with the user's chest cavity.

[0023] It also includes: a display module, a voice input module, a voice output module, and a motion module located within the emotional companion robot;

[0024] The display module is connected to the controller and is used to play facial expressions, videos and music according to the control instructions sent by the controller, so as to enhance the robot's interactivity and increase the ways of emotional regulation;

[0025] The voice output module is connected to the controller and is used to play the robot's audio according to the instructions generated by the instruction message generation module.

[0026] The mobile module is connected to the controller and is used to receive control signals from the controller to drive the robot body to move as a whole using differential wheels.

[0027] The emotional companion robot includes: a robot body and a microphone, display, chest display, motion chassis, speaker and video acquisition module installed on the robot body;

[0028] The microphones are arranged in a microphone matrix on the top of the robot body and connected to the voice input module. When used for voice input, the user's position is obtained according to the microphone matrix, and then the robot controller gives a movement command to drive the robot body's motion chassis to rotate so that the robot faces the user.

[0029] The robot body is designed to mimic the human body structure, and a display is embedded in the head of the robot body to facilitate user interaction when standing.

[0030] The controller is located inside the display and connected to the display module; the display and the display module are used to play facial expressions or play videos and music to increase the ways of emotional regulation.

[0031] A chest display is embedded in the middle of the robot body to facilitate interaction when the user is sitting or lying down; the chest display is connected to the display module to play facial expressions, or to play videos and music, and also to display the interactive interface during video calls.

[0032] A non-contact vital sign acquisition module is installed between the chest monitor and the monitor;

[0033] The motion chassis is located at the bottom of the robot body and is connected to the moving module. The moving module drives the motion chassis to move the robot body in a differential wheel manner to achieve the overall movement of the robot.

[0034] The speakers are located on both sides of the robot's head and are connected to the voice output module for playing the robot's audio feedback.

[0035] There are two video acquisition modules, both connected to the controller. One is camera A located on the non-contact vital sign acquisition module, and the other is camera B located on the head of the robot body. The video acquisition modules are used to acquire user facial images and send them to the controller.

[0036] A method for providing companionship using an emotional companion robot system includes the following steps:

[0037] 1) The voice input module acquires the user's voice and converts the user's voice into a text message through the controller. The controller then sends the voice text and voice audio to the cloud server for further analysis.

[0038] 2) The instruction analysis module in the cloud server analyzes the voice text to determine whether the user's voice is an instruction message. If it is an instruction message, it is sent to the instruction message generation module, and no further analysis is performed. The instruction information is returned to the controller in the robot. The controller controls the robot to perform the corresponding action based on the returned instruction. Otherwise, the voice text, user voice, user facial image, breathing rate, and heart rate are sent to their respective sub-modules of the analysis and processing module.

[0039] 3) The respective sub-modules of the analysis and processing module in the cloud server receive the voice text, user voice, user facial image, breathing frequency, and heart rate, and perform their own analysis. The recognition results X1, X2, X3, X4, and X5 obtained after analysis are sent to the multimodal fusion module, and the fusion results are sent to the instruction message generation module.

[0040] 4) The instruction message generation module generates corresponding control instructions and feeds them back to the controller, which then controls the emotional companion robot to execute the relevant instructions.

[0041] 8. The companionship method of the emotional companionship robot system according to claim 7, characterized in that step 3) specifically comprises:

[0042] 1-1) The text analysis module performs intent recognition and text-based sentiment recognition on the speech text and generates a result X1. The range of this content is divided into [-1,1], that is, the generated result X1 is within the divided interval, which represents the sentiment tendency of the current text.

[0043] 1-2) The expression recognition module identifies the current expression in the user's video and generates a result X2. The range of this content is divided into [-1,1]. That is, the generated result X2 is within the divided interval, which represents the emotional tendency of the current expression.

[0044] 1-3) The speech recognition module processes the speech audio and identifies the tone and voiceprint of the current speech, and generates a result X3. The range of this content is divided into [-1,1]; that is, the generated result X3 within the divided interval represents the emotional tendency of the current speech audio.

[0045] 1-4) The heartbeat analysis module analyzes the current heartbeat frequency and generates a result X4, where X4 represents the current real-time heartbeat frequency;

[0046] 1-5) The respiratory analysis module analyzes the current respiratory rate and generates a result X5, where X5 represents the current real-time respiratory rate.

[0047] The multimodal fusion module assigns weights to the recognition results X1, X2, and X3 sent by the controller to obtain the fusion result of X1, X2, and X3, specifically as follows:

[0048] 2-1) Based on the X4 and X5 values ​​provided by the current heart rate and respiration analysis module, where X4 is the real-time heart rate and X5 is the real-time respiratory rate, the resting heart rate is set to 60-100 beats / minute, with 70 as the baseline value for X4. 基准 Corresponding to point 0 in [-1,1], with 130 as the upper limit of X4. 上限 , corresponding to ±1;

[0049] Set the resting respiratory rate to 12-20 breaths per minute, using 13 as the baseline value multiplied by 5. 基准 Corresponding to point 0 in [-1,1], with 28 as the upper limit of X5. 上限 , corresponding to ±1;

[0050] The generation coefficients for real-time heart rate X4 and real-time respiratory rate X5 are respectively:

[0051]

[0052]

[0053] 2-2) Based on X1, X2, X3, the generation coefficients a4 and a5 of real-time heart rate X4 and real-time respiratory rate X5 are generated. At this time, a1, a2, and a3 are introduced. The default values ​​of a1, a2, and a3 are 5, 2, and 3, respectively. Users can set a1, a2, and a3 through the human-computer interaction interface according to the current usage scenario.

[0054] The values ​​of a1, a2, and a3 all range from 1, 2, 3, 4, and 5, where 1 represents the weakest confidence level and 5 represents the strongest confidence level. These values ​​are combined with X1, X2, and X3 to generate the weighted result Y, i.e.:

[0055]

[0056] 2-3) Obtain the segmentation threshold a6 of the multimodal fusion result Z based on the coefficients a4 and a5 generated from the heart rate X4 and respiratory rate X5, i.e.:

[0057]

[0058] 2-4) Based on the segmented threshold a6, the multimodal fusion result Z can be expressed as:

[0059]

[0060] The present invention has the following beneficial effects and advantages:

[0061] 1. This invention integrates a camera, a heart rate sensor, and a respiratory rate sensor into a non-contact vital signs sensor. It is controlled by two drive motors, has two degrees of freedom, and completes facial tracking by cooperating with facial data collected by the camera. This allows the respiratory rate sensor and heart rate sensor to be aligned with the user's chest cavity to accurately collect relevant data.

[0062] 2. This invention integrates and analyzes five types of information generated from the user's voice: text information, facial expression information, audio information, heart rate information, and respiratory rate information, thereby improving the accuracy of emotion recognition and ensuring the effectiveness of the adjustment process.

[0063] 3. This invention can effectively improve the accuracy of emotion recognition in emotional companion robots, enhance their applicability in different usage scenarios, and thus improve the adjustment method of emotional companion robots, making them more natural, more in line with user expectations, and improving the user experience. Attached Figure Description

[0064] Figure 1 A schematic diagram of the emotional companion robot of the present invention;

[0065] Among them, 1 is a microphone, 2 is a head display, 3 is a non-contact vital sign acquisition module, 4 is a chest display, 5 is a motion chassis, 6 is a speaker, and 7 is camera B.

[0066] Figure 2 Schematic diagram of the operating principle of the emotional companion robot of the present invention;

[0067] Figure 3 A schematic diagram of the principle of the cloud server polymorphic fusion module of the present invention. Detailed Implementation

[0068] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0069] like Figure 2 The diagram shows the operating principle of the emotional companion robot of the present invention. The user interacts with the emotional companion robot using voice. At the same time, the robot continuously uses non-contact vital sign sensors to collect the user's facial expressions, breathing rate and heart rate, and sends all the above information to the cloud. Specifically, the present invention provides an emotional companion robot system, including: a controller, a cloud server, and a non-contact vital sign collection module installed on the emotional companion robot.

[0070] The controller receives user voice and facial images, respiratory rate, and heart rate collected by the non-contact vital sign acquisition module. It converts user voice into voice-text messages and sends all of these messages to the cloud server for processing. At the same time, it receives control commands returned by the cloud server after processing and controls the emotional companion robot to complete the corresponding actions.

[0071] The controller functions as described above. It sends data to a cloud server for processing. The cloud server processes the collected data and returns control commands. Upon receiving these commands, the controller controls devices such as the screen and speakers to perform corresponding actions. Additionally, the controller includes a step of converting voice messages into voice-to-text messages, receiving control commands from the cloud server, and then controlling local hardware to perform actions.

[0072] The non-contact vital sign acquisition module is used to recognize the user's facial image and send the recognition result to the cloud server through the controller; at the same time, it monitors the user's breathing and heart rate and sends them to the cloud server through the controller.

[0073] In this embodiment, the non-contact vital sign acquisition module can realize user facial recognition, and can control the drive motors of the respiratory rate sensor and heart rate sensor according to the recognition result to ensure that the respiratory rate sensor and heart rate sensor can be aligned with the user's chest cavity.

[0074] The cloud server receives user voice (audio content), voice-text messages (text content), user facial images (video content), breathing rate (text content), and heart rate (text content) from the controller. First, the voice-text messages are sent to the instruction analysis module. If it is an instruction message, it is sent to the instruction message generation module to generate instructions that can control the robot's actions. Otherwise, the voice-text messages, user voice, user facial images, breathing rate, and heart rate are analyzed, and the resulting control instructions are fed back to the controller.

[0075] like Figure 3 The diagram shown is a schematic diagram of the principle of the cloud server multi-modal fusion module of the present invention. The cloud server includes: an instruction analysis module, an analysis and processing module, a multi-modal fusion module, and an instruction message generation module.

[0076] The instruction analysis module receives user audio content, converted voice-text messages, user facial images, breathing frequency, and heart rate from the controller, and analyzes the voice-text messages. If the message is an instruction, it is sent to the instruction message generation module to generate instructions that can control the robot's actions; otherwise, the voice-text, user voice, user facial images, breathing frequency, and heart rate are sent to the analysis and processing module for analysis.

[0077] The multimodal fusion module, as the core of the cloud server, is used to send the results of weight assignment of X1, X2, and X3 to the instruction message generation module.

[0078] The instruction message generation module is used to process the content received from the instruction analysis module or the multimodal fusion module and generate control instructions that the robot can recognize.

[0079] The analysis and processing module is used to receive voice text, user voice, user facial image, respiratory rate and heart rate sent from the instruction analysis module, analyze them, obtain analysis results X1, X2, X3, X4 and X5, and send them to the multimodal fusion module.

[0080] Among them, Figure 3 The analysis and processing module includes: a text analysis module, an expression recognition module, a speech recognition module, a heartbeat analysis module, and a breathing analysis module, all connected to the multimodal fusion module.

[0081] The text analysis module is used to receive voice text, perform intent recognition and text-based emotion recognition, and generate result X1;

[0082] The facial expression recognition module is used to receive user videos sent by the video capture module, recognize the current facial expression, and generate result X2;

[0083] The speech recognition module is used to receive speech audio, recognize the tone and voiceprint of the current speech, and generate result X3;

[0084] The heartbeat analysis module analyzes the current heartbeat frequency and generates a result X4, where X4 represents the current real-time heartbeat frequency.

[0085] The respiratory analysis module analyzes the current respiratory rate and generates a result X5, where X5 represents the current real-time respiratory rate.

[0086] The present invention provides an emotional companion robot system, which further includes: a display module, a voice input module, a voice output module, and a movement module disposed within the emotional companion robot;

[0087] The display module is connected to the controller and is used to play facial expressions, videos and music according to the control instructions sent by the controller, so as to enhance the robot's interactivity and increase the ways of emotional regulation;

[0088] The voice output module is connected to the controller and is used to execute and play the robot's audio according to the instructions generated by the instruction message generation module;

[0089] The mobile module is connected to the controller and is used to receive control signals from the controller to drive the robot body to move as a whole using differential wheels.

[0090] The robot acquires the user's voice through a voice input module and converts it into a text message. Data collected by a non-contact vital sign acquisition module is then transmitted to the cloud for analysis. First, it analyzes whether the message is a command. If it is, the companion robot provides specific feedback, such as moving forward, backward, turning left, turning right, playing audio or video, or making a video call with a designated contact. If it is not a command, all collected data is sent to the core module in the cloud for fusion processing to analyze the user's current emotion. Based on this emotion, the companion robot provides corresponding feedback to regulate the emotion, such as engaging in multi-turn dialogue, offering professional advice, making interactive facial expressions, or playing audio or video.

[0091] This invention employs a humanoid robot to provide emotional companionship to users. It features excellent human-computer interaction capabilities and a highly integrated non-contact vital sign acquisition module to monitor the user's respiratory rate and heart rate. In addition, the companion robot can also collect the user's voice and facial expressions. All of this data is sent to the core module in the cloud for fusion processing, thereby improving the accuracy of emotion recognition. Based on the results of emotion recognition, the robot can adjust emotions to complete the companionship process, which has practical value.

[0092] like Figure 1 The diagram shown is a schematic of the structure of the emotional companion robot of the present invention; wherein, the non-contact vital sign acquisition module includes: camera A, respiratory rate sensor, heart rate sensor, and drive motor; camera A, respiratory rate sensor, and heart rate sensor are connected to the controller and are used to send the acquired user facial image, respiratory rate, and heart rate to the controller respectively.

[0093] The respiratory rate sensor and heart rate sensor form an adjustable angle structure with camera A to ensure that the controller can control the drive motor so that camera A can track the face in real time. During the tracking process, the non-contact vital sign acquisition module is controlled by a pair of drive motors, which have two degrees of freedom. By cooperating with the facial data collected by the camera, it completes the facial tracking function, thereby enabling the respiratory rate sensor and heart rate sensor to be aligned with the user's chest cavity to complete the collection of relevant data.

[0094] like Figure 1 As shown, based on an emotional companion robot system, this invention designs an emotional companion robot as a carrier, which includes: a robot body and a microphone 1, a display 2, a chest display 4, a motion chassis 5, a speaker 6 and a video acquisition module 7 installed on the robot body.

[0095] Microphone 1 is arranged in a microphone matrix on the top of the robot body and connected to the voice input module. When used for voice input, the user's position is obtained according to the microphone matrix, and then the robot controller gives a movement command to drive the robot body's motion chassis 5 to rotate so that the robot faces the user.

[0096] The robot body is designed to mimic the human body structure. In this embodiment, it is 1.5 meters tall. A display screen 2 is embedded in the head of the robot body to facilitate user interaction when standing.

[0097] The controller is located inside the display 2 and connected to the display module; the display 2 is connected to the display module and is used to display the content of the onboard Android system, play emoticons, or play videos and music to increase the emotional adjustment methods and facilitate user interaction when standing.

[0098] A chest display 4 is embedded in the middle of the robot body to facilitate interaction when the user is sitting or lying down. The chest display 4 is connected to the display module to display the content of the onboard Windows system. The displayed content is part of the content displayed by Android. In addition, it can also display the interactive interface during video calls, which is convenient for the user to use when sitting or lying down.

[0099] A non-contact vital sign acquisition module is provided between the chest display 4 and the display 2;

[0100] The motion chassis 5 is located at the bottom of the robot body and is connected to the moving module. The moving module drives the motion chassis 5 to move the robot body in a differential wheel manner to achieve the overall movement of the robot.

[0101] Speakers 6 are located on both sides of the robot's head and are connected to the voice output module to play the robot's audio feedback.

[0102] There are two video acquisition modules 7, both of which are connected to the controller. One is camera A located on the non-contact vital sign acquisition module, and the other is camera B located on the head of the robot body. The video acquisition module 7 is used to capture images of the user's face and send them to the controller.

[0103] like Figure 3 The diagram shown is a schematic diagram of the principle of the cloud server polymorphic fusion module of the present invention. The present invention discloses a companionship method of an emotional companionship robot system, including the following steps:

[0104] 1) The voice input module acquires the user's voice and converts the user's voice into a text message through the controller. The controller then sends the voice text and voice audio to the cloud server for further analysis.

[0105] 2) The instruction analysis module in the cloud server analyzes the voice text to determine whether the user's voice is an instruction message. If it is an instruction message, it is sent to the instruction message generation module, and no further analysis is performed. The instruction information is returned to the controller in the robot. The controller controls the robot to perform the corresponding action based on the returned instruction. Otherwise, the voice text, user voice, user facial image, breathing rate, and heart rate are sent to their respective sub-modules of the analysis and processing module.

[0106] 3) The respective sub-modules of the analysis and processing module in the cloud server receive the voice text, user voice, user facial image, breathing frequency, and heart rate, and perform their own analysis. The recognition results X1, X2, X3, X4, and X5 obtained after analysis are sent to the multimodal fusion module, and the fusion results are sent to the instruction message generation module.

[0107] Specifically, the analyzed recognition results X1, X2, X3, X4, and X5 are as follows:

[0108] a. The text analysis module performs intent recognition and text-based sentiment recognition on the spoken text, generating a result X1. This result is divided into the range [-1, 1] (-1 to 1 represents a sentiment shift from negative to positive, with 0 representing no state), for example, X1 = 0.2 (how this value is generated is not the main content of the text; the key lies in how to perform the subsequent multimodal fusion process). This number represents the sentiment tendency of the current text.

[0109] b. The facial expression recognition module identifies the current facial expression in the user's video and generates a result X2. This result is divided into the range [-1, 1] (-1 to 1 represents the emotion from negative to positive, and 0 represents no state). For example, X2 = 0.2 (how this value is generated is not the main content of the text; the key is how to perform the subsequent multimodal fusion process). This number represents the emotional tendency of the current facial expression.

[0110] c. The speech recognition module processes the audio and identifies the tone and voiceprint of the current speech, generating a result X3. This result is divided into the range [-1, 1] (-1 to 1 represents the emotion from negative to positive, 0 represents no state), for example, X3 = 0.2 (how this value is generated is not the main content of the text; the key lies in how to perform the subsequent multimodal fusion process). This number represents the emotional tendency of the current audio.

[0111] d. The heart rate analysis module analyzes the current heart rate and generates a result X4, where X4 represents the current real-time heart rate (beats / minute), for example: X4 = 70.

[0112] e. The respiratory analysis module analyzes the current respiratory rate and generates a result X5, where X5 represents the current real-time respiratory rate (breaths / minute), for example: X5 = 17.

[0113] 4) The instruction message generation module generates corresponding control instructions and feeds them back to the controller, which then controls the emotional companion robot to execute the relevant instructions.

[0114] In step 4), the multimodal fusion module assigns weights to the recognition results X1, X2, and X3 sent by the controller to obtain the fusion results X1, X2, and X3, specifically:

[0115] 2-1) Based on the X4 and X5 values ​​provided by the current heart rate and respiration analysis module, where X4 is the real-time heart rate and X5 is the real-time respiratory rate, the resting heart rate is set to 60-100 beats / minute, with 70 as the baseline value for X4. 基准 Corresponding to point 0 in [-1,1], with 130 as the upper limit of X4.上限 , corresponding to ±1;

[0116] Set the resting respiratory rate to 12-20 breaths per minute, using 13 as the baseline value multiplied by 5. 基准 Corresponding to point 0 in [-1,1], with 28 as the upper limit of X5. 上限 , corresponding to ±1;

[0117] The generation coefficients for real-time heart rate X4 and real-time respiratory rate X5 are respectively:

[0118]

[0119]

[0120] 2-2) Based on X1, X2, X3, the generation coefficients a4 and a5 of real-time heart rate X4 and real-time respiratory rate X5 are generated. At this time, a1, a2, and a3 are introduced. The default values ​​of a1, a2, and a3 are 5, 2, and 3, respectively. Users can set a1, a2, and a3 through the human-computer interaction interface according to the current usage scenario.

[0121] The values ​​of a1, a2, and a3 all range from 1, 2, 3, 4, and 5, where 1 represents the weakest confidence level and 5 represents the strongest confidence level. These values ​​are combined with X1, X2, and X3 to generate the weighted result Y, i.e.:

[0122]

[0123] 2-3) Obtain the segmentation threshold a6 of the multimodal fusion result Z based on the coefficients a4 and a5 generated from the real-time heart rate X4 and real-time respiratory rate X5, i.e.:

[0124]

[0125] 2-4) Based on the segmented threshold a6, the multimodal fusion result Z can be expressed as:

[0126]

[0127] Example:

[0128] The first step is to process the voice data, dividing it into voice text and voice audio. First, it analyzes whether the voice text content is instructional. After analysis, if it is instructional, the cloud generates the corresponding control command and returns it to the emotional companion robot. The robot completes the corresponding behavior according to the command, thereby completing the emotional regulation. If it is not instructional, the voice text, user voice, user facial image, breathing rate, and heart rate are sent to their respective sub-modules in the analysis and processing module.

[0129] Next, the text information generated from the speech is sent to the natural language processing module to complete intent recognition and text-based emotion recognition, generating result X1. The collected facial expression information is sent to the facial expression recognition module to complete the recognition of the current facial expression, generating result X2. The collected speech audio information is sent to the speech recognition module to recognize the pitch and voiceprint of the current speech, generating result X3. The collected heart rate information is sent to the heart rate analysis module to complete the analysis of the current heart rate, generating result X4. The collected respiratory rate information is sent to the respiratory analysis module to complete the analysis of the current respiratory rate, generating result X5.

[0130] The next step is to send all the X1, X2, X3, X4, and X5 obtained above to the multimodal fusion module. This module will analyze the input content and assign weights to X1, X2, and X3. The specific assignment process in this embodiment is as follows:

[0131] First, based on the X4 and X5 values ​​provided by the current heart rate and respiration analysis module, where X4 is the real-time heart rate and X5 is the real-time respiratory rate, the resting heart rate is determined to be 60-100 beats / minute using empirical values. Then, 70 (X4... 基准 Using ) as the baseline value (this value can be changed by the user according to their actual situation to ensure the accuracy of the result), corresponding to the 0 point in [-1,1], with 130 (X4) as the baseline value. 上限 The upper limit is (this value can be changed by the user according to their actual situation to ensure the accuracy of the result), corresponding to ±1. Assuming the current X4 value is 95, then a4 = (95-70) / (130-70) = 0.41; the resting respiratory rate is determined by experience to be 12-20 breaths / minute, with 13(X5) as the upper limit. 基准 Using ) as the baseline value (this value can be changed by the user according to their actual situation to ensure the accuracy of the result), corresponding to the 0 point in [-1,1], with 28 (X5) 上限 The upper limit is (this value can be changed by the user according to their actual situation to ensure the accuracy of the result), corresponding to ±1. Assuming the current X5 value is 20, then a5 = (20-13) / (28-13) = 0.47.

[0132]

[0133]

[0134] Based on X1, X2, X3, the generation coefficients a4 and a5 of real-time heart rate X4 and real-time respiratory rate X5 are then introduced. The default values ​​of a1, a2, and a3 are 5, 2, and 3, respectively. Users can set a1, a2, and a3 through the human-computer interaction interface according to the current usage scenario.

[0135] The values ​​of a1, a2, and a3 all range from 1, 2, 3, 4, and 5, where 1 represents the weakest confidence level and 5 represents the strongest confidence level. These values ​​are combined with X1, X2, and X3 to generate the weighted result Y, i.e.:

[0136]

[0137] The segmentation threshold a6 of the multimodal fusion result Z is obtained from the coefficients a4 and a5 generated based on the real-time heart rate X4 and real-time respiratory rate X5, i.e.:

[0138]

[0139] Based on the segmentation threshold a6, the multimodal fusion result Z can be expressed as:

[0140]

[0141] The following example uses scenario I (chatting with a robot) as an example (all the following content occurs within the multimodal fusion module). During this chat phase, the controller transmits voice and text data, user facial expression data, voice audio data, heart rate data, and respiratory rate data. Let X1 be 0.4, X2 be 0.3, X3 be 0.6, heart rate X4 = 95 beats / minute, and respiratory rate X5 = 20 breaths / minute. Decomposing the current scenario, a1 = 4, a2 = 5, a3 = 4, a4 ​​is calculated to be a4 = 0.41, a5 is calculated to be 0.47, then Y = 0.42, a6 = 0.43, and Z = 0.

[0142] Z=0 indicates a neutral emotion, which will not be regulated in any way.

[0143] This invention employs a humanoid robot to provide emotional companionship to users. It features excellent human-computer interaction capabilities and a highly integrated non-contact vital sign acquisition module to monitor the user's respiratory rate and heart rate. In addition, the companion robot can also collect the user's voice and facial expressions. All of this data is sent to the core module in the cloud for fusion processing, thereby improving the accuracy of emotion recognition. Based on the results of emotion recognition, the robot can adjust emotions to complete the companionship process, which has practical value.

[0144] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, extensions, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. An emotional companion robot system, characterized in that, include: The controller, cloud server, and non-contact vital sign collection module installed on the emotional companion robot; The controller receives user voice and facial images, respiratory rate, and heart rate collected by the non-contact vital sign acquisition module. It converts user voice into voice-text messages and sends all of these messages to the cloud server for processing. At the same time, it receives control commands returned by the cloud server after processing and controls the emotional companion robot to complete the corresponding actions. The non-contact vital sign acquisition module is used to recognize the user's facial image and send the recognition result to the cloud server through the controller; at the same time, it monitors the user's breathing and heart rate and sends them to the cloud server through the controller. The cloud server receives user audio content, converted voice-text messages, user facial images, breathing rate, and heart rate from the controller. It then judges the voice-text messages. If they are instruction-type messages, it generates instructions that can control the robot's actions. Otherwise, it analyzes the voice-text, user voice, user facial images, breathing rate, and heart rate, and feeds back the resulting control instructions to the controller. The cloud server includes: an analysis and processing module and a multimodal fusion module; The analysis and processing module is used to receive voice text, user voice, user facial image, respiratory rate and heart rate sent from the instruction analysis module, analyze them, obtain analysis results X1, X2, X3, X4 and X5, and send them to the multimodal fusion module. The multimodal fusion module, as the core of the cloud server, is used to send the results of weight assignment of X1, X2, and X3 to the instruction message generation module. The multimodal fusion module assigns weights to the recognition results X1, X2, and X3 sent by the controller to obtain the fusion result of X1, X2, and X3, specifically as follows: 2-1) Based on the X4 and X5 values ​​provided by the current heart rate and respiration analysis module, where X4 is the real-time heart rate and X5 is the real-time respiratory rate, the resting heart rate is set to 60-100 beats / minute, with 70 as the baseline value for X4. 基准 Corresponding to point 0 in [-1,1], with 130 as the upper limit of X4. 上限 , corresponding to ±1; Set the resting respiratory rate to 12-20 breaths per minute, using 13 as the baseline value multiplied by 5. 基准 Corresponding to point 0 in [-1,1], with 28 as the upper limit of X5. 上限 , corresponding to ±1; The generation coefficients for real-time heart rate X4 and real-time respiratory rate X5 are respectively: ; 2-2) Based on X1, X2, X3, the generation coefficients a4 and a5 of real-time heart rate X4 and real-time respiratory rate X5 are generated. At this time, a1, a2, and a3 are introduced. The default values ​​of a1, a2, and a3 are 5, 2, and 3, respectively. Users can set a1, a2, and a3 through the human-computer interaction interface according to the current usage scenario. The values ​​of a1, a2, and a3 all range from 1, 2, 3, 4, and 5, where 1 represents the weakest confidence level and 5 represents the strongest confidence level. These values ​​are combined with X1, X2, and X3 to generate the weighted result Y, i.e.: ; 2-3) Obtain the segmentation threshold a6 of the multimodal fusion result Z based on the coefficients a4 and a5 generated from the heart rate X4 and respiratory rate X5, i.e.: ; 2-4) Based on the segmented threshold a6, the multimodal fusion result Z can be expressed as: 。 2. The emotional companion robot companion system according to claim 1, characterized in that, The cloud server also includes: an instruction analysis module and an instruction message generation module; The instruction analysis module receives user audio content, converted voice-text messages, user facial images, breathing frequency, and heart rate from the controller, and analyzes the voice-text messages. If the message is an instruction, it is sent to the instruction message generation module to generate instructions that can control the robot's actions; otherwise, the voice-text, user voice, user facial images, breathing frequency, and heart rate are sent to the analysis and processing module for analysis. The instruction message generation module is used to process the content received from the instruction analysis module or the multimodal fusion module and generate control instructions that the robot can recognize.

3. The emotional companion robot companion system according to claim 2, characterized in that, The analysis and processing module includes: a text analysis module, an expression recognition module, a speech recognition module, a heartbeat analysis module, and a breathing analysis module connected to the multimodal fusion module; The text analysis module is used to receive voice text, perform intent recognition and text-based emotion recognition, and generate result X1; The facial expression recognition module is used to receive user videos sent by the video capture module, recognize the current facial expression, and generate result X2; The speech recognition module is used to receive speech audio, recognize the tone and voiceprint of the current speech, and generate result X3; The heartbeat analysis module analyzes the current heartbeat frequency and generates a result X4, where X4 represents the current real-time heartbeat frequency. The respiratory analysis module analyzes the current respiratory rate and generates a result X5, where X5 represents the current real-time respiratory rate. The sub-modules in the analysis and processing module send the corresponding generated results X1, X2, X3, X4, and X5 to the multimodal fusion module.

4. The emotional companion robot companion system according to claim 1, characterized in that, The non-contact vital sign acquisition module includes: camera A, respiratory rate sensor, heart rate sensor, and drive motor; Camera A, respiratory rate sensor, and heart rate sensor are connected to the controller to send the collected user facial image, respiratory rate, and heart rate to the controller, respectively. The respiratory rate sensor and heart rate sensor form an adjustable angle structure with camera A, which is used to ensure that the controller can control the drive motor so that camera A can track the face in real time, and during the tracking process, the respiratory rate sensor and heart rate sensor can be aligned with the user's chest cavity.

5. The emotional companion robot companion system according to claim 1, characterized in that, Also includes: The display module, voice input module, voice output module, and motion module are located within the emotional companion robot; The display module is connected to the controller and is used to play facial expressions, videos and music according to the control instructions sent by the controller, so as to enhance the robot's interactivity and increase the ways of emotional regulation; The voice output module is connected to the controller and is used to play the robot's audio according to the instructions generated by the instruction message generation module. The mobile module is connected to the controller and is used to receive control signals from the controller to drive the robot body to move as a whole using differential wheels.

6. The emotional companion robot companion system according to claim 1 or 5, characterized in that, The emotional companion robot includes: a robot body and a microphone (1), a display (2), a chest display (4), a motion chassis (5), a speaker (6), and a video acquisition module (7) installed on the robot body. The microphone (1) is arranged on the top of the robot body in the form of a microphone matrix and connected to the voice input module. When used for voice input, the user's position is obtained according to the microphone matrix, and then the robot controller gives a movement command to drive the robot body's motion chassis (5) to rotate so that the robot faces the user. The robot body is a human-like structure, and a display (2) is embedded in the head of the robot body to facilitate user interaction when standing. The controller is located inside the display (2) and connected to the display module; the display (2) is connected to the display module and is used to play facial expressions or complete the playback of videos and music to increase the ways of emotional regulation; A chest display (4) is embedded in the middle of the robot body to facilitate interaction when the user is sitting or lying down; the chest display (4) is connected to the display module to play expressions, or complete the playback of videos and music, and also display the interactive interface during video calls. A non-contact vital sign acquisition module is provided between the chest display (4) and the display (2); The motion chassis (5) is located at the bottom of the robot body and is connected to the moving module. The moving module drives the motion chassis (5) to move the robot body in a differential wheel manner to achieve the overall movement of the robot. The speaker (6) is located on both sides of the robot's head and is connected to the voice output module for playing the robot's audio feedback. There are two video acquisition modules (7), both of which are connected to the controller. The video acquisition module (7) includes a camera A located on the non-contact vital sign acquisition module and a camera B located on the head of the robot body. The video acquisition module (7) is used to acquire user facial images and send them to the controller.

7. The emotional companion robot companion system according to claim 1, characterized in that, The system's caregiving method includes the following steps: 1) The voice input module acquires the user's voice and converts the user's voice into a text message through the controller. The controller then sends the voice text and voice audio to the cloud server for further analysis. 2) The command analysis module in the cloud server analyzes the voice text to determine whether the user's voice is a command message. If it is a command message, it is sent to the command message generation module, and no further analysis is performed. The command information is returned to the controller in the robot. The controller controls the robot to perform the corresponding action based on the returned command. Otherwise, the voice text, user voice, user facial image, breathing rate, and heart rate are sent to their respective sub-modules of the analysis and processing module. 3) The respective sub-modules of the analysis and processing module in the cloud server receive the voice text, user voice, user facial image, respiratory rate, and heart rate, and perform their own analysis. The recognition results X1, X2, X3, X4, and X5 obtained after analysis are sent to the multimodal fusion module, and the fusion results are sent to the instruction message generation module. 4) The instruction message generation module generates corresponding control instructions and feeds them back to the controller, which then controls the emotional companion robot to execute the relevant instructions.

8. The emotional companion robot companion system according to claim 7, characterized in that, Step 3) specifically refers to: 1-1) The text analysis module performs intent recognition and text-based sentiment recognition on the speech text and generates a result X1. The range of this content is divided into [-1,1], that is, the generated result X1 is within the divided interval, which represents the sentiment tendency of the current text. 1-2) The expression recognition module identifies the current expression in the user's video and generates a result X2. The range of this content is divided into [-1,1]. That is, the generated result X2 is within the divided range, which represents the emotional tendency of the current expression. 1-3) The speech recognition module processes the speech audio and identifies the tone and voiceprint of the current speech, and generates a result X3. The range of this content is divided into [-1,1]; that is, the generated result X3 within the divided interval represents the emotional tendency of the current speech audio. 1-4) The heartbeat analysis module analyzes the current heartbeat frequency and generates a result X4, where X4 represents the current real-time heartbeat frequency; 1-5) The respiratory analysis module analyzes the current respiratory rate and generates a result X5, where X5 represents the current real-time respiratory rate.