Accompanying type intelligent robot system based on large model and emotion recognition

By combining multimodal perception and edge computing with cloud-edge collaborative control, a biomimetic tentacle actuator has been developed, which solves the problem of insufficient emotional understanding and physical interaction capabilities in existing companion robots, and achieves the effects of low-latency deep emotional understanding and fine operation.

CN121598057APending Publication Date: 2026-03-03GUANGXI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511804279.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing companion robots suffer from insufficient emotional understanding capabilities, high system response latency, and weak hardware interaction capabilities, making it difficult to meet the needs for low-latency, deep emotional understanding and anthropomorphic physical interaction.

Method used

Lightweight emotional feature extraction is achieved by using a multimodal perception and edge computing module, hierarchical response is achieved by combining a cloud-edge collaborative control module, fine operation is achieved by using a bionic tentacle actuator, and deep emotional representation and language generation are achieved through a cross-modal attention model.

Benefits of technology

It achieves low-latency emotional response and deep understanding, improves the accuracy of emotional interaction, and can perform precise physical operations, thus expanding application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121598057A_ABST
    Figure CN121598057A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and robots, in particular to an accompanying type intelligent robot system based on a large model and emotion recognition, and the system comprises a multi-mode perception and edge calculation module which is configured to be used for collecting visual and auditory data of a user, and carrying out the lightweight emotion feature extraction and initial emotion judgment locally on a robot; and the cloud large-scale model and fusion decision module is configured to receive the feature data from the edge calculation module, perform multi-modal emotion feature fusion through a cross-modal attention mechanism, and generate response content adaptive to the emotion based on the fused emotion representation. The feedback delay of a high-intensity emotional scene is reduced to a millisecond level through a cloud-edge collaborative hierarchical response mechanism, and meanwhile, the depth and the accuracy of common interaction are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and robotics, and in particular to an intelligent companion robot system and interaction method that combines a large language model, multimodal emotion recognition, cloud-edge collaborative computing and bionic actuators. Background Technology

[0002] With the aging population and the increasing number of people living alone, the market demand for intelligent companion robots with emotional interaction capabilities is becoming increasingly urgent. Most existing companion robot products suffer from the following technical shortcomings:

[0003] First, their emotional understanding is superficial, mostly based on pre-set scripts or single-modal emotion classification, and they cannot conduct in-depth analysis by combining context, micro-expressions, and tone of voice, resulting in a lack of empathy in their interactions.

[0004] Secondly, the system has high response latency. To achieve complex semantic understanding, all perceptual data usually needs to be uploaded to the cloud for processing. The queuing of network transmission and computing leads to excessively long response times, making it difficult to meet the real-time requirements of emotional interaction.

[0005] Secondly, the hardware interaction capabilities are weak. The actuator has a single function and cannot complete life assistance tasks that require delicate operation, such as gently soothing or flexibly handing over objects.

[0006] Therefore, there is an urgent need in this field for an integrated solution that can achieve low latency, deep emotional understanding, and anthropomorphic physical interaction. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of the prior art and provide a biomimetic companion robot system and method with rapid response, deep emotional understanding, and sophisticated physical interaction capabilities, so as to solve the problems mentioned in the background art.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a companion intelligent robot system based on large model and emotion recognition, comprising:

[0009] The multimodal perception and edge computing module is configured to collect the user's visual and auditory data and perform lightweight emotional feature extraction and preliminary emotional judgment locally on the robot.

[0010] The cloud-based large-scale model and fusion decision module is configured to receive feature data from the edge computing module, perform multimodal emotional feature fusion through a cross-modal attention mechanism, and generate emotionally adapted response content based on the fused emotional representation.

[0011] The cloud-edge collaborative control module is configured to implement a tiered response strategy: when the edge computing module's initial sentiment assessment indicates a high-intensity negative emotion or an emergency, it immediately triggers a local rapid response protocol and simultaneously uploads the data asynchronously to the cloud-fusion decision module; otherwise, it waits for the decision result from the cloud-fusion decision module before executing the response.

[0012] The biomimetic tentacle actuator module includes a segmented tentacle structure based on a carbon fiber skeleton and flexible silicone, and a control unit configured to perform motion trajectory planning based on a logarithmic spiral model and to achieve compliant force control using an adaptive PID controller with an integrated pressure sensor.

[0013] Preferably, the multimodal sensing and edge computing module includes:

[0014] The visual acquisition unit uses an RGB-D camera;

[0015] The audio acquisition unit uses a microphone array;

[0016] The lightweight emotion feature extraction includes: using a lightweight convolutional neural network to extract facial motion unit intensity values ​​and head pose angles from the video stream in real time; and extracting Mel frequency cepstral coefficients, fundamental frequency, and energy feature sequences from the audio stream.

[0017] Preferably, the multimodal sensing and edge computing module includes:

[0018] The visual acquisition unit uses an RGB-D camera;

[0019] The audio acquisition unit uses a microphone array;

[0020] The lightweight emotion feature extraction includes: using a lightweight convolutional neural network to extract facial motion unit intensity values ​​and head pose angles from the video stream in real time; and extracting Mel frequency cepstral coefficients, fundamental frequency, and energy feature sequences from the audio stream.

[0021] Preferably, the multimodal sensing and edge computing module includes:

[0022] The visual acquisition unit uses an RGB-D camera;

[0023] The audio acquisition unit uses a microphone array;

[0024] The lightweight emotion feature extraction includes: using a lightweight convolutional neural network to extract facial motion unit intensity values ​​and head pose angles from the video stream in real time; and extracting Mel frequency cepstral coefficients, fundamental frequency, and energy feature sequences from the audio stream.

[0025] Preferably, the multimodal sensing and edge computing module includes:

[0026] The visual acquisition unit uses an RGB-D camera;

[0027] The audio acquisition unit uses a microphone array;

[0028] The lightweight emotion feature extraction includes: using a lightweight convolutional neural network to extract facial motion unit intensity values ​​and head pose angles from the video stream in real time; and extracting Mel frequency cepstral coefficients, fundamental frequency, and energy feature sequences from the audio stream.

[0029] A companion-type intelligent robot interaction method based on large model and emotion recognition includes the following steps:

[0030] The robot collects multimodal data from the user through its own sensors;

[0031] The multimodal data is processed at the edge to extract lightweight sentiment features and perform preliminary sentiment judgment.

[0032] Based on the results of the preliminary sentiment assessment, a tiered response strategy is executed: if the sentiment state is determined to require a rapid response, the pre-set rapid response action on the edge side is executed immediately, and a cloud-based deep analysis request is initiated asynchronously; otherwise, the feature data is uploaded to the cloud and the system awaits its decision.

[0033] In the cloud, a cross-modal attention model is used to deeply fuse the received multimodal features to generate deep sentiment representations;

[0034] Based on the deep sentiment representation, a conditional large-scale language model is used to generate sentiment-adapted text responses.

[0035] Based on instructions sent from the cloud or triggered at the edge, the bionic tentacle actuator is controlled to complete the corresponding physical action. The motion trajectory of the physical action is planned based on a logarithmic spiral model and compliant force control is achieved through adaptive PID control.

[0036] Preferably, the lightweight emotion features include facial action unit (AU) intensity sequences extracted from visual data and acoustic feature sequences extracted from audio data; the preliminary emotion judgment is achieved through a temporal classification model.

[0037] Preferably, the cross-modal attention model calculates the weights among text, visual, and audio features through a self-attention mechanism, and encodes the weighted fused features to output the deep sentiment representation.

[0038] Preferably, the conditionalized large language model uses the deep emotional representation as a generation condition constraint to guide it in generating dialogue content that matches the user's current emotional state.

[0039] Preferably, when controlling the bionic tentacle actuator, the adaptive PID control dynamically adjusts the output torque of the servo motor based on the error between the real-time pressure feedback at the tentacle end and the desired contact force.

[0040] Compared with the prior art, the beneficial effects of the present invention are:

[0041] It achieves a balance between response speed and depth of understanding: through a hierarchical response mechanism that combines cloud and edge, the feedback latency in high-intensity emotional scenarios is reduced to the millisecond level, while ensuring the depth and accuracy of ordinary interactions.

[0042] It improves the accuracy of emotion insight: through feature-level fusion based on cross-modal attention mechanism, it can understand complex and contradictory emotional expressions, and realize the leap from "identifying emotions" to "understanding states".

[0043] Breaking through the bottleneck of physical interaction: the bionic tentacles, combined with biomechanical trajectory planning and compliant force control, enable the robot to perform human-like fine operations, greatly expanding its application scenarios.

[0044] Highly scalable: The modular design facilitates the integration of new sensors and algorithms, and can be adapted to diverse scenarios such as family companionship, elderly care, education and rehabilitation. Attached Figure Description

[0045] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0046] Figure 1 This is a block diagram of a companion intelligent robot interaction system based on a large model and emotion recognition according to the present invention;

[0047] Figure 2 This is a flowchart of a companion-type intelligent robot interaction method based on large model and emotion recognition according to the present invention. Detailed Implementation

[0048] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0049] Please see Figures 1 to 2 To achieve the above objectives, the present invention adopts the following technical solution:

[0050] In a first aspect, this invention provides a biomimetic companion robot system based on cloud-edge collaboration and multimodal emotion fusion. The system includes:

[0051] Multimodal perception and edge computing module: responsible for collecting user visual (such as RGB-D camera) and auditory (such as microphone array) data, running a lightweight model locally on the robot, extracting facial motion units, acoustic features, etc. in real time, and making preliminary emotion judgments.

[0052] The cloud-based large-scale model and fusion decision module receives feature data uploaded from the edge, performs deep feature fusion through a cross-modal attention mechanism, generates accurate deep sentiment representations, and drives a large-scale language model to generate sentiment-adapted responses.

[0053] The cloud-edge collaborative control module acts as the system's "dispatch center," implementing a tiered response strategy. When the edge determines a high-intensity negative emotion (e.g., a sadness probability > 0.8) or an emergency, it immediately triggers a rapid local response (e.g., playing music or gently approaching), while simultaneously uploading the data asynchronously to the cloud; otherwise, it awaits the cloud's in-depth decision-making results. This effectively resolves the conflict between latency and deep understanding.

[0054] Bionic tentacle actuator module: It adopts a segmented bionic structure of carbon fiber skeleton and flexible silicone. Its control unit plans a smooth and natural motion trajectory based on a logarithmic spiral model, and achieves safe and compliant force control through an adaptive PID controller with integrated pressure sensor, thereby performing delicate daily living assistance tasks.

[0055] Secondly, this invention provides an interaction method based on the aforementioned system. This method mainly includes the following steps: multimodal data acquisition, lightweight edge processing and preliminary judgment, cloud-edge collaborative hierarchical response, cloud-to-cloud multimodal fusion and deep decision-making, emotion-adaptive content generation, and compliant control of a bionic tentacle.

[0056] Example 1: System Hardware and Software Configuration

[0057] Edge computing unit: Uses NVIDIA Jetson series embedded development board as main controller to run lightweight neural network models (such as MobileNetV3, SqueezeBERT).

[0058] Perception sensors: An Intel RealSense series RGB-D camera is used for visual perception, and a ring microphone array is used for audio acquisition and sound source localization.

[0059] Cloud services: Deploy large-scale language models (such as LLaMA, ChatGLM) and cross-modal fusion models based on cloud computing platforms (such as Alibaba Cloud, AWS).

[0060] Bionic tentacles: They use a 3D-printed carbon fiber skeleton covered with medical-grade flexible silicone. Each joint segment is driven by a micro servo motor, and a thin-film pressure sensor is integrated at the end.

[0061] Example 2: Interaction Flow Example

[0062] User A sighed and said "I'm fine" (text), but visual analysis showed that the corners of his mouth were drooping (high AU15 intensity), and audio analysis showed that his tone was low and his speech was slow.

[0063] Fast edge detection: The lightweight edge-side model integrates visual and audio features to quickly calculate the probability of high-intensity sadness (0.85), triggering the LEVEL1 response.

[0064] Immediate Action: The cloud-edge collaborative control module instructs the robot to slowly approach user A and play locally stored soothing music. (Latency < 100ms)

[0065] Deep cloud-based analysis: Simultaneously, data is uploaded to the cloud. A cross-modal attention model comprehensively analyzes text, AU features, and acoustic features to identify the deep emotional state of "forced smiles."

[0066] Generate an empathetic response: Based on this deep state, the conditionalized large language model generates a response: "You sound like you might be a little tired today. I'll keep you company. Would you like me to get you a glass of water?"

[0067] Compliant execution: The cloud sends a reply text and a "pour water" command. The touch control unit plans the cup-grabbing trajectory based on a logarithmic spiral and uses adaptive PID control to achieve safe grasping and delivery.

[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A companion intelligent robot system based on large-scale modeling and emotion recognition, characterized in that, include: The multimodal perception and edge computing module is configured to collect the user's visual and auditory data and perform lightweight emotional feature extraction and preliminary emotional judgment locally on the robot. The cloud-based large-scale model and fusion decision module is configured to receive feature data from the edge computing module, perform multimodal emotional feature fusion through a cross-modal attention mechanism, and generate emotionally adapted response content based on the fused emotional representation. The cloud-edge collaborative control module is configured to implement a tiered response strategy: when the edge computing module's initial sentiment assessment is a high-intensity negative emotion or an emergency state, it immediately triggers a local rapid response protocol and simultaneously uploads the data asynchronously to the cloud-based fusion decision module. Otherwise, wait for the decision result from the cloud-integrated decision module before executing the response; The biomimetic tentacle actuator module includes a segmented tentacle structure based on a carbon fiber skeleton and flexible silicone, and a control unit configured to perform motion trajectory planning based on a logarithmic spiral model and to achieve compliant force control using an adaptive PID controller with an integrated pressure sensor.

2. The companion intelligent robot system based on large model and emotion recognition according to claim 1, characterized in that, The multimodal sensing and edge computing module includes: The visual acquisition unit uses an RGB-D camera; The audio acquisition unit uses a microphone array; The lightweight emotion feature extraction includes: using a lightweight convolutional neural network to extract facial motion unit intensity values ​​and head pose angles from the video stream in real time; and extracting Mel frequency cepstral coefficients, fundamental frequency, and energy feature sequences from the audio stream.

3. The companion intelligent robot system based on large model and emotion recognition according to claim 1, characterized in that, The multimodal sensing and edge computing module includes: The visual acquisition unit uses an RGB-D camera; The audio acquisition unit uses a microphone array; The lightweight emotion feature extraction includes: using a lightweight convolutional neural network to extract facial motion unit intensity values ​​and head pose angles from the video stream in real time; and extracting Mel frequency cepstral coefficients, fundamental frequency, and energy feature sequences from the audio stream.

4. The companion intelligent robot system based on large model and emotion recognition according to claim 1, characterized in that, The multimodal sensing and edge computing module includes: The visual acquisition unit uses an RGB-D camera; The audio acquisition unit uses a microphone array; The lightweight emotion feature extraction includes: using a lightweight convolutional neural network to extract facial motion unit intensity values ​​and head pose angles from the video stream in real time; and extracting Mel frequency cepstral coefficients, fundamental frequency, and energy feature sequences from the audio stream.

5. The companion intelligent robot system based on large model and emotion recognition according to claim 1, characterized in that, The multimodal sensing and edge computing module includes: The visual acquisition unit uses an RGB-D camera; The audio acquisition unit uses a microphone array; The lightweight emotion feature extraction includes: using a lightweight convolutional neural network to extract facial motion unit intensity values ​​and head pose angles from the video stream in real time; and extracting Mel frequency cepstral coefficients, fundamental frequency, and energy feature sequences from the audio stream.

6. A companion-style intelligent robot interaction method based on large-scale models and emotion recognition, characterized in that, Includes the following steps: The robot collects multimodal data from the user through its own sensors; The multimodal data is processed at the edge to extract lightweight sentiment features and perform preliminary sentiment judgment. Based on the results of the preliminary sentiment assessment, a tiered response strategy is executed: if the sentiment state is determined to require a rapid response, the pre-set rapid response action on the edge side is executed immediately, and a cloud-based deep analysis request is initiated asynchronously; otherwise, the feature data is uploaded to the cloud and the system awaits its decision. In the cloud, a cross-modal attention model is used to deeply fuse the received multimodal features to generate deep sentiment representations; Based on the deep sentiment representation, a conditional large-scale language model is used to generate sentiment-adapted text responses. Based on instructions sent from the cloud or triggered at the edge, the bionic tentacle actuator is controlled to complete the corresponding physical action. The motion trajectory of the physical action is planned based on a logarithmic spiral model and compliant force control is achieved through adaptive PID control.

7. The companion intelligent robot interaction method based on large model and emotion recognition according to claim 6, characterized in that, The lightweight emotion features include facial action unit (AU) intensity sequences extracted from visual data and acoustic feature sequences extracted from audio data; the preliminary emotion judgment is achieved through a temporal classification model.

8. The companion intelligent robot interaction method based on large model and emotion recognition according to claim 6, characterized in that, The cross-modal attention model calculates the weights among text, visual, and audio features through a self-attention mechanism, and encodes the weighted fused features to output the deep sentiment representation.

9. The companion intelligent robot interaction method based on large model and emotion recognition according to claim 6, characterized in that, The conditionalized large-scale language model uses the deep emotional representation as a generation condition constraint to guide it in generating dialogue content that matches the user's current emotional state.

10. The companion intelligent robot interaction method based on large model and emotion recognition according to claim 6, characterized in that, When controlling the bionic tentacle actuator, the adaptive PID control dynamically adjusts the output torque of the servo motor based on the error between the real-time pressure feedback at the tentacle end and the desired contact force.

Citation Information

Cited By

  • Family accompanying spherical unmanned aerial vehicle family accompanying method and system

    CN121812103A