A home service robot with integrated voice, gesture, and facial expression control.
The recognition module, which integrates voice, gestures, and facial expressions, utilizes convolutional neural networks to identify joint coordinates and skeletal features. This solves the problem that existing robots cannot accurately recognize multiple interaction methods, enabling efficient multimodal interaction and emotion-driven home services.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Utility models(China)
- Current Assignee / Owner
- SUZHOU ZHONGKE ADVANCED TECH RES INST CO LTD
- Filing Date
- 2025-08-29
- Publication Date
- 2026-07-31
AI Technical Summary
Existing home service robots cannot accurately recognize various methods such as voice commands, gesture recognition, and facial expressions, resulting in a high rate of misrecognition, unnatural responses, and a lack of emotional interaction, making it difficult to meet users' personalized and natural home companionship and service needs.
It employs a recognition module that integrates voice, gesture, and facial expression control. It recognizes user actions and voice through a camera and microphone, uses a convolutional neural network architecture to identify joint coordinates and skeletal features, constructs a gesture framework, and combines it with an emotional state fusion feedback module to achieve multimodal interaction and dynamic control.
It improves the accuracy and efficiency of gesture recognition, enhances the robot's ability to interact with users, better meets users' personalized needs, and provides natural emotional interaction and home services.
Smart Images

Figure CN224575679U_ABST
Abstract
Description
Technical Field
[0001] This utility model relates to the field of robotics technology, and in particular to an emotionally-coordinated home service robot with integrated voice, gesture and facial expression control. Background Technology
[0002] Existing home service robots mostly rely on a single interaction method, such as voice or gesture control, which has problems such as a high rate of interaction misrecognition, unnatural response, and lack of emotional interaction. Although some robots have emotion recognition capabilities, the integration of multimodal interaction and emotion-driven dynamic control strategies are not yet mature, making it difficult to meet users' personalized and natural home companionship and service needs.
[0003] This is because existing robots cannot recognize multiple methods such as voice commands, gesture recognition, and facial expressions, meaning they cannot accurately capture the user's emotional state, thus failing to improve the interaction between the robot and the user. Utility Model Content
[0004] To address the problem that existing robots cannot accurately recognize multiple interaction methods, this invention proposes an emotionally-coordinated home service robot with integrated voice, gesture, and facial expression control.
[0005] This utility model is achieved through the following technical solution:
[0006] This utility model proposes an emotion-cooperative home service robot with integrated voice, gesture, and facial expression control. Its features include a recognition module, an analysis module, a processing module, a feedback control module, and a learning module disposed within the robot's main body, wherein:
[0007] The processing module is connected to the recognition module, the analysis module, the feedback control module, and the learning module, respectively, and the analysis module is connected to the feedback control module and the learning module, respectively.
[0008] The identification module is used to receive identification signals and transmit the signals to the processing module;
[0009] The processing module is used to convert and process the received signal;
[0010] The analysis module is used to analyze the converted signal;
[0011] The feedback control module is used to receive and analyze the signals and then control the robot accordingly.
[0012] Furthermore, the recognition module includes a voice recognition module and an image recognition module. The image recognition module is connected to at least one camera disposed on the robot body, and the voice module is connected to at least one microphone disposed on the robot body. The voice recognition module and the image recognition module are connected.
[0013] Furthermore, the feedback control module includes a drive module and a voice module for controlling the robot body.
[0014] Furthermore, it also includes a sensing module and an adaptive module, wherein the sensing module is connected to multiple sensors disposed around the robot body, and the adaptive module is connected to the sensing module and the feedback module respectively.
[0015] Furthermore, the image recognition module is used to recognize the images captured by the camera, and the recognition steps of the image recognition module include:
[0016] S1. Perform frame-by-frame segmentation and recognition of the video, and establish multiple sets of joint coordinate features;
[0017] S2. Establish a skeletal feature set based on the joint coordinate feature set and the recognition image;
[0018] S3. Construct a gesture framework based on the set of joint coordinate features and the set of skeletal features;
[0019] S4. Arrange the gesture frames in each frame in sequence and determine the gesture characteristics.
[0020] Furthermore, the establishment of multiple joint coordinate feature sets in step S1 includes:
[0021] S11. Identifying joint coordinates of different sizes based on a convolutional neural network architecture;
[0022] S12. Arrange the joint coordinates according to their different sizes and directions;
[0023] S13. Establish a set of joint coordinate features based on the set of joint coordinates and arrangement directions.
[0024] Furthermore, the step of constructing the gesture framework based on the set of joint coordinate features and skeletal features in step S2 includes:
[0025] S21. Identification of multiple skeletal curves of different sizes based on a convolutional neural network architecture;
[0026] S22. Establish a set of skeletal features based on multiple skeletal curves of different sizes;
[0027] S23. Determine the bone orientation based on skeletal features and correct the orientation;
[0028] S24. Incorporate the bone orientation into the bone feature set.
[0029] Furthermore, the step of constructing the gesture framework based on joint coordinate features and the skeleton set in step S3 includes:
[0030] S31. Substitute the skeleton feature set into the corresponding joint coordinate feature set;
[0031] S32. Perform algorithmic correction on skeletal features and joint coordinate features to remove incorrect joint coordinate arrangement directions;
[0032] S33. Connect the skeletal features and joint coordinate features to construct a gesture framework.
[0033] Furthermore, it also includes a smart home linkage module, which is connected to the feedback module.
[0034] Furthermore, it also includes an emotional state fusion feedback module, which is connected to the analysis module.
[0035] The beneficial effects of this utility model are:
[0036] (1) The emotion-coordinated home service robot recognition module with voice, gesture and facial expression fusion control proposed in this utility model receives and recognizes external signals, and then transmits them to the processing module. The processing module converts and processes the signals and sends them to the analysis module for analysis. After analysis, the analysis module transmits the signals to the feedback module and controls the robot to make corresponding control. The learning module can learn the user's behavioral habits to make it easier to accurately recognize them. By using multiple modules, it can better interact with the user to complete interaction and home control.
[0037] (2) The present invention proposes an emotionally coordinated home service robot with voice, gesture and facial expression fusion control. The present invention splits the screen and re-identifies it according to the set of joint coordinate features and the set of skeletal features. Compared with the screen recognition module, which recognizes the human body's raising hand movement as a complete gesture, it can improve the accuracy and efficiency of gesture recognition. Attached Figure Description
[0038] Figure 1 This is a structural diagram of the emotionally-coordinated home service robot of this utility model, which integrates voice, gesture, and facial expression control.
[0039] In the diagram: Recognition module 1, Speech recognition module 12, Image recognition module 13, Analysis module 2, Processing module 3, Feedback control module 4, Voice module 41, Driving module 42, Learning module 5, Smart home linkage module 6, Emotional state fusion feedback module 7, Perception module 8, Adaptive module 9;
[0040] The purpose, features, and advantages of this utility model will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0041] To more clearly and completely illustrate the technical solution of this utility model, the following description, in conjunction with the accompanying drawings, will further explain this utility model.
[0042] Please refer to Figure 1 This utility model proposes an emotionally-coordinated home service robot with integrated voice, gesture, and facial expression control, comprising a recognition module 1, an analysis module 2, a processing module 3, a feedback control module 4, and a learning module 5, all housed within the robot's main body.
[0043] The processing module 3 is connected to the identification module 1, the analysis module 2, the feedback control module 4, and the learning module 5, respectively; the analysis module 2 is connected to the feedback control module 4 and the learning module 5, respectively.
[0044] The identification module 1 is used to receive the identification signal and transmit the signal to the processing module 3;
[0045] The processing module 3 is used to convert and process the received signal;
[0046] The analysis module 2 is used to analyze the converted signal;
[0047] The feedback control module 4 is used to receive and analyze the signals and then control the robot accordingly.
[0048] In a specific implementation, the identification module 1 receives and identifies external signals, and then transmits them to the processing module 3. The processing module 3 converts and processes the signals and sends them to the analysis module 2 for analysis. After analysis, the analysis module 2 transmits the signals to the feedback module and controls the robot to make corresponding control. The learning module 5 can learn the user's behavioral habits to make it easier to accurately identify them. This utility model uses multiple modules to better interact with the user and complete interaction and home control.
[0049] Furthermore, the recognition module 1 includes a voice recognition module 121 and an image recognition module 131. The image recognition module 131 is connected to at least one camera disposed on the robot body, and the voice module 41 is connected to at least one microphone disposed on the robot body. The voice recognition module 121 and the image recognition module 131 are connected.
[0050] In a specific implementation, the voice recognition module 121 recognizes the user's voice through a microphone, and the image recognition module 131 recognizes the user's actions through a camera. The voice recognition module 121 and the image recognition module 131 transmit the received voice and animation to the processing module 3. The processing module 3 then processes the user's voice and actions and transmits it to the analysis module 2 for analysis.
[0051] Furthermore, the feedback control module 4 includes a drive module 42 and a voice module 41 for controlling the robot body.
[0052] In a specific implementation, the drive module 42 controls the robot to provide feedback, and the robot interacts with the user by emitting sounds through the voice module 41.
[0053] Furthermore, it also includes a sensing module 8 and an adaptive module 9, wherein the sensing module 8 is connected to multiple sensors disposed around the robot body, and the adaptive module 9 is connected to the sensing module 8 and the feedback module respectively.
[0054] In a specific implementation, the robot is equipped with multiple sensors around it, including infrared sensors and depth cameras. The robot can perceive its surrounding environment in real time through its built-in sensors and automatically avoid obstacles, walls, etc. through intelligent path planning algorithms, allowing it to move freely in the environment. The adaptive module 9 can also adjust the robot's working trajectory according to the spatial layout and the user's activity area.
[0055] Furthermore, the image recognition module 131 is used to recognize the image captured by the camera, and the recognition steps of the image recognition module 131 include:
[0056] S1. Perform frame-by-frame segmentation and recognition of the video, and establish multiple sets of joint coordinate features;
[0057] S2. Establish a skeletal feature set based on the joint coordinate feature set and the recognition image;
[0058] S3. Construct a gesture framework based on the set of joint coordinate features and the set of skeletal features;
[0059] S4. Arrange the gesture frames in each frame in sequence and determine the gesture characteristics.
[0060] In a specific implementation, each identified frame is split for identification. For example, if a 15-second frame is identified, it is split into 15 seconds and each non-hand frame is removed. Then, the frame is identified frame by frame to establish a set of joint coordinate features. After establishing the set of joint coordinate features, the frame is identified and a set of skeletal features is established in conjunction with the set of joint coordinate features. Finally, the set of joint coordinate features and the set of skeletal features are filled into the same frame, and a gesture frame is constructed. Finally, the gesture features are determined based on the gesture frame of each frame.
[0061] In one embodiment, errors can easily occur when recognizing gestures. For example, a user may be too thin to make a complete gesture, but the image recognition module 131 may recognize the user's hand raising as a complete gesture. This invention splits the image and re-recognizes the gesture according to the set of joint coordinate features and the set of skeletal features, which can improve the accuracy and efficiency of gesture recognition.
[0062] Furthermore, the establishment of multiple joint coordinate feature sets in step S1 includes:
[0063] S11. Identifying joint coordinates of different sizes based on a convolutional neural network architecture;
[0064] S12. Arrange the joint coordinates according to their different sizes and directions;
[0065] S13. Establish a set of joint coordinate features based on the set of joint coordinates and arrangement directions.
[0066] In a specific implementation, joint coordinates refer to the connection points of human joints and bones. The size of the connection points of each bone in the human body is different. A convolutional neural network architecture is used to identify each joint coordinate. After identification, the joint coordinates are arranged according to their different sizes. Since the joint coordinates of the fingers are smaller than those of the wrist, and the joint coordinates of the wrist are smaller than those of the elbow, the approximate direction from the elbow to the fingers of the entire hand can be determined by the size of the joint coordinates. Then, a set of joint coordinate features is established by combining the joint coordinates and the arrangement direction.
[0067] Furthermore, the step of constructing the gesture framework based on the set of joint coordinate features and skeletal features in step S2 includes:
[0068] S21. Identification of multiple skeletal curves of different sizes based on a convolutional neural network architecture;
[0069] S22. Establish a set of skeletal features based on multiple skeletal curves of different sizes;
[0070] S23. Determine the bone orientation based on skeletal features and correct the orientation;
[0071] S24. Incorporate the bone orientation into the bone feature set.
[0072] In a specific implementation, the skeletal curve is the human skeleton. The skeleton is identified using a convolutional neural network architecture, and a set of skeletal features is built. The skeletal direction is then determined based on the skeletal features and corrected using algorithms, such as the direction of the fingers and the direction of the arms, to make it conform to the human body structure. Finally, the skeletal direction is added to the skeletal set.
[0073] Furthermore, the step of constructing the gesture framework based on joint coordinate features and the skeleton set in step S3 includes:
[0074] S31. Substitute the skeleton feature set into the corresponding joint coordinate feature set;
[0075] S32. Perform algorithmic correction on skeletal features and joint coordinate features to remove incorrect joint coordinate arrangement directions;
[0076] S33. Connect the skeletal features and joint coordinate features to construct a gesture framework.
[0077] In a specific implementation, the skeletal feature set is input into the corresponding joint coordinates, the skeletal curves and joint coordinates are placed on the same screen and connected, and a complete gesture feature is formed after connection. Then, the algorithm is used to correct it based on the skeletal direction and the incorrect joint coordinate direction, and the incorrect joint coordinate arrangement direction data is removed. Finally, the joint coordinate direction and skeletal direction are filled in the joint coordinates and skeletal curves, thereby constructing a gesture framework. Through the gesture framework combined with voice, better interaction with the user can be achieved.
[0078] In one embodiment, the user's gestures are recognized by a camera, and deep learning algorithms are used to improve recognition accuracy. For example, the user can control the robot to turn lights on and off by waving their hand, or instruct the robot to perform a task by pointing to a specific location.
[0079] Furthermore, it also includes a smart home linkage module 6, which is connected to the feedback module.
[0080] In a specific implementation: the smart home linkage module 6 supports linkage control with smart home devices to achieve environmental and emotional adaptation; the robot controls smart home lighting, curtains and audio devices through wireless protocols, and adjusts the environmental atmosphere according to the user's emotional state and needs, such as automatically dimming the lights and playing soothing music when the user is tired. At the same time, the robot perceives the surrounding environment in real time through built-in sensors and automatically avoids obstacles, walls and other obstacles through intelligent path planning algorithms, and moves freely in the environment. The adaptive module 9 can also adjust the robot's working trajectory according to the spatial layout and the user's activity area.
[0081] Furthermore, it also includes an emotional state fusion feedback module 7, which is connected to the analysis module 2.
[0082] In a specific implementation: the emotional state fusion feedback module 7 includes an emotional analysis engine and an emotional feedback system. The emotional analysis engine can combine emotional information from voice, gestures and facial expressions to analyze the user's emotional state, such as pleasure, anxiety, fatigue, etc., and generate appropriate feedback based on the analysis results. The emotional feedback system can interact with the user through various means such as the robot's voice, actions, and lighting effects based on the user's emotional feedback to enhance emotional resonance.
[0083] In one embodiment, this invention can be applied to elderly care. The robot can recognize and determine the needs of the elderly through voice, gestures and facial expressions, such as reminding them to take their medicine or providing food. When the elderly show signs of fatigue or anxiety, the robot will adjust its behavior through emotion analysis, providing soothing voices or gentle movements, while avoiding disturbing the elderly's daily activities.
[0084] Of course, there may be other implementations of this utility model. Based on this implementation, other implementations obtained by those skilled in the art without any creative effort are all within the scope of protection of this utility model.
Claims
1. An emotion-matching home service robot with voice, gesture and expression fusion control, characterized in that, This includes a recognition module, an analysis module, a processing module, a feedback control module, and a learning module, all housed within the robot's main body. The processing module is connected to the recognition module, the analysis module, the feedback control module, and the learning module, respectively, and the analysis module is connected to the feedback control module and the learning module, respectively. The identification module is used to receive identification signals and transmit the signals to the processing module; The processing module is used to convert and process the received signal; The analysis module is used to analyze the converted signal; The feedback control module is used to receive and analyze the signals and then control the robot accordingly. 2.The emotion-matching home service robot with voice, gesture and expression fusion control according to claim 1, wherein, The recognition module includes a voice recognition module and an image recognition module. The image recognition module is connected to at least one camera mounted on the robot body, and the voice module is connected to at least one microphone mounted on the robot body. The voice recognition module and the image recognition module are connected. 3.The emotion-matching home service robot with voice, gesture and expression fusion control according to claim 1, wherein, The feedback control module includes a drive module and a voice module for controlling the robot body. 4.The emotion-matching home service robot with voice, gesture and expression fusion control according to claim 1, wherein, It also includes a perception module and an adaptive module, wherein the perception module is connected to multiple sensors disposed around the robot body, and the adaptive module is connected to the perception module and the feedback module respectively. 5.The emotion-matching home service robot with voice, gesture and expression fusion control according to claim 1, wherein, It also includes a smart home linkage module, which is connected to the feedback module. 6.The emotion-matching home service robot with voice, gesture and expression fusion control according to claim 1, wherein, It also includes an emotional state fusion feedback module, which is connected to the analysis module.