Remote emotion interaction robot system based on real-time pose projection

Through real-time posture projection technology and multi-sensor fusion, the shortcomings of remote interactive robots in slow movement response and facial expression display are solved, efficient emotional interaction and scene adaptation are achieved, and the user experience and practicality of the robot are improved.

CN120606400APending Publication Date: 2025-09-09NORTHEASTERN UNIV AT QINHUANGDAO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511002756.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing remote interactive robots have problems in cross-space action interaction, such as slow response and large delay, and are unable to truly and fully display human expressions and three-dimensional features. They lack emotional state feedback and have difficulty providing a deep and realistic interactive experience, which limits their applicability in scenarios such as family, elderly care, and children's education.

Method used

A remote emotional interaction robot system based on real-time pose projection is adopted. It combines the user-side system and the robot-side, utilizes multi-sensor fusion, precise calculation and intelligent filtering technology, and performs multi-model fusion through MediaPipe's Face Mesh and Pose models to capture the user's posture and facial data in real time, achieving accurate motion mapping and emotional feedback.

Benefits of technology

It improves the naturalness and realism of the interaction, enhances the robot's scene adaptability and practical performance, can adaptively adjust the interaction mode, provide more real psychological support and emotional companionship, and enhance the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120606400A_ABST
    Figure CN120606400A_ABST
Patent Text Reader

Abstract

The invention provides a remote emotion interaction robot system based on real-time pose projection, and relates to the technical field of robots. Limb actions, facial micro-expressions and voice information of a user are captured in real time through a multi-sensor fusion technology (a camera, a somatosensory glove, a pressure sensor and the like), 468 facial key points and 33 body joint points are extracted in combination with a MediaPipe model, and low-delay and high-precision action smooth mapping is realized by adopting Kalman filtering and an exponential weighted moving average (EMA) algorithm. Modularized bionic design is adopted for the robot, the bionic arms, the head and the breathing and heartbeat bionic module driven by flexible steering engines are arranged, the action rigidity, speed and contact strength can be dynamically adjusted, and human natural body languages and vital signs are simulated. A real-time communication link is constructed through a TCP / IP protocol, dual-mode interaction of real-time projection and a preset action library is supported, and the system is suitable for scenes such as remote nursing accompanying, parent-child interaction and psychotherapy healing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robotics technology, and in particular to a remote emotional interaction robot system based on real-time posture projection. Background Art

[0002] Against the backdrop of an aging population, a more specialized social division of labor, and an increasing number of people living alone, the demand for emotional companionship is growing. Emotional companion robots are becoming crucial in settings such as families, elderly care, and childcare. These robots mimic human body language, expressions, and movement habits to establish emotional connections with users, providing companionship, communication, and psychological comfort. For example, in family settings, robots can enhance interactive fun by mimicking the user's gestures, body movements, or dance. In elderly care settings, remote children can project real-time motions onto robots to perform intimate gestures like hugs and caresses on their behalf, alleviating loneliness. In partner settings, these robots offer a more authentic emotional interaction experience. However, current emotional companion robots face significant challenges in cross-spatial motion interaction. In particular, technical bottlenecks exist in capturing the user's emotional movements in real time and naturally mapping them to the robot itself, limiting the depth and realism of the emotional interaction.

[0003] To address these challenges, researchers have developed a number of new technologies to enhance robots' emotional expression capabilities and interactive effects. (1) Humanoid design and emotional expression: Existing humanoid companion interactive robots mostly adopt a humanoid appearance design, which can better convey emotional information. Robots are usually equipped with components such as facial expressions, eye movements, and body movements to simulate human natural behavior and achieve richer emotional expression. (2) Preset action interaction library based on keyword extraction: Robots select and execute appropriate actions or behaviors by recognizing the user's voice or text input. They can not only respond to emotional changes but also enhance the emotional connection with the user.

[0004] However, while the field of remote interactive robots has made some progress, many significant problems still exist. Currently, most remote interactive robots are extremely limited in their interactive presentation methods. They rely solely on simple display devices such as screens to achieve facial image mapping, remaining at a two-dimensional level. They completely lack three-dimensional bionic-level mapping, and are unable to truly and comprehensively display the rich expressions and three-dimensional features of the human face, making it difficult to provide users with an immersive interactive experience. In terms of action response, they generally have shortcomings such as slow action response and large delays. They are unable to follow user instructions or imitate human actions in a timely and accurate manner, which makes the interaction process stiff and unsmooth, seriously affecting the user experience. At the same time, traditional robot remote interaction must rely on professional equipment, which severely limits the conditions and scenarios for the robot's use.

[0005] At the same time, remote interactive robots with basic voice conversation capabilities still have significant flaws. They often lack effective emotional state feedback mechanisms, are unable to perceive and understand users' emotional changes, and are unable to respond accordingly. This makes it difficult to provide users with a truly "real-time presence" interactive experience.

[0006] The resulting chain reaction of these issues has left existing robots severely deficient in expressing complex body language or subtle nonverbal cues. Complex body language and subtle nonverbal cues play a crucial role in human communication, conveying a rich tapestry of emotions, intentions, and attitudes. This, to a certain extent, limits their suitability for high-quality companionship. Whether chatting with the elderly to relieve boredom, accompanying children in their learning and growth, or providing emotional support to those with special needs, existing robots struggle to meet users' demands for deep, authentic interactive experiences. They are unable to truly integrate into human life scenarios and deliver the ideal companionship value. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to address the deficiencies of the above-mentioned existing technologies and provide a remote emotional interaction robot system based on real-time posture projection, which significantly improves the interactive realism and scene adaptability, and provides an efficient and humanized solution for the field of emotional companionship.

[0008] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0009] A remote emotional interaction robot system based on real-time posture projection, including a user-side system, a user wearable module and a robot side;

[0010] The user-side system is responsible for processing, identifying and calculating posture and facial data, including a personal computer-based video and audio receiving module, a limb motion recognition and angle calculation module, a pressure recognition module, a hand micro-motion recognition module, a facial expression motion recognition module, a voice recognition module, and a command recognition processing and decision-making module; the video and audio receiving module is used to collect the user's real-time full-body motion video stream mov1 and the audio stream wav1 that supports voice commands and real-time transmission through the camera and microphone on the personal computer, and store them in the pre-processing buffer area; the limb motion recognition and angle calculation module is used to process the video stream collected by the video receiving module, mark the response joint point information, and identify the limb motion to calculate the key joint angles, including the head posture angle, the eye and mouth opening and closing rate, and the arm posture angle; the pressure The force recognition module is used to transmit the data of the pressure sensor in the user-worn module in real time; the hand micro-motion recognition module transmits the hand micro-motion in real time through the lightweight finger motion sensor worn by the user; the facial expression and movement recognition module performs high-precision facial algorithm recognition through the video stream mov2 intercepted by the second camera on the robot side; the voice recognition module is responsible for identifying the user's conversation voice based on the audio stream feedback data in real time, and screening the content of the user's voice command response, and responding to the corresponding prompt content to guide the robot side to express; the command recognition and processing module is responsible for integrating and judging the data of multiple module sensors, identifying according to the limb position status, pressure, and acceleration, and generating command data for the corresponding operation control data, encapsulating the relevant commands into a fixed-format command package, and sending it to the robot side;

[0011] The user-worn module is used to sense various subtle hand movements and pressure changes between the body and external objects. It includes a pressure sensing module and a lightweight finger motion sensor. The pressure sensing module is worn on the user's arm and joints to measure movement force and contact pressure. The lightweight finger motion sensor is used to monitor micro-movements of the hand. The user-worn module is connected to the user-end system via TCP / IP or I2C.

[0012] The robot is responsible for playing audio and executing actions, and includes a Raspberry Pi-based central processing module and several execution modules. The execution modules include two bionic arms, two bionic hands, a bionic head consisting of ten servos, an information acquisition module consisting of two cameras and two microphones, a voice module, a breathing and heartbeat bionic module, and a pressure feedback sensor. The central processing module parses instructions and distributes them to the execution modules. Each execution module is connected to the central processing module via PWM or I2C to implement instruction execution and interaction.

[0013] Each of the two bionic arms contains three servos, which respectively control the forward and backward movement of the shoulder, the left and right rotation, and the bending and straightening of the elbow;

[0014] The bionic head contains eyes, mouth, and neck modules, and includes 26 motors that are position-controlled. The flexible facial skin is easily connected to the hardware mechanism via magnets. The neck module supports three-axis rotation: roll, pitch, and yaw. Three motors control the movement of the neck on these three axes. Twelve motors control the upper face, including the eyeballs, eyelids, and eyebrows. Eleven motors control the mouth mechanism and mandible. The magnetic connection design allows the robot's facial skin to be easily replaced. The eye module has a magnetically connected linkage that controls the eyebrows, including the upper eyelid, lower eyelid, eyeball linkage, and eyeball frame. The mouth module uses six passive mouth linkages.

[0015] The second camera is used to detect user distance and environmental perception; the second microphone cooperates with the voice module to support voice interaction and emotional voice feedback; the breathing and heartbeat bionic module is used to simulate chest rise and fall, vibration, and heartbeat; the pressure feedback sensor is used to detect the strength of the hug and adjust the softness of the action;

[0016] The robot and user systems communicate via TCP / IP.

[0017] Furthermore, the user-side system calculates posture and facial data using a real-time posture projection method, using a combination of multi-sensor fusion, precise calculation, intelligent filtering, preset actions, and multi-threading technology, and uses MediaPipe's Face Mesh and Pose models for multi-model fusion. The calculation steps are as follows:

[0018] S1: Initialize model parameters and wait for the user to appear and the interaction to begin;

[0019] S2: When the user first appears, the user-side system defines the initial arm length and facial key line lengths at the user's spatial location, uses this information as a reference for subsequent angle calculations, and simultaneously announces a welcome message through the voice module to initiate interaction.

[0020] S3: During the interaction process, the camera 1 and microphone 1 of the video and audio receiving module collect the user's movements, expressions, and voice information, and transmit them to the body movement recognition and angle calculation module in the user-side system for real-time calculation of human-robot dynamic mapping;

[0021] Adaptive adjustments are made to user movements based on the robot's joint range of motion, degrees of freedom of movement, and flexible limb characteristics to avoid dangerous movements that exceed mechanical limits.

[0022] S4: The user-side system transmits the calculated head posture angle, eye and mouth opening and closing rate, and arm posture angle to the robot-side central processing module based on the Raspberry Pi. The central processing module parses the instructions and distributes them to the execution modules to control the robot's posture movements and play real-time voice.

[0023] S5: Repeat steps S1 to S4 until the interaction ends.

[0024] Furthermore, when performing real-time calculation of the human-robot dynamic mapping in the S3 step, the limb motion recognition and angle calculation module in the user-end system constructs a camera projection model based on the perspective projection principle, establishes a mapping relationship between 3D head model points and 2D image points, combines the camera matrix and distortion parameters, iteratively optimizes and solves the head rotation and displacement vectors, converts the rotation matrix into Euler angle representation, and at the same time limits the angle to a physically feasible range through a custom range judgment function.

[0025] Furthermore, when performing real-time calculation of the human-robot dynamic mapping in step S3, in terms of limb projection, the limb motion recognition and angle calculation module in the user-side system uses the key points of the upper body shoulders, elbows, wrists, neck, limb contours, and facial contours identified by MediaPipe's Pose model to achieve accurate analysis and projection of the arm posture; the limb motion recognition and angle calculation module extracts the key point information of the right shoulder, right elbow, and right wrist of the right body and the left shoulder, left elbow, and left wrist of the left body from the detection results of the Pose model. In the first frame image, the distance between the right shoulder and elbow and between the elbow and wrist is calculated as the arm length reference value, and the corresponding arm length on the left is also calculated. These reference values ​​provide an important reference basis for subsequent angle calculations; for the angle calculation of the arm posture, the limb motion recognition and angle calculation module uses geometric and trigonometric functions to calculate the angles required for multiple servo control; at the same time, in terms of calculation, the limb motion recognition and angle calculation module uses geometric decomposition to decompose 3D motion into multiple 2D planar motions; after the decomposition is completed, the limb motion recognition and angle calculation module uses trigonometric functions to calculate the angles of each joint.

[0026] In the servo control phase, the robot-side system's command recognition, processing, and decision-making module compares the calculated control angle with a preset physical safety range. If the angle exceeds the range, it is adjusted to within the safety range. The command recognition, processing, and decision-making module introduces an exponentially weighted moving average method to smooth the calculated angle by setting a smoothing factor. The smoothed angle is then calculated and output based on historical angle values ​​and current measurements.

[0027] In estimating the head pitch angle, the Euclidean distance between the midpoint of the left and right inner canthus and the tip of the nose is calculated, combined with the inverse sine function and different scaling factors to achieve accurate estimation. For the roll angle, a dual verification mechanism is adopted, which is the calculation of the horizontal and vertical difference between the two eyes and the iterative optimization result based on the perspective projection model. The specific implementation is as follows: first, the pixel coordinates of the left and right inner canthus and the tip of the nose are located on the image plane, the coordinates of the midpoint of the left and right inner canthus are calculated, and then the pixel distance between the midpoint coordinate and the tip of the nose is solved. The reference distance and physical size information obtained in the calibration stage are combined with the trigonometric function relationship to estimate the head pitch angle. Different scaling factors are applied to the movement characteristics of looking up and looking down to improve the estimation accuracy. For the roll angle, on the one hand, the ratio of the vertical pixel difference to the horizontal pixel difference of the left and right eye centers is calculated, and the inverse tangent function is used to solve the preliminary roll angle. On the other hand, based on the correspondence between the 3D head model points and the 2D image points, a nonlinear optimization problem is constructed. The rotation matrix is ​​iteratively solved by minimizing the reprojection error, and the roll angle component is extracted from it. Finally, the roll angles obtained by the two methods are verified for consistency and fused.

[0028] Six key points out of 468 key points on the face are extracted through MediaPipe, which are located at the edge of the eyes, the tip of the nose, the corner of the mouth and the chin. 2D and 3D head pose arrays are constructed. The camera matrix K is calculated based on the Euler rotation theorem according to the following matrix formula (1), where the simulated camera focal length is taken as the image width, the image center is set as the center of the coordinate system, and the tilt parameter of the camera matrix is ​​set to 0;

[0029] (1);

[0030] Among them, P c is a point in the camera coordinate system, expressed as a two-dimensional coordinate (X C , Y C );P w is a point in the world coordinate system, expressed as a three-dimensional coordinate ; K is the camera's internal parameter matrix, including the focal length f x 、f y and light center c x 、c y ; R is the rotation matrix, which represents the rotation from the world coordinate system to the camera coordinate system; T is the translation matrix, which represents the translation from the world coordinate system to the camera coordinate system; f x 、f y is the focal length of the camera, corresponding to the x and y directions respectively; c x 、c y is the principal point coordinate of the camera, i.e. the optical center; r ij and t x ,t y ,t zare the elements of the rotation matrix and the elements of the translation vector, respectively, representing the transformation from the world coordinate system to the camera coordinate system;

[0031] Given the 3D and 2D coordinates and the camera matrix, the perspective n-point PnP problem is used to solve the above equation (1) to obtain the rotation matrix R. Then, the direction angle of the head posture is obtained by performing RQ decomposition on the rotation matrix R, which is used to control the robot's head angle to imitate human head movement.

[0032] Furthermore, the facial expression and action recognition module of the user-side system has the ability to calibrate the reference distance and angle in real time, and can dynamically adjust the scaling factors used for head posture angle estimation, which are the head-up angle scaling factor UP_SCALE_FACTOR, the head-down angle scaling factor DOWN_SCALE_FACTOR, the head-right turn scaling factor RIGHT_SCALE_FACTOR, and the head-left turn scaling factor LEFT_SCALE_FACTOR. After the facial expression and action recognition module completes the calibration of the reference distance and angle, the solvePnP algorithm and RQDecomp3 are used to calculate the head posture angle. The Euler angles of the head are obtained by x3 decomposition, and then the radians are converted to degrees. Finally, the corresponding scaling factor is applied according to the direction of head movement for linear adjustment to obtain the final head pitch and roll angle control data; the facial expression and action recognition module will dynamically apply the corresponding scaling factor to adjust the estimated angle according to different head movements. By default, the values ​​of UP_SCALE_FACTOR, RIGHT_SCALE_FACTOR and LEFT_SCALE_FACTOR are 1.5, and the value of DOWN_SCALE_FACTOR is 0.5. The default values ​​are dynamically adjusted according to actual usage scenarios and needs.

[0033] Furthermore, the receiving and identifying processes of each module of the user terminal system are as follows:

[0034] The video and audio stream receiving process of the video and audio receiving module is as follows: the user-side system collects video and audio by calling the camera 1 and microphone 1 in the personal computer. After collection, the audio wav1 and videos mov1 and mov2 are first processed into tracks, and the timestamps are saved during processing to ensure the consistency of the video and audio timing; the processed video and audio streams are subjected to noise reduction processing, and the processed video and audio streams are placed in the buffer area and wait for the system program to process them;

[0035] The limb movement recognition process of the limb movement recognition and angle calculation module is as follows: the video stream mov1 returned by the video and audio receiving module is passed into the limb movement recognition and angle calculation module, and the CV2 algorithm based on the OpenCV library is used to identify 33 key points of the torso, and the real-time torso projection distance is measured according to the real-time measured key point spacing. In each subsequent frame, the projection distance Ldet of the key point in the X-axis and Y-axis directions is calculated: for the left arm, the vertical projection distance L_arm_left_2 is the absolute value of the coordinate difference of the left arm key point in the Y-axis direction, and the calculation formula is L_arm_left_2=|y2-y1|; where y1 and y2 are the Y-axis coordinates of the two adjacent key points of the left arm respectively; the horizontal projection distance L_arm_left_3 is the absolute value of the coordinate difference of the left arm key point in the X-axis direction, and the calculation formula is L_arm_left_3=|x2-x1|; where x1 and x2 are the X-axis coordinates of the two adjacent key points of the left arm respectively;

[0036] The projection ratio is calculated using the trigonometric relationship ratio = Ldet / Lreal, where Lreal represents the actual physical reference length of the corresponding part of the user's limb, which is obtained through initialization calibration during the user's first interaction.

[0037] The control angle of different servos is calculated using the following specific formula:

[0038] Shoulder left and right rotation servo angle θ1: θ1 = arcsin (L_arm_left_3 / Lreal_arm_left), where Lreal_arm_left is the actual physical reference length of the corresponding part of the left arm;

[0039] Elbow flexion and extension servo angle θ2: θ2=arccos (L_arm_left_2 / Lreal_arm_left);

[0040] Shoulder forward and backward motion servo angle θ3: θ3 = arcsin (L_arm_right_3 / Lreal_arm_right), where L_arm_right_3 is the horizontal projection distance of the right arm, and Lreal_arm_right is the actual physical reference length of the corresponding part of the right arm;

[0041] At the same time, a timestamp is added to each frame of action data after recognition to facilitate subsequent instruction recognition processing and decision module processing;

[0042] The pressure recognition process of the pressure recognition module is as follows: the voltage values ​​of the multi-touch pressure sensors and acceleration sensors worn by the user on the arms and joints are converted by ADC to measure the speed and force data of the user when performing the corresponding action ( ), and send the corresponding data to the instruction recognition and processing module;

[0043] The hand micro-motion recognition module's hand micro-motion recognition process is as follows: A lightweight finger motion sensor responds to changes in finger joint motion data in real time and calculates the corresponding finger joint angle rotation data; the gravity sensor and geomagnetic sensor in the lightweight finger motion sensor are used to calculate the wrist rotation angle and pitch angle data based on the corresponding voltage data, and the finger and wrist data are transmitted to the command recognition processing and decision module;

[0044] The facial expression recognition process of the facial expression action recognition module is as follows: through the intercepted video stream mov2, combined with the camera intrinsic parameter matrix and distortion parameters, the solvePnP algorithm is used to solve the rotation vector rot_vec and the translation vector trans_vec, and the corresponding head joint angle data is calculated and sent back to the command recognition processing and decision module;

[0045] The speech recognition module's speech recognition process is as follows: the user-side system uses the processed audio stream wav1 to identify audio command keywords, and then performs corresponding processing on the recognized preset keyword commands to guide the robot's expression. At the same time, the user-side system also encapsulates the entire video stream based on whether the user selects real-time speech transmission and transmits it to the robot.

[0046] The instruction recognition and processing process of the instruction recognition, processing and decision-making module is as follows: during the instruction processing process, the data previously input by the limb movement recognition and angle calculation module, pressure sensor, hand micro-movement recognition module, and facial expression movement recognition module are integrated as a whole to sort out the stiffness coefficient, damping coefficient, torque coefficient, speed, and angle of the corresponding action joint; at the same time, the time sequence is aligned according to the torso movement and audio timestamp information, and the data information of the same time sequence is organized into a dictionary, encapsulated in json format, and transmitted to the robot end using the TCP / IP protocol.

[0047] Furthermore, the collaboration process between the robot side, the user wearable module and the user side system is as follows:

[0048] Step 1-1: The entire emotional interaction robot system is initialized and placed in an idle state. After the robot is powered on, each module is initialized in sequence: the joint motors are reset to a neutral position, with the cervical spine at 0° and the arms hanging naturally. Camera 2 initiates a self-test and establishes a connection with the user system via the TCP / IP protocol. Simultaneously, in the idle state, the two bionic arms open and close slightly in a cyclical manner, and the voice module plays low-frequency breathing sounds at a specific frequency to simulate human breathing rhythms.

[0049] Step 1-2: Once the robot's camera detects the human silhouette, the user's system video and audio receiving module extracts the coordinates of the human hip joint and calculates the user's horizontal angle relative to the robot. A cervical rotation command is sent, causing the cervical joint to rotate to the corresponding angle, so that the robot's line of sight faces the user. Simultaneously, the voice module plays a preset welcome sound effect, replacing complex voice interaction. Real-time tracking of all joints is also enabled, and joint status feedback from the robot hardware is received via the local area network.

[0050] Steps 1-3: During the interaction process, the robot uses the command recognition processing and decision-making module to calculate the human body's motion posture and transmit visual posture data to the robot side. Through the preset action library, it realizes the dual-mode interaction of "real-time projection" and "specific action triggering";

[0051] Steps 1-4: After receiving the instructions from the command recognition and decision-making modules, the central processing module on the robot drives the motors to synchronize with the user's movements in real time. When executing the "wave" or "hug" commands, the joints coordinate to complete the corresponding actions according to the preset trajectory or trigger conditions.

[0052] Steps 1-5: When the user leaves the camera's field of view for 30 consecutive seconds, the robot triggers the exit process: the command recognition and decision-making module stops posture data processing and sends a "reset" command. The robot returns its head to its original position, lowers its arms, and the display returns to its soft breathing lighting effect. At the same time, all components on the robot enter a low-power standby state, maintain a LAN connection, and maintain basic functions in low-power mode while waiting to be woken up. When the user leaves the camera's field of view, that is, no human body is detected for 30 consecutive seconds, the robot enters the exit process.

[0053] Step 1-6: Repeat steps 1-2 to 1-5 until the interaction ends.

[0054] Furthermore, the steps 1-3 specifically include the following steps:

[0055] Step 2-1: Use the cv2.VideoCapture(0) function of the OpenCV library to capture the video stream from the camera and resize the video frame to 640×480 pixels; use MediaPipe Face Mesh technology to detect 468 facial key points and 33 body key points;

[0056] Step 2-2: Posture calculation;

[0057] First, the head pose is accurately calculated. By applying the solvePnP algorithm, key points in the two-dimensional image are mapped to three-dimensional space. During this process, the rotation vector rot_vec and translation vector trans_vec are solved by combining the camera intrinsic parameter matrix and distortion parameters. Then, using the Rodrigues transform, the rotation vector is converted into a rotation matrix, which is further decomposed into Euler angles, including pitch, yaw, and roll, to fully describe the rotation state of the head.

[0058] Secondly, the calculation of limb posture; by using trigonometric geometric relationships, the rotation angles of the shoulder, elbow and wrist are calculated in detail;

[0059] Finally, the calculated angle data is processed using Kalman filtering technology. At the same time, the data is smoothed using the exponential moving average method (EMA) to ensure that more stable and accurate angle information can be obtained in real-time applications.

[0060] Step 2-3: In terms of data format and communication protocol, the user first accurately maps the required control angle information and converts it into a specific command format suitable for servo control. Then, using Socket communication technology and the reliable TCP / IP protocol, the converted servo control commands are stably and efficiently sent to the Raspberry Pi, ensuring the accuracy and real-time nature of command transmission.

[0061] The beneficial effects of adopting the above technical solution are as follows: the remote emotional interaction robot system based on real-time posture projection provided by the present invention can collect multi-dimensional motion data such as human posture, gesture details and facial micro-expressions in real time with the help of camera visual recognition technology. Lightweight finger motion sensors are used to accurately convey hand micro-movements and strength. Compared with traditional video stream time interval measurement, this method can provide more accurate and low-latency force sensing. By extracting limb and facial movements through the MediaPipe facial mesh model, the remote robot can accurately capture the user's emotional body language and synchronously generate interactive movements, greatly enhancing the naturalness and realism of companionship interaction. The visual processing algorithm designed by the present invention, based on MediaPipe multi-model fusion and spatiotemporal sequence analysis, extracts facial and body key points through the Face Mesh and Pose models, combines them with the OpenCV library for video stream processing, and uses spatiotemporal sequence analysis to effectively distinguish target users from interfering objects, solving the problem of pose recognition in scenarios such as occlusion, lighting changes, and multi-target coexistence. It ensures the continuous and accurate capture of user movements, avoids interaction interruptions or movement dislocations caused by environmental interference, improves the coherence and reliability of posture mapping interactions, and enhances the practical performance of robots in complex environments. The human-robot dynamic mapping model established by the present invention adaptively adjusts user movements based on the robot's joint range of motion, degrees of freedom of movement, and flexible limb characteristics. It uses an inverse kinematics algorithm to avoid dangerous movements that exceed mechanical limits. It adjusts the speed, acceleration, and contact force parameters of the robot's movements based on pressure sensors on the body to determine the user's emotional needs and emotional state. Based on these physiological signals, the robot can adaptively adjust its interaction mode to more accurately respond to the user's emotional state and help users alleviate negative emotions. The present invention incorporates bionic breathing and heartbeat modules, enabling the robot to simulate human vital signs, responding to natural breathing and heartbeats, giving the user a more lifelike companionship experience. This provides a more authentic psychological support, particularly for soothing and alleviating emotions such as anxiety and loneliness. The modular bionic mechanical structure designed in this invention supports both real-time and preset interaction strategies based on different emotional companionship scenarios, enabling a differentiated emotional companionship experience and enhancing the robot's practicality and scope of use. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 This is an architecture diagram of the interactive system between the user-side system, user-worn module, and robot-side provided by an embodiment of the present invention;

[0063] Figure 2 A diagram showing the appearance of the robot side of the emotional interaction robot provided in an embodiment of the present invention;

[0064] Figure 3 A flowchart of human-robot posture capture and mapping based on real-time posture projection provided by an embodiment of the present invention;

[0065] Figure 4 Schematic diagram of key parameters of joint freedom for upper body movement provided by an embodiment of the present invention.

[0066] In the picture, 1. Camera 2; 2. Neck module; 3. Right shoulder servo; 4. Right elbow servo; 5. Right manipulator; 6. Chassis; 7. Left manipulator; 8. Left elbow servo; 9. Left shoulder servo. DETAILED DESCRIPTION

[0067] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0068] This embodiment builds a robot system for cross-space natural emotional interaction, focusing on the core path of "high-precision posture analysis - emotional action reproduction - lightweight interactive deployment" to achieve technical breakthroughs. The user end relies on a personal computer and a lightweight wearable module (camera 1, inertial sensor) to directly connect to the robot end through the TCP / IP protocol, avoiding the dependence on professional interactive equipment. Figure 1 The system uses MediaPipe's Face Mesh model to extract six key facial features (eye rims, nose tip, mouth corners, and chin). Combined with the Pose model, it captures the 3D coordinates of limb joints (shoulder, elbow, and wrist), constructing a spatiotemporal sequence of head and limb poses.

[0069] Limb movements are decomposed into 2D planar motions using a geometric decomposition algorithm. Trigonometric functions are used to calculate servo control angles. A ratio constraint mechanism is introduced to physically constrain joint angles, reproducing emotional details such as hug strength and movement speed. This aligns with the angle correction logic of the limb movement recognition and angle calculation modules. Head posture is inferred using a PnP algorithm. Euler angles are extracted based on simulated camera matrix K and facial keypoint pairs, driving the robot head to mimic human movement. The posture calculation process is described in Equation (1).

[0070] Furthermore, this embodiment also creates a human-machine fusion interaction mode, allowing users to experience a certain sense of "avatar" when operating the robot. Through scene adaptability and personalized design, the method and content of gesture projection are optimized to meet the needs of different users and scenarios, ultimately creating a natural cross-space emotional interaction experience. The details are as follows.

[0071] A remote emotional interaction robot system based on real-time posture projection includes a user-side system, a user wearable module and a robot side.

[0072] The user-side system is responsible for processing, identifying and calculating posture and facial data, including a video and audio receiving module, a body movement recognition and angle calculation module, a pressure recognition module, a hand micro-movement recognition module, a facial expression movement recognition module, a voice recognition module, and a command recognition processing and decision-making module; the video and audio receiving module is used for the camera and microphone on the personal computer to receive the user's real-time movement video stream mov1 and audio stream wav1, and store them in the pre-processing buffer area; the body movement recognition and angle calculation module is used to process the video stream collected by the video receiving module, mark the response joint point information, and identify the body movement and its amplitude; the pressure recognition module is used to transmit the data of the pressure sensor in the user's wearable module in real time. ; The hand micro-motion recognition module transmits hand micro-motions in real time through the lightweight finger motion sensor worn by the user; the facial expression and movement recognition module performs high-precision facial algorithm recognition through the video stream mov2 intercepted by the robot's camera 21; the voice recognition module is responsible for real-time recognition of the user's conversation voice based on the audio stream feedback data, and screens the content of the user's voice command response, and responds to the corresponding prompt content to guide the robot's expression; the command recognition and processing module is responsible for integrating and judging the data of multiple module sensors, identifying according to the limb position status, pressure, and acceleration, and generating command data for the corresponding operation control data, encapsulating the relevant commands into a fixed-format command package, and sending it to the robot.

[0073] The user-worn module is used to sense various subtle movements of the hand and changes in pressure between the body and external objects, including a pressure sensing module and a lightweight finger motion sensor; the pressure sensing module is worn on the user's arm and joints to measure movement force and contact pressure; the lightweight finger motion sensor is used to monitor micro-movements of the hand; the user-worn module is connected to the user-end system via TCP / IP or I2C.

[0074] The robot is responsible for playing audio and executing actions. It includes a Raspberry Pi-based central processing module and several execution modules. These modules include two bionic arms, two bionic hands (i.e., the left manipulator 7 and the right manipulator 5), a bionic head composed of ten servos, an information acquisition module consisting of camera 2 (1) and microphone 2 (2), a voice module, a bionic breathing and heartbeat module, and a pressure feedback sensor. The central processing module parses commands and distributes them to the execution modules. Each execution module is connected to the central processing module via PWM or I2C to implement command execution and interaction. The robot and user systems communicate via TCP / IP.

[0075] The appearance of the robot is as follows Figure 2As shown in Figure 2, each bionic arm contains three servos, which control the forward and backward movement of the shoulder, left and right rotation, and elbow flexion and extension. Specifically, the forward and backward movement of the left and right shoulders is controlled by the left shoulder servo 9 and the right shoulder servo 3, respectively, allowing the arm to push forward or pull backward. Another servo controls the left and right rotation of the shoulder, allowing the arm to rotate horizontally, similar to the external and internal rotation of the shoulder. The elbow is equipped with a servo (right elbow servo 4 and left elbow servo 8) to control elbow flexion and extension.

[0076] The bionic head contains eyes, mouth, and neck module 2, which contains 26 motors and adopts position control; the flexible facial skin is easily connected to the hardware mechanism through magnets; the neck module 2 supports three-axis rotation, namely roll, pitch and yaw, and three motors control the movement of the neck on the three axes respectively; 12 motors control the upper face, including eyeballs, eyelids and eyebrows; 11 motors control the mouth mechanism and mandible; the magnetic connection design makes the robot's facial skin easy to replace; the eye module has a magnetically connected link to control the eyebrows, including the upper eyelid, lower eyelid, eyeball link, and eyeball frame; the mouth module uses 6 passive mouth links.

[0077] Camera 2 (1) detects user proximity and environmental awareness. Microphone 2 works in conjunction with the voice module to support voice interaction and emotional voice feedback. A breathing and heartbeat bionic module, located within the robot's body shell, simulates chest movement and vibrations, as well as heartbeats. A pressure feedback sensor detects hug strength and adjusts the softness of the hug. The robot's body is fixed to base (6).

[0078] like Figure 3 The figure below shows the flow chart for human-robot posture capture and mapping based on real-time posture projection, covering four major steps: data acquisition, feature extraction, data processing and analysis, and motion control. The data acquisition stage uses a camera to collect data through MediaPipe's multi-task model inference processing, combined with lightweight finger motion sensors. The feature extraction stage obtains 468 3D facial coordinates, 33 skeletal coordinates of posture, and hand micro-motions (changes in finger joints and wrist rotation angles). The data processing and analysis stage uses a Kalman filter for time series smoothing and a posture conversion algorithm to achieve coordinate transformation for human-robot mapping. The motion control stage generates motion commands, calculates servo motion parameters, and transmits these data via TCP to the robot receiver. The servo control system (PCA9685 for PWM control) ultimately executes the action.

[0079] The user-side system calculates posture and facial data using a real-time pose projection method, which integrates multi-sensor fusion, precise calculation, intelligent filtering, preset actions, and multi-threading technology, and uses MediaPipe's Face Mesh and Pose models for multi-model fusion.

[0080] In actual operation, the system leverages multi-sensor fusion technology to integrate data from different sensor types, providing a richer and more accurate information foundation for subsequent pose analysis. Precise calculation technology ensures accurate calculation of various pose parameters, enabling the robot to more meticulously mimic human movements. Intelligent filtering technology effectively removes noise and interference from the data, improving data reliability and stability and making the robot's movements smoother and more natural. Preset motion technology provides the system with templates for common movements that can be quickly invoked when specific scenarios are detected, enhancing the system's responsiveness and practicality. Face Mesh accurately detects the 3D mesh and key points of the face. By analyzing facial landmarks, the system accurately calculates head pitch, yaw, and roll angles, enabling precise capture of head pose. The Pose model precisely identifies key points on the upper body, such as the shoulders, elbows, and wrists. Based on this key point location information, the system can conduct in-depth analysis of posture movements such as arm extension, flexion, and rotation. The two models work together to improve the accuracy and comprehensiveness of pose detection, enabling the robot to more realistically mimic a variety of human movements.

[0081] In terms of processing efficiency, multi-threading technology parallelizes camera image capture, head pose communication, and upper body pose communication tasks, ensuring real-time processing of large amounts of image data and avoiding lag. The body movement recognition and angle calculation module, based on the principle of perspective projection, constructs a camera projection model and establishes a mapping relationship between 3D head model points and 2D image points. Combining the camera matrix with distortion parameters, iteratively optimizes and solves the head rotation and displacement vectors, converting the rotation matrix into Euler angle representation. A custom range judgment (boundary judgment) function is used to constrain the angles to within the physically feasible range. Smoothing uses an exponentially weighted moving average method to calculate and output smoothed angles based on historical angle values ​​and current measurements. A smoothing factor is also introduced in the head movement point recognition class for weighted averaging, effectively reducing angle fluctuations. Inter-thread synchronization mechanisms are also used to ensure data consistency.

[0082] The specific calculation steps are as follows:

[0083] S1: Initialize model parameters and wait for the user to appear and the interaction to begin.

[0084] S2: When the user first appears, the user-side system defines the initial arm length and facial key line length at the user's spatial position, and uses this information as a reference for subsequent angle calculations. At the same time, a welcome message is broadcast through the voice module to initiate interaction.

[0085] S3: During the interaction, the video and audio receiving module's camera 1 and microphone 1 capture the user's movements, expressions, and voice information, and transmit it to the body movement recognition and angle calculation module in the user-side system for real-time calculation of the human-robot dynamic mapping. Based on the robot's joint range of motion, degrees of freedom, and flexible limb characteristics, the module adaptively adjusts the user's movements to avoid dangerous movements that exceed mechanical limits.

[0086] S4: The user-side system transmits the calculated head posture angle, eye and mouth opening and closing rate, and arm posture angle to the robot-side central processing module based on Raspberry Pi, which parses the instructions and distributes them to each execution module to control the robot's posture movements and play real-time voice.

[0087] S5: Repeat steps S1 to S4 until the interaction ends.

[0088] When performing real-time calculation of human-robot dynamic mapping in step S3, the limb motion recognition and angle calculation module in the user-side system builds a camera projection model based on the principle of perspective projection, establishes a mapping relationship between 3D head model points and 2D image points, combines the camera matrix and distortion parameters, iteratively optimizes and solves the head rotation and displacement vectors, converts the rotation matrix into Euler angle representation, and limits the angle to a physically feasible range through a custom range judgment function.

[0089] Regarding body projection, the limb motion recognition and angle calculation module in the user-side system uses MediaPipe's Pose model to identify key points of the upper body—the shoulders, elbows, wrists, neck, body contours, and facial outlines—to accurately analyze and project arm posture. From the Pose model's detection results, the module extracts key points of the right shoulder, elbow, and wrist on the right side of the body, and the left shoulder, elbow, and wrist on the left side. In the first frame, it calculates the distances between the right shoulder and elbow, and between the elbow and wrist, as arm length benchmarks. It also calculates the corresponding arm length on the left side. These benchmarks provide important references for subsequent angle calculations. To calculate arm angles, the module uses geometric and trigonometric methods to calculate the angles required for multiple servo control. Furthermore, the module employs geometric decomposition to decompose 3D motion into multiple 2D planar motions. This is because 2D planar motion is relatively simple to analyze and calculate, and mathematically, it has mature and efficient methods for processing it. After the disassembly is completed, the limb movement recognition and angle calculation module uses trigonometric functions to calculate the angles of each joint.

[0090] To ensure the rationality of the calculation results and avoid angle values ​​that are inconsistent with actual physical conditions, the system introduces ratio constraints. The calculated joint angles can be filtered and corrected to ensure that the final joint angles are within a reasonable range. Within the servo control process, a series of measures are implemented to ensure safe and stable servo operation. The robot-side system command recognition, processing, and decision-making module compares the calculated control angle with a preset physical safety range. If the angle exceeds the range, it is adjusted to within the safe range. To reduce fluctuations in angle calculations, the robot-side system command recognition, processing, and decision-making module introduces an exponentially weighted moving average method. This smoothing factor is used to smooth the calculated angles. The smoothed angle is then output based on historical angle values ​​and current measurements. This method combines historical angle values ​​with current measurements to achieve more stable output angles, avoiding drastic fluctuations in angles caused by measurement errors or sudden changes in movement, and improving the stability and accuracy of limb projection.

[0091] To estimate the pitch angle, the Euclidean distance between the midpoint of the left and right inner canthus and the tip of the nose is calculated, combined with the inverse sine function and different scaling factors to achieve accurate estimation. For the roll angle, a dual verification mechanism is used: the horizontal and vertical differences between the two eyes are calculated, and an iterative optimization method based on a perspective projection model is used. Specifically, the pixel coordinates of the left and right inner canthus and the tip of the nose are located on the image plane. The coordinates of the midpoint of the left and right inner canthus are calculated, and then the pixel distance between this midpoint and the tip of the nose is calculated. The pitch angle is estimated using the reference distance and physical dimensions obtained during the calibration phase, combined with trigonometric relationships. Different scaling factors are applied to improve estimation accuracy based on the characteristics of head-up and head-down movements. For the roll angle, the ratio of the vertical pixel difference to the horizontal pixel difference between the left and right eye centers is calculated, and the inverse tangent function is used to determine the preliminary roll angle. Furthermore, based on the correspondence between 3D head model points and 2D image points, a nonlinear optimization problem is constructed. The rotation matrix is ​​iteratively solved by minimizing the reprojection error, and the roll angle component is extracted from it. Finally, the roll angles obtained by the two methods are verified for consistency and fused.

[0092] Six key points out of 468 facial key points are extracted through MediaPipe, which are located at the edge of the eyes, the tip of the nose, the corner of the mouth and the chin. 2D and 3D head pose arrays are constructed. The camera matrix K is calculated based on the Euler rotation theorem according to the following matrix formula (1), where the simulated camera focal length is taken as the image width, the image center is set as the center of the coordinate system, and the tilt parameter of the camera matrix is ​​set to 0.

[0093] (1);

[0094] Among them, P c is a point in the camera coordinate system, expressed as a two-dimensional coordinate (X C, Y C );P w is a point in the world coordinate system, expressed as a three-dimensional coordinate ; K is the camera's internal parameter matrix, including the focal length f x 、f y and light center c x 、c y ; R is the rotation matrix, which represents the rotation from the world coordinate system to the camera coordinate system; T is the translation matrix, which represents the translation from the world coordinate system to the camera coordinate system; f x 、f y is the focal length of the camera, corresponding to the x and y directions respectively; c x 、c y is the principal point coordinate of the camera, i.e. the optical center (usually the center of the image); r ij and t x ,t y ,t z The elements of the rotation matrix and the elements of the translation vector, respectively, represent the transformation from the world coordinate system to the camera coordinate system.

[0095] Given the 3D and 2D coordinates and the camera matrix, the perspective n-point PnP problem is used to solve the above equation (1) to obtain the rotation matrix R. Then, the direction angle of the head posture is obtained by performing RQ decomposition on the rotation matrix R, which is used to control the robot's head angle to imitate human head movement.

[0096] The facial expression and action recognition module of the user-side system has the ability to calibrate the reference distance and angle in real time, and can dynamically adjust the scaling factors used for head posture angle estimation. These scaling factors are the head-up angle scaling factor UP_SCALE_FACTOR, the head-down angle scaling factor DOWN_SCALE_FACTOR, the head-right turn scaling factor RIGHT_SCALE_FACTOR, and the head-left turn scaling factor LEFT_SCALE_FACTOR. After the facial expression and action recognition module completes the calibration of the reference distance and angle, it obtains the Euler angle of the head through the solvePnP algorithm and RQDecomp3x3 decomposition, then converts the radian system to the degree system. Finally, the corresponding scaling factor is applied according to the direction of head movement for linear adjustment to obtain the final head pitch angle and swing angle control data. The facial expression and action recognition module dynamically applies the corresponding scaling factor to adjust the estimated angle based on different head movements. By default, the values ​​of UP_SCALE_FACTOR, RIGHT_SCALE_FACTOR, and LEFT_SCALE_FACTOR are 1.5, and the value of DOWN_SCALE_FACTOR is 0.5. The default values ​​are dynamically adjusted according to actual usage scenarios and needs to improve the system's adaptability to different users and environments and enhance the accuracy of head posture recognition.

[0097] The receiving and identification processes of each module in the user-side system are as follows:

[0098] The video and audio stream receiving process of the video and audio receiving module is as follows: the user-end system calls the camera 1 and microphone 1 based on the personal computer to collect video and audio. After collection, the audio wav1, video mov1 and mov2 are first processed into tracks, and the timestamps are saved during processing to ensure the consistency of video and audio timing; the processed video stream and audio stream are subjected to noise reduction processing, and are placed in the cache area after processing, waiting for the system program to process.

[0099] The limb motion recognition process of the limb motion recognition and angle calculation module is as follows: the video stream mov1 returned by the video and audio receiving module is passed into the limb motion recognition and angle calculation module, and the 33 key points of the torso are identified based on the CV2 algorithm of the OpenCV library, and the real-time torso projection distance is measured according to the real-time measured key point spacing. In each subsequent frame, the projection distance Ldet of the key point in the X-axis and Y-axis directions is calculated: for the left arm, the vertical projection distance L_arm_left_2 is the absolute value of the coordinate difference of the left arm key point in the Y-axis direction, and the calculation formula is L_arm_left_2=|y2-y1|, where y1 and y2 are the Y-axis coordinates of the two adjacent key points of the left arm respectively; the horizontal projection distance L_arm_left_3 is the absolute value of the coordinate difference of the left arm key point in the X-axis direction, and the calculation formula is L_arm_left_3=|x2-x1|, where x1 and x2 are the X-axis coordinates of the two adjacent key points of the left arm respectively.

[0100] The projection ratio is calculated using the trigonometric relationship ratio = Ldet / Lreal. Lreal represents the actual physical reference length of the corresponding part of the user's limb, obtained through initial calibration during the user's first interaction.

[0101] The control angle of different servos is calculated using the following specific formula:

[0102] Shoulder left and right rotation servo angle θ1: θ1 = arcsin (L_arm_left_3 / Lreal_arm_left), where Lreal_arm_left is the actual physical reference length of the corresponding part of the left arm;

[0103] Elbow flexion and extension servo angle θ2: θ2=arccos (L_arm_left_2 / Lreal_arm_left);

[0104] Shoulder forward and backward motion servo angle θ3: θ3 = arcsin (L_arm_right_3 / Lreal_arm_right), where L_arm_right_3 is the horizontal projection distance of the right arm, and Lreal_arm_right is the actual physical reference length of the corresponding part of the right arm.

[0105] At the same time, a timestamp is added to each frame of action data after recognition to facilitate subsequent instruction recognition processing and decision module processing.

[0106] The pressure recognition process of the pressure recognition module is as follows: the voltage values ​​of the multi-touch pressure sensors and acceleration sensors worn by the user on the arms and joints are converted by ADC to measure the speed and force data of the user when performing the corresponding action ( ) and sends the corresponding data to the instruction recognition and processing module.

[0107] The hand micro-motion recognition module uses a lightweight finger motion sensor to respond to changes in finger joint motion data in real time and calculate the corresponding finger joint angle rotation data. The gravity sensor and geomagnetic sensor in the lightweight finger motion sensor calculate the wrist rotation and pitch angle data based on the corresponding voltage data, and then transmit the finger and wrist data to the command recognition processing and decision module.

[0108] The facial expression recognition process of the facial expression action recognition module is as follows: through the intercepted video stream mov2, combined with the camera intrinsic parameter matrix and distortion parameters, the solvePnP algorithm is used to solve the rotation vector rot_vec and the translation vector trans_vec, and the corresponding head joint angle data is calculated and returned to the command recognition processing and decision module.

[0109] The speech recognition process of the speech recognition module is as follows: the user-side system recognizes the audio command keywords through the processed audio stream wav1, and performs corresponding processing on the recognized preset keyword instructions to guide the robot-side expression; at the same time, the user-side system also encapsulates the video stream of the entire paragraph and transmits it to the robot side according to whether the user chooses real-time voice transmission.

[0110] The instruction recognition and processing process of the instruction recognition, processing and decision-making module is as follows: during the instruction processing process, the data previously input by the limb movement recognition and angle calculation module, pressure sensor, hand micro-movement recognition module, and facial expression movement recognition module are integrated as a whole to sort out the stiffness coefficient, damping coefficient, torque coefficient, speed, and angle of the corresponding action joint; at the same time, the time sequence is aligned according to the torso movement and audio timestamp information, and the data information of the same time sequence is organized into a dictionary, encapsulated in json format, and transmitted to the robot end using the TCP / IP protocol.

[0111] In terms of biomimetic signals, the robot in this embodiment breaks the limitations of traditional robots in emotional interaction. High-precision sensors collect heartbeat and respiratory signals from both the user and the user terminal in real time, accurately reproducing these physiological signal characteristics. This simulates extremely realistic human interaction scenarios in various emotional interaction scenarios, providing users with a more natural and intimate interaction experience and offering multi-dimensional emotional expression modalities that differ from the single interaction mode of traditional robots.

[0112] In terms of mechanical bionics, the robot in this embodiment abandons the traditional remote interactive robot method of simply using a screen for 2D projection mapping. The mechanical bionic module utilizes flexible servos to drive joints, achieving more flexible mechanical movements that are closer to human joint motion. Furthermore, by integrating the specific physical structure of the face, it can precisely control the movement of facial joints to achieve a rich variety of expressions.

[0113] The working process of each module of the robot is as follows:

[0114] Data processing and calculation process: Uses TCP / IP communication protocol to communicate with the client. Responsible for receiving processed dictionary data frames from the client system. Each frame of data is split into the corresponding format and sent to the corresponding robot control module for execution.

[0115] Trunk and limb movement process: The corresponding part of the trunk and limb execution module receives the corresponding servo channel number and related coefficient parameters after segmentation from the central processing module and angles According to different stiffness coefficients, damping coefficients, torque coefficients, speeds, and angles, the corresponding PWM signals are sent to the channel servos of the corresponding control parts for output control.

[0116] Voice processing: After receiving instructions from the central processing module, the voice module issues corresponding prompt interaction words according to different instructions. Or it transmits real-time voice data according to user needs.

[0117] Bionic face simulation process: According to the corresponding facial expression recognition signal (jaw opening and closing angle, eye rotation angle, eyelid opening and closing degree) split out by the control system, the corresponding data signal is output to the corresponding module of the bionic face using RX / TX. The bionic face driver board receives the target angle and Output PWM signals to the corresponding facial servos to control facial movements.

[0118] Bionic signal processing: The breathing and heartbeat bionic module sends sensor data to the corresponding simulated parts of the robot based on the breathing and heartbeat signals sent by the control system. Simultaneously, the robot's pressure feedback sensors detect whether a person is hugging or approaching. When an approaching or emotionally engaging hug is detected, the robot's breathing and heartbeat sensors automatically activate and transmit the collected bionic signals in a two-way data transmission. If no approach is detected for more than three seconds, the corresponding transmission module automatically enters standby mode.

[0119] The interactive robot hardware and the user-side system communicate in real time via the TCP / IP protocol on the local area network. The robot hardware is responsible for action execution (joint actuation), while the user-side system is responsible for data collection and command generation. The collaborative process between the robot, the user wearable module, and the user-side system is as follows:

[0120] Step 1-1: The entire emotional interaction robot system is initialized and put into standby mode. After the robot is powered on, each module is initialized in turn: Figure 4 As shown, the joint motor is reset to a neutral posture, that is, the cervical spine is at 0° and the arm is naturally hanging down. Camera 2 starts self-test and establishes a connection with the user-end system through the TCP / IP protocol. At the same time, in the idle state, the ends of the two bionic arms open and close slightly in a cycle, and the voice module plays low-frequency breathing sounds at a specific frequency to simulate the human breathing rhythm.

[0121] Step 1-2: Once the user-side camera detects the human outline, the user-side system video and audio receiving module extracts the coordinates of the human hip joint and calculates the horizontal angle of the user relative to the robot. A cervical rotation command is sent: the cervical joint rotates to the corresponding angle so that the robot's line of sight faces the user. At the same time, the voice module plays a preset welcome sound effect to replace complex voice interaction. Real-time tracking of all joints is also enabled, and joint status feedback from the robot hardware is received through the local area network.

[0122] Step 1-3: During the interaction, the robot uses the command recognition processing and decision module to calculate the human body's movement posture and transmit it to the robot side. Figure 4 The included posture data information realizes the dual-mode interaction of "real-time projection" and "specific action triggering" through the preset action library. The specific steps include:

[0123] Step 2-1: Use the cv2.VideoCapture(0) function of the OpenCV library to capture the video stream from the camera and resize the video frame to 640×480 pixels. Use MediaPipe Face Mesh technology to detect 468 facial key points and 33 body key points, which cover major joints such as shoulders, elbows, and wrists.

[0124] Step 2-2: Posture calculation;

[0125] First, the head pose is accurately calculated. By applying the solvePnP algorithm, key points in the two-dimensional image are mapped to three-dimensional space. During this process, the rotation vector rot_vec and translation vector trans_vec are solved by combining the camera intrinsic parameter matrix and distortion parameters. Then, using the Rodrigues transform, the rotation vector is converted into a rotation matrix, which is further decomposed into Euler angles, including pitch, yaw, and roll, to fully describe the rotation state of the head.

[0126] Secondly, the calculation of limb posture; by using trigonometric geometric relationships, the rotation angles of the shoulder, elbow and wrist are calculated in detail;

[0127] Finally, to improve data stability and real-time performance, the calculated angle data is processed using Kalman filtering technology, which effectively removes noise and improves data reliability. At the same time, the exponential moving average method (EMA) is used to smooth the data, ensuring more stable and accurate angle information in real-time applications.

[0128] Step 2-3: In terms of data format and communication protocol, the user first accurately maps the required control angle information and converts it into a specific command format suitable for servo control. Then, using Socket communication technology and the reliable TCP / IP protocol, the converted servo control commands are stably and efficiently sent to the Raspberry Pi, ensuring the accuracy and real-time nature of command transmission.

[0129] Steps 1-4: After receiving instructions from the command recognition and decision-making module, the central processing module on the robot drives the motor to synchronize the user's actions in real time. When executing the "wave" or "hug" instructions, the joints coordinate to complete the corresponding actions according to the preset trajectory or trigger conditions.

[0130] Steps 1-5: When the user leaves the camera's field of view for 30 consecutive seconds, the robot triggers the exit process: the command recognition processing and decision-making module stops posture data processing and sends a "reset" command. The robot performs the head return and arms drooping movements, and the display screen resumes the soft breathing lighting effect; at the same time, all components on the robot side enter a low-power standby state, maintain a LAN connection, and maintain basic functions in low-power mode waiting to be awakened. When the user leaves the camera's field of view, that is, no human body is detected for 30 consecutive seconds, the robot enters the exit process.

[0131] Step 1-6: Repeat steps 1-2 to 1-5 until the interaction ends.

[0132] To reflect the robot's sense of life, the robot's camera will also identify the distance between the remote interlocutor and the robot. When the distance is less than 0.5m, the breathing and heartbeat bionic module will be activated, and the volume will be increased accordingly as the distance gets closer.

[0133] To fully describe the characteristics of user movements and embody the robot's sense of life, the robot performs bionic motion simulation based on a series of pressure data sent by the pressure sensor. At the same time, it dynamically adjusts the stiffness and flexibility coefficients of the motor.

[0134] Regarding the actual robot operation mode, considering that there may be scenarios in which the user is not present or the remote interactor wants to use the robot without real-time user action input, this embodiment uses a dual interaction strategy of real-time and preset. The preset library contains a series of multiple action mappings, expression expressions and voice interaction strategies for custom scenarios, such as:

[0135] Preset actions include basic social actions such as waving, nodding, and hugging, as well as specific action combinations designed for different emotions;

[0136] The preset voices include greetings, responses, emotional interjections, etc., which can be selected and played according to the needs of the scene and emotional expression.

[0137] Examples of practical scenarios of the remote emotional interaction robot system described in this embodiment are as follows.

[0138] The first one is remote elderly care companionship.

[0139] In an aging society, many elderly people may feel lonely because their children are busy with work or live far away. Through the remote emotional interaction robot system of this embodiment, children can interact with the elderly in a more intimate and authentic way through the robot system anywhere, alleviating their loneliness and strengthening their emotional connection.

[0140] Children interact with the robot system via real-time video via a remote camera. The system synchronizes the operator's movements to the robot, transmitting intimate gestures like hugs and caresses to the elderly remotely. By synchronizing movements and expressions, the elderly not only see their children's faces but also feel their "body language," such as gestures and expressions, greatly enhancing the sense of authenticity and emotional connection with their companionship.

[0141] The second type is remote parent-child interaction.

[0142] For busy parents, the emotionally interactive robot system can become an emotional bridge between parents and children, especially when parents are unable to be with their children due to work reasons. The robot system can make children feel the closeness and care of their parents.

[0143] Parents can remotely control the robot through a camera to interact with their children. The robot synchronizes the parent's movements (such as nodding, gesturing, and clapping) in real time, helping parents and children connect emotionally. Especially when parents are unable to accompany their children for play or study, the robot system can compensate for their absence by "replacing" their movements.

[0144] The third type is remote interaction between long-distance lovers.

[0145] In modern society, many couples are separated by distance due to work or school, which limits emotional communication and can easily lead to emotional alienation. The emotional interaction robot system, through remote posture projection technology, can help long-distance lovers maintain intimacy and strengthen emotional connection.

[0146] Long-distance couples interact through an emotionally interactive robot system. The robot synchronizes their body language, movements, and expressions in real time, allowing them to share each other's emotional expressions, such as hugs and holding hands. The robot system not only supports traditional voice and video calls, but also enhances the emotional authenticity and intimacy through body language, enhancing the interactive experience between couples.

[0147] The fourth type is psychological healing and emotional expression.

[0148] For people with psychological problems or emotional stress (such as depression patients, anxiety patients, etc.), emotional interactive robots can provide support through body language and emotional synchronization, engage in emotional resonance and express themselves, and play a role in psychological healing.

[0149] Users engage in remote conversations with therapists through the robot system. The robot synchronizes the user's body language and movements, helping the therapist better understand the user's emotional state. The robot system can convey the user's emotional changes through non-verbal cues such as movement and expression, making psychological therapy more intuitive and profound.

[0150] Fifth, emotional support and interaction in distance education.

[0151] In distance education, interaction between students and teachers is often limited by screens. Furthermore, limited audio interaction can make classes inefficient, hindering teachers from receiving timely student feedback. Emotionally interactive robots can help teachers engage with students more vividly and naturally through body language, interacting with students in real time and more easily understanding classroom situations. This can increase student motivation.

[0152] During distance learning, teachers can interact with students through emotionally interactive robots. The robots can synchronize with the teacher's movements, expressions, and body language in real time, helping teachers demonstrate teaching content and providing emotional support to students. This improves online teaching efficiency and reduces the distraction experienced in traditional online classes. Furthermore, when students encounter learning difficulties, the robots can enhance their confidence by simulating the teacher's encouraging gestures or comforting body language (such as pats and hugs).

[0153] Sixth, family gatherings and remote interactions.

[0154] In modern families, family members may be scattered across cities or countries due to geography or work, making it difficult to gather together during holidays. Emotionally interactive robots can use remote synchronization technology to allow family members to interact with each other through avatar-like movements, compensating for the emotional gap caused by physical distance.

[0155] During holidays or family gatherings, family members who are unable to attend in person can interact with their families through emotionally interactive robots, synchronizing their movements and expressions in real time to perform intimate gestures such as hugs and handshakes. In this way, remote interactions between family members are no longer limited to screens and voice, but can express emotions through more natural body language.

[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the present invention.

Claims

1. A remote emotional interaction robot system based on real-time posture projection, characterized by: The remote emotional interaction robot system includes a user-side system, a user-wearable module and a robot side; The user-side system is responsible for processing, identifying, and calculating posture and facial data, and includes a personal computer-based video and audio receiving module, a body movement recognition and angle calculation module, a pressure recognition module, a hand micro-movement recognition module, a facial expression movement recognition module, a voice recognition module, and a command recognition processing and decision-making module. The video and audio receiving module is used to capture the user's real-time full-body movement video stream mov1 and the audio stream wav1 that supports voice commands and real-time transmission through the camera and microphone on the personal computer, and store them in a pre-processing buffer area. The limb movement recognition and angle calculation module is used to process the video stream collected by the video receiving module, mark the response joint point information, and identify the limb movement to calculate the key joint angles, including the head posture angle, eye and mouth opening and closing rate, and arm posture angle; the pressure recognition module is used to transmit the data of the pressure sensor in the user wearable module in real time; the hand micro-movement recognition module transmits the hand micro-movement in real time through the lightweight finger motion sensor worn by the user; the facial expression movement recognition module performs high-precision facial algorithm recognition through the video stream mov2 intercepted by the second camera on the robot side; the voice recognition module is responsible for real-time recognition of the user's conversation voice based on the audio stream return data, and screens the content of the user's voice command response, and responds to the corresponding prompt content to guide the robot side to express; the command recognition and processing module is responsible for integrating and judging the data of multiple module sensors, identifying according to the limb position status, pressure, and acceleration, and generating command data for the corresponding operation control data, encapsulating the relevant commands into a fixed format command package, and sending it to the robot side; The user-worn module is used to sense various subtle hand movements and pressure changes between the body and external objects. It includes a pressure sensing module and a lightweight finger motion sensor. The pressure sensing module is worn on the user's arm and joints to measure movement force and contact pressure. The lightweight finger motion sensor is used to monitor micro-movements of the hand. The user-worn module is connected to the user-end system via TCP / IP or I2C. The robot is responsible for playing audio and executing actions, and includes a Raspberry Pi-based central processing module and several execution modules. The execution modules include two bionic arms, two bionic hands, a bionic head consisting of ten servos, an information acquisition module consisting of two cameras and two microphones, a voice module, a breathing and heartbeat bionic module, and a pressure feedback sensor. The central processing module parses instructions and distributes them to the execution modules. Each execution module is connected to the central processing module via PWM or I2C to implement instruction execution and interaction. The robot and user systems communicate via TCP / IP. Each of the two bionic arms contains three servos, which respectively control the forward and backward movement of the shoulder, the left and right rotation, and the bending and straightening of the elbow; The bionic head contains eyes, mouth, and neck modules, and includes 26 motors that are position-controlled. The flexible facial skin is easily connected to the hardware mechanism via magnets. The neck module supports three-axis rotation: roll, pitch, and yaw. Three motors control the movement of the neck on these three axes. Twelve motors control the upper face, including the eyeballs, eyelids, and eyebrows. Eleven motors control the mouth mechanism and mandible. The magnetic connection design allows the robot's facial skin to be easily replaced. The eye module has a magnetically connected linkage that controls the eyebrows, including the upper eyelid, lower eyelid, eyeball linkage, and eyeball frame. The mouth module uses six passive mouth linkages. The second camera is used to detect user distance and environmental perception; the second microphone cooperates with the voice module to support voice interaction and emotional voice feedback; The breathing and heartbeat bionic module is used to simulate the chest rise and fall, vibration, and heartbeat; the pressure feedback sensor is used to detect the hugging strength and adjust the softness of the movement.

2. A remote emotional interaction robot system based on real-time posture projection according to claim 1, characterized in that: The user-side system calculates posture and facial data using a real-time posture projection method, using a combination of multi-sensor fusion, precise calculation, intelligent filtering, preset actions, and multi-threading technology, and uses MediaPipe's Face Mesh and Pose models for multi-model fusion. The calculation steps are as follows: S1: Initialize model parameters and wait for the user to appear and the interaction to begin; S2: When the user first appears, the user-side system defines the initial arm length and facial key line lengths at the user's spatial location, uses this information as a reference for subsequent angle calculations, and simultaneously announces a welcome message through the voice module to initiate interaction. S3: During the interaction process, the camera 1 and microphone 1 of the video and audio receiving module collect the user's movements, expressions, and voice information, and transmit them to the body movement recognition and angle calculation module in the user-side system for real-time calculation of human-robot dynamic mapping; Adaptive adjustments are made to user movements based on the robot's joint range of motion, degrees of freedom of movement, and flexible limb characteristics to avoid dangerous movements that exceed mechanical limits. S4: The user-side system transmits the calculated head posture angle, eye and mouth opening and closing rate, and arm posture angle to the robot-side central processing module based on the Raspberry Pi. The central processing module parses the instructions and distributes them to the execution modules to control the robot's posture movements and play real-time voice. S5: Repeat steps S1 to S4 until the interaction ends.

3. The remote emotional interaction robot system based on real-time posture projection according to claim 2, characterized in that: When performing real-time calculation of the human-robot dynamic mapping in the S3 step, the limb motion recognition and angle calculation module in the user-side system constructs a camera projection model based on the perspective projection principle, establishes a mapping relationship between 3D head model points and 2D image points, combines the camera matrix and distortion parameters, iteratively optimizes and solves the head rotation and displacement vectors, converts the rotation matrix into Euler angle representation, and limits the angle to a physically feasible range through a custom range judgment function.

4. The remote emotional interaction robot system based on real-time posture projection according to claim 3, characterized in that: When performing real-time calculation of human-robot dynamic mapping in the S3 step, in terms of limb projection, the limb motion recognition and angle calculation module in the user-side system uses the key points of the upper body shoulders, elbows, wrists, neck, limb contours, and facial contours identified by MediaPipe's Pose model to achieve accurate analysis and projection of arm posture; the limb motion recognition and angle calculation module extracts key point information of the right shoulder, right elbow, and right wrist of the right body and the left shoulder, left elbow, and left wrist of the left body from the detection results of the Pose model, and calculates the distance between the right shoulder and elbow, and between the elbow and wrist in the first frame image as the arm length reference value, and also calculates the corresponding arm length on the left side. These reference values ​​provide an important reference basis for subsequent angle calculations; for the angle calculation of the arm posture, the limb motion recognition and angle calculation module uses geometric and trigonometric functions to calculate the angles required for multiple servo control; at the same time, in terms of calculation, the limb motion recognition and angle calculation module uses geometric decomposition to decompose 3D motion into multiple 2D planar motions; after the decomposition is completed, the limb motion recognition and angle calculation module uses trigonometric functions to calculate the angles of each joint; In the servo control phase, the robot-side system's command recognition, processing, and decision-making module compares the calculated control angle with a preset physical safety range. If the angle exceeds the range, it is adjusted to within the safety range. The command recognition, processing, and decision-making module introduces an exponentially weighted moving average method to smooth the calculated angle by setting a smoothing factor. The smoothed angle is then calculated and output based on historical angle values ​​and current measurements. In estimating the head pitch angle, the Euclidean distance between the midpoint of the left and right inner canthus and the tip of the nose is calculated, combined with the inverse sine function and different scaling factors to achieve accurate estimation. For the roll angle, a dual verification mechanism is adopted, which is the calculation of the horizontal and vertical difference between the two eyes and the iterative optimization result based on the perspective projection model. The specific implementation is as follows: first, the pixel coordinates of the left and right inner canthus and the tip of the nose are located on the image plane, the coordinates of the midpoint of the left and right inner canthus are calculated, and then the pixel distance between the midpoint coordinate and the tip of the nose is solved. The reference distance and physical size information obtained in the calibration stage are combined with the trigonometric function relationship to estimate the head pitch angle. Different scaling factors are applied to the movement characteristics of looking up and looking down to improve the estimation accuracy. For the roll angle, on the one hand, the ratio of the vertical pixel difference to the horizontal pixel difference of the left and right eye centers is calculated, and the inverse tangent function is used to solve the preliminary roll angle. On the other hand, based on the correspondence between the 3D head model points and the 2D image points, a nonlinear optimization problem is constructed. The rotation matrix is ​​iteratively solved by minimizing the reprojection error, and the roll angle component is extracted from it. Finally, the roll angles obtained by the two methods are verified for consistency and fused. Six key points out of 468 key points on the face are extracted through MediaPipe, which are located at the edge of the eyes, the tip of the nose, the corner of the mouth and the chin. 2D and 3D head pose arrays are constructed. The camera matrix K is calculated based on the Euler rotation theorem according to the following matrix formula (1), where the simulated camera focal length is taken as the image width, the image center is set as the center of the coordinate system, and the tilt parameter of the camera matrix is ​​set to 0; (1); Among them, P c is a point in the camera coordinate system, expressed as a two-dimensional coordinate (X C , Y C );P w is a point in the world coordinate system, expressed as a three-dimensional coordinate ; K is the camera's internal parameter matrix, including the focal length f x 、f y and light center c x 、c y ; R is the rotation matrix, which represents the rotation from the world coordinate system to the camera coordinate system; T is the translation matrix, which represents the translation from the world coordinate system to the camera coordinate system; f x 、f y is the focal length of the camera, corresponding to the x and y directions respectively; c x 、c y is the principal point coordinate of the camera, i.e. the optical center; r ij and t x ,t y ,t z are the elements of the rotation matrix and the elements of the translation vector, respectively, representing the transformation from the world coordinate system to the camera coordinate system; Given the 3D and 2D coordinates and the camera matrix, the perspective n-point PnP problem is used to solve the above equation (1) to obtain the rotation matrix R. Then, the direction angle of the head posture is obtained by performing RQ decomposition on the rotation matrix R, which is used to control the robot's head angle to imitate human head movement.

5. The remote emotional interaction robot system based on real-time posture projection according to claim 1, characterized in that: The facial expression and action recognition module of the user-side system has the ability to calibrate the reference distance and angle in real time, and can dynamically adjust the scaling factors used for head posture angle estimation. These scaling factors are the head-up angle scaling factor UP_SCALE_FACTOR, the head-down angle scaling factor DOWN_SCALE_FACTOR, the head-right turn scaling factor RIGHT_SCALE_FACTOR, and the head-left turn scaling factor LEFT_SCALE_FACTOR. After the facial expression and action recognition module completes the calibration of the reference distance and angle, the solvePnP algorithm and RQDecomp3x3 are used to calculate the head posture angle. The Euler angles of the head are decomposed and converted into degrees. Finally, the corresponding scaling factor is applied according to the direction of head movement for linear adjustment to obtain the final head pitch and roll angle control data. The facial expression and action recognition module will dynamically apply the corresponding scaling factor to adjust the estimated angle according to different head movements. By default, the values ​​of UP_SCALE_FACTOR, RIGHT_SCALE_FACTOR, and LEFT_SCALE_FACTOR are 1.5, and the value of DOWN_SCALE_FACTOR is 0.

5. The default values ​​are dynamically adjusted according to actual usage scenarios and needs.

6. The remote emotional interaction robot system based on real-time posture projection according to claim 1, characterized in that: The receiving and identifying processes of each module of the user end system are as follows: The video and audio stream receiving process of the video and audio receiving module is as follows: the user-side system collects video and audio by calling the camera 1 and microphone 1 in the personal computer. After collection, the audio wav1 and videos mov1 and mov2 are first processed into tracks, and the timestamps are saved during processing to ensure the consistency of the video and audio timing; the processed video and audio streams are subjected to noise reduction processing, and the processed video and audio streams are placed in the buffer area and wait for the system program to process them; The limb movement recognition process of the limb movement recognition and angle calculation module is as follows: the video stream mov1 returned by the video and audio receiving module is passed into the limb movement recognition and angle calculation module, and the CV2 algorithm based on the OpenCV library is used to identify 33 key points of the torso, and the real-time torso projection distance is measured according to the real-time measured key point spacing. In each subsequent frame, the projection distance Ldet of the key point in the X-axis and Y-axis directions is calculated: for the left arm, the vertical projection distance L_arm_left_2 is the absolute value of the coordinate difference of the left arm key point in the Y-axis direction, and the calculation formula is L_arm_left_2=|y2-y1|; where y1 and y2 are the Y-axis coordinates of the two adjacent key points of the left arm respectively; the horizontal projection distance L_arm_left_3 is the absolute value of the coordinate difference of the left arm key point in the X-axis direction, and the calculation formula is L_arm_left_3=|x2-x1|; where x1 and x2 are the X-axis coordinates of the two adjacent key points of the left arm respectively; The projection ratio is calculated using the trigonometric relationship ratio = Ldet / Lreal, where Lreal represents the actual physical reference length of the corresponding part of the user's limb, which is obtained through initialization calibration during the user's first interaction. The control angle of different servos is calculated using the following specific formula: Shoulder left and right rotation servo angle θ1: θ1 = arcsin (L_arm_left_3 / Lreal_arm_left), where Lreal_arm_left is the actual physical reference length of the corresponding part of the left arm; Elbow flexion and extension servo angle θ2: θ2=arccos (L_arm_left_2 / Lreal_arm_left); Shoulder forward and backward motion servo angle θ3: θ3 = arcsin (L_arm_right_3 / Lreal_arm_right), where L_arm_right_3 is the horizontal projection distance of the right arm, and Lreal_arm_right is the actual physical reference length of the corresponding part of the right arm; At the same time, a timestamp is added to each frame of action data after recognition to facilitate subsequent instruction recognition processing and decision module processing; The pressure recognition process of the pressure recognition module is as follows: the voltage values ​​of the multi-touch pressure sensors and acceleration sensors worn by the user on the arms and joints are converted by ADC to measure the speed and force data of the user when performing the corresponding action ( ), and send the corresponding data to the instruction recognition and processing module; The hand micro-motion recognition module's hand micro-motion recognition process is as follows: A lightweight finger motion sensor responds to changes in finger joint motion data in real time and calculates the corresponding finger joint angle rotation data; the gravity sensor and geomagnetic sensor in the lightweight finger motion sensor are used to calculate the wrist rotation angle and pitch angle data based on the corresponding voltage data, and the finger and wrist data are transmitted to the command recognition processing and decision module; The facial expression recognition process of the facial expression action recognition module is as follows: through the intercepted video stream mov2, combined with the camera intrinsic parameter matrix and distortion parameters, the solvePnP algorithm is used to solve the rotation vector rot_vec and the translation vector trans_vec, and the corresponding head joint angle data is calculated and sent back to the command recognition processing and decision module; The speech recognition module's speech recognition process is as follows: the user-side system uses the processed audio stream wav1 to identify audio command keywords, and then performs corresponding processing on the recognized preset keyword commands to guide the robot's expression. At the same time, the user-side system also encapsulates the entire video stream based on whether the user selects real-time speech transmission and transmits it to the robot. The instruction recognition and processing process of the instruction recognition, processing and decision-making module is as follows: during the instruction processing process, the data previously input by the limb movement recognition and angle calculation module, pressure sensor, hand micro-movement recognition module, and facial expression movement recognition module are integrated as a whole to sort out the stiffness coefficient, damping coefficient, torque coefficient, speed, and angle of the corresponding action joint; at the same time, the time sequence is aligned according to the torso movement and audio timestamp information, and the data information of the same time sequence is organized into a dictionary, encapsulated in json format, and transmitted to the robot end using the TCP / IP protocol.

7. The remote emotional interaction robot system based on real-time posture projection according to claim 1, characterized in that: The collaboration process between the robot side, user wearable module and user side system is as follows: Step 1-1: The entire emotional interaction robot system is initialized and placed in an idle state. After the robot is powered on, each module is initialized in sequence: the joint motors are reset to a neutral position, with the cervical spine at 0° and the arms hanging naturally. Camera 2 initiates a self-test and establishes a connection with the user system via the TCP / IP protocol. Simultaneously, in the idle state, the two bionic arms open and close slightly in a cyclical manner, and the voice module plays low-frequency breathing sounds at a specific frequency to simulate human breathing rhythms. Step 1-2: Once the robot's camera detects the human silhouette, the user's system video and audio receiving module extracts the coordinates of the human hip joint and calculates the user's horizontal angle relative to the robot. A cervical rotation command is sent, causing the cervical joint to rotate to the corresponding angle, so that the robot's line of sight faces the user. Simultaneously, the voice module plays a preset welcome sound effect, replacing complex voice interaction. Real-time tracking of all joints is also enabled, and joint status feedback from the robot hardware is received via the local area network. Steps 1-3: During the interaction process, the robot uses the command recognition processing and decision-making module to calculate the human body's motion posture and transmit visual posture data to the robot side. The preset action library is used to achieve dual-mode interaction of "real-time projection" and "specific action triggering"; Steps 1-4: After receiving instructions from the command recognition and decision-making modules, the robot's central processing module drives the motors to synchronize with the user's movements in real time. When executing "wave" or "hug" commands, the joints coordinate to complete the corresponding actions according to the preset trajectory or trigger conditions. Steps 1-5: If the user leaves the camera's field of view for 30 consecutive seconds, the robot triggers the exit process: the command recognition and decision-making module stops processing posture data and sends a "reset" command. The robot returns its head to its original position, lowers its arms, and the display returns to its soft breathing lighting effect. At the same time, all components on the robot enter a low-power standby state, maintain a LAN connection, and maintain basic functions in low-power mode while waiting to be woken up. If the user leaves the camera's field of view, that is, if no human body is detected for 30 consecutive seconds, the robot enters the exit process. Step 1-6: Repeat steps 1-2 to 1-5 until the interaction ends.

8. The remote emotional interaction robot system based on real-time posture projection according to claim 7, characterized in that: The steps 1-3 specifically include the following steps: Step 2-1: Use the cv2.VideoCapture(0) function of the OpenCV library to capture the video stream from the camera and resize the video frame to 640×480 pixels; use MediaPipe Face Mesh technology to detect 468 facial key points and 33 body key points; Step 2-2: Posture calculation; First, the head pose is accurately calculated. By applying the solvePnP algorithm, key points in the two-dimensional image are mapped to three-dimensional space. During this process, the rotation vector rot_vec and translation vector trans_vec are solved by combining the camera intrinsic parameter matrix and distortion parameters. Then, using the Rodrigues transform, the rotation vector is converted into a rotation matrix, which is further decomposed into Euler angles, including pitch, yaw, and roll, to fully describe the rotation state of the head. Secondly, the calculation of limb posture; by using trigonometric geometric relationships, the rotation angles of the shoulder, elbow and wrist are calculated in detail; Finally, the calculated angle data is processed using Kalman filtering technology. At the same time, the data is smoothed using the exponential moving average method (EMA) to ensure that more stable and accurate angle information can be obtained in real-time applications. Step 2-3: In terms of data format and communication protocol, the user first accurately maps the required control angle information and converts it into a specific command format suitable for servo control. Then, using Socket communication technology and the reliable TCP / IP protocol, the converted servo control commands are stably and efficiently sent to the Raspberry Pi, ensuring the accuracy and real-time nature of command transmission.