A Multimodal Fusion and Control Platform and System for Human-Humanoid Robot Interaction
Through multimodal fusion and control platform, the stability and accuracy of action recognition and prediction in the interaction between humans and humanoid robots is solved, efficient human-computer interaction in complex environments is achieved, and the robot response speed and accuracy is improved.
Patent Information
- Application Number
- CN202510164218.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-17
- Filing Date
- 2025-02-14
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-02-14
AI Technical Summary
In the prior art, the movement recognition and prediction of human-like robot interactions are insufficient in complex environments, the multi-sensor layout and data acquisition are complex, the accuracy of the algorithm needs to be improved in practical applications, and it is difficult to achieve efficient multimodal perception and collaborative operation.
The multi-modal fusion and control platform is adopted, including the human signal fusion perception platform and the two-arm robot mirror control platform. Through multi-sensor arrangement, stable multi-sensor coordinate calibration, adaptive filtering technology and generalized embodied intelligent human-computer interaction algorithm model, the interaction quality of action recognition and prediction is improved.
It significantly improves the stability and reliability of data acquisition, reduces the complexity of platform logic, enhances the robot's action recognition and prediction capabilities in complex environments, and achieves more natural and accurate human-computer interaction.
Smart Images

Figure CN119839888B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot interaction, and particularly to a multi-modal fusion and control platform and system for human-anthropomorphic robot interaction. Background Art
[0002] In the existing human-computer interaction technology, the interaction between humans and anthropomorphic robots faces many technical challenges. In order to achieve accurate, comprehensive, and continuous action recognition and prediction in complex human-robot interaction tasks, it is necessary to arrange a multi-sensor system for more multi-modal and comprehensive perception. Multi-sensors can capture information from different dimensions in different forms, such as video image signals, inertial signals, electromyogram signals, etc. These information complement each other and can provide more comprehensive interaction data.
[0003] However, it is extremely difficult to collect stable and continuous data on human interaction actions. The limitations of the sensors themselves and external environmental interference will lead to insufficient stability and reliability of data collection. Moreover, each of the sensors related to pose in the multi-sensor has its own independent reference coordinate system, and spatial coordinate calibration and unification are required. This process is complex and time-consuming because the position and orientation of each sensor need to be accurately calibrated to ensure the accuracy of the data. This increases the logical complexity of the platform, especially in the case of multiple sensors.
[0004] The multi-sensor layout on the human body and the robot needs to be carefully screened and designed. A reasonable layout can not only improve the accuracy of data collection but also effectively reduce interference and redundant information between sensors. However, in practical applications, how to arrange these sensors to obtain a more accurate action prediction function is still a difficult problem.
[0005] The current action recognition and prediction algorithms still need to be improved in terms of accuracy. Although there are already some advanced algorithms that can achieve a high recognition rate under ideal conditions, in practical applications, action recognition and prediction in complex interaction scenarios still face huge challenges. Complex interaction scenarios not only include the recognition and prediction of multiple actions but also involve factors such as environmental changes and the diversity of interaction objects, which pose higher requirements for the adaptability and generalization ability of the algorithms.
[0006] Therefore, the present invention aims to provide a human-anthropomorphic robot interaction platform to solve the above problems. Summary of the Invention
[0007] The purpose of the present invention is to solve the above problems and to provide a multimodal fusion and control platform and system for the interaction between humans and humanoid robots, including a human signal fusion perception platform and a dual-arm robot mirror control platform. Through multi-sensor arrangement and fusion, a stable multi-sensor coordinate calibration system, adaptive filtering technology and a generalized embodied intelligent human-computer interaction algorithm model, the interaction quality between humans and humanoid robots in motion recognition and prediction is improved.
[0008] In order to achieve the above object, the technical solution of the present invention is as follows:
[0009] In a first aspect, the present invention provides a multimodal fusion and control platform for human-humanoid robot interaction, the multimodal fusion and control platform comprising a human signal fusion perception platform and a dual-arm robot mirror control platform profile frame;
[0010] The human body signal fusion sensing platform includes a human body signal fusion sensing platform profile frame, a surface electromyography multi-lead sensing bracelet, and motion sensing wearable gloves;
[0011] The human signal fusion perception platform profile frame and the dual-arm robot mirror control platform profile frame are both rectangular frames. A data center is provided at the bottom of the human signal fusion perception platform profile frame. The data center includes a display and several depth camera industrial computers.
[0012] A dual-arm robot mirror control platform controller is provided on the top of the dual-arm robot mirror control platform profile frame, and humanoid mechanical arms are provided on opposite sides of the dual-arm robot mirror control platform controller. A humanoid mechanical arm is provided at one end of the humanoid mechanical arm that is not connected to the dual-arm robot mirror control platform controller. The humanoid mechanical arm includes three shoulder joints, one elbow joint and three wrist joints. Each joint of the humanoid mechanical arm adopts a high-precision servo motor and a multi-stage reducer. The humanoid mechanical arm adopts a series configuration and operates based on the form of multi-joint connecting rod follow-up.
[0013] The humanoid robotic dexterous hand comprises six degrees of freedom, wherein the thumb structure comprises two degrees of freedom, including bending degrees of freedom and joint degrees of freedom, which are used for precise grasping and manipulation of complex objects, and the remaining four fingers each comprise one degree of freedom of bending. The humanoid robotic dexterous hand is used for performing finger bending and grasping actions based on a multi-joint linkage follow-up form by integrating a high-precision micro-servo motor and a linkage mechanism, and the humanoid robotic dexterous hand has a built-in force tactile sensor.
[0014] In some embodiments, the human signal fusion perception platform profile frame and the top of the dual-arm robot mirror control platform controller are both equipped with multiple depth cameras.
[0015] In some embodiments, the motion-sensing wearable glove includes an inertial sensor and a bending sensor for capturing hand movements;
[0016] There are eleven inertial sensors, which are respectively distributed at the palm position, the metacarpophalangeal joint of each finger, and the proximal interphalangeal joint of each finger;
[0017] There are fourteen bending sensors, which are respectively distributed at the metacarpophalangeal joint of each finger, the proximal interphalangeal joint of each finger, and the distal interphalangeal joint of the four fingers other than the thumb. The bending sensors are used to sense the angle of the finger bending towards the palm and capture the second bending movement of the finger.
[0018] In some embodiments, the depth camera is connected to the profile frame of the human signal fusion perception platform and the controller of the dual-arm robot mirror control platform through a sliding device.
[0019] In some embodiments, universal wheels are provided at the bottoms of the profile frame of the human signal fusion perception platform and the profile frame of the dual-arm robot mirror control platform.
[0020] In a second aspect, the present invention provides a multi-modal fusion and control system for human-anthropomorphic robot interaction. The multi-modal fusion and control system includes a multi-sensor fusion module, a spatial coordinate calibration fusion module, an adaptive filtering module, and a human-computer interaction module;
[0021] The multi-sensor fusion module is used to collect and fuse data from multiple sensors to obtain multi-sensor fusion data, and send the multi-sensor fusion data to the spatial coordinate calibration fusion module. Among them, the data of the multiple sensors includes surface electromyography data collected by a surface electromyography multi-lead sensing bracelet, triaxial acceleration data, triaxial attitude data collected by the inertial sensor of the motion-sensing wearable glove, and bending degree data collected by the bending degree sensor. The multi-sensor fusion data includes fused hand muscle state data and hand movement attitude data;
[0022] The spatial coordinate calibration fusion module is used to correct and synchronize human behavior data in different coordinate systems collected by the depth camera of the profile frame of the human signal fusion perception platform, and perform coordinate correction on the depth visual information collected by the depth camera of the controller of the dual-arm robot mirror control platform to obtain multi-sensor fusion calibration data, and the multi-sensor fusion calibration data has the same spatial coordinate system;
[0023] The adaptive filtering module is used to perform adaptive filtering on the human behavior data after correction and synchronization processing in the multi-sensor fusion calibration data and the pressure data collected by the built-in force-tactile sensor of the anthropomorphic robot dexterous hand (9) to obtain filtered human behavior data and pressure data;
[0024] The human-computer interaction module is used to determine a robot response strategy based on the multi-sensor fusion data, the filtered human behavior data, and the pressure data, and control the humanoid robotic arm and the humanoid robotic dexterous hand to perform interactive movements with the user based on the robot response strategy, where the robot response strategy includes the positions and postures of the humanoid robotic arm and the humanoid robotic dexterous hand.
[0025] In some embodiments, in the multi-sensor fusion module,
[0026] Determine the hand movement posture of the user according to the following posture estimation formula:
[0027]
[0028] where and are the attitude data in quaternion form at the (k - 1)-th moment and the k-th moment respectively, represents the quaternion multiplication symbol, is the time interval between the (k - 1)-th moment and the k-th moment, is the angular velocity data at the k-th moment;
[0029] Transform the acceleration data and the attitude data according to the following formula to obtain the linear acceleration in the global coordinate system:
[0030]
[0031] where is the linear acceleration in the global coordinate system, is the rotation matrix generated by is the acceleration data;
[0032] Determine the position trajectory of the hand movement according to the following formula:
[0033]
[0034]
[0035] where is the velocity of the hand movement at the k-th moment, is the position of the hand movement at the k-th moment, is the velocity of the hand movement at the (k - 1)-th moment, is the position of the hand movement at the (k - 1)-th moment;
[0036] Determine the multi-sensor fusion data according to the following formula:
[0037]
[0038]
[0039] Wherein, is the hand motion posture data in the multi-sensor fusion data at the k-th moment, is the fusion transformation function, is the bending degree data of the i-th finger joint at the k-th moment, is the resistance change of the i-th bending sensor at the k-th moment, is the change mapping function;
[0040] The surface electromyography data collected by the surface electromyography multi-lead sensing bracelet (3) is expressed as , includes the left hand surface electromyography data and the right hand surface electromyography data, where / represents the data of the i-th lead in the left / right hand surface electromyography data. According to the following formula, the hand muscle state data in the multi-sensor fusion data is determined:
[0041]
[0042] Wherein, and respectively represent the hand muscle state data of the left hand and the right hand, is the mapping function from the surface electromyography data to the hand muscle state data.
[0043] In some embodiments, in the spatial coordinate calibration fusion module,
[0044] According to the following formula, the human behavior data after correction and synchronization processing and the depth visual information after coordinate correction in the multi-sensor fusion calibration data are determined:
[0045]
[0046]
[0047] Wherein, W is the defined world coordinate system, is the coordinate system of the depth camera itself on the i-th profile frame of the human signal fusion perception platform, represents the transformation matrix from the coordinate system to the W coordinate system, represents the human behavior data in the W coordinate system, represents the human behavior data in the coordinate system, represents the human behavior data after correction and synchronization processing in the multi-sensor fusion calibration data in the W coordinate system, is a spatial coordinate fusion function, and n represents the label of the sensor.
[0048] In some embodiments, in the adaptive filtering module,
[0049]
[0050] )
[0051] In the formula, is the filtered human behavior data, is the hand movement posture data in the multi-sensor fusion data at the k-th moment, represents the human behavior data after calibration and synchronization processing in the multi-sensor fusion calibration data in the W coordinate system, is the fusion function, is the mapping function for mapping the pressure data matrix before filtering to the pressure data after filtering, is the pressure data matrix before filtering at the i-th finger of the humanoid robot dexterous hand , represents the pressure data of the i-th finger of the humanoid robot dexterous hand obtained by filtering.
[0052] In some embodiments, in the human-computer interaction module, the robot response strategy is determined through the embodied intelligent core algorithm of the generalization large model; the angles of each joint of the humanoid robotic arm and the humanoid robot dexterous hand are determined through the robot inverse kinematics solution method.
[0053] Compared with the prior art, the beneficial effects of this solution are:
[0054] Through multi-sensor fusion and arrangement, the present invention significantly improves the stability and reliability of data acquisition. Stable multi-sensor coordinate calibration reduces the logical complexity of the platform and improves the accuracy of sensor data. The adaptive filtering technology is adopted to effectively filter noise and provide more accurate action data. The generalization embodied intelligent human-computer interaction algorithm model enhances the action recognition, prediction and anticipation capabilities of the robot in interaction, enabling it to interact with humans more naturally and accurately. Brief Description of the Drawings
[0055] Figure 1 is a schematic structural diagram of the multi-modal fusion and control platform for human-robot interaction in the embodiment of the present invention;
[0056] Figure 2 is the overall data flow and analysis flowchart of the multi-modal fusion and control platform for human-robot interaction in the embodiment of the present invention;
[0057] Figure 3This is the control logic flow block diagram of the multi-modal fusion and control system for human-anthropomorphic robot interaction in the embodiments of the present invention.
[0058] In the figure: 1. Profile frame of the human body signal fusion and perception platform; 2. Profile frame of the double-arm robot mirror control platform; 3. Surface electromyogram multi-lead sensing bracelet; 4. Motion perception wearable glove; 5. Depth camera; 6. Display; 7. Industrial control computer of the depth camera; 8. Universal wheel; 9. Dexterous hand of the anthropomorphic robot; 10. Controller of the double-arm robot mirror control platform; 11. Anthropomorphic robot arm. Specific implementation manners
[0059] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solution of the present invention will be further described in detail below in conjunction with the embodiments and drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0060] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below in conjunction with the embodiments.
[0061] In the first aspect, the present invention provides a multi-modal fusion and control platform for human-anthropomorphic robot interaction, as Figure 1 shown, including a human body signal fusion and perception platform and a profile frame 2 of the double-arm robot mirror control platform.
[0062] The human body signal fusion and perception platform is the platform for users to perform movements, and the profile frame 2 of the double-arm robot mirror control platform is the platform where the anthropomorphic robot arm 11 and the dexterous hand 9 of the anthropomorphic robot for interacting with users are located.
[0063] The human signal fusion perception platform includes the profile frame 1 of the human signal fusion perception platform, the surface electromyogram multi-lead sensing bracelet 3, and the motion perception wearable glove 4. The profile frame 1 of the human signal fusion perception platform can provide a motion space for users and an installation and sliding position for sensors. The surface electromyogram multi-lead sensing bracelet 3 can collect muscle information of users, and the motion perception wearable glove 4 can capture the hand motion postures of users. The user wears the motion perception wearable glove 4 for contact acquisition of the hand motion posture information of the user, enhancing the robustness of the fine control of the system. The surface electromyogram multi-lead sensing bracelet 3 is worn at the forearm and upper arm positions of the two arms of the user, for perceiving the motion state information of the upper limbs of the human body and simultaneously obtaining the electromyogram signals of the upper limbs of the human body, for monitoring the fatigue degree of the user and preventing fatigue operation. With such a setting, the response speed and accuracy of the humanoid robot to human actions can be improved, thereby enhancing the naturalness and fluency of human-computer interaction.
[0064] This multimodal fusion and control platform has a high degree of flexibility and adaptability, and can be configured and adjusted according to different task requirements, thereby improving the operation efficiency and the quality of task completion. In some embodiments, the motion perception wearable glove 4 includes inertial sensors and bending sensors; the number of inertial sensors is eleven, and the inertial sensors are respectively distributed at the palm position, the metacarpophalangeal joints of each finger, and the proximal interphalangeal joints of each finger; the number of bending sensors is fourteen, and the bending sensors are respectively distributed at the metacarpophalangeal joints of each finger, the proximal interphalangeal joints of each finger, and the distal interphalangeal joints of the four fingers other than the thumb. The bending sensors are used to sense the angle of the finger bending towards the palm and capture the second-stage bending action of the finger.
[0065] The profile frame 1 of the human signal fusion perception platform and the profile frame 2 of the dual-arm robot mirror control platform are both rectangular frames. A data center is provided at the inner bottom end of the profile frame 1 of the human signal fusion perception platform. The data center includes a display 6 and several depth camera industrial computers 7. The display 6 is used for the debugging of the depth camera industrial computer 7 and the visual display of data. The depth camera industrial computer 7 as the terminal is used to connect to the depth camera 5 to collect the collected depth data and perform synchronous sending and processing of the collected data.
[0066] At the top of the profile frame 2 of the dual-arm robot mirror control platform, there is a dual-arm robot mirror control platform controller 10 for controlling the humanoid robotic arm 11 and the humanoid robotic dexterous hand 9. On the opposite sides of the dual-arm robot mirror control platform controller 10, there are humanoid robotic arms 11. At the end of the humanoid robotic arm 11 not connected to the dual-arm robot mirror control platform controller 10, there is a humanoid robotic dexterous hand 9. The humanoid robotic arm 11 includes three shoulder joints, one elbow joint, and three wrist joints. Each joint of the humanoid robotic arm 11 adopts a high-precision servo motor and a multi-stage reducer. The humanoid robotic arm 11 adopts a serial configuration and operates based on the form of multi-joint link following.
[0067] The anthropomorphic robotic dexterous hand 9 has six degrees of freedom. Among them, the thumb structure has two degrees of freedom, including the bending degree of freedom and the joint degree of freedom, which are used for precise grasping and operation of complex objects. The remaining four fingers each have one bending degree of freedom. The anthropomorphic robotic dexterous hand 9 is used for finger bending and grasping actions based on the form of multi-joint link following, through integrating high-precision micro servo motors and link mechanisms. The anthropomorphic robotic dexterous hand 9 is built-in with force and tactile sensors.
[0068] The dual-arm robot mirror control platform integrates multi-modal perception technology to receive and fuse motion data of the user's hand, forearm, upper arm, etc., as well as other various sensing information (such as vision, voice, and force feedback data). The platform adopts the same degree-of-freedom configuration as the human's two arms, with a robotic arm design of 7 degrees of freedom, and can accurately imitate the natural movements of the human upper limb. This highly biomimetic structure not only supports the precise replication of the user's actions to achieve refined remote control of the robotic arm, but also can generate interaction actions adapted to the human collaboration needs through algorithm optimization. The robotic arm of the platform adopts a serial joint structure, including 3 degrees of freedom at the shoulder (front-back swing, left-right swing, and rotation), 1 degree of freedom at the elbow (bending), and 3 degrees of freedom at the wrist (pitch, yaw, and rotation), restoring the movement range and flexibility of the human arm to the greatest extent. At the same time, the anthropomorphic dexterous hand has a five-finger design, in which the thumb has 2 degrees of freedom, and the remaining four fingers each have 1 degree of freedom. Adopting the form of multi-joint link following, it can efficiently complete fine grasping and complex operations. Force and tactile sensors are integrated at each fingertip joint, which can sense the corresponding pressure data information during grasping and operation. This design enables the dual-arm robot to achieve naturalness and high precision of actions in the teleoperation scenario, and accurately reproduce human actions in complex environments. In addition, the anthropomorphic structure of the dual-arm robot shows extremely high adaptability in the scenarios of human-robot interaction and collaboration. By combining the fusion processing of multi-modal perception data, the platform can collaborate with humans in a form more similar to human-to-human interaction, truly realizing human-robot coexistence. This design is not only applicable to medical surgery assistance, operations in dangerous environments such as nuclear power plants or chemical plants, but also can show strong adaptability and collaboration capabilities in high-complexity tasks such as external repairs of spacecraft.
[0069] In some embodiments, the provided space size of the profile frame 2 of the dual-arm robot mirror control platform can ensure that the anthropomorphic robotic arm 11 and the anthropomorphic robotic dexterous hand 9 have sufficient movement space to ensure the safety of movement.
[0070] In some embodiments, multiple depth cameras 5 are provided on the tops of the profile frame 1 of the human body signal fusion perception platform and the controller 10 of the dual-arm robot mirror control platform.
[0071] The depth camera 5 is used to sense the information of the interaction environment and the interaction object (human body behavior data and robot depth vision information). In some embodiments, the depth camera 5 is connected to the profile frame 1 of the human body signal fusion perception platform through a sliding device, so as to timely track and capture the movement of the user when the user moves. With such a setting, multi-angle non-contact acquisition of the user's hand movement information can be achieved, and the occlusion problem can be avoided. A plurality of depth cameras 5 are driven by motors on the profile frame 1 of the human body signal fusion perception platform and the double-arm robot mirror control platform controller 10, and the acquisition position of the depth camera 5 can be adjusted based on application requirements.
[0072] The double-arm robot mirror control platform controller 10 can complete the integration, connection, and processing of all components, and is also the center for human interaction. It is mainly used to receive the movement information of the user's hands, forearms, and upper arms and the perception data of each sensor obtained by the human body signal fusion perception platform, so as to realize the refined control of the remote operation of the robotic arms on both sides of the double-arm robot mirror control platform controller 10. Through the fusion processing of these signals and data, the double-arm robot mirror control platform controller 10 can accurately imitate the user's actions, or realize interaction and cooperation with the user, and is used to remotely complete specified tasks. The double-arm robot mirror control platform controller 10 is applicable to various complex environments that require remote operation, including but not limited to medical surgery assistance, operations in dangerous environments (such as nuclear power plants, chemical plants, etc.), and external repairs of spacecraft.
[0073] In some embodiments, universal wheels 8 are provided at the bottoms of the profile frame 1 of the human body signal fusion perception platform and the profile frame 2 of the double-arm robot mirror control platform, so as to facilitate assembly, debugging, and movement.
[0074] In some embodiments, a wheeled intelligent perception chassis is provided at the bottoms of the profile frame 1 of the human body signal fusion perception platform and the profile frame 2 of the double-arm robot mirror control platform, including an inertial sensor and a torque sensor, which are used to monitor and adjust the movement state of the platform in real time to ensure the stability and flexibility of the platform in different terrains and environments.
[0075] In some embodiments, the user wears a motion perception wearable glove 4 and a surface electromyography multi-lead sensing bracelet 3, and moves the upper limb into the motion space provided by the profile frame 1 of the human body signal fusion perception platform to collect relevant signals, electromyography information, and upper limb movement information.
[0076] Through the above design, the present invention provides a multi-modal fusion and control platform for efficient, flexible, and precise human-anthropomorphic robot interaction, which is applicable to various application scenarios that require refined operation and complex environment adaptation.
[0077] When using the multi-modal fusion and control platform through the above embodiments of the present invention, it can be configured diversely according to the needs of users. If the user needs to achieve high-precision and complete long-term hand interaction, after entering the motion capture space and wearing the device, a series of hand movements are completed in the space. In addition, if the motion actions involve using other instruments, such as holding other tools, different sensing devices can complement each other. The depth camera 5 and the motion sensing wearable glove 4 can complement each other in motion mode recognition. For example, the acceleration and angular velocity data provided by the motion sensing wearable glove 4 can supplement the deficiencies of the depth camera 5 in the case of occlusion. At the same time, in a specific motion, the platform can focus on the body parts involved in the action. For example, when making gestures, the pose and motion data of the palm can be focused on, reducing unnecessary data redundancy.
[0078] According to the specific requirements of the human-computer interaction task, the humanoid robotic arm 11 and the humanoid robotic dexterous hand 9 can respond to the user's interaction, including the response of the robotic arms and the response of the five-finger dexterous hand. In addition, in the robot platform, the data information sensed by the depth camera 5 and the tactile information of the five-finger dexterous hand are fed back to the embodied intelligent human-computer interaction algorithm model to realize a process closed-loop and achieve a real interaction effect.
[0079] In summary, the present invention provides a multi-modal fusion and control platform for human-robot interaction, which improves the interaction quality of humans and humanoid robots in action recognition, prediction, and tracking, increases the response speed and accuracy of humanoid robots to human actions, thereby enhancing the naturalness and fluency of human-computer interaction, and ensuring the interaction quality and experience in various complex environments.
[0080] In a second aspect, the present invention provides a multi-modal fusion and control system for human-robot interaction, which includes a multi-sensor fusion module, a spatial coordinate calibration and fusion module, an adaptive filtering module, and a human-computer interaction module;
[0081] The multi-sensor fusion module is used to collect and fuse the data of multiple sensors to obtain multi-sensor fusion data, and send the multi-sensor fusion data to the spatial coordinate calibration and fusion module. Among them, the data of multiple sensors include the surface electromyography data collected by the surface electromyography multi-lead sensing bracelet 3, the three-axis acceleration data, three-axis attitude data collected by the inertial sensors of the motion sensing wearable glove 4, and the curvature data collected by the curvature sensors. The multi-sensor fusion data includes the fused surface electromyography data and hand motion attitude data.
[0082] The spatial coordinate calibration and fusion module is used to correct and synchronize the human behavior data in different coordinate systems collected by the depth camera 5 of the profile frame 1 of the human signal fusion perception platform, and correct the coordinate of the depth vision information collected by the depth camera 5 of the controller 10 of the dual-arm robot mirror control platform, so as to obtain multi-sensor fusion calibration data, and the multi-sensor fusion calibration data has the same spatial coordinate system.
[0083] The adaptive filtering module is used to perform adaptive filtering on the human behavior data after correction and synchronization in the multi-sensor fusion calibration data and the pressure data collected by the built-in force and tactile sensor of the humanoid robotic dexterous hand 9. Sensor devices have their respective limitations: the depth camera 5 of the profile frame 1 of the human signal fusion perception platform may be affected by the light in the environment, resulting in errors, and data loss or blurring may also occur due to insufficient frame rate when the hand moves quickly; the motion sensing wearable glove 4 is worn on the human hand in the form of a wearable device, and data deviation may be caused by the different hand sizes of different users and the differences in wearing methods. The adaptive filtering technology adaptively adjusts the parameters of the filter and the fusion weights of different data according to the real-time states of different data, and obtains the filtered human behavior data and pressure data, so as to reduce noise and improve the accuracy of the data, and significantly improve the tracking performance of the hand in a variety of complex scenarios.
[0084] The human-computer interaction module is used to determine the robot response strategy based on the multi-sensor fusion data, the filtered human behavior data and the pressure data, and control the humanoid robotic arm 11 and the humanoid robotic dexterous hand 9 to interact with the user based on the robot response strategy. Among them, the robot response strategy includes the positions and postures of the humanoid robotic arm 11 and the humanoid robotic dexterous hand 9.
[0085] Exemplarily, as Figure 2 and Figure 3 shown, signal preprocessing (such as segmentation), correction and synchronization processing, coordinate correction processing, adaptive filtering processing, etc. are performed on the human behavior data and depth vision information, surface electromyogram data, finger bending degree data and inertial navigation unit information that respectively reflect the user's movement collected by the depth camera 5, the surface electromyogram multi-lead sensing bracelet 3, the motion sensing wearable glove 4 and the tactile sensor of the humanoid robotic dexterous hand 9, so as to obtain the filtered data, and then control the humanoid robotic arm 11 and the humanoid robotic dexterous hand 9 to complete human-computer interaction. In some embodiments, different human-computer interaction intention perception models are selected and built according to actual application requirements (such as trajectory planning, human-machine interaction, human-machine collaboration and human-machine coexistence).
[0086] In some embodiments, in the multi-sensor fusion module,
[0087] According to the following pose estimation formula, determine the hand movement pose of the user:
[0088]
[0089] In the formula, and are the attitude data in quaternion form at the (k - 1)-th moment and the k-th moment respectively, represents the quaternion multiplication symbol, is the time interval between the (k - 1)-th moment and the k-th moment, is the angular velocity data at the k-th moment;
[0090] According to the following formula, the acceleration data and the attitude data are transformed to obtain the linear acceleration in the global coordinate system:
[0091]
[0092] where is the linear acceleration in the global coordinate system, is the generated rotation matrix, is the acceleration data;
[0093] According to the following formula, the position trajectory of the hand movement is determined:
[0094]
[0095]
[0096] In the formula, is the velocity of the hand movement at the k-th moment, is the position of the hand movement at the k-th moment, is the velocity of the hand movement at the (k - 1)-th moment, is the position of the hand movement at the (k - 1)-th moment;
[0097] According to the following formula, the multi-sensor fusion data is determined:
[0098]
[0099]
[0100] In the formula, the number of hand movement postures in the multi-sensor fusion data at the k-th moment, is the fusion transformation function, is the bending degree data of the i-th finger joint at the k-th moment, is the resistance change of the i-th bending sensor at the k-th moment, is the change mapping function;
[0101] The surface electromyography data collected by the multi - lead sensing bracelet 3 for surface electromyography is expressed as , including the surface electromyography data of the left hand and the surface electromyography data of the right hand, where / represents the data of the i - th lead in the surface electromyography data of the left / right hand. According to the following formula, the hand muscle state data in the multi - sensor fusion data is determined:
[0102]
[0103] In the formula, and represent the hand muscle state data of the left hand and the right hand respectively, is the mapping function from the surface electromyography data to the hand muscle state data.
[0104] In order to ensure that the sensors with direction - sensing functions can perform precise spatial coordinate calibration, the present invention has developed a set of stable calibration procedures. This calibration procedure can perform high - precision coordinate calibration and transformation on multiple depth cameras through the provided spatial pose information and spatial confidence, realize the unification of the world coordinate system, and improve the overall performance of the invention.
[0105] In some embodiments, in the spatial coordinate calibration and fusion module,
[0106] According to the following formula, the human behavior data after correction and synchronization processing and the depth vision information after coordinate correction in the multi - sensor fusion calibration data are determined:
[0107]
[0108]
[0109] In the formula, W is the defined world coordinate system, is the coordinate system of the depth camera 5 itself on the i - th human signal fusion perception platform profile frame 1, represents the transformation matrix from the coordinate system to the W coordinate system, represents the human behavior data in the W coordinate system, represents the human behavior data in the coordinate system, represents the human behavior data after correction and synchronization processing in the multi - sensor fusion calibration data in the W coordinate system,
[0110] In some embodiments, the formula for determining the depth visual information after coordinate correction in the multi-sensor fusion calibration data may be the same as the formula for determining the human behavior data after correction and synchronization processing in the multi-sensor fusion calibration data.
[0111] In some embodiments, in the adaptive filtering module,
[0112]
[0113]
[0114] In the formula, The filtered human behavior data, The hand movement posture data in the multi-sensor fusion data at the k-th moment, Represents the human behavior data after correction and synchronization processing in the multi-sensor fusion calibration data in the W coordinate system, Is the fusion function, Is the mapping function used to map the pressure data matrix before filtering to the pressure data after filtering, Is the pressure data matrix before filtering at the i-th finger of the anthropomorphic robotic dexterous hand (9) , Represents the pressure data of the i-th finger of the anthropomorphic robotic dexterous hand (9) obtained by filtering.
[0115] In some embodiments, in the human-computer interaction module, the robot response strategy is determined through the embodied intelligence core algorithm of the generalization large model; the angles of each joint of the anthropomorphic robotic arm 11 and the anthropomorphic robotic dexterous hand 9 are determined through the robot inverse kinematics solution method.
[0116] In some embodiments, the embodied intelligence core algorithm of the generalization large model: Based on the generalization large model and the embodied intelligence theory, a perception system capable of perceiving human complex behaviors and environmental data is constructed. Through the combination of pre-training the large model and real-time environmental perception, the transfer of generalization ability is realized, so that the algorithm is not only applicable to specific scenarios but also can cope with complex and changeable environments.
[0117] The input of this algorithm mainly includes the following multi-modal data, denoted as:
[0118]
[0119] Among them, Is determined according to the actual task requirements, Denote the total task. In different actual tasks, the requirements for multimodal data are different. For example, in a certain task, there is no need for the pressure data of the humanoid robotic dexterous hand 9. In this scenario, tactile force data can be not collected, thereby reducing the data load of the system. Single-modal data cannot meet the capture of multiple levels of human behavior. For example, in hand activities, there is not only hand gesture motion information but also surface electromyography information. Therefore, different tasks are determined according to the actual task requirements.
[0120] Use a pre-trained large model to process the fused features, denoted as M(X), whose goal is to generate task-related outputs from the input features. For different tasks (such as robotic arm motion planning, dexterous hand operation, etc.), different decodings are used to dynamically adapt to the data characteristics of different scenarios. At the same time, utilize the understanding and generalization ability of the large model to output the required robot responses (i.e., the response behaviors of the humanoid robotic arm 11 and the humanoid robotic dexterous hand 9) in different scenarios.
[0121] Reinforcement learning framework: With the support of the generalized large model, construct a reinforcement learning framework. Based on user interaction data, environmental feedback, and the robot responses output by the large model, continuously optimize the robot response strategy to enable the robot to dynamically adapt to complex tasks in practice by adjusting actions, paths, and language outputs in real time. The robot strategy is represented by a parameterized model, and by optimizing the strategy parameters, the cumulative reward (expected return) is maximized. The purpose of reinforcement learning is to enable the robot to learn under insufficient sample sizes, learn the user's behavior, and thus achieve the same motion form as the user. It can be continuously trained according to the mechanism of reinforcement learning.
[0122] Robot control algorithm: Based on the robot response strategy given by the embodied intelligence core algorithm of the generalized large model, use the robot inverse kinematics algorithm to realize the conversion of control information, and then use the method of adaptive control to realize robot control.
[0123] The robot response strategy given by the embodied intelligence core algorithm of the generalized large model provides the output of the humanoid robotic arm 11 and the positions and postures of the humanoid robotic dexterous hand 9. By solving the robot inverse kinematics, the angles of each joint of the humanoid robotic arm and the humanoid robotic dexterous hand are obtained, so as to carry out further control.
[0124] According to the robot inverse kinematics, for a seven-degree-of-freedom humanoid robotic arm, the position of the end effector can be described as:
[0125]
[0126] where represents the transformation matrix of the i-th joint, T is the position of the end of the robotic arm, represents the joint angle.
[0127] The solution of joint angles can be achieved through the numerical solution of inverse kinematics , specifically using the iterative method, which iterates based on the joint angle error. The iteration expression is as follows:
[0128]
[0129] where is the current joint angle; is the Jacobian matrix, which describes the influence of joint changes on the end effector; is the error between the current end position and the target position; is the step factor, used to control the iteration step. Through multiple iterations, when the error is small enough, a solution that meets the accuracy is obtained.
[0130] For the anthropomorphic robotic dexterous hand 9, the four fingers except the thumb each have only one degree of freedom, and can be directly converted; the thumb uses the same inverse kinematics method as the anthropomorphic robotic arm 11 to achieve the conversion of joint angles.
[0131] The adaptive control algorithm combines the joint angle information provided by inverse kinematics and the feedback signal of the robot to achieve real-time parameter estimation of the robot system model, online update of controller parameters, ensure that the robot can maintain an accurate response, and overcome the uncertainty of the system model and external disturbances. Through the adaptive control algorithm, it is possible to quickly approach the control target at the beginning of control; as the error between the input and output gradually decreases, the system parameters and controller parameters change accordingly, approaching the control target with smaller output changes, achieving a fast and accurate response process. At the same time, in the presence of external disturbances, the adaptive control algorithm can automatically adjust the parameters to maintain stable control.
[0132] The specific control process is as follows:
[0133] The adaptive control algorithm dynamically estimates the parameters of the control system (such as the inertia and damping coefficients of the robot joints), and the estimation of these parameters usually depends on the input signal, output signal, and error signal. The basic form of control parameter update can be expressed as:
[0134]
[0135] where is the parameter estimation of the system; is the adaptive gain matrix; is the regression vector containing the input; It is the difference between the target value and the current output. The parameter estimation of the system will be continuously adjusted according to the current error to learn an accurate system model. At the same time, when affected by external disturbances, the system parameter estimation is automatically updated to facilitate subsequent stable control.
[0136] The controller of the adaptive control algorithm generates a control signal based on the current parameter estimation and the regression vector. The control signal is in the following form:
[0137]
[0138] where is the control signal; is the current parameter estimation; is the regression vector containing the input, which reflects the left and right of the system input. When there is an error between the system output and the target value, the parameter update rule automatically adjusts the control parameters through a feedback mechanism, making the control signal more accurate, so as to approximate the target output.
[0139] In some embodiments, the depth visual information collected by the depth camera 5 of the controller 10 of the dual-arm robot mirror control platform and the pressure data collected by the force and touch sensor of the humanoid robotic dexterous hand 9 can be sent to the multi-sensor fusion module and the spatial coordinate calibration and fusion module in real time for real-time feedback, so as to improve the accuracy of human-machine interaction.
[0140] In summary, the present invention provides a multi-modal fusion and control system for human-robot interaction, which improves the interaction quality between humans and humanoid robots in action recognition, prediction and tracking, increases the response speed and accuracy of humanoid robots to human actions, thereby enhancing the naturalness and fluency of human-machine interaction, and ensuring the interaction quality and experience in various complex environments.
[0141] The above specific embodiments are only explanations of the present invention, and they are not limitations of the present invention. Those skilled in the art can make modifications to these embodiments without creative contributions according to needs after reading this specification, but as long as they are within the scope of the claims of the present invention, they are protected by the patent law.
Claims
1. A multimodal fusion and control system for human-anthropomorphic robot interaction, characterized in that: The multi-modal fusion and control system includes a multi-sensor fusion module, a spatial coordinate calibration and fusion module, an adaptive filtering module, and a human-computer interaction module; The multi-sensor fusion module is used to collect and fuse data from multiple sensors to obtain multi-sensor fusion data, and send the multi-sensor fusion data to the spatial coordinate calibration and fusion module. Among them, the data of the multiple sensors includes surface electromyography data collected by the surface electromyography multi-lead sensing bracelet (3), triaxial acceleration data, triaxial attitude data, and curvature data collected by the inertial sensors of the motion perception wearable glove (4). The multi-sensor fusion data includes fused hand muscle state data and hand motion attitude data; In the multi-sensor fusion module, According to the following attitude estimation formula, determine the hand motion attitude of the user: wherein, and are attitude data in quaternion form at the (k - 1)-th moment and the k-th moment respectively, represents the quaternion multiplication symbol, is the time interval between the (k - 1)-th moment and the k-th moment, is the angular velocity data at the k-th moment; According to the following formula, transform the acceleration data and attitude data to obtain the linear acceleration in the global coordinate system: Among them, is the linear acceleration in the global coordinate system, is the generated rotation matrix, is the acceleration data; According to the following formula, determine the position trajectory of hand motion: Wherein, is the velocity of the hand movement at the k-th moment, is the position of the hand movement at the k-th moment, is the velocity of the hand movement at the (k - 1)-th moment, is the position of the hand movement at the (k - 1)-th moment; According to the following formula, determine the multi-sensor fusion data: In the formula, the hand motion posture data in the multi-sensor fusion data at the k-th moment, is the fusion transformation function, is the bending degree data of the i-th finger joint at the k-th moment, is the resistance change of the i-th bending sensor at the k-th moment, is the change mapping function; The surface electromyography data collected by the multi-lead sensing bracelet for surface electromyography (3) is expressed as , including the surface electromyography data of the left hand and the surface electromyography data of the right hand, where / represents the data of the i-th lead in the left / right hand surface electromyography data. According to the following formula, the hand muscle state data in the multi-sensor fusion data is determined: In the formula, and respectively represent the hand muscle state data of the left and right hands, is the mapping function from surface electromyography data to hand muscle state data; The spatial coordinate calibration and fusion module is used to correct and synchronize the human behavior data in different coordinate systems collected by the depth camera (5) of the human signal fusion perception platform profile frame (1), and correct the coordinate of the depth visual information collected by the depth camera (5) of the double-arm robot mirror control platform controller (10) to obtain multi-sensor fusion calibration data with the same spatial coordinate system; The adaptive filtering module is used to perform adaptive filtering on the human behavior data after correction and synchronization in the multi-sensor fusion calibration data and the pressure data collected by the built-in force and touch sensor of the humanoid robotic dexterous hand (9) to obtain filtered human behavior data and pressure data; The human-computer interaction module is used to determine the robot response strategy based on the multi-sensor fusion data, the filtered human behavior data, and the pressure data, and control the humanoid robotic arm (11) and the humanoid robotic dexterous hand (9) to interact with the user based on the robot response strategy. Among them, the robot response strategy includes the position and attitude of the humanoid robotic arm (11) and the humanoid robotic dexterous hand (9).
2. The multimodal fusion and control system for human-robot interaction according to claim 1, characterized in that: in In the spatial coordinate calibration and fusion module, According to the following formula, determine the human behavior data after correction and synchronization in the multi-sensor fusion calibration data and the depth visual information after coordinate correction: Where W is the defined world coordinate system, is the coordinate system of the depth camera (5) itself on the i-th human signal fusion perception platform profile frame (1), represents the transformation matrix from coordinate system to the W coordinate system, represents the human body behavior data in the W coordinate system, represents the human body behavior data in the coordinate system, represents the human body behavior data after calibration and synchronization processing in the multi-sensor fusion calibration data in the W coordinate system, is the spatial coordinate fusion function, and n represents the label of the sensor.
3. The multimodal fusion and control system for human-robot interaction according to claim 1, characterized in that: in In the adaptive filtering module, Wherein, Filtered human behavior data Hand movement posture data in the multi-sensor fusion data at the k-th moment Represents the human behavior data after calibration and synchronization processing in the multi-sensor fusion calibration data in the W coordinate system Is the fusion function Is the mapping function used to map the pressure data matrix before filtering to the pressure data after filtering Is the pressure data matrix before filtering at the i-th finger of the humanoid robot dexterous hand (9) , Represents the pressure data of the i-th finger of the humanoid robot dexterous hand (9) obtained by filtering 4. The multimodal fusion and control system for human-anthropomorphic robot interaction according to claim 1, characterized in that: in In the human-computer interaction module, determine the robot response strategy through the embodied intelligent core algorithm of the generalization large model; determine the angles of each joint of the humanoid robotic arm (11) and the humanoid robotic dexterous hand (9) through the robot inverse kinematics solution method.
5. A multimodal fusion and control platform for human-anthropomorphic robot interaction, which is applied to a multimodal fusion and control system for human-anthropomorphic robot interaction according to any one of claims 1-4, and is characterized in that: It includes a human signal fusion perception platform and a double-arm robot mirror control platform profile frame (2); The human signal fusion perception platform includes a human signal fusion perception platform profile frame (1), a surface electromyography multi-lead sensing bracelet (3), and a motion perception wearable glove (4); The human body signal fusion perception platform profile frame (1) and the dual-arm robot mirror control platform profile frame (2) are both rectangular frames. A data center is provided at the bottom of the human body signal fusion perception platform profile frame (1), and the data center includes a display (6) and a plurality of depth camera industrial computers (7); A dual-arm robot mirror control platform controller (10) is provided on the top of the dual-arm robot mirror control platform profile frame (2); humanoid mechanical arms (11) are provided on opposite sides of the dual-arm robot mirror control platform controller (10); a humanoid mechanical hand (9) is provided at one end of the humanoid mechanical arm (11) that is not connected to the dual-arm robot mirror control platform controller (10); the humanoid mechanical arm (11) comprises three shoulder joints, one elbow joint and three wrist joints; each joint of the humanoid mechanical arm (11) adopts a high-precision servo motor and a multi-stage reducer; the humanoid mechanical arm (11) adopts a series configuration and operates in the form of multi-joint connecting rod follow-up; The humanoid robot dexterous hand (9) comprises six degrees of freedom, wherein the thumb structure comprises two degrees of freedom, including a bending degree of freedom and a joint degree of freedom, which are used for the precise grasping and operation of complex objects, and the remaining four fingers each comprise a bending degree of freedom. The humanoid robot dexterous hand (9) is used to perform finger bending and grasping actions based on a multi-joint connecting rod follow-up form by integrating a high-precision micro-servo motor and a connecting rod mechanism, and the humanoid robot dexterous hand (9) has a built-in force tactile sensor.
6. A multimodal fusion and control platform for human-robot interaction with a humanoid robot as claimed in claim 5, characterized in that: The human signal fusion perception platform profile frame (1) and the dual-arm robot mirror control platform controller (10) are both provided with a plurality of depth cameras (5) on the top.
7. The multimodal fusion and control platform for human-robot interaction according to claim 5, characterized in that: The motion sensing wearable glove (4) comprises an inertial sensor and a bending sensor, which are used to capture hand movements; There are eleven inertial sensors, which are respectively distributed at the palm position, the metacarpophalangeal joint of each finger, and the proximal phalangeal joint of each finger; There are fourteen bending sensors, which are respectively distributed at the metacarpophalangeal joint of each finger, the proximal knuckle of each finger, and the distal knuckles of the four fingers other than the thumb. The bending sensors are used to sense the bending angle of the finger toward the palm and capture the second bending movement of the finger.
8. The multimodal fusion and control platform for human-robot interaction according to claim 6, characterized in that: The depth camera (5) is connected to the human signal fusion perception platform profile frame (1) and the dual-arm robot mirror control platform controller (10) via a slidable device.
9. A multimodal fusion and control platform for human-robot interaction as described in claim 5, characterized in that: The bottoms of the human body signal fusion perception platform profile frame (1) and the dual-arm robot mirror control platform profile frame (2) are both provided with universal wheels (8).
Citation Information
Patent Citations
Robot control system based on field depth sensing mechanism and working method of robot control system
CN106965183A
Multi-modal information fusion perception system for upper limb rehabilitation robot
CN112022619A