Industrial humanoid robot teleoperation cooperation system and method based on multi-modal motion capture fusion

By using a multimodal motion capture fusion system that combines optical and inertial motion capture modules, the problems of occlusion and cumulative errors are solved, enabling high-precision and stable remote operation of industrial robots that can adapt to complex industrial environments.

CN120839787APending Publication Date: 2025-10-28广州里工实业有限公司

Patent Information

Application Number
CN202511084310.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In existing industrial robot teleoperation technologies, optical motion measurement systems are susceptible to occlusion and may fail, while inertial motion capture systems suffer from cumulative errors, resulting in insufficient accuracy and unstable operation.

Method used

A multimodal motion capture fusion system is adopted, which combines an optical motion measurement module and an inertial motion capture module. The working mode is dynamically switched under different occlusion rates through a mode switching module. Combined with an environmental perception module and a force feedback device, multimodal data fusion and collaborative control are realized.

Benefits of technology

It improves the robot's operational stability and environmental adaptability in high-precision assembly and heavy-duty handling scenarios, enhances the continuity and accuracy of operations under occlusion conditions, and reduces assembly errors and the risk of load tipping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120839787A_ABST
    Figure CN120839787A_ABST
Patent Text Reader

Abstract

The invention provides an industrial humanoid robot teleoperation cooperation system and method based on multi-modal motion capture fusion, and belongs to the technical field of industrial humanoid robots. According to the invention, an optical motion measurement module is used for collecting whole body position data of an operator in an unshielded scene; the inertial motion capture module is used for collecting joint angle data of an operator in a shielding scene; the mode switching module is used for switching working modes; the environment sensing module is used for collecting environment data, workpiece images and robot tail end contact force information and sending the information to the cooperative control unit. The force feedback device is used for converting the contact force information of the tail end of the robot into tactile vibration and collecting grip strength data of an operator; the cooperative control unit is used for fusing the multi-modal dynamic capture pose data and generating an execution instruction in combination with an obstacle avoidance algorithm and a force control model; the industrial humanoid robot is used for receiving the execution instruction. The operation stability and environmental adaptability of the robot in scenes such as high-precision assembly and heavy carrying can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of industrial humanoid robot technology, and in particular to a collaborative system and method for remote operation of industrial humanoid robots based on multimodal motion capture fusion. Background Technology

[0002] In related technologies, motion capture methods in industrial robot remote operation technology have significant limitations:

[0003] 1. Optical motion measurement systems (such as VR teleoperation) rely on infrared cameras and reflective markers. They have high accuracy (≤0.1mm) but are easily affected by the obstruction of robotic arms and workpieces in industrial scenarios. When the obstruction rate is >30%, data is interrupted, leading to operation failure.

[0004] 2. Inertial motion capture systems (such as the Noitom solution) achieve full-body motion capture through wearable IMU sensors, which have strong anti-occlusion capabilities, but long-term use results in cumulative errors (≥1° / min), which can easily lead to positioning deviations in high-precision assembly;

[0005] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention

[0006] The main objective of this application is to propose a collaborative remote operation system and method for industrial humanoid robots based on multimodal motion capture fusion. This system can solve the problems of occlusion failure and insufficient accuracy of single motion capture methods in industrial scenarios, and improve the operational stability and environmental adaptability of robots in high-precision assembly, heavy handling and other scenarios.

[0007] To achieve the above objectives, one aspect of this application proposes a remote operation collaborative system for an industrial humanoid robot based on multimodal motion capture fusion. The system includes a multimodal motion capture unit, an environmental perception module, a force feedback device, a collaborative control unit, and an industrial humanoid robot.

[0008] The multimodal motion capture unit includes an optical motion measurement module, an inertial motion capture module, and a mode switching module;

[0009] The optical motion measurement module includes an infrared camera array and reflective markers; the optical motion measurement module is used to collect the operator's full-body position data in unobstructed scenes;

[0010] The inertial motion capture module includes a wearable IMU sensor network; the inertial motion capture module is used to collect joint angle data of the operator in occluded scenarios;

[0011] The mode switching module is used to switch the working mode;

[0012] The environmental perception module includes a 3D LiDAR, a binocular industrial camera, and a six-dimensional force sensor; the environmental perception module is used to collect environmental data, workpiece images, and robot end-effector contact force information, and send them to the collaborative control unit.

[0013] The force feedback device is used to convert the contact force information of the robot end into tactile vibration and to collect the operator's grip force data;

[0014] The collaborative control unit is used to receive multimodal motion capture pose data, environmental data, and grip force data, fuse the multimodal motion capture pose data, combine obstacle avoidance algorithms and force control models to generate execution commands, and convert the robot end-effector contact force information into force feedback signals; the multimodal motion capture pose data includes the whole-body pose data and the joint angle data;

[0015] The industrial humanoid robot is used to receive the execution command, control the movement of the waist, arms, hands and wheel-leg composite walking mechanism, and feed back the movement status to the collaborative control unit through body sensors.

[0016] In some embodiments, the mode switching module is further configured to switch the working mode according to the optical marker occlusion rate; the optical marker occlusion rate is equal to the number of occluded markers divided by the total number of markers;

[0017] When the optical marker occlusion rate is less than 30%, the mode switching module enables the optical-dominated fusion mode; the optical-dominated fusion mode includes setting the optical data weight to 0.8, setting the inertial data weight to 0.2, and calibrating the optical data for jittering through Kalman filtering;

[0018] When the optical marker occlusion rate is not less than 30% and less than 60%, the mode switching module enables the equal fusion mode; the equal fusion mode includes setting the optical data weight and the inertial data weight to be dynamically allocated within a preset weight range, and adaptively adjusting the size of the optical data weight based on the optical marker recognition rate;

[0019] The mode switching module enables the inertial independent mode when the optical marker occlusion rate is not less than 60%; the inertial independent mode includes suppressing the accumulated error of the inertial sensor through a preset zero-drift calibration model and calibrating the error to a preset error range;

[0020] The mode switching module is also used to perform a smooth transition using a preset time during mode switching.

[0021] In some embodiments, the optical motion measurement module and the inertial motion capture module are further configured to synchronize via timestamps and transmit data to the mode switching module via a USB 3.0 interface.

[0022] In some embodiments, the collaborative control unit is further configured to employ extended Kalman filtering when fusing the multimodal motion capture pose data, wherein the formula used includes:

[0023]

[0024] in, The pose data after fusion at time k; the pose data includes position data and attitude data; The predicted pose data at time k is based on the data at time k-1; K k The Kalman gain is calculated using the covariance matrix; z k The observed values ​​are optical or inertial data; H k This is the observation matrix.

[0025] In some embodiments, the optical motion measurement module is further configured to distribute the reflective markers on the operator's head, torso, shoulders, elbows, wrists, and finger joints, and to set the positioning accuracy and sampling frequency of the reflective markers;

[0026] The inertial motion capture module's IMU sensor incorporates a three-axis accelerometer and a three-axis gyroscope.

[0027] In some embodiments, the six-dimensional force sensor of the environmental perception module is installed on the fingertip of the robot's end effector to collect X, Y, and Z axis force components and torque components around the three axes.

[0028] In some embodiments, the force feedback device is a wearable glove with a built-in 12-channel vibration motor and pressure sensor.

[0029] In some embodiments, the formulas used in the force control model include:

[0030] F ref =K p ·(X d -X f )+K d ·(v d -v f )+K i ·∫(X d -X f )dt;

[0031] Among them, F ref For the desired output force; K p K is the proportionality coefficient. d K is the damping coefficient; i X is the integral coefficient; d X is the desired position; f For actual location; vd For the desired speed, V f This refers to the actual speed.

[0032] To achieve the above objectives, another aspect of this application proposes a collaborative teleoperation method for industrial humanoid robots based on multimodal motion capture fusion, which is implemented through the aforementioned collaborative teleoperation system for industrial humanoid robots based on multimodal motion capture fusion. The method includes:

[0033] Multimodal motion capture units are used to collect multimodal motion capture pose data of operators;

[0034] Based on the multimodal motion capture pose data, the mode switching module outputs fused pose data according to the occlusion rate of optical markers;

[0035] The environmental perception module is used to synchronously collect environmental information, workpiece images, and robot end-effector contact force information.

[0036] The collaborative control unit integrates the multimodal motion capture pose data with the environmental information, and generates execution instructions through obstacle avoidance algorithms and force control models;

[0037] The execution command is received by an industrial humanoid robot, and the contact force information of the robot's end effector is converted into tactile vibration feedback to the operator to obtain robot motion state feedback data.

[0038] Based on the robot's motion state feedback data, the motion capture fusion weights and control parameters are dynamically adjusted to form a closed-loop control.

[0039] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.

[0040] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.

[0041] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.

[0042] The embodiments of this application include at least the following beneficial effects: This application provides a collaborative method, system, electronic device, storage medium, and program product for remote operation of industrial humanoid robots based on multimodal motion capture fusion. The invention includes: an optical motion measurement module for collecting full-body positional data of the operator in unobstructed scenarios; an inertial motion capture module for collecting joint angle data of the operator in obstructed scenarios; a mode switching module for switching working modes; an environmental perception module for collecting environmental data, workpiece images, and robot end-effector contact force information, and sending them to the collaborative control unit; a force feedback device for converting the robot end-effector contact force information into tactile vibrations and collecting the operator's grip force data; a collaborative control unit for fusing multimodal motion capture pose data, combining obstacle avoidance algorithms and force control models to generate execution commands; and an industrial humanoid robot for receiving execution commands. This invention can improve the operational stability and environmental adaptability of robots in high-precision assembly, heavy-duty handling, and other scenarios. Attached Figure Description

[0043] Figure 1 This is a schematic diagram of a remote operation collaborative system for industrial humanoid robots based on multimodal motion capture fusion provided in an embodiment of this application;

[0044] Figure 2 This is a system architecture diagram of the industrial humanoid robot teleoperation collaborative system based on multimodal motion capture fusion provided in the embodiments of this application;

[0045] Figure 3 This is a flowchart of the workflow of the industrial humanoid robot teleoperation collaborative system based on multimodal motion capture fusion provided in the embodiments of this application;

[0046] Figure 4 This is a block diagram of the motion capture data fusion algorithm for a remote operation collaborative system for industrial humanoid robots based on multimodal motion capture fusion, provided in an embodiment of this application. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0048] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”

[0049] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0051] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained first. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.

[0052] 1) IMU (Inertial Measurement Unit): A combination of sensors that measure the three-axis angular velocity and acceleration of an object;

[0053] 2) VR (Virtual Reality);

[0054] 3) EKF (Extended Kalman Filter): An extended version of the Kalman filter, used for state estimation of nonlinear systems;

[0055] 4) FPGA (Field-Programmable Gate Array), a programmable logic device;

[0056] 5) Roll, the angle of rotation of an object around its front and rear axes (X-axis);

[0057] 6) Pitch, the angle of rotation of an object around its left and right axes (Y-axis);

[0058] 7) yaw, the angle of rotation of an object around its vertical axis (Z-axis);

[0059] 8) RRT (Rapidly-exploring Random Tree), a random search algorithm for path planning, suitable for high-dimensional spaces;

[0060] 9) PID (Proportional-Integral-Derivative) is a closed-loop control algorithm that uses proportional, integral, and derivative control errors.

[0061] In related technologies, motion capture methods in industrial robot remote operation technology have significant limitations:

[0062] 1. Optical motion measurement systems (such as VR teleoperation) rely on infrared cameras and reflective markers. They have high accuracy (≤0.1mm) but are easily affected by the obstruction of robotic arms and workpieces in industrial scenarios. When the obstruction rate is >30%, data is interrupted, leading to operation failure.

[0063] 2. Inertial motion capture systems (such as the Noitom solution) achieve full-body motion capture through wearable IMU sensors, which have strong anti-occlusion capabilities, but long-term use results in cumulative errors (≥1° / min), which can easily lead to positioning deviations in high-precision assembly;

[0064] Neither approach has achieved deep adaptation to industrial scenarios: the optical system lacks an occlusion tolerance mechanism, and the inertial system does not incorporate environmental perception for error compensation, making it difficult to meet the operational needs of complex industrial environments.

[0065] In view of this, this application provides a collaborative system and method for remote operation of an industrial humanoid robot based on multimodal motion capture fusion. The system includes a multimodal motion capture unit (containing an optical motion measurement module and an inertial motion capture module), an environmental perception module, a force feedback device, a collaborative control unit, and an industrial humanoid robot. The multimodal motion capture unit dynamically switches or fuses optical and inertial motion capture data by adjusting the occlusion rate, outputting a high-precision operator pose. The environmental perception module collects three-dimensional information and contact force data of the working environment. The force feedback device converts the robot's end effector force information into tactile vibrations. The collaborative control unit fuses multi-source data to achieve "remote operation-autonomous control" collaboration. This invention solves the problems of occlusion failure and insufficient accuracy of single motion capture methods in industrial scenarios, improving the operational stability and environmental adaptability of robots in high-precision assembly, heavy-duty handling, and other scenarios.

[0066] The technical problems solved by this invention include:

[0067] 1. The single motion capture method has insufficient scene adaptability: the optical system is prone to occlusion failure, and the inertial system has cumulative errors;

[0068] 2. In industrial settings, low remote operation accuracy (assembly error > 0.5mm) and lack of force information lead to overpressure damage to workpieces;

[0069] 3. The robot's posture stability is poor when handling heavy loads, and it lacks a collaborative control mechanism based on multi-source motion capture data.

[0070] The embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0071] On one hand, embodiments of the present invention provide a remote operation collaborative system for industrial humanoid robots based on multimodal motion capture fusion, with reference to... Figure 1 The system includes a multimodal motion capture unit, an environmental perception module, a force feedback device, a collaborative control unit, and an industrial humanoid robot.

[0072] The multimodal motion capture unit includes an optical motion measurement module, an inertial motion capture module, and a mode switching module;

[0073] The optical motion measurement module includes an infrared camera array and reflective markers; the optical motion measurement module is used to collect the operator's full-body position data in unobstructed scenes;

[0074] The inertial motion capture module includes a wearable IMU sensor network; the inertial motion capture module is used to collect joint angle data of the operator in occluded scenarios;

[0075] The mode switching module is used to switch working modes;

[0076] The environmental perception module includes a 3D LiDAR, a binocular industrial camera, and a six-dimensional force sensor; the environmental perception module is used to collect environmental data, workpiece images, and robot end-effector contact force information, and send them to the collaborative control unit.

[0077] Force feedback devices are used to convert contact force information at the robot's end effector into tactile vibrations and to collect the operator's grip force data;

[0078] The collaborative control unit receives multimodal motion capture pose data, environmental data, and grip force data, fuses the multimodal motion capture pose data, combines obstacle avoidance algorithms and force control models to generate execution commands, and converts the robot's end-effector contact force information into force feedback signals; the multimodal motion capture pose data includes whole-body pose data and joint angle data;

[0079] Industrial humanoid robots are used to receive and execute commands, control the movement of the waist, arms, hands, and wheel-leg composite walking mechanism, and feed back the movement status to the collaborative control unit through body sensors.

[0080] Furthermore, the optical motion measurement module consists of an infrared camera array (≥8 units) and reflective markers, used to acquire the operator's full-body pose in an unobstructed scene (positioning accuracy ≤0.1mm, sampling frequency 120Hz);

[0081] The inertial motion capture module consists of a wearable IMU sensor network (≥16 nodes, sampling rate 1000Hz) used to collect operator joint angles (accuracy ±0.5°) in occluded scenarios.

[0082] The mode switching module is used to switch the working mode according to the optical marker occlusion rate α (α = number of occluded markers / total number of markers): when α < 30%, the optical-dominated fusion mode is enabled; when 30% ≤ α < 60%, the equal fusion mode is enabled; and when α ≥ 60%, the inertial independent mode is enabled.

[0083] Furthermore, the environmental perception module includes a 3D LiDAR, a binocular industrial camera, and a six-dimensional force sensor, which are used to collect three-dimensional point clouds of the working environment, high-definition images of the workpiece, and contact force information (Fx, Fy, Fz, Mx, My, Mz) of the robot end effector, and send them to the collaborative control unit.

[0084] Furthermore, the collaborative control unit fuses motion capture data (multimodal motion capture pose data) through Kalman filtering.

[0085] In some embodiments, the mode switching module disclosed in this invention is further configured to switch the working mode according to the optical marker occlusion rate; the optical marker occlusion rate is equal to the number of occluded markers divided by the total number of markers;

[0086] When the optical marker occlusion rate is less than 30%, the mode switching module enables the optical-dominated fusion mode. The optical-dominated fusion mode includes setting the weight of optical data to 0.8 and the weight of inertial data to 0.2, and performing jitter on the optical data by calibrating it through Kalman filtering.

[0087] When the optical marker occlusion rate is not less than 30% and less than 60%, the mode switching module enables the equal fusion mode. The equal fusion mode includes setting the weight of optical data and the weight of inertial data to be dynamically allocated within a preset weight range, and adaptively adjusting the weight of optical data based on the optical marker recognition rate.

[0088] When the optical marker occlusion rate is not less than 60%, the mode switching module enables the inertial independent mode. The inertial independent mode includes suppressing the accumulated error of the inertial sensor by using a preset zero-drift calibration model and calibrating the error to a preset error range.

[0089] The mode switching module is also used to perform a smooth transition using a preset time during mode switching.

[0090] Furthermore, the fusion logic of the mode switching module includes:

[0091] Optical-dominated fusion mode: optical data weight 0.8, inertial data weight 0.2, and high-frequency jitter of optical data is calibrated by Kalman filtering;

[0092] Equal fusion mode: The weights of optical and inertial data are dynamically allocated (0.6-0.4), and are adaptively adjusted based on the optical marker recognition rate (0.6 for optical weight when ≥90%, and linearly reduced to 0.4 when <90%).

[0093] Inertial Independent Mode: Suppresses the cumulative error of the inertial sensor by using a preset zero-drift calibration model (triggered once every 5 minutes), with a calibration error ≤0.1° / h;

[0094] A smooth 0.5s transition is used when switching modes to ensure that the pose data does not jump.

[0095] In some embodiments, the optical motion measurement module and the inertial motion capture module disclosed in this invention are further used to synchronize data via timestamps and transmit the data to the mode switching module via a USB 3.0 interface.

[0096] Furthermore, the optical motion measurement module and the inertial motion capture module are synchronized via timestamps (error ≤ 1ms), and the data is transmitted to the mode switching module via a USB 3.0 interface.

[0097] In some embodiments, the cooperative control unit disclosed in this invention is further configured to employ extended Kalman filtering when fusing multimodal motion capture pose data, and the formula used includes:

[0098]

[0099] in, The pose data after fusion at time k; the pose data includes position data and attitude data. The predicted pose data at time k is based on the data at time k-1; K k This represents the Kalman gain; the Kalman gain is calculated using the covariance matrix; z k For observations; observations are optical data or inertial data; H k This is the observation matrix.

[0100] Furthermore, the motion capture data fusion algorithm of the collaborative control unit adopts extended Kalman filtering (EKF);

[0101] The fused pose data at time k includes position (x, y, z) and attitude (roll, pitch, yaw);

[0102] The observations include optical data z optor inertial data z inert ;

[0103] The observation matrix includes optical mode H opt =0.8, inertial mode H inert =0.2.

[0104] In some embodiments, the optical motion measurement module disclosed in this invention is also used to distribute reflective markers on the operator's head, torso, shoulders, elbows, wrists, and finger joints, and to set the positioning accuracy and sampling frequency of the reflective markers;

[0105] The inertial motion capture module's IMU sensor incorporates a three-axis accelerometer and a three-axis gyroscope.

[0106] Furthermore, the reflective markers of the optical motion measurement module are distributed on the operator's head, torso, shoulders, elbows, wrists, and finger joints, with a positioning accuracy of ≤0.1mm and a sampling frequency of 120Hz; the IMU sensor of the inertial motion capture system has a built-in three-axis accelerometer (range ±16g) and a three-axis gyroscope (range ±2000° / s), with a joint angle measurement accuracy of ±0.5°.

[0107] In some embodiments, the six-dimensional force sensor of the environmental perception module disclosed in this invention is installed on the fingertip of the robot's end effector to collect X, Y, and Z axis force components and torque components around the three axes.

[0108] Furthermore, the six-dimensional force sensor of the environmental perception module is installed on the fingertip of the robot's end effector to collect the force components (Fx, Fy, Fz) of the X, Y, and Z axes and the torque components (Mx, My, Mz) around the three axes, with a sampling frequency ≥1kHz and a measurement range of ±500N.

[0109] In some embodiments, the force feedback device disclosed in this invention is a wearable glove with 12 built-in vibration motors and pressure sensors.

[0110] Furthermore, the force feedback device is a wearable glove with 12 built-in vibration motors (frequency adjustable from 50-500Hz) and pressure sensors, used to convert contact force information (robot end contact force information) into tactile vibration and collect operator grip force data.

[0111] In some embodiments, the formulas used in the force control model disclosed in this invention include:

[0112] F ref =K p ·(X d -X f )+K d ·(v d -v f )+Ki ·∫(X d -X f )dt;

[0113] Among them, F ref For the desired output force; K p K is the proportionality coefficient. d K is the damping coefficient; i X is the integral coefficient; d X is the desired position; f For actual location; v d For the desired speed, v f This refers to the actual speed.

[0114] Furthermore, K p The proportionality coefficient is (500-2000 N / m); K d The damping coefficient is (50-200 N·s / m); K i The integral coefficient is (10-50 N·s / m).

[0115] On the other hand, embodiments of the present invention also provide a collaborative teleoperation method for industrial humanoid robots based on multimodal motion capture fusion, which is implemented through the aforementioned collaborative teleoperation system for industrial humanoid robots based on multimodal motion capture fusion. The method includes:

[0116] Multimodal motion capture units are used to collect multimodal motion capture pose data of operators;

[0117] Based on the multimodal motion capture pose data, the mode switching module outputs fused pose data according to the occlusion rate of optical markers;

[0118] The environmental perception module is used to synchronously collect environmental information, workpiece images, and robot end-effector contact force information.

[0119] The collaborative control unit integrates multimodal motion capture pose data with environmental information, and generates execution commands through obstacle avoidance algorithms and force control models;

[0120] The industrial humanoid robot receives and executes commands, and uses a force feedback device to convert the contact force information of the robot's end effector into tactile vibration feedback to the operator, thereby obtaining robot motion state feedback data.

[0121] Based on the robot's motion state feedback data, the motion capture fusion weights and control parameters are dynamically adjusted to form a closed-loop control.

[0122] Furthermore, collaborative remote operation methods for industrial humanoid robots based on multimodal motion capture fusion include:

[0123] The multimodal motion capture unit collects the operator's pose in real time, and the mode switching module outputs fused pose data based on the occlusion rate α.

[0124] The environmental perception module synchronously collects 3D point cloud data of the working environment, workpiece images, and contact force information;

[0125] The collaborative control unit integrates motion capture data and environmental information, and generates execution commands through obstacle avoidance algorithms and force control models;

[0126] The industrial humanoid robot executes commands, and the force feedback device converts the contact force into tactile vibrations and feeds them back to the operator.

[0127] Based on the robot's motion state feedback, the motion capture fusion weights and control parameters are dynamically adjusted to form a closed-loop control.

[0128] As an optional implementation, the core components of the industrial humanoid robot teleoperation collaborative system based on multimodal motion capture fusion in this embodiment of the invention include:

[0129] Multimodal motion capture unit:

[0130] Optical motion measurement module: 8 infrared cameras (120° field of view, 120Hz sampling rate) are arranged in a ring in the work area, and 20 reflective markers are affixed to key nodes on the operator's body to achieve sub-millimeter pose measurement when there is no obstruction;

[0131] Inertial motion capture module: 16 IMU sensors form a wearable network (1 for the head, 2 for the torso, and 13 for the limbs), with a sampling rate of 1000Hz, and real-time output of joint angles (accuracy ±0.5°);

[0132] Mode switching module: The occlusion rate α (number of occluded markers / total number of markers) is calculated in real time through the optical system, and the fusion mode is dynamically switched (α < 30% → optical dominance, 30% ≤ α < 60% → equal fusion, α ≥ 60% → inertial independence).

[0133] Environmental perception module:

[0134] 3D LiDAR (scanning frequency 10Hz, ranging accuracy ±2cm): Constructs a 3D point cloud of the working environment;

[0135] Binocular industrial camera (12MP resolution, 30fps): identifies workpiece pose (positioning error <0.1mm);

[0136] Six-dimensional force sensor (measurement range ±500N, accuracy 0.1N): installed on the robot's fingertip to collect contact force information.

[0137] Force feedback device:

[0138] The wearable glove integrates 12 vibration motors (frequency 50-500Hz) and pressure sensors (range 0-50N), mapping the magnitude of contact force to vibration intensity (1N corresponds to 50Hz, 50N corresponds to 500Hz), while also collecting the operator's grip force to control the robot's end effector clamping force.

[0139] Cooperative control unit:

[0140] A GPU+FPGA architecture is adopted to achieve real-time fusion of multi-source data (processing latency <50ms);

[0141] Core algorithms: Extended Kalman filter motion capture fusion algorithm, point cloud-based RRT obstacle avoidance algorithm, force control PID adjustment model, and load adaptive attitude compensation algorithm.

[0142] Industrial humanoid robots:

[0143] The whole body has 28 degrees of freedom, the end effector has a load capacity of 30kg, and the waist uses a harmonic reducer (±120° rotation);

[0144] Bipedal or wheeled walking mechanism: walking speed 0.8m / s (load > 20kg), wheel speed 1.5m / s (load < 20kg), the body integrates a six-axis IMU and foot force sensor for posture stability control.

[0145] The beneficial effects of the embodiments of the present invention include:

[0146] 1. Multimodal motion capture fusion: When the occlusion rate is <60%, the pose accuracy is ≤0.1mm, the cumulative error is <0.1° / h, and the operation stability is improved by 100% compared with the single motion capture method;

[0147] 2. Obstruction adaptability: In scenarios with strong obstruction such as engine compartment (α=70%), the inertial independent mode still maintains operational continuity, and the operating efficiency is improved by 60% compared with the optical single mode;

[0148] 3. Operational accuracy: Combining force feedback and multimodal motion capture, assembly error is reduced to the 0.1mm level, and contact force control error is <1N;

[0149] 4. Load stability: The wheel-leg composite structure combined with the attitude compensation algorithm reduces the risk of tipping over by 60% during heavy handling.

[0150] As an optional implementation, the system architecture reference of the embodiments of the present invention is as follows: Figure 2 The system is divided into a three-layer architecture:

[0151] 1. Perception layer:

[0152] Multimodal motion capture unit: optical camera array (arranged around the work area, 2.5m high) and inertial sensor network (wearable vest + gloves + leg sensors);

[0153] Environmental sensing devices: LiDAR (mounted on the robot's shoulder, scanning range 360°×90°), binocular camera (robot head, focal length 12mm), and six-dimensional force sensor (fingertip).

[0154] 2. Control Layer:

[0155] Collaborative control unit: realizes three core functions:

[0156] Motion capture data fusion: Optical and inertial data are fused using an extended Kalman filter algorithm. Key parameters:

[0157] State vector x = [x, y, z, roll, pitch, yaw] (position + orientation);

[0158] Process noise covariance Q: Q = 0.01 for optically dominant mode and Q = 0.1 for inertial independent mode;

[0159] Observation noise covariance R: R = 0.001 for optical data, R = 0.01 for inertial data;

[0160] Obstacle avoidance algorithm: RRT path planning based on octree, node expansion step size 0.05m, safety distance 0.1-0.5m (dynamically adjustable);

[0161] Force control model: Using the above force control model formula, K p =1500N / m, K d =100 N·s / m, K i =30N·s / m, and the contact force is stabilized within ±1N of the target value through PID adjustment.

[0162] 3. Execution layer:

[0163] Industrial humanoid robot: The joint drive uses a 200N·m servo motor (with a 16-bit absolute encoder), and the end effector is a two-finger gripper (gripping force adjustable from 0-1000N, surface rubber material with a friction coefficient of 0.8);

[0164] Force feedback gloves: 12 vibration motors (distributed at the fingertips and knuckles), pressure sensor range 0-50N (accuracy 0.1N), data transmission via Bluetooth 5.0.

[0165] As an optional implementation, this embodiment of the invention takes the assembly of aero-engine blades (in a scenario with strong obstruction) as an example, and the process is as follows: Figure 3 As shown:

[0166] 1. Initialization and Calibration (S1):

[0167] The operator wears an inertial sensor network and attaches optical reflective markers. The system establishes the optical-inertial coordinate system transformation matrix T (rotation matrix R + translation vector t) using the "three-point calibration method".

[0168] The lidar scans the working environment to generate a 3D point cloud. The collaborative control unit uses a Euclidean clustering algorithm to identify the engine block (clustering threshold 0.2m) and the blade storage position (positioning error <0.5mm).

[0169] 2. Mode Selection (S2):

[0170] The system detects the optical marker occlusion rate α = 10% (no obvious occlusion) in real time, and activates the "optical-dominated fusion mode" with an optical data weight of 0.8 and an inertial data weight of 0.2.

[0171] The VR headset displays a virtual scene (environmental point cloud + robot digital twin + blade model), and the operator adjusts the viewing angle by rotating their head (collected by inertial sensors, with a viewing angle range of ±150°).

[0172] 3. Grab the blade (S3):

[0173] The operator's hand movements are captured by an inertial glove and converted into robot finger opening and closing commands (opening and closing speed 50mm / s);

[0174] When in contact with the blade, the six-dimensional force sensor detects F. z =20N, the force feedback glove's fingertip motor vibrates at 200Hz to indicate that contact has occurred;

[0175] The collaborative control unit automatically avoids blade edge protrusions based on lidar point clouds (safe distance 0.05m).

[0176] 4. Scene switching due to occlusion (S4):

[0177] When the robotic arm carrying the blades enters the engine compartment, the optical marker occlusion rate α = 65%, and the system automatically switches to "inertial independent mode";

[0178] Inertial data is used to correct accumulated errors through a zero-drift calibration model (triggered every 5 minutes) to ensure the continuity of operations inside the cabin;

[0179] The operator controls the robot's posture using inertial sensors, while the VR headset displays a magnified view of the cabin (captured by a binocular camera).

[0180] 5. High-precision assembly (S5):

[0181] When the blade approaches the mounting surface, the force control model is activated, using formula F.ref =K p ·(X d -X f )+K d ·(v d -v f )+K i ·∫(X d -X f Adjust the end position to stabilize the contact force at 15±1N;

[0182] The binocular camera identifies the deviation in the mounting hole position (Δx = 0.2mm, Δy = 0.1mm), and the collaborative control unit corrects the trajectory, compensating for the deviation to <0.1mm;

[0183] The operator senses the contact state by observing changes in the vibration intensity of the force feedback glove, thus completing the final assembly.

[0184] 6. Assignment completed (S6):

[0185] The system records assembly trajectory, force parameters, and motion capture fusion weights to generate autonomous operation templates. Subsequent similar operations can be executed with "one-click reproduction".

[0186] The robot exits the cabin, with an occlusion rate of α = 20%, and automatically switches back to the "optical-dominated fusion mode".

[0187] Key parameter settings for embodiments of the present invention:

[0188] Motion capture mode switching thresholds: occlusion rate α = 30% (optical-dominated → equal fusion), α = 60% (equal fusion → inertial-independent);

[0189] Force feedback vibration mapping: 0-10N→50-200Hz (linear mapping), 10-50N→200-500Hz (linear mapping), >50N triggers 3 strong vibration alarms;

[0190] Inertial calibration cycle: triggered once every 5 minutes, with calibration conditions being that the operator remains stationary for ≥10 seconds (joint angle change <0.5°);

[0191] Obstacle avoidance safety distance: 0.3m for static obstacles, 0.5m for dynamic obstacles (such as mobile forklifts).

[0192] As an optional implementation method, refer to Figure 4 The block diagram of the motion capture data (multimodal motion capture pose data) fusion algorithm of this invention includes four modules: optical data preprocessing (Gaussian filtering), inertial data preprocessing (low-pass filtering), extended Kalman filter fusion, and output calibration. The input is the original optical and inertial data, and the output is the fused pose.

[0193] On the other hand, embodiments of the present invention also provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned method.

[0194] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0195] Another embodiment of the hardware structure of the electronic device, the electronic device including:

[0196] The processor can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to achieve the technical solutions provided in the embodiments of this application.

[0197] The memory can be implemented in the form of read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory and called by the processor to execute the methods described in the embodiments of this application.

[0198] Input / output interfaces are used to implement information input and output;

[0199] The communication interface is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0200] A bus is used to transfer information between various components of a device, such as processors, memory, input / output interfaces, and communication interfaces.

[0201] The processor, memory, input / output interfaces, and communication interfaces communicate with each other within the device via a bus.

[0202] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0203] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0204] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0205] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0206] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0207] This application provides a method, system, electronic device, storage medium, and program product for collaborative teleoperation of industrial humanoid robots based on multimodal motion capture fusion. It belongs to the field of humanoid robot teleoperation control technology in industrial environments. The method includes a multimodal motion capture unit (containing an optical motion measurement module and an inertial motion capture module), an environmental perception module, a force feedback device, a collaborative control unit, and an industrial humanoid robot. The multimodal motion capture unit dynamically switches or fuses optical and inertial motion capture data by adjusting the occlusion rate, outputting a high-precision operator pose. The environmental perception module collects three-dimensional information and contact force data of the working environment. The force feedback device converts the robot's end effector force information into tactile vibration. The collaborative control unit fuses multi-source data to achieve "teleoperation-autonomous control" collaboration. This invention is applicable to industrial scenarios such as high-precision assembly, heavy-load handling, and complex occluded environments, and particularly addresses the insufficient adaptability of single motion capture methods in industrial scenarios by proposing a solution.

[0208] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0209] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0210] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0211] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0212] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0213] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0214] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0215] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0216] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0217] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0218] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A collaborative remote operation system for industrial humanoid robots based on multimodal motion capture fusion, characterized in that, The system includes a multimodal motion capture unit, an environmental perception module, a force feedback device, a collaborative control unit, and an industrial humanoid robot. The multimodal motion capture unit includes an optical motion measurement module, an inertial motion capture module, and a mode switching module; The optical motion measurement module includes an infrared camera array and reflective markers; the optical motion measurement module is used to collect the operator's full-body position data in unobstructed scenes; The inertial motion capture module includes a wearable IMU sensor network; the inertial motion capture module is used to collect joint angle data of the operator in occluded scenarios; The mode switching module is used to switch the working mode; The environmental perception module includes a 3D LiDAR, a binocular industrial camera, and a six-dimensional force sensor; the environmental perception module is used to collect environmental data, workpiece images, and robot end-effector contact force information, and send them to the collaborative control unit. The force feedback device is used to convert the contact force information of the robot end into tactile vibration and to collect the operator's grip force data; The collaborative control unit is used to receive multimodal motion capture pose data, environmental data, and grip force data, fuse the multimodal motion capture pose data, combine obstacle avoidance algorithms and force control models to generate execution commands, and convert the robot end-effector contact force information into force feedback signals; the multimodal motion capture pose data includes the whole-body pose data and the joint angle data; The industrial humanoid robot is used to receive the execution command, control the movement of the waist, arms, hands and wheel-leg composite walking mechanism, and feed back the movement status to the collaborative control unit through body sensors.

2. The system according to claim 1, characterized in that, The mode switching module is also used to switch the working mode according to the optical marker occlusion rate; the optical marker occlusion rate is equal to the number of occluded markers divided by the total number of markers. When the optical marker occlusion rate is less than 30%, the mode switching module enables the optical-dominated fusion mode; the optical-dominated fusion mode includes setting the optical data weight to 0.8, setting the inertial data weight to 0.2, and calibrating the optical data for jittering through Kalman filtering; When the optical marker occlusion rate is not less than 30% and less than 60%, the mode switching module enables the equal fusion mode; the equal fusion mode includes setting the optical data weight and the inertial data weight to be dynamically allocated within a preset weight range, and adaptively adjusting the size of the optical data weight based on the optical marker recognition rate; The mode switching module enables the inertial independent mode when the optical marker occlusion rate is not less than 60%; the inertial independent mode includes suppressing the accumulated error of the inertial sensor through a preset zero-drift calibration model and calibrating the error to a preset error range; The mode switching module is also used to perform a smooth transition using a preset time during mode switching.

3. The system according to claim 1, characterized in that, The optical motion measurement module and the inertial motion capture module are also used to synchronize data via timestamps and transmit the data to the mode switching module via a USB 3.0 interface.

4. The system according to claim 1, characterized in that, The collaborative control unit is also used to employ extended Kalman filtering when fusing the multimodal motion capture pose data, and the formula used includes: in, The pose data after fusion at time k; the pose data includes position data and attitude data; The predicted pose data at time k is based on the data at time k-1; K k The Kalman gain is calculated using the covariance matrix; z k The observed values ​​are optical or inertial data; H k This is the observation matrix.

5. The system according to claim 1, characterized in that, The optical motion measurement module is also used to distribute the reflective markers on the operator's head, torso, shoulders, elbows, wrists, and finger joints, and to set the positioning accuracy and sampling frequency of the reflective markers. The inertial motion capture module's IMU sensor incorporates a three-axis accelerometer and a three-axis gyroscope.

6. The system according to claim 1, characterized in that, The six-dimensional force sensor of the environmental perception module is installed on the fingertip of the robot's end effector to collect the force components of the X, Y, and Z axes and the torque components around the three axes.

7. The system according to claim 1, characterized in that, The force feedback device is a wearable glove with 12 built-in vibration motors and pressure sensors.

8. The system according to claim 1, characterized in that, The formulas used in the force control model include: F ref =K p ·(X d -X f )+K d ·(v d -v f )+K i ·∫(X d -X f )dt; Among them, F ref K represents the desired output force. p K is the proportionality coefficient. d K is the damping coefficient; i X is the integral coefficient; d X is the desired position. f For actual location; v d For the desired speed, v f This refers to the actual speed.

9. A collaborative teleoperation method for industrial humanoid robots based on multimodal motion capture fusion, used to implement a collaborative teleoperation system for industrial humanoid robots based on multimodal motion capture fusion as described in any one of claims 1 to 8, characterized in that, The method includes: Multimodal motion capture units are used to collect multimodal motion capture pose data of operators; Based on the multimodal motion capture pose data, the mode switching module outputs fused pose data according to the occlusion rate of optical markers; The environmental perception module is used to synchronously collect environmental information, workpiece images, and robot end-effector contact force information. The collaborative control unit integrates the multimodal motion capture pose data with the environmental information, and generates execution instructions through obstacle avoidance algorithms and force control models; The execution command is received by an industrial humanoid robot, and the contact force information of the robot's end effector is converted into tactile vibration feedback to the operator to obtain robot motion state feedback data. Based on the robot's motion state feedback data, the motion capture fusion weights and control parameters are dynamically adjusted to form a closed-loop control.

10. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method of claim 9.

Citation Information

Patent Citations

  • Humanoid remote control method for immersion type dual-arm robot

    CN111319026A

  • Motion capture system and method based on laser large-space positioning and optical inertia complementation

    CN112256125A

  • Multi-mode wearable humanoid robot data acquisition and remote operation system

    CN119416153A

  • Multi-modal human-computer interaction interface and behavior prediction system

    CN120066277A

  • Mixed motion capture system and method

    US20180089841A1

Cited By

  • Human body motion capture system based on multi-mode synchronous acquisition

    CN121512506A