Robot imitation learning system and method based on multi-modal somatosensory data
By using multimodal somatosensory data acquisition and point cloud diffusion strategies, the problems of low efficiency and insufficient accuracy of traditional imitation learning methods are solved, enabling robots to efficiently and accurately imitate human movements.
Patent Information
- Application Number
- CN202411597564.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2044-11-11
AI Technical Summary
Traditional imitation learning methods are cumbersome, costly, and difficult to accurately capture the details of human operation, resulting in limited learning effectiveness.
A multimodal motion data acquisition module is used, combined with inverse kinematics and forward kinematics algorithms, to relocalize and convert human motion data into point cloud data. The robot model is then trained using a point cloud diffusion strategy to generate precise motion command sequences.
It simplifies the data collection process, improves the efficiency of imitation learning, and enables robots to accurately reproduce complex human movements.
Smart Images

Figure CN119260729B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of imitation learning technology, and in particular to a robot imitation learning system and method based on multimodal somatosensory data. Background Technology
[0002] In the field of robotics, in order for robots to perform corresponding actions according to instructions, they need to learn from human actions. Traditional imitation learning methods typically rely on remotely controlling the robot's arm or hand to collect operational data. This process is not only cumbersome and costly, but also inefficient. Furthermore, this method often struggles to accurately capture the details and subtle movements of human actions, resulting in limited learning effectiveness. Summary of the Invention
[0003] This invention provides a robot imitation learning system and method based on multimodal somatosensory data, which not only simplifies the data acquisition process and improves efficiency, but also maintains high imitation learning accuracy, enabling robots to more accurately reproduce complex human movements.
[0004] A first aspect of the present invention provides a robot imitation learning system based on multimodal somatosensory data, comprising:
[0005] The somatosensory data acquisition module is used to collect multimodal somatosensory data;
[0006] The data processing and conversion module is used to reposition the modal somatosensory data so that the robot's body joint movements are consistent with human movements in space, convert the repositioned multimodal somatosensory data into point cloud data, and add the robot's perception state into the point cloud data for point cloud diffusion to obtain the robot's action strategy and the robot's action instruction sequence.
[0007] The training module is used to train the robot model using the robot's motion strategy and the instruction sequence of robot motion, so that the trained robot model can imitate human actions.
[0008] Optionally, in one embodiment of the present invention, the somatosensory data acquisition module includes a data acquisition glove and a full-body somatosensory device;
[0009] The data acquisition glove is equipped with a variety of sensors to collect motion data of various joints in the hand;
[0010] The whole-body sensing device includes a variety of sensors and cameras for collecting motion data of the whole body's movements and postures.
[0011] Optionally, in one embodiment of the present invention, the data processing and conversion module is specifically used to map full multimodal somatosensory data to the robot's posture using an inverse kinematics algorithm, and to map the fingertip position and six-degree-of-freedom palm posture of the human hand to the robot's hand.
[0012] Optionally, in one embodiment of the present invention, the data processing and conversion module is specifically used to convert the repositioned multimodal somatosensory data into point cloud data, and to integrate the virtual point cloud of the robot model into the converted point cloud data through forward kinematics.
[0013] Optionally, in one embodiment of the present invention, the data processing and conversion module is specifically used to generate an action strategy based on the processed point cloud data and the robot's perception state, convert the point cloud observation and the robot's state data into a series of preset target positions, and output the instruction sequence as the robot's action.
[0014] A second aspect of the present invention provides a robot imitation learning method based on multimodal somatosensory data, comprising the following steps:
[0015] Collect multimodal somatosensory data;
[0016] The modal somatosensory data is repositioned so that the robot's joint movements are consistent with human movements in space. The repositioned multimodal somatosensory data is converted into point cloud data, and the robot's perception state is added to the point cloud data for point cloud diffusion to obtain the robot's action strategy and the robot's action instruction sequence.
[0017] Robot models are trained using robot motion strategies and instruction sequences to mimic human actions.
[0018] Optionally, in one embodiment of the present invention, the acquisition of multimodal somatosensory data includes acquiring motion data of the joints of the hand and motion data of the whole body's movements and postures.
[0019] Optionally, in one embodiment of the present invention, repositioning the modal somatosensory data includes:
[0020] The inverse kinematics algorithm is used to map full multimodal somatosensory data into the robot's pose, and to map the fingertip position and six-DOF hand pose of the human hand into the robot's hand.
[0021] Optionally, in one embodiment of the present invention, converting the repositioned multimodal somatosensory data into point cloud data includes:
[0022] The repositioned multimodal somatosensory data is converted into point cloud data, and the virtual point cloud of the robot model is incorporated into the converted point cloud data through forward kinematics.
[0023] Optionally, in one embodiment of the present invention, the robot's perception state is added to point cloud data for point cloud diffusion to obtain the robot's action strategy and the instruction sequence for the robot's actions, including:
[0024] Based on the processed point cloud data and combined with the robot's perception state, a motion strategy is generated. The point cloud observation and the robot's state data are transformed into a series of preset target positions, and the output is a sequence of instructions for the robot's actions.
[0025] The robot imitation learning system and method based on multimodal somatosensory data in this invention simplifies the data acquisition process and improves efficiency by using multimodal somatosensory data for imitation learning, while maintaining high imitation learning accuracy, enabling the robot to more accurately reproduce complex human movements.
[0026] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0027] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0028] Figure 1 This is a schematic diagram of a robot imitation learning system based on multimodal somatosensory data according to an embodiment of the present invention;
[0029] Figure 2 This is a schematic diagram illustrating the robot imitation learning process using multimodal somatosensory data according to an embodiment of the present invention.
[0030] Figure 3 This is a schematic diagram illustrating the robot imitation application of multimodal somatosensory data according to an embodiment of the present invention;
[0031] Figure 4 This is a flowchart of a robot imitation learning method based on multimodal somatosensory data according to an embodiment of the present invention. Detailed Implementation
[0032] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0033] Figure 1 This is a schematic diagram of a robot imitation learning system based on multimodal somatosensory data according to an embodiment of the present invention.
[0034] like Figure 1 As shown, the robot imitation learning system 10 based on multimodal somatosensory data includes: a somatosensory data acquisition module 100, a data processing and conversion module 200, and a training module 300.
[0035] The system includes a motion data acquisition module 100, used to collect multimodal motion data. A data processing and conversion module 200 is used to relocalize the modal motion data to ensure that the robot's joint movements are consistent with human movements in space. It converts the relocalized multimodal motion data into point cloud data and adds the robot's perception state to the point cloud data for point cloud diffusion, obtaining the robot's motion strategy and the instruction sequence for robot movements. A training module 300 is used to train the robot model using the robot's motion strategy and the instruction sequence for robot movements, enabling the trained robot model to mimic human movements.
[0036] In one embodiment of the present invention, the motion data acquisition module includes a data acquisition glove and a full-body motion sensing device; the data acquisition glove is equipped with a variety of sensors to collect motion data of the joints of the hand; the full-body motion sensing device includes a variety of sensors and a camera to collect motion data of the whole body's movements and postures.
[0037] Specifically, the data acquisition gloves are equipped with multiple highly sensitive sensors that can accurately record the movements of each joint in the hand, including fingertip pressure and detailed angles of finger flexion. The whole-body sensing device integrates multiple sensors and cameras to comprehensively capture the wearer's full-body movements and postures, and can record in detail various movements from head to toe, such as gait, body tilt angles, and relative positions between limbs, thereby providing comprehensive data support for imitation learning.
[0038] In one embodiment of the present invention, the data processing and conversion module is specifically used to map full multimodal somatosensory data to the robot's pose using an inverse kinematics algorithm, and to map the fingertip position and six-degree-of-freedom palm pose of the human hand to the robot's hand.
[0039] The key to hand and body motion relocalization lies in adjusting and transforming the collected human hand and full-body motion data to suit the robot's operational needs. Due to structural differences between humans and robots, directly transferring this motion data to the robot is challenging. Hand movements, in particular, require precise mapping of the fingertip positions and six-degree-of-freedom (6-DoF) hand postures to the robot. Furthermore, full-body motion data is processed using inverse kinematics (IK) algorithms to ensure that the robot's joint movements are spatially consistent with human movements. This includes adjusting the optimal positions and orientations of joints such as the wrist, elbow, and shoulder to ensure continuity and coordination. This redirected motion data will serve as labels for robot training, ensuring that the robot can accurately mimic the details of human movements.
[0040] In one embodiment of the present invention, the data processing and conversion module is specifically used to convert the repositioned multimodal somatosensory data into point cloud data, and to integrate the virtual point cloud of the robot model into the converted point cloud data through forward kinematics.
[0041] In the post-processing of observation data, RGB-D images captured by a LiDAR camera are converted into point cloud data. This conversion improves the stability of the observation data and enhances its flexibility and operability. To ensure that the point cloud data is consistent with the visual representation of human hand and body movements, a virtual point cloud of the robot model is incorporated into the real observation point cloud through forward kinematics. This point cloud data is then used to train the robot's motion policy, ensuring that the captured movements are executable within the robot's operational range. Furthermore, all RGB-D frames are processed into point clouds aligned with the robot's space, and task-irrelevant elements, such as desktop points, are removed, thereby optimizing the input observation data for the robot's policy.
[0042] In one embodiment of the present invention, the data processing and conversion module is specifically used to generate an action strategy based on the processed point cloud data and the robot's perception state, convert the point cloud observation and the robot's state data into a series of preset target positions, and output the instruction sequence as the robot's action.
[0043] In embodiments of the present invention, a point cloud diffusion strategy is a key component, designed to train the robot's motion strategy to accurately mimic and execute complex human movements. This strategy utilizes processed and transformed point cloud data, combined with the robot's current self-aware state (including positional information of the hands, wrists, and other body joints), to generate a motion strategy. The point cloud diffusion strategy transforms point cloud observations and the robot's state data into a series of preset target positions, outputting a sequence of instructions for the robot's movements. This strategy not only generates coherent motion trajectories but also optimizes the smoothness and naturalness of the movements, thereby significantly improving training efficiency and execution accuracy.
[0044] The main advantage of the point cloud diffusion strategy lies in its ability to effectively handle high-dimensional multimodal motion commands. This strategy is particularly suitable for complex tasks in the field of robot learning, such as fine manipulation and multi-joint coordination, because it can: comprehensively process high-dimensional data: by integrating high-dimensional point cloud data with the robot's multimodal perception states, the strategy can generate more complex and detailed motion sequences; adapt to complex motion spaces: by adjusting the point cloud and motion trajectory through data augmentation techniques (such as random 2D translation), the point cloud diffusion strategy enhances the robot's adaptability and generalization to complex motion spaces; and improve the flexibility and accuracy of learning: the diffusion strategy model allows for flexible adjustment of parameters and states during training, thereby accurately mimicking human actions and improving the accuracy of task execution.
[0045] like Figure 2 and Figure 3 As shown, the robot imitation learning system based on multimodal somatosensory data of the present invention comprises three parts:
[0046] 1. Data Collection:
[0047] Hand and body motion capture involves operators wearing data gloves and full-body sensing devices to perform specific tasks. These devices capture the operator's hand movements and full-body posture data, recording every subtle movement and corresponding spatial position.
[0048] 2. Data processing and conversion
[0049] Hand and body motion relocalization uses inverse kinematics (IK) technology to process hand and body motion data to adapt to the robot's motion execution system;
[0050] Post-processing of the observation data involves converting RGB-D images into point clouds to stabilize the data and optimize its performance within the robot's operating space. The point cloud data is further refined to eliminate any task-irrelevant elements, providing clear, focused input data.
[0051] 3. Strategy training and execution:
[0052] Point cloud diffusion strategy training: Using the transformed and optimized point cloud data, combined with the robot's self-perceived state, the robot's action execution strategy is trained through the point cloud diffusion strategy.
[0053] In embodiments of the present invention, in order to map human hand and body posture data to the robot's joint configuration, inverse kinematics (IK) is used to calculate the robot's joint angles.
[0054] \[\textbf{q}_t=\text{IK}(\textbf{p}_h)\]
[0055] Here, \(\textbf{p}_h\) represents the position and orientation data of the human end effector, and \(\textbf{q}_t\) represents the joint angle that the robot should reach at time \(t\).
[0056] Point cloud conversion:
[0057] The data acquired through RGB-D images is converted into point clouds (\textbf{O}_t\), which are then subjected to data augmentation processing to serve as the robot's visual perception.
[0058] \[\textbf{O}_t=\text{PC}(\text{RGB-D}_t)\]
[0059] Here, the function \(\text{PC}\) processes the RGB-D image \(\text{RGB-D}_t\) captured by the camera and converts it into point cloud data.
[0060] Point cloud diffusion strategy training: Based on the point cloud data \(\textbf{O}_t\) and the robot's current joint state \(\textbf{q}_t\), the policy model \(\pi\) generates the next action sequence \(\textbf{a}_t\):
[0061] \[\textbf{a}_t=\pi(\textbf{q}_t,\textbf{O}_t)\]
[0062] \(\textbf{q}_t\) represents the robot's current joint state, \(\textbf{O}_t\) is the processed point cloud data, and \(\textbf{a}_t\) is the action command output by the policy model \(\pi\) based on the current state and observations.
[0063] Position control and motion output:
[0064] Based on the action sequence generated by the policy model \(\pi\), the robot executes the corresponding action:
[0065] \[\textbf{q}_{t+1}=\text{Execute}(\textbf{a}_t)\]
[0066] Here, the function \(\text{Execute}\) updates the robot's joint state to the next time step \(\textbf{q}_{t+1}\) based on the provided motion instruction \(\textbf{a}_t\).
[0067] The robot imitation learning system based on multimodal somatosensory data of the present invention will be described below through specific embodiments.
[0068] 1. System Configuration
[0069] Hardware equipment: inspection robot, motion data acquisition module and high-resolution camera (used to capture the operator's hand movements when pressing buttons and detailed location data).
[0070] Imitation learning module: Deep learning models deployed on robots or cloud servers to learn and replicate human fine hand movements.
[0071] Execution control system: The control unit integrated on the robot is responsible for outputting and executing actions based on the learned model.
[0072] 2. Inspection Scenario Description
[0073] The robot is deployed in a power control station equipped with multiple operating buttons to perform routine operational tasks. The objective is to mimic the actions of a human technician, precisely pressing designated buttons to execute specific power system control commands.
[0074] 3. Operating Procedures
[0075] Task Initiation: When the operation task begins, the technician demonstrates the action of pressing the button through the data glove, while the robot captures detailed data of these actions in real time through its observation module.
[0076] Data Acquisition: The motion data acquisition module captures the technician's hand movements and button position data, recording every detail of the movement, such as the bending angle of the fingers and the pressure applied.
[0077] Imitation learning training: The imitation learning module analyzes the collected hand movement data and trains the robot to imitate these movements through a point cloud diffusion strategy.
[0078] Action execution: Based on the output of the imitation learning module, the execution control system controls the robot arm and hand to perform the same button pressing action. The system dynamically adjusts the robot's hand position and force to ensure the accuracy of the pressing.
[0079] Performance feedback and optimization: If an execution error is observed, such as a button not being pressed successfully, the robot will automatically record the parameters of this operation and adjust its action strategy to optimize the subsequent execution effect.
[0080] 4. Model Application
[0081] In this power line inspection application scenario, the key to imitation learning is the high-precision replication of human hand movements. By accurately capturing and learning how technicians press buttons, the robot can autonomously perform complex control tasks, reducing direct human intervention and improving operational safety and efficiency. Furthermore, by continuously collecting operational data and feedback, the system can continuously optimize its action execution strategies, improving the robot's accuracy and adaptability.
[0082] This specific embodiment demonstrates how advanced imitation learning technology can be applied to real-world industrial scenarios, effectively solving the fine-grained operation problems that traditional automation systems struggle to handle, and significantly improving the operational capabilities and intelligence level of automation systems.
[0083] The robot imitation learning system based on multimodal somatosensory data proposed in this embodiment of the invention simplifies the data acquisition process and improves efficiency by using multimodal somatosensory data for imitation learning, while maintaining high imitation learning accuracy, enabling the robot to more accurately reproduce complex human movements.
[0084] Figure 4 This is a flowchart of a robot imitation learning method based on multimodal somatosensory data according to an embodiment of the present invention.
[0085] like Figure 4 As shown, this robot imitation learning method based on multimodal somatosensory data includes the following steps:
[0086] In step S101, multimodal somatosensory data is collected.
[0087] In step S102, the modal somatosensory data is repositioned so that the robot's joint movements are consistent with human movements in space. The repositioned multimodal somatosensory data is converted into point cloud data, and the robot's perception state is added to the point cloud data for point cloud diffusion to obtain the robot's action strategy and the robot's action instruction sequence.
[0088] In step S103, the robot model is trained using the robot's motion strategy and the sequence of instructions for robot motion, so that the trained robot model can imitate human motion.
[0089] Optionally, in one embodiment of the present invention, the acquisition of multimodal somatosensory data includes acquiring motion data of the joints of the hand and motion data of the whole body's movements and postures.
[0090] Optionally, in one embodiment of the present invention, repositioning the modal somatosensory data includes:
[0091] The inverse kinematics algorithm is used to map full multimodal somatosensory data into the robot's pose, and to map the fingertip position and six-DOF hand pose of the human hand into the robot's hand.
[0092] Optionally, in one embodiment of the present invention, converting the repositioned multimodal somatosensory data into point cloud data includes:
[0093] The repositioned multimodal somatosensory data is converted into point cloud data, and the virtual point cloud of the robot model is incorporated into the converted point cloud data through forward kinematics.
[0094] Optionally, in one embodiment of the present invention, the robot's perception state is added to point cloud data for point cloud diffusion to obtain the robot's action strategy and the instruction sequence for the robot's actions, including:
[0095] Based on the processed point cloud data and combined with the robot's perception state, a motion strategy is generated. The point cloud observation and the robot's state data are transformed into a series of preset target positions, and the output is a sequence of instructions for the robot's actions.
[0096] The robot imitation learning method based on multimodal somatosensory data proposed in this embodiment of the invention simplifies the data acquisition process and improves efficiency by using multimodal somatosensory data for imitation learning, while maintaining high imitation learning accuracy, enabling the robot to more accurately reproduce complex human movements.
[0097] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0098] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0099] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
Claims
1. A robot imitation learning system based on multimodal somatosensory data, characterized in that, include: The somatosensory data acquisition module is used to collect multimodal somatosensory data; The data processing and conversion module is used to reposition the modal somatosensory data so that the robot's body joint movements are consistent with human movements in space, convert the repositioned multimodal somatosensory data into point cloud data, and add the robot's perception state into the point cloud data for point cloud diffusion to obtain the robot's action strategy and the robot's action instruction sequence. The training module is used to train the robot model using the robot's motion strategy and the instruction sequence of the robot's motion, so that the trained robot model can imitate human actions. The somatosensory data acquisition module includes a data acquisition glove and a full-body somatosensory device; The data acquisition glove is equipped with a variety of sensors to collect motion data of various joints in the hand; The whole-body sensing device includes a variety of sensors and cameras for collecting motion data of the whole body's movements and postures.
2. The system according to claim 1, characterized in that, The data processing and conversion module is specifically used to map full multimodal somatosensory data to the robot's posture using inverse kinematics algorithms, and to map the fingertip position and six-degree-of-freedom hand posture of the human hand to the robot's hand.
3. The system according to claim 1, characterized in that, The data processing and conversion module is specifically used to convert the repositioned multimodal somatosensory data into point cloud data, and to integrate the virtual point cloud of the robot model into the converted point cloud data through forward kinematics.
4. The system according to claim 3, characterized in that, The data processing and conversion module is specifically used to generate action strategies based on the processed point cloud data and the robot's perception state, converting the point cloud observation and the robot's state data into a series of preset target positions, and outputting a sequence of instructions as robot actions.
5. The robot imitation learning method for a robot imitation learning system based on multimodal somatosensory data as described in claim 1, characterized in that, Includes the following steps: Collect multimodal somatosensory data; The modal somatosensory data is repositioned so that the robot's joint movements are consistent with human movements in space. The repositioned multimodal somatosensory data is converted into point cloud data, and the robot's perception state is added to the point cloud data for point cloud diffusion to obtain the robot's action strategy and the robot's action instruction sequence. Robot models are trained using robot motion strategies and instruction sequences to mimic human actions.
6. The method according to claim 5, characterized in that, The collection of multimodal somatosensory data includes the collection of motion data of various joints in the hand and motion data of the whole body's movements and postures.
7. The method according to claim 5, characterized in that, Relocation of the modal somatosensory data includes: The inverse kinematics algorithm is used to map full multimodal somatosensory data into the robot's pose, and to map the fingertip position and six-DOF hand pose of the human hand into the robot's hand.
8. The method according to claim 5, characterized in that, The relocalized multimodal somatosensory data is converted into point cloud data, including: The repositioned multimodal somatosensory data is converted into point cloud data, and the virtual point cloud of the robot model is incorporated into the converted point cloud data through forward kinematics.
9. The method according to claim 8, characterized in that, The robot's perceived state is added to the point cloud data for point cloud diffusion to obtain the robot's action strategy and the sequence of instructions for robot actions, including: Based on the processed point cloud data and combined with the robot's perception state, a motion strategy is generated. The point cloud observation and the robot's state data are transformed into a series of preset target positions, and the output is a sequence of instructions for the robot's actions.
Citation Information
Patent Citations
Humanoid robot action simulation method and device based on 3D human body posture estimation
CN116079727A
Humanoid robot real-time control system and method integrating electroencephalogram, myoelectricity and monocular vision
CN117532609A