Dynamic dexterous grabbing benchmark test system in man-machine handover

By proposing a benchmark test system for dynamic and flexible grabbing in the human-computer handover task, including the DexH2R data set and DynamicGrasp method, the problem of lack of high-quality data and effective solutions in the existing technology is solved, and an efficient and reliable human-computer object handover task is achieved.

CN120170740APending Publication Date: 2025-06-20SHANGHAI TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510446943.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art lacks high-quality real-world data and effective solutions to realize human-computer object handover tasks, especially in dynamic interactive scenarios and complex handover tasks.

Method used

A benchmark test system for dynamic and dexterous grabbing in human-computer handover is proposed, including the DexH2R and DynamicGrasp method of the dexterous hand-over data set of human-computer object handover. The DynamicGrasp method divides the task into three stages: grasping pose preparation, proximity motion generation and target pose alignment. It uses technology such as conditional variational autoencoder model and diffusion strategy to generate stable and practical grasping poses and motion trajectories.

Benefits of technology

It realizes a high-quality, realistic collection of smart H2R data sets, which can effectively process multimodal perceptual data and generate human-like dynamic grabbing behaviors, improving the performance and reliability of human-computer object handover tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120170740A_ABST
    Figure CN120170740A_ABST
Patent Text Reader

Abstract

The technical scheme of the invention discloses a benchmark test system for dynamic and flexible grabbing in man-machine handover. The overall architecture of the method comprises a high-quality data set, a baseline method comparison method in a solution and a task performance evaluation method, an excellent object handover generation result is shown on the data set, multi-modal sensing data can be effectively processed, and human-like dynamic grabbing behaviors can be generated. And a better handover generation result is provided for the data in the data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a benchmarking framework for studying human-robot object handover tasks. Background Art

[0002] Enabling a robot to naturally receive an object passed by a human is a fundamental operation in the field of human-robot interaction and cooperation. The realization of its safety, human-like behavior, and practicality relies on high-quality real-world data and effective solutions.

[0003] However, the following limitations exist in the prior art:

[0004] 1. Existing datasets are mostly based on static grasping scenarios and lack the dynamic interaction characteristics required for human-to-robot handover (H2R), and thus cannot meet the requirements of dynamic handover tasks;

[0005] 2. Some studies have attempted to put static grasping data into a simulation environment to generate H2R data, but there are significant differences between the givers, receivers, and visual data generated by the simulation and the real world, making it difficult to directly apply them to actual scenarios;

[0006] 3. Current dynamic H2R grasping research mostly focuses on parallel-jaw robots, and its methods are difficult to be migrated to five-fingered dexterous hands, which limits their application in complex handover tasks. Summary of the Invention

[0007] The object of the present invention is to provide a high-quality, real-world collected dexterous hand H2R dataset and a feasible solution to promote the development of human-robot object handover tasks.

[0008] To achieve the above object, the technical solution of the present invention discloses a benchmarking system for dynamic dexterous grasping in human-robot handover, which is characterized by including:

[0009] A dexterous hand human-robot object handover dataset DexH2R, which contains multi-modal perception data, real human-robot interaction data, and human-like dynamic grasping behaviors;

[0010] A DynamicGrasp method implementation module, wherein the DynamicGrasp method divides the task into three stages according to human behavior habits: grasping posture preparation, approaching motion generation, and target posture alignment, where:

[0011] In the grasping posture preparation stage:

[0012] Use the conditional variational autoencoder model as the backbone network of the grasping pose generation model. Extract features from the normalized object point cloud and the ShadowHand dexterous hand point cloud through the PointNet network to generate grasping poses that completely cover the object surface.

[0013] In the approaching motion generation stage, it includes:

[0014] Predict the subsequent hand states based on the current object point cloud features extracted by the PointNet network, the current predicted dexterous hand grasping target pose output by the grasping pose generation model, and the historical observation data.

[0015] Adopt a method based on the diffusion strategy, and through the noise addition and denoising processes of the diffusion strategy, iteratively predict the subsequent hand states.

[0016] In the target pose alignment stage:

[0017] When the distance between the current global displacement of the dexterous hand and the global displacement of the target grasping pose is less than the set threshold, perform linear interpolation on the current dexterous hand pose and the target pose, and through dynamic iteration, gradually align the generated dexterous hand trajectory to the reliable grasping target pose.

[0018] The task performance evaluation module is used to evaluate the task performance based on the grasping pose generation result evaluation index and the approaching motion stage result evaluation index.

[0019] Preferably, the following steps are used to construct the dexterous hand human-robot object handover dataset DexH2R:

[0020] Build a vision system using 18 cameras to provide RGB information, depth information, and accurate hand-object annotation data, and the data of each modality is accurately calibrated and aligned.

[0021] During data collection, the action trajectories of the human hand passing the object are dynamic and diverse, ensuring the dynamics and authenticity of the interaction object and the interaction process.

[0022] Use the teleoperation system to obtain the real data of the humanoid dexterous hand in the H2R task to ensure the practicality and transferability of the data.

[0023] Preferably, first pre-train the grasping pose generation model on a large amount of simulation data, and fine-tune the grasping pose generation model using the grasping data in the real world.

[0024] Preferably, in the grasping posture preparation stage, in the simulation environment Isaac Gym, physical filters are used to screen out grasping results that are successful in all six directions, and the hand data of the transmitter is combined as a geometric filter to refine the posture candidates generated by the grasping posture generation model, so as to ensure the generation of stable and practical grasping postures.

[0025] Preferably, the diffusion strategy-based method includes the DP method and the DP3 method.

[0026] Preferably, the evaluation metrics for the grasping posture generation results include:

[0027] One-direction success rate: The success rate under a single random force applied to the object;

[0028] Six-direction success rate: The success rate under random forces in six directions;

[0029] Penetration distance: Measure the penetration distance of the hand into the object surface;

[0030] Diversity: The standard deviation of the pose parameters when grasping successfully.

[0031] Preferably, the evaluation metrics for the approaching motion stage results include:

[0032] Success rate: The percentage of trajectories that reach the target grasping posture;

[0033] Motion trajectory length: The total length of the generated trajectory;

[0034] Total number of inference frames of the model: The average number of frames required to infer the trajectory;

[0035] Penetration distance: The maximum depth of the hand inserted into the object during the approaching stage;

[0036] Penetration frames: The number of frames during the approaching stage when the penetration force exceeds the safety threshold;

[0037] Safety rate: The percentage of trajectories with an average penetration rate lower than the threshold.

[0038] The overall architecture of the present invention includes a high-quality dataset, a comparison of baseline methods in the solution method, and a task performance evaluation method. Compared with the existing technical solutions, it has the following beneficial effects:

[0039] 1. Excellent performance on the dataset: The present invention shows excellent object handover generation results on the dataset, can effectively process multi-modal perception data and generate human-like dynamic grasping behaviors. It has relatively good handover generation results for the data in the dataset.

[0040] 2. Real machine experiment implementation: a) Object grasping and trajectory generation: The Giver picks up an object from the tabletop and generates a random dynamic trajectory. The vision system performs 6D pose estimation on the templated object and applies the estimated pose to the pre-generated dexterous hand target grasping action to screen for suitable grasping postures.

[0041] b) Approach motion generation: Input the historical observation data and the current object point cloud into the approach motion generation module (MotionNet is used here) to generate a stable and smooth motion trajectory.

[0042] c) Approach motion generation: Input the historical observation data and the current object point cloud into the approach motion generation module to generate a stable and smooth motion trajectory. Brief Description of the Drawings

[0043] Figure 1 is the flowchart of the present invention;

[0044] Figure 2 is an example illustration. Detailed Implementation Manner

[0045] The following further elaborates the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0046] The overall architecture of a benchmark test system for dynamic dexterous grasping in human-robot handover disclosed by the present invention includes the following three parts:

[0047] I) Dataset: The present invention proposes the first real-world collected dexterous hand human-robot object handover (H2R) dataset DexH2R, which contains rich multi-modal perception data, real human-robot interaction data, and human-like dynamic grasping behaviors. Specifically:

[0048] 1. Build a vision system using 18 cameras to provide high-quality RGB information, depth information, and accurate hand-object annotation data. Each modal data is accurately calibrated and aligned.

[0049] 2. During data collection, the action trajectories of the human hand passing the object are dynamic and diverse, ensuring the dynamics and authenticity of the interaction object and the interaction process.

[0050] 3. Use a teleoperation system to obtain real data of the human-like dexterous hand in the H2R task to ensure the practicality and transferability of the data.

[0051] (II) Solution: An efficient and feasible solution, DynamicGrasp, is proposed for the human-robot object handover task of a dexterous robot hand. This method divides the task into three stages according to human behavior habits: grasping pose preparation, approaching motion generation, and target pose alignment.

[0052] 1. Grasping pose preparation: The conditional variational autoencoder (cVAE) model is used as the backbone network. PointNet is used to extract features from the normalized object point cloud and the ShadowHand dexterous hand point cloud, and the extracted features are input into the cVAE model to generate a large number of grasping poses that completely cover the object surface. To ensure the generalization ability of the model, the present invention first pre-trains the grasping pose generation model on a large-scale simulation data and fine-tunes it using real-world grasping data. To ensure the stability and safety of the dexterous hand grasping pose, in the simulation environment Isaac Gym, a physical filter is used to screen out the grasping results that are successful in all six directions, and the human hand data of the giver is combined as a geometric filter to refine the generated pose candidates to ensure the generation of stable and practical grasping poses.

[0053] 2. Approaching motion generation: Two paradigms, the autoregressive method and the diffusion policy (Diffusion Policy), are explored. By inputting historical observation data such as RGB images, depth maps, object point clouds, and dexterous hand state parameters, the future approaching action trajectory is generated. Specifically:

[0054] a) The autoregressive-based MotionNet network: Input the current object point cloud features extracted by PointNet, the current predicted dexterous hand grasping target pose (goal pose), and historical observation data to predict the subsequent hand state (handstate).

[0055] b) The methods based on the diffusion policy, including two methods, DP and DP3. Among them, DP requires RGB data, while DP3 only requires point cloud information. To adapt to the current task, depth data is additionally added to the input of the DP method. Through the noise addition and denoising process of the diffusion policy, the subsequent hand state is iteratively predicted.

[0056] 3. Target pose alignment: A simple and effective linear interpolation method is adopted in this stage. When the distance between the current global displacement of the dexterous hand and the global displacement of the target grasping pose is less than the set threshold, linear interpolation is performed on the current dexterous hand pose and the target pose. Through dynamic iteration, the generated dexterous hand trajectory is gradually aligned to a reliable grasping target pose, thereby ensuring an accurate and physically reasonable grasp. While maintaining accuracy, this method strictly adheres to environmental constraints and finally achieves a robust and reliable grasping result.

[0057] (III) Task Performance Evaluation: A comprehensive set of evaluation metrics are proposed for attributes such as safety, accuracy, and reliability, specifically including:

[0058] 1. Evaluation Metrics for Grasping Pose Generation Results:

[0059] a) Success Rate in One Direction: The success rate under a single random force applied to the object;

[0060] b) Success Rate in Six Directions: The success rate under random forces in six directions;

[0061] c) Penetration Distance: Measure the penetration distance of the hand into the object surface;

[0062] d) Diversity: The standard deviation of pose parameters during successful grasping.

[0063] 2. Evaluation Metrics for the Approach Motion Phase Results:

[0064] a) Success Rate: The percentage of trajectories that reach the target grasping pose;

[0065] b) Length of Motion Trajectory: The total length of the generated trajectory (ideal range: 1 - 2 meters);

[0066] c) Total Inference Frames of the Model: The average number of frames required to infer the trajectory;

[0067] d) Penetration Distance: The maximum depth that the hand inserts into the object during the approach phase;

[0068] e) Number of Penetration Frames: The number of frames during the approach phase when the penetration force exceeds the safety threshold;

[0069] f) Safety Rate: The percentage of trajectories with an average penetration rate lower than the threshold.

Claims

1. A benchmark test system for dynamic and dexterous grasping in human-machine interaction, characterized in that: include: DexH2R, a dataset for dexterous human-machine object handover, contains multimodal perception data, real human-machine interaction data, and human-like dynamic grasping behaviors; DynamicGrasp method implementation module, the DynamicGrasp method divides the task into three stages according to human behavior habits: grasping posture preparation, approach motion generation and target posture alignment, where: In the grasping posture preparation phase: The conditional variational autoencoder model is used as the backbone network of the grasping posture generation model. The PointNet network is used to extract features from the normalized object point cloud and the ShadowHand dexterous hand point cloud to generate a grasping posture that completely covers the object surface. In the approach motion generation phase, including: The point cloud features of the current object extracted by the PointNet network, the current predicted dexterous hand grasping target posture output by the grasping posture generation model, and the historical observation data are used to predict the subsequent hand state; A diffusion strategy-based method is used to iteratively predict the subsequent hand state through the denoising and denoising process of the diffusion strategy; During the target pose alignment phase: When the distance between the current global displacement of the dexterous hand and the global displacement of the target grasping posture is less than the set threshold, the current dexterous hand posture and the target posture are linearly interpolated, and the generated dexterous hand trajectory is gradually aligned to the reliable grasping target posture through dynamic iteration; The task performance evaluation module is used to evaluate the task performance based on the grasping posture generation result evaluation index and the approach motion stage result evaluation index.

2. A benchmark test system for dynamic and dexterous grasping in human-machine interaction as claimed in claim 1, characterized in that: The following steps are used to construct the dexterous hand human-machine object handover dataset DexH2R: A visual system is built using 18 cameras to provide RGB information, depth information, and accurate hand and object annotation data. Each modality data is precisely calibrated and aligned. During data collection, the motion trajectories of human hands passing objects are dynamic and diverse, ensuring the dynamics and authenticity of the interactive objects and the interactive process; Use a teleoperation system to obtain real data of the humanoid dexterous hand in the H2R mission to ensure the practicality and transferability of the data.

3. A benchmark test system for dynamic dexterous grasping in human-machine interface as claimed in claim 1, characterized in that: First, the grasping posture generation model is pre-trained on large-scale simulation data, and then the grasping posture generation model is fine-tuned using real-world grasping data.

4. A benchmark test system for dynamic and dexterous grasping in human-machine interaction as claimed in claim 1, characterized in that: In the grasping posture preparation stage, a physical filter is used in the simulation environment Isaac Gym to screen out successful grasping results in six directions, and the hand data of the handover is used as a geometric filter to refine the posture candidates generated by the grasping posture generation model to ensure the generation of a stable and practical grasping posture.

5. A benchmark test system for dynamic dexterous grasping in human-machine interaction as claimed in claim 1, characterized in that: The methods based on diffusion strategy include DP method and DP3 method.

6. A benchmark test system for dynamic and dexterous grasping in human-machine interaction as claimed in claim 1, characterized in that: The evaluation indexes of the grasping posture generation result include: One-direction success rate: the success rate under a single random force applied to the object; Six-direction success rate: success rate under random forces in six directions; Penetration distance: measures the penetration distance of the hand on the surface of the object; Diversity: The standard deviation of posture parameters during successful grasps.

7. A benchmark test system for dynamic and dexterous grasping in human-machine interaction as claimed in claim 1, characterized in that: The approach movement stage result evaluation indicators include: Success rate: the percentage of trajectories that achieve the target grasping posture; Motion trajectory length: the total length of the generated trajectory; Total model inference frames: the average number of frames required to infer the trajectory; Penetration distance: the maximum depth of the hand inserted into the object during the approach phase; Penetration frames: the number of frames in which the penetration exceeds the safety threshold during the approach phase; Safety rate: The percentage of trajectories with an average penetration rate below the threshold.