Robot dexterous hand tool body deployment system based on manual data capture
By utilizing a robotic dexterous hand embodied deployment system based on human-captured data, and combining large-scale human operation data and physical constraints with a modular converter architecture, the problem of data acquisition scale limitations of robotic dexterous hands is solved, enabling efficient and stable operation strategy learning and cross-platform migration.
Patent Information
- Application Number
- CN202511805304.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-02-27
AI Technical Summary
Existing data acquisition methods for robotic dexterous hands are limited by scale, resulting in suboptimal grasping poses with poor success rates. Furthermore, cross-body pre-training is difficult, and the quality of datasets varies, making it difficult to achieve universal operation.
A robotic dexterous hand embodied deployment system based on human hand capture data is adopted, including data acquisition, simulation filtering, model training and embodied deployment modules. It is pre-trained using large-scale human finger operation data, combined with physical constraints and prior information on human operation, and a modular converter architecture is designed. A two-stage training paradigm is used for transfer learning.
In scenarios with few samples, it significantly improves the operational proficiency and accuracy of the robot's dexterous hand, reduces data collection and annotation costs, improves the efficiency and generalization ability of the learning strategy, and ensures the stability and reliability of the strategy.
Smart Images

Figure CN121580848A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robots, in particular to a robot dexterous hand grasping strategy based on human hand motion capture data. BACKGROUND
[0002] Dexterous grasping, as a basic ability of robot operation tasks, has attracted widespread attention. Compared with simple end effectors such as parallel grippers and vacuum cups, five-fingered dexterous hands have significant advantages in terms of operational flexibility, precision, and multi-task adaptability due to their human-like structure. With the deployment of robots in human environments, dexterous hands are increasingly critical as they can manipulate diverse objects and use artificial tools like humans. Therefore, accurate, robust, and versatile dexterous grasping methods have become the core of embodied intelligent interaction.
[0003] Dexterous grasping, as a core basic ability for robots to achieve complex operation tasks, has become the focus of research in the field of intelligent robots. Traditional end effectors such as parallel grippers and vacuum cups have advantages such as simple structure and low cost, but their single function and poor environmental adaptability make it difficult to meet the operational needs of diverse objects in open scenarios. In contrast, five-fingered dexterous hands, through bionics design, replicate the multi-degree-of-freedom structure and fine motor characteristics of human hands, exhibiting significant advantages in terms of operational flexibility, enabling complex grasping postures and micro-object manipulation, millimeter-level positioning control precision, and multi-task adaptability.
[0004] With the rapid development of human-robot integration environments, robots are gradually expanding from industrial closed scenarios to open human environments such as home services and medical rehabilitation. In such scenarios, dexterous hands, due to their human-like operation characteristics, can not only adaptively grasp irregularly shaped and large-sized daily objects, but also accurately operate artificial tools such as screwdrivers and scissors, becoming a key enabling technology for robots to integrate into human life and perform complex service tasks.
[0005] Early dexterous grasping research mainly relied on analytical methods, i.e., optimizing grasping poses to satisfy specific physical constraints. However, this approach has low success rates due to the large search space and high complexity of high-degree-of-freedom optimization. In contrast, data-driven methods use large-scale datasets to learn effective priors, reducing the search space and providing strong guidance for initialization. Regression-based methods directly predict grasping parameters from object inputs, often resulting in insufficient pose diversity due to pattern collapse and averaging.
[0006] Recently, large language models and visual models have attracted attention, providing inspiration for robot dexterous hand grasping: I. Special action strategies represented by ACT, diffusion strategy, and 3D diffusion strategy; II. Visual-Linguistic-Action (VLA) models represented by RT-2, OpenVLA, RD, pi, and pi0.5. These grasping strategies are paradigms for learning from massive data, which are much more effective than other methods. However, when this paradigm is applied to the long-existing field of autonomous robot operation, there is a key challenge: how to achieve the required scale of data collection.
[0007] The current mainstream data collection method for robot imitation learning is to directly control the robot hardware by human operators for demonstration. OpenX-Embodiment and DROID and other researches first launched community-level cooperation to collect hundreds of hours of robot teleoperation data. Although such data sets can effectively pre-train control strategies, physical robot operation forms a bottleneck, and this paradigm is difficult to break through the existing scale limit. Diffusion models show strong ability to model the complexity of dexterous grasping by iteratively transforming simple distributions (such as Gaussian distribution) into complex high-dimensional distributions. However, existing diffusion-based methods often generate suboptimal grasping poses, resulting in insufficient hand-object penetration or contact, and poor success rate. This problem is due to the lack of physical rule constraints. VLA models are usually pre-trained across bodies on robot data sets such as OpenX-Embodiment and AgiBotWorldColosseo. This method has two major limitations: the diversity of robot morphology and action space makes it difficult to train uniformly, and the limited size and uneven quality of existing data sets fundamentally limit the availability of data and the generalization ability required for general robot dexterous hand operation. Other researches explore learning visual representations from Internet videos or images, which have the advantage of data size, but unstructured videos lack the precise annotations necessary for learning dexterous operation.
[0008] In sharp contrast, human operation behavior constitutes a vast and easily accessible demonstration data source. The recent emergence of large-scale embodied video data sets (such as EgoDex containing 829 hours of operation videos) with fine-grained hand pose annotations provides an unprecedented opportunity to learn rich behavioral priors. Human demonstrations naturally contain object functional characteristics, operation strategies, and task decomposition patterns, which can serve as strong inductive biases for robot learning. SUMMARY
[0009] In order to overcome the technical problems in the prior art, the present application proposes a human hand-robot dexterous hand diffusion system, which uses large-scale human hand motion capture data to enhance the operation ability of the robot dexterous hand.
[0010] A robot dexterous hand embodied deployment system based on human hand motion capture data, characterized by comprising a data acquisition module, a simulation filtering module, a model training module, and an embodied deployment module, and the specific use method is: Step 1: Collect motion capture data through the data acquisition module to obtain raw motion capture data, and perform data cleaning and preprocessing on the collected data; Step 2: Perform simulation filtering on the preprocessed data using the simulation filtering module; Step 3: Train the model using the model training module to build a two-stage grasping strategy, and obtain the optimal grasping strategy through iterative calculation; Step 4: Deploy the grasping strategy onto the real robot using the embodied deployment module.
[0011] The data acquisition module collects raw motion capture data by capturing human finger motion, and then preprocesses the collected data. The simulation filtering module performs redirection and physical consistency filtering on the preprocessed data; The model building module builds a general diffusion grasping strategy for the robot's dexterous hand and performs iterative calculations to obtain the optimal diffusion grasping strategy model. The embodied deployment module enables the robot's dexterous hand to perform grasping tasks in a real environment based on the optimal general diffusion grasping strategy.
[0012] The data acquisition module collects data through a hand motion capture device, which consists of an optical motion capture camera, optical reflective markers, and an optical motion capture glove. The hand motion capture device is configured with at least 8 points, with markers placed approximately 2cm behind the web of the thumb, the middle of the thumb, the middle of the index finger, and the middle of the ring finger. The optical motion capture glove can replace or work in conjunction with the markers to improve the stability of motion capture data acquisition.
[0013] The specific steps of step 1 are as follows: Determine the hand model requirements and prepare a motion capture glove or a dot-matrix pattern (should it be the same as the hand model?) (Does the dot-matrix pattern refer to where to apply optical reflective markers?) Use the calibration box to complete the spatial calibration of the camera array, ensuring that the hand area is completely covered; keep the hand naturally open and perform static and dynamic calibration to build a skeletal model; execute hand movements, and the system captures data such as position, attitude, velocity, and acceleration in real time; correct flypoints or drifts, perform data smoothing and supplementary acquisition; export data in the corresponding format.
[0014] The specific steps of the simulation filtering in step 2 are as follows: (1) After obtaining the URDF file of the robot with a specific configuration, the URDF is converted into MJCF format through the Mujoco official toolchain. The "mj" is loaded to obtain the robot simulation configuration with specific inertia, collision and vision data. The MJCF content is modified accordingly and the controller is loaded to control the robot in the simulation environment. (2) Projecting and redirecting the human hand pose onto the robot's dexterous hand with a specific configuration: First, place the human hand and the robot's dexterous hand in the same world coordinate system to avoid scale and rotation differences. Then, establish a semantic mapping table from human fingers to robot fingers. For each finger, use forward kinematics to calculate the fingertip pose in the current state. Obtain the fingertip target pose through the key points of the human hand. Then, use inverse kinematics to calculate the joint angle of the robot's fingers. Finally, trim the calculated angle to the robot's own joint limit range (joint-limit projection) to "translate" human actions into actions that the robot can actually perform. (3) Redirection and physical consistency filtering are performed in the Mujoco simulation environment. The hand movements obtained through the data acquisition module are directed to the specific robot execution space. Using physical simulation, frames that cannot be realized in the real environment are filtered out to avoid the model learning action data that cannot be reproduced in the real environment in subsequent training. The simulation filtering module takes into account the number of human joints and the number of robot joints in a specific configuration, preserves the robot's reachable postures, ensures physical consistency, avoids robot hand clipping, object flying away, and finger collisions, and ensures the consistency of training set data with the execution data dynamics of real robot dexterous hand, laying the foundation for the embody deployment of robot dexterous hand in the fourth step (sim-to-real).
[0015] The two-stage crawling strategy in step 3 is as follows: Phase 1: Based on the prior information of human-robot actions after matching, a general diffusion strategy is trained. With the assistance of human hand joint data, 6D object data, point cloud information and LLM text, the training model in the model training module learns the general prior of "how to grasp any object", obtains the general diffusion strategy, and retains the weights of the pre-trained Transformer backbone and semantic encoder. Phase 2 involves training a modular hand adapter based on the robot's hand configuration. Without real robot teaching data, the Mujoco-based simulated robot dexterity hand learns the mapping from human action space to simulated robot joint space, thus obtaining a general strategy that can be deployed with zero samples. At this stage, the Transformer backbone and semantic encoder from Phase 1 are still used, and only the StateAdapter and ActionDecoder are trained. The StateAdapter maps the simulated robot joint angles to human action space, and the ActionDecoder maps the diffusion output to the simulated robot joint angle dimension. Before embodied deployment, domain randomization is used to improve the model's transferability.
[0016] Step 4, the embodied deployment of the robot's dexterous hand, refers to placing it in a real environment to perform grasping actions, compensating for sim-to-real errors through actual interaction, recording the sequence of actions of successful grasping, and using this data to update the modular hand adapter, ultimately achieving data backflow and cross-platform migration.
[0017] The specific steps for deploying the robot's dexterous hand in step 4 are as follows: (1) Hardware preparation: Robot dexterous hand, driver, control system, target object and data acquisition equipment are required; (2) Software preparation: Deploy the model, deploy H-RDT-G to the robot's control system, and develop the corresponding interface to convert the model's output into instructions that the robot control system can understand; (3) Online optimization: Lightweight physical guided sampling is used to compensate for sim-to-real errors without changing any parameters of H-RDT-G; (4) Data feedback: The actual robot execution results are automatically written back to stage 2 of the H-RDT-G model training. The modular hand adapter parameters are adjusted to improve the grasping accuracy with a small amount of real machine data. (5) Cross-platform migration: Replace the URDF file in the Mujoco simulation environment and retrain the modular hand adapter to achieve cross-platform migration of dexterous hands of robots with different configurations.
[0018] The data acquisition module collects raw motion capture data by capturing human finger motion, and then preprocesses the collected data. The simulation filtering module performs redirection and physical consistency filtering on the preprocessed data; The model building module builds a general diffusion grasping strategy for the robot's dexterous hand and performs iterative calculations to obtain the optimal diffusion grasping strategy model. The robot dexterous hand embodied deployment module enables the robot dexterous hand to perform grasping tasks in a real environment based on the optimal general diffusion grasping strategy.
[0019] A computer-readable storage medium having stored thereon computer-executable instructions, which, when executed by a processor, are used to implement a robotic dexterous hand embodied deployment system based on human hand-captured data.
[0020] A computer device includes a memory and a processor, the memory storing computer-executable instructions, and the processor executing the computer-executable instructions to implement a robotic dexterous hand embodied deployment system based on human hand-captured data.
[0021] The methodology of this application focuses on three core aspects: I. Data scarcity: Using data with 6D hand pose annotations, we extract rich behavioral priors containing natural operation strategies, object functional characteristics, and task decomposition patterns.
[0022] Second, cross-embodied transfer: design a modular converter architecture and equip it with a dedicated motion codec to realize knowledge transfer from human demonstrations to heterogeneous robot platforms, while retaining the learned operational knowledge.
[0023] Third, training efficiency: a two-stage training paradigm based on flow matching is adopted, which first pre-trains on large-scale human data and then performs cross-embodied fine-tuning to ensure stable and efficient policy learning throughout the entire process.
[0024] The beneficial effects of this invention are as follows: This invention achieves innovative knowledge transfer between human fingers and robotic dexterity hands through pre-training with human finger manipulation data, in terms of structural design and training mechanism. The system utilizes large-scale human finger manipulation data to enhance the robot's learning strategy. A diffusion converter architecture equipped with modular human finger-robotic dexterity hand transfer components enables efficient cross-embodied knowledge transfer. Simultaneously, physical constraints are considered in the simulation environment, revealing the core value of prior human manipulation information for efficient robotic dexterity hand learning strategies (especially in low-sample scenarios).
[0025] This invention introduces physical constraints into a simulation environment and fully integrates prior information about human operation. This allows the robot's dexterous hand to quickly master operational skills in scenarios with few samples, significantly shortening the learning time and achieving high operational proficiency and accuracy with less sample data, thus greatly improving the efficiency of the learning strategy.
[0026] This invention considers physical constraints to ensure that the strategies learned by the robot in the simulation environment can better adapt to the physical laws of the real world, while prior information from human operations provides more general guidance for the strategies. This enables the robot's dexterous hand to flexibly apply its learned knowledge when faced with different but related tasks and scenarios, demonstrating stronger generalization ability. In few-sample scenarios, even if the details of the task change, the robot's dexterous hand can still quickly adjust and complete the task based on the learned strategies.
[0027] In scenarios with few samples, traditional learning strategies often require a large amount of sample data to achieve good results. However, this method reduces the dependence on large amounts of sample data by utilizing physical constraints and prior information from human operations. The robot dexterous hand can quickly learn effective operation strategies based on limited sample data, reducing the cost and time of data collection and annotation.
[0028] The physical constraints of this invention provide reasonable boundaries and rules for the robot's operation, avoiding unreasonable actions and strategies during the learning process, while prior information from human operation provides effective operation patterns that have been verified in practice. The combination of the two makes the strategies learned by the robot's dexterous hand more stable and reliable, maintaining good performance even in scenarios with few samples, and reducing the probability of strategy failure and anomalies. Attached Figure Description
[0029] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the process of the present invention; Figure 2 This is a schematic diagram of a human hand motion capture data acquisition device according to the present invention; Figure 3 This is a schematic diagram of the training logic of the H-RDT-G model of the present invention; Figure 4 This is a schematic diagram of the H-RDT-G model stage 2 Adapter training of the present invention; Figure 5 This is a schematic diagram illustrating the deployment of the human finger into the robot's dexterous hand according to the present invention. Detailed Implementation
[0030] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. The components of the embodiments of the present application described and shown in the accompanying drawings can typically be arranged and designed in various different configurations.
[0031] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application. Example 1
[0032] A robot dexterous hand embodied deployment system based on human hand-captured data is characterized by comprising a data acquisition module, a simulation filtering module, a model training module, and an embodied deployment module, wherein the data acquisition module, simulation filtering module, model training module, and embodied deployment module are connected. Its specific usage method is as follows: Step 1: Collect motion capture data through the data acquisition module to obtain raw motion capture data, and perform data cleaning and preprocessing on the collected data; Motion capture data is obtained by wearing a motion capture suit on the human hand and attaching appropriate markers to the fingers and target object based on the specific data acquisition situation. The camera captures the information from these markers to obtain 6D data of the human finger and the object. The 6D data includes 3D position data of the joint relative to the global coordinate system and the parent joint coordinate system, as well as 4D quaternion pose data. In addition, it also includes velocity, angular velocity, and acceleration data, along with corresponding timestamps.
[0033] The raw motion capture data is converted into a format that the model can directly use. To address the issue of missing frames, cubic spline interpolation can be used to fill in the missing frames. If a segment of missing frames accounts for more than 10%, this segment of data should be removed. To address the coordinate system drift issue, static object markers are used for correction, and the system can be recalibrated every 10 seconds.
[0034] Step 2: Perform simulation filtering on the preprocessed data using the simulation filtering module; Redirection and physical consistency filtering are performed in the Mujoco simulation environment. Human-captured actions are redirected to the robot's specific executable space. Frames that are not feasible in the real environment are filtered out through physical simulation. Furthermore, human finger movements are converted into robot dexterity hand movements to prevent the model from learning action data that cannot be reproduced in the real environment during subsequent training. Specifically, the number of joints in the human model and the number of joints in the robot with a specific configuration are considered, preserving the robot's reachable poses.
[0035] Ensure physical consistency to prevent the robotic hand from passing through the mold, objects from flying away, and fingers from colliding. Ensure dynamic consistency between the training set data and the execution data of the real robotic dexterous hand, laying the foundation for the embody deployment of the robotic dexterous hand in the fourth step (sim-to-real).
[0036] Step 3: Train the model using the model training module to build a two-stage grasping strategy, and obtain the optimal grasping strategy through iterative calculation; H-RDT-G model training. A general model is pre-trained using human hand motion capture data, then fine-tuned using simulated robot data and subsequent data feedback. Through a modular hand adapter, cross-platform, low-sample, high-success-rate grasping is achieved. The H-RDT-G model includes a two-stage training logic: Phase 1: Pre-training with motion capture data. Using human hand joint data, 6D object data, point cloud information, and LLM text assistance, the model learns the general prior of "how to grasp any object", obtains the general diffusion strategy, and retains the weights of the Transformer backbone and semantic encoder after pre-training.
[0037] Phase 2: Simulation Form Adaptation. Without real robot teaching data, the simulation robot in Mujoco learns the mapping from the human hand motion space to the simulation robot's joint space, thus obtaining a general strategy that can be deployed with zero samples. At this stage, the Transformer backbone and semantic encoder from Phase 1 are still used, and only the StateAdapter and ActionDecoder are trained. The StateAdapter maps the simulation robot's joint angles to the human motion space, and the ActionDecoder maps the diffusion output to the simulation robot's joint angle dimensions. Domain randomization is used to improve the model's transferability before embodied deployment.
[0038] A modular hand adapter, introduced in Phase 2 fine-tuning of model training, is used to adapt to the specific physical characteristics and joint space of the robotic hand. It ensures that the output of the pre-trained model can be correctly adapted to the specific robotic hand. In the absence of robot teaching data, simulated data can be generated based on URDF (Ultra-Ultra-Fractional Data Generation). Virtual robotic hand motion data can be generated in a simulation environment using URDF files, and then the adapter can be trained using this data. Alternatively, training can be performed through self-supervised learning, incorporating a redirection error penalty during the diffusion sampling phase, allowing the model to self-supervisedly learn how to map the joint angles of the human hand to the joint angles of the robotic hand.
[0039] Step 4: Deploy the grasping strategy onto the real robot using the embodied deployment module.
[0040] The trained model is then applied to the actual dexterous hand of a robot to perform real object grasping tasks.
[0041] The specific method for step four is as follows: Hardware preparation: Robotic dexterity hand, actuator, control system and target object are required.
[0042] Software preparation: Model deployment is required. H-RDT-G needs to be deployed to the robot's control system, and corresponding interfaces need to be developed to convert the model's output into instructions that the robot control system can understand.
[0043] Online optimization: Compensates for sim-to-real errors using lightweight physical guided sampling without altering any parameters of H-RDT-G.
[0044] Data feedback: The actual robot execution results are automatically written back to stage 2 of the H-RDT-G model training. Only the parameters of the modular hand adapter are adjusted, and the grasping accuracy can be improved with a small amount of real machine data.
[0045] Cross-platform migration: Replace the URDF file in the Mujoco simulation environment and retrain the modular hand adapter.
[0046] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0047] The above description is merely an embodiment of this application and is not intended to limit the application. Various modifications and variations can be made to this application by those skilled in the art. All modifications and variations within the spirit and principles of this application are subject to interpretation. Any modifications, equivalent substitutions, improvements, etc., made shall be included within the scope of the claims of this application.
[0048] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. The various components mentioned in this invention are common technologies in the existing field. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A robot dexterous hand body deployment system based on human manual data capture, characterized by, The system comprises a data acquisition module, a simulation filtering module, a model training module and a body deployment module, and the data acquisition module, the simulation filtering module, the model training module and the body deployment module are connected. The specific use method is: Step 1: acquiring motion capture data through the data acquisition module to obtain original motion capture data, and performing data cleaning and preprocessing on the acquired data; Step 2: performing simulation filtering on the preprocessed data through the simulation filtering module; Step 3: training through the model training module, building a two-stage grabbing strategy, and obtaining the optimal grabbing strategy through iterative calculation; Step 4: deploying the grabbing strategy to a real robot through the body deployment module.
2. A robotic hand exoskeleton system based on human hand motion capture data according to claim 1, wherein, The data acquisition module acquires human finger motion capture data to obtain original motion capture data, and preprocesses the acquired data; The simulation filtering module performs redirection and physical consistency filtering on the preprocessed data; The model building module builds a general diffusion grabbing strategy for a robot dexterous hand, and obtains an optimal diffusion grabbing strategy model through iterative calculation; The body deployment module executes a grabbing task for the robot dexterous hand in a real environment according to the optimal general diffusion grabbing strategy.
3. A robotic hand exoskeleton system based on human hand motion capture data according to claim 1, wherein, The data acquisition module acquires data through a hand motion capture device, and the hand motion capture device comprises an optical motion capture camera, optical reflective marker points and an optical motion capture glove; the hand motion capture device is provided with at least 8 points of configuration, and the marker points are placed about 2 cm behind the web space, the middle segment of the thumb, the middle segment of the index finger and the middle segment of the ring finger; The optical motion capture glove can replace the marker points or work cooperatively with the marker points to improve the stability of motion capture data acquisition.
4. The robotic hand anthropomorphization system based on human hand motion capture data of claim 1, wherein, The specific steps of step 1 are: Determine the hand model requirements and prepare the motion capture glove or marker point scheme (whether it should be and) (whether the marker point scheme refers to where to perform optical reflective marker point pasting); Use the calibration frame to complete camera array space calibration to ensure that the hand region is completely covered; keep the hand naturally open, perform static and dynamic calibration to establish a bone model; perform hand movements, and the system captures position, posture, speed, acceleration and other data in real time; repair flying points or drift, perform data smoothing processing and supplementary acquisition; export data in the corresponding format.
5. The anthropomorphic hand body deployment system for a data-capturing robot hand according to claim 1, wherein The specific steps of step 2 simulation filtering are: (1) After obtaining the specific configuration of the robot URDF file, convert the URDF into MJCF format through the Mujoco official tool chain, load "mj" to obtain a robot simulation configuration with specific inertia, collision and vision data, modify the MJCF content accordingly and load the controller to regulate the robot in the simulation environment; (2) Project human hand pose to specific configuration of robot hand: first, put human hand and robot hand in the same world coordinate system to avoid scale difference and rotation difference, then establish semantic mapping table of human finger to robot finger, for each finger single chain, use robot forward kinematics to get the current state of the fingertip pose; get the fingertip target pose through the human hand key points, and then use inverse kinematics to solve the robot finger joint angle, finally clip the calculated angle to the joint limit range of the robot itself (joint-limit projection), so as to "translate" human action into the action that the robot can actually make; (3) Redirect and filter physical consistency in Mujoco simulation environment, direct the human hand action obtained through the data acquisition module to the specific robot executable space, use physical simulation as a means to filter out those frames that cannot be realized in the real environment, to avoid the model learning the action data that cannot be reproduced in the real environment in the subsequent training.
6. A robotic hand exoskeleton system based on human hand motion capture data according to claim 1, wherein, The simulation filtering module acts on the number of human joints and the number of robot joints of specific configuration, retains the reachable pose of the robot, ensures the physical consistency, avoids the robot hand from being out of shape, the object from flying away and the finger from colliding, ensures the consistency of the dynamics of the training set data and the execution data of the real robot hand, and lays a foundation for the sim-to-real of the robot hand in the fourth step; The two-stage grasping strategy of step 3 is specifically: Stage 1, train a general diffusion strategy according to the prior information of matched human-robot action, with the help of human hand joint data, object 6D data, point cloud information and LLM text, the training model in the model training module learns the general priori of "how to grasp any object", gets the general diffusion strategy, and retains the weights of the pre-trained Transformer backbone and semantic encoder; Stage 2, train modular hand Adapter according to robot hand configuration, under the premise of no real robot demonstration data, use the simulation robot hand in Mujoco to complete the mapping learning from human action space to simulation robot joint space, so as to get a general strategy that can be deployed with zero samples; At this time, the Transformer backbone and semantic encoder of stage 1 are still used, only the StateAdapter and ActionDecoder are trained, the StateAdapter maps the simulation robot joint angle to the human action space, and the ActionDecoder maps the diffusion output to the simulation robot joint angle dimension, and the domain randomization is used to improve the model transfer ability before embodiment deployment.
7. The system of claim 1, wherein the system is a hand-based data-capturing robot dexterous hand exoskeleton system. The embodiment deployment of the robot hand in step 4 refers to placing it in a real environment to perform grasping action, compensating for the sim-to-real error through actual interaction, recording the action sequence of successful grasping, and updating the modular hand Adapter using the data, finally realizing data backflow and cross-platform migration; The specific steps of the robot hand embodiment deployment in step 4 are: (1) Hardware preparation: a robot dexterous hand, a driver, a control system, a target object, and data acquisition equipment are required; (2) Software preparation: model deployment is performed, H-RDT-G is deployed into the control system of the robot, and a corresponding interface is developed to convert the output of the model into instructions that can be understood by the robot control system; (3) Online optimization: using lightweight physically guided sampling, sim-to-real error is compensated without changing any parameters of H-RDT-G; (4) Data backflow: the real robot execution result is automatically written back to the H-RDT-G model training stage 2, the modular hand adapter parameters are adjusted, and a small amount of real robot data can improve the grasping precision; (5) Cross-platform migration: replacing the URDF file in the Mujoco simulation environment, retraining the modular hand adapter can realize cross-platform migration of different configurations of robot dexterous hands.
8. The robotic hand anthropomorphization system based on human hand motion capture data of claim 1, wherein, The data acquisition module acquires human finger motion capture data to obtain original motion capture data, and pre-processes the acquired data; The simulation filtering module performs redirection and physical consistency filtering on the pre-processed data; The model building module builds a general diffusion grasping strategy for a robot dexterous hand, and iteratively calculates to obtain an optimal diffusion grasping strategy model; The robot dexterous hand body deployment module executes a grasping task in a real environment according to the optimal general diffusion grasping strategy.
9. A computer readable storage medium having stored thereon computer- executable instructions, characterized in that, The computer executable instructions, when executed by the processor, implement the method of any one of claims 1-6.
10. A computer device comprising a memory and a processor, wherein computer executable instructions are stored on the memory, and the computer device is characterized in that, The processor executes the computer executable instructions to implement the method of any one of claims 1-6.