Functional grabbing synthesis method for dexterous hand robot and training method of corresponding prediction model
By generating 3D asset data and identifying the availability parts of objects using a pre-trained feature extractor, and combining axis alignment and availability grasping strategies, the problem of insufficient grasping diversity of robotic hands is solved, enabling dexterous hands to grasp objects of different scales efficiently and accurately.
Patent Information
- Application Number
- CN202511149105.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-18
AI Technical Summary
In existing technologies, robotic hand grasping methods rely on demonstrations of human hands, resulting in datasets that are biased towards large objects and lack large-scale datasets. This makes it difficult to achieve effective grasping of objects of different sizes by dexterous hands, especially small objects, and also results in insufficient grasping diversity.
Three-dimensional asset data is generated by acquiring two-dimensional image data. A pre-trained feature extractor is used to identify the availability parts of objects. By combining axis alignment and availability grasping strategies, the grasping configuration is optimized, and accurate grasping is achieved using a pre-set encoder and decoder.
It enables training on large-scale datasets with diverse scales, improving the efficiency and generalization ability of dexterous hand functional grasping, and enabling precise grasping of objects of different scales, reducing manual annotation work.
Smart Images

Figure CN120977014A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the field of humanoid robots and dexterous hand robots, and more specifically, to a functional grasping synthesis / prediction method for dexterous hand robots and a training method for the corresponding prediction model. Background Technology
[0002] In related fields, dexterous functional grasping typically refers to a robot's ability to grasp and manipulate objects through flexible hand operations to complete specific tasks (such as grasping, placing, and manipulating). These tasks require robots to possess advanced perception, decision-making, and execution capabilities. The ability to perform dexterous functional grasping is fundamental to robot manipulation, supporting a wide range of applications from domestic services to precision industrial tasks. Achieving this capability in robotic hands, especially those with high degrees of freedom, remains a significant challenge.
[0003] In existing technologies, deep learning methods have been used to improve the generalization ability of robotic hand grasping strategies. However, existing methods typically rely on demonstrations of human hands, which are labor-intensive and limited in number. Moreover, existing methods cause the robotic hand's grasping to inherit biases in human hand size and kinematics, resulting in datasets biased towards larger objects and imprecise functional grasping. Therefore, an improved method is needed to enhance the functional grasping ability of robotic hands. Summary of the Invention
[0004] The embodiments of this disclosure provide a method for predicting the functional grasping of a dexterous hand robot and a training method for the corresponding prediction model, with the aim of achieving efficient functional grasping prediction of a dexterous hand.
[0005] In one general aspect, a functional grasping synthesis method for a dexterous hand robot is provided, the method comprising: acquiring two-dimensional image data of a first object of a specific category; performing three-dimensional asset generation processing on the two-dimensional image data of the first object to obtain three-dimensional asset data of the first object; inputting point cloud data of the first object corresponding to the three-dimensional asset data of the first object into a pre-trained first feature extractor for recognizing corresponding parts between different instances of objects of the same category to obtain first point cloud feature data of the corresponding first object; performing preset template matching processing associated with the specific category on the point cloud feature data of the first object to obtain availability data of the first object; and performing initialization and functional grasping optimization processing on the point cloud data of the first object corresponding to the availability data of the first object using an axis-aligned grasping initialization strategy and an availability-based functional grasping synthesis strategy to obtain grasping configuration data of the first object.
[0006] Optionally, the step of performing three-dimensional asset generation processing on the two-dimensional image data of the first object to obtain the three-dimensional asset data of the first object may include: performing target recognition and image segmentation processing on the two-dimensional image data of the first object to obtain the first processed image data of the first object; and inputting the first processed image data of the first object into a pre-trained three-dimensional generation model to obtain the three-dimensional asset data of the first object.
[0007] Optionally, a preset template associated with the specific category can be obtained by: acquiring instance data of at least one object for each of multiple categories; performing template matching annotation processing on the availability portion for any object in each category to obtain object availability data with availability portion tags for the object; using the object availability data with availability portion tags for the object in each category, obtaining object availability data with availability portion tags for other objects in each category through automatic feature matching processing, thereby obtaining a preset template associated with the specific category, wherein the preset template associated with the specific category is included in the object availability data with availability portion tags for specific objects of the specific category in the multiple categories.
[0008] Optionally, the step of performing initialization and functional grasping optimization processing on the first object point cloud data corresponding to the first object's availability data using an axis-aligned grasping initialization strategy and an availability-based functional grasping synthesis strategy to obtain the first object's grasping configuration data may include: performing axis-aligned grasping initialization processing on the first object point cloud data corresponding to the first object's availability data using an axis-aligned grasping initialization strategy to obtain the first object's initial grasping data; and performing availability-based functional grasping optimization processing on the first object's initial grasping data using an availability-based functional grasping synthesis strategy to obtain grasping configuration data that meets preset constraints, which serves as the first object's grasping configuration data.
[0009] Optionally, the preset constraints may include at least one of the following: functional part distance loss constraints, grasping part distance loss constraints, force closure loss constraints, collision loss constraints between the dexterous hand and the object, and dexterous hand self-collision loss constraints.
[0010] In another general aspect, a training method for a functional grasping prediction model for a dexterous hand robot is provided. The training method includes: for each first object of at least one object of a specific category, performing the following processes: acquiring grasping configuration data for the first object; inputting the grasping configuration data for the first object into a pre-trained second feature extractor for object availability recognition to obtain grasping feature data for the first object; inputting point cloud data of the first object into a pre-trained third feature extractor for extracting preset details to obtain second point cloud feature data for the first object; inputting the point cloud data of the first object and the second point cloud feature data of the first object into a pre-trained fourth feature extractor for extracting preset non-details to obtain third point cloud feature data for the first object; concatenating the grasping feature data of the first object and the third point cloud feature data of the first object and inputting the concatenated data into a preset encoder and a preset decoder to obtain a predicted grasping configuration for both the dexterous hand robot and the first object, so as to control the dexterous hand robot to grasp the first object through the predicted grasping configuration.
[0011] Optionally, the second feature extractor for object availability identification, the third feature extractor for extracting preset details, the fourth feature extractor for extracting preset non-details, the preset encoder, and the preset decoder can constitute a functional grasping prediction model for a dexterous hand robot.
[0012] Optionally, the training method may further include: adjusting model parameters based on the predicted grasping configuration and surface point cloud data included in the 3D asset data associated with the at least one object to obtain a trained functional grasping prediction model for a dexterous hand robot. The step of obtaining the trained functional grasping prediction model for the dexterous hand robot may include: adjusting model parameters by calculating a reconstruction loss and / or KL divergence for the functional grasping prediction model for the dexterous hand robot based on the predicted grasping configuration and the surface point cloud data included in the 3D asset data to obtain the trained functional grasping prediction model for the dexterous hand robot.
[0013] In another general aspect, a functional grasping prediction method for a dexterous hand robot is provided. The method includes: acquiring two-dimensional image data of a first object of a specific category; performing three-dimensional asset generation processing on the two-dimensional image data of the first object to obtain three-dimensional asset data of the first object; inputting point cloud data of the first object corresponding to the three-dimensional asset data of the first object into a pre-trained first feature extractor for recognizing corresponding parts between different instances of objects of the same category to obtain first point cloud feature data of the corresponding first object; performing preset template matching processing associated with a specific category on the point cloud feature data of the first object to obtain availability data of the first object; and performing initialization and functional grasping synthesis on the point cloud data of the first object corresponding to the availability data of the first object using an axis-alignment-based grasping initialization strategy and an availability-based functional grasping synthesis strategy. The grasping optimization process yields grasping configuration data for the first object. This data is then input into a pre-trained second feature extractor for object availability recognition to obtain grasping feature data for the first object. The point cloud data of the first object is input into a pre-trained third feature extractor for extracting preset details to obtain second point cloud feature data for the first object. The point cloud data and second point cloud feature data of the first object are input into a pre-trained fourth feature extractor for extracting preset non-details to obtain third point cloud feature data for the first object. The grasping feature data and third point cloud feature data of the first object are concatenated and input into a preset encoder, latent space, and preset decoder to obtain a predicted grasping configuration for both the dexterous hand robot and the first object. This predicted grasping configuration is used to control the dexterous hand robot to achieve precise grasping of the first object.
[0014] Optionally, the step of performing three-dimensional asset generation processing on the two-dimensional image data of the first object to obtain the three-dimensional asset data of the first object may include: performing image recognition and image segmentation processing on the two-dimensional image data of the first object to obtain the first processed image data of the first object; and inputting the first processed image data of the first object into a pre-trained three-dimensional generation model to obtain the three-dimensional asset data of the first object.
[0015] Optionally, a preset template associated with a specific category can be obtained by: acquiring instance data of at least one object for each of multiple categories; performing availability part labeling processing on any object in each category to obtain object availability data with availability part labels for any object; using the object availability data with availability part labels for any object in each category, obtaining object availability data with availability part labels for other objects in each category through automatic feature matching processing, thereby obtaining a preset template associated with a specific category, wherein the preset template associated with a specific category is included in the object availability data with availability part labels for specific objects in the specific category of the multiple categories.
[0016] Optionally, the step of performing initialization and functional grasping optimization processing on the first object point cloud data corresponding to the first object's availability data using an axis-aligned grasping initialization strategy and an availability-based functional grasping synthesis strategy to obtain the first object's grasping configuration data may include: performing axis-aligned grasping initialization processing on the first object point cloud data corresponding to the first object's availability data using an axis-aligned grasping initialization strategy to obtain the first object's initial grasping data; and performing availability-based functional grasping optimization processing on the first object's initial grasping data using an availability-based functional grasping synthesis strategy to obtain grasping configuration data that meets preset constraints, which serves as the first object's grasping configuration data.
[0017] Optionally, the steps of concatenating the grasping feature data of the first object and the third point cloud feature data of the first object and then inputting them sequentially into a preset encoder, a latent space, and a preset decoder to obtain a predictive grasping configuration for both the dexterous hand robot and the first object, so as to control the dexterous hand robot to achieve accurate grasping of the first object through the predictive grasping configuration, may include: concatenating the grasping feature data of the first object and the third point cloud feature data of the first object and then inputting them into a preset encoder to obtain encoded data associated with the first object; inputting the encoded data associated with the first object and the third point cloud feature data of the first object into the latent space and the preset decoder to obtain a predictive grasping configuration for both the dexterous hand robot and the first object, so as to control the dexterous hand robot to achieve accurate grasping of the first object through the predictive grasping configuration.
[0018] Optionally, the preset constraints may include at least one of the following: functional part distance loss constraint, grasping part distance loss constraint, force closure loss constraint, collision loss constraint between the dexterous hand and the object, and dexterous hand self-collision loss constraint.
[0019] In another general aspect, a functional grasping prediction device for a dexterous hand robot is provided. The device includes: a data acquisition module configured to: acquire two-dimensional image data of a first object of a specific category; a grasping configuration construction module configured to: perform three-dimensional asset generation processing on the two-dimensional image data of the first object to obtain three-dimensional asset data of the first object; input first object point cloud data corresponding to the three-dimensional asset data of the first object into a pre-trained first feature extractor for recognizing corresponding parts between different instances of objects of the same category to obtain first point cloud feature data of the corresponding first object; perform preset template matching processing associated with a specific category on the first object point cloud feature data to obtain first object availability partial data; and perform grasping initialization strategy based on axis alignment and functional grasping synthesis strategy based on availability on the first object point cloud data corresponding to the first object availability partial data. The initialization and functional grasping optimization processes are performed to obtain grasping configuration data for the first object. The grasping prediction module is configured to: input the grasping configuration data of the first object into a pre-trained second feature extractor for object availability recognition to obtain grasping feature data of the first object; input the point cloud data of the first object into a pre-trained third feature extractor for extracting preset details to obtain second point cloud feature data of the first object; input the point cloud data of the first object and the second point cloud feature data of the first object into a pre-trained fourth feature extractor for extracting preset non-details to obtain third point cloud feature data of the first object; and concatenate the grasping feature data of the first object and the third point cloud feature data of the first object and input them into a preset encoder, latent space, and preset decoder to obtain a predicted grasping configuration for both the dexterous hand robot and the first object, so as to control the dexterous hand robot to achieve accurate grasping of the first object through the predicted grasping configuration.
[0020] Optionally, the operation of the capture configuration building module to perform three-dimensional asset generation processing on the two-dimensional image data of the first object to obtain the three-dimensional asset data of the first object may include: performing image recognition and image segmentation processing on the two-dimensional image data of the first object to obtain the first processed image data of the first object; and inputting the first processed image data of the first object into a pre-trained three-dimensional generation model to obtain the three-dimensional asset data of the first object.
[0021] Optionally, a preset template associated with a specific category can be obtained by: acquiring instance data of at least one object for each of multiple categories; performing availability part labeling processing on any object in each category to obtain object availability data with availability part labels for any object; using the object availability data with availability part labels for any object in each category, obtaining object availability data with availability part labels for other objects in each category through automatic feature matching processing, thereby obtaining a preset template associated with a specific category, wherein the preset template associated with a specific category is included in the object availability data with availability part labels for specific objects in the specific category of the multiple categories.
[0022] Optionally, the step of the crawl configuration construction module performing initialization and functional crawl optimization processing on the first object point cloud data corresponding to the first object's availability data using an axis-aligned crawl initialization strategy and an availability-based functional crawl synthesis strategy to obtain the first object's crawl configuration data may include: performing axis-aligned crawl initialization processing on the first object point cloud data corresponding to the first object's availability data using an axis-aligned crawl initialization strategy to obtain the first object's initial crawl data; and performing availability-based functional crawl optimization processing on the first object's initial crawl data using an availability-based functional crawl synthesis strategy to obtain crawl configuration data that meets preset constraints, which serves as the first object's crawl configuration data.
[0023] Optionally, the grasping prediction module concatenates the grasping feature data of the first object and the third point cloud feature data of the first object, and then inputs them sequentially into a preset encoder, a latent space, and a preset decoder to obtain a predicted grasping configuration for both the dexterous hand robot and the first object. The operation of controlling the dexterous hand robot to accurately grasp the first object through the predicted grasping configuration may include: concatenating the grasping feature data of the first object and the third point cloud feature data of the first object and inputting them into a preset encoder to obtain encoded data associated with the first object; inputting the encoded data associated with the first object and the third point cloud feature data of the first object into the latent space and the preset decoder to obtain a predicted grasping configuration for both the dexterous hand robot and the first object, so as to control the dexterous hand robot to accurately grasp the first object through the predicted grasping configuration.
[0024] Optionally, the preset constraints may include at least one of the following: functional part distance loss constraint, grasping part distance loss constraint, force closure loss constraint, collision loss constraint between the dexterous hand and the object, and dexterous hand self-collision loss constraint.
[0025] In another general aspect, a training method is provided for a functional grasping prediction model for a dexterous hand robot. The training method includes: acquiring sample two-dimensional image data of at least one object for each of multiple categories; performing three-dimensional asset generation processing on the sample two-dimensional image data to obtain sample three-dimensional asset data of at least one object; inputting first object point cloud data corresponding to the sample three-dimensional asset data into a pre-trained first feature extractor for identifying corresponding parts between different instances of objects of the same category to obtain corresponding second point cloud feature data; and performing preset template matching processing associated with a specific category on the second object point cloud feature data to obtain at least one object availability component. Data; using an axis-aligned grasping initialization strategy and an availability-based functional grasping synthesis strategy, initialization and functional grasping optimization are performed on the second object point cloud data corresponding to the availability data of at least one object to obtain grasping configuration data for at least one object; the grasping configuration data of at least one object and the second object point cloud data are input into a functional grasping prediction model for a dexterous hand robot to obtain a predicted grasping configuration for both the dexterous hand robot and at least one object; the model parameters are adjusted based on the predicted grasping configuration and surface point cloud data included in the sample 3D asset data to obtain a trained functional grasping prediction model for the dexterous hand robot.
[0026] Optionally, the functional grasping prediction model for the dexterous hand robot may include: a second feature extractor for object availability identification, a third feature extractor for extracting preset details, a fourth feature extractor for extracting preset non-details, a preset encoder, a latent space, and a preset decoder. The step of inputting grasping configuration data of at least one object and point cloud data of a second object into the functional grasping prediction model for the dexterous hand robot to obtain a predicted grasping configuration for both the dexterous hand robot and at least one object may include: inputting the grasping configuration data of at least one object into the second feature extractor to obtain grasping feature data of at least one object; inputting the point cloud data of the second object into the third feature extractor to obtain second point cloud feature data of at least one object; inputting the second object point cloud data and the second point cloud feature data into the fourth feature extractor to obtain third point cloud feature data of at least one object; concatenating the grasping feature data of at least one object and the third point cloud feature data and then inputting the concatenated data into the preset encoder, latent space, and preset decoder to obtain a predicted grasping configuration for both the dexterous hand robot and at least one object. The steps of adjusting model parameters based on the predictive grasping configuration and surface point cloud data included in the sample 3D asset data to obtain a trained functional grasping prediction model for a dexterous hand robot may include: adjusting model parameters by calculating the reconstruction loss and / or KL divergence for the functional grasping prediction model for the dexterous hand robot based on the predictive grasping configuration and surface point cloud data included in the sample 3D asset data to obtain a trained functional grasping prediction model for the dexterous hand robot.
[0027] In another general aspect, a training apparatus for a functional grasping prediction model for a dexterous hand robot is provided. The training apparatus includes: a sample acquisition module configured to: acquire sample two-dimensional image data of at least one object for each of multiple categories; a sample grasping configuration construction module configured to: perform three-dimensional asset generation processing on the sample two-dimensional image data to obtain sample three-dimensional asset data of at least one object; input first object point cloud data corresponding to the sample three-dimensional asset data into a pre-trained first feature extractor for recognizing corresponding parts between different instances of objects of the same category to obtain corresponding second point cloud feature data; and perform preset template matching processing associated with a corresponding specific category on the second object point cloud feature data to obtain at least one pair of... The system takes at least one object and uses the available partial data for the object as an example. It employs an axis-aligned grasping initialization strategy and an availability-based functional grasping synthesis strategy to perform initialization and functional grasping optimization on the second object point cloud data corresponding to the available partial data of at least one object, resulting in grasping configuration data for at least one object. A grasping prediction model tuning module is configured to input the grasping configuration data of at least one object and the second object point cloud data into a functional grasping prediction model for a dexterous hand robot, obtaining predicted grasping configurations for both the dexterous hand robot and at least one object. Based on the predicted grasping configurations and surface point cloud data included in the sample 3D asset data, the model parameters are adjusted to obtain a trained functional grasping prediction model for the dexterous hand robot.
[0028] Optionally, the functional grasping prediction model for the dexterous hand robot may include: a second feature extractor for object availability recognition, a third feature extractor for extracting preset details, a fourth feature extractor for extracting preset non-details, a preset encoder, a latent space, and a preset decoder. The grasping prediction model parameter tuning module inputs grasping configuration data of at least one object and point cloud data of a second object into the functional grasping prediction model for the dexterous hand robot to obtain predicted grasping configurations for both the dexterous hand robot and at least one object. This operation may include: inputting the grasping configuration data of at least one object into the second feature extractor to obtain grasping feature data of at least one object; inputting the point cloud data of the second object into the third feature extractor to obtain second point cloud feature data of at least one object; inputting the second object point cloud data and the second point cloud feature data into the fourth feature extractor to obtain third point cloud feature data of at least one object; and concatenating the grasping feature data of at least one object and the third point cloud feature data and then inputting the concatenated data into the preset encoder, latent space, and preset decoder to obtain predicted grasping configurations for both the dexterous hand robot and at least one object. The parameter tuning module for the grasping prediction model adjusts the model parameters based on the predicted grasping configuration and surface point cloud data included in the sample 3D asset data to obtain a trained functional grasping prediction model for the dexterous hand robot. The operation may include: adjusting the model parameters by calculating the reconstruction loss and / or KL divergence for the functional grasping prediction model for the dexterous hand robot based on the predicted grasping configuration and surface point cloud data included in the sample 3D asset data, to obtain a trained functional grasping prediction model for the dexterous hand robot.
[0029] In another general aspect, a computer program product is provided, the computer program product comprising a computer program / instructions that, when executed by a processor, implement the functional grasping prediction method for a dexterous hand robot as described above and the training method for the functional grasping prediction model for a dexterous hand robot as described above.
[0030] In another general aspect, a computing device is provided, the computing device comprising: at least one processor; at least one memory storing computer-executable instructions, wherein, when executed by the at least one processor, the computer-executable instructions cause the at least one processor to perform the functional grasping prediction / synthesis method for a dexterous hand robot as described above and the training method for a functional grasping prediction model for a dexterous hand robot as described above.
[0031] In another general aspect, a computer-readable storage medium is provided that, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device / server, enables the electronic device / server to perform the functional grasping prediction / synthesis method for a dexterous hand robot as described above, and the training method for a functional grasping prediction model for a dexterous hand robot as described above.
[0032] The functional grasping prediction / synthesis method for dexterous hand robots and the corresponding prediction model training method according to embodiments of this disclosure achieve efficient functional grasping prediction for dexterous hands. Furthermore, the functional grasping prediction / synthesis method for dexterous hand robots and the corresponding prediction model training method according to embodiments of this disclosure train the prediction model on a proposed large-scale dataset with diverse scales, solving the common problem of existing datasets being biased towards large objects due to robot hand size constraints. In addition, the functional grasping prediction / synthesis method for dexterous hand robots and the corresponding prediction model training method according to embodiments of this disclosure also enhance the efficiency and generalization ability of dexterous hand functional grasping generation. Attached Figure Description
[0033] The above and other objects and features of the embodiments of this disclosure will become clearer from the following description taken in conjunction with the accompanying drawings illustrating the embodiments, wherein: Figure 1 This is a flowchart illustrating a functional grasping prediction method for a dexterous hand robot according to an embodiment of the present disclosure; Figure 2 This is a schematic diagram illustrating the availability area of a dexterous hand according to an embodiment of the present disclosure; Figure 3 This is a schematic diagram illustrating the availability area of an object according to an embodiment of the present disclosure; Figure 4 This is a schematic diagram illustrating the grab initialization according to an embodiment of the present disclosure; Figure 5 This is a schematic diagram illustrating a functional grasping prediction model for a dexterous hand robot according to an embodiment of the present disclosure; Figure 6 This shows simulation results comparing the prediction method according to this disclosure with other methods; Figure 7 This is a flowchart illustrating a training method for a functional grasping prediction model for a dexterous hand robot according to an embodiment of the present disclosure; Figure 8 This is a block diagram illustrating a functional grasping prediction device for a dexterous hand robot according to an embodiment of the present disclosure; Figure 9This is a block diagram illustrating a training apparatus for a functional grasping prediction model for a dexterous hand robot according to an embodiment of the present disclosure; Figure 10 This is a block diagram illustrating a computing device according to an embodiment of the present disclosure. Detailed Implementation
[0034] The following detailed embodiments are provided to aid the reader in gaining a comprehensive understanding of the methods, apparatus, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but may be changed as will become clear upon understanding this disclosure, except for operations that must occur in a specific order. Furthermore, for clarity and conciseness, descriptions of features known in the art may be omitted.
[0035] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings, examples of which are illustrated in the drawings, wherein the same reference numerals always refer to the same parts. The embodiments will now be described with reference to the accompanying drawings in order to explain this disclosure.
[0036] Existing functional grasping methods suffer from several problems. For instance, regarding dataset construction, they are costly, inefficient, and lack large-scale datasets. Current methods typically rely on demonstrations of human hands, which are labor-intensive and limited in number. Furthermore, the dependence of current dataset construction methods on existing object datasets or manual 3D (three-dimensional) scans further restricts their scalability and diversity. Regarding the size adaptability of grasped objects, existing methods typically redirect to the robot hand through direct redirection or contact map alignment, directly or indirectly scaling up human hand grasping configurations to the robot hand. This lack of effective mechanisms for handling different object scales makes it difficult to generate functional grasping configurations suitable for the robot hand itself, especially when dealing with small objects. Additionally, regarding the diversity of grasping gestures, existing methods either directly mimic human hand grasping to generate functional grasping for dexterous hands, or assume that human hand and dexterous hand functional grasping share a common contact map and redirect human hand grasping to the dexterous hand via the contact map. However, these methods all suffer from insufficient diversity in functional grasping.
[0037] To address the aforementioned and / or other problems, this disclosure provides a method for predicting functional grasping in dexterous hand robots, as well as a training method for the corresponding prediction model. In one aspect, by utilizing a large-scale dataset specifically designed for scale diversity, this disclosure solves the common problem of existing datasets being biased towards large objects due to robot hand size constraints. In another aspect, by employing an availability-based grasping generation method, this disclosure overcomes the limitation of existing methods (e.g., contact map methods) being sensitive to the hand-to-object scale ratio. The prediction and training methods of this disclosure, by combining automated processing and efficiency improvements of pre-trained models, significantly reduce manual annotation work and enhance the efficiency and generalization ability of dexterous hand functional grasping generation.
[0038] The following reference Figures 1 to 9 This disclosure provides a detailed description of a functional grasping prediction method for dexterous hand robots, a training method for the corresponding prediction model, and a corresponding apparatus, according to embodiments of the present disclosure. The methods and apparatus of this disclosure can be applied to the field of robot grasping technology.
[0039] First, refer to Figures 1 to 5 , Figure 7 A detailed description of a functional grasping prediction method for a dexterous hand robot according to embodiments of the present disclosure. Figure 1 This is a flowchart illustrating a functional grasping prediction method 100 for a dexterous hand robot according to an embodiment of the present disclosure. Figure 2 This is a schematic diagram illustrating the availability area of a dexterous hand according to an embodiment of the present disclosure. Figure 3 This is a schematic diagram illustrating the availability area of an object according to an embodiment of the present disclosure. Figure 4 This is a schematic diagram illustrating the grab initialization according to an embodiment of the present disclosure. Figure 5 This is a schematic diagram illustrating a functional grasping prediction model for a dexterous hand robot according to an embodiment of the present disclosure. Figure 7 This is a flowchart illustrating a training method 700 for a functional grasping prediction model for a dexterous hand robot according to an embodiment of the present disclosure.
[0040] Reference Figure 1 According to an embodiment of the present disclosure, in step S101, two-dimensional image data of a first object of a specific category is acquired.
[0041] According to an embodiment of this disclosure, in step S102, a three-dimensional asset generation process is performed on the two-dimensional image data of the first object to obtain the three-dimensional asset data of the first object.
[0042] As an example, step S102 may further include steps S1021 and S1022: performing image recognition and image segmentation processing on the two-dimensional image data of the first object to obtain the first processed image data of the first object (S1021); inputting the first processed image data of the first object into a pre-trained three-dimensional generative model to obtain the three-dimensional asset data of the first object (S1022).
[0043] According to this disclosure, by generating a large-scale three-dimensional asset dataset using acquired two-dimensional images, data diversity can be improved, and the subsequent successful capture results are better than those of real datasets.
[0044] According to an embodiment of this disclosure, in step S103, the point cloud data of the first object corresponding to the three-dimensional asset data of the first object is input into a pre-trained first feature extractor for identifying corresponding parts between different instances of the same type of object, so as to obtain the first point cloud feature data of the corresponding first object.
[0045] According to an embodiment of this disclosure, in step S104, a preset template matching process associated with a specific category is performed on the point cloud feature data of the first object to obtain the availability data of the first object.
[0046] Here, it should be noted that "first object availability data" encompasses both object availability data and dexterity hand availability data. The term "first object" indicates both types of data specific to the first object, rather than only including data related to the first object. That is, "first object availability data" includes not only the availability data of the first object itself but also the corresponding dexterity hand availability data. Furthermore, other data described using "first object" can be understood, depending on the need and context, to include corresponding data for both the dexterity hand and the object.
[0047] Here, in this disclosure, object (or feature) availability refers to the operable properties of an object component that imply potential interaction. By reasoning about object availability rather than mimicking human demonstrations, this disclosure enables dexterous hand robots to achieve more efficient, intuitive, adaptive, and diverse grasping strategies.
[0048] For example, a preset template associated with a specific category can be obtained through the following processes S1041 to S1043: In process S1041, instance data of at least one object for each of the multiple categories is obtained.
[0049] In processing S1042, the availability part labeling process is performed on any object in each category to obtain the object availability data with availability part labels for any object.
[0050] In processing S1043, using the object availability data with availability partial labels of any object in each category, automatic feature matching is used to obtain the object availability data with availability partial labels of other objects in each category, thereby obtaining a preset template associated with a specific category.
[0051] Here, the preset template associated with a specific category is included in the object availability data of a specific object in a specific category across multiple categories, with the availability part labeled.
[0052] According to this disclosure, availability detection of objects at the category level is achieved by identifying availability components (e.g., availability parts) among objects of the same category through dense feature matching, and generalization and efficiency are effectively improved by using a first feature extractor.
[0053] Furthermore, the availability-based algorithm proposed in this disclosure more closely resembles the human cognitive process in object interaction, providing a more fundamental logic for grasping prediction. In addition, the availability-based algorithm can fully utilize the capabilities of the pre-trained base model, thereby helping to improve the adaptability and prediction efficiency of grasping prediction methods.
[0054] According to an embodiment of this disclosure, in step S105, using an axis-aligned grasping initialization strategy and an availability-based functional grasping synthesis strategy, initialization and functional grasping optimization processes are performed on the first object point cloud data corresponding to the first object availability data to obtain the first object grasping configuration data.
[0055] As an example, step S105 may further include steps S1051 and S1052: using an axis-aligned grasping initialization strategy, performing axis-aligned grasping initialization processing on the first object point cloud data corresponding to the first object availability data to obtain the first object's initial grasping data (S1051); using an availability-based functional grasping synthesis strategy, performing availability-based functional grasping optimization processing on the first object's initial grasping data to obtain grasping configuration data that meets preset constraints, which is used as the first object's grasping configuration data (S1052).
[0056] For example, the preset constraints may include at least one of the following: functional part distance loss constraints, grasping part distance loss constraints, force closure loss constraints, collision loss constraints between the dexterous hand and the object, and dexterous hand self-collision loss constraints.
[0057] According to this disclosure, by adopting the above-mentioned axis-aligned grasping initialization strategy, the availability-based functional grasping synthesis strategy, and various preset constraints, the adaptability to changes in object scale is greatly improved.
[0058] For example, considering common real-world sizes for each category, a scale interval from a first scale S_low to a second scale S_high is defined, and a predetermined number (e.g., 15) of scales are randomly sampled within this interval for grasping synthesis. At extreme scales (e.g., smaller sizes), the robot hand is relatively large and prone to hand self-collision, leading to unreasonable grasping. Therefore, the method disclosed herein achieves robust grasping of smaller objects through gesture variations (rather than just fine-tuning the finger configuration).
[0059] According to an embodiment of this disclosure, in step S106, the grasping configuration data of the first object is input into a pre-trained second feature extractor for object availability recognition to obtain the grasping feature data of the first object.
[0060] According to an embodiment of this disclosure, in step S107, the point cloud data of the first object is input into a pre-trained third feature extractor for extracting preset details to obtain the second point cloud feature data of the first object.
[0061] As an example, the third feature extractor here can be the same as the first feature extractor above.
[0062] According to an embodiment of this disclosure, in step S108, the first object point cloud data and the second point cloud feature data of the first object are input into a pre-trained fourth feature extractor for extracting preset non-details to obtain the third point cloud feature data of the first object.
[0063] According to an embodiment of this disclosure, in step S109, the grasping feature data of the first object and the third point cloud feature data of the first object are concatenated and input into a preset encoder, a latent space (or hidden space, which includes a latent code) and a preset decoder to obtain a predicted grasping configuration for both the dexterous hand robot and the first object, so as to control the dexterous hand robot to achieve accurate grasping of the first object through the predicted grasping configuration.
[0064] For example, the step of "obtaining the predicted grasping configuration for both the dexterous hand robot and the first object" in step S109 can specifically include: first obtaining the predicted grasping configuration of the dexterous hand robot for the first object, obtaining the hand surface point cloud data of the dexterous hand under the predicted grasping configuration through forward kinematics (i.e., the predicted grasping configuration for the dexterous hand robot), and obtaining the corresponding predicted grasping configuration for the first object.
[0065] As an example, step S109 may further include steps S1091 and S1092: concatenating the grasping feature data of the first object and the third point cloud feature data of the first object and inputting them into a preset encoder to obtain coded data associated with the first object (S1091); inputting the coded data associated with the first object and the third point cloud feature data of the first object into the latent space and a preset decoder to obtain a predicted grasping configuration for both the dexterous hand robot and the first object, so as to control the dexterous hand robot to achieve accurate grasping of the first object through the predicted grasping configuration (S1092).
[0066] According to this disclosure, by adopting a lightweight network architecture for smart functional crawling prediction, the high computational burden caused by complex models is avoided, while exhibiting good generalization ability and high-quality crawling results.
[0067] The functional grasping prediction method disclosed herein demonstrates outstanding performance in terms of efficiency in constructing dexterous hand functional grasping datasets and adaptability to diverse object scales.
[0068] Specifically, this disclosure proposes an efficient method for constructing a functional grasping dataset for dexterous hands. Through automated object generation and automated availability and grasping annotation, a large-scale functional grasping dataset for dexterous hands is constructed. This disclosure employs an availability-based grasping synthesis algorithm to synthesize diverse objects (e.g., everyday items, tools, etc.) using grasping configurations without expert demonstration, achieving a flexible hand-to-object size ratio. This enables larger robotic hands (relative to human hands) to effectively grasp objects. For example, a robotic hand larger than a human hand can grasp everyday objects suitable for human grasping (e.g., objects on the same order of magnitude as a human hand, such as objects that can generally be picked up by a human hand, rather than objects much larger than a human hand, such as refrigerators or wardrobes), and can grasp objects that are relatively small compared to the robotic hand.
[0069] For example, the dataset disclosed herein may contain five object categories, each containing more than 1,000 unique shapes with 15 scale variations. After data processing (e.g., filtering), the dataset provides more than 60,000 synthetic grasps to two dexterous hand robots respectively.
[0070] Furthermore, the grasping prediction method of this disclosure can significantly reduce the limitations of the relative size between the hand and the object, and can adapt to objects of a wider range of sizes. The method of functional grasping with a dexterous hand based on object availability synthesis can break free from the limitations of human demonstrations, instead relying on the correlation between the availability components of the dexterous hand and the corresponding availability components of the object to achieve precise grasping. For example, the grasping in the synthesis method of this disclosure originates from the availability analysis of the object, rather than merely mimicking examples of humans grasping a particular object of that category, thereby reducing the constraints of the relative scale between the hand and the object in grasping synthesis.
[0071] Furthermore, the grasping prediction method disclosed herein can significantly enhance the diversity of functional grasping. For example, since the grasping method based on object availability synthesis of dexterous hand functional grasping is derived from the availability analysis of each object, rather than merely imitating a number of human grasping actions, each grasping action is specific, increasing the diversity of grasping.
[0072] Next, refer to Figure 7 A detailed description is provided of a training method for a functional grasping prediction model for a dexterous hand robot according to embodiments of the present disclosure.
[0073] Reference Figure 7 In step S701, sample two-dimensional image data of at least one object in each of the multiple categories are obtained.
[0074] In step S702, a three-dimensional asset generation process is performed on the sample two-dimensional image data to obtain sample three-dimensional asset data of at least one object.
[0075] For example, a large number of 2D images can be collected, preferably through internet keyword searches or any other suitable method. Then, image processing techniques (such as, but not limited to, Grounded SAM2) are used to filter and segment the 2D images. Afterward, multiple pre-trained generative models (such as, but not limited to, Instantmesh, Triposr, and TRELLIS) are used to generate 3D object assets (or 3D assets) based on the 2D images. Furthermore, mesh diversity can be ensured by randomly using the generative models.
[0076] In addition, the training method disclosed herein may further perform the following processes (1) to (3) to filter out anomalies brought about by the pre-trained model: (1) Images that are not related to the category after segmentation are quickly screened manually by thumbnails, with each category controlled within a preset time (e.g., ten minutes); (2) Degenerate (e.g., unstructured surface or poor geometric quality) whose grid size is not within a preset range (e.g., the range of 1M to 10M) will be eliminated; (3) Objects are further screened based on grasping quality (such as indicators including simulation stability, hand-object corresponding availability region distance, hand-object collision and hand self-collision).
[0077] Furthermore, even if anomalous objects exist (such as incomplete objects, objects with extra components, or illusory objects generated by the generative model), as long as the capture is successful and of high quality, the object is retained to further enhance the diversity of the dataset, thus providing a greater advantage in data diversity than the real dataset.
[0078] In step S703, the first object point cloud data corresponding to the sample 3D asset data is input into a pre-trained first feature extractor for identifying corresponding parts between different instances of the same type of object, to obtain the corresponding second point cloud feature data.
[0079] Here, the first feature extractor can be, for example, but not limited to, the PTv3-object model in SAMPart3D, whose training process is distilled from FeatUp-DINOv2. By employing a pre-trained first feature extractor, since the extracted features have the property that "parts" can be well matched between different instances of the same object, it is possible to identify corresponding availability parts within a category based on point clouds, thereby improving the generalization and efficiency in availability retrieval.
[0080] The training method disclosed herein utilizes visual models and abundant 2D image resources to efficiently generate a large number of 3D object assets through generative models. Then, features extracted from pre-trained perceptual models are used to identify the availability parts of 3D objects, thereby enabling scalable and automated dataset construction and annotation.
[0081] In step S704, a preset template matching process associated with the corresponding specific category is performed on the second object point cloud feature data to obtain at least one object availability data.
[0082] In step S705, using an axis-aligned grasping initialization strategy and an availability-based functional grasping synthesis strategy, initialization and functional grasping optimization are performed on the second object point cloud data corresponding to at least one object availability data to obtain grasping configuration data for at least one object.
[0083] Here, through the grab synthesis method disclosed herein, functional grabs can be synthesized to obtain a grab configuration that performs specific grab post-operations, integrating object availability and task purpose.
[0084] The following reference Figures 2 to 4 To illustrate the specific process of category-level object availability detection: For each category, an object is randomly selected as a template, and its availability components are pre-labeled. This template serves as the current instance, and other instances are automatically identified through feature matching to determine their corresponding availability components. This labeling process is extremely lightweight, requiring only one template per category, but this disclosure is not limited to this. Figure 3 and Figure 4 In the diagram, the green part represents the functional part, and the blue part represents the stable gripping part.
[0085] For example, the grasping initialization of this disclosure uses an axis alignment-based initialization strategy. Taking the case where the index finger of a dexterous hand is the functional finger as an example, the grasping-functional axis and the force axis are defined for the dexterous hand and the object, respectively. The initialization process first precisely aligns the primary axis (grasping-functional axis), then aligns the secondary axis (force axis) in the orthogonal plane, then translates the dexterous hand until it makes initial contact with the object but does not penetrate it, and finally initializes the finger joints of the dexterous hand to a slightly flexed state to facilitate subsequent optimization.
[0086] Next, in the process of optimizing grasping synthesis, the training method of this disclosure can identify available parts among objects of the same category through dense feature matching. Regarding the functional grasping synthesis of a dexterous hand, this disclosure generates dexterous functional grasps based on an availability strategy, identifying functional parts and grasping parts in the object and the dexterous hand, respectively. During grasping synthesis, similar available parts in the object and the dexterous hand are constrained to remain close, ensuring both the functionality and stability of the grasp. Furthermore, minimizing collisions for physically plausible grasps is considered, including hand-object penetration and hand-to-hand penetration. Based on these constraints, the method of this disclosure can synthesize dexterous functional grasps that possess functionality, stability, and physical plausibility.
[0087] As an example, the optimization objective of the above-mentioned dexterous hand functional grasping synthesis method is to minimize the weighted sum (corresponding to equation (1) below), grasping component distance loss (corresponding to equation (2) below), force closure loss (corresponding to equation (3) below), hand-object collision loss (corresponding to equation (4) below), and hand-self-collision loss (corresponding to equation (5) below) (corresponding to equation (6) below). Equations (1) to (6) below show how the specific loss functions corresponding to the above optimization objective are calculated: (1) in, This indicates functional distance loss. This represents two sets of point clouds. These represent the points in the two point clouds, Indicates the number of points. A point cloud representing a functional part or component (i.e., the area of an object that performs a specific operation, such as the button on a drill). This indicates the specific anchor point of the dexterous hand selected for the distance loss of functional components.
[0088] Here, by employing a functional partial distance loss, the following can be achieved: and The shorter the distance between them, the more likely it is to encourage the corresponding available components to move closer together.
[0089] (2) in, This indicates a loss in distance during capture. Anchor points on the grasping components of a dexterous hand. Point cloud representing gripping parts (typically located on the handle or body of an object, such as the handle of a drill).
[0090] By employing a grasping distance loss, the proximity between the corresponding hand parts of a dexterous hand and the object grasping parts can be effectively balanced.
[0091] (3) in, Indicates force closure loss, Indicates the friction cone shaft, Represents matrix multiplication and , , Indicates the contact point for achieving the grabbing action. The number of contact points The corresponding friction cone shaft can be expressed as: .
[0092] Here, through the aforementioned force-closure loss, we can measure the deviation between the actual force-torque balance and the ideal state, as well as the degree of violation of friction constraints, ensuring that the applied force and torque can prevent the object from sliding or becoming unstable.
[0093] (4) in, Indicates loss due to hand-object collision. Represents a signed distance field and represents Distance to the nearest surface in space and set of symbolic distinguishing points In the surface ( ) or outside ( ), Representing point clouds of objects, A grid representing the hand of a dexterous hand. This refers to the various parts of a dexterous hand (also simply called the hand) (e.g., including the fingers and the palm).
[0094] By employing hand-object collision loss to minimize hand-object penetration, a physically plausible grasp can be generated.
[0095] (5) in, Indicates damage from a manual-automatic collision. and To indicate different parts of the hand, This represents the surface point cloud of different parts of the hand.
[0096] By employing manual collision loss, manual penetration can be limited and minimized.
[0097] (6) in, This represents the total loss after weighted summation of the above losses. These are the weighting coefficients for each loss.
[0098] It should be noted that in this disclosure, the formulas mentioned above and below refer to... and All of them are natural numbers greater than zero, and will not be elaborated further.
[0099] Furthermore, this disclosure also enables the selection of synthetic crawlers based on multiple metrics, retaining only high-quality samples to ensure high-quality datasets. For example, by using a simulation environment to filter and refine the synthetic crawlers, crawlers with severe collisions can be removed, retaining only high-quality crawlers, thereby further optimizing the synthetic results.
[0100] In simulation experiments, through the high-quality initialization and batch optimization processing disclosed herein, the optimization process can converge in less than 200 steps, and a single RTX 3090 Ti GPU can synthesize 150 smart functional grasps simultaneously within 30 seconds, with an average grasp time of less than 0.2 seconds per grasp, which is far superior to the existing ContactGrasp method (which takes more than 0.5 hours per grasp) and also shows improvement compared to the existing DexFG method based on real object datasets (which takes about 0.25 seconds per grasp) (for example, see the simulation results below).
[0101] The 3D features extracted using the availability checking method disclosed herein still exhibit strong robustness when the object point cloud undergoes significant rotational changes.
[0102] In step S706, the grasping configuration data of at least one object and the point cloud data of the second object are input into the functional grasping prediction model for the dexterous hand robot to obtain the predicted grasping configuration for both the dexterous hand robot and at least one object.
[0103] In step S707, the model parameters are adjusted based on the predicted grasping configuration and the corresponding actual point cloud data included in the sample 2D image data to obtain the trained functional grasping prediction model for the dexterous hand robot. For example, the model parameters can be adjusted based on the predicted grasping configuration and the corresponding surface point cloud data included in the sample 3D asset data to obtain the trained functional grasping prediction model for the dexterous hand robot.
[0104] For example, the functional grasping prediction model for dexterous hand robots described above may include: a second feature extractor for object availability recognition, a third feature extractor for extracting preset details, a fourth feature extractor for extracting preset non-details, a preset encoder, a latent space, and a preset decoder.
[0105] In this case, step S706 may specifically include steps S7061 to S7064: In step S7061, the grasping configuration data of at least one object is input into the second feature extractor to obtain the grasping feature data of at least one object.
[0106] In step S7062, the second object point cloud data is input into the third feature extractor to obtain the second point cloud feature data of at least one object.
[0107] In step S7063, the second object point cloud data and the second point cloud feature data are input into the fourth feature extractor to obtain the third point cloud feature data of at least one object.
[0108] In step S7064, the grasping feature data of at least one object and the third point cloud feature data are concatenated and then input into the preset encoder, latent space and preset decoder to obtain the predicted grasping configuration for the dexterous hand robot and at least one object.
[0109] For example, refer to Figure 5 For example, a functional grasping prediction model for a dexterous hand robot can be described. The functional grasping prediction model for a dexterous hand robot may include a first preset feature extractor 503, a second preset feature extractor 505, a third preset feature extractor 506, an encoder 507, a latent space 508, and a decoder 509.
[0110] like Figure 5As shown, from the perspective of data flow, the following processing is included: 1) The capture configuration 501 is subjected to first feature extraction by a first preset feature extractor 503 (e.g., but not limited to a multilayer perceptron MLP module) to obtain the extracted first feature; 2) The object point cloud 502 is subjected to second feature extraction by a second preset feature extractor 505 (e.g., but not limited to a pre-trained object feature extractor PTv3-Object from SAMPart3D) to obtain the extracted second feature (i.e., point cloud feature 504); 3) The second feature and the object point cloud are combined and input into a third preset feature extractor 506 (e.g., but not limited to a pre-trained object feature extractor PointNet from SAMPart3D, or PointNet++, Pointnext, Pointtransformer, Pointvit, etc.) to obtain the third feature; 4) The first feature and the third feature are combined and input into an encoder 507, and the third feature is also input into the latent space 508. All these data are finally processed by a decoder 509 to obtain the prediction result 510. In other words, the object point cloud and its principal components are concatenated and input into the encoder, then mapped to the latent space. The encoder then predicts the grasping configuration and uses differentiable positive kinematics to predict the point cloud of both the dexterous hand and the object.
[0111] This disclosure achieves efficient grasping prediction by integrating a pre-trained object feature extractor using a lightweight conditional variational autoencoder architecture. Furthermore, this disclosure employs a lightweight network architecture for smart, functional grasping prediction, avoiding the high computational burden of complex models while exhibiting good generalization ability and high-quality grasping results.
[0112] In the above case, step S707 may specifically include: adjusting the model parameters by calculating the reconstruction loss (corresponding to Equation (7) below) and / or KL divergence (corresponding to Equation (8) below) for the functional grasping prediction model for the dexterous hand robot based on the predicted grasping configuration and the corresponding actual point cloud data included in the sample two-dimensional image data (e.g., by calculating the corresponding surface point cloud data included in the sample three-dimensional asset data based on the predicted grasping configuration) to obtain the trained functional grasping prediction model for the dexterous hand robot.
[0113] (7) in, Indicates the reconstruction loss. Represents the number of points in the point cloud. The point cloud represents the hand of a dexterous hand that is predicted to grasp. This represents the actual point cloud corresponding to the predicted hand point cloud.
[0114] By using point clouds instead of joint angles as the alignment target, the above prediction model can learn spatial structure more intuitively, reduce the risk of overfitting, and improve generalization ability.
[0115]
[0116] (8) in, Denotes KL divergence, This represents the dimension of the latent space of the autoencoder. Represents the mean and (It is predicted by the encoder). Represents variance and , Indicates a Gaussian distribution. Represents the identity matrix.
[0117] By employing KL divergence, the latent space can be regularized to conform to a Gaussian distribution. This improves the model's generalization ability.
[0118] Furthermore, by removing the encoder from the prediction model (or prediction network) during the inference phase, directly from... By sampling latent variables, the predictive model can generate functional captures of novel objects and objects of different scales, greatly improving the generalization and zero-shot transfer capabilities from synthetic to real-world scenes.
[0119] As an experiment, refer to Tables 1 to 4 below and Figure 6 (It shows simulation results comparing the prediction method of this disclosure with other methods.) By comparing the above-described method of this disclosure with other existing methods, the simulation results demonstrate that the above-described method of this disclosure can achieve the significant beneficial effects mentioned in this disclosure.
[0120] Table 1
[0121] Table 1 shows a comparison of the simulation results related to grasping on multiple datasets. In Table 1, the dataset of the method of this disclosure (shown as ScaleADFG in Table 1) is compared with the datasets of other methods. As can be seen from Table 1, the ScaleADFG method of this disclosure has significant improvements over other methods in terms of functionality, multiple dexterous hands, and scale variability. Furthermore, , , The numbers represent the number of categories, the number of objects, and the number of items grabbed, respectively. In Table 1, the values in bold represent the optimal values, and the values in underline represent the second-best values.
[0122] Table 2
[0123] Table 2 shows a comparison of the evaluation results regarding the quality of the datasets. As can be seen from Table 2, the datasets generated according to the method of this disclosure are of high quality.
[0124] Table 3
[0125] Table 3 shows a comparison of the evaluation results of various grasping-related metrics under different data volumes (i.e., training set) (volume in Table 3 represents data volume). Here, grasses / cat. represents the number of grasps per category, i.e., the number of grasps per category. Table 3 shows that the larger the data volume, the higher the success rate SR, used to measure stability, and the higher the distance d for the grasping component in contact. G (The smaller the distance d (used to measure the coordination / consistency between the hand's grasping parts and the object's grasping parts), the more functional the distance between the functional parts.) F (The smaller the value of d, which measures the coordination / consistency between the functional parts of the hand and the functional parts of an object, the deeper the hand-object interaction.) IP (The smaller the distance field used to represent the maximum negative sign) and the deeper the hand self-penetration d, the better. SP The smaller the value (used to measure the maximum penetration of different hand parts), the better. This indicates that the prediction method disclosed herein can accurately grasp various appropriate categories of objects.
[0126] Table 4
[0127] Table 4 shows the differences between the dataset of this disclosure and the DexFG dataset across various metrics (including...). Figure 3 A similar comparison table of evaluation results (stability, contact, functionality, and penetration) is provided. As can be seen from Table 4, the simulation results of the dataset presented in this publication are superior to those of the existing DexFG dataset in terms of the four metrics for predictive crawling.
[0128] Next, refer to Figure 8 A detailed description of a functional grasping prediction device for a dexterous hand robot according to embodiments of the present disclosure.
[0129] Figure 8 This is a block diagram illustrating a functional grasping prediction device 800 for a dexterous hand robot according to an embodiment of the present disclosure.
[0130] Reference Figure 8A functional grasping prediction device 800 for a dexterous hand robot according to an embodiment of the present disclosure may include a data acquisition module 810, a grasping configuration construction module 820, and a grasping prediction module 830.
[0131] According to an embodiment of this disclosure, the data acquisition module 810 can perform the following: acquire two-dimensional image data of a first object of a specific category.
[0132] According to embodiments of this disclosure, the capture configuration building module 820 may perform the following operations 821) to 824): In operation 821), a 3D asset generation process is performed on the 2D image data of the first object to obtain the 3D asset data of the first object.
[0133] In operation 822), the point cloud data of the first object corresponding to the three-dimensional asset data of the first object is input into the pre-trained first feature extractor for recognizing corresponding parts between different instances of the same type of object, so as to obtain the first point cloud feature data of the corresponding first object.
[0134] In operation 823), a preset template matching process associated with a specific category is performed on the point cloud feature data of the first object to obtain the availability data of the first object.
[0135] In operation 824, using an axis-aligned grasping initialization strategy and an availability-based functional grasping synthesis strategy, the point cloud data of the first object corresponding to the availability data of the first object is initialized and functional grasping is optimized to obtain the grasping configuration data of the first object.
[0136] According to embodiments of this disclosure, the capture prediction module 830 may perform the following operations 831) to 834): In operation 831), the grasping configuration data of the first object is input into the pre-trained second feature extractor for object availability recognition to obtain the grasping feature data of the first object.
[0137] In operation 832), the point cloud data of the first object is input into a pre-trained third feature extractor for extracting preset details to obtain the second point cloud feature data of the first object.
[0138] In operation 833), the point cloud data of the first object and the second point cloud feature data of the first object are input into a pre-trained fourth feature extractor for extracting preset non-details, to obtain the third point cloud feature data of the first object.
[0139] In operation 834, the grasping feature data of the first object and the third point cloud feature data of the first object are concatenated and input into the preset encoder, latent space and preset decoder to obtain the predictive grasping configuration for both the dexterous hand robot and the first object, so as to control the dexterous hand robot to achieve accurate grasping of the first object through the predictive grasping configuration.
[0140] It should be noted that the operations performed on the above structural frames can be compared with those in the reference section. Figure 1 The related content is similar, so I will not repeat it here.
[0141] Here, refer to Figure 9 A detailed description is provided of a training apparatus for a functional grasping prediction model for a dexterous hand robot according to embodiments of the present disclosure.
[0142] Figure 9 This is a block diagram illustrating a training apparatus 900 for a functional grasping prediction model for a dexterous hand robot according to an embodiment of the present disclosure.
[0143] Reference Figure 9 The training apparatus 900 for a functional grasping prediction model for a dexterous hand robot according to an embodiment of the present disclosure may include a sample acquisition module 910, a sample grasping configuration construction module 920, and a grasping prediction model parameter tuning module 930.
[0144] According to an embodiment of this disclosure, the sample acquisition module 910 can perform the following: acquire sample two-dimensional image data for at least one object in each of a plurality of categories.
[0145] According to embodiments of this disclosure, the sample capture configuration building module 920 may perform the following operations 921) to 924). In operation 921), a 3D asset generation process is performed on the sample 2D image data to obtain sample 3D asset data of at least one object.
[0146] In operation 922), the first object point cloud data corresponding to the sample 3D asset data is input into a pre-trained first feature extractor for identifying corresponding parts between different instances of the same type of object, and the corresponding second point cloud feature data is obtained.
[0147] In operation 923), a preset template matching process associated with the corresponding specific category is performed on the point cloud feature data of the second object to obtain at least one object availability data.
[0148] In operation 924, using an axis-aligned grasp initialization strategy and an availability-based functional grasp synthesis strategy, initialization and functional grasp optimization are performed on the second object point cloud data corresponding to at least one object availability data to obtain grasp configuration data for at least one object.
[0149] According to embodiments of this disclosure, the prediction model tuning module 930 can perform the following operations 931) and 932). In operation 931), grasping configuration data of at least one object and point cloud data of the second object are input into a functional grasping prediction model for a dexterous hand robot to obtain a predicted grasping configuration for both the dexterous hand robot and at least one object.
[0150] In operation 932, the model parameters are adjusted based on the predicted grasping configuration and the corresponding actual point cloud data included in the sample 2D image data to obtain a trained functional grasping prediction model for a dexterous hand robot. For example, the model parameters are adjusted based on the predicted grasping configuration and the corresponding surface point cloud data included in the sample 3D asset data to obtain a trained functional grasping prediction model for a dexterous hand robot.
[0151] As an example, a functional grasping prediction model for a dexterous hand robot may include: a second feature extractor for object availability recognition, a third feature extractor for extracting preset details, a fourth feature extractor for extracting preset non-details, a preset encoder, a latent space, and a preset decoder.
[0152] In this case, the operation 931) performed by the prediction model parameter tuning module 930 may further include the following operations 9311) to 9314): In operation 9311), the crawling configuration data of at least one object is input into the second feature extractor to obtain the crawling feature data of at least one object.
[0153] In operation 9312), the point cloud data of the second object is input into the third feature extractor to obtain the second point cloud feature data of at least one object.
[0154] In operation 9313), the second object point cloud data and the second point cloud feature data are input into the fourth feature extractor to obtain the third point cloud feature data of at least one object.
[0155] In operation 9314, the grasping feature data of at least one object and the third point cloud feature data are concatenated and input into the preset encoder, latent space and preset decoder to obtain the predictive grasping configuration for the dexterous hand robot and at least one object.
[0156] Furthermore, in this case, the operation 932 performed by the grasping prediction model tuning module may further include the following operation: adjusting the model parameters by calculating the reconstruction loss and / or KL divergence for the functional grasping prediction model for the dexterous hand robot based on the predictive grasping configuration and the corresponding actual point cloud data included in the sample two-dimensional image data (e.g., by calculating the corresponding surface point cloud data included in the sample three-dimensional asset data based on the predictive grasping configuration), thereby obtaining the trained functional grasping prediction model for the dexterous hand robot.
[0157] It should be noted that the operations performed on the above structural frames can be compared with those in the reference section. Figure 7 The related content is similar, so I will not repeat it here.
[0158] Figure 10 This is a block diagram illustrating a computing device 1000 according to an embodiment of the present disclosure.
[0159] Reference Figure 10 The computing device 1000 according to embodiments of the present disclosure may include a processor 1010 and a memory 1020. The processor 1010 may include (but is not limited to) a central processing unit (CPU), a digital signal processor (DSP), a microcomputer, a field-programmable gate array (FPGA), a system-on-a-chip (SoC), a microprocessor, an application-specific integrated circuit (ASIC), etc. The memory 1020 may store computer-executable instructions to be executed by the processor 1010. The memory 1020 includes high-speed random access memory and / or a non-volatile computer-readable storage medium. When the processor 1010 executes the computer-executable instructions stored in the memory 1020, the functional grasping prediction method for a dexterous hand robot described above and the training method for a functional grasping prediction model for a dexterous hand robot described above can be implemented.
[0160] The functional grasping prediction method for a dexterous hand robot and the training method for a functional grasping prediction model for a dexterous hand robot according to embodiments of this disclosure can be written as a computer program / instructions to form a computer program product and stored on a computer-readable storage medium. When the computer program / instructions are executed by a processor, the functional grasping prediction method for a dexterous hand robot and the training method for a functional grasping prediction model for a dexterous hand robot as described above can be implemented. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device / server, the electronic device / server is enabled to perform the functional grasping prediction method for a dexterous hand robot and the training method for a functional grasping prediction model for a dexterous hand robot as described above. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid-state drive (SSD), card storage (such as multimedia cards, secure digital (SD) cards, or ultra-fast digital (XD) cards), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, and any other device configured to store computer programs and any associated data, data files, and data structures in a non-transitory manner and to provide the computer programs and any associated data, data files, and data structures to a processor or computer so that the processor or computer can execute the computer programs. In one example, the computer programs and any associated data, data files, and data structures are distributed across a networked computer system, such that the computer programs and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner through one or more processors or computers.
[0161] The functional grasping prediction method for dexterous hand robots and the training method for the corresponding prediction model according to the embodiments of this disclosure achieve efficient functional grasping prediction for dexterous hand robots.
[0162] On the other hand, the functional grasping prediction method for dexterous hand robots and the corresponding prediction model training method according to the embodiments of this disclosure, through the prediction model training method, trains the proposed large-scale dataset for scale diversity, solving the common problem that existing datasets are biased towards large objects due to the size constraints of robot hands.
[0163] On the other hand, the functional grasping prediction method for dexterous hand robots and the training method for the corresponding prediction model according to the embodiments of this disclosure can also enhance the efficiency and generalization ability of dexterous hand functional grasping generation.
[0164] While some embodiments of this disclosure have been disclosed and described, those skilled in the art will understand that modifications and variations may be made to these embodiments without departing from the concept and spirit of this disclosure, which is defined by the claims and their equivalents.
Claims
1. A functional grasping synthesis method for a dexterous hand robot, characterized in that, The functional capture and synthesis method includes: Acquire two-dimensional image data of a first object of a specific category; Perform 3D asset generation processing on the 2D image data of the first object to obtain the 3D asset data of the first object; The point cloud data of the first object corresponding to the three-dimensional asset data of the first object is input into a pre-trained first feature extractor for identifying corresponding parts between different instances of the same type of object, so as to obtain the first point cloud feature data of the corresponding first object. Perform a preset template matching process associated with the specific category on the point cloud feature data of the first object to obtain the availability data of the first object; By using an axis-aligned grasping initialization strategy and an availability-based functional grasping synthesis strategy, the point cloud data of the first object corresponding to the availability data of the first object is initialized and functional grasping is optimized to obtain the grasping configuration data of the first object.
2. The functional grasping and synthesis method according to claim 1, characterized in that, The step of performing 3D asset generation processing on the 2D image data of the first object to obtain the 3D asset data of the first object includes: Target recognition and image segmentation are performed on the two-dimensional image data of the first object to obtain the first processed image data of the first object; The first processed image data of the first object is input into a pre-trained 3D generative model to obtain the 3D asset data of the first object.
3. The functional grasping and synthesis method according to claim 1, characterized in that, The preset template associated with the specific category is obtained in the following way: Retrieves instance data of at least one object for each of multiple categories; For any object in each category, perform template matching annotation on the availability portion to obtain object availability data with availability portion labels for that object; Using the object availability data with availability partial labels for any object in each category, automatic feature matching is used to obtain the object availability data with availability partial labels for other objects in each category, thereby obtaining a preset template associated with the specific category. The preset template associated with the specific category is included in the object availability data of the specific object of the specific category with availability part labels in multiple categories.
4. The functional grasping and synthesis method according to claim 1, characterized in that, The steps of performing initialization and functional grasping optimization on the point cloud data of the first object corresponding to the first object's availability data, using an axis-aligned grasping initialization strategy and an availability-based functional grasping synthesis strategy to obtain the grasping configuration data of the first object, include: Using an axis-aligned grasping initialization strategy, the first object point cloud data corresponding to the first object availability data is subjected to axis-aligned grasping initialization processing to obtain the first object's initial grasping data; Using an availability-based functional crawling synthesis strategy, the initial crawling data of the first object is subjected to availability-based functional crawling optimization processing to obtain crawling configuration data that meets preset constraints, which is then used as the crawling configuration data of the first object.
5. The functional grasping and synthesis method according to claim 4, characterized in that, The preset constraints include at least one of the following: functional part distance loss constraint, grasping part distance loss constraint, force closure loss constraint, collision loss constraint between dexterous hand and object, and dexterous hand self-collision loss constraint.
6. A training method for a functional grasping prediction model for a dexterous hand robot, characterized in that, The training method includes: For each first object among at least one object of a specific category, perform the following processing: Retrieve the crawling configuration data for the first object; The grasping configuration data of the first object is input into the pre-trained second feature extractor for object availability recognition to obtain the grasping feature data of the first object; The point cloud data of the first object is input into a pre-trained third feature extractor for extracting preset details to obtain the second point cloud feature data of the first object. The first object's point cloud data and the first object's second point cloud feature data are input into a pre-trained fourth feature extractor for extracting preset non-details, to obtain the first object's third point cloud feature data. The grasping feature data of the first object and the third point cloud feature data of the first object are concatenated and then input into a preset encoder and a preset decoder to obtain a predictive grasping configuration for both the dexterous hand robot and the first object, so as to control the dexterous hand robot to grasp the first object through the predictive grasping configuration.
7. The training method according to claim 6, characterized in that, A second feature extractor for object availability identification, a third feature extractor for extracting preset details, a fourth feature extractor for extracting preset non-details, a preset encoder, and a preset decoder constitute a functional grasping prediction model for a dexterous hand robot.
8. The training method according to claim 7, characterized in that, The training method further includes: adjusting model parameters based on the predicted grasping configuration and surface point cloud data included in the 3D asset data associated with the at least one object, to obtain a trained functional grasping prediction model for the dexterous hand robot. The steps for obtaining the trained functional grasping prediction model for the dexterous hand robot include: The model parameters are adjusted by calculating the reconstruction loss and / or KL divergence for the functional grasping prediction model for the dexterous hand robot based on the predicted grasping configuration and the surface point cloud data included in the 3D asset data, to obtain the trained functional grasping prediction model for the dexterous hand robot.
9. A computer program product, characterized in that, The computer program product includes a computer program / instruction that, when executed by a processor, implements the functional grasping synthesis method for a dexterous hand robot as described in any one of claims 1 to 5, and the training method for a functional grasping prediction model for a dexterous hand robot as described in any one of claims 7 and 8.
10. A computing device, characterized in that, The computing device includes: at least one processor; at least one memory storing computer-executable instructions, wherein, when executed by the at least one processor, the computer-executable instructions cause the at least one processor to perform the functional grasping synthesis method for a dexterous hand robot as claimed in any one of claims 1 to 5 and the training method for a functional grasping prediction model for a dexterous hand robot as claimed in any one of claims 7 and 8.
Citation Information
Cited By
Grabbing posture generation method and device and electronic equipment
CN121696954A