Grabbing type prior-based dexterous hand grabbing method and equipment and medium
By introducing grab type discrimination, contact heat map generation and pose optimization modules in the dexterous hand grab system, the problems of poor crawling adaptability and insufficient stability in the prior art are solved, and a more efficient and stable crawling process is achieved.
Patent Information
- Application Number
- CN202510254534.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The existing dexterous hand-grabbing method is difficult to accurately judge the appropriate type of grab when facing objects with complex shapes or irregular surfaces, resulting in inaccurate grasping posture and even failure.
By introducing a grab type discrimination module, a contact heat map generation module and a grab pose optimization module, combining point cloud information and grab type prior information, a suitable grab pose is generated. The specific steps include: obtaining point cloud information of the object, judging the grab type, generating a contact heat map, generating a preliminary grab pose, and generating the final grab pose through optimization strategies.
It improves the applicability and stability of dexterous hand grabbing, and can more accurately generate grasping postures suitable for different objects shapes and characteristics, significantly improving the accuracy and success rate of grabbing.
Smart Images

Figure CN120095812A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot grasping, and in particular to a grasping type prior-based dexterous hand grasping method, equipment and medium. Background Art
[0002] With the rapid development of artificial intelligence and robotics, dexterous hands are playing an increasingly important role in fields such as industrial automation, service robots, and medical rehabilitation. The core function of dexterous hands is to be able to flexibly adjust the grasping strategy according to the shape, material, and task requirements of the object to achieve stable and efficient grasping. The core issue of dexterous hand grasping technology is how to accurately predict the grasping posture of an object based on the three-dimensional representation of a given object. The grasping posture of an object usually includes three parts: (1) the displacement of the dexterous hand relative to the center of the object to be grasped; (2) the rotation of the dexterous hand relative to the center of the object to be grasped; and (3) the angles of each joint of the dexterous hand.
[0003] Currently, most dexterous hand grasping methods generate grasping postures based on the geometric information of the object (such as point clouds). However, existing methods often ignore the prior information of the grasping type of the object. For example, when faced with an object with a convex surface (such as a vase), existing methods may generate inappropriate grasping postures, such as using the palm to grasp the body of the vase instead of using the fingers to grasp the top of the vase. When grasping an object, humans can quickly select an appropriate grasping posture based on the shape, size, and characteristics of the object. Therefore, if the robot can draw on the prior information of this grasping type, it will greatly improve the stability and accuracy of the grasping. Summary of the invention
[0004] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide a dexterous hand grasping method, device and medium based on grasping type prior.
[0005] The first technical solution adopted by the present invention is:
[0006] A dexterous hand grasping method based on grasping type prior includes the following steps:
[0007] Obtaining point cloud information of the object to be grasped, and determining the grasping type according to the point cloud information; wherein the grasping type includes a palm-type grasping type and a finger-type grasping type;
[0008] Generate a contact heat map between the object to be grasped and the dexterous hand based on the point cloud information and the obtained grasping type;
[0009] The contact heat map is used as a constraint to generate a preliminary grasping posture based on the point cloud information, namely the displacement, rotation and joint angle of the dexterous hand relative to the object to be grasped;
[0010] According to the grasping type, the corresponding optimization strategy is used to optimize the preliminary grasping posture to generate the final grasping posture. Furthermore, a pre-trained grasping type discrimination module is used to predict the grasping type. The input of the grasping type discrimination module is the point cloud information of the object, and the output is the corresponding grasping type;
[0011] The PointNet++ model is used to extract the local features of the object point cloud, the Transformer model is used to capture the global features of the object point cloud, the graph convolutional network (GNN) is used to extract the topological features from the topological structure of the point cloud, the various features are fused through the self-attention mechanism, and the output grasping type is predicted.
[0012] Furthermore, the expression of the local feature is:
[0013] F local =f PointNet++ (P)
[0014] In the formula, F local Represents the local features of the object, and P is the point cloud of the object;
[0015] The expression of the global feature is:
[0016] F global =f Transformer (P)
[0017] In the formula, F global Represents the local features of the object, and P is the point cloud of the object;
[0018] The expression of the topological feature is:
[0019] F topo =f GNN (P)
[0020] In the formula, F topo Represents the local features of the object, and P is the point cloud of the object;
[0021] The expression of the features after fusion using the self-attention mechanism is:
[0022] F final =Attention(F local ,F global ,F topo )
[0023] In the formula, F final Represents the final feature after fusion;
[0024] Through the classification head (MLP), the grasping type of the object is output based on the fused features:
[0025]
[0026] In the formula, is the predicted grasp type, i.e., palm grasp type or finger grasp type.
[0027] Furthermore, a pre-trained contact heatmap generation module is used to generate the contact heatmap;
[0028] For palm-type grasping and finger-type grasping, two conditional variational autoencoders (CVAEs) with the same structure are trained to generate contact heat maps of finger-type and palm-type grasping objects.
[0029] The conditional variational autoencoder generates a contact heat map based on the object point cloud information, and the expression is:
[0030]
[0031] In the formula, (H contacy |P) is the conditional probability distribution of the contact heat map, H contact is the contact heat map, P is the object point cloud, μ is the mean, and σ is the standard deviation;
[0032] The contact heat map is used to constrain the key parts of the contact area on the surface of the object to optimize the grasping posture of the dexterous hand. The objective function is:
[0033]
[0034] Where z is a hidden variable, p(H|z) is the output probability distribution of the generator; KL represents the Kullback-Leibler divergence; q(z|P) is, p(z) is; the generated contact heat map H = {h 1 ,h 2 ,…,h m}, where h j ∈[0,1] represents the contact strength of each contact point in the point cloud, which is used to constrain the grasping posture of the dexterous hand; For expected operation.
[0035] Furthermore, a pre-trained grasping posture generation module is used to generate preliminary grasping postures;
[0036] The working mode of the grasping posture generation module is as follows:
[0037] The point cloud information of the object is serialized and embedded to obtain a high-dimensional feature representation, which is expressed as:
[0038] F embedding =f embedding (P)
[0039] In the formula, F embedding is the embedded feature, P is the object point cloud;
[0040] The embedded features are spatially compressed through the grid pooling layer, and the expression is:
[0041] F pooled =f GridPool (F embedding )
[0042] In the formula, F pooled is the feature after pooling;
[0043] Conditional position encoding and attention mechanism are used to optimize the spatial representation of features. The expression is:
[0044] F att =f Attention (F pooled )
[0045] In the formula, F att Features optimized by the attention mechanism;
[0046] The grasping pose is output by the prediction head, and the expression is:
[0047]
[0048] In the formula, The predicted grasping pose includes the object’s displacement, rotation, and joint angles of the dexterous hand.
[0049] Furthermore, the loss function is calculated and back-propagated:
[0050]
[0051] In the formula, λ 1 ,λ 2 ,λ 3 is the weight coefficient, which is used to balance the impact of different loss terms; represents displacement loss, T is the real displacement, To predict displacement, ∥∥ represents the Euclidean norm; represents the rotation displacement, θ is the angle between the predicted rotation and the actual rotation; represents the joint angle loss, J i is the true joint angle, To predict the joint angles, M is the number of joints of the dexterous hand.
[0052] Furthermore, according to the grasping type, the initial grasping posture is optimized by using a corresponding optimization strategy to generate a final grasping posture, including:
[0053] For palm-type grasping, first optimize the displacement and rotation so that the palm is as close to the grasping part as possible to provide sufficient grasping force; then optimize the joint angle of the fingers so that the fingers assist in grasping; finally, reduce unreasonable hand-object penetration through optimization;
[0054] For finger-type grasping, the joint angles of the fingers are directly optimized to make the fingers as close to the object as possible to ensure sufficient grasping force; unreasonable hand-object penetration is reduced through optimization.
[0055] Furthermore, the formula for optimizing the palm grip type is:
[0056] E palm =Distance(palm,object)
[0057]
[0058] In the formula, E palm represents the distance measurement between the palm and the object, palm represents the palm part of the dexterous hand, object represents the object to be grasped; R is the rotation of the dexterous hand relative to the center of the object, t is the rotation of the dexterous hand relative to the center of the object, Δ(R, t) represents the relative position change between the palm and the object under the combined action of rotation R and displacement t; E pen represents the hand-object penetration measure, Indicated in posture The dexterous hand area under t,q, p represents a sampling point on the palm or finger surface in the dexterous hand model, d p,object represents the distance from point p to the surface of the object point cloud; q is the joint angle of the dexterous hand, and Δ(R, t, q) represents the rotation The relative position change between the palm and the object under the combined effect of displacement t and joint angle q; η is the learning rate;
[0059] The formula for optimizing the finger-type grip type is:
[0060] E finger{i} =Distance(finger{i},object)
[0061]
[0062] In the formula, E finger{i} Represents the distance measurement between the finger and the object, finger{i} represents the i-th finger, a total of 5 fingers; qpos finger{i} represents the joint angle of the i-th finger in the dexterous hand, Δ(qpos finger{i} ) represents the joint angle qpos finger{i} The relative position change between the finger and the object under the action.
[0063] The second technical solution adopted by the present invention is:
[0064] An electronic device comprises a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by the processor to implement a dexterous hand grasping method based on grasping type prior as described above.
[0065] The third technical solution adopted by the present invention is:
[0066] A computer-readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, wherein the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a dexterous hand grasping method based on grasping type prior as described above.
[0067] The fourth technical solution adopted by the present invention is:
[0068] A computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above method.
[0069] The beneficial effects of the present invention are as follows: the present invention improves the applicability and stability of dexterous hand grasping by guiding the generation of grasping posture by grasping type priori, and is suitable for the application of robot dexterous hands in various common grasping scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the embodiments of the present invention or the drawings of related technical solutions in the prior art are introduced below. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0071] Figure 1 is a flowchart of the steps of a dexterous hand grasping method based on grasping type prior in an embodiment of the present invention;
[0072] Figure 2 is a structural diagram of a crawl type discrimination module in an embodiment of the present invention;
[0073] Figure 3It is a structural schematic diagram of a grasping posture generation module in an embodiment of the present invention. DETAILED DESCRIPTION
[0074] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0075] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., and orientations or positional relationships indicated are based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.
[0076] In the description of the present invention, "several" means one or more, "more" means more than two, "greater than", "less than", "exceed" etc. are understood as not including the number itself, and "above", "below", "within" etc. are understood as including the number itself. If there is a description of "first" or "second", it is only used for the purpose of distinguishing the technical features, and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features.
[0077] In the description of the present invention, unless otherwise clearly defined, terms such as setting, installing, connecting, etc. should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the present invention based on the specific content of the technical solution.
[0078] At present, most of the existing dexterous hand grasping methods rely on the geometric shape and point cloud data of the object to directly generate the grasping posture. However, these methods have the problems of poor grasping adaptability and insufficient grasping stability. Specifically, when faced with objects with complex shapes or irregular surfaces, existing methods often cannot effectively determine the appropriate grasping type (such as finger-type grasping or palm-type grasping), resulting in inaccurate grasping posture or even failure. At the same time, most methods also ignore the reasonable planning of the contact area between the object and the dexterous hand, resulting in uneven distribution of grasping force, which further affects the stability of grasping.
[0079] In order to solve these problems, the present invention proposes a grasping method for dexterous hands based on grasping type prior, which fundamentally improves the adaptability and stability of grasping by introducing grasping type discrimination, contact heat map generation and grasping posture optimization modules. First, through the grasping type discrimination module, the grasping type (finger type or palm type) of the object can be accurately judged to ensure that the dexterous hand selects the most appropriate grasping method. Then, using the contact heat map generation module, accurate contact area information is provided for the grasping posture, so that the grasping force can be reasonably distributed during the grasping process, avoiding unnecessary hand-object penetration. Finally, combined with the grasping posture optimization module, the present invention further optimizes the displacement, rotation and joint angle to ensure the stability and efficiency of the grasping posture. Compared with the traditional method, the innovation of the present invention lies in modular design and multi-level optimization, which not only improves the grasping adaptability under different object types, but also significantly improves the stability and success rate during the grasping process, especially when facing objects with complex shapes and irregular surfaces, it has significant advantages.
[0080] Example 1
[0081] like Figure 1 As shown, this embodiment provides a dexterous hand grasping method based on grasping type prior, which combines the point cloud information of the object and the grasping type prior information to achieve efficient and stable grasping posture generation. The method specifically includes the following steps:
[0082] S1. Obtain point cloud information of the object to be grasped, and determine the grasping type according to the point cloud information; the grasping types include palm-type grasping type and finger-type grasping type.
[0083] In this embodiment, the grasping type is determined by constructing a grasping type determination module, the input of which is the point cloud information of the object to be grasped, and the output is its grasping type (finger type or palm type).
[0084] Specifically, see Figure 2 , construct a PointNet++ model to extract the local features of the object point cloud, the Transformer model to capture the global features, and construct a graph convolutional network (GNN) to learn features from the topological structure of the point cloud. Finally, the various features are fused through the self-attention mechanism and combined into a grasping type discrimination module to output the grasping type.
[0085] As an implementation method, the crawl type identification module works as follows:
[0086] S21. Extract features based on point cloud information.
[0087] 1) Use the PointNet++ network to extract local features of the object point cloud. PointNet++ extracts local features of different scales through hierarchical aggregation operations. The formula is as follows:
[0088] F local =f PointNet++ (P)
[0089] Among them, F local represents the local features of the object, and P is the point cloud of the object.
[0090] 2) Use the Transformer model to capture global features and long-range dependencies in point cloud information. The formula is as follows:
[0091] F global =f Transformer (P)
[0092] Among them, F global represents the local features of the object, and P is the point cloud of the object.
[0093] 3) Use graph convolutional network (GNN) to process the topological structure of point cloud and extract topological features. The formula is as follows:
[0094] F topo =f GNN (P)
[0095] Among them, F topo represents the local features of the object, and P is the point cloud of the object.
[0096] S22, feature fusion, using the self-attention mechanism to fuse the features from PointNet++, Transformer and GNN. The formula is as follows:
[0097] F final =Attention(F local ,F global ,F topo )
[0098] Among them, F final Represents the final feature after fusion.
[0099] S23, prediction output: Through the classification head (MLP), the object grasping type is output based on the fused features:
[0100]
[0101] in is the predicted grasp type (palm or finger).
[0102] S24. Calculate the loss function and backpropagate, using cross-entropy loss to measure the difference between the predicted category and the true category:
[0103]
[0104] Where N is the number of training samples, y i is the true category label of the i-th sample, is the corresponding predicted probability.
[0105] In this embodiment, the role of PointNet++ is to extract the local geometric features of the point cloud, retain the detailed information of the object surface, and is suitable for processing the complex shapes of the object surface; the role of Transformer is to capture the global contextual relationship of the object point cloud, and is especially suitable for modeling long-distance dependent features, such as the overall shape and structure of the object; the role of GNN is to learn the morphological features of the object from the topological structure of the point cloud, such as the connection relationship between points, which is particularly critical for the structured feature extraction of complex objects; finally, the role of introducing the self-attention mechanism for feature fusion is to fuse local, global and topological features, ensure the information complementarity between different features, and improve the classification module's ability to discriminate the grasping type.
[0106] S2. Generate a contact heat map between the object to be grasped and the dexterous hand based on the point cloud information and the obtained grasping type.
[0107] In this embodiment, the contact heat map is generated by the contact heat map generation module. Training the contact heat map generation module: Two conditional variational autoencoders (CVAEs) with the same structure are trained for palm-type grasping objects and finger-type grasping objects respectively. Through the input of object point cloud information, the contact heat map of the corresponding grasping type object and the dexterous hand is predicted respectively as a constraint for the subsequent grasping posture generation and post-optimization process.
[0108] The contact heat map generation module uses two conditional variational autoencoders (CVAEs) to generate contact heat maps for finger-type and palm-type grasping objects, respectively. The CVAE model generates contact heat maps based on the object point cloud information. The formula is expressed as:
[0109]
[0110] Among them, p(H contact |P) is the conditional probability distribution of the contact heat map, μ is the mean, and σ is the standard deviation. The training goal is to maximize the variational lower bound (ELBO), the formula is:
[0111]
[0112] in, For the expected operation, D KL is the Kullback-Leibler divergence, which measures the difference between the true distribution and the generated distribution.
[0113] S3. Using the contact heat map as a constraint, generate a preliminary grasping posture based on the point cloud information, i.e., the displacement, rotation, and joint angle of the dexterous hand relative to the object to be grasped.
[0114] In this embodiment, a preliminary grasping posture is generated by a grasping posture generation module. Training the grasping posture generation module: Based on the point cloud information of the object, it sequentially passes through the serialization layer, embedding layer, grid pooling layer, order randomization layer, conditional position encoding, layer normalization, attention layer, layer normalization, multi-layer perceptron, prediction head and other operations to predict and output the grasping posture, that is, the displacement and rotation of the dexterous hand relative to the object and the joint angle of the dexterous hand.
[0115] As an optional implementation, see Figure 3 ,The working steps of the grasping pose generation module include:
[0116] S41. The object point cloud data is processed by serialization and embedding to obtain a high-dimensional feature representation. The formula is:
[0117] F embedding =f embedding (P)
[0118] Among them, F embedding is the embedded feature.
[0119] S42, spatially compress the point cloud data through the grid pooling layer:
[0120] F pooled =f GridPool (F embedding )
[0121] S43. Optimizing the spatial representation of features using conditional position encoding and attention mechanism:
[0122] F att =f Attention (F pooled )
[0123] Among them, F att are the features optimized by the attention mechanism.
[0124] S44, output the grasping posture through the prediction head:
[0125]
[0126] in, The predicted grasping pose includes the object’s displacement, rotation, and joint angles of the dexterous hand.
[0127] S45, calculate the loss function and back propagate:
[0128]
[0129] In the formula, λ 1 ,λ 2 ,λ 3 is the weight coefficient, which is used to balance the impact of different loss terms; represents displacement loss, T is the real displacement, To predict displacement, ∥∥ represents the Euclidean norm; represents the rotation displacement, θ is the angle between the predicted rotation and the actual rotation; represents the joint angle loss, J i is the true joint angle, To predict the joint angles, M is the number of joints of the dexterous hand.
[0130] S4. According to the grasping type, the corresponding optimization strategy is used to optimize the preliminary grasping posture to generate the final grasping posture.
[0131] Exemplarily, the generated preliminary grasping posture is post-optimized: different post-optimization strategies are applied to optimize the posture according to the grasping type:
[0132] 1) For the palm type, first optimize the displacement and rotation so that the palm is as close to the grasping part as possible to provide sufficient grasping force; then optimize the joint angles of the fingers to assist grasping; finally, optimize to reduce unreasonable hand-object penetration; the specific formula is as follows:
[0133] E palm =Distance(palm,object)
[0134]
[0135] In the formula, E palm represents the distance measurement between the palm and the object, palm represents the palm part of the dexterous hand, object represents the object to be grasped; R is the rotation of the dexterous hand relative to the center of the object, t is the rotation of the dexterous hand relative to the center of the object, Δ(R, t) represents the relative position change between the palm and the object under the combined action of rotation R and displacement t; E pen represents the hand-object penetration measure, Indicated in posture The dexterous hand area under t,q, p represents a sampling point on the palm or finger surface in the dexterous hand model, d p,object represents the distance from point p to the surface of the object point cloud; q is the joint angle of the dexterous hand, and Δ(R, t, q) represents the rotation The relative position change between the palm and the object under the combined effect of displacement t and joint angle q; η is the learning rate.
[0136] 2) For the finger type, directly optimize the joint angles of the fingers to make them as close to the object as possible to ensure sufficient gripping force; after that, optimize the hand-object penetration problem as well. The specific formula is as follows:
[0137] E finger{i} =Distance(finger{i},object)
[0138]
[0139] In the formula, E finger{i} Represents the distance measurement between the finger and the object, finger{i} represents the i-th finger, a total of 5 fingers; qpos finger{i} represents the joint angle of the i-th finger in the dexterous hand, Δ(qpos finger{i} ) represents the joint angle qpos finger{i} The relative position change between the finger and the object under the action.
[0140] In this embodiment, the benefit of the post-optimization strategy is that it is tailored for different grasping types, especially for the displacement and rotation optimization of palm-type grasping, which significantly improves the rationality of the grasping force distribution and effectively reduces the hand-object penetration problem in grasping failures.
[0141] As a preferred implementation, the method of this embodiment also includes the step of constructing a training set: selecting an existing dexterous grasping data set, generating a contact heat map by analyzing the contact between the object point cloud and various parts of the dexterous hand, and labeling the grasping type according to the main grasping force providing part in the grasping data, including finger-type grasping and palm-type grasping, that is, the main grasping force is provided by the fingers or palm respectively, to prepare data for subsequent training of various modules.
[0142] Optionally, the above-mentioned dexterous grasping dataset is an image dataset publicly available on the Internet, such as DexGraspNet. Specifically, DexGraspNet is a simulation dataset that provides 1126 types of objects and their corresponding grasping postures.
[0143] Specifically, the step of constructing a training set includes the following steps A1-A3:
[0144] A1. Point cloud data acquisition: By acquiring the point cloud data of the object, the 3D shape information of each object is used for model training in subsequent steps.
[0145] A2. Contact heat map generation: The contact heat map is calculated using the contact conditions between the object point cloud and the dexterous hand. The contact heat map represents the contact intensity between different areas of the object surface and the dexterous hand. Specifically, the following formula is used:
[0146]
[0147] Among them, H contact is the contact heat map, N is the number of points on the surface of the object, P i is the coordinate of the i-th point cloud point, S j is the jth part of the dexterous hand, and σ(·) is the function for calculating the contact strength.
[0148] A3. Grasping type annotation: According to the main part that provides the grasping force of the object in the hand-object grasping in the dataset, the object grasping type is annotated, which is mainly divided into palm-type grasping and finger-type grasping, that is, the force of grasping the object is mainly borne by the palm or fingers respectively.
[0149] As a preferred implementation, the method of this embodiment also includes the step of testing each trained module. A test set is selected from the dexterous grasping data set, and the point cloud information of the object to be grasped in the test set is input into each module that has been trained as described above according to the preset logic, and the corresponding preliminary grasping posture is output: During the test, the object point cloud information is first input, and the grasping type of the object is judged by the grasping type discrimination module. According to the grasping type, the corresponding contact heat map generation module is selected to generate the corresponding contact heat map. The contact heat map is then input as a constraint into the grasping posture generation module to predict the appropriate grasping posture.
[0150] Specifically, the test step includes the following steps B1-B5:
[0151] B1. Input point cloud data: Select the point cloud data P of the object to be grasped from the test set test as input.
[0152] B2. Grasp type identification: Input the point cloud data into the trained grasp type identification module and output the predicted grasp type:
[0153] y test =TypeClassifier(P test )
[0154] Among them, y test ∈{Finger,Palm}.
[0155] B3. Select the contact heat map generation module: According to the predicted grasping type y test , select the corresponding contact heat map generation module (CVAE model), and convert the point cloud data P test Input the CVAE model and generate the contact heat map:
[0156] H test =CVAE ytest (P test )
[0157] Among them, Htest is the resulting contact heat map.
[0158] B4. Grasping posture prediction: Contact heat map H test and point cloud data P test Together, they are input into the grasping posture generation module to predict the initial grasping posture:
[0159]
[0160] in, Includes predicted displacements, rotations, and joint angles.
[0161] B5. Use the generated contact heat map to optimize the grasping posture. First calculate the actual contact heat map between the current grasping posture and the object:
[0162]
[0163] E contact =MSE(H test ,H gt )
[0164]
[0165] In the formula, H gt represents the contact heat map between the actual dexterous hand and the object to be grasped; E contact represents the contact error metric, MSE represents the mean square error, which is used to measure the difference between the predicted contact heat map and the actual contact heat map; Represents the current dexterous hand posture parameters, including the rotation of the dexterous hand relative to the center of the object to be grasped The displacement t and joint angle q relative to the center of the object to be grasped, and η is the learning rate.
[0166] In summary, the method of the present invention solves the problems of poor grasping adaptability and insufficient grasping stability in the prior art by introducing modules such as grasping type discrimination, contact heat map generation, and grasping posture optimization. Through the synergistic effect of the grasping type discrimination module and the contact heat map generation module, the grasping posture generation process is optimized, ensuring the distribution of grasping force and the reasonable configuration of the dexterous hand. At the same time, the accuracy and efficiency of grasping are further improved through a more reasonable post-optimization strategy. Compared with the prior art, it includes at least the following advantages and beneficial effects:
[0167] (1) The present invention improves the rationality and stability of the grasping posture and can generate an appropriate grasping posture according to the shape and characteristics of the object.
[0168] (2) The present invention optimizes the grasping posture generation process through the synergistic effect of the grasping type discrimination module and the contact heat map generation module, ensuring the distribution of grasping force and the reasonable configuration of the dexterous hand.
[0169] (3) The method of the present invention can adapt to the grasping of different types of objects and further improve the grasping accuracy and efficiency through post-optimization strategy.
[0170] Example 2
[0171] An embodiment of the present invention further provides an electronic device, the electronic device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the following Figure 1 A dexterous hand grasping method based on grasp type prior is shown.
[0172] It is understood that the memory may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.
[0173] The processor may include one or more processing cores. The processor uses various interfaces and lines to connect the various parts of the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory, and calling data stored in the memory. Optionally, the processor can be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor can integrate one or a combination of a central processing unit (CPU) and a modem. Among them, the CPU mainly processes the operating system and application programs; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor, but implemented separately through a chip.
[0174] Since the electronic device is an electronic device corresponding to the dexterous hand grasping method based on grasping type prior of an embodiment of the present invention, and the principle of solving the problem by the electronic device is similar to that of the method, the implementation of the electronic device can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0175] Example 3
[0176] The embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, the at least one program, the code set or instruction set is loaded and executed by a processor to implement the following Figure 1 A dexterous hand grasping method based on grasp type prior is shown.
[0177] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0178] Since the storage medium is a storage medium corresponding to a dexterous hand grasping method based on grasping type prior in an embodiment of the present invention, and the principle of solving the problem by the storage medium is similar to that of the method, the implementation of the storage medium can refer to the implementation process of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0179] Example 4
[0180] In some possible implementations, various aspects of the method of the embodiment of the present invention may also be implemented in the form of a program product, which includes a program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of a dexterous hand grasping method based on grasping type prior according to various exemplary embodiments of the present application described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0181] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0182] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0183] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable ordinary technicians in the field to understand the content of the present invention and implement it accordingly, and they cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made based on the essence of the content of the present invention should be included in the protection scope of the present invention.
Claims
1. A dexterous hand grasping method based on grasping type prior, characterized in that: The following steps are involved: Obtain the point cloud information of the object to be grasped, and determine the grasping type based on the point cloud information; The grasping types include palm grasping type and finger grasping type; Generate a contact heat map between the object to be grasped and the dexterous hand based on the point cloud information and the obtained grasping type; The contact heat map is used as a constraint to generate a preliminary grasping posture based on the point cloud information, namely the displacement, rotation and joint angle of the dexterous hand relative to the object to be grasped; According to the grasping type, the corresponding optimization strategy is used to optimize the preliminary grasping posture and generate the final grasping posture.
2. The dexterous hand grasping method based on grasping type prior according to claim 1, characterized in that: A pre-trained grasping type discrimination module is used to predict the grasping type, wherein the input of the grasping type discrimination module is the point cloud information of the object, and the output is the corresponding grasping type; The PointNet++ model is used to extract the local features of the object point cloud, the Transformer model is used to capture the global features of the object point cloud, the graph convolutional network is used to extract the topological features from the topological structure of the point cloud, the various features are fused through the self-attention mechanism, and the output grasping type is predicted.
3. The dexterous hand grasping method based on grasping type prior according to claim 2 is characterized in that: The expression of the local feature is: F local =f PointNet++ (P) In the formula, F local Represents the local features of the object, and P is the point cloud of the object; The expression of the global feature is: F global =f Transformer (P) In the formula, F global Represents the local features of the object, and P is the point cloud of the object; The expression of the topological feature is: F topo =f GNN (P) In the formula, F topo Represents the local features of the object, and P is the point cloud of the object; The expression of the features after fusion using the self-attention mechanism is: F final =Attention(F local ,F global ,F topo ) In the formula, F fimal Represents the final feature after fusion; Through the classification head, the grasping type of the object is output based on the fused features: In the formula, The predicted crawl type.
4. The dexterous hand grasping method based on grasping type prior according to claim 1 is characterized in that: A pre-trained contact heatmap generation module is used to generate contact heatmaps; For palm-type grasping and finger-type grasping, two conditional variational autoencoders with the same structure are trained to generate contact heat maps of finger-type and palm-type grasping objects. The conditional variational autoencoder generates a contact heat map based on the object point cloud information, and the expression is: In the formula, (H contact |P) is the conditional probability distribution of the contact heat map, H contact is the contact heat map, P is the object point cloud, μ is the mean, and σ is the standard deviation; The contact heat map is used to constrain the key parts of the contact area on the surface of the object to optimize the grasping posture of the dexterous hand. The objective function is: Where z is a hidden variable, p(H|z) is the output probability distribution of the generator; KL represents the Kullback-Leibler divergence; q(z|P) is, p(z) is; the generated contact heat map H = {h1,h2,…,h m }, where h j ∈[0,1] represents the contact strength of each contact point in the point cloud, which is used to constrain the grasping posture of the dexterous hand; For expected operation.
5. The dexterous hand grasping method based on grasping type prior according to claim 1, characterized in that: A pre-trained grasping posture generation module is used to generate preliminary grasping postures; The working mode of the grasping posture generation module is as follows: The point cloud information of the object is serialized and embedded to obtain a high-dimensional feature representation, which is expressed as: F embedding =f embedding (P) In the formula, F embedding is the embedded feature, P is the object point cloud; The embedded features are spatially compressed through the grid pooling layer, and the expression is: F pooled =f GridPool (F embedding ) In the formula, F pooled is the feature after pooling; Conditional position encoding and attention mechanism are used to optimize the spatial representation of features. The expression is: F att =f Attention (F pooled ) In the formula, F att Features optimized by the attention mechanism; The grasping pose is output by the prediction head, and the expression is: In the formula, The predicted grasping pose includes the object’s displacement, rotation, and joint angles of the dexterous hand.
6. The dexterous hand grasping method based on grasping type prior according to claim 5, characterized in that: Calculate the loss function and backpropagate: Where λ1, λ2, λ3 are weight coefficients used to balance the impact of different loss terms; represents displacement loss, T is the real displacement, To predict displacement, || || represents the Euclidean norm; represents the rotation displacement, θ is the angle between the predicted rotation and the actual rotation; represents the joint angle loss, J i is the true joint angle, To predict the joint angles, M is the number of joints of the dexterous hand.
7. The dexterous hand grasping method based on grasping type prior according to claim 1 is characterized in that: According to the grasping type, the corresponding optimization strategy is used to optimize the preliminary grasping posture to generate the final grasping posture, including: For palm-type grasping, first optimize the displacement and rotation so that the palm is as close to the grasping part as possible to provide sufficient grasping force; then optimize the joint angle of the fingers so that the fingers assist in grasping; finally, reduce unreasonable hand-object penetration through optimization; For finger-type grasping, the joint angles of the fingers are directly optimized to make the fingers as close to the object as possible to ensure sufficient grasping force; unreasonable hand-object penetration is reduced through optimization.
8. The dexterous hand grasping method based on grasping type prior according to claim 7, characterized in that: The formula for optimizing the palm grip type is: E palm =Distance(palm,object) In the formula, E palm It represents the distance measurement between the palm and the object, palm represents the palm part of the dexterous hand, and object represents the object to be grasped; R is the rotation of the dexterous hand relative to the center of the object, t is the rotation of the dexterous hand relative to the center of the object, Δ(R, t) represents the relative position change between the palm and the object under the combined action of rotation R and displacement t; E pen represents the hand-object penetration measure, Indicated in posture The dexterous hand area under t,q, p represents a sampling point on the palm or finger surface in the dexterous hand model, d p,object represents the distance from point p to the surface of the object point cloud; q is the joint angle of the dexterous hand, and Δ(R, t, q) represents the rotation The relative position change between the palm and the object under the combined effect of displacement t and joint angle q; η is the learning rate; E finger{i} =Distance(finger{i},object) In the formula, E finger{i} Represents the distance measurement between the finger and the object, finger{i} represents the i-th finger, a total of 5 fingers; qpos finger{i} represents the joint angle of the i-th finger in the dexterous hand, Δ(qpos finger{i} ) represents the joint angle qpos finger{i} The relative position change between the finger and the object under the action.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Control method for five-finger dexterous hand of intelligent explosive-handling robot
CN108638054A
Grasping of an object by a robot based on grasp strategy determined using machine learning model(s)
CN111788041A
Robot grabbing mode selection method and system based on multiple constraint conditions
CN112809680A
Method and system for generating grabbing posture of dexterous hand guided by limited grabable area
CN116704160A
Intra-class multi-specification tool oriented dexterous hand grabbing posture generation method
CN119089989A
Cited By
Dexterous hand self-adaptive grabbing method, dexterous hand control system and storage medium
CN121132718A
Grabbing posture generation method and device and electronic equipment
CN121696954A