Grasping pose generation method and apparatus, and electronic device

CN121696954BActive Publication Date: 2026-09-11北京中科慧灵机器人技术有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511930568.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-09-11
Estimated Expiration
2045-12-19

AI Technical Summary

Technical Problem

[0003]本申请实施例提供一种抓取姿态生成方法、装置及电子设备,以解决抓取的稳定性的问题

Benefits of technology

[0018] In this embodiment, by constructing a mapping relationship between the functional areas of the hand and the surface areas of the object, the dexterous hand can understand the parts of the object that the fingers should contact when grasping, thereby achieving a unity of hand coordination and grasping stability, and improving the success rate and stability of grasping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121696954B_ABST
    Figure CN121696954B_ABST
Patent Text Reader

Abstract

The application provides a grasping posture generation method and device and electronic equipment, and relates to the technical field of robots. The method comprises the following steps: acquiring first point cloud data of a first object to be grasped, wherein the first point cloud data comprises data for representing a surface region of the first object and constraint data of a second object on the first object; acquiring geometric structure data of a dexterous hand to be used for grasping the first object, wherein the geometric structure data comprises data of a plurality of functional regions of the dexterous hand; determining a mapping relationship between the plurality of functional regions of the dexterous hand and the surface region of the first object based on the first point cloud data and the geometric structure data; and determining a grasping posture of the dexterous hand according to the mapping relationship and constraint information of the dexterous hand and geometric information of the second object. According to the embodiment of the application, the mapping relationship between the functional regions of the hand and the surface region of the object is constructed, so that the dexterous hand can understand the part of the object that should be contacted by the fingers during grasping, and the success rate and stability of grasping are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotics, and in particular to a method, apparatus and electronic device for generating grasping postures. Background Technology

[0002] In the field of robotics, precise grasping is a prerequisite for ensuring grasping stability and operational flexibility. Currently, robot grasping posture generation mostly involves identifying the graspable area of ​​the target object and then using a dexterous hand to grasp that area. However, in actual grasping processes, issues such as fingertip slippage and object dropping can easily occur, leading to poor grasping stability. Summary of the Invention

[0003] This application provides a grasping posture generation method, apparatus, and electronic device to solve the problem of grasping stability.

[0004] To solve the above-mentioned technical problems, this application is implemented as follows:

[0005] In a first aspect, embodiments of this application provide a method for stabilizing the capture process, the method comprising:

[0006] Acquire first point cloud data of a first object to be captured. The first point cloud data includes data for characterizing the surface area of ​​the first object and constraint data of the second object on the first object. The first object is placed on the second object.

[0007] Obtain geometric structure data of the dexterous hand to grasp the first object, the geometric structure data including data of multiple functional areas of the dexterous hand, the multiple functional areas being regions divided according to grasping function;

[0008] Based on the first point cloud data and the geometric structure data, the mapping relationship between the multiple functional areas of the dexterous hand and the surface area of ​​the first object is determined;

[0009] Based on the mapping relationship, as well as the pre-acquired constraint information of the dexterous hand and the geometric information of the second object, the grasping posture of the dexterous hand is determined.

[0010] Secondly, embodiments of this application provide a stabilizing device for grasping, the device comprising:

[0011] The first acquisition module is used to acquire the first point cloud data of the first object to be captured. The first point cloud data includes data for characterizing the surface area of ​​the first object and constraint data of the second object on the first object. The first object is placed on the second object.

[0012] The second acquisition module is used to acquire the geometric structure data of the dexterous hand that is to grasp the first object. The geometric structure data includes data of multiple functional areas of the dexterous hand, and the multiple functional areas are areas divided according to the grasping function.

[0013] The first determining module is used to determine the mapping relationship between the multiple functional areas of the dexterous hand and the surface area of ​​the first object based on the first point cloud data and the geometric structure data.

[0014] The second determining module is used to determine the grasping posture of the dexterous hand based on the mapping relationship, the pre-acquired constraint information of the dexterous hand, and the geometric information of the second object.

[0015] Thirdly, embodiments of this application provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the grasping stability method described in the first aspect.

[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the grasping stability method described in the first aspect.

[0017] Fifthly, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps of the grasping stability method as described in the first aspect.

[0018] In this embodiment, by constructing a mapping relationship between the functional areas of the hand and the surface areas of the object, the dexterous hand can understand the parts of the object that the fingers should contact when grasping, thereby achieving a unity of hand coordination and grasping stability, and improving the success rate and stability of grasping. Attached Figure Description

[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart of a method for ensuring the stability of data capture provided in an embodiment of this application;

[0021] Figure 2 This is a schematic diagram of the structure of a gripping stability device provided in an embodiment of this application;

[0022] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] This application provides a method, apparatus, and electronic device for improving the stability of object grasping, thereby solving the problem of poor object grasping stability.

[0025] See Figure 1 , Figure 1 This is a flowchart of a method for ensuring the stability of data capture provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following steps:

[0026] Step 101: Obtain the first point cloud data of the first object to be captured. The first point cloud data includes data for characterizing the surface area of ​​the first object and constraint data of the second object on the first object. The first object is placed on the second object.

[0027] Step 102: Obtain the geometric structure data of the dexterous hand to grasp the first object. The geometric structure data includes data of multiple functional areas of the dexterous hand, which are areas divided according to the grasping function.

[0028] Step 103: Based on the first point cloud data and the geometric structure data, determine the mapping relationship between the multiple functional areas of the dexterous hand and the surface area of ​​the first object;

[0029] Step 104: Determine the grasping posture of the dexterous hand based on the mapping relationship, the pre-acquired constraint information of the dexterous hand, and the geometric information of the second object.

[0030] When the first object is placed on the second object, the second object constrains the first object. During the grasping process of the first object, the surface shape of the first object needs to be considered, and the interference of the constraint of the second object on the grasping needs to be considered.

[0031] The system acquires point cloud data of object points on the surface of a first object and constraint data of the second object on the first object. An environmental safety mask is then constructed based on a preset safety threshold. This environmental safety mask is a binary mask constructed based on desktop constraints and is used to mark whether points on the object's surface are within a safe distance from the desktop to avoid collision.

[0032] For example, estimate the plane of the tabletop (the second object) using plane fitting, and calculate the signed distance from each object point of the first object to the tabletop.

[0033] Furthermore, the geometric data of the dexterous hand used to grasp the first object can be acquired. The dexterous hand can pre-divide the area according to the grasping function of the object being grasped:

[0034] Fingertip Region: Used for grasping the edges and protrusions of objects;

[0035] Finger Pad Region: Used for grasping curved and flat surfaces of objects;

[0036] Palm Region: Used to support objects.

[0037] The above is just one way to divide the area. Further divisions can be made based on the grasping function, such as fingers and palm.

[0038] Different parts of a dexterous hand are used to perform different grasping functions.

[0039] The above-mentioned areas can be divided according to the functions in the actual grasping process, and the above-mentioned multiple functional areas can also be adjusted according to the shape and position of the object.

[0040] The first point cloud data and the geometric structure data of the dexterous hand are input into a pre-trained network model. Based on the output of the network model, the mapping relationship between multiple functional areas of the dexterous hand and the surface area of ​​the first object is determined. During training, the network model can learn the mapping relationship between the surface area of ​​the object and the functional areas of the hand during the grasping process, so that each object point not only has accessibility, but also has semantic correspondence information with the hand area.

[0041] For example, the region mapping module learns typical matching patterns such as "fingertip-edge", "finger pad-curved surface", and "palm-support surface" to match the mapping relationship between the functional areas of a dexterous hand and the surface areas of an object.

[0042] In some implementations, a matching algorithm can be used to determine the matching relationship between the functional area of ​​the dexterous hand and the surface area of ​​the object. The optimal matching degree can then be determined by designing reward functions such as successful grasping.

[0043] Pre-acquired constraint information of the dexterous hand, such as the physical constraint information of the fingers.

[0044] The grasping posture of the dexterous hand is determined based on the mapping relationship between the functional area of ​​the dexterous hand and the surface area of ​​the object, the physical constraint information of the dexterous hand, the constraint information of the second object on the first object, and the geometric information of the second object (including size information, shape information, position information, etc.).

[0045] In some implementations, an energy function is constructed to minimize energy loss (contact loss, collision loss, attitude naturalness loss, stability loss) during the grasping process. By minimizing energy loss, the grasping attitude is determined based on a mapping relationship.

[0046] In some implementations, a multi-objective optimization function is constructed to determine the grasping posture of the dexterous hand by minimizing the energy consumption of joint movements and maximizing grasping stability.

[0047] In this embodiment, by constructing a semantic mapping between the functional areas of the hand and the surface of an object, the dexterous hand can understand the parts of the object that the fingers are in contact with, achieving a unity of multi-finger coordination and grasping stability, and significantly improving the success rate and stability of grasping. By constructing constraints on a second object and an environmental safety mask, the problem of mutual interference between the traditional grasping posture and the support surface is solved, allowing the robotic hand to naturally adapt to the desktop environment.

[0048] Optionally, determining the mapping relationship between the multiple functional areas of the dexterous hand and the surface area of ​​the first object based on the first point cloud data and the geometric structure data includes:

[0049] Determine the environmental security mask of the first object based on the first point cloud data;

[0050] The environmental safety mask, the first point cloud data, and the geometric structure data are input into a pre-trained network model, and the mapping relationship between the multiple functional areas of the dexterous hand and the surface area of ​​the first object is determined based on the output of the network model.

[0051] Based on security threshold Constructing an environmental security mask

[0052]

[0053] Where Di is the signed distance from each point of the first object to the second object, calculated based on the first point cloud data. When the environmental safety mask is 1, it indicates that there is a safe area that can be captured; when the environmental safety mask is 0, it indicates that there is a danger area on the object surface and collision may occur.

[0054] The first point cloud data, the geometric structure data of the dexterous hand, and the environmental safety mask are input into a pre-trained network model. The network model outputs the contact probability distribution between the palm and each finger of the dexterous hand (i.e., the probability of the dexterous hand making point contact with the object). Based on the output of the network model (e.g., at least one of the contact probability and mapping relationship), the mapping relationship between multiple functional areas of the dexterous hand and the surface area of ​​the first object is obtained.

[0055] Because the network model learns the correspondence between hand functional regions and object regions through pre-training, that is, the semantic correspondence between hand regions and object regions, the network region can perform region matching based on the geometric structure data of the hand and the point cloud data of the object.

[0056] In some implementations, when training the network model, it can be trained to grasp objects of different shapes and orientations, so that the network model can learn the mapping relationship between the surface areas of objects of different shapes and the functional areas of the dexterous hand.

[0057] In some implementations, local aggregate features of points on the body surface are extracted, point features are aggregated into region features, and region features are matched with functional areas of the hand.

[0058] Optionally, the multiple functional areas of the dexterous hand include a first functional area, and the surface area of ​​the first object includes a first surface area;

[0059] The process of determining the mapping relationship between the multiple functional areas of the dexterous hand and the surface area of ​​the first object based on the output of the network model includes:

[0060] The contact probability between the first functional area of ​​the dexterous hand and the first surface area of ​​the first object is determined using the output of the network model, and the mapping relationship between the first functional area of ​​the dexterous hand and the first surface area of ​​the first object is determined based on the contact probability.

[0061] The first functional area is any one of the multiple functional areas of the dexterous hand, and the first surface area is any sub-area of ​​the surface area of ​​the first object.

[0062] The contact probability between the first functional area of ​​the dexterous hand and the first surface area of ​​the first object is the probability that the first functional area of ​​the hand contacts the first surface area when the dexterous hand grasps the first object. This example only uses the first functional area of ​​the dexterous hand and the first surface area of ​​the first object; the mapping relationship between other functional areas of the hand and other sub-surface areas of the first object can be described as above.

[0063] For example, output the probability distribution of contact between the palm and each finger:

[0064]

[0065] in, This is the first point of cloud data.

[0066] In some implementations, a Transformer-based contact map prediction network model outputs a multi-channel contact probability vector for each point on the object surface, representing the correspondence between each object point and different hand regions.

[0067] For example, when the contact probability between the first functional area and the first surface area is greater than a preset threshold, it indicates that there is a mapping relationship between the first functional area and the first surface area.

[0068] By obtaining the mapping relationship between the functional areas of the hand and the surface area of ​​the first object in the above manner, the object can be grasped according to the semantic correspondence between the functional areas of the hand and the surface of the object, thereby improving the grasping stability.

[0069] Furthermore, a predicted map of hand-object contact can be output based on the correspondence.

[0070] Optionally, the network model is used for:

[0071] Based on the environmental security mask, the first point cloud data, and the geometric structure data, determine the similarity between the first functional region and the first surface region;

[0072] Based on the similarity, the contact probability between the first functional area and the first surface area is determined.

[0073] After inputting the environmental safety mask, first point cloud data, and geometric structure data into the trained network model, the network model uses hand region embedding... Embedding with object points Define the similarity between multiple functional areas of the hand and the surface of the first object:

[0074]

[0075] The matching probability matrix (i.e., the contact probability between each functional region and each sub-surface region) is obtained by normalizing based on similarity:

[0076]

[0077] The contact probability (i.e. the probability of a mapping relationship) between the functional area of ​​the hand and the surface area of ​​the object can be obtained through the above method.

[0078] Optionally, the method further includes:

[0079] Acquire second point cloud data of a third object, the second point cloud data including data for characterizing the surface area of ​​the third object and constraint information of a fourth object on the third object, the third object being placed on the fourth object;

[0080] The second point cloud data and the geometric structure data of the dexterous hand that is to grasp the third object are input into the network model to be trained to obtain the contact probability between the functional area of ​​the dexterous hand and the surface area of ​​the third object.

[0081] Based on the difference between the contact probability and the pre-labeled data, the network model to be trained is trained to obtain the pre-trained network model, wherein the label data is used to characterize whether the functional area of ​​the dexterous hand is in contact with the object point of the third object.

[0082] Before using the network model to predict mapping relationships, the network model is trained. Second-point cloud data is acquired under various different scenarios.

[0083] The third object can include objects of various shapes, and the fourth object can also include objects of various shapes.

[0084] For example, a dexterous hand can be used to grasp a cylindrical cup (third object) placed on a table (fourth object); a dexterous hand can be used to grasp a square cup (third object) placed on a grid (fourth object). When the shape of the fourth object is different, multiple constraints need to be considered; and when the shape of the third object is different, the contact method between the hand area and the object surface area is different.

[0085] The surface shape of a third object is used to construct a second point cloud data. Complex surface features can be captured based on the point cloud structure by using a contact map prediction network that fuses local and global features.

[0086] Obtain the geometric structure data of the aforementioned dexterous hand, and input the data into the network model for training.

[0087] The network model employs cross-entropy supervised loss:

[0088]

[0089] Where Lce represents the cross-entropy loss function, which measures the difference between the predicted probability distribution and the true distribution;

[0090] k represents the index of the functional areas of the hand (such as fingertips, finger pads, palms, etc.);

[0091] i represents the index of a point on the object's surface;

[0092] yk,i The label indicates whether the hand region k is in contact with the object point i (1 indicates yes, 0 indicates no).

[0093] P k (i): Predicted probability, representing the probability that hand region k comes into contact with object point i.

[0094] By minimizing the loss function, the model's prediction results are made closer to the true labels, and the mapping relationship between the hand region and the object surface region is learned.

[0095] Through the above training process, the grasping generation is expanded from "prediction of the accessible area of ​​the object" to "semantic correspondence between the functional area of ​​the hand and the surface of the object", which enables the network model to learn the semantic contact between the hand area and the surface area of ​​the object during the grasping process, thereby obtaining the mapping relationship between the hand area and the surface area of ​​the object.

[0096] By fusing a unified representation of hand functional semantics and object surface geometric properties, the model's generalization ability across different object categories and placement conditions is enhanced. The entire framework enables end-to-end training, with contact prediction and pose optimization converging collaboratively, significantly improving the efficiency and interpretability of grasp generation.

[0097] Optionally, determining the grasping posture of the dexterous hand based on the mapping relationship, the pre-acquired constraint information of the dexterous hand, and the geometric information of the second object includes:

[0098] An objective function is constructed based on the pre-acquired constraint information of the dexterous hand, the geometric information of the second object, and the mapping relationship. The objective function includes at least one of contact matching loss, collision constraint loss, posture naturalness loss, and stability loss.

[0099] The grasping posture of the dexterous hand is determined based on the objective function.

[0100] In the attitude generation stage, an energy optimization-based approach is adopted, which combines contact map prediction results, dexterous hand kinematic constraints, and tabletop geometric information to optimize and solve the hand attitude.

[0101] The objective function consists of contact matching loss, collision constraint loss, pose naturalness loss, and stability loss. The contact matching loss is used to maintain the predicted contact. Figure 1 Consistency; Collision constraint loss is used to avoid collisions between the hand and the table or object; Posture naturalness loss is used to maintain a natural hand grip; Stability loss is used to improve the mechanical stability of the grip.

[0102] Based on the above four losses, an objective function is constructed to obtain the total energy loss:

[0103]

[0104] The four types of losses will be described separately below.

[0105] 1. Contact matching loss (bringing the fingertip close to the corresponding prediction point or candidate point):

[0106]

[0107] in, This indicates that it is determined based on the contact probability (i.e., the probability of having a mapping relationship). The set of candidate contact points (which can be selected) (The points with the highest probability in the middle) Indicates the corresponding weight (e.g., by probability). (Settings); k represents the configuration of the joint angles of the dexterous hand; t represents the coordinate position of the entire dexterous hand; t represents the position of the target object.

[0108] 2. Loss of naturalness in posture (preventing the production of extremely unnatural hand shapes):

[0109]

[0110] in, To relax the reference posture.

[0111] 3. Collision penalty (used to avoid intersections with the desktop or non-target areas), often using a smoothing penalty:

[0112]

[0113] in, Representing object points The distance to the nearest obstacle (object not in the target area or on the table). For safety margin. A function that indicates no collision between the hand and the table. This represents the predefined closest distance between the hand and the tabletop. This information can be obtained based on the constraints of the dexterous hand and the geometric information of the second object (e.g., dimensions, coordinates). By using the constraints of the second object and a collision loss function, the generated hand pose becomes feasible in physical space.

[0114] 4. Stability loss can be simplified to the loss at the desired point. Distance:

[0115]

[0116] in, c represents the coordinate position of the entire dexterous hand. kThis indicates the desired position of the object to be grasped, and a stable state is formed after the target point of the object is grasped.

[0117] By minimizing the objective function, the final hand grasping posture is obtained. Considering factors such as collision avoidance, contact matching, natural posture, and stability, this grasping posture improves stability when grasping objects. By introducing a matching probability matrix to determine the contact point, the grasping posture accurately matches the semantic mapping relationship, achieving a matching pattern grasping between the hand and the object. Solving for a stable posture in high-dimensional joint space using an energy optimization framework achieves a balance between contact accuracy and dynamic stability.

[0118] Through joint optimization, a grasping posture that conforms to physical constraints can be obtained in a high-dimensional joint space. Finally, the object with different shapes, sizes and placement states (e.g., vertical, horizontal, and tilted) is validated in a simulation platform (such as IsaacSim) to evaluate the executability of the generated posture and the grasping success rate, thus constructing a closed-loop system from generation to validation.

[0119] See Figure 2 , Figure 2 This is a schematic diagram of the structure of a grasping posture generation device provided in an embodiment of this application, as shown below. Figure 2 As shown, the grasping posture generation device 200 includes:

[0120] The first acquisition module 201 is used to acquire the first point cloud data of the first object to be captured. The first point cloud data includes data for characterizing the surface area of ​​the first object and constraint data of the second object on the first object. The first object is placed on the second object.

[0121] The second acquisition module 202 is used to acquire the geometric structure data of the dexterous hand to grasp the first object. The geometric structure data includes data of multiple functional areas of the dexterous hand, and the multiple functional areas are areas divided according to the grasping function.

[0122] The first determining module 203 is used to determine the mapping relationship between the multiple functional areas of the dexterous hand and the surface area of ​​the first object based on the first point cloud data and the geometric structure data.

[0123] The second determining module 204 is used to determine the grasping posture of the dexterous hand based on the mapping relationship, the pre-acquired constraint information of the dexterous hand, and the geometric information of the second object.

[0124] Optionally, the first determining module is specifically used for:

[0125] Determine the environmental security mask of the first object based on the first point cloud data;

[0126] The environmental safety mask, the first point cloud data, and the geometric structure data are input into a pre-trained network model, and the mapping relationship between the multiple functional areas of the dexterous hand and the surface area of ​​the first object is determined based on the output of the network model.

[0127] Optionally, the multiple functional areas of the dexterous hand include a first functional area, and the surface area of ​​the first object includes a first surface area;

[0128] The first determining module is specifically used for:

[0129] The contact probability between the first functional area of ​​the dexterous hand and the first surface area of ​​the first object is determined using the output of the network model, and the mapping relationship between the first functional area of ​​the dexterous hand and the first surface area of ​​the first object is determined based on the contact probability.

[0130] Optionally, the network model is used for:

[0131] Based on the environmental security mask, the first point cloud data, and the geometric structure data, determine the similarity between the first functional region and the first surface region;

[0132] Based on the similarity, the contact probability between the first functional area and the first surface area is determined.

[0133] Optionally, the device further includes:

[0134] The third acquisition module is used to acquire the second point cloud data of the third object. The second point cloud data includes data for characterizing the surface area of ​​the third object and constraint information of the fourth object on the third object. The third object is placed on the fourth object.

[0135] The input module is used to input the second point cloud data and the geometric structure data of the dexterous hand that is to grasp the third object into the network model to be trained, so as to obtain the contact probability between the functional area of ​​the dexterous hand and the surface area of ​​the third object.

[0136] The training module is used to train the network model to be trained based on the difference between the contact probability and the pre-labeled data, so as to obtain the pre-trained network model, wherein the label data is used to characterize whether the functional area of ​​the dexterous hand is in contact with the object point of the third object.

[0137] Optionally, the second determining module is specifically used for:

[0138] An objective function is constructed based on the pre-acquired constraint information of the dexterous hand and the geometric information of the second object. The objective function includes at least one of contact matching loss, collision constraint loss, posture naturalness loss, and stability loss.

[0139] Based on the objective function and the mapping relationship, the grasping posture of the dexterous hand is determined.

[0140] The grasping posture generation device can achieve Figure 1 The various processes implemented in the method embodiments can achieve the same technical effect, and will not be described again here to avoid repetition.

[0141] like Figure 3 As shown, this application embodiment also provides an electronic device 300, including: a processor 301, a memory 302, and a program stored in the memory 302 and executable on the processor 301. When the program is executed by the processor 301, it implements the various processes of the above-described grasping posture generation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0142] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described grasping posture generation method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0143] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described... Figure 1 The various processes of the method embodiments shown can achieve the same technical effect, and will not be described again here to avoid repetition.

[0144] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0145] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0146] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A grasp pose generation method, characterized by, The method includes: Acquire first point cloud data of a first object to be captured. The first point cloud data includes data for characterizing the surface area of ​​the first object and constraint data of the second object on the first object. The first object is placed on the second object. Obtain geometric structure data of the dexterous hand to grasp the first object, the geometric structure data including data of multiple functional areas of the dexterous hand, the multiple functional areas being regions divided according to grasping function; Based on the first point cloud data and the geometric structure data, the mapping relationship between the multiple functional areas of the dexterous hand and the surface area of ​​the first object is determined; Based on the mapping relationship, as well as the pre-acquired constraint information of the dexterous hand and the geometric information of the second object, the grasping posture of the dexterous hand is determined; The step of determining the mapping relationship between the multiple functional areas of the dexterous hand and the surface area of ​​the first object based on the first point cloud data and the geometric structure data includes: Determine the environmental security mask of the first object based on the first point cloud data; The environmental safety mask, the first point cloud data, and the geometric structure data are input into a pre-trained network model, and the mapping relationship between the multiple functional areas of the dexterous hand and the surface area of ​​the first object is determined based on the output of the network model. The dexterous hand has multiple functional regions including a first functional region, and the surface region of the first object includes a first surface region; determining the mapping relationship between the multiple functional regions of the dexterous hand and the surface region of the first object based on the output of the network model includes: The contact probability between the first functional area of ​​the dexterous hand and the first surface area of ​​the first object is determined using the output of the network model, and the mapping relationship between the first functional area of ​​the dexterous hand and the first surface area of ​​the first object is determined based on the contact probability. The network model is used for: Based on the environmental security mask, the first point cloud data, and the geometric structure data, determine the similarity between the first functional region and the first surface region; Based on the similarity, the contact probability between the first functional area and the first surface area is determined.

2. The method of claim 1, wherein, The method further includes: Acquire second point cloud data of a third object, the second point cloud data including data for characterizing the surface area of ​​the third object and constraint information of a fourth object on the third object, the third object being placed on the fourth object; The second point cloud data and the geometric structure data of the dexterous hand that is to grasp the third object are input into the network model to be trained to obtain the contact probability between the functional area of ​​the dexterous hand and the surface area of ​​the third object. Based on the difference between the contact probability and the pre-labeled data, the network model to be trained is trained to obtain the pre-trained network model, wherein the label data is used to characterize whether the functional area of ​​the dexterous hand is in contact with the object point of the third object.

3. The method of any one of claims 1-2, wherein, Determining the grasping posture of the dexterous hand based on the mapping relationship, the pre-acquired constraint information of the dexterous hand, and the geometric information of the second object includes: An objective function is constructed based on the pre-acquired constraint information of the dexterous hand, the geometric information of the second object, and the mapping relationship. The objective function includes at least one of contact matching loss, collision constraint loss, posture naturalness loss, and stability loss. Based on the objective function, the grasping posture of the dexterous hand is determined.

4. A grasp pose generation apparatus characterized by comprising: include: The first acquisition module is used to acquire the first point cloud data of the first object to be captured. The first point cloud data includes data for characterizing the surface area of ​​the first object and constraint data of the second object on the first object. The first object is placed on the second object. The second acquisition module is used to acquire the geometric structure data of the dexterous hand that is to grasp the first object. The geometric structure data includes data of multiple functional areas of the dexterous hand, and the multiple functional areas are areas divided according to the grasping function. The first determining module is used to determine the mapping relationship between the multiple functional areas of the dexterous hand and the surface area of ​​the first object based on the first point cloud data and the geometric structure data. The second determining module is used to determine the grasping posture of the dexterous hand based on the mapping relationship, the pre-acquired constraint information of the dexterous hand, and the geometric information of the second object. The first determining module is specifically used for: Determine the environmental security mask of the first object based on the first point cloud data; The environmental safety mask, the first point cloud data, and the geometric structure data are input into a pre-trained network model, and the mapping relationship between the multiple functional areas of the dexterous hand and the surface area of ​​the first object is determined based on the output of the network model. The dexterous hand has multiple functional areas including a first functional area, and the surface area of ​​the first object includes a first surface area; the first determining module is specifically used for: The contact probability between the first functional area of ​​the dexterous hand and the first surface area of ​​the first object is determined using the output of the network model, and the mapping relationship between the first functional area of ​​the dexterous hand and the first surface area of ​​the first object is determined based on the contact probability. The network model is used for: Based on the environmental security mask, the first point cloud data, and the geometric structure data, determine the similarity between the first functional region and the first surface region; Based on the similarity, the contact probability between the first functional area and the first surface area is determined.

5. An electronic device, comprising: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the grasping posture generation method as described in any one of claims 1 to 3.

6. A computer readable storage medium characterized by, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the grasping posture generation method as described in any one of claims 1 to 3.

7. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the grasping posture generation method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Dexterous hand self-adaptive grabbing method, dexterous hand control system and storage medium

    CN121132718A