Grabbing pose estimation method and related device

By using Gaussian heat map guidance method, combined with the YOLOv10 object detection model and the HGGD model, the target artifacts are identified and positioned to generate the optimal grasping point, which solves the problem of low precision in grasping pose estimation in the existing technology and achieves more efficient grasping accuracy.

CN120047398APending Publication Date: 2025-05-27WUYI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510094830.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art lacks consideration of global and local features in grasping pose estimation, resulting in low efficiency and low registration accuracy, especially in the absence of target features, the grasping accuracy is not high.

Method used

Using the Gaussian heat map guidance method, the target artifact is identified and positioned to generate the optimal crawling point through the YOLOv10 object detection model and the six-degree of freedom crawling based on efficient heat map guidance, thereby improving the crawling accuracy.

Benefits of technology

More accurate target workpiece identification and positioning is achieved, grasping accuracy is improved, and the problems of low efficiency and low registration accuracy in the existing technology are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047398A_ABST
    Figure CN120047398A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a grabbing pose estimation method and a related device, and the method comprises the steps: obtaining a target workpiece image which is an RGBD image; inputting the target workpiece image into a YOLOv10 target detection model for target identification to obtain a target detection result; the target detection result is input into an HGGD model for grabbing pose estimation, the estimated grabbing pose of the target workpiece is obtained, the HGGD model comprises GHM and NMG, the GHM is used for processing RGBD images, efficient CNN is used for extracting semantic features, and multiple grabbing heat maps are generated to serve as guidance of a grabbing area; the NMG aggregates the local points to the grabbing area by using the grabbing heat map, and detects grabbing in the grabbing area through a lightweight point encoder. On the basis, the target workpiece can be recognized and positioned more accurately, the area where the target workpiece is located can be selected, the optimal grabbing point is calculated, and therefore the grabbing precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of visual grasping technology, and in particular to a grasping posture estimation method and related devices. Background Art

[0002] Due to the rapid development of depth cameras, 6D grasp pose estimation has begun to be applied in various fields. People are increasingly interested in the development and application of algorithms for 6D grasp pose estimation, such as image recognition, monitoring and interpretation of object pose and motion. 6D grasp pose estimation is an important technology for pose estimation of unordered workpieces and has been applied to tasks such as grasp detection and target detection, but it lacks consideration of global and local features.

[0003] Since the surface features of objects are difficult to characterize, the locked objects have reflections and textures, and lack consideration of global and local features. Therefore, the existing technology has low efficiency and low registration accuracy. In the absence of target features, the registration performance is poor, resulting in low grasping accuracy. Therefore, how to improve grasping accuracy has become a technical problem that needs to be solved urgently. Summary of the invention

[0004] The embodiment of the present invention provides a grasping posture estimation method and related devices, which can more accurately identify and locate the target workpiece. Using Gaussian heat map guidance, the area where the target workpiece is located can be framed and the optimal grasping point can be calculated, thereby improving the grasping accuracy.

[0005] In a first aspect, an embodiment of the present invention provides a grasping posture estimation method, comprising:

[0006] Acquire a target workpiece image, wherein the target workpiece image is an RGBD image;

[0007] Input the target workpiece image into a pre-trained YOLOv10 target detection model for target recognition to obtain a target detection result;

[0008] The target detection result is input into the six-degree-of-freedom grasping detection HGGD model based on efficient heat map guidance to estimate the grasping pose, so as to obtain the estimated grasping pose of the target workpiece, wherein the HGGD model includes a grasping point heat map model GHM and a non-uniform multi-grasp point generator NMG, wherein the GHM is used to process the RGBD image, extract semantic features using efficient CNN, and generate multiple grasping heat maps as guidance for the grasping area; the NMG uses the grasping heat map to aggregate local points into the grasping area, and detects grasping within the grasping area through a lightweight point encoder.

[0009] In some embodiments, inputting the target workpiece image into a pre-trained YOLOv10 target detection model for target recognition to obtain a target detection result includes:

[0010] Performing image preprocessing on the target workpiece image to obtain a preprocessed image, wherein the image preprocessing includes image size adjustment and image normalization processing;

[0011] Input the preprocessed image into the YOLOv10 target detection model for feature extraction to obtain target image features;

[0012] Dividing the preprocessed image into a plurality of grids, and predicting a plurality of bounding boxes and confidences corresponding to the bounding boxes for each of the grids, wherein the bounding boxes are used to locate the target in the image, and the confidences are used to represent the probability that the target exists in the bounding boxes;

[0013] Filtering target bounding boxes that are higher than a confidence threshold, and adjusting the filtered target bounding boxes to obtain the adjusted target bounding boxes;

[0014] The adjusted target bounding box and the corresponding category information are output as the target detection result.

[0015] In some embodiments, inputting the target detection result into a six-degree-of-freedom grasping detection HGGD model based on efficient heat map guidance to estimate the grasping pose, and obtaining the estimated grasping pose of the target workpiece, comprises:

[0016] Inputting the target detection result into the GHM for processing to generate a capture heat map, wherein the capture heat map includes a confidence heat map and a grid attribute heat map;

[0017] The NMG efficiently aggregates the plurality of the grasping regions under the guidance of the grasping heat map;

[0018] The NMG predicts the remaining grip point rotation attributes using the grip point attributes of each grid in the grid attribute heat map;

[0019] The NMG processes the captured heat map and point cloud into useful local features;

[0020] The NMG refines the rotation attribute of the gripping point through the local features to generate a plurality of gripping points;

[0021] An estimated grasping posture of the target workpiece is obtained based on the multiple grasping points.

[0022] In some embodiments, the GHM is an encoder-decoder model comprising two output branches, wherein the two output branches are a confidence branch and an attribute branch, the confidence branch is used to construct the confidence heat map, and the attribute branch is used to generate the grid attribute heat map.

[0023] In some embodiments, the NMG includes two parts: heat map guided region aggregation and non-uniform multi-gripper generator. Under the guidance of the grasping heat map, the NMG aggregates local points into graspable regions for use by the non-uniform multi-gripper generator.

[0024] In some embodiments, obtaining the estimated grasping posture of the target workpiece based on the multiple grasping points includes:

[0025] Generate grasping points in the heat map-guided area where the target workpiece is located;

[0026] Using a three-dimensional point cloud generation technology to convert multiple grasping points generated by the guide heat map of the target workpiece into grasping points under the three-dimensional point cloud;

[0027] Generate the corresponding 3D point cloud gripper according to the grasping point, and select the target 3D point cloud gripper with the highest grasping quality score;

[0028] Convert the image coordinate system of the target 3D point cloud gripper into a base coordinate system based on the robot;

[0029] The base coordinate system is transmitted to the robot via Ethernet to execute the corresponding object grasping.

[0030] In some embodiments, the training method of the YOLOv10 target detection model includes:

[0031] Constructing a target loss function, wherein the target loss function is determined according to a classification loss function, a coordinate loss function, and a confidence loss function, wherein the coordinate loss function is used to measure the position difference between the predicted bounding box and the true bounding box, and the confidence loss function is calculated using a binary cross entropy loss;

[0032] The YOLOv10 target detection model is trained based on the target loss function to obtain the trained YOLOv10 target detection model.

[0033] In a second aspect, an embodiment of the present invention further provides a grasping posture estimation device, the device comprising:

[0034] An acquisition module is used to acquire a target workpiece image, wherein the target workpiece image is an RGBD image;

[0035] A recognition module is used to input the target workpiece image into a pre-trained YOLOv10 target detection model for target recognition to obtain a target detection result;

[0036] A grasping module is used to input the target detection result into the six-degree-of-freedom grasping detection HGGD model based on efficient heat map guidance to estimate the grasping pose, so as to obtain the estimated grasping pose of the target workpiece, wherein the HGGD model includes a grasping point heat map model GHM and a non-uniform multi-grasp point generator NMG, the GHM is used to process the RGBD image, use efficient CNN to extract semantic features, and generate multiple grasping heat maps as guidance for the grasping area; the NMG uses the grasping heat map to aggregate local points into the grasping area, and detects grasping within the grasping area through a lightweight point encoder.

[0037] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the grasping pose estimation method as described in the first aspect when executing the computer program.

[0038] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the grasping pose estimation method as described in the first aspect.

[0039] According to the grasping pose estimation method and related device provided by the embodiment of the present invention, the grasping pose estimation method includes: obtaining a target workpiece image, which is an RGBD image; inputting the target workpiece image into a pre-trained YOLOv10 target detection model for target recognition to obtain a target detection result; inputting the target detection result into a six-degree-of-freedom grasping detection HGGD model guided by an efficient heat map to estimate the grasping pose, and obtain an estimated grasping pose of the target workpiece, wherein the HGGD model includes a grasping point heat map model GHM and a non-uniform multi-grasping point generator NMG, GHM is used to process the RGBD image, extract semantic features using an efficient CNN, and generate multiple grasping heat maps as a guide for the grasping area; NMG uses the grasping heat map to aggregate local points into the grasping area, and detects grasping in the grasping area through a lightweight point encoder. Based on this, the embodiment of the present invention can more accurately identify and locate the target workpiece, and with the guidance of the Gaussian heat map, the area where the target workpiece is located can be framed, and the optimal grasping point can be calculated, thereby improving the grasping accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a flow chart of a grasping posture estimation method provided by an embodiment of the present invention;

[0041] Figure 2 It is a schematic diagram of the YOLOv10 target detection model framework provided by an embodiment of the present invention;

[0042] Figure 3 It is a HGGD model framework diagram provided by an embodiment of the present invention;

[0043] Figure 4 It is a schematic diagram of HGGD model input and output provided by an embodiment of the present invention;

[0044] Figure 5 is a 6D pose estimation framework diagram provided by an embodiment of the present invention;

[0045] Figure 6 is a schematic diagram of a grasping posture estimation device provided by an embodiment of the present invention;

[0046] Figure 7 is a schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0048] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification and claims and the following drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0049] In the embodiments of the present invention, the words "further", "exemplarily" or "optionally" are used to indicate examples, illustrations or descriptions, and should not be interpreted as being more preferred or more advantageous than other embodiments or designs. The use of the words "further", "exemplarily" or "optionally" is intended to present related concepts in a specific way.

[0050] In order to more conveniently describe the working principle of the embodiment of the present invention later, an introduction to relevant technical scenarios is first given below.

[0051] Due to the rapid development of depth cameras, 6D grasp pose estimation has begun to be applied in various fields. People are increasingly interested in the development and application of algorithms for 6D grasp pose estimation, such as image recognition, monitoring and interpretation of object pose and motion. 6D grasp pose estimation is an important technology for pose estimation of unordered workpieces and has been applied to tasks such as grasp detection and target detection, but it lacks consideration of global and local features.

[0052] Since the surface features of objects are difficult to characterize, the locked objects have reflections and textures, and lack consideration of global and local features. Therefore, the existing technology has low efficiency and low registration accuracy. In the absence of target features, the registration performance is poor, resulting in low grasping accuracy. Therefore, how to improve grasping accuracy has become a technical problem that needs to be solved urgently.

[0053] Based on this, the present invention provides a grasping pose estimation method and related devices. Among them, the grasping pose estimation method includes: obtaining a target workpiece image, which is an RGBD image; inputting the target workpiece image into a pre-trained YOLOv10 target detection model for target recognition to obtain a target detection result; inputting the target detection result into a six-degree-of-freedom grasping detection HGGD model guided by an efficient heat map to estimate the grasping pose, and obtain an estimated grasping pose of the target workpiece, wherein the HGGD model includes a grasping point heat map model GHM and a non-uniform multi-grasping point generator NMG, GHM is used to process the RGBD image, extract semantic features using an efficient CNN, and generate multiple grasping heat maps as a guide for the grasping area; NMG uses the grasping heat map to aggregate local points into the grasping area, and detects grasping within the grasping area through a lightweight point encoder. Based on this, the embodiment of the present invention can more accurately identify and locate the target workpiece, and with the guidance of the Gaussian heat map, the area where the target workpiece is located can be framed, and the optimal grasping point can be calculated, thereby improving the grasping accuracy.

[0054] The embodiments of the present invention are further described below in conjunction with the accompanying drawings.

[0055] like Figure 1 As shown, Figure 1 is a flowchart of a grasping pose estimation method provided by an embodiment of the present invention. The grasping pose estimation method may include but is not limited to steps S101 to S103.

[0056] Step S101: Acquire a target workpiece image, where the target workpiece image is an RGBD image;

[0057] Step S102: input the target workpiece image into a pre-trained YOLOv10 target detection model for target recognition to obtain a target detection result;

[0058] Step S103: Input the target detection result into the six-degree-of-freedom grasping detection HGGD model based on efficient heat map guidance to estimate the grasping pose, and obtain the estimated grasping pose of the target workpiece, wherein the HGGD model includes a grasping point heat map model GHM and a non-uniform multi-grasp point generator NMG. GHM is used to process RGBD images, use efficient CNN to extract semantic features, and generate multiple grasping heat maps as guidance for the grasping area; NMG uses the grasping heat map to aggregate local points into the grasping area, and detects grasping within the grasping area through a lightweight point encoder.

[0059] It is understandable that in order to solve the accuracy of 6D pose estimation technology registration in the scene of stacking, texture and reflective objects, the present invention realizes feature point and template registration of noisy and incomplete picture data with higher accuracy. The proposed network encodes the information of RGB image and RGBD image into feature vectors, an effective local grab generator, and combines the grab Gaussian heat map as a guide to infer in a global to local and semantic to point manner. Specifically, the present invention uses Gaussian coding and a grid-based strategy to predict the grab heat map as guidance information, aggregates local points into grabbable areas, and provides global semantic information. In addition, the present invention also designs a novel non-uniform anchor sampling mechanism to improve the accuracy and diversity of grabbing. In terms of target recognition technology, the yo lov10 target recognition framework with the highest adaptability is adopted, which can realize rapid recognition of workpiece objects, and has high accuracy, which can ensure accuracy in multi-target recognition.

[0060] It is understandable that since two technologies need to be called at the same time, a script is used to mobilize the two technologies to avoid conflicts between the two technical environments. The script mobilizes the two technologies as follows: install two algorithm operating environments on the Ubuntu 20.04 version, namely hggd and yo lov10, and then call the main file in the yo lov10 algorithm in the script file. After startup, the recognized RGB image and the coordinate file of the grabbing coordinate frame are stored in a specific folder. Then the script starts the main file of the HGGD algorithm, starts the generation of the gripper, and selects the gripper with the highest grasping quality score, and converts the gripper information into the grasping information sent to the robot.

[0061] It is understandable that the present invention uses a pre-trained YOLOv10 target detection model to identify the target workpiece. The YOLOv10 target detection model is as follows: Figure 2As shown. The recognition process of YOLOv10 mainly involves image preprocessing, feature extraction, bounding box prediction, and post-processing steps (such as non-maximum suppression, but YOLOv10 avoids the use of NMS during inference). The following is a detailed explanation of the YOLOv10 recognition process:

[0062] 1. Image Preprocessing

[0063] Input resizing: The input image is resized to the size required by the model to fit the input layer of the model.

[0064] Normalization: Normalize the image and scale the pixel values ​​to a specific range (such as 0 to 1) to ensure the stability and convergence of the model training.

[0065] 2. Feature Extraction

[0066] Backbone network: YOLOv10 uses an enhanced version of CSPNet (Cross-Stage Partial Network) as the backbone network, which is responsible for extracting features from the input image. These features contain key information in the image for subsequent object detection.

[0067] Neck network: It is designed to aggregate features of different scales and pass them to the head network. It contains a PAN (Path Aggregation Network) layer for effective multi-scale feature fusion to further enhance the expressiveness of features.

[0068] 3. Bounding Box Prediction

[0069] Grid partitioning: The input image is divided into multiple grids, and each grid is responsible for predicting the object in its area.

[0070] Predicting bounding boxes: For each grid, YOLOv10 predicts multiple bounding boxes and corresponding confidences. These bounding boxes are used to locate objects in the image, and the confidence indicates the probability of the object existing in the bounding box.

[0071] Label allocation:

[0072] One-to-Many Head: During training, multiple predictions are generated for each object to provide rich supervision signals and improve learning accuracy.

[0073] One-to-One Head: During inference, a single best prediction is generated for each object, thus avoiding the use of NMS, reducing latency and improving efficiency.

[0074] Dual assignment strategy: combines one-to-one and many-to-one label assignment methods to optimize the training process so that the model can maintain high efficiency during inference.

[0075] 4. Post-processing

[0076] Although YOLOv10 avoids the traditional NMS step during inference, it still requires post-processing of the prediction results to filter out the final detection results. This may include:

[0077] Confidence filtering: Filter out high-confidence bounding boxes based on the confidence threshold.

[0078] Bounding box adjustment: Further adjust the filtered bounding box to improve positioning accuracy.

[0079] 5. Output test results

[0080] The processed bounding box and the corresponding category information are output as the final detection result.

[0081] It can be understood that for the pose estimation framework, the monocular RGBD image and camera characteristics c are used to efficiently generate high-quality and rich grip G in a novel global-to-local semantic-to-point manner. The method of the present invention does not directly process the observed point cloud, but inputs the RGBD image to encode the grip heat map as a graspable area guidance. Under the guidance of the key heat map, only the semantic and geometric representations of these areas are extracted and fused. On this basis, a novel local grasp generator is designed to realize heat map-guided six-degree-of-freedom grasp detection (HGGD), realizing high-quality and diverse real-time detection of grasps. The six-degree-of-freedom grasp detection HGGD model framework based on efficient heat map guidance of the present invention is as follows Figure 3 As shown in Figure 2, the model consists of two submodules: the grasp heatmap model (GHM) and the non-uniform multi-grasp generator (NMG). The GHM preprocesses the input RGBD image, extracts semantic features using an efficient CNN, and further generates four grasp heatmaps as guidance. Gaussian coding is used to encode the grasp ground truth to assist in more accurately locating the graspable area. The grid-based strategy transforms the discontinuous pixel-by-pixel regression into a regression based on predictions about neighborhood similarity, which makes the heatmap generation more robust for each local area.

[0082] It can be understood that the input and output of the HGGD model framework of the six-degree-of-freedom grasping detection based on efficient heat map guidance of the present invention are as follows: Figure 4As shown. GHM takes a monocular RGBD image as input and generates a grasp confidence heatmap Qc and a mesh attribute heatmap (Qθ, Qw, Qd). Then, guided by the heatmap, NMG transfers the depth image to the point cloud through the camera essence c for regional aggregation. Feature fusion and point encoder extract regional features that incorporate semantic information from GHM. Finally, the multi-grasp point generator combined with a novel non-uniform anchor sampling mechanism outputs grasp points using fused features. NMG uses the heatmap generated by GHM as a guide to aggregate local points into graspable areas and detect grasps through a lightweight point encoder for each area. The present invention proposes an NMG non-uniform anchor sampling mechanism to improve the grasp quality by better fitting the ground truth distribution. In addition, the present invention adopts a novel semantic-point feature fusion module to detect grasps more robustly.

[0083] It can be understood that GHM is an encoder-decoder model with two output branches, the confidence branch is intended to construct the confidence heat map Qc, and the attribute branch is intended to generate the attribute heat map (Qθ, Qw, Qd). The present invention adopts Gaussian coding and grid-based strategies to decouple the different characteristics between heat maps. Figure 4 As shown in Figure 2, the ground truth 6-DOF grasp is projected onto the image plane and encoded as a heatmap (Q^c, Q^θ, Q^w, Q d). The Gaussian encoding strategy uses a 2D Gaussian kernel to encode the projected ground truth grasp center before training. This approach effectively highlights the grasp center without ignoring nearby pixels, as nearby pixels will also serve as useful guidance for further grasp detection. The value of a pixel (u,v) in the confidence heatmap used for training can be obtained by the following formula:

[0084]

[0085] where (u0, v0) represents the center point of a grasp ground truth and σg is the standard deviation that depends on the width of each grasp. Under the supervision of Q^c, the confidence branch applies pixel-wise classification to predict Qc.

[0086] The grid-based strategy proposed in this paper encodes and predicts the grasp attributes (θ, w, d) within a specific local grid, instead of directly performing pixel-by-pixel regression. Due to similar geometric structures, grasp attributes usually have high similarity in these areas. Therefore, by fully utilizing the similarity of adjacent grasp points, more robust grasp point attribute prediction can be achieved. Specifically, the full-size image is divided into Hr×Wr grid cells with a side length of r. Based on the oriented anchor box mechanism, for each grid cell, ka introduces multiple oriented anchors with uniformly sampled angles. Therefore, the ground truth θ can be assigned to the nearest anchor point. The number distribution of anchor points in each grid is obtained, and then the sigmoid function is applied to it to obtain Q^θ. In addition, the average normalized w, d are calculated in the grid to obtain the ground truth attribute heat map (Q^w, Q^d). Under the supervision of these heat maps, in a patch-like manner, the attribute branch predicts Qθthrough a combination of anchor classification and offset regression, and estimates (Qw, Qd) through direct regression.

[0087] It should be noted that encoding grasp points as pixel-by-pixel rectangles has two drawbacks. First, they do not highlight the importance of the most significant grasp probability at the center point. Second, the ground truth attribute heatmaps (Q^θ, Q^w, Q^d) are not as smooth as the confidence heatmap Q^c due to the relatively dense grasp annotations in cluttered scenes. In contrast, the designed GHM can highlight the grasp center and predict more robust grasp attributes, especially in cluttered scenes.

[0088] It can be understood that for grasping pose estimation, the 6D pose estimation framework is as follows Figure 5 As shown, NMG takes heatmaps and scene point clouds as inputs, and efficiently aggregates multiple graspable local regions under the guidance of heatmaps. Then, using the grasping point attributes in each grid, NMG predicts the remaining grasping point rotation attributes, and refines the previously generated grasping point rotation attributes through local features to generate multiple grasping points. The non-uniform anchor point sampling mechanism proposed in this invention improves the grasping quality, and the novel semantic-to-point feature fusion contributes to the robustness of the detected grasps. According to different functions, the overall structure of NMG can be divided into two parts: heatmap-guided region aggregation and non-uniform multi-gripper generator.

[0089] For heatmap-guided region aggregation: The first part of NMG processes heatmaps and point clouds into useful local features, which includes two steps: region aggregation and feature fusion.

[0090] First, region aggregation, guided by the GHM heatmap, aggregates local points into graspable regions for use by the subsequent multi-grasp generator. Specifically, the grasp confidence heatmap is downsampled to Hr×Wr by bilinear interpolation, with the same shape as the attribute heatmap. Then the top kcenter grid with the highest prediction confidence is selected, containing a total of kcenter local peaks as the region center. This grid-based selection suppresses the center density to reduce the aggregation of duplicate regions. During training, kcenter is set to a large number to ensure that most graspable local regions are extracted. During inference, kcenter can be easily adjusted to achieve grasp detection with different coverage.

[0091] After the above grasping detection, grasping points will be generated in the heat map guidance area where the workpiece is located, and the multiple grasping points generated by the workpiece's guiding heat map will be converted into grasping points under the three-dimensional point cloud through the three-dimensional point cloud generation technology, and the corresponding 3D point cloud gripper will be generated according to the grasping points, and the 3D point cloud gripper with the highest grasping quality score will be selected. The image coordinate system of the gripper will be converted into a base coordinate system based on the robot, and the base coordinate system will be transmitted to the robot through Ethernet to execute the corresponding object for grasping.

[0092] It is understandable that the YOLOv10 network structure is used in the target detection training, which uses a network training method that combines the anchor-free idea. The initial learning rate is 10 -6 , the batch size is 64. The experimental platform is configured with an NVIDIA GeForce RTX 3090GPU and an Intel i5 CPU for training, the CUDA version is 11.8, the Python version is 3.8, and the networks are all implemented on the PyTorch architecture. The loss function of the YOLOv10 network plays a vital role in the target detection task. It guides the learning process of the model to ensure that the model can accurately predict the category, location, and confidence of the target. The loss function includes classification loss function, coordinate loss function, and confidence loss function. In terms of classification loss function, BCE Loss is used for classification. The formula of BCE Loss is:

[0093]

[0094] The coordinate loss function is mainly used to measure the position difference between the predicted bounding box and the real bounding box, including the center point coordinates and width and height. YOLOv10 may combine the center point coordinate loss and width and height loss into a unified coordinate loss and pass a weight coefficient (such as λ coord ) to balance the impact of these two losses on the total loss. Therefore, the overall formula of coordinate loss is as follows:

[0095] L corrd =λ coord (L cent +L wh ) (3)

[0096] Among them, λ coord Is a hyperparameter used to adjust the weight of coordinate loss in the total loss. During training, since the true IoU (i.e., the IoU between the labeled box and the predicted box) is known, the confidence loss function focuses on the prediction accuracy of the object. The confidence can be expressed as:

[0097]

[0098] However, during training, since the IoU is known, the model mainly learns P(Object), which is the probability that the object exists in the predicted box. The confidence loss function is usually calculated using binary cross entropy loss. For each predicted box, if it is a positive sample (that is, it has sufficient IoU overlap with a real box), the object label is 1; if it is a negative sample, the object label is 0. The confidence loss formula can be roughly expressed as:

[0099]

[0100] Where is the number of cells in the grid, B is the number of predicted boxes per cell, cij is the actual objectness label (0 or 1) of the Jth predicted box in the Ith cell, and cij is the objectness probability predicted by the model, expressed as:

[0101] L=L Qc +a×L cls +b×L reg +L anchor +c×L offset (6)

[0102] L QC represents the pixel-wise cross entropy loss C between the predicted confidence QC and the encoded ground truth. Another focal loss L cls For supervised multi-label classification, theta learning is performed. In addition, a masked smooth L1 loss L reg All regression problems in GHM are adopted. L anchor represents the local grasp rotation anchor classification loss calculated using focal loss, L anchor is a smooth L1 loss used to predict the grasp center offset of candidate points with different rotations.

[0103] It should be noted that in the simulation experiment, the gripper quality score is used to determine the grasping success rate, while in the real experiment, the grasping success rate is determined by whether the grasping is successful.

[0104] Based on this, compared with the prior art, the grasping posture estimation method of the present invention has at least the following beneficial effects:

[0105] 1. Improve the accuracy and success rate of grasping

[0106] 6DOF pose data: Compared with 2D vision, 3D vision technology can provide 6DOF (six degrees of freedom) pose data of the target object, including position (x, y, z) and posture (rotation around the x-axis, rotation around the y-axis, rotation around the z-axis), which enables the robot to more accurately identify and locate the target object.

[0107] Depth information and point cloud: 3D vision technology can provide depth information of the target object or point cloud information of the object surface. This information helps the robot to determine the grasping point more accurately, thereby improving the grasping accuracy.

[0108] 2. Adapt to complex scenarios

[0109] Multi-target recognition: 3D vision-based methods can simultaneously identify multiple target objects, which is particularly important in complex scenarios, such as automated sorting, assembly and other production line operations.

[0110] Occlusion and shadow processing: 3D vision technology can better handle occlusion and shadow problems. Even in the case of limited camera viewing angles or the presence of occlusion and shadows, a more complete picture of the surface of the grasped target can be obtained, thereby improving the ability to recognize the grasped target.

[0111] 3. Optimize crawling planning

[0112] Grasping posture evaluation: By considering the impact of grasping posture on grasping time, a grasping posture quality evaluation method can be proposed to screen out high-quality grasping postures, thereby reducing the error rate in the grasping process and improving grasping efficiency.

[0113] Reduce the time spent on grasping: Combining the neural network inference acceleration framework and the efficient point cloud processing algorithm can significantly reduce the time spent on grasping and improve production efficiency.

[0114] 4. Improve the level of automation

[0115] Improved intelligence: The 3D vision-based multi-target recognition and grasping pose estimation method enables the robot to have a higher level of intelligence, and can independently complete a series of tasks such as target perception, motion planning, grasping planning, etc.

[0116] Wide application: This method has broad application prospects in industrial robots, home service robots, clinical surgical robots and other fields, and can promote the improvement of automation levels in these fields.

[0117] In addition, if Figure 6As shown, one embodiment of the present invention further discloses a grasping posture estimation device, the device comprising:

[0118] An acquisition module 110 is used to acquire a target workpiece image, where the target workpiece image is an RGBD image;

[0119] The recognition module 120 is used to input the target workpiece image into the pre-trained YOLOv10 target detection model for target recognition to obtain a target detection result;

[0120] The grasping module 130 is used to input the target detection result into the six-degree-of-freedom grasping detection HGGD model based on efficient heat map guidance to estimate the grasping pose, so as to obtain the estimated grasping pose of the target workpiece, wherein the HGGD model includes a grasping point heat map model GHM and a non-uniform multi-grasp point generator NMG. GHM is used to process the RGBD image, use efficient CNN to extract semantic features, and generate multiple grasping heat maps as guidance for the grasping area; NMG uses the grasping heat map to aggregate local points into the grasping area, and detects grasping within the grasping area through a lightweight point encoder.

[0121] The grasping posture estimation device of the embodiment of the present invention is used to execute the grasping posture estimation method in the above embodiment. Its specific processing process is the same as the grasping posture estimation method in the above embodiment, and will not be repeated here.

[0122] In addition, if Figure 7 As shown, an embodiment of the present invention also discloses an electronic device, including: at least one processor 210; at least one memory 220, for storing at least one program; when the at least one program is executed by the at least one processor 210, a grasping pose estimation method as in any of the previous embodiments is implemented.

[0123] In addition, an embodiment of the present invention further discloses a computer-readable storage medium, in which computer-executable instructions are stored. The computer-executable instructions are used to execute the grasping pose estimation method as described in any of the previous embodiments.

[0124] The system architecture and application scenarios described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those skilled in the art will appreciate that with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.

[0125] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0126] In hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0127] The terms "component", "module", "system", etc. used in this specification are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, or a computer. By way of illustration, both applications running on a computing device and a computing device can be components. One or more components may reside in a process or an execution thread, and a component may be located on a computer or distributed between two or more computers. In addition, these components may be executed from various computer-readable media having various data structures stored thereon. Components may communicate, for example, through local or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, or a network, such as the Internet interacting with other systems through signals).

Claims

1. A grasping pose estimation method, characterized in that: include: Acquire a target workpiece image, wherein the target workpiece image is an RGBD image; Input the target workpiece image into a pre-trained YOLOv10 target detection model for target recognition to obtain a target detection result; The target detection result is input into the six-degree-of-freedom grasping detection HGGD model based on efficient heat map guidance to estimate the grasping pose, so as to obtain the estimated grasping pose of the target workpiece, wherein the HGGD model includes a grasping point heat map model GHM and a non-uniform multi-grasp point generator NMG, wherein the GHM is used to process the RGBD image, extract semantic features using efficient CNN, and generate multiple grasping heat maps as guidance for the grasping area; the NMG uses the grasping heat map to aggregate local points into the grasping area, and detects grasping within the grasping area through a lightweight point encoder.

2. The grasping posture estimation method according to claim 1, characterized in that: The target workpiece image is input into a pre-trained YOLOv10 target detection model for target recognition to obtain a target detection result, including: Performing image preprocessing on the target workpiece image to obtain a preprocessed image, wherein the image preprocessing includes image size adjustment and image normalization processing; Input the preprocessed image into the YOLOv10 target detection model for feature extraction to obtain target image features; Dividing the preprocessed image into a plurality of grids, and predicting a plurality of bounding boxes and confidences corresponding to the bounding boxes for each of the grids, wherein the bounding boxes are used to locate the target in the image, and the confidences are used to represent the probability that the target exists in the bounding boxes; Filtering target bounding boxes that are higher than a confidence threshold, and adjusting the filtered target bounding boxes to obtain the adjusted target bounding boxes; The adjusted target bounding box and the corresponding category information are output as the target detection result.

3. The grasping posture estimation method according to claim 1, characterized in that: The target detection result is input into the six-degree-of-freedom grasping detection HGGD model based on efficient heat map guidance to estimate the grasping posture, so as to obtain the estimated grasping posture of the target workpiece, including: Inputting the target detection result into the GHM for processing to generate a capture heat map, wherein the capture heat map includes a confidence heat map and a grid attribute heat map; The NMG efficiently aggregates the plurality of the grasping regions under the guidance of the grasping heat map; The NMG predicts the remaining grip point rotation attributes using the grip point attributes of each grid in the grid attribute heat map; The NMG processes the captured heat map and point cloud into useful local features; The NMG refines the rotation attribute of the gripping point through the local features to generate a plurality of gripping points; An estimated grasping posture of the target workpiece is obtained based on the multiple grasping points.

4. The grasping posture estimation method according to claim 3, characterized in that: The GHM is an encoder-decoder model comprising two output branches, wherein the two output branches are a confidence branch and an attribute branch, wherein the confidence branch is used to construct the confidence heat map, and the attribute branch is used to generate the grid attribute heat map.

5. The grasping posture estimation method according to claim 3, characterized in that: The NMG includes two parts: heat map guided region aggregation and non-uniform multi-gripper generator. Under the guidance of the grasping heat map, the NMG aggregates local points into graspable regions for use by the non-uniform multi-gripper generator.

6. The grasping posture estimation method according to claim 3, characterized in that: The step of obtaining an estimated grasping posture of the target workpiece based on the multiple grasping points includes: Generate grasping points in the heat map-guided area where the target workpiece is located; Using a three-dimensional point cloud generation technology to convert multiple grasping points generated by the guide heat map of the target workpiece into grasping points under the three-dimensional point cloud; Generate the corresponding 3D point cloud gripper according to the grasping point, and select the target 3D point cloud gripper with the highest grasping quality score; Convert the image coordinate system of the target 3D point cloud gripper into a base coordinate system based on the robot; The base coordinate system is transmitted to the robot via Ethernet to execute the corresponding object grasping.

7. The grasping posture estimation method according to claim 1, characterized in that: The training method of the YOLOv10 target detection model includes: Constructing a target loss function, wherein the target loss function is determined according to a classification loss function, a coordinate loss function, and a confidence loss function, wherein the coordinate loss function is used to measure the position difference between the predicted bounding box and the true bounding box, and the confidence loss function is calculated using a binary cross entropy loss; The YOLOv10 target detection model is trained based on the target loss function to obtain the trained YOLOv10 target detection model.

8. A grasping posture estimation device, characterized in that: The device comprises: An acquisition module is used to acquire a target workpiece image, wherein the target workpiece image is an RGBD image; A recognition module is used to input the target workpiece image into a pre-trained YOLOv10 target detection model for target recognition to obtain a target detection result; A grasping module is used to input the target detection result into a six-degree-of-freedom grasping detection HGGD model based on efficient heat map guidance to estimate the grasping pose, so as to obtain the estimated grasping pose of the target workpiece, wherein the HGGD model includes a grasping point heat map model GHM and a non-uniform multi-grasp point generator NMG, the GHM is used to process the RGBD image, use efficient CNN to extract semantic features, and generate multiple grasping heat maps as guidance for the grasping area; the NMG uses the grasping heat map to aggregate local points into the grasping area, and detects grasping within the grasping area through a lightweight point encoder.

9. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the grasping posture estimation method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the grasping pose estimation method according to any one of claims 1 to 7.