Object grasping pose detection method, device, equipment and storage medium

By acquiring seed point datasets and candidate grasping point datasets, and combining deep neural networks and clustering algorithms, the problems of low efficiency and limited applicability of robotic arm grasping poses in existing technologies are solved, and efficient and accurate target foreground object pose grasping is achieved in cluttered scenes.

CN115564832BActive Publication Date: 2025-11-21INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211216621.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-30
Publication Date
2025-11-21
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

Existing robotic arm grasping pose methods require a clear understanding of the shape and category of the target object, resulting in low efficiency and limited applicability in simple, structured scenarios.

Method used

By acquiring a seed point dataset of the target foreground object, determining a candidate grasping point dataset based on the seed point dataset, and predicting the target grasping pose, the deep neural network and clustering algorithm can efficiently and accurately grasp the pose of the target foreground object in a cluttered scene without needing to know the shape and category of the target foreground object.

Benefits of technology

It improves the efficiency and applicability of capturing the pose of target foreground objects, enabling efficient and accurate capture of the pose of target foreground objects in cluttered scenes without needing to know the shape and category of the target foreground object.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115564832B_ABST
    Figure CN115564832B_ABST
Patent Text Reader

Abstract

The application provides an object grasping pose detection method and device, electronic equipment and storage medium, comprising: obtaining seed point data set of a target foreground object, different seed point data in the seed point data set representing different initial local features of the target foreground object in different initial positions in a cluttered scene; based on the seed point data set, determining a candidate grasping point data set of the target foreground object, the candidate grasping point data set representing that there are candidate grasping point data and candidate positions and candidate local features of the candidate grasping point data in the target number of regions corresponding to the seed point data; based on the candidate grasping point data set, predicting a target grasping pose of the target foreground object, the target grasping pose representing that the candidate grasping pose is predicted for the candidate grasping point data cluster divided from the candidate grasping point data set and the quality score of the grasping pose meets the preset quality requirement. The application can improve the efficiency of grasping the target foreground object pose and expand the application range of grasping the target foreground object pose.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot vision technology, and in particular to a method, apparatus, device, and storage medium for detecting the grasping pose of an object. Background Technology

[0002] Currently, due to the wide range of applications and development prospects of robotic arm grasping operations in fields such as industry, home services, healthcare, and space exploration, how to flexibly and intelligently grasp the pose of target objects has become a key issue that urgently needs to be addressed.

[0003] In related technologies, robotic arm grasping pose methods typically first construct and store target object models containing different shapes and categories of target objects arranged in an orderly manner, then match the current scene with the target object models in terms of category and 3D pose, and grasp the target object based on a predefined grasping pose when the match is successful.

[0004] However, existing robotic arm pose grasping methods require a clear understanding of the shape and category of the target object before they can perform pose grasping operations. Moreover, pose grasping operations are only applicable to simple, rule-based structured scenarios, resulting in low efficiency and limited applicability in grasping the pose of the target object. Summary of the Invention

[0005] This invention provides a method, apparatus, device, and storage medium for detecting object grasping pose, which solves the shortcomings of existing technologies that require knowing the shape and category of the target foreground object before performing pose grasping operations, and that pose grasping operations are only applicable to simple, structured scenes, resulting in low efficiency and limited applicability of target foreground object pose. It achieves the goal of efficiently and accurately grasping the pose of target foreground objects in cluttered scenes without needing to know the shape and category of the target foreground object, which not only improves the efficiency of grasping target foreground object pose but also greatly expands the applicability of grasping target foreground object pose.

[0006] This invention provides a method for detecting the grasping pose of an object, comprising:

[0007] Obtain a seed point dataset of the target foreground object, wherein different seed point data in the seed point dataset represent different initial local features of the target foreground object at different initial positions in a cluttered scene;

[0008] Based on the seed point dataset, a candidate grab point dataset for the target foreground object is determined. The candidate grab point dataset represents the number of regions corresponding to the seed point data where candidate grab point data exists, as well as the candidate positions and candidate local features of the candidate grab point data.

[0009] Based on the candidate grasp point dataset, the target grasp pose of the target foreground object is predicted. The target grasp pose represents the predicted grasp pose for the candidate grasp point data clusters divided from the candidate grasp point dataset, and the quality score of the grasp pose meets the preset quality requirements.

[0010] According to the object grasping pose detection method provided by the present invention, the step of obtaining a seed point dataset of the target foreground object includes:

[0011] Acquire depth image information containing objects in cluttered scenes;

[0012] The depth image information is converted into a point cloud to obtain the original point cloud data of the object;

[0013] The original point cloud data is input into a preset deep neural network model to obtain the seed point dataset of the target foreground object output by the preset deep neural network model;

[0014] The preset deep neural network model is used to extract features from the original point cloud data, and then to filter foreground objects from the candidate seed point dataset of the object obtained by feature extraction, so as to obtain the seed point dataset of the target foreground object. The candidate seed point dataset represents the different initial local features of the global point cloud of the object at different initial positions.

[0015] According to the object grasping pose detection method provided by the present invention, the step of determining a candidate grasping point dataset for the target foreground object based on the seed point dataset includes:

[0016] Determine the first number of regions corresponding to the seed point data in the seed point dataset;

[0017] Based on the initial position and initial local features of the seed point data, a first number of region scores, a first number of candidate grasp point data, and candidate positions and candidate local features of each of the first number of regions are predicted, representing that each of the first number of regions contains a feasible grasping pose.

[0018] Based on the first number of region scores, a second number of candidate crawling point data that meet the preset region score requirements are determined from the first number of candidate crawling point data.

[0019] Based on the candidate position and candidate local features of each candidate grab point in the second number of candidate grab point data, a candidate grab point dataset of the target foreground object is determined, and the product of the second number and the total number of seed point data is the target number.

[0020] According to the object grasping pose detection method provided by the present invention, the step of predicting the target grasping pose of the target foreground object based on the candidate grasping point dataset includes:

[0021] Based on the different candidate positions of different candidate crawl point data in the candidate crawl point dataset, the candidate crawl point dataset is clustered to determine multiple candidate crawl point data clusters;

[0022] Based on the different candidate local features of the different candidate grasping point data, the grasping pose of each candidate grasping point data cluster and the quality score of the grasping pose are predicted.

[0023] Based on the quality score, the target grasping pose of the target foreground object is predicted from multiple grasping poses.

[0024] According to the object grasping pose detection method provided by the present invention, the step of filtering foreground objects from the candidate seed point dataset of the target foreground object obtained by feature extraction to obtain the seed point dataset of the target foreground object includes:

[0025] Obtain the object surface shape information for the target foreground object;

[0026] Based on the object surface shape information, the candidate seed point dataset of the target foreground object obtained by feature extraction is used to filter foreground objects and obtain the seed point dataset of the target foreground object.

[0027] According to the object grasping pose detection method provided by the present invention, determining the first number of regions corresponding to the seed point data in the seed point dataset includes:

[0028] Based on the preset region division model and the seed point dataset, a first number of regions corresponding to the seed point data in the seed point dataset are predicted; wherein, the preset region division model is used to first establish a spherical coordinate system with each seed point data in the seed point dataset as the center and a preset length as the radius, and to divide the spherical coordinate system into regions based on a preset azimuth angle and a preset zenith angle, thereby determining the first number of regions corresponding to the seed point data in the seed point dataset.

[0029] According to the object grasping pose detection method provided by the present invention, the step of predicting, based on the seed point data, a first number of region scores, a first number of candidate grasping point data, a candidate position and candidate local features of each of the first number of regions containing feasible grasping poses, and based on the seed point data, includes:

[0030] Based on the preset neural network model and the different initial local features and different initial positions of different seed point data in the seed point dataset, the first number of region scores, the first number of candidate grasping point data, the candidate position and candidate local features of each of the first number of regions containing feasible grasping poses are predicted.

[0031] The preset neural network model is used to predict the position offset and local feature offset between candidate grab point data and corresponding seed point data in different regions based on the different initial local features; and to determine the candidate position and candidate local features of each candidate grab point data based on the position offset, the local feature offset, the initial position and the initial local features; and to predict the region score containing feasible grab poses in different regions based on the different initial positions.

[0032] The present invention also provides a device for detecting the grasping pose of an object, comprising:

[0033] The acquisition module is used to acquire a seed point dataset of the target foreground object. Different seed point data in the seed point dataset represent different initial local features of the target foreground object at different initial positions in a cluttered scene.

[0034] The determination module is used to determine a candidate grasping point dataset of the target foreground object based on the seed point dataset. The candidate grasping point dataset represents the existence of candidate grasping point data in a number of regions corresponding to the seed point data, as well as the candidate positions and candidate local features of the candidate grasping point data.

[0035] The prediction module is used to predict the target grasping pose of the target foreground object based on the candidate grasping point dataset. The target grasping pose represents the predicted grasping pose for the candidate grasping point data clusters divided from the candidate grasping point dataset, and the quality score of the grasping pose meets the preset quality requirements.

[0036] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the object grasping pose detection method as described above.

[0037] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the object grasping pose detection method as described above.

[0038] The present invention provides a method, apparatus, electronic device, and storage medium for detecting object grasping pose. The method for detecting object grasping pose acquires a seed point dataset of the target foreground object. Since different seed point data in the seed point dataset represent different initial local features of the target foreground object at different initial positions in a cluttered scene, combining this with object foreground object localization technology improves the intelligence and accuracy of locating the target foreground object, eliminating the need to locate the target foreground object by category and shape. Furthermore, based on the seed point dataset, a candidate grasping point dataset of the target foreground object is determined. Then, based on the candidate grasping point dataset, the target grasping pose of the target foreground object is predicted. Since the candidate grasping point dataset represents the existence of candidate grasping point data and candidate positions and candidate local features in the target area corresponding to the seed point data, the target grasping pose representation is specific to the target foreground object. The candidate grasping point dataset is divided into candidate grasping point data clusters, which predict the grasping pose and the quality score of the grasping pose meets the preset quality requirements. Therefore, it is only necessary to select candidate grasping point data in the vicinity of each seed point data and determine the candidate position and candidate local features of the candidate grasping point data. By first predicting the candidate pose of the candidate grasping point data cluster obtained by clustering, and then judging whether the quality score of the predicted candidate pose meets the preset quality requirements, the target grasping pose of the target foreground object can be determined. This achieves the purpose of grasping the pose of the target foreground object in cluttered scenarios such as daily life and personalized orders. Moreover, it can efficiently and accurately grasp the pose of the target foreground object in cluttered scenarios without knowing the shape and category of the target foreground object. This not only improves the efficiency of grasping the pose of the target foreground object, but also greatly expands the applicable scope of grasping the pose of the target foreground object. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0040] Figure 1 This is one of the flowcharts illustrating the object grasping pose detection method provided by the present invention;

[0041] Figure 2 This is a schematic diagram of the robotic arm grasping posture detection system provided by the present invention;

[0042] Figure 3 This is a schematic diagram of the 6-DOF grasping pose coordinate system representation method and simplified gripper control points provided by the present invention.

[0043] Figure 4This is a schematic diagram of the grasping pose generation method provided by the present invention;

[0044] Figure 5 This is a schematic diagram of the region division method near seed point data provided by the present invention;

[0045] Figure 6 This is a schematic diagram of the method for calculating the difference between two grasping poses during the training process provided by the present invention;

[0046] Figure 7 This is the second flowchart of the object grasping pose detection method provided by the present invention;

[0047] Figure 8 This is a schematic diagram of the system configuration for implementing the object grasping pose detection method provided by the present invention;

[0048] Figure 9 This is a schematic diagram of the object grasping posture detection device provided by the present invention;

[0049] Figure 10 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0051] The following is combined Figures 1-10 This invention describes a method, apparatus, electronic device, and storage medium for detecting object grasping pose. The execution entity of the object grasping pose detection method can be a terminal device or a server. The terminal device can be a personal computer (PC), portable device, laptop, or other electronic device. The server can be a standalone server or a server cluster composed of multiple servers. For example, the server can be a physical server containing independent hosts, a virtual server hosted by a host cluster, a cloud server, etc. This invention does not limit the specific form of the terminal device, nor does it specifically limit the specific form of the server.

[0052] It should be noted that the following method embodiments are described using a terminal device as the execution subject, and the execution subject of the following method embodiments can be part or all of the terminal device.

[0053] Figure 1This is a flowchart illustrating the object grasping pose detection method provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method for detecting the object's grasping pose includes:

[0054] Step 110: Obtain the seed point dataset of the target foreground object. Different seed point data in the seed point dataset represent different initial local features of the target foreground object at different initial positions in a cluttered scene.

[0055] Among them, the cluttered scene can be a scene in which multiple stacked objects are placed; the target foreground object can be a stacked object of different shapes and types, and the number of target foreground objects can be multiple.

[0056] Specifically, the object grasping pose detection method provided in this embodiment of the invention is applicable to, for example, Figure 2 In the robotic arm grasping pose detection system shown, in Figure 2 In this design, the robotic arm's workspace is located within a pre-defined grasping plane. This pre-defined grasping plane can be a planar area of ​​a physical object, such as a box or shelf. The workspace contains multiple stacked objects of different shapes and types. The depth camera is positioned directly above these stacked objects, capturing depth images and transmitting them to a terminal device via a data cable. The terminal device uses the depth images to determine a seed point dataset of the target foreground object, then determines the target grasping pose based on the seed point dataset. It then calculates the feasible motion path information for the robotic arm based on the target grasping pose and sends this information to the robotic arm control box. The control box uses this feasible motion path information to control the robotic arm and two-finger gripper to execute grasping commands, completing the grasping operation of the target foreground object. The depth camera can be an RGB-D camera, such as an Intel RealSense camera. The camera is L515; the robotic arm can be the Aubo-i5 robotic arm, and the two-finger gripper can be the DH-AG95 two-finger gripper.

[0057] Based on this, the terminal device can receive depth image information sent by the depth camera and parse the depth image information to extract different initial local features of the target foreground object at different initial positions in a cluttered scene, and display them as a seed point dataset. Each seed point in the seed point dataset can be represented as {s}. i =(x i ,f i |i=1,…,I}, It is a 3-dimensional position vector space. Let x be a C-dimensional eigenvector space. iFor the i-th seed point data in the seed point dataset s i The initial position vector, f i For the i-th seed point data in the seed point dataset s i The initial local feature vector; I is the total number of seed points in the seed point dataset, and I is a positive integer greater than 1, such as I = 256.

[0058] Step 120: Based on the seed point dataset, determine the candidate grab point dataset of the target foreground object. The candidate grab point dataset represents the number of regions corresponding to the seed point data where candidate grab point data exists, as well as the candidate positions and candidate local features of the candidate grab point data.

[0059] Specifically, the terminal device makes predictions based on the seed point dataset and candidate crawling point data. That is, it first determines a number of target regions around and / or near each seed point data based on a preset candidate crawling point data prediction algorithm. Then, it predicts candidate crawling point data for each of the determined target regions, and obtains a number of candidate prediction points corresponding to each seed point data. Furthermore, it predicts the position and local features of each of the target candidate prediction points, thereby determining the number of candidate crawling point data corresponding to each seed point data and the candidate position vector and candidate local feature vector of each candidate crawling point data.

[0060] Step 130: Based on the candidate grasp point dataset, predict the target grasp pose of the target foreground object. The target grasp pose characterizes the predicted grasp pose of the candidate grasp point data clusters divided for the candidate grasp point dataset, and the quality score of the grasp pose meets the preset quality requirements.

[0061] Specifically, for the candidate grasping point dataset, the terminal device first clusters the candidate grasping point dataset based on a preset clustering algorithm to determine multiple candidate grasping point data clusters. Then, based on a preset grasping pose prediction algorithm, it predicts the grasping pose for each candidate grasping point data cluster, determining multiple grasping poses and a quality score for each grasping pose. Finally, based on the matching relationship between the quality scores and preset quality requirements, it determines the target grasping pose from the multiple grasping poses. Each grasping pose is a 6-DOF grasping pose, including 3 three-dimensional positional degrees of freedom and 3 three-dimensional rotational degrees of freedom. Figure 2 The robotic arm in the image is naturally a 6-DOF robotic arm, and the representation of the 6-DOF grasping pose is as follows: Figure 3 The transformation matrix shown is the local gripper coordinate system relative to the world coordinate system. Figure 3 In the diagram, the x-axis represents the direction of finger closure. The z-axis represents the gripping direction of the gripper. The origin of the local gripper coordinate system is located at the midpoint between the tips of the two fingers and represents the gripping pose position t.g .

[0062] The object grasping pose detection method provided by this invention acquires a seed point dataset of the target foreground object. Since different seed point data in the seed point dataset represent different initial local features of the target foreground object at different initial positions in a cluttered scene, combining this with object foreground object localization technology improves the intelligence and accuracy of locating the target foreground object, eliminating the need to locate the target foreground object by category and shape. Furthermore, based on the seed point dataset, a candidate grasping point dataset of the target foreground object is determined. Then, based on the candidate grasping point dataset, the target grasping pose of the target foreground object is predicted. Since the candidate grasping point dataset represents the presence of candidate grasping point data and candidate positions and candidate local features in the target area corresponding to the seed point data, the target grasping pose represents the candidate grasping pose divided based on the candidate grasping point dataset. The system predicts the grasping pose from clusters of seed data, and the quality score of the grasping pose meets the preset quality requirements. Therefore, by selecting candidate grasping point data in the vicinity of each seed point data and determining the candidate position and candidate local features of the candidate grasping point data, the target grasping pose of the target foreground object can be determined by first predicting the candidate pose of the clustered candidate grasping point data cluster and then judging whether the quality score of the predicted candidate pose meets the preset quality requirements. This achieves the goal of grasping the pose of the target foreground object in cluttered scenarios such as daily life and personalized orders. Moreover, it can efficiently and accurately grasp the pose of the target foreground object in cluttered scenarios without knowing the shape and category of the target foreground object. This not only improves the efficiency of grasping the pose of the target foreground object but also greatly expands the applicable scope of grasping the pose of the target foreground object.

[0063] Optionally, the specific implementation process of step 110 may include:

[0064] First, acquire depth image information of the target foreground object in a cluttered scene; then, perform point cloud conversion on the depth image information to obtain the original point cloud data of the target foreground object; then, input the original point cloud data into a preset deep neural network model to obtain the seed point dataset of the target foreground object output by the preset deep neural network model.

[0065] The pre-defined deep neural network model is used to extract features from the original point cloud data. Then, the candidate seed point dataset of the target foreground object obtained from the feature extraction is used to filter foreground objects, thus obtaining a seed point dataset of the target foreground object. The candidate seed point dataset represents the different initial local features of the global point cloud of the target foreground object at different initial positions. The pre-defined deep neural network model is the model obtained after training the initial neural network model.

[0066] Specifically, the terminal device can acquire depth image information containing target foreground objects in a cluttered scene. This can be achieved by first instructing the depth camera to capture images of multiple stacked objects in the workspace, and then having the depth camera feed back the depth image information to the terminal device. Alternatively, the terminal device can acquire depth image information that matches the current workspace from a pre-stored set of multiple depth image information sets. No specific limitations are specified here.

[0067] Furthermore, by inputting the original point cloud data into a preset deep neural network model, a seed point dataset for the target foreground object is obtained. The preset deep neural network model contains a preset PointNet++ feature extraction network model, a preset foreground-background segmentation network model, and a preset Hough voting mechanism model. The preset PointNet++ feature extraction network model is used to sample the farthest point from the original point cloud data and then perform local feature extraction after fusing the local point cloud data obtained from the farthest point sampling with the original point cloud data to determine the candidate seed point dataset for the target foreground object. The candidate seed point dataset can contain N candidate seed point data. The preset foreground-background segmentation network model is used to determine whether each candidate seed point data belongs to the target foreground object or the background based on the initial local feature vector of each candidate seed point data in the candidate seed point dataset using a self-attention mechanism, that is, to predict the probability that each candidate seed point data belongs to the target foreground object, thus obtaining the seed point dataset for the target foreground object. The background can be any other non-target object such as a desktop, a robotic arm, or a camera stand.

[0068] The object grasping pose detection method provided by this invention involves the terminal device first performing point cloud conversion on the depth image containing the target foreground object in a cluttered scene, and then outputting the original point cloud data of the target foreground object obtained by point cloud conversion to a preset deep neural network model to obtain the seed point dataset of the target foreground object. This method, combined with the PointNet++ feature extraction mechanism and self-attention mechanism, significantly improves the reliability and accuracy of obtaining the seed point dataset of the target foreground object, and also significantly improves the success rate of subsequent grasping pose.

[0069] Optionally, the specific implementation process of step 120 may include:

[0070] First, a first number of regions corresponding to the seed point data in the seed point dataset are determined. Then, based on the initial position and initial local features of the seed point data, a first number of region scores, a first number of candidate grasping point data, and candidate positions and candidate local features of each candidate grasping point data are predicted, representing the feasible grasping poses in the first number of regions. Next, based on the first number of region scores, a second number of candidate grasping point data that meet the preset region score requirements are determined from the first number of candidate grasping point data. Finally, based on the candidate positions and candidate local features of each candidate grasping point data in the second number of candidate grasping point data, the candidate grasping point dataset of the target foreground object is determined.

[0071] Specifically, the terminal device first selects a first number of regions around and / or near each seed point in the seed point dataset. Then, based on the initial feature vector of each seed point, it predicts the candidate grasping point data contained in each of the first number of regions, as well as the candidate position and candidate local features of each candidate grasping point data. It also predicts a region score representing a feasible grasping pose for each of the first number of regions. Based on a preset Hough voting mechanism, it maps each seed point in the seed point dataset to the corresponding candidate grasping point data. That is, it sorts the predicted first number of region scores from largest to smallest, and then selects the second number of region scores that meet the preset region score requirements from the sorted first number of region scores. This determines the second number of candidate grasping point data corresponding to the second number of region scores. Since each seed point data corresponds to the second number of candidate grasping point data, the candidate grasping point dataset of the target foreground object contains the target number of candidate grasping point data, and the target number of candidate grasping point data are distributed in the target number of regions. The product of the second number and the total number of seed point data is the target number. For example, when the second quantity is G and the seed point dataset contains I seed point data, the target quantity can be G×I.

[0072] The object grasping pose detection method provided by this invention involves a terminal device predicting a first number of region scores and a first number of candidate grasping point data for each seed point data in a seed point dataset. Based on the first number of region scores, a second number of candidate grasping point data are selected from the first number of candidate grasping point data. Finally, based on the candidate position and candidate local features of each candidate grasping point data in the second number of candidate grasping point data, a candidate grasping point dataset for the target foreground object is determined. This method, combined with a voting mechanism, selects candidate grasping point data from the regions with the highest scores around each seed point data, thereby improving the accuracy and reliability of determining the candidate grasping point dataset for the target foreground object.

[0073] Optionally, the specific implementation process of step 130 may include:

[0074] Based on the different candidate positions of different candidate grasp point data in the candidate grasp point dataset, the candidate grasp point dataset is clustered to determine multiple candidate grasp point data clusters; based on the different candidate local features of different candidate grasp point data, the grasp pose of each candidate grasp point data cluster and the quality score of the grasp pose are predicted; based on the quality score, the target grasp pose of the target foreground object is predicted from multiple grasp poses.

[0075] Specifically, the terminal device, based on the different candidate positions and local features of different candidate grasping points in the candidate grasping point dataset, can use a pre-defined fully convolutional neural network model to predict the target grasping pose of the foreground object. That is, the pre-defined fully convolutional neural network model clusters the candidate grasping point dataset based on the different candidate positions, determining multiple candidate grasping point clusters. Then, based on the different local features of the different candidate grasping points, it predicts the grasping pose and quality score of each candidate grasping point cluster, thus obtaining multiple quality scores. Finally, the highest quality score among these is selected, and the grasping pose corresponding to the highest quality score is determined as the target grasping pose. The pre-defined fully convolutional neural network model is a model obtained after training an initial fully convolutional neural network model.

[0076] It should be noted that, based on the different local features of different candidate grasping point data, the grasping pose of each candidate grasping point data cluster is predicted. Specifically, multi-head neural networks and pooling methods can be used to predict the location of the grasping pose, the finger closing direction, the gripper approach direction, and the quality score of the grasping pose. Furthermore, the Schmidt orthogonalization method can be used to ensure that the finger closing direction is perpendicular to the gripper approach direction.

[0077] For example, the candidate grasp point dataset is clustered to identify 256 candidate grasp point data clusters, and a neural network with four prediction heads is used to predict feasible grasp poses and their quality scores, such as... Figure 4 As shown, the four prediction heads output the capture position t respectively. g The gripper approaches the grasping direction. finger closing direction And the score s of the capture pose. cls ; and using the Schmidt orthogonalization method to transform the predicted direction vector and Adjustments were made to ensure they were perpendicular to each other.

[0078] The object grasping pose detection method provided by this invention involves a terminal device first clustering a candidate grasping point dataset, then predicting the grasping pose and its quality score for the clustered candidate grasping point data clusters, and finally determining the target grasping pose from multiple grasping poses based on the highest quality score among multiple quality scores. This method combines the clustering and prediction mechanisms in a neural network to improve the accuracy of determining the target grasping pose, providing a reliable guarantee for the subsequent successful grasping of the target foreground object.

[0079] Optionally, determining a first number of regions corresponding to the seed point data in the seed point dataset, the implementation process may include:

[0080] Based on the preset region division model and the seed point dataset, the first number of regions corresponding to the seed point data in the seed point dataset are predicted; wherein, the preset region division model is used to first establish a spherical coordinate system with each seed point data in the seed point dataset as the center and a preset length as the radius, and to divide the spherical coordinate system into regions based on the preset azimuth angle and the preset zenith angle, thereby determining the first number of regions corresponding to the seed point data in the seed point dataset.

[0081] Specifically, the terminal device uses a preset region partitioning model to partition the seed point dataset, that is, it can perform region partitioning as follows: Figure 5 As shown, a spherical coordinate system is established with each seed point data point as the center and a preset length as the radius. The preset length can be the average diameter of the workspace × 0.05, for example, a preset length of 5cm. The spherical coordinate system is then divided into m regions according to a preset azimuth angle φ∈[0,2π), and into n regions using a preset zenith angle θ∈[0,π]. This yields M regions corresponding to the seed point data points in the seed point dataset, i.e., the first set of regions is M. Furthermore, when m = 8 and n = 4, the first set of regions can be 32.

[0082] The object grasping pose detection method provided by the present invention involves a terminal device first establishing a spherical coordinate system for each seed point data using a preset region division model, and then dividing the spherical coordinate system into regions based on preset azimuth angles and preset zenith angles. This determines the first number of regions corresponding to the seed point data in the seed point dataset. By combining this with the preset region division model, the intelligence and convenience of region division are improved, thereby making the determined first number of regions more accurate.

[0083] Optionally, based on seed point data, a first number of region scores, a first number of candidate grasping point data, candidate positions and candidate local features of each candidate grasping point data are predicted, representing a first number of regions containing feasible grasping poses. The implementation process may include:

[0084] Based on a pre-defined neural network model and different initial local features and positions of different seed point data in the seed point dataset, the system predicts a first number of region scores, a first number of candidate grasping point data, candidate positions, and candidate local features for each candidate grasping point data, representing a first number of regions containing feasible grasping poses. Specifically, the pre-defined neural network model is used to predict the positional offset and local feature offset between candidate grasping point data and corresponding seed point data in different regions based on different initial local features. Based on the positional offset, local feature offset, initial position, and initial local features, it determines the candidate position and candidate local features for each candidate grasping point data. Furthermore, it is used to predict the region scores containing feasible grasping poses in different regions based on different initial positions. The pre-defined neural network model is a model obtained by training the initial neural network model.

[0085] Specifically, since each seed point in the seed point dataset can be represented as s i =(x i ,f i ), x i For the i-th seed point data in the seed point dataset s i The initial position vector, f i For the i-th seed point data in the seed point dataset s i The initial local feature vector; therefore, when the seed point dataset is input into the preset neural network model, the terminal device can obtain the preset neural network model's prediction of the position offset and local feature offset of different seed point data in each region based on different initial local feature vectors. That is, different seed point data are used as voting points to predict the offset of the seed point data in terms of position and local features, thereby predicting (Δx) i,k ,Δf i,k ) = Vote(s i ), Δx i,k For the i-th seed point data s i The position offset Δf in the k-th region i,k For the i-th seed point data s i The local feature offset in the k-th region, where k is the region number and its maximum value can be M. Based on this, the data s of the i-th seed point can be determined. i Candidate position vector in the k-th region and the data s of the i-th seed point i Candidate local feature vectors in the k-th region

[0086] At the same time, the terminal device can also obtain the region scores containing feasible grasping poses in different regions predicted by the preset neural network model based on different initial positions. When the seed point dataset contains I seed point data and the first number of regions corresponding to each seed point data is M regions, then M region scores can be predicted for each seed point data. After sorting the M region scores from largest to smallest, the candidate grasping point data contained in the G regions corresponding to the top G region scores are selected as the G candidate grasping point data corresponding to each seed point data, where G < M. At this time, the candidate position vector and candidate local feature vector of each candidate grasping point data in the G candidate grasping point data can be determined.

[0087] The object grasping pose detection method provided by this invention involves a terminal device inputting different initial local features and different initial positions of different seed point data into a preset neural network model to predict a first number of region scores, a first number of candidate grasping point data, and candidate positions and candidate local features for each candidate grasping point data. By combining this with the preset neural network model to first predict the position and local feature offsets, and then predict the candidate positions and candidate local features, the accuracy of determining each candidate grasping point data can be effectively improved, and the candidate positions and candidate local features of each candidate grasping point data can also be accurately determined.

[0088] Optionally, the candidate seed point dataset of the object obtained from feature extraction can be used to filter foreground objects to obtain the seed point dataset of the target foreground objects. The implementation process may include:

[0089] First, obtain the surface shape information of the target foreground object; then, based on the surface shape information, filter the candidate seed point dataset of the object obtained by feature extraction to obtain the seed point dataset of the target foreground object.

[0090] Among them, the surface shape information of the object can characterize the three-dimensional geometric information of the target foreground object.

[0091] Specifically, the terminal device acquires the surface shape information of the target foreground object. This can be done by the user inputting the surface shape information into the terminal device, or by the user inputting the surface shape information into an application on another device connected to the terminal device. The specific method by which the terminal device acquires the surface shape information of the target foreground object is not limited here. Based on this, the terminal device uses a preset deep neural network model containing a preset foreground / background segmentation network model and a preset Hough voting mechanism model to determine the seed point dataset. This can be combined with the surface shape information of the target foreground object. Specifically, the preset foreground / background segmentation network model uses a self-attention mechanism and the surface shape information of the target foreground object to determine whether each candidate seed point belongs to the foreground or background based on the initial local feature vector of each candidate seed point in the candidate seed point dataset, i.e., predicting the probability that each candidate seed point belongs to the foreground object. Thus, the terminal device can use the preset Hough voting mechanism to first sort the predicted probabilities of each candidate seed point belonging to the foreground object from largest to smallest, and then determine the position vector and local feature vector corresponding to the top I probabilities from the sorted results as the seed point dataset for the target foreground object. It should be noted that the self-attention mechanism is constructed in the following way in this embodiment of the invention:

[0092]

[0093] Among them, X s This is a set of candidate local feature vectors that contains data for each candidate seed point. Let C be the feature vector of N candidate seed points, and σ(·) be the soft-max function. * (·) represents the multi-layer convolution operation in the key generation, query generation, and value generation modules, M K M is the key matrix. Q For querying the matrix, For N candidate seed point data 3D feature vectors, M V For value matrices, S represents the probability that the data from the N candidate seed points belong to the foreground object.

[0094] Furthermore, obtaining the surface shape information of the target foreground object can be achieved by acquiring RGB image information and / or depth image information containing the target foreground object, and then performing image segmentation on the RGB image information and / or depth image information. The acquisition of RGB image information and / or depth image information can be achieved by instructing an RGB camera and / or a depth camera to capture RGB image information and / or depth image information of multiple stacked objects in the workspace, and then sending the captured RGB image information and / or depth image information to the terminal device via a data cable. Alternatively, it can be achieved by instructing other devices connected to the terminal device to send RGB image information and / or depth image information containing the target foreground object. The specific method by which the terminal device acquires RGB image information and / or depth image information containing the target foreground object is not specifically limited here.

[0095] The object grasping pose detection method provided by this invention involves a terminal device filtering foreground objects by combining the object surface shape information of the target foreground object with the candidate seed point dataset of the target foreground object obtained by feature extraction. This improves the reliability and accuracy of obtaining the seed point dataset and provides sufficient basis for improving the success rate of target grasping pose in the future.

[0096] It should be noted that the training methods for the pre-defined deep neural network model, the pre-defined neural network model, and the pre-defined fully convolutional neural network model can all use supervised learning. The training process can be divided into three stages: feature extraction and foreground / background classification training for the initial deep neural network model, voting position training for the initial neural network model, and pose capture training for the initial fully convolutional neural network model. The total number of iterations is set to Q, and the first 10% of Q iterations are the training iterations for the initial deep neural network model, with the training stopping when the foreground / background loss function L is reached. obj The first preset loss requirement is achieved; the subsequent 30%Q iterations are the training rounds for training the initial deep neural network model, and the training stops when the voting loss function L of the model training is reached. vote The second preset loss requirement is met; the final 60%Q iterations are the training iterations performed on the initial fully convolutional neural network model, and the training stops when the total loss function L of the model training reaches the second preset loss requirement. The loss function involved is:

[0097]

[0098]

[0099]

[0100] L = wobj L obj +w vote L vote +w grasp L grasp L grasp =w add-s L add-s +w sym L sym +w bce L bce ;

[0101]

[0102]

[0103] Where N is the total number of candidate seed point data in the candidate seed point dataset, and y p Let the data of the p-th candidate seed point belong to the label of the foreground object and y p ∈{0,1}, Let p be the predicted probability that the data of the p-th candidate seed point belongs to the foreground object and log(·) is the logarithmic function with the natural constant e as the base, L vote-reg Let w be the first binary cross-entropy loss function. vote-reg The first binary cross-entropy loss function L vote-reg The weight, L vote-cls Let w be the second binary cross-entropy loss function. vote-cls The second binary cross-entropy loss function L vote-cls The weight, The score for the j-th region corresponding to the i-th seed point data. Let be the position vector of the j-th region corresponding to the i-th seed point data. Let y' be the coordinate of the kth actual capture location. p′ For each seed point data, does the p'-th region contain more than one label for a feasible crawling location and y'? p' ∈{0,1}, The probability value that there is a feasible crawling location in the p'-th region corresponding to each seed point data and L grasp To capture the pose prediction loss function, w obj Foreground / background loss function L obj The weight, w vote The voting loss function L vote The weight, w grasp The total loss function L for capturing pose grasp The weight, L add-s To capture the pose loss function, w add-sTo capture the pose loss function L add-s The weight, L sym For symmetric capture loss function, w sym For symmetric grasping loss function L sym The weight, L bce Let w be the pose scoring loss function. bce The weights of the pose scoring loss function are... Let σ(·) be the predicted score for the l-th grasping pose, S be the total number of predicted grasping poses, and σ(·) be the predicted score for the l-th grasping pose. + The first value of the soft-max function result, g l For the predicted l-th grasp pose, g gt For the l-th real capture pose, p c For the predicted l-th grasping pose g l The coordinates of the c-th control point, p gt,c Let c be the coordinates of the l-th actual capture pose. For a set of real-time captured poses, L bce Predicted score for the l-th grasp pose Binary cross-entropy loss with long-tailed distribution.

[0104] It should be noted that, during the training process for grasping pose detection in this invention, methods such as... Figure 6 The method shown measures the difference between two grasp poses, i.e., the difference described by the aforementioned ADD(·,·) function; according to Figure 3 The method shown simplifies the gripper to 5 control points p1 to p5, and calculates the average distance between the corresponding control points in the two poses. Considering the symmetry of the gripper, {p1, p3} is equivalent to {p2, p4}. Therefore, when calculating the difference, the influence of this symmetry must be taken into account, and the minimum difference value between the original pose and the symmetrical pose should be taken as the final difference value.

[0105] Reference Figure 7 The flowchart shown is a method for detecting the object grasping pose. Figure 7 As shown, feature extraction is performed on the input raw point cloud data to obtain a candidate seed point dataset. Then, the candidate seed point dataset is processed by a self-attention mechanism and a foreground object selection method to obtain a target foreground object seed point dataset. At this point, a voting mechanism is used to predict the position offset and local feature offset of different seed point data in each region, thereby determining the candidate grasping point dataset representing the voting points. Finally, the candidate grasping point dataset is aggregated and classified into multiple candidate grasping point data clusters for pose prediction. A 6-DOF grasping pose is predicted for each candidate grasping point data cluster. The specific process involved is the same as the previous process and will not be repeated here.

[0106] Reference Figure 8 The diagram shows the system configuration for implementing the object grasping pose detection method. Figure 8 As shown, after performing feature extraction and seed point selection on the depth image information, a candidate grasping point dataset representing voting points is determined. Then, the object surface shape information of the target foreground object obtained from image instance segmentation is combined to predict the grasping pose of the candidate grasping point dataset. Finally, the target grasping pose is selected from the multiple predicted grasping poses, and the grasping operation is performed. The specific processes involved are the same as those described above, and will not be repeated here.

[0107] The object grasping pose detection device provided by the present invention will be described below. The object grasping pose detection device described below can be referred to in correspondence with the object grasping pose detection method described above.

[0108] Reference Figure 9 This is a schematic diagram of the object grasping pose detection device provided by the present invention, as shown below. Figure 9 As shown, the object grasping pose detection device 900 includes:

[0109] The acquisition module 910 is used to acquire the seed point dataset of the target foreground object. Different seed point data in the seed point dataset represent different initial local features of the target foreground object at different initial positions in a cluttered scene.

[0110] The determination module 920 is used to determine the candidate grasping point dataset of the target foreground object based on the seed point dataset. The candidate grasping point dataset represents the number of regions corresponding to the seed point data where candidate grasping point data exists, as well as the candidate positions and candidate local features of the candidate grasping point data.

[0111] The prediction module 930 is used to predict the target grasping pose of the target foreground object based on the candidate grasping point dataset. The target grasping pose characterizes the predicted grasping pose of the candidate grasping point data clusters divided for the candidate grasping point dataset, and the quality score of the grasping pose meets the preset quality requirements.

[0112] Optionally, the acquisition module 910 can be used to acquire depth image information containing objects in a cluttered scene; perform point cloud conversion on the depth image information to acquire the original point cloud data of the objects; input the original point cloud data into a preset deep neural network model to obtain the seed point dataset of the target foreground objects output by the preset deep neural network model; wherein, the preset deep neural network model is used to extract features from the original point cloud data, and then to filter foreground objects from the candidate seed point dataset of the target foreground objects obtained by feature extraction, thereby obtaining the seed point dataset of the target foreground objects, and the candidate seed point dataset represents the different initial local features of the global point cloud of the object at different initial positions.

[0113] Optionally, the determining module 920 can be used to determine a first number of regions corresponding to the seed point data in the seed point dataset; based on the initial position and initial local features of the seed point data, predict a first number of region scores, a first number of candidate grasping point data, and candidate positions and candidate local features of each candidate grasping point data, representing the first number of regions containing feasible grasping poses; based on the first number of region scores, determine a second number of candidate grasping point data that meet the preset region score requirements from the first number of candidate grasping point data; based on the candidate positions and candidate local features of each candidate grasping point data in the second number of candidate grasping point data, determine the candidate grasping point dataset of the target foreground object, and the product of the second number and the total number of seed point data is the target number.

[0114] Optionally, the prediction module 930 can be used to cluster the candidate grasping point dataset based on the different candidate positions of different candidate grasping point data in the candidate grasping point dataset to determine multiple candidate grasping point data clusters; predict the grasping pose and grasping pose quality score of each candidate grasping point data cluster based on the different candidate local features of different candidate grasping point data; and predict the target grasping pose of the target foreground object from multiple grasping poses based on the quality score.

[0115] Optionally, the acquisition module 910 can also be used to acquire the object surface shape information of the target foreground object; based on the object surface shape information, the candidate seed point dataset of the target foreground object obtained by feature extraction is used to filter foreground objects and acquire the seed point dataset of the target foreground object.

[0116] Optionally, the acquisition module 910 can also be used to acquire RGB image information containing the target foreground object; perform image segmentation on the RGB image information to acquire the object surface shape information of the target foreground object.

[0117] Optionally, the determining module 920 can also be used to predict a first number of regions corresponding to the seed point data in the seed point dataset based on a preset region division model and the seed point dataset; wherein, the preset region division model is used to first establish a spherical coordinate system with each seed point data in the seed point dataset as the center and a preset length as the radius, and to divide the spherical coordinate system into regions based on a preset azimuth angle and a preset zenith angle, thereby determining the first number of regions corresponding to the seed point data in the seed point dataset.

[0118] Optionally, the determining module 920 can also be used to predict, based on different initial local features and different initial positions of different seed point data in a preset neural network model and seed point dataset, a first number of region scores representing feasible grasping poses in a first number of regions, a first number of candidate grasping point data, candidate positions and candidate local features of each candidate grasping point data; wherein, the preset neural network model is used to predict the position offset and local feature offset between the candidate grasping point data and the corresponding seed point data in different regions based on different initial local features, and to determine the candidate position and candidate local features of each candidate grasping point data based on the position offset, local feature offset, initial position and initial local features; and to predict the region scores containing feasible grasping poses in different regions based on different initial positions.

[0119] Figure 10 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 10 As shown, the electronic device 1000 may include: a processor 1010, a communication interface 1020, a memory 1030, and a communication bus 1040, wherein the processor 1010, the communication interface 1020, and the memory 1030 communicate with each other through the communication bus 1040. The processor 1010 can call logical instructions in the memory 1030 to execute a method for detecting the object grasping pose, the method including:

[0120] Obtain the seed point dataset of the target foreground object. Different seed point data in the seed point dataset represent different initial local features of the target foreground object at different initial positions in a cluttered scene.

[0121] Based on the seed point dataset, a candidate grab point dataset for the target foreground object is determined. The candidate grab point dataset represents the number of regions corresponding to the seed point data where candidate grab point data exists, as well as the candidate positions and candidate local features of the candidate grab point data.

[0122] Based on the candidate grasp point dataset, the target grasp pose of the target foreground object is predicted. The target grasp pose characterizes the predicted grasp pose for the candidate grasp point data clusters divided from the candidate grasp point dataset, and the quality score of the grasp pose meets the preset quality requirements.

[0123] Furthermore, the logical instructions in the aforementioned memory 1030 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0124] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the object grasping pose detection method provided by the above methods, the method comprising:

[0125] Obtain the seed point dataset of the target foreground object. Different seed point data in the seed point dataset represent different initial local features of the target foreground object at different initial positions in a cluttered scene.

[0126] Based on the seed point dataset, a candidate grab point dataset for the target foreground object is determined. The candidate grab point dataset represents the number of regions corresponding to the seed point data where candidate grab point data exists, as well as the candidate positions and candidate local features of the candidate grab point data.

[0127] Based on the candidate grasp point dataset, the target grasp pose of the target foreground object is predicted. The target grasp pose characterizes the predicted grasp pose for the candidate grasp point data clusters divided from the candidate grasp point dataset, and the quality score of the grasp pose meets the preset quality requirements.

[0128] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for detecting object grasping pose provided by the methods described above, the method comprising:

[0129] Obtain the seed point dataset of the target foreground object. Different seed point data in the seed point dataset represent different initial local features of the target foreground object at different initial positions in a cluttered scene.

[0130] Based on the seed point dataset, a candidate grab point dataset for the target foreground object is determined. The candidate grab point dataset represents the number of regions corresponding to the seed point data where candidate grab point data exists, as well as the candidate positions and candidate local features of the candidate grab point data.

[0131] Based on the candidate grasp point dataset, the target grasp pose of the target foreground object is predicted. The target grasp pose characterizes the predicted grasp pose for the candidate grasp point data clusters divided from the candidate grasp point dataset, and the quality score of the grasp pose meets the preset quality requirements.

[0132] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0133] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting a grasping pose of an object, characterized in that, The method comprises the following steps: obtaining a seed point data set of a target foreground object, wherein different seed point data in the seed point data set represent different initial local features of the target foreground object in different initial positions in a cluttered scene; based on the seed point data set, determining a candidate grasping point data set of the target foreground object, wherein the candidate grasping point data set represents that there are candidate grasping point data in a target number of regions corresponding to the seed point data, and the candidate position and candidate local feature of the candidate grasping point data; based on the candidate grasping point data set, predicting a target grasping pose of the target foreground object, wherein the target grasping pose represents that a candidate grasping point data cluster divided from the candidate grasping point data set predicts a grasping pose, and the quality score of the grasping pose meets a preset quality requirement; the method further comprises the following steps: determining a first number of regions corresponding to the seed point data in the seed point data set; based on the initial position and initial local feature of the seed point data, predicting a first number of region scores, a first number of candidate grasping point data, a candidate position and a candidate local feature of each candidate grasping point data, wherein the first number of region scores represent that the first number of regions respectively contain a feasible grasping pose; based on the first number of region scores, determining a second number of candidate grasping point data from the first number of candidate grasping point data, wherein the second number of candidate grasping point data meets a preset region score requirement; based on the candidate position and candidate local feature of each candidate grasping point data in the second number of candidate grasping point data, determining a candidate grasping point data set of the target foreground object, wherein the product of the second number and the total number of seed point data is the target number.

2. The method of object grasp pose detection according to claim 1, characterized in that, the method further comprises the following steps: obtaining depth image information of an object in a cluttered scene; performing point cloud conversion on the depth image information to obtain original point cloud data of the object; inputting the original point cloud data into a preset deep neural network model to obtain a seed point data set of the target foreground object output by the preset deep neural network model; wherein the preset deep neural network model is used for feature extraction on the original point cloud data, and then performs foreground object screening on a candidate seed point data set obtained by the feature extraction to obtain a seed point data set of the target foreground object, wherein the candidate seed point data set represents different initial local features of the global point cloud of the object in different initial positions.

3. The method of claim 1, wherein the method further comprises the following steps: based on different candidate positions of different candidate grasping point data in the candidate grasping point data set, clustering the candidate grasping point data set to determine a plurality of candidate grasping point data clusters; based on different candidate local features of the different candidate grasping point data, predicting a grasping pose of each candidate grasping point data cluster and a quality score of the grasping pose; based on the quality score, predicting a target grasping pose of the target foreground object from a plurality of grasping poses.

4. The method of claim 2, wherein The foreground object screening is performed on the candidate seed point data set of the object obtained through the feature extraction, and a seed point data set of the target foreground object is obtained. Obtain object surface shape information of the target foreground object; The foreground object screening is performed on the candidate seed point data set of the object obtained through the feature extraction based on the object surface shape information, and a seed point data set of the target foreground object is obtained.

5. The method of claim 1, wherein The determination of the first quantity of regions corresponding to the seed point data in the seed point data set comprises: Based on the preset region division model and the seed point data set, the first quantity of regions corresponding to the seed point data in the seed point data set is predicted; wherein the preset region division model is used to first establish a spherical coordinate system based on each seed point data in the seed point data set as a spherical center and a preset length as a radius, and then divide the spherical coordinate system based on a preset azimuth angle and a preset zenith angle to determine the first quantity of regions corresponding to the seed point data in the seed point data set.

6. The method of claim 1, wherein The first quantity of region scores, the first quantity of candidate grasp point data, the candidate position and the candidate local feature of each candidate grasp point data in the first quantity of regions which respectively contain a feasible grasp pose are predicted based on the seed point data, comprising: Based on the preset neural network model and the different initial local features and the different initial positions of the different seed point data in the seed point data set, the first quantity of region scores, the first quantity of candidate grasp point data, the candidate position and the candidate local feature of each candidate grasp point data in the first quantity of regions which respectively contain a feasible grasp pose are predicted; Wherein, the preset neural network model is used to predict the position offset and the local feature offset between the candidate grasp point data in different regions and the corresponding seed point data based on the different initial local features, to determine the candidate position and the candidate local feature of each candidate grasp point data based on the position offset, the local feature offset, the initial position and the initial local feature; and to predict the region score of the region containing a feasible grasp pose in different regions based on the different initial positions.

7. A device for detecting the posture of an object grasping action, characterized in that, Comprise: An acquisition module is configured to acquire a seed point data set of a target foreground object, wherein different seed point data in the seed point data set represent different initial local features of the target foreground object at different initial positions in a cluttered scene. A determination module is configured to determine, based on the seed point data set, a candidate grasp point data set of the target foreground object, wherein the candidate grasp point data set represents that candidate grasp point data exist in a target quantity of regions corresponding to the seed point data and candidate positions and candidate local features of the candidate grasp point data. A prediction module is configured to predict, based on the candidate grasp point data set, a target grasp pose of the target foreground object, wherein the target grasp pose represents that a predicted grasp pose of a candidate grasp point data cluster divided for the candidate grasp point data set meets a preset quality requirement in quality score. The determination module is specifically configured to: determine a first quantity of regions corresponding to seed point data in the seed point data set; based on the initial position and the initial local feature of the seed point data, predict a first quantity of region scores, a first quantity of candidate grasp point data, a candidate position and a candidate local feature of each of the candidate grasp point data, which represent that the first quantity of regions contain the feasible grasp poses respectively; based on the first quantity of region scores, determine a second quantity of candidate grasp point data from the first quantity of candidate grasp point data, which meet a preset region score requirement; based on the candidate position and the candidate local feature of each of the second quantity of candidate grasp point data, determine a candidate grasp point data set of the target foreground object, and a product of the second quantity and a total quantity of the seed point data is the target quantity.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the object grasp pose detection method of any one of claims 1-6 when executing the program. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program implements the object grasp pose detection method of any one of claims 1-6 when executed by the processor.

Citation Information

Patent Citations

  • Real-time pose estimation method and positioning grabbing system for three-dimensional target object

    CN110648361A