A pose estimation method based on contour extraction point pair features
By using the pose estimation method based on contour extraction points in disordered grab recognition, the problem of low recognition rate in close adjacent and stacked scenarios is solved, and higher matching accuracy and accuracy are achieved, and the success rate of the disordered grab process is improved.
Patent Information
- Application Number
- CN202211318552.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-10-26
AI Technical Summary
The prior art in disordered grab recognition, especially in closely adjacent and stacked scenarios, has a low recognition rate, resulting in insufficient pose estimation accuracy and accuracy, affecting the success rate of the entire disordered grab process.
The pose estimation method based on the contour extraction point pair features is adopted. Point clouds are photographed through a three-dimensional camera and contour extraction is performed. The features are extracted and stored in a hash table. The scene is contour extraction is performed on the online recognition stage, point-to-feature features are calculated, and the pose estimation is completed through the hash table matching, and the pose estimation is completed in combination with ICP iterative processing.
By performing contour extraction of both templates and scenes, the matching accuracy and accuracy are improved, the matching time is reduced, and the efficiency and accuracy of disorderly grabbing and recognition are enhanced.
Smart Images

Figure CN115527202B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of disordered grasping and recognition, and more specifically, to a pose estimation method based on contour extraction point pair features. Background Art
[0002] In the case of disordered grasping and recognition, the Point Pair Feature (PPF) algorithm and the LCCP algorithm are commonly used. The LCCP algorithm cannot recognize objects in closely adjacent scenes when the distance is less than 4mm or when the point cloud is incomplete, and has high requirements for imaging quality. The Point Pair Feature (PPF) algorithm is divided into two parts: offline modeling and online recognition. In offline modeling, a point cloud template is made by shooting a point cloud with a 3D camera, and then feature extraction is performed and stored in a hash table. During online recognition, a certain proportion of points are selected as reference points in the scene point cloud. According to the given angle threshold and distance limit, the four-dimensional features of the point pairs formed by the reference points and the scene points are calculated. Then, the hash table obtained by offline modeling is searched, and the model point pair features that are similar to the scene point pair features are extracted. The transformation relationship between the two point pairs is calculated, and then pose voting is performed. The results are sorted by the number of votes, and each result must be compared with other results in turn. Poses with higher votes are more likely to become candidate poses. The coordinates of points with higher votes are the optimal solution. Finally, ICP precise matching is performed to complete pose estimation. The PPF algorithm can handle many complex scenes, but in closely adjacent and stacked scenes, it is difficult for the template to find the corresponding point cloud match in the scene, and the correct grasping point cannot be calculated, resulting in a low recognition rate in such application scenarios, reducing the success rate of the entire disordered grasping process, and the matching rate is low.
[0003] The prior art discloses a method for improving the recognition and positioning accuracy of automobile sheet metal workpieces, including: 1. Obtaining a complete scene image of the workpiece to be grasped and preprocessing it, completing instance segmentation of the workpiece in the image, and extracting the edge of the two-dimensional image according to the instance segmentation result; 2. Preprocessing the point cloud data of the scene of the workpiece to be grasped, further extracting the edge of the point cloud data according to the edge of the extracted two-dimensional image as an index; and calculating the point pair features in the edge of the point cloud to establish a global model description; 3. Performing online model matching, using a voting method based on the Hough voting principle to obtain candidate poses, and using a connectivity density clustering algorithm to cluster the candidate poses, and using an ICP registration algorithm to optimize the poses. In this scheme, only two-dimensional image extraction is performed on the scene image, and finally the scene contour and template are used for matching, which increases the time consumption, and has a high matching error rate and a low accuracy rate. Summary of the invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for pose estimation based on contour extraction point pair features to improve the efficiency and accuracy of pose estimation.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0006] A method for posture estimation based on contour extraction point pair features is provided, comprising the following steps:
[0007] S1: Use a 3D camera to shoot point clouds to create template point clouds, extract contours, extract features, and store them in a hash table;
[0008] S2: Obtain the scene point cloud of the workpiece to be grasped through a 3D camera, extract the contour of the scene point cloud, select a reference point in the scene point cloud, and calculate the four-dimensional features of the scene point pair composed of the reference point and all scene points according to the angle threshold and distance limit;
[0009] S3: Search the hash table in step S1, extract the model point pair features that are similar to the four-dimensional features of the field scenic spot pair in step S2, and obtain the candidate pose;
[0010] S4: Vote for the candidate poses and obtain the voting results;
[0011] S5: Perform ICP iteration processing using the voting result in step S4 as the initial pose to complete pose estimation.
[0012] The method for posture estimation based on contour extraction point pair features of the present invention comprises two stages: an offline modeling stage and an online recognition stage. In the offline modeling stage, the contour of the template is extracted, and then the feature is extracted and stored in a hash table; in the online recognition stage, the contour of the scene is extracted, and then the feature is extracted, and point pair features are established with the features stored in the hash table by offline modeling, and then posture voting is performed, and then the posture estimation is completed through ICP iterative processing. The present invention extracts contours from both the template and the scene, and uses the contours to replace the original template and scene point cloud for point pair feature matching, which saves time and improves matching accuracy and correctness.
[0013] Preferably, in step S1 and step S2, when performing contour extraction on the template point cloud and the scene point cloud, the PCL contour extraction method is adopted, the angle threshold limit and the search point limit are given, the normal and search method of the boundary estimation are set, and the contour extraction of the template and the scene is performed.
[0014] Preferably, in step S1, in order to achieve fast matching, the four intervalized feature values are used as indexes after operation, and the point pair features of any two points in the template are calculated and stored in a hash table.
[0015] Preferably, the calculation process of the point pair feature between any two points in the template is specifically as follows: feature extraction is performed on the template point cloud contour, and two arbitrary points m in space are 1 、m 2 The normal vectors are n 1 、n 2 , then m 1 、m 2 The point pair feature F(m 1 , m 2 ):
[0016] F(m 1 , m 2 )=(||d|| 2 ,∠(n 1 , d), ∠(n 2 , d), ∠(n 1 , n 2 ))
[0017] Where d represents m 1 、m 2 The distance between them; ||d|| 2 is m 1 With m 2 The Euclidean distance between 1 , d) represent n 1 With m 1 、m 2 The angle between the lines; ∠(n 2 , d) represents n 2 With m 1 、m 2 The angle between the lines; ∠(n 1 , n 2 ) represents n 1 and n 2 The angle between n 1 and n 2 The angle range is between [0, π].
[0018] Preferably, in step S4, the specific process of voting for the candidate poses is: voting for the four-dimensional feature F(S r , S i ) for discretization;
[0019] Using the discretized result as the index key, search for the model point pair feature F(m r , m i ), calculate the scene pair (S r , S i ) and model point pairs (m r , m i )’s conversion relationship:
[0020]
[0021] Among them, S i are all scene points, S r is the scene reference point, R x (α) is the model point m i Rotate the scene point S around the x-axis by an angle α i The coincident matrix, T m→g They are the transformation relationship between the reference point normal vector and the model point normal vector after they are rotated to align with the x-axis of the local reference coordinate system;
[0022] Calculate the rotation angle α and add 1 to the established two-dimensional accumulator;
[0023] Traverse all points after contour extraction and repeat the above steps.
[0024] Preferably, in step S5, before performing ICP optimization, a hierarchical clustering algorithm is used to cluster the voting results to reduce the number of voting results.
[0025] Preferably, before performing ICP optimization, the template and the scene are downsampled.
[0026] Preferably, the specific steps of downsampling are: given a template point cloud threshold and a scene point cloud threshold, when the template point cloud and the scene point cloud exceed the template point cloud threshold and the scene point cloud threshold respectively, downsampling is performed, and a part of the data is selected from the majority set in proportion and recombined with the minority set into a new data set.
[0027] Preferably, the specific process of ICP iterative processing is: the downsampled scene point cloud and the template point cloud are applied to the initial pose together, and the error is reduced by continuously minimizing the distance between the model point and the scene point through iterative processing.
[0028] Preferably, a threshold is set, and the accuracy of the pose is judged by whether the distance between the model point cloud and the scene point cloud is less than the given threshold; when the distance between the model point cloud and the scene point cloud is less than the threshold, the correctness of the pose is evaluated by the ratio of the number of successful matching points to the number of scene object points.
[0029] Compared with the background technology, the pose estimation method based on contour extraction point pair features of the present invention has the following beneficial effects:
[0030] By extracting contours from both the template and the scene, the contours are used to replace the original template and scene point cloud for point-to-point feature matching, which saves time and improves matching accuracy and correctness. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1This is a flow chart of a method for posture estimation based on contour extraction point pair features in Embodiment 1 of the present invention;
[0032] Figure 2 This is a schematic diagram of a closely adjacent scene in Embodiment 3 of the present invention;
[0033] Figure 3 This is a schematic diagram of a stacking scenario in Embodiment 3 of the present invention;
[0034] Figure 4 is a schematic diagram of a metal sheet in Embodiment 3 of the present invention;
[0035] Figure 5 Schematic diagram of the outline of closely adjacent scenes in Embodiment 3 of the present invention;
[0036] Figure 6 This is a schematic diagram of the outline of the stacking scene in the third embodiment of the present invention;
[0037] Figure 7 Schematic diagram of the outline of the metal sheet in the third embodiment of the present invention. DETAILED DESCRIPTION
[0038] The present invention is further described below in conjunction with specific implementation methods. The accompanying drawings are only used for exemplary descriptions and are only schematic diagrams, not actual drawings, and cannot be understood as limiting this patent; in order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.
[0039] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if the terms "upper", "lower", "left", "right" and the like indicate directions or positional relationships based on the directions or positional relationships shown in the drawings, it is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limitations on this patent. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.
[0040] Embodiment 1
[0041] A pose estimation method based on contour extraction point pair features, such as Figure 1 As shown, the following steps are included:
[0042] S1: Use a 3D camera to shoot point clouds to create template point clouds, extract contours, extract features, and store them in a hash table;
[0043] S2: Obtain the scene point cloud of the workpiece to be grasped through a 3D camera, extract the contour of the scene point cloud, select a reference point in the scene point cloud, and calculate the four-dimensional features of the scene point pair composed of the reference point and all scene points according to the angle threshold and distance limit;
[0044] S3: Search the hash table in step S1, extract the model point pair features that are similar to the four-dimensional features of the field scenic spot pair in step S2, and obtain the candidate pose;
[0045] S4: Vote for the candidate poses and obtain the voting results;
[0046] S5: Perform ICP iteration processing using the voting result in step S4 as the initial pose to complete pose estimation.
[0047] The above-mentioned pose estimation method based on contour extraction point pair features includes two stages: an offline modeling stage and an online recognition stage. In the offline modeling stage, the template is contour extracted, and then feature extraction is performed and stored in a hash table; in the online recognition stage, the scene is contour extracted, and then feature extraction is performed, and point pair features are established with the features stored in the hash table by offline modeling, and then pose voting is performed, and then ICP iterative processing is performed to complete the pose estimation. In this embodiment, contour extraction is performed on both the template and the scene, and the contour is used to replace the original template and scene point cloud for point pair feature matching, which is time-saving and improves the matching accuracy and accuracy.
[0048] Although the contour extraction method in 2D images is very mature and has high extraction accuracy, the 3D point cloud has more information than the 2D image. The feature relationship established between the 2D image and the 3D point cloud is not accurate enough. Converting the 3D point cloud into a 2D image will cause serious information loss. Therefore, in step S1 and step S2, when performing contour extraction on the template point cloud and the scene point cloud, the PCL contour extraction method is used, and the angle threshold limit and the search point limit are given. The normal and search method of the boundary estimation are set to extract the contour of the template and the scene.
[0049] In step S1, for fast matching, the four intervalized feature values are used as indexes after operation, and the point pair features of any two points in the template are calculated and stored in a hash table.
[0050] The calculation process of the point pair feature between any two points in the template is as follows: feature extraction is performed on the template point cloud contour, and two arbitrary points m in space are 1 、m 2 The normal vectors are n 1 、n 2 , then m 1 、m 2 The point pair feature F(m 1 , m2 ):
[0051] F(m 1 , m 2 )=(||d|| 2 ,∠(n 1 , d), ∠(n 2 , d), ∠(n 1 , n 2 ))
[0052] Where d represents m 1 、m 2 The distance between them; ||d|| 2 is m 1 With m 2 The Euclidean distance between 1 , d) represent n 1 With m 1 、m 2 The angle between the lines; ∠(n 2 , d) represents n 2 With m 1 、m 2 The angle between the lines; ∠(n 1 , n 2 ) represents n 1 and n 2 The angle between n 1 and n 2 The angle range is between [0, π].
[0053] In step S4, the specific process of voting for candidate poses is as follows: the four-dimensional feature F(S r , S i ) for discretization;
[0054] Using the discretized result as the index key, search for the model point pair feature F(m r , m i ), calculate the scene pair (S r , S i ) and model point pairs (m r , m i )’s conversion relationship:
[0055]
[0056] Among them, S i are all scene points, S r is the scene reference point, R x (α) is the model point m i Rotate the scene point S around the x-axis by an angle α i The coincidence matrix, Ts→g 、T m→g They are the transformation relationship between the reference point normal vector and the model point normal vector after they are rotated to align with the x-axis of the local reference coordinate system;
[0057] Calculate the rotation angle α and add 1 to the established two-dimensional accumulator;
[0058] Traverse all points after contour extraction and repeat the above steps.
[0059] In step S5, before performing ICP optimization, a hierarchical clustering algorithm is used to cluster the voting results to reduce the number of voting results.
[0060] Embodiment 2
[0061] This embodiment is similar to the first embodiment, except that before performing ICP optimization, the template and the scene are downsampled to reduce the number of template point clouds and scene point clouds, avoid computational redundancy, and reduce matching time.
[0062] The specific steps of downsampling are as follows: given the template point cloud threshold and the scene point cloud threshold, when the template point cloud and the scene point cloud exceed the template point cloud threshold and the scene point cloud threshold respectively, downsampling is performed, and a part of the data is selected from the majority set in proportion and recombined with the minority set into a new data set.
[0063] The specific process of ICP iterative processing is: the downsampled scene point cloud and template point cloud are applied to the initial pose together, and the distance between the model point and the scene point is continuously minimized through iterative processing to reduce the error.
[0064] A threshold is set, and the accuracy of the pose is judged by whether the distance between the model point cloud and the scene point cloud is less than the given threshold; when the distance between the model point cloud and the scene point cloud is less than the threshold, the correctness of the pose is evaluated by the ratio of the number of successful matching points to the number of scene object points.
[0065] Embodiment 3
[0066] This embodiment is for pose estimation of closely adjacent scenes and stacked scenes.
[0067] like Figure 2 , Figure 3 As shown, the closely adjacent scene and the stacked scene each contain 5 metal sheets. Figure 4 As shown in , the contours of closely adjacent scene point clouds, stacked scene point clouds and template point clouds are extracted respectively, as shown in Figures 5 to 7 As shown, the point cloud information before and after contour extraction is shown in Table 1:
[0068] Table 1 Point cloud information before and after contour extraction
[0069] Point cloud type Number of original scene point clouds Number of scene point clouds after contour extraction Template point cloud 9194 607 Closely adjacent scene point cloud 70305 2072 Stacking scene point clouds 58494 3290
[0070] PPF, LCCP and contour PPF are used to match the point clouds of closely adjacent scenes and stacked scenes respectively. In the closely adjacent scenes, PPF and LCCP cannot obtain accurate matching structures, while contour PPF based on contour extraction can identify the metal flakes on the edge; in the stacked scenes, PPF identifies 6 metal flakes, which interferes with the disordered grasping, while LCCP cannot identify the metal flakes in the stacked scenes, while contour PPF can accurately obtain matching results. Contour PPF solves the problem of difficult identification of closely adjacent scenes and stacked scenes in disordered grasping.
[0071] In the process of pose estimation by contour PPF, the threshold is set to 3000. As can be seen from Table 1, the number of template point clouds, closely adjacent point clouds and stacked scene point clouds is greater than 3000, and the template point clouds, closely adjacent point clouds and stacked scene point clouds are downsampled.
[0072] The metal sheets in closely adjacent scenes and stacked scenes are randomly grasped. The results of random grasping are shown in Table 2:
[0073] Table 2 Success rate and accuracy of metal sheet grabbing
[0074] algorithm Crawl times Crawl success rate Average single crawling time (s) Recognizable accuracy PPF 200 86% 22.33 8mm LCCP 61 93.4% 21.57 5mm Contour PPF 331 96.37% 21.81 ≤1mm
[0075] As shown in Table 2, PPF and contour PPF can stably grasp for more than one hour, while the LCCP algorithm cannot recognize the metal flakes after grasping 61 times, resulting in a pause in grasping; compared with the PPF algorithm, the grasping success rate of contour PPF is increased by 10.37%, and the average single grasping retrieval time is 0.52s; compared with the LCCP algorithm, the grasping success rate of contour PPF is increased by 2.97%, but the average single grasping time is increased by 0.24s. However, the LCCP algorithm can only recognize metal flakes with a spacing greater than 5mm, and has high requirements on the algorithm imaging quality and poor stability. The original PPF can only recognize metal flakes with a spacing greater than 8mm, and the contour PPF can recognize metal flakes with a spacing less than 1mm, with high accuracy.
[0076] In the specific contents of the above-mentioned specific implementation methods, the various technical features can be combined in any non-contradictory manner. In order to make the description concise, not all possible combinations of the above-mentioned technical features are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0077] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the embodiments here. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A pose estimation method based on contour extraction point pair features, It is characterized in that The following steps are involved: S1: Use a 3D camera to shoot point clouds to create template point clouds, extract contours, extract features, and store them in a hash table; S2: Obtain the scene point cloud of the workpiece to be grasped through a 3D camera, extract the contour of the scene point cloud, select a reference point in the scene point cloud, and calculate the four-dimensional features of the scene point pair composed of the reference point and all scene points according to the angle threshold and distance limit; S3: Search the hash table in step S1, extract the model point pair features that are similar to the four-dimensional features of the field scenic spot pair in step S2, and obtain the candidate pose; S4: Vote for the candidate poses and obtain the voting results; S5: Perform ICP iteration processing using the voting result in step S4 as the initial pose to complete pose estimation; In step S1 and step S2, when performing contour extraction on the template point cloud and the scene point cloud, the PCL contour extraction method is adopted, the angle threshold limit and the search point limit are given, the normal line and the search method of the boundary estimation are set, and the contour extraction of the template and the scene is performed; In step S1, in order to quickly match, the four intervalized feature values are used as indexes after operation, and the point pair features of any two points in the template are calculated and stored in the hash table; In step S1, the calculation process of the point pair feature between any two points in the template is as follows: feature extraction is performed on the template point cloud contour, and two arbitrary points in space are , The normal vectors are , ,but , Point pair features between : in, express , The distance between yes and The Euclidean distance between Respectively and , The angle between the lines; express and , The angle between the lines; express and The angle between and The angle range is between; In step S4, the specific process of voting for candidate poses is as follows: the four-dimensional features of the scene point pairs Discretize Using the discretized result as the index key, search for model point pair features that are similar to the four-dimensional features of the scene point pair , calculate the appearance of the scenic spot And model point pair The conversion relationship is: in, are all scene points, is the scene reference point, It is a model point Rotate around the x-axis Angle and scene point The coincident matrix, , They are the transformation relationship between the reference point normal vector and the model point normal vector after they are rotated to align with the x-axis of the local reference coordinate system; Calculate the rotation angle , add 1 to the established two-dimensional accumulator; Traverse all points after contour extraction and repeat the above steps.
2. The method for posture estimation based on contour extraction point pair features according to claim 1, It is characterized in that In step S5, before performing ICP optimization, a hierarchical clustering algorithm is used to cluster the voting results to reduce the number of voting results.
3. The method for posture estimation based on contour extraction point pair features according to claim 1 or 2, It is characterized in that In step S5, before performing ICP optimization, the template and the scene are downsampled.
4. The method for posture estimation based on contour extraction point pair features according to claim 3, It is characterized in that The specific steps of downsampling are as follows: given the template point cloud threshold and the scene point cloud threshold, when the template point cloud and the scene point cloud exceed the template point cloud threshold and the scene point cloud threshold respectively, downsampling is performed, and a part of the data is selected from the majority set in proportion and recombined with the minority set into a new data set.
5. The method for posture estimation based on contour extraction point pair features according to claim 3, It is characterized in that In step S5, the specific process of ICP iterative processing is: the downsampled scene point cloud and the template point cloud are applied to the initial pose together, and the distance between the model point and the scene point is continuously minimized through iterative processing to reduce the error.
6. The method for posture estimation based on contour extraction point pair features according to claim 5, It is characterized in that In step S5, a threshold is set, and the accuracy of the pose is judged by whether the distance between the model point cloud and the scene point cloud is less than the given threshold; when the distance between the model point cloud and the scene point cloud is less than the threshold, the correctness of the pose is evaluated by the ratio of the number of successfully matched points to the number of scene object points.
Citation Information
Patent Citations
Object pose measurement method and device and storage medium
CN111598946A
Method for improving recognition and positioning precision of automobile sheet metal workpiece
CN113538486A