A method for constructing a 6D pose dataset for common industrial parts
By combining a depth camera with a pose tracking board, the 6D pose of common industrial parts can be automatically labeled, solving the time-consuming, labor-intensive, and low-precision labeling issues in industrial scenarios. This generates a high-precision dataset suitable for robotic automated grasping and virtual assembly.
Patent Information
- Application Number
- CN202211472379.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-23
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-11-23
AI Technical Summary
Existing technologies make it difficult to efficiently construct 6D pose datasets for common industrial parts in industrial scenarios. In particular, automated labeling in cluttered scenes is inaccurate and time-consuming and labor-intensive, and traditional methods are easily affected by occlusion and weak texture features.
By combining a depth camera with a pose tracking board, removing abnormal information through plane fitting and the RANSAC-PnP algorithm, and using augmented reality for virtual-reality interactive registration, the 6D pose of common industrial parts is automatically labeled to generate a high-precision dataset.
It realizes the automatic and accurate labeling of the 6D pose of industrial parts, reduces the cost of manual labeling, and improves calibration accuracy. It is suitable for industrial scenarios such as robot automated grasping and virtual assembly.
Smart Images

Figure CN115761407B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method for constructing a 6D pose dataset of general industrial parts. Background Art
[0002] Estimating the 6D pose of common industrial parts from videos and images has broad and important application value in the Industry 4.0 era. The ground-truth poses in 6D pose datasets provide ample validation for pose estimation and tracking algorithms. More importantly, deep learning-based pose estimation and tracking algorithms have proliferated in recent years, driven by the rapid development of machine learning in computer vision. Therefore, in addition to validation, pose datasets also provide a wealth of training data for deep learning-based pose estimation and tracking algorithms. However, currently, few 6D pose datasets built on common industrial parts are available, significantly limiting the development of key industrial technologies such as robotic automated grasping and virtual assembly.
[0003] Traditional manual pose labeling methods cannot meet the needs of the large number and variety of industrial parts. Manually labeling dozens of parts in thousands of images often consumes a lot of manpower and resources. Therefore, there is an urgent need for an automated labeling method that can generate large 6D pose datasets in a short time with minimal manpower and resources.
[0004] At present, a widely used 6D pose automatic labeling method is to use the precise tracking of planar markers such as AprilTag to indirectly obtain the 6D pose of the object. Specifically, the relative pose between the object and the AprilTag is first determined, and then the AprilTag is tracked in the video stream using the marker detection toolkit to indirectly obtain the 6D labeling pose of the object. However, for industrial scenarios, there are often cases where the AprilTag is blocked by parts or cluttered scenes: when the AprilTag is blocked by a large area, this automatic labeling method will not be able to proceed; if the AprilTag is blocked by a small area, the marker detection toolkit may provide abnormal AprilTag information, which seriously reduces the accuracy of the part labeling pose.
[0005] On the other hand, this automatic annotation method often uses a PnP algorithm to determine the pose of the part relative to the landmark by finding corresponding feature points on the image and the object. The accuracy of this method is highly susceptible to human annotation errors and a small number of feature point pairs. Furthermore, parts have the unique property of weak textures compared to other objects, and many parts, such as rotary shafts, lack sharp corners. Therefore, when constructing a 6D pose dataset for common industrial parts, it is often difficult to find enough feature point pairs for some parts to use the PnP algorithm. Summary of the Invention
[0006] In order to solve the above problems, the present invention proposes a method for constructing a 6D pose dataset of general industrial parts, which realizes the automatic annotation of the 6D poses of multiple parts in a large number of images in industrial cluttered scenes, and solves the problem of time-consuming and labor-intensive construction and poor accuracy of pose datasets.
[0007] The specific technical solutions created by the present invention are as follows:
[0008] 1. A method for constructing a 6D pose dataset for common industrial parts
[0009] S1: Fix the depth camera and pose tracking board in a specified industrial scene, use the depth camera to shoot the pose tracking board in the current scene, and obtain the initial RGB detection image;
[0010] S2: Extract and obtain the pose information of each plane marker AprilTag in the initial RGB detection image, and use the tracking board pose solution algorithm to process the pose information of each plane marker AprilTag to obtain the initial pose cMb of the pose tracking board 0 ;
[0011] S3: Initial pose cMb of the tracking board according to the pose 0 , using the virtual-real interactive registration method based on prior pose specification and augmented reality to place multiple industrial general parts on the pose tracking board and obtain the corresponding placement poses bMo of multiple industrial general parts i ;
[0012] S4: Use the depth camera in S1 in RGB-D video recording mode to record RGB-D videos from multiple angles that include all common industrial parts and the full view of the pose tracking board;
[0013] S5: Repeat S2 to solve and obtain the tracking board pose cMb of all RGB-D images in the RGB-D video. J , and then according to the corresponding placement postures of multiple industrial general parts bMo i , calculate and obtain the calibration pose cMo corresponding to multiple industrial general parts of all frames of RGB-D images i J ;
[0014] S6: Based on the 3D models of multiple common industrial parts and their corresponding calibration poses cMo in all RGB-D image frames, obtain the part numbers, minimum bounding boxes, and segmentation masks in all RGB-D image frames to form a 6D pose dataset for the current scene. This 6D pose dataset is used to train and validate a deep learning-based 6D pose estimation / tracking algorithm. The trained algorithm can be applied to industrial scenarios such as robotic automated grasping and virtual assembly.
[0015] In S1, the posture tracking board includes an original tracking board and a plurality of planar markers AprilTags, which are fixedly mounted on the four edges of the surface of the original tracking board at equal intervals, and all of the plurality of planar markers AprilTags belong to the same family.
[0016] The S2 is specifically:
[0017] S21: Use the AprilTag detection toolkit to extract the pose information of each plane marker AprilTag in the initial RGB detection image and record it as the camera coordinate system feature of the pose tracking board. The pose information of each plane marker AprilTag includes the number of each plane marker AprilTag, its own corner point sequence number, corner point pixel coordinates [u corn l , v corn l ] and 6D pose matrix;
[0018] S22: Extract the coordinates of the AprilTag center and its unit vector in the z-axis direction from the 6D pose matrix of each plane marker AprilTag The center point coordinates with normal vectors are composed, and the plane marker AprilTag point cloud is generated by the center point coordinates with normal vectors of all plane marker AprilTags;
[0019] S23: Based on the AprilTag point cloud of the planar marker, the plane fitting method based on principal component analysis and least squares is used to obtain the spatial fitting plane where the pose tracking board is located in the camera coordinate system.
[0020] S24: Fitting the plane in the space where the tracking board is located according to the camera coordinate system Calculate the center of each plane marker AprilTag and the spatial fitting plane where the pose tracking board is located in the camera coordinate system The distance d between them is calculated, and the plane markers AprilTag whose distance d exceeds the preset distance threshold are removed to obtain the AprilTag set Γ1 of each plane marker after removing the abnormal distance;
[0021] S25: Calculate the spatial fitting plane The unit normal vector Then calculate the z-axis unit vector of each plane marker AprilTag in the camera coordinate system in the AprilTag set Γ1 after removing the abnormal distance Fitting a plane to space The unit normal vector The angle θ between them is removed, and the AprilTag of the plane markers whose angle θ exceeds the preset angle threshold is removed to obtain the AprilTag set Γ2 of each plane marker after removing the abnormal angle, and then the camera coordinate system features of the pose tracking board are updated according to the AprilTag set Γ2 of each plane marker after removing the abnormal angle;
[0022] S26: Establishing the object coordinate system of the pose tracking board, determining the number of each plane marker AprilTag in the pose tracking board, the corner point sequence number and the coordinates of all corner points in the object coordinate system according to the object coordinate system of the pose tracking board, and recording this information as the object coordinate system feature of the pose tracking board;
[0023] S27: Based on the updated camera coordinate system features and object coordinate system features of the pose tracking board, establish 2D-3D feature relationships for all corner points of each plane marker AprilTag, obtain corresponding 2D-3D feature point pairs, and form a pose feature set from all 2D-3D feature point pairs;
[0024] S28: Based on the pose feature set, the RANSAC-PnP relative pose solution algorithm is used to calculate and obtain the pose of the pose tracking board relative to the camera coordinate system, which is recorded as the initial pose cMb of the pose tracking board. 0 .
[0025] The AprilTag detection toolkit includes AprilTag 3, AprilTags C++, and VISP.
[0026] The S3 is specifically:
[0027] S31: Generate 3D models of multiple common industrial parts and their corresponding 3D model point clouds, and specify the prior placement poses bMo corresponding to the 3D models of multiple common industrial parts i p ;
[0028] S32: bMo based on the prior placement of multiple common industrial parts i p and the initial pose cMb of the pose tracking board 0 Calculate the prior pose cMo of multiple common industrial parts in the camera coordinate system ip ;
[0029] S33: Based on the prior pose cMo of multiple common industrial parts in the camera coordinate system i p , transform the corresponding 3D model point cloud to obtain the 3D model point cloud P in the camera coordinate system i vir ;
[0030] S34: Use augmented reality methods to transform the three-dimensional models of multiple common industrial parts in S31 into the corresponding prior poses cMo in the camera coordinate system i p Rendering: After obtaining the corresponding rendered image, it is presented in real time to the virtual-reality interactive interface of the pose tracking board. Multiple common industrial parts are placed on the pose tracking board and the poses of the corresponding parts are adjusted according to the rendered images of the parts, thus achieving coarse virtual-reality registration of the placement poses of multiple common industrial parts.
[0031] S35: Use the depth camera to obtain the depth point cloud of the general industrial parts placed on the pose tracking board, and after the coordinate conversion of the depth point cloud, obtain the three-dimensional actual point cloud P corresponding to multiple general industrial parts in the camera coordinate system. i real ;
[0032] S36: Using the iterative closest point ICP algorithm to calculate the 3D model point cloud P of multiple common industrial parts in the camera coordinate system i vir And the three-dimensional actual point cloud P in the camera coordinate system i real Perform virtual-real precise registration to obtain the real parts and prior placement poses bMo corresponding to multiple common industrial parts i p The relative pose between the three-dimensional models under i p Mo i , and then calculate the placement pose bMo corresponding to multiple industrial general parts after precise registration i , recorded as the placement pose bMo corresponding to multiple industrial general parts i .
[0033] 2. A storage medium
[0034] A computer program is stored, and when the computer program is executed by a processor, the method described is implemented.
[0035] 3. A storage medium
[0036] The computer program described herein is an instruction corresponding to the method.
[0037] The beneficial effects of the present invention are:
[0038] (1) The present invention realizes the automatic annotation of 6D positioning poses of industrial parts, solving the time-consuming and labor-intensive problems caused by the large amount of images in the data set and the large number of parts in most industrial application scenarios;
[0039] (2) The present invention utilizes the designed AprilTag posture tracking board to successfully solve the problem of automatic labeling being impossible due to a small number of AprilTags being blocked by a large area in cluttered industrial scenes;
[0040] (3) When solving the tracking board's pose, the present invention first uses a plane fitting method to eliminate abnormal AprilTag estimation information, avoiding the introduction of corner point information with excessive errors in the subsequent calculation of the tracking board's pose, and greatly reducing the error impact caused by partial occlusion of the AprilTag in industrial cluttered scenes;
[0041] (4) The present invention uses the RANSAC-PnP algorithm to solve the tracking plate pose in the process of solving the tracking plate pose, reducing the influence of AprilTag estimation error caused by factors such as image blur and the long distance between the camera and the tracking plate, and indirectly improving the accuracy of the part calibration pose;
[0042] (5) The present invention uses virtual-reality interactive registration based on prior pose specification and augmented reality to obtain the part placement pose, avoiding the problems of difficult, slow and inaccurate part placement pose calibration caused by the weak texture characteristics of industrial parts and the lack of sufficient obvious features of some parts.
[0043] In summary, the present invention can avoid time-consuming and labor-intensive manual labeling when constructing a 6D pose dataset for general industrial parts, and realize automatic and accurate labeling of the 6D position and pose of parts in video streams, which has good engineering practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is a flowchart for constructing the 6D pose dataset of industrial general parts in the present invention.
[0045] Figure 2 These are eight common industrial parts used in the embodiments of the present invention.
[0046] Figure 3 It is an AprilTag posture tracking board designed in an embodiment of the present invention.
[0047] Figure 4 Schematic diagram of the AprilTag posture information detected by the AprilTag detection toolkit of the present invention.
[0048] Figure 5This is a schematic diagram of the extreme errors encountered when detecting the AprilTag in the present invention, as well as a rendering of the part marking pose before and after the implementation of the tracking board pose solution algorithm.
[0049] Figure 6 This is a schematic diagram of the coarse registration process based on prior pose specification and augmented reality in the present invention.
[0050] Figure 7 It is a set of RGB-D images and their annotation information in the dataset constructed in this invention.
[0051] Figure 8 These are the other two scenarios of the data set constructed in the embodiment of the present invention. DETAILED DESCRIPTION
[0052] The present invention is further explained below by describing the process of constructing a 6D pose dataset for eight common industrial parts in a robotic arm operating console scenario:
[0053] The present invention is applied to the process of constructing a 6D pose dataset of industrial general parts with a robot arm operating table as the scene. The flowchart of this embodiment is as follows: Figure 1 As shown, the following steps are included:
[0054] S1: Fix the depth camera and pose tracking board in a specified industrial scene, use the depth camera to shoot the pose tracking board in the current scene, and obtain an initial RGB detection image. The initial RGB detection image contains the full image of the pose tracking board;
[0055] like Figure 3 As shown, in S1, the posture tracking board includes an original tracking board and multiple planar marker AprilTags. The multiple planar marker AprilTags are printed at equal intervals on the four edges of the surface of the original tracking board. The side lengths of the multiple planar marker AprilTags are the same. The side length of each planar marker AprilTag depends on the side length of the original tracking board. The side length of the original tracking board is determined by the number and external dimensions of the parts placed in the original tracking board. The original tracking board is square. The multiple planar marker AprilTags all belong to the same family, but each planar marker AprilTag is unique, and the multiple planar marker AprilTags are numbered one by one.
[0056] Specifically, this example uses an Intel Realsense D455i depth camera to capture RGB-D images. Its intrinsic parameters have been calibrated before leaving the factory, and the intrinsic parameters are obtained using the pyrealsense2 toolkit in the Python environment. The current step uses the highest resolution of 1280*800 pixels to capture RGB images. At this time, the camera's intrinsic parameter matrix is:
[0057]
[0058] This example constructs a 6D pose dataset for 8 parts. SolidWorks is used to construct the 3D model of the parts. Figure 2 The 3D models of eight parts were displayed, including a fixture, a bearing end cap, a bearing seat, a universal positioning tool, a frame, a rod locator, a planetary gear carrier, and a gear shaft. Parts are numbered 1 through 8. These are common components found in industrial scenarios and exhibit common features of common industrial parts, such as metal and 3D printing materials, simple and complex structures, symmetrical and asymmetrical structures, and all have a lightly textured surface.
[0059] In this example, we imported 3D models of eight components into Keyshot software and arranged them in a virtual environment. Based on this, we estimated the required field size and determined the tracking pad's side length to be 0.5m. To enrich the scene information, we added a pure black square background area with a side length of 0.4m in the center of the tracking pad.
[0060] According to the side length determined in S12, 24 Tag36h11 family markers with a side length of 48 mm and numbered from 0 to 23 are used as the AprilTag used in this embodiment. The final designed AprilTag posture tracking board is as follows Figure 3 shown.
[0061] Place the AprilTag pose tracking board on the robotic arm console, secure the camera with a tripod, and adjust the camera to 1280*800 resolution RGB image capture mode to capture an RGB image.
[0062] S2: Extract and obtain the pose information of each plane marker AprilTag in the initial RGB detection image, and use the tracking board pose solution algorithm to process the pose information of each plane marker AprilTag to obtain the initial pose cMb of the pose tracking board 0 ;
[0063] S2 is specifically:
[0064] S21: Use the AprilTag detection toolkit to extract the pose information of each plane marker AprilTag in the initial RGB detection image and record it as the camera coordinate system feature of the pose tracking board. The pose information of each plane marker AprilTag includes the number of each plane marker AprilTag, its own corner point sequence number, corner point pixel coordinates [u corn l , v corn l](l=1, 2, 3, 4, the pixel coordinates of each AprilTag on the detection image have four corner points) and the 6D pose matrix (the pose of the coordinate system of the center of each AprilTag relative to the camera coordinate system);
[0065] The AprilTag detection toolkit includes AprilTag 3, AprilTags C++, and VISP.
[0066] Because the industrial scenes we created are extremely cluttered, some AprilTag corners may be partially obscured in the captured images. When the AprilTag detection toolkit estimates these AprilTag information, there's a chance it will obtain incorrect corner number and coordinates, severely impacting the detected AprilTag pose and, in turn, the estimated pose of the tracking board. Therefore, we use S2's least-squares plane fitting to eliminate these outliers.
[0067] S22: Extract the coordinates of the AprilTag center and its unit vector in the z-axis direction from the 6D pose matrix of each plane marker AprilTag The center point coordinates with normal vectors are formed (perpendicular to the plane AprilTag and outward), and the plane marker AprilTag point cloud is generated by the center point coordinates with normal vectors of all plane marker AprilTags;
[0068] Specifically, if Figure 4 As shown in (a), the vpDetectorAprilTag::detect() function in the VISP toolkit in the C++ environment is used to detect the AprilTag information in the image. This information includes the number corresponding to each AprilTag, the sequence number of the detected corner point (VISP numbers the four corner points from 1 to 4 according to the prescribed spatial order), the pixel coordinates of the corner point, and the homogeneous transformation matrix expressing the 6D pose of the AprilTag. Wherein, k refers to the number of AprilTag detected in the current image. Get the coordinates of each AprilTag in the camera coordinate system [x c k ,y c k , z c k ] and the unit normal vector in the z direction Then generate the AprilTag point cloud file in .ply format.
[0069] S23: Based on the AprilTag point cloud of the planar marker, the plane fitting method based on principal component analysis and least squares is used to obtain the spatial fitting plane where the pose tracking board is located in the camera coordinate system.
[0070] Space fitting plane The equation is:
[0071] Ax+By+Cz+D=0
[0072] Among them, A, B, C, and D are the first to fourth parameters of the fitting plane, respectively.
[0073] S24: Fitting the plane in the space where the tracking board is located according to the camera coordinate system Calculate the center of each plane marker AprilTag and the spatial fitting plane where the pose tracking board is located in the camera coordinate system The distance d between them is calculated, and the plane markers AprilTag whose distance d exceeds the preset distance threshold are removed to obtain the AprilTag set Γ1 of each plane marker after removing the abnormal distance;
[0074] The distance d is calculated as follows:
[0075]
[0076] S25: Calculate the spatial fitting plane The unit normal vector Then calculate the z-axis unit vector of each plane marker AprilTag in the camera coordinate system in the AprilTag set Γ1 after removing the abnormal distance Fitting a plane to space The unit normal vector The angle θ between them is removed, and the AprilTag of the plane markers whose angle θ exceeds the preset angle threshold is removed to obtain the AprilTag set Γ2 of each plane marker after removing the abnormal angle, and then the camera coordinate system features of the pose tracking board are updated according to the AprilTag set Γ2 of each plane marker after removing the abnormal angle;
[0077] Specifically, this step is used to remove Figure 5 (a) The two circles in the AprilTag are abnormal information. These AprilTags are incorrectly estimated by the detection toolkit for the corner point sequence number or corner point coordinates, which leads to completely abnormal AprilTag 6D pose estimation. If these abnormal information are not removed, the final part pose annotation will have huge errors. Figure 5(b) shows the calibration pose projection without any AprilTag abnormal information removal and optimization operations. In the figure, there is a large deviation between the projection of the part 3D model and the imaging of the real part.
[0078] In this embodiment, in order to eliminate the influence of these extreme abnormal information on the subsequent tracking board posture solution, the AprilTag information with an angle θ>30° is eliminated.
[0079] S26: Establish the object coordinate system of the posture tracking board. In the specific implementation, the center of the posture tracking board is taken as the origin, the vertical direction outward from the tracking board is the Z axis, and the horizontal direction to the right is the X axis, and a right-handed coordinate system is established. According to the object coordinate system of the posture tracking board, the number of the AprilTag of each plane marker in the posture tracking board, the sequence number of the corner point, and the coordinates of all corner points in the object coordinate system [x corn l,k ,y corn l,k , z corn l,k ], l = 1, 2, 3, 4, k refers to the number of the AprilTag detected in the current image, and this information is recorded as the object coordinate system feature of the pose tracking board;
[0080] Specifically, the tracking board object coordinate system established in this embodiment is as follows: Figure 4 As shown in the center point coordinate system in (a), the coordinates of all AprilTag corner points in the tracking board object coordinate system are obtained based on this coordinate system, for example Figure 4 In (b), the coordinates of the four corner points of AprilTag ID: 0 are (unit: mm):
[0081] [-23.4, 23.4, 0]
[0082] [-23.4, 18.6, 0]
[0083] [-18.6, 18.6, 0]
[0084] [-18.6, 23.4, 0]
[0085] S27: According to the updated camera coordinate system features and object coordinate system features of the pose tracking board, a 2D-3D feature relationship is established for all corner points of each plane marker AprilTag, and the corresponding 2D-3D feature point pairs are obtained, that is, the pixel coordinates of the corner points [u corn l,k , v corn l,k ] and the corner coordinates [x corn l,k ,y corn l,k , zcorn l,k ] One-to-one correspondence, all 2D-3D feature point pairs constitute the pose feature set;
[0086] S28: Based on the pose feature set, the RANSAC-PnP relative pose solution algorithm is used to calculate and obtain the pose of the pose tracking board relative to the camera coordinate system, which is recorded as the initial pose cMb of the pose tracking board. 0 .
[0087] Specifically, the cv::solvePnPRansac() function in OpenCV is used to implement the RANSAC-PnP algorithm, and its input parameters are as follows:
[0088] 3D corner coordinates [x corn l,k ,y corn l,k , z corn l,k ]array;
[0089] 2D pixel coordinate u corn l,k , v corn l,k array;
[0090] Camera intrinsic parameter matrix;
[0091] Camera distortion coefficient matrix. In this embodiment, the Realsense depth camera is a distortion-free camera.
[0092] The number of iterations is set to 50 in this embodiment;
[0093] The RANSAC threshold is set to 5 in this embodiment;
[0094] The confidence level is set to 0.99 in this embodiment.
[0095] right Figure 5 After the above optimization process is carried out for the scene shown in (a), the final part positioning posture projection diagram is as follows Figure 5 As shown in (c), the projection of the three-dimensional model of the part completely coincides with the imaging of the part in the image.
[0096] S3: Initial pose cMb of the tracking board according to the pose 0 , using the virtual-real interactive registration method based on prior pose specification and augmented reality to place multiple industrial general parts on the pose tracking board and obtain the corresponding placement poses bMo of multiple industrial general parts i ;
[0097] S3 specifically:
[0098] S31: Generate 3D models of multiple common industrial parts and their corresponding 3D model point clouds, and specify the prior placement poses bMo corresponding to the 3D models of multiple common industrial parts i p , where i represents the ordinal number of the part among all parts, and p represents the a priori specified one.
[0099] S32: bMo based on the prior placement of multiple common industrial parts i p and the initial pose cMb of the pose tracking board 0 Calculate the prior pose cMo of multiple common industrial parts in the depth camera coordinate system of the initial RGB detection image in S1 i p The calculation formula is as follows:
[0100] cMo i p =cMb 0 ×bMo i p
[0101] S33: Based on the prior pose cMo of multiple common industrial parts in the camera coordinate system i p , transform the corresponding 3D model point cloud to obtain the 3D model point cloud P in the camera coordinate system i vir ;
[0102] The specific formula is as follows:
[0103]
[0104] Among them, p i,j vir Represents the point cloud P of the 3D model of part No. i i vir The jth discrete point in Represents the 3D model point cloud P i vir Each discrete point p in i,j vir The components of the X, Y, and Z axes in the camera coordinate system, It is the coordinate of the 3D model point cloud of each part in its own object coordinate system, and T represents transposition.
[0105] S34: Use the existing augmented reality method to transform the 3D models of multiple common industrial parts in S31 into the corresponding prior poses cMo in the camera coordinate system. i pRendering: After obtaining the corresponding rendered image, it is presented in real time to the high-resolution virtual-reality interactive interface of the pose tracking board. Multiple common industrial parts are placed on the pose tracking board and the poses of the corresponding parts are adjusted according to the rendered images of the parts, thus achieving coarse virtual-reality registration of the placement poses of multiple common industrial parts.
[0106] Specifically, this embodiment uses the toolkit VisPy in the Python environment to achieve part rendering. Figure 6 (a) and (b) show the manual coarse registration process and coarse registration results of part No. 3, respectively. When the projection of the actual part in the virtual-reality interaction interface coincides with the projection of the 3D model rendered by VisPy, the placement of part No. 3 is completed.
[0107] S35: Use the depth camera to obtain the depth point cloud of the general industrial parts placed on the pose tracking board at the current viewing angle. After the depth point cloud is converted to coordinates, the conversion formula is as follows to obtain the three-dimensional actual point cloud P corresponding to multiple general industrial parts in the camera coordinate system. i real ;
[0108]
[0109] Among them, p i,q real Represents the actual three-dimensional point cloud P of part No. i i real The qth discrete point in Represents the three-dimensional actual point cloud P i real Each point p i,q real The components of the X, Y, and Z axes in the camera coordinate system, For each point p i,q real The corresponding homogeneous pixel coordinates in the image plane are, is the p obtained from the part depth point cloud i,q real The depth value of the point, K -1 is the inverse of the camera intrinsic parameter matrix.
[0110] Specifically, a depth image of the current viewing angle is captured by a depth camera, and the part projection rendered by VisPy is used to obtain the ROI area of the current part on the depth image; the depth point cloud of the part in the ROI area is converted into the three-dimensional actual point cloud P in the camera coordinate system through the above formula i real .
[0111] S36: Using the iterative closest point ICP algorithm to calculate the 3D model point cloud P of multiple common industrial parts in the camera coordinate system ivir And the three-dimensional actual point cloud P in the camera coordinate system i real Perform virtual-real precise registration and obtain P i real With P i vir The relative pose transformation relationship between them, that is, the relative pose between the real parts and the virtual projection corresponding to multiple industrial general parts i p Mo i , and then calculate the placement pose bMo corresponding to multiple industrial general parts after precise registration i , recorded as the placement pose bMo corresponding to multiple industrial general parts i .
[0112] Placement pose after precise alignment bMo i The calculation formula is as follows:
[0113] bMo i =bMo i p ×o i p Mo i
[0114] Specifically, the open3d toolkit in Python environment is used to achieve precise registration of virtual and real interaction. i real With P i vir As input, through the registration_icp function under open3d i p Mo i In this embodiment, the threshold of the ICP algorithm is set to 0.5 mm, and the maximum number of iterations is 35.
[0115] S4: Use the depth camera in S1 in RGB-D video recording mode to record RGB-D videos from multiple angles that include all common industrial parts and the full view of the pose tracking board;
[0116] Specifically, during the video recording phase, the Intel Realsense D455i depth camera was removed from the fixed bracket and the RGB-D video stream was recorded using handheld recording. In this phase, the camera resolution was adjusted to 640*480 pixels, and the camera internal parameters became:
[0117]
[0118] In this embodiment, 1174 frames of RGB-D images are recorded for the scene, and each image contains 8 parts and the full view of the tracking board.
[0119] S5: Repeat S2 to solve and obtain the tracking board pose cMb of all RGB-D images in the RGB-D video. J , J represents the frame number of the RGB-D image, and then according to the placement posture bMo corresponding to multiple industrial general parts i , calculate and obtain the calibration pose cMo corresponding to multiple industrial general parts of all frames of RGB-D images i J ;
[0120] The calculation formula is as follows:
[0121] cMo i J =cMb J ×bMo i
[0122] S6: According to the three-dimensional models of multiple common industrial parts and the corresponding calibration poses cMo of multiple common industrial parts in all frames of RGB-D images, obtain the numbers, minimum bounding boxes and segmentation masks of multiple common industrial parts in all frames of RGB-D images and form a 6D pose dataset of the current scene;
[0123] Specifically, this embodiment uses the toolkit BOP Toolkit in the Python environment to generate the part number, 2D minimum bounding box, and segmentation mask in each image.
[0124] The camera's intrinsic parameters and the corresponding calibration pose cMo for each part in each image are used to construct a .json file. This file, along with the part's 3D model, is then used as input to the BOP Toolkit. The BOP Toolkit then automatically generates a .json file containing the minimum bounding box of each part in each image, along with a segmentation mask for each part. Figure 7 (a) and (b) show a RGB-D image pair in the scene and the calibration information obtained according to the process of this invention. Figure 7 (c) shows a rendering diagram of the part's 6D calibration pose, including the projection rendering of each part, the part object coordinate system, and the 3D minimum bounding box. Figure 7 (d) shows an image containing 8 part segmentation masks. The masks of the 8 parts are distinguished by different pixel values, and the image also annotates the two-dimensional minimum bounding box of each part.
[0125] Place the pose tracking board in different scenes, change the color and texture of the AprilTag pose tracking board background, repeat S1-S6 to obtain 6D pose datasets for different industrial scenes, and the final 6D pose dataset is composed of the 6D pose datasets of all scenes.
[0126] Based on the 6D pose dataset, pose estimation / tracking is achieved using a deep learning-based pose estimation / tracking algorithm. The 6D pose dataset can also be used to implement robotic automated grasping and virtual assembly.
[0127] Specifically, in addition to the above-mentioned robot operation platform scenario, this embodiment also constructs part 6D pose datasets in two other different scenarios, such as Figure 8 As shown in (a) and (b).
[0128] The present invention is not limited to the above-described embodiments and is intended only to facilitate understanding of the methods and core concepts of the present invention. It should be noted that those skilled in the art may, without departing from the principles of the present invention, make various improvements and modifications to the present invention, and such improvements and modifications fall within the scope of protection of the claims of the present invention. Any material not described in detail in this specification is prior art known to those skilled in the art.
Claims
1. A method for constructing a 6D pose dataset of common industrial parts, characterized in that: The following steps are involved: S1: Fix the depth camera and pose tracking board in a specified industrial scene, use the depth camera to shoot the pose tracking board in the current scene, and obtain the initial RGB detection image; S2: Extract and obtain the pose information of each plane marker AprilTag in the initial RGB detection image, and use the tracking board pose solution algorithm to process the pose information of each plane marker AprilTag to obtain the initial pose cMb of the pose tracking board 0 ; S3: Initial pose cMb of the tracking board according to the pose 0 , using the virtual-real interactive registration method based on prior pose specification and augmented reality to place multiple industrial general parts on the pose tracking board and obtain the corresponding placement poses bMo of multiple industrial general parts i ; The S3 is specifically: S31: Generate 3D models of multiple common industrial parts and their corresponding 3D model point clouds, and specify the prior placement poses bMo corresponding to the 3D models of multiple common industrial parts i p ; S32: bMo based on the prior placement of multiple common industrial parts i p and the initial pose cMb of the pose tracking board 0 Calculate the prior pose cMo of multiple common industrial parts in the camera coordinate system i p ; S33: Based on the prior pose cMo of multiple common industrial parts in the camera coordinate system i p , transform the corresponding 3D model point cloud to obtain the 3D model point cloud P in the camera coordinate system i vir ; S34: Use augmented reality methods to transform the three-dimensional models of multiple common industrial parts in S31 into the corresponding prior poses cMo in the camera coordinate system i p Rendering: After obtaining the corresponding rendered image, it is presented in real time to the virtual-reality interactive interface of the pose tracking board. Multiple common industrial parts are placed on the pose tracking board and the poses of the corresponding parts are adjusted according to the rendered images of the parts, thus achieving coarse virtual-reality registration of the placement poses of multiple common industrial parts. S35: Use the depth camera to obtain the depth point cloud of the general industrial parts placed on the pose tracking board, and after the coordinate conversion of the depth point cloud, obtain the three-dimensional actual point cloud P corresponding to multiple general industrial parts in the camera coordinate system. i reali ; S36: Using the iterative closest point ICP algorithm to calculate the 3D model point cloud P of multiple common industrial parts in the camera coordinate system i viri And the three-dimensional actual point cloud P in the camera coordinate system i reali Perform virtual-real precise registration to obtain the real parts and prior placement poses bMo corresponding to multiple common industrial parts i p The relative pose between the three-dimensional models under i p Mo i , and then calculate the placement pose bMo corresponding to multiple industrial general parts after precise registration i , recorded as the placement pose bMo corresponding to multiple industrial general parts i ; S4: Use the depth camera in S1 in RGB-D video recording mode to record RGB-D videos from multiple angles that include all common industrial parts and the full view of the pose tracking board; S5: Repeat S2 to solve and obtain the tracking board pose cMb of all RGB-D images in the RGB-D video. J , and then according to the corresponding placement postures of multiple industrial general parts bMo i , calculate and obtain the calibration pose cMo corresponding to multiple industrial general parts of all frames of RGB-D images i J ; S6: Based on the three-dimensional models of multiple common industrial parts and the corresponding calibration poses cMo of multiple common industrial parts in all frame RGB-D images i J , obtain the numbers, minimum bounding boxes and segmentation masks of multiple common industrial parts in all frame RGB-D images and form a 6D pose dataset of the current scene.
2. The method for constructing a 6D pose dataset of a general industrial part according to claim 1, characterized in that: In S1, the posture tracking board includes an original tracking board and a plurality of planar markers AprilTags, which are fixedly mounted on the four edges of the surface of the original tracking board at equal intervals, and all of the plurality of planar markers AprilTags belong to the same family.
3. The method for constructing a 6D pose dataset of a general industrial part according to claim 1, characterized in that: The S2 is specifically: S21: Use the AprilTag detection toolkit to extract the pose information of each plane marker AprilTag in the initial RGB detection image and record it as the camera coordinate system feature of the pose tracking board. The pose information of each plane marker AprilTag includes the number of each plane marker AprilTag, its own corner point sequence number, corner point pixel coordinates [u corn l , v corn l ] and 6D pose matrix; S22: Extract the coordinates of the AprilTag center and its unit vector in the z-axis direction from the 6D pose matrix of each plane marker AprilTag The center point coordinates with normal vectors are composed, and the plane marker AprilTag point cloud is generated by the center point coordinates with normal vectors of all plane marker AprilTags; S23: Based on the AprilTag point cloud of the planar marker, the plane fitting method based on principal component analysis and least squares is used to obtain the spatial fitting plane where the pose tracking board is located in the camera coordinate system. S24: Fitting the plane in the space where the pose tracking board is located in the camera coordinate system Calculate the center of each plane marker AprilTag and the spatial fitting plane where the pose tracking board is located in the camera coordinate system The distance d between them is calculated, and the plane markers AprilTag whose distance d exceeds the preset distance threshold are removed to obtain the AprilTag set Γ1 of each plane marker after removing the abnormal distance; S25: Calculate the spatial fitting plane The unit normal vector Then calculate the z-axis unit vector of each plane marker AprilTag in the camera coordinate system in the AprilTag set Γ1 after removing the abnormal distance Fitting a plane to space The unit normal vector The angle θ between them is removed, and the AprilTag of the plane markers whose angle θ exceeds the preset angle threshold is removed to obtain the AprilTag set Γ2 of each plane marker after removing the abnormal angle, and then the camera coordinate system features of the pose tracking board are updated according to the AprilTag set Γ2 of each plane marker after removing the abnormal angle; S26: Establishing the object coordinate system of the pose tracking board, determining the number of each plane marker AprilTag in the pose tracking board, the corner point sequence number, and the coordinates of all corner points in the object coordinate system according to the object coordinate system of the pose tracking board, and recording them as the object coordinate system features of the pose tracking board; S27: Based on the updated camera coordinate system features and object coordinate system features of the pose tracking board, establish 2D-3D feature relationships for all corner points of each plane marker AprilTag, obtain corresponding 2D-3D feature point pairs, and form a pose feature set from all 2D-3D feature point pairs; S28: Based on the pose feature set, the RANSAC-PnP relative pose solution algorithm is used to calculate and obtain the pose of the pose tracking board relative to the camera coordinate system, which is recorded as the initial pose cMb of the pose tracking board. 0 .
4. The method for constructing a 6D pose dataset of industrial general parts according to claim 3, characterized in that: The AprilTag detection toolkit includes AprilTag 3, AprilTags C++, and VISP.
5. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Low-texture industrial part pose estimation method based on deep learning
CN110910452A
6D pose labeling method and system and storage medium
CN113034593A