A 3D Visual Pose Estimation Method for Unordered Workpieces Based on Deep Learning
Through a deep learning-based method, combined with image instance segmentation, stacking estimation and pose estimation calculation method, the problem of missing information in three-dimensional visual pose estimation of disordered workpieces is solved, and the accurate pose estimation and grasping of workpieces is realized, and the success rate of robot operations is improved.
Patent Information
- Application Number
- CN202111373613.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-11-19
AI Technical Summary
The prior art lacks information in the three-dimensional visual position estimation of disordered workpieces, which affects the accuracy of the estimation, making it difficult to accurately grasp and load and unload multiple specifications of disordered workpieces.
Using a deep learning-based method, by collecting color images and depth information of disordered workpieces, combining image instance segmentation algorithm, stacking estimation algorithm and pose estimation algorithm, instance segmentation of workpiece point cloud, stacking relationship estimation and precise pose estimation.
It effectively improves the accuracy and speed of estimation of poses of disordered workpieces, reduces the difficulty of estimation of workpieces, improves the success rate of robot grasping, and is suitable for accurate grasping and loading and unloading of multiple specifications of disordered workpieces.
Smart Images

Figure CN114140526B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of workpiece pose estimation, and specifically relates to a three-dimensional visual pose estimation method for disordered workpieces based on deep learning, which is a deep learning method for estimating the position and pose of disordered workpieces on a production line. Background Art
[0002] The precise grasping and loading and unloading of disordered workpieces has always been one of the key research topics in the field of intelligent industrial robots. This link is generally equipped with visual sensors, which estimate the position and posture of the workpiece by identifying the visual information collected by the visual sensor, and then realize grasping. According to the relative position relationship of the workpieces, the placement of disordered workpieces can be divided into disordered discrete and disordered stacking. Disordered discrete means that the workpieces are placed on a horizontal plane without contact or stacking with each other; disordered stacking means that the workpieces are randomly and disorderly placed, and there is overlap or contact between the workpieces. The precise grasping of disordered workpieces of multiple specifications requires the system to be able to identify the disordered workpieces within the field of view, judge the stacking relationship between the workpieces, estimate the position and posture of the workpieces that are non-overlapping and easy to grasp, and plan the robot's movement path.
[0003] In recent years, with the continuous improvement of computer performance and the rapid development of visual sensors and related algorithms, the workpiece pose estimation technology based on two-dimensional vision has become mature and has been widely used in various automatic loading and unloading systems. However, using only two-dimensional images to represent three-dimensional workpieces will inevitably cause information loss, thereby affecting the accuracy of the pose estimation of disordered workpieces. Therefore, the mixed-line production of multi-specification products inevitably requires the study of three-dimensional visual pose estimation technology for disordered workpieces, and further realizes the precise grasping and loading and unloading of disordered workpieces. It can be said that it is crucial to improve the system's environmental perception ability, study the intelligent recognition and pose estimation technology of disordered workpieces, and develop intelligent industrial robot systems suitable for the precise grasping and loading and unloading of multi-specification disordered workpieces. Summary of the invention
[0004] In order to solve the problems existing in the prior art, the present invention provides a method for estimating the three-dimensional visual pose of disordered workpieces based on deep learning, which is used to estimate the position and pose of disordered workpieces on a production line relative to the robot base coordinate system. When the estimation method of the present invention is used in conjunction with an industrial robot, the loading and unloading of disordered workpieces can be realized.
[0005] A method for estimating three-dimensional visual pose of disordered workpieces based on deep learning comprises the following steps:
[0006] (1) Collect color images and depth information of disordered artifacts;
[0007] (2) Use the constructed image instance segmentation algorithm to process the color image and obtain target detection information and instance segmentation information;
[0008] (3) Using the target detection information, crop the color image to obtain the detection image of each workpiece;
[0009] (4) inputting the inspection image into the constructed stack estimation algorithm to obtain the stack estimation information of all workpieces;
[0010] (5) selecting a workpiece with the lowest stacking degree from the instance segmentation information according to the stacking estimation information to form a mask image of the workpiece;
[0011] (6) Segmenting the workpiece point cloud from the depth information based on the mask image of the workpiece;
[0012] (7) The workpiece point cloud is input into the constructed pose estimation algorithm to estimate the pose information of the grasping part of the workpiece relative to the robot base coordinate system.
[0013] In the above step (1), a three-dimensional vision sensor is used to collect color images and depth information (three-dimensional point cloud) of disordered workpieces within the field of view.
[0014] In step (2), the target detection information is the bounding box of each artifact in the color image, and the instance segmentation information is the pixel set of each artifact in the color image.
[0015] Preferably, the image instance segmentation algorithm consists of a deep convolutional network, a feature pyramid network, a result prediction network and a post-processing module;
[0016] The deep convolutional network extracts high-dimensional feature vectors from color images, and is composed of five sets of convolutional layers + pooling layers in series, each set of composite structures generates a set of feature vectors, namely feature vector 1, feature vector 2, feature vector 3, feature vector 4 and feature vector 5;
[0017] The feature pyramid network combines convolution operation and upsampling operation to process feature vectors generated by the deep convolution network, wherein feature vector 5 generates feature vector 6 after convolution operation, feature vector 4 is added to feature vector 6 after upsampling operation to form feature vector 7 after convolution operation, feature vector 3 is added to feature vector 7 after upsampling operation to form feature vector 8 after convolution operation, feature vector 2 is added to feature vector 8 after upsampling operation to form feature vector 9 after convolution operation, and feature vector 6, feature vector 7, feature vector 8 and feature vector 9 are sequentially generated into feature vector 10, feature vector 11, feature vector 12 and feature vector 13 after convolution operation;
[0018] The result prediction network consists of two network branches, which share weights for feature vectors 10, 11, 12, and 13. The first network branch consists of multiple deep convolutional layers and multiple fully connected layers in series, which regresses and predicts the bounding box of the artifact in the color image to form preliminary target detection information; the second network branch consists of multiple deep convolutional layers in series, which predicts the probability (with a value of 0 to 1.0) that each pixel in the color image belongs to a specific artifact, forming preliminary instance segmentation information.
[0019] The post-processing module consists of a non-maximum suppression unit and a threshold filtering unit. The non-maximum suppression unit processes the preliminary target detection information, eliminates redundant workpiece bounding boxes, and forms target detection information. The threshold filtering unit uses a threshold of 0.5 to filter the preliminary instance segmentation information to form instance segmentation information.
[0020] Preferably, the stacked estimation algorithm is composed of a plurality of deep convolutional layers and a plurality of fully connected layers connected in series.
[0021] Preferably, the stacking estimation information is a one-dimensional matrix, the number of elements in the matrix is equal to the number of workpiece detection images, each element represents the probability of a workpiece being stacked (with a value of 0 to 1.0), and the larger the probability value of stacking (the closer to 1), the lower the degree of stacking of the corresponding workpiece.
[0022] Preferably, the pose estimation algorithm comprises:
[0023] A data pre-processing module, which performs statistical filtering and grid down-sampling pre-processing on the workpiece point cloud;
[0024] A point cloud classification unit, which classifies the preprocessed workpiece point cloud according to the type and placement posture of the workpiece, and outputs a point cloud category;
[0025] A point cloud fusion unit, wherein the point cloud fusion unit fuses the preprocessed workpiece point cloud with the point cloud category to form a point cloud vector;
[0026] A posture estimation unit estimates the posture information of the workpiece grasping part relative to the robot base coordinate system according to the point cloud-like vector.
[0027] As a further preference, the point cloud classification unit comprises:
[0028] A sampling module, wherein the sampling module randomly samples a fixed number of point clouds from the preprocessed workpiece point cloud;
[0029] Normalization module, which maps the three-dimensional coordinate value of each point in the point cloud obtained by the sampling module to [-a 1 ,b 1]; where a 1 ∈[0.5~1.5], b 1 ∈[0.5~1.5];
[0030] The point cloud classification network is composed of a multi-layer perceptron, a maximum pooling layer and a fully connected layer with shared weights connected in series, and predicts the point cloud category according to the floating point number output by the normalization module.
[0031] As a further preferred method, the specific method of forming a point cloud-like vector is:
[0032] First, the point cloud categories are converted into one-hot encodings, and then combined with the coordinate values in the processed workpiece point cloud in sequence.
[0033] As a further preference, the pose estimation unit includes a position estimation unit and a pose estimation unit, and the pose information includes position information (x, y, z) and pose information (rx, ry, rz).
[0034] As further preferred, the position estimation unit includes:
[0035] A sampling module, wherein the sampling module samples the point cloud-like vector and forms a vector of fixed dimension;
[0036] A normalization module calculates the mean of each dimension of the vector collected by the sampling module and maps each value in the vector to [-a 2 ,b 2 ]; where a 2 ∈[0.5~1.5], b 2 ∈[0.5~1.5];
[0037] The position estimation network obtains the position information (x, y, z) of the workpiece grasping part relative to the robot base coordinate system according to the mean value of the vector calculated by the normalization module and the output floating point number.
[0038] Preferably, the position estimation network consists of two network branches, one of which is composed of a multi-layer perceptron, a maximum pooling layer and a fully connected layer with shared weights connected in series, and forms a first position estimation component (x1, y1, z1) according to the floating-point number output by the normalization module; the other network branch is a fully connected layer, and forms a second position estimation component (x2, y2, z2) according to the vector mean calculated by the normalization module; the first position estimation component is added to the second position estimation component to obtain the position information (x, y, z) of the workpiece grasping part relative to the robot base coordinate system.
[0039] As a further preferred embodiment, the posture estimation unit comprises:
[0040] A sampling module, wherein the sampling module samples the point cloud-like vector and forms a vector of fixed dimension;
[0041] Normalization module, the normalization module processes the vector obtained by the sampling module and maps each value in the vector to [-a 3 ,b 3 ]; where a 3 ∈[0.5~1.5], b 3 ∈[0.5~1.5];
[0042] The posture estimation network obtains the posture information (rx, ry, rz) of the workpiece grasping part relative to the robot base coordinate system according to the floating point number output by the normalization module.
[0043] Preferably, the posture estimation network consists of two network branches, both of which are composed of a multi-layer perceptron with shared weights, a maximum pooling layer and a fully connected layer connected in series;
[0044] One of the network branches estimates the absolute values of the rotation angles around the X-axis, the Y-axis, and the Z-axis of the workpiece grasping part relative to the robot base coordinate system based on the floating-point numbers output by the normalization module; the other network branch estimates the direction of rotation around the Z-axis of the workpiece grasping part relative to the robot base coordinate system based on the floating-point numbers output by the normalization module; the output results of the two network branches of the posture estimation network are combined to form the posture information (rx, ry, rz) of the workpiece grasping part relative to the robot base coordinate system.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] 1. The three-dimensional visual pose estimation method of disordered workpieces based on deep learning of the present invention adopts the idea of deep learning to realize point cloud instance segmentation, stacking estimation and pose estimation of disordered workpieces, and is suitable for the positioning and loading and unloading of disordered workpieces on industrial assembly lines.
[0047] 2. The pose estimation method of the present invention combines the 3D reconstruction process of the 3D vision sensor with the image instance segmentation algorithm to realize the instance segmentation of the workpiece point cloud, greatly reducing the difficulty of point cloud instance segmentation and effectively improving the speed and accuracy of point cloud instance segmentation.
[0048] 3. The pose estimation method of the present invention proposes for the first time to use a deep learning algorithm to estimate the stacking relationship of workpieces to determine the grasping priority of disordered workpieces. In the pose estimation process, only the pose of non-stacked and easy-to-grasp workpieces needs to be estimated, which effectively reduces the difficulty of workpiece pose estimation and improves the success rate of robot grasping.
[0049] 4. The pose estimation method of the present invention can be widely used in the actual production of the automotive industry, electrical and electronic industry, metal machinery industry and other industries, has broad market application prospects, and has extremely important practical significance for improving the digitalization and intelligence level of my country's manufacturing industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a schematic diagram of a flow chart of an embodiment of the present invention;
[0051] Figure 2 Schematic diagram of the process of image instance segmentation algorithm in an embodiment of the present invention;
[0052] Figure 3 Schematic diagram of the flow of the pose estimation algorithm in an embodiment of the present invention;
[0053] Figure 4 Schematic diagram of the process of the point cloud classification unit in an embodiment of the present invention;
[0054] Figure 5 is a schematic diagram of a flow chart of a position estimation unit in an embodiment of the present invention;
[0055] Figure 6 Schematic diagram of the process of the posture estimation unit in the embodiment of the present invention. DETAILED DESCRIPTION
[0056] The technical solution of the present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0057] like Figure 1 As shown, a method for estimating the three-dimensional visual pose of an unordered workpiece based on deep learning comprises the following steps:
[0058] ①Set the 3D vision sensor just above the workpiece grasping area to collect visual information of disordered workpieces within the field of view and obtain color images and depth information;
[0059] ② Input the color image into the constructed image instance segmentation algorithm to obtain target detection information and instance segmentation information. The target detection information is the bounding box of each artifact in the color image, and the instance segmentation information is the pixel set of each artifact in the color image;
[0060] ③ Use the target detection information to crop the color image to form a detection image of each workpiece with the same number of workpieces;
[0061] ④ Input the inspection image of the workpiece into the constructed stacking estimation algorithm to obtain the stacking estimation information. The stacking estimation information is a one-dimensional matrix. The number of elements in the matrix is equal to the number of workpiece inspection images. Each element represents the probability of a workpiece being stacked. The larger the probability value, the lower the degree of stacking of the corresponding workpiece.
[0062] ⑤ According to the stacking estimation information, the workpiece with the lowest stacking degree is selected from the instance segmentation information to form a mask image of the workpiece with the lowest stacking degree;
[0063] ⑥ Segment the workpiece point cloud from the depth information based on the mask image of the workpiece;
[0064] ⑦ Input the workpiece point cloud into the constructed pose estimation algorithm to estimate the pose information of the grasping part of the workpiece relative to the robot base coordinate system, where the pose information includes position information (x, y, z) and posture information (rx, ry, rz).
[0065] like Figure 2 As shown in the figure, the image instance segmentation algorithm consists of a deep convolutional network, a feature pyramid network, a result prediction network and a post-processing module; the deep convolutional network extracts high-dimensional feature vectors from color images, which is composed of five groups of convolutional layers + pooling layers in series. Each group of composite structures produces a group of feature vectors, which are feature vector 1, feature vector 2, feature vector 3, feature vector 4 and feature vector 5 respectively.
[0066] The feature pyramid network combines convolution operation and upsampling operation to process the feature vector generated by the deep convolution network, wherein feature vector 5 generates feature vector 6 after convolution operation, feature vector 4 is added to feature vector 6 after upsampling operation to form feature vector 7 after convolution operation, feature vector 3 is added to feature vector 7 after upsampling operation to form feature vector 8 after convolution operation, feature vector 2 is added to feature vector 8 after upsampling operation to form feature vector 9 after convolution operation, and feature vector 6, feature vector 7, feature vector 8 and feature vector 9 are sequentially generated into feature vector 10, feature vector 11, feature vector 12 and feature vector 13 after convolution operation;
[0067] The result prediction network consists of two network branches, which share weights for feature vectors 10, 11, 12, and 13. The first network branch is composed of several deep convolutional layers and fully connected layers in series, which regresses and predicts the bounding box of the artifact in the color image to form preliminary target detection information; the second network branch is composed of several deep convolutional layers in series, which predicts the probability (with a value of 0 to 1.0) that each pixel in the color image belongs to a specific artifact, forming preliminary instance segmentation information.
[0068] The post-processing module consists of a non-maximum suppression unit and a threshold filtering unit. The non-maximum suppression unit processes the preliminary target detection information, eliminates redundant workpiece bounding boxes, and forms target detection information. The threshold filtering unit uses a threshold of 0.5 to filter the preliminary instance segmentation information to form instance segmentation information.
[0069] The stacking estimation algorithm consists of multiple deep convolutional layers and multiple fully connected layers in series to predict the probability of the workpiece being stacked (with a value of 0 to 1.0); if the output probability value is closer to 1, the degree of stacking of the workpiece is lower.
[0070] like Figure 3 As shown in the figure, the pose estimation algorithm consists of a data pre-processing module, a point cloud classification unit, a point cloud fusion unit and a pose estimation unit; the data pre-processing unit performs statistical filtering and grid downsampling pre-processing operations on the workpiece point cloud, and outputs the processed workpiece point cloud; the point cloud classification unit receives the pre-processed workpiece point cloud, classifies it according to the type and placement posture of the workpiece, and outputs the point cloud category;
[0071] The point cloud fusion unit fuses the processed workpiece point cloud with the point cloud category to form a point cloud vector. The specific method is to first convert the point cloud category into a one-hot encoding, and then combine it with the coordinate values in the processed workpiece point cloud in sequence; the pose estimation unit is composed of a position estimation unit and a pose estimation unit, which estimates the pose information of the workpiece grasping part relative to the robot base coordinate system according to the point cloud vector.
[0072] like Figure 4 As shown in the figure, the point cloud classification unit consists of a sampling module, a normalization module and a point cloud classification network; the sampling module randomly samples a fixed number of point clouds from the processed workpiece point cloud; the normalization module maps the three-dimensional coordinate value of each point in the point cloud obtained by the sampling module to a floating point number between [-1.0, 1.0]; the point cloud classification network is composed of a multi-layer perceptron with shared weights, a maximum pooling layer and a fully connected layer connected in series, and predicts the point cloud category according to the floating point number output by the normalization module.
[0073] like Figure 5 As shown, the position estimation unit consists of a sampling module, a normalization module and a position estimation network; the sampling module samples the point cloud vector to form a vector of fixed dimension; the normalization module processes the vector obtained by the sampling module, calculates the mean of each dimension of the vector and maps each value in the vector to a floating point number between [-1.0, 1.0];
[0074] The position estimation network consists of two network branches. One network branch is composed of a multi-layer perceptron with shared weights, a maximum pooling layer and a fully connected layer in series. The first position estimation component (x1, y1, z1) is formed according to the normalized value (floating point number) output by the normalization module; the other network branch is a fully connected layer, which forms the second position estimation component (x2, y2, z2) according to the vector mean calculated by the normalization module; the first position estimation component is added to the second position estimation component to obtain the position information (x, y, z) of the workpiece grasping part relative to the robot base coordinate system.
[0075] like Figure 6As shown, the posture estimation unit consists of a sampling module, a normalization module and a posture estimation network; the sampling module samples the point cloud vector to form a vector of fixed dimension; the normalization module processes the vector obtained by the sampling module and maps each value in the vector to a floating point number between [-1.0, 1.0];
[0076] The posture estimation network consists of two network branches. One network branch is composed of a multi-layer perceptron with shared weights, a maximum pooling layer and a fully connected layer in series. The absolute value of the rotation angle around the X-axis, the rotation angle around the Y-axis and the rotation angle around the Z-axis of the workpiece grasping part relative to the robot base coordinate system is estimated according to the normalized value (floating point number) output by the normalization module; the other network branch is composed of a multi-layer perceptron with shared weights, a maximum pooling layer and a fully connected layer in series. The direction of rotation around the Z-axis of the workpiece grasping part relative to the robot base coordinate system is estimated according to the normalized value (floating point number) output by the normalization module; the outputs of the two network branches of the posture estimation network are integrated to form the posture information (rx, ry, rz) of the workpiece grasping part relative to the robot base coordinate system.
[0077] This embodiment is applicable to the positioning and loading and unloading of disordered workpieces on an industrial assembly line, and its specific implementation process includes a training phase and an implementation phase.
[0078] The training process of the training phase of this embodiment is as follows:
[0079] 1. Build a robot 3D visual grasping system, which consists of a robot, a 3D visual sensor, a workbench, a host computer and a gripper; the workbench is placed in the robot's workspace to place disordered workpieces to be grasped; the 3D visual sensor is installed directly above the workbench to collect visual information of disordered workpieces (including color images, depth information (3D point cloud)); the gripper is installed at the end of the robot to grasp disordered workpieces; various algorithms described in the present invention are set in the host computer, and interact with the 3D visual sensor and the robot;
[0080] 2. Construction of image instance segmentation algorithm: Select several workpieces and place them on the workbench in disorder, and use a 3D vision sensor to take several color images of the disordered workpieces; the number and placement of the workpieces need to be adjusted each time they are taken; outline the outer contour of each workpiece on the color image (including target detection information and instance segmentation information), use the color image as input, target detection information and instance segmentation information as output, and form a training data set for the image instance segmentation algorithm; use the training data set to train the image instance segmentation algorithm;
[0081] 3. Construction of stacking estimation algorithm: Input the color image collected in step 2 into the image instance segmentation algorithm to generate target detection information, and then crop the color image according to the target detection information to form multiple workpiece detection images; mark the stacking degree of the workpieces in each workpiece detection image, if the workpieces are stacked, it is marked as 0, if not stacked, it is marked as 1, with the workpiece detection image as input and the stacking degree of the workpiece as output, to form a training data set for the stacking estimation algorithm; use the training data set to train the workpiece stacking estimation algorithm;
[0082] 4. Construction of pose estimation algorithm: Select several workpieces and place them on the workbench in disorder, use a three-dimensional vision sensor to collect the three-dimensional point cloud of the workpiece, teach the robot to the grasping position (grasping part) of the workpiece and record the robot's pose, then pre-process the three-dimensional point cloud and extract the workpiece point cloud; repeat several times to form a training data set for the pose estimation algorithm, where each set of data contains a set of workpiece point clouds (as input) and its corresponding robot pose (as output); use the training data set to train the pose estimation algorithm.
[0083] The implementation process of the implementation stage of this embodiment is as follows:
[0084] 1. Select several workpieces and place them on the workbench in disorder, and use a 3D vision sensor to collect the color image and depth information of the workpieces;
[0085] 2. Input the color image and depth information into the three-dimensional visual pose estimation method of disordered workpieces based on deep learning of this embodiment to obtain the pose information (including position information and posture information) of a grasping part of the workpiece relative to the robot base coordinate system;
[0086] 3. The robot grasps the workpiece from the disordered workpieces based on the position and posture information of the workpiece estimated by this embodiment;
[0087] 4. Repeat steps 1 to 3 to complete the grabbing of all disordered workpieces.
Claims
1. A three-dimensional vision pose estimation method for disordered workpieces based on deep learning, characterized in that, it includes the following steps: (1) Collect the color image and depth information of the disordered workpiece; (2) Use the constructed image instance segmentation algorithm to process the color image to obtain target detection information and instance segmentation information; (3) Crop the color image using the target detection information to obtain the detection image of each workpiece; (4) Input the detection image into the constructed stacking estimation algorithm to obtain the stacking estimation information of all workpieces; (5) According to the stacking estimation information, select the workpiece with the lowest stacking degree from the instance segmentation information to form the mask image of the workpiece; (6) Segment the workpiece point cloud from the depth information according to the mask image of the workpiece; (7) Input the workpiece point cloud into the constructed pose estimation algorithm to estimate the pose information of the grasping part of the workpiece relative to the robot base coordinate system; The image instance segmentation algorithm is composed of a deep convolutional network, a feature pyramid network, a result prediction network and a post-processing module; The deep convolutional network extracts high-dimensional feature vectors from the color image and is composed of a series of five sets of convolutional layer + pooling layer composite structures. Each set of composite structures generates a set of feature vectors, namely feature vector 1, feature vector 2, feature vector 3, feature vector 4 and feature vector 5 in sequence; The feature pyramid network combines convolutional operations and upsampling operations to process the feature vectors generated by the deep convolutional network. Among them, feature vector 5 generates feature vector 6 after convolutional operation, and feature vector 4 generates feature vector 7 after convolutional operation and is added to the upsampled feature vector 6. Feature vector 3 generates feature vector 8 after convolutional operation and is added to the upsampled feature vector 7. Feature vector 2 generates feature vector 9 after convolutional operation and is added to the upsampled feature vector 8. Feature vectors 6, 7, 8 and 9 generate feature vectors 10, 11, 12 and 13 in sequence after convolutional operation; The result prediction network consists of two network branches, sharing weights for feature vectors 10, 11, 12 and 13. The first network branch is composed of a series of multiple deep convolutional layers and multiple fully connected layers, and regresses and predicts the bounding box of the workpiece in the color image to form preliminary target detection information. The second network branch is composed of a series of multiple deep convolutional layers, predicting the probability that each pixel in the color image belongs to a specific workpiece to form preliminary instance segmentation information; The post-processing module is composed of a non-maximum suppression unit and a threshold filtering unit. The non-maximum suppression unit processes the preliminary target detection information to eliminate redundant workpiece bounding boxes to form target detection information. The threshold filtering unit filters the preliminary instance segmentation information with a threshold of 0.5 to form instance segmentation information; The stacking estimation algorithm is composed of a series of multiple deep convolutional layers and multiple fully connected layers; The pose estimation algorithm includes: A data preprocessing module, and the data preprocessing module performs statistical filtering and voxel downsampling preprocessing on the workpiece point cloud; A point cloud classification unit that classifies the preprocessed workpiece point cloud and outputs the point cloud category; A class point cloud fusion unit that fuses the preprocessed workpiece point cloud with the point cloud category to form a class point cloud vector; A pose estimation unit that estimates the pose information of the workpiece grasping part relative to the robot base coordinate system according to the class point cloud vector.
2. The three-dimensional visual pose estimation method for disordered workpieces based on deep learning according to claim 1, characterized in that The stacked estimation information is a one-dimensional matrix, and the number of elements in the matrix is equal to the number of workpiece detection images. Each element represents the probability that a workpiece is stacked. The larger the probability value of being stacked, the lower the degree of stacking of the corresponding workpiece.
3. The three-dimensional visual pose estimation method for disordered workpieces based on deep learning according to claim 1, characterized in that The point cloud classification unit includes: A sampling module that randomly samples a fixed number of point clouds from the preprocessed workpiece point cloud; Normalization module, which maps the three-dimensional coordinate values of each point in the point cloud obtained by the sampling module to floating-point numbers between [-a 1 , b 1 ; A point cloud classification network that is formed by cascading a multi-layer perceptron with shared weights, a max pooling layer, and a fully connected layer, and predicts the point cloud category according to the floating-point numbers output by the normalization module.
4. The three-dimensional visual pose estimation method for disordered workpieces based on deep learning according to claim 1, characterized in that The pose estimation unit includes a position estimation unit and an attitude estimation unit, and the pose information includes position information and attitude information.
5. The three-dimensional visual pose estimation method for disordered workpieces based on deep learning according to claim 4, characterized in that The position estimation unit includes: A sampling module that samples the class point cloud vector and forms a vector with a fixed dimension; Normalization module, which calculates the mean of each dimension of the vector collected by the sampling module and maps each value in the vector to a floating-point number between [-a 2 , b 2 ; A position estimation network that obtains the position information of the workpiece grasping part relative to the robot base coordinate system according to the mean value of the vector calculated by the normalization module and the output floating-point numbers.
6. The three-dimensional visual pose estimation method for disordered workpieces based on deep learning according to claim 5, characterized in that The position estimation network consists of two network branches. One network branch is formed by cascading a multi-layer perceptron with shared weights, a max pooling layer, and a fully connected layer, and forms a first position estimation component according to the floating-point numbers output by the normalization module. The other network branch is a fully connected layer that forms a second position estimation component according to the mean value of the vector calculated by the normalization module. Adding the first position estimation component and the second position estimation component gives the position information of the workpiece grasping part relative to the robot base coordinate system.
7. The three-dimensional visual pose estimation method for disordered workpieces based on deep learning according to claim 4, characterized in that The attitude estimation unit includes: A sampling module that samples the class point cloud vector and forms a vector with a fixed dimension; Normalization module, the normalization module processes the vectors obtained by the sampling module and maps each value in the vectors to floating-point numbers within [-a 3 , b 3 ; An attitude estimation network that obtains the attitude information of the workpiece grasping part relative to the robot base coordinate system according to the floating-point numbers output by the normalization module.
8. The three-dimensional visual pose estimation method for disordered workpieces based on deep learning according to claim 7, characterized in that The pose estimation network consists of two network branches, and both network branches are formed by connecting a multi-layer perceptron with shared weights, a max pooling layer, and a fully connected layer in series; One of the network branches estimates the absolute values of the angles of rotation about the X-axis, Y-axis, and Z-axis of the workpiece grasping part relative to the robot base coordinate system based on the floating-point numbers output by the normalization module; the other network branch estimates the direction of rotation about the Z-axis of the workpiece grasping part relative to the robot base coordinate system based on the floating-point numbers output by the normalization module; the output results of the two network branches of the comprehensive pose estimation network are combined to form the pose information of the workpiece grasping part relative to the robot base coordinate system.