A pose estimation method based on region-level feature fusion
By fusing depth information and color features in the pose estimation method and using neural network to process the features, the problem of poor prediction performance of existing methods in heavy occlusion and complex backgrounds is solved, and higher pose estimation accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202111414301.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-11-25
AI Technical Summary
The existing deep learning pose estimation method has poor prediction performance under heavy occlusion and complex backgrounds, and cannot effectively eliminate interfering information, resulting in inaccurate accuracy.
The pose estimation method based on region-level feature fusion is adopted to fuse depth information and color features, use neural networks to process the fusion features, and generate global features through symmetric reduction functions to enhance the details and multi-scaleness of the features, thereby improving the robustness of the algorithm.
In the case of background chaos and severe occlusion, the method shows good robustness, improves pose estimation accuracy, faster calculation speed, and more timely response, exceeding DenseFusion's 6D pose estimation performance.
Smart Images

Figure CN114155406B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of artificial intelligence and relates to a posture estimation method based on region-level feature fusion. Background Art
[0002] 6D pose estimation is to estimate the rotation and translation of an object in 3D space. Specifically, the 6D pose is represented by a rigid transformation [R|t], where R represents the 3D rotation and t represents the 3D translation. 6D pose estimation is an important part of many real-world applications, such as robotic grasping and manipulation, autonomous navigation, and augmented reality.
[0003] Traditionally, the problem of 6D object pose estimation is solved by matching feature points between 3D models and images. However, these methods require objects to have rich textures in order to detect feature points to match. Therefore, they cannot handle objects without textures because the surface of textureless objects does not provide enough information to extract 2D key points. With the advent of depth cameras, several methods have been proposed to recognize objects with less texture using RGB-D data.
[0004] In recent years, due to the great success of deep learning in many other fields, such as object detection and semantic segmentation, deep learning methods have also been tried to be applied in the field of pose estimation. These methods can be divided into three categories. The first category uses deep learning to find the correspondence of 2D-3D or 3D-3D feature points, including explicit and implicit methods, such as the displayed method BB8, the implicit method 3DFeat-Net, to solve the problem that traditional methods are not applicable to textureless objects, however, this method is still sensitive to occlusion; the second category finds the correspondence between the current input and the existing template with 6D pose, including explicit and implicit methods, such as the displayed method PoseCNN, the implicit method SSD6D, however, the accuracy of pose estimation is usually not accurate enough, so these methods require time-consuming post-processing steps (such as ICP) to optimize the pose estimation results; the third category obtains key points by indirect voting of pixels or 3D points and then uses algorithms to obtain 6D poses or directly votes to obtain 6D poses, such as indirect voting PVNet, direct voting DenseFusion.
[0005] The pose estimation method based on deep learning can be applied to textureless objects and can extract effective information from depth images and color images. Therefore, the pose estimation method based on deep learning has achieved satisfactory experimental results in some pose estimation dataset tests. However, the current algorithm still has the following defects:
[0006] In the case of heavy occlusion and complex backgrounds, only part of the object can be seen due to the occlusion. Most of the existing deep learning pose estimation methods based on RGB-D data source usually extract information from color and depth data separately. When extracting features, it is impossible to effectively eliminate interference and obtain effective information, which seriously limits the prediction performance of the algorithm under heavy occlusion and complex backgrounds. Summary of the invention
[0007] The purpose of the present invention is to overcome the defect of poor prediction performance of existing pose estimation methods under heavy occlusion and complex background, and provide a pose estimation method based on regional feature fusion, which fuses depth information and color features together, and then uses the regional fusion features obtained by neural network processing, and uses a symmetric reduction function to process multiple regional fusion features to generate a global feature, and then adds the global feature to each regional fusion feature, so as to obtain color and depth regional fusion features with more details and multi-scales, so that the algorithm has good robustness under background clutter and severe occlusion. At the same time, the pose estimation is divided into two steps, first calculating the three-dimensional translation prediction, and then combining the regional fusion features to more accurately obtain the three-dimensional rotation prediction, so that the network's attention is more focused, the scope of the solution space is narrowed, it is easier to solve, the calculation speed is faster, and the response is more timely.
[0008] The present invention can be achieved through the following technical solutions:
[0009] A pose estimation method based on region-level feature fusion includes the following steps:
[0010] S1. Acquire an image of the object to be inspected through a three-dimensional camera, including a color image and a depth image;
[0011] S2, inputting the color image into a first neural network to extract color features of the object to be inspected;
[0012] S3, converting the corresponding area of the object to be inspected in the depth image into a point cloud image, and then inputting the point cloud image into a second neural network to extract geometric features of the image to be inspected and generate a three-dimensional translation prediction;
[0013] S4, fusing the color features and geometric features pixel by pixel to generate multiple region-level fusion features, and then inputting the multiple region-level fusion features into a multi-layer perceptron to generate multiple three-dimensional rotation predictions and their corresponding confidences;
[0014] S5. Combining the three-dimensional translation prediction and the three-dimensional rotation prediction with the maximum confidence to generate a 6D pose estimation.
[0015] Furthermore, the color features and geometric features are input into a third neural network for pixel-by-pixel fusion to generate multiple region-level fusion features, the multiple region-level fusion features are processed using a symmetric reduction function to generate a global feature, the global feature is then added to each region-level fusion feature, and the multiple added region-level fusion features are then input into a multi-layer perceptron to generate multiple three-dimensional rotation predictions and their corresponding confidence levels.
[0016] Furthermore, the second neural network is set to a PoinNet-like network, including a five-layer network structure. The point cloud image is input into the PoinNet-like network for feature extraction, and the output results of the first layer and the second layer of the network are spliced together to form the geometric features of the object to be inspected. The final output result of the PoinNet-like network is used as the three-dimensional translation prediction of the object to be inspected.
[0017] Furthermore, the PointNet-like network is set to N*3-mlp(3, 640)-mlp(64, 128)-mlp(128, 512)-mlp(512, 1024)-average pooling-mlp(1024, 512, 128, 3), wherein N*3 indicates that the input layer is a point cloud image, N=h*w, h and w respectively represent the height and width of the corresponding area of the object to be inspected in the depth image, mlp represents a multilayer perceptron, the weights are shared by all point cloud points, and average pooling represents an average pooling layer.
[0018] Furthermore, the distance between the sampling point on the object model under the real pose and the corresponding point on the object model under the estimated pose is defined as the pose estimation loss. After training and learning, the pose estimation loss is continuously reduced, and the pose estimation loss with the smallest pose estimation loss is selected as the final 6D pose estimation.
[0019] For asymmetric objects, the pose estimation loss is calculated using the following equation:
[0020]
[0021] Among them, x j represents a 3D point randomly selected from the model of the object to be inspected in the real pose, there are M of them, p = [R|t] represents the real pose, represents the pose estimation after the i-th iteration;
[0022] For symmetric objects, the pose estimation loss is calculated using the following equation:
[0023]
[0024] Among them, x jrepresents a 3D point randomly selected from the model of the object to be inspected in the real pose, there are M of them, p = [R|t] represents the real pose, represents the pose estimation after the i-th iteration; x k represents the distance x from the M three points randomly selected on the model of the object to be inspected under pose estimation j The closest 3D point.
[0025] Furthermore, the network structure corresponding to the final 6D pose estimation is set as the main network. When the pose estimation loss is less than a predetermined value, the optimization network is loaded to optimize the final 6D pose estimation of the main network. The structure of the optimization network is the same as that of the main network.
[0026] Further, the result of the final 6D pose estimation is transformed to obtain a point cloud image, which is then used together with the original color image as the input of the optimization network to calculate the estimated pose residual. Then, the point cloud image obtained after the estimated pose residual is transformed and the original color image are used as the next input of the optimization network until the specified number of iterations is reached. The optimal pose estimation is calculated using the following equation:
[0027] p2=p1·T4·T3·T2·T1
[0028] Among them, p2 is the optimal pose estimate, p1 is the final 6D pose estimate estimated by the main network, and T1, T2, T3, and T4 are the estimated pose residuals calculated by each iteration of the optimization network.
[0029] Further, the preset value is set to 0.013, and the specified number of times is set to 100.
[0030] Beneficial effects:
[0031] First, the present invention proposes regional fusion features for the first time: the network merges depth information into color features, and then uses a neural network to process the resulting fusion features to obtain further detailed color and depth regional fusion features, making the algorithm highly robust in the case of background clutter and severe occlusion.
[0032] Second, the present invention improves the accuracy of pose estimation by using a decoupling method, and divides the pose estimation into two steps: first, evaluating the three-dimensional translation, and then combining the regional fusion features to more accurately obtain the three-dimensional rotation.
[0033] Third, the present invention can achieve 6D pose estimation performance that exceeds DenseFusion on the LINEMOD dataset. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0035] Figure 2 It is a schematic diagram of the structure of the PoinNet-like network of the present invention;
[0036] Figure 3 A schematic diagram of the process of further performing pose estimation on the loading optimization network of the present invention. DETAILED DESCRIPTION
[0037] The specific implementation modes of the present invention are further described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0038] In the description of the present invention, it is necessary to understand that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0039] The present invention proposes a pose estimation method based on region-level feature fusion, such as Figure 1 As shown, the specific steps include:
[0040] S1. Acquire an image of the object to be inspected through a three-dimensional camera, including a color image and a depth image;
[0041] S2, semantic segmentation and color feature extraction:
[0042] The color image is input into the first neural network, such as the PSPNet network. The PSPNet network is an encoder-decoder architecture that takes a color image as input and outputs a segmentation map of N+1 channels, where the first channel describes the background and the other N channels describe objects of N known classes. The bounding box of the object to be tested is found in the output segmentation map, and the color image in the segmented bounding box is used as the color feature. The bounding box is applied to the cropping of the depth image in the following step.
[0043] S3, geometric feature extraction and three-dimensional translation prediction: converting the corresponding area of the object to be inspected in the depth image into a point cloud map, and then inputting the point cloud map into the second neural network to extract the geometric features of the image to be inspected and generate a three-dimensional translation prediction;
[0044] The depth image is cropped according to the bounding box obtained by S2 to obtain a depth image block of size h×w×depth containing only one object to be detected. Then, the depth image block is converted into a point cloud using the following equation to convert the image point [u,v] to the world coordinate point [x w ,y w ,zw ],
[0045]
[0046] Among them, s is the camera zoom parameter, u0,v0 is the center coordinate of the camera, and f x , f y They are the lengths of the x and y axes of the camera respectively.
[0047] Use an architecture similar to the PointNet network to extract geometric features and predict 3D translation: Input the converted point cloud image into the second neural network such as a PoinNet-like network. The specific structure of the PointNet-like network can be N*3-mlp(3, 640)-mlp(64, 128)-mlp(128, 512)-mlp(512, 1024)-average pooling-mlp(1024, 512, 128, 3), such as Figure 2 As shown, N*3 represents the input layer, which is the point cloud map converted from the depth image block, N=h*w, mlp is a multi-layer perceptron, the weights are shared by all point cloud points, average pooling is the average pooling layer, and the final output result has two branches. One is the three-dimensional rotation prediction, which is the final output result of the network, and the other is the multi-stage geometric feature composed of the N*64 feature map extracted by mlp (3*64) and the N*128 feature map extracted by mlp (64*128), which is the output result of the first and second layers of the network spliced together.
[0048] S4, the color features and geometric features are fused pixel by pixel to generate multiple region-level fusion features, and then the multiple region-level fusion features are input into a multi-layer perceptron to generate multiple three-dimensional rotation predictions and their corresponding confidences, as follows:
[0049] a) Feature fusion and region-level fusion feature extraction: So far, geometric features have been extracted from the depth map, and color features have been extracted from the color map. Since the geometric features and color features are one-to-one corresponding, pixel-by-pixel fusion can be used, that is, the geometric features can be fused into the color features point-to-point. The geometric features and color features can be input into the CNN network for point-to-point fusion, and multiple region-level fusion features can be output. The structure of the CNN network can be Input-Conv-LRN-Pooling-Conv-LRN–Pooling, where Input represents the input layer, Conv represents the convolution layer, LRN represents the nonlinear normalization layer, and Pooling represents the downsampling layer.
[0050] b) Three-dimensional rotation prediction: We use a symmetric reduction function to convert multiple region-level fusion features into a global feature from the local feature map, and then add the global feature to each region-level fusion feature. The added multiple region-level fusion features are then input into a multi-layer perceptron to generate multiple three-dimensional rotation predictions. At the same time, the multi-layer perceptron also outputs a prediction confidence ci for each three-dimensional rotation prediction, achieving a self-supervision effect.
[0051] S5. Combining the above three-dimensional translation prediction and the three-dimensional rotation prediction with the highest confidence to generate a 6D pose estimate.
[0052] In order to continuously optimize the results of 6D pose estimation, we can use the LINEMOD dataset to continuously train and learn the above network. The LINEMOD dataset is a widely used benchmark dataset for 6D pose estimation. During training, 2373 training sets can be loaded, but the training will be repeated 20 times, so there will be 2373X20=47460 frames. We set batch_size to 8, 47460 / 8=5932.5, and there are less than 8 in the end, so there will only be 5932 batches loaded, that is: 5932X8=47456 frames.
[0053] Then, the distance between the sampling point on the object model under the real pose and the corresponding point on the object model under the estimated pose is defined as the pose estimation loss. After training and learning, the pose estimation loss is continuously reduced, and the pose with the smallest pose estimation loss is selected as the final 6D pose estimation.
[0054] For asymmetric objects, the pose estimation loss is calculated using the following equation:
[0055]
[0056] Among them, x j represents a 3D point randomly selected from the model of the object to be inspected in the real pose, there are M of them, p = [R|t] represents the real pose, represents the pose estimation after the i-th iteration;
[0057] For symmetric objects, the pose estimation loss is calculated using the following equation:
[0058]
[0059] Among them, x j represents a 3D point randomly selected from the model of the object to be inspected in the real pose, there are M of them, p = [R|t] represents the real pose, represents the pose estimation after the i-th iteration; x k represents the distance x from the M three points randomly selected on the model of the object to be inspected under pose estimationj The closest 3D point.
[0060] In addition, considering the calculation speed and the structure of the network itself, we can use the network mentioned above as the main network and load the optimized network to further optimize the final 6D pose estimation result, such as Figure 3 As shown, the optimization network is not trained together with the main network because it is difficult to converge. Therefore, the main network is first trained until it converges, and then the main network is set fixed before starting to train the optimization network.
[0061] The details are as follows:
[0062] When the pose estimation loss is less than the preset value, such as 0.013, the optimized network will be loaded for training;
[0063] First, transform the result of the final 6D pose estimation to obtain a point cloud image, and then use it and the original color image as the input of the optimization network to calculate the estimated pose residual. Then, the point cloud image obtained after the estimated pose residual is transformed and the original color image is used as the next input of the optimization network until the specified number of iterations, such as 100. If epoch is set to 100, the optimal pose estimate is calculated using the following equation:
[0064] p2=p1·T4·T3·T2·T1
[0065] Among them, p2 is the optimal pose estimate, p1 is the final 6D pose estimate estimated by the main network, and T1, T2, T3, and T4 are the estimated pose residuals calculated by each iteration of the optimization network.
[0066] In order to verify the performance of the present invention, experiments were conducted on a public dataset LINEMOD, and the main network of the present invention, i.e., the regional-level fusion pose estimation network, was analyzed and compared with other benchmark networks. The experiments were trained and tested according to the experimental regulations of the corresponding dataset.
[0067] Among them, Table 1 shows the ADD(-S) metric results of the method of the present invention and other benchmark methods on the LINEMOD dataset. It can be seen from Table 1 that without optimizing the network, the method of the present invention achieves an accuracy of 82.9%, which is higher than other methods without optimizing the network; with the optimized network, the accuracy of the present invention is improved by 13.65%, reaching 96.50%, and the accuracy of the method is also higher than other methods with optimized networks, 7.9% and 2.2% higher than PoseCNN and DenseFusion, respectively.
[0068] Table 2 shows the 2D projection measurement results of the method of the present invention and other benchmark methods on the LINEMOD dataset. It can be seen from Table 2 that these methods are compared according to the presence and absence of the optimization network. Without the optimization network, the method of the present invention achieves an accuracy of 97.47%, which is higher than the accuracy of other methods; with the optimization network, the method of the present invention achieves an accuracy of 97.82%, which is higher than the accuracy of other methods with the optimization network.
[0069] It is worth noting that the method of the present invention achieves better performance than DenseFusion in both ADD(-s) metric and 2D projection metric on the LINEMOD dataset.
[0070] Table 1 - Accuracy of our method and the baseline method on the linemod dataset using the ADD metric
[0071]
[0072] Table 2-2D projection measurement method, the accuracy of this method and the baseline method on the linemod dataset
[0073]
[0074] Although specific embodiments of the present invention are described above, those skilled in the art should understand that these are merely examples and that various changes or modifications may be made to these embodiments without violating the principles and essence of the present invention.
Claims
1. A pose estimation method based on region-level feature fusion, characterized in that The following steps are involved: S1. Acquire an image of the object to be inspected through a three-dimensional camera, including a color image and a depth image; S2, inputting the color image into a first neural network to extract color features of the object to be inspected; S3, converting the corresponding area of the object to be inspected in the depth image into a point cloud image, and then inputting the point cloud image into a second neural network to extract geometric features of the image to be inspected and generate a three-dimensional translation prediction; S4, fusing the color features and geometric features pixel by pixel to generate multiple region-level fusion features, and then inputting the multiple region-level fusion features into a multi-layer perceptron to generate multiple three-dimensional rotation predictions and their corresponding confidences; S5, combining the three-dimensional translation prediction and the three-dimensional rotation prediction with the maximum confidence to generate a 6D pose estimation; Inputting the color features and geometric features into a third neural network for pixel-by-pixel fusion to generate multiple region-level fusion features, processing the multiple region-level fusion features using a symmetric reduction function to generate a global feature, then adding the global feature to each region-level fusion feature, and then inputting the added multiple region-level fusion features into a multi-layer perceptron to generate multiple three-dimensional rotation predictions and their corresponding confidence levels; The second neural network is set to be a PoinNet-like network, including a five-layer network structure. The point cloud image is input into the PoinNet-like network for feature extraction, and the output results of the first layer network and the second layer network are spliced together to form the geometric features of the object to be inspected. The final output result of the PoinNet-like network is used as the three-dimensional translation prediction of the object to be inspected.
2. The method for pose estimation based on regional feature fusion according to claim 1, characterized in that: The PointNet-like network is set to N*3-mlp(3,640)-mlp(64,128)-mlp(128,512)-mlp(512,1024)-average pooling-mlp(1024,512,128,3), where N*3 indicates that the input layer is a point cloud image, N=h*w, h and w respectively represent the height and width of the corresponding area of the object to be inspected in the depth image, mlp represents a multi-layer perceptron, the weights are shared by all point cloud points, and average pooling represents the average pooling layer.
3. The method for pose estimation based on region-level feature fusion according to claim 1, characterized in that: The distance between the sampling point on the object model under the real pose and the corresponding point on the object model under the estimated pose is defined as the pose estimation loss. After training and learning, the pose estimation loss is continuously reduced, and the pose estimation loss with the smallest pose estimation loss is selected as the final 6D pose estimation. For asymmetric objects, the pose estimation loss is calculated using the following equation: Among them, x j represents a 3D point randomly selected from the model of the object to be inspected in the real pose, there are M of them, p = [R|t] represents the real pose, represents the pose estimation after the i-th iteration; For symmetric objects, the pose estimation loss is calculated using the following equation: Among them, x j represents a 3D point randomly selected from the model of the object to be inspected in the real pose, there are M of them, p = [R|t] represents the real pose, represents the pose estimation after the i-th iteration; x k represents the distance x from the M three points randomly selected on the model of the object to be inspected under pose estimation j The closest 3D point.
4. The method for posture estimation based on regional feature fusion according to claim 3, characterized in that: The network structure corresponding to the final 6D pose estimation is set as the main network. When the pose estimation loss is less than a predetermined value, the optimization network is loaded to optimize the final 6D pose estimation of the main network. The structure of the optimization network is the same as that of the main network.
5. The method for posture estimation based on regional feature fusion according to claim 4, characterized in that: The result of the final 6D pose estimation is transformed to obtain a point cloud image, which is then used together with the original color image as the input of the optimization network to calculate the estimated pose residual. Then, the point cloud image obtained after the estimated pose residual is transformed and the original color image are used as the next input of the optimization network until the specified number of iterations is reached. The optimal pose estimation is calculated using the following equation: p2=p1·T4·T3·T2·T1 Among them, p2 is the optimal pose estimate, p1 is the final 6D pose estimate estimated by the main network, and T1, T2, T3, and T4 are the estimated pose residuals calculated by each iteration of the optimization network.
6. The method for posture estimation based on regional feature fusion according to claim 5, characterized in that: The preset value is set to 0.013, and the specified number of times is set to 100.
Citation Information
Patent Citations
Neural network system and neural network method for six-dimensional attitude estimation
CN110956663A
6D pose estimation method based on monocular RGB camera regression depth information
CN113393522A