Global feature fusion 6D pose estimation method based on residual connection

By combining the feature extraction and fusion method of SE-ResNet and PointNet networks with the RealFormer module, the problem of decreased target recognition accuracy of the DenseFusion network in occluded scenes is solved, and higher pose estimation accuracy and robustness are achieved.

CN120747211APending Publication Date: 2025-10-03SHANGHAI NORMAL UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510700595.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The existing DenseFusion network has reduced target recognition accuracy when objects are occluded, and loses spatial structure and detail information during the global feature generation process, making it impossible to effectively capture cross-layer information.

Method used

The SE-ResNet feature extraction network is combined with the PointNet network, and the RealFormer module is used to extract global context features. The squeeze excitation module SE is used to adjust the feature weights, and the multi-head attention mechanism and residual connection structure are used to fuse features to enhance the model's adaptability in complex environments.

Benefits of technology

The accuracy and robustness of 6D pose estimation are improved, and it can maintain local information and capture global context in occluded scenes, thereby improving the recognition accuracy and stability of the model in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747211A_ABST
    Figure CN120747211A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a global feature fusion 6D pose estimation method based on residual connection, and the method comprises the steps: constructing a pose estimation network model; the pose estimation network model is based on a DenseFusion network architecture, an SE-ResNet feature extraction network is adopted to carry out color feature extraction on an RGB image, a Point Net network is adopted to carry out geometric feature extraction on a depth image, after the color features and the geometric features are fused, global context feature extraction is carried out through a RealFormer module, and thus 6D pose estimation is carried out; and training the pose estimation network model by using the data set, and performing pose estimation on the to-be-detected image by using the trained pose estimation network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence and relates to a global feature fusion 6D pose estimation method based on residual connection. Background Art

[0002] Against the backdrop of economic globalization and rapid technological development, the status and role of manufacturing, as the cornerstone of a modern industrial system, have become increasingly prominent. Manufacturing not only provides stable support for economic growth but also plays a vital role in promoting employment, advancing technological progress, and enhancing national competitiveness.

[0003] With the rapid development of computer vision and deep learning algorithms, 6D pose estimation technology has become a key research direction. Applying these advanced technologies to the manufacturing industry will not only significantly improve production efficiency and product quality, but also drive the manufacturing industry towards intelligence and digitalization, further enhancing the industry's global competitiveness. 6D pose estimation technology primarily focuses on determining the position and orientation of objects in three-dimensional space. Its core task is to extract 3D geometric information from 2D images or videos to infer the object's specific position and pose in space. This not only enables robots to accurately identify and locate work objects, but also ensures precise and stable operation.

[0004] The DenseFusion pose estimation algorithm based on direct voting proposes an end-to-end learning method that fuses RGB and depth maps. By combining the features of RGB images and depth maps, it leverages their complementary strengths to form a more comprehensive and robust object representation. As a result, this technique has achieved satisfactory experimental results in pose estimation for some image classification problems. However, this technique currently has the following drawbacks:

[0005] First, the network used by DenseFusion to extract color features employs an encoder-decoder architecture. The encoder uses ResNet18 as its base network, responsible for extracting high-level feature information from the input RGB image. ResNet18 fuses spatial and channel information within the local receptive field of each layer to construct a feature map rich in color information. However, when objects are partially occluded, the color features extracted by ResNet18 may be lost due to the occluded portion of the target object, resulting in reduced target recognition accuracy.

[0006] Second, DenseFusion uses a pixel-level density fusion method to fuse color and geometric features, independently estimating the pose of each pixel or point cloud, and then selecting the optimal pose from all predictions as the final prediction. However, the generation of global features in DenseFusion is completed by performing an average pooling operation on the feature map. This operation aggregates the information of all points into a single global representation, focusing only on the contextual information of the current layer, losing spatial structure and detail information, and ignoring the accumulation of cross-layer information. Summary of the Invention

[0007] This paper proposes a 6D pose estimation method based on global feature fusion of residual connections, which aims to improve the accuracy and robustness of pose estimation by strengthening the feature extraction and fusion process. By combining color features and point cloud features, it can effectively capture the global context while maintaining local information, thereby enhancing the model's adaptability to complex environments and ensuring precise positioning and reliable operation in grasping tasks, in order to solve technical problems such as the loss of target recognition accuracy due to occlusion of part of the target object in the existing DenseFusion network structure.

[0008] The present invention can be achieved through the following technical solutions:

[0009] A 6D pose estimation method based on global feature fusion of residual connections, comprising the following steps:

[0010] Step 1: Build a pose estimation network model;

[0011] The pose estimation network model is based on the DenseFusion network architecture. It uses the SE-ResNet feature extraction network to extract color features from RGB images and the PointNet network to extract geometric features from depth images. After fusing the color and geometric features, the Realformer module extracts global context features to perform 6D pose estimation.

[0012] Step 2: Use the data set to train the pose estimation network model, and use the trained pose estimation network to estimate the pose of the image to be inspected.

[0013] Furthermore, when fusing color features and geometric features, two convolution operations are performed to connect and fuse the color features and geometric features corresponding to each pixel position, generating the intermediate feature map pointfeat1 and the refined feature map pointfeat2 in turn. The refined feature map pointfeat2 is then globally average pooled to extract a preliminary global feature representation. Finally, the preliminary global features are concatenated with the refined feature map pointfeat2 and sent to the RealFormer module to extract global context features.

[0014] Furthermore, when using the SE-ResNet feature extraction network to extract color features, the encoder part uses the ResNet network to extract high-level semantic features, and adaptively adjusts the importance of each channel by introducing the Squeeze-and-Excitation module SE; the decoder part gradually restores the spatial resolution through four upsampling and outputs color features.

[0015] Furthermore, the RealFormer module includes a multi-head attention mechanism module, a feedforward neural network module, and a residual connection structure; wherein the multi-head attention mechanism module is used to perform multi-dimensional attention calculations on input features to capture the dependencies between different feature channels; the feedforward neural network module is used to perform nonlinear transformations on the features after attention calculation to enhance the expressiveness of the features; the residual connection structure is used to directly add the input features to the features processed by the multi-head attention mechanism module and the feedforward neural network module to output global context features;

[0016] The SE-ResNet feature extraction network includes multiple residual blocks and a squeeze excitation module SE; each residual block consists of multiple convolutional layers, batch normalization layers and activation function layers, and is constructed by residual connection to extract high-level semantic features of the image; each squeeze excitation module SE is embedded in the corresponding residual block, and global information at the channel level is obtained by performing a global average pooling operation on the feature map, and then the channel weights are recalibrated through the fully connected layer and activation function layer to adaptively adjust the importance of each channel feature.

[0017] Furthermore, a semantic segmentation network is used to perform target recognition on the RGB image to be inspected, and the corresponding semantic labels are output. According to the semantic labels corresponding to the target objects, the RGB images and depth images are cropped to obtain image crops and mask point clouds of uniform size, which are used as the input of the pose estimation network model.

[0018] Beneficial effects:

[0019] First, we proposed for the first time the use of the squeeze excitation module SE in color feature extraction in the DenseFusion network framework. The squeeze excitation module SE can automatically adjust the feature weights according to the importance of the features, so that the model can more effectively focus on important color features and suppress invalid or noisy features, thereby improving feature expression capabilities.

[0020] Second, for the first time, the Realformer module is used to extract and fuse global context features during the fusion stage of color features and geometric features in the DenseFusion network framework. This Realformer module can simultaneously capture global information and effectively combine feature information from different viewpoints and different sensors, thereby obtaining a more comprehensive feature representation in complex scenes and improving the accuracy and robustness of object pose estimation.

[0021] Third, the accuracy of the 6D pose estimation of the present invention is better than that of mainstream algorithms on the Linemod and OcclusionLinemod datasets. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0023] Figure 2 Schematic diagram of the structure of the extrusion excitation module SE of the present invention;

[0024] Figure 3 Schematic diagram of the structure of the RealFormer module of the present invention. DETAILED DESCRIPTION

[0025] The specific implementation methods of the present invention will be further described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0026] In the description of the present invention, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside", "outside", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore should not be understood as limiting the present invention.

[0027] like Figure 1 As shown in the figure, the present invention proposes a global feature fusion 6D pose estimation method based on residual connection, which aims to improve the accuracy and robustness of pose estimation by strengthening the feature extraction and fusion process. By combining color features and point cloud features, it can effectively capture global context features while maintaining local information, thereby enhancing the model's adaptability to complex environments and ensuring precise positioning and reliable operation in grasping tasks.

[0028] The specific steps include:

[0029] Step 1: Build a dataset, including an image training set and an image test set;

[0030] The dataset is divided into training and test sets according to the standard dataset partitioning method commonly used in most existing models, where 15% of the data is used as the training set and the rest is used for testing;

[0031] i) extracting the target area in the depth image and generating point cloud data;

[0032] Based on the semantic mask corresponding to the target object, the area corresponding to the object in the depth image is extracted. Using the known camera intrinsic parameters, the extracted depth information is converted into a dense point cloud representation of the target object surface, preserving the geometric shape information of the target object;

[0033] ii) Crop the RGB image and depth image to generate input data pairs:

[0034] Crop the RGB image and depth image according to the semantic mask corresponding to the target object to obtain a uniformly sized image crop and masked point cloud, providing a unified input data pair for subsequent feature extraction and pose regression.

[0035] Step 2: Build a pose estimation model and train it, as follows:

[0036] a) Set the structure of the pose estimation model, including the color feature extraction subnetwork, the geometric feature extraction subnetwork, the pixel-level feature fusion module, the RealFormer global feature generation module, the pose regression module, and the pose iterative optimization module; during training, the batch size is fixed to 1, the initial learning rate is 0.0001, the threshold for starting the optimization model is set to 0.013, and the upper limit T for the number of training times is set;

[0037] b) Extract color features from the RGB image and map each pixel information to generate color embedding;

[0038] The SE-ResNet feature extraction network is used as the color feature extraction subnetwork, and the image cropping corresponding to the RGB image is input into the SE-ResNet feature extraction network for color feature extraction. The encoder part uses the ResNet network structure to extract high-level semantic features of the image. However, when the object is partially occluded, the color features extracted by the ResNet network structure may cause a decrease in target recognition accuracy due to the loss of the occluded part of the target object. Therefore, the present invention introduces a Squeeze-and-Excitation (SE) excitation module to adaptively adjust the importance of each channel, thereby enhancing the network's modeling ability for key color features, especially for dealing with occlusion scenes.

[0039] Specifically, the SE-ResNet feature extraction network includes multiple residual blocks and squeeze-and-excitation modules (SE); each residual block consists of multiple convolutional layers, batch normalization layers and activation function layers, and is constructed through residual connections to extract high-level semantic features of the image and solve the problem of difficult training of deep networks; each squeeze-and-excitation module SE is embedded in the corresponding residual block, and global information at the channel level is obtained by performing a global average pooling operation on the feature map, and then the channel weights are recalibrated through the fully connected layer and activation function layer, and the importance of each channel feature is adaptively adjusted, thereby improving the network's ability to extract key features and achieving more effective image feature extraction.

[0040] like Figure 2 As shown in the figure, the squeeze excitation module SE recalibrates the residual features and extracts the global information of each channel through global average pooling, so that the network can adaptively learn which channels are more important in the current task. In this way, the squeeze excitation module SE helps the network adjust the feature weights during training, so that effective features have larger weights, while invalid or less effective features are suppressed, thereby improving the network's robustness and recognition accuracy in the case of occlusion;

[0041] The decoder consists of four upsampling layers, which gradually restore spatial resolution through four upsampling steps, outputting high-quality color feature maps that provide rich local information for subsequent feature fusion. Specifically, each upsampling layer gradually expands the low-resolution feature map, enabling the network to better restore the image's spatial information while maintaining the richness and accuracy of color features. Through this gradual upsampling approach, the decoder can effectively generate high-quality color feature maps, further supporting subsequent object recognition and pose estimation tasks.

[0042] c) extracting geometric features from the mask point cloud and generating geometric embedding;

[0043] The PointNet network is used as a geometric feature extraction subnetwork to process mask point cloud data. It encodes the features of each point through a multi-layer perceptron (MLP) and uses a maximum pooling operation to extract the global feature vector of the overall geometric shape, effectively overcoming the disorder of point cloud data. It finally outputs geometric features that correspond one-to-one to color features and maps them into a geometric embedding with geometric information.

[0044] d) Use pixel-level feature fusion module to fuse color features and geometric features at pixel level

[0045] The color embedding and geometric embedding corresponding to each pixel position are connected and fused. That is, through two convolution operations, the first fusion generates the intermediate feature map pointfeat1, and the second fusion generates the refined feature map pointfeat2. This process achieves a close integration of local spatial information and 3D geometric information, improving the accuracy of subsequent pose regression.

[0046] e) Extract global context features through the RealFormer global feature generation module

[0047] Perform global average pooling on the refined feature map pointfeat2 to extract preliminary global feature representation. The preliminary global features are concatenated with the refined feature map pointfeat2 and sent to the RealFormer module to extract global context features.

[0048] like Figure 3 As shown in the figure, the RealFormer module includes a Multi-Head Attention module, a Feed Forward neural network module, and an Add residual connection structure, which can effectively extract the long-range dependency between point cloud features and image features. The multi-head attention module is used to perform multi-dimensional attention calculations on the input features to capture the dependencies between different feature channels. The feedforward neural network module is used to perform nonlinear transformations on the features after attention calculation to enhance the expressiveness of the features. The residual connection structure is used to directly add the input features to the features processed by the multi-head attention module and the feedforward neural network module to output global context features, avoiding the problem of gradient vanishing or gradient exploding when the network depth increases, thereby improving the stability and convergence speed of network training.

[0049] In the multi-head self-attention module, the model learns the diverse dependencies between point cloud features in parallel, capturing point cloud information at different scales and establishing relationships between different points, regardless of their location, helping the model understand the spatial structure and global semantics of the point cloud. At the same time, historical attention is introduced through residual connections, directly adding the attention results of the previous layer to the current layer. In this way, each layer can obtain accumulated contextual information. Subsequently, the feedforward neural network (FFN) performs a nonlinear transformation on the features processed by the self-attention mechanism, thereby enhancing the expressive power of the feature space and extracting higher-level global semantic information. Through the RealFormer module, the model can comprehensively consider the detailed information of the point cloud and the global contextual relationships, better understanding the semantics and geometric structure of the target object in complex scenes.

[0050] That is to say, the RealFormer module retains historical attention through residual connections and uses a multi-head self-attention mechanism to model the long-distance dependency between point cloud features and image features, thereby effectively capturing global semantic information across local regions and improving the robustness of pose estimation in complex environments.

[0051] f) Perform preliminary 6D pose estimation based on global context features and output confidence;

[0052] The global context features are input into the pose regression module, and the 6D pose (rotation matrix and translation vector) of each pixel is independently predicted, and the preliminary 6D pose estimate is output, and the confidence c is output at the same time. i , used to measure the reliability of the prediction results of each pixel for subsequent optimal posture screening;

[0053] g) Perform K iterations of optimization on the preliminary 6D pose estimate;

[0054] The pose optimization module iterates K times on the initial 6D pose estimate. In each iteration, the point cloud is re-transformed using the current estimated pose, fused with image features, and regressed again, gradually converging to a more accurate final pose. This module can be trained jointly with the main network, but is typically initiated after the main network has converged to avoid early noise affecting the optimization results.

[0055] h) Calculate the training loss function based on the posture prediction error

[0056] The goal of pose estimation is to define the loss function by calculating the distance between the sampling points in the model of the target object in the actual pose and the points transformed by the predicted pose, that is, for each sampling point, calculate the distance error between the predicted pose and the actual pose.

[0057] The pose estimation task in computer vision usually obtains the pose of the object relative to the camera by predicting the rotation matrix and translation vector of the object. However, for symmetrical objects, due to the symmetry of the object, some rotation changes may look the same, resulting in multiple solutions for the predicted pose. In order to adapt to this situation of symmetrical objects, Densefusion adjusts the standard loss function and solves the problem of canonical frame diversity by calculating the minimum distance between each point in the predicted model and the true model point. That is, in the pose estimation process, when the network predicts a preliminary pose, Densefusion does not directly use this pose to calculate the loss, but calculates the distance from each model point to all possible rotation positions in the predicted model, thereby selecting the minimum distance value for loss calculation. The corresponding formula is as follows:

[0058]

[0059] Among them, x j represents the jth point among M points randomly sampled from the 3D model of the target object, p = [R / t] is the true pose of the target object, is the predicted pose of the target object.

[0060] Taking into account the confidence weight and regularization term, during the optimization process, by summing up all dense pixel pose losses, the total loss function is obtained as:

[0061]

[0062] Where N represents the number of dense pixels randomly selected from the feature segment of the target object, c i represents the confidence of each pixel, and w represents the balancing hyperparameter.

[0063] i) Determine whether the training is complete and save the model;

[0064] When the cumulative number of training rounds reaches the set upper limit T or the loss function converges to the preset threshold, the final trained 6D pose estimation model is saved and the training phase ends; otherwise, return to step 2 b) to continue training.

[0065] Step 3: Use the trained pose estimation model to estimate the pose of the image to be inspected.

[0066] For the image to be inspected, first, a semantic segmentation network is used to perform target recognition on the RGB image to be inspected, and the corresponding semantic label is output. According to the semantic label corresponding to the target object, the RGB image and depth image are cropped to obtain the image cropping and mask point cloud of uniform size corresponding to the target object, which is used as the input of the pose estimation network model.

[0067] To verify the feasibility of the method of the present invention, we compared it with existing methods. This example conducted experiments on two public datasets (Linemod and OcclusionLinemod), and analyzed and compared two methods with RGB input and RGB-D input. The experiments were trained and tested according to the experimental specifications of the corresponding datasets.

[0068] Table 1

[0069]

[0070] Table 1 shows the experimental results of the proposed algorithm compared with other baseline algorithms on the Linemod dataset based on the ADD-(S) evaluation metric. The proposed algorithm demonstrates significant performance advantages under the ADD-(S) metric on the Linemod dataset, achieving an average accuracy of 97.6%, outperforming existing baseline methods across multiple evaluation targets. Specifically, the proposed algorithm achieves a 3.3% improvement in accuracy over DenseFusion, a 10.1% improvement over PoseMatcher, a 2.5% improvement over PVN3D, a 0.7% improvement over Trans6D+, a 5.6% improvement over GS-Pose, and a 23.93% improvement over PointFusion. Taking specific object categories as an example, for the cam category, the algorithm of the present invention achieved an accuracy of 98.7%, surpassing the 94.4% of Densefusion and 94.2% of PVN3D, with improvements of 4.3% and 4.5% respectively. For the eggbox category, the algorithm of the present invention, like other methods such as Densefusion and PVN3D, reached 100%, demonstrating excellent robustness. In the lamp category, the algorithm of the present invention reached 98.7%, an improvement of 5.0% compared to 93.7% of PoseMatcher and 3.4% compared to 95.3% of Densefusion.

[0071] Table 2

[0072]

[0073] Table 2 shows the experimental results of our algorithm and other baseline algorithms on the Occlusion Linemod dataset using the ADD-(S) evaluation metric. Based on the experimental data, our algorithm achieved an accuracy of 58.8%, which is 0.9% higher than Trans6D+, 18.0% higher than PVNet, and 33.9% higher than PoseCNN.

[0074] In summary, the overall performance of the proposed algorithm is significantly improved compared with existing methods, further verifying its superiority and robustness in occlusion scenarios.

[0075] Although specific embodiments of the present invention have been described above, those skilled in the art will appreciate that these are merely examples and that various changes or modifications may be made to these embodiments without departing from the principles and essence of the present invention.

Claims

1. A 6D pose estimation method based on global feature fusion of residual connections, characterized by The following steps are involved: Step 1: Build a pose estimation network model; The pose estimation network model is based on the DenseFusion network architecture. It uses the SE-ResNet feature extraction network to extract color features from RGB images and the PointNet network to extract geometric features from depth images. After fusing the color and geometric features, the Realformer module extracts global context features to perform 6D pose estimation. Step 2: Use the data set to train the pose estimation network model, and use the trained pose estimation network to estimate the pose of the image to be inspected.

2. The method for 6D pose estimation based on global feature fusion and residual connection according to claim 1, characterized in that: When fusing color features and geometric features, two convolution operations are performed to connect and fuse the color features and geometric features corresponding to each pixel position, generating the intermediate feature map pointfeat1 and the refined feature map pointfeat2 in turn. The refined feature map pointfeat2 is then globally average pooled to extract a preliminary global feature representation. Finally, the preliminary global features are concatenated with the refined feature map pointfeat2 and sent to the RealFormer module to extract global context features.

3. The method for 6D pose estimation based on global feature fusion and residual connection according to claim 2, characterized in that: When using the SE-ResNet feature extraction network to extract color features, the encoder part uses the ResNet network to extract high-level semantic features and adaptively adjusts the importance of each channel by introducing the Squeeze-and-Excitation module SE; the decoder part gradually restores the spatial resolution through four upsampling and outputs color features.

4. The method for 6D pose estimation based on global feature fusion and residual connection according to claim 2, characterized in that: The RealFormer module includes a multi-head attention mechanism module, a feedforward neural network module, and a residual connection structure; wherein the multi-head attention mechanism module is used to perform multi-dimensional attention calculations on input features to capture the dependencies between different feature channels; the feedforward neural network module is used to perform nonlinear transformations on the features after attention calculation to enhance the expressiveness of the features; the residual connection structure is used to directly add the input features to the features processed by the multi-head attention mechanism module and the feedforward neural network module to output global context features; The SE-ResNet feature extraction network includes multiple residual blocks and a squeeze excitation module SE; each residual block consists of multiple convolutional layers, batch normalization layers and activation function layers, and is constructed by residual connection to extract high-level semantic features of the image; each squeeze excitation module SE is embedded in the corresponding residual block, and global information at the channel level is obtained by performing a global average pooling operation on the feature map, and then the channel weights are recalibrated through the fully connected layer and activation function layer to adaptively adjust the importance of each channel feature.

5. The method for 6D pose estimation based on global feature fusion and residual connection according to claim 1, characterized in that: A semantic segmentation network is used to identify the target in the RGB image to be inspected, and the corresponding semantic label is output. According to the semantic label corresponding to the target object, the RGB image and depth image are cropped to obtain the image cropping and mask point cloud of uniform size, which are used as the input of the pose estimation network model.

Citation Information

Cited By

  • Six-degree-of-freedom pose estimation system and method based on Mamba network

    CN121505037A