Complex workpiece 6D pose estimation method and device, electronic equipment and storage medium
Through the Mask R-CNN and Transformer encoder combined with EdgeConv and the furthest point sampling algorithm, the robustness and accuracy problems of point cloud deep learning in complex industrial scenarios are solved, and efficient pose estimation is achieved.
Patent Information
- Application Number
- CN202510902091.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-01
AI Technical Summary
In complex industrial scenarios, the existing point cloud deep learning methods still have room for improvement in robustness and accuracy when facing problems such as multi-object occlusion, stacking, and point cloud sparseness. It is difficult to efficiently extract the multi-level features of point clouds and improve the global information modeling capabilities.
The Mask R-CNN instance segmentation network is used to process RGB images, and the point cloud coordinate system mapping is established in combination with the 3D camera depth map. EdgeConv and the furthest point sampling algorithm are introduced to extract local features, and global feature modeling is used to model the Transformer encoder, and the translation and rotation parameters are optimized using a dual-branch regression structure.
It significantly improves the robustness and accuracy of pose estimation in complex scenarios, reduces the computational complexity, improves feature representation and global modeling capabilities, and achieves the synchronous improvement of pose regression accuracy and robustness.
Smart Images

Figure CN120411245A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of workpiece pose estimation, and particularly to a method, device, electronic device, and storage medium for 6D pose estimation of complex workpieces. Background Art
[0002] In the intelligent manufacturing system, the robot autonomous grasping technology is a key link to achieve flexible production and efficient logistics. Especially in a complex industrial environment, with diverse workpieces, disordered stacking, serious occlusion and overlap, higher requirements are put forward for the perception and operation capabilities of the robot. 6D pose estimation (i.e., position and orientation in three-dimensional space) is the basis for achieving high-precision grasping and assembly, and its accuracy directly affects the success rate of the robotic arm path planning and operation. Traditional RGB- or RGB-D-based vision methods often struggle to ensure robustness and accuracy in complex scenarios such as occlusion, illumination changes, and object overlap. In recent years, point cloud deep learning methods have become a research hotspot in the field of robot perception due to their powerful modeling ability for spatial geometric features.
[0003] Point cloud data has natural advantages in 6D pose estimation, but how to efficiently extract multi-level features of the point cloud, improve the global information modeling ability, and balance the computational efficiency and generalization ability remains a key problem to be solved urgently.
[0004] With the development of depth sensors and point cloud learning technologies, point cloud- or depth-based methods have gradually received attention. CloudPose extracts the features of point cloud data by using a point cloud processing network (such as PointNet), and directly regresses the 6D pose parameters from the point cloud features using a fully connected layer. CloudAAE proposes a 6D pose estimation method based on point cloud, and achieves high-precision pose estimation through an enhanced autoencoder and a lightweight data synthesis pipeline. Although the above methods have made certain progress in point cloud feature extraction and global information modeling, in complex industrial scenarios, in the face of problems such as multi-object occlusion, stacking, and sparse point clouds, the robustness and accuracy of existing methods still have room for improvement. Summary of the Invention
[0005] Based on this, in view of the deficiencies of existing point cloud methods in terms of feature expression, global modeling, and computational efficiency, it is necessary to provide a method, device, electronic device, and storage medium for 6D pose estimation of complex workpieces.
[0006] A method for 6D pose estimation of complex workpieces, the method comprising: Obtain the RGB image and the corresponding depth map of the target workpiece collected by a 3D camera.
[0007] Process the RGB image using a Mask R-CNN instance segmentation network to obtain a segmentation processing result.
[0008] Based on the segmentation processing results, the corresponding depth map and the camera intrinsic parameter matrix, a mapping relationship between the similarity coordinate system and the three-dimensional point cloud coordinate system is established, and the pixel points in the depth map are converted into three-dimensional point cloud data to obtain the point cloud of the target workpiece.
[0009] The point cloud feature extraction module is used to extract features from the point cloud of the target workpiece to obtain local features and position features; the point cloud feature extraction module is used to extract features from the point cloud using EdgeConv and the farthest point sampling algorithm.
[0010] After the local features and position features are positionally encoded, the global feature extraction module is used to extract features to obtain global features; the global feature extraction module is used to model the global features of the point cloud using the Transformer encoder.
[0011] The global features are processed using the pose regression module to obtain the 6D pose parameters of the target workpiece in three-dimensional space.
[0012] A 6D pose estimation device for a complex workpiece, comprising: The point cloud construction unit of the target workpiece is used to obtain the RGB image and corresponding depth map of the target workpiece captured by the 3D camera; the Mask R-CNN instance segmentation network is used to process the RGB image to obtain the segmentation processing result; based on the segmentation processing result, the corresponding depth map, and the camera intrinsic parameter matrix, a mapping relationship between the similarity coordinate system and the 3D point cloud coordinate system is established, and the pixels in the depth map are converted into 3D point cloud data to obtain the point cloud of the target workpiece.
[0013] The point cloud feature extraction unit is used to extract features from the point cloud of the target workpiece using the point cloud feature extraction module to obtain local features and position features; the point cloud feature extraction module is used to extract features from the point cloud using EdgeConv and the farthest point sampling algorithm.
[0014] The global feature extraction unit is used to position-encode local features and position features and then use the global feature extraction module to extract features to obtain global features; the global feature extraction module is used to model the global features of the point cloud using the Transformer encoder.
[0015] The workpiece pose estimation unit is used to process the global features using the pose regression module to obtain the 6D pose parameters of the target workpiece in three-dimensional space.
[0016] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any of the above methods when executing the computer program.
[0017] A computer-readable storage medium has a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0018] The above complex workpiece 6D pose estimation method, device, electronic device and storage medium. The method introduces a farthest point sampling strategy in the point cloud feature extraction module, effectively reducing point cloud redundancy, enhancing feature representativeness, effectively reducing computational complexity and retaining key geometric features, and further enhancing the global feature modeling ability by combining a Transformer encoder, significantly improving the pose estimation robustness and accuracy in complex scenarios; adopting a dual-branch regression structure to optimize translation and rotation parameters respectively, achieving synchronous improvement of pose regression accuracy and robustness. Brief Description of the Drawings
[0019] Figure 1 It is a schematic flowchart of the complex workpiece 6D pose estimation method in an embodiment; Figure 2 It is a structural diagram of the complex workpiece 6D pose estimation network in an embodiment; Figure 3 It is a schematic diagram of the architecture of the point cloud feature extraction module in an embodiment; Figure 4 It is a schematic diagram of the framework of the global feature extraction module in another embodiment; Figure 5 It is a schematic diagram of the framework of the pose regression module in another embodiment; Figure 6 It is an example diagram of an object model in another embodiment; Figure 7 It is a CAD model diagram of a workpiece in another embodiment; Figure 8 It is a curve diagram of the translation loss function in another embodiment; Figure 9 It is a curve diagram of the rotation loss function in another embodiment; Figure 10 It is a curve diagram of the total loss function in another embodiment; Figure 11 It is a success rate and ADD metric result diagram of the validation set in another embodiment, where (a) is a success rate curve diagram of this method under the ADD distance threshold (10% of the object diameter), and (b) is a schematic diagram of the change of the ADD average value with the number of iterations; Figure 12 It is a displacement loss curve diagram of a finished flange in another embodiment; Figure 13 It is a rotation loss curve diagram of a finished flange in another embodiment; Figure 14Total loss curve graph of the finished flange in another embodiment; Figure 15 Success index result graph of the finished flange in another embodiment, where (a) is the success rate result graph of ADD-S(0.1d), and (b) is the ADD-S result graph; Figure 16 Schematic diagram of the robotic arm grasping process in another embodiment; Figure 17 Internal structure diagram of the electronic device in another embodiment. Detailed implementation manners
[0020] To make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not used to limit this application.
[0021] In one embodiment, as Figure 1 shown, a 6D pose estimation method for complex workpieces is provided, and this method should include the following steps: Step 100: Obtain the RGB image and the corresponding depth map of the target workpiece collected by the 3D camera; process the RGB image using the Mask R-CNN instance segmentation network to obtain the segmentation result; according to the segmentation result, the corresponding depth map, and the camera intrinsic matrix, establish the mapping relationship from the similarity coordinate system to the three-dimensional point cloud coordinate system, and convert the pixel points in the depth map into three-dimensional point cloud data to obtain the point cloud of the target workpiece.
[0022] Specifically, step 100 is the first stage of this method, namely the preprocessing stage. First, take the RGB image obtained by the 3D camera as the input, use the Mask R-CNN instance segmentation network to obtain the bounding box, class label, and instance segmentation mask (mask) of a single workpiece, and then combine the depth map collected by the 3D camera and the camera intrinsic matrix to establish the mapping relationship from the pixel coordinate system to the three-dimensional point cloud coordinate system, and convert the pixel points in the depth map into three-dimensional point cloud data to obtain the point cloud of the target workpiece.
[0023] This method is based on the Mask R-CNN instance segmentation network. By extracting point cloud data from the depth map and combining local and global feature extraction modules, the 6D pose estimation of the target workpiece is realized.
[0024] Step 102: Use the point cloud feature extraction module to extract features from the point cloud of the target workpiece to obtain local features and position features; the point cloud feature extraction module is used to extract features from the point cloud using the EdgeConv and farthest point sampling algorithms.
[0025] Specifically, traditional CloudAAE only uses dynamic point cloud feature aggregation, which is prone to losing spatial distribution information. Therefore, in this method, the farthest point sampling (FPS) strategy is introduced, which can greatly reduce feature redundancy and computational complexity while ensuring uniform spatial coverage of the point cloud, thus enhancing the sensitivity to the structural details of the object.
[0026] Rich features are extracted from the point cloud through a feature extraction module composed of EdgeConv and FPS (Farthest Point Sampling).
[0027] Step 104: After performing position encoding on the local features and position features, a global feature extraction module is used for feature extraction to obtain global features; the global feature extraction module is used to model the global features of the point cloud using a Transformer encoder.
[0028] Specifically, after feature aggregation, multiple layers of Transformer encoders are introduced to efficiently capture the global geometric relationships between distant points in the point cloud using the self-attention mechanism, effectively solving the problem that traditional convolution / pooling methods are difficult to model complex spatial dependencies.
[0029] Before the features extracted by the point cloud feature extraction module are fed into the Transformer encoder, the feature expression ability is further enhanced using learnable position encoding, and then the global features of the point cloud are modeled through the Transformer encoder.
[0030] Step 106: The global features are processed using a pose regression module to obtain the 6D pose parameters of the target workpiece in the three-dimensional space.
[0031] Specifically, the pose regression module is a module used to predict the translation and rotation parameters of an object in 3D space. Using the obtained global geometric features, the final pose parameters of the object can be predicted. The complex workpiece 6D pose estimation network adopted by the complex workpiece 6D pose estimation method is as Figure 2 shown. The complex workpiece 6D pose estimation network includes: a Mask R-CNN instance segmentation network, a point cloud feature extraction module, a global feature extraction module, and a pose regression module.
[0032] In the above complex workpiece 6D pose estimation method, the method effectively reduces the point cloud redundancy by introducing the farthest point sampling strategy in the point cloud feature extraction module, improves the feature representativeness, effectively reduces the computational complexity and retains the key geometric features, and further enhances the global feature modeling ability by combining the Transformer encoder, significantly improving the pose estimation robustness and accuracy in complex scenarios; a double-branch regression structure is adopted to optimize the translation and rotation parameters respectively, realizing the synchronous improvement of the pose regression accuracy and robustness.
[0033] In one embodiment, the point cloud feature extraction module includes: 1 convolutional layer, four EdgeConv layers, and two downsampling operations using the FPS algorithm; step 102 includes: inputting the point cloud of the target workpiece into the convolutional layer to obtain high-dimensional features; inputting the high-dimensional features into the first EdgeConv layer to obtain the first local features and their position features; downsampling the first local features and their position features using the FPS algorithm and then inputting them into the second EdgeConv layer to obtain the second local features and their position features; inputting the second local features into the third EdgeConv layer to obtain the third local features and their position features; downsampling the third local features and their position features using the FPS algorithm and then inputting them into the fourth EdgeConv layer to obtain the local features and their position features.
[0034] In one embodiment, the EdgeConv layer includes: a feature extraction module, a first two-dimensional convolutional layer, a first regularization layer, a max pooling layer, a concatenation layer, a second two-dimensional convolutional layer, and a second regularization layer; wherein, the feature extraction module is used to utilize the coordinate information of the point cloud itself and combine the KNN algorithm to construct a dynamic graph structure.
[0035] Specifically, in order to enable the deep learning method to perform feature extraction on point cloud data, its unique representation problem needs to be solved. Different from image data, point cloud data has two key characteristics: permutation invariance and rotation invariance. Permutation invariance means that the arrangement order of points in the point cloud does not affect its description of the shape and features of the object, that is, the point cloud data can be arbitrarily arranged, and the resulting is still the same object; rotation invariance means that after the point cloud undergoes a rotation transformation, although the spatial coordinates of the points change, its local features should remain unchanged and it is still the same object. These characteristics require that the feature extraction network must design a representation method with geometric transformation robustness. This embodiment elaborates the local feature extraction mechanism from two dimensions of dynamic graph convolution and sampling strategy. By using the FPS algorithm in the process of constructing the graph structure, the point cloud graph structure features of different scales are obtained, ensuring that the important features of the point cloud can be captured at different scales. Optimize the fusion of point cloud feature expressions.
[0036] (1) Dynamic graph convolutional neural network Unlike PointNet and PointNet++, which directly use point clouds as network input, the Dynamic Graph Convolutional Neural Network (DGCNN) adopts a more sophisticated approach, capturing the geometric relationships between points by constructing a graph structure. In DGCNN, each point in the point cloud is treated as a node in the graph, and the relationships between points are represented by edges. This allows the network to capture the geometric relationships between points by performing convolution operations on the graph structure.
[0037] The calculation process of EdgeConv, the core module of the dynamic graph convolutional neural network, consists of three stages: 1) Construct the graph structure. For input n points F dimensional point cloud , where pi represents the characteristics of each point (3D coordinates, color, normal, etc.), if F =3, then it can be expressed as the three-dimensional coordinates of each point . Use the distance metric to calculate its K nearest neighbors for each point ,in is a set of vertices, is the edge set.
[0038] 2) Calculate edge features, for each edge connecting the center point and neighbors , calculate the edge features. The edge features are calculated as follows: ; in, Typically a multi-layer perceptron (MLP), is a learnable parameter.
[0039] 3) Aggregate edge features, use maximum pooling to aggregate neighborhood information, and update center point features: ; in, is the updated central feature.
[0040] (2) Farthest point sampling algorithm The Farthest Point Sampling (FPS) algorithm achieves point cloud downsampling by iteratively selecting the maximum distance points. Its core is to reduce the number of points while preserving the key geometric features of the point cloud as much as possible. The specific process of the algorithm includes: randomly selecting a starting point from the point cloud set P and adding it to the sampling point set S. This starting point can be selected arbitrarily because the iterative nature of the algorithm will ensure the consistency of the final result. Calculate the distance from all points in the point cloud top the distance from the 0 point, select the point with the largest distance as p 1, and add it to the sampling point set S. Repeat the above steps until the required N number of points is sampled.
[0041] (3) Point cloud feature extraction module The point cloud feature extraction module is as Figure 3 shown. The input of the point cloud feature extraction module is the original point cloud data. To map the point cloud data to a higher-dimensional feature space, this method adopts a one-dimensional convolution operation. Through this operation, the coordinate information of each point is converted into a vector with 8-dimensional features, thus laying a foundation for the subsequent feature extraction process. This preliminary feature mapping step enables the point cloud data to be effectively processed and analyzed in a higher-dimensional space. After completing the initial feature mapping, using the coordinate information of the point cloud itself, combined with the KNN (K-Nearest Neighbors) algorithm to construct a dynamic graph structure. Specifically, for each center point in the graph, find its k number of neighboring points through the KNN algorithm. To further enhance the feature expression ability, the feature difference between the center point and the neighboring edge points is used as the feature of the edge. In this way, a feature tensor containing point features and edge features is obtained. Subsequently, a convolution operation is performed on this feature tensor, and the convolutional kernel slides on the feature tensor to extract the feature information in the local neighborhood. To prevent overfitting and ensure the stability and generalization ability of the features, the convolved features are regularized. Finally, a max-pooling operation is performed on the features along the neighborhood dimension, so as to obtain the aggregated features. This aggregated feature can effectively capture the key information in the local neighborhood.
[0042] After completing the above steps, a complete point cloud feature extraction process is achieved. In the original CloudAEE network, the above operation is adopted 4 times, changing the point cloud dimension from 3 to 128 while keeping the number of point clouds unchanged. This design can effectively extract multi-level features of the point cloud, but the computational complexity is relatively high, especially when dealing with large-scale point cloud data. In order to further reduce the computational complexity and extract more representative features, this method uses the FPS algorithm to downsample the point cloud twice. The first downsampling samples the number of point clouds from 2048 to 512. The second downsampling samples the number of point clouds from 512 to 128. The above point cloud feature extraction process is repeated before and after downsampling. The extraction operation after downsampling can be regarded as a secondary feature extraction of the point cloud data, further enhancing the hierarchy and representativeness of the features. After the original network extracts the point cloud features, it also splices the point cloud features of each dimension for subsequent input. However, due to the downsampling operation in this method, the dimensions of each layer are no longer aligned with the number of point clouds and cannot be directly spliced. Therefore, this method selects the features and their coordinates obtained from the last layer for subsequent input. The features of the last layer have higher hierarchy and representativeness after multiple downsampling and feature extraction processes.
[0043] Through this improved point cloud feature extraction network architecture, multi-level features of the point cloud can be effectively extracted, and the complexity of the point cloud can be gradually reduced through a reasonable sampling strategy, so as to provide high-quality input features for subsequent modules, thereby improving the performance and efficiency of the entire feature network.
[0044] In one embodiment, the global feature extraction module includes: a feature fusion module, a Transformer encoder, an MLP, a max pooling layer, and a fully connected layer; step 104 includes: performing positional encoding on the local feature and the positional feature and then inputting them into the feature fusion module to obtain a fused input feature; inputting the fused input feature into the Transformer encoder to obtain an encoded feature; processing the encoded feature through the MLP, the max pooling layer, and the fully connected layer in sequence to obtain the global feature.
[0045] Specifically, (1) Transformer structure The Transformer was originally proposed by Vaswani et al. It was initially used for natural language processing (NLP) tasks. Later, it was extended to computer vision (CV), point clouds, multi-modal learning, and many other tasks. A standard Transformer encoder mainly consists of five parts.
[0046] Input Embedding: The main function of input embedding is to map the original input data into a high-dimensional vector space so that the model can process and learn. Specifically, the role of input embedding varies for different input sources. For text data, through a learnable embedding matrix, each word is mapped to a vector of a fixed dimension. It can convert discrete data (such as words) into continuous data and can also capture the semantic information of the input data (such as the semantic similarity of words). The formula is: ; where, E is the embedding matrix, i is the index of the word, E i is the corresponding word embedding vector. Positional Encoding: In the Transformer model, since the self-attention mechanism itself does not have the ability to explicitly model the order of the input sequence, positional encoding is needed to introduce the position information of the elements in the sequence. Self-Attention Mechanism: The self-attention mechanism calculates the similarity between each element in the sequence and other elements to generate attention weights. These weights are used to weighted aggregate the information of all elements in the sequence to capture the global context relationship. The calculation formula of the self-attention mechanism is as follows: ; where Q, K, and V represent the query, key, and value matrices respectively, and
[0047] is the dimension of the key vector. Feed-Forward Network (FFN): After the self-attention mechanism, Transformer uses a two-layer feed-forward neural network to perform non-linear transformation on the features. Residual Connection and Layer Normalization: To alleviate the problem of gradient vanishing, Transformer adds residual connection and layer normalization operations after each sub-layer (self-attention and FFN).
[0047] (2) Global Feature Extraction Module In the original network, for global features, only a global pooling operation is used to extract global features, but the ability to capture global context information is limited. To make up for this deficiency, this application introduces the Transformer architecture on the basis of the CloudAEE network to enhance the global feature extraction ability. The global feature extraction module of this method is as Figure 4As shown, specifically, the input of the Transformer encoder is the feature processed by the point cloud feature extraction module. To further enhance the expression ability of the feature, the position information of the point cloud is fused with the local feature. The position information is first upsampled through a fully connected layer. Then, the upsampled position information is added to the local feature to form the fused input feature. After layer normalization, the fused input feature is fed into the multi-head self-attention processing unit. After the attention mechanism, a residual is used to connect with the input and regularize it, and then a feed-forward neural network is used to perform a non-linear transformation on the feature. Residual connections and layer normalization operations are added. And a downsampling mapping is used to restore the feature to the original size. Then, the feature will be connected with the attention mechanism feature after residual and regularization, and finally the final global feature of the point cloud is obtained.
[0048] In one embodiment, the local feature and the position feature are position-encoded and then input into the feature fusion module to obtain the fused input feature, including: after the local feature and the position feature are position-encoded, they are respectively processed by a fully connected layer and then added, and then layer normalization is performed to obtain the fused input feature.
[0049] In one embodiment, the pose regression module includes: two independent branch prediction networks, each branch prediction network includes three layers of fully connected layers; the first branch prediction network is used to estimate the translation parameter, and the second branch prediction network is used to estimate the rotation parameter; step 106 includes: inputting the global feature into the first branch prediction network to obtain the three-dimensional coordinate offset of the target workpiece; inputting the global feature into the second branch prediction network to obtain the axis-angle representation regression target of the target workpiece.
[0050] Specifically, the structure of the pose regression module is as Figure 5 shown. This module uses two independent branch prediction networks to estimate the translation and rotation parameters respectively. Each sub-branch consists of three layers of fully connected layers. The translation branch directly regresses the three-dimensional coordinate offset. The global feature is processed through three layers of fully connected layers, and finally the translation offset is output t pred . The rotation branch uses the axis-angle representation as the regression target. The axis-angle representation describes the rotation geometric characteristics through a three-dimensional vector : its direction represents the unit vector of the rotation axis a , and the modulus corresponds to the rotation angle . Based on the Rodrigue formula, the axis-angle parameter can be mapped to the rotation matrix R .
[0051] In the pose regression stage, a dual-branch structure is adopted to optimize the translation and rotation parameters respectively. The rotation branch uses quaternion regression and introduces pose consistency loss, while the translation branch combines the point cloud center and geometric constraints to improve the positioning accuracy. Through the above improvements, this method not only significantly enhances the robustness of pose estimation in complex scenarios such as occlusion and stacking, but also takes into account the inference efficiency.
[0052] In one embodiment, the overall loss function of the pose regression module is defined as the weighted sum of the rotation loss and the translation loss; the expression of the overall loss function of the pose regression module is: ; ; ; where, is the overall loss function of the pose regression module, 、 are the weights of the translation loss and the rotation loss respectively, 、 are the rotation loss and the translation loss respectively, represents the trace of the relative transformation matrix of two rotations, is a function to output the minimum rotation angle between two rotations, R gt and R pred are the true rotation matrix and the predicted rotation matrix respectively, is the predicted translation value in the global coordinate system, is the true translation value.
[0053] Specifically, to quantify the difference between the predicted rotation true values, a rotation loss function and a translation loss function are designed respectively. Let R gt and R pred be the true rotation matrix and the predicted rotation matrix respectively, and the rotation matrix loss function is defined as: ; where, represents the trace of the relative transformation matrix of two rotations, is a function to output the minimum rotation angle between two rotations.
[0054] For the prediction task of the translation offset, since the predicted translation value is obtained based on the local coordinate system after the input point cloud is centered, it is necessary to remap the predicted value back to the original coordinate system and then compare it with the true value. Let the centroid of the input point cloud be , and the predicted translation offset be t pred, the complete predicted translation value should be calculated by the following formula: ; where is the centroid of the point cloud, and the calculation formula is: ; where N represents the number of points in the point cloud, represents the i -th point coordinate. After obtaining the predicted translation value in the global coordinate system, it can be compared with the true translation value , and the translation loss function is calculated. The translation loss function usually uses the L2 norm to measure the difference between the two: .
[0055] For the overall loss function of the pose regression module, it can be defined as: ; where , are the weights of the translation loss and the rotation loss respectively. They are used to balance the influence of the rotation loss and the translation loss. By adjusting , , the optimization focus of the model on rotation and translation prediction can be controlled.
[0056] Experimental results show that this method has achieved excellent performance on both the standard LineMod dataset and the self-made industrial workpiece dataset, verifying its application value in actual industrial scenarios.
[0057] It should be understood that although each step in the flowchart of Figure 1 is shown in sequence according to the arrow indication, these steps do not necessarily execute in the order indicated by the arrow. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in
[0058] In a verification example, the experimental hardware platform uses an NVIDIA GeForce RTX 4090D graphics card (24G video memory), the processor is an Intel(R) Xeon(R) Platinum 8474C (32 cores and 64 threads, main frequency 3.60 GHz), and the memory is 64GB of DDR4; the system runs on Ubuntu20.04.
[0059] At the software level, the algorithm is implemented based on Python3.8. The deep learning framework selects PyTorch1.10.0 and relies on the CUDA11.3 and cuDNN8.2 acceleration libraries. The specific configuration is shown in Table 1.
[0060] Table 1 Experimental Environment
[0061] To optimize the model training process, the following training strategies and parameter configurations are adopted in this section of the experiment: The optimizer selects AdamW, the initial learning rate is set to 0.0003, and the learning rate is adjusted in combination with the StepLR strategy; the weight decay parameter is 0.0001 to prevent overfitting; the BatchSize is set to 64; the total number of training epochs is 200 to ensure that the model converges sufficiently. The specific parameter configuration is shown in Table 2.
[0062] Table 2 Training Parameters
[0063] (1) Dataset 1) Benchmark Datasets This embodiment is mainly verified on two datasets, the LineMod dataset and the self-made workpiece dataset. LineMod contains 13 low-texture daily objects (such as beverage cans, toy ducks, glue bottles, etc.), and each object corresponds to about 12,00 groups of data units. Each data unit contains strictly aligned RGB images, depth maps, camera intrinsics, and accurate 6D pose annotations (rotation matrix and translation vector). The LineMod data is collected in a real scene, and the target object surface lacks significant texture features, and the background environment is complex with dynamic light changes. In this embodiment, following the domain consensus, 15% of each class of objects is selected for training, and the remaining 85% of the data is used for testing. Figure 6 Shows 13 object models selected from the LineMod dataset.
[0064] 2) Self-made dataset Although the LineMod dataset provides a standardized evaluation benchmark for 6D pose estimation algorithms, there are still significant differences in its application to specific industries. To effectively transfer the algorithm to the actual production environment, a self-made dataset for the workpiece stacking scenario is constructed in this embodiment. This dataset consists of two core components: Real annotation data: 100 annotated samples covering both processed and unprocessed workpieces, collected through multi-view RGB-D images and manually calibrated; CAD models of the workpieces as shown in Figure 7 shown.
[0065] (2)Evaluation metrics To evaluate the performance of the 6D pose estimation network, the ADD and ADD-S evaluation metrics are used in this embodiment. For asymmetric objects, the ADD evaluation metric calculates the average deviation of each point distance between the predicted pose and the 3D model points of the true pose, and determines whether the distance accuracy is less than a certain threshold of the object diameter (e.g., ADD-0.1d), defined as follows: ; where: R and t are the true rotation and translation, and are the estimated rotation and translation.
[0066] For symmetric objects, the network prediction results and the true values may differ significantly on the rotation axis, but this does not necessarily mean that the actual effect is not good. Therefore, the average distance of the nearest points ADD-S is used to evaluate the network. Defined as follows: .
[0067] (3)Experimental comparison and analysis 1)LineMod dataset To verify the effectiveness of the improved model, 13 classes on the LineMod dataset are trained in this embodiment. Due to space limitations, only the training process of the Cam class on the LineMod dataset is analyzed in this embodiment. As shown in Figure 8 , Figure 9 , Figure 10 shown, the translation loss, rotation loss, and overall loss of the Cam class in the improved model are respectively shown.
[0068] The experimental results show that the rotation loss and translation loss of the training set (red curve) and the validation set (blue curve) decrease sharply in the first 50 epochs and then gradually tend to stabilize and remain at a low level. When the number of training iterations reaches 75 epochs, each loss function tends to converge. When the network iteration is completed, the translation loss and rotation loss of the final model converge to 1.2076×10-6 and 0.103408 respectively on the training set. On the validation set, the translation loss and rotation loss converge to 2.4004×10-5 and 0.136103 respectively.
[0069] In terms of evaluation metrics, the success rate and average ADD value of this method under the ADD distance threshold (10% of the object diameter) vary with the number of iterations as shown in Figure 11 shown below, where Figure 11 Figure (a) is the success rate curve of this method under the ADD distance threshold (10% of the object diameter), Figure 11 and figure (b) is the schematic diagram of the change of the average ADD value with the number of iterations.
[0070] Through 200 epochs of training, the model achieves an accuracy of 95.4% and an average ADD value of 0.009989 in terms of categories.
[0071] This method is experimentally verified on the LineMod dataset. According to the differences in object symmetry, the ADD (for asymmetric objects) and ADD-S (for symmetric objects) evaluation metrics are respectively used. For asymmetric objects, the symmetric objects in the LineMod dataset are Egg Box and Glue. The experimental results are shown in Table 3, and the results show that this method has achieved significant improvements in most object categories, and the specific analysis is as follows.
[0072] Table 3 Comparison of accuracy on the LineMod dataset
[0073] The above table compares the performance of each method on the LineMod dataset. After improving the CloudAEE network, among the comparisons with the original network in 13 categories, the cam category has the highest improvement amplitude, and the success rate increases from 88.9% to 95.4% (+6.5%). Although there are slight fluctuations in the iron and lamp categories, the iron decreases slightly by 1.3% and the lamp decreases by 0.8%. The improved method achieves positive improvements in 11 out of 13 categories, and the average success rate increases by 1.6% (95.5%→97.1%), which proves the effectiveness of the improvement.
[0074] 2) Self-made dataset Based on the verification results of the benchmark dataset, in this embodiment, the proposed pose estimation framework is further applied to a specific scenario. First, the instance segmentation technology is used to solve the positioning problem of the target workpiece. Then, the point cloud of the target object is extracted and used as the input of the pose estimation network. The improved pose estimation network is used to calculate the pose of the workpiece. The setting of the training parameters is still the same as that in Table 2.
[0075] Two types of industrial workpieces (blank flanges and finished flanges) were selected for verification in the experiment. This embodiment focuses on presenting the training process data of the finished flanges. The convergence curves of the translation loss, rotation loss, and total loss are as Figure 12 , Figure 13 , Figure 14 shown.
[0076] The experimental results show that when the number of iterations reaches 75 epochs, each loss function tends to converge, which is consistent with the performance on the benchmark dataset. This indicates that the model gradually stabilizes during the training process and the optimization effect is significant. When the network iteration is completed, the translation loss and rotation loss of the final model converge to 3.130049×10-6 and 1.482548 respectively on the training set. While on the validation set, the translation loss and rotation loss converge to 1.010853×10-6 and 1.538814 respectively. These results prove that the improved network shows good convergence and stability during both the training and validation processes, and can effectively estimate the pose of the workpiece in the actual grasping scenario.
[0077] Both the blank flanges and finished flanges in the self-made dataset are symmetric objects. In this embodiment, the ADD-S metric accuracy rate is used. The success rate on the finished flange validation set during its training process is as Figure 15 shown, where Figure 15 in (a) is the success rate result graph of ADD-S(0.1d), Figure 15 in (b) is the ADD-S result graph.
[0078] The overall ADD-S accuracy rate on the finished flange validation set shows an upward trend. When the number of training times reaches 200 times, the accuracy rate of the improved model reaches 100%, and ADD-S is 0.002906.
[0079] Table 4 Comparison of accuracy rates on the self-made workpiece dataset
[0080] As can be seen from Table 4, the accuracy of the improved network has increased on the self-made data. The accuracy on the rough workpieces has increased by 0.3% to reach 100%; the accuracy on the finished workpieces has increased by 0.5% to reach 100%; the average ADD-S has increased from 99.6% to 100%, an increase of 0.4%. The experimental results prove that the improvement of the algorithm is effective.
[0081] (4)Grasping experiment of the workpiece grasping platform To verify the effectiveness of the pose estimation algorithm, this method was deployed on an actual industrial workpiece grasping platform, and its grasping performance in typical industrial scenarios such as a small number and dense stacking was systematically evaluated.
[0082] To verify the effectiveness of the algorithm, workpiece grasping experiments were designed under different stacking degrees, namely the scenario of stacking a small number of workpieces (8 workpieces of each type) and the scenario of mixed stacking a large number of workpieces (14 workpieces of each type).
[0083] 1) Scenario of stacking a small number of workpieces The experimental process of the scenario of stacking a small number of workpieces is as Figure 16 shown. The experimental objective is to place all the workpieces in the workpiece frame into the material library. The grasping process is (the serial numbers correspond to those Figure 16 in the figure): 1. The robotic arm is in the initial position, and the workpieces in the workpiece frame are photographed and recognized through the vision system to obtain the pose information of the workpieces; 2. According to the type of the target workpiece, the robotic arm switches to a suitable fixture. For example, if it is a finished flange, it switches to an internal support fixture; if it is a rough flange, it switches to a suction cup; 3. After the fixture is switched, the robotic arm returns to the initial position; 4. The robotic arm grabs the target workpiece; 5. The robotic arm places the grabbed workpiece on the turning table; 6. To perform the turning operation, the robotic arm switches to an external support fixture; 7. The robotic arm completes the turning action; 8. The robotic arm switches back to the internal support fixture to grab the workpiece; 9. The robotic arm places the target workpiece into the material library. Repeat the above process until all the workpieces in the workpiece frame are grabbed.
[0084] For the experiment of the scenario of stacking a small number of workpieces, there are 8 rough flanges and 8 finished flanges in the workpiece frame, and 80 grasping experiments are carried out for each type of workpiece. The experimental results are shown in the following table.
[0085] Table 5 Experimental results of the scenario of stacking a small number of workpieces
[0086] For the rough flanges and finished flanges, when grabbed 80 times, the number of successful grabs is 76 times (95% success rate) and 74 times (92.5% success rate) respectively. The experimental results show that the pose estimation algorithm proposed in this application can meet the grasping requirements when facing a small number of stacked situations.
[0087] 2) Scenario of stacking a large number of workpieces For the experiment on the scenario of mixing and stacking a large number of workpieces, its corresponding work process is the same as that of the experiment on the scenario of stacking a small number of workpieces. There are 14 rough flange plates and 14 finished flange plates in the workpiece box, and 140 grasping experiments are carried out for each type of workpiece. The experimental results are shown in Table 6.
[0088] Table 6 Experimental results of the scenario of stacking a small number of workpieces
[0089] The rough flange plates and the finished flange plates are grasped 140 times, and the number of successful grasps is 130 times (success rate of 92.9%) and 125 times (success rate of 89.3%) respectively. The data shows that although the success rate in the case of densely stacked workpieces is not as high as that in the case of a small number of stacked workpieces, a relatively high grasping success rate is still achieved. Generally speaking, the success rate of the rough flange plates is higher than that of the finished flange plates because the fixture for grasping the rough flange plates is a suction cup, while the fixture for grasping the finished flange plates is an internal support fixture. When the internal support fixture grasps the flange plate close to the edge of the workpiece box, due to the characteristics of the internal support fixture itself, an outward force will be applied to the workpiece during grasping, resulting in the workpiece colliding with the workpiece box, thus triggering the collision protection of the robotic arm and causing the grasping to fail. This is also the main reason for the failure of grasping the finished flange plates. The main reason for the failure of grasping the rough flange plates is mostly the deviation in height when obtaining the pose of the target workpiece, resulting in poor contact between the suction cup and the rough flange plate and the failure of workpiece grasping.
[0090] This method has been fully verified on the standard LineMod dataset and the self-made industrial workpiece dataset. The results show that the average ADD(-S) accuracy of the proposed method on the LineMod dataset reaches 97.1%, and the grasping success rate for both types of workpieces on the self-made dataset reaches 100%, which is significantly better than the original CloudAEE network and other mainstream methods, verifying the effectiveness and engineering application potential of the method. Further, in this embodiment, the algorithm is deployed on an actual industrial workpiece grasping platform, and system experiments are carried out respectively in the scenarios of a small number of stacked and densely stacked workpieces, and the grasping success rates reach 93.8% and 91.1% respectively. The experimental results show that the proposed method not only has good algorithm performance, but also can effectively support the actual industrial automation grasping task and has strong engineering application potential.
[0091] In one embodiment, a complex workpiece 6D pose estimation device is provided, including: a point cloud construction unit for the target workpiece, a point cloud feature extraction unit, a global feature extraction unit, and a workpiece pose estimation unit, where: The point cloud construction unit of the target workpiece is used to obtain the RGB image and the corresponding depth map of the target workpiece collected by the 3D camera; use the Mask R-CNN instance segmentation network to process the RGB image to obtain the segmentation result; according to the segmentation result, the corresponding depth map and the camera internal parameter matrix, establish the mapping relationship from the similarity coordinate system to the three-dimensional point cloud coordinate system, convert the pixel points in the depth map into three-dimensional point cloud data, and obtain the point cloud of the target workpiece.
[0092] The point cloud feature extraction unit is used to extract features from the point cloud of the target workpiece by using the point cloud feature extraction module to obtain local features and position features; the point cloud feature extraction module is used to extract features from the point cloud by using EdgeConv and the farthest point sampling algorithm.
[0093] The global feature extraction unit is used to perform position encoding on the local features and position features and then use the global feature extraction module to extract features to obtain global features; the global feature extraction module is used to model the global features of the point cloud by using the Transformer encoder.
[0094] The workpiece pose estimation unit is used to process the global features by using the pose regression module to obtain the 6D pose parameters of the target workpiece in the three-dimensional space.
[0095] In one embodiment, the point cloud feature extraction module includes: 1 convolutional layer, four EdgeConv layers and two downsampling operations using the FPS algorithm; the point cloud feature extraction unit is also used to input the point cloud of the target workpiece into the convolutional layer to obtain high-dimensional features; input the high-dimensional features into the first EdgeConv layer to obtain the first local feature and its position feature; downsample the first local feature and its position feature using the FPS algorithm and then input them into the second EdgeConv layer to obtain the second local feature and its position feature; input the second local feature into the third EdgeConv layer to obtain the third local feature and its position feature; downsample the third local feature and its position feature using the FPS algorithm and then input them into the fourth EdgeConv layer to obtain the local feature and its position feature.
[0096] In one embodiment, the EdgeConv layer in the point cloud feature extraction unit includes: a feature extraction module, a first two-dimensional convolutional layer, a first regularization layer, a max pooling layer, a concatenation layer, a second two-dimensional convolutional layer and a second regularization layer; among them, the feature extraction module is used to utilize the coordinate information of the point cloud itself and combine the KNN algorithm to construct a dynamic graph structure.
[0097] In one embodiment, the global feature extraction module includes: a feature fusion module, a Transformer encoder, an MLP, a max pooling layer, and a fully connected layer; the global feature extraction unit is further configured to perform positional encoding on the local feature and the positional feature and then input them into the feature fusion module to obtain fused input features; input the fused input features into the Transformer encoder to obtain encoded features; and obtain global features after processing the encoded features through the MLP, the max pooling layer, and the fully connected layer in sequence.
[0098] In one embodiment, the global feature extraction unit is further configured to perform positional encoding on the local feature and the positional feature, respectively process them through fully connected layers, add them, and then perform layer normalization processing to obtain fused input features.
[0099] In one embodiment, the pose regression module includes: two independent branch prediction networks, each branch prediction network includes three layers of fully connected layers; the first branch prediction network is used to estimate the translation parameters, and the second branch prediction network is used to estimate the rotation parameters; the workpiece pose estimation unit is further configured to input the global features into the first branch prediction network to obtain the three-dimensional coordinate offset of the target workpiece; and input the global features into the second branch prediction network to obtain the axis-angle representation regression target of the target workpiece.
[0100] In one embodiment, the overall loss function of the pose regression module in the workpiece pose estimation unit is defined as the weighted sum of the rotation loss and the translation loss; the overall loss function of the pose regression module is as shown in the expression of the overall loss function of the pose regression module above.
[0101] For the specific limitations of the complex workpiece 6D pose estimation device, reference can be made to the limitations of the complex workpiece 6D pose estimation method in the above text, which will not be elaborated here. Each module in the above complex workpiece 6D pose estimation device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0102] In one embodiment, an electronic device is provided. The electronic device can be a terminal, and its internal structure diagram can be as Figure 17As shown. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for estimating the 6D pose of a complex workpiece. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0103] Those skilled in the art can understand that Figure 17 the structure shown in [the figure] is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0104] In one embodiment, an electronic device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps in the above method embodiment.
[0105] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in the above method embodiment.
[0106] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0107] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0108] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application.
Claims
1. A method for 6D pose estimation of complex workpieces, characterized in that, The method includes: Obtaining the RGB image and the corresponding depth map of the target workpiece collected by a 3D camera; Processing the RGB image by using a Mask R-CNN instance segmentation network to obtain a segmentation processing result; According to the segmentation processing result, the corresponding depth map, and the camera intrinsic matrix, establishing a mapping relationship from a similarity coordinate system to a three-dimensional point cloud coordinate system, converting the pixel points in the depth map into three-dimensional point cloud data, and obtaining the point cloud of the target workpiece; Using a point cloud feature extraction module to extract features from the point cloud of the target workpiece to obtain local features and position features; the point cloud feature extraction module is used to extract features from the point cloud by using EdgeConv and the farthest point sampling algorithm; Performing position encoding on the local features and the position features and then using a global feature extraction module to extract features to obtain global features; the global feature extraction module is used to model the global features of the point cloud by using a Transformer encoder; Processing the global features by using a pose regression module to obtain the 6D pose parameters of the target workpiece in three-dimensional space.
2. The 6D pose estimation method for complex workpieces according to claim 1, wherein, The point cloud feature extraction module includes: 1 convolutional layer, four EdgeConv layers, and two downsampling operations using the FPS algorithm; Using a point cloud feature extraction module to extract features from the point cloud of the target workpiece to obtain local features and position features, including: Inputting the point cloud of the target workpiece into the convolutional layer to obtain high-dimensional features; Inputting the high-dimensional features into the first EdgeConv layer to obtain first local features and their position features; Downsampling the first local features and their position features by using the FPS algorithm and then inputting them into the second EdgeConv layer to obtain second local features and their position features; Inputting the second local features into the third EdgeConv layer to obtain third local features and their position features; Downsampling the third local features and their position features by using the FPS algorithm and then inputting them into the fourth EdgeConv layer to obtain local features and their position features.
3. The 6D pose estimation method for complex workpieces according to claim 2, characterized in that The EdgeConv layer includes: a feature extraction module, a first two-dimensional convolutional layer, a first regularization layer, a max pooling layer, a concatenation layer, a second two-dimensional convolutional layer, and a second regularization layer; Wherein, the feature extraction module is used to utilize the coordinate information of the point cloud itself and combine it with the KNN algorithm to construct a dynamic graph structure.
4. The 6D pose estimation method for complex workpieces according to claim 1, characterized in that The global feature extraction module includes: a feature fusion module, a Transformer encoder, an MLP, a max pooling layer, and a fully connected layer; Performing position encoding on the local features and the position features and then using a global feature extraction module to extract features to obtain global features, including: Performing position encoding on the local features and the position features and then inputting them into the feature fusion module to obtain fused input features; Inputting the fused input features into the Transformer encoder to obtain encoded features; Processing the encoded features through an MLP, a max pooling layer, and a fully connected layer in sequence to obtain global features.
5. The complex workpiece 6D pose estimation method according to claim 4, wherein After performing position encoding on the local feature and the position feature, input them into the feature fusion module to obtain the fused input feature, including: After performing position encoding on the local feature and the position feature, process them respectively through fully connected layers, add them together, and then perform layer normalization to obtain the fused input feature.
6. The 6D pose estimation method for complex workpieces according to claim 1, characterized in that The pose regression module includes: two independent branch prediction networks, each branch prediction network includes three layers of fully connected layers; the first branch prediction network is used to estimate the translation parameters, and the second branch prediction network is used to estimate the rotation parameters; Process the global feature using the pose regression module to obtain the 6D pose parameters of the target workpiece in the three-dimensional space, including: Input the global feature into the first branch prediction network to obtain the three-dimensional coordinate offset of the target workpiece; Input the global feature into the second branch prediction network to obtain the axis-angle representation regression target of the target workpiece.
7. The method for estimating the 6D pose of a complex workpiece according to claim 1, characterized in that, The overall loss function of the pose regression module is defined as the weighted sum of the rotation loss and the translation loss; the overall loss function of the pose regression module is: Among them, is the overall loss function of the pose regression module, , are the weights of the translation loss and the rotation loss respectively, , are the rotation loss and the translation loss respectively, represents the trace of the relative transformation matrix of two rotations, is a function to output the minimum rotation angle between two rotations, R gt and R pred are the true rotation matrix and the predicted rotation matrix respectively, is the predicted translation value in the global coordinate system, is the true translation value.
8. An apparatus for estimating the 6D pose of a complex workpiece, characterized in that, The device includes: A point cloud construction unit for the target workpiece, which is used to obtain the RGB image and the corresponding depth map of the target workpiece collected by the 3D camera; process the RGB image using the Mask R-CNN instance segmentation network to obtain the segmentation result; according to the segmentation result, the corresponding depth map, and the camera intrinsic matrix, establish the mapping relationship from the similarity coordinate system to the three-dimensional point cloud coordinate system, and convert the pixel points in the depth map into three-dimensional point cloud data to obtain the point cloud of the target workpiece; A point cloud feature extraction unit, which is used to extract features from the point cloud of the target workpiece using the point cloud feature extraction module to obtain local features and position features; the point cloud feature extraction module is used to extract features from the point cloud using the EdgeConv and farthest point sampling algorithms; A global feature extraction unit, which is used to perform position encoding on the local feature and the position feature and then extract features using the global feature extraction module to obtain global features; the global feature extraction module is used to model the global features of the point cloud using the Transformer encoder; A workpiece pose estimation unit, which is used to process the global feature using the pose regression module to obtain the 6D pose parameters of the target workpiece in the three-dimensional space.
9. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the complex workpiece 6D pose estimation method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the complex workpiece 6D pose estimation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Intelligent visual system
CN112651423A
6D pose estimation method fusing point cloud local features
CN113221647A
6D pose estimation method based on monocular RGB camera regression depth information
CN113393522A
Weak texture target pose estimation method based on feature fusion
CN114821263A
Semantic component attitude estimation method based on deep learning
CN117218343A