Transform model-based end-to-end point cloud registration method and system
Through the end-to-end point cloud registration method based on the Transformer model, the problem of low point cloud registration accuracy at low overlap rate is solved, and higher registration performance and robustness are achieved.
Patent Information
- Application Number
- CN202510189406.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-20
AI Technical Summary
The prior art has low accuracy of point cloud registration at low overlap rate, resulting in unsatisfactory registration results.
The end-to-end point cloud registration method based on the Transformer model is adopted, and the point cloud rigid transformation matrix and predicted position coordinates are generated through deep feature extraction, cross-coding of the GNF feature fusion module and Transformer Encoder.
The performance of point cloud registration at low overlap rate is improved, the robustness of the model is enhanced, the random rotation error and translation error are reduced, and the accuracy of registration is significantly improved.
Smart Images

Figure CN120125625A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial 3D vision, and specifically to an end-to-end point cloud registration method and system based on a Transformer model. Background Art
[0002] 3D point cloud registration is an important part of many fields such as scene reconstruction, autonomous vehicle navigation, and robotics. Recently, with the development of deep learning technology, the need to build an end-to-end registration algorithm that meets industrial requirements has been put on the agenda. For object-level point cloud data such as the data of the middle frame of a mobile phone, the number of its points is often relatively dense. Since a high-precision line laser scanner is used for data acquisition, a series of downsamplings of the dense point cloud are required during the preprocessing process. At the same time, it also has characteristics such as symmetry and high overlap rate, so it increases the difficulty of 3D registration in industry. In practical applications, due to factors such as sensor viewing angle limitations and object occlusion, the low-overlap point cloud to be registered is incomplete, and only some areas overlap, that is, the low overlap rate. Therefore, there is a problem that the accuracy of the existing technology in registration at low overlap rates is low. Summary of the Invention
[0003] The purpose of the present invention is to provide an end-to-end point cloud registration method and system based on a Transformer model for the problem of low accuracy in registration at low overlap rates in the existing technology.
[0004] The technical solution adopted by the present invention to solve the above technical problems is:
[0005] An end-to-end point cloud registration method based on a Transformer model includes the following steps:
[0006] Step 1: Deep feature extraction is respectively performed on the source key points (0) x i and the target key points (0) y i to obtain source unconditional features and target unconditional features;
[0007] Step 2: The source key points (0) x i , the target key points (0) y i , the source unconditional features and the target unconditional features are input into the GNF feature fusion module to obtain the fused features;
[0008] The specific steps executed by the GNF feature fusion module are as follows:
[0009] Step 2-1: Use K-NN to respectively (0) x i and the target key points(0) y i are processed, and the results obtained after processing are respectively compared with the source key points (0) x i and the target key points (0) y i After splicing, they are successively subjected to normalization, dimensionality reduction, activation function, and max pooling processing to obtain the source feature description in the first stage (1) x i and the target feature description in the first stage (1) y i , (1) x i and (1) y i are respectively expressed as:
[0010]
[0011] where h θ represents normalization, dimensionality reduction, and activation function processing, max represents max pooling processing, cat represents the splicing operation, and setting x i ∈R b represents the feature encoding of the source point cloud superpoint set, b represents the dimension of the feature matrix, and i, j ∈ ε represents the graph edge information between the source point cloud superpoint set and the target point cloud superpoint set;
[0012] Step 2: Use K-NN to process the source feature description in the first stage (1) x i and the target feature description in the first stage (1) y i After processing, the results obtained after processing are respectively compared with the source feature description in the first stage (1) x i and the target feature description in the first stage (1) y i After splicing, they are successively subjected to normalization, dimensionality reduction, activation function, and max pooling processing to obtain the source feature description in the second stage (2) x i and the target feature description in the second stage (2) y i , (2) x i and (2) y i are respectively expressed as:
[0013]
[0014] Step 23: Use the multi-head attention mechanism for the source feature description in the second stage (2) x i and the target feature description in the second stage (2) yi Enhance to obtain enhanced feature x i GNF and enhanced feature y i GNF , x i GNF and y i GNF are respectively expressed as:
[0015] x i GNF = (2) x i + MLP[cat(S i , v i )]
[0016] y i GNF = (2) y i + MLP[cat(S i , v i )]
[0017]
[0018] where MLP represents dimensionality reduction, S i and a i represent intermediate variables, b represents the value of the head, represents the learnable weight matrix, q i represents Query, k i represents Key, v i represents Value;
[0019] Step Two and Four: Normalize and apply activation functions to x i GNF and y i GNF respectively to obtain enhanced feature x i ' and enhanced feature y i ', x i ' and y i ' are respectively expressed as:
[0020] x i ' = w θ (x i GNF )
[0021] y i ' = w θ (y i GNF )
[0022] where w θ represents the normalization and activation functions;
[0023] Step 25: Concatenate x i ' and y i ' to obtain the fully connected feature It is expressed as:
[0024]
[0025] Step 26: Concatenate the source unconditional feature and the target unconditional feature to obtain the fully connected feature F, and F is expressed as:
[0026]
[0027] Step 27: Fuse and F to obtain the finally fused feature It is expressed as:
[0028]
[0029] Among them, S θ represents the splitting function, and add represents adding according to the matrix dimension;
[0030] Step 3: After cross-encoding the fused feature The source key points and the target key points in the Transformer Encoder, the output point cloud rigid transformation matrix and the predicted corresponding position coordinates are obtained through a three-layer Mlp network.
[0031] Furthermore, the deep feature extraction in the first step is performed by an improved resampled point convolution network. The improved resampled point convolution network includes 11 downsampling point convolution modules and 1 resampling module. The 11 downsampling point convolution modules are connected in sequence. Between the second downsampling point convolution module and the tenth downsampling point convolution module, there is a skip connection every three layers. The eleventh downsampling point convolution module is connected to the resampling module.
[0032] Furthermore, the multi-head attention mechanism is a 4-head attention mechanism.
[0033] Furthermore, the predicted corresponding position coordinates are expressed as:
[0034]
[0035] Among them, and respectively represent the positions of the predicted source point cloud and the target point cloud, and respectively represent the positions of the actual source point cloud and the target point cloud, represents and The point set obtained after splicing, denotes and The point set obtained after splicing, denotes the indication matrix.
[0036] Furthermore, the point cloud rigid transformation matrix is expressed as:
[0037]
[0038] where R and t respectively denote the rotation matrix and the translation matrix, and respectively denote the optimal rotation matrix and the translation vector, M' and N' respectively represent the superpoint sets of the source point cloud and the target point cloud after feature extraction, and respectively denote and the i-th row of.
[0039] An end-to-end point cloud registration system based on the Transformer model, the system includes a feature extraction module, a feature fusion module, and a cross-encoding module;
[0040] The feature extraction module is used to perform deep feature extraction on the source key points (0) x i and the target key points (0) y i respectively to obtain the source unconditional feature and the target unconditional feature;
[0041] The feature fusion module is used to input the source key points (0) x i , the target key points (0) y i , the source unconditional feature and the target unconditional feature into the GNF feature fusion module to obtain the fused feature;
[0042] The GNF feature fusion module specifically performs the following steps:
[0043] Step 1: Use K-NN to process the source key points (0) x i and the target key points (0) y i respectively, and splice the results obtained after processing with the source key points (0) x i and the target key points (0) y i respectively, and then successively perform normalization, dimensionality reduction, activation function, and max pooling processing to obtain the first-stage source feature description(1) x i and the first-stage target feature description (1) y i , (1) x i and (1) y i are respectively expressed as:
[0044]
[0045] where h θ represents normalization, dimensionality reduction, and activation function processing, max represents max pooling processing, cat represents concatenation operation, set x i ∈R b represents the feature encoding of the source point cloud superpoint set, b represents the dimension of the feature matrix, i, j ∈ ε represents the graph edge information between the source point cloud superpoint set and the target point cloud superpoint set;
[0046] Step 2: Use K-NN to process the first-stage source feature description (1) x i and the first-stage target feature description (1) y i respectively, and then concatenate the processed results with the first-stage source feature description (1) x i and the first-stage target feature description (1) y i respectively, and then successively perform normalization, dimensionality reduction, activation function, and max pooling processing to obtain the second-stage source feature description (2) x i and the second-stage target feature description (2) y i , (2) x i and (2) y i are respectively expressed as:
[0047]
[0048] Step 3: Use the multi-head attention mechanism to enhance the second-stage source feature description (2) x i and the second-stage target feature description (2) y i to obtain the enhanced feature x i GNF and the enhanced feature y i GNF , x i GNF and y i GNF are respectively expressed as:
[0049] x i GNF = (2) x i + MLP[cat(S i , v i )]
[0050] y i GNF = (2) y i + MLP[cat(S i , v i )]
[0051]
[0052] where MLP represents dimensionality reduction, S i and a i represent intermediate variables, b represents the value of the head, represents the learnable weight matrix, q i represents Query, k i represents Key, v i represents Value;
[0053] Step 4: Normalize and apply activation functions to x i GNF and y i GNF respectively to obtain enhanced features x i ' and enhanced feature y i '. x i ' and y i ' are respectively expressed as:
[0054] x i ' = w θ (x i GNF )
[0055] y i ' = w θ (y i GNF )
[0056] where w θ represents the normalization and activation functions;
[0057] Step 5: Concatenate x i ' and y i ' to obtain the fully connected feature expressed as:
[0058]
[0059] Step 6: Concatenate the source unconditional feature and the target unconditional feature to obtain the fully connected feature F, which is expressed as:
[0060]
[0061] Step 7: Fuse and F to obtain the finally fused feature which is expressed as:
[0062]
[0063] where S θ represents the splitting function, and add represents addition according to matrix dimensions;
[0064] The cross - encoding module is used to cross - encode the fused feature After cross - encoding the source key points and the target key points in the Transformer Encoder, the output point cloud rigid transformation matrix and the predicted corresponding position coordinates are obtained through a three - layer Mlp network.
[0065] Furthermore, in the feature extraction module, deep - level feature extraction is performed through an improved resampled point convolution network. The improved resampled point convolution network includes 11 downsampling point convolution modules and 1 resampling module. The 11 downsampling point convolution modules are connected in sequence. Between the second downsampling point convolution module and the tenth downsampling point convolution module, there is a skip connection every three layers. The eleventh downsampling point convolution module is connected to the resampling module.
[0066] Furthermore, the multi - head attention mechanism is a 4 - head attention mechanism.
[0067] Furthermore, the predicted corresponding position coordinates are expressed as:
[0068]
[0069] where and respectively represent the positions of the predicted source point cloud and the target point cloud, and respectively represent the positions of the actual source point cloud and the target point cloud, represents and the point set obtained after concatenation, represents and the point set obtained after concatenation, represents the indicator matrix.
[0070] Furthermore, the point cloud rigid transformation matrix is expressed as:
[0071]
[0072] Among them, R and t represent the rotation matrix and the translation matrix respectively. and represent the optimal rotation matrix and translation vector respectively. M' and N' represent the superpoint sets of the source point cloud and the target point cloud after feature extraction respectively. and represent respectively and the i-th row of.
[0073] The beneficial effects of the present invention are:
[0074] This application uses the Resample KPConv as the backbone, which improves the registration performance of the model under low overlap rates. And a GNF feature fusion module that can fuse and enhance the neighborhood information of adjacent points is designed. It is located in the bottleneck layer and can enhance the robustness of the network when facing sparse and noisy data. Finally, this application uses an 8-layer Transformer Encoder network to perform deeper interaction on the information of the source point cloud and the target point cloud, so as to obtain more accurate conditional features. Finally, the accuracy of point cloud registration under low overlap rates can be improved according to the simple Mlp output, and the random rotation error and translation error can be effectively reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 is the flow chart of this application;
[0076] Figure 2 is the GNFTR structure diagram;
[0077] Figure 3 is the Resample KPConv backbone network structure diagram;
[0078] Figure 4 is the structure diagram of the GNF feature fusion module;
[0079] Figure 5 is the Transformer encoding layer structure diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0080] It should be specifically noted that, without conflict, the various embodiments disclosed in this application can be combined with each other.
[0081] Specific Embodiment 1: The end-to-end point cloud registration method based on the Transformer model described in this embodiment first inputs the processed source point cloud and target point cloud into the ResampleKPConvBackbone. Through four layers of downsampling, the initial input point cloud data is sampled into a superpoint set. Subsequently, the superpoint set and its corresponding features are fed into the GNF feature fusion module designed by us. The missing neighborhood information in the features is compensated through the self-attention mechanism and the K-NN algorithm. By the advantages of the dynamic weight allocation of the attention mechanism, the influence of noise points and outlier data points is suppressed. Subsequently, the features and the corresponding point set are input into the Transformer model for cross-encoding. Utilizing the powerful interaction ability of the Transformer, the features in the source point cloud and the target point cloud are deeply interacted. Finally, the corresponding points in the source point cloud and the target point cloud are predicted through 3 layers of Mlp, and finally the rigid transformation matrix and the finally registered point cloud data are output. The overall flow chart of this application is as shown in Figure 1 shown.
[0082] Specifically:
[0083] Step 1: Make point cloud datasets of different categories, including scene-level and object-level point cloud data respectively. Through preprocessing, operations such as random rotation and translation are performed on the datasets, and Gaussian noise is also added. And they are divided into training sets, validation sets and test sets in different proportions, divided into training sets, validation sets and test sets according to 7:2:1. Among them, the scene-level dataset is similar to 3DMatch, which includes laboratory space point cloud sets, classroom point cloud sets, etc. The object-level point cloud data includes daily necessities, common small objects such as mobile phones, cups, chairs, etc. In actual applications, 3D point cloud data can be captured by a DeepVision intelligent camera. Through preprocessing operations such as random data augmentation and adding Gaussian noise to the point cloud, and uniformly converting it into the point cloud format of ply.
[0084] Step 2. Apply the training set obtained in Step 1 to the network model for training. Input the source point cloud data and target point cloud data in the mobile phone frame dataset into the improved resampled point convolutional network for feature extraction. It goes through 11 downsampling point convolutional modules and 1 resampling module in total. Introducing the resampling module at the lowest level can improve the robustness and registration recall rate of the model in low overlap rate registration, and effectively reduce errors. The resampled point convolutional module is connected in a residual form as the backbone network, increasing the diversity and effectiveness of the features obtained by the network, enabling the network to obtain deeper features, and improving the robustness and registration performance of the network. Among them, a skip connection is required after every 3 downsampling layers, and the specific connection method is in the form of the Resnet residual form to obtain the deep features of the source point cloud and target point cloud, which are respectively called the source unconditional feature and the target unconditional feature. Among them, the resampled point convolutional module (Resample KPConv Block) is a variant of the point convolutional module (KPConv Block, as Figure 3 shown), which can dynamically adjust the positions of feature points based on multi-scale feature extraction, so as to realize the redistribution of points during the decoding process. This dynamic adjustment mechanism can not only effectively alleviate the interference of noise points on feature extraction, but also provide a consistent feature representation between different density regions, ensuring that the features in sparse regions can also be described with high quality.
[0085] In the scene data matching task, since the mutual correlations between many point cloud features often exceed the local neighborhood range, the attention mechanism can take into account both local details and global geometric relationships at the same time, thereby improving the overall matching accuracy. In addition, when processing 3D point cloud data that usually contains noise and is sparsely distributed, according to Step 3 of the input, the corresponding key points and corresponding features obtained by downsampling are fed into the GNF feature fusion module (the structure is as Figure 4 shown) to balance the performance of the model and use the attention mechanism to take into account both local details and global geometric relationships at the same time; feed it into the designed GNF feature fusion module for depth compensation of neighborhood information. The following formula briefly describes the GNF representation methods of X and Y. Among them, the sparse point set generated by downsampling is selected, and then the K-NN algorithm is used to connect several sparse points into a three-dimensional space graph.
[0086] Step 3. Feed the source key points, source unconditional features, target key points, and target unconditional features obtained by downsampling into the GNF feature fusion module at the same time to balance the performance of the model and use the attention mechanism to take into account both local details and global geometric relationships at the same time. The following formula briefly describes the GNF representation methods of X (source point cloud) and Y (target point cloud). Among them, the source key points and target key points generated by downsampling are selected, and then the K-NN algorithm is used to connect several sparse points into a three-dimensional space graph. The feature encoding of the sparse point set is xi ∈R b It means that subsequently, a normalization operation is used to normalize the feature matrix formed by connecting K-NN (in the K-NN algorithm of this application, self-attention is introduced. The main idea is to connect two layers of K-NN modules in series, introduce three ports between the first and second layers of K-NN, project the information of the third layer into target vectors through the q, k, and v of the multi-head self-attention mechanism, and finally fuse them into the relevant information of the final GNF through the softmax algorithm). Subsequently, dimensionality reduction operations and max pooling operations need to be performed on the K-NN representations of the connected source point cloud and target point cloud respectively. Thus, the overall formula for the feature description method in the first stage that can be obtained can be written as:
[0087]
[0088] Where (0) x i , (0) y i represent the initial input source key points and target key points. The operations in the formula mean that in each calculation process of the K-NN algorithm, there is a meaning of crossing the original key points and target key points. The (1) x i in the formula represents the first-stage K-NN representation of the source point cloud, and it can be extended to (1) x i being the representation of the first stage of the target point cloud, where max is the max pooling operation, and h θ represents a series of operations such as normalization, dimensionality reduction, and activation functions. Subsequently, the second K-NN description method will be performed, and the same will be updated to:
[0089]
[0090] The symbols therein represent various operations and remain unchanged, all representing the same operation meaning. In the subsequent calculations, this application introduces a multi-head self-attention mechanism to enhance the features described by K-NN, which can be specifically represented as:
[0091]
[0092] The formula described here is the general formula for the multi-head attention mechanism. The number of heads we selected is 4, where W O represents the learnable weight matrix. Where:
[0093]
[0094] Where q i , k i , v iDifferent from the traditional Q, K, and V, we use the K-NN description methods at different stages as multiple projections of the multi-head self-attention mechanism. The following calculations also need to be performed according to the above operations:
[0095]
[0096] In the above formula The formula for this part is the fixed calculation step in the attention mechanism, which replaces the parameters originally used as the random projection matrix with the features of different stages of the k-NN description method. Therefore, this application encapsulates the above part into a new GNF representation method, which can be respectively represented as x i GNF , y i GNF . The specific operation is completed according to the following formula:
[0097] x i GNF = (2) x i + MLP[cat(S i , v i )]
[0098] y i GNF = (2) y i + MLP[cat(S i , v i )]
[0099] Among them, x i GNF represents the GNF description method corresponding to the source point cloud, and the GNF representation method in the target point cloud is the same as above. After finishing this part, I will perform distribution fusion on the useful feature part.
[0100] Subsequently, it is necessary to pass through the feature fusion part to finally enhance the source key point features and the target key point features. The first step of fusion is performed as follows:
[0101] x i ' = w θ (x i GNF )
[0102] y i ' = w θ (y i GNF )
[0103] Among them, w θ represents a series of operations, including normalization and activation functions, etc. Subsequently, the feature fusion in the first stage is shown in the following formula:
[0104]
[0105] In the formula What is fused is the GNF description of the source point cloud and the GNF description of the target point cloud. Among them, F fuses the features of the two point clouds before entering the GNF module, concatenates the unconditional features of the source point cloud and the unconditional features of the target point cloud into an overall vector, and this part can be equivalent to a fully connected operation. Then the last two steps can be expressed together by the following formula:
[0106]
[0107] Where S θ represents the splitting function, which can split the long array matrix formed by a whole fully connected operation according to the original position, restore the fully connected operation. Add represents adding according to the matrix dimension, and MLP is to use only one layer of convolution for dimensionality reduction operation.
[0108] This application is inspired by the GNN and Transformer modules (the structure diagram of the Transformer encoding layer is as shown in Figure 5 ) and designs the GNF feature fusion module. It introduces the self-attention mechanism in the k-nn algorithm, which can dynamically adjust the weights of each point according to the context of the input features, focus on the points crucial to the task, and at the same time suppress the influence of noise or irrelevant points, thereby enhancing the robustness of the network when facing sparse and noisy data.
[0109] This application uses the GNF feature fusion module as the transitional structure of the bottleneck layer to capture the multi-scale information of the point cloud. This module can enhance the ability to capture local and global features, which is the key factor for its excellent performance on the mobile phone middle frame dataset. By dynamically adjusting the weights for each feature and establishing connections between points globally, this module enables the network to focus on those points that carry important geometric information.
[0110] Step 4. Input the fused feature data and point information into the TransformerEncoder for cross-encoding,
[0111] Use the TransformerEncoder layer to perform deep cross-encoding on the features and key points corresponding to the source point cloud and the target point cloud. First, each passes through the self-attention module for self-feature enhancement, and at the same time performs sine encoding according to the matrix of key points, encodes the information of the key points and embeds it into the corresponding feature information. Then, deep information interaction is performed through cross-attention. Finally, information aggregation is performed according to the fully connected FFN. Finally, the point cloud rigid transformation matrix and the predicted corresponding position coordinates are output through a three-layer Mlp network.
[0112] The Transformer Encoder module in step 4 has some differences from that in NLP tasks. It is a variant for the format of point cloud data and can take data as dual input channels simultaneously. To fuse global information and local feature details, feature encoding is particularly important. It is more suitable for point cloud-related tasks and overcomes problems such as insufficient global position information and loss of absolute position information of points.
[0113] Step 5. Finally, output the rigid transformation matrix of the point cloud and the predicted corresponding position coordinates through a three-layer Mlp network. At the same time, use a single fully connected layer with sigmoid activation to predict the overlap confidence respectively. This can be used to mask the influence of correspondences that are not accurately predicted outside the overlapping region. And use the validation set to adjust the hyperparameters of the model. After determining the best hyperparameters, use the test set to evaluate the registration performance of the model. The main evaluation metrics include registration recall rate (RR), random translation error (RTE), and random rotation error (RRE), etc.
[0114] The various symbols in the formula are as follows in this example. Here, it is used to connect the predicted transformation positions in two directions to obtain the final set of M'+N' correspondences. This part uses a weighted variant of the Kabsch-Umeyama algorithm to calculate and obtain a closed form:
[0115]
[0116] Among them, and represent the positions of the predicted source point cloud and target point cloud respectively. and represent the actual positions of the source point cloud and target point cloud respectively. represents and the point set obtained after splicing. represents and the point set obtained after splicing. represents the indicator matrix used to mark which points are valid correspondences.
[0117] The rigid transformation matrix of the point cloud is expressed as:
[0118]
[0119] Among them, R and t represent the rotation matrix and translation matrix respectively. and respectively represent the optimal rotation matrix and translation vector. M' and N' respectively represent the superpoint sets of the source point cloud and the target point cloud after feature extraction. and respectively represent and the i-th row of.
[0120] The point cloud rigid transformation matrix is a commonly used optimization objective function in point cloud registration, usually used in the Iterative Closest Point (ICP) algorithm or its variants. The goal of the formula is to find an optimal rotation matrix and translation vector to minimize the distance between the source point cloud and the target point cloud.
[0121] Objective function: The goal of the formula is to find an optimal rotation matrix and translation vector to minimize the objective function. The objective function is the sum of the squares of the distances between all corresponding point pairs.
[0122] Summation range: The summation range is from 1 to M′ + N′, where M′ and N′ are the numbers of points participating in registration in the source point cloud and the target point cloud respectively.
[0123] Distance term: is an indicator function used to mark the point in the source point cloud whether it successfully matches the point in the target point cloud. If the match is successful,
[0124] Rotation and translation: represents the coordinates of the point in the source point cloud after rotation and translation.
[0125] Target point: is the point in the target point cloud corresponding to the point in the source point cloud
[0126] Distance: is the square of the Euclidean distance between the point in the source point cloud after rotation and translation and the corresponding point in the target point cloud.
[0127] This application uses the Transformer Encoder module to perform deep cross-encoding on data. By integrating the self-attention and cross-attention mechanisms, it realizes the effective learning and fusion of point cloud features, improving the performance of the point cloud registration task. At the same time, its end-to-end design avoids the cumbersome steps in traditional methods, making the model more concise and easier to train.
[0128] In the last step 5, the point cloud rigid transformation matrix and the predicted corresponding position coordinates are output through a three-layer Mlp network. After using the validation set to adjust the hyperparameters of the model and determining the optimal hyperparameters, the registration performance of the model is evaluated using the test set. At this point, the registration task can be perfectly completed. To make the registration effect more beautiful, we provide the source code that can replace the weight file at any time and observe the registration visualization effect. The error values for the source point cloud and the target point cloud will also be output during operation.
[0129] Experiments show that GNFTR not only has excellent registration performance on the mobile phone middle frame dataset, but also has good performance on other scene-level and object-level datasets. The point cloud registration algorithm based on GNFTR in this application realizes efficient, accurate, and rapid registration, with excellent registration effects. It provides strong support for industrial inspection.
[0130] It should be noted that the specific implementation manners are only explanations and illustrations of the technical solutions of the present invention, and the scope of the right protection cannot be limited thereby. Those that are only partial changes made according to the claims and the specification of the present invention should still fall within the protection scope of the present invention.
Claims
1. End-to-end cloud registration method based on Transformer model, characterized by The following steps are involved: Step 1: Targeting the source key points (0) x i and target key points (0) y i Perform deep-level feature extraction respectively to obtain source unconditional features and target unconditional features; Step 2: Source key points (0) x i , target key points (0) y i , the source unconditional features and the target unconditional features are input into the GNF feature fusion module to obtain the fused features; The GNF feature fusion module specifically performs the following steps: Step 21: Use K-NN to analyze the source key points (0) x i and target key points (0) y i Processing is performed and the processed results are compared with the source key points (0) x i and target key points (0) y i After splicing, it is processed by normalization, dimensionality reduction, activation function and maximum pooling in turn to obtain the first stage source feature description (1) x i and the first stage target characterization (1) y i , (1) x i and (1) y i Respectively expressed as: Among them, h θ represents normalization, dimensionality reduction and activation function processing, max represents maximum pooling processing, cat represents concatenation operation, and sets x i ∈R b It represents the feature encoding of the source point cloud super point set, b represents the dimension of the feature matrix, and i,j∈ε represents the edge information between the source point cloud super point set and the target point cloud super point set; Step 22: Use K-NN to describe the source features of the first stage (1) x i and the first stage target characterization (1) y i Processing is performed and the processed results are compared with the source feature description of the first stage (1) x i and the first stage target characterization (1) y i After splicing, it is processed by normalization, dimensionality reduction, activation function and maximum pooling in turn to obtain the second stage source feature description (2) x i and second stage target characterization (2) y i , (2) x i and (2) y i Respectively expressed as: Step 23: Use the multi-head attention mechanism to describe the second stage source features (2) x i and second stage target characterization (2) y i Enhance and obtain enhanced feature x i GNF and enhanced feature y i GNF , x i GNF and i GNF Respectively expressed as: x i GNF = (2) x i +MLP[cat(S i ,v i )] and i GNF = (2) and i +MLP[cat(S i ,v i )] Among them, MLP represents dimensionality reduction, S i and a i represents the intermediate variable, b represents the value of head, represents the learnable weight matrix, q i represents Query, k i Indicates Key, v i Indicates Value; Step 24: x i GNF and i GNF Normalization and activation function processing are performed respectively to obtain the enhanced feature x i ' and enhanced feature y i ', x i ' and y i 'Respectively expressed as: x i '=w θ (x i GNF ) y i '=w θ (y i GNF ) Among them, w θ represents normalization and activation function; Step 25: x i ' and y i 'Concatenate to get fully connected features It is expressed as: Step 26: Concatenate the source unconditional features and the target unconditional features to obtain the fully connected feature F, which is expressed as: Step 27: F is fused with F to obtain the final fused features It is expressed as: Among them, S θ represents the splitting function, and add represents addition according to the matrix dimension; Step 3: The fused features After the source key points and the target key points are cross-encoded in the Transformer Encoder, the output point cloud rigid transformation matrix and the predicted corresponding position coordinates are obtained through a three-layer MLP network.
2. The end-to-end cloud registration method based on the Transformer model according to claim 1 is characterized in that The deep feature extraction in the step 1 is performed by an improved resampling point convolutional network, which comprises 11 layers of downsampling point convolutional modules and 1 layer of resampling modules. The 11 layers of downsampling point convolutional modules are connected in sequence, and the second layer of downsampling point convolutional modules and the tenth layer of downsampling point convolutional modules have a jump connection every three layers, and the eleventh layer of downsampling point convolutional modules are connected to the resampling module.
3. The end-to-end cloud registration method based on the Transformer model according to claim 1 is characterized in that The multi-head attention mechanism is a 4-head attention mechanism.
4. The end-to-end cloud registration method based on the Transformer model according to claim 3 is characterized in that The corresponding position coordinates of the prediction are expressed as: in, and Respectively represent the positions of the predicted source point cloud and target point cloud, and Respectively represent the positions of the actual source point cloud and target point cloud, express and The point set obtained after splicing is express and The point set obtained after splicing is represents the indicator matrix.
5. The end-to-end cloud registration method based on the Transformer model according to claim 4 is characterized in that The point cloud rigid transformation matrix is expressed as: Among them, R and t represent the rotation matrix and translation matrix respectively, and Represent the optimal rotation matrix and translation vector respectively, M' and N' represent the super point set of the source point cloud and the super point set of the target point cloud after feature extraction respectively, and Respectively and The i-th row of .
6. End-to-end cloud registration system based on Transformer model, characterized by The system includes a feature extraction module, a feature fusion module and a cross encoding module; The feature extraction module is used to extract the source key points (0) x i and target key points (0) y i Perform deep-level feature extraction respectively to obtain source unconditional features and target unconditional features; The feature fusion module is used to transform the source key points (0) x i , target key points (0) y i , the source unconditional features and the target unconditional features are input into the GNF feature fusion module to obtain the fused features; The GNF feature fusion module specifically performs the following steps: Step 1: Use K-NN to analyze the source key points (0) x i and target key points (0) y i Processing is performed and the processed results are compared with the source key points (0) x i and target key points (0) y i After splicing, it is processed by normalization, dimensionality reduction, activation function and maximum pooling in turn to obtain the first stage source feature description (1) x i and the first stage target characterization (1) y i , (1) x i and (1) y i Respectively expressed as: Among them, h θ represents normalization, dimensionality reduction and activation function processing, max represents maximum pooling processing, cat represents concatenation operation, and sets x i ∈R b It represents the feature encoding of the source point cloud super point set, b represents the dimension of the feature matrix, and i,j∈ε represents the edge information between the source point cloud super point set and the target point cloud super point set; Step 2: Use K-NN to describe the source features of the first stage (1) x i and the first stage target characterization (1) y i Processing is performed and the processed results are compared with the source feature description of the first stage (1) x i and the first stage target characterization (1) y i After splicing, it is processed by normalization, dimensionality reduction, activation function and maximum pooling in turn to obtain the second stage source feature description (2) x i and second stage target characterization (2) y i , (2) x i and (2) y i Respectively expressed as: Step 3: Use the multi-head attention mechanism to describe the second-stage source features (2) x i and second stage target characterization (2) y i Enhance and obtain enhanced feature x i GNF and enhanced feature y i GNF , x i GNF and i GNF Respectively expressed as: x i GNF = (2) x i +MLP[cat(S i ,v i )] and i GNF = (2) and i +MLP[cat(S i ,v i )] Among them, MLP represents dimensionality reduction, S i and a i represents the intermediate variable, b represents the value of head, represents the learnable weight matrix, q i represents Query, k i Indicates Key, v i Indicates Value; Step 4: x i GNF and i GNF Normalization and activation function processing are performed respectively to obtain the enhanced feature x i ' and enhanced feature y i ', x i ' and y i 'Respectively expressed as: x i '=w θ (x i GNF ) y i '=w θ (y i GNF ) Among them, w θ represents normalization and activation function; Step 5: Place x i ' and y i 'Concatenate to get fully connected features It is expressed as: Step 6: Concatenate the source unconditional features and the target unconditional features to obtain the fully connected feature F, which is expressed as: Step 7: F is fused with F to obtain the final fused features It is expressed as: Among them, S θ represents the splitting function, and add represents addition according to the matrix dimension; The cross encoding module is used to transform the fused features After the source key points and the target key points are cross-encoded in TransformerEncoder, the output point cloud rigid transformation matrix and the predicted corresponding position coordinates are obtained through a three-layer Mlp network.
7. The end-to-end cloud registration system based on the Transformer model according to claim 6, characterized in that The deep-level feature extraction in the feature extraction module is performed through an improved resampling point convolution network, which includes 11 layers of downsampling point convolution modules and 1 layer of resampling module. The 11 layers of downsampling point convolution modules are connected in sequence, and the second layer of downsampling point convolution module and the tenth layer of downsampling point convolution module have a jump connection every three layers, and the eleventh layer of downsampling point convolution module is connected to the resampling module.
8. The end-to-end cloud registration system based on the Transformer model according to claim 6, characterized in that The multi-head attention mechanism is a 4-head attention mechanism.
9. The end-to-end cloud registration system based on the Transformer model according to claim 8, characterized in that The corresponding position coordinates of the prediction are expressed as: in, and Respectively represent the positions of the predicted source point cloud and target point cloud, and Respectively represent the positions of the actual source point cloud and target point cloud, express and The point set obtained after splicing is express and The point set obtained after splicing is represents the indicator matrix.
10. The end-to-end cloud registration system based on the Transformer model according to claim 9, characterized in that The point cloud rigid transformation matrix is expressed as: Among them, R and t represent the rotation matrix and translation matrix respectively, and Represent the optimal rotation matrix and translation vector respectively, M' and N' represent the super point set of the source point cloud and the super point set of the target point cloud after feature extraction respectively, and Respectively and The i-th row of .
Citation Information
Patent Citations
Three-dimensional point cloud registration method based on feature interaction and reliable corresponding relation estimation
CN116128944A
End-to-end point cloud registration method based on multi-scale fusion and mixed position coding
CN117058203A
Transform-based low-overlap point cloud registration method
CN118096843A
WGAN-based unsupervised multi-view three-dimensional point cloud joint registration method
WO2022165876A1
Neural network techniques for appliance creation in digital oral care
WO2024127315A1