Point cloud registration method and system
The soft points are generated through the Transformer model and the rotation translation parameters are directly estimated, which solves the problems of many RANSAC iterations and insufficient feature descriptions in the prior art, and efficient and accurate point cloud registration is achieved.
Patent Information
- Application Number
- CN202510614985.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The prior art has large time consumption due to the large number of iterations of RANSAC in point cloud registration, and insufficient local feature description leads to many feature matching errors, which is difficult to effectively solve.
The Transformer model is used to process self-attention and cross-attention on the characteristics of the source point cloud and the target point cloud, and generate k-group source point cloud soft points and target point cloud soft points. The rotation translation parameters are directly estimated through the Kabsch-Umeyama algorithm to avoid feature matching and external point elimination processes.
It improves the efficiency of point cloud registration, reduces time consumption, enhances the corresponding robustness in low overlap areas, avoids feature matching errors, and achieves fast and accurate point cloud registration.
Smart Images

Figure CN120125629A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional vision, and particularly to a point cloud registration method and system. Background Art
[0002] Point cloud registration is an important 3D vision task, and is often applied to three-dimensional reconstruction, robot pose estimation, and remote sensing observation. Point cloud registration aims to find the optimal pose to align two point clouds to obtain a complete object or scene. To estimate the pose transformation parameters between two point clouds, the common approach is to select key points, then describe the features of the local regions where the key points are located, perform feature matching based on the feature similarity of the regions where the key points of the scene point cloud and the target point cloud are located to obtain a set of corresponding points, and estimate the rotation matrix and translation vector based on the corresponding relationship between the corresponding points.
[0003] In the prior art, some methods for describing the features at key points have been proposed, such as methods like PFH, FPFH, SHOT, RoPS, Spin Image, etc. The features extracted by these methods have a certain rigid transformation invariance and perform well on small datasets. However, since the features extracted by these methods are all manually designed low-level features, there are problems such as poor robustness and insufficient description ability. In recent years, more and more large-scale labeled 3D datasets can be obtained, and methods based on deep neural networks (3DMatch, CGF, PerfectMatch, FCGF, DIP, D3Feat, SpinNet, etc.) have achieved certain results in learning local feature descriptions. However, in practice, the source point cloud and the target point cloud usually have only a small overlapping area, and the geometric structures at the corresponding points obtained through feature matching are not necessarily exactly the same, which results in a large number of incorrect corresponding points in explicit local feature matching.
[0004] A large number of incorrect corresponding points lead to a waste of a large amount of time in estimating the rotation and translation parameters. Therefore, it is necessary to eliminate the existing incorrect corresponding points through other methods. RANSAC (RAndom SAmple Consensus) is the most commonly used outlier elimination method, which finds the optimal transformation parameters through multiple iterations. In each iteration, RANSAC randomly selects a part of the corresponding points to estimate the parameters, and finally selects the parameters with the highest inlier rate among multiple groups of parameters as the final estimation result. Currently, some more excellent outlier elimination methods have also emerged, such as: first constructing a graph for the corresponding points, and then finding the maximum clique in the graph to eliminate incorrect correspondences; reducing the input correspondences to a smaller number but with a higher inlier rate based on the boundary rule, and the time complexity of this method is the square of the number of point correspondences. Although these methods can effectively eliminate incorrect correspondences, they all consume a large amount of time.
[0005] To avoid excessive time consumption caused by outlier rejection, some RANSAC-free registration methods have emerged recently. They are usually based on implicit feature matching. DCP, RPMNet, CoFiNet, and GeoTransformer are typical implicit feature matching methods. They usually construct a matching matrix based on feature similarity and optimal transport theory. This matrix does not describe one-to-one correspondence but many-to-many correspondence. The rotation and translation parameters are calculated by weighted SVD, and the weights are provided by the matching matrix. Compared with explicit feature matching methods, such methods improve the robustness of feature matching and can avoid the RANSAC process, realizing an end-to-end registration algorithm. However, the construction of the matching matrix still depends on the similarity of local features, which may still lead to incorrect correspondences due to inconsistent local geometric structures at corresponding points in some registration tasks. Both explicit and implicit feature matching methods have certain limitations in registration tasks with a small overlapping area.
[0006] Therefore, how to avoid the large amount of time consumption caused by the large number of iterations required by RANSAC and solve the problem of a large number of incorrect matches generated during the feature matching process due to insufficient local feature description is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0007] In view of this, the present invention provides a point cloud registration method and system to avoid the large amount of time consumption caused by the large number of iterations required by RANSAC and solve the problem of a large number of incorrect matches generated during the feature matching process due to insufficient local feature description.
[0008] One aspect of the present invention provides a point cloud registration method applied to a point cloud registration system. The point cloud registration system includes a feature extraction and interaction module, a soft point generation module, and a rotation and translation parameter estimation module. The method includes: Step S1, the feature extraction and interaction module takes the source point cloud and the target point cloud as inputs, linearly projects the source point cloud and the target point cloud into source point cloud features and target point cloud features respectively, and then transports the source point cloud features and the target point cloud features into a Transformer model. The Transformer model includes 12 Transformer layers. The first 6 Transformer layers use self-attention layers, and the last 6 Transformer layers use cross-attention layers. The source point cloud features and the target point cloud features are processed through the self-attention layer and the cross-attention layer to obtain the interacted source point cloud features and the interacted target point cloud features; Step S2, in the soft point generation module, all points in the interacted source point cloud features are weighted according to k a set of source point cloud weights to obtain kA source point cloud soft point, which is obtained by weighting all points in the target point cloud features after interaction according to k a set of target point cloud weights k to obtain k a set of source point cloud weights is predicted from the source point cloud features after interaction, k and a set of target point cloud weights is predicted from the target point cloud features after interaction; Step S3, in the rotation and translation parameter estimation module, the Kabsch-Umeyama algorithm is used to directly estimate the rotation parameters and translation parameters from k the source point cloud soft points and k the target point cloud soft points through singular value decomposition, and then point cloud registration is performed based on the rotation parameters and translation parameters.
[0009] Another aspect of the present invention provides a point cloud registration system, including a feature extraction and interaction module, a soft point generation module, and a rotation and translation parameter estimation module; The feature extraction and interaction module takes the source point cloud and the target point cloud as inputs, linearly projects the source point cloud and the target point cloud into source point cloud features and target point cloud features respectively, and then sends the source point cloud features and the target point cloud features to the Transformer model. The Transformer model includes 12 Transformer layers. The first 6 Transformer layers use self-attention layers, and the last 6 Transformer layers use cross-attention layers. The source point cloud features and the target point cloud features are processed through the self-attention layers and the cross-attention layers to obtain the source point cloud features after interaction and the target point cloud features after interaction; In the soft point generation module, all points in the source point cloud features after interaction are weighted according to k a set of source point cloud weights k to obtain k a set of target point cloud weights k to obtain k a set of source point cloud weights is predicted from the source point cloud features after interaction, k and a set of target point cloud weights is predicted from the target point cloud features after interaction; In the rotation and translation parameter estimation module, the Kabsch-Umeyama algorithm is used to directly estimate the rotation parameters and translation parameters from k the source point cloud soft points and k the target point cloud soft points through singular value decomposition, and then point cloud registration is performed based on the rotation parameters and translation parameters.
[0010] According to the point cloud registration method and system provided by the present invention, the source point cloud features and target point cloud features are processed through the self-attention layer and the cross-attention layer to obtain the interacted source point cloud features and the interacted target point cloud features. Instead of relying on the similarity of local features when finding corresponding points, the present invention generates k source point cloud soft points and k target point cloud soft points for the source point cloud and the target point cloud respectively, which can form k groups of corresponding soft points, and the number of generated soft points is less than that of the point cloud to be registered. Subsequently, the rotation parameters and translation parameters can be directly estimated from k groups of corresponding soft points. Different from the existing methods, the present invention does not establish correspondences based on explicit or implicit feature matching at all, but automatically generates k groups of soft points to form correspondences, which makes the present invention have no feature matching process and avoids a large number of incorrect matching problems caused by feature matching in low-overlap regions. In addition, the present invention does not require an outlier rejection process, which can effectively improve the running speed and shorten the point cloud registration time. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a flowchart of the point cloud registration method provided by an embodiment of the present invention; Figure 2 is an exemplary registration effect diagram; Figure 3 is a structural block diagram of the point cloud registration system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the embodiments of the present invention, but should not be construed as limiting the present invention.
[0013] Please refer to Figure 1 , an embodiment of the present invention provides a point cloud registration method, which is applied to a point cloud registration system. The point cloud registration system includes a feature extraction and interaction module, a soft point generation module, and a rotation and translation parameter estimation module. The core of this method is soft point registration, and the point cloud registration system focuses on finding k groups of reliable corresponding points for quickly estimating the rotation and translation parameters. Since both explicit and implicit feature matching are based on the similarity of local features, but in low-overlap scenarios, the local geometric structures at the corresponding points are not exactly the same, the present invention does not rely on the similarity of local features when finding corresponding points, but uses the attention mechanism to generate k soft points for the source point cloud and the target point cloud respectively, forming ka set of corresponding points. Subsequently, the rotation parameters and translation parameters can be directly estimated from k the set of corresponding points.
[0014] Specifically, the method includes steps S1 to S3: Step S1, the feature extraction and interaction module takes the source point cloud and the target point cloud as inputs, linearly projects the source point cloud and the target point cloud into source point cloud features and target point cloud features respectively, and then transports the source point cloud features and the target point cloud features into the Transformer model. The Transformer model includes 12 Transformer layers. The first 6 Transformer layers use self-attention layers, and the last 6 Transformer layers use cross-attention layers. The source point cloud features and the target point cloud features are processed through the self-attention layers and the cross-attention layers to obtain the interacted source point cloud features and the interacted target point cloud features.
[0015] Among them, the feature extraction and interaction module takes the source point cloud and the target point cloud as inputs, , , represents the set of real numbers; is the total number of points in the source point cloud and also the total number of points in the target point cloud. The self-attention layers and the cross-attention layers in the Transformer are used to extract and interact the features of the source point cloud and the target point cloud.
[0016] Among them, the self-attention layer is used to perceive the geometric structures of the source point cloud and the target point cloud and obtain point-wise features, and the cross-attention layer is used to let the features of the source point cloud interact with the features of the target point cloud so that they can perceive each other's geometric structures. Specifically, first and are linearly projected into d dimensional embeddings to obtain the source point cloud features and the target point cloud features , , , in the present invention, d is set to 256. Then, and After being added to the positional encoding, it is fed into the Transformer model. The Transformer model consists of 12 Transformer layers, and each Transformer layer contains two sub-layers: a multi-head attention layer and a point-wise feed-forward network. The self-attention layer is used in the first 6 Transformer layers, which allows points to interact with other points in the same point cloud, thereby perceiving the geometric structures of the source point cloud and the target point cloud respectively. The cross-attention layer is used in the last 6 Transformer layers, which is used to interact the features of the two point clouds, enabling them to exchange information, and thus perceiving the geometric structures of each other. The multi-head attention operation in each Transformer layer is defined as follows: ; ; where, represents the multi-head attention mechanism, , and are the query matrix, the key matrix, and the value matrix respectively, , , represent the first, the h th, and the H th heads respectively, represents concatenation along the channel dimension, represents the weight matrix that can be learned through the network, , , represent the weight matrices corresponding to the query matrix, the key matrix, and the value matrix corresponding to the h th head respectively.
[0017] , , . The present invention sets H = 8, d head = d / H = 32.
[0018] In traditional single-head attention, the dot product between and introduces a computational cost that grows quadratically with the length of the input sequence. In the case of a large number of points in the point cloud, directly applying the vanilla version of single-head attention is impractical. The present invention uses the attention layer proposed in the linear Transformer to reduce the computational complexity: ; represents the attention mechanism, is an improved activation function, , , is an activation function, represents transpose.
[0019] In the first 6 Transformer layers, the self-attention layer processes the source point cloud features and the target point cloud features as follows: ; ; where, is the source point cloud feature processed by the self-attention layer, is the target point cloud feature processed by the self-attention layer; In the last 6 Transformer layers, the cross-attention layer processes and as follows: ; ; where, is the source point cloud feature after interaction, is the target point cloud feature after interaction.
[0020] Step S2, in the soft point generation module, all points in the source point cloud feature after interaction are weighted by k groups of source point cloud weights to obtain k source point cloud soft points, and all points in the target point cloud feature after interaction are weighted by k groups of target point cloud weights to obtain k target point cloud soft points. k The groups of source point cloud weights are predicted from the source point cloud feature after interaction, k The groups of target point cloud weights are predicted from the target point cloud feature after interaction.
[0021] The present invention uses and to generate k soft points for the source point cloud and the target point cloud respectively. The th source point cloud soft point and the th target point cloud soft point are denoted as and , , . Each pair of and should have the ability to automatically find a corresponding local part of the source point cloud and the target point cloud. Taking Taking the calculation process of as an example (the calculation principle is the same). It is obtained by weighting all points in the source point cloud, and its weight is denoted as , and is predicted by . Specifically, is obtained by being mapped to 1D by MLP and passing through the softmax calculation. Since has been updated with information from , when assigning weights to each point in the source point cloud, it can automatically give a higher weight to the local area jointly concerned by and , and assign a low weight to points with inconsistent local geometric structures. In this embodiment, The calculation formula of is: ; Among them, represents the th soft point of the source point cloud, represents the weight of the th point in the source point cloud, represents the th point in the source point cloud, is the total number of points in the source point cloud.
[0022] In this embodiment, the weights of k groups and k soft points can be calculated simultaneously: ; ; Among them, represents k groups of source point cloud weights, , is the normalization function, represents the operation of reducing the feature dimension from d dimensions to k dimensions through a multi-layer perceptron, represents k soft points of the source point cloud, , represents the source point cloud.
[0023] It can be understood that k soft points of the target point cloud can be calculated in the same way.
[0024] Step S3, in the rotation and translation parameter estimation module, the Kabsch-Umeyama algorithm is used to obtain from k soft points of the source point cloud and kThe rotation parameters and translation parameters are directly estimated from the soft points of the target point cloud through singular value decomposition, and then point cloud registration is performed based on the rotation parameters and translation parameters.
[0025] Among them, the rotation and translation parameter estimation module satisfies the following formula: ; ; ; ; ; Wherein, represents k soft points of the target point cloud, represents the th soft point of the target point cloud, represents the centroid of, represents the centroid of, , , respectively represent the left singular vector, singular value, and right singular vector, represents the rotation parameter, represents the translation parameter.
[0026] In this embodiment, the rotation and translation error is used as the loss function to minimize the and L1 error between, is the result obtained by transforming the source point cloud according to the estimated rotation parameters and translation parameters, is the result obtained by transforming the source point cloud according to the true rotation parameters and true translation parameters.
[0027] Specifically, the loss function of the point cloud registration system is: ; Wherein, represents the result obtained by transforming the th point in the source point cloud according to the estimated rotation parameter and translation parameter , represents the result obtained by transforming the th point in the source point cloud according to the true rotation parameters and true translation parameters.
[0028] The method of the present invention is tested below. Commonly used evaluation indicators for point cloud registration effects include: Registration Recall (RR, registration recall), Relative Rotation Errors (RRE, relative rotation error), and Relative Translation Errors (RTE, relative translation error). The present invention selects the dataset 3DloMatch with a lower difficulty to compare the method of the present invention with other methods.The specific results are shown in Table 1. In Table 1, 3DSN (The Perfect Match: 3D Point Cloud Matching with Smoothed Densities) is a point cloud matching method based on smoothed densities; FCGF (Feature Correspondence Graph for Fast Point Cloud Registration) is a point cloud matching method based on the feature correspondence graph for fast point cloud registration; D3Feat (Joint Learning of Dense Detection and Description of 3D Local Features) is a point cloud matching method based on the joint learning of dense detection and description of 3D local features; Predator (Registration of 3D Point Clouds with Low Overlap) is a 3D point cloud registration method under low overlap, with two different versions, Predator-5k and Predator-1k, where 1k and 5k represent the number of key points; OMNet (Learning Overlapping Mask for Partial-to-Partial Point Cloud Registration) is an overlapping mask learning method for partial point cloud to partial point cloud registration; DGR (Deep Global Registration) is a deep global registration method; PCAM (Product of Cross-Attention Matrices for Rigid Registration of Point Clouds) is a point cloud rigid registration method based on the product of cross-attention matrices; REGTR (End-to-end Point Cloud Correspondences with Transformers) is a registration method based on end-to-end point cloud correspondences with Transformers; Geo Transformer (Fast and robust point cloud registration with geometric transformer) is a fast and robust point cloud registration method based on Transformers.
[0029] Table 1 Comparison of effects on the dataset 3DloMatch
[0030] As can be seen from Table 1, on 3DLoMatch, the method of the present invention has the best performance in all three metrics of RR, RRE, and RTE. Since the overlap degree of point cloud pairs in the 3DLoMatch dataset is very low, less than 30%, many corresponding point pairs in the overlapping area do not have similar geometric structures, resulting in incorrect point correspondences unable to be generated by feature matching. The method of the present invention does not find point correspondences based on feature matching, but forms correspondences with soft points generated by the weights automatically assigned by the attention mechanism. The weights automatically assigned by the network can automatically focus on similar local geometric structures and filter out inconsistent geometric structures, which greatly improves the robustness of correspondences in low-overlap areas and makes the method of the present invention more prominent in low-overlap registration datasets.
[0031] Table 2 shows the comparison of the time required for the present invention and other methods to complete the registration of two point clouds. In Table 2, SpinNet (Learning a General Surface Descriptor for 3D Point Cloud Registration) is a 3D point cloud registration method based on learning a general surface descriptor.
[0032] Table 2 Time required for different methods to register two point clouds
[0033] As can be seen from Table 2, compared with the method using RANSAC, the method of the present invention can significantly speed up the running speed. Since the point correspondences formed by soft points are reliable enough, the method of the present invention does not rely on RANSAC to eliminate incorrect correspondences, but directly estimates the transformation parameters according to the soft point correspondences. Therefore, the method of the present invention belongs to a method without RANSAC and has a faster running speed.
[0034] In addition, the method of the present invention is an end-to-end registration method, which only needs to input the source point cloud and the target point cloud, and finally obtains the rotation and translation transformation from the source point cloud to the target point cloud. As Figure 2 shown, only the source point cloud and the target point cloud with an overlapping area need to be input, and finally the registered point cloud is obtained.
[0035] In summary, according to the point cloud registration method of the above embodiment, the source point cloud features and the target point cloud features are processed through the self-attention layer and the cross-attention layer to obtain the source point cloud features after interaction and the target point cloud features after interaction. The present invention no longer bases on the similarity of local features when finding corresponding points, but respectively generates k source point cloud soft points and k target point cloud soft points, and can form ka set of corresponding soft points, and the number of generated soft points is less than the point cloud to be registered. Subsequently, the rotation parameters and translation parameters can be directly estimated from k the corresponding soft points in the set. Different from the existing methods, the present invention does not establish correspondences based on explicit or implicit feature matching at all, but automatically generates k a set of soft points to form correspondences, which makes the present invention have no feature matching process and avoids a large number of incorrect matching problems caused by feature matching in low-overlap regions. In addition, the present invention does not require an outlier rejection process, which can effectively improve the running speed and shorten the point cloud registration time.
[0036] Please refer to Figure 3 , another embodiment of the present invention provides a point cloud registration system, including a feature extraction and interaction module, a soft point generation module, and a rotation and translation parameter estimation module; The feature extraction and interaction module takes the source point cloud and the target point cloud as inputs, linearly projects the source point cloud and the target point cloud into source point cloud features and target point cloud features respectively, and then transports the source point cloud features and the target point cloud features into the Transformer model. The Transformer model includes 12 Transformer layers. The first 6 Transformer layers use self-attention layers, and the last 6 Transformer layers use cross-attention layers. The source point cloud features and the target point cloud features are processed through the self-attention layers and the cross-attention layers to obtain the interacted source point cloud features and the interacted target point cloud features; In the soft point generation module, all points in the interacted source point cloud features are weighted according to k a set of source point cloud weights to obtain k a number of source point cloud soft points, and all points in the interacted target point cloud features are weighted according to k a set of target point cloud weights to obtain k a number of target point cloud soft points, k the set of source point cloud weights is predicted from the interacted source point cloud features, k the set of target point cloud weights is predicted from the interacted target point cloud features; In the rotation and translation parameter estimation module, the Kabsch-Umeyama algorithm is used to directly estimate the rotation parameters and translation parameters from k a number of source point cloud soft points and k a number of target point cloud soft points through singular value decomposition, and then point cloud registration is performed based on the rotation parameters and translation parameters.
[0037] In this embodiment, the Transformer model satisfies the following formula: ; ; ; Among them, represents the multi-head attention mechanism, , and are the query matrix, the key matrix, and the content matrix respectively, , , respectively represent the first, the h th, and the H th heads, represents concatenation along the channel dimension, represents the weight matrix that can be learned through the network, , , respectively represent the weight matrices corresponding to the query matrix, the key matrix, and the content matrix corresponding to the h th head, represents the attention mechanism, is the improved activation function, , , is the activation function, represents transpose; In the first 6 Transformer layers, the self-attention layer processes the source point cloud feature and the target point cloud feature as follows: ; ; Among them, is the source point cloud feature processed by the self-attention layer, is the target point cloud feature processed by the self-attention layer; In the last 6 Transformer layers, the cross-attention layer processes and as follows: ; ; Among them, is the source point cloud feature after interaction, is the target point cloud feature after interaction.
[0038] In this embodiment, the soft point generation module satisfies the following formula: ; Among them, represents the th source point cloud soft point, represents the weight of the th point in the source point cloud, denotes the th point in the source point cloud, and
[0039] d is the total number of points in the source point cloud. In this embodiment, the feature dimensions of the source point cloud feature and the target point cloud feature are both ; ; where denotes k groups of source point cloud weights, is a normalization function, denotes the operation of reducing the feature dimension from d dimensions to k dimensions through a multi-layer perceptron, denotes k source point cloud soft points, denotes the source point cloud.
[0040] In this embodiment, the rotation and translation parameter estimation module satisfies the following formula: ; ; ; ; ; where denotes k target point cloud soft points, denotes the th target point cloud soft point, denotes centroid of denotes centroid of , , denote the left singular vector, singular value, and right singular vector respectively, denotes the rotation parameter, denotes the translation parameter.
[0041] In this embodiment, the loss function of the point cloud registration system is: ; where denotes the th point in the source point cloud according to the estimated rotation parameter And the translation parameter The result obtained after transformation Indicates the source point cloud The th point in the source point cloud after transformation according to the true values of the rotation parameter and the translation parameter.
[0042] In summary, according to the point cloud registration system of the above embodiment, the source point cloud features and the target point cloud features are processed through the self-attention layer and the cross-attention layer to obtain the source point cloud features after interaction and the target point cloud features after interaction. In the present invention, when searching for corresponding points, it is no longer based on the similarity of local features, but respectively generates k source point cloud soft points and k target point cloud soft points, and can form k groups of corresponding soft points, and the number of generated soft points is less than that of the point cloud to be registered. Subsequently, the rotation parameter and the translation parameter can be directly estimated from k groups of corresponding soft points. Different from the existing methods, the present invention does not establish correspondence based on explicit or implicit feature matching at all, but automatically generates k groups of soft points to form correspondence, which makes the present invention have no feature matching process and avoids a large number of incorrect matching problems caused by feature matching in low-overlap regions. In addition, the present invention does not require an outlier rejection process, can effectively improve the running speed, and shorten the point cloud registration time.
[0043] The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. A point cloud registration method, characterized in that: Applied to a point cloud registration system, the point cloud registration system includes a feature extraction and interaction module, a soft point generation module, and a rotation and translation parameter estimation module, and the method includes: Step S1, the feature extraction and interaction module takes the source point cloud and the target point cloud as input, linearly projects the source point cloud and the target point cloud into source point cloud features and target point cloud features respectively, and then transmits the source point cloud features and the target point cloud features to the Transformer model. The Transformer model includes 12 Transformer layers, the first 6 Transformer layers use self-attention layers, and the last 6 Transformer layers use cross-attention layers. The source point cloud features and the target point cloud features are processed by the self-attention layers and the cross-attention layers to obtain the source point cloud features after interaction and the target point cloud features after interaction; Step S2, in the soft point generation module, all points in the source point cloud features after interaction are generated according to k The group source point cloud weights are weighted to obtain k source point cloud soft points, all points in the target point cloud features after interaction are based on k The group target point cloud weights are weighted to obtain k Target point cloud soft points, k The group source point cloud weight is predicted by the interactive source point cloud features. k The weight of the group target point cloud is predicted by the target point cloud features after interaction; Step S3, in the rotation and translation parameter estimation module, the Kabsch-Umeyama algorithm is used to obtain k Source point cloud soft points and k The rotation parameters and translation parameters are directly estimated from the soft points of the target point cloud by singular value decomposition, and then the point cloud registration is performed based on the rotation parameters and translation parameters.
2. The point cloud registration method according to claim 1, characterized in that: The Transformer model satisfies the following formula: ; ; ; in, represents the multi-head attention mechanism, , and They are query matrix, keyword matrix and content matrix respectively. , , Respectively represent the first and h , H Head, represents splicing along the channel dimension, represents the weight matrix that can be learned through the network, , , Respectively represent h The weight matrix corresponding to the query matrix, keyword matrix, and content matrix of each head, represents the attention mechanism, is an improved activation function, , , is the activation function, represents transpose; In the first 6 Transformer layers, the self-attention layer performs the following calculation on the source point cloud features: and target point cloud features To process: ; ; in, is the source point cloud feature after processing by the self-attention layer, It is the target point cloud feature after processing by the self-attention layer; In the last 6 Transformer layers, the cross attention layer is used as follows and To process: ; ; in, is the source point cloud feature after interaction, is the target point cloud feature after interaction.
3. The point cloud registration method according to claim 2, characterized in that: The soft point generation module satisfies the following formula: ; in, Indicates Source point cloud soft points, Indicates the source point cloud The weight of the point, Represents the first Points, is the total number of points in the source point cloud.
4. The point cloud registration method according to claim 3, characterized in that: The feature dimensions of the source point cloud features and the target point cloud features are d dimension; The soft point generation module also satisfies the following formula: ; ; in, express k Group source point cloud weights, is the normalization function, It means that the feature dimension is transformed from d Dimensionality reduction k Dimensional operations, express k Source point cloud soft points, Represents the source point cloud.
5. The point cloud registration method according to claim 4, characterized in that: The rotation and translation parameter estimation module satisfies the following formula: ; ; ; ; ; in, express k Target point cloud soft points, Indicates Target point cloud soft points, express The center of mass, express The center of mass, , , They represent left singular vectors, singular values, and right singular vectors, respectively. represents the rotation parameter, Represents the translation parameter.
6. The point cloud registration method according to claim 5, characterized in that: The loss function of the point cloud registration system for: ; in, Represents the source point cloud The The points are rotated according to the estimated rotation parameters and translation parameters The result after transformation is Represents the source point cloud The The result obtained by transforming the points according to the true value of the rotation parameter and the true value of the translation parameter.
7. A point cloud registration system, characterized in that: It includes feature extraction and interaction module, soft spot generation module, rotation and translation parameter estimation module; The feature extraction and interaction module takes the source point cloud and the target point cloud as input, linearly projects the source point cloud and the target point cloud into source point cloud features and target point cloud features respectively, and then transmits the source point cloud features and the target point cloud features to the Transformer model. The Transformer model includes 12 Transformer layers. The first 6 Transformer layers use self-attention layers, and the last 6 Transformer layers use cross-attention layers. The source point cloud features and the target point cloud features are processed through the self-attention layers and cross-attention layers to obtain the source point cloud features after interaction and the target point cloud features after interaction. In the soft point generation module, all points in the source point cloud feature after interaction are k The group source point cloud weights are weighted to obtain k source point cloud soft points, all points in the target point cloud features after interaction are based on k The weight of the target point cloud group is weighted to obtain k Target point cloud soft points, k The group source point cloud weight is predicted by the interactive source point cloud features. k The weight of the group target point cloud is predicted by the target point cloud features after interaction; In the rotation and translation parameter estimation module, the Kabsch-Umeyama algorithm is used to obtain k Source point cloud soft points and k The rotation parameters and translation parameters are directly estimated from the soft points of the target point cloud by singular value decomposition, and then the point cloud registration is performed based on the rotation parameters and translation parameters.
8. The point cloud registration system according to claim 7, characterized in that: The Transformer model satisfies the following formula: ; ; ; in, represents the multi-head attention mechanism, , and They are query matrix, keyword matrix and content matrix respectively. , , Respectively represent the first and h , H Head, represents splicing along the channel dimension, represents the weight matrix that can be learned through the network, , , Respectively represent h The weight matrix corresponding to the query matrix, keyword matrix, and content matrix of each head, represents the attention mechanism, is an improved activation function, , , is the activation function, represents transpose; In the first 6 Transformer layers, the self-attention layer performs the following calculation on the source point cloud features: and target point cloud features To process: ; ; in, is the source point cloud feature after processing by the self-attention layer, It is the target point cloud feature after processing by the self-attention layer; In the last 6 Transformer layers, the cross attention layer is used as follows and To process: ; ; in, is the source point cloud feature after interaction, is the target point cloud feature after interaction.
9. The point cloud registration system according to claim 8, characterized in that: The soft point generation module satisfies the following formula: ; in, Indicates Source point cloud soft points, Indicates the source point cloud The weight of the point, Represents the first Points, is the total number of points in the source point cloud.
10. The point cloud registration system according to claim 9, characterized in that: The feature dimensions of the source point cloud features and the target point cloud features are d dimension; The soft point generation module also satisfies the following formula: ; ; in, express k Group source point cloud weights, is the normalization function, It means that the feature dimension is transformed from d Dimensionality reduction k Dimensional operations, express k Source point cloud soft points, Represents the source point cloud.
Citation Information
Patent Citations
End-to-end three-dimensional point cloud registration method
CN114332176A
Laser point cloud registration method and system based on virtual super point
CN115661218A
Position-enhanced attention mechanism-based point cloud registration method
CN116912296A
Mask learning-based partially overlapped point cloud registration method
CN117455965A
FPP-based high-precision point cloud registration method guided by local-to-global structure
CN117475170A