Coarse-to-fine point cloud automatic registration method and system for overlapping mask optimization
By introducing an automatic point cloud registration method optimized by overlap mask, the problem of inaccurate point correspondence between point cloud registration methods in the existing technology in the coarse-scale stage in complex indoor scenarios is solved, and high-precision and robust point cloud registration effect is achieved, which is suitable for virtual reality, cultural heritage protection and reverse engineering and other fields.
Patent Information
- Application Number
- CN202510741535.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
When facing complex indoor three-dimensional scenarios, the existing point cloud registration method based on deep learning is insufficient in the accuracy and robustness of the point correspondence relationship in the rough-scale stage, resulting in low reliability and accuracy of the point correspondence relationship, making it difficult to meet the requirements of high-quality point cloud data registration.
An automatic point cloud registration method for overlap mask optimization from coarse to fine is introduced. Superpoint visual features are extracted through the KPConv encoder, feature expression is enhanced with the position perception attention module, non-overlapping areas are eliminated using the overlap mask, and the similarity matrix is optimized with the sample deviation correction loss function. Finally, the rigid transformation parameters are estimated through the RANSAC algorithm to realize automatic registration of indoor scene point clouds.
The accuracy of over-point matching and the quality of point correspondence in the rough-scale stage are improved, the accuracy and robustness of point cloud registration are improved, and a reliable data foundation is provided for subsequent applications.
Smart Images

Figure CN120259393A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and three-dimensional scene reconstruction, and particularly relates to a method and system for automatically registering point clouds from coarse to fine with overlapping mask optimization. Background Technique
[0002] With the rapid development of lidar scanning technology, the acquisition of high-resolution and high-precision point cloud data has become more efficient, and three-dimensional point clouds have gradually become a key data form in three-dimensional applications. However, due to factors such as environmental occlusion and sensor viewing angle limitations, the single-view data acquisition process usually only obtains partial point cloud data of the target scene. Point cloud registration aligns partial point cloud data located in different coordinate systems to the same coordinate system by solving rigid transformation parameters, so as to obtain a complete three-dimensional scene, and has been widely used in fields such as three-dimensional reconstruction, autonomous driving, virtual reality, and simultaneous localization and mapping (SLAM).
[0003] In point cloud registration, establishing a reliable point correspondence is a prerequisite for solving rigid transformation parameters. In traditional point cloud registration algorithms, the ways of establishing point correspondences can be roughly divided into two categories. One category is methods mainly represented by the classic ICP algorithm (Iterative Closest Point) and its related variants (such as point-to-plane ICP, GICP (Generalized-ICP), VGICP (Voxelized Generalized-ICP), etc.). These methods use the spatial distance between point clouds and establish point correspondences through an iterative manner. The other category establishes point correspondences through artificially designed three-dimensional feature descriptors (such as PFH, FPFH, SHOT, etc.). These methods model the local geometric information of point clouds, generate three-dimensional feature descriptors with rotational invariance, and use these descriptors to find homologous points with the same features to establish point correspondences. Although the above traditional registration methods show good generality and efficiency in the registration task, when facing actual scenes with a large point cloud scale, high repeatability, or strong symmetry, the established point correspondences are not reliable enough, cannot provide effective help for subsequent parameter estimation, and it is difficult to balance registration efficiency and accuracy.
[0004] In recent years, due to its powerful feature expression and learning ability, deep learning has gradually gained attention in the field of point cloud registration. The point cloud registration method based on deep learning learns the semantic features of the point cloud through a neural network, establishes the corresponding relationship between the source point cloud and the target point cloud based on this, and then combines a robust pose estimation algorithm (such as RANSAC) to calculate the rigid transformation parameters. Currently, most of the point cloud registration methods based on deep learning rely on key point detection to construct point correspondences. Key point detection samples points with strong geometric features to avoid ambiguity during matching, thereby improving the registration accuracy. Although key point detection can reduce the interference of redundant data and lower the computational cost in the registration tasks of real scenes, when the overlap degree between two point clouds is low, key point detection often fails to effectively extract key points with significant geometric characteristics, resulting in a low inlier ratio of the constructed point correspondences and being unable to provide effective support for parameter estimation. To address the limitations in constructing point correspondences, some studies have proposed a coarse-to-fine matching strategy. By dividing the construction of point correspondences into two stages: a coarse scale and a fine scale, the registration accuracy and robustness are gradually improved.
[0005] Through the above analysis, the problems and defects of the existing technology are as follows: In recent years, many three-dimensional point cloud registration algorithms based on deep learning have been proposed and have shown stronger adaptability in complex indoor three-dimensional scene matching using a coarse-to-fine strategy. However, these methods rely to a large extent on the accuracy of superpoint correspondences in the coarse scale stage. The existence of repetitive and similar structures in real indoor three-dimensional scenes makes these methods still have problems such as poor robustness and low accuracy. Therefore, on the basis of fully mining the local features and global information of the three-dimensional point cloud, introducing an overlap mask into the coarse-to-fine correspondence construction mechanism to improve the accuracy of superpoint correspondences in the coarse scale stage, and then establishing reliable point feature correspondences to improve the quality of automatic registration of point cloud data in indoor scenes has practical significance. Summary of the Invention
[0006] To overcome the problems in the related technology, the disclosed embodiments of the present invention provide an overlap mask optimized coarse-to-fine point cloud automatic registration method and system.
[0007] The technical solution is as follows: The overlap mask optimized coarse-to-fine point cloud automatic registration method includes the following steps: S1, obtain the three-dimensional point cloud data of the indoor scene. In the KPConv encoding layer, through the kernel point-guided local feature aggregation mechanism and downsampling operation, extract the superpoint visual features at the coarse scale; S2. Introduce a position-aware attention module to inject position information into the superpoint features, providing geometric constraints for the subsequent attention mechanism. Through self-attention and cross-attention operations, adaptively fuse the local geometric features and global context information of the superpoints, thereby effectively enhancing the expressive power of the superpoint features. S3. Input the enhanced superpoint features into the superpoint overlap prediction network, generate an overlap mask through multiple layers of convolution, and eliminate non-overlapping regions. On this basis, construct a similarity matrix using the remaining overlapping superpoints, and optimize it in combination with the sample bias correction loss function to establish superpoint correspondence relationships at the coarse scale. S4. In the decoding layer of the KPConv network, obtain the fine-scale point cloud visual features through nearest neighbor interpolation and upsampling operations, and then combine the nearest neighbor superpoint grouping strategy to extend the superpoint correspondence to the local point cloud block correspondence. S5. Under the guidance of the density adaptive mechanism, construct a fine-scale point correspondence set based on the matching relationship between local point cloud blocks. On this basis, estimate the rigid transformation parameters between the source point cloud and the target point cloud through the RANSAC algorithm to achieve automatic registration of indoor scene point clouds.
[0008] In step S1, the extracting of the coarse-scale superpoint visual features through the kernel point-guided local feature aggregation mechanism and downsampling operation in the KPConv encoding layer includes: Using the encoding layer of the KPConv network to downsample the source point cloud superpoints and target point cloud superpoints of the obtained indoor scene to the uniformly distributed downsampled source point cloud superpoints and downsampled target point cloud superpoints , and learn the visual features between the downsampled source point cloud superpoints and downsampled target point cloud superpoints and ; Input the downsampled and into the position encoding network; wherein, are both the original input source point cloud superpoints and target point cloud superpoints, is the set of real numbers, is the number of point clouds included in, are the downsampled source point cloud superpoints and downsampled target point cloud superpoints; is the number of superpoints included in, is the dimension of each superpoint feature, is the visual feature between the downsampled downsampled source point cloud superpoints and downsampled target point cloud superpoints.
[0009] In step S2, the introduced position-aware attention module injects position information into the superpoint features, providing geometric constraints for the subsequent attention mechanism; through self-attention and cross-attention operations, it adaptively fuses the local geometric features and global context information of the superpoints, thereby effectively enhancing the expressive ability of the superpoint features, including: Calculate the centroids of the input source point cloud superpoints and target point cloud superpoints and ; eliminate the overall spatial offsets of the source point cloud and target point cloud through the centering operation; use the position encoding network to map the position information of the source point cloud and target point cloud to the high-dimensional feature space and fuse it with the original features to obtain the source point cloud superpoint features with position perception ability and the target point cloud superpoint features with position perception ability ; ; and are generated by the following expressions: (1) (2) In the formula, is the position encoding function, and " " is the feature fusion operation.
[0010] Input the superpoint features and embedded with position information into the attention mechanism, perform self-attention and cross-attention operations on the superpoint features to achieve the aggregation of global context information and enhance the expression of the features. The attention mechanism uses multi-head attention, and the multi-head attention includes: (3) (4) In the formula, is the output of the multi-head attention mechanism, is to concatenate multiple vectors by dimension, is the calculation result of the th attention head, is the output of the standard single-head attention mechanism, are the query, key, and value in the attention mechanism, is the linear transformation matrix, , is the projection matrix of each attention head, is the number of heads, is the dimension of the input feature, is the dimension to be processed by a single attention head; , each attention head adopts single - head dot - product attention, and the expression is: (5) In the formula, is a normalization function, which is used to convert the similarity between the query and the key into attention weights, and then weight the value ; is the key transpose; The target feature and the attention result are fused through a multi - layer perceptron MLP, and the expression is: (6) In the formula, is the fused feature, is the multi - layer perceptron, is to concatenate multiple vectors by dimension, is the original input feature, is the feature after being processed by the attention mechanism; When performing self - attention processing, is set to the features of the same point cloud, and is set to: and in the source point cloud and the target point cloud respectively; when performing cross - attention processing, is set to the features of different point clouds, and is set to: and .
[0011] In step S3, the enhanced super - point features are input into the super - point overlap prediction network, and overlapping masks are generated through multi - layer convolutions to remove non - overlapping regions; on this basis, a similarity matrix is constructed using the remaining overlapping super - points, and optimization is carried out in combination with the sample bias correction loss function, and the super - point correspondence relationship at the coarse scale is established based on this: The backbone network of the super - point matching module based on the overlapping mask consists of 4 one - dimensional convolutional layers and activation functions. For each layer, point - wise convolution operations are performed using the convolutional kernel, and then group normalization is carried out. Except for the last layer, the activation function is used for non - linear transformation after each convolutional layer; the super - point matching module based on the overlapping mask outputs the point - wise features for each super - point in the last layer, which are processed by formula (7): (7) In the formula, is the convolution operation in the network, is the super - point feature after being processed by the previous layer. When it represents the source point cloud super - point feature with position perception ability of the input , is the classification result of superpoints in the source point cloud output by the last layer of the network, including the probabilities of superpoints belonging to overlapping points and non-overlapping points; Perform the same operation on the superpoint features of the target point cloud with position perception ability to obtain the classification result ; For and process to generate an overlap mask for superpoints, and select the category of each point, including: (8) (9) In the formula, is operation, and are the overlap masks of the source point cloud and the target point cloud respectively, and each mask element takes a value of 0 or 1; For apply the masks and , only retain the features located in the overlapping area, perform mask screening, and obtain the superpoint features of the source point cloud after masking and the superpoint features of the target point cloud after masking ; Use and to construct a similarity matrix that matches the superpoint features of the target point cloud with position perception ability and the superpoint features of the target point cloud with position perception ability , to quantify the similarity relationship between superpoints, and combine the Sinkhorn algorithm to perform optimal transport iterative optimization on the similarity matrix ; The Sinkhorn algorithm performs regularization operations on the rows and columns of the similarity matrix , and continuously iteratively optimizes to obtain the best match between superpoints. Based on the optimized similarity matrix , generate a set of superpoint correspondence relationships in the coarse matching stage .
[0012] Furthermore, take It is divided into a confidence deviation area and a confidence matching area to distinguish the performance deviation of the model in prediction. Among them, the confidence deviation area reflects that there are significant differences between the model prediction results and the true matching, which are easily overlooked but have strong supervision value and are regarded as positive samples; while the confidence matching area corresponds to the area where the model prediction accuracy is relatively high and is regarded as negative samples. Since the number of positive samples is much less than that of negative samples in the actual matching process, it is necessary to design a sample correction loss function for sample imbalance to mitigate the training deviation caused by imbalance. The sample correction loss function includes the local point cloud true matching loss Ture Loss and the local point cloud matching deviation correction loss Correct Loss, which is expressed as: (10) In the formula, is the sample correction loss function, is the local point cloud true matching loss Ture Loss, is the local point cloud matching deviation correction loss Correct Loss; Both Ture Loss and Correct Loss are applicable to the local point cloud sample imbalance. Only using Ture Loss makes the local point cloud block matching unstable, and adding Correct Loss is used to assist the matching; 1 Loss is used to strengthen the model's ability to recognize true matching point pairs, which is defined as: (11) Where: (12) In the formula, , are all hyperparameters for controlling the sample weights, with the default value of , is the probability, is the node.
[0013] Correct Loss focuses on compensating and optimizing the areas where the model prediction deviates, which is defined as: (13) Where, is the confidence deviation value, is the confidence matching value, and the two jointly participate in the regularization adjustment of the loss.
[0014] In step S4, in the decoding layer of the KPConv network, through the nearest neighbor interpolation and upsampling operations, the fine-scale point cloud visual features are obtained, and then combined with the nearest neighbor superpoint grouping strategy, the superpoint correspondence is extended to the local point cloud block correspondence, including: Obtain the correspondence set of superpoints After that, through the nearest superpoint grouping strategy (hereinafter referred to as the point-to-node grouping strategy), the superpoint correspondence is extended to the correspondence between each pair of local point cloud patches, and the point cloud with a smaller scale in the two local point cloud patches is matched within a local range to extract point-level correspondences; for any superpoint , the grouping process of point-to-node can be implemented by formula (14): (14) In the formula, is the Euclidean distance, is each point cloud to be assigned, and are respectively the point set and visual features on the fine scale related to the superpoint after grouping, is the th superpoint to which it is to be assigned, is the th superpoint to which it is to be assigned, is each feature to be assigned.
[0015] In step S5, under the guidance of the density adaptive mechanism, a fine-scale point correspondence set is constructed based on the matching relationship between local point cloud patches; on this basis, the rigid transformation parameters between the source point cloud and the target point cloud are estimated by the RANSAC algorithm to achieve the automatic registration of indoor scene point clouds: The density adaptive mechanism is used to calculate the similarity matrix of each pair of local point cloud patches. The density adaptive mechanism marks the invalid points caused by repeated sampling as infinitesimal negative values, and iteratively eliminates the matches between invalid points through optimal transport to generate a point-level similarity matrix ; use as the confidence matrix of candidate matches, and the final point correspondences are obtained by selecting the point pairs with higher confidence scores ; the RANSAC algorithm estimates the rigid transformation parameters by randomly sampling the point correspondences in, and calculates the number of local inliers based on these estimated transformations. The rigid transformation parameters with the most local inliers are selected as the optimal solution to complete the automatic registration of indoor scene point clouds.
[0016] Another object of the present invention is to provide an overlapping mask optimized coarse-to-fine point cloud automatic registration system, which implements the overlapping mask optimized coarse-to-fine point cloud automatic registration method. The system includes: The visual feature acquisition module is used to acquire the three-dimensional point cloud data of the indoor scene. The three-dimensional point cloud data includes source point cloud superpoints and target point cloud superpoints. The source point cloud superpoints and target point cloud superpoints are respectively downsampled to downsampled source point cloud superpoints and downsampled target point cloud superpoints with uniform distribution, and the visual features between the downsampled source point cloud superpoints and downsampled target point cloud superpoints are learned. The visual feature acquisition module is used to acquire the three-dimensional point cloud data of the indoor scene. The three-dimensional point cloud data includes source point cloud and target point cloud. The source point cloud and target point cloud are respectively downsampled to superpoints with uniform distribution, and the related visual features are learned. The position-aware attention module is used to explicitly inject position information into the superpoint features, enabling the features to have geometric sensitivity while expressing semantics, and providing geometric constraints for the subsequent attention mechanism. Subsequently, the self-attention mechanism is used to mine the correlation between the internal features of the superpoints, strengthen the expression ability of the local geometry, and at the same time combine the cross-attention mechanism to realize the context feature interaction between different point clouds. Through this process, the local geometric structure and the global relationship are jointly modeled, thereby significantly improving the discriminability of the superpoint features.
[0017] The overlapping mask generation module is used to judge the overlap of the enhanced superpoint features, filter out the superpoints located in the non-overlapping regions, and generate a set of superpoint correspondence relationships in the coarse matching stage based on the remaining superpoints in the overlapping regions. The superpoint correspondence set refinement and rigid transformation parameter estimation module is used to refine the superpoint correspondence set into a point correspondence set by using the density adaptive mechanism. Based on the point correspondence set, the rigid transformation parameters between the source point cloud and the target point cloud are further estimated to complete the automatic registration of the indoor scene point cloud.
[0018] Combining all the above technical solutions, the beneficial effects of the present invention are as follows: The present invention aims to improve the establishment process of the superpoint correspondence relationship in the coarse scale stage of the coarse-to-fine point cloud automatic registration scheme, thereby improving the quality and robustness of the registration result. The present invention can effectively improve the accuracy of superpoint matching in the coarse scale stage, and at the same time improve the quality of the construction of point correspondence relationships, providing a reliable guarantee for the accuracy and robustness of point cloud registration. Introducing an overlapping mask in the coarse scale stage of the coarse-to-fine point cloud registration method effectively eliminates the influence of the superpoints located in the non-overlapping regions on the superpoint matching, improves the quality of the superpoint correspondence relationship, and thus can establish a more reliable point correspondence relationship, improving the accuracy and robustness of point cloud registration. Brief Description of the Drawings
[0019] The drawings here are incorporated into the specification and form a part of this specification, showing the embodiments in line with the present disclosure, and are used together with the specification to explain the principles of the present disclosure; Figure 1It is the flowchart of the coarse-to-fine point cloud automatic registration method for overlapping mask optimization provided by the embodiments of the present invention; Figure 2 It is the flowchart of the superpoint matching module based on overlapping mask provided by the embodiments of the present invention; Figure 3 It is the visualization effect diagram of the registration result on the indoor scene provided by the embodiments of the present invention. Detailed implementation manners
[0020] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings. Many specific details are set forth in the following description to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific implementations disclosed below.
[0021] The innovation of the present invention lies in: an overlapping mask is introduced in the coarse-scale superpoint correspondence construction stage of the coarse-to-fine point cloud registration method to guide the model to focus on the superpoints within the overlapping area at the coarse scale, so as to improve the quality of the superpoint correspondence relationship; at the same time, a sample correction function is used to further optimize the similarity matrix between superpoints, further improving the reliability of the superpoint correspondence. In addition, on the basis of fully integrating the position information and geometric features of the superpoints, the position-aware attention mechanism aggregates the superpoint context information under geometric constraints and enhances the features to efficiently generate the overlapping mask in a single inference manner.
[0022] The present invention aims to improve the establishment process of superpoint correspondence relationships in the coarse matching stage of the coarse-to-fine point cloud automatic registration scheme, thereby enhancing the quality and robustness of the registration results. The method includes the following steps: Step 1, obtain the three-dimensional point cloud data of the indoor scene, and use the KPConv network to downsample the three-dimensional point cloud data to uniformly distributed superpoints and learn their visual features; Step 2, introduce a position-aware attention module to embed position information in the superpoint features, provide geometric constraints for the attention mechanism, and adaptively fuse local geometric structures and global context information through self-attention and cross-attention operations to enhance the feature expression ability; Step 3, input the enhanced superpoint features into the overlap prediction network to generate an overlap mask to eliminate non-overlapping regions, construct a similarity matrix based on the retained superpoints, and introduce a sample correction loss function for optimization to complete the establishment of superpoint correspondence relationships at the coarse scale; Step 4, in the KPConv decoder, restore the fine-scale point features through nearest neighbor interpolation and upsampling, and combine the nearest neighbor superpoint grouping strategy to extend the matching relationship from superpoints to local point cloud patches; Step 5, under the guidance of the density adaptive mechanism, construct a fine-scale point correspondence set, and use the RANSAC algorithm to estimate the rigid transformation parameters to complete the automatic registration of the indoor scene point cloud. The present invention can effectively improve the accuracy of superpoint matching in the coarse matching stage, and at the same time improve the construction quality of point correspondence relationships, providing a reliable guarantee for the accuracy and robustness of point cloud registration.
[0023] Example 1, as Figure 1 shown, the method for optimizing the overlap mask for coarse-to-fine point cloud automatic registration provided by the embodiment of the present invention includes the following steps: S1, obtain the three-dimensional point cloud data of the indoor scene, and in the KPConv encoding layer, extract the coarse-scale superpoint visual features through the kernel point-guided local feature aggregation mechanism and downsampling operation; S2, introduce a position-aware attention module to inject position information into the superpoint features, providing geometric constraints for the subsequent attention mechanism; through self-attention and cross-attention operations, adaptively fuse the local geometric features and global context information of the superpoints, thereby effectively enhancing the expression ability of the superpoint features; S3, input the enhanced superpoint features into the superpoint overlap prediction network, generate an overlap mask through multiple layers of convolution and eliminate non-overlapping regions; on this basis, construct a similarity matrix using the retained overlapping superpoints, and optimize it in combination with the sample correction loss function, and establish superpoint correspondence relationships at the coarse scale based on this; S4, in the KPConv network decoding layer, obtain the fine-scale point cloud visual features through nearest neighbor interpolation and upsampling operations, and then combine the nearest neighbor superpoint grouping strategy to extend the superpoint correspondence to the local point cloud patch correspondence; S5. Under the guidance of the density adaptive mechanism, construct a fine-scale point correspondence set based on the matching relationship between local point cloud blocks. On this basis, estimate the rigid transformation parameters between the source point cloud and the target point cloud through the RANSAC algorithm to achieve automatic registration of indoor scene point clouds.
[0024] Embodiment 2. As another implementation manner of the present invention, the automatic point cloud registration system with overlapping mask optimization from coarse to fine provided by the embodiment of the present invention mainly includes the following: A visual feature acquisition module for acquiring three-dimensional point cloud data of an indoor scene, where the three-dimensional point cloud data includes a source point cloud and a target point cloud; downsample the source point cloud and the target point cloud to uniformly distributed superpoints respectively, and learn relevant visual features. A position-aware attention module for explicitly injecting position information into the superpoint features, enabling the features to have geometric sensitivity while expressing semantics, and providing geometric constraints for the subsequent attention mechanism; then, use the self-attention mechanism to mine the correlation between the internal features of the superpoints, strengthen the expression ability of the local geometry, and at the same time combine the cross-attention mechanism to realize the context feature interaction between different point clouds. Through this process, jointly model the local geometric structure and the global relationship, thereby significantly improving the discriminability of the superpoint features.
[0025] An overlapping mask generation module for performing overlapping judgment on the enhanced superpoint features, filtering out the superpoints located in the non-overlapping regions, and generating a set of superpoint correspondence relationships in the coarse matching stage based on the remaining superpoints in the overlapping regions. A superpoint correspondence refinement and rigid transformation parameter estimation module for refining the set of superpoint correspondence relationships into a set of point correspondence relationships by using the density adaptive mechanism; further estimating the rigid transformation parameters between the source point cloud and the target point cloud based on the set of point correspondence relationships to complete the automatic registration of the indoor scene point clouds.
[0026] Embodiment 3. As another implementation manner of the present invention, the automatic point cloud registration method with overlapping mask optimization from coarse to fine provided by the embodiment of the present invention further includes the following: Superpoint matching is sparse and loose, and the geodesic distance between two spatially close superpoints may be far. At the same time, superpoint matching has a higher dependence on global context information. Therefore, it is necessary to capture more comprehensive and deep feature representations to achieve effective matching. The attention mechanism can be used to mine the correlation between point features and improve the neural network's ability to aggregate global context information. However, the ordinary attention mechanism will ignore the geometric structure of the point cloud when capturing the context information of the point cloud and lacks the perception ability of the position information of the input point cloud, which will result in poor discriminability of the learned point cloud features in the geometric structure and easily form a large number of abnormal matches.
[0027] Therefore, the present invention designs a position-aware attention module, which adds geometric constraints to the attention mechanism by embedding position information in the point cloud features, helps the network model better understand the relationship between points, improves the network model's ability to aggregate global context information, and further enhances the features of the point cloud.
[0028] (1) Calculate the centroid of the input superpoint and , and eliminate the overall spatial offset of the point cloud through a centering operation; subsequently, use a position encoding network to map the point cloud position information to a high-dimensional feature space and fuse it with the original features to obtain the superpoint features and embedded with position information. and can be generated by the following expressions: (1) (2) In the formula, is the position encoding function, and " " is the feature fusion operation. and are the superpoint features embedded with position information.
[0029] (2) Input the superpoint features and embedded with position information into the attention mechanism, perform self-attention and cross-attention operations on the superpoint features to achieve the aggregation of global context information, and enhance the expression of features.
[0030] The attention mechanism is mainly implemented based on the multi-head attention mechanism module, and the information generated by the attention mechanism is deeply fused with the original features through a multi-layer perceptron (MLP) to further improve the feature expression ability. The present invention innovatively proposes the operation process of multi-head attention as follows: (3) (4) In the formula, is the output of the multi-head attention mechanism, is to concatenate multiple vectors by dimension, is the calculation result of the -th attention head, is the output of the standard single-head attention mechanism, are the query, key, and value in the attention mechanism, is the linear transformation matrix, , is the projection matrix for each attention head, is the number of heads, is the dimension of the input features, is the dimension to be processed by a single attention head; Set the number of heads H to 4, , is the dimension of the input features. The present invention innovatively proposes that each attention head adopts single-head dot product attention: (5) In the formula, is the normalization function, which is used to convert the similarity (dot product) of the query and the key into attention weights, and then weight the value for weighting, is the key transpose; The target feature and the attention result are fused through a multi-layer perceptron (MLP): (6) In the formula, is the fused feature, is the multi-layer perceptron, is to concatenate multiple vectors by dimension, is the original input feature, is the feature processed by the attention mechanism; The self-attention mechanism enables the point cloud to focus on other parts within the same point cloud, is set to the features of the same point cloud, and is set to in the source point cloud and the target point cloud respectively: and ; The cross-attention mechanism allows each point cloud to interact with the point clouds from other point cloud sets, so , , is set to the features of other point clouds, and is set to in the source point cloud and the target point cloud respectively: and .
[0031] After being processed by the attention mechanism, the model pays more attention to the local feature information of the superpoints, the information between the corresponding points in the source point cloud and the target point cloud is further interacted, and the point cloud features are further enhanced.
[0032] The superpoint matching module 2 based on the overlapping mask includes: In the actual point cloud registration process, due to environmental occlusion and partial data loss of the point cloud, not all superpoints are located within the overlapping region. Superpoints located in non-overlapping regions are difficult to form effective matches due to the lack of corresponding homologous points. In the coarse-scale stage, these non-overlapping superpoints will introduce interference in the calculation of the similarity matrix, easily leading to incorrect coarse correspondence relationships between superpoints located in the overlapping region, thereby reducing the quality of the refined point correspondence relationships and further affecting the registration accuracy and robustness. To address this issue, the present invention proposes a superpoint matching module based on an overlap mask (denoted as the overlap-mask module), which introduces an overlap mask into the mechanism for constructing point correspondence relationships from coarse to fine to improve the accuracy of superpoint correspondence relationships in the coarse-scale stage.
[0033] The superpoint matching module based on the overlap mask, i.e., the overlap-mask module, converts the prediction task of the overlapping region into a binary classification problem. By optimizing the classification objective, it accelerates the convergence of the model, fully utilizes the point-by-point features and local details of the superpoints, and efficiently generates an overlap mask for the superpoints without the need for iteration. Using the generated overlap mask, the interference of superpoints in non-overlapping regions on the matching process can be effectively eliminated, ensuring that only the effective superpoint features in the overlapping region participate in the matching.
[0034] Specifically, the overlap-mask backbone network consists of 4 one-dimensional convolutional layers and activation functions. The output feature dimension of each layer is defined by parameters. For each layer, a pointwise convolution operation is performed using a convolutional kernel of size 1, followed by optional group normalization (GN) to stabilize the training process. Except for the last layer, a activation function is used for non-linear transformation after each convolutional layer. The overlap-mask module outputs the classification of each superpoint in the last layer, and the output result contains the probabilities of the superpoint belonging to overlapping points and non-overlapping points. The point-by-point features of the superpoints are mainly processed by formula (7): (7) In the formula, represents the convolutional operation in the network, represents the superpoint features after being processed by the previous layer, and when it represents the original input features , is the classification result of the superpoints in the source point cloud output by the last layer of the network, which contains the probabilities of the superpoints belonging to overlapping points and non-overlapping points. The same operation is performed on the superpoint features in the target point cloud to obtain the classification result .
[0035] After that, for and are processed to generate an overlap mask for the superpoints, and the most likely class (overlap or non - overlap) of each point is selected. That is: (8) (9) In the formula, represents operation, and respectively represent the overlap region masks of the source point cloud and the target point cloud. Each mask element takes a value of 0 or 1.
[0036] For the original superpoint features , apply the masks and , only retain the features located within the overlap region, and denote the superpoint features filtered by the masks as and . Subsequently, use and to construct a similarity matrix of the superpoint features to quantify the similarity relationship between superpoints, and combine the Sinkhorn algorithm to perform optimal transport iterative optimization on the similarity matrix . Introduce a regularization mechanism in the Sinkhorn algorithm to ensure the efficiency and stability of the optimal transport process. Based on the optimized similarity matrix , generate the corresponding relationship set of the superpoints.
[0037] For the obtained similarity matrix of the superpoint features, divide into a confidence deviation region and a confidence matching region to distinguish the performance deviation of the model in prediction. Among them, the confidence deviation region reflects that there are significant differences between the model prediction results and the true matches, which are easily overlooked but have strong supervision value and are regarded as positive samples; while the confidence matching region corresponds to the region where the model prediction accuracy is relatively high and is regarded as negative samples. Since the number of positive samples is much less than that of negative samples in the actual matching process, it is necessary to design a sample correction loss function for sample imbalance to mitigate the training bias caused by the imbalance. The sample correction loss function includes the local point cloud true matching loss Ture Loss and the local point cloud matching deviation correction loss CorrectLoss, which is expressed as: (10) In the formula, is the sample correction loss function, is the local point cloud true matching loss Ture Loss, It is the deviation correction loss for local point cloud matching, namely Correct Loss; Both Ture Loss and Correct Loss are applicable to the imbalance of local point cloud samples. Using only Ture Loss makes the matching of local point cloud blocks unstable. Adding Correct Loss is used to assist in the matching; 1 Loss is used to strengthen the model's ability to recognize true matching point pairs and is defined as: (11) Where: (12) In the formula, , Both are hyperparameters for controlling the sample weights, with the default value being , is the probability, is the node.
[0038] Correct Loss focuses on compensating and optimizing the areas where the model prediction deviates and is defined as: (13) Wherein, is the confidence deviation value, is the confidence matching value, and the two jointly participate in the regularization adjustment of the loss.
[0039] The point cloud pair rigid transformation parameter estimation module 3 includes: After obtaining the corresponding relationship set of superpoints , the present invention expands the superpoint correspondence into the corresponding relationship between each pair of local point cloud blocks (including the points and after upsampling by the KPConv network and their related visual features and of the patch set) through the nearest superpoint grouping strategy (hereinafter referred to as the point-to-node grouping strategy), and matches the smaller-scale point clouds in the two local point cloud blocks within a local range to extract point-level corresponding relationships. For any superpoint , the grouping process of point-to-node can be realized by formula (14): (14) In the formula, is the Euclidean distance, is each point cloud to be assigned, and are respectively the point set and visual features on the fine scale related to the superpoint after grouping, For the th superpoint to be assigned, For the th superpoint to be assigned, For each feature to be assigned.
[0040] The present invention adopts a "density adaptive" mechanism to calculate the similarity matrix of each pair of local point cloud blocks. The density adaptive mechanism marks the invalid points caused by repeated sampling as infinitesimal values, and eliminates the matches between invalid points through optimal transport and iterative optimization to generate a point-level similarity matrix . Using as the confidence matrix of candidate matches, the final point correspondence is obtained by selecting the point pairs with higher confidence scores . Finally, based on the point correspondence , the rigid transformation parameters between two point clouds are estimated by using the RANSAC algorithm, so as to complete the automatic registration of indoor scene point clouds. As Figure 3 shown by the visualization effect of the registration result on the indoor scene provided by the embodiment of the present invention.
[0041] It can be seen from the above embodiments that the present invention combines the feature extraction ability of a deep learning network and uses a position-aware attention mechanism to achieve efficient interaction between local features and global context information of point clouds. By introducing a superpoint matching module based on an overlapping mask, the establishment process of superpoint correspondences in the coarse-scale stage of the coarse-to-fine point cloud automatic registration scheme is optimized, significantly improving the accuracy of superpoint correspondences in the coarse-scale stage, and at the same time enhancing the construction quality of point correspondences in the fine-scale stage, thereby significantly improving the quality and robustness of subsequent point cloud registration, and providing a reliable data basis for deep application fields such as virtual and augmented reality, cultural heritage protection, and reverse engineering.
[0042] To verify the effect of the present invention, comparative experiments were conducted between the proposed method of the present invention and multiple classic point cloud registration models such as 3DSN, FCGF, D3Feat, and Predator. The relative translation error (RTE), relative rotation error (RRE), and registration recall rate (RR) were used as evaluation indicators. The registration results of the present invention on the indoor scene datasets 3DMatch and 3DLoMatch are shown in Table 1.
[0043] Table 1 Registration results on 3DMatch and 3DLoMatch
[0044] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be covered by the protection scope of the present invention.
Claims
1. An automatic point cloud registration method for overlapping mask optimization from coarse to fine, characterized in that, The method comprises the following steps: S1, obtain the 3D point cloud data of the indoor scene, and extract the coarse-scale super-point visual features through the kernel-guided local feature aggregation mechanism and downsampling operation in the KPConv coding layer; S2, introduces the position-aware attention module to inject position information into the super-point features, providing geometric constraints for the subsequent attention mechanism; through self-attention and cross-attention operations, the local geometric features of the super-points and the global context information are adaptively fused, thereby effectively enhancing the expressiveness of the super-point features; S3, input the enhanced super-point features into the super-point overlap prediction network, generate overlapping masks through multi-layer convolution and remove non-overlapping areas; on this basis, use the retained overlapping super-points to build a similarity matrix, and optimize it in combination with the sample correction loss function, based on which the super-point correspondence on a coarse scale is established; S4, in the KPConv network decoding layer, the fine-scale point cloud visual features are obtained through nearest neighbor interpolation and upsampling operations, and then combined with the nearest super-point grouping strategy, the super-point correspondence is extended to the local point cloud block correspondence; S5, under the guidance of the density adaptive mechanism, a set of precise scale point correspondences is constructed based on the matching relationship between local point cloud blocks; on this basis, the rigid transformation parameters between the source point cloud and the target point cloud are estimated through the RANSAC algorithm to achieve automatic registration of indoor scene point clouds.
2. The coarse-to-fine automatic point cloud registration method for overlapping mask optimization according to claim 1, characterized in that In step S1, in the KPConv coding layer, the coarse-scale super-point visual features are extracted through the kernel-guided local feature aggregation mechanism and downsampling operation: Use the encoding layer of the KPConv network to superpoint the obtained source point cloud and target point cloud of the indoor scene and superpoint downsample to uniformly distributed superpoints and , and learn the relevant visual features and ; Input the downsampled and into the position encoding network; Among them, are all the input original source point cloud superpoints and target point cloud superpoints, is the set of real numbers, is the number of point clouds included in, are the source point cloud superpoint and the target point cloud superpoint; represents the number of superpoints, is the dimension of each superpoint feature, is the visual feature of the downsampled superpoints.
3. The coarse-to-fine automatic point cloud registration method for overlapping mask optimization according to claim 1, characterized in that In step S2, the position-aware attention module is introduced to inject position information into the super-point features to provide geometric constraints for the subsequent attention mechanism, including: Calculate the centroids of the input source point cloud superpoints and target point cloud superpoints and eliminate the overall spatial offset of the source point cloud and target point cloud through a decentralization operation; use a position encoding network to map the position information of the source point cloud and target point cloud to a high-dimensional feature space and fuse it with the original features to obtain the source point cloud superpoint features with position perception ability and the target point cloud superpoint features with position perception ability ; and is generated by the following expression: (1) (2) In the formula, is the position encoding function, " " is the feature fusion operation.
4. The coarse-to-fine automatic point cloud registration method with overlapping mask optimization according to claim 1, wherein In step S2, the local geometric features of the superpoint and the global context information are adaptively integrated through self-attention and cross-attention operations, thereby effectively enhancing the expressiveness of the superpoint features, including: The self-attention and cross-attention operations use multi-head attention, which includes: (3) (4) In the formula, is the output of the multi-head attention mechanism, is to concatenate multiple vectors by dimension, is the computation result of the th attention head, is the output of the standard single-head attention mechanism, are the query, key, and value in the attention mechanism, is the linear transformation matrix, , is the projection matrix for each attention head, is the number of heads, is the dimension of the input features, is the dimension to be processed by a single attention head; , and each attention head uses single-head dot-product attention, and the expression is: (5) In the formula, is a normalization function, which is used to convert the similarity between the query and the key into attention weights, and then is weighted, is the key is the transpose; When performing self-attention processing, Set as the features of the same point cloud, and set them in the source point cloud and the target point cloud respectively as: and ; When performing cross-attention processing, Set as the features of different point clouds, and set them in the source point cloud and the target point cloud respectively as: and .
5. The coarse-to-fine automatic point cloud registration method for overlapping mask optimization according to claim 1, characterized in that, In step S2, the step of enhancing the expressiveness of the super-point feature further includes: The original features are fused with the attention results through a multi-layer perceptron (MLP), and the expression is as follows: and the attention results are fused, and the expression is: (6) In the formula, is the fused feature, is the multi-layer perceptron, is to splice multiple vectors by dimension, is the original input feature, is the feature after being processed by the attention mechanism.
6. The coarse-to-fine automatic point cloud registration method with overlapping mask optimization according to claim 1, characterized in that, In step S3, the enhanced super-point features are input into the super-point overlap prediction network, and overlapping masks are generated through multi-layer convolution and non-overlapping areas are eliminated; On this basis, the similarity matrix is constructed using the retained overlapping super-points, and optimized in combination with the sample correction loss function. Based on this, the super-point correspondence relationship on the coarse scale is established, including: The backbone network of the superpoint matching module based on overlapping masks consists of 4 one-dimensional convolutional layers and activation functions. For each layer, pointwise convolution operations are performed using convolutional kernels, followed by group normalization. Except for the last layer, an activation function is used for non-linear transformation after each convolutional layer; the per-point features of each superpoint output by the superpoint matching module based on overlapping masks are processed by formula (7): (7) In the formula, is the convolution operation in the network, is the superpoint feature after being processed by the previous layer. When , it represents the input superpoint feature , is the classification result of the superpoints in the source point cloud output by the last layer of the network, including the probabilities of the superpoints belonging to overlapping points and non-overlapping points; Perform the same operation on the superpoint features of the target point cloud to obtain the classification result ; Pair And are processed to generate an overlap mask for super points, and the category of each point is selected, including: (8) (9) In the formula, is an operation, and are respectively the overlap masks of the source point cloud and the target point cloud, and each mask element takes a value of 0 or 1; For , apply a mask and , only retain the features located within the overlapping region, perform mask screening, and obtain the masked source point cloud superpoint features and the masked target point cloud superpoint features ; Utilize and to construct a similarity matrix between superpoint features , to quantify the similarity relationship between superpoints, and combine the Sinkhorn algorithm to perform optimal transport iterative optimization on the similarity matrix . At the same time, a sample correction loss function is used to alleviate the problem of sample imbalance; the Sinkhorn algorithm performs regularization operations on the rows and columns of the similarity matrix , and continuously iteratively optimizes to obtain the best match between superpoints. Based on the optimized similarity matrix , a set of superpoint correspondence relationships in the rough matching stage is generated .
7. The coarse-to-fine automatic point cloud registration method with overlapping mask optimization according to claim 6, characterized in that Similarity matrix for super points , divide into a confidence deviation region and a confidence matching region to distinguish the performance deviation of the model in prediction; among them, the confidence deviation region reflects that there is a significant difference between the model prediction result and the true match, which is easy to be ignored but has strong supervision value and is regarded as a positive sample; while the confidence matching region corresponds to the region with higher prediction accuracy of the model and is regarded as a negative sample; the sample correction loss function includes the local point cloud true matching loss Ture Loss and the local point cloud matching deviation correction loss Correct Loss, which is expressed as: (10) In the formula, is the sample deviation correction loss function, is the local point cloud true matching loss, is the local point cloud matching deviation correction loss; Both True Loss and Correct Loss are suitable for imbalanced local point cloud samples. Using only True Loss makes the matching of local point cloud blocks unstable, and Correct Loss is added for auxiliary matching. 1 Loss is used to enhance the model's ability to identify true matching point pairs and is defined as: (11) in: (12) In the formula, , are both hyperparameters for controlling the sample weights, with the default value being , is the probability, is the node; Correct Loss focuses on compensating and optimizing areas where the model prediction deviates, and is defined as: (13) Among them, is the confidence deviation value, is the confidence matching value, and both jointly participate in the regularization adjustment of the loss.
8. The coarse-to-fine automatic point cloud registration method for overlapping mask optimization according to claim 1, characterized in that In step S4, the decoding layer of the KPConv network obtains the fine-scale point cloud visual features through nearest neighbor interpolation and upsampling operations, and then combines the nearest neighbor superpoint grouping strategy to extend the superpoint correspondence to the local point cloud patch correspondence, including: Obtain the corresponding relationship set of super points After that, through the nearest super point grouping strategy (hereinafter referred to as the point-to-node grouping strategy), the corresponding relationship of super points is extended to the corresponding relationship between each pair of local point cloud blocks, and the point clouds with a smaller scale in the two local point cloud blocks are matched within the local range to extract point-level corresponding relationships; for any super point , the grouping process of point-to-node is implemented by formula (14): (14) Wherein, is the Euclidean distance, is each point cloud to be assigned, and are respectively the point set and visual feature on the fine scale related to the superpoint after grouping, is the th superpoint to be assigned to, is the th superpoint to be assigned to, is each feature to be assigned.
9. The coarse-to-fine automatic point cloud registration method for overlapping mask optimization according to claim 1, characterized in that, In step S5, under the guidance of the density adaptive mechanism, a fine-scale point correspondence set is constructed based on the matching relationship between local point cloud patches; on this basis, the rigid transformation parameters between the source point cloud and the target point cloud are estimated by the RANSAC algorithm to achieve the automatic registration of the indoor scene point cloud, including: The density adaptive mechanism is used to calculate the similarity matrix of each pair of local point cloud patches. The density adaptive mechanism marks the invalid points caused by repeated sampling as infinitesimal negative values, and eliminates the matches between invalid points through optimal transport and iterative optimization to generate a point-level similarity matrix. ; Use as the confidence matrix of candidate matches, and obtain the final point correspondence by selecting the point pairs with higher confidence scores. ; The RANSAC algorithm estimates the rigid transformation parameters by randomly sampling the point correspondences in , calculates the number of inliers of the local region based on these estimated transformations, and selects the rigid transformation parameters with the most inliers as the optimal solution to complete the automatic registration of the indoor scene point cloud.
10. An automatic point cloud registration system for overlapping mask optimization from coarse to fine, characterized in that, The system implements the coarse-to-fine point cloud automatic registration method with overlapping mask optimization as described in any one of claims 1-9. The system includes: A visual feature acquisition module for acquiring the three-dimensional point cloud data of the indoor scene, where the three-dimensional point cloud data includes a source point cloud and a target point cloud; downsampling the source point cloud and the target point cloud to uniformly distributed superpoints respectively, and learning the relevant visual features; A position-aware attention module for explicitly injecting position information into the superpoint features, making the features geometrically sensitive while having semantic expression, and providing geometric constraints for the subsequent attention mechanism; subsequently, using the self-attention mechanism to mine the correlation between the internal features of the superpoints, strengthening the expression ability of the local geometry, and at the same time combining the cross-attention mechanism to achieve context feature interaction between different point clouds; through this process, jointly model the local geometric structure and the global relationship, thereby significantly improving the discriminability of the superpoint features; An overlapping mask generation module for judging the overlap of the enhanced superpoint features, filtering out the superpoints located in the non-overlapping regions, and generating a set of superpoint correspondence relationships in the coarse matching stage based on the remaining superpoints in the overlapping regions; A superpoint correspondence set refinement and rigid transformation parameter estimation module for refining the superpoint correspondence set into a point correspondence set using the density adaptive mechanism; further estimating the rigid transformation parameters between the source point cloud and the target point cloud based on the point correspondence set to complete the automatic registration of the indoor scene point cloud.
Citation Information
Patent Citations
Point cloud registration method based on geometric embedding of significant anchor points
CN116228825A
Mask learning-based partially overlapped point cloud registration method
CN117455965A
Three-dimensional point cloud registration method based on micro-surface fusion and alignment
CN117876447A
Coarse-to-fine indoor scene point cloud automatic registration method
CN118154651A
System and method for semantic segmentation of images
US20190057507A1
Cited By
Robust point cloud registration method and device based on common-view query embedding
CN120543608A
Multi-view radar point cloud registration system
CN121010635A
Multi-view radar point cloud registration system
CN121010635B
Automobile pneumatic flow field prediction method, device and system
CN122113698A