An end-to-end point cloud registration method based on multi-scale fusion and hybrid position coding

By employing an end-to-end point cloud registration method that combines multi-scale fusion and hybrid position coding, the accuracy problem of point cloud registration in scenarios with scale variations and symmetrical structures is solved, achieving higher accuracy and robustness in point cloud registration.

CN117058203BActive Publication Date: 2025-11-07TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310931831.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-27
Publication Date
2025-11-07
Estimated Expiration
2043-07-27

AI Technical Summary

Technical Problem

Existing point cloud registration methods are sensitive to the initial pose and have difficulty in accurately registering in scenes with scale variations or symmetrical structures. Furthermore, deep learning-based methods often ignore local details, leading to reduced registration accuracy.

Method used

An end-to-end point cloud registration method using multi-scale fusion and hybrid location coding is adopted. Multi-scale features are obtained through a weight-sharing multi-scale feature extraction network, and combined with a hybrid location coding information interaction network, the correspondence and rigid transformation matrix of point clouds are predicted.

Benefits of technology

It improves the accuracy and robustness of point cloud registration, performs well in real-world application scenarios, and effectively utilizes the parallel computing capabilities of GPUs to enhance registration speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058203B_ABST
    Figure CN117058203B_ABST
Patent Text Reader

Abstract

This invention relates to the fields of computer vision and point cloud registration, specifically an end-to-end point cloud registration method based on multi-scale fusion and hybrid position coding. It includes: S1: processing the input source point cloud... X With target point cloud Y Key points and their corresponding features are obtained through a multi-scale feature extraction network with weight sharing, respectively; S2: The key points and their features are input into an information interaction network with hybrid positional encoding to obtain depth features containing spatial location information; S3: After the depth features are processed by a prediction network, the corresponding positions of points in the source point cloud and points in the target point cloud are output, and the corresponding positions of points in the target point cloud and the source point cloud are predicted, while the overlap score is predicted; S4: A rigid transformation matrix is ​​calculated based on the correspondence of key points in the overlapping area predicted in S3, and the registration of the source point cloud and the target point cloud is achieved based on this transformation matrix. This invention can achieve high registration accuracy and has good performance in real-world application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer vision and point cloud registration, and particularly relates to an end-to-end point cloud registration method based on multi-scale fusion and hybrid position coding. BACKGROUND

[0002] Point cloud is a data set composed of a large number of discrete three-dimensional points, which can be used to represent objects and scenes in the real world. However, due to the differences in position, pose and scale of point cloud data obtained at different times and from different perspectives, point cloud registration technology is needed to achieve the integrity and consistency of point cloud data, and thus to achieve fine modeling of the complete three-dimensional scene. As an important computer vision technology, point cloud registration has shown great application potential in many fields. In robot navigation, by registering point cloud data from different perspectives, the robot can complete accurate environmental perception and establish an accurate scene map, thereby achieving autonomous navigation and obstacle avoidance functions in complex environments. In geological exploration, point cloud registration can fuse multiple collected geological data to obtain a more comprehensive and accurate geological model. In the field of autonomous driving, point cloud registration can register point cloud data collected by vehicle-mounted laser radar or cameras with high-precision maps, which can help autonomous vehicles accurately locate and navigate in complex road environments.

[0003] The purpose of point cloud registration is to convert point cloud data from two or more camera coordinate systems to the world coordinate system and complete splicing. However, due to factors such as partial overlap and symmetry, point cloud registration of real scenes is still a challenging problem.

[0004] Although traditional methods represented by ICP, NDT and 4PCS and their variants have been widely used in various fields. However, most of these methods are sensitive to the initial pose and are difficult to generalize to real-world three-dimensional point clouds. When encountering scale changes or scenes with symmetrical structures, these methods may obtain incorrect point correspondences.

[0005] On the other hand, point registration methods based on deep learning have attracted attention due to their robustness to three-dimensional point clouds. Point cloud registration methods based on deep learning can be divided into two categories: feature-based methods and methods. Feature-based methods usually proceed in two stages, first learning local descriptors of down-sampled sparse points to find correspondences, and then using a robust pose estimator such as RANSAC to estimate the transformation matrix. Feature-based methods can be performed without an initial pose, and are generally robust to outliers. However, the use of RANSAC-like estimators to reject outliers makes this method time-consuming, and the use of down-sampling operations when extracting features loses a lot of detailed information.

[0006] Recently, more attention is paid to registration methods. These methods complete the learning of point features and transformation estimation in one forward propagation. The point cloud registration method maximizes the advantages of deep learning methods and better utilizes the parallel computing power of GPUs to obtain faster speed. However, these methods often ignore the local detailed information of the point cloud. For repeated structures or changes in scale, they may produce incorrect correspondence estimation. Effective and accurate registration remains a problem to be solved.

[0007] Most point cloud registration methods need to continuously down-sample the input point cloud, and use the features obtained by the last down-sampling for registration, but this will lose the fine-grained information of the underlying features, resulting in unclear features for finding corresponding relationships, causing the performance of the model to decrease. Moreover, there are repeated or symmetrical structures in the point cloud, for example, there may be multiple identical chairs in a point cloud, and the left and right hands of a person, and similar situations. These repeated and ambiguous structural information will greatly affect the uniqueness of the extracted features, resulting in a significant decrease in registration accuracy, so solving these two problems is of great significance to improve the accuracy of point cloud registration. SUMMARY

[0008] The present application provides an end-to-end point cloud registration method based on multi-scale fusion and hybrid position coding to solve the above problems.

[0009] The present application adopts the following technical scheme: an end-to-end point cloud registration method based on multi-scale fusion and hybrid position coding, comprising:

[0010] S1: inputting the source point cloud X and the target point cloud Y respectively through a weight-shared multi-scale feature extraction network to obtain key points , and corresponding features , ;

[0011] S2: inputting the key points and their features into an information interaction network embedded with hybrid position coding to obtain deep features containing spatial position information and ;

[0012] S3: after the deep features pass through a prediction network, outputting the corresponding positions of the points in the source point cloud in the target point cloud , and the corresponding positions of the points in the target point cloud in the source point cloud , and simultaneously predicting the overlap score , ;

[0013] S4: Calculate the rigid transformation matrix [R, t] according to the correspondence of the key points in the overlapping region predicted in S3, where R represents the rotation matrix and t represents the translation vector, and the registration of the source point cloud and the target point cloud is realized according to the transformation matrix.

[0014] In some embodiments, the multi-scale feature extraction network in step S1 includes the following steps:

[0015] S11: The input source point cloud and target point cloud are respectively subjected to three times of downsampling to obtain three features of different scales, which are respectively denoted as first scale feature , second scale feature , and third scale feature .

[0016] S12: The feature obtained after the last sampling is upsampled to obtain a feature , which is fused with the second scale feature . The fused feature is again upsampled to obtain a feature , which is fused with the first scale feature to obtain the deepest feature .

[0017] S13: Finally, the is aggregated with the to obtain a feature with multi-scale information, which is used for subsequent registration.

[0018] In some embodiments, the information interaction network with mixed position encoding in step S2 includes the following steps:

[0019] S21: Embed the position information in the feature , to obtain a key point feature with spatial position information to reject the plausible matching in the point cloud pair;

[0020] S22: The point cloud feature after position information embedding is input into the Transformer network to output an enhanced feature.

[0021] In some embodiments, the position information of the key point is modeled in step S21 using the following formula,

[0022]

[0023] wherein represents the absolute position encoding of , represents the relative position encoding of absolute position encoding, which embeds the position information of one point into the feature and interacts with another point feature, i.e., two features interact with each other only one of which carries the position information, and represents relative position encoding between , which indicates that both features carry position information interact with each other, is a position encoding function, represents position information, represents relative position information between

[0024] Step S22 includes:

[0025] S221: The key point feature embedded with position information is associated with the global context information in the same point cloud through a self-attention module;

[0026] S222: The feature obtained through the self-attention module is associated with the global context information from two different point clouds through a cross-attention module;

[0027] S223: The feature obtained through the cross-attention module is output as a deep feature and through two layers of feedforward networks.

[0028] Step S3 includes: the prediction network takes and obtained by the information interaction network as input, respectively passes through two branches, one branch predicts the coordinate position and of the corresponding key point; the other branch predicts the overlap score and .

[0029] The first branch predicts the position of the key point through a three-layer MLP network;

[0030] The second branch directly predicts the overlap score through a linear layer, and the overlap score represents the probability that the point is located in the overlap region.

[0031] In step S4, the rigid transformation matrix [R, t] is calculated according to the following formula,

[0032]

[0033] wherein, represents the position of the key point in the target point cloud in the source point cloud; represents the position of the key point in the source point cloud in the target point cloud; represents the probability that the key point in the source point cloud is located in the overlap region. represent the probability that the key point in the target point cloud is located in the overlapping area, , , and obtained by the correspondence prediction network; and obtained by the multi-scale feature extraction network, respectively representing the key points of the source point cloud and the key points of the target point cloud.

[0034] Compared with the prior art, the present application has the following beneficial effects:

[0035] 1. The present application proposes a multi-scale feature extraction network that uses a feature pyramid structure to extract multi-scale features of point clouds, so that the point cloud features used for registration contain both low-level structure information and high-level semantic information.

[0036] 2. The present application designs a hybrid position encoding method that considers both the absolute position and the relative position relationship of points in space, and embeds the hybrid position encoding method into the Transformer information interaction network to constrain the learned features through position information, so as to encourage the network to mine more effective point cloud feature representations.

[0037] 3. The present application integrates the multi-scale feature extraction network, the information interaction network embedded with the hybrid position encoding, and the correspondence prediction network into an end-to-end point cloud registration network, and simultaneously trains on a large number of point cloud datasets to obtain more accurate and robust point correspondences.

[0038] 4. Compared with existing deep learning-based point cloud registration methods, the present application can achieve high registration accuracy due to the use of multi-scale feature extraction methods and hybrid position encoding methods, and has good performance in real application scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0039] Figure 1 is a schematic diagram of the registration process of the present application;

[0040] Figure 2 is a schematic diagram of the multi-scale feature extraction module of the present application;

[0041] Figure 3 is a schematic diagram of the information interaction module with hybrid position encoding;

[0042] Figure 4 is a registration result diagram on 3DMatch;

[0043] Figure 5 is a registration result diagram on 3DLoMatch;

[0044] Figure 6To register the results in real scenes. DETAILED DESCRIPTION

[0045] In order to make the objects, technical solutions and advantages of the present application clearer, further detailed description will be given below in combination with the drawings and examples. It should be understood that the examples described herein are only used to illustrate and explain the present application, and are not used to limit the present application.

[0046] The present application relates to an end-to-end point cloud registration method based on multi-scale fusion and hybrid position coding, which is used to solve the point cloud registration task in different scenes and improve the point cloud registration accuracy. Most point cloud registration methods directly use the sampled points to extract features. Because the number of points obtained after sampling is greatly reduced compared to the original points, a lot of detailed information is lost when extracting features. The present application proposes a multi-scale feature extraction module to fuse shallow features and deep features. The features used for registration contain multi-scale information. In view of the repeated or symmetrical structures widely existing in point clouds, such as two identical chairs, people observe things not only by considering the appearance of things, but also by considering their positions in real scenes. The present application proposes hybrid position coding to consider the absolute position and relative position of points at the same time, and to reject the spurious matching of point clouds.

[0047] Specifically, as shown in Figure 1 , an end-to-end point cloud registration method based on multi-scale fusion and hybrid position coding comprises the following steps:

[0048] S1: input source point cloud X and target point cloud Y respectively through a weight-shared multi-scale feature extraction network to obtain key points , and corresponding features , ;

[0049] S2: input the key points and their features into a hybrid position coding embedded information interaction network to obtain deep features containing spatial position information and ;

[0050] S3: after the deep features pass through a prediction network, output the corresponding positions of the points in the source point cloud in the target point cloud , and the corresponding positions of the points in the target point cloud in the source point cloud , and simultaneously predict the overlap score , ;

[0051] S4: calculating a rigid transformation matrix [R, t] according to the correspondence of the key points in the overlapping region predicted in S3, wherein R represents a rotation matrix and t represents a translation vector, and the registration of the source point cloud and the target point cloud is realized according to the transformation matrix.

[0052] The definitions are as follows:

[0053]

[0054] wherein, , , , obtained by the correspondence prediction network; and obtained by the multi-scale feature extraction network. The final rigid transformation matrix [R, t] can be obtained by minimizing the above formula, and the registration of the source point cloud and the target point cloud can be realized by the transformation matrix, wherein R represents a rotation matrix and t represents a translation vector.

[0055] Figure 2 A network flowchart of multi-scale feature extraction in step S1 is shown, and the specific steps are as follows:

[0056] S11: the input source point cloud and target point cloud are respectively subjected to three times of downsampling to obtain three features of different scales, which are respectively denoted as first scale feature , second scale feature , and third scale feature , Figure 2 The three times of downsampling are shown in the dashed box, a residual block structure is used to extract the features of the points after each sampling, the feature dimension is increased from 64 to 1024, and KPConv is used as the convolution.

[0057] S12: the feature obtained after the last sampling is upsampled to obtain a feature , which is fused with the second scale feature , and the fused feature is upsampled again to obtain a feature , which is fused with the first scale feature to obtain the deepest feature , wherein the fusion of the deep and shallow features is a point-by-point addition manner.

[0058] S13: finally, the is aggregated with the to obtain a feature with multi-scale information, which is used for subsequent registration, wherein the fusion operation of the features of different scales uses an aggregation splicing manner.

[0059] The information interaction module with mixed position coding in step S2 takes the key points and the corresponding features obtained in S1 as inputs, first embeds position information in the features through position coding, and then interacts the features through a Transformer to aggregate global context information. Meanwhile, the application designs a network structure that simultaneously embeds absolute position information and relative position information in the Transformer. Figure 3 A flowchart of the information interaction module with mixed position coding is shown:

[0060] S21: Embedding position information in the features , of the key points to obtain key point features with spatial position information to reject false matches in the point cloud. The application designs mixed position coding that simultaneously considers the absolute position and relative position information of the points. The following formula is used:

[0061]

[0062] wherein represents the absolute position coding of , represents the absolute position coding of . The absolute position coding embeds the position information of one point in the feature and interacts with another point feature, that is, two features interact with each other, only one of which carries position information, and represents the relative position coding between and . It represents the interaction of two features both carrying position information, is a position coding function, represents the position information of , represents the relative position information between .

[0063] S22: Input the point cloud features with embedded position information into the Transformer network to output enhanced features.

[0064] The Transformer in step S22 includes the following steps:

[0065] S221: The key point features with embedded position information aggregate global context information in the same point cloud through a self-attention module.

[0066] S222: The features obtained through the self-attention module correlate global context information from two different point clouds through a cross-attention module.

[0067] S223: The features obtained through the cross-attention module are output as deep features through two layers of feedforward networks and .

[0068] Step S3 includes predicting the correspondence between the key points in the source point cloud and the target point cloud by the correspondence prediction network. and As input, two branches are respectively passed through, one branch predicts the coordinate position of the corresponding key point and ; the other branch predicts the overlap score and ; the first branch predicts the position of the key point through a three-layer MLP network; the second branch directly predicts the overlap score through a linear layer, and the overlap score represents the probability that the point is located in the overlap region.

[0069] Taking the source point cloud as an example, the corresponding positions of the key points in the target point cloud are predicted:

[0070]

[0071] wherein and represent the learnable weights and biases, respectively. Similarly, the same method is used to predict the corresponding positions of the key points of the target point cloud in the source point cloud.

[0072] In step S4, the rigid transformation matrix is calculated according to the correspondence of the key points in the overlap region, which is defined as follows:

[0073]

[0074] wherein, represents the position of the key point in the target point cloud in the source point cloud; represents the position of the key point in the source point cloud in the target point cloud; represents the probability that the key point in the source point cloud is located in the overlap region; represents the probability that the key point in the target point cloud is located in the overlap region, , , and are obtained by the correspondence prediction network; and are obtained by the multi-scale feature extraction network, and respectively represent the key points of the source point cloud and the key points of the target point cloud. The final rigid transformation matrix [R, t] can be obtained by minimizing the above formula, and the registration of the source point cloud and the target point cloud can be realized through the transformation matrix, wherein R represents the rotation matrix and t represents the translation vector.

[0075] To verify the effectiveness of the present invention, it was compared with several other point cloud registration algorithms on the 3DMatch and ModelNet datasets, as well as their corresponding low-overlap datasets 3DLoMatch and ModelLoNet, to verify the superiority of the present invention in different application scenarios.

[0076] The 3DMatch dataset consists of 62 real-world indoor scenes, including 46 training sets, 8 validation sets, and 8 test sets. These scenes are derived from 3D reconstruction datasets such as SUN3D and 7-Scene. ModelNet40 is a synthetic dataset. It comprises 12,311 CAD (computer-aided design) mesh models without noise or outliers, covering 40 common categories. This dataset is primarily used for point cloud classification and retrieval, and is also a major synthetic dataset in the field of deep learning point cloud registration in recent years.

[0077] This invention uses common point cloud registration evaluation metrics, namely registration recall, relative rotation error, and relative translation error.

[0078] The registration performance on the 3DMatch dataset is evaluated using Translation Error, while the registration performance on the ModelNet dataset is evaluated using Chamfer Distance, Relative Rotation Error, and Relative Translation Error.

[0079] The specific experimental procedure includes the following steps:

[0080] (1) Model training: The method of this invention is trained on the training set of the public dataset 3DMatch and ModelNet. The AdamW optimizer is used with an initial learning rate of 0.0001. After 20 epochs of training, the learning rate will decrease by half.

[0081] (2) Model testing: The model was tested on the 3DMatch test set and the ModelNet dataset to verify the effectiveness of the model on these two datasets.

[0082] (3) Experimental results: The performance on the two evaluation indicators is shown in the table below:

[0083] Table 1. Comparison of registration performance on the two datasets

[0084]

[0085]

[0086] From the above experimental results, it can be seen that, whether on the 3DMatch or ModelNet dataset, the method of the present application achieves better registration accuracy compared with some advanced comparative algorithms, and also obtains lower RRE and RTE. In particular, in Table 1, the method of the present application can achieve 94.9% on 3DMatch. This shows that the present method can well improve the accuracy of point cloud registration. Figure 3 , Figure 3 For the qualitative results of the comparison between the present method and some advanced methods, it can be seen that the present method has better registration performance

[0087] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An end-to-end point cloud registration method based on multi-scale fusion and hybrid location encoding, characterized in that, Comprise: S1: input source point cloud X and target point cloud Y respectively through a weight-shared multi-scale feature extraction network to obtain key points , and corresponding features , ; S2: input the key points and their features into the information interaction network embedded with the mixed position coding to obtain deep features containing spatial position information with ; The information interaction network with mixed location encoding in step S2 comprises the following steps: S21: embedding the feature , in the middle of the position information, obtaining the key point feature with spatial position information to reject the false match of the point cloud The position information of the key points is modeled in step S21 using the following formula, wherein represents an absolute position encoding of represents an absolute position encoding of represents a relative position encoding between and is a position encoding function, represents a position information of represents a relative position information between S22: The point cloud features embedded with position information are input into a Transformer network, and enhanced features are output. S3: After the deep features pass through the prediction network, the corresponding positions of the points in the source point cloud in the target point cloud are output , and the corresponding positions of the points in the target point cloud in the source point cloud are also predicted , ; S4: A rigid transformation matrix [R, t] is calculated according to the correspondence of the key points in the overlapping region predicted in S3, where R represents a rotation matrix and t represents a translation vector, and the registration of the source point cloud and the target point cloud is realized according to the transformation matrix.

2. The method of claim 1, wherein, The multi-scale feature extraction network in step S1 comprises the following steps: S11: input source point cloud and target point cloud are respectively subjected to three times of down-sampling to obtain three features of different scales, denoted as first scale feature , second scale feature , and third scale feature ; S12: the features obtained by the last sampling are up-sampled to obtain features are fused with the second scale features are fused to obtain features are up-sampled again to obtain features are fused with the first scale features are fused to obtain the deepest level features ; S13: Finally, the with aggregation operation to obtain features with multi-scale information for subsequent registration.

3. The method of claim 1, wherein, The step S22 comprises: S221: The key point features embedded with position information are aggregated with the global context information in the same point cloud through a self-attention module; S222: The features obtained through the self-attention module are associated with the global context information from two different point clouds through a cross-attention module; S223: The features obtained through the cross-attention module pass through a two-layer feedforward network to output deep features and .

4. The method of claim 1, wherein, The step S3 comprises: predicting information obtained by the interaction network of the network and As input, two branches are passed respectively, one branch predicts the coordinate position of the corresponding key point and The other branch predicts the overlap score and ; The first branch predicts the position of the key point through a three-layer MLP network; The second branch directly predicts the overlap score through a linear layer, and the overlap score represents the probability that the point is located in the overlapping region.

5. The method of claim 1, wherein, The rigid transformation matrix [R, t] in step S4 is calculated according to the following formula, wherein, representing that a key point in the target point cloud is located at a position of the source point cloud; representing that a key point in the source point cloud is located at a position in the target point cloud; representing a probability that a key point in the source point cloud is located at an overlapping region; representing a probability that a key point in the target point cloud is located at an overlapping region, , , and obtained by the correspondence prediction network; and obtained by the multi-scale feature extraction network, respectively representing a key point of the source point cloud and a key point of the target point cloud.