A 3D Point Cloud Registration Method Based on Dual Attention Mechanism

By introducing a dual attention mechanism and weight ratio strategy in three-dimensional point cloud registration, the problem of point cloud registration accuracy and efficiency in the existing technology is solved, and a higher accuracy and stronger practical point cloud registration effect is achieved.

CN116862962BActive Publication Date: 2025-06-13FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310867586.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-15
Publication Date
2025-06-13
Estimated Expiration
2043-07-15

AI Technical Summary

Technical Problem

Existing 3D point cloud registration methods have accuracy and efficiency issues when dealing with noise, occlusion and unstructured data, especially when fast real-time registration is required.

Method used

A three-dimensional point cloud registration method based on the dual attention mechanism is adopted to obtain point clouds of different sparseness through multi-layer downsampling, and the influence of different levels of features is captured in combination with the dual attention mechanism, and a weight ratio strategy is designed to avoid error accumulation.

Benefits of technology

It improves the accuracy and efficiency of point cloud registration, enhances attention to channel and spatial characteristics, avoids error accumulation during the precise registration process, and thus adapts to a variety of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116862962B_ABST
    Figure CN116862962B_ABST
Patent Text Reader

Abstract

The present invention relates to a three-dimensional point cloud registration method based on a dual attention mechanism. Step S1: Use multi-layer downsampling to operate on the input two frames of point clouds to obtain point clouds with different sparsities; Step S2: The point clouds with different sparsities calculate the corresponding relationship between the two frames of point clouds according to the corresponding point matching module; Step S3: The feature maps of the matching point pairs pass through the dual attention mechanism to the descriptor feature module to obtain the descriptor features; Step S4: The confidence matrix output by the corresponding point matching module and the output of the descriptor feature module are subjected to singular value decomposition module to obtain the rotation transformation matrix; Step S5: The last downsampling layer calculates the rotation transformation matrix as the coarse registration process, and then refines the coarse registration result through the weight ratio strategy of the fine registration process. In the point cloud registration framework based on deep learning, the present invention introduces a dual attention mechanism for the feature effects of different levels of attention, designs a weight ratio strategy to solve the accumulation of registration errors, and avoids the occurrence of error accumulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional data, and in particular to a three-dimensional point cloud registration method based on a dual attention mechanism. Background Art

[0002] Three-dimensional data is an important expression of the real physical world and has various forms, including depth images, polygon meshes, and point clouds. The maturity and popularization of three-dimensional point cloud data acquisition devices such as lidar and the improvement of computer computing power have enabled the wide application of point cloud data. Point cloud registration technology is a very challenging research direction in point cloud data processing because point cloud is an unstructured data form, with occlusion, noise, and a large amount of data, and it needs to be completed accurately and quickly.

[0003] Research on point cloud registration algorithms has been carried out since the 1970s. ICP (iterative closest point) is one of the most classic point cloud registration algorithms, which estimates the transformation matrix by iteratively searching for corresponding points and minimizing the overall distance between points. Since it is non-convex, it is easy to fall into local minima. When large rotations and translations are required for point cloud registration, ICP often fails to obtain the correct result. NDT (normal distributions transform) is another classic fine registration method, which performs registration by maximizing the score of the probability density of the normal distribution calculated after voxelizing the target points with the source points. Simple point-to-point correspondences are not robust, so much subsequent work has been devoted to finding more robust correspondences, such as point-to-plane, plane-to-plane, and feature-to-feature. The registration parameters are estimated through these correspondence relationships, which improve the correct rate of point cloud registration and reduce the dependence on the correct initial pose. Feature detection reduces the total number of points and retains the parts with high distinctiveness; feature description aims to convert high-dimensional feature information into low-dimensional descriptors that are still easy to distinguish and compare; feature matching finds the correspondence relationships between features. Once the correspondence relationships are established, the registration problem becomes a convex problem and the transformation estimation becomes easier. Another coarse registration method is based on congruent four-point sets. Under rigid transformation, the transformation of some points is also the transformation of all points, and the best transformation is obtained by verifying multiple congruent sets. Traditional non-learning point cloud registration methods have low hardware requirements, are easy to implement, have strong interpretability, and do not require a time-consuming training process, but they have problems such as being easy to fall into local minima or having poor generality. Using manual features to distinguish correspondence relationships is greatly affected by the designer's experience and parameter adjustment ability. In addition, coarse registration methods are often time-consuming and may encounter bottlenecks in some real-time applications.

[0004] In recent years, deep learning has made a lot of progress. More and more algorithms have achieved performance improvements through deep learning, and deep learning has also been widely applied in the task of point cloud registration. Many algorithms have achieved speed improvement and shown robustness to noise, low overlap, and occlusion. The potential shown by learning-based methods in performance improvement makes them a popular research direction in the current field of point cloud registration. Learning on point clouds requires overcoming noise and occlusion of point clouds. A greater challenge lies in the unstructured, disordered, and irregular nature of point clouds, which is not conducive to promoting mature structured image methods to point clouds. Therefore, voxel-based methods, multi-view-based methods, and methods that directly learn on the original point cloud are used for learning on point clouds. Learning-based point cloud registration has many commonalities with non-learning methods. Many learning-based point cloud registration algorithms are based on non-learning registration algorithms, replacing a part of them with a network, or completely consisting of an end-to-end algorithm composed of a network. Therefore, learning-based methods are often a combination of non-learning methods and deep learning concepts. Most learning-based point cloud registration methods are feature-based, but the learned features have better generality, robustness, and descriptive ability compared to manually designed features. Learning-based methods can also learn more advanced and hierarchical descriptors, and are more robust to feature selection, feature representation, and differences in point density and measurement angles. Through the parallel computing ability of the GPU (graphics processing unit), learning-based point cloud registration methods have an advantage in speed, especially in coarse registration. However, learning-based methods often require more computing resources compared to non-learning methods. When there are real-time requirements but the computing resources are very limited, it is not suitable to use learning-based methods. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a three-dimensional point cloud registration method based on a dual attention mechanism, which can better achieve the efficiency of point cloud registration, has stronger practicability; and can achieve the registration of radar point cloud maps with higher accuracy, adapt to various application scenarios, capture the influence of different levels of features by introducing a dual attention mechanism, and establish a weight ratio strategy to solve the problem of error accumulation.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions: A three-dimensional point cloud registration method based on a dual attention mechanism, including the following steps:

[0007] Step S1: Use multi-layer downsampling to operate on the input two frames of point clouds to obtain point clouds with different sparsities;

[0008] Step S2: The point clouds with different sparsities calculate the corresponding relationship between the two frames of point clouds according to the corresponding point matching module;

[0009] Step S3: The feature maps of the matching point pairs pass through the dual attention mechanism to the descriptor feature module to obtain the descriptor features;

[0010] Step S4: The confidence matrix output by the corresponding point matching module and the output of the descriptor feature module are processed by the singular value decomposition module to obtain the rotation transformation matrix;

[0011] Step S5: The last layer of downsampling layer calculates the rotation transformation matrix as the coarse registration process, and then refines the coarse registration result through the weight ratio strategy of the fine registration process to obtain the optimal rotation transformation matrix.

[0012] In a preferred embodiment, the specific implementation process of the step S1 is as follows:

[0013] The point cloud downsampling module first randomly selects any point as the initial point and puts it into the set of candidates. Then, it calculates the distances from the initial point to all the remaining points, and selects the farthest point as another point to put into the set. Then, it continues to calculate the farthest distances from the points in the set to the remaining points as the points in the set, and repeats this process until the number of points in the set meets the required number.

[0014] In a preferred embodiment, the specific implementation process of the step S2 is as follows:

[0015] First, for each sampled point, find K nearest points in the upper layer of the point cloud by spatial distance to obtain a cluster;

[0016] For each cluster, calculate the geometric feature F D ; It is obtained by concatenating the relative coordinate positions, distances from the nearest points to the center, and the descriptors D calculated for the nearest points in the upper layer;

[0017] F D = torch.cat(r rela , r dist , knn_feature)

[0018] r rela represents the relative coordinate position, r dist represents the distance, and knn_feature is the neighboring feature;

[0019] These features are fed into the perceptron to obtain the feature map, and then the weights of each neighboring point are obtained through a max layer and a normalization layer;

[0020] Finally, the weighted sum of these nearest points is used as the final sampled point.

[0021] In a preferred embodiment, the specific implementation process of the step S3 is as follows:

[0022] Send the geometric features into a three-layer shared perceptron to obtain a feature map, and send the feature map into a dual attention mechanism:

[0023] F″ = M S (M C (F map ) × F map ) × (M C (F map ) × F map )

[0024] where F map is the feature map, and M C (·) represents the channel attention mechanism, and M S (·) represents the spatial attention mechanism.

[0025] In a preferred embodiment, the specific implementation process of step S4 is as follows:

[0026] Step S41: Solve the rotation transformation matrix;

[0027] First, to solve the optimal rotation transformation matrix, the following function needs to be satisfied:

[0028]

[0029] where c i represents the confidence matrix, and x i , y i represent the matching point pairs in the source point cloud and the target point cloud respectively; K represents the number of matching point pairs;

[0030] It is obtained by singular value decomposition, and its representation is as follows:

[0031] R = VU T

[0032] where R, V, and U are all orthogonal matrices;

[0033] Step S42: Loss function;

[0034] The loss function consists of a rotation loss function and a translation loss function, and its calculation method is as follows:

[0035] L = L trans + αL rot

[0036] where L rot represents the rotation loss function, L trans represents the translation loss function, and α is a preset weight value, and their calculation methods are as follows:

[0037]

[0038]

[0039] Where R and R respectively represent the true value and the estimated value of rotation, t and t respectively represent the true value and the estimated value of translation, I represents the identity matrix, and ||·|| 2 represents calculating the Euclidean distance;

[0040] Step S43: Use the stochastic gradient descent algorithm to minimize the loss function to update the local model parameters.

[0041] In a preferred embodiment, the results of the first-layer feature extraction and the second-layer feature extraction are used to cross-match the feature weights of different layers to avoid the matching error affected by individual matching error points; among them, the weight matching parameter satisfies the condition: w 1 +w 2 +w 3 +w 4 = 1.

[0042] Compared with the prior art, the present invention has the following beneficial effects: The present invention proposes a point cloud registration framework based on deep learning, introduces a dual attention mechanism to capture the influence of different-level features, improves the attention to channel and spatial features, and thus improves the registration accuracy. In addition, a weight matching strategy is designed in the fine registration process, rather than using the steps of a single fine registration process iteration, thereby avoiding the occurrence of error accumulation in the fine registration process to improve the point cloud registration degree. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is the structural diagram of the point cloud registration network based on deep learning in the preferred embodiment of the present invention.

[0044] Figure 2 is the point cloud registration recall rate under different relative translation error thresholds in the preferred embodiment of the present invention.

[0045] Figure 3 is the point cloud registration recall rate under different relative rotation error thresholds in the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0046] The present invention will be further described below in conjunction with the drawings and embodiments.

[0047] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0048] Note that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0049] The present invention provides a point cloud registration network framework based on deep learning. Refer to Figures 1 to 3 , including the following steps:

[0050] Step S1: Use multi-layer downsampling to operate on the input two frames of point clouds to obtain point clouds with different sparsities;

[0051] Step S2: Calculate the corresponding relationship between the two frames of point clouds according to the corresponding point matching module for point clouds with different sparsities;

[0052] Step S3: The feature maps of the matching point pairs pass through a dual attention mechanism to the descriptor feature module to obtain descriptor features;

[0053] Step S4: Calculate the rotation transformation matrix through the singular value decomposition module for the confidence matrix output by the corresponding point matching module and the output of the descriptor feature module;

[0054] Step S5: The last downsampling layer calculates the rotation transformation matrix as the coarse registration process, and then refines the coarse registration result through the weight ratio strategy of the fine registration process to obtain the optimal rotation transformation matrix.

[0055] Furthermore, the specific implementation process of step S1 is as follows:

[0056] The point cloud downsampling module first randomly selects any point as the initial point and puts it into the set to be selected. Then, it calculates the distances from the initial point to all the remaining points, and selects the farthest point as another point to put into the set. Then, it continues to calculate the farthest distance from the points in the set to the remaining points as the points in the set, and repeats this process until the number of points in the set meets the required number.

[0057] Furthermore, the specific implementation process of step S2 is as follows:

[0058] First, for each sampled point, find K nearest points in the upper layer of the point cloud by spatial distance to obtain a cluster; for each cluster.

[0059] For each cluster, calculate the geometric feature F D . It is obtained by concatenating the relative coordinate positions, distances from the nearest points to the center, and the descriptors D calculated for the nearest points in the upper layer.

[0060] FD = torch.cat(r rela , r dist , knn_feature)

[0061] r rela represents the relative coordinate position, r dist represents the distance, and knn_feature is the neighboring feature.

[0062] These features are fed into a perceptron to obtain a feature map, and then through a max layer and a normalization layer to obtain the weights of each neighboring point.

[0063] Finally, the weighted sum of these nearest points is used as the final sampling point.

[0064] Furthermore, the specific implementation process of step S3 is as follows:

[0065] The geometric features are fed into a three-layer shared perceptron to obtain a feature map, and the feature map is fed into a dual attention mechanism:

[0066] F″ = M S (M C (F map ) × F map ) × (M C (F map ) × F map )

[0067] where F map is the feature map, M C (·) represents the channel attention mechanism, and M S (·) represents the spatial attention mechanism.

[0068] Furthermore, in step S4, step S41, solve the rotation transformation matrix;

[0069] First, in order to solve the optimal rotation transformation matrix, the following function needs to be satisfied:

[0070]

[0071] where c i represents the confidence matrix, x i , y i represent the matching point pairs in the source point cloud and the target point cloud respectively. K represents the number of matching point pairs.

[0072] In order to solve the rotation transformation matrix, it needs to be obtained according to the singular value decomposition function, which is expressed as follows:

[0073] R = VU T

[0074] Among them, R, V, and U are all orthogonal matrices.

[0075] Step S42, loss function;

[0076] Since the main goal of registration is to solve the correct rotation transformation angle and translation distance, this paper mainly considers the rotation error and translation error as the main loss functions. This paper believes that its loss function is composed of a rotation loss function and a translation loss function, and its calculation method is as follows:

[0077] L = L trans + αL rot

[0078] where L rot represents the rotation loss function, L trans represents the translation loss function, and α is a preset weight value. Their calculation methods are as follows:

[0079]

[0080]

[0081] where R represents the true value and the estimated value of rotation respectively, t represents the true value and the estimated value of translation respectively, I represents the identity matrix, and ||·|| 2 represents calculating the Euclidean distance.

[0082] Step S43, adopt the stochastic gradient descent algorithm to update the local model parameters by minimizing the loss function.

[0083] Furthermore, in the said step S5, in order to further improve the accuracy of point cloud registration, the results of the first-layer feature extraction and the results of the second-layer feature extraction are used to cross-match the feature weights of different layers to avoid the matching error affected by individual mis-matched points.

[0084] The present invention proposes a point cloud registration method based on a dual attention mechanism, which further improves the accuracy of 3D point cloud registration. In order to improve the registration efficiency, considering the huge amount of point cloud data, multi-layer downsampling operations are adopted to obtain point cloud sets with different sparsities, and the point cloud sets with different sparsities are used to calculate matching points; at the same time, in order to improve the attention to the influence of different-level features, the present invention introduces a dual attention mechanism to capture the attention of different channel and spatial features; finally, in order to avoid the problem of error accumulation caused by the iteration in the fine registration process, the present invention designs a weight ratio strategy to assign weights to the point cloud sets with different sparsities for calculating the rotation matrix, avoiding the occurrence of the error accumulation problem.

[0085] The accuracy of the method of the present invention in point cloud registration is improved. By comparing with other traditional multi-modal feature fusion methods, it is found that compared with the existing algorithms, the present solution can achieve better accuracy in the point cloud registration task.

[0086] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention, when the functions and effects produced do not exceed the scope of the technical solution of the present invention, shall fall within the protection scope of the present invention.

Claims

1. A 3D point cloud registration method based on a dual attention mechanism, characterized in that, it includes the following steps: Step S1: Use multi-layer downsampling to operate on the input two frames of point clouds to obtain point clouds with different sparsities; Step S2: The point clouds with different sparsities calculate the corresponding relationship between the two frames of point clouds according to the corresponding point matching module; Step S3: The feature maps of the matching point pairs pass through the dual attention mechanism to the descriptor feature module to obtain descriptor features; Step S4: The confidence matrix output by the corresponding point matching module and the output of the descriptor feature module are passed through the singular value decomposition module to obtain the rotation transformation matrix; Step S5: The last downsampling layer calculates the rotation transformation matrix as the coarse registration process, and then refines the coarse registration result through the weight ratio strategy of the fine registration process to obtain the optimal rotation transformation matrix; The specific implementation process of the above-mentioned Step S1 is as follows: The point cloud downsampling module first randomly obtains any point as the initial point and puts it into the set to be selected. Then, it calculates the distances from the initial point to all the remaining points, and puts the farthest point into the set as another point. Then, it continues to calculate the farthest distance from the points in the set to the remaining other points as the points in the set, and repeats this process until the number of points in the set meets the required number; The specific implementation process of the above-mentioned Step S2 is as follows: For each sampled point, find K nearest points in the upper layer of the point cloud by spatial distance to obtain a cluster; For each cluster, calculate the geometric feature F D ; it is obtained by concatenating the relative coordinate position, distance from the nearest point to the center, and the descriptor D calculated in the previous layer for the nearest point; F D = torch.cat(r rela , r dist , knn_feature) r rela represents the relative coordinate position, and r dist represents the distance, and knn_feature is the neighboring feature; Send these features into the perceptron to obtain the feature map, and then obtain the weights of each neighboring point through a maximum layer and a normalization layer; Finally, the weighted sum of these nearest points is used as the final sampled point; The specific implementation process of the above-mentioned Step S4 is as follows: Step S41: Solve the rotation transformation matrix; First, in order to solve the optimal rotation transformation matrix, the following function needs to be satisfied: Among them, c i represents the confidence matrix, x i represents the matching point pairs in the source point cloud, y i represents the matching point pairs in the target point cloud; K represents the number of matching point pairs; Obtained according to the singular value decomposition, which is expressed as follows: R = VU T Among them, R, V, and U are all orthogonal matrices; Step S42: Loss function; The loss function consists of a rotation loss function and a translation loss function, and its calculation method is as follows: L = L trans + αL rot Among which, L rot represents the rotation loss function, and L trans represents the translation loss function. α is a preset weight value, and their calculation methods are as follows: where represents the true value of rotation, R represents the estimated value of rotation, represents the true value of translation, t represents the estimated value of translation, I represents the identity matrix, ||·|| 2 represents the calculation of Euclidean distance; Step S43: Adopt the stochastic gradient descent algorithm to minimize the loss function to update the local model parameters.

2. A 3D point cloud registration method based on a dual attention mechanism according to claim 1, characterized in that, the specific implementation process of the above-mentioned Step S3 is as follows: Send the geometric features into a three-layer shared perceptron to obtain the feature map, and send the feature map into the dual attention mechanism: F″ = M S (M C (F map ) × F map ) × (M C (F map ) × F map ) Among them, F map is the feature mapping, M C (·) represents the channel attention mechanism, M S (·) represents the spatial attention mechanism.

3. A 3D point cloud registration method based on a dual attention mechanism according to claim 1, characterized in that, Use the results of the first-layer feature extraction and the results of the second-layer feature extraction to cross-match the feature weights of different layers, avoiding the matching error affected by individual matching error points; among them, the weight matching parameter satisfies the condition: w 1 +w 2 +w 3 +w 4 = 1.

Citation Information

Patent Citations

  • Low-overlap 3D dynamic point cloud registration method and system based on attention mechanism

    CN114332175A

  • Apparatus and method for searching for global minimum of point cloud registration error

    US20220254095A1