Point cloud registration method based on local feature enhancement of isotropic graph neural network
By using the local feature enhancement method of the equivariant graph neural network, the problems of low success rate and high complexity of point cloud registration in complex scenes are solved, achieving higher registration accuracy and robustness, especially with excellent performance under large angle transformations.
Patent Information
- Application Number
- CN202511064831.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-11
AI Technical Summary
Existing point cloud registration methods have low success rates and high complexity in highly challenging real-world scenarios. In particular, local feature extraction is difficult in point cloud pairs with low overlap or repetitive structures, and rotation and other variations are not given sufficient attention, resulting in low registration accuracy.
A local feature enhancement method based on equivariant graph neural network is adopted. By using the LCF-Feature local context fusion module, the GB-global bilinear regularization module and the SE(3)-Transformer equivariant graph neural network module, local geometric information and global perception are fused to enhance the robustness and accuracy of point cloud registration.
It significantly improves the success rate of point cloud registration and reduces complexity, especially in complex geometric transformation scenarios, with higher registration accuracy and robustness, and can effectively handle large-angle transformations.
Smart Images

Figure CN120931697A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of point cloud registration technology, specifically relating to a point cloud registration method based on local feature enhancement of equivariant graph neural networks. Background Technology
[0002] 3D point cloud registration (PCR) is a technique for aligning two point clouds from different scans of the same scene. It is an important and fundamental problem in 3D computer vision, widely used in tasks such as localization, 3D reconstruction, and LiDAR SLAM mapping. It aims to estimate the six degrees of freedom (6-DoF) pose transformations of two 3D scans of the same object or scene to accurately align the two input point clouds. Traditional point cloud registration methods typically consist of three basic parts: feature extraction, keypoint matching, and rigid transformation estimation based on the corresponding keypoints. The latter's result mainly depends on the correspondence output by the former. The efficient and sufficient utilization of point cloud structural information plays a crucial role in the final rigid transformation estimation result. With the development of deep learning-based registration methods, using feature correspondences between points to solve the PCR problem has become a popular and reliable solution. However, in highly challenging real-world scenarios (e.g., point cloud pairs with low overlap or structural repetition), many existing point cloud registration processing networks still face some pressing problems that need to be addressed.
[0003] (1) Unlike keypoint matching in well-organized 2D images, point cloud data in 3D space suffers from a lack of texture, uneven density, and disordered organization, which increases the difficulty for neural networks to extract local features from the input data, making 3D matching tasks extremely challenging. Recent research data shows that in many existing networks, the feature information of given point cloud data is not fully utilized, resulting in a small number of inliers or a low inlier ratio during point cloud registration, thus reducing the success rate of point cloud registration.
[0004] (2) Capturing local geometric features of 3D point clouds is a crucial step in establishing the correspondence between source and target point clouds. However, in current point cloud local feature extraction methods, the inherent symmetry of point cloud data, including rotational isomorphism, has not received sufficient attention. Point cloud rotational isomorphism can be derived from... Figure 1 In simple terms, this hinders the model's effective learning, leading to a need for more training data and increasing model complexity to achieve better results. Summary of the Invention
[0005] The purpose of this invention is to provide a point cloud registration method based on local feature enhancement of equivariant graph neural networks, which solves the problems of low success rate and high complexity in existing point cloud registration technologies.
[0006] The technical solution adopted in this invention is a point cloud registration method based on local feature enhancement of equivariant graph neural networks, which is implemented according to the following steps:
[0007] Step 1: Construct an LCF-Feature local context fusion module based on point coordinates and point features;
[0008] Step 2: Construct the GB-global bilinear regularization module;
[0009] Step 3: Construct an equivariant graph neural network module based on SE(3)-Transformer;
[0010] Step 4: Construct the loss function and similarity matching.
[0011] The invention is further characterized in that,
[0012] Step 1 is implemented in the following steps:
[0013] The LCF-Feature local context fusion module, based on point coordinates and point features, collects more comprehensive local information by fusing local geometric information and feature context from the point cloud, thereby adding more local uniqueness to the point feature representation. The feature context is captured in parallel using the same constraints and formulas derived from the inherent 3D coordinates of the point cloud. The input is a point cloud containing N points and the basic point features within that point cloud. The 3D coordinates are represented by formula (2-1):
[0014]
[0015] Where P represents the coordinates of all scattered points in the point cloud, p i This represents the coordinates of any point in the point cloud. The dimension representing point coordinates is three-dimensional;
[0016] The point cloud feature map is represented by formula (2):
[0017]
[0018] Where F represents the point cloud feature map, f i Represents the point features of any point in a point cloud. The dimension representing the point feature is C. In summary, the point p in the point cloud... i The corresponding point features are express;
[0019] Based on the KNN algorithm or the ball-query algorithm, calculate and measure the 3D Euclidean distance between scattered points in the point cloud, and find a specific point p among the scattered points. i and all its neighboring nodes in This allows for the regional division of scattered points in the point cloud;
[0020] Then, by constructing point p i and neighboring nodes The edge relationship information between them is used to construct point p in three-dimensional space. i The local geometry of the graph, whose components are represented by formula (3):
[0021]
[0022] in Point p i Local geometry, Point p i Neighboring nodes within the region, Indicates the dimension of the graph;
[0023] In summary, the local geometry of all scattered points P in the point cloud. This can be expressed by formula (4):
[0024]
[0025] in, Representing the local geometry of a point cloud The dimension is N×k×6;
[0026] Then, the local geometry is encoded by a multilayer perceptron module (MLP) to obtain the local geometric context of the point cloud, as shown in formula (5):
[0027]
[0028] Where P represents the local geometric context of the point cloud, max represents the max pooling function, and M... (H) This represents a multilayer perceptron module (MLP), where C represents the feature dimension of a point in the point cloud.
[0029] Similarly, point p in the point cloud i Corresponding point features f i The local feature map in C-dimensional space is represented by formula (6):
[0030]
[0031] in, Representing point features f i Local feature map, This represents the point characteristics of node k, a neighbor of point i. The dimension of the local feature map is k×2C;
[0032] In the description above, and There is a corresponding relationship, that is yes Corresponding features;
[0033] According to formula (6), the local feature map of the point feature map F is obtained. The formula is shown in (7):
[0034]
[0035] Consistent with the processing flow of formula (5), the final point cloud local feature context F is obtained, as shown in formula (8):
[0036]
[0037] Among them, M (I) This represents another multilayer perceptron module (MLP) in the network, which encodes the local feature map F;
[0038] Finally, the obtained local geometric context P of the point cloud and the local feature context F of the point cloud are concatenated. The concatenated result is used as the output of the local context fusion module. The concatenation operation is shown in Equation (9):
[0039]
[0040] Where concat represents the concatenation operation, F L This represents the final output of the local context fusion module. F represents L The dimension is N×C.
[0041] The MLP module in step 1 includes a 1×1 convolution, a batch normalization (BN) layer, and an activation layer. It aggregates the local geometry graph by applying the maximum pooling function (MP) to the k neighboring nodes, and finally obtains the local geometric context of the point cloud.
[0042] Step 2 is implemented in the following steps:
[0043] The GB-Global Bilinear Regularization module was designed and proposed. The GB-Global Bilinear Regularization module utilizes global channel information and point descriptors to achieve global perception through element-level interdependence in the point cloud feature map. It aims to refine local features by considering the global perception of the entire point cloud.
[0044] First, the output F of the local context fusion module L Dimensionality reduction is performed using a weight matrix, which is represented by formula (10):
[0045]
[0046] Among them, W c This represents a weight matrix, where r represents the reduction factor;
[0047] Next, the ReLU activation function is used to process the dimension-reduced matrix to ensure the nonlinearity of the dimension-reduced matrix and simultaneously satisfy the nonnegativity requirement of formula (16). The mathematical definition of the ReLU activation function is described by formula (11):
[0048] f(x) = max(0, x) (11)
[0049] Where x represents the input to the ReLU activation function, and max represents the maximum value between the input x and 0, that is, when the input x is greater than 0, the output f(x) = x; when the input x is less than or equal to 0, the output f(x) = 0.
[0050] Then, average pooling is performed on the N elements along the space axis of the dimensionality-reduced matrix after ReLU activation, and the compressed spatial information is used as the global channel information descriptor g. c The operation process is represented by formula (12):
[0051]
[0052] Among them, F L W c F represents L and W c The matrix product, where avg represents average pooling, r represents the reduction factor, and r must satisfy r≥2;
[0053] Specifically, the global channel information descriptor g c The global mean response of each channel in the point cloud feature map is used as the representation, as shown in Equation (13):
[0054]
[0055] Where, μ j The global mean response of the j-th channel in the point cloud feature map;
[0056] Similarly, using another weight matrix W p The ReLU activation function and average pooling process C / r elements along the channel-axis to construct a global point descriptor g. p As shown in formula (14):
[0057]
[0058] Among them, F LW p F represents L and W p The matrix product, where avg represents average pooling, r represents the reduction factor, and r must satisfy r≥2;
[0059] Global point descriptor g p Similarly, the global mean response of each point in the point cloud feature map is used, as shown in Equation (15):
[0060]
[0061] Where, λ i The global mean response of the i-th point in the point cloud feature map;
[0062] Then, the already obtained global channel information descriptor g... c and global point descriptor g p Perform the outer product operation to fully integrate g. c With g p The square root operation is then performed on the outer product matrix to generate a low-rank global bilinear response. The entire process is shown in Equation (16):
[0063]
[0064] Where G represents the low-rank global bilinear response, and sqrt represents the square root operation. Indicates the calculation of the outer product;
[0065] In summary, the element η located in the i-th row and j-th column of the global bilinear response G... ij Calculate using formula (17):
[0066]
[0067] By recovering the channel dimension through two residual structures and a multilayer perceptron (MLP) module, a full-size global feature perception map F is generated. G As shown in formula (18):
[0068]
[0069] Among them, M ψ W represents a multilayer perceptron module. C and W P These represent two different weight matrices;
[0070] The average pooling operations in formulas (12) and (14) compress global information from the channel space and point space, respectively. The pooling mean usually represents the common patterns in the feature map. This invention compresses global information from the output F of the local context fusion module.L Subtract F G The learned features are sharpened in a way that is expressed by formula (19):
[0071]
[0072] Among them, F out σ represents the final output of the module, and σ represents an activation function.
[0073] Step 3 is implemented in the following steps:
[0074] The SE(3)-Transformer-based equivariant graph neural network module introduces an equivariant graph neural network to promote the aggregation of key point neighborhood features. At the same time, based on graph message passing technology, the SE(3) equivariant properties are embedded into the learned key point geometric feature descriptors to calculate the similarity between pairs of features in two frames for relative transformation prediction.
[0075] First, based on the key points of the paired source and target point clouds and their respective k neighboring points, a sphere query graph construction algorithm is applied to construct a graph G = (V, E) with vertices V and edges E. Here, the features of a single point... and single point coordinates These are respectively used as the graph node features and edge features in graph G, and the graph convolutional layer updates the edge equivariance information in each equivariance layer. Graph node features and embedded coordinate information The update methods are represented by formulas (20), (21), and (22), respectively:
[0076]
[0077] Where, φ m h represents a one-dimensional convolutional layer that handles edge-variable information message passing. i l This represents the feature information of node i in the l-layer graph of the network. This represents the feature information of node k in the l-layer graph of the network. This represents the coordinate information of node k in the l-layer graph of the network. This represents the coordinate information of node i in the l-layer graph of the network:
[0078]
[0079] Where, φ x This represents a one-dimensional convolutional layer that embeds coordinate information for message passing, where C represents the normalization factor, N(i) represents all neighboring nodes of graph node i, and proj F Indicates the local isotropic projection operation:
[0080]
[0081] Where, φ h A one-dimensional convolutional layer representing the message passing of graph node feature information.
[0082] In step 3, in formula (21), in order to limit the scope of information aggregation between graph nodes to the local context, for graph nodes... A neighborhood search is performed within a specific radius to find N(i) neighborhood feature descriptors for constructing edge features, thereby reducing the complexity of the graph feature adjacency matrix from O(ni) to O(ni). 2 The complexity is reduced to O(n), where... m ik The processing results in the local isovariant projection help to maintain the SO(3) invariance of the features, where F ik Part of the construction uses the pairwise coordinate embedding method proposed in ClofNet, as shown in Equation (23):
[0083]
[0084] Where, "||" represents the modulo operation on the vector;
[0085] In summary, the edge information m ik Projection information obtained after local isomorphic operations It is represented as a linear combination, consisting of scalar coefficients. To make adjustments, thereby maintaining the projection information. The SO(3) characteristic invariance is expressed by formula (24):
[0086]
[0087] Therefore, the projection information in formula (22) The summation operation is still based on vector summation, after passing through φ h One-dimensional convolutional layers retain their equivariance after processing.
[0088] Step 3 also includes the following steps:
[0089] The node features h that need to be processed i l and edge features m ik The feature descriptors are mapped to aggregate descriptors for similarity calculation between the source and target point clouds. In this process, it is first necessary to retain the feature h of each graph node in the source point cloud. i l Features of each graph node in the target point cloud map Then, each graph node feature is connected to the average coordinates of all nodes within its neighborhood. According to formula (21), the node coordinates have already embedded the edge information of the graph. Then, the new features connected in the source point cloud graph and the target point cloud graph are stacked to form the feature matrix of the source point cloud. and the feature matrix of the target point cloud
[0090] The feature matrix H of the source point cloud is processed using two stacked linear forward layers with low-rank constraints. src The feature matrix H of the target point cloud tar This reduces the number of features, resulting in a compressed feature matrix. and Effectively aggregate local information.
[0091] Step 4 is implemented in the following steps:
[0092] In the similarity matching process, the compressed feature matrix H' is used. src and H' tar To calculate the feature similarity score matrix and establish the feature correspondence between the source point cloud and the target point cloud, we first need to calculate H' src and H' tar The elements in the feature matrix are normalized, and then the dot product of the processed feature elements is performed to complete the feature similarity score matrix. The construction of the score matrix is performed, and each row of the score matrix is normalized to obtain... This involves viewing the matrix. The method of using traces to determine whether there are ambiguous point correspondences, similarity matrix Each row should contain a value close to 1, indicating that there is a strong matching relationship between the corresponding point features;
[0093] To eliminate outlier feature correspondences, each matched element Each step is verified by a full-rank check of a 7×7 submatrix centered on itself. After the full-rank check, to mask invalid similarity features in subsequent calculations, any invalid rows are... Set to zero, and according to the rank theorem, Rank(AB) ≤ min(Rank(A), Rank(B)), the similarity matrix... The final effective rank should satisfy r ≤ 35, based on the verified feature similarity score matrix. After pooling and fully connected layer processing, the translation vector is predicted. and rotation matrix
[0094] In geometric calculations, a multi-objective optimization approach is used to separate the error propagation paths of rotation and translation, and the rotation error RE loss L is used. rot Constraining the geometric consistency of the rotation matrix, using translation error TE loss L trans The spatial alignment error of the translation vector is monitored, and the regularization strength is balanced by the weighting coefficient β, as shown in Equation (25):
[0095]
[0096] The specified rotational error loss L rot Translation error loss L trans Measurements were taken in radians and meters respectively, with the rotational error loss L. rot The calculation formula is expressed by formula (26), where the translation error loss L trans The calculation formula is expressed by formula (27):
[0097]
[0098] Among them, R * and t * These represent the rotation matrix and translation vector used to transform the source point cloud into the target point cloud, respectively.
[0099] The beneficial effects of this invention are that, based on the point cloud registration method with local feature enhancement of equivariant graph neural networks, a point cloud registration method integrating local feature enhancement and global perception is proposed, which significantly improves the success rate and reduces the complexity. To generate more discriminative local descriptors, the local geometry and feature context are uniquely constrained by the parallel fusion of point coordinates and point features. Furthermore, the high-order space-channel dependency is captured by utilizing the bilinear combination of global channels and point descriptors, effectively injecting global perception, suppressing local noise, and enhancing feature coherence. At the same time, SE(3) equivariance is directly embedded in the key point feature aggregation, so that the learned descriptors have intrinsic invariance to rotation / translation, which greatly improves the robustness under large angle transformations and reduces training complexity. Compared with methods that rely on traditional graph convolution or heavy self-attention, this invention achieves higher registration accuracy and success rate with lower computational cost, especially in complex geometric transformation scenarios. Attached Figure Description
[0100] Figure 1 This is a schematic diagram of point cloud rotational isomorphism in the background art of this invention;
[0101] Figure 2 This is an architecture diagram of the point cloud registration method based on local feature enhancement of equivariant graph neural networks in this invention;
[0102] Figure 3 This is a flowchart of the point cloud registration method based on local feature enhancement of equivariant graph neural network according to the present invention;
[0103] Figure 4 This is the LCF-Feature local context fusion module in this invention;
[0104] Figure 5 This is the GB-global bilinear regularization module in this invention;
[0105] Figure 6 This is the equivariant graph neural network module based on SE(3)-Transformer in this invention;
[0106] Figure 7 This is the normalized feature similarity score matrix in this invention;
[0107] Figure 8(a) shows the input point cloud before registration in Scene 1 of the point cloud registration method based on local feature enhancement of equivariant graph neural network in this invention.
[0108] Figure 8(b) shows the input point cloud before registration in Scene 2 of the point cloud registration method based on local feature enhancement of equivariant graph neural network in this invention.
[0109] Figure 8(c) shows the input point cloud before registration in scene 3 of the point cloud registration method based on local feature enhancement of equivariant graph neural network in this invention;
[0110] Figure 8(d) shows the ground truth registration effect of scene 1 in the point cloud registration method based on local feature enhancement of equivariant graph neural network of the present invention;
[0111] Figure 8(e) shows the ground truth registration effect of scene 2 in the point cloud registration method based on local feature enhancement of equivariant graph neural network of the present invention;
[0112] Figure 8(f) shows the ground truth registration effect of scene 3 in the point cloud registration method based on local feature enhancement of equivariant graph neural network of the present invention;
[0113] Figure 8(g) shows the predicted registration effect of the method of the present invention in scenario 1 of the point cloud registration method based on local feature enhancement of equivariant graph neural network.
[0114] Figure 8(h) shows the predicted registration effect of the method of the present invention in scenario 2 of the point cloud registration method based on local feature enhancement of equivariant graph neural network.
[0115] Figure 8(i) shows the predicted registration effect of the method of the present invention in scenario 3 of the point cloud registration method based on local feature enhancement of equivariant graph neural network.
[0116] Figure 9(a) shows the module input in the point cloud registration method based on local feature enhancement of equivariant graph neural network in this invention;
[0117] Figure 9(b) is a t-SNE visualization of the module output features in Figure 8(a) of the point cloud registration method based on local feature enhancement of equivariant graph neural network in this invention.
[0118] Figure 9(c) shows the module rotation input in the point cloud registration method based on local feature enhancement of equivariant graph neural network in this invention;
[0119] Figure 9(d) is a t-SNE visualization of the module output features in Figure 8(c) of the point cloud registration method based on local feature enhancement of equivariant graph neural network in this invention. Detailed Implementation
[0120] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0121] This invention relates to a point cloud registration method based on local feature enhancement of equivariant graph neural networks, which is implemented according to the following steps:
[0122] Step 1: Construct an LCF-Feature local context fusion module based on point coordinates and point features;
[0123] Step 1 is implemented in the following steps:
[0124] In many point cloud registration tasks, finding the geometric correspondence between the source and target point clouds is a crucial step. Current research largely focuses on designing 3D features that capture discriminative local geometry from point clouds to establish this correspondence. While some previous methods used self-attention structures to estimate element dependencies, the additional computational burden of self-attention is practically undesirable due to the inherent complexity of point cloud data and the network itself.
[0125] To more fully utilize the feature information of the downsampled point cloud and obtain more accurate and representative local geometric 3D features, this invention proposes and designs an LCF-Feature local context fusion module based on point coordinates and point features, such as... Figure 4 As shown. The LCF-Feature local context fusion module, based on point coordinates and point features, collects more comprehensive local information by fusing local geometric information and feature context from the point cloud, thereby adding more local uniqueness to the point feature representation. The feature context is captured in parallel using the same constraints and formulas derived from the inherent 3D coordinates of the point cloud. The input is a point cloud containing N points and the basic point features within that point cloud. The 3D coordinates are represented by formula (2-1):
[0126]
[0127] Where P represents the coordinates of all scattered points in the point cloud, p iThis represents the coordinates of any point in the point cloud. The dimension representing point coordinates is three-dimensional;
[0128] The point cloud feature map is represented by formula (2):
[0129]
[0130] Where F represents the point cloud feature map, f i Represents the point features of any point in a point cloud. The dimension representing the point feature is C. In summary, the point p in the point cloud... i The corresponding point features are express;
[0131] Based on the KNN algorithm or the ball-query algorithm, calculate and measure the 3D Euclidean distance between scattered points in the point cloud, and find a specific point p among the scattered points. i and all its neighboring nodes in This allows for the regional division of scattered points in the point cloud;
[0132] Then, by constructing point p i and neighboring nodes The edge relationship information between them is used to construct point p in three-dimensional space. i The local geometry of the graph, whose components are represented by formula (3):
[0133]
[0134] in Point p i Local geometry, Point p i Neighboring nodes within the region, Indicates the dimension of the graph;
[0135] In summary, the local geometry of all scattered points P in the point cloud. This can be expressed by formula (4):
[0136]
[0137] in, Representing the local geometry of a point cloud The dimension is N×k×6;
[0138] Then, the local geometry is encoded by a multi-layer perceptron (MLP) module to obtain the local geometric context of the point cloud, as shown in formula (5):
[0139]
[0140] Where P represents the local geometric context of the point cloud, max represents the max pooling function, and M... (H) This represents a multilayer perceptron module (MLP), where C represents the feature dimension of a point in the point cloud.
[0141] Similarly, point p in the point cloud i Corresponding point features f i The local feature map in C-dimensional space is represented by formula (6):
[0142]
[0143] in, Representing point features f i Local feature map, This represents the point characteristics of node k, a neighbor of point i. The dimension of the local feature map is k×2C.
[0144] In the description above, and There is a corresponding relationship, that is yes The corresponding features.
[0145] According to formula (6), the local feature map of the point feature map F is obtained. The formula is shown in (7):
[0146]
[0147] Consistent with the processing flow of formula (5), the final point cloud local feature context F is obtained, as shown in formula (8):
[0148]
[0149] Among them, M (I) This represents another multilayer perceptron module (MLP) in the network, which encodes the local feature map F.
[0150] Finally, the obtained local geometric context P of the point cloud and the local feature context F of the point cloud are concatenated. The concatenated result is used as the output of the local context fusion module. The concatenation operation is shown in Equation (9):
[0151]
[0152] Where concat represents the concatenation operation, F L This represents the final output of the local context fusion module. F represents L The dimension is N×C.
[0153] Compared with the Edge Conv graph convolution method, the local context fusion module proposed in this invention uses the inherent three-dimensional geometric relationships of the point cloud to constrain the local geometric context and local feature context. It can simultaneously consider the geometric characteristics and feature information of points in the point cloud, improve the feature representation of each point in the point cloud, and thus have stronger expressive power in subsequent tasks.
[0154] The MLP module in step 1 includes a 1×1 convolution, a batch normalization (BN) layer, and an activation layer. It aggregates the local geometry graph by applying the max pooling function (MP) to k neighboring nodes, and finally obtains the local geometric context of the point cloud.
[0155] Step 2: Construct the GB-global bilinear regularization module;
[0156] Step 2 is implemented in the following steps:
[0157] The local context fusion module enhances the feature representation of points by focusing on local geometric and feature contexts in the point cloud, collecting more local details for each point's feature representation. However, while this module effectively improves the local feature representation of the point cloud, it still lacks awareness of global geometric structure and feature distribution, and local feature information may be affected by local noise, resulting in insufficient coherence of the overall features. To address these issues, this invention designs and proposes a GB-global bilinear regularization module, such as... Figure 5 As shown. The GB-Global Bilinear Regularization module utilizes global channel information and point descriptors to achieve global perception through element-level interdependencies in the point cloud feature map, aiming to refine local features by considering the global perception of the entire point cloud;
[0158] First, the output F of the local context fusion module L Dimensionality reduction is performed using a weight matrix, which is represented by formula (10):
[0159]
[0160] Among them, W c This represents a weight matrix, where r represents the reduction factor;
[0161] Next, the ReLU activation function is used to process the dimension-reduced matrix to ensure the nonlinearity of the dimension-reduced matrix and simultaneously satisfy the nonnegativity requirement of formula (16). The mathematical definition of the ReLU activation function is described by formula (11):
[0162] f(x) = max(0,x) (11)
[0163] Where x represents the input to the ReLU activation function, and max represents the maximum value between the input x and 0, that is, when the input x is greater than 0, the output f(x) = x; when the input x is less than or equal to 0, the output f(x) = 0.
[0164] Then, average pooling is performed on the N elements along the space axis of the dimensionality-reduced matrix after ReLU activation, and the compressed spatial information is used as the global channel information descriptor g. c The operation process is represented by formula (12):
[0165]
[0166] Among them, F L W c F represents L and W c The matrix product, where avg represents average pooling, r represents the reduction factor, and r must satisfy r≥2;
[0167] Specifically, the global channel information descriptor g c The global mean response of each channel in the point cloud feature map is used as the representation, as shown in Equation (13):
[0168]
[0169] Where, μ j The global mean response of the j-th channel in the point cloud feature map;
[0170] Similarly, similar to formula (12), another weight matrix W is used. p The ReLU activation function and average pooling process C / r elements along the channel-axis to construct a global point descriptor g. p As shown in formula (14):
[0171]
[0172] Among them, F L W p F represents L and W p The matrix product, where avg represents average pooling, r represents the reduction factor, and r must satisfy r≥2;
[0173] Global point descriptor g p Similarly, the global mean response of each point in the point cloud feature map is used, as shown in Equation (15):
[0174]
[0175] Where, λ i The global mean response of the i-th point in the point cloud feature map;
[0176] Then, the already obtained global channel information descriptor g... c and global point descriptor g p Perform the outer product operation to fully integrate g. c With g p The square root operation is then performed on the outer product matrix to generate a low-rank global bilinear response. The entire process is shown in Equation (16):
[0177]
[0178] Where G represents the low-rank global bilinear response, and sqrt represents the square root operation. Indicates the calculation of the outer product;
[0179] In summary, the element η located in the i-th row and j-th column of the global bilinear response G... ij Calculate using formula (17):
[0180]
[0181] The reason for using the global bilinear response can be explained from two aspects. Firstly, the global channel information descriptor captures channel-dependent feature information, while the global point descriptor contains contextual information about the overall shape of the point cloud. Using the outer product to compute the bilinear combination of the two global descriptors can fully preserve both types of global information. Secondly, due to λ... i and μ j These are the mean global responses of the i-th point and the j-th channel in the point cloud feature map, respectively. Therefore, the element η in the i-th row and j-th column of the global bilinear response... ij It is λ i and μ j The geometric mean is used to provide a high-order average response based on spatial and channel-related information.
[0182] By recovering the channel dimension through two residual structures and a multilayer perceptron (MLP) module, a full-size global feature perception map F is generated. G As shown in formula (18):
[0183]
[0184] Among them, M ψ W represents a multilayer perceptron module. C and W P These represent two different weight matrices;
[0185] The average pooling operations in formulas (12) and (14) compress global information from the channel space and point space, respectively. The pooling mean usually represents the common patterns in the feature map. This invention compresses global information from the output F of the local context fusion module. L Subtract F G The learned features are sharpened in a way that is expressed by formula (19):
[0186]
[0187] Among them, F out σ represents the final output of the module, and σ represents an activation function.
[0188] Step 3: Construct an equivariant graph neural network module based on SE(3)-Transformer;
[0189] Step 3 is implemented in the following steps:
[0190] To fully utilize the inherent symmetry and rotational equivariance of point cloud data and enhance the robustness of the local feature extraction process, this invention designs an equivariant graph neural network module based on SE(3)-Transformer, such as... Figure 6 As shown. The equivariant graph neural network module based on SE(3)-Transformer introduces the equivariant graph neural network proposed by Satorras et al. to promote the aggregation of neighborhood features of key points. At the same time, based on graph message passing technology, the SE(3) equivariant property is embedded into the learned key point geometric feature descriptor to calculate the similarity between pairs of features in two frames for relative transformation prediction.
[0191] First, based on the key points of the paired source and target point clouds and their respective k neighboring points, a sphere query graph construction algorithm is applied to construct a graph G = (V, E) with vertices V and edges E. Here, the features of a single point... and single point coordinates These are respectively used as the graph node features and edge features in graph G, and the graph convolutional layer updates the edge equivariance information in each equivariance layer. Graph node features and embedded coordinate information The update methods are represented by formulas (20), (21), and (22), respectively:
[0192]
[0193] Where, φ m h represents a one-dimensional convolutional layer that handles edge-variable information message passing. i l This represents the feature information of node i in the l-layer graph of the network. This represents the feature information of node k in the l-layer graph of the network. This represents the coordinate information of node k in the l-layer graph of the network. This represents the coordinate information of node i in the l-layer graph of the network:
[0194]
[0195] Where, φ x This represents a one-dimensional convolutional layer that embeds coordinate information for message passing, where C represents the normalization factor, N(i) represents all neighboring nodes of graph node i, and proj F Indicates the local isotropic projection operation:
[0196]
[0197] Where, φ h A one-dimensional convolutional layer representing the message passing of graph node feature information.
[0198] In step 3, in formula (21), in order to limit the scope of information aggregation between graph nodes to the local context, for graph nodes... A neighborhood search is performed within a specific radius to find N(i) neighborhood feature descriptors for constructing edge features, thereby reducing the complexity of the graph feature adjacency matrix from O(ni) to O(ni). 2 The computation time is reduced to O(n), effectively preventing information overflow during the aggregation of neighboring node information. In the formula... m ik The processing results in the local isovariant projection help to maintain the SO(3) invariance of the features, where F ik Part of the construction uses the pairwise coordinate embedding method proposed in ClofNet, as shown in Equation (23):
[0199]
[0200] Where, "||" represents the modulo operation on the vector;
[0201] In summary, the edge information m ik Projection information obtained after local isomorphic operations It is represented as a linear combination, consisting of scalar coefficients. To make adjustments, thereby maintaining the projection information. The SO(3) characteristic invariance is expressed by formula (24):
[0202]
[0203] Therefore, the projection information in formula (22) The summation operation is still based on vector summation, after passing through φ hOne-dimensional convolutional layers retain their equivariance after processing.
[0204] Step 3 also includes the following steps:
[0205] The node features h that need to be processed i l and edge features m ik The feature descriptors are mapped to aggregate descriptors for similarity calculation between the source and target point clouds. In this process, it is first necessary to retain the feature h of each graph node in the source point cloud. i l Features of each graph node in the target point cloud map Then, each graph node feature is connected to the average coordinates of all nodes within its neighborhood. According to formula (21), the node coordinates have already embedded the edge information of the graph. Then, the new features connected in the source point cloud graph and the target point cloud graph are stacked to form the feature matrix of the source point cloud. and the feature matrix of the target point cloud
[0206] To improve the reliability and computational efficiency of similarity matching, this invention uses two stacked linear forward layers with low-rank constraints to process the feature matrix H of the source point cloud. src The feature matrix H of the target point cloud tar This reduces the number of features, resulting in a compressed feature matrix. and Effectively aggregate local information.
[0207] Step 4: Construct the loss function and similarity matching.
[0208] Step 4 is implemented in the following steps:
[0209] In the similarity matching process, the compressed feature matrix H' is used. src and H' tar To calculate the feature similarity score matrix and establish the feature correspondence between the source point cloud and the target point cloud, we first need to calculate H' src and H' tar The elements in the feature matrix are normalized, and then the dot product of the processed feature elements is performed to complete the feature similarity score matrix. The construction of the score matrix is performed, and each row of the score matrix is normalized to obtain... This involves viewing the matrix. The method of using traces to determine whether there are ambiguous point correspondences, similarity matrix Each row should contain a value close to 1, indicating a strong matching relationship between the corresponding point features, such as... Figure 7 As shown:
[0210] To eliminate outlier feature correspondences, each matched element Each step is validated using a full-rank check of a 7×7 submatrix centered on itself. This validation process ensures local consistency in similarity matching, helping to identify globally consistent and reliable matching results from the similarity matrix. After the full-rank check, to mask invalid similarity features in subsequent calculations, any invalid rows are... Set to zero, and according to the rank theorem, Rank(AB) ≤ min(Rank(A), Rank(B)), the similarity matrix... The final effective rank should satisfy r ≤ 35, based on the verified feature similarity score matrix. After pooling and fully connected layer processing, the translation vector is predicted. and rotation matrix
[0211] To optimize the geometric transformation estimation capability of the prediction model and maintain its generalization ability in rigid body transformation scenarios, this invention employs a multi-objective optimization approach in geometric computation, separating the error propagation paths of rotation and translation, and using the rotation error RE loss L... rot Constraining the geometric consistency of the rotation matrix, using translation error TE loss L trans The spatial alignment error of the translation vector is monitored, and the regularization strength is balanced by the weighting coefficient β, as shown in Equation (25):
[0212]
[0213] Considering that the transformation between the source point cloud and the target point cloud is relative, in order to avoid numerical stability issues during training, the rotation error loss L is specified. rot Translation error loss L trans Measurements were taken in radians and meters respectively, with the rotational error loss L. rot The calculation formula is expressed by formula (26), where the translation error loss L trans The calculation formula is expressed by formula (27):
[0214]
[0215]
[0216] Among them, R * and t * These represent the rotation matrix and translation vector used to transform the source point cloud into the target point cloud, respectively.
[0217] Example 1
[0218] This invention relates to a point cloud registration method based on local feature enhancement of equivariant graph neural networks, which is implemented according to the following steps:
[0219] Step 1: Construct an LCF-Feature local context fusion module based on point coordinates and point features;
[0220] Step 2: Construct the GB-global bilinear regularization module;
[0221] Step 3: Construct an equivariant graph neural network module based on SE(3)-Transformer;
[0222] Step 4: Construct the loss function and similarity matching.
[0223] Example 2
[0224] The specific steps of the entire method are as follows:
[0225] (1) The input source point cloud and target point cloud are preprocessed using the voxel grid downsampling method, which reduces the computational load of the network while largely preserving the shape features of the original point cloud;
[0226] (2) The downsampled point cloud is processed using the GBE-Feature Global Bilinear Local Feature Enhancement Module, which enhances the basic point cloud feature representation by fusing more local context information;
[0227] (3) Use the FCGF (Fully Convolutional Geometric Features) network to process the point cloud after the basic point cloud features are enhanced, and complete the extraction of point cloud geometric feature descriptors;
[0228] (4) The Ball Query Graph Construction algorithm is used to generate the corresponding adjacency graph based on the coordinate information of the source point cloud and the target point cloud. The graph shows the relationship between each point in the point cloud and its k neighboring points.
[0229] (5) Use EGNN (Equivariant Graph Neural Networks) E(n) to process the relevant data of the source point cloud and the target point cloud respectively. Perform graph convolution operation on the extracted point cloud geometric features, point cloud coordinate information, adjacency graph edge feature information and edge information to ensure the equivariance of point cloud features.
[0230] (6) The feature information and coordinate information of the source point cloud and the target point cloud are spliced and fused to obtain new point cloud features. Then, the latest point cloud features are reduced in dimensionality by low-rank constraint to effectively remove redundant and irrelevant features and optimize the model's expressive ability.
[0231] (7) The feature information of the source point cloud and the target point cloud is fused, and the similarity between pairs of features is calculated to perform relative transformation prediction;
[0232] (8) Apply the transformation prediction results to the source point cloud, perform rotation and translation operations on the source point cloud, and finally complete the registration of the source point cloud and the target point cloud; the specific process is as follows: Figure 3 As shown.
[0233] Example 3
[0234] Experimental comparison and analysis
[0235] To verify the effectiveness of the method proposed in this invention, relevant experiments were conducted and the experimental results were analyzed.
[0236] Experimental setup
[0237] (1) Dataset
[0238] The experimental portion of this invention utilizes the 3DMatch public dataset and follows standard evaluation protocols to prepare training and testing data. This dataset, collected by a high-precision 3D scanning device, comprises data from 62 scenes. It is divided into training, validation, and testing sets, with data from 46 scenes used for training, 8 scenes for validation, and 8 scenes for evaluation. The original 3DMatch dataset includes two .txt files and multiple seq folders. Each seq folder contains multiple frames of .color.png, .depth.png, and .pose.txt files, but does not inherently include point cloud data. This invention uses the point cloud generation method in 3DLocalMultiViewDesc, fusing 50 consecutive depth frames to obtain point cloud fragments, which are then saved as .ply files. Taking the 8 scenes in the test set as an example, most scan pairs in the official 3DMatch test set have an overlap greater than 30%, with only a small number of extreme cases having an overlap less than 7%. The distribution of its point cloud samples is shown in Table 1.
[0239] Table 1. Sample distribution of the 3DMatch test set
[0240] name Point cloud quantity Number of Pairs 7-scenes-redkitchen 60 506 sun3d-home_at-home_at_scan1_2013_jan_1 60 156 sun3d-home_md-home_md_scan9_2012_sep_30 60 208 sun3d-hotel_uc-scan3 55 226 sun3d-hotel_umd-maryland_hotel1 57 104 sun3d-hotel_umd-maryland_hotel3 37 54 sun3d-mit_76_studyroom-76-1studyroom2 66 292 sun3d-mit_lab_hj-lab_hj_tea_nov_2_2012_scan1_erika 38 77
[0241] (2) Evaluation indicators
[0242] The experiment uses translation error (TE) and rotation error (RE) in PointDSC to evaluate the accuracy of the predicted pose error in successful registration. The formulas for calculating TE and RE can be expressed by formula (25):
[0243]
[0244] Among them, R * and t * Let represent the actual rotation matrix and translation vector, respectively. Tr(·) is the trace of the matrix, and ||·||2 represents the Euclidean norm, which is used to calculate the Euclidean distance between two translation vectors.
[0245] In addition, the experiment also included Registration Recall (RR) and F1 score as evaluation indicators for registration performance. RR is mainly used to measure the proportion of correctly matched point pairs in the registration algorithm. Root Mean Square Error (RMSE) is used as the standard for judging correct registration and is a core indicator in the RR calculation process. Only when RMSE is lower than a specified threshold τ will it be recorded as a successful match. The formula for calculating the registration recall value δ is shown in formula (26):
[0246]
[0247] Among them, Ω * Indicates the number of matched point pairs, (x * ,y * ) represents a pair of real matching points, x * y represents a point in the source point cloud. * This represents the corresponding point in the target point cloud. Represents the predicted rotation matrix. Represents the predicted translation vector, with the symbol... An indicator that shows the degree to which conditions are met.
[0248] Remove the condition check part from formula (26) and convert it to the standard root mean square error (RMSE), as shown in formula (27).
[0249]
[0250] The F1 score reflects the overall matching quality, and its calculation formula is shown in formula (28):
[0251]
[0252] Where Precision represents the percentage of correct matches among all predicted matching pairs, and Recall represents the percentage of matching pairs successfully found by the algorithm among all correct matches.
[0253] (3) Parameter settings
[0254] In this experiment, a voxel grid downsampling method was proposed for point cloud preprocessing. A voxel size of 5cm was used for initial downsampling of the point cloud, generating 1024 sampling points. Translation error (TE) of 30cm and rotation error (RE) of 15° were set as thresholds for successful registration, and the threshold τ in formula (26) was set to 10cm. The dimension of the initially extracted feature descriptors was set to 32 for subsequent graph neural network feature learning, and the number of neighboring nodes for each node was set to 16 to construct the 3DMatch ball query graph. During the graph neural network feature learning process, the graph node feature dimension was set to 32, and the edge feature dimension was set to 3.
[0255] (4) Experimental Environment
[0256] This section demonstrates the experimental results of the proposed method through experimental verification and performance analysis. To ensure the accuracy and reproducibility of the results, the specific environmental configuration required for the experiment is shown in Table 2. The experiments in this section used Python 3.8.20 to write the program and were conducted on a computer configured with a 2.50GHz CPU, 16GB of RAM, a GeForce RTX 3050 graphics card, and a Windows 10 operating system.
[0257] Table 2 Experimental Environment Configuration
[0258] Configuration Specification operating system Windows 10 GPU NVIDIA GeForce RTX 3050 CPU Intel(R)Core(TM)i7-12400F CPU@2.50GHz Memory 16 GB Python 3.8.20 CUDA 11.8 PyTorch 2.0.0
[0259] Example 4
[0260] Experimental comparison and analysis
[0261] To verify the effectiveness of the proposed method in point cloud registration tasks, this section presents experimental evaluations on the pairwise registration task of the 3DMatch dataset. This includes using different feature descriptors to extract local and global features from point clouds, including the deep learning-based 3D feature descriptor FCGF and the traditional hand-designed feature descriptor FPFH (Fast Point Feature Histogram). This invention first selects six representative traditional methods: FCR, SM, RANSAC, GC-RANSAC, Go-ICP, and Super4PCS, as well as the state-of-the-art geometry-based method TEASER. For deep learning-based methods, this invention selects DGR, 3DRegNet, D3Feat, SpinNet, PointDSC, and RoReg for comparison. The evaluation results of all the above registration methods on the 3DMatch dataset are shown in Table 3. In the table, * indicates that the original model lacks support for feature descriptors, bold text represents the best performance among all compared methods, and underlined text represents the second-best performance.
[0262] Table 3 Evaluation results of different registration methods based on FCGF on the 3DMatch dataset.
[0263] method RR (%)↑ RE(°)↓ TE(cm)↓ F1(%)↑ Time(s)↓ FGR 78.56 2.82 8.36 - 0.76 SM 86.57 2.29 7.07 48.21 0.03 RANSAC-1k 86.57 3.16 9.67 76.62 0.08 RANSAC-10k 90.70 2.69 8.25 80.76 0.58 RANSAC-100k 91.50 2.49 7.54 81.43 5.50 RANSAC-100k+refine 92.30 2.17 6.76 81.43 5.51 GC-RANSAC-100k 92.05 2.33 7.11 75.69 0.47 Go-ICP 22.95 5.38 14.70 20.08 771.0 Super4PCS 21.6 5.25 14.10 19.86 4.55 TEASER 85.77 2.73 8.66 73.96 0.11 DGR 91.30 2.40 7.48 89.76 1.36 3DRegNet 77.76 2.74 8.13 58.33 0.05 D3Feat* 89.79 2.57 8.16 87.40 0.14 SpinNet* 93.74 1.93 6.24 92.07 2.84 PointDSC 93.28 2.06 6.55 89.35 0.09 RoReg* 93.70 1.84 6.28 91.60 2226 VRHCF 96.21 1.81 6.02 87.76 - Method of the present invention 96.13 1.67 5.54 94.25 0.12
[0264] As shown in Table 3, the method of this invention exhibits excellent performance under the FCGF feature descriptor-based setting, demonstrating the lowest translation error (TE) of 5.54 cm, the lowest rotation error (RE) of 1.67°, and the best F1 score of 94.25%. In the point cloud registration experiment on the 3DMatch dataset, the method of this invention outperforms all the comparative methods proposed in this invention in most evaluation metrics, only being 0.09 seconds slower on average compared to the traditional SM model. SpinNet and RoReg, ranking second and third in F1 score, have the smallest registration error and the highest registration score, demonstrating the necessity of fusing rotation invariance to point cloud features during point cloud registration and highlighting the advantages of fusing rotation features.
[0265] To more fully illustrate the registration performance of the method of this invention under different feature descriptors, this invention further replaces the deep learning-based 3D feature descriptor FCGF with the traditional hand-designed feature descriptor FPFH, and conducts experimental evaluations on the 3DMatch dataset against all comparison methods. The experimental results are shown in Table 4. Unlike the comparison methods shown in Table 3, the traditional methods Go-ICP and Super4PCS, as well as the deep learning-based methods D3Feat, SpinNet, RoReg, and VRHCF, did not have evaluation results on the 3DMatch dataset based on the feature descriptor FPFH found in relevant point cloud registration papers.
[0266] Table 4 Evaluation results of different registration methods based on FPFH on the 3DMatch dataset
[0267]
[0268]
[0269] As shown in Table 4, under the traditional manually designed feature descriptor FPFH, this invention still demonstrates reasonable performance compared to other methods, exhibiting minimal rotation error (RE) and translation error (TE). Specifically, RANSAC-1k and RANSAC-10k show significant performance degradation when using the feature descriptor FPFH to establish input correspondences. This is because FPFH is sensitive to local geometric changes and easily affected by outliers, resulting in poor robustness. RANSAC-100k, however, minimizes performance loss at the cost of higher computation time. In summary, when using descriptors with limited feature representation capabilities, other methods exhibit a significant performance decrease compared to the method of this invention, verifying the strong robustness of this invention compared to the weaker feature descriptor FPFH and effectively reducing erroneous correspondence interference caused by feature descriptor ambiguity.
[0270] Different downsampling input points have a significant impact on the geometric representation ability and robustness of point cloud registration methods. To verify the registration performance of the method of the present invention under different sampling input point numbers, the present invention constructs sparse comparison experiments on input point cloud downsampling levels under the setting of FCGF feature descriptor. Sparse tests are performed on deep learning-based registration models D3Feat, PointDSC, and SpinNet under the settings of 4096, 2048, 1024, 512, and 256 point cloud downsampling points, respectively. The experimental results are shown in Table 5.
[0271] Table 5. Evaluation results of the number of 3DMatch point cloud downsampled samples on the registration recall (RR) of different registration methods.
[0272] method 4096 2048 1024 512 256 average value D3Feat 91.9 90.4 89.8 86.0 82.5 88.1 PointDSC 94.9 95.1 93.2 90.4 86.5 92.0 SpinNet 93.8 93.6 93.7 89.5 85.7 91.3 VRHCF 96.3 95.9 96.2 93.7 88.3 94.1 Method of the present invention 96.5 96.3 96.1 92.4 89.2 94.1
[0273] As shown in Table 5, the method of this invention can enhance the registration performance of the model by increasing the marginal benefit when using more points. It achieves excellent results for different numbers of sampling points, exhibiting optimal performance when the number of sampling points in the point cloud is 4096, 2048, and 256, which fully demonstrates that the method of this invention is robust to the number of sampling points in the point cloud. However, the method of this invention is not as good as the VRHCF registration method when the number of sampling points is 1024 and 512. This is because the sparsity of the point cloud weakens the local feature aggregation ability, while VRHCF maintains stronger spatial consistency at low sampling rates through hierarchical filtering.
[0274] Considering the need to strike an optimal balance between accuracy and computational efficiency, the method of this invention ultimately selects 1024 as the number of point cloud downsampling points, because compared with the number of point cloud downsampling points of 2048 and 4096, the performance difference of 1024 point cloud downsampling points is only within 1%.
[0275] Example 5
[0276] Visualization effect comparison analysis
[0277] To more intuitively evaluate the registration effect of the method of the present invention in complex scenarios, the present invention uses real pose data provided in the 3DMatch benchmark test set to perform a visual comparative analysis with the rotation matrix and translation vector predicted by the method of the present invention.
[0278] The comparison results are shown in Figure 8. In Figure 8, column (a) represents the two input point clouds before registration, including the source point cloud and the target point cloud. The points in the source point cloud and the target point cloud are rendered in yellow and blue, respectively; column (b) represents the effect of transforming the input point clouds according to the registration ground truth provided by the 3DMatch benchmark set. The source point cloud P src By superimposing rotation matrices Translation vector Spatial transformation is performed to generate registered point clouds. Target point cloud preservation P tar The original coordinates remain unchanged, while the geometric curvature is increased by coloring the normal direction. Points in the source point cloud and points transformed into the target point cloud by rotation and translation are rendered in yellow. Column (c) shows the registration effect of the input point cloud after the predicted transformation by the method of this invention, and the positions that are significantly different from the true value registration effect are highlighted with a solid red line.
[0279] As can be seen from the experimental registration results, the point cloud registration method based on local feature enhancement of equivariant graph neural networks proposed in this invention achieves excellent registration results in the point cloud registration task of the 3DMatch dataset. It maintains accurate pose alignment even in regions with complex geometric structures, showing only a slight difference compared to the actual registration results. This further verifies the effectiveness and practicality of the proposed method in solving the problem of point cloud registration from different perspectives within the same scene, providing strong support for research and application in the field of point cloud registration.
[0280] To more intuitively demonstrate the learning ability of the method of this invention regarding isovariant features, this invention uses t-SNE graph visualization to examine the output features of the isovariant graph neural network module, thereby proving the rotational isovariance of the module's output features, such as... Figures 9(a) to 9(d) As shown, Figure 9(a) is the module input in this invention; Figure 9(b) is the t-SNE visualization effect of the module output feature of Figure 9(a) in this invention; Figure 9(c) is the module rotation input in this invention; Figure 9(d) is the t-SNE visualization effect of the module output feature of Figure 9(c) in this invention.
[0281] Example 6
[0282] ablation experiment
[0283] To verify the effectiveness of each module proposed in the method of this invention, this invention uses 3DMatch as the test dataset and designs ablation experiments with different configurations and combinations of modules in the proposed model.
[0284] (1) Ablation experiments of feature descriptor module and equivariant graph neural network module
[0285] This experiment aims to verify the effectiveness of the corresponding modules by removing or replacing the feature descriptor module and the equivariant graph neural network module, and to determine the contribution of the module based on the model performance reflected by relevant indicators. The experimental results are shown in Table 6.
[0286] Table 6 Ablation experiments of the feature descriptor module and the isotropic graph neural network module on the 3DMatch dataset.
[0287]
[0288]
[0289] As shown in Table 6, removing the feature descriptor module leads to a significant drop in model performance (RR = 61.39%, RE = 10.26°), demonstrating the importance of point cloud geometric feature extraction for subsequent correspondence establishment and illustrating the effectiveness of the geometric feature extraction by the feature descriptor module in the method of this invention. However, removing the equivariant graph neural network module only reduces the registration time by 0.06s while sacrificing 35% of registration performance, exhibiting a significant performance-efficiency imbalance. Replacing it with a regular graph convolutional neural network only partially recovers the performance (RR = 68.52%, RE = 8.32°). The regular graph convolutional neural network impairs the rotational convergence in rigid transformations, highlighting the unique advantages of equivariant embedding features in rigid transformations. Although the equivariant layer introduces a slight computational delay, the overall performance gain is significant. In summary, the equivariant graph neural network module and the feature descriptor module are key to the performance improvement of the method of this invention. Through their synergistic effect, local feature discriminability and global geometric consistency are achieved during point cloud feature extraction.
[0290] (2) Comparative analysis of KNN graph construction algorithm and ball query graph construction algorithm
[0291] This experiment aims to compare and analyze the impact of two different graph construction algorithms in the equivariant graph neural network module on the model registration performance of the method of this invention, in order to find the most effective graph construction algorithm. Among them, the ball query graph construction algorithm is the graph construction algorithm used in the method of this invention. The comparative experimental results are shown in Table 7.
[0292] Table 7 compares the KNN graph construction algorithm and the ball query graph construction algorithm on the 3DMatch dataset.
[0293] method RR (%)↑ RE(°)↓ TE(cm)↓ F1(%)↑ Time(s)↓ KNN graph construction algorithm 92.37 1.72 5.31 93.74 0.16 Ball Query Graph Construction Algorithm 94.60 1.67 5.68 94.35 0.12
[0294] As shown in Table 7, the sphere query graph construction algorithm outperforms the KNN graph construction algorithm in both efficiency and performance. This is primarily because the fixed-radius neighborhood search strategy of the sphere query graph construction algorithm significantly reduces computational complexity, maintains spatial scale consistency of local geometry, facilitates the learning of stable features, and exhibits strong robustness to noise and density variations. In contrast, the neighborhood range of the KNN graph construction algorithm changes with point cloud density, thus affecting its generalization ability. Therefore, the sphere query graph construction algorithm used in this invention demonstrates superior geometric modeling capabilities in point cloud registration tasks.
[0295] (3) Ablation experiments of local context fusion module and global linear regularization module
[0296] This experiment aims to analyze and evaluate the contributions of the local context fusion module and the global linear regularization module to the model performance in the point cloud registration method of this invention. The experimental results are shown in Table 8, which shows the performance changes of the model after removing the two feature processing modules in the registration method of this invention.
[0297] Table 8 Ablation experiments of the local context fusion module and the global linear regularization module on the 3DMatch dataset.
[0298] method RR (%)↑ RE(°)↓ TE(cm)↓ F1(%)↑ Time(s)↓ Remove local context blending module 76.45 6.41 7.92 78.09 0.10 Remove global linear regularization module 87.76 2.58 6.93 88.36 0.07 Method of the present invention 96.13 1.67 5.54 94.25 0.12
[0299] As shown in Tables 2-8, removing the local context fusion module significantly reduced model performance (RR = 76.45%, RE = 6.41°). This indicates that the module effectively enhanced the local geometric structure modeling capability through multi-scale neighborhood feature aggregation, and its absence led to an overall decrease in the registration performance of the model in this invention. Further removing the global linear regularization module separately, although the model performance also decreased significantly compared to the complete model (RR = 87.76%, RE = 2.58°), was still significantly better than removing only the local context fusion module. This suggests that the global linear regularization module mainly optimizes global features by constraining the geometric consistency of the local feature space, and its performance depends on the discriminative basis of the local features. In summary, the local context fusion module and the global linear regularization module jointly solve the feature enhancement problem from local features to global features, providing a foundation for high-precision point cloud registration using the method of this invention.
[0300] In the theoretical framework of point cloud registration algorithms, the completeness and discriminativeness of local feature extraction from point cloud data constitute key constraints for improving subsequent registration accuracy. However, current local feature extraction methods do not fully utilize the rotational equivariance and input feature information of point cloud data itself, thus limiting the generalization ability of local features in registration tasks. To address this issue, this invention proposes a point cloud registration method based on local feature enhancement using equivariant graph neural networks. On the one hand, a dual-stream collaborative mechanism based on 3D coordinate constraints and a global bilinear regularization module are introduced, significantly enhancing feature representation capability and spatial distribution consistency. On the other hand, an SE(3)-equivariant graph neural network model is introduced, which learns local features from two frames through feature descriptors, embeds equivariant features through layer fusion, and calculates similarity scores to estimate the point cloud pose transformation matrix. Comparative experiments on the 3DMatch dataset demonstrate the good performance of this method.
Claims
1. A point cloud registration method based on local feature enhancement of equivariant graph neural networks, characterized in that, The specific steps are as follows: Step 1: Construct an LCF-Feature local context fusion module based on point coordinates and point features; Step 2: Construct the GB-global bilinear regularization module; Step 3: Construct an equivariant graph neural network module based on SE(3)-Transformer; Step 4: Construct the loss function and similarity matching.
2. The point cloud registration method based on local feature enhancement of equivariant graph neural networks according to claim 1, characterized in that, Step 1 is implemented in the following steps: The LCF-Feature local context fusion module, based on point coordinates and point features, collects more comprehensive local information by fusing local geometric information and feature context from the point cloud, thereby adding more local uniqueness to the point feature representation. The feature context is captured in parallel using the same constraints and formulas derived from the inherent 3D coordinates of the point cloud. The input is a point cloud containing N points and the basic point features within that point cloud. The 3D coordinates are represented by formula (2-1): Where P represents the coordinates of all scattered points in the point cloud, p i This represents the coordinates of any point in the point cloud. The dimension representing point coordinates is three-dimensional; The point cloud feature map is represented by formula (2): Where F represents the point cloud feature map, f i Represents the point features of any point in a point cloud. The dimension representing the point feature is C. In summary, the point p in the point cloud... i The corresponding point features are express; Based on the KNN algorithm or the ball-query algorithm, calculate and measure the 3D Euclidean distance between scattered points in the point cloud, and find a specific point p among the scattered points. i and all its neighboring nodes in This allows for the regional division of scattered points in the point cloud; Then, by constructing point p i and neighboring nodes The edge relationship information between them is used to construct point p in three-dimensional space. i The local geometry of the graph, whose components are represented by formula (3): in Point p i Local geometry, Point p i Neighboring nodes within the region, Indicates the dimension of the graph; In summary, the local geometry of all scattered points P in the point cloud. This can be expressed by formula (4): in, Representing the local geometry of a point cloud The dimension is N×k×6; Then, the local geometry is encoded by a multilayer perceptron module (MLP) to obtain the local geometric context of the point cloud, as shown in formula (5): Where P represents the local geometric context of the point cloud, max represents the max pooling function, and M... (H) This represents a multilayer perceptron module (MLP), where C represents the feature dimension of a point in the point cloud. Similarly, point p in the point cloud i Corresponding point features f i The local feature map in C-dimensional space is represented by formula (6): in, Representing point features f i Local feature map, This represents the point characteristics of node k, a neighbor of point i. The dimension of the local feature map is k×2C; In the description above, and There is a corresponding relationship, that is yes Corresponding features; According to formula (6), the local feature map of the point feature map F is obtained. The formula is shown in (7): Consistent with the processing flow of formula (5), the final point cloud local feature context F is obtained, as shown in formula (8): Among them, M (I) This represents another multilayer perceptron module (MLP) in the network, which processes local feature maps. Encode; Finally, the obtained local geometric context P of the point cloud and the local feature context F of the point cloud are concatenated. The concatenated result is used as the output of the local context fusion module. The concatenation operation is shown in Equation (9): Where concat represents the concatenation operation, F L This represents the final output of the local context fusion module. F represents L The dimension is N×C.
3. The point cloud registration method based on local feature enhancement of equivariant graph neural networks according to claim 2, characterized in that, The MLP module in step 1 includes a 1×1 convolution, a batch normalization (BN) layer, and an activation layer. It aggregates the local geometry graph by applying the maximum pooling function (MP) to k neighboring nodes, and finally obtains the local geometric context of the point cloud.
4. The point cloud registration method based on local feature enhancement of equivariant graph neural networks according to claim 2, characterized in that, Step 2 is implemented in the following steps: The GB-Global Bilinear Regularization module was designed and proposed. The GB-Global Bilinear Regularization module utilizes global channel information and point descriptors to achieve global perception through element-level interdependence in the point cloud feature map. It aims to refine local features by considering the global perception of the entire point cloud. First, the output F of the local context fusion module L Dimensionality reduction is performed using a weight matrix, which is represented by formula (10): Among them, W c This represents a weight matrix, where r represents the reduction factor; Next, the ReLU activation function is used to process the dimension-reduced matrix to ensure the nonlinearity of the dimension-reduced matrix and simultaneously satisfy the nonnegativity requirement of formula (16). The mathematical definition of the ReLU activation function is described by formula (11): Where x represents the input to the ReLU activation function, and max represents the maximum value between the input x and 0, that is, when the input x is greater than 0, the output f(x) = x; when the input x is less than or equal to 0, the output f(x) = 0. Then, average pooling is performed on the N elements along the space axis of the dimensionality-reduced matrix after ReLU activation, and the compressed spatial information is used as the global channel information descriptor g. c The operation process is represented by formula (12): Among them, F L W c F represents L and W c The matrix product, where avg represents average pooling, r represents the reduction factor, and r must satisfy r≥2; Specifically, the global channel information descriptor g c The global mean response of each channel in the point cloud feature map is used as the representation, as shown in Equation (13): Where, μ j The global mean response of the j-th channel in the point cloud feature map; Similarly, using another weight matrix W p The ReLU activation function and average pooling process C / r elements along the channel-axis to construct a global point descriptor g. p As shown in formula (14): Among them, F L W p F represents L and W p The matrix product, where avg represents average pooling, r represents the reduction factor, and r must satisfy r≥2; Global point descriptor g p Similarly, the global mean response of each point in the point cloud feature map is used, as shown in Equation (15): Where, λ i The global mean response of the i-th point in the point cloud feature map; Then, the already obtained global channel information descriptor g... c and global point descriptor g p Perform the outer product operation to fully integrate g. c With g p The square root operation is then performed on the outer product matrix to generate a low-rank global bilinear response. The entire process is shown in Equation (16): Where G represents the low-rank global bilinear response, and sqrt represents the square root operation. Indicates the calculation of the outer product; In summary, the element η located in the i-th row and j-th column of the global bilinear response G... ij Calculate using formula (17): By recovering the channel dimension through two residual structures and a multilayer perceptron (MLP) module, a full-size global feature perception map F is generated. G As shown in formula (18): Among them, M ψ W represents a multilayer perceptron module. C and W P These represent two different weight matrices; The average pooling operations in formulas (12) and (14) compress global information from the channel space and point space, respectively. The pooling mean usually represents the common patterns in the feature map. This invention compresses global information from the output F of the local context fusion module. L Subtract F G The learned features are sharpened in a way that is expressed by formula (19): Among them, F out σ represents the final output of the module, and σ represents an activation function.
5. The point cloud registration method based on local feature enhancement of equivariant graph neural networks according to claim 4, characterized in that, Step 3 is implemented in the following steps: The SE(3)-Transformer-based equivariant graph neural network module introduces an equivariant graph neural network to promote the aggregation of key point neighborhood features. At the same time, based on graph message passing technology, the SE(3) equivariant properties are embedded into the learned key point geometric feature descriptors to calculate the similarity between pairs of features in two frames for relative transformation prediction. First, based on the key points of the paired source and target point clouds and their respective k neighboring points, a sphere query graph construction algorithm is applied to construct a graph G = (V, E) with vertices V and edges E. Here, the features of a single point... and single point coordinates These are respectively used as the graph node features and edge features in graph G, and the graph convolutional layer updates the edge equivariance information in each equivariance layer. Graph node features and embedded coordinate information The update methods are represented by formulas (20), (21), and (22), respectively: Where, φ m h represents a one-dimensional convolutional layer that handles edge-variable information message passing. i l This represents the feature information of node i in the l-layer graph of the network. This represents the feature information of node k in the l-layer graph of the network. This represents the coordinate information of node k in the l-layer graph of the network. This represents the coordinate information of node i in the l-layer graph of the network: Where, φ x This represents a one-dimensional convolutional layer that embeds coordinate information for message passing, where C represents the normalization factor, N(i) represents all neighboring nodes of graph node i, and proj F Indicates the local isotropic projection operation: Where, φ h A one-dimensional convolutional layer representing the message passing of graph node feature information.
6. The point cloud registration method based on local feature enhancement of equivariant graph neural networks according to claim 5, characterized in that, In step 3, in formula (21), in order to limit the scope of information aggregation between graph nodes to the local context, for graph nodes... A neighborhood search is performed within a specific radius to find N(i) neighborhood feature descriptors for constructing edge features, thereby reducing the complexity of the graph feature adjacency matrix from O(ni) to O(ni). 2 The complexity is reduced to O(n), where... m ik The processing results in the local isovariant projection help to maintain the SO(3) invariance of the features, where F ik Part of the construction uses the pairwise coordinate embedding method proposed in ClofNet, as shown in Equation (23): Here, "||" represents the modulo operation on the vector; In summary, the edge information m ik Projection information obtained after local isomorphic operations It is represented as a linear combination, consisting of scalar coefficients. To make adjustments, thereby maintaining the projection information. The SO(3) characteristic invariance is expressed by formula (24): Therefore, the projection information in formula (22) The summation operation is still based on vector summation, after passing through φ h One-dimensional convolutional layers retain their equivariance after processing.
7. The point cloud registration method based on local feature enhancement of equivariant graph neural networks according to claim 6, characterized in that, Step 3 further includes the following steps: The node features h that need to be processed i l and edge features m ik The feature descriptors are mapped to aggregate descriptors for similarity calculation between the source and target point clouds. In this process, it is first necessary to retain the feature h of each graph node in the source point cloud. i l Features of each graph node in the target point cloud map Then, each graph node feature is connected to the average coordinates of all nodes within its neighborhood. According to formula (21), the node coordinates have already embedded the edge information of the graph. Then, the new features connected in the source point cloud graph and the target point cloud graph are stacked to form the feature matrix of the source point cloud. and the feature matrix of the target point cloud The feature matrix H of the source point cloud is processed using two stacked linear forward layers with low-rank constraints. src The feature matrix H of the target point cloud tar This reduces the number of features, resulting in a compressed feature matrix. and Effectively aggregate local information.
8. The point cloud registration method based on local feature enhancement of equivariant graph neural networks according to claim 7, characterized in that, Step 4 is implemented in the following steps: In the similarity matching process, the compressed feature matrix H' is used. src and H' tar To calculate the feature similarity score matrix and establish the feature correspondence between the source point cloud and the target point cloud, we first need to calculate H' src and H' tar The elements in the feature matrix are normalized, and then the dot product of the processed feature elements is performed to complete the feature similarity score matrix. The construction of the score matrix is performed, and each row of the score matrix is normalized to obtain... This involves viewing the matrix. The method of using traces to determine whether there are ambiguous point correspondences, similarity matrix Each row should contain a value close to 1, indicating that there is a strong matching relationship between the corresponding point features; To eliminate outlier feature correspondences, each matched element Each step is verified by a full-rank check of a 7×7 submatrix centered on itself. After the full-rank check, to mask invalid similarity features in subsequent calculations, any invalid rows are... Set to zero, and according to the rank theorem, Rank(AB) ≤ min(Rank(A), Rank(B)), the similarity matrix... The final effective rank should satisfy r ≤ 35, based on the verified feature similarity score matrix. After pooling and fully connected layer processing, the translation vector is predicted. and rotation matrix In geometric calculations, a multi-objective optimization approach is used to separate the error propagation paths of rotation and translation, and the rotation error RE loss L is used. rot Constraining the geometric consistency of the rotation matrix, using translation error TE loss L trans The spatial alignment error of the translation vector is monitored, and the regularization strength is balanced by the weighting coefficient β, as shown in Equation (25): The specified rotational error loss L rot Translation error loss L trans Measurements were taken in radians and meters respectively, with the rotational error loss L. rot The calculation formula is expressed by formula (26), where the translation error loss L trans The calculation formula is expressed by formula (27): Among them, R * and t * These represent the rotation matrix and translation vector used to transform the source point cloud into the target point cloud, respectively.
Citation Information
Cited By
Retina OCT three-dimensional image registration fusion method and system
CN121304754A