A point cloud registration method and system based on deep feature consistency

By combining multi-scale graph feature fusion, corresponding point weights and deep feature matching modules, the problem of robustness and low success rate of point cloud registration in complex scenarios is solved, and a more efficient point cloud registration effect is achieved.

CN113963040BActive Publication Date: 2025-08-22YUNNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111287533.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-02
Publication Date
2025-08-22
Estimated Expiration
2041-11-02

AI Technical Summary

Technical Problem

Existing point cloud registration methods are robust and successful in complex real scenarios, especially end-to-end registration networks do not perform well in complex scenarios.

Method used

The point cloud registration method based on depth feature consistency is adopted. Through the combination of multi-scale graph feature fusion module, corresponding point weight module and deep feature matching module, the fusion feature matrix of the corresponding point set is extracted, outlier points are filtered, candidate inner point set is determined, and the optimal rigid transform estimation matrix is ​​selected based on the internal point number maximization condition.

Benefits of technology

It improves the robustness and success rate of point cloud registration in complex actual scenarios, and improves the accuracy and efficiency of registration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113963040B_ABST
    Figure CN113963040B_ABST
Patent Text Reader

Abstract

The present invention relates to a point cloud registration method and system based on deep feature consistency. The method comprises: matching a source point cloud and a target point cloud to obtain a corresponding point set; extracting a fused feature matrix of the corresponding point set using a multi-scale graph feature fusion module; inputting the fused feature matrix of the corresponding point set into a corresponding point weight module to filter outliers and obtain a candidate inlier set; determining an inlier subset corresponding to each candidate inlier in the candidate inlier set, and determining a rigid transformation estimation matrix corresponding to each inlier subset based on a deep feature matching module; and selecting an optimal rigid transformation estimation matrix from the rigid transformation estimation matrices corresponding to each of the inlier subsets based on a condition of maximizing the number of inliers. The present invention combines the multi-scale graph feature fusion module, the corresponding point weight module, and the deep feature matching module to improve the robustness and success rate of point cloud registration in complex real-world scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a point cloud registration method and system based on depth feature consistency. Background Art

[0002] Iterative Closet Point (ICP) is currently the most well-known algorithm for solving rigid registration problems. This algorithm alternately performs the search for corresponding points and the least squares optimization algorithm to update the registration state. The performance of the ICP algorithm is highly dependent on the accuracy of the initial rigid transformation estimate, but the initial estimate obtained from the timer is not very reliable, which makes the ICP algorithm easily fall into a local optimal solution. In order to find an optimal transformation method, Yang et al. proposed the Go-ICP algorithm based on Branch and Bound (BnB) to determine the global optimal pose. When point cloud registration needs to provide a global optimal solution, the Go-ICP method is better than the ICP algorithm. Other algorithms include those based on convex relaxation, Riemann optimization, mixed integer programming, etc. to determine the global optimal pose estimate. These methods have relatively high computational costs and cannot meet the needs of practical applications well.

[0003] In recent years, the application of deep learning in point cloud registration and pose estimation has made great progress. For example, PointNetLK combines the global feature descriptor based on PointNet with the Lucas / Kanade optimization algorithm, and then solves the relative rigid transformation in an iterative calculation manner. DCP (Deep Closet Point) uses the DGCNN network to extract local features, and then first calculates the soft corresponding points before using the Kabsch algorithm to estimate the parameters of the rigid transformation. PointDSC proposes a non-local feature aggregation module to complete the feature embedding of the input corresponding points, and then uses the neural spectral matching method to calculate the rigid transformation of each seed. Predator adopts a parallel encoding-decoding structure and proposes a method of using a deep attention mechanism for the overlapping areas of the point cloud to exchange information of the point cloud to be registered.

[0004] Point cloud registration is a critical task in autonomous driving and robotics applications. Point cloud registration is primarily used to estimate the relative rigid transformation between two point clouds. It is also extremely important in target detection, tracking, and pose estimation. To effectively improve the accuracy and speed of registration, researchers have proposed methods such as end-to-end registration networks. In recent years, end-to-end registration networks such as DCP, PointNetLK, and VCR-Net have emerged. Compared with other classic registration networks, the efficiency of end-to-end neural networks has been fully verified. However, in some complex real-world scenarios, end-to-end registration networks suffer from poor robustness and low success rates. Summary of the Invention

[0005] The purpose of the present invention is to provide a point cloud registration method and system based on depth feature consistency to improve the robustness and success rate of point cloud registration in complex practical scenes.

[0006] To achieve the above object, the present invention provides a point cloud registration method based on depth feature consistency, the method comprising:

[0007] Step S1: Match the source point cloud and the target point cloud to obtain a set of corresponding points;

[0008] Step S2: extracting a fusion feature matrix of the corresponding point set using a multi-scale graph feature fusion module; the multi-scale graph feature fusion module includes a graph neural network and a multi-scale feature fusion unit;

[0009] Step S3: Inputting the fusion feature matrix of the corresponding point set into the corresponding point weight module to filter outliers and obtain a candidate inlier set; the corresponding point weight module includes a multi-layer perceptron layer and a candidate inlier sampling layer; the candidate inlier set includes multiple candidate inliers;

[0010] Step S4: determining an inlier subset corresponding to each candidate inlier point set, and determining a rigid transformation estimation matrix corresponding to each inlier subset based on a deep feature matching module;

[0011] Step S5: selecting an optimal rigid transformation estimation matrix from the rigid transformation estimation matrices corresponding to each of the inlier point subsets based on a condition of maximizing the number of inliers.

[0012] Optionally, the extracting a fusion feature matrix of a corresponding point set by using a multi-scale graph feature fusion module specifically includes:

[0013] Step S21: inputting the corresponding point set into the graph neural network for feature extraction to obtain graph features of the corresponding point set;

[0014] Step S22: Inputting the graph features into the multi-scale feature fusion unit to perform multi-scale feature fusion to obtain a fusion feature matrix of the corresponding point set.

[0015] Optionally, the step of inputting the fusion feature matrix of the corresponding point set into a corresponding point weight module to filter outliers and obtain a candidate inlier set specifically includes:

[0016] Step S31: inputting the fusion feature matrix of the corresponding point set into the multi-layer perceptron layer for confidence estimation to obtain the confidence of each corresponding point;

[0017] Step S32: Use the candidate inlier sampling layer to select the front The corresponding points with the highest confidence constitute the candidate inlier set.

[0018] Optionally, determining an inlier subset corresponding to each candidate inlier point set, and determining a rigid transformation estimation matrix corresponding to each inlier subset based on a deep feature matching module, specifically includes:

[0019] Step S41: using the k-NN method, searching in the feature space according to the corresponding point condition to form an inlier subset corresponding to each candidate inlier point set;

[0020] Step S42: constructing a feature consistency matrix for each of the inlier subsets;

[0021] Step S43: Calculate the principal component weights corresponding to each of the feature consistency matrices using principal component analysis;

[0022] Step S44: inputting the weights of each principal component and each inlier subset into a weighted singular value decomposition unit to perform rigid transformation estimation, and obtaining a rigid transformation estimation matrix corresponding to each inlier subset.

[0023] Optionally, matching the source point cloud and the target point cloud to obtain a corresponding point set specifically includes:

[0024] Step S11: downsampling the source point cloud and the target point cloud and extracting point features respectively to obtain the points corresponding to each point cloud and the features corresponding to each point;

[0025] Step S12: randomly selecting a set number of points from the points corresponding to the source point cloud using a random sampling method to form a first point matrix, and selecting features corresponding to each point in the first point matrix to form a first feature matrix;

[0026] Step S13: randomly selecting a set number of points from the points corresponding to the target point cloud using a random sampling method to form a second dot matrix, and selecting features corresponding to each point in the second dot matrix to form a second feature matrix;

[0027] Step S14: using a nearest neighbor search algorithm to construct a corresponding point set according to the first feature matrix and the second feature matrix.

[0028] The present invention also provides a point cloud registration system based on depth feature consistency, the system comprising:

[0029] Point cloud matching module, used to match the source point cloud and the target point cloud to obtain the corresponding point set;

[0030] Multi-scale graph feature fusion module, used to extract the fusion feature matrix of the corresponding point set;

[0031] A corresponding point weight module is used to filter outliers based on the fusion feature matrix of the corresponding point set to obtain a candidate inlier set; the candidate inlier set includes multiple candidate inliers;

[0032] A deep feature matching module is used to determine the inlier subset corresponding to each candidate inlier set, and the rigid transformation estimation matrix corresponding to each inlier subset;

[0033] The optimal selection module is used to select the optimal rigid transformation estimation matrix from the rigid transformation estimation matrices corresponding to each of the interior point subsets based on the condition of maximizing the number of interior points.

[0034] Optionally, the multi-scale graph feature fusion module specifically includes:

[0035] Graph neural network, used to extract features of corresponding point sets and obtain graph features of corresponding point sets;

[0036] The multi-scale feature fusion unit is used to perform multi-scale feature fusion on the graph features to obtain a fusion feature matrix of the corresponding point set.

[0037] Optionally, the corresponding point weight module specifically includes:

[0038] The multi-layer perceptron layer is used to estimate the confidence of the fused feature matrix of the corresponding point set and obtain the confidence of each corresponding point;

[0039] Candidate interior point sampling layer, used to select the front The corresponding points with the highest confidence constitute the candidate inlier set.

[0040] Optionally, the deep feature matching module specifically includes:

[0041] A search unit is configured to use a k-NN method to search in the feature space according to corresponding point conditions to form an inlier subset corresponding to each candidate inlier point set;

[0042] a feature consistency matrix determination unit, configured to construct a feature consistency matrix for each of the interior point subsets;

[0043] A principal component weight determination unit, configured to calculate the principal component weights corresponding to each of the feature consistency matrices using a principal component analysis method;

[0044] The weighted singular value decomposition unit is used to perform rigid transformation estimation based on the weights of each principal component and each inlier subset, and obtain a rigid transformation estimation matrix corresponding to each inlier subset.

[0045] Optionally, the point cloud matching module specifically includes:

[0046] The feature extraction unit is used to downsample the source point cloud and the target point cloud and extract point features to obtain the points corresponding to each point cloud and the features corresponding to each point;

[0047] A first random selection unit is configured to randomly select a set number of points from the points corresponding to the source point cloud using a random sampling method to form a first point matrix, and select features corresponding to each point in the first point matrix to form a first feature matrix;

[0048] A second random selection unit is used to randomly select a set number of points from the points corresponding to the target point cloud using a random sampling method to form a second dot matrix, and select features corresponding to each point in the second dot matrix to form a second feature matrix;

[0049] The corresponding point set construction unit is configured to construct a corresponding point set according to the first feature matrix and the second feature matrix by using a nearest neighbor search algorithm.

[0050] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0051] The present invention relates to a point cloud registration method and system based on deep feature consistency. The method comprises: matching a source point cloud and a target point cloud to obtain a corresponding point set; extracting a fused feature matrix of the corresponding point set using a multi-scale graph feature fusion module; inputting the fused feature matrix of the corresponding point set into a corresponding point weight module to filter outliers and obtain a candidate inlier set; determining an inlier subset corresponding to each candidate inlier in the candidate inlier set, and determining a rigid transformation estimation matrix corresponding to each inlier subset based on a deep feature matching module; and selecting an optimal rigid transformation estimation matrix from the rigid transformation estimation matrices corresponding to each of the inlier subsets based on a condition of maximizing the number of inliers. The present invention combines the multi-scale graph feature fusion module, the corresponding point weight module, and the deep feature matching module to improve the robustness and success rate of point cloud registration in complex real-world scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0053] Figure 1 This is a flow chart of the point cloud registration method based on depth feature consistency of the present invention;

[0054] Figure 2 Schematic diagram of the specific structure of the multi-scale graph feature extraction module of the present invention;

[0055] Figure 3 This is a schematic diagram of the rigid transformation estimation matrix corresponding to each inlier subset determined based on the deep feature matching module of the present invention;

[0056] Figure 4 This is a structural diagram of the point cloud registration system based on depth feature consistency of the present invention;

[0057] Figure 5 This is an example of the registration visualization of the kitchen scene in the 3DMatch dataset;

[0058] Figure 6 This is an example visualization of the registration of the living room (Home1) scene in the 3DMatch dataset. DETAILED DESCRIPTION

[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0060] The purpose of the present invention is to provide a point cloud registration method and system based on depth feature consistency to improve the robustness and success rate of point cloud registration in complex practical scenes.

[0061] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0062] Example 1

[0063] The present invention discloses a point cloud registration method based on depth feature consistency, the method comprising:

[0064] Step S1: Match the source point cloud and the target point cloud to obtain a set of corresponding points.

[0065] Step S2: Using a multi-scale graph feature fusion module to extract a fusion feature matrix of a corresponding point set; the multi-scale graph feature fusion module includes a graph neural network and a multi-scale feature fusion unit.

[0066] Step S3: Input the fusion feature matrix of the corresponding point set into the corresponding point weight module to filter outliers and obtain a candidate inlier set; the corresponding point weight module includes a multi-layer perceptron layer and a candidate inlier sampling layer; the candidate inlier set includes multiple candidate inliers.

[0067] Step S4: Determine the inlier subset corresponding to each candidate inlier point set, and determine the rigid transformation estimation matrix corresponding to each inlier subset based on the deep feature matching module.

[0068] Step S5: selecting an optimal rigid transformation estimation matrix from the rigid transformation estimation matrices corresponding to each of the inlier point subsets based on a condition of maximizing the number of inliers.

[0069] The following describes each step in detail:

[0070] Step S1: Match the source point cloud and the target point cloud to obtain a set of corresponding points, specifically including:

[0071] Step S11: down-sample the source point cloud and the target point cloud and extract point features respectively to obtain the points corresponding to each point cloud and the features corresponding to each point; in this embodiment, the FCGF feature description operator is used to extract the features corresponding to each point in the point cloud.

[0072] Step S12: Randomly select a set number of points from the points corresponding to the source point cloud using a random sampling method to form a first dot matrix, select the features corresponding to each point in the first dot matrix to form a first feature matrix, and each point in the first dot matrix is ​​called a key point corresponding to the source point cloud; that is, the first dot matrix P = {x i ,i=1,2,…,N}, the first characteristic matrix Where N is 1000, x i represents the i-th key point corresponding to the source point cloud, Represents the Nth key point x N Corresponding features.

[0073] Step S13: Randomly select a set number of points from the points corresponding to the target point cloud using a random sampling method to form a second dot matrix, select the features corresponding to each point in the second dot matrix to form a second feature matrix, and call each point in the second dot matrix the key point corresponding to the target point cloud; that is, the second dot matrix Q = {y i ,i=1,2,…,N}, the second characteristic matrix Where N is 1000, y i represents the i-th key point corresponding to the target point cloud, Represents the Nth key point y N Corresponding features.

[0074] Step S14: Using a nearest neighbor search algorithm, a corresponding point set is constructed according to the first feature matrix and the second feature matrix. The specific calculation formula is:

[0075]

[0076] in, represents the set of corresponding points, Represents the i-th key point x in the first feature matrix i The corresponding features, Represents the jth key point y in the second feature matrix j Corresponding features, N represents the total number of key points.

[0077] Step S2: Use the multi-scale graph feature fusion module to extract the fusion feature matrix of the corresponding point set. Figure 2 As shown in the figure, the multi-scale graph feature fusion module mainly consists of two parts: a graph neural network (GNN) and a multi-scale feature fusion unit (Multi-Scale Feature Merging, Multi-Scale FM).

[0078] Step S2 specifically includes:

[0079] Step S21: Set the corresponding points Input graph neural network for feature extraction to obtain corresponding point set Graph features The size of is 12×N×100, where N=1000, indicating the number of corresponding points.

[0080] Step S22: The graph features Input the multi-scale feature fusion unit to perform multi-scale feature fusion and obtain the fusion feature matrix of the corresponding point set The size is N×D, where D=256, indicating The dimension of f i Indicates the i corresponding points (x i ,y i )’s fusion features.

[0081] like Figure 2As shown in the figure, the multi-scale feature fusion unit consists of three scale layers, each of which consists of two layers of convolution functions, BatchNormalization (BN), and ReLU activation functions. The first scale layer has 64 output channels, so ×2 is marked in the upper right corner of the first dotted box to indicate that the first scale layer consists of two layers of convolution, BN, and ReLU activation functions. After three scaling layers, the number of dimensions becomes 64, 64, and 128 respectively. It is worth noting that After the first scale layer, the size remains unchanged, but after the second and third scale layers, its size is halved layer by layer. In order to smoothly fuse the output features of the three scale layers, the invention uses an upsampling method to perform upsampling processing on the output features of the second and third scale layers in parallel during the subsequent processing, so that the output feature size of the last two scale layers is doubled to be consistent with the output feature size of the first scale layer. The graph feature outputs of the three scale layers are fused and the fused feature outputs are sequentially sent to the convolution function, BatchNormalization (BN) and ReLU activation function layers to obtain the feature fusion matrix of the corresponding point set.

[0082] Step S3: Input the fusion feature matrix of the corresponding point set into the corresponding point weight module to filter outliers and obtain the candidate inlier set in, And the total number of samples of candidate inliers The value is 200. The corresponding point weight module includes a multi-layer perceptron layer and a candidate internal point sampling layer. The multi-layer perceptron layer consists of three fully connected layers. The first two fully connected layers are composed of convolution functions and ReLU activation functions. The last fully connected layer lacks the ReLU activation function and is composed only of the convolution function.

[0083] Step S3 specifically includes:

[0084] Step S31: Fusion feature matrix of corresponding point set The confidence level of each corresponding point is obtained by inputting the confidence level into the multi-layer perceptron layer.

[0085] Step S32: Use the candidate inlier sampling layer to select the front The corresponding points with the highest confidence constitute the candidate inlier set Specifically, all corresponding points are arranged in descending order according to the size of the confidence level, and the top The corresponding points with the highest confidence constitute the candidate inlier set The remaining corresponding points are determined as candidate outliers.

[0086] Step S4: Determine the inlier subset corresponding to each candidate inlier set, and determine the rigid transformation estimation matrix corresponding to each inlier subset based on the deep feature matching module. The specific process is as follows: Figure 3 As shown, specifically including:

[0087] Step S41: Using the k-NN method, search in the feature space according to the corresponding point condition to form an inlier subset corresponding to each candidate inlier point in the candidate inlier point set; the corresponding point condition is the continuous and adjacent k corresponding points of the candidate inlier point, specifically, k = 40. Because the total number of candidate inlier point samples is Therefore, the total number of internal point subsets is also

[0088] Step S42: constructing a feature consistency matrix for each of the inlier subsets.

[0089] Step S43: Calculate the principal component weights corresponding to each of the feature consistency matrices using principal component analysis (PCA). Since the feature consistency matrix is ​​constructed using an inlier subset, it is natural to consider the principal component weights as the inlier probability of the inlier subset. As the core of the deep feature matching module, the feature consistency matrix can be solved using the following formula:

[0090]

[0091] Here, σ represents a hyperparameter used to control the sensitivity to feature differences. represents f i The L2-normalized vector of , M represents the feature consistency matrix.

[0092] Step S44: Input the weights of each principal component and each inlier subset into a weighted singular value decomposition unit for rigid transformation estimation, and obtain the rigid transformation estimation matrix [R, t] corresponding to each inlier subset; wherein R∈SO(3) represents the rotation angle matrix, represents the translation matrix, and the two together constitute the rigid transformation estimation matrix.

[0093] The process of predicting the rigid transformation estimate matrix [R, t] by weighted singular value decomposition unit is to minimize the mean square error function This is achieved by predicting the rigid transformation estimation matrix [R, t] using the following mathematical formula:

[0094]

[0095] Among them, w i Represents the corresponding point (x i ,yi ) is the probability of the inliers, k represents the number of corresponding points in each inlier subset, R∈SO(3) represents the rotation angle matrix, Represents the translation matrix.

[0096] By performing the above deep feature matching process on each inlier subset in parallel, the rigid transformation prediction corresponding to each inlier subset can be obtained, that is, the total number of rigid transformation predictions and the total number of inlier subsets are also

[0097] Step S5: Selecting the optimal rigid transformation estimation matrix from the rigid transformation estimation matrices corresponding to each of the inlier point subsets based on the inlier number maximization condition; Specifically, the criterion for determining the rigid transformation is to maximize geometric consistency, i.e., to maximize the number of inliers. Specifically, the number of inliers corresponding to each rigid transformation that satisfies a given inlier threshold is calculated, and the rigid transformation that satisfies the inlier number maximization condition is established as the optimal rigid transformation estimation matrix [R * ,t * ]. The process of calculating the number of interior points corresponding to a rigid transformation estimation matrix can be achieved by maximizing the objective function E2. The specific formula is as follows:

[0098]

[0099] Among them, τ is the inlier threshold. For the 3DMatch dataset and KITTI Odometry, the value of τ is set to 10cm and 60cm respectively. i ,y i ) satisfies ||Rx i +ty i When ||<τ, it can be considered that the corresponding point (x i ,y i ) is a pair of interior points, otherwise, the corresponding point (x i ,y i ) are a pair of outliers.

[0100] The present invention proposes a multi-scale graph feature fusion module for extracting features of corresponding point sets, making full use of graph feature representations and semantic information at different scales. After the corresponding point weight module calculates the confidence of each corresponding point, it selects a certain number of corresponding points as candidate inliers, indirectly completing the coarse filtering of outliers. The selected candidate inliers have the characteristics of high confidence and wide distribution, and have a greater probability of becoming inliers, which further ensures the accuracy of the registration model. In addition, the present invention also designs a deep feature matching module to calculate the inlier probability of each inlier subset. In theory, the more obvious the features of the corresponding point, the higher the probability assigned to the corresponding point. Conversely, the probability assigned to the corresponding point with unclear features will be relatively low. This step effectively avoids the adverse effects that the corresponding points with unclear features may have on the registration results, and improves the performance of the registration method proposed by the present invention.

[0101] The method disclosed in the present invention is not only applicable to point cloud registration in complex scenes and with low overlap rates, but is also applicable to fields such as autonomous driving, digital presentation of cultural relics, three-dimensional reconstruction, and robotics applications.

[0102] Example 2

[0103] like Figure 4 As shown, the present invention also discloses a point cloud registration system based on depth feature consistency, the system comprising:

[0104] The point cloud matching module 401 is used to match the source point cloud and the target point cloud to obtain a corresponding point set.

[0105] Multi-scale graph feature fusion module 402, used to extract a fusion feature matrix of a corresponding point set;

[0106] The corresponding point weight module 403 is used to filter outliers according to the fusion feature matrix of the corresponding point set to obtain a candidate inlier set; the candidate inlier set includes multiple candidate inliers.

[0107] The deep feature matching module 404 is used to determine the inlier subsets corresponding to each candidate inlier point set and the rigid transformation estimation matrix corresponding to each inlier subset.

[0108] The optimal selection module 405 is configured to select an optimal rigid transformation estimation matrix from the rigid transformation estimation matrices corresponding to each of the inlier point subsets based on a condition of maximizing the number of inliers.

[0109] The following is a detailed discussion of each module:

[0110] As an optional implementation, the multi-scale graph feature fusion module 402 of the present invention specifically includes:

[0111] Graph neural network is used to extract features of corresponding point sets and obtain graph features of corresponding point sets.

[0112] The multi-scale feature fusion unit is used to perform multi-scale feature fusion on the graph features to obtain a fusion feature matrix of the corresponding point set.

[0113] As an optional implementation, the corresponding point weight module 403 of the present invention specifically includes:

[0114] The multi-layer perceptron layer is used to perform confidence estimation on the fused feature matrix of the corresponding point set to obtain the confidence of each corresponding point.

[0115] Candidate interior point sampling layer, used to select the front The corresponding points with the highest confidence constitute the candidate inlier set.

[0116] As an optional implementation, the depth feature matching module 404 of the present invention specifically includes:

[0117] The search unit is used to use the k-NN method to search in the feature space according to the corresponding point condition to form an inlier subset corresponding to each candidate inlier point in the candidate inlier point set.

[0118] The feature consistency matrix determination unit is used to construct the feature consistency matrix of each of the interior point subsets.

[0119] The principal component weight determination unit is used to calculate the principal component weight corresponding to each of the feature consistency matrices using a principal component analysis method.

[0120] The weighted singular value decomposition unit is used to perform rigid transformation estimation based on the weights of each principal component and each inlier subset, and obtain a rigid transformation estimation matrix corresponding to each inlier subset.

[0121] As an optional implementation, the point cloud matching module 401 of the present invention specifically includes:

[0122] The feature extraction unit is used to downsample the source point cloud and the target point cloud and extract point features respectively to obtain the key points corresponding to each point cloud and the features corresponding to each key point.

[0123] The first random selection unit is used to randomly select a set number of points from the points corresponding to the source point cloud using a random sampling method to form a first point matrix, and select features corresponding to each point in the first point matrix to form a first feature matrix.

[0124] The second random selection unit is used to randomly select a set number of points from the points corresponding to the target point cloud using a random sampling method to form a second dot matrix, and select features corresponding to each point in the second dot matrix to form a second feature matrix.

[0125] The corresponding point set construction unit is configured to construct a corresponding point set according to the first feature matrix and the second feature matrix by using a nearest neighbor search algorithm.

[0126] The same contents as those in Example 1 are not repeated here.

[0127] Example 3

[0128] The present invention is developed based on the deep learning framework Pytorch and runs on a graphics workstation with a GPU graphics card of NVIDIA GeForce GTX2080Ti. In order to make a quantitative comparison with existing methods, the present invention uses the public 3DMatch dataset and KITTI Odometry dataset to evaluate the performance of the registration network proposed in the present invention. The 3DMatch dataset collects datasets from 62 indoor scenes, of which data from 54 scenes are used for training and data from 8 scenes are used for evaluation. The KITTI dataset was jointly created by the Karlsruhe Institute of Technology in Germany and Toyota Research Institute of America. It is currently the world's largest computer vision algorithm evaluation dataset for autonomous driving scenarios. This dataset is used to evaluate the performance of computer vision technologies such as stereo images (Stereo), optical flow (Optical Flow), visual odometry (Visual Odometry), 3D object detection (Object detection) and 3D tracking (tracking) in a vehicle-mounted environment.

[0129] In order to correctly evaluate the performance of the algorithm proposed in this invention, we adopted the following three evaluation bodies: Rotation Error (RE), Translation Error (TE) and Registration Recall (RR). Among them, the rotation matrix error and the translation matrix error are mainly used to measure the error between the estimated pose and the ground truth pose, while the registration recall is used to evaluate the number of successfully registered point cloud pairs. When RE and TE are less than a given threshold, we can consider that the result of the point cloud registration is successful. For example, for the 3DMatch dataset, when RE < 15° and TE < 30cm, the pairwise registration result of 3DMatch can be considered successful.

[0130] Figure 5Figure a is an example image of kitchen scene registration obtained by using the ICP method, Figure b is an example image of kitchen scene registration obtained by using the RANSAC method, Figure c is an example image of kitchen scene registration obtained by using the DGR (Deep Global Registration) method, Figure d is an example image of kitchen scene registration obtained by using the PointDSC method, and Figure e is an example image of kitchen scene registration obtained by using the method of the present invention; Figure 6 Figure a is an example of a living room scene registration obtained by the ICP method, Figure b is an example of a living room scene registration obtained by the RANSAC method, Figure c is an example of a living room scene registration obtained by the DGR method, Figure d is an example of a living room scene registration obtained by the PointDSC method, and Figure e is an example of a living room scene registration obtained by the method of the present invention; Figure 5 and Figure 6 From the visualization results, it can be seen that compared with the traditional registration methods ICP, RANSAC and the learning-based methods DGR (Deep Global Registration) and PointDSC, for the same source point cloud and target point cloud input, the local point cloud image obtained by the registration method proposed in the present invention is more complete and delicate, and has a higher success rate.

[0131] Table 1 shows the test results of three traditional methods, four learning-based methods, and the method of the present invention evaluated on the 3DMatch dataset under the same experimental parameter settings. The time in the last column does not include the time for constructing the corresponding point set. From Table 1, we can conclude that the method for robust point cloud registration based on feature consistency proposed by the present invention has the best overall performance, among which the registration recall rate RR has reached the most advanced level. Compared with other learning-based registration methods DGR and PointDSC, the method of the present invention has improved by 2.1% and 0.19% respectively. In addition, compared with the other methods listed in the table, the two evaluation indicators of rotation matrix error RE and translation matrix error TE of the registration method proposed by the present invention have reached the lowest level, which are 1.67° and 6.04cm respectively. This fully proves that the registration method proposed by the present invention has extremely strong robustness.

[0132] Table 2 shows the evaluation results of traditional registration methods, the learning-based methods DGR and PointDSC, and the registration method of our invention on the large-scale outdoor scene KITTI Odometry dataset. The time in the last column does not include the time to construct the corresponding point set. The results in Table 2 show that our method still has significant advantages over other registration methods. The registration recall rate remains very high, and both the rotation matrix error RE and the translation matrix error TE evaluation metrics reach the lowest levels. This demonstrates that our method maintains strong robustness in complex outdoor registration scenarios.

[0133] Table 1 Quantitative evaluation results on the 3DMatch dataset

[0134]

[0135] Table 2 Quantitative evaluation results on the KITTI Odometry dataset

[0136]

[0137]

[0138] This paper conducted a series of comparative experiments and quantitative and qualitative analyses on the indoor scene dataset 3DMatch and the outdoor scene KITTI Odometry. The experimental results show that compared with traditional algorithms and learning-based algorithms, the registration method based on deep feature matching proposed in this paper has superior performance and strong robustness. It can still effectively and quickly estimate the rigid transformation and successfully align the two point clouds in complex outdoor registration scenarios. It has extremely broad application prospects in fields such as autonomous driving, digital presentation of cultural relics, smart home, and virtual fitting.

[0139] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0140] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A point cloud registration method based on deep feature consistency, characterized in that: The method comprises: Step S1: Match the source point cloud and the target point cloud to obtain a set of corresponding points, specifically including: Step S11: downsampling the source point cloud and the target point cloud and extracting point features respectively to obtain the points corresponding to each point cloud and the features corresponding to each point; Step S12: randomly selecting a set number of points from the points corresponding to the source point cloud using a random sampling method to form a first point matrix, and selecting features corresponding to each point in the first point matrix to form a first feature matrix; Step S13: randomly selecting a set number of points from the points corresponding to the target point cloud using a random sampling method to form a second dot matrix, and selecting features corresponding to each point in the second dot matrix to form a second feature matrix; Step S14: using a nearest neighbor search algorithm to construct a corresponding point set according to the first feature matrix and the second feature matrix; Step S2: extracting a fusion feature matrix of the corresponding point set using a multi-scale graph feature fusion module; the multi-scale graph feature fusion module includes a graph neural network and a multi-scale feature fusion unit; Step S3: Inputting the fusion feature matrix of the corresponding point set into the corresponding point weight module to filter outliers and obtain a candidate inlier set; the corresponding point weight module includes a multi-layer perceptron layer and a candidate inlier sampling layer; the candidate inlier set includes multiple candidate inliers; Step S4: Determine the inlier subset corresponding to each candidate inlier point set, and determine the rigid transformation estimation matrix corresponding to each inlier subset based on the deep feature matching module, specifically including: Step S41: using the k-NN method, searching in the feature space according to the corresponding point condition to form an inlier subset corresponding to each candidate inlier point set; Step S42: constructing a feature consistency matrix for each of the inlier subsets; Step S43: Calculate the principal component weights corresponding to each of the feature consistency matrices using principal component analysis; Step S44: inputting the principal component weights and the inlier subsets into a weighted singular value decomposition unit to perform rigid transformation estimation, and obtaining a rigid transformation estimation matrix corresponding to each inlier subset; Step S5: selecting an optimal rigid transformation estimation matrix from the rigid transformation estimation matrices corresponding to each of the inlier point subsets based on a condition of maximizing the number of inliers.

2. The point cloud registration method based on depth feature consistency according to claim 1, characterized in that: The method of extracting a fusion feature matrix of a corresponding point set by using a multi-scale graph feature fusion module specifically includes: Step S21: inputting the corresponding point set into the graph neural network for feature extraction to obtain graph features of the corresponding point set; Step S22: Inputting the graph features into the multi-scale feature fusion unit to perform multi-scale feature fusion to obtain a fusion feature matrix of the corresponding point set.

3. The point cloud registration method based on depth feature consistency according to claim 1, characterized in that: Inputting the fusion feature matrix of the corresponding point set into the corresponding point weight module to filter outliers and obtain the candidate inlier set specifically includes: Step S31: inputting the fusion feature matrix of the corresponding point set into the multi-layer perceptron layer for confidence estimation to obtain the confidence of each corresponding point; Step S32: Using the candidate inlier sampling layer, select the first NS corresponding points with the highest confidence to form a candidate inlier set.

4. A point cloud registration system based on deep feature consistency, characterized in that: The system comprises: The point cloud matching module is used to match the source point cloud and the target point cloud to obtain a corresponding point set. The point cloud matching module specifically includes: The feature extraction unit is used to downsample the source point cloud and the target point cloud and extract point features to obtain the points corresponding to each point cloud and the features corresponding to each point; A first random selection unit is configured to randomly select a set number of points from the points corresponding to the source point cloud using a random sampling method to form a first point matrix, and select features corresponding to each point in the first point matrix to form a first feature matrix; A second random selection unit is used to randomly select a set number of points from the points corresponding to the target point cloud using a random sampling method to form a second dot matrix, and select features corresponding to each point in the second dot matrix to form a second feature matrix; a corresponding point set construction unit, configured to construct a corresponding point set according to the first feature matrix and the second feature matrix using a nearest neighbor search algorithm; Multi-scale graph feature fusion module, used to extract the fusion feature matrix of the corresponding point set; A corresponding point weight module is used to filter outliers based on the fusion feature matrix of the corresponding point set to obtain a candidate inlier set; the candidate inlier set includes multiple candidate inliers; A deep feature matching module is used to determine the inlier subset corresponding to each candidate inlier point set and the rigid transformation estimation matrix corresponding to each inlier subset. The deep feature matching module specifically includes: A search unit is configured to use a k-NN method to search in the feature space according to corresponding point conditions to form an inlier subset corresponding to each candidate inlier point set; a feature consistency matrix determination unit, configured to construct a feature consistency matrix for each of the interior point subsets; A principal component weight determination unit, configured to calculate the principal component weights corresponding to each of the feature consistency matrices using a principal component analysis method; A weighted singular value decomposition unit is used to perform rigid transformation estimation based on the weights of the principal components and the inlier subsets to obtain a rigid transformation estimation matrix corresponding to each inlier subset; The optimal selection module is used to select the optimal rigid transformation estimation matrix from the rigid transformation estimation matrices corresponding to each of the interior point subsets based on the condition of maximizing the number of interior points.

5. The point cloud registration system based on depth feature consistency according to claim 4, characterized in that: The multi-scale graph feature fusion module specifically includes: Graph neural network, used to extract features of corresponding point sets and obtain graph features of corresponding point sets; The multi-scale feature fusion unit is used to perform multi-scale feature fusion on the graph features to obtain a fusion feature matrix of the corresponding point set.

6. The point cloud registration system based on depth feature consistency according to claim 4, characterized in that: The corresponding point weight module specifically includes: The multi-layer perceptron layer is used to estimate the confidence of the fused feature matrix of the corresponding point set and obtain the confidence of each corresponding point; The candidate inlier sampling layer is used to select the first NS corresponding points with the highest confidence to form a candidate inlier set.

Citation Information

Patent Citations

  • Complex component point cloud splicing method and system based on feature fusion

    CN113160287A

  • Method for automatically detecting building structure and generating 3D model based on laser radar

    WO2019242174A1