A Point Cloud Registration Method and System Based on the Generation of Deep Virtual Corresponding Points
By building feature coding modules, virtual corresponding point generation layer and outlier point filtering modules, combined with weighted singular value decomposition method, the problem of point cloud registration is solved, and the rapid and effective registration is achieved in complex scenarios, which is suitable for fields such as autonomous driving and digital presentation of cultural relics.
Patent Information
- Application Number
- CN202210398269.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-04-15
AI Technical Summary
In the absence of initial transformation information, the existing point cloud registration methods have problems such as long registration time and low success rate, especially in complex scenarios and low overlap rate, which are difficult to effectively perform point cloud registration.
Using a method based on deep virtual corresponding point generation, a feature encoding module, a virtual corresponding point generation layer, a corresponding point weighting layer and an outlier point filtering module are constructed, combined with a weighted singular value decomposition method, the rigid transformation information between point clouds, including rotation matrix and translation vectors.
In the absence of initial transformation information, rigid transformation approximate solutions can be solved quickly and effectively, and the efficiency and robustness of point cloud registration can be improved. It is suitable for autonomous driving, digital presentation of cultural relics and robot applications.
Smart Images

Figure CN114708315B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of three-dimensional point cloud registration, and in particular to a point cloud registration method and system based on the generation of depth virtual corresponding points. Background Art
[0002] With the popularization and price reduction of sensors such as lidar (light detection and ranging, abbreviated as LiDAR), RGB-D cameras (RealSense, Kinect), and 3D scanners, the latest 3D acquisition technologies have made great leaps. Different from the widely used 2D data, 3D data has rich scale and geometric information, which can help machines better understand the environment. Point cloud registration is the process of unifying point cloud data collected from different perspectives into the same coordinate system. The ultimate goal is to obtain a complete point cloud description of the object surface. The key to the problem is how to obtain the rigid transformation parameters of the point cloud coordinates so that the distance between the corresponding points measured from different perspectives is minimized after coordinate transformation. Point cloud registration is an indispensable part of automatic 3D reconstruction and is also the basis of virtual display technology.
[0003] Iterative Closest Point (abbreviated as ICP) is currently the most well-known algorithm for solving the rigid registration problem. This algorithm alternately executes the search for corresponding points and the least squares optimization algorithm to update the registration state. The point-to-point cost function uses the coordinate distance or feature distance between points to find the nearest point pair as the corresponding point. The performance of the ICP algorithm highly depends on the accuracy of the initial rigid transformation estimate. However, the initial estimate obtained from the timer is not so reliable, which leads to the ICP algorithm being easily trapped in a local optimal solution. In order to find an optimal transformation method, Yang et al. proposed the Go-ICP algorithm based on Branch and Bound (abbreviated as BnB) to determine the global optimal pose. When point cloud registration requires providing a global optimal solution, the Go-ICP method is superior to the ICP algorithm. Other algorithms use methods such as convex relaxation, Riemannian optimization, and mixed integer programming to determine the global optimal pose estimate, but these methods have a relatively high computational cost and cannot well meet the actual application requirements.
[0004] In recent years, the application of deep learning in point cloud registration pose estimation has made great progress. PointNetLK uses PointNet to extract the global features of two input point clouds, and then uses the inverse compositional (IC) algorithm to estimate the transformation matrix. By estimating the transformation matrix, the registration target becomes minimizing the feature difference between the two features. Deep Closest Point (DCP) uses the DGCNN network to extract local features, and then calculates soft corresponding points first before using the Kabsch algorithm to estimate the parameters of the rigid transformation. PointDSC proposes a non-local feature aggregation module to complete the feature embedding of the input corresponding points, and then uses the neural spectral matching method to calculate the rigid transformation of each seed. Deep Global Registration (DGR) proposes a 6D convolutional network architecture for likelihood prediction of inliers / outliers. However, the above methods generally have problems such as long registration time and low registration success rate in scenarios lacking initial rigid motion information. Summary of the Invention
[0005] The purpose of the present invention is to provide a point cloud registration method and system based on the generation of deep virtual corresponding points, so as to quickly and effectively solve the approximate solution of the rigid transformation in the case of lacking initial transformation information, and improve the efficiency and robustness of point cloud registration.
[0006] To achieve the above purpose, the present invention provides the following solutions:
[0007] A point cloud registration method based on the generation of deep virtual corresponding points, comprising:
[0008] Construct a feature encoding module; the input of the feature encoding module is the source point cloud and the target point cloud to be registered, and the output includes feature vectors and cross-attention matrices;
[0009] In the feature space, use the feature encoding module to extract the feature vectors of the key points in the source point cloud and the target point cloud, and construct a global feature attention matrix according to the feature vectors;
[0010] Construct a virtual corresponding point generation layer, and obtain corresponding point pairs by combining the global feature attention matrix;
[0011] Construct a corresponding point weighting layer for predicting the confidence that the corresponding point pairs become inliers;
[0012] Construct an outlier filtering module to remove the corresponding point pairs with confidence lower than the inlier threshold, and obtain the remaining corresponding point pairs;
[0013] Using the weighted singular value decomposition method with the remaining corresponding point pairs as the input, solve the rigid transformation information between the source point cloud and the target point cloud; the rigid transformation information includes a rotation matrix and a translation vector;
[0014] Register the source point cloud and the target point cloud according to the rigid transformation information.
[0015] Optionally, the constructing feature encoding module specifically includes:
[0016] Construct a feature encoding module composed of L layers of repeated cross-attention networks; each layer of the cross-attention network contains a residual neural network block The residual neural network block Is sequentially connected by 4 layers of exactly the same basic residual neural network units.
[0017] Optionally, the using the feature encoding module in the feature space to extract the feature vectors of the key points in the source point cloud and the target point cloud, and constructing a global feature attention matrix according to the feature vectors specifically includes:
[0018] The input of the l-th layer cross-attention network in the feature encoding module is the feature vector output by the (l - 1)-th layer cross-attention network and where 1 < l ≤ L;
[0019] The feature vector and respectively pass through the same residual neural network block to obtain intermediate feature vectors and
[0020] According to the intermediate feature vectors and construct the l-th layer cross-attention matrix A (l) ;
[0021] According to the l-th layer cross-attention matrix A (l) Exchange the information between the source point cloud and the target point cloud to obtain the feature vectors output by the l-th layer cross-attention network and
[0022] Perform an accumulation operation on the L attention matrices A (l) to obtain the global feature attention matrix A.
[0023] Optionally, the constructing virtual corresponding point generation layer, combining the global feature attention matrix to obtain corresponding point pairs, specifically includes:
[0024] Construct a virtual corresponding point generation layer. Based on the global feature attention matrix A, use the formula to weight the key points in the target point cloud to obtain the virtual corresponding points of each key point in the source point cloud; where is a normalization pointer, which defines the mapping relationship from the key point x in the source point cloud i to the virtual corresponding point in the target point cloud; M and N are the numbers of key points in the source point cloud and the target point cloud respectively; A is the i-th row of the global feature attention matrix A; A i is the element in the i-th row and j-th column of the global feature attention matrix A; ij is the element in the i-th row and j-th column of the global feature attention matrix A;
[0025] Form a corresponding point pair (x in the source point cloud i and its virtual corresponding point ), that is, (x i , ).
[0026] Optionally, the construction of the corresponding point weighting layer is used to predict the confidence that the corresponding point pair becomes an inlier, specifically including:
[0027] Construct a corresponding point weighting layer, which is composed of 9 residual neural network basic units with different output channel numbers, a multi-layer perceptron, and a Softmax function connected in sequence;
[0028] Use the corresponding point weighting layer to predict the confidence pi that the corresponding point pair (x i , ) becomes an inlier.
[0029] A point cloud registration system based on deep virtual corresponding point generation, including:
[0030] A feature encoding module construction module, used to construct a feature encoding module; the input of the feature encoding module is the source point cloud and the target point cloud to be registered, and the output includes a feature vector and a cross-attention matrix;
[0031] A feature vector extraction module, used to extract the feature vectors of the key points in the source point cloud and the target point cloud in the feature space, and construct a global feature attention matrix according to the feature vectors;
[0032] A virtual corresponding point generation layer construction module, used to construct a virtual corresponding point generation layer, and obtain a corresponding point pair in combination with the global feature attention matrix;
[0033] Confidence prediction module, used to construct a corresponding point weighting layer for predicting the confidence that the corresponding point pair becomes an inlier;
[0034] Corresponding point pair rejection module, used to construct an outlier filtering module to reject the corresponding point pairs with confidence lower than the inlier threshold, obtaining the remaining corresponding point pairs;
[0035] Rigid transformation information solving module, used to use the weighted singular value decomposition method with the remaining corresponding point pairs as input to solve the rigid transformation information between the source point cloud and the target point cloud; the rigid transformation information includes a rotation matrix and a translation vector;
[0036] Point cloud registration module, used to register the source point cloud and the target point cloud according to the rigid transformation information.
[0037] Optionally, the feature encoding module construction module specifically includes:
[0038] Feature encoding module construction unit, used to construct a feature encoding module composed of L layers of repeated cross-attention networks; each layer of the cross-attention network contains a residual neural network block The residual neural network block Consists of 4 layers of exactly the same residual neural network basic units connected in sequence.
[0039] Optionally, the feature vector extraction module specifically includes:
[0040] Cross-attention matrix construction unit, used to construct the l-th layer cross-attention matrix A according to the intermediate feature vector and ; the input of the l-th layer cross-attention network in the feature encoding module is the feature vector output by the (l - 1)-th layer cross-attention network (l) ; where 1 < l ≤ L; the feature vectors and respectively pass through the same residual neural network block and to obtain intermediate feature vectors and and
[0041] Feature vector output unit, used to exchange information between the source point cloud and the target point cloud according to the l-th layer cross-attention matrix A (l) to obtain the feature vector output by the l-th layer cross-attention network and
[0042] Global feature attention matrix construction unit, used to process L of the attention matrices A (l)Perform an accumulation operation to obtain the global feature attention matrix A.
[0043] Optionally, the virtual corresponding point generation layer construction module specifically includes:
[0044] A virtual corresponding point generation layer construction unit for constructing a virtual corresponding point generation layer. Based on the global feature attention matrix A, use the formula to weight the key points in the target point cloud to obtain the virtual corresponding points of each key point in the source point cloud; where is a normalization pointer that defines the mapping relationship from the key point x in the source point cloud i to the virtual corresponding point in the target point cloud; M and N are the numbers of key points in the source point cloud and the target point cloud respectively; A is the i-th row of the global feature attention matrix A; A i is the element in the i-th row and j-th column of the global feature attention matrix A; ij is the element in the i-th row and j-th column of the global feature attention matrix A;
[0045] A corresponding point pair generation unit for pairing the key point x in the source point cloud i with its virtual corresponding point to form a corresponding point pair (x i , ).
[0046] Optionally, the confidence prediction module specifically includes:
[0047] A corresponding point weighting layer construction unit for constructing a corresponding point weighting layer, which is composed of 9 residual neural network basic units with different output channel numbers, a multi-layer perceptron, and a Softmax function connected in sequence;
[0048] A confidence prediction unit for using the corresponding point weighting layer to predict the confidence p i that the corresponding point pair (x ) becomes an inlier. i .
[0049] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:
[0050] The present invention provides a point cloud registration method and system based on the generation of depth virtual corresponding points. The method includes: constructing a feature encoding module; the input of the feature encoding module is the source point cloud and the target point cloud to be registered, and the output is a feature vector and a cross-attention matrix; extracting the feature vectors of key points in the source point cloud and the target point cloud using the feature encoding module in the feature space, and constructing a global feature attention matrix according to the feature vectors; constructing a virtual corresponding point generation layer to obtain corresponding point pairs in combination with the global feature attention matrix; constructing a corresponding point weighting layer for predicting the confidence that the corresponding point pairs become inliers; constructing an outlier filtering module to eliminate the corresponding point pairs with a confidence lower than the inlier threshold, and obtaining the remaining corresponding point pairs; using the weighted singular value decomposition method with the remaining corresponding point pairs as the input to solve the rigid transformation information between the source point cloud and the target point cloud; the rigid transformation information includes a rotation matrix and a translation vector; registering the source point cloud and the target point cloud according to the rigid transformation information. The method of the present invention can quickly and effectively solve the approximate solution of the rigid transformation in the absence of initial transformation information, improve the efficiency and robustness of point cloud registration, and is applicable to many fields such as autonomous driving, digital presentation of cultural relics, 3D reconstruction, and robot applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is a flowchart of a point cloud registration method based on the generation of depth virtual corresponding points according to the present invention;
[0053] Figure 2 It is a schematic structural diagram of the DeepVCG network adopted by the point cloud registration method of the present invention;
[0054] Figure 3 It is a schematic structural diagram of the feature encoding module adopted in the point cloud registration method of the present invention;
[0055] Figure 4 It is a schematic structural diagram of the basic unit of the residual neural network in the feature encoding module provided by the present invention;
[0056] Figure 5 It is a schematic structural diagram of the corresponding point weighting layer adopted in the point cloud registration method of the present invention;
[0057] Figure 6It is a schematic diagram for visual comparison of registration in the hotel scenario of the 3DMatch dataset provided by the embodiments of the present invention. Detailed implementation manners
[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] The purpose of the present invention is to provide a point cloud registration method and system based on the generation of depth virtual corresponding points, aiming to quickly and effectively solve the approximate solution of the rigid transformation in the case of lack of initial transformation information. This method can effectively and quickly align two point clouds to be registered, achieving a good balance between the registration time consumption and the registration success rate.
[0060] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0061] Figure 1 It is a flowchart of a point cloud registration method based on the generation of depth virtual corresponding points of the present invention. Figure 2 It is a schematic diagram of the DeepVCG network structure adopted by a point cloud registration method based on the generation of depth virtual corresponding points of the present invention. Refer to Figure 2 , the present invention designs a new robust point cloud registration method based on the generation of depth virtual correspondences (abbreviated as DeepVCG). This registration method aims to provide a method for correctly recovering the mapping relationship and classifying inliers / outliers in a scenario lacking initial rigid motion information. The specific network structure of DeepVCG is as Figure 2 shown. The overall structure of DeepVCG consists of the following parts: a feature encoding module, a virtual corresponding point generation layer, a corresponding point weighting layer, and an outlier filtering module. Among them, the virtual corresponding point generation layer constructs the mapping relationship, which stipulates the corresponding points between the source point cloud and the target point cloud . The corresponding point weighting layer and the outlier filtering module assign inlier weights to each corresponding point pair and remove the corresponding points with smaller weight values to avoid expanding the error of the rigid transformation estimation.
[0062] Refer to Figure 1 , a point cloud registration method based on DeepVCG of the present invention includes:
[0063] Step 101: Construct a feature encoding module.
[0064] The present invention proposes a feature encoding module to extract the feature descriptions of key points in the source and target point clouds. The core part of the feature encoding module consists of residual neural network blocks which are composed of four exactly the same basic units of the residual neural network.
[0065] Figure 3 is the structural schematic diagram of the feature encoding module provided by the present invention. Refer to Figure 3 , the present invention proposes to use a deep feature encoding module to extract the features of each key point in the source point cloud and the target point cloud , where and are spaces of M×3 and N×3 dimensions respectively, and M and N are the numbers of key points in the source point cloud and the target point cloud respectively. The specific structure of the deep feature encoding module is as shown in Figure 3 , this module takes the source point cloud to the target point cloud as the input, and the output includes a feature vector and a cross-attention matrix.
[0066] As shown in Figure 3 , the feature encoding module of the present invention consists of L layers of repeated cross-attention networks. The specific number of layers of L is set according to the actual situation and will be specifically discussed in subsequent experiments. Each layer of the cross-attention network contains a residual neural network block The residual neural network block is sequentially connected by four exactly the same basic units of the residual neural network. The basic unit of the residual neural network of the present invention is based on kernel point convolution (abbreviation: KPConv) and feature-kernel alignment (abbreviation: FKAConv), and its specific structure is as shown in Figure 4 . Refer to Figure 4 , the basic unit of the residual neural network of the present invention is sequentially composed of a one-dimensional convolution Conv1D, a k-nearest neighbor algorithm (abbreviation: kNN) and KPConv, and two layers of FKAConv. Among them, k takes 40, instance normalization (abbreviation: IN) and the activation function ReLU (rectified linear units) are added after the Conv1D and FKAConv convolutional layers, and IN and the activation function LeakyReLU are added after the KPConv convolutional layer to improve the stability of training.
[0067] Step 102: In the feature space, use the feature encoding module to extract the feature vectors of the key points in the source point cloud and the target point cloud, and construct a global feature attention matrix according to the feature vectors.
[0068] In the feature space, use the feature encoding module proposed by the present invention to extract the feature vectors of the key points in the source point cloud and the target point cloud, and at the same time construct a global attention matrix according to the obtained feature vectors. As Figure 3 shown, in the feature encoding module of the present invention, the input of the l-th layer of cross-attention convolutional network is the output feature vector of the previous layer and The l-th layer of cross-attention convolutional network outputs an attention matrix and a new feature vector and Specifically, the residual neural network block maps the output feature vector of the previous layer into a high-dimensional feature space, that is and respectively pass through the same residual neural network block to obtain the intermediate feature vector of the l-th layer and For the convenience of description, the mapping process is briefly described as:
[0069]
[0070] where C (l-1) is the number of output channels of the residual neural network block in the (l-1)-th layer, and the dimension of the feature vector is (2C (l-1) ). It should be noted that the input of the first layer (l = 1) of cross-attention network is the key points in the source point cloud and the target point cloud , that is
[0071] Adopt a differentiable approximation expression to construct the cross-attention matrix A (l) :
[0072]
[0073]
[0074] where and are the L2-normalization feature vectors of the intermediate feature vectors and respectively. a ij measures the key point x iand y j The similarity probability between, swap x i and y j The high-level semantic information between. σ takes 0.2, which is a temperature parameter used to control the sensitivity to feature similarity. [·] + is the max(·,0) operator, used to ensure the non-negativity of a ij . (A (l) ) ij is the element in the i-th row and j-th column of the cross-attention matrix A (l) , x i and y j are the key points in the source point cloud and the target point cloud respectively.
[0075] Exchange the information between the source point cloud and the target point cloud through the cross-attention matrix A (l) and obtain a new feature description of the key points:
[0076]
[0077] where cat[·, ·] represents the per-channel fusion function, can also be obtained in a similar way. and represent the output feature vectors of the l-th layer cross-attention network and are also the input features of the next layer network.
[0078] According to the cross-attention matrix A (l) output by the l-th layer cross-attention network, construct the global feature attention matrix A. The l-th layer cross-attention network will output the attention matrix The present invention accumulates the L attention matrices A (l) to obtain the global feature attention matrix A through an accumulation operation:
[0079]
[0080] The present invention proposes to use a feature encoding module to extract the features of the key points in the source and target point clouds, and construct the global feature attention matrix A based on the features of these key points. The feature attention matrix A can fully assist in exchanging the context information between the source and target point clouds and enhance the structural information of the point cloud.
[0081] Step 103: Construct a virtual corresponding point generation layer, and obtain corresponding point pairs in combination with the global feature attention matrix.
[0082] Construct a virtual corresponding point generation layer, and synthesize virtual corresponding points by combining with the global feature attention matrix. In the present invention, by constructing a virtual corresponding point generation layer, for each key point x in the source point cloud i generate corresponding virtual corresponding points The structure of the virtual corresponding point generation layer is as Figure 2 shown. Given the global feature attention matrix A, define a differentiable soft mapping relationship:
[0083]
[0084] wherein, plays an important role as a normalized pointer, and this pointer defines the mapping relationship from the key point x in the source point cloud i to the corresponding point in the target point cloud ; A i is the i-th row of matrix A, and A ij is the element in the i-th row and j-th column of matrix A.
[0085] Predict a virtual target point cloud through weighted operations
[0086]
[0087] The predicted virtual target point cloud is represented in the form of a tensor of size M×3. Equation (7) estimates the virtual corresponding point The process of generates a matching mapping relationship from x i to , which is generated by the weighted sum of the real 3D points located in the target point cloud .
[0088] In the virtual corresponding point generation layer of the present invention, the global feature attention matrix is used to weight the key points in the target point cloud, so as to obtain the virtual corresponding points of each key point in the source point cloud. The key point x in the source point cloud i and its virtual corresponding point can form a corresponding point pair
[0089] The present invention proposes a virtual corresponding point generation layer to solve the difficulty of finding exact corresponding points in the target point cloud, and uses the global feature attention matrix to weight a series of candidate key points in the target point cloud to generate virtual corresponding points.
[0090] Step 104: Construct a corresponding point weighting layer for predicting the confidence that the corresponding point pair becomes an inlier.
[0091] Construct a corresponding point weighting layer and an outlier filtering module. The corresponding point weighting layer is used to predict the probability (confidence) of a corresponding point pair being an inlier, and the outlier filtering module removes some corresponding point pairs (outliers) with lower confidence according to the confidence of the corresponding point pairs. The virtual corresponding point generation layer in the previous step 103 is the point cloud All key points x in i Generated paired corresponding points Since the point cloud And Are not completely overlapping, which means And Only a part of the key points in can be successfully matched. Therefore, it is necessary to take measures to filter out the corresponding point pairs (outliers) that fail to match. The present invention constructs a corresponding point weighting layer and an outlier filtering module to assign an inlier probability (confidence) weight to each corresponding point pair and remove some corresponding point pairs with smaller weight values.
[0092] Figure 5 Is the structural schematic diagram of the corresponding point weighting layer provided by the present invention. See Figure 5 , The input of the corresponding point weighting layer of the present invention is M corresponding point pairs The output is the inlier probability of each corresponding point pair (that is, the confidence of each corresponding point pair being an inlier) p i The larger the value, the greater the possibility that the corresponding point pair (x i , ) becomes an inlier. As Figure 5 Shown, the corresponding point weighting layer is composed of 9 residual neural network basic units with different numbers of output channels, a multi-layer perceptron (MLP for short), and a Softmax function connected in sequence. The internal structure of each residual neural network basic unit is the same as Figure 4 Exactly the same, only the number of output channels is different. The output channels of these 9 residual neural network basic units are specifically 16, 16, 32, 32, 64, 64, 64, 128, 128 in sequence. The input dimension of the first residual neural network basic unit is equal to 6, indicating the dimension of the input corresponding point pair (x i , ). The multi-layer perceptron MLP is composed of three fully connected layers. The first two layers are composed of a convolutional function Conv1D and an activation function ReLU, and the third layer is only composed of the convolutional function Conv1D. The Softmax function converts the output result of the multi-layer perceptron into a probability estimate value.
[0093] Step 105: Construct an outlier filtering module to remove the corresponding point pairs with a confidence lower than the inlier threshold, and obtain the remaining corresponding point pairs.
[0094] In the corresponding point weighting layer, an inlier probability, i.e., a confidence p i , , is predicted for each corresponding point pair (x i , p i ∈ (0, 1) represents the likelihood that a corresponding point pair (x i , ) is an inlier. The outlier filtering module compares the confidence p i of the corresponding point pair (x ) with the inlier threshold τ. When the probability value p i is less than the given threshold τ, the value of γ i is set to 0 (in this case, the corresponding point pair is classified as an outlier). Conversely, the value of γ i is equal to p i (in this case, the corresponding point pair is classified as an inlier), and the corresponding point pair is thus classified as an inlier / outlier:
[0095]
[0096] where is the Iverson bracket, and γ i is the confidence of the remaining corresponding point pair (x i , , ).
[0097] When each corresponding point pair (x i , ) performs the above comparison operation in parallel, the corresponding point pairs with low confidence can be filtered out, effectively improving the registration accuracy. In the present invention, the outlier filtering module is only used on the validation set, and the threshold τ is set to 0.5. In this way, about 50% of the corresponding point pairs are preprocessed as outliers, and the remaining corresponding points are used as inliers.
[0098] Step 106: Use the weighted singular value decomposition method with the remaining corresponding point pairs as input to solve for the rigid transformation information between the source point cloud and the target point cloud.
[0099] The corresponding point weighting layer of the present invention is successively composed of 9 residual neural network basic units with different numbers of output channels, a multi-layer perceptron, and a Softmax function, and is used to estimate the inlier probability value (confidence) of the corresponding point pair; the outlier filtering module eliminates some corresponding point pairs with low confidence based on the inlier probability. These corresponding point pairs are called outliers, and the remaining corresponding point pairs are called inliers. The weighted singular value decomposition method adopted in the present invention uses the inliers and the corresponding inlier probability values as input to solve for the rigid transformation information between the source and target point clouds.
[0100] Specifically, the present invention uses the weighted singular value decomposition method with the remaining corresponding point pairs (inliers) as the input to solve the rigid transformation prediction between two point clouds. Similar to most registration algorithms, after obtaining the correctly matched corresponding point pairs (inliers), the present invention uses the weighted singular value decomposition method to solve the relative rigid transformation prediction value as the rigid transformation information; the relative rigid transformation prediction value includes the prediction values of the rotation matrix and the translation vector. Given the weight value (confidence) γ of each remaining corresponding point pair i , the least squares fitting method is used to estimate the rotation matrix R * and the translation vector t * :
[0101]
[0102] Equation (9) can be solved by the weighted singular value decomposition method to obtain the prediction values of the relative rigid transformation information (rotation matrix R * and translation vector t * ). The specific prediction process is as follows:
[0103] Define the weighted centroids of the source point cloud and the virtual target point cloud respectively as and
[0104]
[0105] Calculate the covariance matrix H:
[0106]
[0107] Perform singular value decomposition on the matrix H:
[0108] H = USV T (12)
[0109] Then the rotation matrix R * and the translation vector t * can be obtained as approximate values through Equations (13) and (14):
[0110] R * = Vdiag(1, 1,..., det(VU T ))U T (13)
[0111]
[0112] Step 107: Register the source point cloud and the target point cloud according to the rigid transformation information.
[0113] After obtaining the rotation matrix R *and the translation vector t * After obtaining the predicted value of * and t * the source point cloud and the target point cloud data collected from different perspectives are unified into the same coordinate system according to the rigid transformation parameters R
[0114] The present invention designs a new point cloud rigid registration method, and the design goal is to provide a solution for point cloud registration in complex scenes with a low overlap rate. The overall process of the registration method described in the present invention is as follows: First, in the feature space, the feature encoding module proposed by the present invention is used to extract the feature vectors of each key point from the source point cloud to the target point cloud and at the same time, an attention matrix is constructed according to the obtained feature vectors; Secondly, a virtual corresponding point generation layer is constructed, and at the same time, combined with the global feature attention matrix A, virtual corresponding points are synthesized; Then, a corresponding point weighting layer and an outlier filtering module are constructed. The corresponding point weighting layer is used to predict the probability (confidence) of a corresponding point pair becoming an inlier, and the outlier filtering module removes some corresponding point pairs (outliers) with lower confidence according to the confidence of the corresponding point pairs; Finally, the weighted singular value decomposition method is used with the remaining corresponding point pairs (inliers) as the input to solve the rigid transformation information between the two point clouds; According to the rigid transformation information, the source point cloud and the target point cloud can be registered.
[0115] The following provides experiments to verify the technical effects of a point cloud registration method based on deep virtual corresponding point generation of the present invention.
[0116] In order to make a fair comparison with existing methods, the publicly available 3DMatch dataset and ModelNet40 dataset are used to evaluate the performance of the DeepVCG registration network proposed by the present invention. The 3DMatch and ModelNet40 datasets adopted in the embodiments of the present invention are publicly available benchmark datasets commonly used by most point cloud registration methods. The 3DMatch dataset collects datasets from 62 indoor scenes, of which the data of 54 scenes are used for training and the data of 8 scenes are used for evaluation. 3000 points are randomly sampled from each piece of point cloud data as key points, that is, M = N = 3000. The ModelNet40 dataset contains 40 synthetic 3D object categories, a total of 12,311 CAD models (each model contains 2048 points). In order to simulate local point clouds, the present invention adopts the following sampling strategy: First, the farthest-point sampling (FPS) method is used to sample the first 1024 (M = 1024) points from the outer surface of each CAD model to construct the source point cloud Then for each piece of source point cloud Perform a random rigid transformation to generate the corresponding target point cloud (N = 1024): The rotation angle is arbitrarily set within [0, 45°] along three axes, and the translation vector is arbitrarily set within [-0.5 m, 0.5 m]. Finally, randomly select a 3D point from the source point cloud and the target point cloud and sample 768 adjacent points around this 3D point respectively.
[0117] To correctly evaluate the performance of the point cloud registration method proposed in the present invention, the following three evaluation metrics are used in the 3DMatch dataset in the embodiments of the present invention: Rotation Error (RE for short), Translation Error (TE for short), and Registration Recall (RR for short). Among them, the rotation error RE and the translation error TE are mainly used to measure the error between the estimated pose and the ground truth pose, while the registration recall RR is used to evaluate the proportion of successfully registered point cloud pairs. When RE and TE are less than a given threshold, the result of point cloud registration can be considered successful. For example, for the 3DMatch dataset, when RE < 15° and TE < 30 cm, the pairwise registration result of 3DMatch can be considered successful. In addition, for the ModelNet40 dataset, the embodiments of the present invention adopt two evaluation metrics, root mean square error (RMSE for short) and mean absolute error (MAE for short), to evaluate the prediction errors of the rotation matrix and the translation vector.
[0118] The present invention is developed based on the deep learning framework Pytorch. The AdamW optimizer is used for all training models on the two datasets. The weight decay and the initial learning rate are both set to 0.001, and the batch size is set to 1. The experiments on the 3DMatch dataset are run on a graphics workstation with an Intel i9-10900X CPU and an NVIDIA GeForce GTX 3090 GPU. L = 6, and the number of output channels C of the residual neural network block in each layer of the cross-attention network l)As shown in Table 1, a total of 120 rounds of training were conducted. During testing, random sample consensus (RANSAC) and ICP were used to jointly optimize the prediction of the initial rigid transformation. The validation experiment of the ModelNet40 dataset was run on a graphics workstation with an Intel i9-10900X CPU and an NVIDIA GeForce GTX 3090 GPU, where L = 4, and the residual neural network blocks in each cross-attention network layer The number of output channels C( l )As shown in Table 1, a total of 20 rounds of training were conducted.
[0119] Table 1 The specific number of output channels C of the residual neural network blocks in each attention layer of (l)
[0120]
[0121] From Figure 6 the visualization results (white represents the source point cloud, and grayish-black represents the target point cloud), it can be seen that compared with traditional registration methods (such as ICP and RANSAC) and learning-based methods (such as the DGR method), for the same input of source and target point clouds, the images obtained by registering using the method proposed in the present invention are more complete and delicate.
[0122] Table 2 shows the test results evaluated on the 3DMatch dataset for 4 traditional methods (FGR, ICP, GC-RANSAC, and RANSAC), 2 learning-based methods (DGR and PCAM), and the method of the present invention under the same experimental parameter settings.
[0123] Table 2 Quantitative evaluation results of traditional and learning-based methods on the 3DMatch dataset
[0124]
[0125] It can be concluded from Table 2 that the method for robust registration of point clouds proposed in the present invention based on the generation of virtual corresponding points has relatively superior overall performance. Among them, the registration recall rate RR reaches the highest level. Compared with the other two learning-based registration methods, DGR and PCAM-soft, the method of the present invention is 1.31% higher than the DGR method and 0.75% higher than the PCAM-soft method. In addition, compared with the other methods listed in Table 2, the proposed registration method of the present invention reaches the lowest levels in both evaluation indicators of rotation error RE and translation error TE, which are 1.57° and 0.06m respectively, which fully demonstrates that the proposed registration method of the present invention has strong robustness in the real point cloud registration scenario.
[0126] Table 3 shows the registration results of the traditional registration method, the learning-based method DGR, and the registration method proposed in the present invention on the unseen local point cloud data (ModelNet40).
[0127] Table 3 Quantitative evaluation and comparison results of the present invention with traditional and learning-based methods on unseen local point clouds (ModelNet40)
[0128]
[0129] As can be seen from the results in Table 3, the RMSE indexes of the rotation matrix and translation vector of the present invention reach the lowest level, and still have great advantages compared with other registration methods, which indicates that the method of the present invention still maintains strong robustness in the synthetic point cloud registration scenario.
[0130] Table 4 shows the registration results of the traditional registration method, the learning-based method DGR, and the registration method proposed in the present invention on the noisy local point cloud data (ModelNet40).
[0131] Table 4 Quantitative evaluation and comparison results of the present invention with traditional and learning-based methods on noisy local point clouds (ModelNet40)
[0132]
[0133] As can be seen from the results in Table 4, the RMSE indexes of the rotation matrix and translation vector of the present invention reach the lowest level, and still have great advantages compared with other registration methods, which indicates that the method of the present invention has relatively strong resistance to noise.
[0134] The present invention proposes a method for robust point cloud registration based on the idea of generating virtual corresponding points. This method focuses on solving the pain point of difficultly finding corresponding points, and the design goal is to provide a solution for predicting an approximate solution of the rigid transformation between the point clouds to be registered. The present invention has carried out a series of comparative experiments and quantitative or qualitative analyses on the real-scene dataset 3DMatch and the synthetic dataset ModelNet40. The bold in Tables 2-4 represents the optimal values of different indexes. The experimental results in Tables 2-4 show that the registration method based on virtual corresponding point generation proposed in the present invention performs better than the traditional algorithm and the learning-based algorithm. The registration algorithm proposed in the present invention has strong robustness and has extremely broad application prospects in the fields of autonomous driving, geographical exploration, 3D reconstruction, virtual fitting, etc.
[0135] Based on the method provided by the present invention, the present invention also provides a point cloud registration system based on the generation of depth virtual corresponding points, including:
[0136] A feature encoding module construction module for constructing a feature encoding module; the input of the feature encoding module is the source point cloud and the target point cloud to be registered, and the output includes a feature vector and a cross-attention matrix;
[0137] A feature vector extraction module for extracting the feature vectors of key points in the source point cloud and the target point cloud in the feature space using the feature encoding module, and constructing a global feature attention matrix according to the feature vectors;
[0138] A virtual corresponding point generation layer construction module for constructing a virtual corresponding point generation layer and obtaining corresponding point pairs in combination with the global feature attention matrix;
[0139] A confidence prediction module for constructing a corresponding point weighting layer to predict the confidence that the corresponding point pair becomes an inlier;
[0140] A corresponding point pair rejection module for constructing an outlier filtering module to reject the corresponding point pairs with a confidence lower than the inlier threshold, and obtaining the remaining corresponding point pairs;
[0141] A rigid transformation information solving module for using the weighted singular value decomposition method with the remaining corresponding point pairs as input to solve the rigid transformation information between the source point cloud and the target point cloud; the rigid transformation information includes a rotation matrix and a translation vector;
[0142] A point cloud registration module for registering the source point cloud and the target point cloud according to the rigid transformation information.
[0143] Among them, the feature encoding module construction module specifically includes:
[0144] A feature encoding module construction unit for constructing a feature encoding module composed of L layers of repeated cross-attention networks; each layer of the cross-attention network contains a residual neural network block ; the residual neural network block Consists of 4 layers of exactly the same residual neural network basic units connected in sequence.
[0145] The feature vector extraction module specifically includes:
[0146] A cross-attention matrix construction unit for constructing the l-th layer cross-attention matrix A according to the intermediate feature vector and ; the input of the l-th layer cross-attention network in the feature encoding module is the feature vector output by the (l - 1)-th layer cross-attention network (l) ; where 1 < l ≤ L; the feature vector and where 1 < l ≤ L; the feature vector and respectively pass through the same residual neural network block to obtain an intermediate feature vector and
[0147] a feature vector output unit, configured to exchange information between the source point cloud and the target point cloud according to the l-th layer cross-attention matrix A (l) to obtain the feature vector output by the l-th layer cross-attention network and
[0148] a global feature attention matrix construction unit, configured to perform an accumulation operation on L of the attention matrices A (l) to obtain the global feature attention matrix A
[0149] The virtual corresponding point generation layer construction module specifically includes:
[0150] a virtual corresponding point generation layer construction unit, configured to construct a virtual corresponding point generation layer, and based on the global feature attention matrix A, use the formula to weight the key points in the target point cloud to obtain the virtual corresponding points of each key point in the source point cloud; where is a normalization pointer, which specifies the mapping relationship from the key point x in the source point cloud i to the virtual corresponding point in the target point cloud ; M and N are respectively the numbers of key points in the source point cloud and the target point cloud; A i is the i-th row of the global feature attention matrix A; A ij is the element in the i-th row and j-th column of the global feature attention matrix A
[0151] a corresponding point pair generation unit, configured to form a corresponding point pair (x in the source point cloud i with its virtual corresponding point as (x i , )
[0152] The confidence prediction module specifically includes:
[0153] a corresponding point weighting layer construction unit, configured to construct a corresponding point weighting layer, which is composed of 9 residual neural network basic units with different output channel numbers, a multi-layer perceptron, and a Softmax function connected in sequence
[0154] a confidence prediction unit, configured to use the corresponding point weighting layer to predict the corresponding point pair (xi , ) The confidence p of becoming an inlier i .
[0155] In a point cloud registration method and system based on the generation of depth virtual corresponding points proposed by the present invention, first, a feature encoding module is used to extract the feature vectors of key points in the source and target point clouds, and a global feature attention matrix is constructed according to the obtained feature vectors. The feature attention matrix can achieve the goal of exchanging context information between the source and target point clouds and enhancing structural information; secondly, a virtual corresponding point generation layer is constructed, and the virtual corresponding points of each key point in the source point cloud are synthesized in combination with the global feature attention matrix; then, a corresponding point weighting layer and an outlier filtering module are constructed. The corresponding point weighting layer is used to predict the probability (confidence) of a corresponding point pair becoming an inlier, and the outlier filtering module is used to remove some corresponding point pairs with low confidence; finally, the weighted singular value decomposition method (weighted singular value decomposition, abbreviated as weighted SVD) is used to solve the predicted information of the rigid transformation of the source and target point clouds.
[0156] The evaluation experiments of the algorithm proposed by the present invention on the open-source datasets 3DMatch and ModelNet40 show that the method of the present invention is superior to the previous traditional registration methods and deep learning-based registration methods in terms of robustness, accuracy, registration quality, etc., and has broad application prospects in the fields of autonomous driving, digital modeling of cultural relics, smart home, etc.
[0157] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0158] In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A point cloud registration method based on the generation of depth virtual corresponding points, characterized in that, including: a feature encoding module; the input of the feature encoding module is the source point cloud and the target point cloud to be registered, and the output includes a feature vector and a cross-attention matrix; in the feature space, use the feature encoding module to extract the feature vectors of key points in the source point cloud and the target point cloud, and construct a global feature attention matrix according to the feature vectors; construct a virtual corresponding point generation layer, and combine the global feature attention matrix to obtain corresponding point pairs; construct a corresponding point weighting layer for predicting the confidence that the corresponding point pair becomes an inlier; construct an outlier filtering module to eliminate the corresponding point pairs with a confidence lower than the inlier threshold, and obtain the remaining corresponding point pairs; use the weighted singular value decomposition method with the remaining corresponding point pairs as the input to solve the rigid transformation information between the source point cloud and the target point cloud; the rigid transformation information includes a rotation matrix and a translation vector; register the source point cloud and the target point cloud according to the rigid transformation information.
2. The method according to claim 1, wherein The construction of the feature encoding module specifically includes: Construct a feature encoding module composed of L layers of repeated cross-attention networks; each layer of the cross-attention network contains a residual neural network block The residual neural network block It is composed of 4 layers of exactly the same basic residual neural network units connected in sequence.
3. The method according to claim 2, wherein The use of the feature encoding module in the feature space to extract the feature vectors of key points in the source point cloud and the target point cloud, and construct a global feature attention matrix according to the feature vectors specifically includes: The input of the cross-attention network in the $l$-th layer of the feature encoding module is the feature vector output by the cross-attention network in the $(l - 1)$-th layer and where $1 \lt l \leq L$; The feature vector and respectively pass through the same residual neural network block to obtain intermediate feature vectors and According to the intermediate feature vector and construct the cross-attention matrix A of the l-th layer (l) ; According to the l-th layer cross-attention matrix A (l) Exchange information between the source point cloud and the target point cloud to obtain the feature vector output by the l-th layer cross-attention network and Perform an accumulation operation on the L attention matrices A (l) to obtain the global feature attention matrix A.
4. The method according to claim 3, wherein The construction of the virtual corresponding point generation layer, and combining the global feature attention matrix to obtain corresponding point pairs specifically includes: Construct a virtual corresponding point generation layer based on the global feature attention matrix A, using the formula The key points in the target point cloud are weighted to obtain virtual corresponding points of each key point in the source point cloud; For the normalized pointer, it is specified that the source point cloud The key point x in i Point to target point cloud Virtual corresponding point in M and N are the number of key points in the source point cloud and the target point cloud respectively; A i is the i-th row of the global feature attention matrix A; A ij is the i-th row and j-th column element of the global feature attention matrix A; Take the source point cloud The key point x i in it and its virtual corresponding point to form a corresponding point pair 5. The method according to claim 4, characterized in that, The construction of the corresponding point weighting layer for predicting the confidence that the corresponding point pair becomes an inlier specifically includes: Construct a corresponding point weighting layer, which is composed of 9 residual neural network basic units with different output channel numbers, a multi-layer perceptron, and a Softmax function connected in sequence; Predict the corresponding point pairs using the corresponding point weighting layer Confidence p of becoming an inlier i .
6. A point cloud registration system based on the generation of depth virtual corresponding points, characterized in that, including: a feature encoding module construction module for constructing a feature encoding module; the input of the feature encoding module is the source point cloud and the target point cloud to be registered, and the output includes a feature vector and a cross-attention matrix; a feature vector extraction module for using the feature encoding module in the feature space to extract the feature vectors of key points in the source point cloud and the target point cloud, and constructing a global feature attention matrix according to the feature vectors; a virtual corresponding point generation layer construction module for constructing a virtual corresponding point generation layer and combining the global feature attention matrix to obtain corresponding point pairs; a confidence prediction module for constructing a corresponding point weighting layer for predicting the confidence that the corresponding point pair becomes an inlier; a corresponding point pair elimination module for constructing an outlier filtering module to eliminate the corresponding point pairs with a confidence lower than the inlier threshold, and obtaining the remaining corresponding point pairs; a rigid transformation information solving module for using the weighted singular value decomposition method with the remaining corresponding point pairs as the input to solve the rigid transformation information between the source point cloud and the target point cloud; the rigid transformation information includes a rotation matrix and a translation vector; a point cloud registration module for registering the source point cloud and the target point cloud according to the rigid transformation information.
7. The system according to claim 6, characterized in that, The feature encoding module construction module specifically includes: A feature encoding module construction unit for constructing a feature encoding module composed of L layers of repeated cross-attention networks; each layer of the cross-attention network contains a residual neural network block The residual neural network block It is composed of 4 layers of exactly the same basic residual neural network units connected in sequence.
8. The system according to claim 7, characterized in that, The feature vector extraction module specifically includes: The cross-attention matrix construction unit is used to construct the l-th layer cross-attention matrix A according to the intermediate feature vectors and ; the input of the l-th layer cross-attention network in the feature encoding module is the feature vectors output by the (l-1)-th layer cross-attention network (l) where 1 < l ≤ L; the feature vectors and respectively pass through the same residual neural network block and to obtain intermediate feature vectors and and A feature vector output unit, configured to obtain, according to the l-th layer cross-attention matrix A (l) the feature vector output by the l-th layer cross-attention network by exchanging information between the source point cloud and the target point cloud and Global feature attention matrix construction unit, which is used to perform an accumulation operation on the L attention matrices A (l) to obtain the global feature attention matrix A.
9. The system according to claim 8, wherein The virtual corresponding point generation layer construction module specifically includes: A virtual corresponding point generation layer construction unit is used to construct a virtual corresponding point generation layer based on the global feature attention matrix A, using the formula The key points in the target point cloud are weighted to obtain virtual corresponding points of each key point in the source point cloud; For the normalized pointer, it is specified that the source point cloud The key point x in i Point to target point cloud Virtual corresponding point in M and N are the number of key points in the source point cloud and the target point cloud respectively; A i is the i-th row of the global feature attention matrix A; A ij is the i-th row and j-th column element of the global feature attention matrix A; A corresponding point pair generation unit, configured to generate a corresponding point pair from the key point x in the source point cloud i and its virtual corresponding point to form a corresponding point pair 10. The system according to claim 9, wherein The confidence prediction module specifically includes: A corresponding point weighting layer construction unit for constructing a corresponding point weighting layer, where the corresponding point weighting layer is composed of 9 residual neural network basic units with different numbers of output channels, a multi-layer perceptron, and a Softmax function connected in sequence; A confidence prediction unit, configured to use the corresponding point weighting layer to predict the confidence p of the corresponding point pair becoming an inlier i .
Citation Information
Patent Citations
Eyeglasses.
US1040105A