Point Cloud Registration Method Based on Deep Learning and Non-Corresponding Point Estimation

By introducing feature interaction modules and non-corresponding point estimation modules in the point cloud registration method, the problem that point cloud registration method in the prior art fails to effectively aggregate information and consider the impact of non-overlapping areas, and achieves high-precision and efficient point cloud registration.

CN116630381BActive Publication Date: 2025-06-17HEBEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310610044.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-29
Publication Date
2025-06-17
Estimated Expiration
2043-05-29

AI Technical Summary

Technical Problem

The existing point cloud registration method based on global features failed to effectively aggregate the information of the source point cloud and the target point cloud during the feature extraction stage, and did not consider the impact of non-overlapping areas on the global features, resulting in the failure of partial overlap point cloud registration.

Method used

The feature interaction module is introduced to enhance information exchange between point clouds during the feature extraction stage, and a non-corresponding point estimation module is built to estimate the global features of overlapping areas, and the rigid transformation parameters are calculated by minimizing the global feature difference.

Benefits of technology

It effectively reduces the impact of non-overlapping areas on registration, improves the accuracy of partial overlapping point cloud registration, and achieves high-precision point cloud registration while ensuring efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116630381B_ABST
    Figure CN116630381B_ABST
Patent Text Reader

Abstract

The present invention is a point cloud registration method based on deep learning and non-corresponding point estimation. The method includes the following steps: obtaining a target point cloud and a source point cloud; constructing a point cloud registration model based on deep learning and non-corresponding point estimation: the point cloud registration model based on deep learning and non-corresponding point estimation is set in an iterative form, and each iteration performs a processing process including a feature extraction module, a non-corresponding point estimation module, and a rigid transformation calculation module, and outputs a predicted rigid transformation matrix after iteration; the non-corresponding point estimation module includes a softmax function, a weighting operation, and a pooling operation. When estimating the correspondence relationship, the distance between the local feature and the global feature is used as the weight, effectively reducing the influence of noise on the global feature and showing strong robustness to noise; the model does not need to construct a complex network to calculate the exact correspondence relationship, the computational amount is significantly reduced, and the registration efficiency is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of 3D point cloud registration, and particularly relates to a point cloud registration method based on deep learning and non-corresponding point estimation. Background Art

[0002] The 3D point cloud registration technology is a key link in Simultaneous Localization and Mapping (SLAM) and 3D reconstruction. Its purpose is to solve the matching problem of point clouds under different views so that they can finally be stitched into a complete object. With the progress of acquisition technology, the scale of point cloud data has gradually increased, and traditional methods can no longer meet the requirements of registration in terms of accuracy and efficiency. Since deep neural networks have significant advantages in processing large-scale data, researchers have tried to use deep learning for registration. The point cloud registration method based on deep neural networks can solve the disadvantages of traditional methods, such as long time consumption and low accuracy in registering large-scale point cloud data, and provides a good solution for accurate and efficient point cloud registration. The existing registration methods based on deep neural networks can be divided into two types: based on correspondence and based on global features.

[0003] Methods based on correspondence, such as CentroidReg, PREDATOR, PCAM, etc., use deep neural networks to predict the overlapping area of point clouds and generate corresponding local features, and achieve registration by establishing the correspondence between local features. However, this method highly depends on the accuracy of the correspondence. If the local features of the input point cloud are not significant, it is easy to produce the problem of feature mis-matching, which will seriously reduce the accuracy of the algorithm for estimating the rigid body transformation parameters.

[0004] The registration method based on global features avoids the above mis-matching problem by aggregating the global information of point clouds. The PointNetLK algorithm is one of the classic registration algorithms based on global features, which has the advantages of high efficiency, high accuracy, and easy implementation of the network architecture, and has achieved extremely high performance in the registration problem of completely overlapping point clouds. However, most of the current registration methods based on global features do not aggregate the information of the source point cloud and the target point cloud in the feature extraction stage, so they cannot provide additional clues to enhance the feature descriptor. In addition, for the registration of partially overlapping point clouds, such methods do not consider the influence of non-overlapping areas on global features and lack the analysis of the correspondence between point clouds, so they cannot obtain the matching information between overlapping areas, resulting in registration failure. Summary of the Invention

[0005] Aiming at the deficiencies in the above-mentioned registration method based on global features, the technical problem to be solved by the present invention is to provide a point cloud registration method based on deep learning and non-corresponding point estimation. First, a feature interaction module is introduced into the feature extraction module for extracting features to enhance the information exchange between point clouds. The feature extraction module generates local features and global features of the point cloud. Then, a non-corresponding point estimation module is constructed to estimate the overlapping region by using the local features and global features of the point cloud, and the global features of the overlapping region of the point cloud are re-obtained. Finally, the rigid transformation parameters are calculated by minimizing the global feature difference of the overlapping region. The present invention can better adapt to the partial overlapping point cloud registration task, effectively reduce the influence generated by the non-overlapping region during the registration process, and achieve better performance. In addition, the method of the present invention does not need to calculate a very precise corresponding relationship, and can achieve high-precision partial point cloud registration while ensuring efficiency.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] In a first aspect, the present invention provides a point cloud registration method based on deep learning and non-corresponding point estimation, and the method includes the following contents:

[0008] Obtain a target point cloud and a source point cloud;

[0009] Construct a point cloud registration model based on deep learning and non-corresponding point estimation: The point cloud registration model based on deep learning and non-corresponding point estimation is set in an iterative form, and each iteration performs a processing process including a feature extraction module, a non-corresponding point estimation module, and a rigid transformation calculation module, and outputs a predicted rigid transformation matrix after iteration;

[0010] The feature extraction module includes two feature extraction branches composed of the same number of convolutional layers. The inputs of the two feature extraction branches are the source point cloud and the target point cloud respectively. A feature interaction module is inserted between different convolutional layers in each feature extraction branch. The specific insertion method is as follows: all convolutional layers in a single feature extraction branch are divided into m convolutional parts, and the number of convolutional layers in each convolutional part can be the same or different. m is an integer greater than 1 and less than the total number of convolutional layers; a feature interaction module is inserted between two adjacent convolutional parts, and the number of feature interaction modules in a single feature extraction branch is m - 1; the output of the m-th convolutional part is recorded as the local feature of the point cloud, and the output of the m-th convolutional part after passing through the second max pooling layer is recorded as the global feature of the point cloud;

[0011] The feature interaction module includes a first max pooling layer and a concatenation operation. After the output of the convolutional part on another branch is processed by the first max pooling layer, it is repeated to the same number as the number of points in the input point cloud through a repetition function to obtain the repeated features. Then, the repeated features are concatenated with the output of the previous convolutional part on the current branch at the dimension level to obtain the output of the feature interaction module, and the output of the feature interaction module is connected to the next convolutional part;

[0012] The non-corresponding point estimation module includes a softmax function, a weighting operation, and a pooling operation. The Euclidean distance d is calculated using the local feature f and the global feature F of the same point cloud output by the feature extraction module, and then the non-corresponding point probability w of the point cloud is obtained after being processed by the 1-softmax function; the local feature f of the same point cloud is weighted according to the non-corresponding point probability of the point cloud and then pooled to obtain the global feature Φ of the point cloud output by the non-corresponding point estimation module;

[0013] The rigid transformation calculation module is used to calculate the rigid transformation matrix for the global features Φ of the target point cloud and the source point cloud output by the non-corresponding point estimation module.

[0014] The source point cloud is transformed using the rigid transformation matrix T obtained in each iteration to obtain the updated source point cloud. The target point cloud and the updated source point cloud are used as the input of the feature extraction module for the next iteration. After obtaining the global features Φ(P) and Φ(Q) of the point cloud output by the non-corresponding point estimation module again, the rigid transformation increment is calculated; after performing the set number of iterations, the rigid transformation increments obtained in each iteration are applied to the set initial rigid transformation matrix T0 to obtain the final rigid transformation matrix T between the source point cloud and the target point cloud, which is the predicted value. The partial point cloud registration of the source point cloud and the target point cloud is realized using the final rigid transformation matrix T.

[0015] The loss function L of the point cloud registration model based on deep learning and non-corresponding point estimation consists of two parts: the mean square error between the predicted rigid transformation matrix and the true value of the rigid transformation, and the mean square error between the global features of the point cloud output by the non-corresponding point estimation module in the last iteration. Specifically:

[0016]

[0017] where T -1 represents the inverse of the predicted rigid transformation matrix T, T gt represents the true value of the rigid transformation, I4 represents the 4*4 identity matrix; Φ(P) and Φ(Q) are the global features of the source point cloud and the target point cloud output by the non-corresponding point estimation module.

[0018] The point cloud registration method can be used for the registration of partially overlapping point clouds and also for the registration of completely overlapping point clouds.

[0019] In a second aspect, the present invention provides a point cloud registration method based on deep learning and non-corresponding point estimation. The specific steps of the point cloud registration method are as follows:

[0020] Step 1: Read the point cloud data from the dataset to obtain the three-dimensional coordinates of the points contained in each sample. Perform random rotation and random translation on it to generate two point clouds that are completely overlapping but have different initial positions and the ground truth of the rigid transformation. Perform a cropping operation on the two point clouds to generate an incompletely overlapping source point cloud and target point cloud;

[0021] Step 2: Randomly sample the source point cloud and the target point cloud, and divide them into a training set and a test set according to the class labels of the samples. Each sample corresponds to the ground truth of the rigid transformation between the source point cloud and the target point cloud;

[0022] Step 3: Construct a point cloud registration model based on deep learning and non-corresponding point estimation, and set the overall model in an iterative form;

[0023] The point cloud registration model based on deep learning and non-corresponding point estimation includes a feature extraction module, two non-corresponding point estimation modules, and a rigid transformation calculation module; the input of the feature extraction module is the target point cloud and the source point cloud. The local features and global features obtained after the target point cloud is processed by the feature extraction module are input into one non-corresponding point estimation module, and the local features and global features obtained after the source point cloud is processed by the feature extraction module are input into another non-corresponding point estimation module; the outputs of the two non-corresponding point estimation modules are connected to the rigid transformation calculation module, and the rigid transformation matrix between the source point cloud and the target point cloud is obtained through the rigid transformation calculation module;

[0024] Use the rigid transformation matrix T obtained in each iteration to transform the source point cloud to obtain the updated source point cloud. Use the target point cloud and the updated source point cloud as the input of the feature extraction module for the next iteration. After obtaining the global features Φ(P) and Φ(Q) of the point cloud output by the non-corresponding point estimation module again, calculate the rigid transformation increment; after performing the set number of iterations, apply the rigid transformation increments obtained in each iteration to the set initial rigid transformation matrix T0 to obtain the final rigid transformation matrix T between the source point cloud and the target point cloud, which is the predicted value. Use the final rigid transformation matrix T to achieve partial point cloud registration of the source point cloud and the target point cloud;

[0025] Step 4: Construct the loss function of the point cloud registration model;

[0026] The loss function L of the point cloud registration model is formula (37):

[0027]

[0028] where T -1 represents the inverse of the rigid transformation matrix T, and T gt represents the true value of the rigid transformation, and I4 represents the 4×4 identity matrix;

[0029] Step 5: Use the training set to complete the supervised training of the point cloud registration model, use the test set to test the effectiveness of the trained point cloud registration model, and obtain the trained point cloud registration model based on deep learning and non-corresponding point estimation for point cloud registration.

[0030] The feature extraction module is based on the PointNet network and has two identical feature extraction branches. Each feature extraction branch includes seven one-dimensional convolutional layers, two feature interaction modules, and a max pooling layer; the first two of the seven one-dimensional convolutional layers are combined into a convolutional part, denoted as the h1 part, with the number of convolutional kernels being (64, 64), the next three layers are combined into the third convolutional part, denoted as the h3 part, with the number of convolutional kernels being (256, 256, 512), and the middle two layers are combined into the second convolutional part, denoted as the h2 part, with the number of convolutional kernels being (128, 128). A feature interaction module is inserted between the h1 part and the h2 part, and between the h2 part and the h3 part respectively. The max pooling layer is connected after the h3 part;

[0031] The source point cloud P and the target point cloud Q with a data shape of 1024×3 are first fed as inputs into the h1 part. The input of the h1 part is 3D and the output is 64D; after passing through the h1 part, the source point cloud P and the target point cloud Q obtain a local feature matrix of 1024×64 dimensions:

[0032] f1 a (P)=h1(P) (13)

[0033] f1 a (Q)=h1(Q) (14)

[0034] where f1 a (P) and f1 a (Q) respectively represent the local features obtained after the source point cloud P and the target point cloud Q pass through the h1 part;

[0035] The 1024×64-dimensional local features obtained after the two point clouds pass through the h1 part are input into the two feature interaction modules. The inputs of the feature interaction modules are f1 a (P) and f1 a (Q). On the target point cloud feature extraction branch, f1 on the source point cloud feature extraction branch a(P) The global features processed by the max pooling layer are repeated in the point number dimension through the repeat function, so that the repeated point number dimension is the same as the point number of the input point cloud; the repeated features are then combined with f1 on the target point cloud feature extraction branch a (Q) A connection operation is performed at the dimension level to obtain the local feature f1 output by the feature interaction module on the target point cloud feature extraction branch b (Q);

[0036] On the source point cloud feature extraction branch, f1 on the target point cloud feature extraction branch a (Q) The global features processed by the max pooling layer are repeated in the point number dimension through the repeat function, so that the repeated point number dimension is the same as the point number of the input point cloud; the repeated features are then combined with f1 on the source point cloud feature extraction branch a (P) A connection operation is performed at the dimension level to obtain the local feature f1 output by the feature interaction module on the source point cloud feature extraction branch b (P), f1 b (Q) and f1 b (P) are 1024 * 128 - dimensional local feature matrices:

[0037]

[0038]

[0039] Among them, f1 b (P) and f1 b (Q) respectively represent the local features obtained after f1 a (P) and f1 a (Q) are processed by the first feature interaction module. represents the connection operation at the dimension level, repeat[·] is the repeat function, (1024,1) means repeating 1024 times in the first dimension, and sym represents the max pooling function;

[0040] The 1024 * 128 - dimensional local features output by the first feature interaction module are used as the input of the h2 part; the input dimension of the h2 part is 128, and the output dimension is 128. After passing through the h2 part, a local feature matrix of size 1024 * 128 is still obtained;

[0041]

[0042]

[0043] Among them, and respectively represent f1 b (P) and f1 b(Q) The local features obtained after passing through the h2 part;

[0044] Input the 1024*128-dimensional local feature matrix obtained after passing through the h2 part into the second feature interaction module to obtain 1024*256-dimensional local features:

[0045]

[0046]

[0047] Among them, and respectively represent and the local features obtained after being processed by the second feature interaction module;

[0048] Take the 1024*256-dimensional local features obtained after being processed by the second feature interaction module as the input of the h3 part; the input dimension of the h3 part is 256, and the output dimension is 512. After passing through the h3 part, a local feature matrix of size 1024*512 is obtained:

[0049]

[0050]

[0051] Among them, f n (P) and f n (Q) respectively represent the local features of the source point cloud P and the target point cloud Q obtained at the nth iteration. The subscript n refers to the nth iteration; in the above, the superscript a represents the output of the convolutional part, and the superscript b represents the output of the feature interaction module;

[0052] Aggregate the 1024*512 local features obtained after passing through the h3 part through the max pooling layer after the h3 part to obtain global feature vectors of size 1024*1 respectively:

[0053] F n (P) = sym(f n (P)) (23)

[0054] F n (Q) = sym(f n (Q)) (24)

[0055] Among them, F n (P) and F n (Q) respectively represent the global features of the source point cloud P and the target point cloud Q obtained at the nth iteration.

[0056] The process of the non-corresponding point estimation module is as follows: The local features f n (P), f n (Q) and the global feature vectors F n (P), F n (Q) output by the feature extraction module at the nth iteration are used as inputs to calculate the Euclidean distance; the Euclidean distance is normalized by the 1-softmax function into a probability, representing the probability that each point is located in the non-overlapping region. The greater the distance, the smaller the weight assigned to the point, and vice versa. The mathematical definition is shown in equations (25)-(26); the local feature matrix is weighted and updated using the weights and passed through the max pooling layer to obtain the global features of the point cloud output by the non-corresponding point estimation module as shown in equations (27)-(28).

[0057] w(P) = 1 - softmax(||f n (P) - F n (P)||2) (25)

[0058] w(Q) = 1 - softmax(||f n (Q) - F n (Q)||2) (26)

[0059] Φ n (P) = sym(w(P)⊙f n (P)) (27)

[0060] Φ n (Q) = sym(w(Q)⊙f n (Q)) (28)

[0061] Among them, w(P) and w(Q) are 1024*1-dimensional column vectors, representing the probabilities of non-corresponding points of the source point cloud P and the target point cloud Q respectively. ||||2 represents the operation of calculating the Euclidean distance from the local feature to the current global feature. Φ n (P) and Φ n (Q) are the global features of the source point cloud P and the target point cloud Q obtained at the nth iteration. The subscript n refers to the nth iteration, and ⊙ represents the element-wise multiplication operation.

[0062] The dataset is the ModelNet40 dataset, and the clipping rate cp for the clipping operation is set to 0.8.

[0063] Compared with the prior art, the present invention has at least the following beneficial effects:

[0064] (1) The present invention constructs a point cloud registration model based on deep learning and non-corresponding point estimation. The neural network is used to adaptively extract the local features and global features of the point cloud without the need for artificial design of feature descriptors. The feature interaction module completes the information interaction between the source point cloud and the target point cloud in the early stage of registration (feature extraction stage) by means of feature connection, effectively improving the credibility of the global features. The non-corresponding point estimation module is constructed to predict and mask the non-corresponding points in the two point clouds, suppressing the information in the non-overlapping regions of the point cloud and effectively improving the registration accuracy of partially overlapping point clouds.

[0065] (2) For the point cloud registration model based on deep learning and non-corresponding point estimation constructed by the present invention, the distance between the local features and the global features is used as the weight when estimating the corresponding relationship, effectively reducing the influence of noise on the global features and showing strong robustness to noise. The model does not need to construct a complex network to calculate the exact corresponding relationship, significantly reducing the computational amount and significantly improving the registration efficiency. Description of the Drawings

[0066] Figure 1 It is a schematic structural diagram of the point cloud registration model based on deep learning and non-corresponding point estimation in the point cloud registration method based on deep learning and non-corresponding point estimation of the present invention.

[0067] Figure 2 It is a schematic input-output structural diagram of two feature interaction modules.

[0068] Figure 3 It is a schematic structural diagram of the non-corresponding point estimation module.

[0069] Figure 4 It is a comparison chart of the registration results of the dataset without noise using the method of the present invention and different existing registration methods.

[0070] Figure 5 It is a comparison chart of the registration results of the dataset with noise using the method of the present invention and different existing registration methods. Detailed Embodiments

[0071] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0072] It should be understood that the embodiments described in the present invention are exemplary, and the specific parameters used in the description of the embodiments are only for the convenience of describing the present invention and are not used to limit the present invention.

[0073] The point cloud registration method of the present invention can be used for the registration of partially overlapping point clouds, especially in 3D modeling, and can also be used for the registration of completely overlapping point clouds. During the acquisition process of the source point cloud and the target point cloud, the three-dimensional coordinates of each point in the point cloud data set are (x, y, z); the read point cloud data is rotated at a random angle and translated at a random distance; the rotation angle and translation distance of each obtained source point cloud corresponding to the target point cloud are set as the true value T of the rigid transformation gt ; perform a cropping operation on two point clouds that are completely overlapping but have different initial positions to obtain an incompletely overlapping source point cloud and target point cloud, and the cropping rate is represented by cp:

[0074]

[0075] wherein, M, M c respectively represent the number of points in the point cloud before and after cropping.

[0076] Randomly sample the source point cloud and the target point cloud. The random sampling operation refers to randomly sampling K points from the points included in the source point cloud and the target point cloud.

[0077] The rigid transformation calculation module is completed by the improved Lucas-Kanade algorithm. The global feature difference between the source point cloud and the target point cloud is used as the objective function to be optimized, and the objective function is minimized iteratively to calculate an appropriate rigid transformation matrix to minimize the distance of the global features; update the position of the source point cloud, and input the target point cloud and the updated source point cloud into the feature extraction module again; after performing N iterations, the overall rigid transformation T between the final source point cloud and the target point cloud is obtained by combining all the rigid transformation increments in each iteration est ;

[0078] In the rigid transformation calculation, the objective function to be optimized by the improved Lucas-Kanade algorithm is:

[0079]

[0080] wherein, T -1 represents the inverse of the rigid transformation matrix T, ξ represents the parameter of the rigid transformation matrix T, represents the operation of rotating and translating the point cloud;

[0081] The final rigid transformation matrix T between the source point cloud and the target point cloud is expressed as:

[0082] T = ΔT N ×ΔT N-1 ×(…)×T0 (11)

[0083] wherein, ΔT N is the increment of the rigid transformation matrix obtained in the Nth iteration, ΔTN-1 is the increment of the rigid transformation matrix obtained in the (N - 1)-th iteration, and T0 represents the set initial rigid transformation matrix.

[0084] The method for training the point cloud registration model using the training set is as follows: taking the source point cloud and the target point cloud of the training set as the input of the model, the predicted rigid transformation matrix of the two point clouds and the global features of the point cloud output by the non-corresponding point estimation module as the output of the point cloud registration model, reducing the error through the gradient descent method, performing backpropagation on the error, optimizing the network parameters of the point cloud registration model, and obtaining the trained point cloud registration model.

[0085] Embodiment

[0086] The point cloud registration method based on deep learning and non-corresponding point estimation in this embodiment is as follows:

[0087] Step 1: Read the point cloud data from the dataset, obtain the three-dimensional coordinates of the points contained in each sample, perform random rotation and random translation on it, generate two point clouds with different initial positions but completely overlapping and the true value of the rigid transformation, and perform a cropping operation on the two point clouds to generate an incompletely overlapping source point cloud and target point cloud;

[0088] This embodiment uses the ModelNet40 dataset, reads the point cloud data from the dataset, each point cloud data contains 2048 points, and the three-dimensional coordinates of each point are (x, y, z); perform random angle rotation and random distance translation on the read point cloud data, the angle range is [0, 45°], and the distance range is [0, 0.3]; the rotation angle and translation distance of the point cloud are processed to obtain the true value T of the rigid transformation matrix gt ; perform a cropping operation on the two point clouds with different initial positions but completely overlapping to obtain an incompletely overlapping source point cloud and target point cloud, and the cropping rate cp is set to 0.8.

[0089] Step 2: Randomly sample the source point cloud and the target point cloud, divide them into a training set and a test set according to the class labels of the samples, and each sample corresponds to the true value of the rigid transformation between the source point cloud and the target point cloud;

[0090] Randomly sample 1024 points from the source point cloud and the target point cloud obtained in Step 1; the ModelNet40 dataset contains 40 object categories, and the first 20 object categories are selected as the training set to train the network, containing 5112 training data. The remaining 20 object categories are selected as the test set to test the network, containing 1266 data.

[0091] Step 3: As Figure 1As shown, a point cloud registration model based on deep learning and non-corresponding point estimation is constructed. The overall model is set in an iterative form, with the number of iterations being 10 times. In the figure, the subscript n refers to the situation at the nth iteration.

[0092] The point cloud registration model based on deep learning and non-corresponding point estimation includes a feature extraction module, two non-corresponding point estimation modules, and a rigid transformation calculation module. The input of the feature extraction module is the target point cloud and the source point cloud. The local features and global features obtained after the target point cloud is processed by the feature extraction module are input into one non-corresponding point estimation module, and the local features and global features obtained after the source point cloud is processed by the feature extraction module are input into the other non-corresponding point estimation module. The outputs of the two non-corresponding point estimation modules are connected to the rigid transformation calculation module, and the rigid transformation matrix T between the source point cloud and the target point cloud is obtained through the rigid transformation calculation module.

[0093] In this embodiment, the feature extraction module is based on the PointNet network and has two identical feature extraction branches. Each feature extraction branch includes seven one-dimensional convolutional layers, two feature interaction modules, and a max pooling layer. The first two layers of the seven one-dimensional convolutional layers are combined into a convolutional part, denoted as the h1 part, with the number of convolutional kernels being (64, 64). The next three layers are combined into the third convolutional part, denoted as the h3 part, with the number of convolutional kernels being (256, 256, 512). The middle two layers are combined into the second convolutional part, denoted as the h2 part, with the convolutional size being (128, 128). Here, h represents the function approximated by the multi-layer perceptron composed of multiple convolutional layers. A feature interaction module is inserted between the h1 part and the h2 part, and between the h2 part and the h3 part respectively. The max pooling layer is connected behind the h3 part.

[0094] The source point cloud P and the target point cloud Q with a data shape of 1024*3 are first fed as inputs into the h1 part. The input of the h1 part is 3-dimensional, and the output is 64-dimensional. After passing through the h1 part, the source point cloud P and the target point cloud Q obtain a local feature matrix of 1024*64 dimensions:

[0095] f1 a (P) = h1(P) (13)

[0096] f1 a (Q) = h1(Q) (14)

[0097] Among them, f1 a (P) and f1 a (Q) respectively represent the local features obtained after the source point cloud P and the target point cloud Q pass through the h1 part.

[0098] The 1024*64-dimensional local features obtained after the two point clouds pass through the h1 part are input into as Figure 2In the two feature interaction modules shown, the input of one feature interaction module is f1 a (P) and f1 a (Q). On the target point cloud feature extraction branch, the global feature of f1 a (P) on the source point cloud feature extraction branch after being processed by the max pooling layer is repeated in the number of points dimension through the repeat function, so that the repeated number of points dimension is the same as the number of points of the input point cloud; the repeated feature is then concatenated with f1 a (Q) on the target point cloud feature extraction branch at the dimension level to obtain the local feature f1 b (Q) output by the feature interaction module on the target point cloud feature extraction branch;

[0099] On the source point cloud feature extraction branch, the global feature of f1 a (Q) on the target point cloud feature extraction branch after being processed by the max pooling layer is repeated in the number of points dimension through the repeat function, so that the repeated number of points dimension is the same as the number of points of the input point cloud; the repeated feature is then concatenated with f1 a (P) on the source point cloud feature extraction branch at the dimension level to obtain the local feature f1 b (P) output by the feature interaction module on the source point cloud feature extraction branch. f1 b (Q) and f1 b (P) are 1024 * 128 - dimensional local feature matrices:

[0100]

[0101]

[0102] Among them, f1 b (P) and f1 b (Q) respectively represent the local features obtained after f1 a (P) and f1 a (Q) are processed by the first feature interaction module. represents the concatenation operation at the dimension level. repeat[·] is the repeat function, (1024, 1) means repeating 1024 times in the first dimension, and sym represents the max pooling function.

[0103] The 1024 * 128 - dimensional local feature output by the first feature interaction module is used as the input of the h2 part. The input dimension of the h2 part is 128, and the output dimension is 128. After passing through the h2 part, a local feature matrix of size 1024 * 128 is still obtained.

[0104]

[0105]

[0106] Among them, and respectively represent the local features obtained after f1 b (P) and f1 b (Q) pass through the h2 part.

[0107] Input the 1024*128-dimensional local feature matrix obtained after passing through the h2 part into the second feature interaction module to obtain a 1024*256-dimensional local feature:

[0108]

[0109]

[0110] Among them, and respectively represent and the local features obtained after being processed by the second feature interaction module.

[0111] Use the 1024*256-dimensional local feature obtained after being processed by the second feature interaction module as the input of the h3 part. The input dimension of the h3 part is 256, and the output dimension is 512. After passing through the h3 part, a local feature matrix of size 1024*512 is obtained:

[0112]

[0113]

[0114] Among them, f n (P) and f n (Q) respectively represent the local features of the source point cloud P and the target point cloud Q obtained in the nth iteration; the superscript a above represents the output of the convolutional part, and the superscript b represents the output of the feature interaction module.

[0115] Aggregate the 1024*512 local features obtained after passing through the h3 part through the max pooling layer after the h3 part to obtain global feature vectors of size 1024*1 respectively:

[0116] F n (P) = sym(f n (P)) (23)

[0117] F n (Q) = sym(f n (Q)) (24)

[0118] Among them, F n (P) and Fn (Q) respectively represent the global features of the source point cloud P and the target point cloud Q obtained at the nth iteration, and the subscript n refers to the nth iteration.

[0119] The process of the non-corresponding point estimation module is as Figure 3 shown, where f represents the local feature, the square is used to represent the local feature space, and the points in it represent the feature vectors of each point in the point cloud. F represents the global feature, d represents the Euclidean distance between the local feature and the global feature, w represents the probability after 1-softmax, the subscript i represents the Euclidean distance and probability of the ith point, and w⊙f represents the operation of weighting the local feature f with the probability w. The weighted local feature passes through the max pooling layer sym to obtain a new global feature Φ. The local feature matrix f n (P), f n (Q) and the global feature vector F n (P), F n (Q) are used as inputs to calculate the Euclidean distance; the Euclidean distance is normalized to a probability through the 1-softmax function, representing the probability that each point is located in the non-overlapping area. The greater the distance of a point, the smaller the weight assigned to it, and vice versa. The mathematical definition is shown in Eqs. (25)-(26); the local feature matrix is weighted and updated using the weight and passes through the max pooling layer to obtain the global feature shown in Eqs. (27)-(28);

[0120] w(P) = 1 - softmax(||f n (P) - F n (P)||2) (25)

[0121] w(Q) = 1 - softmax(||f n (Q) - F n (Q)||2) (26)

[0122] Φ n Φ(P) = sym(w(P)⊙f n (P)) (27)

[0123] Φ n Φ(Q) = sym(w(Q)⊙f n (Q)) (28)

[0124] where w(P) and w(Q) are 1024*1-dimensional column vectors, representing the probabilities of non-corresponding points of the source point cloud P and the target point cloud Q respectively. ||||2 represents the operation of calculating the Euclidean distance from the local feature to the current global feature. Φ n (P) and Φ n(Q) is the global feature of the source point cloud P and the target point cloud Q obtained in the nth iteration, and ⊙ represents the operation of element-wise corresponding multiplication.

[0125] In the rigid transformation calculation module, the improved Lucas-Kanade algorithm updates the rigid transformation matrix T by calculating the transformation parameters. The relationship between the transformation parameter ξ = [ξ1, ξ2, ξ3, ξ4, ξ5, ξ6] and the rigid transformation matrix T is as follows:

[0126]

[0127] where exp() represents the exponential mapping function, G l is the generator matrix, and l is the number of elements in the transformation parameter;

[0128] The optimal rigid transformation matrix T is calculated by minimizing the difference between the global features of the source point cloud and the target point cloud output by the non-corresponding point estimation module. Its objective function is expressed as:

[0129]

[0130] where T -1 represents the inverse of the rigid transformation matrix T, ξ represents the parameter of the rigid transformation matrix T, represents the operation of performing rotation and translation on the point cloud;

[0131] The specific process of solving the objective function is as follows: perform the first-order Taylor expansion on the objective function; approximate the Jacobian matrix of the transformation parameter ξ = [ξ1, ξ2, ξ3, ξ4, ξ5, ξ6] by the finite difference method; for the nth iteration, obtain its approximate solution by taking the partial derivative of the rigid transformation parameter increment; calculate the rigid transformation increment using the obtained rigid transformation parameter increment. The mathematical expressions of the specific process are shown in equations (31)-(35);

[0132]

[0133]

[0134] Δξ n =-(J T ·J) -1 ·J T ·[Φ n (P)-Φ n (Q)] (33)

[0135]

[0136] T n =T n-1 ×ΔT n (35)

[0137] where \(J\) is the Jacobian matrix of \(\varPhi\) with respect to the transformation parameter \(\xi\), \(\Delta\xi\) n is the increment at the \(n\)th iteration, \(\Delta t\) is the infinitesimal perturbation of \(\xi\), set to 0.01, \(\Delta T\) n represents the rigid transformation increment obtained at the \(n\)th iteration, \(T\) n is the rigid transformation matrix obtained at the \(n\)th iteration, \(T\) n-1 is the rigid transformation matrix obtained at the \((n - 1)\)th iteration.

[0138] Using the \(T\) obtained in each iteration n to transform the source point cloud so that it gets closer and closer to the target point cloud. The target point cloud and the updated source point cloud will be used as the input of the feature extraction module for the next iteration (the \((n + 1)\)th iteration). After obtaining the global features \(\varPhi(P)\) and \(\varPhi(Q)\) again, the rigid transformation increment is calculated. After 10 iterations, the predicted rigid transformation matrix \(T\) between the source point cloud and the target point cloud is obtained by applying the rigid transformation increments obtained in each iteration to the set initial rigid transformation matrix \(T_0\) (a given initial value):

[0139] \(T=\Delta T\) 10 \(\times\Delta T_9\times(\cdots)\times T_0\ (36)\)

[0140] Step 4: Construct the loss function of the point cloud registration model;

[0141] In the said Step 4, the loss function \(L\) of the point cloud registration model consists of two parts: the mean square error between the predicted rigid transformation matrix \(T\) and the true value of the rigid transformation, and the mean square error between the global features. Specifically:[[]]

[0142]

[0143] where \(T\) -1 represents the inverse of the rigid transformation matrix \(T\), \(T\) gt represents the true value of the rigid transformation (obtained in the data processing stage, a known value), and \(I_4\) represents the \(4\times4\) identity matrix.

[0144] Step 5: Use the training set to complete the supervised training of the point cloud registration model, and use the test set to test the effectiveness of the trained point cloud registration model;

[0145] Use the training set divided in Step 2 to train the point cloud registration model. The entire training process uses the Adam optimizer to optimize the neural network parameters. The training batch size is 16, the initial learning rate is 0.001, and the number of training epochs is 200;

[0146] The trained point cloud registration model is directly used to predict the rigid transformation between the target point cloud and the source point cloud in the test set divided in step 2. To make the effectiveness and advantages of the embodiments of the present invention clearer, the traditional method ICP (calculating the rigid transformation parameters of the point cloud using the correspondence between points) and three deep learning-based registration methods are used for comparison with the present invention. The three deep learning algorithms are DCP (calculating the soft mapping relationship using a deep neural network and then obtaining the rigid transformation parameters), PRNet (on the basis of DCP, realizing the solution process of the rigid transformation using iteration), and PointNetLK (minimizing the global features between two point clouds to solve the rigid transformation). The root mean square error (RMSE) and mean absolute error (MAE) of rotation and translation of the above four methods and the present invention 1 (five convolutional layers and one feature interaction module, and the feature interaction module is inserted after the second convolutional layer), the present invention 2 (seven convolutional layers and two feature interaction modules, and the feature interaction modules are inserted after the second and fourth convolutional layers respectively), and the present invention 3 (seven convolutional layers and three feature interaction modules, and the feature interaction modules are inserted after the second, fourth, and sixth convolutional layers respectively) on the noiseless test set and the noisy test set are shown in Table 1 and Table 2 respectively. To more intuitively display the superior performance of the present invention, two data are randomly selected from the test set for visualizing the registration results, as shown in Figure 4 and Figure 5 respectively.

[0147] Table 1 Results of the noiseless test set

[0148]

[0149]

[0150] Table 2 Results of the noisy test set

[0151]

[0152] It can be seen from Table 1, Table 2 and the visualization results of the registration that compared with the existing methods, the method of the present invention has a greater improvement in registration accuracy due to the introduction of the feature interaction module and the non-corresponding point estimation module. In addition, when the non-corresponding point estimation module estimates the correspondence relationship, the distance between features is used as the weight, which can reduce the influence of noise on the global features. Therefore, the present invention shows strong robustness.

[0153] Those not described in the present invention are applicable to the prior art.

Claims

1. A point cloud registration method based on deep learning and non-corresponding point estimation, characterized in that, The method includes the following steps: Obtain a target point cloud and a source point cloud; Construct a point cloud registration model based on deep learning and non-corresponding point estimation: The point cloud registration model based on deep learning and non-corresponding point estimation is set in an iterative form. Each iteration performs a processing procedure including a feature extraction module, a non-corresponding point estimation module, and a rigid transformation calculation module, and outputs a predicted rigid transformation matrix after iteration; The feature extraction module includes two feature extraction branches composed of the same number of convolutional layers. The inputs of the two feature extraction branches are the source point cloud P and the target point cloud Q respectively. A feature interaction module is inserted between different convolutional layers in each feature extraction branch. The specific insertion method is as follows: All convolutional layers in a single feature extraction branch are divided into m convolutional parts. The number of convolutional layers in each convolutional part can be the same or different. m is an integer greater than 1 and less than the total number of convolutional layers; A feature interaction module is inserted between two adjacent convolutional parts. The number of feature interaction modules in a single feature extraction branch is m - 1; The output of the m-th convolutional part is denoted as the local feature of the point cloud, and the output of the m-th convolutional part after passing through the second max pooling layer is denoted as the global feature of the point cloud; The feature interaction module includes a first max pooling layer and a concatenation operation. After the output of the convolutional part on the other branch is processed by the first max pooling layer, it is repeated to the same number as the number of points in the input point cloud through a repeating function to obtain the repeated feature. The repeated feature and the output of the previous convolutional part on the current branch are concatenated at the dimension level to obtain the output of the feature interaction module. The output of the feature interaction module is connected to the next convolutional part; The non-corresponding point estimation module includes a softmax function, a weighting operation, and a pooling operation. The Euclidean distance d is calculated using the local feature f and the global feature F of the same point cloud output by the feature extraction module, and then the non-corresponding point probability w of the point cloud is obtained after being processed by the softmax function; The local feature f of the same point cloud is weighted according to the non-corresponding point probability of the point cloud and then pooled to obtain the global feature Φ of the point cloud output by the non-corresponding point estimation module; The rigid transformation calculation module is used to calculate the rigid transformation matrix for the global features Φ of the target point cloud and the source point cloud output by the non-corresponding point estimation module.

2. The point cloud registration method based on deep learning and non-corresponding point estimation according to claim 1, characterized in that, The source point cloud is transformed using the rigid transformation matrix T obtained in each iteration to obtain an updated source point cloud. The target point cloud and the updated source point cloud are used as the inputs of the feature extraction module for the next iteration. After obtaining the global features Φ(P) and Φ(Q) of the point cloud output by the non-corresponding point estimation module again, the rigid transformation increment is calculated; After performing the set number of iterations, the rigid transformation increments obtained in each iteration are applied to the set initial rigid transformation matrix T0 to obtain the final rigid transformation matrix T between the source point cloud and the target point cloud. This is the predicted value, and the partial point cloud registration of the source point cloud and the target point cloud is realized using the final rigid transformation matrix T.

3. The point cloud registration method based on deep learning and non-corresponding point estimation according to claim 1, characterized in that, The loss function L of the point cloud registration model based on deep learning and non-corresponding point estimation consists of two parts: the mean square error between the predicted rigid transformation matrix and the true value of the rigid transformation, and the mean square error between the global features of the point clouds output by the non-corresponding point estimation module in the last iteration. Specifically: Among them, T -1 represents the inverse of the predicted rigid transformation matrix T, and T gt represents the true value of the rigid transformation, and I4 represents the 4×4 identity matrix; Φ(P) and Φ(Q) are the global features of the source point cloud and the target point cloud output by the non-corresponding point estimation module.

4. The point cloud registration method based on deep learning and non-corresponding point estimation according to claim 1, characterized in that, The point cloud registration method can be used for the registration of partially overlapping point clouds and also for the registration of completely overlapping point clouds.

5. A point cloud registration method based on deep learning and non-corresponding point estimation, characterized in that, The specific steps of the point cloud registration method are as follows: Step 1: Read the point cloud data from the dataset to obtain the three-dimensional coordinates of the points contained in each sample. Randomly rotate and translate it to generate two point clouds that are completely overlapping but have different initial positions and the true value of the rigid transformation. Perform a cropping operation on the two point clouds to generate an incompletely overlapping source point cloud and target point cloud; Step 2: Randomly sample the source point cloud P and the target point cloud Q, and divide them into a training set and a test set according to the class labels of the samples. Each sample corresponds to the true value of the rigid transformation between the source point cloud and the target point cloud; Step 3: Construct a point cloud registration model based on deep learning and non-corresponding point estimation, and set the overall model in an iterative form; The point cloud registration model based on deep learning and non-corresponding point estimation includes a feature extraction module, two non-corresponding point estimation modules, and a rigid transformation calculation module; the input of the feature extraction module is the target point cloud and the source point cloud. The local features and global features obtained after the target point cloud is processed by the feature extraction module are input into one non-corresponding point estimation module, and the local features and global features obtained after the source point cloud is processed by the feature extraction module are input into the other non-corresponding point estimation module; the outputs of the two non-corresponding point estimation modules are connected to the rigid transformation calculation module, and the rigid transformation matrix between the source point cloud and the target point cloud is obtained through the rigid transformation calculation module; Use the rigid transformation matrix T obtained in each iteration to transform the source point cloud to obtain the updated source point cloud. Use the target point cloud and the updated source point cloud as the input of the feature extraction module for the next iteration. After obtaining the global features Φ(P) and Φ(Q) of the point clouds output by the non-corresponding point estimation module again, calculate the rigid transformation increment; after performing the set number of iterations, apply the rigid transformation increments obtained in each iteration to the set initial rigid transformation matrix T0 to obtain the final rigid transformation matrix T between the source point cloud and the target point cloud, which is the predicted value. Use the final rigid transformation matrix T to achieve partial point cloud registration of the source point cloud and the target point cloud; Step 4: Construct the loss function of the point cloud registration model; The loss function L of the point cloud registration model is formula (37): where T -1 represents the inverse of the rigid transformation matrix T, and T gt represents the true value of the rigid transformation, and I4 represents the 4×4 identity matrix; Step 5: Use the training set to complete the supervised training of the point cloud registration model, use the test set to test the effectiveness of the trained point cloud registration model, and obtain the trained point cloud registration model based on deep learning and non-corresponding point estimation for point cloud registration.

6. The point cloud registration method based on deep learning and non-corresponding point estimation according to claim 5, characterized in that, The feature extraction module is based on the PointNet network and has two identical feature extraction branches. Each feature extraction branch includes seven one-dimensional convolutional layers, two feature interaction modules, and a max pooling layer. The first two of the seven one-dimensional convolutional layers are combined into a convolutional part, denoted as the h1 part, with the number of convolutional kernels being (64, 64). The next three layers are combined into the third convolutional part, denoted as the h3 part, with the number of convolutional kernels being (256, 256, 512). The middle two layers are combined into the second convolutional part, denoted as the h2 part, with the number of convolutional kernels being (128, 128). A feature interaction module is inserted between the h1 part and the h2 part, and between the h2 part and the h3 part respectively. The max pooling layer is connected behind the h3 part. The source point cloud P and the target point cloud Q with a data shape of 1024*3 are first fed as inputs into the h1 part. The input of the h1 part is 3-dimensional and the output is 64-dimensional. After passing through the h1 part, the source point cloud P and the target point cloud Q obtain a local feature matrix of 1024*64 dimensions: f1 a (P) = h1(P) (13) f1 a (Q) = h1(Q) (14) Among them, f1 a (P) and f1 a (Q) respectively represent the local features obtained after the h1 part of the source point cloud P and the target point cloud Q; The 1024×64 - dimensional local features obtained after passing two point clouds through the h1 part are input into two feature interaction modules. The inputs of the feature interaction modules are f1 a (P) and f1 a (Q). On the target point cloud feature extraction branch, the global feature of f1 a (P) on the source point cloud feature extraction branch after being processed by the max - pooling layer is repeated in the number of points dimension through a repeating function, making the repeated number of points dimension the same as the number of points of the input point cloud; the repeated feature is then concatenated with f1 a (Q) on the target point cloud feature extraction branch at the dimension level to obtain the local feature f1 b (Q) output by the feature interaction module on the target point cloud feature extraction branch; On the source point cloud feature extraction branch, f1 on the target point cloud feature extraction branch a (Q) The global feature processed by the max pooling layer is repeated in the number of points dimension through the repeat function, so that the repeated number of points dimension is the same as the number of points of the input point cloud; the repeated feature is then combined with f1 on the source point cloud feature extraction branch a (P) A connection operation is performed at the dimension level to obtain the local feature f1 output by the feature interaction module on the source point cloud feature extraction branch b (P), f1 b (Q) and f1 b (P) is a local feature matrix of 1024 * 128 dimensions: Among them, f1 b (P) and f1 b (Q) respectively represent the local features obtained after f1 a (P) and f1 a (Q) are processed by the first feature interaction module. represents the connection operation at the dimension level, repeat[·] is the repetition function, (1024,1) means repeating 1024 times in the first dimension, and sym represents the max pooling function; The 1024*128-dimensional local features output by the first feature interaction module are used as the input of the h2 part. The input dimension of the h2 part is 128 and the output dimension is 128. After passing through the h2 part, a local feature matrix of size 1024*128 is still obtained. Among them, and respectively represent the local features obtained after f1 b (P) and f1 b (Q) pass through the h2 part; The 1024*128-dimensional local feature matrix obtained after passing through the h2 part is input into the second feature interaction module to obtain 1024*256-dimensional local features: Among them, and respectively represent and the local features obtained after being processed by the second feature interaction module; The 1024*256-dimensional local features obtained after being processed by the second feature interaction module are used as the input of the h3 part. The input dimension of the h3 part is 256 and the output dimension is 512. After passing through the h3 part, a local feature matrix of size 1024*512 is obtained: Among them, f n (P) and f n (Q) respectively represent the local features of the source point cloud P and the target point cloud Q obtained in the n-th iteration, where the subscript n refers to the n-th iteration; in the above formula, the superscript a represents the output of the convolutional part, and the superscript b represents the output of the feature interaction module; The 1024*512 local features obtained after passing through the h3 part are aggregated through the max pooling layer after the h3 part to obtain global feature vectors of size 1024*1 respectively: F n (P) = sym(f n (P)) (23) F n (Q) = sym(f n (Q)) (24) Among them, F n (P) and F n (Q) respectively represent the global features of the source point cloud P and the target point cloud Q obtained in the nth iteration.

7. The point cloud registration method based on deep learning and non-corresponding point estimation according to claim 5, wherein, The process of the non-corresponding point estimation module is as follows: The local features f n (P), f n (Q) and the global feature vectors F n (P), F n (Q) are used as inputs to calculate the Euclidean distance; the Euclidean distance is normalized by the 1-softmax function into a probability, representing the probability that each point is located in the non-overlapping region. The greater the distance, the smaller the weight assigned to the point, and vice versa. The mathematical definition is shown in Equations (25)-(26); the local feature matrix is weighted and updated using the weights and passed through the max pooling layer to obtain the global features of the point cloud output by the non-corresponding point estimation module as shown in Equations (27)-(28). w(P) = 1 - softmax(||f n (P) - F n (P)||²) (25) w(Q) = 1 - softmax(||f n (Q) - F n (Q)||²) (26) Φ n (P) = sym(w(P) ⊙ f n (P)) (27) Φ n (Q) = sym(w(Q) ⊙ f n (Q)) (28) Among them, w(P) and w(Q) are 1024×1 dimensional column vectors, representing the probabilities of non-corresponding points of the source point cloud P and the target point cloud Q respectively. || ||2 represents the operation of calculating the Euclidean distance from the local feature to the current global feature, and Φ n (P) and Φ n (Q) are the global features of the source point cloud P and the target point cloud Q obtained in the nth iteration. The subscript n refers to the nth iteration, and ⊙ represents the operation of element-wise corresponding multiplication.

8. The point cloud registration method based on deep learning and non-corresponding point estimation according to claim 5, wherein, The dataset is the ModelNet40 dataset, and the clipping rate cp for the clipping operation is set to 0.8.

Citation Information

Patent Citations

  • Visual SLAM method based on deep learning

    CN113313238A

  • 3D point cloud registration network model based on double-branch feature interaction and registration method

    CN113538535A