A multi-temporal point cloud data registration method, device, and storage medium

Through the point cloud local invariant feature extraction model based on CNN neural network, the problem of deformation impact and structural matching in multi-time phase point cloud data registration of power transmission and transformation equipment is solved, and high-precision and efficient point cloud registration are achieved, which is suitable for the status evaluation of power transmission and transformation equipment.

CN115359102BActive Publication Date: 2025-07-25TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211014222.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-23
Publication Date
2025-07-25
Estimated Expiration
2042-08-23

AI Technical Summary

Technical Problem

The existing multi-time phase point cloud data registration algorithm for power transmission and transformation equipment cannot effectively eliminate the impact of deformation, cannot match the structural characteristics of the equipment, and it is difficult to take into account the registration accuracy and processing efficiency.

Method used

The point cloud local invariant feature extraction model based on CNN neural network is adopted, and the point cloud data conversion, neighborhood information supplement, resolution improvement and feature extraction modules are combined with coarse registration and refined registration algorithms to achieve high-precision registration of point cloud data.

Benefits of technology

It improves the accuracy and efficiency of point cloud registration, can effectively match the structural characteristics of power transmission and transformation equipment, and has good universality and anti-deformation interference capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359102B_ABST
    Figure CN115359102B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-temporal point cloud data registration method, device and storage medium based on local invariant feature extraction. The method includes: acquiring point cloud data; point cloud feature extraction: establishing a point cloud local invariant feature extraction model based on a CNN neural network and training it to extract the local invariant features of the point cloud data. Among them, the point cloud local invariant feature extraction model includes: a point cloud data conversion module for converting three-dimensional point clouds into three-dimensional tensors, a neighborhood information supplementation module for supplementing data in the three-dimensional tensors, a resolution enhancement module for expanding the size of the three-dimensional tensors, a feature extraction module for learning the distributed feature representation of the point cloud, and a fully connected output module for realizing feature classification and outputting the local invariant features of the point cloud; feature point matching; rough registration of the point cloud; and refined registration of the point cloud. Compared with the prior art, the present invention has the advantages of high registration accuracy, high processing efficiency, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of state assessment of power transmission and transformation equipment, and in particular to a multi-temporal point cloud data registration method, device and storage medium based on local invariant feature extraction. Background Art

[0002] The 4D assessment of the morphology of power transmission and transformation equipment can accurately describe the evolution process of equipment deformation from the spatio-temporal scale, meeting the refined requirements for equipment geometric morphology assessment. As a key technology for the 4D assessment of morphology, the multi-temporal point cloud data registration algorithm needs to meet the requirements of high accuracy and high universality.

[0003] Mainstream point cloud registration methods can be roughly divided into feature-based point cloud registration methods, point-based point cloud registration methods, mathematical model-based point cloud registration methods, and deep learning-based point cloud registration methods. Feature-based point cloud registration methods have strong interpretability and fast registration speed, and are often used for rough registration of point clouds; however, the calculation of point cloud features in such methods is based on the estimation of normal vectors or fitting surfaces, and their registration accuracy is limited by the point cloud density and the uniformity of point cloud distribution. The point-based point cloud registration method takes the ICP algorithm as the most classic representative, with the significant advantage of high registration accuracy. This type of method is sensitive to the initial position relationship of point cloud data, is prone to falling into local optima, and has a large computational cost, making it unsuitable for large-scale and high-density point cloud objects. Mathematical model-based point cloud registration methods use various mathematical models to describe the spatial distribution of point cloud data, thereby realizing the registration of point cloud data, with the advantage of fast operation speed, and are suitable for point cloud registration tasks in large scenes and high real-time requirements; however, limited by the fitting accuracy of the point cloud mathematical model, the registration accuracy of such methods is usually not high, and it is difficult to meet the requirements of high-precision point cloud registration. Deep learning-based point cloud registration methods achieve the registration of point cloud data by constructing a deep neural network, and have achieved better performance than the traditional ICP algorithm on specific point cloud datasets (ModelNet40, Kitti or 3DMatch); however, the interpretability, reproducibility and generalization ability of such methods are poor, and their computational cost will increase rapidly with the increase of point cloud density, which limits the application of such methods in actual engineering.

[0004] For the multi-temporal point cloud data registration algorithm for the 4D assessment of the morphology of power transmission and transformation equipment, on the one hand, in order to ensure the accuracy and reliability of the subsequent calculation of the difference value of the point cloud model, the algorithm needs to have excellent registration accuracy; on the other hand, since power transmission and transformation equipment often has similarity and symmetry in geometric morphology, the algorithm needs to efficiently distinguish point cloud data with similar geometric characteristics and achieve accurate matching of corresponding points; in addition, considering the possibility of stress deformation of power transmission and transformation equipment, the algorithm needs to have good anti-deformation interference ability. From the above perspectives, the existing multi-temporal point cloud data registration algorithms for the 4D assessment of the morphology of power transmission and transformation equipment have the following difficult-to-solve problems:

[0005] 1) It is impossible to eliminate the influence of the deformation of power transmission and transformation equipment;

[0006] 2) It is impossible to match the structural characteristics of power transmission and transformation equipment;

[0007] 3) It is impossible to balance the registration accuracy and processing efficiency.

[0008] Traditional point cloud registration methods only focus on the registration algorithm itself and do not consider the characteristics and registration requirements of the objects to be processed; therefore, when dealing with the multi-temporal point cloud data of power transmission and transformation equipment, traditional point cloud registration methods are difficult to achieve ideal registration effects. Summary of the Invention

[0009] The purpose of the present invention is to provide a multi-temporal point cloud data registration method, device and storage medium based on local invariant feature extraction, which solves the influence of the geometric feature similarity of the point cloud of power transmission and transformation equipment on the point cloud data registration algorithm, and balances the registration accuracy and processing efficiency, and has good universality and point cloud density robustness.

[0010] The purpose of the present invention can be achieved by the following technical solutions:

[0011] A multi-temporal point cloud data registration method based on local invariant feature extraction includes the following steps:

[0012] Obtain point cloud data;

[0013] Point cloud feature extraction: Establish a point cloud local invariant feature extraction model based on a CNN neural network and train it to extract the local invariant features of the point cloud data. The point cloud local invariant feature extraction model is a CNN neural network structure, including:

[0014] Point cloud data conversion module: used to convert three-dimensional point cloud into three-dimensional tensor,

[0015] Neighborhood information supplement module: used to supplement the data in the three-dimensional tensor,

[0016] Resolution improvement module: used to expand the size of the three-dimensional tensor,

[0017] Feature extraction module: used to learn the distributed feature representation of the point cloud,

[0018] Fully connected output module: used to map the learned distributed feature representation into the sample label space to complete feature classification and output the point cloud local invariant features;

[0019] Feature point matching: Complete feature point matching by measuring the similarity degree of the point cloud local invariant features to obtain matching point pairs;

[0020] Coarse registration of point cloud: Based on the coarse registration algorithm, using the one-to-one correspondence of matching point pairs, select the top M groups of matching point pairs with the highest to lowest similarity of local invariant features to calculate the corresponding rigid body transformation matrix, and realize the preliminary coincidence of the point cloud model;

[0021] Fine registration of point cloud: Based on the fine registration algorithm on the basis of coarse registration, obtain the optimal registration result of point cloud data.

[0022] The point cloud data includes three-dimensional point cloud and feature observation values. Among them, the three-dimensional point cloud is a three-dimensional point cloud data set composed of the target point cloud and its neighborhood points. The selection range of neighborhood points is limited within the r-neighborhood of the target point cloud, and the feature observation value is the FPFH feature.

[0023] The point cloud data conversion module includes 1 point cloud data conversion layer; the input is a three-dimensional point cloud data set P = [p i , where p0 = (x0, y0, z0) represents the target point cloud and its three-dimensional coordinates; the output is a three-dimensional tensor I of (2N + 1) × (2N + 1) × 3, where N is the three-dimensional tensor size calculation parameter, and the size of the three-dimensional tensor I is (H, W, C), H represents height, W represents width, and C represents the number of channels;

[0024] The data processing process of the point cloud data conversion module includes the following steps:

[0025] Step 2-1) Perform principal component analysis on the input point cloud data set, solve the eigenvalues λ0 ≥ λ1 ≥ λ2 and their corresponding three eigenvectors v0, v1, v2, and construct a dimensionality reduction matrix V = [v0, v1] to realize the planarization of the point cloud data Q = V T P = [q i , where q0 = (u0, v0) represents the target point cloud after planarization processing;

[0026] Step 2-2) Perform rasterization encoding processing on the planarized point cloud to establish the mapping relationship T V (q i ):

[0027]

[0028] Among them, r represents the radius for neighborhood search with the target point cloud as the center, and the function ceil(x) takes the smallest integer not less than x;

[0029] Step 2-3) According to the mapping relationship T V (q i) Numerically fill and normalize the two-dimensional slices of the three-dimensional tensor I with the coordinate values of the point cloud data. If a tensor element corresponds to multiple point cloud data, an averaging operation is performed. Among them, the number of channels of the three-dimensional tensor is configured to be 3, and there are 3 two-dimensional slices corresponding to the channels of the three-dimensional tensor, namely I(H,W,1), I(H,W,2), and I(H,W,3). The numerical filling is to fill I(H,W,1), I(H,W,2), and I(H,W,3) with the coordinate values x, y, and z of the point cloud data respectively.

[0030] The neighborhood information supplement module includes 1 grouped convolutional layer, 1 BN layer, 1 ReLU layer, and 1 average pooling layer. Among them, the grouped convolutional layer sequentially performs convolutional operations on the two-dimensional slices corresponding to each channel of the three-dimensional tensor, learns the data within the receptive field using the convolutional kernel, and completes feature mapping to achieve the augmentation of the data in the three-dimensional tensor. The average pooling layer is used to perform filtering and smoothing processing on the three-dimensional tensor data after grouped convolution.

[0031] The resolution enhancement module includes 1 grouped transposed convolutional layer, 1 BN layer, 1 ReLU layer, and 1 average pooling layer. Among them, the grouped transposed convolutional layer sequentially performs deconvolution operations on the two-dimensional slices corresponding to each channel of the three-dimensional tensor to expand the size of the three-dimensional tensor.

[0032] Among them, the discrete form of the deconvolution formula is:

[0033]

[0034] Among them, I represents the two-dimensional slice of the input three-dimensional tensor; K represents the two-dimensional convolutional kernel.

[0035] The feature extraction module includes 3 convolutional layers, 3 BN layers, and 3 ReLU layers. Using the feature extraction ability of the convolutional layers, the data stored in the three-dimensional tensor is mapped to the hidden layer feature space to achieve the learning of distributed feature representation.

[0036] The fully connected output module includes 2 fully connected layers and 1 Dropout layer.

[0037] The training and optimization of the point cloud local invariant feature extraction model adopt the Adam algorithm. The training and optimization process includes the following steps:

[0038] Step S1: Define the loss function of the CNN neural network as the mean square error of the features:

[0039]

[0040] Among them, f(P; θ) is the predicted feature value output by the CNN neural network when the input is P; F is the observed feature value; ||A|| represents the L2 norm of vector A; t i is the i-th element of the predicted feature value; y i is the i-th element of the observed feature value; n represents the number of elements in the vector;

[0041] Step S2: Define the cost function of the CNN neural network:

[0042]

[0043] Among them, is the empirical distribution, and m is the number of point cloud data samples participating in the training;

[0044] Step S3: Set the step size ε, the exponential decay rates ρ1 and ρ2 of the moment estimation, and a small constant δ for numerical stability; Initialize the parameters θ of the CNN neural network, set the first-order moment variable s = 0, the second-order moment variable r = 0, and the time step t = 0;

[0045] Step S4: Train and update the parameters of the CNN neural network;

[0046] Step S4-1: Select a mini-batch containing m data samples from the training dataset and calculate the gradient:

[0047]

[0048] Step S4-2: Accumulate the time step, t = t + 1;

[0049] Step S4-3: Update the biased first-order moment estimation s, update the biased second-order moment estimation r, and correct the bias of the first-order moment Correct the bias of the second-order moment

[0050]

[0051] Step S4-4: Update the parameters θ of the CNN neural network element by element:

[0052]

[0053] Step S5: Determine whether the time step is less than the pre-configured maximum time step. If so, return to Step S4; if not, execute Step S6;

[0054] Step S6: Complete the training of the CNN neural network to obtain the optimal parameter solution.

[0055] A multi-temporal point cloud data registration device based on local invariant feature extraction, comprising a memory, a processor, and a program stored in the memory, wherein when the processor executes the program, the method described above is implemented.

[0056] A storage medium, on which a program is stored, and when the program is executed, the method described above is implemented.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] (1) The present invention uses a point cloud local invariant feature extraction model based on a CNN neural network to extract point cloud features. Instead of overly focusing on the overall and local geometric characteristics of the point cloud, it focuses on describing the essential attributes of the point cloud. The obtained point cloud features have better specific information description capabilities and do not rely on the calculation of geometric attributes. Thus, it well solves the influence of the similarity of the point cloud structure geometric features of power transmission and transformation equipment on the point cloud data registration algorithm, and also provides help for the implementation of the local reference area registration idea.

[0059] (2) The point cloud local invariant feature extraction model based on a CNN neural network of the present invention is different from traditional point cloud data registration algorithms that require all point cloud data to participate in the registration operation. Feature point matching only needs to select the N groups of matching point pairs with the highest local invariant feature similarity to complete the calculation of the rigid body transformation matrix. This means that the execution time of the multi-temporal point cloud data registration algorithm affected by the total amount of point cloud data will be greatly reduced, and it is expected to balance registration accuracy and registration efficiency.

[0060] (3) Point cloud data presents an irregular discrete distribution in three-dimensional space, so the convolution kernel cannot be directly applied to point cloud data; the point cloud data conversion module of the present invention can convert three-dimensional point clouds into three-dimensional tensors that can be processed by the convolution kernel, which not only reflects the relative position relationship between the target point cloud and its neighboring points, but also well carries the three-dimensional coordinate information of the point cloud data, maximally retaining the individual and group characteristics of the point cloud data set. And the point cloud data conversion module does not require any predefined parameters, which can effectively reduce the experience requirements for technicians in engineering applications.

[0061] (4) The present invention realizes the supplementation of data in the three-dimensional tensor through the neighborhood information supplementation module to cope with the problem of sparse data stored in the three-dimensional tensor, and uses the average pooling layer to improve the statistical efficiency of the subsequent network for feature information.

[0062] (5) Due to the limitation of the neighborhood range, the number of input point clouds is relatively limited, which directly leads to the opposition between the spatial resolution and data sparsity of the three-dimensional tensor. To avoid the three-dimensional tensor being too sparse and affecting the feature extraction effect, the initial size of the three-dimensional tensor must be constrained. The present invention expands the size of the three-dimensional tensor through the resolution improvement module, improves the spatial resolution, is conducive to increasing the receptive field, and thus suppresses overfitting.

[0063] (6) Dropout in the point cloud local invariant feature extraction model based on the CNN neural network of the present invention can randomly discard neurons according to a given probability, ensuring that neurons can exist independently of specific other neurons, effectively reducing the co-adaptation between neurons. Both BN and ReLU can avoid the problems of gradient explosion or gradient disappearance during the training of the neural network and can accelerate the training and convergence of the neural network; BN and Dropout can, to a certain extent, avoid the phenomenon of overfitting in the neural network and improve the generalization ability of the neural network.

[0064] (7) The point cloud registration algorithm of the present invention avoids the deformed areas of power transmission and transformation equipment from participating in the registration operation through the point cloud registration idea for the reference area, thereby eliminating the influence of the deformation of power transmission and transformation equipment and being able to match the structural characteristics of power transmission and transformation equipment, and has wide applicability. Brief Description of the Drawings

[0065] Figure 1 is the flowchart of the method of the present invention;

[0066] Figure 2 is the structural schematic diagram of the point cloud local invariant feature extraction model based on the CNN neural network of the present invention;

[0067] Figure 3 is the internal structural schematic diagram of the point cloud data conversion layer;

[0068] Figure 4 is the test data display diagram of an embodiment;

[0069] Figure 5 is the cost function - iteration step number curve diagram of an embodiment;

[0070] Figure 6 is the RMSE - iteration step number curve diagram of an embodiment;

[0071] Figure 7 is the point cloud data registration effect diagram, where (a) is the initial position relationship diagram of the point cloud data, (b) is the feature point matching result diagram, (c) is the rough registration result diagram, and (d) is the fine registration result diagram;

[0072] Figure 8It is a performance analysis diagram of the experimental results of the point cloud data registration method. Among them, (a) is the performance analysis diagram of EE1-1, (b) is the performance analysis diagram of EE2-1, and (c) is the performance analysis diagram of EE3-1;

[0073] Figure 9 It is a diagram of the experimental results of the point cloud density robustness of the point cloud data registration method. Among them, (a) is the curve diagram of the RE-density change amount, (b) is the curve diagram of the TE-density change amount, and (c) is the curve diagram of the RMSE-density change amount. Specific implementation manner

[0074] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manner and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0075] A multi-temporal point cloud data registration method based on local invariant feature extraction, as Figure 1 shown, includes the following steps:

[0076] 1) Obtain point cloud data;

[0077] The point cloud data includes three-dimensional points and feature observation values.

[0078] The three-dimensional point cloud is a three-dimensional point cloud data set composed of the target point cloud and its neighborhood points, and the selection range of the neighborhood points is limited within the r-neighborhood of the target point cloud.

[0079] In this embodiment, the feature observation value is the FPFH feature. The FPFH feature can use features with lower dimensions to efficiently describe the local geometric characteristics of the point cloud, has rigid body transformation invariance, good distinguishability and lower computational complexity, and has been widely used in fields such as registration, reconstruction, and segmentation. The FPFH feature calculates the angular difference between the normal vectors of the target point and the neighborhood points; through data normalization processing and statistical analysis, the aforementioned angular difference can be converted into a parameterized vector, and the corresponding simplified point feature histogram SPFH can be obtained therefrom; by weighted summing the SPFH of the target point cloud and the points within its neighborhood, the fast point feature histogram FPFH of the target point can be obtained.

[0080] 2) Point cloud feature extraction

[0081] Establish a point cloud local invariant feature extraction model based on a CNN neural network and train it to extract the local invariant features of the point cloud data.

[0082] The point cloud local invariant feature extraction model is a CNN neural network structure, and its structure schematic diagram is as Figure 2 shown, including:

[0083] ① Point cloud data conversion module: used to convert 3D point cloud into 3D tensor, and its internal structure is as Figure 3 shown.

[0084] Different from the uniform and ordered arrangement of image pixel matrix, point cloud data presents an irregular discrete distribution in 3D space, so the convolution kernel cannot be directly applied to point cloud data; the point cloud data conversion module can convert 3D point cloud into 3D tensor that can be processed by the convolution kernel.

[0085] The point cloud data conversion module includes 1 point cloud data conversion layer (Point Cloud T-Layer); the input is a 3D point cloud data set P = [p i , where p0 = (x0, y0, z0) represents the target point cloud and its 3D coordinates; the output is a 3D tensor I of (2N + 1) × (2N + 1) × 3, where N is the 3D tensor size calculation parameter, and in this embodiment, N takes 5. The size of the 3D tensor I is (H, W, C), where H represents height, W represents width, and C represents the number of channels.

[0086] The data processing process of the point cloud data conversion module includes the following steps:

[0087] Step 2-1) Perform principal component analysis on the input point cloud data set, solve the eigenvalues λ0 ≥ λ1 ≥ λ2 and their corresponding three eigenvectors v0, v1, v2, and construct a dimensionality reduction matrix V = [v0, v1] to realize the planarization of the point cloud data Q = V T P = [q i , where q0 = (u0, v0) represents the target point cloud after planarization processing;

[0088] Step 2-2) Perform rasterization encoding processing on the planarized point cloud to establish the mapping relationship T V (q i ):

[0089]

[0090] where r represents the radius of neighborhood search with the target point cloud as the center, and the function ceil(x) takes the smallest integer not less than x;

[0091] Step 2-3) According to the mapping relationship T V (q i) Numerically fill and normalize the two-dimensional slices of the three-dimensional tensor I with the coordinate values of the point cloud data. If a tensor element corresponds to multiple point cloud data, an averaging operation is performed. Among them, the number of channels of the three-dimensional tensor is configured to be 3, and there are 3 two-dimensional slices corresponding to the channels of the three-dimensional tensor, namely I(H,W,1), I(H,W,2), and I(H,W,3). The numerical filling is to fill I(H,W,1), I(H,W,2), and I(H,W,3) with the coordinate values x, y, and z of the point cloud data respectively.

[0092] ② Neighborhood information supplement module: used to supplement the data in the three-dimensional tensor.

[0093] The neighborhood information supplement module includes 1 group convolution layer (Group Convolution, G-Conv), 1 BN layer, 1 ReLU layer, and 1 average pooling layer (Average Pooling, AvgPool).

[0094] The point cloud data is obtained by sampling with a lidar along the horizontal and vertical directions at a certain angular interval step by step. Discreteness is its inherent property; while the three-dimensional tensor generated by the point cloud data conversion module corresponds to a continuous surface area; therefore, the data stored in the three-dimensional tensor is relatively sparse. The group convolution layer can sequentially perform convolution operations on the two-dimensional slices corresponding to each channel of the three-dimensional tensor, learn the data within the receptive field using the convolution kernel and complete feature mapping to achieve the supplementation of the data in the three-dimensional tensor; the average pooling layer is used to filter and smooth the data of the three-dimensional tensor after group convolution, which can effectively improve the statistical efficiency of the subsequent network for feature information.

[0095] ③ Resolution improvement module: used to expand the size of the three-dimensional tensor.

[0096] The resolution improvement module includes 1 group transposed convolution layer (Group Deconvolution, G-Deconv), 1 BN layer, 1 ReLU layer, and 1 average pooling layer (Average Pooling, AvgPool).

[0097] Due to the limitation of the neighborhood range, the number of input point clouds is relatively limited; this directly leads to the opposition between the spatial resolution and data sparsity of the three-dimensional tensor. In order to avoid the three-dimensional tensor being too sparse and affecting the effect of feature extraction, the initial size of the three-dimensional tensor must be restricted. The group transposed convolution layer sequentially performs deconvolution operations on the two-dimensional slices corresponding to each channel of the three-dimensional tensor to expand the size of the three-dimensional tensor.

[0098] Deconvolution is also known as Fractially Strided Convolution or Transposed Convolution. Similar to the convolution operation, the deconvolution operation can also be abstracted into a discrete form of the convolution formula.

[0099]

[0100] Among them, I represents the two-dimensional slice of the input three-dimensional tensor, and K represents the two-dimensional convolution kernel.

[0101] ④ Feature extraction module: used to learn the distributed feature representation of the point cloud.

[0102] The feature extraction module includes 3 Convolution (Conv) layers, 3 Batch Normalization (BN) layers, and 3 Rectified Linear Unit (ReLU) layers. Utilizing the feature extraction ability of the convolution layers, the data stored in the three-dimensional tensor is mapped to the hidden layer feature space to achieve the learning of the distributed feature representation.

[0103] ⑤ Fully connected output module: used to map the learned distributed feature representation into the sample label space, complete feature classification, and output the local invariant features of the point cloud.

[0104] The fully connected output module includes 2 Fully Connected Layer (FC) layers and 1 Dropout layer. The fully connected layer can map the learned distributed feature representation into the sample label space and act as a "classifier".

[0105] In the above functional modules, BN (Batch Normalization) mentioned represents the batch normalization layer, which can normalize the input data of the activation layer; ReLU (Rectified Linear Unit) represents the rectified linear function f(x) = max(0, x), which belongs to the activation layer function and can introduce non-linear factors into the neural network; Dropout represents the dropout layer, which can randomly discard neurons according to a given probability to ensure that neurons can exist independently of specific other neurons, effectively reducing the co-adaptation between neurons. Both BN and ReLU can avoid the problems of gradient explosion or gradient disappearance during the training of the neural network and can accelerate the training and convergence of the neural network; BN and Dropout can, to a certain extent, avoid the phenomenon of overfitting in the neural network and improve the generalization ability of the neural network.

[0106] In this embodiment, the specific structural parameters of the CNN neural network are shown in Table 1.

[0107] Table 1 Point Cloud Local Invariant Feature Extraction Network

[0108] Neural network layer Parameter Grouped convolution Size = 3×3; Stride = 1×1; Number of groups = 3 Grouped transposed convolution Size = 3×3; Stride = 2×2; Number of groups = 3 Average pooling layer Size = 2×2; Stride = 1×1 Convolutional layer 1 Size = 3×3; Stride = 2×2; Number of kernels = 16 Convolutional layer 2 Size = 3×3; Stride = 2×2; Number of kernels = 32 Convolutional layer 3 Size = 3×3; Stride = 2×2; Number of kernels = 64 Fully connected layer 1 Number of inputs = 256; Number of outputs = 100 Fully connected layer 2 Number of inputs = 100; Number of outputs = 33

[0109] The training process of the CNN convolutional neural network is a process of optimizing the algorithm to find the optimal solution of the cost function J(θ) and improve the performance metric P. Therefore, the choice of the optimization algorithm directly affects the final training result of the neural network. The Stochastic Gradient Descent (SGD) algorithm and its variants are the most widely used optimization algorithms in deep learning. In convex optimization problems, SGD can descend along the gradient direction of randomly selected small batches of data to obtain the optimal solution. In the actual neural network training process, non-convex situations are often encountered. If the hyperparameters are not properly selected, SGD may fall into a local optimal solution. At the same time, SGD uses the same learning rate for the update of each parameter, ignoring the different update requirements of each parameter, which limits the final parameter optimization result. Therefore, in this embodiment, the Adaptive Moment Estimation (Adam) algorithm is used to optimize the training of the neural network. The Adam algorithm is an adaptive learning rate algorithm that can set different learning rates for each parameter and automatically adapt to these learning rates during the training process of the neural network, thereby accelerating the convergence of the neural network and avoiding falling into a local optimum.

[0110] The training optimization process includes the following steps:

[0111] Step S1: Define the loss function of the CNN neural network as the mean square error of the features:

[0112]

[0113] where f(P; θ) is the feature prediction value output by the CNN neural network when the input is P; F is the feature observation value; ||A|| represents the L2 norm of vector A; t i is the i-th element of the feature prediction value; y i is the i-th element of the feature observation value; n represents the number of elements of the vector;

[0114] Step S2: Define the cost function of the CNN neural network:

[0115]

[0116] where is the empirical distribution, and m is the number of point cloud data samples participating in the training;

[0117] Step S3: Set the step size ε, the exponential decay rates ρ1 and ρ2 of the moment estimation, and a small constant δ for numerical stability; initialize the CNN neural network parameters θ, set the first-order moment variable s = 0, the second-order moment variable r = 0, and the time step t = 0;

[0118] Step S4: Train and update the parameters of the CNN neural network;

[0119] Step S4-1: Select a mini-batch containing m data samples from the training dataset and calculate the gradient:

[0120]

[0121] Step S4-2: Accumulate in time steps, t = t + 1;

[0122] Step S4-3: Update the biased first moment estimate s, update the biased second moment estimate r, and correct the bias of the first moment Correct the bias of the second moment

[0123]

[0124] Step S4-4: Update the parameters θ of the CNN neural network element by element:

[0125]

[0126] Step S5: Determine whether the time step is less than the pre-configured maximum time step. If so, return to Step S4; if not, execute Step S6;

[0127] Step S6: Complete the training of the CNN neural network to obtain the optimal parameter solution.

[0128] 3) Feature point matching: Measure the similarity of the local invariant features of the point cloud and complete the feature point matching to obtain the matching point pairs.

[0129] 4) Coarse registration of the point cloud: Based on the coarse registration algorithm, using the one-to-one correspondence of the matching point pairs, select the top M groups of matching point pairs with the highest similarity of local invariant features from high to low to calculate the corresponding rigid body transformation matrix, and realize the preliminary coincidence of the point cloud models.

[0130] The coarse registration algorithm includes the SAC-IA algorithm.

[0131] 5) Fine registration of the point cloud: Based on the fine registration algorithm on the basis of coarse registration, perform fine registration to obtain the optimal point cloud data registration result.

[0132] The fine registration algorithm includes the ICP algorithm.

[0133] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0134] In order to verify the registration effect of a multi-temporal point cloud data registration method based on local invariant feature extraction described above, this embodiment conducts a case test on it. The test platform uses an Intel(R) Core(TM) i7-6700HQ CPU@2.60GHz, 8.00GB of memory, an NVIDIA GeForce GTX960M, and a Windows10 64-bit operating system; the algorithm is built based on the MATLAB R2021b application program.

[0135] The test data set and environment of this embodiment are as follows:

[0136] There are various types of power transmission and transformation equipment with different functions, and the equipment of each category is completely different in geometric appearance. In order to ensure that the test data has good representativeness in terms of geometric shape, this article divides the experimental data into 3 groups: 1) Base-type equipment (such as Figure 4 EE1-1, EE1-2, EE1-3 in Figure 4 ): It can be clearly distinguished from the geometric shape into two parts: the base and the non-base. Among them, the base structure usually has characteristics such as adjacent to the base surface, large rigid strength, and not easy to deform; 2) Columnar-type equipment (such as

[0137] EE2-1, EE2-2, EE2-3 in Figure 4 ): It is composed of several columnar structures from the geometric shape, and there is no significant boundary between each part, and its end structure is relatively not easy to deform;

[0138] During the training process of the CNN neural network, a total of 28,317 training sample data from 5 electrical devices were selected. The number of batch samples processed in each iteration was 256, and the number of iterations was 3,300 steps. The downward trend of the cost function is as Figure 5 shown. The cost function is the average of the mean square error (MSE) of the mini-batch samples, which reflects the convergence and fitting degree of the model training. From Figure 5 it can be seen that when the number of iterations reaches 1,000 steps, the model has tended to converge, and the final value of the cost function is stable at about 0.9. This means that the trained model already has the ability to describe point cloud features similar to the FPFH features.

[0139] During the training process of the neural network model, the change trend of the root mean square error (RMSE) of the sample values is as Figure 6 shown. From Figure 6 it can be seen that the final value of the RMSE after the model converges is stable at about 1.35. This means that the model has characteristics different from the FPFH features in the feature description of the point cloud.

[0140] To test the processing effect and universality of the algorithm proposed in this paper, experimental studies were carried out on 3 groups of test data (a total of 9 electrical devices), and the experimental results are as Figure 7 shown. Figure 7 (a) successively shows the relative position relationship between the reference point cloud and the target point cloud of 9 electrical devices in three-dimensional space. From Figure 7 (a), it can be seen that there are large translation and rotation deviations between the original reference point cloud and the target point cloud.

[0141] The results of point cloud feature extraction and feature point matching in this embodiment are as Figure 7 (b) shown. From Figure 7 (b), it can be obtained from the matching results that the feature points in the reference point cloud (only located in the benchmark area of the reference point cloud) can well find the corresponding matching points (located in the target point cloud). The rough registration result of this embodiment is shown in Figure Figure 7 (c). From Figure 7 (c), it can be seen that the rough registration can well optimize the relative position relationship between the target point cloud and the reference point cloud. The refined registration result of this embodiment is as Figure 7 (d) shown. From Figure 7 (d), it can be seen that the overlap degree between the target point cloud and the reference point cloud in three-dimensional space has been further improved, and an ideal point cloud registration effect has been achieved. Therefore, the algorithm proposed in this paper has both an ideal registration effect and universality for various types of power transmission and transformation equipment.

[0142] To further test the registration accuracy and processing efficiency of the point cloud registration method described in the present invention, one test sample (EE1-1, EE2-1, EE3-1) was selected from each of the 3 groups of test data for quantitative comparative analysis of the algorithm; the FPFH algorithm and the SIFT algorithm were selected as comparison objects, and the method proposed in the present invention was named the CNN algorithm. For the registration accuracy of the algorithm, it will be evaluated from three aspects: 1) the root mean square error RMSE of registration; 2) the rotation error RE of registration; 3) the translation error TE of registration; the calculation formulas for these three registration accuracy evaluation indicators are as follows,

[0143]

[0144] where N is the number of reference point clouds, p i represents the reference point cloud, and q i represents the set of nearest points corresponding to p i ; re x and re y and re z represent the rotation errors generated during rotation around the x, y, and z axes during the rigid body transformation process respectively; te x and te y and te z represent the translation errors generated during translation along the x, y, and z axes during the rigid body transformation process respectively. For the processing efficiency of the algorithm, it is also evaluated from three aspects: 1) the time t0 for feature point matching; 2) the time t1 for initial registration; 3) the time t2 for fine registration. The performance analysis diagram of the experimental results for testing the registration accuracy and processing efficiency of the algorithm is as Figure 8 shown.

[0145] Analysis Figure 8 From the RE index in Figure 8 , it can be obtained that the final value of RE generated by the CNN algorithm is much smaller than that of the FPFH algorithm and the SIFT algorithm. The accuracy of the rigid body rotation transformation of the CNN algorithm is improved by about 93.5% - 99.5% compared with the FPFH algorithm and by about 93.0% - 99.5% compared with the SIFT algorithm. Analysis Figure 8From the RMSE index in it, it can be seen that the final RMSE value generated by the CNN algorithm is much smaller than that of the FPFH algorithm and the SIFT algorithm. The overall registration accuracy of the CNN algorithm is improved by about 74.5% - 96.0% compared with the FPFH algorithm and about 73.5% - 96.0% compared with the SIFT algorithm. To sum up, the CNN algorithm can significantly improve the calculation accuracy of the rigid body transformation matrix, optimize the spatial position of the target point cloud, and effectively improve the overall registration accuracy.

[0146] Analysis Figure 8 From the T0 index in it, it can be seen that the time taken by the CNN algorithm for feature point matching is about 2 - 4 times more than that of the FPFH algorithm and about 5 - 7 times more than that of the SIFT algorithm. The possible reason is that limited by the performance of GPU resources, the efficiency of the CNN algorithm in extracting local invariant features of the point cloud is relatively low, thus increasing the time for feature point matching. Analysis Figure 8 From the T1 index in it, it can be seen that the time for initial registration of the three algorithms is not much different, staying between 1 - 2.5 s. Analysis Figure 8 From the T2 index in it, it can be seen that the time for fine registration of the three algorithms is not much different, staying between 0.45 - 1.5 s. Although the overall processing efficiency of the CNN algorithm is slightly inferior to that of the FPFH algorithm and the SIFT algorithm, the total time required for the CNN algorithm to register power transmission and transformation equipment does not exceed 175 s, meeting the needs of actual engineering applications.

[0147] To test the robustness of the CNN algorithm to point cloud density transformation, device 1 - 1 is selected from the test dataset for experimental research. The original number of points in the point cloud of device 1 - 1 is 221774, and after uniform downsampling, it is 177419. The average number of points in the r - neighborhood (r = 0.05 m) of it is 99.44. The experimental results of the point cloud density robustness are as Figure 9 shown.

[0148] When the change amount of point cloud density Δρ (point cloud density ρ = 1 - Δρ) increases from 0 to 0.3, from Figure 9 (a), it can be seen that the registration rotation error generated by the CNN algorithm remains at about 5.5×10 -5 rad, and the average value of its change amplitude is about 1.75×10 - 5 rad; while the registration rotation errors generated by the FPFH algorithm and the SIFT algorithm show a violent fluctuation characteristic, and the average value of their change amplitudes is about 6.4 times and 8.4 times that of the CNN algorithm. From Figure 9 (b), it can be seen that the registration translation error generated by the CNN algorithm remains at about 3.39×10 -3 m, and the average value of its change amplitude is about 6.95×10 -4m; while the registration translation errors generated by the FPFH algorithm and the SIFT algorithm show a fluctuating upward trend, and the average value of their change amplitudes is about 13 times and 15 times that of the algorithm proposed in this paper. From Figure 9 (c), it can be seen that the registration root mean square error generated by the CNN algorithm can be divided into two segments of change curves. The root mean square error in the first curve (corresponding to the gray box 1) remains at about 5.27×10 -4 m, and the average value of its change amplitude is about 1.74×10 - 4 rad; the root mean square error in the second curve (corresponding to the gray box 2) remains at about 2.91×10 -3 m, and the average value of its change amplitude is about 2.35×10 -4 rad; although there is a certain mutation between the two curves, the registration accuracy corresponding to each curve is significantly better than that of the FPFH algorithm and the SIFT algorithm.

[0149] When the change amount of the point cloud density Δρ (point cloud density ρ = 1 - Δρ) increases from 0.3 to 0.5, from Figure 9 (a) and Figure 9 (c), it can be seen that the registration rotation error and root mean square error generated by the CNN algorithm are already close to those of the FPFH algorithm and the SIFT algorithm in terms of numerical value and change trend. From Figure 9 (b), it can be seen that the registration translation error generated by the CNN algorithm is also gradually approaching that of the FPFH algorithm and the SIFT algorithm in terms of numerical value, but its change trend is still better than that of the FPFH algorithm and the SIFT algorithm.

[0150] In summary, the CNN algorithm has good robustness to the point cloud density and can ensure to a certain extent that the overall registration accuracy does not decrease with the reduction of the point cloud density; in other words, on the premise of meeting a certain registration accuracy (RMSE < 3mm), the minimum value of the r-neighborhood (r = 0.05m) data volume required by the CNN algorithm is 25.5% less than that of the traditional algorithm, which means that the algorithm proposed in this paper can reduce the requirements for the point cloud density, is beneficial to improving the operation efficiency and reducing the point cloud scanning and acquisition cost.

[0151] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. A multi-temporal point cloud data registration method based on local invariant feature extraction, characterized in that It includes the following steps: Obtain point cloud data; Point cloud feature extraction: Establish and train a point cloud local invariant feature extraction model based on a CNN neural network to extract the local invariant features of the point cloud data. Among them, the point cloud local invariant feature extraction model is a CNN neural network structure, including: Point cloud data conversion module: Used to convert three-dimensional point cloud into three-dimensional tensor, Neighborhood information supplement module: Used to supplement the data in the three-dimensional tensor, Resolution improvement module: Used to expand the size of the three-dimensional tensor, Feature extraction module: Used to learn the distributed feature representation of the point cloud, Fully connected output module: Used to map the learned distributed feature representation into the sample label space to complete feature classification and output the point cloud local invariant features; Feature point matching: Complete feature point matching by measuring the similarity of the point cloud local invariant features to obtain matching point pairs; Coarse registration of the point cloud: Based on the coarse registration algorithm, using the one-to-one correspondence of the matching point pairs, select the pre-configured number of matching point pairs in the order from high to low according to the similarity of the local invariant features, and calculate the corresponding rigid body transformation matrix to achieve the preliminary coincidence of the point cloud model; Fine registration of the point cloud: Based on the fine registration algorithm on the basis of coarse registration to perform fine registration to obtain the optimal point cloud data registration result.

2. A multi-temporal point cloud data registration method based on local invariant feature extraction according to claim 1, characterized in that, The point cloud data includes three-dimensional point cloud and feature observation values. Among them, the three-dimensional point cloud is a three-dimensional point cloud dataset composed of the target point cloud and its neighborhood points. The selection range of the neighborhood points is limited within the r-neighborhood of the target point cloud, and the feature observation value is the FPFH feature.

3. A multi-temporal point cloud data registration method based on local invariant feature extraction according to claim 2, characterized in that The point cloud data conversion module includes one point cloud data conversion layer; the input is a three-dimensional point cloud data set P = [p i , where p0 = (x0, y0, z0) represents the target point cloud and its three-dimensional coordinates; the output is a three-dimensional tensor I of size (2N + 1) × (2N + 1) × 3, where N is a three-dimensional tensor size calculation parameter, and the size of the three-dimensional tensor I is (H, W, C), H represents height, W represents width, and C represents the number of channels; The data processing process of the point cloud data conversion module includes the following steps: Step 2-1) Perform principal component analysis on the input point cloud dataset, solve the eigenvalues λ0≥λ1≥λ2 and their corresponding three eigenvectors v0, v1, v2, construct the dimensionality reduction matrix V = [v0, v1], and realize the planarization of the point cloud data Q = V T P = [q i , where q0 = (u0, v0) represents the target point cloud after planarization processing; Step 2-2) Perform rasterization encoding on the flattened point cloud to establish the mapping relationship T between the point cloud data and the three-dimensional tensor V (q i ): Among them, r represents the radius of neighborhood search with the target point cloud as the center, and the function ceil(x) takes the smallest integer not less than x; Step 2-3) According to the mapping relationship T V (q i ), numerically fill and normalize the two-dimensional slices of the three-dimensional tensor I with the coordinate values of the point cloud data respectively. If multiple point cloud data correspond to the same tensor element, an averaging operation is performed. Among them, the number of channels of the three-dimensional tensor is configured to be 3, and there are 3 two-dimensional slices corresponding to the channels of the three-dimensional tensor, namely I(H,W,1), I(H,W,2), and I(H,W,3). The numerical filling is to fill I(H,W,1), I(H,W,2), and I(H,W,3) with the coordinate values x, y, and z of the point cloud data respectively.

4. A multi-temporal point cloud data registration method based on local invariant feature extraction according to claim 3, characterized in that The neighborhood information supplement module includes 1 grouped convolution layer, 1 BN layer, 1 ReLU layer and 1 average pooling layer. Among them, the grouped convolution layer sequentially performs convolution operations on the two-dimensional slices corresponding to each channel of the three-dimensional tensor, uses the convolution kernel to learn the data within the receptive field and completes feature mapping to achieve the supplementation of the data in the three-dimensional tensor; the average pooling layer is used to perform filtering and smoothing processing on the three-dimensional tensor data after grouped convolution.

5. A multi-temporal point cloud data registration method based on local invariant feature extraction according to claim 4, characterized in that The resolution improvement module includes 1 grouped transposed convolution layer, 1 BN layer, 1 ReLU layer and 1 average pooling layer. Among them, the grouped transposed convolution layer sequentially performs deconvolution operations on the two-dimensional slices corresponding to each channel of the three-dimensional tensor to expand the size of the three-dimensional tensor, Among them, the discrete form of the deconvolution formula is: Among them, I represents the two-dimensional slice of the input three-dimensional tensor; K represents the two-dimensional convolution kernel.

6. A multi-temporal point cloud data registration method based on local invariant feature extraction according to claim 1, characterized in that The feature extraction module includes 3 convolution layers, 3 BN layers and 3 ReLU layers. Using the feature extraction ability of the convolution layer, the data stored in the three-dimensional tensor is mapped into the hidden layer feature space to achieve the learning of the distributed feature representation.

7. A multi-temporal point cloud data registration method based on local invariant feature extraction according to claim 1, characterized in that The fully connected output module includes 2 fully connected layers and 1 Dropout layer.

8. A multi-temporal point cloud data registration method based on local invariant feature extraction according to claim 2, characterized in that The training optimization of the point cloud local invariant feature extraction model adopts the Adam algorithm, and the training optimization process includes the following steps: Step S1: Define the loss function of the CNN neural network as the mean square error of features: where f(P; θ) is the predicted feature value output by the CNN neural network when the input is P; F is the observed feature value; ||A|| represents the L2 norm of vector A; t i is the i-th element of the predicted feature value; y i is the i-th element of the observed feature value; n represents the number of elements in the vector; Step S2: Define the cost function of the CNN neural network: Among them, is the empirical distribution, and m is the number of point cloud data samples participating in the training; Step S3: Set the step size ε, the exponential decay rates ρ1 and ρ2 of the moment estimation, and a small constant δ for numerical stability; initialize the parameters θ of the CNN neural network, set the first-order moment variable s = 0, the second-order moment variable r = 0, and the time step t = 0; Step S4: Train and update the parameters of the CNN neural network; Step S4-1: Select a mini-batch containing m data samples from the training dataset and calculate the gradient: Step S4-2: Accumulate the time step, t = t + 1; Step S4-3: Update the biased first moment estimate s, update the biased second moment estimate r, and correct the bias of the first moment Correct the bias of the second moment Step S4-4: Update the parameters θ of the CNN neural network element-wise: Step S5: Determine whether the time step is less than the pre-configured maximum time step. If so, return to Step S4; if not, execute Step S6; Step S6: Complete the training of the CNN neural network to obtain the optimal parameter solution.

9. A multi-temporal point cloud data registration device based on local invariant feature extraction, comprising a memory, a processor, and a program stored in the memory, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1-8.

10. A storage medium having a program stored thereon, characterized in that, When the program is executed, it implements the method described in any one of claims 1-8.