Building point cloud reconstruction method of DGCNN model based on improved KNN
By improving the KNN algorithm combined with KD tree to optimize the DGCNN model, DGCNN is solved with the problem of high computational complexity in building point cloud reconstruction, achieving faster processing speed and higher accuracy, and is suitable for point cloud data reconstruction of large-scale complex buildings.
Patent Information
- Application Number
- CN202510094550.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-21
AI Technical Summary
In the prior art, the calculation complexity of the KNN search method in the reconstruction of building point clouds is high, resulting in a sharp increase in computing time, increased hardware costs and accuracy, and it is difficult to process point cloud data of large-scale complex buildings.
The improved KNN algorithm is used to combine KD trees to construct the neighborhood relationship of point cloud data. By recursively dividing the data space, the complexity of KNN search is reduced, and the adjacency graph is dynamically adjusted during network training to optimize memory management.
It significantly improves the speed and accuracy of large-scale point cloud data processing, reduces computing and storage costs, improves the performance and scalability of the DGCNN algorithm, and can better process point cloud data in complex buildings.
Smart Images

Figure CN120070746A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of point cloud data processing, and particularly relates to a building point cloud reconstruction method based on an improved KNN DGCNN model. Background Art
[0002] Point cloud data processing technology is a key technology in the fields of modern computer vision and robotics, and is widely used in fields such as 3D reconstruction and virtual reality. In the modern architecture field, building point cloud 3D reconstruction technology is of crucial significance for aspects such as architectural design, historical building protection, and urban planning. By obtaining the point cloud information of a building, the three-dimensional form of the building can be restored, providing a reliable basis for subsequent analysis and transformation. Point cloud information processing technology is the key to achieving this goal. Among them, the method based on graph convolutional neural network (GCN), especially the dynamic graph convolutional neural network (DGCNN), has shown an important role in the processing and analysis of point cloud information. DGCNN analyzes the neighborhood graph structure of each point and performs convolutional operations on the graph structure to extract local features, thereby realizing the classification, segmentation, and feature extraction of building point cloud data, providing technical support for 3D reconstruction.
[0003] In the prior art, the KNN search method used in DGCNN mainly relies on brute-force search, and the computational complexity of this method is O(N 2 ), where N is the number of points. This brute-force search method exposes many serious problems when processing building point cloud data. With the continuous expansion of building scale, increasing details, and improvement of measurement technology accuracy, the amount of point cloud data shows a rapid growth trend. For example, when collecting point clouds of building groups such as large commercial complexes, dense residential communities, or historical and cultural blocks, the number of data points often easily exceeds millions or even tens of millions. Under such a huge information scale, the computational time of brute-force search will increase exponentially. For each additional data point, distance calculations and comparisons need to be made with all existing points, which makes the training and inference processes extremely slow, seriously restricting the speed and application scope of building point cloud 3D reconstruction.
[0004] Moreover, this high computational complexity also brings high hardware costs and energy consumption. To complete the nearest neighbor search of large-scale building point cloud data, powerful computing devices and a large amount of memory resources are required. This not only increases the equipment procurement and maintenance costs of enterprises and research institutions, but also makes it difficult for some units with limited resources to carry out relevant work. At the same time, the long computational process also causes a large amount of energy waste, which does not conform to the current development concept of energy conservation and environmental protection.
[0005] In addition, it is difficult to guarantee the accuracy of the brute-force search method when dealing with complex building structures. Buildings have diverse shapes and complex internal structures, and the point cloud data is unevenly distributed and interfered by noise. When facing these situations, the brute-force search is easily affected by local extrema and noise points, resulting in the found nearest neighbor points not being truly representative and relevant points, thus affecting the subsequent feature extraction and 3D reconstruction quality, making the reconstruction results deviate, details be lost, or even the overall structure be distorted, and unable to achieve the goal of high-precision reconstruction in the construction field. Therefore, optimizing the KNN search algorithm is crucial for improving the performance of building point cloud 3D reconstruction. Summary of the Invention
[0006] Aiming at the problems of large computational resource occupancy and complexity existing in the KNN algorithm in the dynamic graph convolutional neural network (DGCNN) in the scenario of processing large-scale point cloud data in 3D reconstruction, the present invention provides a building point cloud reconstruction method based on an improved KNN-based DGCNN model. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0007] The present invention provides a building point cloud reconstruction method based on an improved KNN-based DGCNN model, including:
[0008] S1: Normalize the original building point cloud to be reconstructed to obtain the normalized point cloud data;
[0009] S2: Construct a DGCNN network based on the improved KNN algorithm, and train the DGCNN network to obtain a trained DGCNN model. The DGCNN network based on the improved KNN algorithm includes a spatial transformation layer, four graph convolutional layers, a max pooling layer, a first multi-layer perceptron, and a second multi-layer perceptron connected in sequence. Among them, the spatial transformation layer is used to perform spatial transformation on the input normalized point cloud data to make the point cloud data invariant in space; the graph convolutional layer is used to use the improved KNN algorithm to obtain k nearest neighbor points of each point in the building point cloud, and use the k nearest neighbor points for feature aggregation and output the aggregated feature vector; the max pooling layer is used to receive the feature vector output by the last graph convolutional layer and perform a max pooling operation in the dimension of points to obtain a global feature vector; the first multi-layer perceptron is used to abstract and transform the global feature vector to extract higher-level semantic information; the second multi-layer perceptron includes multiple hidden layers and is used to output the predicted category and the corresponding predicted score;
[0010] S3: Input the normalized point cloud data into the trained DGCNN model to obtain the corresponding prediction result.
[0011] In an embodiment of the present invention, the S1 includes:
[0012] Normalize the 6D feature vectors of each point in the original building point cloud respectively. The 6D feature vector includes 3D spatial coordinates (x, y, z) and 3D color information (R, G, B).
[0013] In one embodiment of the present invention, the graph convolutional layer is specifically used for:
[0014] Construct a KD tree according to the data features input to the current graph convolutional layer;
[0015] Use the constructed KD tree to search for the K nearest neighbor points of each point in the point cloud;
[0016] Aggregate the information of the K nearest neighbor points using a predefined weight rule to obtain the updated data features of the current point;
[0017] Process the aggregated data features through an activation function to improve the recognition accuracy of the building point cloud data.
[0018] In one embodiment of the present invention, constructing a KD tree according to the data features input to the current graph convolutional layer includes:
[0019] For k-dimensional data features, create index arrays for each dimension respectively to form a super key (a 1 , a 2 , a 2 , …, a k ), where a 1 , a 2 , a 2 , …, a k respectively represent the data values of the 1st to kth dimensions of the current point;
[0020] Compare and sort the points in sequence according to the front-back order of the super keys of each point to obtain the sorted index array of all points;
[0021] Based on the sorted index array of all points, partition all data features according to the median element to obtain a left subtree and a right subtree, and then progressively construct a KD tree structure.
[0022] In one embodiment of the present invention, aggregating the information of the K nearest neighbor points using a predefined weight rule to obtain the updated data features of the current point includes:
[0023] In the graph convolutional layer, use an improved KNN algorithm to obtain the K nearest neighbor points of each point in the input feature vector for feature aggregation, and the aggregation formula is:
[0024]
[0025] Among them, H () represents the input feature vector of the l-th layer graph convolutional layer, with a size of n×d l , where n is the number of points, and d l is the feature dimension of the input feature vector H () , \(N(i)\) represents the set of K nearest neighbor points found by point i through the KD tree, and W () is the learnable weight matrix of the l-th layer graph convolutional layer, and σ is the activation function. represents the feature vector of point j in the l-th layer graph convolutional layer, \(N(j)\) represents the set of K nearest neighbor points found by point j through the KD tree, and H (+1) represents the aggregated feature vector.
[0026] In one embodiment of the present invention, the loss function of the DGCNN model is the cross-entropy loss, and the expression is:
[0027]
[0028] where M is the number of training samples used in the training process, Q is the number of class labels in the training samples, and y mq represents the probability that the m-th training sample belongs to the q-th category, represents the probability that the predicted value of the m-th training sample belongs to the q-th category.
[0029] In one embodiment of the present invention, training the DGCNN network to obtain a trained DGCNN model includes:
[0030] Set the initial learning rate, coefficient parameters, and training batches, collect building point clouds to form a training data set, and each building point cloud is labeled with a target type label;
[0031] During each batch training, for each point cloud data sample in the training data set, first perform normalization processing, input the normalized point cloud data sample into the DGCNN network to be trained, calculate the prediction result through forward propagation, and calculate the loss value according to the loss function;
[0032] Calculate the gradient according to the loss value through the backpropagation algorithm, and update the parameters of the DGCNN network;
[0033] Every 5 batches of training, evaluate the accuracy of the DGCNN network on the validation data set. If the accuracy does not improve for 3 consecutive batches, then reduce the learning rate according to the formula η new = old ×0.1 for subsequent training, where η old represents the current learning rate, and η new represents the updated learning rate.
[0034] Another aspect of the present invention provides a storage medium in which a computer program is stored, and the computer program is used to execute the steps of the building point cloud reconstruction method based on the improved KNN-based DGCNN model described in any one of the above embodiments.
[0035] Another aspect of the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the building point cloud reconstruction method based on the improved KNN-based DGCNN model described in any one of the above embodiments are implemented.
[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0037] 1. The building point cloud reconstruction method based on the improved KNN-based DGCNN model of the present invention provides an optimization scheme for KNN search based on the KD tree. The neighborhood relationship of point cloud data is constructed through the KD tree data structure. The KD tree recursively divides the data space into hyper-rectangular regions, so that the query time complexity of KNN search is reduced from the complexity O(N 2 ) to O(logN). The present invention has brought extremely significant beneficial effects in the field of building point cloud three-dimensional reconstruction.
[0038] 2. Through the local update mechanism of the KD tree, the present invention can efficiently and dynamically adjust the adjacency graph during the network training process, avoiding the high computational cost of reconstructing the entire search tree. At the same time, the present invention adopts an optimized memory management strategy, effectively reducing memory consumption. This method not only improves the processing speed of large-scale point cloud data, but also can maintain the query efficiency of high-dimensional data, solves the bottleneck of existing KNN search in terms of calculation and storage, improves the overall performance and scalability of the DGCNN algorithm, and this method can also be extended to point cloud processing neural networks with similar structures.
[0039] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. Description of the Drawings
[0040] Figure 1 is a flowchart of a building point cloud reconstruction method based on an improved KNN-based DGCNN model provided by an embodiment of the present invention;
[0041] Figure 2 is a schematic structural diagram of a DGCNN network based on an improved KNN algorithm for processing point cloud data provided by an embodiment of the present invention;
[0042] Figure 3It is a schematic flowchart of constructing and classifying a KD tree for large-scale point cloud data provided by an embodiment of the present invention;
[0043] Figure 4 It is a schematic diagram of dividing subtrees for constructing a KD tree for a set of specific three-dimensional data provided by an embodiment of the present invention;
[0044] Figure 5 It is a schematic diagram of the effect of successfully realizing point cloud semantic segmentation visualization using an optimized nearest neighbor point algorithm provided by an embodiment of the present invention;
[0045] Figure 6 It is a schematic diagram of the feature extraction route of a dynamic graph convolutional neural network based on the inserted optimized KD tree algorithm provided by an embodiment of the present invention. Specific embodiments
[0046] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in combination with the accompanying drawings and specific embodiments, details a building point cloud reconstruction method based on an improved KNN DGCNN model proposed according to the present invention.
[0047] The foregoing and other technical contents, features, and effects of the present invention can be clearly presented in the following detailed description in conjunction with the accompanying drawings. Through the description of the specific embodiments, a more in-depth and specific understanding of the technical means and effects adopted by the present invention to achieve the predetermined purpose can be obtained. However, the attached drawings are only for reference and illustration, and are not used to limit the technical solution of the present invention.
[0048] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant is intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the said element.
[0049] Aiming at the problems of large computational resource occupancy and high complexity of the KNN algorithm in the dynamic graph convolutional neural network (DGCNN) for building point cloud three-dimensional reconstruction, the present invention proposes a building point cloud reconstruction method based on an improved KNN DGCNN model. Please refer to Figure 1 , and this method includes the following steps:
[0050] S1: Normalize the original building point cloud to be reconstructed to obtain the normalized point cloud data.
[0051] In this step, feature extraction and normalization are performed on the original building point cloud to be reconstructed, and the feature vector of each point in the point cloud is used as a data point. For point cloud data with n points and k-dimensional features, calculate its statistical features such as mean and variance, and normalize the features to have a similar scale.
[0052] Specifically, in the 3D reconstruction of buildings, the present invention classifies the point cloud data collected by lidar. In this embodiment, each point in the point cloud has 3D spatial coordinates (x, y, z) and 3D color information (R, G, B), forming a total of 6D feature vectors to describe each point. To make the point cloud data have better consistency and stability in the subsequent process, each feature dimension of the point is normalized. For the spatial coordinates (x, y, z), assuming that the x-coordinate range of all points in the current point cloud scene is [x min , x max , then through the formula normalize the x-coordinate of each point to the interval [0, 1], where x represents the original x-coordinate of the current point, and x normalized represents the normalized x-coordinate of the current point. Similarly, the y and z coordinates of all points are processed in the same way; for the color values, directly divide them by 255 (assuming the color value range is 0-255) to normalize them to the interval [0, 1].
[0053] S2: Construct a DGCNN network based on the improved KNN algorithm, train the DGCNN network, and obtain a trained DGCNN model.
[0054] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a DGCNN network based on the improved KNN algorithm for processing point cloud data provided by an embodiment of the present invention. The DGCNN network includes a sequentially connected spatial transformation layer, four graph convolutional layers, a max pooling layer, a first multi-layer perceptron, and a second multi-layer perceptron. Among them, the spatial transformation layer is used to perform spatial transformation on the input normalized point cloud data to make the point cloud data invariant in space; the graph convolutional layer is used to obtain k nearest neighbor points of each point in the building point cloud using the improved KNN algorithm, and perform feature aggregation using the k nearest neighbor points to output the aggregated feature vector; the max pooling layer is used to receive the feature vector output by the last graph convolutional layer and perform max pooling operation in the dimension of points to obtain the global feature vector; the first multi-layer perceptron is used to abstract and transform the global feature vector to extract higher-level semantic information; the second multi-layer perceptron includes multiple hidden layers and is used to output the predicted category and the corresponding predicted score.
[0055] Specifically, the input of the DGCNN network is the point cloud data normalized in step S1. The normalized point cloud data is input into the spatial transformation layer, which performs spatial transformation on the input point cloud data, possibly including operations such as rotation, translation, and scaling. The purpose is to make the point cloud data have certain invariance in space so that the subsequent network can better learn features without being affected by the initial position and pose of the data. It should be noted that Figure 2 the n×3 in
[0056] subsequently, the data output by the spatial transformation layer is input into the first graph convolutional layer ( Figure 2 the first EdgeConvmlp 64 in Figure 2 ). The first graph convolutional layer receives the point cloud data after spatial transformation, performs edge convolution operations, extracts and updates features, and strengthens the understanding of local structures, making the feature dimension of the output data become 64; subsequently, it is input into the second graph convolutional layer ( Figure 2 the second EdgeConv mlp 64 in Figure 2 ), where edge convolution operations are performed again to further extract and update features and strengthen the understanding of local structures. The output is still the data feature of n×64. Subsequently, it is input into the third graph convolutional layer (
[0057] the third EdgeConv mlp 64 in Figure 2 ), and convolution continues based on the n×64 data feature of the previous layer to further explore the details of the local features of the point cloud, and the output dimension remains n×64. Subsequently, in the fourth graph convolutional layer ( Figure 2 the EdgeConv mlp 128 in ), with the input being the n×64 data feature of the previous layer, through the edge convolution of this layer, the feature dimension of the output data is increased to 128, which can capture more complex local geometric patterns and relationships and provide richer information for subsequent global feature aggregation.
[0057] Subsequently, the max pooling layer (Max Pooling) receives the n×128 feature vector output by the last graph convolutional layer and performs max pooling operations on the point dimension, that is, for each feature dimension, the maximum value of this dimension among all points is selected to obtain a 1×128 global feature vector. This global feature vector synthesizes the information of the entire point cloud and can reflect the overall geometric characteristics and patterns of the point cloud. Then, the 1×128 global feature vector is passed through a multi-layer perceptron (the first multi-layer perceptron, corresponding to Figure 2 the MLP1024 in ), and the output dimension becomes 1×1024, further abstracting and transforming the global features to extract higher-level semantic information for better classification.
[0058] Finally, it passes through a multi-layer perceptron (the second multi-layer perceptron, corresponding to the Figure 2 MLP in
[0059] ), which contains multiple hidden layers with dimensions of 512 and 256 in sequence, and finally outputs a dimension of c, where c represents the number of categories. This step maps the previously extracted features to the specific category space and can obtain the prediction scores for each category.
[0060] Please refer to Figure 3 , when constructing the KD tree based on the data features input to the current graph convolutional layer, use the KD tree building algorithm to build the KD tree for the normalized data features of the building point cloud data. For the k-dimensional data features, create index arrays for each dimension respectively to form a super key (a 1 , a 2 , a 2 , …, a k ), where a 1 , a 2 , a 2 , …, a k respectively represent the data values of the 1st to kth dimensions of the current point. For example, for 6-dimensional data features (x, y, z, R, G, B), during the construction process, create index arrays for each dimension respectively. During the merge sort process, use the super key (such as x:y:z:R:G:B) to determine the order of the points, so as to ensure that duplicate point feature vectors can be correctly processed.
[0061] Subsequently, compare and sort the points in sequence according to the order of the super keys of each point to obtain the index array of all sorted points. Specifically, assume that there are two points P 1 (x 1 , y 1 , z 1 , R 1 , G 1 , B 1 ) and P 2 (x 2 , y 2 , z 2 , R 2 , G2 , B 2 ), when comparing their superkeys, first compare x 1 and x 2 . If x 1 < x 2 , then rank P 2 (x 2 , y 2 , z 2 , R 2 , G 2 , B 2 ) behind P 1 (x 1 , y 1 , z 1 , R 1 , G 1 , B 1 ). If x 1 and x 2 are equal, then continue to compare y 1 and y 2 , and so on until the order is determined.
[0062] Next, according to the sorted index array, partition all data features based on the median element to obtain the left subtree and the right subtree, and recursively construct the KD-tree structure. Specifically, as Figure 4 shown, exemplarily, for an index array of one dimension, after determining the value of the median element, classify the elements smaller than this value into the left subtree, and classify the elements larger than this value into the right subtree, and recursively construct the tree structure.
[0063] Subsequently, in the current graph convolutional layer, use the constructed KD-tree to search for the K nearest neighbors of each point in the point cloud and use the information of the K nearest neighbors for feature aggregation to obtain the updated data of the current point.
[0064] It should be noted that for the traditional KNN algorithm, when searching for the K nearest neighbors (i.e., the K nearest neighbor points), it is necessary to calculate the Euclidean distance between the selected point and each point in the current point cloud data. The Euclidean distance formula is:
[0065]
[0066] where d(p, q) represents the distance between point p and point q in the point cloud, p j and q j respectively represent the j-th dimensional feature values of point p and point q, and n represents the number of dimensions. After obtaining the Euclidean distance between the selected point and each point in the current point cloud data, compare the Euclidean distances and select the K points with the smallest Euclidean distance as the target points, that is, the K nearest neighbor points. This type of method has a high computational complexity and low efficiency in point cloud data with a large information density.
[0067] In the present invention, KNN search is optimized based on the constructed KD-tree. For a selected point, starting from the root node of the KD-tree, according to the comparison between the feature value of the selected point in the current dimension and the median of the points, the subtree that may contain the nearest neighbor points can be quickly located. For example, if the feature value of the k-th dimension of the selected point is less than the median of the current points in this dimension, then search the left subtree; otherwise, search the right subtree. During the search process, a priority queue is used to store the k nearest neighbor points found currently and their distances. The priority queue is sorted in ascending order of distance. If the distance of the newly searched point is less than the distance of the farthest point in the priority queue, then the farthest point is popped out and the new point is inserted, as Figure 6 shown. This can avoid calculating the distances between the selected point and all points in the dataset, greatly reducing the amount of calculation.
[0068] In the actual application process, during the calculation of each graph convolutional layer of the DGCNN network, when it is necessary to find the K nearest neighbors of the points in the building point cloud, the constructed KD-tree is used for searching. Starting from the root node, according to the comparison between the value of the point feature of the building point cloud in the current dimension and the median (the feature value of the root point in this dimension), the subtree that may contain the nearest neighbor points can be quickly located. For example, if a certain feature value of the current point (such as the height dimension) is less than the feature value of the root node in this dimension, then search the subtree area less than this value (left subtree); otherwise, search the subtree area greater than this value (right subtree). During the search process, this comparison and subtree selection process is continuously repeated until reaching the leaf point or finding a sufficient number (for example, the set number of nearest neighbor points is 20, that is, K = 20) of nearest neighbor points.
[0069] Subsequently, after obtaining the K nearest neighbor points of the current point of the building point cloud using the improved KNN algorithm, the data features of these K nearest neighbor points are aggregated. For example, for each point of the building point cloud, after finding its K nearest neighbor points through the KD-tree, according to the features of the nearest neighbor points and the predefined weight rules (such as weighted average considering the building structure features), the features of the K nearest neighbor points are aggregated to obtain a new feature representation of the current point. Then the aggregated features are processed through an activation function to improve the recognition accuracy of the building point cloud data.
[0070] Specifically, in the graph convolutional layer, the improved KNN algorithm is used to obtain the neighbor information of the points and perform feature aggregation. Let the input feature matrix of the l-th graph convolutional layer be H () (as Figure 4 shown, including a total of 4 layers), with a size of n × d l , where n is the number of points of the point cloud to be reconstructed input to the network, and d lis the feature dimension of the l-th layer graph convolutional layer. Its i-th row represents the feature vector of the i-th point. Then its feature aggregation formula is:
[0071]
[0072] where N(i) represents the set of K nearest neighbor points found by point i through the KD tree, and W () represents the learnable weight matrix of the l-th layer, σ is the activation function, represents the feature vector of point j in the l-th layer graph convolutional layer, N(j) represents the set of K nearest neighbor points found by point j through the KD tree, and H (+1) represents the aggregated feature vector.
[0073] Furthermore, the loss function of this DGCNN network selects the cross-entropy loss, and the formula is:
[0074]
[0075] where M is the number of training samples used in the training process, Q is the number of class labels in the training samples, and y mq represents the probability that the m-th training sample belongs to the q-th category, represents the probability that the predicted value of the m-th training sample belongs to the q-th category.
[0076] When the Dynamic Graph Convolutional Neural Network (DGCNN) performs point cloud classification, it first performs data preprocessing. It reads the original point cloud data, represents each point with xyz coordinates and possible additional features, and then normalizes the coordinate values to eliminate scale and position differences. Subsequently, it constructs a dynamic graph structure, uses the K-nearest neighbor parameter to search for the neighboring points of each point, and builds an initial graph based on spatial proximity; as the network propagates forward, it dynamically updates the neighborhood relationship according to the point features. In the feature extraction and convolution operation stage, it focuses on extracting edge features from the relative relationship of adjacent points, and then sends them into the dynamic graph convolutional layer, allowing the convolutional kernel to slide on the irregular graph structure to aggregate neighborhood information and achieve hierarchical feature extraction. After that, it performs global feature aggregation through a pooling operation, adopts the max pooling strategy, summarizes the local maximum features to obtain a global feature vector reflecting the overall geometric characteristics of the point cloud, and finally outputs the corresponding classification category.
[0077] Furthermore, after constructing the DGCNN network, before actual use, the DGCNN network also needs to be trained. During the training process, first, a large number of building point clouds are collected to form a dataset, and each building point cloud is labeled with a target type label.
[0078] The present invention divides 80% of the dataset into the training set, another 10% into the validation set, and 10% into the test set part. During the training process, the stochastic gradient descent algorithm is used, and the initial learning rate is set to η0 = 0.01, the coefficient parameter μ = 0.9, and train for 100 epochs. During each epoch training, for each point cloud data sample in the training dataset, first perform normalization processing, input the normalized point cloud data sample into the DGCNN network to be trained, calculate the prediction result through forward propagation, and calculate the loss value according to the loss function. Then, calculate the gradient according to the loss value through the backpropagation algorithm and update the parameters of the DGCNN network. Every 5 batches of training, evaluate the accuracy of the DGCNN network on the validation dataset. If the accuracy does not improve for 3 consecutive batches, then reduce the learning rate according to the formula η new = old × 0.1 for subsequent training, where η old represents the current learning rate, and η new represents the updated learning rate.
[0079] After the training is completed, on the corresponding test set, according to the points in the input point cloud scene and the features of their nearest neighbors found through the KD tree, output the class prediction of the point cloud scene. For each point cloud scene in the test set, pass the point cloud data to be tested through the inference of the trained model, and select the class with the highest score as the final result. It is also possible to color different objects belonging to different objects in a point cloud data to achieve the visualization of point cloud semantic segmentation, as shown in Figure 5 shown.
[0080] Furthermore, indicators such as accuracy, recall, and F1 value can be used to evaluate the performance of the model. The formula for calculating accuracy is:
[0081]
[0082] where TP (True Positive) represents the number of true positive examples, TN (True Negative) represents the number of true negative examples, FP (False Positive) represents the number of false positive examples, and FN (False Negative) represents the number of false negative examples.
[0083] True positive example, that is, the number of samples that are actually positive examples and are predicted as positive examples by the model. For example, in a classification task of building point cloud data, if it is actually a certain type of building (such as an ancient building) and the model correctly predicts it as an ancient building, then the number of such samples is a true positive example. True negative example, that is, the number of samples that are actually negative examples and are predicted as negative examples by the model. False positive example, that is, the number of samples that are actually negative examples but are predicted as positive examples by the model. False negative example, that is, the number of samples that are actually positive examples but are predicted as negative examples by the model.
[0084] Through comparative experiments, compared with the original KNN-DGCNN structure, the recognition accuracy of this method has been improved by 4%, while the recognition precision of the target point cloud data has been increased by 11%, and the recall rate has been increased by 12%, significantly improving the recognition effect of the corresponding building data.
[0085] S3: Input the normalized point cloud data into the trained DGCNN model to obtain the corresponding prediction results.
[0086] Apply the trained DGCNN model to the prediction task of the point cloud data that actually needs to be processed, such as point classification. For the point classification task, the DGCNN model outputs the class prediction of the point according to the input point and the features of the neighboring points determined by the improved KNN algorithm. During the application process, preprocess the new graph structure data, construct a KD tree, and then perform prediction through the trained DGCNN model to obtain the final prediction result.
[0087] To illustrate the performance evaluation process of the 3D point cloud reconstruction method based on the improved KNN DGCNN model of the present invention, on the basis of the above embodiment, this embodiment further illustrates the performance evaluation of the improved nearest neighbor point sampling algorithm and the comparison process with the original algorithm.
[0088] During the test, randomly select 1000 test points, and use the traditional KNN algorithm and the optimized KNN algorithm based on the KD tree of the present invention to perform K-nearest neighbor search respectively, and calculate the average search time of the two methods. Under the same hardware environment, when the traditional KNN algorithm processes these 1000 test points, the average search time is T for each test point KNN = 0.5 seconds. It can be seen that for the traditional nearest neighbor point sampling method, for a dataset of 20,000 sample points, the computational complexity is huge. While the average search time of the optimized KNN algorithm based on the KD tree of the present invention is reduced to T for each test point KD-tree = 0.05 seconds. This is because the KD tree can quickly locate the area where the possible nearest neighbor points are located through the division of the tree structure, reducing unnecessary distance calculations.
[0089] Compare the computational acceleration ratio of the original KNN algorithm and the improved KD tree algorithm. The formula is This shows that the DGCNN model based on the improved KNN is ten times faster than the traditional KNN algorithm. At the same time, evaluate the accuracy of the nearest neighbor points searched by the two methods, and measure it by calculating the average relative error. Let be the distance from the test point i to its true nearest neighbor point, while is the distance from the nearest neighbor point found by the algorithm. The formula for the average relative error is
[0090] To calculate the average relative error more precisely, for each test point, first find its true nearest neighbor point through a brute-force search method and calculate the true distance Then use the traditional KNN algorithm and the KD-tree-based optimized algorithm respectively to find the nearest neighbor points and calculate the corresponding distances After measurement, the ARE of using the original KNN KNN = 0.02, while after replacing it with the KD tree, the ARE of the KD-tree-based optimized algorithm KD-tree = 0.025. The errors of the two are similar, indicating that while significantly improving the search speed, the optimized algorithm can maintain an accuracy similar to that of the traditional KNN algorithm
[0091] The building point cloud reconstruction method of the DGCNN model based on the improved KNN proposed by the present invention has achieved a qualitative leap in the operation speed of the DGCNN model in terms of computational efficiency. Through an efficient space partitioning strategy, the KD-tree structure significantly reduces the computational amount during KNN search and greatly accelerates the processing process of building point cloud data. Compared with traditional methods, when processing building point cloud data of the same scale, the present invention can significantly shorten the processing time, improve the efficiency of the reconstruction work, effectively reduce the project cycle, and provide more timely data support for building design, planning, protection, etc
[0092] In terms of reconstruction accuracy, the method of the present invention demonstrates excellent performance improvement in the three-dimensional reconstruction task of building point clouds. By more precisely constructing the neighborhood relationship of point cloud data, it can effectively capture the fine structure and complex features of buildings and reduce the reconstruction deviation caused by the nearest neighbor point search error. Whether it is the delicate decorations such as carvings and cornices of ancient buildings or the unique external shapes and internal structures of modern high-rise buildings, they can all be restored more accurately, thus providing highly reliable three-dimensional models for building research, historical building protection, digital display, etc
[0093] In actual application scenarios, the present invention performs excellently in the data registration of buildings, can more accurately match and fuse building point cloud data collected from different perspectives or at different times, and ensure the integrity and consistency of the reconstruction results. During the compression and reconstruction process, the optimized algorithm can effectively reduce the data storage amount, improve the efficiency of data transmission and storage, and reduce the requirements for hardware devices while ensuring the reconstruction quality, providing a more economical, efficient, and reliable solution for the application of building point cloud data, and having broad application prospects in many fields such as urban planning, digital protection of architectural heritage, and real estate development
[0094] In several embodiments provided by the present invention, it should be understood that the devices and methods disclosed in the present invention can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0095] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.
[0096] Another embodiment of the present invention provides a storage medium in which a computer program is stored, and the computer program is used to execute the steps of the three-dimensional point cloud reconstruction method of the improved KNN-based DGCNN model described in the above embodiments. Another aspect of the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the three-dimensional point cloud reconstruction method of the improved KNN-based DGCNN model described in the above embodiments are implemented. Specifically, the above integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium and include several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0097] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A building point cloud reconstruction method based on the DGCNN model of improved KNN, characterized in that: include: S1: normalizing the original building point cloud to be reconstructed to obtain normalized point cloud data; S2: Construct a DGCNN network based on an improved KNN algorithm, and train the DGCNN network to obtain a trained DGCNN model. The DGCNN network based on the improved KNN algorithm includes a spatial transformation layer, four graph convolution layers, a maximum pooling layer, a first multi-layer perceptron, and a second multi-layer perceptron connected in sequence, wherein the spatial transformation layer is used to perform spatial transformation on the input normalized point cloud data so that the point cloud data is spatially invariant; the graph convolution layer is used to obtain the k nearest neighbor points of each point in the building point cloud using the improved KNN algorithm, perform feature aggregation using the k nearest neighbor points, and output the aggregated feature vector; the maximum pooling layer is used to receive the feature vector output by the last graph convolution layer, perform a maximum pooling operation on the point dimension, and obtain a global feature vector; the first multi-layer perceptron is used to abstract and transform the global feature vector to extract higher-level semantic information; the second multi-layer perceptron includes multiple hidden layers, which are used to output prediction categories and corresponding prediction scores; S3: Input the normalized point cloud data into the trained DGCNN model to obtain the corresponding prediction result.
2. The building point cloud reconstruction method based on the DGCNN model of improved KNN according to claim 1, characterized in that: The S1 includes: The 6-dimensional feature vector of each point in the original building point cloud is normalized respectively, and the 6-dimensional feature vector includes 3-dimensional spatial coordinates (x, y, z) and 3-dimensional color information (R, G, B).
3. The building point cloud reconstruction method based on the DGCNN model of improved KNN according to claim 1, characterized in that: The graph convolution layer is specifically used for: Construct a KD tree based on the data features of the current graph convolutional layer input; Use the constructed KD tree to search for the K nearest neighbor points of each point in the point cloud; Using a predefined weight rule to perform feature aggregation on the information of the K nearest neighbor points, to obtain updated data features of the current point; The aggregated data features are processed through activation functions to improve the recognition accuracy of building point cloud data.
4. The building point cloud reconstruction method based on the DGCNN model of improved KNN according to claim 3 is characterized in that: Construct a KD tree based on the data features of the current graph convolution layer input, including: For k-dimensional data features, create index arrays for each dimension to form super keys (a1, a2, a2, ..., a k ), where a1,a2,a2,…,a k Respectively represent the data values of the 1st to kth dimensions of the current point; Compare and sort the points in order according to the order of the super key of each point, and obtain the index array of all the points after sorting; Based on the sorted index array of all points, all data features are partitioned according to the median element to obtain the left subtree and the right subtree, and then the KD tree structure is constructed progressively.
5. The building point cloud reconstruction method based on the DGCNN model of improved KNN according to claim 3, characterized in that: The information of the K nearest neighbor points is aggregated using a predefined weight rule to obtain the updated data features of the current point, including: In the graph convolution layer, the improved KNN algorithm is used to obtain the K nearest neighbor points of each point in the input feature vector for feature aggregation. The aggregation formula is: Among them, H (l) Represents the input feature vector of the lth graph convolutional layer, with a size of n×d l , n is the number of points, d l is the input feature vector H (l) The feature dimension of N(i) is the set of K nearest neighbor points found by the KD tree for point i, and W (l) is the learnable weight matrix of the lth graph convolutional layer, σ is the activation function, represents the feature vector of point j in the lth graph convolutional layer, N(j) represents the set of K nearest neighbor points found by point j through the KD tree, and H (l+1) Represents the aggregated feature vector.
6. The building point cloud reconstruction method based on the DGCNN model of improved KNN according to claim 1, characterized in that: The loss function of the DGCNN model is the cross entropy loss, expressed as: Among them, M is the number of training samples used in the training process, Q is the number of category labels in the training samples, and y mq represents the probability that the mth training sample belongs to the qth category, It represents the probability that the predicted value of the mth training sample belongs to the qth category.
7. The building point cloud reconstruction method based on the DGCNN model of improved KNN according to claim 1, characterized in that: The DGCNN network is trained to obtain a trained DGCNN model, including: Set the initial learning rate, coefficient parameters and training batches, collect building point clouds to form a training data set, and each building point cloud is annotated with a target type label; During each batch training, each point cloud data sample in the training data set is first normalized, and the normalized point cloud data sample is input into the DGCNN network to be trained, and the prediction result is calculated by forward propagation, and the loss value is calculated according to the loss function; Calculate the gradient through the back propagation algorithm according to the loss value, and update the parameters of the DGCNN network; After every 5 training batches, the accuracy of the DGCNN network is evaluated on the validation dataset. If the accuracy does not improve for 3 consecutive batches, the network is trained according to the formula η new =η old ×0.1 to reduce the learning rate for subsequent training, where η old represents the current learning rate, η new Represents the updated learning rate.
8. A storage medium storing a computer program, characterized in that: The computer program is used to execute the steps of the building point cloud reconstruction method based on the improved KNN DGCNN model as described in any one of claims 1 to 7.
9. An electronic device, comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the building point cloud reconstruction method based on the DGCNN model of the improved KNN as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Photogrammetry point cloud semantic segmentation method based on deep learning
CN113449736A
Three-dimensional point cloud classification method based on sparse graph convolution
CN114373099A
Three-dimensional point cloud target detection algorithm based on graph neural network
CN114998890A
Deep learning-based high-precision point cloud completion method and apparatus
WO2024060395A1
Method, apparatus, and medium for point cloud coding
WO2024175011A1