A method for semantic segmentation of vegetation point clouds in urban areas based on the HPCT model

Through the HPCT model combined with Hierarchical Structure and Transformer Block, the accuracy of semantic segmentation of vegetation point clouds in urban areas is solved, accurate vegetation point cloud segmentation and heterologous point cloud data adaptation are achieved, and the degree of automation in the field of forestry remote sensing is improved.

CN116630622BActive Publication Date: 2025-07-11UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310577928.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-22
Publication Date
2025-07-11
Estimated Expiration
2043-05-22

AI Technical Summary

Technical Problem

The existing semantic segmentation method of vegetation point clouds has poor accuracy when migrating to point clouds. Especially in the complex landform distribution in urban areas, it lacks effective point cloud organization methods and cannot effectively extract tree point clouds.

Method used

The semantic segmentation method of urban area vegetation point clouds based on HPCT model is adopted. By unifying the spatial scale processing of heterologous point clouds and combining Hierarchical Structure and Transformer Block, a deep learning model suitable for semantic segmentation of vegetation point clouds is constructed. The self-attention mechanism is used to capture the relative relationship between different spatial scales and land objects, and the expression ability of the model is improved.

Benefits of technology

It realizes the precise segmentation of vegetation point clouds in urban areas, improves the degree of automation, is suitable for heterologous point cloud data, and the segmentation accuracy reaches the level of manual point-by-point segmentation, which can assist in the statistics of vegetation information such as three-dimensional green quantity and urban greening rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116630622B_ABST
    Figure CN116630622B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of forestry remote sensing technology, and specifically relates to a method for semantic segmentation of vegetation point clouds in urban areas based on a deep learning model. The present invention combines deep learning technology to construct the HPCT model: the model is constructed into multiple layers, each layer can process different levels of features of three-dimensional point clouds, and then capture semantic features related to vegetation at different spatial scales and the relative relationships of ground objects; the self-attention mechanism is adopted to perform different weightings according to the importance of each part of the input data, and capture the long-distance dependence relationships between different positions in the three-dimensional point clouds, thereby significantly improving the expression ability and understanding ability of the HPCT model; at the same time, the present invention collects the point cloud data of urban areas from three data sources and performs semantic annotation on them to form a vegetation point cloud data set, which is used for model training and prediction, providing a data basis for HPCT model training and also alleviating the current situation of scarce point cloud data sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of forestry remote sensing, and particularly relates to a method for semantic segmentation of urban area vegetation point clouds based on a deep learning model. Background Technique

[0002] The purpose of semantic segmentation of vegetation point clouds is to extract tree point clouds from the complex ground object distribution in urban areas. In forestry remote sensing, semantic segmentation of vegetation point clouds is an important pre-task and a prerequisite for subsequent processing. With the rapid development of deep learning technology, deep neural networks have become the mainstream technology for image processing tasks due to their powerful high-dimensional feature extraction capabilities. However, point cloud data has the characteristics of being spatially scattered and disordered, making it impossible to directly transfer models based on two-dimensional images to point clouds. Therefore, it is necessary to study the organization method of point clouds. There are three commonly used point cloud organization methods for point cloud processing tasks, namely multi-view projection, three-dimensional reconstruction (Voxel), and direct processing of point clouds (Point-set).

[0003] Multi-view projection is to reduce the dimension of point clouds and represent them as multiple two-dimensional depth images, and then two-dimensional convolutional neural networks can be used for processing. For example, SnapNet reduces the number of point clouds through preprocessing, calculates features, and generates grids. Then, multi-view images of the grids are generated through a virtual camera, and finally, image semantic segmentation technology is used to complete point cloud extraction and project it back to three-dimensional space. There are two obvious defects when multi-view projection is applied to point cloud extraction. One is that it will cause geometric structure loss because multi-view projection is only an approximation of point clouds. The other is that for complex scenes, it is difficult for multi-view projection images to contain all point clouds.

[0004] Three-dimensional reconstruction is to re-establish a regular three-dimensional data structure for point clouds and then use a three-dimensional fully convolutional neural network for processing. For example, SegCloud reconstructs point clouds into a regular three-dimensional pixel matrix in the preprocessing stage, and then completes the point cloud extraction task through a three-dimensional fully convolutional neural network, interpolation, and fully connected conditional random fields. However, due to its redundant three-dimensional structure, it causes a huge computational and memory burden. To solve this problem, OctNet and O-CNN use a more structured octree to complete the reconstruction, greatly reducing redundancy. VV-Net reconstructs point clouds into a structure with richer information than a three-dimensional pixel matrix based on an autoencoder structure. MinkowskiNets provides a four-dimensional convolutional network to extract spatio-temporal features for processing three-dimensional point cloud videos, and it can also be used in point cloud extraction tasks. There are still two relatively serious defects in the methods based on three-dimensional reconstruction. One is the loss of spatial information during the reconstruction process, and the other is the redundancy of the space after reconstruction.

[0005] Directly processing point clouds can make full use of the semantic information of point clouds, which is a very popular and potential research direction. PointNet is the most influential framework in this direction and is designed based on three aspects. First, to solve the problem of the disorder of point clouds, a symmetric structure is made for the whole network based on the max-pooling layer and multi-layer perceptron (the input order of point clouds does not affect the result); second, for feature extraction, a combination mechanism of global features and individual features of point clouds is set for the point cloud extraction task; third, to eliminate the influence of geometric transformations (such as translation and rotation, etc.) on the result, a small network T-net is designed to perform affine coordinate transformation, and the affine transformation matrix is made to approximate an orthogonal matrix by setting a loss function. However, PointNet also has an obvious defect that it does not consider regional semantic features when extracting features. To solve this problem, PointNet++ adopts a hierarchical network structure to extract regional semantic features on its basis. There is a limitation in this type of framework because it completes feature extraction at a fixed spatial scale (sampling rate) and cannot effectively extract semantic information at different levels. Summary of the Invention

[0006] In view of the above existing problems or deficiencies, to solve the problem of poor accuracy of the point cloud organization method when migrating a two-dimensional image model to point clouds in the existing semantic segmentation of vegetation point clouds, the present invention provides a method for semantic segmentation of urban area vegetation point clouds based on a deep learning model.

[0007] The specific technical solution of the present invention is as follows:

[0008] A method for semantic segmentation of urban area vegetation point clouds based on the HPCT model, comprising the following steps:

[0009] Step 1: For the target urban area, collect three-source point cloud data at the same time phase, namely vehicle-mounted lidar point cloud, point cloud reconstructed from unmanned aerial vehicle (UAV) oblique images, and UAV-mounted lidar point cloud;

[0010] Step 2: Use point cloud processing software (such as CloudCompare) to label the semantic tags of each pixel for the three-source point cloud data collected in Step 1, extract the vegetation point clouds and label them as 1, and label other points as 0, thereby obtaining a point cloud data set for model training and testing;

[0011] Step 3: Since the heterogeneous point clouds have different spatial scales, the present invention proposes a method for processing to unify the spatial scales of the heterogeneous data inputs. The training data set in Step 2 is processed by this method, and the specific process is as follows:

[0012] Step 3.1: Convert all point cloud coordinates to the world coordinate system in meters. The world coordinate system is a geocentric coordinate system with the origin at the Earth's center of mass, and the three axes are defined relative to the Earth's shape and orientation. The world coordinate system belongs to the Cartesian coordinate system, that is, it uses rectangular coordinates to represent points in space;

[0013] Step 3.2: Input the training point cloud dataset obtained in Step 2, and simultaneously specify the region segmentation parameter bs, the number of sampling points ns, and the sampling parameter sr;

[0014] Step 3.3: Normalize the coordinates of the point cloud dataset (coordinate range to 0 - 1), and then randomly select a point (x c , y c , z c ) within the region as the sampling center, and obtain all points within the range of bs / 2 around the x and y coordinates of this point to form a set P:

[0015]

[0016] Step 3.4: Randomly sample ns points from the point set P as the input of the HPCT model. The sampling parameter sr is used to control the number of samples N in each epoch during each iteration round iteration ;

[0017] N iteration = N training / ns × sr (2)

[0018] where N training represents the total number of points in the training set.

[0019] Step 4: The present invention proposes an HPCT model suitable for semantic segmentation of vegetation point clouds; construct the HPCT model by combining the deep learning techniques Hierarchical Structure and Transformer Block, set the model parameters, and input the processed training dataset in Step 3 into the model for training;

[0020] Among them, Hierarchical Structure means building the model into multiple levels, each level can process different levels of features of the 3D point cloud, and then capture the semantic features related to vegetation at different spatial scales and the relative relationships of ground objects; the Transformer Block adopts the self-attention mechanism (Self-Attention), which can perform different weightings according to the importance of each part of the input data, capture the long-range dependencies between different positions in the 3D point cloud, thereby significantly improving the model's expression ability and understanding ability for the characteristics of the 3D point cloud in urban areas such as large spatial scales and complex relationships between ground objects. At the same time, the Transformer has the characteristic of invariant sequence arrangement and is applicable to the 3D point cloud with the characteristics of dispersion and disorder in space.

[0021] Step 5: Input the test data set into the model trained in Step 4, automatically segment the vegetation point cloud from the input training data set, and view the semantic segmentation result of the vegetation point cloud.

[0022] Furthermore, the structure of the HPCT model in Step 4 is specifically as follows (as Figure 2 shown):

[0023] It is composed of three cascaded feature extraction modules with different spatial scales. In the downsampling link of each module, the Linear Embedding and Grid Merging layers are implemented through the combination of Farthest Point Sampling (FPS), Ball Query and linear layers.

[0024] Let the input be FPS will select a subset {x i1 , x i2 , …, x im}, where x ij are the m points with the farthest distance in the linear space composed of all channels. Compared with random sampling, FPS can better represent the original point set under the condition of the same number of sampling points.

[0025] After that, Ball Query first calculates and finds all the points within the specified radius r of x ij in the linear space, and randomly samples K of them for channel superposition to obtain

[0026] Finally, the linear layer downsamples the output through channel transformation After the downsampling transformation, the dimension of the point cloud changes from to

[0027] Point Transformer Blocks at the same spatial scale are connected in a cascaded pattern, which can be expressed by Equation (3):

[0028] F i = AT i (F i-1 ), i = 1, 2,..., M j , (3)

[0029] where AT i represents the i-th Attention layer, and F i-1 represents the output of the (i - 1)-th layer, where the output of the 0-th layer is the output of the LinearEmbedding or Grid Merging layer. Since the point cloud itself carries coordinate information, the Positional Embedding module of the Transformer is deprecated in this design.

[0030] Furthermore, as Figure 3 shown, the Attention layer adopts Offset-Attention to calculate the semantic similarity between different point cloud features for semantic modeling. At the same time, predicting the residual block rather than the feature itself can achieve better training results. Let Query, Key, and Value be Q, K, and V respectively. The principle of Offset-Attention is as shown in Equation (4):

[0031] Attention(Q, K, V) = F in · (W q , W k , W v )(4)

[0032] where is the learnable linear transformation shared by this layer; d e = C j , d a = d e / R, where R is an adjustable hyperparameter. N j and C ) are the number of feature points and the number of dimensions of each spatial scale layer respectively. The input F out to the Attention layer is calculated as shown in Equation (5):

[0033] A = Softmax(Attention · K T T

[0034] F sa = A · V(5)

[0035] F out = LBR(F ib - F sa ) + F in

[0036] A represents the Attention Score, and LBR represents the combination of a linear layer, a BathNorm layer, and a ReLU layer.

[0037] The tunable hyperparameters in the HPCT model are: the number of input points N per layer i , the number of Point Transformer Blocks M per layer i , the number of FPS neighboring points K per layer i (i = 1, 2, 3), the number of input channels C in the first layer, and the adjustable parameter R of Offset-Attention.

[0038] The classification Head is a combination of a max pooling layer and a linear layer. The categories use one-hot encoding, and the loss function uses cross-entropy loss, as shown in Equation (6):

[0039]

[0040] where y i) represents the label, and h θ (x i ) j represents the predicted value.

[0041] The segmentation Head part adopts a skip connection form to transmit the semantic information of the same layer to the decoding part to enhance the segmentation effect. The upsampling link uses a Point Interpolate layer;

[0042] First, the point cloud upsampling is completed by the distance-weighted features of k nearest neighbor points, as shown in Equation (7):

[0043]

[0044] where d(x, x i ) is the Euclidean distance from point x to the i-th nearest neighbor point.

[0045] Then, the decoder of the same layer is combined in the feature dimension to obtain the output. The feature combination uses the same Offset-Attention as the backbone network. The loss function uses cross-entropy loss. It should be noted that HPCT is a general backbone network and can be used in different vision tasks (such as classification, detection, and semantic segmentation, etc.).

[0046] Furthermore, the specific process of step 4 is as follows:

[0047] Step 4.1: Build the code of the above HPCT model and use a GPU graphics card for model training and testing;

[0048] Step 4.2: Set the model training parameters: initial learning rate, weight decay rate, batch size, number of input points per layer, number of Point Transformer Blocks per layer, number of FPS neighboring points per layer, number of input channels in the first layer, adjustable parameters of Offset-Attention; when inputting the training dataset, ignore the point cloud RGB information and only input the coordinates xyz; set the region segmentation parameters and the number of sampling points, and use cross-entropy loss with label smoothing as the loss function.

[0049] Step 4.3: Data augmentation will greatly affect the accuracy of the semantic segmentation task. To improve the segmentation accuracy as much as possible, the present invention uses three data augmentation methods in sequence, including: the data augmentation scheme used by the original PointNet++, namely random rotation, spatial scaling, and displacement; Point-BERT uses a resampling mechanism for scale transformation and randomly samples 1024 points from the original point cloud; RandLA-Net and Point Transformer load the entire scene during the training of the semantic segmentation task.

[0050] Step 4.4: Adopting a more effective optimization method during training also helps to improve the model performance. The improvements in the optimization method of the present invention mainly include using AdamW instead of Adam, using cosine annealing learning rate decay instead of step learning rate decay, and using label smoothing.

[0051] Step 4.5: Input the training dataset for training to obtain a trained model file, which records various parameters of the model.

[0052] Applying the technologies in computer vision to the semantic segmentation of vegetation point clouds mainly has three difficulties: (1) Compared with the point clouds of CAD (Computer-Aided Design) models (such as ModelNet40 and ShapeNet-Part datasets) and the point clouds of enclosed rooms (such as S3DIS dataset) commonly processed in the field of computer vision, the spatial scale of the static point cloud data in urban areas is larger and the relative relationships between ground objects are more complex. Therefore, the model is required to have a stronger ability to understand different-level semantic features at a large spatial scale; (2) There are three ways to obtain urban point cloud data: vehicle-mounted lidar point cloud, drone oblique image reconstruction point cloud, and drone-mounted lidar point cloud. Heterogeneous point clouds have different acquisition devices (such as drones, cars, etc.) and sensors (such as lidar and cameras, etc.), resulting in heterogeneous point clouds having different spatial scales and distribution characteristics; (3) There is a lack of urban area vegetation semantic segmentation datasets with rich scenes.

[0053] Combining deep learning techniques such as Hierarchical Structure and Transformer Block, the present invention proposes an HPCT model (Hierarchical Point Cloud Transformer, HPCT) suitable for semantic segmentation of vegetation point clouds. Among them, Hierarchical Structure refers to constructing the model into multiple levels, each level can process features of different levels of 3D point clouds, and then capture semantic features related to vegetation at different spatial scales and the relative relationships of ground objects; the Transformer Block adopts the self-attention mechanism (Self-Attention), which can perform different weightings according to the importance of each part of the input data, capture the long-range dependencies between different positions in the 3D point cloud, thereby significantly improving the model's expression ability and understanding ability for characteristics such as large spatial scales and complex relationships between ground objects in the 3D point cloud of urban areas. At the same time, the Transformer has the characteristic of sequence permutation invariance and is suitable for 3D point clouds with scattered and disordered characteristics in space.

[0054] Spatial scale is crucial for the model to understand the characteristics of ground objects themselves and the distribution relationships between ground objects in the 3D point cloud of urban areas. In order to achieve training and prediction of heterogeneous point clouds based on a unified model, the present invention proposes a sampling method for unifying the spatial scale of heterogeneous data input. Annotate the self-collected vehicle-mounted lidar point cloud, airborne lidar point cloud, and oblique photogrammetry reconstruction point cloud, and make a dataset for semantic segmentation of vegetation in urban areas to provide a data basis for the training of the HPCT model. Finally, the present invention applies the research results of computer vision to the field of forestry remote sensing, broadens the application scenario of semantic segmentation of vegetation point clouds, and improves the accuracy to the level of manual point-by-point segmentation.

[0055] In summary, based on the existing deep learning technology, the present invention innovatively proposes an HPCT model for solving the semantic segmentation of vegetation point clouds in urban areas, realizes the accurate segmentation of vegetation point clouds in the urban area point cloud, improves the automation degree, and the model can be applied to heterogeneous point cloud data, and the point cloud data obtained in different ways does not affect its training and prediction. The method for semantic segmentation of vegetation point clouds in urban areas based on the HPCT model has many uses, such as it can be used to assist in statistics and calculation of vegetation information such as three-dimensional green volume and urban greening rate, and has an important role in the field of forestry remote sensing. At the same time, due to the lack of a dataset for semantic segmentation of vegetation point clouds in urban areas, the present invention collects urban area point clouds and makes a dataset of vegetation point clouds before training the HPCT model to provide a data basis for other related research. Brief Description of the Drawings

[0056] Figure 1 is the flowchart of the present invention;

[0057] Figure 2 It is a schematic diagram of the HPCT model architecture;

[0058] Figure 3 It is a schematic diagram of the Offset-Attention principle;

[0059] Figure 4 It is a schematic diagram of the point cloud data of the vehicle-mounted lidar in the embodiment;

[0060] Figure 5 It is a schematic diagram of the point cloud data reconstructed from the oblique images of the unmanned aerial vehicle in the embodiment;

[0061] Figure 6 It is a schematic diagram of the point cloud data of the unmanned aerial vehicle-mounted lidar in the embodiment;

[0062] Figure 7 It is a schematic diagram of the visualization result of the three-source point cloud semantic segmentation case in the HPCT model survey areas 1 and 3 of the embodiment.

[0063] Figure 8 It is a schematic diagram of the visualization result of the three-source point cloud semantic segmentation case in the HPCT model survey area 2 of the embodiment. Detailed implementation manners

[0064] To intuitively express the advantages of the present invention, in combination with actual data and experimental result drawings, an implementation case of urban area vegetation point cloud semantic segmentation based on the HPCT model is described as follows Figure 1 as shown, and the specific implementation process is as follows:

[0065] Step 1: Select three typical urban areas and denote them as survey area 1, survey area 2, and survey area 3, and collect three-source point cloud data at the same time phase, namely vehicle-mounted lidar point cloud, point cloud reconstructed from oblique images of the unmanned aerial vehicle, and unmanned aerial vehicle-mounted lidar point cloud.

[0066] Step 1.1: Acquisition of on-vehicle lidar point cloud data. The 128-line iScan-S-Z lidar is used to collect data. The lidar is fixed on the roof of the collection vehicle, and the supporting equipment includes a vehicle speed sensor, a point cloud box, and a static differential base station deployed on the ground. The collection vehicle drives along the road to collect the original point cloud data, satellite positioning data GNSS (Global Navigation Satellite System), odometer data, and inertial navigation data IMU (Inertial Measurement Unit). The original GNSS data is converted using the StaticToRinex64 software. The experimental parameters are input into the Inertial Explorer software, including the geographical coordinates of the control points, the installation data of the on-vehicle equipment, and the POS (Position and Orientation System) sampling interval, etc. The IE solution (Inertial-Exterior Solution) is performed using the converted GNSS data, IMU data, and odometer data to obtain the POS file of the driving path. Then, the POS file and the original point cloud data are input into the mmsconvert software, and the corresponding parameters are adjusted according to the experiment to obtain the final static full-scene point cloud of the entire scene, as Figure 4 shown.

[0067] Step 1.2, UAV oblique image reconstruction point cloud data collection, first use DJI M300 RTK equipped with Zenmuse P1 image sensor to collect image data, and then use the image data to perform 3D reconstruction to obtain point cloud data. When the drone collects oblique images, use the supporting operation software DJ Pro to set the flight operation area and set related flight parameters. The key parameters include speed 3m / s, heading overlap rate 80%, lateral overlap rate 70%, flight altitude 60m, etc. The operation software will automatically generate flight plans and photo plans based on the above parameters. After that, you can start the operation with one click to complete data collection. The final image resolution is 8192×5460, containing information such as camera internal parameters and image point GPS (Global Positioning System) coordinates. In order to improve the quality of 3D reconstruction, additional image control points are arranged during data collection, and the spatial position of the image control points is determined using the RTK (Real-Time Kinematic) base station. The image control points in the area are distributed as evenly as possible in the test area and the calibrated points are selected. After data collection is completed, Context Capture software is used to complete 3D reconstruction to obtain a static point cloud. First, input the drone image, POS information and camera parameters into the software, and manually mark the image control points in the drone image and add their GPS information. After aerial triangulation, a sparse point cloud is obtained, and then a dense reconstruction task is submitted to obtain the final static point cloud. When submitting the task, please note: select regular plane grid blocks in the spatial architecture and adjust the tile size to reduce memory usage. Set the hole filling in the processing settings to fill all holes except the tile boundaries, set the expected product to 3D point cloud in the new reconstruction project, and set the sampling interval to 0.05m in the point cloud sampling in the format. The result is as follows: Figure 5 shown.

[0068] Step 1.3, drone-mounted lidar point cloud data collection, use DJI M300 RTK equipped with Zenmuse L1 to collect data. The data collection process is similar to the oblique image reconstruction point cloud in the previous section. After obtaining the raw data, use DJI Zhitu software to process the lidar raw file collected by the Zenmuse L1 sensor to generate 3D point cloud data in LAS format. The basic process of generating a 3D point cloud includes three steps: importing raw data, setting reconstruction parameters, and reconstructing the point cloud. First, import the dynamic point cloud file generated by the Zenmuse L1 sensor, set the base station center point to the coordinates of the base station center point during data collection, and then set the point cloud density to the midpoint cloud density to take into account time and accuracy. Then use the accuracy check to import the previously collected image control points and check the accuracy of the point cloud. After the inspection is completed, an accuracy report will be automatically generated. Finally, select the WGS 84 coordinate system for the output coordinate system. Select the LAS format for the output format. Start point cloud reconstruction to get static point cloud data, such as Figure 6 shown.

[0069] Step 2: Use CloudCompare software to perform per-pixel semantic label annotation on the three-source point cloud data collected in Step 1. Extract the vegetation point cloud and label it as 1, and label other points as 0, thus obtaining a dataset for model training and testing.

[0070] Step 2.1: Open CloudCompare software to read the collected vehicle-mounted lidar point cloud data.

[0071] Step 2.2: Manually segment all the tree point clouds in the vehicle-mounted lidar point cloud data using the Segment function in the software, add a new attribute and assign it a value of 1, and assign a value of 0 to the other segmented point clouds.

[0072] Step 2.3: Merge the point clouds with labels added after segmentation in Step 2.2 to form a complete dataset. Repeat the above steps for the point cloud data reconstructed from the UAV oblique images and the UAV-mounted lidar point cloud data to complete the production of the point cloud dataset.

[0073] Step 3: Process the dataset in Step 2 to unify the spatial scale of the heterogeneous data input. The specific process is shown in Step 3 of the invention content.

[0074] Step 4: Build an HPCT model, set the model parameters, and input the processed dataset in Step 3 into the model for training.

[0075] Step 4.1: Build the code of the above HPCT model based on PyTorch, and use a 40GB NVIDIA A100 graphics card for model training and testing.

[0076] Step 4.2: Set the model training parameters. The initial learning rate lr = 0.001, the weight decay rate is 10 -4 , the batch size (Batch Size) is 128. Set the number of input points per layer N0 = 2N1 = 4N2 = 1024, the number of Point TransformerBlock per layer M1 = M2 = M3 = 2, the number of FPS neighboring points per layer K1 = K2 = K3 = 32, the number of input channels in the first layer C = 64, the adjustable parameter R of Offset-Attention = 4. Ignore the RGB information of the point cloud and only input the coordinates xyz. The regional segmentation parameter bs takes 10m, the number of sampling points ns takes 4098, the loss function uses cross-entropy loss with label smoothing, and run for 50 epochs.

[0077] Step 4.3: Data augmentation significantly affects the accuracy of the semantic segmentation task. To improve the segmentation accuracy as much as possible, the three data augmentation methods used in this embodiment include the data augmentation scheme used in the original PointNet++, namely random rotation, spatial scaling, and displacement; Point-BERT uses a resampling mechanism for scale transformation, randomly sampling 1024 points from the original point cloud; RandLA-Net and Point Transformer load the entire scene during the training of the semantic segmentation task.

[0078] Step 4.4: Adopting a more effective optimization method during training is also helpful for improving the model performance. The improvements in the optimization method of the present invention mainly include using AdamW instead of Adam, using cosine learning rate decay instead of step learning rate decay, and using label smoothing.

[0079] Step 4.5: Input the training data set for training to obtain a trained model file, which records various parameters of the model.

[0080] Step 5: Input the test data set into the model trained in Step 4.5 to view the semantic segmentation results of the vegetation point cloud.

[0081] After the above steps, the task of semantic segmentation of the vegetation point cloud in the urban area can be completed. Then, train PointNet++ and PCT with the same experimental configuration for horizontal comparison of the HPCT effect. The test results are shown in Table 1. It can be seen that: (1) The average Precision of the HPCT model in the three-source data all exceeds 96%, which are 96.90%, 99.42%, and 97.42% respectively; the average IoU all exceeds 95%, which are 96.21%, 98.37%, and 95.75% respectively, and the segmentation accuracy is very high. (2) The average Precision and average IoU of the HPCT model in the three-source data exceed those of PointNet++ and PCT, and the Precision and IoU in most individual measurement areas exceed those of PointNet++ and PCT.

[0082] Table 1 Semantic segmentation results of three-source point clouds based on HPCT, PCT, and PointNet++

[0083]

[0084] Figure 7 and Figure 8Lists the visualization results of the three-source point cloud semantic segmentation experiment of the HPCT model. From top to bottom are survey areas 1, 2, and 3; inside each survey area, from top to bottom represent vehicle-mounted lidar, unmanned aerial vehicle (UAV)-mounted lidar, and UAV oblique photography reconstruction; from left to right are the original scene, true label, PointNet++ prediction result, PCT prediction result, and HPCT prediction result; inside the prediction results, the gray objects and the black background represent the correctly predicted tree point clouds and non-tree point clouds respectively, and the boxes indicate the positions. It can be clearly seen that compared with the PointNet++ and PCT prediction results, the HPCT prediction result has a more complete prediction of tree point clouds, that is, fewer tree point clouds are missed or other feature point clouds are misidentified as tree point clouds.

[0085] As can be seen from the above embodiments, the HPCT model constructed by the present invention has improved the accuracy of the vegetation point cloud semantic segmentation result in the three-source data to the level of manual point-by-point segmentation, with an average Precision exceeding 96% and an average IoU exceeding 95%. In the field of forestry remote sensing, this method can play an important role in calculating the three-dimensional green volume, statistical urban greening rate, etc.

Claims

1. A method for semantic segmentation of urban area vegetation point cloud based on HPCT model, characterized in that It includes the following steps: Step 1: For the target urban area, collect three-source point cloud data at the same time phase, namely vehicle-mounted lidar point cloud, drone oblique image reconstruction point cloud, and drone-mounted lidar point cloud; Step 2: Use point cloud processing software to label the three-source point cloud data collected in Step 1 with per-pixel semantic labels. Extract the vegetation point cloud and label it as 1, and label other points as 0, thus obtaining a point cloud data set for model training and testing; Step 3: Process the training data set in Step 2 to unify the spatial scale of heterogeneous data input; Step 3.1: Convert all point cloud coordinates to the world coordinate system, in meters, and use rectangular coordinates to represent points in space; Step 3.2: Input the training point cloud data set obtained in Step 2, and at the same time specify the region segmentation parameter bs, the number of sampling points ns, and the sampling parameter sr; Step 3.3: Normalize the coordinates of the point cloud dataset so that the coordinate range is from 0 to 1; then randomly select a point (x c , y c , z c ) within the area as the sampling center, and obtain a set P consisting of all points within the range of bs / 2 around the x and y coordinates of this point: Step 3.4: Randomly sample ns points from the point set P as the input of the HPCT model, and the sampling parameter sr is used to control the number of sampling times N of epochs in each iteration round iteration ; N iteration = N training / ns×sr (2) where N training represents the total number of points in the training set; Step 4: Combine the deep learning technologies Hierarchical Structure and Transformer Block to build the HPCT model, set the model parameters, and input the processed training data set in Step 3 into the HPCT model for training; Hierarchical Structure means constructing the model into multiple levels. Each level can process different levels of features of the three-dimensional point cloud, thereby capturing semantic features related to vegetation at different spatial scales and the relative relationships of ground objects; Transformer Block adopts the self-attention mechanism Self-Attention, performs different weightings according to the importance of each part of the input data, and captures the long-range dependence relationships between different positions in the three-dimensional point cloud; Step 5: Input the test data set into the model trained in Step 4, automatically segment the vegetation point cloud of the input training data set, and view the semantic segmentation result of the vegetation point cloud.

2. The method for semantic segmentation of urban area vegetation point cloud based on the HPCT model as described in claim 1, wherein: The structure of the HPCT model in Step 4 is specifically as follows: It consists of three cascaded feature extraction modules with different spatial scales. In the downsampling link of each module, the LinearEmbedding and Grid Merging layers are implemented through the combination of FPS, Ball Query, and linear layers; Let the input be The FPS will select a subset {x i1 , x i2 , …, x im}, where x ij are the m points with the farthest distances in the linear space composed of all channels; After that, Ball Query first calculates to find all points within the specified radius r in the linear space for x ij and randomly samples K of them for channel superposition to obtain Finally, the linear layer downsamples the output of the channel transformation After the downsampling transformation, the point cloud dimension changes from to The Point Transformer Blocks at the same spatial scale are connected in a cascaded mode, which can be represented by Equation (3): F i = AT i (F i-1 ), 1 = 1, 2, ..., M j , (3) Among them, AT i represents the i-th Attention layer, and F i-1 represents the output of the (i - 1)-th layer, where the output of the 0-th layer is the output of the LinearEmbedding or Grid Merging layer.

3. The method for semantic segmentation of urban area vegetation point cloud based on the HPCT model as described in claim 2, wherein: The Attention layer adopts Offset-Attention to calculate the semantic similarity between different point cloud features to achieve semantic modeling; Let Query, Key, and Value be Q, K, and V respectively. The principle of Offset-Attention is as shown in Equation (4): (6, K, V) = F in ·(W q , W % , W v )(4) Among them, is a learnable linear transformation shared by this layer; C e = C j , C a = C e / R, where R is an adjustable hyperparameter; N j and C j are the number of feature points and the number of dimensions of each spatial scale layer respectively. The input F out of the Attention layer is calculated as shown in Equation (5): A = Softmax(6·K T ) F sa = A·V F out = LBR(F on - F sa ) + F in (5) A represents the Attention Score, and LBR represents the combination of a linear layer, a BathNorm layer, and a ReLU layer; The tunable hyperparameters in the HPCT model are: the number of input points N per layer i , the number of Point Transformer Blocks M per layer i , the number of FPS neighboring points K per layer i , i = 1, 2, 3; the number of channels C input to the first layer, the tunable parameter R of Offset-Attention The classification head is a combination of a maximum pooling layer and a linear layer. The category is encoded using one-hot encoding, and the loss function uses cross-loss entropy, as shown in formula (6): where y ij represents a label, h θ (x i ) j represents the predicted value; The segmentation head part uses skip connection to pass the semantic information of the same layer to the decoding part to enhance the segmentation effect. The upsampling stage uses the Point Interpolate layer. First, the point cloud upsampling is completed through the distance weighted features of the k nearest neighbor points, as shown in formula (7): where d(x, x i ) is the Euclidean distance from point x to its i-th nearest neighbor; Then the decoders in the same layer are combined in the feature dimension to get the output. The feature combination uses the same Offset-Attention as the backbone network; the loss function uses cross loss entropy.

4. The method for semantic segmentation of urban area vegetation point cloud based on the HPCT model according to claim 3, wherein, The specific process of step 4 is as follows: Step 4.1: Build the code of the above HPCT model and use the GPU graphics card for model training and testing; Step 4.2, set the model training parameters: initial learning rate, weight decay rate, batch size, number of input points per layer, number of Point Transformer Blocks per layer, number of FPS neighboring points per layer, number of channels for the first layer input, and adjustable Offset-Attention parameters; ignore the point cloud RGB information when inputting the training data set, and only input the coordinates xyz; set the region segmentation parameters and number of sampling points, and use the cross loss entropy with label smoothing as the loss function; Step 4.3: Use three data augmentation methods in turn: the original PointNet++ uses a data augmentation scheme, i.e., random rotation, spatial scaling, and displacement; Point-BERT uses a resampling mechanism to perform scale transformation, randomly sampling 1024 points from the original point cloud; RandLA-Net and Point Transformer load the entire scene when training the semantic segmentation task; Step 4.4: Use the optimization method Optimization during training, including using AdamW instead of Adam, using cosine learning rate descent instead of step learning rate descent, and using label smoothing; Step 4.5: Input the training data set for training to obtain a trained model file, which records various parameters of the model.

Citation Information

Patent Citations

  • Method for constructing multi-level three-dimensional terrain model through multi-source data fusion

    CN111724477A

  • Multi-view point cloud registration and point cloud fusion method based on multi-scale feature extraction

    CN115423854A