Single tree segmentation-biological parameter estimation method based on urban vehicle-mounted LiDAR point cloud
Through the single tree segmentation and biological parameter estimation method of vehicle-mounted LiDAR point cloud, combined with reflection intensity and clustering algorithms, the accuracy and generalization problems of urban tree segmentation and parameter estimation are solved, and efficient and accurate single tree information extraction and biological parameter estimation are achieved, supporting urban greening management and ecological monitoring.
Patent Information
- Application Number
- CN202510760781.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies for urban tree segmentation and biological parameter estimation suffer from insufficient segmentation accuracy, inadequate feature representation, and poor model generalization, resulting in inefficient and costly urban greening resource monitoring.
A single tree segmentation method based on vehicle-mounted LiDAR point cloud is adopted, combined with reflection intensity, clustering algorithm and main direction index. The trunk extraction algorithm and gravity model are fused through DBSCAN-PCA to achieve single tree segmentation, and random forest regression and pigeon flock optimization algorithm are used to estimate biological parameters.
It improves the accuracy of single tree segmentation and biological parameter estimation, enhances the efficiency and reliability of urban greening management and ecological monitoring, and provides high-precision support for carbon sink assessment.
Smart Images

Figure CN120635455A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of urban greening intelligent monitoring, and particularly relates to a single tree segmentation and biological parameter estimation method based on urban vehicle-mounted LiDAR point cloud. Background Art
[0002] As a vital component of urban ecosystems, urban greening trees play a key role in enhancing carbon sequestration, mitigating the urban heat island effect, and improving the human living environment. Current urban greening resource surveys rely primarily on manual field measurements combined with empirically derived biological parameter estimation methods. These methods suffer from inherent drawbacks such as low efficiency and high costs, making them inadequate for the refined monitoring required for smart city development.
[0003] In recent years, LiDAR technology, with its high-precision 3D point cloud acquisition capabilities, has provided a new technical path for extracting urban tree information. However, practical applications still face the following technical bottlenecks:
[0004] In terms of single tree segmentation, traditional methods are mainly divided into two technical routes: those based on the canopy height model (CHM) and those directly based on point clouds. Although the CHM-based method has the advantage of computational efficiency, it is easily affected by the rasterization process and is prone to crown morphological distortion. In addition, it generally suffers from over-segmentation or under-segmentation problems when dealing with dense crown overlap scenes. For example, the watershed algorithm is overly sensitive to local features, and the seed region growing method is highly dependent on resolution. The segmentation method directly based on point clouds can retain complete three-dimensional structural information, but the existing methods have obvious deficiencies in the point cloud feature extraction level. Traditional geometric feature descriptions are difficult to effectively capture the fine-grained structural features of trees, and deep learning models are limited by practical problems such as uneven point cloud density and scale changes, resulting in insufficient feature learning. In addition, existing algorithms such as voxelization processing and iterative moving feature point algorithms have technical defects such as high parameter sensitivity and high computational complexity, making it difficult to adapt to the precise segmentation needs of high-density, irregular-shaped trees in complex urban environments.
[0005] When it comes to estimating biological parameters, traditional simulation equations rely excessively on manually measured data, while geometric models based on the "crown width-crown height" relationship experience significantly increased errors in complex canopy structures. While the application of machine learning methods has improved parameter estimation accuracy, existing algorithms still suffer from inherent issues such as blind feature selection and poor model interpretability. Specifically, given the diverse nature of urban tree species, existing technical solutions lack effective mechanisms for integrating tree species features, resulting in insufficient model generalization across tree species scenarios. Overall, existing methods still have significant limitations in calculating three-dimensional green volume, inverting leaf area index, and assessing carbon storage.
[0006] It's important to note that the accuracy of individual tree segmentation in LiDAR point cloud processing directly impacts the reliability of subsequent biological parameter estimation. In other words, errors in individual tree segmentation can easily restrict the adaptability of biological parameter models. Current solutions often employ isolated approaches, failing to establish a collaborative optimization mechanism for "segmentation-estimation."
[0007] Therefore, how to achieve more accurate single tree segmentation and estimate biological parameters based on the results of single tree segmentation, break through the technical bottlenecks of existing methods in segmentation accuracy, feature representation and model generalization, and provide high-precision technical support for urban ecological monitoring has become a research focus. Summary of the Invention
[0008] In response to the problems presented by the prior art, the present invention aims to provide a method for individual tree segmentation and biological parameter estimation based on vehicle-mounted LiDAR point clouds. By innovating a single tree segmentation algorithm, optimizing a deep feature extraction network, and establishing a biological parameter estimation model that fuses multi-source information, this method constructs an efficient and automated system for individual tree information extraction and biological parameter estimation. This system aims to improve the accuracy of individual tree extraction, the reliability of tree species identification, and the accuracy of biological parameter estimation, thereby providing technical support for urban greening management, ecological monitoring, and carbon sink assessment.
[0009] To achieve the above object, the technical solution of the present invention is as follows:
[0010] The single tree segmentation method based on vehicle-mounted LiDAR point cloud includes the following steps:
[0011] Step 1. Use the vehicle-mounted LiDAR to scan and obtain the original scene point cloud data, and pre-process the original point cloud data to obtain the tree point cloud;
[0012] Step 2. Process the trunk point cloud based on the reflection intensity-guided DBSCAN-PCA fusion trunk extraction algorithm to extract the trunk point cloud. The specific process is as follows:
[0013] Step 2.1. Filter crown points based on reflection intensity. Set a reflection intensity threshold and retain points with reflection intensities above the threshold (mainly tree trunks and a small number of low-level crown points). Filter low-intensity points (mainly crowns and edge trunk points).
[0014] Step 2.2. Use the density-based spatial clustering algorithm (DBSCAN) to extract tree trunk points and obtain candidate tree trunk clusters;
[0015] Step 2.3. Use the main direction index to extract the complete tree trunk:
[0016] Step 2.3.1. Use the principal component analysis (PCA) method to calculate the principal axis direction of the trunk point cloud in the trunk candidate cluster to accurately obtain the trunk growth direction of each tree;
[0017] Step 2.3.1.1. Point cloud data preparation:
[0018] Assume that the point cloud data contains N points, and the information of each point is a three-dimensional coordinate, and each point is represented by X i =[x i ,y i ,z i ] T , where i = 1, 2, ..., N;
[0019] The coordinates of all points in a trunk candidate cluster are formed into a matrix X, as shown in Formula 1:
[0020]
[0021] Step 2.3.1.2. Center the point cloud data X, that is, subtract the mean value from each coordinate. The purpose of this step is to move the center of mass of the point cloud to the origin so that PCA can more accurately describe the directional characteristics of the point cloud: Assume that the center of mass of the point cloud is Calculate the average value of each coordinate component, as shown in formula (2):
[0022]
[0023] After centralizing the point cloud data, we get a new data matrix X c As shown in formula (3):
[0024]
[0025] Step 2.3.1.3. Calculate the covariance matrix C. The covariance matrix describes the linear correlation between the coordinate dimensions in the point cloud:
[0026] The calculation formula is shown in formula (4):
[0027]
[0028] Where C is a three-dimensional symmetric matrix whose elements represent the covariance relationship between different coordinate dimensions. The expanded form of the C matrix is shown in formula (5):
[0029]
[0030] Where Var(x), Var(y), and Var(z) are the variances of the x, y, and z coordinates after centering, respectively, indicating the degree of change in the corresponding dimensions; Cov(x,y) is the covariance between the x and y coordinates after centering, indicating their linear correlation; Cov(x,z) is the covariance between the x and z coordinates after centering, and Cov(y,z) is the covariance between the y and z coordinates after centering.
[0031] Step 2.3.1.4. Perform eigenvalue decomposition on the covariance matrix C to obtain eigenvalues and eigenvectors. The eigenvalue decomposition formula is shown in Equation 6:
[0032] Cv i =λ i v i (6)
[0033] Where λ i is the eigenvalue of the covariance matrix C, which represents the variance of the point cloud data in the direction of the corresponding eigenvector; v i is the corresponding eigenvector, indicating the main direction of the point cloud data;
[0034] The result of eigenvalue decomposition is three eigenvalues λ1, λ2, λ3 and their corresponding eigenvectors v1, v2, v3; the larger the eigenvalue, the greater the variance in that direction, that is, the more obvious the change of the point cloud in that direction;
[0035] Step 2.3.1.5. Main direction extraction:
[0036] Compare the sizes of the three eigenvalues and select the eigenvector corresponding to the largest eigenvalue as the first principal component direction of the point cloud; let the largest eigenvalue be λ max , the corresponding eigenvector v max As shown in formula (7):
[0037]
[0038] Eigenvector v max Indicates the main direction of the point cloud, and the directions of the tree trunks are mostly consistent with this main direction;
[0039] Step 2.3.1.6. Determine whether the main direction is reasonable and obtain the cluster:
[0040] In order to determine whether the main direction of the cluster is close to the vertical direction (Z axis), it is necessary to calculate the main direction v max The angle with the vertical vector z = [0, 0, 1]; the angle θ is calculated as shown in formula (8):
[0041]
[0042] Where, v max z represents the dot product of two vectors, ||v max || represents vector v max The modulus of , that is, the length of the vector;
[0043] If the angle θ is less than the set threshold, the main direction of the cluster is considered to be close to the vertical direction, and the cluster point cloud can be identified as a single tree trunk point cloud. Otherwise, the cluster point cloud is put back into the original point cloud set, the cluster radius is readjusted, and step 2.2 is performed again to extract potential tree trunks.
[0044] Step 2.3.2. Using the extracted main direction as the index, a variable cone nearest neighbor search method is designed to gradually expand the trunk point cloud and effectively identify the missed trunk points.
[0045] Step 2.3.3. Dynamically adjust the search range to ensure the continuity and integrity of the tree trunk point cloud structure;
[0046] Step 3. Process the remaining point clouds to achieve accurate extraction of the crown point cloud and ultimately achieve single tree segmentation. The specific process is as follows:
[0047] Step 3.1. Trunk-constrained canopy point cloud segmentation: Using a trunk-constrained clustering method, canopy point cloud segmentation is performed based on the correctly segmented trunk point cloud.
[0048] Step 3.2. Edge canopy point cloud segmentation based on the gravity model to achieve complete single tree point cloud segmentation.
[0049] Furthermore, the preprocessing in step 1 is to filter the original point cloud using the CSF ground filtering method to filter out the ground point cloud and retain non-ground point clouds such as trees, vegetation and buildings, and then perform denoising to reduce the computational load of subsequent algorithms.
[0050] Furthermore, the specific process of step 2.2 is:
[0051] The point cloud obtained in step 2.1 is divided into core points, reachable points, and noise points. The high-density area is divided according to the set neighborhood radius (eps) and the minimum number of neighbor points (min_samples) and independent clusters are identified as candidate tree trunk clusters.
[0052] Among them, a core point refers to a point cloud that contains at least min_samples neighboring points within its eps neighborhood radius. All core points are connected to each other to form a cluster, and non-core points that are only within the neighborhood of the core point are reachable points. The reachable points are also classified into the corresponding clusters. The remaining isolated points cannot be connected to other points through density and are regarded as noise points.
[0053] Furthermore, the specific process of step 2.3.2 is as follows:
[0054] Step 2.3.2.1. Project each cluster along a cross section perpendicular to the principal direction. Then, perform a least-squares fit of a circle on all projected points. The center of the circle is the midpoint of the trunk centerline, and the line passing through this point in the same direction as the principal direction is the trunk centerline. That is, project the xyz coordinates of the points along the direction vector of the principal direction to obtain a planar point set. Then, perform a least-squares fit on all planar point sets, and constrain the fitting results to a circular curve.
[0055] Step 2.3.2.2. Construct a variable cone search area based on the centerline of the tree trunk. The centerline of the tree trunk is the center of the circle obtained in step 2.3.2.1, and the direction is the same as the trunk direction vector:
[0056] The initial search radius is set as the fitting radius of the trunk base, and a linear attenuation model is used. The search radius r(z) gradually decreases with increasing height. The radius calculation formula is shown in formula (9):
[0057] r(z)=r0-α(zz min ) (9)
[0058] Where r0 is the trunk radius of the lowest height layer, α is the adjustable search radius attenuation rate, which determines the cone contraction rate, and z min is the lowest altitude, z is the current search altitude;
[0059] Based on the radius calculation formula, the centerline of the trunk is used as the search axis, and the height is divided into equal intervals from z_min to z_max. A search area is constructed at each height layer, and the KD-Tree is used to query the nearest neighbor points to extract all the trunk point clouds. When r(z) is 0, the search ends and all the searched points are returned, which is the complete single tree trunk point cloud.
[0060] Furthermore, the specific process of step 3.1 is as follows:
[0061] Step 3.1.1. Modeling of the imaginary trunk axis: Use PCA to extract the main direction of the trunk point cloud, take the lowest point of the trunk as the center starting point, and construct the imaginary trunk axis along the main direction to simulate the vertical extension trajectory of the trunk; the axis is from the bottom center point P base and the main direction vector v max Definition, extending the spatial constraints of tree trunks to full canopy height;
[0062] Step 3.1.2. Calculate canopy point attribution: For each unlabeled point P except the trunk point cloud i , calculate its vertical distance to the main axis of all tree trunks, the formula is:
[0063]
[0064] The minimum distance is obtained by vector cross product and modulus ratio, and combined with the height range screening condition (the height of the canopy point must be within the reasonable range of the corresponding trunk), the point is assigned to the trunk with the nearest distance and matching height.
[0065] Furthermore, the specific process of step 3.2 is as follows:
[0066] Step 3.2.1. Edge point screening: According to each unlabeled point P except the trunk point cloud i The vertical distance to the main axis of all tree trunks is used to filter out edge points that are far from the trunks. These points need to be reassigned labels due to insufficient spatial constraints and serve as core objects for subsequent gravity model processing.
[0067] Step 3.2.2. Spatial Cylinder Construction and Neighborhood Query: With each edge point as the center, construct a cylinder search range with radius r0 and height h0. Use KD-Tree to quickly query neighboring points within the cylinder and obtain the index of the neighboring points to provide spatial local data support for gravity calculation;
[0068] Step 3.2.3. Attraction calculation model: Assuming that the attraction of correctly classified points in the neighborhood to unlabeled points is inversely proportional to the square of the distance, the total attraction value is calculated by formula (11)
[0069]
[0070] Among them, d ij is an unmarked point P i To the neighboring point P j The Euclidean distance is a small constant (to prevent division by zero errors). By comparing the total gravity values of different trunk instances, the unclassified points are assigned to the single tree label corresponding to the one with the largest gravity.
[0071] Step 3.2.4. Dynamic search range adjustment: If the number of neighborhood points within the current cylinder is lower than the set threshold, gradually expand the search radius r and height h until enough reference points are obtained; this mechanism balances computational efficiency and matching stability to avoid misclassification due to local data sparsity.
[0072] A parameter estimation method based on single tree segmentation results includes the following steps:
[0073] S1. Based on the single tree segmentation results, perform voxel division to obtain the single tree structure parameters and point cloud structure parameters;
[0074] S2. Use the Pearson correlation analysis method to calculate the correlation between the dependent variable and the independent variable, and select the input variable with the highest characteristic value;
[0075] The calculation formula is shown in formula (12):
[0076]
[0077] Where, Corr(νar,dep) is the Pearson correlation coefficient between the dependent variable νar and the independent variable dep, M is the total number of samples, ζ νar is the sample standard deviation of νar, ζ dep is the sample standard deviation of dep; νar′ is the mean value of νar, dep′ is the mean value of dep;
[0078] The dependent variables include above-ground biomass (AGB), carbon storage (CS), leaf area index (LAI), and living vegetation volume (LVV); the independent variables include 27 point cloud structure parameters and single tree structure parameters, which include three measurement parameters (DBH, crown height, and crown diameter);
[0079] The larger the absolute value of the correlation coefficient, the stronger the correlation. The interval [0, 0.1] indicates no correlation, the interval [0.1, 0.3] indicates weak correlation, the interval [0.3, 0.5] indicates moderate correlation, and the interval [0.5, 1] indicates strong correlation.
[0080] S3. For different tree species, an improved random forest regression algorithm is used, with the tree species' individual tree structure parameters and point cloud structure parameters as input to obtain the biological estimation parameters of the tree species:
[0081] S3.1. Use the Least Absolute Shrinkage and Selection Operator (Lasso) for feature selection to remove variables that have little impact on the results and achieve adaptive feature selection.
[0082] Lasso regression is an L1 regularization method that compresses the regression coefficients so that the coefficients of some features shrink to zero, thereby achieving automatic feature selection. Its optimization objective function is defined as shown in formula (13):
[0083]
[0084] Among them, M is the number of samples, p is the number of features, and y k is the true value of the single tree biological parameter (such as LAI), x k is the input point cloud feature vector, w is the regression coefficient, λ reg is a regularization hyperparameter used to control the degree of feature shrinkage, and the superscript T represents the rank transformation;
[0085] When λ reg When the value is large, the model will force the coefficients of some features to be zero, retaining only important variables. reg When the value is smaller, more features are retained and the model's fitting ability is enhanced. To ensure the rationality of feature selection, the present invention uses cross-validation to automatically optimize the λ value, so that it can avoid overfitting while retaining key features;
[0086] S3.2. Introducing the pigeon swarm optimization algorithm for hyperparameter tuning, using a swarm intelligence search mechanism to optimize the hyperparameter configuration of random forest regression, thereby improving the stability of biological parameter estimation;
[0087] Furthermore, the point cloud structure parameters in S1 include canopy void ratio parameter, height statistics parameter, height quantile parameter and height density parameter;
[0088] The crown void ratio measures the spatial filling of the canopy area, reflecting the sparseness of the crown and the density of branches and leaves. Trees with a higher crown void ratio have less branch and leaf cover, while trees with a lower crown void ratio usually have a dense canopy with more leaf cover and may exhibit stronger photosynthetic capacity.
[0089] First, the point cloud data is spatially divided according to a fixed voxel size. Then, the number of points in each voxel is counted, and low-density voxels are defined (i.e., the number of points is less than 20% of the average number of points in all voxels). Finally, the proportion of all low-density voxels is calculated, which is the canopy void ratio. The number of low-density voxels N is empty The calculation formula is shown in formula (14):
[0090]
[0091] Where μ is the voxel number in space, For the number of points, N voxels is the total number of voxels.
[0092] On this basis, the calculation formula of crown void ratio is shown in formula (15):
[0093]
[0094] Height statistical parameters are statistical features used to describe the height distribution of tree point clouds, reflecting the overall shape, variability, symmetry, and growth status of trees in three-dimensional space. Height statistical parameters include maximum height, average height, height standard deviation, height skewness, height kurtosis, and height median.
[0095] The lowest point height of a single tree point cloud is 0, and the highest height is the highest value of the z coordinate in the single tree point cloud;
[0096] The average height is the mean value of the height of all point clouds, which reflects the overall distribution trend of the point cloud height. Its calculation formula is shown in formula (16):
[0097]
[0098] The height standard deviation is used to measure the degree of discreteness of the point cloud height and represents the uniformity of the tree canopy. Its calculation formula is shown in formula (17):
[0099]
[0100] The height skewness reflects the symmetry of the height distribution. If the skewness is positive, it means that the height distribution is right-skewed, that is, there are more high points; if the skewness is negative, it means that there are more low points. Its calculation formula is shown in formula (18):
[0101]
[0102] The height kurtosis describes the sharpness of the height distribution. A higher kurtosis means that most points are concentrated near the mean, while a lower kurtosis means that the data is flatter. Its calculation formula is shown in formula (19):
[0103]
[0104] The median height represents the median value of the point cloud height, which is less affected by outliers than the mean. The calculation formula is shown in formula (20):
[0105] Z median =median(z i ) (20)
[0106] The height quantile parameter is used to describe the height of different percentage positions and can be effectively used to analyze the canopy thickness and growth layer of trees. The present invention selects 10 quantiles between 10% and 99% to quantify the height characteristics of different levels. For a certain height quantile h, the quantile calculation formula is shown in formula (21):
[0107] P h =percentile(z,h) (21)
[0108] where h∈{10,20,30,...,90,99}.
[0109] The height density parameter is used to measure the point cloud density in different height intervals, so as to analyze the hierarchical structure and vertical distribution pattern of trees. First, the height range of the entire tree is divided into 10 equally spaced intervals, as shown in formula (22):
[0110]
[0111] Then, calculate the proportion of point clouds in each height interval D h , the calculation formula is shown in formula (23):
[0112]
[0113] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0114] 1. This paper addresses the difficulty of segmenting individual trees in urban scenes due to overlapping tree crowns. By designing a point cloud tree segmentation algorithm based on tree geometry, this method uses tree geometry as a constraint, combines point cloud reflectance information with clustering algorithms and principal direction indexing, and accurately extracts tree trunks and crowns from scene point clouds. Furthermore, a boundary point cloud optimization strategy is introduced to improve the accuracy of crown edge point segmentation, ensuring the integrity of individual tree point clouds and the reliability of structural parameter extraction. This algorithm also enables the automatic extraction of key structural features such as tree height, diameter at breast height, and crown width.
[0115] 2. This paper addresses the problems of high feature redundancy and poor model interpretability in existing estimation methods. Based on the random forest model, it introduces an adaptive feature selection algorithm to improve variable screening efficiency, uses pigeon flock optimization to dynamically adjust hyperparameters, enhances model generalization and stability, and combines the SHAP method to carry out feature importance analysis, improves the interpretability and credibility of the model, and finally designs an estimation model for estimating biological parameters such as leaf area index, biomass, and carbon storage. At the same time, model verification and comparative analysis are carried out through measured sample data to ensure the reliability of the estimation results, providing reliable support for the quantitative assessment of urban green resources and carbon storage calculation. BRIEF DESCRIPTION OF THE DRAWINGS
[0116] Figure 1 Schematic diagram of the process of the present invention.
[0117] Figure 2 Schematic diagram of the DBSCAN algorithm principle.
[0118] Figure 3 This is a schematic diagram of the individual tree numbering in the first sub-area of the survey area in Example 1.
[0119] Figure 4 This is the result of Pearson correlation analysis;
[0120] Among them, (a) the correlation between Ginkgo biloba input parameters and AGB (b) the correlation between Metasequoia glyptostroboides input parameters and AGB (c) the correlation between Jacaranda input parameters and LAI (d) the correlation between Phoebe zhennanzi input parameters and LAI. DETAILED DESCRIPTION
[0121] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the implementation methods and drawings.
[0122] A single tree segmentation and parameter estimation method based on vehicle-mounted LiDAR point cloud, the flow chart of which is as follows Figure 1 As shown, the following steps are included:
[0123] Step 1. Use the vehicle-mounted LiDAR to scan and obtain the original scene point cloud data, and pre-process the original point cloud data. The original scene point cloud data usually contains point clouds of multiple categories such as ground, vegetation, and buildings. Therefore, the original point cloud needs to be pre-processed to improve the accuracy and computational efficiency of subsequent single tree segmentation. The specific process is as follows:
[0124] Step 1.1. Filtering the Ground Point Cloud: Use the CSF ground filtering method with the following parameters: Cloth resolution = 2.0, Max iterations = 500, Classification threshold = 0.5. After CSF processing, ground points are filtered out, while non-ground point clouds such as trees, vegetation, and buildings are retained. This provides a denoised data foundation for subsequent processing and reduces the computational load of subsequent algorithms.
[0125] Step 1.2. Extracting vegetation point clouds: After filtering out ground points, semantic segmentation is used to accurately extract vegetation point clouds from buildings and man-made structures that are mixed in the remaining point cloud. Semantic segmentation is performed using the PointNet++ deep learning network. During the training phase, model parameters are optimized by extensively annotating point cloud data, enabling accurate classification of trees, buildings, and grass. During the inference phase, the de-grounded point cloud is fed into the trained model, which performs point-by-point classification and outputs labeled vegetation point cloud data.
[0126] Step 2. Extract and segment single tree trunks:
[0127] Step 2.1. Filter crown points based on reflection intensity. Set a reflection intensity threshold and retain points with reflection intensities above the threshold (mainly tree trunks and a small number of low-level crown points). Filter low-intensity points (mainly crowns and edge trunk points).
[0128] In the lidar point cloud data, the reflection intensity reflects the signal strength returned by the laser pulse, and this value is affected by many factors such as the surface material, roughness, laser incident angle, and laser wavelength of the object. When other conditions are the same, the reflectivity of objects with rough surface materials and poor light transmittance is higher, and the corresponding collected point cloud reflection intensity is also higher. The method of using reflection intensity information to assist in target classification has been widely used in fields such as land cover classification and forest resource surveys, but there is no precedent for its application in tree trunk point cloud data. There is a significant difference in the reflection intensity between the trunk and the crown; the trunk is composed of dense wood fibers and has a high reflection intensity (average value of about 41019), while the crown has a lower reflection intensity (average value of about 28588) due to the complex structure of the leaves and branches, which causes laser scattering and absorption. The reflection intensity is affected by factors such as material, roughness, and incident angle, but the difference in the numerical distribution between the trunk and the crown is stable and can be used as a basis for classification;
[0129] The specific process is as follows:
[0130] Step 2.1.1. Reflection Intensity Data Statistics and Distribution Modeling: We manually separated the trunk and crown point cloud samples and analyzed the reflection intensity distribution. The trunk point cloud reflection intensity was concentrated in the high-value area (Gaussian mean 41019), while the crown point cloud was concentrated in the low-value area (Gaussian mean 28588). Statistics revealed that the 25th percentile of the trunk point cloud reflection intensity and the 75th percentile of the crown point cloud reflection intensity were both close to 35000, indicating that this value is the critical threshold for distinguishing the two.
[0131] Step 2.1.2. Threshold Setting and Optimization: A reflection intensity threshold of 35,000 is initially selected, which retains approximately 75% of the trunk point cloud and filters 75% of the crown point cloud. In practice, if purer trunk data is required, the threshold can be increased appropriately, but a balance must be struck between trunk retention and crown filtering efficiency to avoid excessively removing valid trunk points.
[0132] Step 2.1.3. Crown point filtering implementation: retain the point cloud with reflection intensity above the threshold (mainly tree trunks and a small number of low-level crown points), and filter out low-intensity points (mainly crown and edge trunk points). High reflection intensity points are used as input for the subsequent single tree segmentation algorithm, significantly reducing computational complexity;
[0133] At the same time, the crown point filtering method based on reflection intensity was verified: including multi-survey area verification and robustness testing, anti-interference and error control, and actual application effect evaluation, which demonstrated the excellence of the method of the present invention in filtering crown point clouds.
[0134] Multi-area validation and robustness testing: The universality of the threshold was verified in measurement areas 1-3: the 25th percentile of the trunk point cloud (34671–36721) and the 75th percentile of the crown point cloud (35832–37774) were both close to 35,000, proving that the threshold can stably distinguish between trunk and crown point clouds in different environments.
[0135] Interference prevention and error control: Factors such as scanning distance and incident angle may reduce the reflection intensity of tree trunk edge points. By combining spatial distribution analysis (such as the clustering of tree trunk point clouds), we can correct incorrectly filtered tree trunk points and improve filtering accuracy.
[0136] Practical application evaluation: After filtering with a threshold of 35,000, the crown point cloud removal rate is approximately 75%, and the trunk retention rate is approximately 75%. While ensuring the accuracy of subsequent algorithms, the threshold can be dynamically adjusted to balance data purity and integrity.
[0137] Step 2.2. Use the density-based spatial clustering algorithm (DBSCAN) to extract tree trunk points and obtain candidate tree trunk clusters;
[0138] After crown point filtering, potential trunk point cloud extraction is performed to perform preliminary segmentation of the scene point cloud into individual tree trunks, accurately identifying the number of individual trees within the measurement area and providing a foundation for subsequent complete trunk extraction. Because the base of a tree trunk near the ground has a larger diameter and relatively stable geometric characteristics, its corresponding point cloud is densely distributed and less affected by crown points. Therefore, a clustering method can be used to segment the trunk point cloud from bottom to top, effectively identifying potential clusters of individual tree trunk points and further confirming the number of trees within the measurement area.
[0139] The present invention adopts a density-based clustering algorithm (Density-Based Spatial Clustering of Applications with Noise, DBSCAN) to extract trunk points. The algorithm is suitable for processing point cloud data with irregular shapes, and can effectively deal with noise points and outliers. Compared with traditional K-means or mean shift clustering algorithms, DBSCAN is more suitable for single tree trunk segmentation tasks. First, the number of single trees in different survey areas is uncertain. The algorithm can adaptively identify trunk clusters according to the point cloud density without presetting the number of trees, thereby improving applicability. Secondly, due to the different shapes of trees, the trunk point cloud may present an irregular distribution. DBSCAN ensures segmentation accuracy by identifying density-connected point cloud clusters, which is not affected by the trunk shape. In addition, point cloud data may contain measurement errors or interference points. DBSCAN can effectively improve the reliability of trunk point extraction by automatically eliminating outliers with low density.
[0140] The core idea of the DBSCAN algorithm is to divide data points into core points, reachable points, and noise points, and to divide high-density areas and identify independent clusters by setting the neighborhood radius (eps) and the minimum number of neighboring points (min_samples). A core point refers to a point cloud that contains at least min_samples neighboring points within its eps neighborhood. All core points are connected to form a cluster, while non-core points that are only within the neighborhood of the core point are assigned to the corresponding cluster, and the remaining isolated points are regarded as noise points. The schematic diagram of the DBSCAN algorithm principle is shown below. Figure 2 As shown in the figure, where min_samples = 4 and eps is represented by a circle, all red points are core points because they contain at least four points (including themselves) within the range enclosed by eps. Points B and C are connected through other core points, so even though they are not core points, they are still considered to belong to the same cluster. Point N, on the other hand, is neither a core point nor connected to other points through density, so it is considered a noise point.
[0141] The schematic diagram of the DBSCAN algorithm principle is as follows Figure 2 As shown in the figure, where min_samples = 4 and eps is represented by a circle, all red points are core points because they contain at least 4 points (including themselves) within the range enclosed by eps. Points B and C are density-connected via other core points. Although they are not core points, they are reachable points and therefore belong to the same cluster. Point N is neither a core point nor can it be density-connected to other points, so it is considered a noise point.
[0142] Therefore, eps and min_samples are two important parameters. The rationality of their values directly affects the extraction of potential tree trunks. This step is based on the geometric characteristics of the tree trunk point cloud and sets the two key parameters of DBSCAN as follows:
[0143] ①eps = 0.5m: Since there is usually a certain distance between the trunks of adjacent trees, and the trunk diameter is generally between 0.2m and 0.7m, eps is set to 0.5m to ensure that DBSCAN can correctly cluster the point clouds of the same trunk while avoiding misclassifying adjacent trees into the same cluster.
[0144] ② min_samples = 10: In the low-height area of the tree trunk, the point cloud density is high. To prevent too small clusters from being mistaken for tree trunks, the present invention sets min_samples to 10 to ensure the stability of each tree trunk point cluster and eliminate low-density noise points.
[0145] Based on the above parameters, the potential trunk extraction work based on the clustering algorithm is carried out to obtain the single tree point cloud of the scene and the base trunk point cloud segmentation results;
[0146] Step 2.3. Use the main direction index to extract the complete tree trunk:
[0147] After completing the potential trunk extraction based on the clustering algorithm, although the number of individual trees in the scene was successfully identified and the trunk base point cloud was extracted, some trunk point clouds were still missing, especially the trunk point clouds at higher positions and the trunk edge points that were mistakenly filtered out due to the crown point cloud filtering. On the other hand, the basic clustering algorithm may have segmentation errors, such as misclassifying the trunk point clouds of two or more adjacent individual trees into the same cluster. In order to obtain the most complete individual tree trunk point cloud possible and check the clustering error problem that occurs in the clustering algorithm, thereby supporting the subsequent crown point cloud extraction and individual tree structural parameter calculation, this step further invents a complete trunk extraction algorithm based on the main direction index.
[0148] Considering the geometric characteristics of tree trunks in natural environments, their point cloud distribution typically forms a columnar structure along an approximately vertical direction. Therefore, this step first uses the principal component analysis (PCA) method to calculate the principal axis direction of the trunk point cloud to accurately obtain the trunk growth direction of each tree. Subsequently, using the extracted principal direction as an index, a variable cone nearest neighbor search method is designed to gradually expand the trunk point cloud and effectively identify missed trunk points. At the same time, to adapt to the morphological changes of different trees, this method can dynamically adjust the search range during the expansion process to ensure the continuity and integrity of the trunk point cloud structure.
[0149] First, in each cluster obtained in the previous step, the PCA algorithm is used to extract the main direction of the point cloud. The first principal component of PCA usually corresponds to the axis with the largest change in direction in the point cloud, and the main axis of the tree trunk is usually close to the vertical direction (Z axis). Therefore, the angle between the first principal component of each cluster and the Z axis is the main direction of the trunk. When the angle between the main direction of a cluster and the vertical direction is greater than the set threshold, it can be considered that an incorrect clustering has occurred, that is, multiple close trees may be clustered into one category. At this time, the cluster radius can be reduced and the DBSCAN algorithm in the previous step can be re-called until all single tree trunks are accurately segmented.
[0150] Step 2.3.1. Use the principal component analysis (PCA) method to calculate the principal axis direction of the trunk point cloud in the trunk candidate cluster to accurately obtain the trunk growth direction of each tree;
[0151] Step 2.3.2. Using the extracted main direction as the index, a variable cone nearest neighbor search method is designed to gradually expand the trunk point cloud and effectively identify the missed trunk points.
[0152] Step 2.3.3. Dynamically adjust the search range to ensure the continuity and integrity of the tree trunk point cloud structure;
[0153] Step 2.4. Accurately extract the crown point cloud to achieve single tree segmentation. The specific process is as follows:
[0154] Step 2.4.1. Use a trunk-constrained clustering method to segment the canopy point cloud. The optimized method compensates for missing high-level trunk points by imaginary trunk axes, forcing vertical spatial consistency between canopy points and trunks, thus resolving upper and lower canopy segmentation errors. The matching accuracy between canopy points and corresponding trunks is significantly improved, ensuring the integrity of individual tree instances.
[0155] The nearest neighbor search method is often used in the existing technology. This method is based on the segmented single tree trunk point cloud, calculates the Euclidean distance of each unclassified canopy point to all trunk points, and assigns it to the single tree instance corresponding to the nearest trunk (an independent trunk point cloud instance that has completed clustering and labeling); the canopy points are marked by spatial proximity expansion (marking according to the distance from the unmarked point to the marked trunk, and each unclassified point is classified to the trunk label closest to it, thus becoming a point cloud belonging to the corresponding tree). However, the lack of high-position trunk points causes the upper and lower canopies to be incorrectly divided into different instances, resulting in segmentation confusion.
[0156] Step 2.4.2. Edge point cloud segmentation based on the gravity model. The canopy point cloud at the edge of the tree still suffers from misclassification. Canopy point cloud segmentation methods that only consider the trunk center constraint have limitations. Therefore, the present invention designs a canopy point cloud segmentation method based on the gravity model to re-segment the edge canopy points.
[0157] Determine edge points by dual conditions:
[0158] (1) The vertical distance to all trunk axes is greater than a threshold (e.g., 1.5 times the average crown width)
[0159] (2) Located in the outer space of the tree canopy (identified by alpha-shape algorithm)
[0160] After applying the method of the present invention, most edge canopy points are correctly classified into their respective individual tree instances, significantly improving the integrity and accuracy of canopy segmentation, especially reducing over-segmentation and under-segmentation problems in the interlaced canopy areas;
[0161] Step 4. A series of point cloud structure parameters are introduced:
[0162] Since factors such as tree height, canopy structure, and branch and leaf density can reflect the growth status of trees to a certain extent, the point cloud structure parameters constructed based on this information can be used to establish a data-driven biomass estimation model. In order to comprehensively describe the canopy density, tree height characteristics, and their hierarchical structure information contained in the single tree point cloud, this step designs a total of 27 point cloud structure parameters in four categories: canopy void ratio parameters, height statistical parameters, height quantile parameters, and height density parameters. These parameters are quantified from multiple angles from the perspectives of canopy point cloud structure, overall point cloud height distribution, and characteristics of different height intervals, thereby effectively describing the sparseness of the canopy, the concentration trend and discreteness of tree height distribution, the proportion of point clouds in different height intervals, and other information, providing a data basis for subsequent biological parameter estimation;
[0163] Step 5. Perform Pearson correlation analysis on the point cloud structural parameters: The variables for the Pearson correlation analysis include 27 point cloud structural parameters, namely, crown void ratio, 6 height statistical parameters, 10 height quantile parameters, 10 height density parameters, and 3 individual tree structural parameters, namely, diameter at breast height, crown height, and crown diameter, for a total of 30 parameters.
[0164] Step 6. For a certain tree species, based on the highly correlated point cloud structural parameters and individual tree structural parameters of the tree species, an improved random forest regression algorithm is used to estimate the individual tree parameters of the tree species:
[0165] Step 6.1. Select the most representative input features through AFS to reduce data dimensionality and improve model learning efficiency.
[0166] Estimating individual tree parameters involves a large number of point cloud features, including individual tree structural parameters such as tree height, crown diameter, diameter at breast height, and crown height, as well as various point cloud structural variables such as crown porosity, height statistics, height quantiles, and height density. These variables can describe the three-dimensional morphology of trees from different perspectives. However, in high-dimensional feature spaces, some features may exhibit strong collinearity or redundancy, increasing computational complexity during model training and reducing generalization ability. Therefore, before building an estimation model, it is necessary to filter the input features and retain only the most representative variables to improve modeling efficiency.
[0167] The present invention adopts the Least Absolute Shrinkage and Selection Operator (Lasso) to perform feature selection to remove variables that have little impact on the results and realize adaptive feature selection.
[0168] Step 6.2. Use the Pigeon Occurrence Optimization (PIO) algorithm to optimize the hyperparameters of the random forest model to maintain high prediction accuracy and stability across different tree species and environmental conditions.
[0169] In the training process of the random forest regression model, the selection of hyperparameters plays a decisive role in model performance. The core hyperparameters of random forest include the number of decision trees, the maximum tree depth, the minimum number of split samples, the minimum number of leaf node samples, etc. These parameters affect the fitting ability, computational efficiency and generalization performance of the model. Traditional hyperparameter search methods (such as grid search and random search) have high computational costs in high-dimensional parameter space and are difficult to globally optimize. Therefore, the present invention introduces the pigeon flock optimization algorithm to perform hyperparameter tuning, and utilizes the swarm intelligence search mechanism to optimize the hyperparameter configuration of random forest regression, thereby improving the stability of biological parameter estimation.
[0170] PIO simulates the dynamic behavior of a pigeon colony during foraging and migration, using a pilot pigeon (optimal solution) to lead the flock in searching for the optimal hyperparameter combination. The algorithm consists of two phases: map memory and target navigation. In the map memory phase, the pigeon colony (candidate hyperparameter set) adjusts its search direction based on the historical optimal parameters, bringing the entire colony closer to the optimal solution. The update formula is shown in Equation (24):
[0171]
[0172] Among them, P i t is the hyperparameter position of the t-th generation pigeon, is the global optimal solution of the current population, and η is a random perturbation factor used to maintain the diversity of the search and prevent falling into the local optimum.
[0173] In the target navigation phase, the pigeons further optimize the search path of the hyperparameters based on the information of the neighboring individuals, ensuring that the entire pigeon group maintains a certain convergence speed while performing the global search. The update formula is shown in Equation (25):
[0174]
[0175] in, is the hyperparameter position of another pigeon randomly selected from the population to simulate the information interaction between individual pigeons and improve the search efficiency.
[0176] During the experiment, the present invention uses 5 pigeons (candidate solutions) for optimization, iterates 10 rounds, and conducts 50 experiments with different random seeds to finally select the optimal hyperparameter combination.
[0177] In the practical application of machine learning models, the transparency and interpretability of model prediction results are important criteria for measuring their reliability. For the task of estimating biological parameters of individual trees, understanding the impact of different input features on the prediction results helps optimize the model structure, select key variables, and improve the credibility of the prediction. Therefore, the present invention introduces the SHAP method to analyze the contribution of different input features to biological parameter estimation. SHAP measures the impact of each input variable on the model output by calculating the Shapley value of the feature, which is defined as shown in Equation (26):
[0178]
[0179] where φ j is the SHAP value of eigenvalue j, S is the feature subset that does not contain j, f(S) is the predicted value of the model on S, and f(S∪{j}) is the change in the predicted value after adding eigenvalue j. Represents the entire set of input features.
[0180] SHAP analysis can provide global feature importance and local explanation. Global feature importance analysis is used to measure the contribution of different variables in the overall prediction task and screen the variables that are most critical for biological parameter estimation. Local explanation is used to analyze the prediction process of a single sample, which can help understand why the biological parameter value of a certain tree is higher or lower than the mean, thereby providing a more intuitive prediction basis. In addition, the SHAP method can also be used for model optimization and feature screening. The present invention uses the top 10 features ranked by SHAP value for further feature selection, and combines the results of adaptive feature selection to ultimately determine the optimal feature set. This strategy not only reduces the computational cost of the model, but also improves the prediction stability and generalization ability.
[0181] Example 1
[0182] The following indicators are used to evaluate the effect of single tree segmentation in the present invention: mean Per-Tree Precision (mPTP), mean Per-Tree Recall (mPTR) and mean F1-score (mF1-score) are used as quantitative evaluation indicators to measure the overall effect of single tree segmentation.
[0183] In the task of single tree segmentation, the following two key factors are often focused on:
[0184] (1) Segmentation accuracy: the proportion of the point cloud predicted to be a tree that actually belongs to the tree;
[0185] (2) Segmentation recall: that is, how many of all point clouds belonging to a certain tree are correctly predicted.
[0186] The average single tree precision is used to measure the proportion of all points that are segmented into a certain tree that actually belong to the tree. The calculation formula is shown in formula (27):
[0187]
[0188] Where N is the total number of trees in the data set, TP ν is the number of points correctly predicted by the νth tree, FP ν is the number of points misclassified as the νth tree;
[0189] The average single tree recall rate mPTR is used to measure how many of the true points of all single trees are correctly predicted. The calculation formula is shown in formula (28):
[0190]
[0191] Among them, PTR ν is the number of points in the true points of the νth tree that are incorrectly predicted as other categories;
[0192] The average single-tree F1-score is the harmonic mean of PTP and PTR, which comprehensively measures the segmentation quality of a single tree. The calculation formula is shown in Equation (29):
[0193]
[0194] Based on the above indicators, the performance evaluation of the single tree segmentation algorithm will be carried out in the future.
[0195] To systematically validate the effectiveness of a point cloud tree segmentation algorithm based on tree geometry, we selected a survey area containing 241 ginkgo trees as a representative experimental subject. This area was divided into four contiguous subareas (numbered 1-4) for validation, and the sample size and distribution characteristics were statistically significant. Figure 3 This is a schematic diagram of the numbering of individual trees in the first sub-area of the survey area; different colors in the figure represent independent individual tree instances.
[0196] The experimental process includes data preprocessing and multi-stage segmentation. After optimizing the quality of the original point cloud data through ground filtering and vegetation point separation technology, the reflection intensity threshold is used to filter the crown interference points based on the characteristics of the thick ginkgo trunk. The improved clustering algorithm and main direction indexing technology are combined to complete the trunk point cloud extraction. 46, 70, 64 and 61 independent trunk clusters were obtained in the four sub-areas respectively, which fully matched the field survey data and had no segmentation errors, laying a solid foundation for subsequent processing. In the crown segmentation stage, the main canopy segmentation is first achieved based on the trunk spatial constraints. Then, the edge points of the intersection area are secondary optimized and allocated through the gravity model, effectively reducing the misclassification rate of complex canopy structures.
[0197] Except for a few areas with significant canopy overlap, the segmentation of most individual trees was ideal, and the point clouds of different trees were accurately distinguished overall, validating the effectiveness of the point cloud individual tree segmentation algorithm based on tree geometric feature constraints. To further quantitatively evaluate the algorithm's segmentation performance, this step calculated the mean individual tree precision (mPTP), mean individual tree recall (mPTR), and mean F1-score (mF1-score) for the four subregions of the survey area. The results are shown in Table 1.
[0198] Table 1 Results of quantitative evaluation indicators for single tree segmentation in the survey area
[0199]
[0200] The following conclusions can be drawn from the data in Table 1:
[0201] (1) The overall tree segmentation effect is good. In all regions, mPTP, mPTR, and mF1-score are all over 98.5%, indicating that the segmentation algorithm can maintain high precision and recall in different regions, ensuring the correct identification and attribution of individual trees.
[0202] (2) The segmentation results between the parts are highly consistent. The index values of the four regions are relatively close, indicating that the algorithm has good stability and robustness in different regions and the segmentation performance will not be significantly affected by scene changes.
[0203] (3) The decomposition precision and recall are relatively balanced. The mF1-score of all regions is above 98.78%, and the overall region reaches 99.13%, indicating that the algorithm can achieve a good balance between high precision and high recall, ensuring the overall quality of the segmentation results.
[0204] In general, based on the above qualitative and quantitative evaluation results, the single tree segmentation algorithm based on tree geometric feature constraints proposed in this invention performs excellently on the point cloud data of the selected scene, can efficiently and accurately segment single tree point clouds, and the algorithm has strong adaptability and stability between different regions.
[0205] For biological parameter estimation, backpack LiDAR was used to collect point cloud data from 81 ginkgo trees and 73 metasequoia trees, along with their corresponding AGB and CS measurements. Point cloud data from 38 jacaranda trees and 67 zhennan trees, along with their corresponding LAI measurements, were also collected. All point cloud data were preprocessed for individual tree segmentation and species identification to ensure data integrity and consistency. On this basis, individual tree structural parameters and point cloud structural parameters were extracted from each individual tree point cloud. Input parameter sets for different tree species were constructed, and improved random forest models were trained separately to ultimately obtain the optimal estimation results. The model's performance in different tree species and different biological parameter estimation tasks was evaluated by comparing the predicted values with the measured data, thereby verifying the effectiveness of the method.
[0206] The evaluation indicators used in the invention include root mean square error (RMSE), mean absolute error (MAE), mean absolute percentage error (MAPE), explained variance score (EVS) and determination coefficient R 2 Five types:
[0207] RMSE is one of the key indicators for measuring model prediction error. It represents the average error between the predicted value and the true value. The unit is consistent with the target variable. The calculation formula is shown in formula (30):
[0208]
[0209] RMSE emphasizes the impact of large errors through the sum of squares. Therefore, it is more sensitive to large errors and is suitable for applications where large deviations are of particular concern. A smaller value indicates a smaller prediction error and a better fit. However, due to the presence of the squared term, RMSE is more sensitive to outliers and may overestimate the error when the data contains large outliers.
[0210] MAE is another important indicator for measuring prediction error. It directly calculates the average absolute value of all errors, thereby reflecting the average deviation between the predicted value and the true value. The calculation formula is shown in formula (31):
[0211]
[0212] Unlike RMSE, MAE doesn't amplify the effects of larger errors, making it more stable in the presence of outliers. MAE is easy to understand and directly quantifies the degree of deviation between the model's prediction and the true value. However, because it uses absolute values, it may be less able to distinguish between small and large errors than RMSE. MAE is often used in tasks that require a stable measurement of error size and do not want outliers to have a significant impact on model evaluation.
[0213] MAPE is a measure of the percentage of prediction error relative to the true value. It can effectively evaluate the relative error level of the model. The calculation formula is shown in formula (32):
[0214]
[0215] Because MAPE is expressed as a percentage, it is unit-independent, allowing direct comparison of data across different scales. However, MAPE can be prone to instability when the true value is close to zero, potentially amplifying the impact of errors. Furthermore, it cannot distinguish between overestimations and underestimations; even if the forecast is overestimated or underestimated by the same amount, the MAPE error remains the same.
[0216] EVS is used to measure the proportion of the target variable variance that the model prediction value can explain. The calculation formula is shown in formula (33):
[0217]
[0218] Among them, Var represents variance, and the value range of EVS is between 0 and 1. The closer the value is to 1, the better the model can explain the variance changes of the target variable.
[0219] EVS focuses on the overall variance of the predicted values, rather than the specific error size, and therefore measures the model's stability to the data. While EVS can well reflect the overall fit of the model, it is insensitive to the specific distribution of the errors, so a comprehensive evaluation requires combining metrics such as RMSE and MAE. EVS is suitable for regression tasks and is particularly valuable for assessing the ability of EVS to explain data variance.
[0220] Coefficient of determination R 2 It is used to measure the goodness of fit of the model, that is, the proportion of the total variance of the data that the model can explain. Its calculation formula is shown in formula (34):
[0221]
[0222] R 2The value is between [-∞,1]. The closer the value is to 1, the better the model fit is and the better it can explain the changes in the data. 2 is close to 0, indicating that the model's predictive ability is comparable to that of the simple mean prediction, while R 2 If it is less than 0, it means that the model is even worse than the simple mean prediction. 2 It intuitively measures the overall goodness of fit of the model, but it is sensitive to the number of features, and adding additional variables may increase R 2 value, but it does not necessarily improve the generalization ability of the model.
[0223] It can be seen that each evaluation indicator has its corresponding advantages and disadvantages. The comprehensive use of multiple indicators helps to more accurately evaluate the model effect and parameter estimation level.
[0224] For the AGB or LAI estimation task of the four tree species, we first need to analyze the correlation between the input parameters and the measured values. After calling the Pearson correlation analysis method, the correlation analysis results between the input parameters and the Ginkgo AGB, Metasequoia AGB, Jacaranda LAI, and Phoebe zhennan LAI are obtained as follows: Figure 4 As shown in Figure 2, there is a direct linear relationship between carbon storage CS and AGB, and there is no authoritative way to obtain the true value of LVV, so it will not be evaluated separately.
[0225] According to the results of correlation analysis, Figure 4 In (a), the AGB of ginkgo trees is primarily influenced by tree height, crown diameter, and height standard deviation. Taller trees generally possess higher biomass reserves, while crown diameter reflects the horizontal expansion capacity of the crown and is closely related to biomass accumulation. Notably, the height standard deviation is highly correlated with the AGB, indicating that ginkgo biomass is largely dependent on the complexity of its vertical structure. Low-level point cloud density features generally exhibit a negative correlation, indicating that ginkgo trees with highly concentrated structures are likely to have larger AGBs.
[0226] Figure 4 In (b), the AGB of Metasequoia exhibits a different trait-dependence pattern than that of Ginkgo biloba. Its AGB is more significantly positively correlated with tree height and crown diameter, indicating that, as a typical tall tree, its biomass is primarily determined by height. Furthermore, the kurtosis and skewness of the height distribution also have a strong influence on AGB, reflecting that Metasequoia biomass is closely related to the concentration of its height distribution. Metasequoia's AGB is also highly sensitive to DBH, consistent with traditional allometric growth theory. Compared to Ginkgo biloba, Metasequoia is less dependent on low-level density characteristics, making structural integrity the key to estimating its biomass.
[0227] Comparing ginkgo and metasequoia from an AGB perspective reveals that while both are influenced by tree height and crown diameter, their emphasis differs. Ginkgo relies more on the degree of height variation in the point cloud, emphasizing vertical structural diversity; metasequoia, on the other hand, relies more on tree height itself and the centrality of its structure. This difference reveals fundamental differences in the growth strategies and spatial occupation patterns of the two tree species. Ginkgo grows more dispersed and has a variable structure, while metasequoia tends to maintain a unified, tall, upright structure. This morphological difference is clearly reflected in the point cloud features.
[0228] Figure 4 In (c), the LAI estimation of Jacaranda trees relies primarily on characteristics such as crown diameter, tree height, and canopy void ratio. Larger crown diameters generally indicate a wider canopy coverage, which is positively correlated with leaf area; increased tree height also tends to result in a larger total leaf volume. Canopy void ratio directly reflects the degree of light penetration and leaf density within the canopy, making it a crucial factor in LAI estimation. Furthermore, the influence of height skewness indicates that Jacaranda leaf area is constrained by the asymmetry of its height distribution, reflecting the complexity of its canopy structure.
[0229] Figure 4 In (d), the LAI of Phoebe zhennan is clearly controlled by the statistical characteristics of the point cloud, particularly skewness and kurtosis, indicating that its leaf area index is more dependent on the overall pattern of the height distribution. Compared with Jacaranda, Phoebe zhennan is more dependent on structural parameters such as tree height, canopy height, and diameter at breast height (DBH), indicating that its LAI is more influenced by overall structural morphology rather than simply changes in density. Crown diameter and canopy void fraction also have some influence, but their contributions are slightly lower than those of Jacaranda.
[0230] Comparing the LAI estimates for Jacaranda and Phoebe zhennan reveals that while both are influenced by the standard deviation of crown diameter and height, the LAI of Jacaranda is more dependent on crown density and asymmetry, tending to reflect the actual density of leaf distribution in the crown. Meanwhile, the LAI of Phoebe zhennan focuses more on the stability and layering of structural distribution, demonstrating a comprehensive response to overall morphology. The differences in canopy characteristics between the two tree species not only influence their ecological functions but also reflect different parameter dependencies in LAI estimation.
[0231] In summary, the differences in growth morphology, structural complexity, and biomass distribution patterns among different tree species make the same point cloud features show different sensitivities and contributions in different tree species.
[0232] Parameters with correlations greater than 0.4 for each tree species were selected and trained using five machine learning methods: RFR, Improved-RFR, GBR, KNN, and XGB. The AGB of Ginkgo biloba, Metasequoia glyptostroboides, LAI of Jacaranda, and LAI of Phoebe zhennan were estimated. The evaluation results are shown in Table 2, with the optimal results highlighted in bold.
[0233] Table 2 Evaluation of biological parameter estimation results using machine learning methods
[0234]
[0235]
[0236] From the biological parameter estimation results in Table 2, the improved random forest algorithm (Improved-RFR) of the present invention has achieved significant advantages in both AGB and LAI estimation tasks. Compared with the traditional random forest, Improved-RFR reduces RMSE and MAE in all tasks and improves the determination coefficient R 2 and EVS, indicating that it has stronger adaptability and stability in different tree species and biological parameters. At the same time, in terms of MAPE, Improved-RFR has also reduced to varying degrees compared with RFR, especially in Ginkgo AGB estimation, where MAPE dropped from 9.08% to 7.53%, and in Metasequoia AGB estimation, where MAPE dropped from 10.53% to 10.03%, showing better generalization ability and robustness. Compared with other machine learning methods such as GBR, KNN, and XGB, Improved-RFR has a better R 2 Both the performance of the proposed method and EVS are at a high level, especially in the AGB estimation task, which shows that it has a strong ability to learn complex nonlinear relationships.
[0237] In the AGB estimation task, both Ginkgo and Metasequoia showed a relatively clear contrast trend. For Ginkgo AGB, the RMSE of Improved-RFR (67.27) was significantly lower than that of RFR (81.88), and it also had a clear advantage over GBR (67.83) and XGB (97.74). In addition, KNN performed poorly in AGB estimation, with an RMSE of 100.3 and R 2 It is only 0.5744, which shows that the local nearest neighbor based method is difficult to effectively capture the complex relationship between AGB and point cloud structural features. Similarly, in the estimation of AGB of Metasequoia, Improved-RFR is also better than RFR, GBR and XGB, R 2 It reached 0.6909 and the EVS was the highest, indicating that the improved feature selection and hyperparameter optimization strategy effectively improved the model's fitting ability. In contrast, KNN still performed poorly, with an RMSE of 46.93 and R 2 It is only 0.3462, indicating that the traditional distance-based regression method is difficult to effectively adapt to the high-dimensional feature space of the AGB estimation task.
[0238] In the LAI estimation task, Improved-RFR also showed a significant advantage. Taking Jacaranda as an example, the RMSE of Improved-RFR (0.1296) was lower than that of RFR (0.1412), and significantly better than GBR (0.1842) and XGB (0.2001), indicating that it can still maintain a high accuracy in the smaller-scale leaf area index prediction task. In addition, the R 2 The highest (0.7103), EVS also reached 0.7138, far exceeding GBR (0.415), XGB (0.309) and other models, indicating that it has a strong generalization ability in the LAI estimation task. In the Zhennan LAI estimation task, Improved-RFR still has a clear advantage, R 2 It reaches 0.7282, which is higher than RFR (0.6472) and GBR (0.6086), and the RMSE also drops to 0.5838, showing its stability and adaptability in predicting tree species LAI.
[0239] Further analysis of the estimation effects of different biological parameters shows that the overall R 2 It is higher, indicating that the structural parameters of point cloud data have better explanatory power for biomass estimation, while the R 2 The results are relatively low, indicating that the estimation of leaf area index still faces certain challenges. In particular, the R 2 Even lower than 0.6, indicating that traditional machine learning methods are difficult to accurately describe the nonlinear relationship between LAI and point cloud features. At the same time, Improved-RFR has a better R 2 Both are about 0.05-0.1 higher than RFR, proving that it has stronger stability in biological parameter estimation tasks.
[0240] From the comparison of estimations of different tree species, even for the same biological parameter (such as AGB or LAI), the best feature combination and model performance of different tree species are still quite different. For example, in the AGB task, the best R 2 Much higher than the Metasequoia AGB, while in the LAI task, the best R 2This indicates that the mapping relationship between biological parameters and point cloud features of different tree species is species-specific, indicating that early tree species identification is necessary to provide a prerequisite for targeted feature extraction and model optimization for subsequent biological parameter estimation. Without distinguishing tree species, directly using a unified model for AGB or LAI estimation may lead to a decrease in the model's generalization ability, ultimately affecting estimation accuracy. Therefore, a biological parameter estimation strategy based on tree species identification not only improves the model's adaptability but also increases the reliability of the estimation results.
[0241] The above description is only a specific embodiment of the present invention. Any feature disclosed in this specification, unless otherwise stated, can be replaced by other equivalent or alternative features with similar purposes; all disclosed features, or all steps in the methods or processes, except for mutually exclusive features and / or steps, can be combined in any way.
Claims
1. A single tree segmentation method based on vehicle-mounted LiDAR point cloud, characterized in that: The following steps are involved: Step 1. Use the vehicle-mounted LiDAR to scan and obtain the original scene point cloud data, and pre-process the original point cloud data to obtain the tree point cloud; Step 2. Process the trunk point cloud based on the reflection intensity-guided DBSCAN-PCA fusion trunk extraction algorithm to extract the trunk point cloud. The specific process is as follows: Step 2.
1. Filter the crown points based on reflection intensity. Set a reflection intensity threshold, retain the point cloud with reflection intensity higher than the threshold, and filter out low-intensity points. Step 2.
2. Use the density-based clustering algorithm DBSCAN to extract tree trunk points and obtain candidate tree trunk clusters; Step 2.
3. Use the main direction index to extract the complete tree trunk: Step 2.3.
1. Use the principal component analysis (PCA) method to calculate the principal axis direction of the trunk point cloud in the trunk candidate cluster to accurately obtain the trunk growth direction of each tree. The specific process is as follows: Step 2.3.1.
1. Point cloud data preparation: Assume that the point cloud data contains N points, and the information of each point is a three-dimensional coordinate, and each point is represented by X i =[x i ,y i ,z i ] T , where i = 1, 2, ..., N; The coordinates of all points in a trunk candidate cluster are formed into a matrix X, as shown in Formula 1: Step 2.3.1.
2. Center the point cloud data X by subtracting the mean value from each coordinate. Assume that the center of mass of the point cloud is Calculate the average value of each coordinate component, as shown in formula (2): After centralizing the point cloud data, we get a new data matrix X c As shown in formula (3): Step 2.3.1.
3. Calculate the covariance matrix C. The calculation formula is shown in formula (4): Where C is a three-dimensional symmetric matrix whose elements represent the covariance relationship between different coordinate dimensions. The expanded form of the C matrix is shown in formula (5): Where Var(x), Var(y), and Var(z) are the variances of the x, y, and z coordinates after centering, respectively; Cov(x,y) is the covariance between the x and y coordinates after centering, Cov(x,z) is the covariance between the x and z coordinates after centering, and Cov(y,z) is the covariance between the y and z coordinates after centering. Step 2.3.1.
4. Perform eigenvalue decomposition on the covariance matrix C to obtain eigenvalues and eigenvectors. The eigenvalue decomposition formula is shown in formula (6): Cv i =λ i v i (6) Where λ i is the eigenvalue of the covariance matrix C, which represents the variance of the point cloud data in the direction of the corresponding eigenvector; v i is the corresponding eigenvector, indicating the main direction of the point cloud data; The result of eigenvalue decomposition is three eigenvalues λ1, λ2, λ3 and their corresponding eigenvectors v1, v2, v3; Step 2.3.1.
5. Main direction extraction: Compare the sizes of the three eigenvalues and select the eigenvector corresponding to the largest eigenvalue as the first principal component direction of the point cloud; let the largest eigenvalue be λ max , the corresponding eigenvector v max As shown in formula (7): Eigenvector v max Indicates the main direction of the point cloud, and the directions of the tree trunks are mostly consistent with this main direction; Step 2.3.1.
6. Determine whether the main direction is reasonable and obtain the cluster: In order to determine whether the main direction of the cluster is close to the vertical direction, it is necessary to calculate the main direction v max The angle with the vertical vector z = [0, 0, 1]; the angle θ is calculated as shown in formula (8): Where, v max z represents the dot product of two vectors, ||v max || represents vector v max The modulus of , that is, the length of the vector; If the angle θ is less than the set threshold, the main direction of the cluster is considered to be close to the vertical direction, and the cluster point cloud is identified as a single tree trunk point cloud. Otherwise, the cluster point cloud is put back into the original point cloud set, the cluster radius is readjusted, and step 2.2 is executed again to extract potential tree trunks. Step 2.3.
2. Using the extracted main direction as the index, a variable cone nearest neighbor search method is designed to gradually expand the trunk point cloud and identify the missed trunk points. Step 2.3.
3. Dynamically adjust the search range to ensure the continuity and integrity of the tree trunk point cloud structure; Step 3. Process the remaining point clouds to achieve accurate extraction of the crown point cloud and ultimately achieve single tree segmentation. The specific process is as follows: Step 3.
1. Trunk-constrained canopy point cloud segmentation: Using a trunk-constrained clustering method, canopy point cloud segmentation is performed based on the correctly segmented trunk point cloud. Step 3.
2. Edge canopy point cloud segmentation based on the gravity model to achieve complete single tree point cloud segmentation.
2. The single tree segmentation method based on vehicle-mounted LiDAR point cloud according to claim 1, characterized in that: The preprocessing in step 1 is to filter the original point cloud using the CSF ground filtering method, filter out the ground point cloud, retain the non-ground point cloud, and then perform denoising.
3. The single tree segmentation method based on vehicle-mounted LiDAR point cloud according to claim 1, characterized in that: The specific process of step 2.2 is: The point cloud obtained in step 2.1 is divided into core points, reachable points and noise points. The high-density area is divided according to the set neighborhood radius eps and the minimum number of neighbor points min_samples and independent clusters are identified as candidate tree trunk clusters. Among them, a core point refers to a point cloud that contains at least min_samples neighboring points within its eps neighborhood radius. All core points are connected to each other to form a cluster, and non-core points that are only within the neighborhood of the core point are reachable points. The reachable points are also classified into the corresponding clusters. The remaining isolated points cannot be connected to other points through density and are regarded as noise points.
4. The single tree segmentation method based on vehicle-mounted LiDAR point cloud according to claim 1, characterized in that: The specific process of step 2.3.2 is as follows: Step 2.3.2.
1. Project each cluster along a cross section perpendicular to the principal direction. Then, perform a least squares fit of a circle to all projected points. The center of the circle is the midpoint of the trunk centerline. The line passing through this point in the same direction as the principal direction is the trunk centerline. Step 2.3.2.
2. Construct a variable cone search area based on the centerline of the tree trunk. The centerline of the tree trunk is the center of the circle obtained in step 2.3.2.1, and the direction is the same as the trunk direction vector: The initial search radius is set as the fitting radius of the trunk base, and a linear attenuation model is used. The search radius r(z) gradually decreases with increasing height. The radius calculation formula is shown in formula (9): r(z)=r0-α(zz min ) (9) Where r0 is the trunk radius of the lowest height layer, α is the adjustable search radius attenuation rate, which determines the cone contraction rate, and z min is the lowest altitude, z is the current search altitude; Based on the radius calculation formula, the centerline of the trunk is used as the search axis, and the height is divided into equal intervals from z_min to z_max. A search area is constructed at each height layer, and the KD-Tree is used to query the nearest neighbor points to extract all the trunk point clouds. When r(z) is 0, the search ends and all the searched points are returned, which is the complete single tree trunk point cloud.
5. The single tree segmentation method based on vehicle-mounted LiDAR point cloud according to claim 1, characterized in that: The specific process of step 3.1 is as follows: Step 3.1.
1. Modeling the hypothetical trunk axis: Use PCA to extract the main direction of the trunk point cloud. With the lowest point of the trunk as the center starting point, construct a hypothetical trunk axis along the main direction to simulate the vertical extension trajectory of the trunk. The axis is from the bottom center point P base and the main direction vector v max Definition, extending the spatial constraints of tree trunks to full canopy height; Step 3.1.
2. Calculate canopy point attribution: For each unlabeled point P except the trunk point cloud i , calculate its vertical distance to the main axis of all tree trunks, the formula is: The minimum distance is obtained by vector cross product and modulus ratio, and combined with the height range filtering condition, the point is assigned to the tree trunk with the closest distance and matching height.
6. The single tree segmentation method based on vehicle-mounted LiDAR point cloud according to claim 1, characterized in that: The specific process of step 3.2 is as follows: Step 3.2.
1. Edge point screening: According to each unlabeled point P except the trunk point cloud i The vertical distance to the main axis of all tree trunks is used to filter out edge points that are far from the trunks; Step 3.2.
2. Spatial Cylinder Construction and Neighborhood Query: With each edge point as the center, construct a cylinder search range with radius r0 and height h0. Use KD-Tree to quickly query neighboring points within the cylinder and obtain the index of the neighboring points. Step 3.2.
3. Attraction calculation model: Assuming that the attraction of correctly classified points in the neighborhood to unlabeled points is inversely proportional to the square of the distance, the total attraction value is calculated by formula (11): Among them, d ij is an unmarked point P i To the neighboring point P j The Euclidean distance is a small constant. By comparing the total gravity values of different tree trunk instances, the unclassified points are assigned to the single tree label corresponding to the one with the largest gravity. Step 3.2.
4. Dynamic search range adjustment: If the number of neighborhood points in the current cylinder is lower than the set threshold, gradually expand the search radius r and height h until enough reference points are obtained.
7. A parameter estimation method based on single tree segmentation results, characterized in that: The following steps are involved: S1. Perform voxel segmentation on the single-tree segmentation result obtained by the single-tree segmentation method according to any one of claims 1-6 to obtain single-tree structural parameters and point cloud structural parameters; S2. Use the Pearson correlation analysis method to calculate the correlation between the dependent variable and the independent variable, and select the input variable with the highest characteristic value; The calculation formula is shown in formula (12): Where, Corr(νar,dep) is the Pearson correlation coefficient between the dependent variable νar and the independent variable dep, M is the total number of samples, ζ νar is the sample standard deviation of νar, ζ dep is the sample standard deviation of dep; νar′ is the mean value of νar, dep′ is the mean value of dep; The dependent variables include aboveground biomass (AGB), carbon storage (CS), leaf area index (LAI), and three-dimensional green volume (LVV); the independent variables include 27 point cloud structure parameters and single tree structure parameters, which include three measurement parameters. S3. For different tree species, an improved random forest regression algorithm is used, with the tree species' individual tree structure parameters and point cloud structure parameters as input to obtain the biological estimation parameters of the tree species: S3.
1. Use the least absolute shrinkage and selection operators for feature selection to remove variables that have little impact on the results and achieve adaptive feature selection. Lasso regression is an L1 regularization method that compresses the regression coefficients so that the coefficients of some features shrink to zero, thereby achieving automatic feature selection. Its optimization objective function is defined as shown in formula (13): Among them, M is the number of samples, p is the number of features, and y k is the true value of the biological parameter of a single tree, x k is the input point cloud feature vector, w is the regression coefficient, λ reg is the regularization hyperparameter, and the superscript T indicates the rank transformation; S3.
2. Introduce the pigeon swarm optimization algorithm for hyperparameter tuning, use the swarm intelligence search mechanism to optimize the hyperparameter configuration of random forest regression, and improve the stability of biological parameter estimation.
8. The parameter estimation method according to claim 7, wherein: The larger the absolute value of the correlation coefficient in S2, the stronger the correlation. The interval [0, 0.1] indicates no correlation, the interval [0.1, 0.3] indicates weak correlation, the interval [0.3, 0.5] indicates moderate correlation, and the interval [0.5, 1] indicates strong correlation.
9. The parameter estimation method according to claim 7, wherein: The point cloud structure parameters in S1 include crown void ratio parameter, height statistics parameter, height quantile parameter and height density parameter; The calculation process of the crown void ratio is as follows: first, the point cloud data is spatially divided according to a fixed voxel size, then the number of points in each voxel is counted, and low-density voxels are defined; finally, the proportion of all low-density voxels is calculated, which is the crown void ratio. The number of low-density voxels N is empty The calculation formula is shown in formula (14): Where μ is the voxel number in space, For the number of points, N voxels is the total number of voxels. The calculation formula of crown void ratio is shown in formula (15): Height statistical parameters are statistical features used to describe the height distribution of tree point clouds, reflecting the overall shape, variability, symmetry, and growth status of trees in three-dimensional space. Height statistical parameters include maximum height, average height, height standard deviation, height skewness, height kurtosis, and height median. The lowest point height of a single tree point cloud is 0, and the highest height is the highest value of the z coordinate in the single tree point cloud; The average height is the mean value of the height of all point clouds, which reflects the overall distribution trend of the point cloud height. Its calculation formula is shown in formula (16): The height standard deviation is used to measure the degree of discreteness of the point cloud height and represents the uniformity of the tree canopy. Its calculation formula is shown in formula (17): The height skewness reflects the symmetry of the height distribution. If the skewness is positive, it means that the height distribution is right-skewed, that is, there are more high points; if the skewness is negative, it means that there are more low points. Its calculation formula is shown in formula (18): The height kurtosis describes the sharpness of the height distribution. A higher kurtosis means that most points are concentrated near the mean, while a lower kurtosis means that the data is flatter. Its calculation formula is shown in formula (19): The median height represents the median value of the point cloud height, which is less affected by outliers than the mean. The calculation formula is shown in formula (20): Z median =median(z i ) (20) The height quantile parameter is used to describe the height of different percentage positions and is used to analyze the canopy thickness and growth layer of trees. For a certain height quantile h, the quantile calculation formula is shown in formula (21): P h =percentile(z,h) (21) where h∈{10,20,30,...,90,99}. The height density parameter is used to measure the point cloud density in different height intervals, so as to analyze the hierarchical structure and vertical distribution pattern of trees. First, the height range of the entire tree is divided into 10 equally spaced intervals, as shown in formula (22): Then, calculate the proportion of point clouds in each height interval D h , the calculation formula is shown in formula (23):
Citation Information
Cited By
Airborne laser radar point cloud forest biomass estimation method and device, medium and product
CN121348276A
Single tree point cloud segmentation method and system for gradient closed forest
CN121685563A
Landscape space quality evaluation method and device, terminal and medium
CN121708308A