A canopy-scale urban green vegetation classification method based on remote sensing

By fusing multi-source high-resolution remote sensing data and improving the random forest model, a canopy-scale multidimensional feature set was constructed, which solved the problems of insufficient feature extraction and inefficient sample labeling in traditional methods, and achieved high-precision and adaptive urban green space vegetation classification.

CN120953684BActive Publication Date: 2026-02-10CHINA INST OF WATER RESOURCES & HYDROPOWER RES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511076646.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2026-02-10
Estimated Expiration
2045-08-01

Smart Images

  • Figure CN120953684B_ABST
    Figure CN120953684B_ABST
Patent Text Reader

Abstract

The application discloses a kind of crown layer scale urban green vegetation classification method based on remote sensing, comprising: multi-source remote sensing data acquisition and fusion;Fusion multi-source remote sensing data's crown layer scale multidimensional feature set construction;Urban green vegetation enhanced sample set construction;The construction of low-dimensional feature set after optimization;And the crown layer scale automatic classification method of urban green based on improved random forest.The application constructs multidimensional feature set by multi-source high-resolution remote sensing data;Introduce the automation dynamic feature optimization and weight distribution mechanism based on mRMR, effectively reduce redundancy and highlight discriminant feature;Combining multi-source priori knowledge constructs heterogeneity driven adaptive sample set;And based on the improved random forest algorithm of introducing dynamic weighted node splitting, neighborhood constraint and adaptive class balance, realize the high-precision, crown layer scale, adaptive classification of urban green vegetation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of urban ecological monitoring and intelligent processing of remote sensing images, specifically to a method for refined classification of urban green space vegetation at the canopy scale based on multi-source high-resolution remote sensing data fusion and dynamic feature optimization, aiming to achieve high-precision, adaptive, canopy structure-sensing classification of urban green space vegetation. Background Technology

[0002] With the acceleration of global urbanization, urban green spaces, as an important component of the urban ecosystem, play a crucial role in improving urban ecological quality, addressing climate change, and protecting residents' health through refined monitoring and management. Traditional urban green space classification methods, especially in high-resolution remote sensing image applications, face numerous challenges: limitations of single data sources: traditional methods rely on single optical images, acquiring only two-dimensional spectral and texture information, making it difficult to effectively distinguish subtle differences between different vegetation types at the canopy scale; deficiencies in feature extraction and representation: spectral feature-based methods are susceptible to interference from "same object, different spectrum" and "different object, same spectrum" phenomena, while texture feature-based methods are easily affected by local noise and background interference in high-resolution images; existing feature sets often ignore the three-dimensional structure and spatiotemporal dynamics of urban green space vegetation caused by mixed planting, seasonal changes, and other factors; insufficient potential for multi-source data synergy: using LiDAR alone cannot fully exploit the spectral-structural correlation characteristics of vegetation; simple multi-source data overlay or concatenation processing is insufficient to achieve deep data fusion and synergistic enhancement; insufficient feature selection and model generalization capabilities: existing feature selection methods lack automated and dynamic screening mechanisms for specific classification tasks, resulting in low efficiency; low sample annotation efficiency and quality, making it difficult to capture the spatial heterogeneity of urban green space vegetation, thus limiting the performance of supervised classification models.

[0003] Therefore, there is an urgent need for an urban green space vegetation classification method that can deeply integrate multi-source high-resolution remote sensing data, construct a canopy-scale multi-dimensional feature set, combine multi-source prior knowledge, adopt an automated dynamic feature optimization mechanism, and is based on an improved classification model. This method would enable high-precision, adaptive, and canopy-structure-aware urban green space vegetation classification, providing more reliable and refined data support for urban ecological planning, green space management, and biodiversity assessment. Summary of the Invention

[0004] This invention addresses the shortcomings of existing urban green space vegetation classification methods in terms of feature extraction, feature selection, and model construction, and proposes a canopy-scale urban green space vegetation classification method based on remote sensing.

[0005] The specific solution of the present invention is as follows:

[0006] A remote sensing-based canopy-scale urban green space vegetation classification method includes the following steps:

[0007] Step 1: Acquisition and Fusion of Multi-Source Remote Sensing Data: Acquire multi-source remote sensing data from the same region, multiple time phases, and consistent spatiotemporal references; preprocess the multi-source remote sensing data; and perform spatial registration and fusion to obtain a fused data cube.

[0008] Step 2: Construction of a Canopy-Scale Multidimensional Feature Set from Multi-Source Remote Sensing Data: From the data cube fused in Step 1, extract five categories of multidimensional features: spectral features, time-series dynamic features, three-dimensional structural features, contextual spatial structural features, and multi-source fusion composite features to construct a structured multidimensional feature set D.

[0009]

[0010] Among them, F R F represents the regional feature vector of urban green space areas, where the numerical subscript indicates the ordinal number of the urban green space area, and M is the total number of urban green space areas; R The feature of a certain dimension in the data is represented as F. i , where i represents the dimension number;

[0011] Step 3: Construction of urban green space vegetation enhancement sample set: The spatial distribution probability map of vegetation type is obtained by fusing multi-source prior data. Based on the spatial distribution probability map and green space functional zoning, the multi-dimensional feature set constructed in Step 2 is used as the basic data for adaptive sampling. The samples are then verified and enhanced by human-computer collaboration to obtain the urban green space vegetation enhancement sample set. Each sample includes a multi-dimensional feature vector and the corresponding category label.

[0012] Step 4: Construction of the optimized low-dimensional feature set: For the multi-dimensional feature set constructed in Step 2, the feature subset is first screened based on the maximum correlation minimum redundancy (mRMR) algorithm. Then, the variance of each feature in the enhanced sample set obtained in Step 3 is calculated and subjected to maximum and minimum normalization to obtain the normalized variance as the feature weight. Then, the dimensionality is reduced by principal component analysis (PCA) to generate the low-dimensional feature set D'.

[0013] Step 5: Automatic Green Space Classification: The optimized low-dimensional feature set D' constructed in Step 4 and the urban green space vegetation enhancement sample set constructed in Step 3 are input into the improved random forest model to achieve automatic classification of urban green spaces at the canopy scale. When constructing the decision tree, the improved random forest algorithm uses the weighted Gini coefficient based on the feature weights calculated in Step 4 as the node splitting criterion, and performs spatial consistency optimization based on neighborhood constraints on the classification results. The formula for calculating the weighted Gini coefficient is:

[0014]

[0015] Where S is the sample set of the current node, Values(Fi ) is a feature F i The set of possible values ​​for |S v | represents the number of samples with feature value v, |S| represents the total number of samples in the current node, Classes represents the set of all class labels, and c represents the specific class label. v,c | for feature F i The number of samples with value v and category c, w i The feature F calculated in step four i The weight.

[0016] Further optimization scheme, in step one, the multi-source high-resolution remote sensing data includes: optical multispectral data: spatial resolution better than 5 meters; optical hyperspectral data: spatial resolution better than 30 meters; laser point cloud data: point density not less than 10 points / square meter; the spatial registration and fusion: resample all data to a unified spatial resolution, perform spatial registration, and stack the effective bands and derivative products of different data sources to construct a multi-channel, multi-modal, spatiotemporally consistent fused data cube.

[0017] Furthermore, in step two, the spectral features include: Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), Red Edge Index (REI), and Chlorophyll Absorption Valley Depth (CAI); the time-series dynamic features include: Seasonal Variation Index (SVI), Phenological Phase Features (PPF), and Growing Season Cumulative Greenness (GAG); the three-dimensional structure includes: Canopy Volumetric Density (CVD), Vertical Profile Heterogeneity (VPH), Canopy Height Stratification (CHF), Reflectance Gradient (RIG), and Canopy Surface Roughness (CSR); the contextual spatial structure features include: Spatial Autocorrelation Index (MorI) and Neighborhood Texture Diversity (NTD based on the standard deviation of entropy); and the multi-source fusion composite features include: Spectral-Height Joint Index (SHJI) and Reflectance Intensity-Texture Synergy Features (RITC).

[0018] Furthermore, in step two, for each candidate urban green space area R... k The spectral features, time-series dynamic features, three-dimensional structural features, contextual spatial features, and multi-source fusion composite features of urban green space areas were extracted to construct regional-level features:

[0019]

[0020] Wherein, N(R) k () represents the total number of pixels contained in the k-th urban green space area. Indicates the region R k Summing all pixels j within f jLet be the single-pixel feature vector of the j-th pixel, which ultimately forms the structured feature set D.

[0021] Furthermore, step three specifically includes the following operations:

[0022] 3-1 Multi-source prior data fusion and spatial mapping:

[0023] Collect vector data of green space functional zoning in the study area and existing tree / grass species planting distribution databases or survey data; use kernel density estimation (KDE) to generate spatial distribution probability maps of major tree / grass species; overlay the probability maps with green space functional zoning to identify high-probability distribution areas of specific vegetation types in different functional zones;

[0024] 3-2 Heterogeneity-Driven Adaptive Sampling Method:

[0025] Based on the spatial distribution probability map of vegetation types and green space functional zoning obtained by integrating prior knowledge, a heterogeneity-driven adaptive sampling strategy is designed: in areas with high probability distribution or complex vegetation types, dense grid sampling or object-based sampling is adopted; in areas with low probability distribution or simple vegetation types, sparse grid sampling is adopted; the grid spacing or sampling density is dynamically correlated with the spatial heterogeneity of vegetation types or the probability of specific types in the area.

[0026] 3-3 Human-Machine Collaborative Sample Validation and Enhancement:

[0027] Candidate sample points or regions obtained through adaptive sampling are combined with high-resolution remote sensing imagery and ground survey data for human-machine collaborative verification and correction to ensure the accuracy of the samples. The vegetation type represented by each sample point or region is determined and assigned a corresponding category label. For categories with a small number of samples or uneven distribution, sample augmentation techniques are used to generate synthetic samples to solve the class imbalance problem.

[0028] Multi-source prior knowledge (such as green space functional zoning and tree / grass species planting distribution databases) is used to guide sample selection, aiming to improve sample representativeness and labeling efficiency, especially for complex and highly heterogeneous urban green space vegetation types. This step ultimately yields an urban green space vegetation sample set with precise category labels, which will be used for subsequent training of the improved random forest model.

[0029] Furthermore, in step four, the feature subset selection based on the maximum relevance minimum redundancy (mRMR) algorithm includes the following steps:

[0030] 4-1 Calculate mutual information:

[0031] First, calculate F for each feature dimension. iMutual information I(F) between the target category label Y of the urban green space vegetation enhancement sample obtained in step three and the target category label Y of the target category label Y. i ;Y);Mutual Information I(F) i The higher the value of F(Y), the higher the feature dimension F. i The more information it contains that can distinguish different vegetation types, the greater the correlation.

[0032] Secondly, for any two different feature dimensions F i and F j Calculate the mutual information I(F) between them. i ;F j Mutual information I(F) i ;F j The higher the value of ), the more common information these two feature dimensions contain, i.e., the higher the redundancy.

[0033] 4-2mRMR criterion for sorting:

[0034] The mRMR criterion balances feature relevance and redundancy using the following formula and ranks all candidate features:

[0035]

[0036] Where S is the selected feature subset, and |S| is the number of selected features; the ranking results are arranged from high to low according to the mRMR value, and the features ranked higher are considered more important;

[0037] 4-3 Feature Subset Selection:

[0038] Based on a preset threshold or the number of features, select the top-ranked feature subset S from the sorting results. selected .

[0039] Furthermore, in step four, the specific operation for obtaining feature weights is as follows:

[0040] For the selected feature subset S selected Calculate F for each feature i The variance Var(F) in the augmented sample set i ), and perform max-min normalization to obtain the normalized variance Var′(F i ):

[0041]

[0042] Use the normalized variance as the feature weight:

[0043] w i =Var′(F i )

[0044] wi It is the weight of the i-th feature, which will be used for dynamic weighted node splitting in the subsequent random forest classification model.

[0045] Furthermore, step five also includes the following operations: Spatial consistency optimization of neighborhood constraints: In the final voting stage of the random forest, spatial neighborhood information is introduced to correct the classification results. For each pixel, the class distribution of the random forest voting results within its preset neighborhood is statistically analyzed. If the initial classification result of the center pixel is inconsistent with the voting results of the majority of pixels in its neighborhood, the class of the center pixel is corrected to the class of the majority of pixels in its neighborhood. The calculation formula is as follows:

[0046] Class(i) final =mode({Class(j)} initial |j∈N(i)})

[0047] Where, Class(i) final For the final classification result of pixel i, mode({…}) takes the element that appears most frequently in the set (the mode), and Class(j) initial Let N(i) be the initial classification result (random forest voting result) for pixel j, and let N(i) represent the neighborhood of pixel i.

[0048] The improvements of this invention are mainly in the following two aspects:

[0049] 1. Constructing a canopy-scale multi-source feature set for complex urban green space scenarios to enhance ground feature representation capabilities. Addressing the issues of multi-layered structure and spectral obfuscation in urban green space vegetation, this invention deeply integrates multispectral, hyperspectral, and LiDAR data to construct a canopy-scale multi-dimensional feature set that includes spectral, time-series, three-dimensional structure, contextual space, and multi-source fusion composite information. This set can more comprehensively and precisely characterize the complex features of urban green space vegetation.

[0050] 2. This paper proposes an efficient classification method based on mRMR feature selection and variance weighting to improve the model's classification accuracy and generalization ability. The invention uses the mRMR algorithm to select a subset of features with high relevance to the target class and low redundancy, and innovatively utilizes feature variance for weight allocation, avoiding dependence on model training results. Based on this, by constructing a heterogeneity-driven augmented sample set and employing a random forest classifier with a weighted Gini coefficient, high-precision and robust classification of urban green space vegetation is achieved.

[0051] The beneficial effects of this invention are:

[0052] This invention innovatively integrates multi-source high-resolution remote sensing data, including multispectral, hyperspectral, and LiDAR data, to construct a multi-dimensional feature set encompassing spectral, time-series, refined 3D structure, contextual space, and multi-source fusion. It introduces an automated feature selection mechanism based on mRMR and a variance-based dynamic feature weight allocation mechanism to effectively reduce redundancy and highlight discriminative features. A heterogeneity-driven adaptive sample set is constructed by combining multi-source prior knowledge. Based on an improved random forest algorithm incorporating dynamic weighted node splitting, neighborhood constraints, and adaptive class balancing, high-precision, canopy-scale, and adaptive classification of urban green space vegetation is achieved. Specific advantages are as follows:

[0053] Deep fusion of multi-source data and feature complementarity: It innovatively integrates multiple high-resolution remote sensing data such as multispectral, hyperspectral, and LiDAR, and constructs a multidimensional feature set at the canopy scale covering 16 dimensions. It makes full use of the complementary information of different data sources in terms of spectrum, structure, temperature, and time series, breaking through the limitations of single data source or simple data superposition, and significantly improving the ability to express the complex characteristics of urban green space vegetation.

[0054] Refined feature extraction at the canopy scale: Targeting the multi-layered structure of urban green space vegetation, three-dimensional structural features with clear physical significance, such as canopy volumetric density (CVD), vertical profile heterogeneity (VPH), and canopy height stratification (CHF), were extracted from LiDAR point clouds. Combined with physiological and biochemical features such as hyperspectral red edge index (REI) and chlorophyll absorption valley depth (CAI), as well as time series features such as seasonal variation index (SVI) and phenological phase characteristics (PPF) from multi-temporal optical images, a refined characterization of vegetation canopy structure, physiological state, and phenological dynamics was achieved.

[0055] Automated dynamic feature selection and weight allocation: By adopting an automated feature selection mechanism based on mRMR and a dynamic feature weight allocation mechanism based on variance, the optimal feature subset can be adaptively selected according to data characteristics and classification tasks, and its contribution in the classification model can be adjusted. This reduces reliance on human experience, effectively reduces feature redundancy, and improves the efficiency and robustness of the model.

[0056] Multi-source prior knowledge-driven adaptive sample construction: Innovatively integrates multi-source prior knowledge such as green space functional zoning and tree / grass species planting distribution databases, and constructs a probability map based on kernel density estimation to guide heterogeneity-driven adaptive sampling. This achieves intelligent, efficient, and spatially representative layout of sample points, significantly improving the efficiency and quality of sample labeling, especially improving sample coverage for rare or mixed vegetation types.

[0057] Improved robustness and accuracy of random forest classification model: Based on traditional random forest, a node splitting criterion based on dynamic feature weights, a classification result correction mechanism considering spatial neighborhood information, and an adaptive class balance sampling strategy are introduced. This effectively improves the model's adaptability to complex urban environments, reduces "salt-and-pepper noise", solves the class imbalance problem, and significantly improves the classification accuracy and robustness of urban green space vegetation at the canopy scale.

[0058] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Attached Figure Description

[0059] Figure 1 Results of spectral feature extraction for urban green spaces;

[0060] Figure 2 Extracting dynamic features of urban green space over time series;

[0061] Figure 3 LiDAR 3D structural feature extraction for urban green spaces;

[0062] Figure 4 Extraction of spatial structural features of urban green space context;

[0063] Figure 5 Extraction of multi-source integrated composite features of urban green spaces;

[0064] Figure 6 Classification results of urban green space vegetation at the canopy scale in typical areas of Beijing (I);

[0065] Figure 7 The results of the canopy-scale urban green space vegetation classification in typical areas of Beijing (II). Detailed Implementation

[0066] Example 1

[0067] A remote sensing-based canopy-scale urban green space vegetation classification method includes the following steps:

[0068] Step 1: Multi-source remote sensing data acquisition and fusion: Acquire multi-source remote sensing data from the same region, multiple time phases, and consistent spatiotemporal references; preprocess the multi-source remote sensing data; and perform spatial registration and fusion to obtain a fused data cube. This embodiment specifically includes the following steps:

[0069] (1) Data Acquisition

[0070] Based on the needs of vegetation classification in the study area and urban green spaces, multi-source high-resolution remote sensing data of the same area, multiple time phases, and consistent spatiotemporal references were acquired, including:

[0071] 1) Optical multispectral data: Acquire multispectral images with a spatial resolution better than 5 meters (e.g., Sentinel-2, Landsat-8, Gaofen series, etc.), covering visible light, near infrared, and short-wave infrared bands, for extracting spectral and time series features.

[0072] 2) Optical hyperspectral data: Acquire hyperspectral images with a spatial resolution better than 30 meters (e.g., Hyperion, GF-5, etc.), with more than 200 narrow bands for extracting fine spectral features, especially vegetation red edges and chlorophyll absorption bands.

[0073] 3) Laser point cloud data: Point cloud data is acquired through airborne or ground-based LiDAR, with a point density of no less than 10 points / square meter. It contains accurate three-dimensional coordinates (X,Y,Z) and reflection intensity information, and is used to extract three-dimensional structural features of vegetation.

[0074] (2) Data Preprocessing and Fusion

[0075] The acquired multi-source data undergoes rigorous preprocessing, spatial registration, and fusion.

[0076] 1) Optical data preprocessing: Perform radiometric correction and atmospheric correction (such as FLAASH, 6S model) to eliminate atmospheric effects, perform geometric fine correction and precise registration between multi-temporal images (sub-pixel level) to generate surface reflectance products.

[0077] 2) LiDAR point cloud data processing: Point cloud filtering (such as cloth simulation filtering CSF) is performed to separate ground points and non-ground points, generating a digital elevation model (DEM) and a digital surface model (DSM). The normalized digital surface model (nDSM) is obtained by subtracting the DEM from the DSM, reflecting the height of ground features. The original reflectance intensity values ​​are corrected for distance and incident angle.

[0078] 3) Hyperspectral data preprocessing: Radiometric calibration and atmospheric correction (such as ACORN model) are performed to eliminate the effects of water vapor absorption, and geometric fine correction and multispectral data registration are performed.

[0079] 4) Multi-source data spatial fusion: All data are resampled to a uniform spatial resolution (e.g., 2 meters or 5 meters) and accurately spatially registered. Effective bands and derivative products (such as nDSM, LST) from different data sources are stacked to construct a multi-channel, multi-modal, spatiotemporally consistent fused data cube.

[0080] Step 2: Construction of a Canopy-Scale Multidimensional Feature Set from Multi-Source Remote Sensing Data: From the data cube fused in Step 1, extract five categories of features across multiple dimensions: spectral features, time-series dynamic features, three-dimensional structural features, contextual spatial structure, and multi-source fusion composite features to construct a structured multidimensional feature set D.

[0081]

[0082] Among them, F R The region-level feature vector represents urban green space areas, with its numerical subscript indicating the ordinal number of the urban green space area; M is the total number of urban green space areas; F R The feature of a certain dimension in the data is represented as F. i , where i represents the dimension number;

[0083] For each candidate urban green space area R k The spectral features, time-series dynamic features, three-dimensional structural features, contextual spatial structure, and multi-source fusion composite features of urban green space areas were extracted to construct regional-level features:

[0084]

[0085] Wherein, N(R) k () represents the total number of pixels contained in the k-th urban green space area. Indicates the region R k Summing all pixels j within f j Let be the single-pixel feature vector of the j-th pixel, which ultimately forms the structured feature set D.

[0086] The spectral features include: Normalized Difference Vegetation Index (NDVI), Enhanced Vegetation Index (EVI), Red Edge Index (REI), and Chlorophyll Absorption Valley Depth (SAF); the time-series dynamic features include: Seasonal Variation Index (SVI), Phenological Phase Features (PPF), and Growing Season Cumulative Greenness (GAG); the LiDAR three-dimensional structural features include: Canopy Volumetric Density (CVD), Vertical Profile Heterogeneity (VPH), Canopy Height Stratification (CHF), Reflectance Gradient (RIG), and Canopy Surface Roughness (CSR); the contextual spatial structural features include: Spatial Autocorrelation Index (MorI) and Neighborhood Texture Diversity (NTD based on entropy standard deviation); the multi-source fusion composite features include: Spectral-Height Joint Index (SHJI) and Reflectance Intensity-Texture Synergy Feature (RITC). See details. Figures 1-5 .

[0087] (1) Spectral characteristics

[0088] Based on multispectral and hyperspectral data, spectral features reflecting the physiological state and components of vegetation are extracted.

[0089] The spectral vector of each remote sensing image pixel (x,y) is represented as:

[0090]

[0091] To enhance discriminative power, the spectral characteristics of normalized difference vegetation NDVI(x,y), enhanced vegetation EVI(x,y), red-edge REI(x,y), and chlorophyll absorption valley depth SAF(x,y) were calculated respectively:

[0092]

[0093] Among them, I NIR (x,y) is the near-infrared band value of pixel (x,y), I RED (x,y) are the red band values, I BLUE (x,y) are the blue band values, I 750 (x,y) are the spectral values ​​at a wavelength of 750nm, I 700 (x,y) are spectral values ​​at a wavelength of 700nm, I 680 (x,y) are spectral values ​​at a wavelength of 680nm, I 550 (x,y) are spectral values ​​at a wavelength of 550nm, I 720 (x,y) is the spectral value with a wavelength of 720nm.

[0094] (2) Dynamic characteristics of time series

[0095] To compensate for the shortcomings of traditional static spectral features, dynamic features of time series are extracted by utilizing the phenological differences in different vegetation growth processes, and the seasonal variation index (SVI), phenological phase feature (PPF), and growing season cumulative greenness (GAG) are calculated.

[0096]

[0097] Where, σ NDVI μ represents the standard deviation of NDVI values ​​within a given season. NDVI It is the average NDVI value within that season.

[0098]

[0099] Where f represents the Fourier transform. NDVI represents the phase angle. series This represents an NDVI sequence at consecutive time points.

[0100]

[0101] in, This indicates the summation of all items from the first time point (t=1) to the last time point (T). Here, T represents the total number of time points or the total number of time steps in the remote sensing image observations during the growing season. NDVI tThis represents the value of the Normalized Difference Vegetation Index (NDVI) at time point t. A higher NDVI value indicates more abundant vegetation. t is the time point, and Δt represents the time interval, indicating the NDVI value at each point in the time series. t The value represents the length of the time period, or the time interval between two adjacent NDVI observations.

[0102] Based on multi-temporal NDVI calculations, the cumulative greenness during the vegetation growing season can be used to distinguish between fast-growing herbaceous plants and slow-growing woody plants.

[0103] (3) Three-dimensional structural features of LiDAR

[0104] Urban green spaces are characterized by a multi-layered structure with mixed planting. By analyzing the spatial distribution and reflection characteristics of point cloud data, the vertical and horizontal structural information of the vegetation canopy is quantified, and the characteristics of canopy volume density (CVD), vertical profile heterogeneity (VPH), canopy height stratification (CHF), reflection intensity gradient (RIG), and canopy surface roughness (CSR) are calculated, providing key criteria for the classification of urban green space vegetation.

[0105] Canopy volume density reflects the density of canopy branches and leaves. The calculation method is as follows: First, the vegetation canopy point cloud is divided into a regular three-dimensional voxel grid, and the number of point clouds N in each voxel is counted. voxel Then calculate the point density per unit volume, i.e., the CVD value, using the following formula:

[0106]

[0107] Where, N voxel V is the number of LiDAR point clouds contained within a single voxel. voxel It represents the volume of a voxel; the default voxel is a cube with a side length of 1m.

[0108] Vertical profile heterogeneity reflects the complexity of the canopy's vertical structure. The calculation method is as follows: first, the canopy point cloud is layered by height, and the mean reflection intensity I of each layer is calculated. zk Then, the standard deviation σ and mean μ of the reflection intensity of all layers are statistically analyzed, and the VPH value is calculated using the following formula:

[0109]

[0110] Among them, I z1 I z2 , ...I zn It is a data sequence representing the average reflectance intensity calculated for each layer after the canopy point cloud is layered by height. zk(where k ranges from 1 to n) represents the mean reflection intensity within the kth height layer, n represents the total number of height layers into which the canopy point cloud is divided, σ is the standard deviation function, which calculates the standard deviation of the data set listed in parentheses, and μ is the mean function, which calculates the mean of the data set listed in parentheses.

[0111] Canopy height stratification reflects the height characteristics of the canopy, and the calculation formula is:

[0112]

[0113] Where, N k N is the number of point clouds in the k-th altitude layer. total It is the total number of canopy point clouds. The vertical layering interval needs to match the vegetation type. Based on the experience of lawns, low shrubs, and tall trees, it can be divided into three layers: 0-5m, 5-10m, and >10m. The number of layers can also be adjusted according to the actual ground conditions.

[0114] The reflection intensity gradient reflects the differences in the surface material of the canopy, and the calculation formula is as follows:

[0115]

[0116] Where σ(I) is the standard deviation of the canopy point cloud reflectance intensity, and μ(I) is the mean of the canopy point cloud reflectance intensity.

[0117] The formula for calculating the surface roughness of the canopy is: CSR

[0118]

[0119] Among them, z i It is the height value of a single point cloud. is the average height of the canopy point cloud, and N is the total number of point clouds.

[0120] (4) Context space structure features

[0121] Spatial autocorrelation index quantifies the spatial clustering of vegetation types, avoiding misclassification of isolated pixels. For example, trees in urban parks may exhibit high spatial autocorrelation, while scattered shrubs may show low autocorrelation. The calculation formula is:

[0122]

[0123] Where N is the total number of pixels in the remote sensing image; w ij It is a spatial weight matrix, representing the spatial adjacency relationship between pixel i and pixel j, x i x is the feature value of pixel i. j These are feature values ​​of pixel j, such as NDVI, spectral value, or vegetation type probability. It is the average value of all unit eigenvalues.

[0124] Neighborhood texture diversity reflects the spatial heterogeneity of vegetation distribution. The standard deviation σ of the entropy value E within a local region is calculated based on a 3×3 or 5×5 sliding window. The calculation formula is as follows:

[0125] NTD=σ(E window1 E window2 ,…,E windown )

[0126] The entropy value E of a single window is calculated based on the gray-level co-occurrence matrix (GLCM), using the following formula:

[0127]

[0128] P(i,j) is the joint probability of gray values ​​i and j in GLCM.

[0129] (5) Multi-source fusion composite characteristics

[0130] By innovatively combining the advantages of different data sources, a composite feature that better reflects the comprehensive characteristics of vegetation can be constructed.

[0131] To combine vegetation greenness and vertical structure information, a spectral-height joint index was constructed based on NDVI extracted from optical images and canopy height CH extracted from LiDAR point clouds. The difference in height dimensions between different regions was eliminated using max(CH) to ensure feature comparability. The calculation formula is as follows:

[0132]

[0133] To enhance the ability to distinguish the canopy surface, LiDAR reflectance intensity I is fused. LiDAR With optical image texture energy E GLCM Construct reflection intensity-texture co-features, and calculate them using the following formula:

[0134]

[0135] Where, μ(I LiDAR σ(I) is the mean reflectance intensity of the LiDAR point cloud, reflecting the reflectance characteristics of the canopy surface. LiDAR ) is the standard deviation of the reflectance intensity of LiDAR point clouds, which characterizes the heterogeneity of canopy reflectance.

[0136] Step 3: Construction of Urban Green Space Vegetation Enhancement Sample Set: The spatial distribution probability map of vegetation type is obtained by fusing multi-source prior data. Based on the spatial distribution probability map and green space functional zoning, the multi-dimensional feature set constructed in Step 2 is used as the basic data for adaptive sampling. The samples are then verified and enhanced by human-machine collaboration to obtain the urban green space vegetation enhancement sample set. Each sample includes a multi-dimensional feature vector and its corresponding precise label.

[0137] To improve the efficiency and quality of sample annotation, especially to address the challenges of complex and unevenly distributed vegetation types in urban green spaces, multi-source prior knowledge is integrated to guide sample collection. The specific operation of this embodiment is as follows:

[0138] (1) Multi-source prior data fusion and spatial mapping

[0139] Collect vector data of green space functional zoning in the study area (such as parks, street green belts, residential green spaces, etc.) and existing tree / grass species planting distribution databases or survey data.

[0140] Based on a planting distribution database, kernel density estimation (KDE) was used to generate spatial distribution probability maps of major tree / grass species. These probability maps were then overlaid with green space functional zoning to identify high-probability distribution areas of specific vegetation types within different functional zones.

[0141] Based on the above data, kernel density estimation (KDE) is used to generate tree / grass species distribution probability maps. The calculation formula is as follows:

[0142]

[0143] Where (x,y) represents the coordinates of the position, P c (x, y) represents the probability of planting a certain tree or grass species c at position (x, y), where (x, y) is the probability of planting a certain tree or grass species c at position (x, y). i ,y i ) represents the location coordinates of the i-th sample point, n is the number of sample points, h is the bandwidth, and K is the Gaussian kernel function. The probability density map is overlaid with the green space functional zoning to delineate high-probability planting areas (P...). c >0.7) and low probability planting areas (P c (≤0.7), the high probability area is the key area for sampling.

[0144] (2) Heterogeneity-driven adaptive sampling method

[0145] Based on the spatial distribution probability map of vegetation types and green space functional zoning obtained by integrating prior knowledge, a heterogeneity-driven adaptive sampling strategy is designed. In areas with high probability distribution or complex vegetation types (determined by calculating the local feature diversity index), dense grid sampling or object-based sampling is used; in areas with low probability distribution or uniform vegetation types, sparse grid sampling is used. The grid spacing or sampling density is dynamically correlated with the spatial heterogeneity of vegetation types or the probability of specific types in that area. The formula for calculating the grid spacing d is:

[0146]

[0147] (3) Human-machine collaborative sample verification and enhancement

[0148] Candidate sample points or regions obtained through adaptive sampling will be combined with high-resolution remote sensing imagery and ground survey data for human-machine collaborative verification and correction to ensure sample accuracy. During verification, the vegetation type represented by each sample point or region needs to be determined and assigned a corresponding category label (e.g., trees, shrubs, grasslands). The category label should be consistent with the vegetation classification system defined within the study area.

[0149] For classes with limited or unevenly distributed samples, sample augmentation techniques are employed, such as temporal interpolation of time-series features or generation of synthetic sequences; spatial affine transformations (rotation, scaling, flipping) of images or feature blocks; and oversampling methods like SMOTE (Synthetic Minority Over-sampling Technique) to generate synthetic samples to address class imbalance. When generating synthetic samples, they should be assigned the same class labels as the original samples.

[0150] Step 4: Construction of the optimized low-dimensional feature set: For the multi-dimensional feature set constructed in Step 2, the feature subset is first screened based on the maximum correlation minimum redundancy (mRMR) algorithm. Then, the variance of each feature in the enhanced sample set obtained in Step 3 is calculated and subjected to maximum and minimum normalization to obtain the normalized variance as the feature weight. Then, the dimensionality is reduced by principal component analysis (PCA) to generate the low-dimensional feature set D'.

[0151] (1) Feature subset selection based on the maximum correlation minimum redundancy (mRMR) algorithm includes the following steps:

[0152] 1) Calculate mutual information:

[0153] First, calculate F for each feature dimension. i Mutual information I(F) between the target category label Y of the urban green space vegetation enhancement sample obtained in step three and the target category label Y of the target category label Y. i ;Y);Mutual Information I(F) i The higher the value of F(Y), the higher the feature dimension F. i The more information it contains that can distinguish different vegetation types, the greater the correlation.

[0154] Secondly, for any two different feature dimensions F i and F j Calculate the mutual information I(F) between them. i ;F j Mutual information I(F) i ;F j The higher the value of ), the more common information these two feature dimensions contain, i.e., the higher the redundancy.

[0155] 2) mRMR criterion ranking:

[0156] The mRMR criterion balances feature relevance and redundancy using the following formula and ranks all candidate features:

[0157]

[0158] Where S is the selected feature subset, and |S| is the number of selected features; the ranking results are arranged from high to low according to the mRMR value, and the features ranked higher are considered more important;

[0159] 3) Feature subset selection:

[0160] Based on a preset threshold or the number of features, select the top-ranked feature subset S from the sorting results. selected .

[0161] (2) The selected feature subset S selected Calculate F for each feature i The variance Var(F) in the augmented sample set i ), and perform max-min normalization to obtain the normalized variance Var′(F i ):

[0162]

[0163] Use the normalized variance as the feature weight:

[0164] w i =Var′(F i )w i It is the weight of the i-th feature, which will be used for dynamic weighted node splitting in the subsequent random forest classification model.

[0165] (3) Dimensionality reduction and feature space optimization

[0166] Principal component analysis (PCA) is used to reduce the dimensionality of the optimized feature set, retaining principal components with a cumulative contribution rate of over 95%, and generating a low-dimensional feature subset D′, which further reduces the computational cost and may remove linear correlations between features.

[0167] Step 5: Automatic classification of green space: The optimized low-dimensional feature set D' constructed in Step 4 and the urban green space vegetation enhancement sample set constructed in Step 3 are input into the improved random forest model to achieve automatic classification of urban green space at the canopy scale; when constructing the decision tree, the improved random forest algorithm uses the weighted Gini coefficient based on the feature weights calculated in Step 4 as the node splitting criterion, and performs spatial consistency optimization of the classification results based on neighborhood constraints.

[0168] (1) Model Input

[0169] The preprocessed multi-source remote sensing data and the constructed urban green space vegetation sample set were input into the improved random forest model.

[0170] 1) Optimize the multidimensional feature set: The multidimensional feature vectors after optimization and dimensionality reduction in step four are used as input features for the random forest model.

[0171] 2) Enhanced training sample set: The urban green space vegetation training sample set constructed and enhanced in step three contains the feature vector of each sample and the corresponding precise class label.

[0172] (2) Model Algorithm (Improved Random Forest)

[0173] This invention employs an improved random forest algorithm, the core improvement of which lies in the construction of the decision tree and the integration of the final classification results. It adopts an improved random forest algorithm that introduces dynamic weighted node splitting, neighborhood constraints, and adaptive class balancing, combined with the optimized multidimensional feature set and enhanced sample set constructed in the aforementioned steps, to achieve automatic classification of urban green space at the canopy scale.

[0174] 1) Dynamically weighted node splitting criterion

[0175] To overcome the limitations of traditional decision tree node splitting criteria (such as the Gini coefficient or information gain) in handling multi-source high-dimensional features and class imbalance, this invention proposes a dynamic weighted node splitting criterion. This criterion optimizes the feature selection and splitting threshold determination process by introducing dynamic weights based on feature importance and adaptive weights based on class distribution when constructing each decision tree in a random forest. The specific implementation is as follows:

[0176] The formula for calculating the weighted Gini coefficient is:

[0177]

[0178] Where S is the sample set of the current node, Values(F i ) is a feature F i The set of possible values ​​for |S v | represents the number of samples with feature value v, |S| represents the total number of samples in the current node, Classes represents the set of all class labels, and c represents the specific class label. v,c | for feature F i The number of samples with value v and category c, w i The feature F calculated in step four i The weight.

[0179] By using the aforementioned dynamic weighted node splitting criterion, this invention enables the decision tree to utilize multi-source feature information more effectively during the construction process, prioritizes the selection of features that contribute more to distinguishing different vegetation types, and significantly improves the model's learning ability and recognition accuracy for key vegetation types (such as rare tree species or specific ecological function green spaces) with limited sample size through class weights and weighted sampling mechanisms, thereby enhancing the model's robustness.

[0180] 2) Spatial consistency optimization with neighborhood constraints

[0181] In the final voting stage of the random forest, spatial neighborhood information is introduced to correct the classification results, reducing "salt-and-pepper noise" and enhancing the spatial consistency of the classification results. For each pixel, the class distribution of the random forest voting results within its preset neighborhood (e.g., a 3x3 or 5x5 window) is statistically analyzed. If the initial classification result of the center pixel is inconsistent with the voting results of the majority of pixels in the neighborhood (e.g., exceeding 50% or a higher threshold), the class of the center pixel is corrected to the class of the majority of pixels in the neighborhood. This is a post-processing spatial smoothing mechanism, but its correction logic is based on the majority voting results in the neighborhood rather than simple filtering, preserving more details. The calculation formula is:

[0182] ClasS(i) final =mode({Class(j)} initial |j∈N(i)})

[0183] Where, Class(i) final For the final classification result of pixel i, mode({…}) takes the element that appears most frequently in the set (the mode), and Class(j) initial Let N(i) be the initial classification result (random forest voting result) for pixel j, and let N(i) represent the neighborhood of pixel i.

[0184] (3) Model Training

[0185] The improved random forest model is trained using an enhanced training sample set. Parameters of the random forest are set, such as the number of decision trees (e.g., 100-500), maximum depth, minimum number of leaf node samples, etc. These parameters can be optimized on the training set using methods such as cross-validation. The training process involves building multiple decision trees, each independently learning the relationship between features and categories.

[0186] (4) Accuracy verification

[0187] The classification accuracy of the model is evaluated using an independent test sample set. Evaluation metrics include overall accuracy, Kappa coefficient, confusion matrix, user accuracy, and producer accuracy. Particular attention is paid to user accuracy and producer accuracy for a few classes. Based on the accuracy validation results, model parameters, feature selection strategies, and sample augmentation methods can be iteratively optimized until satisfactory classification accuracy is achieved.

[0188] Through the above steps, this invention enables high-precision, adaptive, and canopy structure-aware classification of urban green space vegetation at the canopy scale based on remote sensing, providing strong data support for urban ecological environment research and management. The method in this embodiment, applied to a typical area of ​​Beijing, yields the following canopy-scale urban green space vegetation classification results: Figure 6 , 7 As shown.

[0189] The above embodiments are only a partial embodiment of the present invention and do not cover all of the present invention. Based on the above embodiments and the accompanying drawings, those skilled in the art can obtain more implementation methods without creative effort. Therefore, all implementation methods obtained without creative effort should be included within the protection scope of the present invention.

Claims

1. A method for classifying urban green space vegetation at the canopy scale based on remote sensing, characterized in that, Includes the following steps: Step 1: Acquisition and Fusion of Multi-Source Remote Sensing Data: Acquire multi-source remote sensing data from the same region, multiple time phases, and consistent spatiotemporal references; preprocess the multi-source remote sensing data; and perform spatial registration and fusion to obtain a fused data cube. Step 2: Construction of a Canopy-Scale Multidimensional Feature Set from Multi-Source Remote Sensing Data: From the data cube fused in Step 1, extract five categories of multidimensional features: spectral features, time-series dynamic features, three-dimensional structural features, contextual spatial structural features, and multi-source fusion composite features to construct a structured multidimensional feature set D. Among them, F R F represents the regional feature vector of urban green space areas, where the numerical subscript indicates the ordinal number of the urban green space area, and M is the total number of urban green space areas; R The feature of a certain dimension in the data is represented as F. i , where i represents the dimension number; Step 3: Construction of urban green space vegetation enhancement sample set: The spatial distribution probability map of vegetation type is obtained by fusing multi-source prior data. Based on the spatial distribution probability map and green space functional zoning, the multi-dimensional feature set constructed in Step 2 is used as the basic data for adaptive sampling. The samples are then verified and enhanced by human-computer collaboration to obtain the urban green space vegetation enhancement sample set. Each sample includes a multi-dimensional feature vector and the corresponding category label. Step 4: Construction of the optimized low-dimensional feature set: For the multi-dimensional feature set constructed in Step 2, the feature subset is first screened based on the maximum correlation minimum redundancy (mRMR) algorithm. Then, the variance of each feature in the enhanced sample set obtained in Step 3 is calculated and subjected to maximum and minimum normalization to obtain the normalized variance as the feature weight. Then, the dimensionality is reduced by principal component analysis (PCA) to generate the low-dimensional feature set D'. Step 5: Automatic Green Space Classification: The optimized low-dimensional feature set D' constructed in Step 4 and the urban green space vegetation enhancement sample set constructed in Step 3 are input into the improved random forest model to achieve automatic classification of urban green spaces at the canopy scale. When constructing the decision tree, the improved random forest algorithm uses the weighted Gini coefficient based on the feature weights calculated in Step 4 as the node splitting criterion, and performs spatial consistency optimization based on neighborhood constraints on the classification results. The formula for calculating the weighted Gini coefficient is: Where S is the sample set of the current node, values(F) i ) is a feature F i The set of possible values ​​for |S v | represents the number of samples with feature value v, |S| represents the total number of samples in the current node, Classes represents the set of all class labels, and c represents the specific class label. v,c | for feature F i The number of samples with value v and category c, w i The feature F calculated in step four i The weight.

2. The method for classifying urban green space vegetation at the canopy scale based on remote sensing according to claim 1, characterized in that, In step one, the multi-source high-resolution remote sensing data includes: optical multispectral data with a spatial resolution better than 5 meters; optical hyperspectral data with a spatial resolution better than 30 meters; and laser point cloud data with a point density of not less than 10 points / square meter. The spatial registration and fusion involves resampling all data to a unified spatial resolution, performing spatial registration, and stacking the effective bands and derivative products from different data sources to construct a multi-channel, multi-modal, spatiotemporally consistent fused data cube.

3. The method for classifying urban green space vegetation at the canopy scale based on remote sensing according to claim 1, characterized in that, In step two, the spectral features include: normalized vegetation index, enhanced vegetation index, red edge index, and chlorophyll absorption valley depth; the time-series dynamic features include: seasonal variation index, phenological phase features, and cumulative greenness during the growing season; the three-dimensional structure includes: canopy volume density, vertical profile heterogeneity, canopy height stratification, reflectance gradient, and canopy surface roughness; the contextual spatial structure features include: spatial autocorrelation index and neighborhood texture diversity; and the multi-source fusion composite features include: spectral-height joint index and reflectance-texture synergistic features.

4. The method for classifying urban green space vegetation at the canopy scale based on remote sensing according to claim 1, characterized in that, In step two, for each candidate urban green space area R k The spectral features, time-series dynamic features, three-dimensional structural features, contextual spatial features, and multi-source fusion composite features of urban green space areas were extracted to construct regional-level features: Wherein, N(R) k () represents the total number of pixels contained in the k-th urban green space area. Indicates the region R k Summing all pixels j within f j Let be the single-pixel feature vector of the j-th pixel.

5. The method for classifying urban green space vegetation at the canopy scale based on remote sensing according to claim 1, characterized in that, Step three specifically includes the following operations: 3-1 Multi-source prior data fusion and spatial mapping: Collect vector data of green space functional zoning in the study area and existing tree / grass species planting distribution databases or survey data; use kernel density estimation (KDE) to generate spatial distribution probability maps of major tree / grass species; overlay the probability maps with green space functional zoning to identify high-probability distribution areas of specific vegetation types in different functional zones; 3-2 Heterogeneity-Driven Adaptive Sampling Method: Based on the spatial distribution probability map of vegetation types and green space functional zoning obtained by fusing prior knowledge, a heterogeneity-driven adaptive sampling strategy is designed: in areas with high probability distribution or complex vegetation types, dense grid sampling or object-based sampling is adopted. In areas with low probability distribution or areas with a single vegetation type, sparse grid sampling is used; the grid spacing or sampling density is dynamically correlated with the spatial heterogeneity of vegetation types or the probability of specific types in the area. 3-3 Human-Machine Collaborative Sample Validation and Enhancement: Candidate sample points or regions obtained through adaptive sampling are combined with high-resolution remote sensing imagery and ground survey data for human-machine collaborative verification and correction to ensure the accuracy of the samples. The vegetation type represented by each sample point or region is determined and assigned a corresponding category label. For categories with a small number of samples or uneven distribution, sample augmentation techniques are used to generate synthetic samples.

6. The method for classifying urban green space vegetation at the canopy scale based on remote sensing according to claim 1, characterized in that, Step four, the feature subset selection based on the maximum relevance minimum redundancy (mRMR) algorithm, includes the following steps: 4-1 Calculate mutual information: Calculate F for each feature dimension i Mutual information I(F) between the target category label Y of the urban green space vegetation enhancement sample obtained in step three and the target category label Y of the target category label Y. i ;Y);Mutual Information I(F) i The higher the value of F(Y), the higher the feature dimension F. i The more information it contains that can distinguish different vegetation types, the greater the correlation. For any two distinct feature dimensions F i and F j Calculate the mutual information I(F) between them. i ;F j Mutual information I(F) i ;F j The higher the value of ), the more common information these two feature dimensions contain, i.e., the higher the redundancy. 4-2mRMR criterion for sorting: The mRMR criterion balances feature relevance and redundancy using the following formula and ranks all candidate features: Where S is the selected feature subset, and |S| is the number of selected features; the ranking results are arranged from high to low according to the mRMR value, and the features ranked higher are considered more important; 4-3 Feature Subset Selection: Based on a preset threshold or the number of features, select the top-ranked feature subset S from the sorting results. selected .

7. The method for classifying urban green space vegetation at the canopy scale based on remote sensing according to claim 6, characterized in that, In step four, the specific operation for obtaining feature weights is as follows: For the selected feature subset S selected Calculate F for each feature i The variance Var(F) in the augmented sample set i ), and perform max-min normalization to obtain the normalized variance Var′(F i ): Use the normalized variance as the feature weight: w i =Var′(F i ) w i It is the weight of the i-th feature, which will be used for dynamic weighted node splitting in the subsequent random forest classification model.

8. The method for classifying urban green space vegetation at the canopy scale based on remote sensing according to claim 1, characterized in that, Step 5 also includes the following operations: Spatial consistency optimization of neighborhood constraints: In the final voting stage of the random forest, spatial neighborhood information is introduced to correct the classification results. For each pixel, the category distribution of the random forest voting results in its preset neighborhood is statistically analyzed. If the initial classification result of the center pixel is inconsistent with the voting results of the majority of pixels in the neighborhood, the category of the center pixel is corrected to the category of the majority of pixels in the neighborhood. The calculation formula is: Class(i) final =mode({Class(j) initial |j∈N(i)}) Where, Class(i) final For the final classification result of pixel i, mode({…}) takes the element that appears most frequently in the set (the mode), and Class(j) initial Let N(i) be the initial classification result (random forest voting result) for pixel j, and let N(i) represent the neighborhood of pixel i.

Citation Information

Patent Citations

  • Comprehensive weighted master data identification method based on machine learning

    CN113920366A

  • Forest vegetation automatic classification method and system fused with multi-source satellite remote sensing data

    CN118015453A