Single tree extraction method and system based on multispectral data
By using conditional generation adversarial networks in the single wood extraction method for radiation correction, combined with multi-scale attention fusion network and terrain adaptive geometric correction, the single wood segmentation accuracy and reliability problems caused by atmospheric scattering, terrain shadowing and geometric distortion in the traditional method are solved, and higher accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510665494.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The traditional single wood extraction method is disturbed by factors such as atmospheric scattering, terrain shadows and geometric distortions during multispectral imaging, resulting in spectral information distortion and spatial position shift, seriously affecting the accuracy and reliability of single wood segmentation.
The end-to-end radiation correction system based on condition-generating adversarial network is adopted, and the atmospheric parameter generator and terrain shadow discriminator are synchronously compensated for the influence of atmospheric and terrain, and a multi-scale attention fusion network is designed to achieve the deep fusion of spectral and texture features, while introducing a terrain-adaptive geometric correction strategy and error transmission evaluation model.
Effectively eliminate the atmospheric and terrain influence in multi-spectral data, improve the accuracy and robustness of single-wood extraction, and significantly improve the single-wood segmentation performance in complex scenarios.
Smart Images

Figure CN120182838A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing and forest resource monitoring, and specifically provides a single-tree extraction method and system based on multi-spectral data. Background Art
[0002] As a core technology for forest resource monitoring and ecological research, single-tree extraction is of great significance for precision forestry management and carbon sink measurement. Traditional single-tree extraction methods mainly rely on optical remote sensing data. However, during the multi-spectral imaging process, it is inevitably affected by factors such as atmospheric scattering, terrain shadows, and geometric distortion, resulting in spectral information distortion and spatial position offset, seriously affecting the accuracy and reliability of single-tree segmentation.
[0003] Currently, a Chinese patent with the application number "CN202410855855.6" discloses a eucalyptus extraction method, its device, equipment, and medium based on hyperspectral data. Among them, the method includes: obtaining the first satellite remote sensing image, eucalyptus sample vector data, and cultivated land range data of the target research area; performing image preprocessing on the first satellite remote sensing image to obtain the second satellite remote sensing image; performing image feature difference analysis and processing based on the second satellite remote sensing image and eucalyptus sample vector data to extract the object interpretation marks of multiple typical objects in the second satellite remote sensing image; performing target feature analysis and processing based on multiple object interpretation marks and the second satellite remote sensing image to obtain the eucalyptus spectral features and eucalyptus vegetation index features; performing regional screening and processing based on the eucalyptus spectral features, eucalyptus vegetation index features, preset thresholds, cultivated land range data, and preset screening conditions to obtain the extraction result of the eucalyptus spatial distribution within the cultivated land. Although it can also identify single trees, it cannot effectively reduce inevitable interferences such as atmospheric scattering and terrain shadows.
[0004] Traditional radiometric correction relies on physical model parameter assumptions and is difficult to adapt to the non-linear interferences of complex atmospheric conditions and terrain shadows. Multi-modal feature fusion mostly uses simple weighting or splicing methods and cannot dynamically explore the deep complementarity of spectral-texture features. Geometric registration lacks terrain adaptability, and the interpolation error in complex areas is significant. The feature selection process relies on empirical thresholds, which easily leads to the curse of dimensionality and low computational efficiency. The error transfer mechanism is not clear, and it is difficult to quantify the influence weights of each processing link on the final result. In addition, the noise interference and edge blurring problems in low-quality remote sensing data further restrict the integrity and spatial positioning accuracy of single-tree segmentation. Therefore, there is an urgent need for an intelligent single-tree extraction method that integrates radiometric correction, feature fusion, geometric registration, and error control to break through the application limitations of existing technologies in complex scenarios. Summary of the Invention
[0005] (I) Technical Problems to be Solved Aiming at the deficiencies of the existing technology, the present invention provides a method and system for extracting individual trees based on multi-spectral data, constructs an end-to-end multi-spectral individual tree extraction framework, synchronously compensates for the influence of atmosphere and terrain through a conditional generative adversarial network, designs a multi-scale attention fusion network to achieve deep feature association, and introduces a terrain-adaptive geometric correction strategy. At the same time, an error transfer evaluation model is constructed to quantify the error contribution degree of each link, providing a reliability guarantee for multi-modal data fusion. This solution effectively solves the core problems such as radiation distortion, inefficient feature fusion, geometric registration error and difficulty in error tracing in traditional methods, and significantly improves the accuracy and robustness of individual tree extraction in complex scenes.
[0006] (II)Technical Solution To achieve the above objectives, the present invention is realized through the following technical solutions: A method for extracting individual trees based on multi-spectral data, the extraction method includes: Construct a radiation correction system to eliminate the influence of atmosphere and terrain in the collected multi-spectral dataset, and extract general spectral data from the multi-spectral dataset after eliminating the influence; Fuse the general spectral data with the texture features extracted from the gray-level co-occurrence matrix, segment the individual tree crowns in the general spectral data to form an independent individual tree dataset; Use the SuperPoint feature detection combined with the SuperGlue matching algorithm to establish corresponding point pairs, adjust the registration parameters through the RANSAC algorithm, and perform geometric distortion correction using corresponding interpolation according to the terrain complexity to form a spatially aligned dataset; Denoise the data after geometric distortion correction, use an attention network model to enhance the target edge features, and train the ReliefF-PCA feature selection framework using the dataset to screen out the optimal combination of spectral-texture-structure features; Construct a quantitative evaluation index of radiation, use the error analysis method to evaluate the error transfer effect of each processing link, and extract the individual tree multi-modal dataset from the individual tree dataset through cross-validation.
[0007] Further, the radiation correction system adopts an end-to-end architecture based on a conditional generative adversarial network, including an atmospheric parameter generator, a terrain shadow discriminator and a joint discriminator; Among them, the atmospheric parameter generator adopts an atmospheric parameter generator based on the U-Net structure, including 5 downsampling layers and 5 upsampling layers. Each layer uses the ReLU activation function and batch normalization. The downsampling layer uses a 3×3 convolution kernel to extract high-level features, and the upsampling layer restores the spatial resolution through transposed convolution. At the same time, skip connections are set to fuse high- and low-level features; The terrain shadow discriminator adopts a terrain shadow discriminator with a PatchGAN structure, including 4 convolutional layers, with a stride of 2 for each layer, and using the LeakyReLU activation function; The joint discriminator consists of 3 fully connected layers and is used to receive the corrected spectrum and real surface reflectance data simultaneously.
[0008] Furthermore, the elimination of the influence of atmosphere and terrain in the collected multi-spectral dataset specifically includes the following steps: Input the atmospheric parameters simulated by MODTRAN and the DEM terrain data into the atmospheric parameter generator, and then map the DEM data to the spectral data coordinate system through the spatial transformation layer; The atmospheric parameter generator reconstructs the atmospheric radiation transmission process through the transposed convolutional layer, combines the downsampling layer, upsampling layer and skip connection, and outputs the corrected spectral data; then uses the terrain shadow compensation module to perform terrain shadow compensation on the corrected spectral data to preliminarily process the influence of terrain on multi-spectral data; The Adam optimizer is used to train the model, and the network parameters are optimized through the generator loss and discriminator loss.
[0009] Furthermore, when segmenting the single-tree crowns in the general spectral data, a multi-scale attention fusion network is used to achieve the deep fusion of spectral and texture features, specifically including: Construct a dual-branch attention network based on the channel-level attention mechanism and the spatial-level attention mechanism, calculate the attention weight matrices of spectral features and texture features through global average pooling and fully connected layers, and dynamically adjust the fusion ratio; Construct a 5-layer U-Net architecture, each layer contains 2 convolutional blocks and 1 max-pooling layer, and perform upsampling through transposed convolution in the decoding stage, and use skip connections to fuse the spectral reflectance mean and gray-level co-occurrence matrix features at different scales; Through the combined loss function , the combined loss coefficient is obtained. In the formula, is the focal loss function, is to calculate the Dice coefficient based on the edge ground truth extracted by the Canny operator, are respectively and corresponding weights.
[0010] Furthermore, when forming the independent single-tree dataset, the morphological reconstruction algorithm is used to remove the hole noise in the segmentation result, then the Delaunay triangulation is used to establish the canopy spatial adjacency relationship, and a confidence evaluation model is constructed by combining the spectral reflectance mean and texture entropy value to screen out the complete single-tree samples.
[0011] Furthermore, establishing corresponding points by combining with the SuperGlue matching algorithm specifically includes the following steps: S1. Use the SuperPoint network for end-to-end key point detection and descriptor extraction. The network includes a feature pyramid module and a differentiable non-maximum suppression layer, and outputs sparse key points and their descriptors; S2. Perform cross-view feature matching through the SuperGlue algorithm, construct a matching matrix using the attention mechanism, and perform optimal matching assignment in combination with the Sinkhorn algorithm; S3. Apply the RANSAC algorithm to purify the initial matching points. Set the maximum number of iterations to 1000 times and the distance threshold to 3 pixels, and retain the matching pairs with an inlier ratio exceeding 85%; S4. Select interpolation methods based on terrain complexity grading: bilinear interpolation for flat areas, cubic spline interpolation for medium areas, and thin plate spline interpolation for complex areas; S5. Use the improved TPS algorithm for geometric distortion correction, calculate the terrain undulation degree in combination with DEM data, and dynamically adjust the control point density.
[0012] Furthermore, for the data denoising, a hybrid denoising method combining the non-local mean denoising algorithm and adaptive median filtering is adopted, and edge information is retained by setting the spectral similarity threshold and spatial neighborhood window; The attention network model adopts a two-branch attention architecture, including a channel attention module and a spatial attention module. The channel attention module generates spectral feature weights through global average pooling and fully connected layers, and the spatial attention module generates position weights through convolution operations, and finally outputs an enhanced edge feature map.
[0013] Furthermore, the training process of the ReliefF-PCA feature selection framework includes: Calculate the ReliefF weights of each feature, set the number of iterations and the number of neighbors, and evaluate the correlation between the feature and the category through mutual information; Use PCA for dimensionality reduction and retain the principal components with a cumulative variance contribution rate ≥ 95%; Construct a multi-modal feature fusion space, and standardize the spectral reflectance, gray-level co-occurrence matrix texture, and morphological structure features; Optimize the feature combination through cross-validation and screen out the optimal feature subset.
[0014] Furthermore, the evaluation of error transfer effects includes: a. Define a set of radiation correction quality evaluation indicators, including root mean square error, mean absolute error, structural similarity, and spectral angle matching, and evaluate the radiation fidelity of the corrected spectral data through multi-index fusion; b. Construct an error propagation model, analyze the error contribution degrees of processing links such as atmospheric correction, terrain compensation, and feature fusion based on the error propagation theory, and quantify the influence weights of errors in each processing link on the final individual tree extraction result; c. Design a cross-validation experiment, divide the multi-spectral data set into a training set, a validation set, and a test set, and calculate the confidence interval of the error propagation path through Monte Carlo simulation; d. Establish a multi-modal data fusion rule, standardize and fuse the spectral data verified for errors with texture and structural features, construct a feature space using principal component analysis, and achieve hierarchical extraction of the individual tree multi-modal data set through K-means clustering; Among them, the error propagation model derives the covariance matrix of the processing link errors through the chain rule, and determines the key error sources in combination with sensitivity analysis; the multi-modal data set includes the multi-spectral reflectance after radiometric correction, the texture parameters of the gray-level co-occurrence matrix, and the morphological structure features.
[0015] An individual tree extraction system based on multi-spectral data, the system includes: Radiometric correction module: Construct an end-to-end radiometric correction system based on a conditional generative adversarial network, and eliminate the atmospheric interference and terrain shadows in the multi-spectral data through the collaborative training of the atmospheric parameter generator, terrain shadow discriminator, and joint discriminator, and output general spectral data; Feature fusion and segmentation module: Design a multi-scale attention fusion network, dynamically fuse the general spectral data with the texture features of the gray-level co-occurrence matrix, achieve accurate segmentation of the individual tree crowns, and generate an independent individual tree data set; Geometric correction module: Use SuperPoint feature detection combined with the SuperGlue matching algorithm to establish corresponding point pairs, and dynamically select an interpolation method for geometric distortion correction based on the terrain complexity to form a spatially aligned data set; Data enhancement module: Construct an attention-enhanced denoising network, combine non-local mean filtering and adaptive median filtering for noise suppression, and enhance the target edge features through channel attention and spatial attention mechanisms; Feature selection module: Establish a ReliefF-PCA feature selection framework, and screen the optimal combination of spectral-texture-structural features through mutual information evaluation and principal component analysis; Error evaluation module: Construct an error propagation evaluation model, quantify the error contribution degrees of each processing link, and generate an individual tree multi-modal data set after cross-validation.
[0016] (III) Beneficial effects The present invention provides an individual tree extraction method and system based on multi-spectral data, having the following beneficial effects: Radiation correction system: An end-to-end model is constructed using a conditional generative adversarial network (cGAN). Through the adversarial training of an atmospheric parameter generator (U-Net structure) and a terrain shadow discriminator (PatchGAN structure), the inverse modeling of the atmospheric radiation transfer process and terrain shadow compensation are realized. This system maps DEM data to the spectral coordinate system through a spatial transformation layer, and reconstructs the radiation transfer path in combination with a transposed convolutional layer, while eliminating the interference of atmospheric scattering and terrain shadows and retaining the hyperspectral feature information.
[0017] Multi-modal feature fusion network: A dual-branch attention mechanism is designed. The feature weight matrix is calculated through global average pooling and a fully connected layer to dynamically adjust the fusion ratio of the spectral reflectance mean and the texture of the gray-level co-occurrence matrix. Combining the morphological reconstruction algorithm and Delaunay triangulation technology, while suppressing hole noise, the spatial adjacency relationship of the tree crowns is established, effectively solving the problem of segmenting overlapping tree crowns in complex forest stands.
[0018] Adaptive geometric registration strategy: Based on the terrain complexity grading, interpolation methods are selected (bilinear interpolation in flat areas and thin plate spline interpolation in complex areas). Combining with an improved TPS algorithm to adjust the control point density, high-precision spatial alignment of multi-view data is achieved.
[0019] ReliefF-PCA feature selection framework: The feature weights are calculated through the ReliefF algorithm (evaluating the correlation between features and classes based on mutual information). Combining with PCA dimensionality reduction to retain the principal component information, and finally the optimal feature combination including spectral reflectance, texture parameters, and morphological features is selected. This framework reduces the redundant feature dimensions while maintaining the classification accuracy, improving the model training efficiency.
[0020] Error propagation quantization analysis: An error propagation model based on the chain rule is established. The error contribution of each processing link is analyzed through the covariance matrix, and the key error sources are determined in combination with Monte Carlo simulation. This mechanism realizes the quantitative control of errors in links such as atmospheric correction and feature fusion, improving the reliability and repeatability of the method.
[0021] Attention-enhanced denoising network: A dual-branch attention module (channel attention + spatial attention) is designed to synergistically enhance the target edge features. Combining a hybrid denoising method of non-local means and adaptive median filtering, while suppressing noise, the image edge details are retained, improving the segmentation performance in low-quality data scenarios. Brief Description of the Drawings
[0022] Figure 1 It is a flowchart of the overall structure of the present invention. Detailed Embodiments
[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0024] Forests play a crucial role in global ecosystem services, supporting carbon sequestration, biodiversity conservation, and ecosystem resilience. At the same time, as an important part of the comprehensive natural resource survey work, accurately estimating stand and individual tree parameters is crucial for effective ecosystem management and forestry practices in this context. The progress of remote sensing technology, combined with the increased availability of remote sensing data and improved methods, has partially solved the problem of forest resource surveys in complex terrains and significantly improved efficiency. These advancements help meet the demand for efficient and large-scale forest resource assessments.
[0025] As the core technology for forest resource monitoring and ecological research, individual tree extraction is of great significance for precision forestry management and carbon sink measurement. Traditional individual tree extraction methods mainly rely on optical remote sensing data, but during the multi-spectral imaging process, they are inevitably affected by factors such as atmospheric scattering, terrain shadows, and geometric distortions, resulting in spectral information distortion and spatial position offset, seriously affecting the accuracy and reliability of individual tree segmentation. Specifically as follows: First, in the field of radiometric correction, traditional atmospheric correction models based on MODTRAN rely on prior parameters such as preset aerosol types and water vapor content, making it difficult to adapt to complex and changing atmospheric environments and unable to synchronously compensate for the non-linear effects of terrain shadows on spectra. For example, the reflectance attenuation in shadow areas in mountainous terrains is often misjudged as differences in vegetation types, leading to subsequent classification errors. In addition, traditional methods adopt a step-by-step correction strategy (atmosphere first and then terrain), ignoring the interaction between atmospheric parameters and terrain features, and the corrected data is prone to spectral discontinuity.
[0026] Second, in the feature fusion and segmentation link, existing technologies mostly use simple spectral-texture feature splicing or linear weighting methods, making it difficult to capture the deep associations between multi-modal data. For example, the gray-level co-occurrence matrix (GLCM) texture features are vulnerable to noise pollution in low-resolution images, and traditional convolutional neural networks lack a dynamic attention allocation mechanism for multi-scale features, resulting in blurred segmentation boundaries between adjacent tree crowns in complex stands and problems of over-segmentation or under-segmentation.
[0027] Thirdly, in the data augmentation and feature selection stage, traditional Gaussian filtering or mean filtering is prone to losing edge details during the noise reduction process, while the denoising network based on deep learning lacks targeted optimization for spectral features. In terms of feature selection, when existing ReliefF or PCA methods are used alone, it is difficult to balance feature discrimination and computational efficiency, and the redundancy between multi-modal features is not considered, resulting in a too high dimensionality of the feature space or information loss.
[0028] Now, aiming at the problems pointed out above, as Figure 1 shown, the specific improvements are as follows: 1. Data acquisition Use a drone system equipped with a multi-spectral camera (LightGene MF620 camera) to obtain image data of 5 bands (blue, green, red, red edge, near-infrared), and synchronously record POS positioning information; the flight parameters are set as follows: flight altitude 150m, ground resolution 0.05m, forward overlap rate 75%, and side overlap rate 60%.
[0029] 2. Radiometric correction 2.1 Network architecture: Construct an end-to-end radiometric correction model based on the conditional generative adversarial network (CGAN), which includes an atmospheric parameter generator (U-Net structure), a terrain shadow discriminator (PatchGAN), and a joint discriminator.
[0030] The atmospheric parameter generator adopts a 6-layer downsampling and upsampling structure (such as satellite images are downsampled from 1024x1024 to 16x16 and then restored), each layer includes a 3×3 convolution, batch normalization, and ReLU activation function, and skip connections are used to fuse high and low-level features (similar to the highway structure, allowing deep features to be directly transmitted to the shallow layer).
[0031] The terrain shadow discriminator includes 4 convolutional layers with a stride of 2, and uses the LeakyReLU activation function to discriminate the terrain shadow compensation effect of the corrected spectrum.
[0032] The joint discriminator consists of 3 fully connected layers (the number of neurons is 1024, 512, 1 respectively), and receives the corrected spectral data and the true surface reflectance data at the same time, and optimizes the network parameters by comparing the differences between the two.
[0033] 2.2 Training data preparation: Atmospheric parameter dataset: Generate 100,000 combinations of atmospheric parameters with different aerosol types (rural, urban, ocean), water vapor content (0.5 - 5cm), and solar zenith angle (0° - 70°) through MODTRAN simulation.
[0034] Terrain Shadow Dataset: Based on ASTER GDEM V3 data, terrain shadow masks with different slopes (0° - 60°) and aspect angles (0° - 360°) are generated and registered with the simultaneously acquired multispectral images.
[0035] 2.3. Training Process: Input the atmospheric parameters simulated by MODTRAN and DEM data, and map the DEM to the spectral coordinate system through a spatial transformation layer (affine transformation matrix).
[0036] The atmospheric parameter generator reconstructs the radiation transfer process (simulating the path of light through the atmosphere) through a deconvolution layer and outputs the corrected spectral data; restoring the blurred image to a clear image.
[0037] The terrain shadow compensation module calculates the slope and aspect (such as sunny side and shady side of the mountain) based on the DEM, and generates a shadow mask for compensation.
[0038] 2.4. Model Training: Adopt the Adam optimizer , set the learning rate of the generator to 1e - 4 and the learning rate of the discriminator to 4e - 4. The loss function includes the generator loss (L1 loss) and the discriminator loss (adversarial loss + terrain shadow compensation loss); during the training process, the correction effect is evaluated every 50 epochs, and indicators such as RMSE and MAE are used to verify the model performance.
[0039] The loss function is: , where is the total loss function, is the adversarial loss, used for the adversarial training between the discriminator and the generator; is the L1 norm loss , measuring the absolute error between the corrected spectrum and the true reflectance; is the shadow compensation loss , calculated based on the DEM slope and aspect.
[0040] 3. Feature Fusion and Segmentation 3.1. Network Design: Construct a dual - branch attention fusion network. The channel attention branch generates weights (8×8×64, similar to assigning different weights to different spectral bands) through global average pooling and fully connected layers, and the spatial attention branch uses a 7×7 convolution to generate position weights.
[0041] The U-Net architecture consists of 4 layers of encoding-decoding structures (compressed from 256×256 to 32×32 and then restored). Each convolutional block contains 2 convolutional layers of 3×3. Skip connections fuse features of different scales (spectral mean, GLCM entropy, contrast, for example: differentiating broad-leaved forests and coniferous forests); in the decoding stage, upsampling is performed through transposed convolution, and skip connections fuse the spectral mean (6-dimensional) and the texture of the gray-level co-occurrence matrix (GLCM) (4-dimensional).
[0042] 3.2. Feature Fusion: Dynamic Fusion Formula: , where is the fused feature vector, is dynamically calculated by the attention weight matrix, with a range of [0.3, 0.7], is the spectral feature vector (6-dimensional) after radiometric correction; is the texture feature vector (4-dimensional) extracted from the gray-level co-occurrence matrix.
[0043] 3.3. Training Strategy: The combined loss function is focal loss (γ = 2, focusing on difficult-to-classify samples) and edge Dice loss (λ = 0.5, protecting boundary information), and the AdamW optimizer (learning rate 1e-4) is used.
[0044] Data augmentation includes random rotation (±30°, simulating different shooting angles), scaling (0.8 - 1.2 times, adapting to shooting at different distances), and Gaussian noise (σ = 0.01, simulating sensor errors).
[0045] 3.4. Design of the Combined Loss Function: The combined loss function is adopted: , where is the combined loss function, is the focal loss function, is to calculate the Dice coefficient based on the edge ground truth extracted by the Canny operator, are respectively and the corresponding weights, = 0.7, = 0.3.
[0046] 3.5. Post-Processing of Segmentation: Morphological reconstruction uses opening and closing operations to eliminate holes (such as filling small gaps in the tree crown), Delaunay triangulation to establish adjacency relationships, and a confidence model to combine the mean spectral reflectance (μ > 0.6, such as a higher mean value in the green band of healthy vegetation), texture entropy value (E > 4.2, such as a higher entropy value for rough tree bark), and the number of adjacent tree crowns (N). After segmenting the general spectral data, general single-tree spectral data is obtained.
[0047] 3.6. Confidence evaluation model: Construct a triple evaluation index: mean spectral reflectance (μ), texture entropy value (E), and the number of adjacent tree crowns (N).
[0048] Confidence formula: , where is the confidence, μ is the mean spectral reflectance of the current tree crown; is the mean spectral reflectance of the general single-tree spectral data; N is the number of adjacent tree crowns of the current tree crown; is the maximum number of adjacent tree crowns in the general single-tree spectral data; E is the texture entropy value of the current tree crown; is the mean texture entropy value of the general single-tree spectral data; are the maximum and minimum values of the texture entropy value respectively; Samples with a confidence ≥ 0.7 are selected as the final single-tree dataset.
[0049] 4. Geometric correction 4.1. Feature matching: The SuperPoint network contains 3 layers of feature pyramids (stride = 2, 4, 8), outputs a 2048-dimensional descriptor, and non-maximum suppression retains the top-1000 key points.
[0050] The SuperGlue matching matrix is optimized by the Sinkhorn algorithm for 50 iterations, and RANSAC is iterated 500 times (random sampling verification to exclude incorrect matches), with the distance threshold set to 3 pixels.
[0051] 4.2. Terrain adaptive interpolation: Based on the DEM, calculate the terrain undulation (standard deviation). In flat areas (<2m, such as the North China Plain), bilinear interpolation is used (simple and fast), in medium areas (2 - 5m, such as the Loess Plateau), cubic spline interpolation is used (smoother), and in complex areas (>5m, such as the Hengduan Mountains).
[0052] Dynamically adjust the control point density (0.5 - 2 per 100m², with denser control points in mountainous areas), and combine the DEM slope to correct the transformation parameters.
[0053] 4.3. Terrain complexity classification: Calculating the Terrain Ruggedness Index (RuggednessIndex) based on DEM data: , where is the terrain ruggedness index (unit: meter); is the elevation difference of the i-th adjacent grid; n is the number of grid neighborhoods (usually taking 8-neighborhood); Classification criteria: flat area (RI < 2m), medium area (2m ≤ RI ≤ 5m), complex area (RI > 5m).
[0054] 4.4. Selection of interpolation method: Flat area: Bilinear interpolation (error ≤ 0.5 pixel); Medium area: Cubic spline interpolation (error ≤ 0.8 pixel); Complex area: Thin plate spline interpolation (TPS), dynamically adjusting the control point density (interval 5 - 20m).
[0055] 5. Data augmentation 5.1. Hybrid denoising: The window size of non-local means filtering is set to 7×7 (considering 49 surrounding pixels), and the similarity threshold τ = 0.05 (differences less than 5% are considered similar); Adaptive median filtering is iterated 3 times (gradually expanding the window to process salt-and-pepper noise), retaining the effective edges in the salt-and-pepper noise (such as wires remaining visible in the noise).
[0056] 5.2. Attention enhancement network: The channel attention module generates spectral weights (64-dimensional, such as giving higher weights to the near-infrared band), and the spatial attention module generates a 256×256 weight map through convolution (such as highlighting the road area); On the forest dataset, the edge retention rate after denoising reaches 94% (similar to erasing photo noise but retaining hair details), and the noise variance drops from 0.03 to 0.008.
[0057] 6. Feature selection 6.1. ReliefF-PCA framework: Calculate the ReliefF weights for 50 iterations and 10 neighbors (similar to voting to select important features), and screen the features with the top 30% of the weights; PCA retains 95% of the variance (such as compressing 32 spectral bands into 8 principal components), and compresses the 32-dimensional features to 8 dimensions.
[0058] 6.2. ReliefF weight calculation: The number of iterations is 1000 times, and the number of neighbors is 50. Calculate the correlation between features and categories: , where is the weight value of the i-th feature; is the number of neighboring samples (taking the value of 50); is the difference between the current sample and the j-th nearest neighbor of the same class in feature i; is the prior probability of class c; is the class to which the current sample belongs; is the difference between the current sample and the j-th nearest neighbor of class c in feature i; Retain the top 20% of the features by weight.
[0059] 6.3. Multimodal Fusion: Spectral features (means of 7 bands, such as red, green, blue, etc.), texture features (5 parameters such as GLCM contrast, entropy, etc., such as bark texture), and morphological features (4 parameters such as area, perimeter, etc., such as crown shape) are fused after standardization, and the classification accuracy of the feature subset reaches 93.6% (such as distinguishing pine trees and fir trees).
[0060] 7. Error Evaluation 7.1. Transfer Model: Derive the error covariance matrix based on the chain rule (such as calculating the impact of the error of each module on the final result). The error contribution of the atmospheric correction link is the largest (28%). Determine the error confidence interval through Monte Carlo simulation (1000 times, random sampling test).
[0061] 7.2. Cross-Validation: The multispectral dataset is divided into training / validation / test sets according to 7:2:1 (such as 1000 images divided into 700 / 200 / 100). The error transfer path analysis shows that the error is amplified 1.5 times in the feature fusion link, and the error confidence interval is compressed from ±5.2% to ±3.1% after optimization.
[0062] In the application, several formulas involved are calculated by taking their numerical values after dimensionless. The establishment of the formulas is obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. Some coefficients or weights in the formulas are set by those skilled in the art according to the actual situation, so no more details will be elaborated here.
[0063] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution.
[0064] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, and it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0065] As described above, the specific implementation manners of the present application are only provided, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application.
Claims
1. A method for extracting individual trees based on multispectral data, characterized in that, The extraction method includes: Construct a radiation correction system to eliminate the atmospheric and topographic effects in the collected multi-spectral dataset, and extract general spectral data from the multi-spectral dataset after eliminating the effects; Fuse the general spectral data with the texture features extracted by the gray-level co-occurrence matrix to segment the individual tree crowns in the general spectral data, forming an independent individual tree dataset; Adopt SuperPoint feature detection, combine with the SuperGlue matching algorithm to establish corresponding point pairs, adjust the registration parameters through the RANSAC algorithm, and perform geometric distortion correction using corresponding interpolation according to the terrain complexity to form a spatially aligned dataset; Denoise the data after geometric distortion correction, use an attention network model to enhance the target edge features, and train the ReliefF-PCA feature selection framework using the dataset to screen out the optimal combination of spectral-texture-structure features; Construct a quantitative evaluation index of radiation, adopt an error analysis method to evaluate the error transfer effect, and extract an individual tree multi-modal dataset from the individual tree dataset through cross-validation.
2. The method for extracting individual trees based on multispectral data according to claim 1, characterized in that: The radiation correction system adopts an end-to-end architecture based on a conditional generative adversarial network, including an atmospheric parameter generator, a terrain shadow discriminator, and a joint discriminator; Among them, the atmospheric parameter generator adopts an atmospheric parameter generator based on the U-Net structure, including a downsampling layer and an upsampling layer. Each layer uses the ReLU activation function and batch normalization. The downsampling layer uses a convolutional kernel to extract high-level features, and the upsampling layer restores the spatial resolution through transposed convolution. At the same time, skip connections are set to fuse high-level and low-level features; The terrain shadow discriminator adopts a terrain shadow discriminator with a PatchGAN structure, including a convolutional layer and uses the LeakyReLU activation function; The joint discriminator consists of fully connected layers and is used to receive the corrected spectrum and real surface reflectance data simultaneously.
3. The method for extracting individual trees based on multispectral data according to claim 2, characterized in that: The specific steps for eliminating the atmospheric and topographic effects in the collected multi-spectral dataset include: Input the atmospheric parameters simulated by MODTRAN and the DEM terrain data into the atmospheric parameter generator, and then map the DEM terrain data to the spectral data coordinate system through the spatial transformation layer; The atmospheric parameter generator reconstructs the atmospheric radiation transmission process through the transposed convolution layer, combines the downsampling layer, the upsampling layer, and the skip connection, and outputs the corrected spectral data; then use the terrain shadow compensation module to perform terrain shadow compensation on the corrected spectral data to preliminarily process the influence of terrain on the multi-spectral data; Adopt the Adam optimizer to train the model, and optimize the network parameters through the generator loss and the discriminator loss.
4. The method for extracting individual trees based on multispectral data according to claim 3, characterized in that: When segmenting the individual tree crowns in the general spectral data, a multi-scale attention fusion network is used to realize the fusion of spectral and texture features. The specific steps include: Construct a dual-branch attention network based on the channel-level attention mechanism and the spatial-level attention mechanism, calculate the attention weight matrices of spectral features and texture features through global average pooling and fully connected layers, and dynamically adjust the fusion ratio; Construct a multi-layer U-Net architecture, with each layer containing 2 convolutional blocks and 1 max-pooling layer. In the decoding stage, upsampling is performed through transposed convolution, and skip connections are used to fuse the mean spectral reflectance and gray-level co-occurrence matrix features at different scales. By , the combined loss function is obtained, where is the focal loss function, is to calculate the Dice coefficient based on the edge ground truth extracted by the Canny operator, are respectively and the corresponding weights.
5. The method for extracting individual trees based on multispectral data according to claim 4, characterized in that: When forming the independent single-tree dataset, the morphological reconstruction algorithm is used to remove the hole noise in the segmentation results. Then, the Delaunay triangulation is used to establish the canopy spatial adjacency relationship, and a confidence evaluation model is constructed by combining the mean spectral reflectance and texture entropy values to screen out complete single-tree samples.
6. The method for extracting individual trees based on multispectral data according to claim 5, characterized in that: The establishment of corresponding point pairs by combining the SuperGlue matching algorithm specifically includes the following steps: S1. Use the SuperPoint network for end-to-end key point detection and descriptor extraction. The network includes a feature pyramid module and a differentiable non-maximum suppression layer, and outputs sparse key points and their descriptors. S2. Perform cross-view feature matching through the SuperGlue algorithm, construct a matching matrix using the attention mechanism, and perform optimal matching assignment in combination with the Sinkhorn algorithm. S3. Apply the RANSAC algorithm to purify the initial matching point pairs, set the maximum number of iterations and the distance threshold, and retain the matching pairs with an inlier ratio exceeding 85%. S4. Select the interpolation method based on the terrain complexity grading: bilinear interpolation is used in flat areas, cubic spline interpolation is used in medium areas, and thin plate spline interpolation is used in complex areas. S5. Use the improved TPS algorithm for geometric distortion correction, calculate the terrain undulation degree in combination with DEM data, and dynamically adjust the control point density.
7. The method for extracting individual trees based on multispectral data according to claim 6, characterized in that: For data denoising, a hybrid denoising method combining the non-local means denoising algorithm and adaptive median filtering is used, and edge information is retained by setting the spectral similarity threshold and the spatial neighborhood window. The attention network model adopts a dual-branch attention architecture, including a channel attention module and a spatial attention module. Among them, the channel attention module generates spectral feature weights through global average pooling and fully connected layers; the spatial attention module generates position weights through convolution operations, and finally outputs an enhanced edge feature map.
8. The method for extracting individual trees based on multispectral data according to claim 7, characterized in that: The training process of the ReliefF-PCA feature selection framework includes: Calculate the ReliefF weights of each feature, set the number of iterations and the number of neighbors, and evaluate the correlation between the features and the categories through mutual information. Perform dimensionality reduction using PCA, and retain the principal components with a cumulative variance contribution rate ≥ 95%. Construct a multi-modal feature fusion space, and standardize the spectral reflectance, gray-level co-occurrence matrix texture, and morphological structure features. Optimize the feature combination through cross-validation, and screen out the optimal feature subset.
9. The method for extracting individual trees based on multi - spectral data according to claim 8, characterized in that: The evaluation of the error transfer effect includes: a. Define a set of radiation correction quality evaluation metrics, including root mean square error, mean absolute error, structural similarity, and spectral angle matching, and evaluate the radiation fidelity of the corrected spectral data through multi-metric fusion. b. Construct an error transfer model, analyze the error contribution degree of each processing link based on the error propagation theory, and quantify the influence weight of the errors in each processing link on the final single-tree extraction result. Among them, the processing links include atmospheric correction, terrain compensation, and feature fusion. c. Design a cross-validation experiment, divide the multi-spectral dataset into a training set, a validation set, and a test set, and calculate the confidence interval of the error transfer path through Monte Carlo simulation; d. Establish a multi-modal data fusion rule, standardize and fuse the error-verified spectral data with texture and structural features, construct a feature space using principal component analysis, and achieve hierarchical extraction of the single-tree multi-modal dataset through K-means clustering; Among them, the error transfer model derives the covariance matrix of the processing link errors through the chain rule and determines the key error sources in combination with sensitivity analysis; the multi-modal dataset includes the multi-spectral reflectance after radiometric correction, the texture parameters of the gray-level co-occurrence matrix, and the morphological structure features.
10. An individual tree extraction system based on multi - spectral data, characterized in that, The system includes: Radiometric correction module: Construct an end-to-end radiometric correction system based on a conditional generative adversarial network. Through the collaborative training of an atmospheric parameter generator, a terrain shadow discriminator, and a joint discriminator, eliminate the atmospheric interference and terrain shadows in the multi-spectral data and output general spectral data; Feature fusion and segmentation module: Design a multi-scale attention fusion network, dynamically fuse the general spectral data with the texture features of the gray-level co-occurrence matrix, achieve accurate segmentation of the single-tree crown, and generate an independent single-tree dataset; Geometric correction module: Use the SuperPoint feature detection combined with the SuperGlue matching algorithm to establish corresponding point pairs, and dynamically select an interpolation method based on the terrain complexity for geometric distortion correction to form a spatially aligned dataset; Data enhancement module: Construct an attention-enhanced denoising network, combine non-local mean filtering and adaptive median filtering for noise suppression, and enhance the target edge features through channel attention and spatial attention mechanisms; Feature selection module: Establish a ReliefF-PCA feature selection framework, and screen the optimal combination of spectral-texture-structural features through mutual information evaluation and principal component analysis; Error evaluation module: Construct an error transfer evaluation model, quantify the error contribution degree of each processing link, and generate a single-tree multi-modal dataset after cross-validation.
Citation Information
Patent Citations
Remote sensing image atmospheric terrain geometry joint correction method based on radiation transfer model
CN117372264A
Hyperspectral image lithology identification method and system based on object-oriented spatial spectrum enhanced convolutional neural network model
CN118298313A
Unmanned aerial vehicle multi-dimensional space area measurement system
CN118602997A
Tree symptom detection method based on airborne multispectral and visible light images of unmanned aerial vehicle
CN119007022A
Rock weathering intelligent detection method based on multispectral imaging technology
CN120014371A
Cited By
Urban green land carbon sink dynamic monitoring method and system based on multi-source remote sensing data
CN120853010A