An ultra-high resolution vegetation canopy height estimation method
By combining deep residual convolutional neural networks and feature pyramid networks with Transformer methods, the problems of edge information loss and category uncertainty in vegetation canopy height estimation are solved, and higher accuracy vegetation canopy height estimation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SUN YAT SEN UNIV
- Filing Date
- 2023-12-18
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies suffer from the loss of edge information, uncertainty in the inclusion of category information and non-boundary related information in the features, resulting in low accuracy of the estimation results. This is especially true in developed urban areas where vegetation distribution is irregular and greatly affected by environmental factors, with complex differences in spectral reflectance, large intra-class differences and small inter-class differences.
A deep residual convolutional neural network is used as the encoder, combined with a feature pyramid network (FPN) and a Transformer as the decoder. A high-resolution vegetation canopy height estimation model is trained using multi-source remote sensing data. Low-, medium-, and high-level features of the image are obtained through nonlinear operations and multiple downsampling, and effective feature selection and fusion are performed.
It improved the overall accuracy of vegetation canopy height prediction, enhanced the model's self-learning ability of image features, reduced redundant features, and improved the accuracy of estimation results.
Smart Images

Figure CN117726943B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent vegetation monitoring technology, and more specifically, to an ultra-high resolution method for estimating vegetation canopy height. Background Technology
[0002] Vegetation plays a vital role in maintaining the basic ecosystem functions of Earth's climate and biosphere, providing a wide range of ecosystem services, and mitigating greenhouse gas emissions and the risks of climate change. Canopy height, a simple yet crucial factor in forest vertical structure, reflects vegetation growth capacity, niche requirements, and forest biomass. Higher canopies occupy more favorable niches, receive more sunlight, and promote photosynthesis. Automatically acquiring canopy height information from remote sensing data is essential for estimating aboveground biomass and timber production, monitoring the impacts of forest degradation, measuring the success of forest restoration, and modeling other key ecosystem variables such as primary productivity and biodiversity.
[0003] Existing research has shown that deep convolutional neural networks (CNNs) can achieve excellent results in processing remote sensing images, such as scene classification and object detection and prediction. CNNs can not only automatically learn low-level and mid-level features of images, but also automatically learn high-level semantic features from the original images. Feature Pyramid Networks (FPNs), a variant of deep convolutional neural networks, has become the mainstream framework for multi-feature layer prediction since its introduction. FPNs are typically top-down neural networks that can predict each pixel simultaneously during image segmentation.
[0004] Nevertheless, FPN-based vegetation canopy prediction methods still have the following issues that need to be properly addressed, including:
[0005] (1) By using CNNs as their encoders for image feature extraction, the output contains high-level semantic features, but it is too coarse and easily loses edge details of the image.
[0006] (2) Although passing low-level features to the FPN decoder by “skipping” connections or by taking advantage of the maximum position of the max pooling layer can optimize the prediction results, this approach is prone to generating redundant features, which reduces the learning efficiency of the network.
[0007] In addition, the output features often contain class uncertainty or non-boundary related information, which can affect the optimization of the estimation results.
[0008] 1) In most scenarios, especially in developed urban areas, vegetation is irregularly distributed and affected by other environmental factors. Its spectral reflectance varies greatly and it is easily blocked by the shadows of surrounding high-rise buildings.
[0009] 2) The large intra-class differences and small inter-class differences in high-resolution remote sensing images make the spectral and geometric features of vegetation more complex. Summary of the Invention
[0010] To overcome the shortcomings of the prior art, such as the loss of edge information and the inclusion of category uncertainty or non-boundary related information in the output features, which leads to low accuracy of the estimation results, this invention provides an ultra-high resolution vegetation canopy height estimation method.
[0011] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0012] Firstly, an ultra-high resolution vegetation canopy estimation method includes:
[0013] Acquire ultra-high resolution aerial RGB imagery data, LiDAR point cloud data, Copernicus digital elevation model data, and synthetic aperture radar satellite data for the sample area, and combine them into the first multi-source dataset;
[0014] The first multi-source dataset is preprocessed to obtain the second multi-source dataset;
[0015] A deep residual convolutional neural network was used as the encoder, and FPN and Transformer were used as the decoders to establish an initial ultra-high resolution vegetation canopy height estimation model. The ultra-high resolution vegetation canopy height estimation model was then trained using the second multi-source dataset.
[0016] Acquire vegetation images of the target area, and use the ultra-high resolution vegetation canopy height estimation model to obtain vegetation canopy height information of the target area.
[0017] Secondly, an ultra-high resolution vegetation canopy height estimation system, applying the method described in the first aspect, includes:
[0018] The multi-source data acquisition module is used to acquire ultra-high resolution aerial RGB imagery data, LiDAR point cloud data, Copernicus digital elevation model data, and synthetic aperture radar satellite data for the sample area, and combine them into a first multi-source dataset; it is also used to preprocess the first multi-source dataset to obtain a second multi-source dataset; and it is also used to acquire vegetation images for the target area.
[0019] The model training module is used to establish an initial ultra-high resolution vegetation canopy height estimation model with a deep residual convolutional neural network as the encoder and FPN and Transformer as the decoder, and to train the ultra-high resolution vegetation canopy height estimation model using the second multi-source dataset.
[0020] The vegetation canopy height estimation module is used to carry the trained ultra-high resolution vegetation canopy height estimation model; it is also used to obtain the vegetation canopy height information of the target area using the ultra-high resolution vegetation canopy height estimation model based on the vegetation image of the target area.
[0021] Thirdly, a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the method as described in the first aspect.
[0022] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0023] This invention discloses an ultra-high resolution canopy height estimation method. It establishes an ultra-high resolution vegetation canopy height estimation model using a deep residual convolutional neural network as the encoder and an FPN and Transformer as the decoder. The model is trained using multi-source remote sensing data, enabling it to possess image feature self-learning capabilities. It obtains low-, medium-, and high-level features of the image through nonlinear operations and multiple downsampling, and uses an FT module (FPN and Transformer) to filter and fuse effective features. Compared to existing technologies, this invention improves the overall accuracy of vegetation canopy height prediction. Attached Figure Description
[0024] Figure 1 This is a flowchart illustrating an ultra-high resolution canopy height estimation method according to Embodiment 1 of the present invention;
[0025] Figure 2 This is a schematic diagram of the data processing flow of the ultra-high resolution vegetation canopy height estimation model in Embodiment 1 of the present invention;
[0026] Figure 3 This is a schematic diagram of the structure of an ultra-high resolution canopy height estimation system in Embodiment 2 of the present invention. Detailed Implementation
[0027] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0028] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent.
[0029] To better illustrate this embodiment, some parts in the accompanying drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions;
[0030] It will be understood by those skilled in the art that certain well-known structures and their descriptions may be omitted in the accompanying drawings.
[0031] To facilitate implementation by those skilled in the art, some concepts involved in the embodiments of the present invention are explained as follows:
[0032] 1. Deep Residual Convolutional Neural Network (ResNet)
[0033] ResNet is a deep convolutional neural network that typically includes convolutional modules (convolutional layers, batch layers, ReLU activation layers, and max pooling layers), four structurally similar residual modules, and a classification module (mean pooling layer, fully connected layer, and Softmax classification layer).
[0034] After the input image passes through the first convolutional module, its size is reduced to half. Each time it passes through a residual module, the image size is also reduced to half, ultimately resulting in an image of the original size. Figure 1 / 32 Feature map
[0035] Specifically, based on network depth, there are ResNet18, ResNet34, ResNet50, ResNet101, and ResNet152.
[0036] For ResNet50, ResNet101, and ResNet152, each residual module consists of a projection shortcut and multiple consecutive identity shortcuts. The projection shortcut can double the number of feature maps and reduce the feature maps by half, while the identity shortcut does not change the input / output size or the number of features.
[0037] Furthermore, a convolutional layer is a feature extractor that learns features to represent the input image. A convolutional layer consists of several convolutional units (neurons), with each feature plane composed of multiple neurons. Neurons on the same feature plane share weights, and the parameters of each unit are optimized and calculated using the network's backpropagation algorithm. Its function is to obtain different features of the input, such as edges, right angles, and texture information. Given a feature map X... l-1 As input to the convolutional layer, the k-th filter is used The input feature map is processed by equation (1) to obtain the output feature map:
[0038]
[0039] In the formula, where This is the feature map obtained after the convolution operation; * represents the convolution operation. It is the k-th bias vector of layer l. Formula (1) can greatly reduce the number of neural network parameters while obtaining the feature results corresponding to each neuron.
[0040] The purpose of batch normalization / batch processing (BN) layers is to prevent gradient vanishing or gradient exploding in neural networks. The normalization process for each batch of input in a BN layer is as follows:
[0041]
[0042] in, The result after scaling and shifting, γ l It is the normalized scaling parameter, β l It is the offset parameter. Through the normalization process of equation (2), all inputs can be concentrated near 0, so that the inputs of each layer will not change too much.
[0043] The activation function layer is used to control the activation level of neurons that perform positive signal transformation. Using the result from the Batch Normalization (BN) layer as input, the Modified Linear Unit (ReLU) activation function is employed to perform a non-linear mapping of the input features.
[0044] Pooling layers are primarily used to abstract input features, typically employing max pooling or average pooling to obtain downsampled feature maps. The purpose of pooling layers is to reduce the feature surface area, simplify model complexity, reduce model parameters, and achieve space invariance. Pooling layers follow convolutional layers and also consist of multiple feature surfaces, each corresponding to a feature surface in the previous layer, without changing the total number of feature surfaces. There are typically two forms: mean pooling and max pooling. In mean pooling, each weight in the convolutional kernel is 0.25. If the stride of the convolutional kernel on the input image (Input X) is 2, mean pooling effectively blurs and reduces the original image to 1 / 4 of its original size. In max pooling, only one weight in the convolutional kernel is 1, while the rest are 0. The position of 1 in the convolutional kernel corresponds to the position of the largest value in the portion of the input image covered by the convolutional kernel. If the convolution kernel has a stride of 2 on the input image and a size of 2*2, the effect of maximum sampling is to reduce the original image to 1 / 4 of its original size while retaining the strongest input in each 2*2 region.
[0045] 2. Copernicus Digital Elevation Model Data
[0046] The Copernicus Digital Elevation Model (DEM) is recognized as one of the best global open-source DEMs. Its absolute elevation and horizontal accuracy are the best among global open-source DEMs, while it also has the best terrain detail representation. Its actual effect is better than 30m resolution, and it has good consistency.
[0047] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0048] Example 1
[0049] This embodiment proposes an ultra-high resolution method for estimating vegetation canopy height. (See reference...) Figure 1 ,include:
[0050] Acquire ultra-high resolution aerial RGB imagery data, LiDAR point cloud data, Copernicus digital elevation model data, and synthetic aperture radar satellite data for the sample area, and combine them into the first multi-source dataset;
[0051] The first multi-source dataset is preprocessed to obtain the second multi-source dataset;
[0052] An initial ultra-high resolution vegetation canopy height estimation model was established using a deep residual convolutional neural network as the encoder and FPN and Transformer as the decoders. The ultra-high resolution vegetation canopy height estimation model was then trained using the second multi-source dataset.
[0053] Acquire vegetation images of the target area, and use the ultra-high resolution vegetation canopy height estimation model to obtain vegetation canopy height information of the target area.
[0054] This embodiment establishes an ultra-high resolution vegetation canopy height estimation model (ARFTNet) with an encoder-decoder structure based on deep learning technology. Specifically, the encoder consists of a deep residual convolutional neural network (ResNet) used to learn low-, medium-, and high-level features of the input image. The decoder includes an FPN and a Transformer for filtering and fusing effective features. The model is trained using multi-source remote sensing data to ultimately generate high-quality vegetation canopy height information. Compared to existing technologies, using the ultra-high resolution vegetation canopy height estimation model described in this embodiment can improve the overall accuracy of vegetation canopy height prediction results.
[0055] In some examples, the synthetic aperture radar satellite data is Sentinel-1 data with a spatial resolution of 10m, which can provide high-resolution radar imagery to characterize surface features and help remove the influence of surface features on vegetation canopy height estimation.
[0056] It should be noted that the vegetation images of the target area include ultra-high resolution aerial RGB image data and LIDAR point cloud data of the target area.
[0057] In some examples, the sample region is part of the target region;
[0058] In a specific implementation process, one-third of the ultra-high resolution aerial RGB imagery data and LIDAR point cloud data of the target area are randomly selected as model training data.
[0059] It should also be noted that the ultra-high resolution mentioned in this embodiment refers to a spatial resolution of 1m.
[0060] In some preferred embodiments, acquiring ultra-high resolution aerial RGB imagery data, LiDAR point cloud data, Copernicus digital elevation model data, and synthetic aperture radar satellite data for the sample area includes:
[0061] Image control points are deployed using a regional network deployment method;
[0062] Based on the preset flight path, the initial aerial RGB image data and initial LIDAR point cloud data of the sample area were collected by aerial photography.
[0063] Based on the control points, the initial ultra-high resolution aerial RGB image data and the initial LIDAR point cloud data are subjected to control point encryption, aerial triangulation calculation and accuracy monitoring, and the exterior orientation elements of each image are calculated at the same time to obtain aerial triangulation encrypted ultra-high resolution aerial RGB image data and LIDAR point cloud data.
[0064] Acquire Copernicus Digital Elevation Model data and Synthetic Aperture Radar satellite data, and combine them with the ultra-high resolution aerial RGB imagery data and the LIDAR point cloud data to form the first multi-source dataset.
[0065] This preferred embodiment uses image control point deployment and aerial triangulation to correct the data collected in the field, ensuring the accuracy of the data collected in the field.
[0066] In a specific implementation process, when setting up image control points (ADCs), an irregular area network should be used when terrain or other conditions restrict their placement. ADCs should be placed at concave or convex angles. The span of ADCs along the flight path should be based on 12 baselines. In special circumstances such as when the principal point or standard point is overwatered, or in waterfront or island areas, ADCs cannot be set up as normally. Instead, they should be set up according to the specific circumstances, prioritizing the requirements of aerial triangulation and mapping. VRS mode is the preferred method for ADC measurement. When network RTK services are unavailable, GPS-based static measurement methods should be used. All ADC point marking and finishing work is done on the original image data, using Photoshop software to add finishing information to the photographic data.
[0067] In a specific implementation process, when laying flight lines, the reference plane height is determined based on the terrain and flight safety of the aerial photography area, and the spacing and direction of flight lines are set according to the basic parameter table of aerial photography technology. Flight line generation software is then used to automatically generate flight lines.
[0068] In a specific implementation process, when conducting aerial triangulation, aerial triangulation is performed indoors using aerial image data and field control point files, based on the principles of photogrammetry. This includes control point densification, aerial triangulation calculations, and accuracy monitoring. Simultaneously, the exterior orientation elements of each image are calculated, ultimately outputting ultra-high resolution aerial RGB image data and LIDAR point cloud data that have undergone aerial triangulation densification.
[0069] In a specific implementation process, when using flight equipment for aerial photography, the flight should be as stable as possible, and the yaw angle and yaw angle should not exceed the standard requirements. The field ultra-high resolution aerial RGB image data and airborne LIDAR point cloud data should be collected in accordance with the preset flight path and aerial photography quality requirements.
[0070] In some alternative embodiments, the second multi-source dataset includes canopy height model data (CHM), digital orthophoto data, resampled Copernicus digital elevation model data, and resampled synthetic aperture radar satellite data;
[0071] The data preprocessing of the first multi-source dataset includes:
[0072] For the LIDAR point cloud data:
[0073] The noise points, ground points and non-ground points in the LIDAR point cloud data are filtered respectively, and then the digital elevation model data and digital surface model data are extracted by natural neighborhood interpolation.
[0074] The canopy height model data (CHM) is obtained based on the difference between the digital elevation model data (DEM) and the digital surface model data (DSM).
[0075] For the aforementioned ultra-high resolution aerial RGB imagery data in the field:
[0076] Based on the Copernicus digital elevation model data and aerial triangulation results, the ultra-high resolution aerial RGB image data in the field is resampled, and the center projection is converted into orthophoto projection according to the error correction of each image and each pixel to obtain digital orthophoto image data.
[0077] Regarding the Copernicus digital elevation model data and the synthetic aperture radar satellite data:
[0078] Spatial interpolation is used to resample the Copernicus digital elevation model data and the synthetic aperture radar satellite data so that the resolution of the Copernicus digital elevation model data and the synthetic aperture radar satellite data is the same as that of the ultra-high resolution aerial RGB image data in the field.
[0079] In this optional embodiment, by correcting the field ultra-high resolution aerial RGB image data frame by frame and pixel, errors caused by factors such as ground undulation and tilt of the aerial photography device can be reduced.
[0080] In some examples, 1m resolution digital elevation model (DEM), digital surface model (DSM), and canopy height model (CHM) data are generated based on acquired airborne LiDAR point cloud data.
[0081] In some examples, the filtering of noise points, ground points and non-ground points in the LIDAR point cloud data includes: separating points or point groups that are significantly below the ground (low points) and points or point groups that are significantly above the ground target (high points); and filtering the ground points and non-ground points in the point cloud.
[0082] In some examples, when producing digital orthophoto data, those skilled in the art may refer to the "Digital Orthophoto Maps of Basic Geographic Information 1:500, 1:1000, and 1:2000".
[0083] In some examples, bilinear interpolation is used to resample Copernicus digital elevation model data and synthetic aperture radar satellite data to generate smoother surfaces. The raster value of a pixel is obtained by interpolating once in the Y direction (or X direction) and then once in the X direction (or Y direction) and weighting the distance from the sampling point to its four neighboring pixels. Based on the resolution of the aerial RGB imagery, the Copernicus digital elevation model data and synthetic aperture radar satellite data are resampled to the same resolution as the aerial RGB imagery.
[0084] Further, training the ultra-high resolution vegetation canopy height estimation model using the second multi-source dataset includes:
[0085] The canopy height model data, resampled Copernicus digital elevation model data, and resampled synthetic aperture radar satellite data are overlaid with the red, green, and blue bands of the digital orthophoto image data to obtain the original feature combination image.
[0086] The canopy height model data is used as a label image and combined with the original feature combination image to form an image pair;
[0087] After performing supervised data augmentation on the image pairs, they are divided into training and validation sets.
[0088] The training set is used as input to the encoder, allowing the ultra-high resolution vegetation canopy height estimation model to learn multi-level features of the image on the training set, and the parameters of the ultra-high resolution vegetation canopy height estimation model are tuned using the validation set.
[0089] In this embodiment, the canopy height model data generated based on LIDAR point cloud data is used as the label data for model training. By performing supervised data augmentation, the soil relative input to the model is given a different combination each time, which increases the diversity of the training sample dataset, avoids overfitting in network training, and increases the generalization ability of the model.
[0090] In some examples, the training and validation sets are divided in a 70%:30% ratio.
[0091] In some examples, the labeled image and the original feature combination image are cropped into 128*128 images respectively to form an image pair.
[0092] In some examples, during the training of the ultra-high resolution vegetation canopy height estimation model, the base learning rate is set to 0.001, the learning rate decay rate is set to 0.7, the learning rate is adaptively updated using the Adam stochastic optimization method, and the maximum number of iterations is set to 800.
[0093] Furthermore, the data augmentation includes geometric transformations, adding noise, padding, erasing, and / or forming data pairs (SamplePairing).
[0094] In some optional embodiments, when acquiring initial aerial RGB image data and initial LIDAR point cloud data using aerial photography, flight path quality control is performed according to the following requirements: the forward coverage of the photographic area boundary must exceed the boundary line of the photographic area by 3 baselines; the lateral coverage must exceed the boundary line of the photographic area by 50% of the image frame; the forward overlap must be no less than 70% and the lateral overlap must be no less than 55%.
[0095] In some preferred embodiments, the establishment of an initial ultra-high resolution vegetation canopy height estimation model is described in reference to... Figure 2 ,include:
[0096] An encoder is built based on any one of the deep residual convolutional neural networks, Resnet50, Resnet101, and Resnet152. The classification module of the deep residual convolutional neural network is replaced with a convolutional layer with 256 output features to learn multi-level features of the input image and output high-level feature results about the height of the vegetation canopy.
[0097] The number of input features of the first convolutional module of the deep residual convolutional neural network is set to 64, and a 3*3 convolutional module with 64 output features and constant size is embedded before the deep residual convolutional neural network to receive the input image.
[0098] A decoder is built based on two FPN modules, two Transformer modules, and multiple upsampling layers. It is used to output the canopy height prediction result as the output of the ultra-high resolution vegetation canopy height estimation model, based on the feature results extracted by different residual modules and the high-level feature results.
[0099] It should be noted that in this preferred embodiment, by performing a 2x upsampling operation on the image through multiple upsampling layers, a canopy height prediction result of the same size as the original input image can be obtained.
[0100] In one specific implementation, an encoder is built based on ResNet101 to learn multi-level features of the image. Multiple convolutional or max-pooling operations are used to obtain a high-dimensional image (i.e., a high-level feature extraction result) that is 1 / 32 the size of the original image. Compared to ResNet101, ResNet50 has a smaller depth, and the extracted deep features are not as obvious, while the ResNet152 model has more network layers and requires more computational resources.
[0101] It should be noted that a new convolutional module is embedded before the deep residual convolutional neural network, enabling the ultra-high resolution vegetation canopy height estimation model to accept multi-band (not limited to three-band) image inputs. Its output features number 64 and are the same size as the input image. The improved encoder obtains low-to-medium-to-high-level features from the input image through nonlinear operations and multiple downsampling.
[0102] Those skilled in the art will understand that the Transformer module uses a self-attention mechanism to learn the encoding and representation of the input sequence, capturing the dependencies between different positions in the sequence.
[0103] In some alternative embodiments, the step of building a decoder based on two FPN modules, two Transformer modules, and multiple upsampling layers includes:
[0104] After the second and third residual modules of the deep residual convolutional neural network, the first FPN module and the second FPN module are connected to perform convolution operations, and the high-level feature results output by the encoder are upsampled and then added and upsampled with the output results of the two FPN modules in sequence to obtain a feature result with 256 features.
[0105] Set up a first Transformer module and a second Transformer module, whose inputs are the upsampled results after passing through the first FPN module and the upsampled results after passing through the second FPN module, respectively; after upsampling the high-level feature results output by the encoder, overlay them sequentially with the outputs of the two Transformer modules; set upsampling layers, convolutional layers and ReLU activation layers to process the overlay results to obtain a canopy height prediction result of the same size as the input image.
[0106] Those skilled in the art should understand that, in this optional embodiment, the decoder section, for the high-level feature results output by the encoder section, first performs convolution operations on the feature results obtained from different residual modules through the FPN module to obtain 256 band feature results, and then adds and upsamples the result after upsampling with the output result of the improved encoder; the Transformer module optimizes the high-level feature results processed by the FPN module; then, the high-level feature results processed by the Transformer module are superimposed with the upsampled result; finally, the superimposed result is upsampled and passed through a convolutional layer and a ReLU activation layer to obtain a canopy height prediction result with the same size as the input image. It should be understood that when predicting the vegetation image of the target area, its output is the vegetation canopy height information of the target area.
[0107] In a specific implementation process, a test set is used to compare and evaluate the accuracy of the aforementioned ultra-high resolution vegetation canopy height estimation model (ARFTNet) with existing schemes UNet and SegNet. Specifically, the coefficient of determination (R²) is used. 2 The four accuracy evaluation indicators are root mean square error (RMSE), mean absolute error (MAE), and deviation (B). The experimental data are shown in Table 1.
[0108] Table 1 Comparative experimental data
[0109]
[0110] R 2 This system can evaluate and validate estimated canopy height and represent the confidence level of predicted tree height. RMSE and MAE are used to represent the deviation between the estimated canopy height and the actual canopy height in the test set. The bias value is used to determine how much the model-estimated canopy height is overestimated or underestimated relative to the actual canopy height in the test set. Specifically, R... 2 Higher values indicate better model estimation results, while lower values for RMSE, MAE, and Bias indicate that the model estimation results are closer to the actual results.
[0111] It can be seen that ARFTNet's RMSE is lower than that of UNET and SegNet, and its R 2 Its performance is higher than UNET and SegNet, but its MAE is lower than UNET and SegNet, and its absolute value of Bias is also lower than UNET and SegNet. Overall, the ultra-high resolution vegetation canopy height estimation model described in this embodiment performs better and has higher accuracy.
[0112] Example 2
[0113] This embodiment proposes an ultra-high resolution vegetation canopy height estimation system, applying the method described in Embodiment 1, see reference. Figure 3 ,include:
[0114] The multi-source data acquisition module is used to acquire ultra-high resolution aerial RGB imagery data, LiDAR point cloud data, Copernicus digital elevation model data, and synthetic aperture radar satellite data for the sample area, and combine them into a first multi-source dataset; it is also used to preprocess the first multi-source dataset to obtain a second multi-source dataset; and it is also used to acquire vegetation images for the target area.
[0115] The model training module is used to establish an initial ultra-high resolution vegetation canopy height estimation model with a deep residual convolutional neural network as the encoder and FPN and Transformer as the decoder, and to train the ultra-high resolution vegetation canopy height estimation model using the second multi-source dataset.
[0116] The vegetation canopy height estimation module is used to carry the trained ultra-high resolution vegetation canopy height estimation model; it is also used to obtain the vegetation canopy height information of the target area based on the vegetation image of the target area using the ultra-high resolution vegetation canopy height estimation model.
[0117] It is understood that the system in this embodiment corresponds to the method in Embodiment 1 above, and the options in Embodiment 1 above are also applicable to this embodiment, so they will not be described again here.
[0118] Example 3
[0119] This embodiment proposes a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor, causing the processor to perform some or all of the steps of the method described in Embodiment 1.
[0120] It is understood that the storage medium can be transient or non-transient. Exemplarily, the storage medium includes, but is not limited to, various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0121] By way of example, the processor may be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0122] In some examples, a computer program product is provided, which can be implemented by hardware, software, or a combination thereof. As a non-limiting example, the computer program product can be embodied in the storage medium, or it can be embodied in a software product, such as an SDK (Software Development Kit).
[0123] In some examples, a computer program is provided, including computer-readable code, wherein, when the computer-readable code is run in a computer device, a processor in the computer device performs some or all of the steps for implementing the method.
[0124] This embodiment also proposes an electronic device, including a memory and a processor. The memory stores at least one instruction, at least one program, code set, or instruction set. When the processor executes the at least one instruction, at least one program, code set, or instruction set, it implements some or all of the steps of the method described in Embodiment 1.
[0125] In some examples, a hardware entity of the electronic device is provided, including: a processor, a memory, and a communication interface; wherein the processor typically controls the overall operation of the electronic device; the communication interface is used to enable the electronic device to communicate with other terminals or servers via a network; the memory is configured to store instructions and applications executable by the processor, and may also cache data to be processed or already processed (including but not limited to image data, audio data, voice communication data, and video communication data) to be processed by the processor and various modules in the electronic device, and may be implemented using flash memory or random access memory (RAM).
[0126] Furthermore, data can be transferred between the processor, communication interface, and memory via a bus, which can include any number of interconnected buses and bridges, connecting various circuits of one or more processors and memories together.
[0127] It is understood that the options in Embodiment 1 above also apply to this embodiment, so they will not be described again here.
[0128] The terms used to describe positional relationships in the accompanying drawings are for illustrative purposes only and should not be construed as limiting this patent.
[0129] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. It should be understood that in the various embodiments of this disclosure, the sequence number of each step / process does not imply the order of execution. The execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments. It should also be understood that the system / device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms. For those skilled in the art, other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the claims of the present invention.
Claims
1. A method for estimating vegetation canopy height at ultra-high resolution, characterized in that, include: Acquire ultra-high resolution aerial RGB imagery data, LiDAR point cloud data, Copernicus digital elevation model data, and synthetic aperture radar satellite data for the sample area, and combine them into the first multi-source dataset; The first multi-source dataset is preprocessed to obtain the second multi-source dataset; An initial ultra-high resolution vegetation canopy height estimation model was established using a deep residual convolutional neural network as the encoder and FPN and Transformer as the decoders. The ultra-high resolution vegetation canopy height estimation model was then trained using the second multi-source dataset. Acquire vegetation images of the target area, and use the ultra-high resolution vegetation canopy height estimation model to obtain vegetation canopy height information of the target area; The establishment of the initial ultra-high resolution vegetation canopy height estimation model includes: An encoder is built based on any one of the deep residual convolutional neural networks, Resnet50, Resnet101, and Resnet152. The classification module of the deep residual convolutional neural network is replaced with a convolutional layer with 256 output features to learn multi-level features of the input image and output high-level feature results about the height of the vegetation canopy. The number of input features of the first convolutional module of the deep residual convolutional neural network is set to 64, and a 3*3 convolutional module with 64 output features and constant size is embedded before the deep residual convolutional neural network to receive the input image. A decoder is built based on two FPN modules, two Transformer modules, and multiple upsampling layers. It is used to output the canopy height prediction result as the output of the ultra-high resolution vegetation canopy height estimation model based on the feature results extracted by different residual modules and the high-level feature results. The decoder, built upon two FPN modules, two Transformer modules, and multiple upsampling layers, includes: After the second and third residual modules of the deep residual convolutional neural network, the first FPN module and the second FPN module are connected to perform convolution operations, and the high-level feature results output by the encoder are upsampled and then added and upsampled with the output results of the two FPN modules in sequence to obtain a feature result with 256 features. Set up a first Transformer module and a second Transformer module, whose inputs are the upsampled results after passing through the first FPN module and the upsampled results after passing through the second FPN module, respectively; after upsampling the high-level feature results output by the encoder, overlay them sequentially with the outputs of the two Transformer modules; set upsampling layers, convolutional layers and ReLU activation layers to process the overlay results to obtain a canopy height prediction result of the same size as the input image.
2. The ultra-high resolution vegetation canopy height estimation method according to claim 1, characterized in that, The acquisition of ultra-high resolution aerial RGB imagery data, LiDAR point cloud data, Copernicus digital elevation model data, and synthetic aperture radar satellite data for the sample area includes: Image control points are deployed using a regional network deployment method; Based on the preset flight path, the initial aerial RGB image data and initial LIDAR point cloud data of the sample area were collected by aerial photography. Based on the control points, the initial ultra-high resolution aerial RGB image data and the initial LIDAR point cloud data are subjected to control point encryption, aerial triangulation calculation and accuracy monitoring, and the exterior orientation elements of each image are calculated at the same time to obtain aerial triangulation encrypted ultra-high resolution aerial RGB image data and LIDAR point cloud data. Acquire Copernicus Digital Elevation Model data and Synthetic Aperture Radar satellite data, and combine them with the ultra-high resolution aerial RGB imagery data and the LIDAR point cloud data to form the first multi-source dataset.
3. The ultra-high resolution vegetation canopy height estimation method according to claim 2, characterized in that, The second multi-source dataset includes canopy height model data, digital orthophoto data, resampled Copernicus digital elevation model data, and resampled synthetic aperture radar satellite data; The data preprocessing of the first multi-source dataset includes: For the LIDAR point cloud data: The noise points, ground points and non-ground points in the LIDAR point cloud data are filtered respectively, and then the digital elevation model data and digital surface model data are extracted by natural neighborhood interpolation. The canopy height model data is obtained based on the difference between the digital elevation model data and the digital surface model data; For the aforementioned ultra-high resolution aerial RGB imagery data in the field: Based on the Copernicus digital elevation model data and aerial triangulation results, the ultra-high resolution aerial RGB image data in the field is resampled, and the center projection is converted into orthophoto projection according to the error correction of each image and each pixel to obtain digital orthophoto image data. Regarding the Copernicus digital elevation model data and the synthetic aperture radar satellite data: Spatial interpolation is used to resample the Copernicus digital elevation model data and the synthetic aperture radar satellite data so that the resolution of the Copernicus digital elevation model data and the synthetic aperture radar satellite data is the same as that of the ultra-high resolution aerial RGB image data in the field.
4. The ultra-high resolution vegetation canopy height estimation method according to claim 3, characterized in that, The step of training the ultra-high resolution vegetation canopy height estimation model using the second multi-source dataset includes: The canopy height model data, resampled Copernicus digital elevation model data, and resampled synthetic aperture radar satellite data are overlaid with the red, green, and blue bands of the digital orthophoto image data to obtain the original feature combination image. The canopy height model data is used as a label image and combined with the original feature combination image to form an image pair; After performing supervised data augmentation on the image pairs, they are divided into training and validation sets. The training set is used as input to the encoder, allowing the ultra-high resolution vegetation canopy height estimation model to learn multi-level features of the image on the training set, and the parameters of the ultra-high resolution vegetation canopy height estimation model are tuned using the validation set.
5. The ultra-high resolution vegetation canopy height estimation method according to claim 4, characterized in that, The data augmentation includes geometric transformations, adding noise, padding, erasing, and forming data pairs.
6. The ultra-high resolution vegetation canopy height estimation method according to claim 2, characterized in that, When acquiring initial aerial RGB image data and initial LiDAR point cloud data using aerial photography, flight path quality control is performed according to the following requirements: the forward coverage of the photographic area boundary must exceed the boundary line by 3 baselines; the lateral coverage must exceed the boundary line by 50% of the image frame; the forward overlap must be no less than 70%; and the lateral overlap must be no less than 55%.
7. A high-resolution vegetation canopy height estimation system, using the method described in any one of claims 1-6, characterized in that, include: The multi-source data acquisition module is used to acquire ultra-high resolution aerial RGB imagery data, LiDAR point cloud data, Copernicus digital elevation model data, and synthetic aperture radar satellite data for the sample area, and combine them into a first multi-source dataset; it is also used to preprocess the first multi-source dataset to obtain a second multi-source dataset; and it is also used to acquire vegetation images for the target area. The model training module is used to establish an initial ultra-high resolution vegetation canopy height estimation model with a deep residual convolutional neural network as the encoder and FPN and Transformer as the decoder, and to train the ultra-high resolution vegetation canopy height estimation model using the second multi-source dataset. The vegetation canopy height estimation module is used to carry the trained ultra-high resolution vegetation canopy height estimation model; it is also used to obtain the vegetation canopy height information of the target area based on the vegetation image of the target area using the ultra-high resolution vegetation canopy height estimation model.
8. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, at least one program, code set, or instruction set, which is loaded and executed by a processor to implement the method as described in any one of claims 1-6.