A canopy height estimation method based on unmanned aerial vehicle remote sensing images
Patent Information
- Application Number
- CN202610944140.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-25
AI Technical Summary
为了解决解码器中全局信息不足的问题,同时需要对编码器输出的多尺度特征提取全局特征
[0041]本发明的有益效果是:本发明提出了一种基于无人机遥感图像的树冠高度估计方法,选取无人机遥感数据集作为实验数据集;基于所述实验数据集,选取视觉变换器,通过所述视觉变换器从所述实验数据集的遥感图像中提取多尺度特征图像;基于所述多尺度特征图像,构建特征金字塔提取所述多尺度特征图像中的全局特征图像;基于所述多尺度特征图像和全局特征图像,构建串联-并联双分支特征解码器对所述多尺度特征图像和全局特征图像进行特征融合得到简易冠层高度图;基于所述简易冠层高度图,需要对尺寸和数值进行放缩,最终获得处理后的实验数据集;构建目标损失函数对所述处理后的实验数据集进行预测。
Smart Images

Figure CN122820799A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of canopy height estimation technology, specifically a canopy height estimation method based on UAV remote sensing images. Background Technology
[0002] Canopy height, as a key parameter of forest ecosystem structure, is crucial for a deeper understanding of forest function and its interaction with the environment. It is not only related to forest growth dynamics, carbon storage assessment, and biodiversity monitoring, but also serves as fundamental data for predicting forest responses to climate change. Traditionally, high-precision and high-resolution three-dimensional technologies, such as Light Detection and Ranging (LiDAR) and Synthetic Aperture Radar (SAR), have been used to acquire height information. However, these technologies are costly and their performance is limited under adverse weather conditions; for example, scattering of the laser beam in heavy rain, fog, or snow reduces measurement accuracy. Furthermore, LiDAR is time-consuming and error-prone, while SAR performs poorly when penetrating dense or complex vegetation canopies, limiting their effectiveness in measuring canopy height.
[0003] Therefore, monocular depth estimation technology has become an alternative due to its low cost and high efficiency. This technology extracts depth information from images captured by a single camera, is not limited by lighting or weather conditions, and is suitable for various environments. However, applying monocular depth estimation in natural environments faces the following challenges:
[0004] First, the aerial views provided by drones reveal the diversity of tree morphology, including cylindrical, conical, and semi-ellipsoidal crown shapes, which increases the complexity of canopy height estimation. Second, in dense forests, the shadow areas created by overlapping canopies can lead to misjudgments of depth discontinuity, especially at canopy edges and in areas where height changes rapidly. Furthermore, the significant differences in vegetation cover density and type, ranging from sparse shrubs to dense canopies, further complicate canopy height estimation. This diversity not only complicates the estimation process but also introduces long-tailed distribution problems due to variations in the data environment, significantly impacting the accuracy of the estimation. Summary of the Invention
[0005] To address the aforementioned problems, this invention proposes a tree canopy height estimation method based on UAV remote sensing images. This is an encoder-decoder architecture that integrates a Transformer branch decoder and a CNN branch decoder to construct a cascaded-parallel dual-branch decoder. To address the issue of insufficient global information in the decoder, global features need to be extracted from the multi-scale features of the encoder output. This structure leverages the advantages of convolutional operations for fine-grained local feature extraction, while utilizing the unique ability of the transformer to capture long-range dependencies. The final result is a highly integrated one.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solution:
[0007] Step S1: Select the UAV remote sensing dataset as the experimental dataset;
[0008] Step S2: Based on the experimental dataset, select a visual transformer and extract multi-scale feature images from the remote sensing images in the experimental dataset using the visual transformer;
[0009] Step S3: Based on the multi-scale feature image, construct a feature pyramid to extract the global feature image from the multi-scale feature image;
[0010] Step S4: Based on the multi-scale feature image and the global feature image, a serial-parallel dual-branch feature decoder is constructed to perform feature fusion on the multi-scale feature image and the global feature image to obtain a simplified canopy height map;
[0011] Step S5: Based on the simplified canopy height map, scale the dimensions and values to obtain the processed experimental dataset;
[0012] Step S6: Construct a target loss function to predict the processed experimental dataset.
[0013] A further improvement of the present invention is that: in step S1, after obtaining the experimental dataset, the remote sensing images in the experimental dataset are preprocessed.
[0014] The data preprocessing includes: processing the remote sensing image according to 480... The image was cropped to 640 pixels to match the normal input size of the model. Considering the inconsistency of tree canopy height in different environments, the canopy height image was scaled proportionally, with the maximum value magnified to 255. The magnification factor depends on the dataset.
[0015] A further improvement of this invention lies in: constructing a feature pyramid from the obtained multi-scale feature images and performing global fitting to obtain a global feature image; the expression for constructing the feature pyramid from the multi-scale feature images is:
[0016]
[0017] In the formula, This is a low-latitude feature image; It is a high-dimensional feature image. It is the original high-dimensional feature image before fusion; It's convolution; and These are learnable coefficients; they enable progressive global feature image fusion from low to high dimensions to obtain a preliminary global feature image.
[0018] Considering the robustness of nonlinear transformation, the initial global feature image is processed nonlinearly, and the nonlinear features are fused to obtain the final global feature image, i.e.:
[0019]
[0020] In the formula, It is the final global feature image; It is a preliminary global feature image. and It is an activation function. and It is the learnable coefficient; Functions can enhance the model's ability to distinguish between tree and non-tree regions, while The function can fine-tune ReLU, especially in the shaded areas between trees;
[0021] A further improvement of the present invention is that: a serial-parallel dual-branch feature decoder is constructed based on multi-scale feature images and global feature images; the dual-branch feature decoder includes a Transformer branch and a CNN branch, and the feature fusion capability of the decoder is improved by utilizing the Transformer branch's ability to capture the dependence of global long-distance relationships and the CNN branch's superiority in local feature extraction; a multi-level serial decoder block structure is constructed within the same branch, and a parallel decoder branch structure is constructed between different branches;
[0022] A further improvement of this invention lies in the following: the core structure of the Transformer branch is a gradient-enhanced conditional random field built upon a fully connected conditional random field (FC-CRF); the working principle of FC-CRF is mainly based on Markov random fields and probabilistic undirected graphical models, predicting the conditional probability distribution of the output random variable sequence using a given input random variable sequence, essentially similar to the Swing Transformer coding block; and the encoder features of the same-level output. Perform feature mapping, calculate the attention score matrix for the derived Q and K, and then apply this score matrix to the global features. Weighting is applied; to characterize the spatial relationships between features, FC-CRF introduces relative position embeddings, further improving the spatial sensitivity and accuracy of feature representation, expressed as:
[0023]
[0024]
[0025] In the formula, It is a query vector. and It is the output feature of the corresponding sibling coding block. The feature matrix; Right now The transpose of is used to calculate and The dot product is used to obtain the attention score matrix; It is a relative positional encoding; It is the final global feature image; It is a dot product; to improve the sensitivity to the canopy edge, convolution operations are first performed using kernels of different sizes to extract local features under different receptive fields. Then, gradient information is introduced using the interpolation method for feature processing to obtain a highly abstract composite feature information map. The specific expression is as follows:
[0026]
[0027]
[0028]
[0029]
[0030]
[0031] In the formula, It is the output of FC-CRF; , , , It is an intermediate variable within the Transformer branch; This is the final output of the Transformer branch; This represents a convolution with a kernel size of k; A represents the average; H represents the horizontal gradient operation; V represents the vertical gradient operation. , , and Represents the learnability coefficient; Represents the learnable bias term;
[0032] Considering the local receptive field problem and the Dead ReLU phenomenon, the core structure of the CNN branch is a composite convolutional block based on transposed convolution, Leaky ReLU activation function, and convolution operation, expressed as:
[0033]
[0034] In the formula, These are the output features of the CNN branch; Represents the final global feature image; This represents the transpose convolution operation; Represents the activation function; Conv represents the convolution operation;
[0035] Both the Transformer and CNN branches consist of a multi-level cascaded structure of four decoder blocks, with parallel decoder blocks at the same level existing between each branch. The outputs of the Transformer and CNN decoder blocks at the same level are fused to form a new global feature image, which is then used as input to the next level decoder block. The expression is as follows:
[0036]
[0037] In the formula, This is the new global feature image after fusion; This is the final output of the Transformer branch; This is the output of the CNN decoder block; and All are learnable coefficients; R is a data rearrangement operation that can make... become This is to match the input of the next level decoder block;
[0038] A further improvement of this invention is that: after the serial-parallel dual-branch feature decoder, a sigmoid activation function mapping is performed; the feature map size is 120. The value is 160, which needs to be upsampled by a factor of 2 to match the input size of the original image. Considering that the sigmoid function maps from 0 to 1, the value needs to be scaled again. The scaling factor depends on the data preprocessing stage, and the expression is:
[0039]
[0040] In the formula, It is the output feature map after the decoder; It is an activation function; It's upsampling. It is the scaling factor in the data preprocessing stage, and the final output is the predicted canopy height map.
[0041] The beneficial effects of this invention are as follows: This invention proposes a tree canopy height estimation method based on UAV remote sensing images, selecting a UAV remote sensing dataset as the experimental dataset; based on the experimental dataset, a visual transformer is selected to extract multi-scale feature images from the remote sensing images of the experimental dataset; based on the multi-scale feature images, a feature pyramid is constructed to extract global feature images from the multi-scale feature images; based on the multi-scale feature images and the global feature images, a cascaded-parallel dual-branch feature decoder is constructed to fuse the features of the multi-scale feature images and the global feature images to obtain a simplified canopy height map; based on the simplified canopy height map, the size and values need to be scaled to finally obtain the processed experimental dataset; a target loss function is constructed to predict the processed experimental dataset. Attached Figure Description
[0042] Figure 1 A flowchart illustrating a tree canopy height estimation method based on UAV remote sensing images provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the effects of DatasetA, DatasetB, and Neon datasets provided in an embodiment of the present invention. Figure 3 This is a schematic diagram illustrating the prediction results of the DatasetA, DatasetB, and Neon datasets provided in this embodiment of the invention. Figure 4 This is a schematic diagram illustrating the MAE effect of the DatasetA dataset provided in an embodiment of the present invention; Figure 5 A schematic diagram illustrating the MAE effect of the DatasetB dataset provided in an embodiment of the present invention; Figure 6 This is a schematic diagram illustrating the MAE effect of the Neon dataset provided in an embodiment of the present invention. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0044] See Figure 1 This embodiment of a tree canopy height estimation method based on UAV remote sensing images includes the following operations:
[0045] The UAV remote sensing dataset was selected as the experimental dataset.
[0046] Based on the experimental dataset, a visual transformer was selected, and multi-scale feature images were extracted from the remote sensing images in the experimental dataset using the visual transformer.
[0047] Based on the multi-scale feature image, a feature pyramid is constructed to extract the global feature image from the multi-scale feature image;
[0048] Based on the multi-scale feature image and the global feature image, a serial-parallel dual-branch feature decoder is constructed to perform feature fusion on the multi-scale feature image and the global feature image to obtain a simplified canopy height map;
[0049] Based on the simplified canopy height map, the size and values need to be scaled to obtain the final processed experimental dataset.
[0050] A target loss function is constructed to predict the processed experimental dataset.
[0051] After obtaining the experimental dataset, the remote sensing images in the dataset underwent data preprocessing: the remote sensing images were cropped, i.e., cropped according to a 480° scale. The image was cropped to 640 pixels to match the normal input size of the model. Considering the inconsistency of tree canopy height in different environments, the canopy height image was scaled proportionally, with the maximum value magnified to 255. The magnification factor depends on the dataset.
[0052] To achieve tree canopy height estimation based on UAV remote sensing images, we constructed a serial-parallel dual-branch feature decoder; by combining the Transformer structure's ability to accurately capture long-distance dependencies with the CNN structure's advantage in local feature extraction, we achieved multi-directional feature fusion.
[0053] The input image is a three-channel image with dimensions of 3*480*640 (C*H*W), and the output image is a single-channel image with dimensions of 1*480*640 (C*H*W).
[0054] First, a multi-scale image is obtained through a Swing Transformer encoder. When processing the visual structure image, the Swing Transformer encoder divides the image into multiple visual feature images using a Patch-Partition layer. These visual feature images are then passed through multiple Swing Transformer coding blocks, generating feature maps of different scales at each layer. Each Swing Transformer coding block is followed by a Patch Merging layer, which reduces the resolution of the feature maps, adjusts the number of channels, and forwards the output data to the next, deeper Swing Transformer coding block. The resulting visual feature image is obtained through the Swing Transformer encoder. The expression is:
[0055]
[0056] In the formula, This is a multi-scale feature image extracted from remote sensing images by the Swing Transformer. For the SwinTransformer feature extractor; The input is a remote sensing image.
[0057] In this embodiment, the Swin Transformer encoder uses the Swin-Tiny version as the backbone network, with a window size of 12×12. The feature map resolutions of the four stages are 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the input image, respectively, with corresponding channel numbers of 96, 192, 384 and 768. The encoder is initialized using ImageNet-1K pre-trained weights.
[0058] Then, the obtained multi-scale feature images are used to construct a feature pyramid for global fitting to obtain the global feature image, i.e., the GSA module. The expression for constructing a feature pyramid from multi-scale feature images is:
[0059]
[0060] In the formula, This is a low-latitude feature image; It is a high-dimensional feature image. It is the original high-dimensional feature image before fusion; It's convolution; and These are learnable coefficients; they enable progressive global feature image fusion from low to high dimensions to obtain a preliminary global feature image. , , and The GSA module only processes the outputs of the four stages of the encoder. , , and Upward integration
[0061] It is preserved as the highest-level semantic feature.
[0062] Considering the robustness of nonlinear transformation, the initial global feature image is processed nonlinearly, and the nonlinear features are fused to obtain the final global feature image, i.e.:
[0063]
[0064] In the formula, It is a preliminary global feature image. and It is an activation function. and It is the learnable coefficient; Functions can enhance the model's ability to distinguish between tree and non-tree regions, while The function can fine-tune ReLU, especially in the shaded areas between trees;
[0065] Subsequently, based on multi-scale feature images and global feature images, a serial-parallel dual-branch feature decoder is constructed. The dual-branch feature decoder includes a Transformer branch and a CNN branch. By leveraging the Transformer branch's ability to capture the dependence of global long-distance relationships and the CNN branch's superiority in local feature extraction, the feature fusion capability of the decoder is improved. A multi-level serial decoder block structure is constructed within the same branch, and a parallel decoder branch structure is constructed between different branches.
[0066] The core structure of the Transformer branch is a gradient-enhanced conditional random field (GE-CRF) module built on a fully connected conditional random field (FC-CRF). The FC-CRF works primarily based on Markov random fields and probabilistic undirected graphical models, predicting the conditional probability distribution of the output random variable sequence from a given input random variable sequence. Essentially, it's similar to the Swing Transformer coding block; it also encoders features from the same level output. Perform feature mapping, calculate the attention score matrix for the derived Q and K, and then apply this score matrix to the global features. Weighting is applied; to characterize the spatial relationships between features, FC-CRF introduces relative position embeddings, further improving the spatial sensitivity and accuracy of feature representation, expressed as:
[0067]
[0068]
[0069] In the formula, It is a query vector. and It is the output feature of the corresponding sibling coding block. eigenmatrix for A single row vector in; It is a relative positional encoding; It is the final global feature image; It is a dot product; to improve the sensitivity to the canopy edge, convolution operations are first performed using kernels of different sizes to extract local features under different receptive fields. Then, gradient information is introduced using the interpolation method for feature processing to obtain a highly abstract composite feature information map. The specific expression is as follows:
[0070]
[0071]
[0072]
[0073]
[0074]
[0075] In the formula, It is the output of FC-CRF; This represents a convolution with a kernel size of k; A represents the average; H represents the horizontal gradient operation; V represents the vertical gradient operation. , , and Represents the learnability coefficient; Represents the learnable bias term;
[0076] Considering the local receptive field problem and the Dead ReLU phenomenon, the core structure of the CNN branch is a composite convolutional block based on transposed convolution, LeakyReLU activation function, and convolution operation, namely the SUFC module, expressed as:
[0077]
[0078] In the formula, Represents the final global feature image; This represents the transpose convolution operation; Represents the activation function; Conv represents the convolution operation;
[0079] Both the Transformer and CNN branches consist of a multi-level cascaded structure of four decoder blocks, with parallel decoder blocks at the same level existing between each branch. The outputs of the Transformer and CNN decoder blocks at the same level are fused to form a new global feature image, which is then used as input to the next level decoder block. The expression is as follows:
[0080]
[0081] In the formula, This is the output of the Transformer decoder block; This is the output of the CNN decoder block; and All are learnable coefficients; R is a data rearrangement operation that can make... become This is to match the input of the next level of decoder blocks; the input resolutions of the four decoder blocks are 1 / 32, 1 / 16, 1 / 8 and 1 / 4 of the input image, respectively, and the corresponding output resolutions are 1 / 16, 1 / 8, 1 / 4 and 1 / 2, respectively; within each decoder block, the feature map resolution is doubled through an upsampling operation;
[0082] Finally, after the cascaded-parallel dual-branch feature decoder, a sigmoid activation function mapping is performed; the feature map size is 120. The value is 160, which needs to be upsampled by a factor of 2 to match the input size of the original image. Considering that the sigmoid function maps from 0 to 1, the value needs to be scaled again. The scaling factor depends on the data preprocessing stage, and the expression is:
[0083]
[0084] In the formula, This is the output feature map after the decoder, denoted as the final output feature map after four stages of cascaded-parallel decoders, i.e., the final fused feature. ; It is an activation function; It's upsampling. This is the scaling factor during the data preprocessing stage; in this embodiment, the maximum canopy height of DatasetA is 9.7m, corresponding to... =255 / 9.7≈26.29; The maximum canopy height in DatasetB is 22.1m, corresponding to =255 / 22.1≈11.54; The maximum height after filtering in the Neon dataset is 40m, corresponding to =255 / 40=6.375; For the new dataset to be predicted, The value is the ratio of the maximum canopy height value in the dataset labels to 255; the final output is the predicted canopy height map; the invention will be further described below with reference to specific embodiments:
[0085] Reference Figure 2 The image dataset used in this invention includes high-resolution remote sensing images: DatasetA, DatasetB, and Neon. DatasetA contains 572 pairs of trees, with a maximum height of 9.7 meters, including 400 pairs in the training set, 114 pairs in the validation set, and 58 pairs in the test set. DatasetB contains 451 pairs of trees, with a maximum height of 22.1 meters, including 315 pairs in the training set, 90 pairs in the validation set, and 46 pairs in the test set. The Neon dataset is selected from the National Network of Ecological Observatories (NEON). Each image in the dataset was repeatedly and randomly cropped to a size of 480×640 pixels, and images with a maximum height of less than 40 meters were selected to create a valid dataset.
[0086] Reference Figure 3 , Figure 4 , Figure 5 and Figure 6 The visualization comparison results of some models are presented from multiple perspectives. The English abbreviations appearing in the figures are explained as follows: GT (Ground Truth) represents the true canopy height map; Predicted Height Map represents the canopy height map predicted by the model; Difference Map represents the pixel-by-pixel difference between the predicted height map and the true height map (the closer the difference is to 0, the more accurate the prediction); MAE (Mean Absolute Error) is the mean absolute error, all in meters (m), and the smaller the value, the higher the prediction accuracy.
[0087] in, Figure 3 This paper presents a visual comparison of the canopy height prediction results of the proposed method (DepthCanopyNet) with other methods such as Depth-Anything, IEBins, NewCRFs, and URCDC on three datasets. The comparison includes the input image, the height maps predicted by each method, and the ground truth (GT), which intuitively reflects the prediction performance of each method in different scenarios. Figure 4 The difference between the heightmaps predicted by each method and the true heightmaps on DatasetA is visualized (difference plot), and the MAE of each method is also labeled. Figure 5 The results of the difference visualization and MAE on DatasetB are shown. Figure 6The results of the difference visualization and MAE on the Neon dataset are presented. The difference map clearly shows that the method of the present invention has smaller errors at the tree crown edge, shaded area and height abrupt change, and the predicted height map is closer to the GT.
[0088] (1) Divide the data in the dataset into three parts according to the proportion: training set, validation set and test set. Use the training set to train the model parameters, use the validation set to select the optimal parameters during the training process, and finally use the test set to test the final performance of the model.
[0089] During model training, the AdamW optimizer was used, with an initial learning rate set to 1× The batch size was set to 8, and the training lasted for a total of 200 epochs. The learning rate adopted a cosine annealing decay strategy, with the minimum learning rate decaying to [value missing]. The loss function uses a weighted combination of smooth L1 loss and gradient loss, with weight coefficients of 1.0 and 0.5, respectively. No additional data augmentation strategies were used during training. All experiments were performed on an NVIDIA RTX 3090 GPU.
[0090] (2) Input the image data into the Swing Transformer to extract multi-scale features;
[0091] (3) Input the extracted features into the subsequent decoder network for feature fusion;
[0092] (4) Experiments were conducted on the dataset, and the results are as follows:
[0093] Table 1 Comparison of quantitative results for DatasetA
[0094]
[0095] Table 2 Comparison of quantitative results for DatasetB dataset
[0096]
[0097] Table 3 Comparison of quantitative results from the Neon dataset
[0098]
[0099] To objectively evaluate the effectiveness of the proposed method (DepthCanopyNet), quantitative comparative experiments were conducted on three datasets: DatasetA, DatasetB, and Neon. Several representative depth estimation methods, including D-Net, URCBC, PixelFormer, Zoe, VA, Depth-Anything, Depth-Anything-V2, IEBins, and NewCRFs, were selected as baselines. AbsRel (mean absolute relative error), SqRel (squared relative error), RMS (root mean square error), LogRMS (log mean square error), and threshold accuracy were used as benchmarks. , , As an evaluation indicator, ↓ indicates that the smaller the indicator, the better, and ↑ indicates that the larger the indicator, the better.
[0100] Table 1 presents the quantitative comparison results on Dataset A (tree species: Catalpa, tallest tree approximately 9.7 meters). The method of this invention achieves significantly better results than all comparative methods in terms of the four error indices: AbsRel, SqRel, RMS, and LogRMS, reaching 0.1697, 0.2105, 0.8764, and 0.2824 respectively. Furthermore, the method demonstrates superior threshold accuracy. , , The values reached 0.8991, 0.9608, and 0.9742 respectively, except... Except for a score slightly lower than 0.9081 for Depth-Anything-V2, all other indicators are at the optimal or suboptimal level; experiments show that this method has high-precision estimation capability in low-to-medium canopy scenarios.
[0101] Table 2 shows the quantitative comparison results on DatasetB (tree species: wet pine, tallest tree approximately 22.1 meters). The method of this invention achieved the best results in AbsRel (0.3075), SqRel (0.7509), RMS (1.7606), LogRMS (0.5922), and three threshold accuracy indices (0.8798, 0.9345, 0.9551). Compared to the second-best method, NewCRFs, AbsRel is reduced by approximately 1.6%, and LogRMS by approximately 4.1%. Experiments show that this method maintains a low prediction error even in high-canopy scenes and has good scale generalization ability.
[0102] Table 3 presents the quantitative comparison results on the Neon dataset. This dataset originates from the National Ecological Observation Station Network, featuring complex scenes, diverse canopy structures, and various vegetation types and terrain conditions. The method of this invention achieved the best results in AbsRel (0.3309), LogRMS (0.3887), and three threshold accuracy metrics (0.7025, 0.8787, 0.9469). RMS (2.5631) and SqRel (1.2273) also ranked first and second, only behind PixelFormer. The overall results demonstrate that the method of this invention exhibits excellent robustness and generalization ability when estimating canopy height in complex natural environments.
[0103] Based on the quantitative results in Tables 1 to 3, the serial-parallel dual-branch feature decoder constructed in this invention can effectively integrate the global long-range dependencies captured by the Transformer branch with the local detail texture features extracted by the CNN branch. It achieves the best or leading quantitative scores on three datasets with different tree species, different height scales, and different environmental complexities, significantly improving the accuracy and stability of canopy height estimation, and verifying the effectiveness and advancement of the technical solution proposed in this invention.
[0104] Through the above embodiments, this invention proposes a canopy height estimation method based on UAV remote sensing images, comprising the following operations: selecting a UAV remote sensing dataset as an experimental dataset; selecting a visual transformer based on the experimental dataset, and extracting multi-scale feature images from the UAV remote sensing images in the experimental dataset using the visual transformer; constructing a feature pyramid based on the multi-scale feature images to extract global feature images from the multi-scale feature images; constructing a cascaded-parallel dual-branch feature decoder based on the multi-scale feature images and the global feature images, and fusing the features of the multi-scale feature images and the global feature images to obtain the processed experimental dataset. This invention's method extracts feature images using a visual model, constructs a module to extract and fuse features, and predicts the processed dataset, thereby improving the accuracy and robustness of canopy height estimation from UAV remote sensing images.
[0105] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.
[0106] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A canopy height estimation method based on UAV remote sensing images, characterized in that, Includes the following steps: Step S1: Select the UAV remote sensing dataset as the experimental dataset; Step S2: Based on the experimental dataset, select a visual transformer and extract multi-scale feature images from the remote sensing images in the experimental dataset using the visual transformer; Step S3: Based on the multi-scale feature image, construct a feature pyramid to extract the global feature image from the multi-scale feature image; Step S4: Based on the multi-scale feature image and the global feature image, a serial-parallel dual-branch feature decoder is constructed to perform feature fusion on the multi-scale feature image and the global feature image to obtain a simplified canopy height map; Step S5: Based on the simplified canopy height map, scale the dimensions and values to obtain the processed experimental dataset; Step S6: Construct a target loss function to predict the processed experimental dataset.
2. The canopy height estimation method based on UAV remote sensing images according to claim 1, characterized in that, In step S1, after obtaining the experimental dataset, the remote sensing images in the experimental dataset are preprocessed.
3. The canopy height estimation method based on UAV remote sensing images according to claim 2, characterized in that, The data preprocessing includes: processing the remote sensing image according to 480... A 640-pixel crop is applied to the canopy height image, which is scaled proportionally to a maximum value of 255. The magnification factor depends on the dataset.
4. The canopy height estimation method based on UAV remote sensing images according to claim 1, characterized in that, In step S2, the visual transformer is a Swing Transformer encoder; When processing visual structure images, the Swin Transformer encoder divides the visual structure images into multiple visual feature images through a Patch-Partition layer. These multiple visual feature images are then sequentially passed through multiple Swin Transformer encoding blocks, generating feature maps of different scales at each layer. Each Swin Transformer encoding block is followed by a Patch Merging layer, which reduces the resolution of the feature maps, adjusts the number of channels, and forwards the output data to the next deeper Swin Transformer encoding block. The visual feature images of the Swin Transformer encoder... The expression is: ; In the formula, This is a multi-scale feature image extracted from remote sensing images by the Swing Transformer. For the SwinTransformer feature extractor; The input is a remote sensing image.
5. The canopy height estimation method based on UAV remote sensing images according to claim 1, characterized in that, In step S3, a feature pyramid is constructed from the obtained multi-scale feature images to perform global fitting and obtain a preliminary global feature image; The expression for constructing a feature pyramid from multi-scale feature images is: ; In the formula, This is a low-latitude feature image; It is a high-dimensional feature image. It is the original high-dimensional feature image before fusion; It's convolution; and These are learnable coefficients; by performing nonlinear processing on the initial global feature image and fusing the nonlinear features, the final global feature image is obtained, i.e.: ; In the formula, It is the final global feature image. It is a preliminary global feature image. and It is an activation function. and It is the learnable coefficient.
6. The canopy height estimation method based on UAV remote sensing images according to claim 1, characterized in that, In step S4, the serial-parallel dual-branch feature decoder includes a Transformer branch and a CNN branch. A multi-level serial decoder block structure is built within the same branch, and a parallel decoder branch structure is built between different branches. The core structure of the Transformer branch is a gradient-enhanced conditional random field built on a fully connected conditional random field (FC-CRF); the FC-CRF is introduced into the relative position embedding, expressed as: ; ; In the formula, It is a query vector. and It is the output feature of the corresponding sibling coding block. The feature matrix; Right now The transpose of is used to calculate and The dot product is used to obtain the attention score matrix; It is a relative positional encoding; It is the final global feature image; By performing convolution operations using kernels of different sizes, local features under different receptive fields are extracted. Gradient information is then introduced using the interpolation method for feature processing, resulting in a highly abstract composite feature map. The specific expression is as follows: ; ; ; ; ; In the formula, It is the output of FC-CRF; , , and It is an intermediate variable within the Transformer branch; This is the final output of the Transformer branch; This represents a convolution with a kernel size of k; A represents the average; H represents the horizontal gradient operation; V represents the vertical gradient operation. , , and Represents the learnability coefficient; Represents the learnable bias term; Considering the local receptive field problem and the Dead ReLU phenomenon, the core structure of the CNN branch is a composite convolutional block based on transposed convolution, Leaky ReLU activation function, and convolution operation, expressed as: ; In the formula, These are the output features of the CNN branch; Represents the final global feature image; This represents the transpose convolution operation; Represents the activation function; Conv represents the convolution operation; Both the Transformer and CNN branches consist of a multi-level cascaded structure of four decoder blocks, with parallel decoder blocks at the same level existing between each branch. The outputs of the Transformer decoder blocks and the CNN decoder blocks at the same level are fused to form a new global feature image, as shown in the following expression: ; In the formula, This is the new global feature image after fusion; This is the final output of the Transformer branch; This is the output of the CNN decoder block; and All are learnable coefficients; R is a data recombination operation.
7. The canopy height estimation method based on UAV remote sensing images according to claim 6, characterized in that, In step S4, after the cascaded-parallel dual-branch feature decoder, a sigmoid activation function mapping is performed; the feature map size is 120. 160 is upsampled to a factor of 2, and then the value is scaled down again. The scaling factor depends on the data preprocessing stage, and the expression is: ; In the formula, It is the output feature map after the decoder; It is an activation function; It's upsampling. It is the scaling factor in the data preprocessing stage, and the final output result. This is the predicted canopy height map.