A method for detecting standing dead trees by using a dual-flow feature integration network based on a UAV remote sensing

By constructing a dead tree detection method based on a dual-stream feature integration network and utilizing a multi-branch module and a separation and enhancement attention module, the detection problem in dense target scenes is solved, and high-precision recognition and generalization capabilities of dead trees are achieved.

CN119540801BActive Publication Date: 2025-10-10NORTHEAST FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411685620.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-23
Publication Date
2025-10-10
Estimated Expiration
2044-11-23

AI Technical Summary

Technical Problem

Existing technologies for dead tree detection in dense target scenes have problems such as severe crown occlusion, sparse crowns that are easily affected by background, incomplete image edge areas, less feature information of small targets, and few samples, resulting in insufficient detection accuracy.

Method used

A method based on a dual-stream feature integration network is adopted. The traditional convolution module is replaced by the multi-branch module (DBB). The separation and enhancement attention module (SEAM) and the CBLinear and CBFuse modules are combined to construct a dual-branch structure backbone network to enhance feature extraction and fusion and solve the problems of dense occlusion and background interference.

Benefits of technology

It significantly improves the recognition accuracy of small targets and image edge areas, reduces the missed detection rate, improves the model's ability to recognize the features of dead trees, and enhances the model's generalization performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119540801B_ABST
    Figure CN119540801B_ABST
Patent Text Reader

Abstract

The application discloses a UAV remote sensing dead standing tree detection method based on a double-flow feature integrated network, and belongs to the technical field of UAV image processing.The application significantly improves the recognition accuracy of small targets and image edge region targets by fusing features of different levels.In addition, the application also introduces a DBB module.The DBB module contains multiple branches of different scales and complexities, enhances the feature expression capability of the convolution block, and effectively reduces the missing detection rate of the dense crown sheltering area.Meanwhile, the effective combination of the double-branch feature fusion main body and the DBB module fully excavates the features of the dead standing tree, enhances the propagation of the features in the network, and overcomes the small sample problem.In addition, considering that the crown is sparse after the dead branches and leaves, the introduction of a large amount of background information reduces the detection accuracy, and the SEAM is introduced.The SEAM module enhances the attention to the detail features of the dead standing tree, such as color, edges and corners and texture, effectively improves the recognition ability of the model to the dead branch features in the crown layer, and reduces the interference of the background information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of unmanned aerial vehicle (UAV) image processing, and in particular relates to a UAV remote sensing dead tree detection method based on a dual-stream feature integration network. Background Art

[0002] Forests are complex ecosystems, and the health and life cycle of trees play a vital role in their composition and function. Tree mortality leads to a reduction in the total leaf area of ​​the community, which affects soil processes such as solar radiation and nutrient cycling, fungal activity, soil erosion and evaporation rates. At the same time, large-scale tree mortality events can have a negative impact on forest carbon storage, biodiversity, and the livelihoods of people who rely on forest resources. In the context of global climate change, extreme events such as droughts and wildfires can cause large-scale tree mortality, further amplifying these impacts. However, trees do not usually fall down immediately after death, but become standing dead trees (SDTs). Therefore, accurate detection and monitoring of standing dead trees is crucial to understanding these impacts, optimizing forest management, and reducing forest fires and their associated carbon emissions.

[0003] There are many reasons for tree death, such as stress, pests and diseases, and interspecific competition. Pests and diseases have the characteristic of rapid outbreaks, so many scholars also identify trees infected by pests and diseases in order to control forest health in advance. For example, Mo Dengkui et al. disclosed a large-scale intelligent detection method for dead standing trees caused by pine wilt disease based on drones and convolutional neural networks, which can realize the rapid detection, counting and evaluation of dead standing trees in a large area; Chen Heng et al. proposed an integrated sky-ground remote sensing monitoring method for pine wilt disease epidemics, which realized comprehensive monitoring of pine wilt disease epidemic areas and improved the accuracy and efficiency of monitoring dead trees caused by pine wilt disease; Chen Xiaohua et al. disclosed a sky-ground integrated monitoring method for pine wilt disease, which has high efficiency in screening epidemic areas and can realize large-scale and efficient monitoring; Zhang Fei et al. disclosed a method, device, medium and equipment for patrolling and identifying dead trees and diseased trees, which can accurately identify dead trees and diseased trees; Jiang Miaomiao et al. proposed an improved Yolov5 dead wood detection method combined with a defogging algorithm, which organically combines the target detection model with the defogging algorithm network, so that it can still maintain a relatively high target detection accuracy when facing foggy images. Accuracy; Xu Guoqing et al. provided a dead tree identification method and equipment based on visible light images. By reducing the bit and image conversion of visible light images, and then clustering and denoising the images, the dead trees in the jungle can be identified without relying on infrared images; Chen Farong et al. provided a pine wood nematode disease standing tree detection method and device based on UAV remote sensing images. The detection accuracy of pine wood nematode disease standing trees was improved by fusing multiple model results; Lan Yubin et al. disclosed a pine wood nematode disease dead tree detection and positioning method based on deep learning. The method can quickly, efficiently and accurately detect diseased pine trees and determine the location of diseased pine trees; Xu Guoqing disclosed a visible light remote sensing image dead tree identification software system and identification method. Through the remote sensing image segmentation module, the advanced deep learning method is combined with tree remote sensing image object recognition to realize the identification of dead trees in remote sensing images.

[0004] Currently, dual-branch structures have been widely used in processing images, videos and other data. For example, Xu Xiaolong et al. disclosed a classroom student posture recognition method based on dual-trunk feature fusion, which effectively mines and fuses multi-scale visual features, and realizes accurate and efficient recognition of multi-target student postures in complex scenes; Zhou Hangxia et al. disclosed a solar radiation prediction method and system based on dual-branch feature extraction, combining multi-scale convolution and bidirectional gated recurrent networks to extract meteorological features and time series features respectively, and combining the attention mechanism to optimize the weighted fusion of each branch, with significantly improved prediction accuracy; Hua Chunjian et al. disclosed an RGB-D saliency detection method based on progressive weighted decoding, which extracts multi-level RGB image features and depth image features respectively through a symmetrical dual-stream feature extraction backbone network, and adopts a cross-modal feature fusion module to enhance cross-modal information interaction between different branches. Experimental results show that the invention has high target detection accuracy for various scenarios.

[0005] The above inventions detect trees infected by pests and diseases and dead trees respectively, and explore some applications of dual-branch networks. However, there is a lack of inventions related to the detection of dead trees in dense scenes. Therefore, the present invention proposes a dual-stream feature integration network to deal with problems such as mutual occlusion between tree crowns in dense scenes, background information introduced by sparse tree crowns, incomplete dead trees in the edge area of ​​the image, small targets and few samples. Summary of the Invention

[0006] The purpose of the present invention is to solve the problems of close connection between the crowns of dead trees in dense target scenes, serious occlusion problem, sparse crowns of dead trees after fallen branches are easily affected by the background, incomplete crowns of dead trees in the edge area of ​​the image, and less feature information and few samples of small targets. A UAV remote sensing dead tree detection method based on a dual-stream feature integration network is provided.

[0007] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0008] A method for detecting dead standing trees using UAV remote sensing based on a dual-stream feature integration network is provided. The method comprises:

[0009] Step 1: Use drones to obtain images of dead trees, and use annotation software combined with field surveys to annotate the dead trees.

[0010] Step 2: Construct a deadwood detection dataset, which includes both the collected dataset and the public dataset. The collected dataset is divided into training, validation, and test sets. During the training process, online data augmentation (including rotation, scaling, flipping, and cropping) is used to augment the training set to enhance the generalization performance of the model. The public dataset is used as an additional test set to verify the generalization performance of the model.

[0011] Step 3: Use the Diverse Branch Block (DBB) to replace the traditional convolution module to fully explore the features of dead trees; use the Separated and Enhanced Attention Module (SEAM) to process the shallow features extracted from the trunk to enhance the attention to the details of the dead trees and obtain the enhanced features F SEAM Combining CBLinear and CBFuse to build a dual-branch backbone network, the left backbone of the network is responsible for extracting basic features, while the right backbone not only extracts features, but also integrates and fuses the features of the left backbone to finally obtain feature F;

[0012] Step 4: The feature F obtained in step 3 is fed into the model detection head for training to reduce the difference between the predicted results and the true labels, and the detection performance of the model is evaluated on the validation set to ensure that the model not only learns the characteristics of the data but can also generalize to unseen data to obtain the optimal model.

[0013] Step 5: Apply the trained model to the new image to perform real-time dead tree detection, and finally obtain an output image containing candidate boxes and confidence scores.

[0014] The beneficial effects of the present invention compared to the prior art are as follows: the present invention designs a novel dual-branch feature extraction and fusion backbone network for detecting dead trees using RGB images. By fusing features at different levels, the recognition accuracy of small targets and targets in the edge areas of the image is significantly improved. In addition, in order to solve the problem of densely connected tree crown occlusion, the present invention introduces a Diverse Branch Block (DBB) module to replace the traditional convolution module. The DBB module contains multiple branches of different scales and complexities, which enhances the feature expression capability of the convolution block and effectively reduces the missed detection rate in areas occluded by dense tree crowns. At the same time, the effective combination of the dual-branch feature fusion backbone and the DBB module fully explores the characteristics of dead trees, effectively enhances the propagation of features in the network, and overcomes the problem of small samples. In addition, considering that the tree crown is sparse after the dead branches and leaves fall, the introduction of a large amount of background information reduces the detection accuracy, and the Separated and Enhanced Attention Module (SEAM) is introduced. The SEAM module effectively improves the model's ability to recognize dead branch features in the canopy by enhancing attention to the detailed features of dead trees, such as color, edges and texture, and reduces the interference of background information. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] 图1 This is the dual-backbone feature extraction and fusion network structure diagram. DETAILED DESCRIPTION

[0016] In order to better understand the content of the present invention, the present invention will be further described below in conjunction with specific examples and drawings. The following examples are implemented based on the technology of the present invention and provide detailed implementation methods and operating steps, but the scope of protection of the present invention is not limited to the following examples.

[0017] Example 1:

[0018] A method for detecting dead standing trees using UAV remote sensing based on a dual-stream feature integration network is provided. The method comprises:

[0019] Step 1: Use drones to obtain images of dead trees, and use annotation software combined with field surveys to annotate the dead trees.

[0020] Step 2: Construct a deadwood detection dataset, which includes both the collected dataset and the public dataset. The collected dataset is divided into training, validation, and test sets. During the training process, online data augmentation (including rotation, scaling, flipping, and cropping) is used to augment the training set to enhance the generalization performance of the model. The public dataset is used as an additional test set to verify the generalization performance of the model.

[0021] Step 3: Use the Diverse Branch Block (DBB) to replace the traditional convolution module to fully explore the features of dead trees; use the Separated and Enhanced Attention Module (SEAM) to process the shallow features extracted from the trunk, enhance the attention to the details of the dead trees, and obtain the enhanced features F SEAM Combining CBLinear and CBFuse to build a dual-branch backbone network, the left backbone of the network is responsible for extracting basic features, while the right backbone not only extracts features, but also integrates and fuses the features of the left backbone to finally obtain feature F;

[0022] Step 4: Feed the feature F obtained in step 3 into the model detection head. Train the model to reduce the difference between the predicted results and the true labels, and evaluate the detection performance of the model on the validation set to ensure that the model not only learns the characteristics of the data but also generalizes to unseen data, thus obtaining the optimal model.

[0023] Step 5: Apply the trained model to the new image to perform real-time dead tree detection, and finally obtain an output image containing candidate boxes and confidence scores.

[0024] In step 3, the multi-branch module adopts a multi-branch topology, combining multi-scale convolution, sequential 1x1 and KxK convolution, average pooling, and branch addition, which greatly enriches the feature space but also increases the complexity of training. It is worth noting that the DBB module can be converted into a regular convolution layer during inference, so there is no additional inference cost. This allows DBB to be used as a replacement for existing convolutional layers, improving performance without affecting inference efficiency. In step 3, the training process is shown in Equation (1):

[0025]

[0026] In step 5, the output during inference is shown in formula (2):

[0027] Y = Conv merged (X)(2)

[0028] In formula (1), X and Y represent the input and output of the DBB module, w1, w2, w3, w4 are the weights of each branch, and Conv 1×1 and Conv K×K They represent convolution layers with kernel sizes of 1×1 and K×K, respectively, BN represents batch normalization layer, and AvgPool represents average pooling layer; in formula (2), Conv merged Represents a new convolution operation, which recalculates the weights and biases of each branch of the original DBB module to ensure that the new convolution operation is as consistent as possible with the original multi-branch output.

[0029] The SEAM module uses multiple residual convolution blocks to enhance the expressiveness of features, and recalibrates feature channels through a channel attention mechanism consisting of an adaptive average pooling layer and a fully connected layer, thereby highlighting important features and suppressing unimportant features. This method not only reduces the number of parameters but also improves computational efficiency. The output of the residual convolution block is added to the input features through a residual connection. This structure helps the network capture the residual between the input features and the transformed features, thereby effectively compensating for the information loss caused by occlusion and sparse branches, enhancing the attention and expression of dead tree features, and improving the model's ability to recognize dead trees. The main functional blocks and related formulas are shown below:

[0030] (1) Residual convolution block

[0031] The residual convolution block consists of a standard convolution operation followed by a GELU activation function and batch normalization. The output is then added to the input to form the final output. The formula is as follows:

[0032] output=Conv(input)+input(3)

[0033] Among them, Conv is a standard convolution operation, input is the input feature map, and output is the output feature map;

[0034] (2) Adaptive average pooling layer

[0035] Reduce the spatial dimension of the input feature map to 1×1, that is, perform global averaging on the features within each channel;

[0036] y=AvgPool(x)(4)

[0037] Where x represents the input feature map, i.e., the output of the residual convolution block in formula (3), and AvgPool represents the adaptive average pooling operation, which reduces the spatial dimension of the feature map to 1×1;

[0038] (3) Fully connected layer

[0039] The feature recalibration is achieved through two fully connected layers. The first fully connected layer compresses the number of channels, the second fully connected layer restores the number of channels, and finally the weight factor is output through the Sigmoid function.

[0040] s=σ(FC2(ReLU(FC1(y))))(5)

[0041] Where y is the output of the adaptive average pooling of formula (4); FC1 represents the first fully connected layer for dimensionality reduction, from C in Down to C in / reduction; ReLU is an activation function used to introduce nonlinearity to help the network learn complex functions; FC2 represents the second fully connected layer, which reduces the dimension from C in / reduction restores to C in ; σ is the Sigmoid activation function, which outputs a weight factor s between 0 and 1 for subsequent feature recalibration;

[0042] (4) Feature recalibration

[0043] Use the weight factor output by the fully connected layer to scale each channel of the input feature map;

[0044] output=x·exp(s)(6)

[0045] Where x represents the input original feature map, s represents the weight factor calculated by formula (5), and exp(s) represents the application of an exponential function to s, which is used to enhance the model's sensitivity to important features.

[0046] The main function of the CBLinear module is to perform convolution operations on the input feature map and split the output results into multiple feature maps according to the predefined number of channels; this design helps to disperse information to different branches, facilitating subsequent feature fusion.

[0047] For the input feature map

[0048] (1) Use a convolutional layer with a kernel size of k, a stride of s, and a padding of p:

[0049] Y=Conv2d(X,k,s,p)(7)

[0050] Among them, X represents the input feature map, C in Indicates the number of channels of the input feature map (channel), H indicates the height of the input feature map (height), and W indicates the width of the input feature map (width). Indicates that X is a C in Channel, a three-dimensional tensor of height H and width W, whose elements belong to the set of real numbers Conv2d represents a two-dimensional convolution operation, Y is the output feature map after the convolution operation, Σc out The sum of the number of channels of all output feature maps is obtained by the sum of the elements in c2s (which contains the number of channels that each output branch should have). H' and W' represent the height and width of the output feature map. It means that Y is a out Channel, a three-dimensional tensor of height H′ and width W′, whose elements belong to the set of real numbers

[0051] (2) Split output channels: Split the convolution output Y into m tensors Y i ,in Each segmented feature map represents the characteristics of different branches of the network, which are used for subsequent fusion (that is, the feature map generated after the convolution operation will be divided into several parts, each part represents the feature information of a branch in the network. This design allows different network branches to focus on processing different types of information, and then these different information will be recombined to enhance the network's understanding and processing capabilities of the overall data); the mathematical expression is:

[0052] Y=[Y1,Y2,...,Y m ](8)

[0053] Among them, Y i The number of channels is and Indicates Y i It is a Channel, a three-dimensional tensor of height H′ and width W′, whose elements belong to the set of real numbers

[0054] The function of the CBFuse module is to fuse multi-scale, multi-channel feature maps into a unified feature map to adapt to the different feature requirements of the network; it first interpolates and adjusts the feature maps of different sizes, and then adds them in the feature map dimension to achieve fusion.

[0055] There are n input feature maps {Y1,Y2,…,Y n}, sizes vary; Y n If the target size is H′×W′, all other feature maps are interpolated and adjusted to this size;

[0056] (1) Interpolation adjustment: For each Y i (where i <n),使用最邻近插值将其调整到目标尺寸H′×W′:

[0057] Y i =Interpolate(Y i ,size=(H′,W′))(9)

[0058] Among them, Y i ' is the feature map after interpolation adjustment, Interpolate represents the interpolation operation;

[0059] (2) All adjusted feature maps (including the last feature map) are added element-wise to obtain the fused output feature map F:

[0060]

[0061] Among them, F is the feature map obtained by fusing all feature maps;

[0062] In step 3, the feature fusion process is as follows:

[0063] (1) Feature extraction and preliminary fusion: The left backbone is used to extract basic features F L And use CBLinear to transform the features of each layer (where i = 1, 3, 5, 7, 9) segmentation for subsequent fusion; the right backbone first processes the image through a convolutional layer and a SEAM module to enhance the representation of key features and obtain feature F SEAM ; Then, through the CBFuse module, F SEAM It is preliminarily fused with the features of the five key levels of the left backbone (i.e., layers 1, 3, 5, 7, and 9) to produce feature FF1; this is the first fusion operation in the network.

[0064] (2) Further feature fusion: FF1 undergoes further feature extraction through a DBB module, and then undergoes a second fusion with the features of the 3rd, 5th, 7th, and 9th layers of the left backbone through a CBFuse module to obtain feature FF2; similarly, FF2 undergoes another c2f and DBB module, and undergoes a third fusion with the features of the 5th, 7th, and 9th layers of the left backbone to generate feature FF3; this process continues, and FF3 undergoes a fourth fusion with the features of the 7th and 9th layers of the left backbone through a c2f and DBB module to obtain feature FF4; finally, FF4 undergoes a fifth fusion with the features of the 9th layer of the left backbone through a c2f and DBB module to generate feature FF5; after passing the last c2f module, FF5 enters the Neck part of the network and undergoes feature fusion again to maximize the utilization efficiency of the features, obtaining feature F.

[0065] Combining CBLinear and CBFuse creates a dual-branch feature extraction and fusion backbone. This dual-branch structure enhances the network's expressiveness and adaptability to different scales, helping to capture and preserve important information at different scales. Furthermore, it mitigates the loss of image features during downsampling and enhances feature propagation throughout the network, resulting in more accurate and comprehensive detection results.

[0066] Ablation experiments

[0067] In order to verify the effectiveness of each module, an ablation comparison experiment was designed on the test set. Since the proposed DSFI-YOLO is based on the YOLOv8n model, YOLOv8n is used as the benchmark for the ablation experiment. The results are shown in Table 1. After fusing the features of each layer using the dual-branch feature extraction fusion backbone structure, the mAP 50 The improvement was 1.0%, which enhanced the recognition effect of small targets and dead trees with incomplete features in the edge area of ​​the image; after replacing the convolution module with DBB, the problem of missed detection caused by occlusion between dead trees was solved, and mAP 50 Improved by 1.3%; In addition, the introduction of SEAM attention mechanism improves the recognition effect of the model on sparse tree crowns, mAP 50 It increased by 0.8%. In summary, the mAP of the proposed algorithm 50 Reaching 92.5%, compared with the YOLOv8n algorithm, mAP 50 Improved mAP by 3.1% 50-95 It increased by 0.9%, and P and R increased by 3.7% and 4.9% respectively.

[0068] Table 1 Comparison of ablation experiment results

[0069] Yolov8n DSFI backbone DBB SEAM P / % R / % <![CDATA[mAP 50 / %]]> <![CDATA[mAP 50-95 / %]]> √ - - - 85.5 80.9 89.4 45.2 - √ - - 87.7 83.0 90.4 45.2 - √ √ - 90.0 85.2 91.7 46.7 - √ √ √ 89.2 85.8 92.5 46.1

[0070] “√” means the module is added, and “-” means no operation is performed.

[0071] Comparative test

[0072] The comparison results of DSFI-YOLO with SSD, FasterR-CNN, LDS-YOLO and ImprovedYOLOv7 on the test set and public datasets are shown in Table 2 and Table 3 respectively.

[0073] Table 2 shows that DSFI-YOLO achieved the best recognition performance in the dead tree test set. Its mAP50 score was 92.5%, improving over SSD, Faster R-CNN, LDS-YOLO, and ImprovedYOLOv7 by 9.0%, 12.0%, 5.5%, and 6.4%, respectively. Its P-value improved by 9.2%, 8.7%, 2.5%, and 3.6% compared to SSD, Faster R-CNN, LDS-YOLO, and ImprovedYOLOv7, respectively. Its R-value improved by 2.0%, 6.9%, and 7.0% compared to Faster R-CNN, LDS-YOLO, and ImprovedYOLOv7, but was 0.9% lower than SSD. Its mAP50-95 score improved by 8.3%, 7.1%, 3.4%, and 5.7% compared to SSD, Faster R-CNN, LDS-YOLO, and ImprovedYOLOv7, respectively. In terms of FPS, it's only slightly lower than LDS-YOLO, but significantly higher than the other comparison models. In terms of Params, it's only slightly higher than LDS-YOLO, but significantly lower than the other comparison models. A comprehensive comparison of average P, R, mAP50, mAP50-95, FPS, and Params reveals that the DSFI-YOLO algorithm exhibits superior performance.

[0074] Table 2 Performance comparison of different target detection models on the test set

[0075] 模型 P / % R / % <![CDATA[mAP 50 / %]]> <![CDATA[mAP 50-95 / %]]> <![CDATA[FPS / F·s -1 ]]> Params / MB SSD 80.0 86.7 83.5 37.8 39.3 181.0 FasterR-CNN 80.5 83.8 80.5 39.0 42.4 460.0 LDS-YOLO 86.7 78.9 87.0 42.7 131.6 8.0 ImprovedYOLOv7 85.6 79.0 86.1 40.4 46.1 71.3 DSFI-YOLO 89.2 85.8 92.5 46.1 83.3 15.9

[0076] Table 3 shows that LDS-YOLO achieves the best recognition results on the public dataset, followed by the proposed algorithm. Based on mAP50, mAP50-95, FPS, and Params analysis, the proposed algorithm outperforms SSD, FasterR-CNN, and Improved YOLOv7. Furthermore, combining Tables 2 and 3 reveals that LDS-YOLO prioritizes P-values, while the proposed algorithm is more stable in terms of R-values, achieving an mAP50 of 81.1%, demonstrating excellent performance. This demonstrates the good generalization of the proposed algorithm for deadwood detection.

[0077] Table 3 Performance comparison of different target detection models on public datasets

[0078] 模型 P / % R / % <![CDATA[mAP 50 / %]]> <![CDATA[mAP 50-95 / %]]> <![CDATA[FPS / F·s -1 ]]> Params / MB SSD 77.2 83.4 77.2 37.7 39.2 181.0 FasterR-CNN 68.7 88.5 68.3 39.9 37.3 460.0 LDS-YOLO 86.4 76.6 82.1 42.8 140.8 8.0 ImprovedYOLOv7 68.9 77.7 74.9 41.7 78.1 71.3 DSFI-YOLO 78.8 82.8 81.1 64.0 87.0 15.9

Claims

1. A method for detecting dead standing trees using UAV remote sensing based on a dual-stream feature integration network, characterized by: The method is: Step 1: Use drones to obtain images of dead trees, and use annotation software combined with field surveys to annotate the dead trees. Step 2: Construct a deadwood detection dataset, which includes the collected dataset and the public dataset. The collected dataset is divided into training set, validation set, and test set. During the training process, the training set is augmented with online data augmentation to enhance the generalization performance of the model. The public dataset is used as an additional test set to verify the generalization performance of the model. Step 3: Use the multi-branch module DBB to replace the traditional convolution module to fully explore the characteristics of dead trees; The separation and enhancement attention module SEAM is used to process the shallow features extracted from the trunk to enhance the attention to the details of the dead trees, and the enhanced features F are obtained. SEAM Combining CBLinear and CBFuse to build a dual-branch backbone network, the left backbone of the network is responsible for extracting basic features, while the right backbone not only extracts features, but also integrates and fuses the features of the left backbone to finally obtain feature F. The feature fusion process is as follows: (1) Feature extraction and preliminary fusion: The left backbone is used to extract basic features F L And use CBLinear to transform the features of each layer Segmentation, used for subsequent fusion, where i=1, 3, 5, 7, 9; The right backbone first processes the image through a convolutional layer and the SEAM attention mechanism to enhance the representation of key features and obtain feature F SEAM ; Then, through the CBFuse module, F SEAM It is preliminarily fused with the features of the five key levels of the left backbone, namely the 1st, 3rd, 5th, 7th, and 9th layers, to generate feature FF1; (2) Further feature fusion: FF1 undergoes further feature extraction through a DBB module, and then undergoes a second fusion with the features of the 3rd, 5th, 7th, and 9th layers of the left backbone through the CBFuse module to obtain feature FF2; FF2 undergoes another c2f and DBB module, and undergoes a third fusion with the features of the 5th, 7th, and 9th layers of the left backbone to generate feature FF3; this process continues, and FF3 undergoes a fourth fusion with the features of the 7th and 9th layers of the left backbone after passing through the c2f and DBB module to obtain feature FF4; finally, FF4 undergoes a fifth fusion with the features of the 9th layer of the left backbone after passing through the c2f and DBB module to generate feature FF5; FF5 passes through the last c2f module and enters the Neck part of the network, where feature fusion is performed again to maximize the efficiency of feature utilization and obtain feature F; Step 4: Feed the feature F obtained in step 3 into the model detection head. Train the model to reduce the difference between the predicted results and the true labels, and evaluate the detection performance of the model on the validation set to ensure that the model not only learns the characteristics of the data but also generalizes to unseen data, thus obtaining the optimal model. Step 5: Apply the trained model to the new image to perform real-time dead tree detection, and finally obtain an output image containing candidate boxes and confidence scores.

2. The method for detecting dead standing trees using UAV remote sensing based on a dual-stream feature integration network according to claim 1, characterized in that: In step 3, the multi-branch module adopts a multi-branch topology, combining multi-scale convolution, sequential 1x1 and KxK convolution, average pooling and branch addition; The training process is shown in formula (1): (1) In step 5, the output during inference is shown in formula (2): (2) In formula (1), X and Y represent the input and output of the DBB module, , , , is the weight of each branch, and They represent convolution layers with kernel sizes of 1×1 and K×K, respectively, BN represents batch normalization layer, and AvgPool represents average pooling layer; in formula (2), Represents a new convolution operation, which recalculates the weights and biases of each branch of the original DBB module to ensure that the new convolution operation is as consistent as possible with the original multi-branch output.

3. The method for detecting dead trees using remote sensing by an unmanned aerial vehicle (UAV) based on a dual-stream feature integration network according to claim 1, characterized in that: In step 3, the SEAM module uses multiple residual convolution blocks to enhance the expressiveness of features and recalibrates feature channels through a channel attention mechanism consisting of an adaptive average pooling layer and a fully connected layer, thereby highlighting important features and suppressing unimportant features. The output of the residual convolution block is added to the input features through a residual connection. The main functional blocks and related formulas are shown below: (1) Residual convolution block The residual convolution block consists of a standard convolution operation followed by a GELU activation function and batch normalization. The output is then added to the input to form the final output. The formula is as follows: (3) Among them, Conv is a standard convolution operation, input is the input feature map, and output is the output feature map; (2) Adaptive average pooling layer Reduce the spatial dimension of the input feature map to 1×1, that is, perform global averaging on the features within each channel; (4) Where x represents the input feature map, i.e., the output of the residual convolution block in formula (3), and AvgPool represents the adaptive average pooling operation, which reduces the spatial dimension of the feature map to 1×1; (3) Fully connected layer The feature recalibration is achieved through two fully connected layers. The first fully connected layer compresses the number of channels, the second fully connected layer restores the number of channels, and finally the weight factor is output through the Sigmoid function. (5) Where y is the output after adaptive average pooling of formula (4); FC1 represents the first fully connected layer for dimensionality reduction, from C in Down to C in / reduction; ReLu is an activation function used to introduce nonlinearity to help the network learn complex functions; FC2 represents the second fully connected layer, which reduces the dimension from C in / reduction restores to C in ; It is a Sigmoid activation function that outputs a weight factor s between 0 and 1 for subsequent feature recalibration; (4) Feature recalibration Use the weight factor output by the fully connected layer to scale each channel of the input feature map; (6) Where x represents the input original feature map, s represents the weight factor calculated by formula (5), and exp(s) represents the application of an exponential function to s, which is used to enhance the model's sensitivity to important features.

4. The method for detecting dead standing trees using UAV remote sensing based on a dual-stream feature integration network according to claim 1, characterized in that: In step 3, the feature map of the original branch is split into multi-channel feature maps using the CBLinear module, and then these maps are fused with the features of the right trunk using the CBFuse module; The main function of the CBLinear module is to perform convolution operations on the input feature map and split the output into multiple feature maps according to the predefined number of channels; For the input feature map : (1) Use a convolutional layer with a kernel size of k, a stride of s, and a padding of p: (7) Among them, X represents the input feature map, C in Indicates the number of channels of the input feature map, H indicates the height of the input feature map, and W indicates the width of the input feature map. Indicates that X is a C in Channel, a three-dimensional tensor with a height of H and a width of W, where the elements belong to the real number set R; Conv2d represents a two-dimensional convolution operation, and Y is the output feature map after the convolution operation. Represents the sum of the number of channels of all output feature maps, which is given by The sum of the elements in is obtained, H′ and W′ represent the height and width of the output feature map, It means that Y is a Channel, a three-dimensional tensor of height H′ and width W′, whose elements belong to the set R of real numbers; (2) Split output channels: Split the convolution output Y into m tensors Y i ,in ; Each segmented feature map represents the characteristics of different branches of the network, which is used for subsequent fusion; the mathematical expression is: (8) Among them, Y i The number of channels is ,and , Represents Y i It is a Channel, a three-dimensional tensor of height H′ and width W′, whose elements belong to the set R of real numbers; The CBFuse module is used to fuse multi-scale and multi-channel feature maps into a unified feature map to meet the different feature requirements of the network. It first interpolates and adjusts the feature maps of different sizes, and then adds them together in the feature map dimension to achieve fusion. There are n input feature maps {Y1,Y2,…,Y n }, sizes vary; Y n If the target size is H′×W′, all other feature maps are interpolated and adjusted to this size; (1) Interpolation adjustment: For each Y i , where i < n, is resized to the target size H′×W′ using nearest neighbor interpolation: (9) in, is the feature map after interpolation adjustment, Represents an interpolation operation; (2) All adjusted feature maps, including the last feature map, are added element-wise to obtain the fused output feature map F: (10) Among them, F is the feature map obtained by fusing all feature maps; Combine CBLinear and CBFuse to build a dual-branch feature extraction and fusion backbone.