Forest disturbance event type identification method, device, equipment, medium and product

By adding convolutional block attention module and deep network structure in the YOLOv10 model and improving the loss function, the problem of low detection efficiency of small and medium-sized forest disturbance type recognition is solved, and efficient and accurate forest disturbance event type recognition is achieved.

CN120495775APending Publication Date: 2025-08-15RES INST OF FOREST RESOURCE INFORMATION TECHN CHINESE ACADEMY OF FORESTRY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510645881.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Existing forest perturbation type recognition methods are inefficient when detecting small targets and are prone to ignore small target samples. Deep learning models such as YOLOv10 are difficult to accurately capture complex forest perturbation features.

Method used

Add a convolutional block attention module to the backbone network of the YOLOv10 model, a deep network structure is added to the neck network, and a small object detection head is added to the output end. At the same time, the loss function is improved to use normalized transmission distance instead of transmission ratio, and improve the small object detection accuracy.

Benefits of technology

It improves the detection efficiency and accuracy of forest disturbance event type recognition, especially when detecting small targets, and enhances the detection ability of targets at different scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495775A_ABST
    Figure CN120495775A_ABST
Patent Text Reader

Abstract

The invention discloses a forest disturbance event type identification method and device, equipment, a medium and a product, and relates to the technical field of forest protection and management. According to the method, a first-stage target detection algorithm with higher efficiency is selected to carry out forest disturbance type identification application, aiming at the condition that a small target sample is easily ignored during target detection of an existing YOLOv10 network model, a convolution block attention module is added in a backbone network of the YOLOv10 model, one or more deep network structures are added in a neck network of the YOLOv10 model, and the target detection efficiency is improved. A small target detection head is additionally arranged at the output end of a YOLOv10 model, a YOLOv10 network model is improved, intersection-union-comparison-union in an original loss function is replaced with a normalized transmission distance, and the loss function is improved, so that the method has more obvious advantages in detection and classification of small sample targets such as forest weak disturbance and the like. And small target detection is realized while high efficiency of a one-stage target detection algorithm is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of forest protection and management, and in particular to a method, device, equipment, medium and product for identifying the type of forest disturbance event. Background Art

[0002] Identifying the types of forest disturbance events can accurately grasp the changes in forest ecosystems under the influence of natural factors (such as fires, pest and disease outbreaks, storms, etc.) and human activities (such as logging, afforestation, land use changes, etc.), understand the succession laws and dynamic equilibrium mechanisms of forest ecosystems, and provide basic data and scientific basis for the protection and management of forest ecosystems. Existing methods for identifying forest disturbance types can be based on machine learning or deep learning models. Traditional machine learning algorithms require manual design when extracting features, and are trained on relatively small data sets and achieve good results. However, for some complex nonlinear relationships and high-dimensional data in forest disturbances, machine learning models cannot accurately capture the complex patterns in the data. Deep learning models can automatically learn features from large amounts of data, avoiding the tedious process of manually designing features, and can learn more complex and abstract feature representations, thereby better adapting to the complex situations in forest disturbance type identification.

[0003] Deep learning target detection algorithms are mainly divided into two types, namely two-stage target detection algorithms and one-stage target detection algorithms. (1) The two-stage target detection algorithm is an important method in the target detection task in the field of computer vision. First, the image is traversed to obtain candidate areas that may contain targets, and then the candidate areas are further analyzed and identified to obtain the final detection results. For example, Lee et al. (2020) used U-Net and SegNet algorithms to classify deforestation areas, with an accuracy of 98.4% for forest areas and 88.5% for non-forest areas; Huang et al. (2022) used SqueezeNet network to identify forest pine wilt disease with an accuracy of 94.90%; Pisl et al. (2024) mapped the driving factors of tropical forest loss based on Convolutional Neural Networks (CNN), and classified the driving factors of forest loss (agriculture, deforestation, wildfires, etc.) with an accuracy of up to 91%, but this method is inefficient. (2) The one-stage target detection algorithm does not require a preset selection area, which significantly improves efficiency. The YOLO (You Only Look Once) series of models is a typical example of a one-stage object detection algorithm. Redmon et al. (2016) first proposed the YOLOv1 algorithm, which pioneered the transformation of the detection problem into a regression problem, ushering in a new era of one-stage object detection. While fast, the YOLOv1 algorithm suffers from a low recall rate, prone to false detections, and struggles to detect small objects. Summary of the Invention

[0004] The purpose of this application is to provide a method, device, equipment, medium and product for identifying the type of forest disturbance event, so as to improve the detection efficiency while realizing the detection of small targets.

[0005] To achieve the above objectives, this application provides the following solutions.

[0006] In a first aspect, the present application provides a method for identifying forest disturbance event types, comprising:

[0007] Construct a sample library of forest disturbance types;

[0008] Constructing an improved YOLOv10 model; the improved YOLOv10 model is obtained by improving the YOLOv10 model, wherein the improvement method is as follows: adding a convolutional block attention module to the backbone network of the YOLOv10 model, adding one or more deep network structures to the neck network of the YOLOv10 model, and adding a small target detection head at the output end of the YOLOv10 model;

[0009] Constructing an improved loss function; the improved loss function is obtained by reconstructing the original loss function including the intersection-over-union ratio, and the improvement method is: replacing the intersection-over-union ratio in the original loss function with the normalized transmission distance;

[0010] Training the improved YOLOv10 model using the forest disturbance type sample library and the improved loss function to obtain a trained YOLOv10 model;

[0011] The trained YOLOv10 model is used to identify the types of forest disturbance events.

[0012] In a second aspect, the present application provides a device for identifying the type of a forest disturbance event, wherein the device applies the above-mentioned method for identifying the type of a forest disturbance event, and comprises:

[0013] Sample library construction module, used to construct a forest disturbance type sample library;

[0014] A model construction module for constructing an improved YOLOv10 model; the improved YOLOv10 model is obtained by improving the YOLOv10 model by adding a convolutional block attention module to the backbone network of the YOLOv10 model, adding one or more deep network structures to the neck network of the YOLOv10 model, and adding a small object detection head at the output end of the YOLOv10 model;

[0015] A loss function construction module is used to construct an improved loss function; the improved loss function is obtained by reconstructing the original loss function containing the intersection-over-union ratio, and the improvement method is: replacing the intersection-over-union ratio in the original loss function with the normalized transmission distance;

[0016] A model training module is used to train the improved YOLOv10 model using the forest disturbance type sample library and the improved loss function to obtain a trained YOLOv10 model;

[0017] The recognition module is used to identify the types of forest disturbance events using the trained YOLOv10 model.

[0018] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned forest disturbance event type identification method.

[0019] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned forest disturbance event type identification method when executed by a processor.

[0020] In a fifth aspect, the present application provides a computer program product, including a computer program, which implements the above-mentioned forest disturbance event type identification method when executed by a processor.

[0021] According to the specific embodiments provided in this application, this application has the following technical effects.

[0022] The present application provides a method, apparatus, equipment, medium and product for identifying the type of forest disturbance event. The present application selects a more efficient one-stage target detection algorithm to carry out the application of identifying the type of forest disturbance. In view of the fact that the existing YOLOv10 network model easily ignores small target samples during target detection, a convolutional block attention module is added to the backbone network of the YOLOv10 model, one or more deep network structures are added to the neck network of the YOLOv10 model, a small target detection head is added to the output end of the YOLOv10 model, the YOLOv10 network model is improved, and the normalized transmission distance is used instead of the intersection-over-union ratio in the original loss function to improve the loss function, so that its advantages in the detection and classification of small sample targets such as weak forest disturbances are more obvious, while maintaining the high efficiency of the one-stage target detection algorithm, small target detection is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 A flowchart of a method for identifying forest disturbance event types provided in one embodiment of the present application.

[0025] Figure 2 A schematic diagram of the structure of a convolutional block attention module provided in one embodiment of the present application.

[0026] Figure 3 This is a structural diagram of the improved YOLOv10 model provided in one embodiment of the present application.

[0027] Figure 4 This is an example diagram of a forest disturbance type sample library provided in one embodiment of the present application.

[0028] Figure 5 Comparison of class activation heatmaps for the YOLOv10 model and the improved YOLOv10 model provided in one embodiment of the present application.

[0029] Figure 6 This is a graph of mAP curves for detecting various types of forest disturbances provided in an embodiment of the present application.

[0030] Figure 7 This is a schematic diagram of the results of identifying the type of local forest disturbance events in Ning'er County, Pu'er City, Yunnan Province, provided in one embodiment of the present application.

[0031] Figure 8 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0033] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0034] In some studies, the YOLO series of algorithms have been continuously improved, and YOLO9000 (Redmon et al., 2017), YOLOX (Ge et al., 2021), and YOLOv2-v10 (Redmon et al., 2018; Bochkovskiy et al., 2020; Ultralytics et al., 2021; Wang et al., 2022; Li et al., 2023; Wang et al., 2024; Wang et al., 2024) have appeared successively. These YOLO series algorithms have continuously improved in accuracy and obtained accurate detection results while maintaining the running speed. For example, Zhang et al. (2023) conducted forest pine wood nematode disease detection in seven sample sites in northwest Beijing, with an identification accuracy of 92.6%; Yang et al. (2024) proposed an improved multi-stage forest fire smoke detection model SIMCB-Yolo based on YOLOv5, achieving an average recognition accuracy of 85.6%; Zhu et al. (2024) improved the YOLOv7 model to identify pine wilt trees, with an accuracy of 92.81%; Xu et al. (2024) used the improved YOLOv7 model for forest fire detection, with an accuracy increase of 2.39%, showing good performance, but still unable to detect small targets.

[0035] In an exemplary embodiment, a method for identifying forest disturbance event types is provided, such as Figure 1 As shown, the process includes the following steps 101 to 105.

[0036] Step 101, constructing a forest disturbance type sample library;

[0037] Step 102: construct an improved YOLOv10 model; the improved YOLOv10 model is obtained by improving the YOLOv10 model by adding a convolutional block attention module to the backbone network of the YOLOv10 model, adding one or more deep network structures to the neck network of the YOLOv10 model, and adding a small object detection head at the output end of the YOLOv10 model;

[0038] Step 103: construct an improved loss function; the improved loss function is obtained by reconstructing the original loss function including the intersection-over-union ratio, and the improvement method is: replacing the intersection-over-union ratio in the original loss function with the normalized transmission distance;

[0039] Step 104: training the improved YOLOv10 model using the forest disturbance type sample library and the improved loss function to obtain a trained YOLOv10 model;

[0040] Step 105: Use the trained YOLOv10 model to identify the type of forest disturbance event.

[0041] The embodiment of the present application adopts an improved YOLOv10 model to identify the types of forest disturbance patches and identify forest disturbance event types such as deforestation, fire, and insect pests. The specific method is as follows: (1) First, a forest disturbance type sample library (including three types: deforestation, fire, and insect pests) is constructed, and the sample data is enhanced and expanded through Generative Adversarial Networks (GAN); (2) Then, a Convolutional Block Attention Module (CBAM) hybrid attention mechanism is used to improve the ability to refine forest disturbance features, and a deeper network structure is added to the existing YOLOv10 model. At the same time, a small target detection head is added to improve the model's detection ability for targets of different scales; (3) Finally, the Normalized Wasserstein Distance (NWD) indicator, which is more suitable for calculating the bounding box distance of small targets, is used to improve the loss function and reduce the missed detection probability of the improved YOLOv10 model.

[0042] In another exemplary embodiment, in the above step 101, a forest disturbance type sample library is constructed, remote sensing images of forest disturbance patches are screened and cropped, and the sample data are enhanced and expanded through generative adversarial networks (GAN). The final sample library contains three types of forest disturbance types: deforestation, fire, and insect pests.

[0043] The original data source of the forest disturbance sample library is Sentinel-2 and GF-2 images. Among them, the 10-meter resolution Sentinel-2 image is used to produce forest strong disturbance (deforestation and fire) samples, and the 2-meter resolution GF-2 image is used to produce forest weak disturbance (insect pest) samples. Both remote sensing images are used to produce undisturbed forest samples. Each data sample is cropped into a 128 pixel × 128 pixel PNG format image. Then, based on GAN, new samples with a certain similarity to the original samples are generated, which together constitute the forest disturbance type sample library. Specifically, the following steps 201 to step 205 are included.

[0044] Step 201: Screen the Sentinel-2 images to obtain Sentinel-2 images containing strong disturbances, and crop the Sentinel-2 images containing strong disturbances to a preset size to obtain real samples of forest strong disturbances; the strong disturbances include deforestation and fire.

[0045] Step 202: Screening the GF-2 images to obtain GF-2 images containing weak disturbances, and cropping the GF-2 images containing weak disturbances to a preset size to obtain a true forest weak disturbance sample; the weak disturbances include insect pests;

[0046] Step 203: Screen the Sentinel-2 image and the GF-2 image to obtain an undisturbed Sentinel-2 image and an undisturbed GF-2 image, and crop the undisturbed Sentinel-2 image and the undisturbed GF-2 image to a preset size to obtain an undisturbed real sample.

[0047] Step 204: Based on the real forest samples with strong disturbance, the real forest samples with weak disturbance, and the real forest samples without disturbance, a generative adversarial network is used to generate virtual forest samples with strong disturbance, virtual forest samples with weak disturbance, and virtual forest samples without disturbance.

[0048] Step 205 : constructing the forest disturbance type sample library based on the real samples of strong forest disturbance, the real samples of weak forest disturbance, the real samples of no disturbance, the virtual samples of strong forest disturbance, the virtual samples of weak forest disturbance and the virtual samples of no disturbance.

[0049] In another exemplary embodiment, in the above step 102, to address the problem of poor small target classification performance, a Convolutional Block Attention Module (CBAM) hybrid attention mechanism is used to improve the forest disturbance feature refinement capability. Secondly, the YOLOv10 model is improved by adding a deeper network structure to the existing network and adding a small target detection head to enhance the model's detection capability for targets of different scales.

[0050] (1) Combining channel attention and spatial attention to form a CBAM hybrid attention mechanism to enhance the model's ability to capture key information, such as Figure 2 CBAM is a typical lightweight hybrid attention mechanism suitable for analyzing small object samples with complex background information. CBAM considers important information in both the channel and spatial dimensions in the feature map, and generates an attention map to enhance the model's focus on key information. CBAM can be seamlessly integrated into any neural network architecture and is suitable for object detection and classification tasks. It can achieve significant performance improvements in both detection and classification accuracy for datasets with diverse data features.

[0051] In the channel attention module, global mean pooling and global maximum pooling operations are first performed on the input feature map to capture the global information of the feature map and generate two different spatial context descriptors. and Then, the pooled result is sent to the shared fully connected layer for feature transformation to generate the channel attention map M c ∈R c×1×1 , use multi-layer perceptron (MLP) to extract features and obtain the weight of each channel through nonlinear transformation. In order to reduce parameter overhead, the hidden activation size is set to R using the reduction rate r. c / r×1×1 In addition, the two feature maps are added together and a channel attention map is generated through the sigmoid activation function, which represents the importance weight of each channel. The obtained weight is multiplied by the original feature map to achieve weighted channel attention. In this way, the features of important channels are enhanced, while the features of useless channels are suppressed. The calculation formula of the channel attention module is shown in Equation (1).

[0052]

[0053] Among them, M c (F1) is the channel attention map output by the channel attention mechanism, F1 is the input feature map of the channel attention mechanism, AvgPool() is the global average pooling operation, MaxPool() is the global maximum pooling operation, MLP() is the multi-layer perceptron, σ() is the sigmoid activation function, is the output feature map of the global mean pooling operation of the channel attention mechanism, is the output feature map of the global maximum pooling operation of the channel attention mechanism, W0 is the multi-layer perceptron based on The first channel weight matrix generated, W1 is the multilayer perceptron based on The generated second channel weight matrix, W0∈R C / r×C , W1∈R C×C / r , r is the reduction rate, C is the number of channels of the channel attention mechanism, R C / r×C represents a C / r×C dimensional dataset, R C×C / r Represents a dataset of C×C / r dimensions.

[0054] In the above equation (1), the MLP weights W0 and W1 are shared for both inputs, and the ReLU activation function is followed by W0

[0055] In the spatial attention module, the input feature map is firstly subjected to mean pooling and maximum pooling operations in the channel dimension, which can capture the feature statistics of the feature map at each position and generate two two-dimensional feature maps. and The two feature maps are then concatenated to form a new feature descriptor, which contains two different feature statistics at each position. Finally, the concatenated feature descriptor is transformed through a convolutional layer to generate a spatial attention map, which reflects the importance of each position in the feature map. The calculation formula of the spatial attention module is shown in Equation (2).

[0056]

[0057] Among them, M S (F2) represents the spatial attention map output by the spatial attention mechanism, F2 represents the input feature map of the spatial attention mechanism, and f 7×7 () represents the convolution operation with a filter size of 7×7, is the output feature map of the global mean pooling operation of the spatial attention mechanism, It is the output feature map of the global maximum pooling operation of the spatial attention mechanism.

[0058] (2) Due to the fragmented nature of deforestation and pest disturbance samples, the phenomenon of missing small targets often occurs during the detection process. The detection head is a key component in the target detection model, responsible for extracting target information from the feature map, including the target's location, category, and confidence, etc. Therefore, the embodiment of the present application improves the YOLOv10 network structure by adding a deeper network structure and detection head to increase the robustness of the model in detecting small targets in complex scenes. The existing YOLOv10 network has 3 detection heads and is located at the end of the network. By adding detection heads, more detection paths and feature extraction methods can be introduced. At the same time, detection heads of different scales are set so that the model can better adapt to detection targets of different sizes. The improved YOLOv10 network structure is as follows: Figure 3 As shown in the figure, the numbers represent the number of network layers, and the highlighted boxes represent improvements made to the existing YOLOv10 model. The improved YOLOv10 adds the CRAM hybrid attention mechanism to layer 9, expands the network structure from layers 15 to 17, adds a small object detection head to layer 20, and expands the network structure again from layers 21 to 23.

[0059] In another exemplary embodiment, in step 103, the loss function is improved by using a normalized Wasserstein distance (NWD) indicator that is more suitable for small object bounding box distance calculation to reduce the missed detection probability.

[0060] The loss function is a key component in optimizing model performance, and the intersection over union (IoU) (the intersection area divided by the union area, where the intersection area refers to the area of overlap between the predicted bounding box and the true bounding box, and the union area refers to the total area covered by the predicted bounding box and the true bounding box) plays an important role in the loss function as an important metric for measuring the degree of overlap between the predicted bounding box and the true bounding box. The IoU metric, commonly used in the existing YOLOv10 model, severely reduces the detection performance of anchor-based detectors for small target samples. NWD is a small target detection evaluation metric based on the Wasserstein distance (NWD), which can replace the IoU metric in the loss function. Therefore, this method uses NWD to replace the IoU metric in the YOLOv10 model loss function, thereby improving the model's detection accuracy for small target samples.

[0061] Use Wasserstein distance to calculate the distribution distance, for two two-dimensional Gaussian distributions and The second-order Wasserstein distance between is defined as:

[0062]

[0063] Among them, ||·|| F is the F-norm.

[0064] Bounding box A = (cx a ,cy a ,w a ,h a ) and bounding box B=(cx b ,cy b ,w b ,h b )'s Gaussian distribution and It can be further simplified:

[0065]

[0066] A new metric for Normalized Wasserstein Distance (NWD):

[0067]

[0068] in, To predict the Gaussian distribution of the perturbed bounding box, is the Gaussian distribution of the perturbed bounding box for the label, To predict the normalized transfer distance between the perturbed bounding box and the label perturbed bounding box, To predict the second-order transmission distance between the perturbed bounding box and the label perturbed bounding box, cx a 、cy a 、w a 、h a are the x-axis coordinate, y-axis coordinate, width, and height of the predicted perturbation bounding box, cx b 、cy b 、w b 、h b are the x-axis coordinate, y-axis coordinate, width, and height of the label perturbation bounding box respectively. The C' empirical constant is a constant closely related to the proportion of small target samples in the dataset and is often an empirical value.

[0069] In the experiment, the model was adjusted and optimized according to the proportion of small objects in the dataset, and the parameters were continuously tested. Finally, the C' value in the formula, that is, the parameter iou_ratio, was set to 0.6 to obtain the best performance.

[0070] The loss function based on NWD is designed as:

[0071] L NWD =1-NWD(N p , N g ) (6)

[0072] Among them, L NWD To improve the loss function.

[0073] In an exemplary embodiment, in order to illustrate the effects of the technical solutions of the above-mentioned various method embodiments, the following process is set up for verification.

[0074] (1) In the embodiment of the present application, the LabelImg tool was selected when annotating the data set, and the yolo format was selected. The rectangular boundary of the forest disturbance was drawn by visual interpretation, and more than 10,000 image samples were manually produced. In order to expand the data set, GAN was used to supplement the sample data. GAN consists of two parts: a generator and a discriminator. The generator is used to receive the input image and generate new enhanced data samples. These data samples will gradually approach the real data during the training process. The discriminator is used to determine whether the input sample is real or generated by the generator, and to judge the authenticity of the sample as accurately as possible. The embodiment of the present application constructs a U-Net structure generator, and the discriminator is constructed using a binary classification network, which can effectively capture the multi-scale information in the image. Before the training begins, some data enhancement methods based on geometric changes such as random rotation, flipping, and displacement are added to increase the diversity of the data. During the training process of the GAN network, the generator and the discriminator are updated alternately. The generator generates an enhanced image that is as close to the real as possible, while the discriminator distinguishes between the real image and the generated image. This process is continuously iterated. The discriminator is trained using Binary Cross-Entropy Loss (BCE), while the generator is trained using L1 loss. After training, the performance of the generator is evaluated using the test set, and the low-quality images are compared with the enhanced high-quality images, such as Figure 4 As shown in the figure, a total of 25,617 sample images were generated, all 128 pixels × 128 pixels in size. Among them, 8,234 were images of deforestation, 7,168 were images of forest fires, 6,215 were images of forest insect infestations, and 4,000 were images of undisturbed forests.

[0075] (2) In order to explore the contribution of each component in the model system constructed in the embodiment of the present application to the model accuracy, the ablation experiment results were analyzed, and the baseline was the original YOLOv10 model. It can be seen from Table 1 that the classification performance of the model can be effectively improved by the multi-detection head forest perturbation type recognition method based on the hybrid attention mechanism. After adding the "Detect" small target detection head, the network depth is also increased accordingly, thereby improving the model's expressiveness for small target classification. At the same time, adding the CBAM attention mechanism and the NWD module can improve the accuracy, which is higher than the performance improved by adding NWD and CBAM alone, especially avoiding the interference of background information, thereby improving the classification ability of small target samples.

[0076] Table 1 Ablation experiment results

[0077] Module Acc(%) Performance improvements Baseline 87.53 +Detect 90.75 +3.22 +Detect+CBAM 92.42 +4.89 +Detect+NWD 91.56 +4.03 +Detect+CBAM+NWD 93.79 +6.26

[0078] Figure 5In the class activation heatmap, the red and yellow highlighted areas represent the model's contribution to the final output. This method's heatmap shows more detail for larger forest fire samples. For insect infestations and fragmented deforestation, the smaller the sample area, the more pronounced the improvement. The original YOLOv10 model can only roughly determine the sample's location, but it struggles to distinguish between foreground and background, and it's unable to perform fine-grained segmentation between samples.

[0079] The accuracy of the method of the embodiment of the present application is compared with that of the Faster R-CNN, CenterNet, YOLOv9, and YOLOv10 models, and the mAP curve is used to evaluate the performance of different algorithms. The mAP combines the information of precision and recall. The higher the mAP value, the better the overall performance of the model in the target detection task. Figure 6 As shown in the results of Table 2, Figure 6 (a), (b), and (c) are respectively the mAP curves for forest fire detection, the mAP curves for deforestation detection, and the mAP curves for forest pest detection. The method of the embodiment of the present application does not improve the classification performance of forest fire samples much, but for deforestation and forest pest scenes, the performance improvement effect of the method of the embodiment of the present application is obvious, especially for forest pest samples, most of which are small targets. The method provided by the embodiment of the present application is further used to identify the types of local forest disturbance events in Ning'er County, Yunnan Province. The identification results are as follows: Figure 7 shown.

[0080] Table 2 Comparison of mAP results

[0081] algorithm forest fires Deforestation forest pests Faster R-CNN 0.992 0.842 0.903 YOLOv9 0.879 0.791 0.902 YOLOv10 0.923 0.936 0.912 CenterNet 0.870 0.774 0.497 Ours 0.989 0.975 0.978

[0082] Based on the same inventive concept, embodiments of the present application further provide a forest disturbance event type identification device for implementing the aforementioned forest disturbance event type identification method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the forest disturbance event type identification device provided below can be found in the above-described limitations of the forest disturbance event type identification method and will not be further elaborated here.

[0083] In an exemplary embodiment, a device for identifying a forest disturbance event type is provided, comprising:

[0084] Sample library construction module, used to construct a forest disturbance type sample library;

[0085] A model construction module for constructing an improved YOLOv10 model; the improved YOLOv10 model is obtained by improving the YOLOv10 model by adding a convolutional block attention module to the backbone network of the YOLOv10 model, adding one or more deep network structures to the neck network of the YOLOv10 model, and adding a small object detection head at the output end of the YOLOv10 model;

[0086] A loss function construction module is used to construct an improved loss function; the improved loss function is obtained by reconstructing the original loss function containing the intersection-over-union ratio, and the improvement method is: replacing the intersection-over-union ratio in the original loss function with the normalized transmission distance;

[0087] A model training module is used to train the improved YOLOv10 model using the forest disturbance type sample library and the improved loss function to obtain a trained YOLOv10 model;

[0088] The recognition module is used to identify the types of forest disturbance events using the trained YOLOv10 model.

[0089] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 8 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for identifying the type of forest disturbance event is implemented.

[0090] Those skilled in the art will understand that Figure 8The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the above-mentioned method embodiments when executing the computer program.

[0091] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0092] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0093] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0094] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0095] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0096] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0097] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for identifying forest disturbance event types, characterized in that: The steps include: Construct a sample library of forest disturbance types; Constructing an improved YOLOv10 model; the improved YOLOv10 model is obtained by improving the YOLOv10 model, wherein the improvement method is as follows: adding a convolutional block attention module to the backbone network of the YOLOv10 model, adding one or more deep network structures to the neck network of the YOLOv10 model, and adding a small target detection head at the output end of the YOLOv10 model; Constructing an improved loss function; the improved loss function is obtained by reconstructing the original loss function including the intersection-over-union ratio, and the improvement method is: replacing the intersection-over-union ratio in the original loss function with the normalized transmission distance; Training the improved YOLOv10 model using the forest disturbance type sample library and the improved loss function to obtain a trained YOLOv10 model; The trained YOLOv10 model is used to identify the types of forest disturbance events.

2. The method for identifying forest disturbance event types according to claim 1, characterized in that: Construct a sample library of forest disturbance types, including: Screening Sentinel-2 images to obtain those containing strong disturbances, and cropping the Sentinel-2 images containing strong disturbances to a preset size to obtain real samples of strong forest disturbances; the strong disturbances include deforestation and fire; Screening the GF-2 images to obtain GF-2 images containing weak disturbances, and cropping the GF-2 images containing weak disturbances to a preset size to obtain real samples of forest weak disturbances; the weak disturbances include insect pests; The Sentinel-2 images and GF-2 images were screened to obtain undisturbed Sentinel-2 images and GF-2 images, and both the undisturbed Sentinel-2 images and GF-2 images were cropped to a preset size to obtain undisturbed real samples. Based on the real samples of forest with strong disturbance, real samples of forest with weak disturbance and real samples without disturbance, a generative adversarial network is used to generate virtual samples of forest with strong disturbance, virtual samples of forest with weak disturbance and virtual samples without disturbance. The forest disturbance type sample library is constructed based on real samples of strong forest disturbance, real samples of weak forest disturbance, real samples of no disturbance, virtual samples of strong forest disturbance, virtual samples of weak forest disturbance and virtual samples of no disturbance.

3. The method for identifying forest disturbance event types according to claim 1, characterized in that: The convolutional block attention module includes a channel attention mechanism and a spatial attention mechanism connected in sequence.

4. The method for identifying forest disturbance event types according to claim 3, characterized in that: The calculation formula of the channel attention mechanism is: Among them, M c (F1) is the channel attention map output by the channel attention mechanism, F1 is the input feature map of the channel attention mechanism, AvgPool() is the global average pooling operation, MaxPool() is the global maximum pooling operation, MLP() is the multi-layer perceptron, σ() is the sigmoid activation function, is the output feature map of the global mean pooling operation of the channel attention mechanism, is the output feature map of the global maximum pooling operation of the channel attention mechanism, W0 is the multi-layer perceptron based on The first channel weight matrix generated, W1 is the multilayer perceptron based on The generated second channel weight matrix, W0∈R C / r×C , W1∈R C×C / r , r is the reduction rate, C is the number of channels of the channel attention mechanism, R C / r×C Represents a C / r×C dimensional dataset, R C×C / r represents a dataset of C×C / r dimensions; The calculation formula of the spatial attention mechanism is: Among them, M S (F2) represents the spatial attention map output by the spatial attention mechanism, F2 represents the input feature map of the spatial attention mechanism, and f 7×7 () represents the convolution operation with a filter size of 7×7, is the output feature map of the global mean pooling operation of the spatial attention mechanism, It is the output feature map of the global maximum pooling operation of the spatial attention mechanism.

5. The method for identifying forest disturbance event types according to claim 1, characterized in that: The calculation formula for normalized transmission distance is: in, To predict the Gaussian distribution of the perturbed bounding box, is the Gaussian distribution of the label perturbation bounding box, To predict the normalized transfer distance between the perturbed bounding box and the label perturbed bounding box, To predict the second-order transmission distance between the perturbed bounding box and the label perturbed bounding box, cx a 、cy a 、w a 、h a are the x-axis coordinate, y-axis coordinate, width, and height of the predicted perturbation bounding box, cx b 、cy b 、w b 、h b are the x-axis coordinate, y-axis coordinate, width, and height of the label perturbation bounding box, respectively, and C' is an empirical constant.

6. The method for identifying forest disturbance event types according to claim 1, characterized in that: There are two deep network structures, namely the first deep network structure and the second deep network structure; The first deep network structure includes an UpSample layer, a Concat layer, and a C2f layer connected in sequence; The second deep network structure includes a Conv layer, a Concat layer and a C2f layer connected in sequence.

7. A device for identifying forest disturbance event types, characterized in that: The forest disturbance event type identification device applies the forest disturbance event type identification method according to any one of claims 1 to 6, and the forest disturbance event type identification device comprises: Sample library construction module, used to construct a forest disturbance type sample library; A model construction module for constructing an improved YOLOv10 model; the improved YOLOv10 model is obtained by improving the YOLOv10 model by adding a convolutional block attention module to the backbone network of the YOLOv10 model, adding one or more deep network structures to the neck network of the YOLOv10 model, and adding a small object detection head at the output end of the YOLOv10 model; A loss function construction module is used to construct an improved loss function; the improved loss function is obtained by reconstructing the original loss function containing the intersection-over-union ratio, and the improvement method is: replacing the intersection-over-union ratio in the original loss function with the normalized transmission distance; A model training module is used to train the improved YOLOv10 model using the forest disturbance type sample library and the improved loss function to obtain a trained YOLOv10 model; The recognition module is used to identify the types of forest disturbance events using the trained YOLOv10 model.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the forest disturbance event type identification method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for identifying the type of forest disturbance event according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for identifying the type of forest disturbance event according to any one of claims 1 to 6 is implemented.