Mining area remote sensing landslide detection method based on improved YOLOv8 model
The improved YOLOv8 model addresses interference and feature extraction challenges in landslide detection by integrating C2f_DCNv4, MSFE, and LSKNet modules, resulting in enhanced precision and speed for landslide detection.
Patent Information
- Application Number
- CN202510528759.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-15
AI Technical Summary
The existing remote sensing landslide detection methods have problems such as poor detection accuracy, difficulty in extracting landslide features, and slow detection speed.
The improved YOLOv8 model is adopted, and the MSFE module is introduced after the SPPF module of the backbone network, the LSKNet module is introduced into the neck network, and the C2f module is replaced with the C2f_DCNv4 module. At the same time, EIoU is used as the loss function to optimize the model structure and calculation method.
It improves the accuracy and speed of landslide detection, can more accurately identify small-target landslides and broken landslide areas, reduces the difficulty of extracting landslide features, and promptly warns when landslides occur.
Smart Images

Figure CN120318693A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of disaster detection, and particularly to a method for remotely sensing landslide detection in mining areas based on an improved YOLOv8 model. Background Art
[0002] Landslides are common geological disasters that cause extensive damage to the global natural environment and infrastructure. Timely and accurate acquisition of post-event landslide inventories is crucial for formulating effective rescue and emergency measures.
[0003] Remote sensing technology has been widely used in the collection of landslide inventories due to its advantages such as low acquisition cost and high acquisition efficiency. Remote sensing technology can obtain a large amount of high-resolution images, thus enabling the training of deep learning models to achieve intelligent detection of landslides.
[0004] Existing remote sensing landslide detection methods have the following defects: Firstly, the change in vegetation coverage around the landslide due to seasonal changes and the change in the terrain where the landslide is located due to geological movements will both interfere with the remote sensing landslide images, thus affecting the detection accuracy; Secondly, most of the landslides that occur in reality have irregular shapes and vary in size, making it difficult to extract landslide features; Thirdly, the model calculation speed is slow, resulting in a slow detection speed. Summary of the Invention
[0005] To overcome the technical defects of poor detection accuracy, difficult extraction of landslide features, and slow detection speed existing in the existing remote sensing landslide detection methods, the present invention provides a method for remotely sensing landslide detection in mining areas based on an improved YOLOv8 model.
[0006] The method for remotely sensing landslide detection in mining areas based on the improved YOLOv8 model provided by the present invention includes the following steps:
[0007] S10. Collect remote sensing images of landslides, annotate the remote sensing images of landslides, and construct a multi-temporal landslide dataset.
[0008] S20. Divide the dataset into a training set, a validation set, and a test set, and introduce negative samples to improve the dataset.
[0009] S30. Construct a YOLOv8 model, replace the C2f module in the backbone network with a C2f_DCNv4 module, introduce an MSFE module after the SPPF module in the backbone network, introduce an LSKNet module at the position corresponding to each prediction head of the head network in the neck network, form an improved YOLOv8 model, and at the same time use EIoU as the loss function of the improved YOLOv8 model.
[0010] S40. Use the dataset to train the improved YOLOv8 model to obtain a target model.
[0011] S50. Use the landslide remote sensing image as the input, and adopt the target model to detect the remote sensing landslide in the mining area.
[0012] Optionally, the C2f_DCNv4 module is obtained by replacing the Bottleneck of the C2f module with DCNv4.
[0013] Optionally, the MSFE module includes two branches. The first branch is a residual connection, and the second branch consists of ECA, average pooling, and depthwise separable convolution.
[0014] Optionally, the expression of the ECA is:
[0015] ;
[0016] ;
[0017] where represents the output feature map, represents the input feature map, represents the attention vector, represents the activation function, represents the one-dimensional convolution operation, represents the global average pooling.
[0018] Optionally, the core component of the LSKNet module is the LSK block, and the LSK block realizes efficient multi-scale feature extraction and fusion through a large kernel convolution sequence and a spatial kernel selection mechanism.
[0019] Optionally, the large kernel convolution sequence is used to expand the receptive field, and the expression is:
[0020] ;
[0021] where represents the dilation rate of the i-th depth convolution, represents the kernel size of the i-th depth convolution, represents the receptive field of the i-th depth convolution, ;
[0022] The constraint condition satisfied is:
[0023] ;
[0024] where .
[0025] Optionally, the spatial kernel selection mechanism performs the following operations:
[0026] 1) Multi-scale feature generation, and the expression is:
[0027] ;
[0028] Among them, represents the i-th feature map, represents the i-th depth convolutional layer;
[0029] Each goes through 1×1 convolution for channel mixing to obtain multi-scale spatial features ;
[0030] 2) Concatenate the features of different scales, and the expression is:
[0031] ;
[0032] 3) Perform average pooling and max pooling on the concatenated features in the channel dimension, and the expression is:
[0033] ;
[0034] Among them, represents the spatial feature descriptor of average pooling, represents the spatial feature descriptor of max pooling;
[0035] 3) Concatenate and , and generate N spatial attention maps through a convolutional layer , and the expression is:
[0036] ;
[0037] Apply the sigmoid activation function to each spatial attention map to obtain N spatial masks, and the expression is:
[0038] ;
[0039] Each spatial mask corresponds to a feature map of the decomposed kernel;
[0040] 4) Multiply the spatial features of each scale with the corresponding spatial mask element by element, and after weighted fusion, obtain the attention feature S, and the expression is:
[0041] ;
[0042] Multiply the input feature X with the attention feature S element by element to finally obtain the output feature.
[0043] Optionally, the expression of the EIoU is:
[0044] ;
[0045] ;
[0046] ;
[0047] ;
[0048] Among them, represents the total loss, represents the overlap loss, represents the center point distance loss, represents the aspect ratio consistency loss, represents the square of the Euclidean distance between the center point of the predicted bounding box and the center point of the target bounding box, represents the square of the difference between the width of the predicted bounding box and the width of the target bounding box, represents the square of the difference between the height of the predicted bounding box and the height of the target bounding box, and respectively represent the width and height of the smallest closed bounding box that can cover the predicted bounding box and the target bounding box.
[0049] Optionally, in step S10, the labelme tool is used to annotate the landslide remote sensing image.
[0050] Optionally, in step S20, the ratio of the training set, the validation set, and the test set is 8:1:1.
[0051] The technical solution provided by the present invention has the following advantages compared with the prior art:
[0052] 1) The mine area remote sensing landslide detection method based on the improved YOLOv8 model provided by the present invention introduces the MSFE module after the SPPF module of the backbone network of the traditional YOLOv8 model, which can hierarchically fuse multi-scale features, effectively enhancing the model's ability to capture features of different scales in the landslide area. Especially for the characteristics of variable shapes and large size differences of landslide bodies in remote sensing images, the model can more accurately identify small target landslides and fragmented landslide areas, reducing the difficulty of landslide extraction;
[0053] 2) The mine area remote sensing landslide detection method based on the improved YOLOv8 model provided by the present invention introduces the LSKNet module in the neck network of the traditional YOLOv8 model, which can dynamically adjust the receptive field and use a large-scale depth and spatial selection mechanism, effectively processing the extensive context information and multi-scale landslide targets in the landslide remote sensing image, improving the landslide positioning ability, and thus being able to improve the detection accuracy;
[0054] 3) The method for detecting remote sensing landslides in mining areas based on the improved YOLOv8 model provided by the present invention, on the one hand, replaces the C2f module of the traditional YOLOv8 model with the C2f_DCNv4 module, which can enhance the dynamic features and expression ability of the model, reduce redundant operations, and optimize memory access; on the other hand, uses EIoU to replace the CIoU of the traditional YOLOv8 model, which can improve the calculation speed and optimize the positioning accuracy; the cooperation of the two aspects can greatly improve the calculation speed of the model, thereby improving the detection speed, and further enabling timely warning when a landslide occurs. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0057] Figure 1 It represents the flow chart of the method for detecting remote sensing landslides in mining areas in the embodiments of the present invention;
[0058] Figure 2 It represents the network structure diagram of the improved YOLOv8 model in the embodiments of the present invention;
[0059] Figure 3 It represents the network structure diagram of the LSKNet module in the embodiments of the present invention;
[0060] Figure 4 It represents the schematic diagram of the spatial kernel selection mechanism in the embodiments of the present invention;
[0061] Figure 5 It represents the network structure diagram of the MSFE module in the embodiments of the present invention;
[0062] Figure 6 It represents the comparison diagram between the DCNv4 module and the DCNv3 module in the embodiments of the present invention;
[0063] Figure 7 It represents the flow chart of data set collection in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] In order to be able to more clearly understand the above-mentioned objects, features and advantages of the present invention, the following will further describe the solutions of the present invention. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0065] In the following description, many specific details are set forth in order to provide a thorough understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0066] The following combines Figures 1 to 7 to elaborate in detail on the specific embodiments of the present invention.
[0067] This embodiment provides a remote sensing landslide detection method for mining areas based on an improved YOLOv8 model, including steps S10 to S50.
[0068] S10. Collect landslide remote sensing images, annotate the landslide remote sensing images, and construct a multi-temporal landslide dataset.
[0069] Specifically, use a DJI drone to collect landslide remote sensing images, and use the labelme tool to annotate the landslide remote sensing images.
[0070] S20. Divide the dataset into a training set, a validation set, and a test set, and introduce negative samples to improve the dataset.
[0071] Specifically, the ratio of the training set, the validation set, and the test set is 8:1:1.
[0072] S30. Construct a YOLOv8 model, replace the C2f module in the backbone network with a C2f_DCNv4 module, introduce an MSFE module after the SPPF module in the backbone network, and introduce an LSKNet module at the position corresponding to each prediction head of the head network in the neck network to form an improved YOLOv8 model. At the same time, use EIoU as the loss function of the improved YOLOv8 model.
[0073] Specifically, the C2f_DCNv4 module is obtained by replacing the Bottleneck of the C2f module with DCNv4. The C2f_DCNv4 module streamlines the operation design, removes unnecessary calculation steps, simplifies the operation process, and improves efficiency. Through multi-channel processing, reducing redundant calculations, vectorized load / storage, using half-precision data types, and removing the softmax normalization operation in spatial aggregation, the C2f_DCNv4 module is made more efficient.
[0074] More specifically, DCNv4 is an efficient dynamic sparse operator, which reduces redundant operations and improves processing speed by optimizing memory access. At the same time, it reduces the memory access cost by reducing unnecessary memory access requests, thereby accelerating the execution speed of the operation.
[0075] Details of the memory access optimization strategy include the following aspects:
[0076] 1) Multi-channel processing: DCNv4 uses one thread to process multiple channels instead of each thread processing a single channel. This can reduce the memory access requests for loading sampling offsets and aggregating weight values, thus saving memory access costs;
[0077] 2) Reducing redundant calculations: By reusing bilinear interpolation coefficients and other methods, some redundant calculations are reduced, saving time. Although the time saved in a single operation may not be much, it can accumulate into a significant efficiency improvement in large-scale operations;
[0078] 3) Vectorized loading / storing: Adopting vectorized loading / storing operations can reduce the workload of each thread, thus accelerating the execution speed of the GPU kernel. By optimizing the data loading and storing methods, the execution efficiency of the operation is improved;
[0079] 4) Using half-precision data types: DCNv4 adopts half-precision data types, reducing the number of bytes that the kernel needs to read and write, improving data throughput. This can further increase the data transfer speed and improve the execution efficiency of the operation.
[0080] In this embodiment, by designing the C2f_DCNv4 module, the detection speed of the model can be improved as a whole.
[0081] Specifically, the MSFE module includes two branches. The first branch is a residual connection, and the second branch consists of ECA, average pooling, and depthwise separable convolution. The first branch is used to alleviate the vanishing gradient and accelerate the model training speed, and the second branch can fully extract landslide feature information.
[0082] More specifically, the expression of ECA is:
[0083] ;
[0084] ;
[0085] Among them, represents the output feature map, represents the input feature map, represents the attention vector, represents the activation function, represents the one-dimensional convolution operation, represents the global average pooling.
[0086] After the ECA adaptively determines the kernel size k after using global average pooling without dimensionality reduction to aggregate features, it then performs a one-dimensional convolution with kernel size k on the feature map. Next, it applies the Sigmoid function to obtain the attention vector. Finally, it scales each channel of the input feature map by multiplying the corresponding elements in the attention vector to obtain the output feature map. The ECA part introduces an attention mechanism that aggregates features through global average pooling and adaptively determines the kernel size to enhance the model's attention to important features.
[0087] More specifically, average pooling calculates the average value of the pixels in the relevant region of the input feature map according to the kernel size. The advantage of average pooling is that there are no parameters to be optimized, so it can avoid overfitting.
[0088] More specifically, spatially separable convolution divides the standard convolution operation into multiple small kernel convolution operations in the spatial dimension.
[0089] In this embodiment, by introducing the MSFE module into the backbone network of the YOLOv8 model to perform hierarchical fusion of multi-scale features, the detection ability of the model for landslides of different sizes and shapes is effectively enhanced.
[0090] Specifically, the core component of the LSKNet module is the LSK block, which realizes efficient multi-size feature extraction and fusion through a large kernel convolution sequence and a spatial kernel selection mechanism.
[0091] It is easy to understand that larger-sized kernel convolutions are constructed by explicitly decomposing them into a depth convolution sequence with a significantly growing kernel and introducing dilation.
[0092] More specifically, the large kernel convolution sequence is used to expand the receptive field, and the expression is:
[0093] ;
[0094] Among them, represents the dilation rate of the i-th depth convolution, represents the kernel size of the i-th depth convolution, represents the receptive field of the i-th depth convolution, ;
[0095] The constraint conditions satisfied are:
[0096] ;
[0097] Among them, .
[0098] It should be noted that the increase in the size of the kernel and the dilation rate ensures that the receptive field expands fast enough, while also setting an upper limit on the dilation rate to ensure that dilated convolutions do not introduce gaps between feature maps. Such a design has two advantages: First, it clearly generates multiple features with different large receptive fields, which makes it easier to select subsequent convolutional kernels; Second, sequential decomposition is more efficient than simply applying a single larger convolutional kernel.
[0099] More specifically, the spatial kernel selection mechanism operates as follows:
[0100] 1) Multi-scale feature generation, with the expression:
[0101] ;
[0102] Among them, represents the i-th feature map, represents the i-th depth convolutional layer;
[0103] Each goes through 1×1 convolution for channel mixing to obtain multi-scale spatial features ;
[0104] This step uses a series of decomposed depthwise separable convolutions with different receptive fields to obtain rich context information features from different ranges of inputs;
[0105] 2) Concatenate features of different scales, with the expression:
[0106] ;
[0107] 3) Perform average pooling and max pooling on the concatenated features in the channel dimension, with the expression:
[0108] ;
[0109] Among them, represents the spatial feature descriptor of average pooling, represents the spatial feature descriptor of max pooling;
[0110] This step is used to effectively extract the space;
[0111] 3) Concatenate and and generate N spatial attention maps through a convolutional layer with the expression:
[0112] ;
[0113] For each spatial attention map Apply the sigmoid activation function to obtain N spatial masks, and the expression is:
[0114] ;
[0115] Each spatial mask corresponds to a feature map of a decomposed kernel;
[0116] 4) Multiply the spatial features at each scale with the corresponding spatial masks element-wise, and after weighted fusion, obtain the attention feature S, and the expression is:
[0117] ;
[0118] Multiply the input feature X with the attention feature S element-wise to finally obtain the output feature.
[0119] It should be noted that the LSKNet module is a repeatable stacking module, including several LSK blocks. Each LSK block includes two residual sub-blocks, namely the large kernel selection sub-block (LKSelection) and the feed-forward network sub-block (FFN): the large kernel selection sub-block is used to dynamically adjust the receptive field of the network, and it includes 2 fully connected layers, a GELU activation function, and an LSK module; the feed-forward network sub-block is used for channel mixing and feature refinement, and it includes 2 fully connected layers, a depth convolution, and a GELU activation function.
[0120] In this embodiment, by introducing LSKNet into the neck network of the YOLOv8 model, the attention to the context area can be effectively improved, which is a supplement to the defects of the channel and spatial attention mechanisms; LSKNet introduces multiple convolutional kernels and aggregates feature information along the channel dimension, making it possible to be perceived flexibly and accurately; the LSKNet module meets the requirements of landslide detection under complex backgrounds, can efficiently focus on landslide-related spatial regions, capture richer information, and improve the positioning ability by using spatial attention, and finally improve the success rate of landslide detection.
[0121] Specifically, the expression of EIoU is:
[0122] ;
[0123] ;
[0124] ;
[0125] ;
[0126] Among them, represents the total loss, represents the overlap loss, Represents the center point distance loss, Represents the aspect ratio consistency loss, Represents the square of the Euclidean distance between the center point of the predicted bounding box and the center point of the target bounding box, Represents the square of the difference between the width of the predicted bounding box and the width of the target bounding box, Represents the square of the difference between the height of the predicted bounding box and the height of the target bounding box, and Represent the width and height of the smallest closed bounding box that can cover the predicted bounding box and the target bounding box respectively.
[0127] It should be noted that the penalty term of EIoU is based on the penalty term of CIoU, and the influence factor of the aspect ratio is split to calculate the length and width of the target bounding box and the predicted bounding box respectively. This loss function contains three parts: overlap loss, center point distance loss, and aspect ratio consistency loss. The first two parts continue the method in CIoU, but the aspect ratio consistency loss directly minimizes the difference between the width and height of the target bounding box and the predicted bounding box, making the convergence speed faster. In addition, existing IoU-based losses, such as CIoU and GIoU, cannot effectively measure the difference between the target box and the anchor, resulting in a slow convergence speed and inaccurate positioning of the BBR (bounding box regression) model optimization; EIoU can comprehensively consider the matching of position and size by adding Focal to focus on high-quality anchor boxes, thus significantly improving the positioning accuracy of the target detection model, especially when the size and shape of the target object change greatly.
[0128] S40. Use the dataset to train the improved YOLOv8 model to obtain the target model.
[0129] S50. Use the landslide remote sensing image as the input and adopt the target model to detect the remote sensing landslide in the mining area.
[0130] The following conducts a detailed demonstration of each improvement point and the overall effect of the mining area remote sensing landslide detection method in this embodiment through several comparative experiments.
[0131] It should be noted that all the following experiments are carried out in the same hardware and software environment, using PyTorch as the deep learning framework and the Linux operating system, and YOLOv8s as the benchmark model. The GPU used is three NVIDIA TITANXp, and the GPU memory size is 36G. All models are trained for 300 epochs, and the IoU and confidence are set to 0.45 during the test and training processes. This dataset is divided into training set, validation set, and test set according to the ratio of 8:1:1. The dataset contains 2500 images. The training set, validation set, and test set contain 2000 images, 250 images, and 250 images respectively.
[0132] All the following experiments evaluate the model using six metrics, namely: Precision - accuracy rate, which represents the proportion of the number of landslides correctly identified (TP) to the total number of landslides identified as such (the number of correctly identified landslides TP + the number of misidentified landslides FP); Recall - recall rate, which represents the proportion of the number of landslides correctly identified (TP) to the total number of actual landslides (the number of correctly identified landslides TP + the number of omitted landslides FN); mAP@0.5 represents the average precision when the IoU threshold is 0.5, used to measure the area under the P-R curve of the model, indicating the comprehensive detection performance of the model; mAP@0.5:0.95 represents the average precision when the IoU threshold varies from 0.5 to 0.95, used to more comprehensively evaluate the performance of the model under different IoU thresholds; GFLOP - floating point operation count, which is used to measure the computational complexity during the execution of the model, and the smaller the value, the smaller the computational amount of the model; FPS - frames per second, which represents the number of images that the model can process in one second, used to quantify the operation speed of the model, and the larger the value, the faster the model processing speed.
[0133] Experiment 1:
[0134] Content: Introduce the MSFE module after the SPPF module in the backbone network of the YOLOv8s model.
[0135] Comparison object: The YOLOv8s model and the YOLOv8s model after introducing the MSFE module.
[0136] Purpose: To demonstrate the effect of the MSFE module.
[0137] Data statistics:
[0138]
[0139] Analysis: Introducing the MSFE module after the SPPF module enables the model to more effectively capture the dependencies between different landslide features, allowing the model to more effectively process and fuse important feature information, while taking into account landslides of different shapes and sizes, thereby improving the overall detection performance.
[0140] Conclusion: The experimental data shows that after introducing the MSFE module, the accuracy rate and recall rate of the model have been significantly improved. However, at the same time, the increase in the number of model parameters has led to a decrease in the detection speed.
[0141] Experiment 2:
[0142] Content: On the basis of Experiment 1, further introduce the LSKNet module after the C2f module in the neck network of the YOLOv8s model.
[0143] Comparison objects: YOLOv8s model, YOLOv8s model with the MSFE module introduced, and YOLOv8s model with the MSFE module and LSKNet module introduced.
[0144] Purpose: To demonstrate the effect of the LSKNet module.
[0145] Data statistics:
[0146]
[0147] Analysis: The LSKNet module meets the requirements of landslide detection in complex backgrounds, can efficiently focus on landslide-related spatial regions, capture richer information, improve the positioning ability using spatial attention, and ultimately increase the success rate of landslide detection. The LSKNet module utilizes the characteristics of multiple depthwise separable convolutions with large convolutional kernels to generate features with a wide receptive field, thereby reducing the number of model parameters, and dynamically adjusting the receptive field to adapt to the complex and changeable landslide background environment.
[0148] Conclusion: Experimental data shows that after introducing the LSKNet module, the accuracy and recall rate increase, the number of parameters decreases, and the detection speed improves.
[0149] Experiment Three:
[0150] Content: Based on Experiment Two, further replace the C2f module in the backbone network of the YOLOv8s model with the designed C2f_DCNv4 module.
[0151] Comparison objects: YOLOv8s model, YOLOv8s model with the MSFE module introduced, YOLOv8s model with the MSFE module and LSKNet module introduced, and YOLOv8s model with the MSFE module, LSKNet module, and C2f_DCNv4 module introduced.
[0152] Purpose: To demonstrate the effect of the C2f_DCNv4 module.
[0153] Data statistics:
[0154]
[0155] Analysis: The C2f_DCNv4 module achieves a faster convergence speed and three times the forward processing speed by removing softmax normalization and memory access optimization, thereby significantly improving the landslide detection speed without losing accuracy.
[0156] Conclusion: Experimental data shows that after introducing the LSKNet module, on the basis of basically unchanged detection accuracy and recall rate, the number of model parameters is significantly reduced, and the landslide detection speed is significantly improved.
[0157] Experiment 4:
[0158] Content: Based on Experiment 3, further replace the CIoU loss function of the YOLOv8s model with EIoU.
[0159] Comparison objects: YOLOv8s model, YOLOv8s model with the MSFE module introduced, YOLOv8s model with the MSFE module and LSKNet module introduced, YOLOv8s model with the MSFE module, LSKNet module and C2f_DCNv4 module introduced, and YOLOv8s model with the loss function of EIoU (i.e., the improved YOLOv8s model).
[0160] Purpose: To demonstrate the effect of EIoU.
[0161] Data statistics:
[0162]
[0163] Analysis: EIoU has a more stable training process. By introducing penalty terms for the center point distance and aspect ratio, even when IoU is 0, EIoU can provide effective gradient information to ensure that the model can continue to learn. And it can improve the localization accuracy. By comprehensively considering the matching of position and size, EIoU can significantly improve the localization accuracy of the object detection model, especially when the size and shape of the target object change greatly, making the model have a faster convergence speed and higher accuracy.
[0164] Conclusion: Experimental data shows that by replacing the loss function with EIOU, all indicators have been improved.
[0165] Experiment 5:
[0166] To further verify the superiority and effectiveness of the improved algorithm, the present invention conducts a comparative experiment with the current mainstream models. In the experiment, the proposed model is compared with lightweight YOLOv5, YOLOv7, YOLOv8, and YOLOv9, which have high accuracy in landslide disaster identification. The data statistics are as follows:
[0167]
[0168] Experimental data show that compared with other models, YOLOv5 has a smaller number of parameters, but its detection effect on small targets and dense tasks is poor; YOLOv7 adopts a special design that converts the object detection task into a single convolutional operation, which can effectively detect various objects; compared with other models, YOLOv7 performs better in the detection of small targets; YOLOv8 achieves a balance between model size, computational efficiency, and performance; YOLOv9 introduces more technical improvements, such as feature pyramid networks and attention mechanisms, etc., which improves the accuracy of object detection, and at the same time, YOLOv9 has a better detection effect on dense targets. The present invention makes overall detection improvements based on YOLOv8. The accuracy and recall rate of the present invention are significantly higher than those of other models, and the number of parameters GFLOPs is only slightly more than that of YOLOv5, YOLOv8, and YOLOv9. The detection speed and accuracy are higher than those of these benchmark models. In summary, the model of the present invention has a more comprehensive improvement in detection speed and accuracy compared with these mainstream models.
[0169] Experiment Six:
[0170] To verify the robustness of the model, in addition to conducting experiments on the self-constructed dataset, the model was also experimented on the landslide public dataset (Guizhou Bijie landslide dataset). The data statistics are as follows:
[0171]
[0172] Experimental data show that when comparing the benchmark YOLOv8 model with the improved YOLOv8 model of the present invention, through a large number of experiments, it is proved that the improved model of the present invention has a greater improvement in accuracy and speed than the benchmark model.
[0173] The above are only specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Although the foregoing embodiments have been described in detail, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments, and they should all be covered by the protection scope of the claims.
Claims
1. A remote sensing landslide detection method for mining areas based on an improved YOLOv8 model, characterized in that, It includes the following steps: S10. Collect landslide remote sensing images, annotate the landslide remote sensing images, and construct a multi-temporal landslide dataset; S20. Divide the dataset into a training set, a validation set, and a test set, and introduce negative samples to improve the dataset; S30. Construct a YOLOv8 model, replace the C2f module in the backbone network with a C2f_DCNv4 module, introduce an MSFE module after the SPPF module in the backbone network, and introduce an LSKNet module at the position corresponding to each prediction head in the head network in the neck network to form an improved YOLOv8 model. At the same time, use EIoU as the loss function of the improved YOLOv8 model; S40. Use the dataset to train the improved YOLOv8 model to obtain a target model; S50. Use the landslide remote sensing image as the input and use the target model to detect remote sensing landslides in the mining area.
2. The method for remotely sensing landslide detection in mining areas based on the improved YOLOv8 model according to claim 1, wherein, The C2f_DCNv4 module is obtained by replacing the Bottleneck of the C2f module with DCNv4.
3. The method for detecting remote sensing landslides in mining areas based on the improved YOLOv8 model according to claim 1, characterized in that, The MSFE module includes two branches. The first branch is a residual connection, and the second branch consists of ECA, average pooling, and depthwise separable convolution.
4. The method for detecting remote sensing landslides in mining areas based on the improved YOLOv8 model according to claim 3, characterized in that, The expression of the ECA is: ; ; Among them, represents the output feature map, represents the input feature map, represents the attention vector, represents the activation function, represents the one-dimensional convolution operation, represents the global average pooling.
5. The method for detecting remote sensing landslides in mining areas based on the improved YOLOv8 model according to claim 1, wherein, The core component of the LSKNet module is the LSK block, and the LSK block realizes efficient multi-scale feature extraction and fusion through a large kernel convolution sequence and a spatial kernel selection mechanism.
6. The method for remotely sensing landslide detection in mining areas based on the improved YOLOv8 model according to claim 5, wherein, The large kernel convolution sequence is used to expand the receptive field, and the expression is: ; Among them, represents the dilation rate of the i-th depth convolution, represents the kernel size of the i-th depth convolution, represents the receptive field of the i-th depth convolution, ; The satisfied constraint condition is: ; Among them, .
7. The method for remotely sensing landslide detection in mining areas based on the improved YOLOv8 model according to claim 5, wherein, The spatial kernel selection mechanism performs the following operations: 1) Multi-scale feature generation, and the expression is: ; Among them, represents the i-th feature map, represents the i-th depth convolutional layer; Each undergoes channel mixing through 1×1 convolution to obtain multi-scale spatial features ; 2) Concatenate the features of different scales, and the expression is: ; 3) Perform average pooling and max pooling on the concatenated features in the channel dimension, and the expression is: ; Among them, represents the spatial feature descriptor of average pooling, represents the spatial feature descriptor of max pooling; 3) Concatenate and , and generate N spatial attention maps through a convolutional layer with the expression: ; For each spatial attention map Apply the sigmoid activation function to obtain N spatial masks, with the expression: ; Each spatial mask corresponds to a feature map of a decomposed kernel; 4) Multiply the spatial features at each scale with the corresponding spatial mask element by element, and after weighted fusion, obtain the attention feature S, with the expression: ; Perform an element-wise product of the input feature X and the attention feature S to finally obtain the output feature.
8. The method for remotely sensed landslide detection in mining areas based on the improved YOLOv8 model according to claim 1, wherein, The expression of the EIoU is: ; ; ; ; Among them, represents the total loss, represents the overlap loss, represents the center point distance loss, represents the aspect ratio consistency loss, represents the square of the Euclidean distance between the center point of the predicted bounding box and the center point of the target bounding box, represents the square of the difference between the width of the predicted bounding box and the width of the target bounding box, represents the square of the difference between the height of the predicted bounding box and the height of the target bounding box, and represent the width and height of the smallest closed bounding box that can cover the predicted bounding box and the target bounding box, respectively.
9. The method for detecting remote sensing landslides in mining areas based on the improved YOLOv8 model according to any one of claims 1 to 8, characterized in that, In step S10, use the labelme tool to annotate the landslide remote sensing images.
10. The method for remotely sensing landslide detection in mining areas based on the improved YOLOv8 model according to any one of claims 1 to 8, characterized in that, In step S20, the ratio of the training set, the validation set, and the test set is 8:1:1.