Mining area remote sensing landslide detection method based on improved YOLOv10 model

The improved YOLOv10 model addresses interference and computation issues in remote sensing landslide detection by using AISFI, LSKNet, and EIoU, enhancing feature extraction and fusion for precise and fast landslide detection.

CN120318694APending Publication Date: 2025-07-15TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510529011.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing remote sensing landslide detection methods have problems such as poor detection accuracy, difficulty in extracting landslide features, and slow detection speed.

Method used

The improved YOLOv10 model is adopted, and the SPPF module of the backbone network is replaced as the AISFI module, the LSKNet module is introduced to the neck network, and EIoU is used as the loss function, combining the C2f_DCNv4 module to optimize the model structure.

Benefits of technology

It improves the accuracy and speed of landslide detection, can timely warn of landslide disasters, enhances the processing ability of landslides of different shapes and sizes, reduces the difficulty of feature extraction and optimizes the calculation speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318694A_ABST
    Figure CN120318694A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of disaster detection, in particular to a mining area remote sensing landslide detection method based on an improved YOLOv10 model, and mainly solves the technical problems that an existing remote sensing landslide detection method is poor in detection precision, difficult in landslide feature extraction and low in detection speed. According to the method, an SPPF module of a backbone network of a traditional YOLOv10 model is replaced by an AISFI module, so that the model can process and fuse important feature information more effectively, landslides with different shapes and sizes are considered, the overall detection performance is improved, and the extraction difficulty of landslide features is reduced; an LSKNet module is introduced into the neck network, so that the landslide positioning capability can be improved, and the detection precision can be improved; in addition, the C2f module is replaced by the C2fDCNv4 module, and the EIoU is adopted to replace the CIoU, so that the calculation speed of the model can be greatly improved, the detection speed is improved, and early warning can be performed in time when the landslide occurs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of disaster detection, and particularly to a method for remote sensing landslide detection in mining areas based on an improved YOLOv10 model. Background Art

[0002] Landslides are common geological disasters that cause extensive damage to the global natural environment and infrastructure. Timely and accurate acquisition of post-event landslide inventories is crucial for formulating effective rescue and emergency measures.

[0003] Remote sensing technology has been widely used in the collection of landslide inventories due to its advantages such as low acquisition cost and high acquisition efficiency. Remote sensing technology can obtain a large amount of high-resolution images, enabling the training of deep learning models to achieve intelligent detection of landslides.

[0004] Existing remote sensing landslide detection methods have the following defects: First, the change in vegetation coverage around the landslide due to seasonal changes and the change in the terrain of the landslide location due to geological movements will both interfere with remote sensing landslide images, thus affecting the detection accuracy; Second, most of the landslides that occur in reality have irregular shapes and vary in size, making it difficult to extract landslide features; Third, the model calculation speed is slow, resulting in a slow detection speed. Summary of the Invention

[0005] To overcome the technical defects of poor detection accuracy, difficult extraction of landslide features, and slow detection speed existing in the existing remote sensing landslide detection methods, the present invention provides a method for remote sensing landslide detection in mining areas based on an improved YOLOv10 model.

[0006] The method for remote sensing landslide detection in mining areas based on the improved YOLOv10 model provided by the present invention includes the following steps:

[0007] S10. Collect remote sensing images of landslides, annotate the remote sensing images of landslides, and construct a dataset of multi-temporal landslides;

[0008] S20. Divide the dataset into a training set, a validation set, and a test set, and introduce negative samples to improve the dataset;

[0009] S30. Construct a YOLOv10 model, replace the C2f module in the backbone network and the neck network with the C2f_DCNv4 module, replace the SPPF module in the backbone network with the AISFI module, introduce the LSKNet module at the position corresponding to each double-label assignment prediction head of the head network in the neck network to form an improved YOLOv10 model, and at the same time use EIoU as the loss function of the improved YOLOv10 model;

[0010] S40. Use the dataset to train the improved YOLOv10 model to obtain a target model;

[0011] S50. Use the target model for remote sensing landslide detection in the mining area with landslide remote sensing images as input.

[0012] Optionally, the C2f_DCNv4 module is obtained by replacing the Bottleneck of the C2f module with DCNv4.

[0013] Optionally, the AISFI module includes AIFI and Conv located before the AIFI. The AIFI is used to automatically learn the correlation between features, and the Conv is used to extract local features.

[0014] Optionally, the core component of the LSKNet module is the LSK block, and the LSK block realizes efficient multi-scale feature extraction and fusion through a large kernel convolution sequence and a spatial kernel selection mechanism.

[0015] Optionally, the large kernel convolution sequence is used to expand the receptive field, and the expression is:

[0016] ;

[0017] Among them, represents the dilation rate of the i-th depth convolution, represents the kernel size of the i-th depth convolution, represents the receptive field of the i-th depth convolution, ;

[0018] The constraint conditions satisfied are:

[0019] ;

[0020] Among them, .

[0021] Optionally, the spatial kernel selection mechanism performs the following operations:

[0022] 1) Multi-scale feature generation, and the expression is:

[0023] ;

[0024] Among them, represents the i-th feature map, represents the i-th depth convolution layer;

[0025] Each goes through 1×1 convolution for channel mixing to obtain multi-scale spatial features ;

[0026] 2) Concatenate features of different scales, with the expression:

[0027] ;

[0028] 3) Perform average pooling and max pooling on the concatenated features in the channel dimension, with the expression:

[0029] ;

[0030] Among them, represents the spatial feature descriptor of average pooling, represents the spatial feature descriptor of max pooling;

[0031] 3) Concatenate and and generate N spatial attention maps through a convolutional layer , with the expression:

[0032] ;

[0033] Apply the sigmoid activation function to each spatial attention map to obtain N spatial masks, with the expression:

[0034] ;

[0035] Each spatial mask corresponds to a feature map of the decomposed kernel;

[0036] 4) Multiply the spatial features of each scale element-wise with the corresponding spatial mask , and after weighted fusion, obtain the attention feature S, with the expression:

[0037] ;

[0038] Multiply the input feature X element-wise with the attention feature S to finally obtain the output feature.

[0039] Optionally, the expression of the EIoU is:

[0040] ;

[0041] ;

[0042] ;

[0043] ;

[0044] Among them, represents the total loss, represents the overlap loss, Indicates the center point distance loss, Indicates the aspect ratio consistency loss, Represents the square of the Euclidean distance between the center point of the predicted bounding box and the center point of the target bounding box, Represents the square of the difference between the width of the predicted bounding box and the width of the target bounding box, Represents the square of the difference between the height of the predicted bounding box and the height of the target bounding box, and Represent the width and height of the smallest closed bounding box that can cover the predicted bounding box and the target bounding box, respectively.

[0045] Optionally, in step S10, the labelme tool is used to annotate the landslide remote sensing image.

[0046] Optionally, in step S20, the ratio of the training set, the validation set, and the test set is 8:1:1.

[0047] The technical solution provided by the present invention has the following advantages compared with the prior art:

[0048] 1) The method for detecting mine area remote sensing landslides based on the improved YOLOv10 model provided by the present invention replaces the SPPF module in the backbone network of the traditional YOLOv10 model with the AISFI module, which can more effectively capture the dependence relationship between different landslide features, enabling the model to more effectively process and fuse important feature information, while taking into account landslides of different shapes and sizes, thereby improving the overall detection performance and reducing the difficulty of extracting landslide features;

[0049] 2) The method for detecting mine area remote sensing landslides based on the improved YOLOv10 model provided by the present invention introduces the LSKNet module into the neck network of the traditional YOLOv10 model, which can dynamically adjust the receptive field and use a large-scale depth and spatial selection mechanism, effectively processing the extensive context information and multi-scale landslide targets in the landslide remote sensing image, enhancing the landslide positioning ability, and thus improving the detection accuracy;

[0050] 3) The method for detecting mine area remote sensing landslides based on the improved YOLOv10 model provided by the present invention, on the one hand, replaces the C2f module of the traditional YOLOv10 model with the C2f_DCNv4 module, which can enhance the dynamic feature and expression ability of the model, reduce redundant operations, and optimize memory access; on the other hand, uses EIoU instead of CIoU of the traditional YOLOv10 model, which can improve the calculation speed and optimize the positioning accuracy; the two aspects cooperate to greatly improve the calculation speed of the model, thereby improving the detection speed, and further enabling timely early warning when a landslide occurs. Description of the Drawings

[0051] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments in line with the present invention, and are used together with the specification to explain the principles of the present invention.

[0052] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the accompanying drawings required for use in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0053] Figure 1 It represents the flowchart of the remote sensing landslide detection method in the mining area in the embodiment of the present invention;

[0054] Figure 2 It represents the network structure diagram of the improved YOLOv10 model in the embodiment of the present invention;

[0055] Figure 3 It represents the network structure diagram of the LSKNet module in the embodiment of the present invention;

[0056] Figure 4 It represents the schematic diagram of the spatial kernel selection mechanism in the embodiment of the present invention;

[0057] Figure 5 It represents the network structure diagram of the AISFI module in the embodiment of the present invention;

[0058] Figure 6 It represents the comparison diagram between the DCNv4 module and the DCNv3 module in the embodiment of the present invention;

[0059] Figure 7 It represents the flowchart of the dataset collection in the embodiment of the present invention. Detailed implementation manners

[0060] In order to be able to more clearly understand the above-mentioned objects, features, and advantages of the present invention, the solutions of the present invention will be further described below. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.

[0061] Many specific details are set forth in the following description in order to fully understand the present invention, but the present invention can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0062] The following combines Figures 1 to 7 to elaborate in detail on the specific embodiments of the present invention.

[0063] This embodiment provides a remote sensing landslide detection method in the mining area based on an improved YOLOv10 model, including steps S10 to S50.

[0064] S10. Collect remote sensing images of landslides, annotate the remote sensing images of landslides, and construct a dataset of multi-temporal landslides.

[0065] Specifically, use a DJI drone to collect remote sensing images of landslides, and use the labelme tool to annotate the remote sensing images of landslides.

[0066] S20. Divide the dataset into a training set, a validation set, and a test set, and introduce negative samples to improve the dataset.

[0067] Specifically, the ratio of the training set, the validation set, and the test set is 8:1:1.

[0068] S30. Construct a YOLOv10 model, replace the C2f module in the backbone network and the neck network with the C2f_DCNv4 module, replace the SPPF module in the backbone network with the AISFI module, introduce the LSKNet module at the position where each dual-label assignment prediction head corresponds to the head network in the neck network, form an improved YOLOv10 model, and use EIoU as the loss function of the improved YOLOv10 model.

[0069] Specifically, the C2f_DCNv4 module is obtained by replacing the Bottleneck of the C2f module with DCNv4. The C2f_DCNv4 module streamlines the operation design, removes unnecessary calculation steps, simplifies the operation process, and improves efficiency. By multi-channel processing, reducing redundant calculations, vectorized load / storage, using half-precision data types, and removing the softmax normalization operation in spatial aggregation, the C2f_DCNv4 module is made more efficient.

[0070] More specifically, DCNv4 is an efficient dynamic sparse operator that reduces redundant operations and improves processing speed by optimizing memory access. At the same time, it reduces the memory access cost by reducing unnecessary memory access requests, thereby accelerating the execution speed of the operation.

[0071] Details of the memory access optimization strategy include the following aspects:

[0072] 1) Multi-channel processing: DCNv4 uses one thread to process multiple channels instead of each thread processing a single channel, which can reduce the memory access requests for loading sampling offsets and aggregating weight values, thus saving memory access costs;

[0073] 2) Reducing redundant calculations: By reusing bilinear interpolation coefficients and other methods, some redundant calculations are reduced, saving time. Although the time saved in a single operation may not be much, it can accumulate into a significant efficiency improvement in large-scale operations;

[0074] 3) Vectorized loading / storing: By adopting vectorized loading / storing operations, the workload of each thread can be reduced, thus accelerating the execution speed of the GPU kernel. By optimizing the data loading and storing methods, the execution efficiency of the operations is improved;

[0075] 4) Using half-precision data types: DCNv4 uses half-precision data types, reducing the number of bytes that the kernel needs to read and write, improving the data throughput. This can further increase the data transmission speed and improve the execution efficiency of the operations.

[0076] In this embodiment, by designing the C2f_DCNv4 module, the detection speed of the model can be improved as a whole.

[0077] Specifically, the AISFI module includes AIFI and Conv located before AIFI. AIFI is used to automatically learn the correlation between features, and Conv is used to extract local features.

[0078] It should be noted that the AISFI module improves the efficiency and effectiveness of feature extraction by introducing internal scale feature interaction based on the attention mechanism. Its core is to use the attention mechanism between features of the same scale to enhance the focusing ability of the network and promote richer feature fusion. Compared with traditional feature fusion which involves features of different scales, AISFI focuses on feature fusion within the same scale, which helps to capture finer-grained feature information. Since the size and shape of landslides vary greatly, the detection of large and small targets needs to be considered. Because AIFI can automatically learn the correlation between features, and Conv can extract local features, replacing the SPPF module with the AISFI module and using AIFI and Conv can more effectively capture the dependence relationship between different landslide features, enabling the model to more effectively process and fuse important feature information, while taking into account landslides of different shapes and sizes, thus improving the overall detection performance.

[0079] More specifically, when AIFI performs feature fusion, it first converts the two-dimensional S5 features into high-dimensional vectors, which are processed by AIFI. The mathematical process is that the output of the multi-head self-attention (MHSA) is added to the input through a residual connection. After layer normalization, it efficiently captures the complex relationships in the input sequence. Subsequently, this feature information is passed to the feed forward network (FFN) for non-linear transformation and feature extraction. Then, after layer normalization processing, the attention scores containing important feature information are output. Finally, the result is converted back to the two-dimensional form, denoted as F5, to complete the subsequent cross-scale feature fusion.

[0080] Specifically, the core component of the LSKNet module is the LSK block, which realizes efficient multi-scale feature extraction and fusion through a large kernel convolution sequence and a spatial kernel selection mechanism.

[0081] It is easy to understand that a larger kernel convolution is constructed by explicitly decomposing it into a depth convolution sequence with a significantly growing kernel and introducing dilation.

[0082] More specifically, the large kernel convolution sequence is used to expand the receptive field, and the expression is:

[0083] ;

[0084] where represents the dilation rate of the i-th depth convolution, represents the kernel size of the i-th depth convolution, represents the receptive field of the i-th depth convolution, ;

[0085] The constraint conditions satisfied are:

[0086] ;

[0087] where .

[0088] It should be noted that the increase in the size and dilation rate of the kernel ensures that the receptive field expands fast enough, while also setting an upper limit on the dilation rate to ensure that the dilated convolution does not introduce gaps between the feature maps. Such a design has two advantages: First, it explicitly generates multiple features with different large receptive fields, which makes it easier to select subsequent convolution kernels; second, sequential decomposition is more efficient than simply applying a single larger convolution kernel.

[0089] More specifically, the spatial kernel selection mechanism operates as follows:

[0090] 1) Multi-scale feature generation, and the expression is:

[0091] ;

[0092] where represents the i-th feature map, represents the i-th depth convolutional layer;

[0093] Each is passed through a 1×1 convolution for channel mixing to obtain multi-scale spatial features ;

[0094] This step uses a series of decomposed depthwise separable convolutions with different receptive fields, which can obtain rich context information features from inputs in different ranges;

[0095] 2) Concatenate the features of different scales, and the expression is:

[0096] ;

[0097] 3) Perform average pooling and max pooling on the concatenated features in the channel dimension, and the expression is:

[0098] ;

[0099] Among them, represents the spatial feature descriptor of average pooling, represents the spatial feature descriptor of max pooling;

[0100] This step is used to effectively extract the space;

[0101] 3) Concatenate and , and generate N spatial attention maps through a convolutional layer , and the expression is:

[0102] ;

[0103] Apply the sigmoid activation function to each spatial attention map , and obtain N spatial masks, and the expression is:

[0104] ;

[0105] Each spatial mask corresponds to a feature map of a decomposed kernel;

[0106] 4) Multiply the spatial features of each scale with the corresponding spatial mask element by element, and after weighted fusion, obtain the attention feature S, and the expression is:

[0107] ;

[0108] Multiply the input feature X and the attention feature S element by element to finally obtain the output feature.

[0109] It should be noted that the LSKNet module is a repeatable stacking module, including several LSK blocks. Each LSK block includes two residual sub-blocks, namely the large kernel selection sub-block (LKSelection) and the feed-forward network sub-block (FFN): The large kernel selection sub-block is used to dynamically adjust the receptive field of the network, which includes 2 fully connected layers, a GELU activation function, and an LSK module; The feed-forward network sub-block is used for channel mixing and feature refinement, which includes 2 fully connected layers, a depth convolution, and a GELU activation function.

[0110] In this embodiment, by introducing LSKNet into the neck network of the YOLOv10 model, it can effectively improve the attention to the context area, which is a supplement to the defects of the channel and spatial attention mechanisms; LSKNet introduces multiple convolutional kernels and aggregates feature information along the channel dimension, making it flexible and accurate to be perceived; The LSKNet module meets the requirements of landslide detection under complex backgrounds, can efficiently focus on the landslide-related spatial areas, capture richer information, improve the positioning ability by using spatial attention, and ultimately improve the success rate of landslide detection.

[0111] Specifically, the expression of EIoU is:

[0112] ;

[0113] ;

[0114] ;

[0115] ;

[0116] Among them, represents the total loss, represents the overlap loss, represents the center point distance loss, represents the aspect ratio consistency loss, represents the square of the Euclidean distance between the center point of the predicted bounding box and the center point of the target bounding box, represents the square of the difference between the width of the predicted bounding box and the width of the target bounding box, represents the square of the difference between the height of the predicted bounding box and the height of the target bounding box, and represent the width and height of the smallest closed bounding box that can cover the predicted bounding box and the target bounding box, respectively.

[0117] It should be noted that the penalty term of EIoU is based on the penalty term of CIoU, where the influence factor of the aspect ratio is split to calculate the length and width of the target box and the predicted box separately. This loss function consists of three parts: overlap loss, center point distance loss, and aspect ratio consistency loss. The first two parts follow the methods in CIoU, but the aspect ratio consistency loss directly minimizes the difference between the width and height of the target box and the predicted box, resulting in a faster convergence speed. Additionally, existing IoU-based losses, such as CIoU and GIoU, cannot effectively measure the difference between the target box and the anchor, leading to a slow convergence speed and inaccurate localization in the optimization of the BBR (bounding box regression) model. EIoU can comprehensively consider the matching of position and size by adding Focal to focus on high-quality anchor boxes, thus significantly improving the localization accuracy of the target detection model, especially when the size and shape of the target object vary greatly.

[0118] S40. Use the dataset to train the improved YOLOv10 model to obtain the target model.

[0119] S50. Take the landslide remote sensing image as the input and use the target model for remote sensing landslide detection in the mining area.

[0120] Next, several comparative experiments are conducted to elaborate and demonstrate in detail each improvement point and the overall effect of the remote sensing landslide detection method in this embodiment.

[0121] It should be noted that all the following experiments are carried out under the same hardware and software environment, using PyTorch as the deep learning framework and the Linux operating system, and YOLOv10s as the baseline model. The GPU used is three NVIDIA TITAN Xp, with a GPU memory size of 36G. All models are trained for 300 epochs, and the IoU and confidence are set to 0.45 during the testing and training processes. This dataset is divided into a training set, a validation set, and a test set in the ratio of 8:1:1. The dataset contains 2500 images. The training set, validation set, and test set contain 2000 images, 250 images, and 250 images respectively.

[0122] All of the following experiments evaluate the model using six metrics, namely: Precision - accuracy rate, which represents the proportion of the number of landslides accurately identified (TP) to the total number of landslides identified as such (the number of accurately identified landslides TP + the number of misidentified landslides FP); Recall - recall rate, which represents the proportion of the number of landslides accurately identified (TP) to the total number of actual landslides (the number of accurately identified landslides TP + the number of omitted landslides FN); mAP@0.5 represents the average precision at an IoU threshold of 0.5, used to measure the area under the P-R curve of the model, indicating the comprehensive detection performance of the model; mAP@0.5:0.95 represents the average precision when the IoU threshold varies from 0.5 to 0.95, used to more comprehensively evaluate the performance of the model at different IoU thresholds; GFLOP - floating point operation count, which is used to measure the computational complexity during the execution of the model. The smaller the value, the smaller the computational amount of the model; FPS - frames per second, which represents the number of images that the model can process in one second, used to quantify the operation speed of the model. The larger the value, the faster the model processing speed.

[0123] Experiment 1:

[0124] Content: Replace the SPPF module in the backbone network of the YOLOv10s model with the AISFI module.

[0125] Comparison object: The YOLOv10s model and the YOLOv10s model after introducing the AISFI module.

[0126] Purpose: To demonstrate the effect of the AISFI module.

[0127] Data statistics:

[0128]

[0129] Analysis: Based on the baseline model YOLOv10, replacing the SPPF module in the head network with the AISFI module enables the model to more effectively capture the dependencies between different landslide features, allowing the model to more effectively process and fuse important feature information, while taking into account landslides of different shapes and sizes, thereby improving the overall detection performance.

[0130] Conclusion: Experimental data shows that after introducing the AISFI module, the accuracy rate and recall rate of the model have been significantly improved. However, at the same time, the increase in the number of model parameters has led to a decrease in the detection speed.

[0131] Experiment 2:

[0132] Content: On the basis of Experiment 1, further introduce the LSKNet module after the C2f module in the neck network of the YOLOv10s model.

[0133] Comparison objects: YOLOv10s model, YOLOv10s model with the AISFI module introduced, and YOLOv10s model with the AISFI module and LSKNet module introduced.

[0134] Purpose: To demonstrate the effect of the LSKNet module.

[0135] Data statistics:

[0136]

[0137] Analysis: The LSKNet module meets the requirements of landslide detection in complex backgrounds, can efficiently focus on landslide-related spatial regions, capture richer information, improve the positioning ability using spatial attention, and ultimately increase the success rate of landslide detection. The LSKNet module utilizes the characteristics of multiple depthwise separable convolutions of large convolutional kernels to generate features with a wide receptive field, thereby reducing the number of model parameters and dynamically adjusting the receptive field to adapt to the complex and changeable landslide background environment.

[0138] Conclusion: Experimental data shows that after introducing the LSKNet module, the accuracy and recall rate increase, the number of parameters decreases, and the detection speed improves.

[0139] Experiment 3:

[0140] Content: On the basis of Experiment 2, further replace the C2f module in the backbone network and neck network of the YOLOv10s model with the designed C2f_DCNv4 module.

[0141] Comparison objects: YOLOv10s model, YOLOv10s model with the AISFI module introduced, YOLOv10s model with the AISFI module and LSKNet module introduced, and YOLOv10s model with the AISFI module, LSKNet module, and C2f_DCNv4 module introduced.

[0142] Purpose: To demonstrate the effect of the C2f_DCNv4 module.

[0143] Data statistics:

[0144]

[0145] Analysis: The C2f_DCNv4 module achieves a faster convergence speed and three times the forward processing speed by removing softmax normalization and memory access optimization, thereby greatly improving the landslide detection speed without losing accuracy.

[0146] Conclusion: The experimental data shows that after introducing the LSKNet module, on the basis of the basically unchanged detection accuracy and recall rate, the number of model parameters is significantly reduced, and the detection speed of landslides is significantly improved.

[0147] Experiment Four:

[0148] Content: On the basis of Experiment Three, further replace the CIoU loss function of the YOLOv10s model with EIoU.

[0149] Comparison objects: YOLOv10s model, YOLOv10s model with the AISFI module introduced, YOLOv10s model with the AISFI module and LSKNet module introduced, YOLOv10s model with the AISFI module, LSKNet module and C2f_DCNv4 module introduced, and YOLOv10s model with the loss function of EIoU (i.e., the improved YOLOv10s model).

[0150] Purpose: To demonstrate the effect of EIoU.

[0151] Data statistics:

[0152]

[0153] Analysis: EIoU has a more stable training process. By introducing penalty terms for the center point distance and aspect ratio, even when IoU is 0, EIoU can provide effective gradient information to ensure that the model can continue to learn. And it can improve the localization accuracy. By comprehensively considering the matching of position and size, EIoU can significantly improve the localization accuracy of the target detection model, especially when the size and shape of the target object change greatly, making the model have a faster convergence speed and higher accuracy.

[0154] Conclusion: The experimental data shows that by replacing the loss function with EIOU, all indicators have been improved.

[0155] Experiment Five:

[0156] To further verify the superiority and effectiveness of the improved algorithm, the present invention conducts a comparative experiment with existing mainstream models. In the experiment, the proposed model is compared with lightweight YOLOv5, YOLOv7, YOLOv10, and YOLOv9, which have high accuracy in landslide disaster identification. The data statistics are as follows:

[0157]

[0158] Experimental data shows that compared with other models, YOLOv5 has a smaller number of parameters, but its detection effect on small targets and dense tasks is poor; YOLOv7 adopts a special design that converts the object detection task into a single convolutional operation, which can effectively detect various objects; compared with other models, YOLOv7 performs better in the detection of small targets; YOLOv10 achieves a balance among model size, computational efficiency, and performance; YOLOv9 introduces more technical improvements, such as feature pyramid networks and attention mechanisms, etc., which improves the accuracy of object detection, and at the same time, YOLOv9 has a better detection effect on dense targets. The present invention makes overall detection improvements based on YOLOv10. The accuracy and recall rate of the present invention are significantly higher than those of other models, and the number of parameters GFLOPs is only slightly more than that of YOLOv5, YOLOv10, and YOLOv9. The detection speed and accuracy are both higher than these baseline models. In summary, the model of the present invention has a more comprehensive improvement in detection speed and accuracy compared with these mainstream models.

[0159] The above are only specific implementation manners of the present invention, enabling those skilled in the art to understand or implement the present invention. Although the foregoing embodiments have been described in detail, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments, and they should all be covered by the protection scope of the claims.

Claims

1. A method for detecting remote sensing landslides in mining areas based on an improved YOLOv10 model, characterized in that, It includes the following steps: S10. Collect remote sensing images of landslides, annotate the remote sensing images of landslides, and construct a dataset of multi-temporal landslides; S20. Divide the dataset into a training set, a validation set, and a test set, and introduce negative samples to improve the dataset; S30. Construct a YOLOv10 model, replace the C2f module in the backbone network and the neck network with the C2f_DCNv4 module, replace the SPPF module in the backbone network with the AISFI module, introduce the LSKNet module at the position corresponding to each dual-label assignment prediction head in the head network in the neck network to form an improved YOLOv10 model, and at the same time use EIoU as the loss function of the improved YOLOv10 model; S40. Use the dataset to train the improved YOLOv10 model to obtain a target model; S50. Take the remote sensing image of the landslide as the input, and use the target model to detect remote sensing landslides in the mining area.

2. The method for remote sensing landslide detection in mining areas based on the improved YOLOv10 model according to claim 1, wherein, The C2f_DCNv4 module is obtained by replacing the Bottleneck of the C2f module with DCNv4.

3. The method for detecting remote sensing landslides in mining areas based on the improved YOLOv10 model according to claim 1, wherein, The AISFI module includes AIFI and Conv located before the AIFI. The AIFI is used to automatically learn the correlation between features, and the Conv is used to extract local features.

4. The method for detecting remote sensing landslides in mining areas based on the improved YOLOv10 model according to claim 1, wherein, The core component of the LSKNet module is the LSK block, and the LSK block realizes efficient multi-scale feature extraction and fusion through a large kernel convolution sequence and a spatial kernel selection mechanism.

5. The method for remotely sensing landslide detection in mining areas based on the improved YOLOv10 model according to claim 4, characterized in that, The large kernel convolution sequence is used to expand the receptive field, and the expression is: ; Among them, represents the dilation rate of the i-th depth convolution, represents the kernel size of the i-th depth convolution, represents the receptive field of the i-th depth convolution, ; The satisfied constraint conditions are: ; Among them, 。 6. The mine area remote sensing landslide detection method based on the improved YOLOv10 model according to claim 5, wherein, The spatial kernel selection mechanism performs the following operations: 1) Multi-scale feature generation, and the expression is: ; Among them, represents the i-th feature map, represents the i-th depth convolution layer; Each undergoes channel mixing through 1×1 convolution to obtain multi-scale spatial features ; 2) Concatenate features of different scales, and the expression is: ; 3) Perform average pooling and max pooling on the concatenated features in the channel dimension, and the expression is: ; Among them, represents the spatial feature descriptor of average pooling, represents the spatial feature descriptor of max pooling; 3) Concatenate and , and generate N spatial attention maps through a convolutional layer with the expression: ; For each spatial attention map Apply the sigmoid activation function to obtain N spatial masks, expressed as: ; Each spatial mask corresponds to a feature map of a decomposed kernel; 4) Multiply the spatial features at each scale with the corresponding spatial mask element by element, and after weighted fusion, obtain the attention feature S, with the expression: ; Perform an element-wise product of the input feature X and the attention feature S, and finally obtain the output feature.

7. The method for remotely sensed landslide detection in mining areas based on the improved YOLOv10 model according to claim 1, wherein, The expression of the EIoU is: ; ; ; ; Among them, represents the total loss, represents the overlap loss, represents the center point distance loss, represents the aspect ratio consistency loss, represents the square of the Euclidean distance between the center point of the predicted bounding box and the center point of the target bounding box, represents the square of the difference between the width of the predicted bounding box and the width of the target bounding box, represents the square of the difference between the height of the predicted bounding box and the height of the target bounding box, and respectively represent the width and height of the smallest closed bounding box that can cover the predicted bounding box and the target bounding box.

8. The method for detecting remote sensing landslides in mining areas based on the improved YOLOv10 model according to any one of claims 1 to 7, characterized in that, In step S10, use the labelme tool to annotate the remote sensing images of landslides.

9. The method for detecting remote sensing landslides in mining areas based on the improved YOLOv10 model according to any one of claims 1 to 7, characterized in that, In step S20, the ratio of the training set, the validation set, and the test set is 8:1:1.