Landslide detection method based on improved YOLOv10n model
Through the improved YOLOv10n model, the Swin Transformer backbone network, the deep separable convolution module and the context aggregation attention mechanism are integrated, which solves the problems of insufficient feature extraction capabilities and complex background interference in landslide detection, and achieves high-precision, real-time and robust landslide detection effects.
Patent Information
- Application Number
- CN202510010553.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
When existing landslide detection technology deals with complex backgrounds and large-scale remote sensing images, there are problems such as insufficient feature extraction capabilities and difficult to achieve real-time and accuracy balance.
The improved YOLOv10n model is adopted to integrate the Swin Transformer backbone network, deep separable convolution module and context aggregation attention mechanism, which improves the model's ability to extract multi-scale landslide features, reduces the computational complexity, and enhances the robustness to complex backgrounds.
It significantly improves the accuracy and robustness of landslide detection, and achieves efficient, accurate and real-time landslide detection, suitable for complex terrain and resource-constrained environments.
Smart Images

Figure CN119942326A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a landslide detection method based on an improved YOLOv10n model, and belongs to the technical field of computer vision and geological disaster monitoring. Background Art
[0002] Landslides are a common geological disaster that poses a serious threat to human society. Traditional landslide detection methods, such as geological surveys and manual visual interpretation, have a certain degree of accuracy, but their detection efficiency is low in large-scale, complex terrain or high-risk areas, and are easily affected by subjective factors. In recent years, with the development of remote sensing technology, landslide detection methods based on remote sensing images have attracted widespread attention due to their real-time and large-scale monitoring potential.
[0003] Current landslide detection technologies, including feature threshold-based, change detection, machine learning and deep learning methods, each have their own limitations. Feature threshold-based methods have poor adaptability to complex backgrounds, change detection methods have strong dependence on temporal data, machine learning methods such as SVM (support vector machine), KNN (K nearest neighbor algorithm) and random forest are sensitive to data noise, and although deep learning methods show great potential, their application in real-time and resource-constrained environments still needs to be optimized.
[0004] Although deep learning algorithms, especially two-stage and one-stage target detection methods, have made significant progress in the field of landslide detection, such as Faster R-CNN (Fast Regional Convolutional Neural Network), Mask R-CNN (Mask Regional Convolutional Neural Network) and YOLO series models, there is still room for improvement in computational efficiency, model complexity and real-time performance. In particular, the YOLO series models, although widely used in landslide detection due to their high efficiency and simple structure, still face challenges in multi-scale target distinction, complex background processing and the balance between real-time performance and accuracy.
[0005] With the evolution of technology, YOLOv10n, as an enhanced version of the YOLO series, has improved the accuracy and real-time performance of small target detection by optimizing feature fusion and convolution operations, but its computational complexity and ability to capture global features when processing large-scale remote sensing images still need to be strengthened. Given the frequency and destructiveness of landslide disasters, the development of efficient and accurate landslide detection methods is crucial for disaster prevention and emergency response.
[0006] Therefore, there is an urgent need for a landslide detection method that can overcome the above technical limitations and has high precision, real-time and robustness to meet the needs of landslide monitoring under complex terrain and background. Based on this background, the present invention proposes a landslide detection method based on an improved YOLOv10n model, which aims to achieve efficient, accurate and real-time landslide detection by integrating the Swin Transformer backbone network, the deep separable convolution module and the contextual aggregation attention mechanism. Summary of the invention
[0007] The purpose of the present invention is to provide a landslide detection method based on an improved YOLOv10n model, aiming to solve the technical problems of insufficient feature extraction capability, complex background interference, and balance between real-time performance and accuracy in landslide detection.
[0008] In order to achieve the above objectives, the present invention adopts the following technical solution: a landslide detection method based on an improved YOLOv10n model, by integrating a Swin Transformer backbone network, a deep separable convolution module, and a contextual aggregation attention mechanism. The specific steps are as follows:
[0009] Step 1: Collect remote sensing image data;
[0010] Step 2: preprocessing the remote sensing image data, wherein the preprocessing includes adaptive scaling, normalization and data enhancement of the remote sensing image data, and dividing the preprocessed remote sensing image data into a training set, a validation set and a test set;
[0011] Step 3: Build an improved YOLOv10n model based on the Swin Transformer backbone network, the depth-wise separable convolutional module, and the context-aggregated attention mechanism, and train it using the training set;
[0012] Step 4: Input the validation set and test set into the trained improved YOLOv10n model to perform landslide detection and output the final target box and category label.
[0013] The step 2 is specifically as follows:
[0014] Adaptive scaling algorithm is used to adjust the remote sensing image data to a resolution consistent with the model input size, and zero padding is used to keep the image scale unchanged;
[0015] The distribution differences between different remote sensing image data are processed through normalization operations;
[0016] Data augmentation methods are applied to generate diverse training samples through random cropping, scaling, and splicing operations to improve the generalization ability of the model.
[0017] The step 3 is specifically as follows:
[0018] The improved YOLOv10n model backbone network is built based on Swin Transformer and consists of multiple CSP2 modules. The backbone network uses the window self-attention mechanism and the cross-window attention mechanism to extract the local features and global features of the landslide area respectively; thereby efficiently extracting features of the landslide area in the remote sensing image;
[0019] The backbone network gradually reduces the feature map resolution through the spatial direction downsampling module, and abstracts and compresses features of different scales. The output of the backbone network is aggregated through the SPPF module for global context features.
[0020] The SPPF module aggregates multi-scale features through spatial pooling operations with different kernel sizes to enhance the network's global perception of the landslide area; and sends the aggregated multi-scale features to the neck network; the CSP2 module in the neck network is used to reduce computational redundancy in the feature extraction process while maintaining feature consistency to ensure efficient transmission of landslide area features;
[0021] The neck network combines an upsampling module and a downsampling module, and utilizes a feature pyramid network structure to achieve fusion of shallow detail features and deep semantic features;
[0022] The feature map after fusion by the neck network is input into the detection head network for target detection. The detection head network adopts a decoupled task head, and optimizes the classification and positioning performance respectively through the classification task branch and the bounding box regression task branch. The classification task branch extracts category information through the convolution layer for category prediction of the landslide area. The regression task branch optimizes the positioning accuracy of the bounding box through the void convolution and attention mechanism. The detection head network is combined with the context enhancement module to extract the context information of the landslide area through void convolution operations of different scales, and dynamically adjusts the attention weight of the target area, thereby further improving the detection accuracy and robustness of the landslide target.
[0023] The improved YOLOv10n model backbone network is built based on Swin Transformer and consists of multiple CSP2 modules. The backbone network uses the window self-attention mechanism and the cross-window attention mechanism to extract the local features and global features of the landslide area respectively, including:
[0024] The input remote sensing image data is divided into windows of fixed size, and the local features are calculated within the window using the window self-attention mechanism; global feature associations are established between windows through the cross-window attention mechanism to enhance the feature expression ability of the landslide area; finally, the multi-scale feature map output by the backbone network is globally pooled through the SPPF module, and effective aggregation of contextual information is achieved through a variety of pooling kernels.
[0025] The aggregation includes: performing multi-scale pooling operations on the input feature map using pooling kernels of different sizes; performing step-by-step feature fusion on the pooling results to obtain context information expression; and reducing the number of channels from 1024 to 512 through feature channel compression operations, thereby reducing computational complexity and improving global feature expression capabilities.
[0026] The neck network combines the upsampling module and the downsampling module, and uses the feature pyramid network structure to achieve the fusion of shallow detail features and deep semantic features, including:
[0027] The neck network adopts a design combining feature pyramid structure and CSP2 module, and enhances the detection capability of landslide targets through multi-scale feature fusion. Among them, the CSP2 module performs feature channel division and fusion mechanism, reduces redundant calculations in the feature extraction process, and improves the information flow between features of different scales. The upsampling module enlarges the low-resolution feature map through deconvolution operation to restore the spatial information of small landslide targets. The downsampling module compresses the spatial dimension of the high-resolution feature map through maximum pooling operation to extract the global semantic features of the landslide area.
[0028] The detection head is specifically:
[0029] The detection head network adopts a decoupled task head design to improve the independence of classification and positioning. The classification task branch predicts the category of the landslide area through multi-layer convolution operations; the regression task branch optimizes the positioning accuracy of the landslide bounding box through hole convolution operations, and the context enhancement module combines 3×3, 5×5 and 7×7 hole convolution kernels to dynamically adjust the attention allocation of the landslide target through multi-scale feature extraction, thereby improving detection accuracy and reducing interference from complex backgrounds;
[0030] The classification and regression paths of the decoupled task head are combined to optimize the target category prediction and bounding box positioning performance respectively; the local and global context information of the landslide area is captured through the context enhancement module; the mixed precision training method is introduced in the training phase to improve the computational efficiency and memory utilization while ensuring that the model performance is not lost. The context aggregation attention mechanism calculates the attention weight of each position in the feature map through the following formula, and dynamically adjusts the attention allocation to the landslide area:
[0031]
[0032] Among them, α i is the attention weight, e i is the attention score of position i and N is the number of positions in the feature map.
[0033] The step 4 is specifically as follows:
[0034] The remote sensing image data is input into the trained improved YOLOv10n model. After the forward propagation of the model, the predicted position, category probability and target existence probability of each candidate box are obtained. The confidence score of each candidate box is calculated based on the classification probability and target existence probability.
[0035] A multi-level non-maximum suppression algorithm is used to optimize the target box screening process. By setting multiple different IoU thresholds, the problem of missed detection in small target landslide areas is reduced. Redundant candidate boxes are gradually removed, and the most representative and reliable prediction boxes are finally retained. The category label and confidence score of the prediction box are output as the final detection result output.
[0036] The IoU threshold is specifically:
[0037] The overlap between the predicted box and the real box is evaluated by the intersection over union (IoU). The calculation formula of IoU is:
[0038]
[0039] Among them, A and B represent the predicted box and the real box respectively. | A∩B | is the area of the intersection of the two boxes, | A∪B is the area of the union of the two boxes. The larger the IoU value, the higher the overlap between the predicted box and the real box, and the higher the accuracy of the detection result. When the IoU is greater than the set threshold, the predicted box is judged to be a valid detection result, thereby deciding whether to retain the predicted box as the final detection output.
[0040] The beneficial effects of the present invention are:
[0041] (1) The integration of the Swin Transformer backbone network significantly enhances the model’s ability to extract multi-scale landslide features and improves detection accuracy and robustness;
[0042] (2) The use of the depthwise separable convolution module reduces computational complexity, improves model real-time performance, and is suitable for resource-constrained environments;
[0043] (3) The implementation of the contextual aggregation attention mechanism effectively suppresses background interference in complex scenes and improves the detection accuracy of landslide areas;
[0044] In summary, the method of the present invention achieves a balance between high precision, real-time and robustness in landslide detection tasks, provides efficient and reliable technical support for landslide disaster monitoring, and has important scientific research and application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a technical flow chart of the present invention;
[0046] Figure 2 is a network architecture diagram of the present invention;
[0047] Figure 3 is a data enhancement graph of the present invention;
[0048] Figure 4 It is a landslide detection effect diagram of the present invention. DETAILED DESCRIPTION
[0049] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods.
[0050] In the remote sensing application of landslide detection, traditional detection methods face two major challenges: insufficient feature extraction capabilities and complex background interference. The high-resolution characteristics of remote sensing images provide rich detailed information for landslide detection, but also increase the difficulty of feature extraction, especially when dealing with multi-scale landslide targets. The model needs to have a strong feature expression ability to accurately distinguish and identify landslide areas of different sizes. In addition, complex terrain and background, such as trees, rocks, soil, etc., will generate a lot of visual noise, which interferes with the model's correct capture of landslide features, resulting in false positives and false negatives in the detection results.
[0051] In order to solve the above challenges and improve the accuracy and robustness of landslide detection, this paper proposes an improved YOLOv10 model. Based on the original YOLOv10, this model introduces a series of innovative improvements, including the SwinTransformer backbone network, the deep separable convolution module and the contextual aggregation attention mechanism.
[0052] Example 1: A landslide detection method based on an improved YOLOv10n model, the technical route is as follows Figure 1 As shown in the network architecture diagram Figure 2 As shown, the following steps are included:
[0053] Step 1: Collect remote sensing image data;
[0054] The data contains abundant landslide examples and is suitable for model training and validation.
[0055] Step 2: preprocessing the remote sensing image data, wherein the preprocessing includes adaptive scaling, normalization and data enhancement of the remote sensing image data, and dividing the preprocessed remote sensing image data into a training set, a validation set and a test set;
[0056] Specifically, in order to solve the problem of uneven distribution of positive and negative samples, mosaic, color jitter and scaling to a fixed ratio are used, such as Figure 3The figure shows the data enhancement after processing. The processed remote sensing image data is divided into training set, validation set and test set with a ratio of 8:1:1, and the annotation tool is used to annotate the real frame and the category of the remote sensing image;
[0057] Step 3: Build an improved YOLOv10n model based on the Swin Transformer backbone network, the depth-wise separable convolutional module, and the context-aggregated attention mechanism, and train it using the training set;
[0058] Specifically, an improved YOLOv10n model was constructed and trained, which included a Swin Transformer backbone network, a deep separable convolution module, and a contextual aggregation attention mechanism. The Swin Transformer backbone network enhanced the ability to extract multi-scale landslide features through the self-attention mechanism and window strategy, especially when dealing with small-scale landslide targets, effectively improving the comprehensiveness and adaptability of the model. The deep separable convolution module significantly reduced the computational complexity and improved the computational efficiency by decomposing the standard convolution into deep convolution and point-by-point convolution, which is particularly suitable for real-time monitoring of landslides in resource-constrained environments. The contextual aggregation attention mechanism enhanced the model's ability to capture global and local information of the landslide area through effective context information aggregation, effectively suppressed background interference in complex scenes, significantly improved the detection accuracy of the landslide area, and reduced false detections and missed detections.
[0059] Furthermore, the processed remote sensing image data is input into the improved YOLOv10n model, multi-scale features are extracted through the SwinTransformer backbone network, the depthwise separable convolution module is used to reduce the computational complexity, and the contextual aggregation attention mechanism is used to enhance the ability to capture global and local information of the landslide area, effectively suppressing background interference in complex scenes.
[0060] Specifically, during the model training process, the intersection-over-union (IoU) calculation formula is used to evaluate the overlap between the predicted box and the true box to optimize the prediction performance of the model. At the same time, Focal Loss is used as the loss function to solve the category imbalance problem and improve the small object detection performance.
[0061] Step 4: Input the validation set and test set into the trained improved YOLOv10n model to perform landslide detection and output the final target box and category label.
[0062] In the experiment, high-resolution remote sensing image data was used for verification. The data covers a variety of terrain and climate conditions, which puts high demands on the detection performance of the model. As shown in Table 1, the experimental results show that the improved YOLOv10n model achieves 91.8%, 72.7%, 81.2% and 81.1% in accuracy, recall, average precision (mAP@0.5) and F1 score, respectively, which are 3.7%, 20.8%, 21.2% and 15.6% higher than the original YOLOv10 model. The actual detection visualization effect is as follows Figure 4 As shown in the figure, the significant performance improvement verifies the detection advantage of the improved YOLOv10n under complex terrain conditions and provides efficient and reliable technical support for landslide disaster monitoring.
[0063] Table 1 Results comparison
[0064]
[0065] In summary, the improved YOLOv10n model effectively solves the challenges of feature extraction capability and complex background interference faced by landslide detection in remote sensing applications by integrating the Swin Transformer backbone network, the deep separable convolution module and the contextual aggregation attention mechanism, and achieves a balance between high precision, real-time and robustness. The method of the present invention achieves significant performance improvement in landslide detection tasks and has important practical application value for early warning and prevention of landslide disasters.
[0066] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.
Claims
1. A landslide detection method based on an improved YOLOv10n model, characterized in that: The method comprises the following steps: Step 1: Collect remote sensing image data; Step 2: preprocessing the remote sensing image data, wherein the preprocessing includes adaptive scaling, normalization and data enhancement of the remote sensing image data, and dividing the preprocessed remote sensing image data into a training set, a validation set and a test set; Step 3: Build an improved YOLOv10n model based on the Swin Transformer backbone network, the depth-wise separable convolutional module, and the context-aggregated attention mechanism, and train it using the training set; Step 4: Input the validation set and test set into the trained improved YOLOv10n model to perform landslide detection and output the final target box and category label.
2. The landslide detection method based on the improved YOLOv10n model according to claim 1 is characterized in that: The step 2 is specifically as follows: Adaptive scaling algorithm is used to adjust the remote sensing image data to a resolution consistent with the model input size, and zero padding is used to keep the image scale unchanged; The distribution differences between different remote sensing image data are processed through normalization operations; Data augmentation methods are applied to generate diverse training samples through random cropping, scaling, and splicing operations.
3. The landslide detection method based on the improved YOLOv10n model according to claim 1 is characterized in that: The step 3 is specifically as follows: The improved YOLOv10n model backbone network is built based on Swin Transformer and consists of multiple CSP2 modules. The backbone network uses the window self-attention mechanism and the cross-window attention mechanism to extract the local features and global features of the landslide area respectively. The backbone network gradually reduces the feature map resolution through the spatial direction downsampling module, and abstracts and compresses features of different scales. The output of the backbone network is aggregated through the SPPF module for global context features. The SPPF module aggregates multi-scale features through spatial pooling operations with different kernel sizes, and sends the aggregated multi-scale features to the neck network; The neck network combines an upsampling module and a downsampling module, and utilizes a feature pyramid network structure to achieve fusion of shallow detail features and deep semantic features; The feature map after fusion by the neck network is input into the detection head network for target detection. The detection head network adopts a decoupled task head, and optimizes the classification and positioning performance respectively through the classification task branch and the bounding box regression task branch. The classification task branch extracts category information through the convolution layer for category prediction of the landslide area. The regression task branch optimizes the positioning accuracy of the bounding box through the void convolution and attention mechanism. The detection head network is combined with the context enhancement module to extract the context information of the landslide area through void convolution operations of different scales, and dynamically adjusts the attention weight of the target area.
4. The landslide detection method based on the improved YOLOv10n model according to claim 3 is characterized in that: The improved YOLOv10n model backbone network is built based on Swin Transformer and consists of multiple CSP2 modules. The backbone network uses the window self-attention mechanism and the cross-window attention mechanism to extract the local features and global features of the landslide area respectively, including: The input remote sensing image data is divided into windows of fixed size, and the local features are calculated within the window using the window self-attention mechanism; global feature associations are established between windows through the cross-window attention mechanism; finally, the multi-scale feature maps output by the backbone network are globally pooled through the SPPF module, and the aggregation of contextual information is achieved through a variety of pooling kernels; The aggregation includes: performing a multi-scale pooling operation on the input feature map using pooling kernels of different sizes; performing step-by-step feature fusion on the pooling results to obtain context information expression; and reducing the number of channels from 1024 to 512 through a feature channel compression operation.
5. The landslide detection method based on the improved YOLOv10n model according to claim 3 is characterized in that: The neck network combines the upsampling module and the downsampling module, and uses the feature pyramid network structure to achieve the fusion of shallow detail features and deep semantic features, including: The neck network adopts a design that combines a feature pyramid structure and a CSP2 module, wherein the CSP2 module performs feature channel division and fusion mechanism, the upsampling module enlarges the low-resolution feature map through a deconvolution operation to restore the spatial information of small landslide targets, and the downsampling module compresses the spatial dimension of the high-resolution feature map through a maximum pooling operation to extract the global semantic features of the landslide area.
6. The landslide detection method based on the improved YOLOv10n model according to claim 3 is characterized in that: The detection head is specifically: The detection head network adopts a decoupled task head design. The classification task branch predicts the category of the landslide area through multi-layer convolution operations. The regression task branch optimizes the positioning accuracy of the landslide bounding box through hole convolution operations. The context enhancement module combines 3×3, 5×5 and 7×7 hole convolution kernels to dynamically adjust the attention allocation of the landslide target through multi-scale feature extraction. The classification and regression paths of the decoupled task head are combined to optimize the target category prediction and bounding box localization performance respectively; the local and global context information of the landslide area is captured through the context enhancement module; The mixed precision training method is introduced in the training stage. The contextual aggregation attention mechanism calculates the attention weight of each position in the feature map through the following formula, and dynamically adjusts the attention allocation to the landslide area: Among them, α i is the attention weight, e i is the attention score of position i and N is the number of positions in the feature map.
7. The landslide detection method based on the improved YOLOv10n model according to claim 1 is characterized in that: The step 4 is specifically as follows: The remote sensing image data is input into the trained improved YOLOv10n model. After the forward propagation of the model, the predicted position, category probability and target existence probability of each candidate box are obtained. The confidence score of each candidate box is calculated based on the classification probability and target existence probability. A multi-level non-maximum suppression algorithm is used to gradually remove redundant candidate boxes by setting multiple different IoU thresholds, and finally retain the most representative and reliable prediction boxes. The category label and confidence score of the prediction box are output as the final detection result output.
8. The landslide detection method based on the improved YOLOv10n model according to claim 7 is characterized in that: The IoU threshold is specifically: The overlap between the predicted box and the real box is evaluated by the intersection over union (IoU). The calculation formula of IoU is: Among them, A and B represent the predicted box and the real box respectively. | A∩B | is the area of the intersection of the two boxes, | A∪B is the area of the union of the two boxes. When the IoU is greater than the set threshold, the predicted box is determined to be a valid detection result.
Citation Information
Cited By
Landslide detection method and device based on partial convolution and exponential moving average
CN120451800A
Landslide detection method and device based on partial convolution and exponential moving average
CN120451800B
Slope high-risk stone identification method based on oblique photography technology
CN120543943A
Automatic detection method and system for apparent quality of remote sensing image
CN120894708A
A remote sensing image apparent quality automatic detection method and system
CN120894708B