Landslide area detection method and system based on multi-level and multi-scale feature enhancement

Through the combination of multi-level feature extraction and efficient multi-scale attention module of pyramid, the problems of scale differences and uncertain shapes in landslide area detection are solved, and landslide area identification and detection are achieved with higher accuracy.

CN119992349AActive Publication Date: 2025-05-13STATE GRID LOCATION BASED SERVICE CO LTD +2
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510474234.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect landslide areas of various sizes and shapes in landslide areas, resulting in the problems of missed inspection of small-scale landslides and incomplete inspection of large-scale areas.

Method used

A multi-level feature extraction model is used to extract the remote sensing image. After unifying the number of channels, the high-efficiency multi-scale attention module of the pyramid is input to enhance the feature, and upsampling is performed step by step to obtain landslide area detection results.

Benefits of technology

Through multi-level and multi-scale feature extraction and enhancement, the accuracy of landslide area identification is improved, the missed detection of small-scale landslides is avoided, and the false alarm rate is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992349A_ABST
    Figure CN119992349A_ABST
Patent Text Reader

Abstract

The invention relates to the field of landslide detection, in particular to a landslide area detection method and system based on multi-level and multi-scale feature enhancement, and the method comprises the steps: carrying out the feature extraction of a remote-sensing image landslide area through employing a multi-level feature extraction model, and obtaining a multi-level feature map; unifying the number of channels of the feature maps of the plurality of layers; respectively inputting the feature maps with unified channel numbers into pyramid efficient multi-scale attention modules with different scales to obtain feature-enhanced attention feature maps with multiple scales; performing step-by-step up-sampling processing on the attention feature maps of the multiple scales to obtain a final remote sensing image landslide area detection result; according to the method, the image containing the space structure, the texture feature and the abstract feature is adopted and enhanced, the landslide area can be better recognized and distinguished, missing detection of the small-scale landslide area is avoided, high-level features and low-level features are fused through step-by-step up-sampling, the effect of fully fusing feature maps is achieved, and the false alarm rate is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of landslide detection, and in particular to a landslide area detection method and system based on multi-level and multi-scale feature enhancement. Background Art

[0002] Landslides are major natural disasters that pose a severe challenge to human survival in all regions of the world. Global climate change, population growth, and urbanization have further increased the frequency and severity of landslides. In the field of emergency management, the importance of the on-site conditions of landslide disasters is self-evident, which can help emergency rescue personnel accurately find the locations where disasters are severe and carry out rescue missions in a timely manner. In addition, landslide area detection can also help experts conduct disaster assessments and formulate subsequent reconstruction plans. Therefore, extracting landslide disaster information from remote sensing images is an important direction in the field of remote sensing applications.

[0003] Landslide refers to an area with obvious signs of deformation and a high probability of disaster in the near future. There are two main common methods for identifying landslides. The first is the field survey method, which has high accuracy and strong controllability, but low efficiency and high cost. The second is remote sensing image processing and recognition methods, including statistical methods, shallow machine learning methods and deep learning methods, commonly used in interferometric synthetic aperture radar (InSAR), laser radar (LiDAR) and optical remote sensing image data. Among them, InSAR has the advantages of high resolution and all-weather data acquisition, but is easily restricted by factors such as observation angle, vegetation coverage, water vapor and data processing technology. Existing technology can use the coherent scatterer InSAR (CSI) method to successfully extract 17 unstable slopes in the basin. There are also technologies that use Stacking InSAR and deep learning technology to realize intelligent identification of geological disaster hazards at a wide scale. LiDAR can provide high-precision three-dimensional information, high-density and high-precision sampling and classification, but due to its relatively high cost, it is difficult to implement on a large scale. For example, using object-oriented methods to extract data from LiDAR Landslides covered by vegetation were identified in the data, and more than 70% of the landslide bodies were successfully identified; optical remote sensing images are also a common method for landslide identification. They have wide coverage and high resolution, and can effectively identify areas with obvious signs of deformation in a large area, but are easily restricted by foggy weather, vegetation coverage and time series. Shallow machine learning methods for optical remote sensing imaging landslide identification require pre-extraction of features and interpretable factors in the image, and then use classifiers to perform classification calculations, which increases the complexity of algorithm design, limits the scalability of the algorithm, and the algorithm does not have the performance required for real-time applications.

[0004] Deep learning methods are also widely used in optical remote sensing target detection, semantic segmentation and image classification. Convolutional neural network (CNN) is the main method of deep learning. This method has multi-layer nonlinear mapping capabilities. Multi-layer mapping can effectively extract texture, pop, edge, color and other information of remote sensing images, and can better fit or learn landslide features. In recent years, fully convolutional neural network (FCN) has been used in many remote sensing image recognition tasks. It can input images of any size and generate outputs of corresponding sizes through effective reasoning and learning. In the FCN architecture, the fully connected layer is replaced by a 1×1 convolutional layer, and a deconvolutional layer is used to restore the size of the feature map. Therefore, FCN has become the standard for high-precision pixel-level labeling. This method provides a more efficient and low-cost landslide extraction strategy in the field of remote sensing. However, it has high noise, so the details of the network output still need to be improved. In the further development of technology, U-net is proposed and has become one of the most popular semantic segmentation methods. U-net is an improved FCN model that accepts the entire image as input and outputs the category of pixel-level positioning. The model consists of two paths: the contraction path is used to capture contextual information, and the symmetric expansion path is used to obtain more accurate pixel-level positioning. In the prior art, DeepLabv3+ arbitrarily changes the resolution of the feature map output by the encoder by controlling the dilated convolution rate, increases the receptive field of the convolution layer, and expands DeepLabv3 by adding a simple and effective decoder module to restore the object boundary and enhance the semantic segmentation ability of the model. The attention mechanism, dilated convolution, etc. help the semantic segmentation network better capture contextual details. How to effectively extract and fuse multi-scale features is the key to improving the model detection capability. A multi-scale feature fusion mechanism based on the encoder-decoder network (MFFENet) is used to extract and fuse the multi-scale features of objects in remote sensing images. A novel and practical adaptive triangle fork (ATF) module is designed to effectively integrate useful features of different scales. The fusion of multi-scale features effectively improves the recognition ability of landslide areas. However, landslides often have no fixed shape in remote sensing images, and the scale of disasters is also different. The area of ​​landslides in remote sensing images can be large or small. The general change detection method cannot detect landslides of various scales well, which has caused the problem of missing small-scale landslides and incomplete detection of large-scale areas. Summary of the invention

[0005] In order to solve the problem of missed detection caused by the large size and shape of landslides in the prior art, the first aspect of the present invention proposes a landslide area detection method based on multi-level and multi-scale feature enhancement, comprising: A multi-level feature extraction model is used to extract features of landslide areas in remote sensing images to obtain feature maps at multiple levels. Unify the number of channels of feature maps at multiple levels; The feature maps with unified channel numbers are input into pyramid efficient multi-scale attention modules of different scales to obtain feature-enhanced attention feature maps of multiple scales; The attention feature maps of multiple scales are upsampled step by step to obtain the final remote sensing image landslide area detection result.

[0006] Preferably, the multi-level feature extraction model is used to extract features of the landslide area in the remote sensing image to obtain feature maps at multiple levels, specifically: Determine the number of output layers and output layer of the multi-level feature extraction model according to the detection accuracy; Inputting the landslide area of ​​the remote sensing image into the multi-level feature extraction model, and obtaining a feature map of the corresponding level at each output layer; The size of the image of each output layer is one quarter of the output layer of the previous layer.

[0007] Preferably, the multi-level feature extraction model is a MobileNetV2 model, and the MobileNetV2 model includes a plurality of bottleneck layers connected in sequence.

[0008] Preferably, the output layer has four layers, and the output layers are respectively the third, sixth, thirteenth and seventeenth layers of the bottleneck layer.

[0009] Preferably, the number of channels of feature maps at multiple levels is unified, specifically: The same convolutional layer is used to unify the number of channels of feature maps at multiple levels.

[0010] Preferably, the feature map with unified channel number is input into the pyramid efficient multi-scale attention module of different scales to obtain the attention feature map of multiple scales with feature enhancement, specifically: The feature maps of multiple levels after channel unification are input into pyramid efficient multi-scale attention modules of different scales respectively; Each pyramid efficient multi-scale attention module performs multi-scale segmentation, efficient multi-scale attention extraction, splicing and convolution on the input feature map in sequence to obtain the output feature map, and superimposes the input feature map with the output feature map to obtain the corresponding attention feature map.

[0011] Preferably, the step of performing step-by-step upsampling processing on the attention feature maps of multiple scales to obtain the final remote sensing image landslide area detection result includes: The upsampling module is used to perform step-by-step upsampling of the attention feature maps of multiple scales from high scale to obtain the final remote sensing image landslide area detection result; Wherein, the upsampling module includes multiple feature fusion units and multiple convolutional layers.

[0012] Preferably, the feature fusion unit comprises: an upsampling layer, Convolutional layer, batch normalization layer, and ReLU activation function.

[0013] A second aspect of the present invention provides a landslide area detection system based on multi-level and multi-scale feature enhancement, comprising: Multi-level feature extraction model, used to extract features from landslide areas in remote sensing images to obtain feature maps at multiple levels; Convolution module, used to unify the number of channels of feature maps at multiple levels; Pyramid efficient multi-scale attention module, used to obtain attention feature maps of multiple scales that capture changes in feature map information; The upsampling module is used to perform step-by-step upsampling of attention feature maps of multiple scales to obtain the final remote sensing image landslide area detection results.

[0014] Optionally, the loss function of the landslide area detection system based on multi-level and multi-scale feature enhancement is:

[0015] in, is the loss function of the landslide area detection system based on multi-level and multi-scale feature enhancement, is the binary cross entropy loss function, , · is the dot multiplication operation, c This is the final remote sensing image landslide area detection result. g is the true label, is the Dice loss function, , It is the first paradigm.

[0016] Optionally, the multi-level feature extraction model extracts features from the landslide area in the remote sensing image to obtain feature maps at multiple levels, specifically: Determine the number of output layers and output layer of the multi-level feature extraction model according to the detection accuracy; Inputting the landslide area of ​​the remote sensing image into the multi-level feature extraction model, and obtaining a feature map of the corresponding level at each output layer; The size of the image of each output layer is one quarter of the output layer of the previous layer.

[0017] Optionally, the multi-level feature extraction model is a MobileNetV2 model, and the MobileNetV2 model includes multiple bottleneck layers connected in sequence.

[0018] Optionally, the output layer of the multi-level feature extraction model has 4 layers, and the output layers are respectively the third layer, the sixth layer, the thirteenth layer and the seventeenth layer of the bottleneck layer.

[0019] Optionally, the convolution module unifies the number of channels of feature maps at multiple levels, specifically: The same convolutional layer is used to unify the number of channels of feature maps at multiple levels.

[0020] Optionally, the pyramid efficient multi-scale attention module obtains attention feature maps of multiple scales for image enhancement, specifically: The feature maps of multiple levels after channel unification are input into pyramid efficient multi-scale attention modules of different scales respectively; Each pyramid efficient multi-scale attention module performs multi-scale segmentation, efficient multi-scale attention extraction, splicing and convolution on the input feature map in sequence to obtain the output feature map, and superimposes the input feature map with the output feature map to obtain the corresponding attention feature map.

[0021] Optionally, the upsampling module performs step-by-step upsampling processing on the attention feature maps of multiple scales to obtain the final remote sensing image landslide area detection result, including: The upsampling module is used to perform step-by-step upsampling of the attention feature maps of multiple scales from high scale to obtain the final remote sensing image landslide area detection result; Wherein, the upsampling module includes multiple feature fusion units and multiple convolutional layers.

[0022] Optionally, the feature fusion unit in the upsampling module includes: an upsampling layer, Convolutional layer, batch normalization layer, and ReLU activation function.

[0023] Compared with the prior art, the present invention has the following beneficial effects: The invention provides a landslide area detection method and system based on multi-level and multi-scale feature enhancement. The method comprises the following steps: using a multi-level feature extraction model to extract features of a landslide area in a remote sensing image to obtain feature maps of multiple levels; unifying the number of channels of the feature maps of multiple levels; inputting the feature maps after the number of channels is unified into pyramid efficient multi-scale attention modules of different scales to obtain feature-enhanced attention feature maps of multiple scales; performing step-by-step upsampling processing on the attention feature maps of multiple scales to obtain a final remote sensing image landslide area detection result; the feature maps of multiple levels contain both spatial structure and texture features and abstract features, and further adopting multi-scale attention to enhance the expression ability of features, so as to better identify and distinguish landslide areas, avoid missing small-scale landslide areas, and step-by-step upsampling to gradually and effectively fuse high-level features with low-level features, so as to achieve the effect of fully fusing feature maps of each level and reduce the false alarm rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 This is a schematic diagram of the structure of the landslide area detection method based on multi-level and multi-scale feature enhancement proposed by the present invention; Figure 2 A schematic diagram of the steps of the landslide area detection method based on multi-level and multi-scale feature enhancement proposed by the present invention; Figure 3 This is a schematic diagram of the bottleneck layer structure proposed by the present invention; Figure 4 The present invention proposes Figure 2 Detailed step diagram of step S3; Figure 5 A schematic diagram of the structure of the pyramid efficient multi-scale attention module corresponding to the lowest layer feature map of the backbone network proposed in the present invention; Figure 6 This is a schematic diagram of the structure of the efficient multi-attention proposed by the present invention; Figure 7 The efficient multi-scale attention module of pyramids of different scales proposed by the present invention; Figure 8 This is a schematic diagram of the structure of the up-sampling module proposed in the present invention; Fig. 9 A schematic diagram showing a comparison of detection accuracy of the landslide area detection method based on multi-level and multi-scale feature enhancement proposed in the present invention; Fig.10 Schematic diagram of the structure of the landslide area detection system based on multi-level and multi-scale feature enhancement proposed for the invention. DETAILED DESCRIPTION

[0025] In order to solve the problem of missed detection caused by the large differences in scale and shape of the above-mentioned landslides, the present invention proposes a landslide area detection method and system based on multi-level and multi-scale feature enhancement, aiming to use the post-disaster drone impact data to achieve rapid extraction of landslide areas. This method can be called a multi-level and multi-scale attention neural network. Its core idea is to improve the accuracy of landslide area identification through multi-level and multi-scale feature extraction, enhancement and fusion. Figure 1 As shown in the figure, the multi-level multi-scale attention neural network uses a multi-level feature extraction model (such as MobileNetV2) as the backbone network. After the input image is unified, the feature maps d2, d3, d4 and d5 of multiple levels are obtained, and then the pyramid efficient multi-scale attention modules of multiple scales are input to obtain the feature-enhanced attention feature maps Z2, Z3, Z4 and Z5. Z2, Z3, Z4 and Z5 are upsampled step by step to obtain the final detection result. In particular, MobileNetV2 is a lightweight convolutional neural network with fewer parameters and computational complexity. It is very suitable for tasks that need to be applied in resource-constrained environments. Using MobileNetV2 for feature extraction of drone remote sensing images can efficiently capture the key features in the image, especially in high-resolution drone remote sensing images, and can effectively extract the texture and boundary features of different objects. After obtaining the multi-level features, the pyramid efficient multi-scale attention module is used to process these features, which can enhance the features of space and channels. Specifically, the pyramid efficient multi-scale attention module is used for sub-regions of different scales of the same feature map to enhance the expressiveness of the features, which can better identify and distinguish different spatial structures and texture features in the landslide area, avoid the omission of small-scale landslide areas, and resist the interference of irrelevant factors such as illumination. Finally, the upsampling module is used to decode the extracted and enhanced features. The upsampling module maps the features back to the size of the original image by gradually fusing features at different levels, and finally generates the prediction results of the landslide area. The design of the upsampling module fully considers the multi-level characteristics of the features, and can effectively fuse features at different levels to achieve accurate positioning of the landslide area.

[0026] Embodiment 1: A landslide area detection method based on multi-level and multi-scale feature enhancement, such as Figure 2 As shown, the method includes steps S1 to S4.

[0027] S1: A multi-level feature extraction model is used to extract features of landslide areas in remote sensing images to obtain feature maps at multiple levels.

[0028] In a further preferred solution, the multi-level feature extraction model is used to extract features of the landslide area in the remote sensing image to obtain feature maps at multiple levels, specifically: Determine the number of output layers and output layer of the multi-level feature extraction model according to the detection accuracy; Inputting the landslide area of ​​the remote sensing image into the multi-level feature extraction model, and obtaining a feature map of the corresponding level at each output layer; The size of the image of each output layer is one quarter of the output layer of the previous layer.

[0029] In a further preferred embodiment, the multi-level feature extraction model is a MobileNetV2 model, such as Figure 1 As shown, the MobileNetV2 model includes multiple bottleneck layers connected in sequence.

[0030] The MobileNetV2 model is a lightweight neural network designed for mobile and embedded devices. It can provide high accuracy while ensuring low parameter count and low computational complexity. The main structure of the MobileNetV2 model is the bottleneck layer (such as Figure 3 As shown), it is an inverse residual structure with two ends The convolutional layer has an activation function ReLU6 at the input (limiting the maximum output to 6), no activation function at the output, and a ReLU6 activation function in the middle. Depthwise separable convolution.

[0031] The MobileNetV2 model can be a pre-trained MobileNetV2 model downloaded from the pytorch official website, and in actual use, only the first 17 bottleneck layers are selected as the backbone network of the model. In this embodiment, the outputs of the third, sixth, thirteenth and seventeenth bottleneck layers are used as feature maps at different levels. The features after the first bottleneck layer are not used because the feature map of the lowest level contains too much noise and a lot of redundant information, and the feature map of this level has a large scale and high processing computational cost; the length and width of the feature map of the third bottleneck layer are one-fourth of the original image, that is, the size is one-sixteenth of the original image, and the number of channels is 24; the size of the feature map of the sixth bottleneck layer is one-fourth of the size of the feature map of the third bottleneck layer, and the number of channels is 32; the size of the feature map of the thirteenth bottleneck layer is one-fourth of the size of the feature map of the sixth bottleneck layer, and the number of channels is 96; the size of the feature map of the seventeenth bottleneck layer is one-fourth of the size of the feature map of the thirteenth bottleneck layer, and the number of channels is 320.

[0032] S2: Unify the number of channels of feature maps at multiple levels.

[0033] The feature maps of multiple levels output by different bottleneck layers are input into the convolutional layer to unify the number of channels. For example, they all pass through a After convolution, the number of channels in the convolutional layer is adjusted to 64, and the feature maps of each level are aligned in the channel dimension to facilitate subsequent operations.

[0034] S3: The feature maps with unified channel numbers are input into pyramid efficient multi-scale attention modules of different scales to obtain feature-enhanced attention feature maps of multiple scales.

[0035] In a further preferred embodiment, Figure 4 As shown, step S3 includes S31 and S32: S31: The feature maps of multiple levels after channel unification are input into pyramid efficient multi-scale attention modules of different scales respectively; S32: Each pyramid efficient multi-scale attention module performs multi-scale segmentation, efficient multi-scale attention extraction, splicing and convolution on the input feature map in sequence to obtain an output feature map, and superimposes the input feature map with the output feature map to obtain a corresponding attention feature map.

[0036] The multi-level multi-scale attention neural network adopts aggregated multi-scale contextual attention to enhance the ability to capture detailed change information. Therefore, based on the introduction of the attention mechanism of efficient multi-scale attention, the multi-level multi-scale attention neural network constitutes the Pyramid Efficient Multi-Scale Attention module (Pyramid Efficient Multi-Scale Attention, PEMA), which can aggregate spatiotemporal attention contexts of different scales to obtain multi-scale attention features.

[0037] like Figure 5 As shown in the figure, it is a schematic diagram of the structure of the pyramid efficient multi-scale attention module corresponding to the lowest layer feature map of the backbone network. It has four different branches, each branch divides the feature map into sub-regions, , representing the four scales of the pyramid. For the input feature map , where C is the number of channels of the feature map, H is the height of the feature map, and W is the width of the feature map; in this branch with a scale of N, the feature Figure X Each sub-region of can be expressed as , , For each sub-region of these four branches, an efficient multi-scale attention (EMA) is used to obtain the feature maps processed by the four branches. , Then the feature maps of the four branches are concatenated in the channel dimension, and a convolution layer is used to reduce the channel to the same level as the input feature. Figure XThe same size, get the feature map Finally, the elements in X and Y are superimposed to obtain the attention feature map Z at the corresponding scale. The superposition formula is as follows:

[0038] Among them, Efficient Multi-Scale Attention (EMA) is a new and efficient attention mechanism that aims to effectively capture multi-scale spatial information and improve the performance of convolutional neural networks in computer vision tasks. This method divides the information of the channel dimension into multiple sub-features by feature grouping to prevent the channel dimension from being overwritten. In addition, Figure 6 As shown in the figure, efficient multi-scale attention adopts a parallel processing strategy and uses three parallel paths to extract attention weight descriptors of grouped feature maps. The use of three parallel paths enables EMA to not only encode cross-channel information to adjust the importance of different channels, but also retain precise spatial structure information in the channel; finally, EMA uses a cross-spatial information aggregation method to model long-range dependencies and embeds precise location information into EMA. The fusion of contextual information at different scales enables the neural network to generate better pixel-level attention for feature maps.

[0039] EMA establishes long-range dependencies through cross-spatial information aggregation and embeds precise location information into them. EMA fuses contextual information of different scales, enabling convolutional neural networks to generate better pixel-level attention for high-level feature maps. EMA can handle short-range and long-range dependencies by using cross-spatial information aggregation methods by parallelizing convolution kernels. Compared with the asymptotic behavior that forms a limited receptive field, it uses and Convolution can better utilize contextual information between intermediate feature maps.

[0040] like Figure 7 As shown in the figure, for feature maps of different levels extracted by the backbone network, the number of branches designed in PEMA is different. The low-level feature map is taken as an example above. It is at the low level of the backbone network and contains more detailed information, so it is divided into 4 branches, and the feature map one level higher than the lowest level is divided into 3 branches. , because the resolution of the feature map at this level is smaller and contains less detailed information, there is no need to divide the area with a smaller scale; the feature map of the next higher level has only two branches, and each sub-area is divided into the size of the original feature map or a quarter, that is, ; And so on, there are up to two branches, each sub-region is divided into the size of the original feature map or a quarter, that is ; Similarly, the feature map of the highest layer contains more abstract features, and its height and width are the smallest, so only one branch is divided, that is, .

[0041] By applying the pyramid efficient multi-scale attention to feature maps at different levels, the spatial information between pixels at different levels can be fully explored, and the changed positions can be captured in the feature maps at different levels, so that the changed pixels can be better located during subsequent upsampling. Therefore, in this embodiment, the pyramid efficient multi-scale attention module adds four branches, which respectively process the feature maps at different levels extracted by the backbone network, locate the pixels of the landslide in the feature maps at each level, and achieve more accurate landslide detection.

[0042] S4: The attention feature maps of multiple scales are upsampled step by step to obtain the final remote sensing image landslide area detection result.

[0043] Step S4 specifically comprises using an upsampling module to perform step-by-step upsampling processing on the attention feature maps of multiple scales from a high scale to obtain the final remote sensing image landslide area detection result; Among them, Figure 8 As shown, the upsampling module includes multiple feature fusion units and multiple convolutional layers.

[0044] The feature fusion unit includes: an upsampling layer, Convolutional layer, batch normalization layer, and ReLU activation function.

[0045] High-level features contain abstract semantic information and can better locate the changed area; low-level features contain more details and can better represent the boundary and texture information of the object. Effective fusion of high-level features with low-level features can more accurately represent the specific situation of the changed area. Therefore, the present invention relates to an upsampling module for gradually fusing multi-level feature maps to achieve the effect of fully fusing the feature maps of each level and reduce the false alarm rate.

[0046] In this embodiment, the upsampling module includes three identical feature fusion units, which are used to fuse feature maps at different levels. The feature fusion unit includes an upsampling layer, a Convolutional layer, a batch normalization layer and ReLU activation function. The high-level feature map is first upsampled to the size of the low-level feature map through double linear interpolation, and then the two are concatenated in the channel dimension and then through The convolution reduces the dimension of the mechanical energy channel of the concatenated feature map to achieve the effect of feature fusion. After the convolution, a batch normalization layer and a ReLU activation function are added to accelerate network training and improve generalization ability. For example, the attention feature map Z5 is first linearly interpolated and upsampled to twice the size of the feature map Z4, and then the two are concatenated. The convolution layer realizes the feature fusion of the two, and then passes through the batch standardization layer and activation function. In the actual model, the number of channels of Z5, Z4, Z3 and Z2 is 64, so The number of channels of the convolution input is 128, and the number of channels of the output is 64. After the feature maps at all levels are fused, the final fused feature map is linearly interpolated and upsampled to the size of the input image. The convolutional layer and a The convolutional layer is further used to perform feature fusion to obtain the final output feature map. Finally, the feature map is passed through a Sigmoid function to obtain the predicted result map. The area with a pixel value greater than 0.5 is a landslide, and the area with a pixel value less than 0.5 is the background.

[0047] The upsampling module can avoid high-level features Figure 1 The multiple times are used to reduce feature distortion and prevent the increase of false changed pixels. By gradually fusing feature maps at different levels, the fusion of features at different levels is better achieved and the false alarm rate is reduced.

[0048] Experimental verification of the accuracy of the method of the present invention (1) Dataset The training and testing used unmanned remote sensing data from the Chinese Academy of Sciences Landslide Dataset, a large multi-sensor dataset for landslide detection based on deep learning developed by the Artificial Intelligence Group of the Institute of Mountain Hazards and Environment (CAS) of the Chinese Academy of Sciences. The dataset aims to address the challenges encountered in landslide identification. With the increasing incidence of landslides caused by climate change and earthquakes, there is a growing need for an accurate and comprehensive dataset to support rapid and effective landslide identification. Compared with existing datasets with limitations in dataset size, coverage, sensor type, and resolution, the CAS landslide dataset contains 20,865 images and integrates satellite and drone data from 9 regions. This example only uses drone data from the Chinese Academy of Sciences Landslide Dataset, totaling 13,443 images, and randomly divides each sub-dataset into a training set and a validation set in a ratio of 7:3.

[0049] (2) Evaluation indicators In order to accurately and effectively evaluate the accuracy of the method, the present invention uses overall accuracy, precision, recall rate and F1 score as evaluation indicators. In the overall accuracy, TP (True Positive) indicates the number of positive class samples correctly predicted as positive class samples, FP (False Positive) indicates the number of negative class samples incorrectly predicted as positive class samples, FN (False Negative) indicates the number of positive class samples incorrectly predicted as negative class samples, and TN (True Negative) indicates the number of negative class samples correctly predicted as negative class samples.

[0050] Precision refers to the proportion of samples that are actually positive among the samples predicted as positive in the detection method. It measures the accuracy of the detection method in predicting positive samples. Recall refers to the proportion of all positive samples that are successfully predicted as positive. It measures the coverage of positive samples, that is, how many positive samples can be detected. A high recall rate means that positive samples can be captured well. The F1 score is the harmonic mean of the precision and recall rate. It takes into account the accuracy and comprehensiveness of the model and is a more comprehensive indicator. The F1 score ranges from 0 to 1. The higher its value, the better the model performance. In summary, the calculation formulas for overall accuracy, precision, recall and F1 score are as follows:

[0051]

[0052] in,

[0053] (3) Experimental equipment In terms of parameter setting, the detection method of this application is trained on an Nvidia RTX4060 GPU, using PyTorch as the framework, version 2.1.1, CUDA version 12.6, Adam optimizer, learning rate 0.005, momentum 0.9, weight decay 0.0005, batch size 8, and epoch 200.

[0054] (4) Ablation analysis The ablation experiment results are shown in Table 1 below. It can be seen that the precision and F1 score of the detection method in this application have been improved, the recall rate has slightly decreased, and the accuracy of the detection of the landslide area has been improved.

[0055] Table 1

[0056] (5) Comparison with other methods Table 2 shows the comparison between the method in this application and other methods. Compared with other methods, the method proposed in this application has superior performance. From the experimental results in Table 2, it can be seen that the method in this application performs best among the above five methods in terms of Precision, Recall, F1 score and overall accuracy (OA), which is significantly better than other methods. This shows that the method in this application has higher accuracy and reliability in dealing with landslide extraction tasks.

[0057] Table 2

[0058] The visualization results of the method in this application are as follows Fig. 9 As shown, the first line is the UAV remote sensing image, the second line is the real label, and the third line is the prediction result. It can be seen that the method in this application can well separate the landslide area from other background areas and keep the missed detection rate and false detection rate low.

[0059] Embodiment 2: The present invention based on the same inventive concept also provides a landslide area detection system based on multi-level and multi-scale feature enhancement, comprising: Multi-level feature extraction model, used to extract features from landslide areas in remote sensing images to obtain feature maps at multiple levels; Convolution module, used to unify the number of channels of feature maps at multiple levels; Pyramid efficient multi-scale attention module, used to obtain attention feature maps of multiple scales that capture changes in feature map information; The upsampling module is used to perform step-by-step upsampling of attention feature maps of multiple scales to obtain the final remote sensing image landslide area detection results.

[0060] In the segmentation task of landslide extraction, the area where the landslide occurs is much smaller than the area where the landslide does not occur, which leads to the problem of class imbalance. In order to alleviate this problem and guide the network to learn from complex scenes, this application adopts a hybrid loss, including binary cross entropy (BCE) loss and Dice loss. Therefore, the loss function of the landslide area detection system based on multi-level and multi-scale feature enhancement is:

[0061] in, is the loss function of the landslide area detection system based on multi-level and multi-scale feature enhancement, is the binary cross entropy loss function, , · is the dot multiplication operation, c This is the final remote sensing image landslide area detection result. g is the true label, is the Dice loss function, , It is the first paradigm.

[0062] Optionally, the multi-level feature extraction model extracts features from the landslide area in the remote sensing image to obtain feature maps at multiple levels, specifically: Determine the number of output layers and output layer of the multi-level feature extraction model according to the detection accuracy; Inputting the landslide area of ​​the remote sensing image into the multi-level feature extraction model, and obtaining a feature map of the corresponding level at each output layer; The size of the image of each output layer is one quarter of the output layer of the previous layer.

[0063] Optionally, the multi-level feature extraction model is a MobileNetV2 model, and the MobileNetV2 model includes multiple bottleneck layers connected in sequence.

[0064] Optionally, the output layer of the multi-level feature extraction model has 4 layers, and the output layers are respectively the third layer, the sixth layer, the thirteenth layer and the seventeenth layer of the bottleneck layer.

[0065] Optionally, the convolution module unifies the number of channels of feature maps at multiple levels, specifically: The same convolutional layer is used to unify the number of channels of feature maps at multiple levels.

[0066] Optionally, the pyramid efficient multi-scale attention module obtains attention feature maps of multiple scales for image enhancement, specifically: The feature maps of multiple levels after channel unification are input into pyramid efficient multi-scale attention modules of different scales respectively; Each pyramid efficient multi-scale attention module performs multi-scale segmentation, efficient multi-scale attention extraction, splicing and convolution on the input feature map in sequence to obtain the output feature map, and superimposes the input feature map with the output feature map to obtain the corresponding attention feature map.

[0067] Optionally, the upsampling module performs step-by-step upsampling processing on the attention feature maps of multiple scales to obtain the final remote sensing image landslide area detection result, including: The upsampling module is used to perform step-by-step upsampling of the attention feature maps of multiple scales from high scale to obtain the final remote sensing image landslide area detection result; Wherein, the upsampling module includes multiple feature fusion units and multiple convolutional layers.

[0068] Optionally, the feature fusion unit in the upsampling module includes: an upsampling layer, Convolutional layer, batch normalization layer, and ReLU activation function.

[0069] The above are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are included in the scope of the claims of the present invention to be approved.

Claims

1. A landslide area detection method based on multi-level and multi-scale feature enhancement, characterized in that: include: A multi-level feature extraction model is used to extract features of landslide areas in remote sensing images to obtain feature maps at multiple levels. Unify the number of channels of feature maps at multiple levels; The feature maps with unified channel numbers are input into pyramid efficient multi-scale attention modules of different scales to obtain feature-enhanced attention feature maps of multiple scales; The attention feature maps of multiple scales are upsampled step by step to obtain the final remote sensing image landslide area detection result.

2. The landslide area detection method based on multi-level and multi-scale feature enhancement according to claim 1 is characterized in that: The multi-level feature extraction model is used to extract features of the landslide area in the remote sensing image to obtain feature maps at multiple levels, specifically: Determine the number of output layers and output layer of the multi-level feature extraction model according to the detection accuracy; Inputting the landslide area of ​​the remote sensing image into the multi-level feature extraction model, and obtaining a feature map of the corresponding level at each output layer; The size of the image of each output layer is one quarter of the output layer of the previous layer.

3. The landslide area detection method based on multi-level and multi-scale feature enhancement according to claim 2 is characterized in that: The multi-level feature extraction model is a MobileNetV2 model, and the MobileNetV2 model includes a plurality of bottleneck layers connected in sequence.

4. The landslide area detection method based on multi-level and multi-scale feature enhancement according to claim 3 is characterized in that: The output layer has four layers, and the output layers are respectively the third, sixth, thirteenth and seventeenth layers of the bottleneck layer.

5. The landslide area detection method based on multi-level and multi-scale feature enhancement according to claim 1 is characterized in that: The number of channels of feature maps at multiple levels is unified, specifically: The same convolutional layer is used to unify the number of channels of feature maps at multiple levels.

6. The landslide area detection method based on multi-level and multi-scale feature enhancement according to claim 1 is characterized in that: The feature map with unified channel number is input into the pyramid efficient multi-scale attention module of different scales to obtain the attention feature map of multiple scales with feature enhancement, specifically: The feature maps of multiple levels after channel unification are input into pyramid efficient multi-scale attention modules of different scales respectively; Each pyramid efficient multi-scale attention module performs multi-scale segmentation, efficient multi-scale attention extraction, splicing and convolution on the input feature map in sequence to obtain the output feature map, and superimposes the input feature map with the output feature map to obtain the corresponding attention feature map.

7. The landslide area detection method based on multi-level and multi-scale feature enhancement according to claim 1 is characterized in that: The step-by-step upsampling of the attention feature maps at multiple scales to obtain the final remote sensing image landslide area detection result includes: The upsampling module is used to perform step-by-step upsampling of the attention feature maps of multiple scales from high scale to obtain the final remote sensing image landslide area detection result; Wherein, the upsampling module includes multiple feature fusion units and multiple convolutional layers.

8. The landslide area detection method based on multi-level and multi-scale feature enhancement according to claim 7 is characterized in that: The feature fusion unit includes: an upsampling layer, Convolutional layer, batch normalization layer, and ReLU activation function.

9. A landslide area detection system based on multi-level and multi-scale feature enhancement, characterized in that: include: Multi-level feature extraction model, used to extract features from landslide areas in remote sensing images to obtain feature maps at multiple levels; Convolution module, used to unify the number of channels of feature maps at multiple levels; Pyramid efficient multi-scale attention module, used to obtain multi-scale attention feature maps for image enhancement; The upsampling module is used to perform step-by-step upsampling of attention feature maps of multiple scales to obtain the final remote sensing image landslide area detection results.

10. The landslide area detection system based on multi-level and multi-scale feature enhancement according to claim 9 is characterized in that: The loss function of the landslide area detection system based on multi-level and multi-scale feature enhancement is: in, is the loss function of the landslide area detection system based on multi-level and multi-scale feature enhancement, is the binary cross entropy loss function, , · is the dot multiplication operation, c This is the final remote sensing image landslide area detection result. g is the true label, is the Dice loss function, , It is the first paradigm.

Citation Information

Patent Citations

  • Landslide detection method and device for small-size model

    CN112883907A

  • Remote sensing image target detection method based on cross-scale feature fusion pyramid network

    CN113128559A

  • Landslide remote sensing information extraction method based on convolutional neural network and category thermodynamic diagram

    CN113408462A

  • Conveyor belt tearing detection method based on improved regional convolutional neural network

    CN114926733A

  • Landslide detection method based on improved lightweight YOLOv7

    CN117746252A