An Automatic Landslide Recognition Method, System, Device and Medium Based on Lightweight Convolutional Neural Network and Dual Attention
By adopting lightweight convolutional neural networks and dual attention methods in landslide recognition, the problem of relying on high-cost sensors and huge computing volume in the existing technology is solved, efficient and real-time landslide recognition is achieved, and the dependence on expert knowledge is reduced.
Patent Information
- Application Number
- CN202310218039.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-08
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-03-08
AI Technical Summary
The existing landslide recognition technology has the problem of relying on high-cost sensors, huge computing volume and difficulty in real-time operation, and is highly dependent on expert knowledge and low efficiency.
The automatic landslide recognition method based on lightweight convolutional neural network and dual attention is adopted. The network model composed of Fused-MBConv module, MBConv module and PDC module is combined with the dual attention mechanism to increase the globality of spatial perception and reduce the complexity of the model.
It improves the accuracy and efficiency of landslide recognition, reduces the complexity and calculation cost of the model, realizes real-time recognition, and reduces the dependence on expert knowledge.
Smart Images

Figure CN116206214B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic landslide recognition, and particularly relates to an automatic landslide recognition method, system, device and medium based on a lightweight convolutional neural network and dual attention. Background Art
[0002] China is a country prone to geological disasters. Among them, landslides are one of the most common geological disasters. Landslide disasters pose a serious threat to the lives and property of local people, important infrastructure, etc. Therefore, it is of great significance to identify landslides. At present, the relatively mature technology for landslides is the InSAR technology for early landslide identification. However, due to its high sensor cost, difficulty in obtaining multiple images, long observation period, high technical requirements, etc., its large-scale application in the field of landslide identification is limited.
[0003] At present, most of the image classification methods applied to landslide detection are improvements based on currently relatively mature common networks. For example, Dawen Yu et al. proposed a method based on the attention mechanism for landslide detection in satellite remote sensing images (Ji S, Yu D, Shen C, et al. Landslide detection from an open satellite imagery and digital elevation model dataset using attention boosted convolutional neural networks[J]. Landslides, 2020, 17(6): 1337-1352), adding the proposed SCAM attention mechanism method to improve the network performance on these image classification networks such as VGG-Net and ResNet; Sameen et al. proposed using ResNet for landslide detection (Sameen M I, Pradhan B. Landslide detection using residual networks and the fusion of spectral and topographic information[J]. IEEE Access, 2019, PP(99): 114363-114373), which fuses spectral and topographic information and improves the detection results by adding features; but the above are all improvements based on models with relatively high network depth and complexity to improve the network performance for landslide recognition.
[0004] Currently, the automatic landslide recognition method still has the following points:
[0005] 1. Early landslide identification through manual visual interpretation: This method requires combining geometric features such as the color and texture of remote sensing images with expert knowledge and non-remote sensing data, with relatively high identification accuracy. However, it has problems such as strong dependence on expert knowledge, high labor costs, and low efficiency.
[0006] 2. Early landslide identification using InSAR technology: Due to reasons such as high sensor costs, difficulty in obtaining multiple images, long observation cycles, and high technical requirements, its large-scale application in the field of landslide monitoring is restricted.
[0007] 3. Technologies for landslide identification using deep learning: Currently, mainstream technologies still improve on models with relatively high network depth and complexity to enhance the landslide identification effect. Although the performance of the network model has been improved, these networks often have huge computational amounts, and it is difficult for detection algorithms relying on these basic networks to meet the requirements of real-time operation.
[0008] EfficientNetV2 published in April 2021 (Tan M, Le Q. Efficientnetv2: Smaller models and faster training[C] / / International conference on machine learning. PMLR, 2021: 10096 - 10106) is a lightweight network model. This paper proposed an improved progressive learning method that dynamically adjusts the regularization method according to the size of the training images. In addition, it was verified that the use of depthwise separable convolutions in shallow networks is very slow, so a new network module Fused-MBConv was proposed and applied to shallow networks to improve the training speed of the network and enhance the accuracy of the network. However, because its network architecture is obtained through neural architecture search, it is not fully applicable to landslide identification. Summary of the Invention
[0009] In order to overcome the above deficiencies of the prior art, the purpose of the present invention is to provide an automatic landslide identification method, system, device, and medium based on a lightweight convolutional neural network and dual attention. The main framework of the network model is designed using the Fused-MBConv module and the MBConv module as the main part of the lightweight convolutional neural network. The PDC module is used in the main framework of the network model to increase the global awareness of the spatial perception of the network model. The dual attention mechanism module is mainly used in the PDC module to perform self-attention mechanisms respectively from the spatial dimension and the channel dimension to achieve global modeling. The Our network framework network model of the present invention combines the advantages of the convolutional neural network and the self-attention mechanism to create a lightweight network model.
[0010] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0011] An automatic landslide recognition method based on a lightweight convolutional neural network and dual attention, comprising the following steps:
[0012] Step 1: Preprocess the remote sensing image to obtain a landslide remote sensing image dataset;
[0013] Step 2: Build an Ournetwork framework network model;
[0014] Step 3: Iteratively train the Ournetwork framework network model, and finally save the weights of the round with the best training effect;
[0015] Step 4: Predict the network model trained in Step 3, and use accuracy Accuracy, balanced F-score F1, recall Recall, and precision Precision as evaluation indicators for evaluating the network performance to evaluate the quality of the network;
[0016] Step 5: Randomly generate a tensor using a deep learning framework, send it into the Ournetwork framework network model obtained in Step 2, and calculate FLOPs, the number of parameters, and the model inference time, where FLOPs represents a measure of the complexity of the model.
[0017] The specific method of Step 1 is as follows:
[0018] Step 1.1: Obtain positive samples of the landslide remote sensing image dataset: Crop the remote sensing image according to the landslide scale and expand the pixel points outward as the background to obtain positive samples;
[0019] Step 1.2: Obtain negative samples of the landslide remote sensing image dataset: Crop the remote sensing image in sizes of 128 or 256 or 512, delete all data containing positive samples, and select special scenes as negative samples;
[0020] Step 1.3: Divide the positive samples obtained in Step 1.1 and the negative samples obtained in Step 1.2 into a training set and a test set according to a certain ratio;
[0021] Step 1.4: Divide the training set obtained in Step 1.3 into a new training set and a validation set according to a certain ratio;
[0022] Step 1.5: Perform data augmentation on the new training set obtained in Step 1.4;
[0023] Step 1.6: Resize the data of the test set in Step 1.3, the validation set in Step 1.4, and the training set after data augmentation in Step 1.5 to a unified size of 224*224 or 300*300.
[0024] The specific method of Step 2 is as follows:
[0025] Step 2.1: Use a deep learning framework to construct a Fused-MBConv module, an MBConv module, and a PDC module respectively;
[0026] Step 2.2: Fuse the Fused-MBConv module, MBConv module, and PDC module constructed in Step 2.1.
[0027] The specific method of Step 2.1 is as follows: The Fused-MBConv module includes a 3*3 upsampling convolution and a 1*1 downsampling convolution, and performs a skip connection between the feature input and feature output of the Fused-MBConv module; The MBConv module includes a 1*1 upsampling convolution, a 3*3 depthwise separable convolution, an SE attention mechanism, and a 1*1 downsampling convolution, and performs a skip connection between the feature input and feature output of the MBConv module; The PDC module consists of a 1*1 convolution, a PatchEmded module, a dual attention mechanism, a feature transformation operation, and a 1*1 convolution, and performs a shortcut link between the input feature map and the final output feature map, uses the PatchEmbed module to adjust the size of the feature map channels, and flattens the feature map into a feature vector suitable for the DualAttention Block;
[0028] The specific method of Step 2.2 is as follows: First, construct a 1*1 standard convolutional layer, send its output feature map into the Fused-MBConv module with a stride of 1, and this time no upsampling is performed. The Fused-MBConv module loops twice, and sends its output features into the Fused-MBConv module with a stride of 2 for upsampling by 4 times the number of input feature channels, loops 4 times, sends its output features into the Fused-MBConv module with a stride of 2 for upsampling by 4 times the number of input feature channels, sends its output features into the MBConv module with a stride of 2 for upsampling by 4 times the number of input feature channels, loops 2 times, sends its output features into the PDC module with a stride of 1, loops 2 times, sends its output features into the MBConv module with a stride of 2 for upsampling by 4 times the number of input feature channels, loops 3 times, sends its output features into the PDC module with a stride of 1, loops 3 times, and finally sends it into a 1*1 convolution for pooling operation and fully connected operation to obtain a binary classification result, completing the construction of the network model.
[0029] The specific method of step 3 is as follows: The training set obtained in step 1.6 is fed into the Our network framework network model built in step 2, and the Our network framework network model is iteratively trained. Set parameters including the number of iteration rounds, batch size, and learning rate, perform network training and record the training time. Then, the validation set of step 1.6 is fed into the Our network framework network model built in step 2 to verify the model effect. Finally, save the weights of the round with the best training effect.
[0030] The specific method of step 4 is as follows: The test set of step 1.6 is fed into the network model trained in step 3 for prediction to evaluate the quality of the network. Use accuracy (Accuracy), balanced F-score (F1), recall (Recall), and precision (Precision) as evaluation indicators for judging network performance. Accuracy is the proportion of correctly classified samples in the total number of samples; Recall represents the probability that a positive sample is predicted as a positive sample among the actual positive samples; Precision represents the probability that a sample predicted as a positive sample is actually a positive sample among all samples predicted as positive samples. Precision and Recall are a pair of contradictory metrics. When Precision is high, Recall is low; when Recall is high, Precision is low. When both Precision and Recall need to be considered, the balanced F-score (F1) is used as the weighted harmonic mean of Precision and Recall. The following are the calculation formulas for each evaluation indicator:
[0031]
[0032] Among them, TP represents the number of correctly predicted positive samples, FP represents the number of incorrectly predicted negative samples, TN represents the number of correctly predicted negative samples, and FN represents the number of incorrectly predicted positive samples;
[0033]
[0034]
[0035]
[0036] Among them, P refers to Precision, R refers to Recall, and F 1 refers to the balanced F-score;
[0037] In step 5, the dimension size of the tensor is the dimension and size of the image after uniformly adjusting the dataset size in step 1.
[0038] The special scenarios in step 1.2 include: roads, cultivated land, lakes, vegetation, cloud occlusion, factories, ridges, and snow-covered scenarios; the training set and test set in step 1.3 are divided according to a ratio of 6:4 or 8:2; the ratio range of the new training set and validation set in step 1.4 is: 7:3 - 9:1; the data augmentation in step 1.5 includes: random cropping, random rotation, and adding noise.
[0039] The present invention also provides an automatic landslide recognition system based on a lightweight convolutional neural network and dual attention, including:
[0040] Mobile flipped bottleneck convolution module: This module consists of a 1*1 dimensionality increase convolution, a 3*3 depthwise separable convolution, an SE attention mechanism, and a 1*1 dimensionality reduction convolution. It makes a skip connection between the feature input and feature output of the MBConv module. The inverted residual structure is used in this module, and at the same time, the depthwise separable convolution is used to reduce the number of parameters required for convolution calculations to achieve the purpose of lightweighting the network model.
[0041] Fused mobile flipped bottleneck convolution module: This module consists of a 3*3 dimensionality increase convolution and a 1*1 dimensionality reduction convolution, and makes a skip connection between the feature input and feature output of the Fused-MBConv module. It is applied in the shallow network to improve the training speed of the network.
[0042] Dual attention mechanism module: This module consists of a window multi-head self-attention mechanism in the spatial dimension and a self-attention mechanism in the channel group. It processes problems from an orthogonal perspective, performs self-attention mechanisms from the spatial dimension and the channel dimension respectively to improve the accuracy of the automatic landslide recognition network model.
[0043] Spatial window multi-head self-attention mechanism module: The spatial dimension part of the dual attention mechanism module uses spatial dimension information to improve local features.
[0044] Channel group self-attention mechanism module: The channel dimension part of the dual attention mechanism module uses channel dimension information to capture global connections and features.
[0045] PDC module: This module consists of a 1*1 convolution, a PatchEmded module, a dual attention mechanism, a feature transformation operation, and a 1*1 convolution, and makes a skip link between the input feature map and the final output feature map. It adjusts the input feature map to adapt to the input of the dual attention mechanism module and adjusts the output features of the dual attention mechanism module to adapt to the input of the convolutional neural network layer.
[0046] The present invention also provides an automatic landslide recognition device based on a lightweight convolutional neural network and dual attention, including:
[0047] Memory: for storing computer programs;
[0048] A processor, configured to implement the automatic landslide recognition method based on a lightweight convolutional neural network and dual attention when executing the computer program.
[0049] The present invention also provides a computer-readable storage medium, including:
[0050] The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement an automatic landslide recognition method based on a lightweight convolutional neural network and dual attention.
[0051] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0052] 1. In the existing technologies for landslide recognition using deep learning, more rely on relatively classic and deep network models with high complexity to improve the accuracy of landslide recognition. However, in the present invention, the Our networkframework network model is constructed, which is composed of Fused-MBConv module, MBConv module and PDC module. A lightweight convolutional neural network more suitable for landslide recognition is designed using the Fused-MBConv module and MBConv module. At the same time, the dual attention mechanism in the PDC module is combined to increase the global perception of the network model's space, thereby improving the performance of the landslide recognition network model. The present invention uses a lightweight model to recognize landslides and simultaneously introduces a dual attention mechanism to improve the accuracy of landslide recognition. Therefore, while reducing the model complexity, the overall accuracy of landslide recognition is also improved. Compared with EfficientNetV2, the present invention has improvements in F1, accuracy, precision, and recall in automatic landslide recognition, and at the same time, the model complexity, the number of parameters, the inference time, and the training time are all reduced. Based on EfficientNetV2, the present invention improves the training speed of the network and the accuracy of the network, making it more suitable for landslide recognition tasks.
[0053] 2. Currently, the relatively mature technology for landslide is the InSAR technology for early landslide recognition. However, due to reasons such as high sensor cost, difficulty in obtaining multiple images, long observation period, and high technical requirements, its large-scale application in the field of landslide recognition is restricted. Compared with the InSAR technology, the automatic landslide recognition technology proposed by the present invention is easier to obtain images and does not require high-cost sensor requirements, etc. Therefore, the present invention can be used for a preliminary screening of landslide recognition to improve the efficiency of landslide recognition.
[0054] 3. Compared with the method of manual visual interpretation, the present invention does not strongly rely on expert knowledge because it mainly uses computer technology, with low labor costs and higher efficiency than manual visual interpretation.
[0055] 4. Compared with InSAR technology, the optical remote sensing images in the present invention have low costs and the data is easier to obtain. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is the network architecture diagram of the present invention.
[0057] Figure 2 is Figure 1 The architecture diagram of the DualAttention Block module in the PDC module. DETAILED DESCRIPTION OF THE INVENTION
[0058] The following takes an automatic landslide recognition method based on a lightweight convolutional neural network and dual attention as an example, and further explains and illustrates the technical solution of the present invention in combination with the drawings.
[0059] Step 1: Preprocess the remote sensing images of the Gaofen-2 satellite in Ningxia to obtain the remote sensing image dataset of landslides in Ningxia; the present invention mainly focuses on the remote sensing dataset of landslides and the dataset of potential landslide points (i.e., those that have not yet experienced landslides but have the tendency to occur).
[0060] Step 1.1: Obtain positive samples of the remote sensing image dataset of landslides in Ningxia: Crop the remote sensing images of the Gaofen-2 satellite in Ningxia according to the landslide scale and expand 20 pixel points outward as the background to obtain positive samples;
[0061] Step 1.2: Obtain negative samples of the remote sensing image dataset of landslides in Ningxia: Crop the remote sensing images of the Gaofen-2 satellite in Ningxia in a size of 512*512, delete all data containing positive samples, and select special scenes that are easily misidentified as positive samples, such as roads, cultivated land, lakes, vegetation, cloud cover, factories, ridges, snow cover, etc., as negative samples;
[0062] Step 1.3: Divide the positive samples obtained in Step 1.1 and the negative samples obtained in Step 1.2 into a training set and a test set according to a ratio of 6:4;
[0063] Step 1.4: Divide the training set obtained in Step 1.3 into a new training set and a validation set according to a ratio of 8:2;
[0064] Step 1.5: Perform simple data augmentation such as random cropping, random rotation, and adding noise to the new training set obtained in Step 1.4;
[0065] Step 1.6: Resize the data of the test set in Step 1.3, the validation set in Step 1.4, and the training set after data augmentation in Step 1.5 to a unified size of 300*300.
[0066] Step 2: Build the Ournetwork framework network model.
[0067] As Figure 1 shown, Step 2.1: Use the pytorch deep learning framework to construct the Fused-MBConv module, the MBConv module, and the PDC module respectively;
[0068] As Figure 2 shown, specifically, the Fused-MBConv module includes a 3*3 upsampling convolution and a 1*1 downsampling convolution, and performs a skip connection (ShortcutConnection) between the feature input and the feature output of the Fused-MBConv module; the MBConv module includes a 1*1 upsampling convolution, a 3*3 depthwise separable convolution, an SE (Squeeze-and-Excitation) attention mechanism, and a 1*1 downsampling convolution, and performs a skip connection between the feature input and the feature output of the MBConv module; the PDC module consists of a 1*1 convolution, a PatchEmded module, a dual attention mechanism, a feature transformation operation, and a 1*1 convolution, and performs a shortcut link between the input feature map and the final output feature map, uses the PatchEmbed module to adjust the size of the feature map channels, and flattens the feature map into a feature vector suitable for the DualAttention Block;
[0069] Step 2.2: Integrate the Fused-MBConv module, the MBConv module, and the PDC module constructed in Step 2.1;
[0070] Specifically, first, a 1*1 standard convolutional layer is constructed, and its output feature map is fed into a Fused-MBConv module with a stride of 1 (no dimensionality increase is performed this time). The Fused-MBConv module loops twice, and its output features are fed into a Fused-MBConv module with a stride of 2 (performing a dimensionality increase that is 4 times the number of input feature channels). This loops 4 times. Then, its output features are fed into a Fused-MBConv module with a stride of 2 (performing a dimensionality increase that is 4 times the number of input feature channels), and then into a MBConv module with a stride of 2 (performing a dimensionality increase that is 4 times the number of input feature channels). This loops 2 times. Then, its output features are fed into a PDC module with a stride of 1 and looped 2 times. Then, its output features are fed into a MBConv module with a stride of 2 (performing a dimensionality increase that is 4 times the number of input feature channels) and looped 3 times. Then, its output features are fed into a PDC module with a stride of 1 and looped 3 times. Finally, it is fed into a 1*1 convolution for pooling operation and fully connected operation to obtain a binary classification result, completing the construction of the network model.
[0071] Step 3: Iteratively train the Our network framework network model and finally save the weights of the round with the best training effect.
[0072] Specifically, the training set obtained in Step 1.6 is fed into the Our network framework network model built in Step 2 for iterative training of the Our network framework network model. The number of experimental iteration rounds (epoch) is set to 60, and the batch size (batchsize) is set to 8. An exponentially decaying learning rate is used to train the network, with the initial learning rate set to 0.01. The network is trained according to the above parameters and the training time (Train-time) is recorded. Then, the validation set in Step 1.6 is fed into the Our network framework network model built in Step 2 to verify the model effect. Finally, the weights of the round with the best training effect are saved.
[0073] Step 4: Predict the network model trained in Step 3, and use accuracy (Accuracy), balanced F-score (F1), recall (Recall), and precision (Precision) as evaluation metrics to evaluate the performance of the network and determine the quality of the network.
[0074] Specifically, the test set in step 1.6 is fed into the network model trained in step 3 for prediction to evaluate the quality of the network. The accuracy (Accuracy), balanced F-score F1, recall (Recall), and precision (Precision) are used as evaluation metrics for the network performance. The accuracy Accuracy is the proportion of correctly classified samples in the total number of samples; the recall Recall represents the probability that a predicted positive sample is an actual positive sample; the precision Precision represents the probability that an actual positive sample is among all predicted positive samples; Precision and Recall are a pair of conflicting metrics. When Precision is high, Recall is low; when Recall is high, Precision is low; when both Precision and Recall need to be considered, the balanced F-score F1 index is used as the weighted harmonic mean of Precision and Recall. The following are the calculation formulas for each evaluation metric:
[0075]
[0076] Where TP represents the number of correctly predicted positive samples, FP represents the number of incorrectly predicted negative samples, TN represents the number of correctly predicted negative samples, and FN represents the number of incorrectly predicted positive samples;
[0077]
[0078]
[0079]
[0080] Where P refers to Precision, R refers to Recall, and F 1 refers to the balanced F-score.
[0081] Step 5: Randomly generate a tensor of (3, 300, 300) using the deep learning framework, and feed it into the Our network framework network model obtained in step 2 to calculate FLOPs (used to measure the complexity of the model), the number of parameters Params, and the model inference time Infer-time.
[0082] Such as Figure 1As shown in the figure, the overall framework of the landslide recognition network model based on the lightweight convolutional neural network and the dual attention mechanism. The overall Ournetwork framework model is composed of the MBConv module, the Fused-MBConv module, and the PDC module. First, the Fused-MBConv module and the MBConv module are used to construct the main framework of the network model, which serves as the main part of the lightweight network. Then, the PDC module is connected after each MBConv module to improve the performance of the network.
[0083] In the present invention, the MBConv module is mainly used for lightweighting to reduce the number of parameters and the amount of computation of the network model. The MBConv module is composed of a 1*1 dimensionality-increasing convolution, a 3*3 depthwise separable convolution, an SE attention mechanism, and a 1*1 dimensionality-reducing convolution, and a skip connection is made between the input feature map and the output feature map of the MBConv module.
[0084] Since using the MBConv module in the shallow layer will cause the training speed to slow down, the Fused-MBConv module is used in the shallow network to accelerate the training speed. This module is composed of a 3*3 dimensionality-increasing convolution and a 1*1 dimensionality-reducing convolution, and a skip connection is made between the input feature map and the output feature map of the Fused-MBConv module.
[0085] Although the convolutional neural network has fewer parameters to be trained compared with the Transformer, it has locality in spatial perception, while global spatial perception is crucial for computer vision tasks such as image classification and semantic segmentation. Although the Transformer can obtain global features, it is heavyweight in terms of the number of parameters. Therefore, how to combine the convolutional neural network and the Transformer to construct a network model with high accuracy and lightweight is of research significance. For this purpose, the present invention designs the lightweight network model Ournetwork framework that combines the convolutional neural network and the Transformer, and improves the accuracy of landslide recognition by adding the PDC module to the lightweight convolutional neural network.
[0086] The PDC module mainly consists of a 1×1 convolution for adjusting the number of channels of the feature map, a PatchEmbed that transforms the feature map into the feature vectors required for the input of the self-attention mechanism, i.e., performs a flattening operation, a dual attention mechanism module that first uses the window multi-head self-attention mechanism in the spatial dimension and then follows it with the channel group self-attention mechanism to perceive global information, then transforms the feature vectors into a feature map according to the size of the feature map before inputting PatchEmbed, and a 1×1 convolution for adjusting the number of channels of the output feature map. Finally, the feature map input to the PDC module is skip-connected to the output feature map of the last 1×1 convolution of the PDC module as the output feature map of the PDC module.
[0087] The above is the introduction of each module of the lightweight network model of the present invention. In the Our network framework network model, the network structure is as follows: a 3×3 convolution with a stride of 2 is used for downsampling; the Fused-MBConv module is used in a loop as follows: the Fused-MBConv module with a stride of 1 is looped 2 times, the Fused-MBConv module with a stride of 2 and an expansion factor of 4 is looped 4 times, and the Fused-MBConv module with a stride of 2 and an expansion factor of 4 is looped 1 time; the MBConv module with a stride of 2 and an expansion factor of 4 is looped 2 times; the PDC module with a stride of 1 is looped 2 times; the MBConv module with a stride of 2 and the MBConv module with an expansion factor of 4 are looped 3 times; the PDC module with a stride of 1 is looped 3 times; a 1×1 convolution, pooling operation, and fully connected operation are used.
[0088] The Our network framework network model constructed in the present invention is composed of the Fused-MBConv module, the MBConv module, and the PDC module. A lightweight convolutional neural network more suitable for landslide recognition is designed using the Fused-MBConv module and the MBConv module, and at the same time, the dual attention mechanism in the PDC module is combined to increase the global awareness of the spatial perception of the network model, so as to improve the performance of the landslide recognition network model.
[0089] Specifically, first, the main framework of the network model is designed using the Fused-MBConv module and the MBConv module, which serves as the main part of the lightweight convolutional neural network, enabling it to reduce the network complexity, running time, and the number of parameters while ensuring accuracy. Since the convolutional neural network is used as the backbone of the network model framework, it has locality in spatial perception, and the self-attention mechanism can effectively obtain global perception. Therefore, the PDC module is used in the main framework of the network model of the present invention to increase the global nature of the network model's spatial perception. In the PDC module, a dual attention mechanism module is mainly used to perform self-attention mechanisms separately from the spatial dimension and the channel dimension to achieve global modeling. Our network framework combines the advantages of the convolutional neural network and the self-attention mechanism to create a lightweight network model.
[0090] The present invention also provides an automatic landslide recognition system based on a lightweight convolutional neural network and dual attention, including:
[0091] Mobile inverted bottleneck convolution (MBConv) module (see appendix Figure 1 ): This module consists of a 1*1 upsampling convolution, a 3*3 depthwise separable convolution, an SE attention mechanism, and a 1*1 downsampling convolution. The feature input and feature output of the MBConv module are skip-connected. Since less information is lost after the high-dimensional information passes through the ReLU activation function, an inverted residual structure is used in this module, and at the same time, depthwise separable convolution is used to reduce the number of parameters required for convolution calculations.
[0092] Fused-Mobile inverted bottleneck convolution (Fused-MBConv) module (see appendix Figure 1 ): This module consists of a 3*3 upsampling convolution and a 1*1 downsampling convolution, and the feature input and feature output of the Fused-MBConv module are skip-connected. It can be applied in the shallow network to improve the training speed of the network, and the overhead of the number of parameters and FLOPs is very small.
[0093] Dual Attention module (see appendix Figure 2 ): This module consists of a window multi-head self-attention mechanism in the spatial dimension and a self-attention mechanism for the channel group. It processes problems from an orthogonal perspective, performs self-attention mechanisms separately from the spatial dimension and the channel dimension, and improves the accuracy of the automatic landslide recognition network model through this method.
[0094] Spatial Window Multihead Self-attention module (see appendixFigure 2 ):The spatial dimension part in the dual attention mechanism module uses spatial dimension information to improve local features.
[0095] Channel Group Self-attention module (see Appendix Figure 2 ):The channel dimension part in the dual attention mechanism module uses channel dimension information to capture global connections and features.
[0096] PDC module (see Appendix Figure 1 ):This module consists of a 1*1 convolution, a PatchEmded module, a dual attention mechanism, a feature transformation operation, and a 1*1 convolution, and performs a skip connection between the input feature map and the final output feature map to adjust the feature map to adapt to the input of the dual attention mechanism module and adjust the output features of the dual attention mechanism module to adapt to the input of the convolutional neural network layer.
[0097] Embodiment
[0098] Two lightweight network models, MobileNetV3 and ShuffleNetV2, the relatively mainstream EfficientNetV2 network model in recent years, and the classic convolutional neural network models ResNet101 and DenseNet169 in recent years, as well as two popular transformer models, Vision transformer and Swin transformer, in the past two years, are used as comparative experiments for the present invention. The training set and test set in step 1.6 are used as the data sets for the training and prediction comparative experiments. The network models are modified into the network models for the comparative experiments, and the above steps 3, 4, 5, and 6 are repeated to obtain the model training results of the comparative experiments, and the advantages and disadvantages of the models are evaluated.
[0099] The final experimental performance evaluation results are shown in Table 1, and the best evaluation values in terms of Accuracy, F1, Recall, and Precision are marked in bold. The final lightweight evaluation results are shown in Table 2. It can be seen from Table 1 that the Our networkframework proposed in the present invention is better than other experimental results in terms of the evaluation indexes of Accuracy, F1, and Precision.
[0100] As can be seen from Table 2, compared with the lightweight network models, the Our network framework network model proposed in the present invention is slightly higher in terms of the number of parameters, computational complexity, inference time, and training time. However, multiple evaluation indicators for evaluating network performance are better than those of MobileNetV3 and ShuffleNetV2. Especially in terms of F1, it is 6.98% higher than MobileNetV3 and 5.53% higher than ShuffleNetV2. Moreover, compared with other classic convolutional neural network models and Transformer models, it is much lower in terms of Params, FLOPs, inference time, and running time, and has better results in terms of Accuracy, F1, and Precision evaluation indicators. The results prove the effectiveness of the Our network framework network model proposed in the present invention in landslide recognition.
[0101] Table 1 Evaluation indicators of MobileNetV3, ShuffleNetV2, ResNet101, DenseNet169, Vision transformer, Swin transformer, EfficientNetV2, and Our network framework on the Ningxia landslide dataset
[0102]
[0103] Table 2 Number of parameters FLOPs of MobileNetV3, ShuffleNetV2, ResNet101, DenseNet169, Vision transformer, Swin transformer, EfficientNetV2, and Our network framework network models, as well as inference time and running time on the Bijie landslide dataset
[0104]
[0105] In the present invention, as shown by the experimental results (Table 2), this method is superior to other methods in terms of performance, model complexity, number of model parameters, and running time. Compared with the lightweight MobileNetV3 and ShuffleNetV2, although this method is higher in terms of inference time, number of parameters, and model complexity, its performance is more excellent. Moreover, compared with other classic networks and the newly proposed Transformer in recent years, its performance is better, and the inference time, number of parameters, and model complexity are lower. This method is not only effective on the dataset of landslides that have occurred, but also effective on the dataset of potential landslide points.
[0106] Compared with the existing EfficientNetV2 network model, the present invention has improved in terms of F1, accuracy, precision, and recall in landslide automatic recognition. At the same time, the model complexity, the number of parameters, the inference time, and the training time have all decreased. As shown in Tables 1 and 2, based on EfficientNetV2, the present invention has improved the training speed of the network and the accuracy of the network.
[0107] The present invention also provides a landslide automatic recognition device based on a lightweight convolutional neural network and dual attention, including:
[0108] A memory: used to store computer programs;
[0109] A processor, which is used to implement the landslide automatic recognition method based on a lightweight convolutional neural network and dual attention when executing the computer program.
[0110] The present invention also provides a computer-readable storage medium, including:
[0111] The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement a landslide automatic recognition method based on a lightweight convolutional neural network and dual attention.
Claims
1. An automatic landslide recognition method based on a lightweight convolutional neural network and dual attention, characterized in that: It includes the following steps: Step 1: Preprocess the remote sensing image to obtain a landslide remote sensing image dataset; Step 2: Build the Our network framework network model; The specific method of step 2 is: Step 2.1: Use the deep learning framework to construct the Fused-MBConv module, MBConv module, and PDC module respectively; The Fused-MBConv module includes a 3*3 upsampling convolution and a 1*1 downsampling convolution, and performs a skip connection between the feature input and feature output of the Fused-MBConv module; the MBConv module includes a 1*1 upsampling convolution, a 3*3 depthwise separable convolution, an SE attention mechanism, and a 1*1 downsampling convolution, and performs a skip connection between the feature input and feature output of the MBConv module; the PDC module consists of a 1*1 convolution, a PatchEmded module, a dual attention mechanism, a feature transformation operation, and a 1*1 convolution, and performs a shortcut link between the input feature map and the final output feature map, uses the PatchEmbed module to adjust the size of the feature map channels, and flattens the feature map into a feature vector suitable for the DualAttention Block; Step 2.2: Fuse the Fused-MBConv module, MBConv module, and PDC module constructed in step 2.1; First, construct a 1*1 standard convolutional layer, send its output feature map into the Fused-MBConv module with a stride of 1, and this time do not perform upsampling. The Fused-MBConv module loops twice, and sends its output feature into the Fused-MBConv module with a stride of 2 for upsampling by 4 times the number of input feature channels, loops 4 times, sends its output feature into the Fused-MBConv module with a stride of 2 for upsampling by 4 times the number of input feature channels, sends its output feature into the MBConv module with a stride of 2 for upsampling by 4 times the number of input feature channels, loops 2 times, sends its output feature into the PDC module with a stride of 1, loops 2 times, sends its output feature into the MBConv module with a stride of 2 for upsampling by 4 times the number of input feature channels, loops 3 times, sends its output feature into the PDC module with a stride of 1, loops 3 times, and finally sends it into a 1*1 convolution for pooling operation and full connection operation to obtain a binary classification result and complete the construction of the network model; Step 3: Iteratively train the Ournetwork framework network model and finally save the weights of the round with the best training effect; Step 4: Predict the network model trained in step 3, and use the accuracy Accuracy, balanced F-score F1, recall Recall, and precision Precision as evaluation indicators to evaluate the network performance and evaluate the quality of the network. Step 5: Randomly generate a tensor using a deep learning framework and feed it into the Our network framework network model obtained in Step 2 to calculate FLOPs, the number of parameters, and the model inference time. The FLOPs represents a measure of the complexity of the model.
2. The automatic landslide recognition method based on a lightweight convolutional neural network and dual attention according to claim 1, characterized in that: The specific method of Step 1 is: Step 1.1: Obtain positive samples of the landslide remote sensing image dataset: Crop the remote sensing image according to the landslide scale and expand the pixel points outward as the background to obtain positive samples; Step 1.2: Obtain negative samples of the landslide remote sensing image dataset: Crop the remote sensing image in sizes of 128 or 256 or 512, delete all data containing positive samples, and select special scenarios as negative samples; Step 1.3: Divide the positive samples obtained in Step 1.1 and the negative samples obtained in Step 1.2 into a training set and a test set according to a certain ratio; Step 1.4: Divide the training set obtained in Step 1.3 into a new training set and a validation set according to a certain ratio; Step 1.5: Perform data augmentation on the new training set obtained in Step 1.4; Step 1.6: Adjust the data of the test set in Step 1.3, the validation set in Step 1.4, and the training set after data augmentation in Step 1.5 to a unified size of 224*224 or 300*300.
3. The automatic landslide recognition method based on a lightweight convolutional neural network and dual attention according to claim 1, characterized in that: The specific method of Step 3 is: Feed the training set obtained in Step 1.6 into the Our network framework network model built in Step 2, perform iterative training on the Our network framework network model, set parameters including the number of iterative rounds, batch size, and learning rate, conduct network training and record the training time. Then, feed the validation set in Step 1.6 into the Our network framework network model built in Step 2 to verify the model effect. Finally, save the weights of the round with the best training effect.
4. The automatic landslide recognition method based on a lightweight convolutional neural network and dual attention according to claim 1, characterized in that: The specific method of Step 4 is: Feed the test set in Step 1.6 into the network model trained in Step 3 for prediction and evaluate the quality of the network. Use accuracy (Accuracy), balanced F-score (F1), recall (Recall), and precision (Precision) as evaluation indicators for judging the network performance. Accuracy is the proportion of correctly classified samples to the total number of samples; Recall represents the probability that a positive sample is predicted as a positive sample among the actual positive samples; Precision represents the probability that a sample is actually a positive sample among all samples predicted as positive samples. Precision and Recall are a pair of conflicting metrics. When Precision is high, Recall is low; when Recall is high, Precision is low. When it is necessary to balance Precision and Recall, the balanced F-score F1 is used as the weighted harmonic mean of Precision and Recall. The following are the calculation formulas for each evaluation metric: Among them, TP represents the number of correctly predicted positive samples, FP represents the number of incorrectly predicted negative samples, TN represents the number of correctly predicted negative samples, and FN represents the number of incorrectly predicted positive samples; Among them, P refers to Precision, R refers to Recall, and F 1 refers to the balanced F-score; The dimensional size of the tensor described in step 5 is the dimension and size of the image after uniformly adjusting the dataset size in step 1.
5. A method for automatically identifying landslides based on a lightweight convolutional neural network and dual attention according to claim 2, characterized in that: The special scenarios in step 1.2 include: roads, cultivated land, lakes, vegetation, cloud cover, factories, ridges, and snow-covered scenarios; the training set and the test set in step 1.3 are divided according to a ratio of 6:4 or 8:2; the ratio range of the new training set and the validation set in step 1.4 is: 7:3 - 9:1; the data augmentation in step 1.5 includes: random cropping, random rotation, and adding noise.
6. A system for automatically identifying landslides based on a lightweight convolutional neural network and dual attention, characterized in that: including: Mobile flip bottleneck convolution module: This module consists of a 1*1 dimensionality increase convolution, a 3*3 depthwise separable convolution, an SE attention mechanism, and a 1*1 dimensionality reduction convolution. It makes a skip connection between the feature input and the feature output of the MBConv module. The inverted residual structure is used in this module, and at the same time, the depthwise separable convolution is used to reduce the number of parameters required for convolution calculations to achieve the purpose of lightweighting the network model; Fused mobile flip bottleneck convolution module: This module consists of a 3*3 dimensionality increase convolution and a 1*1 dimensionality reduction convolution, and makes a skip connection between the feature input and the feature output of the Fused-MBConv module. It is applied in the shallow network to improve the training speed of the network; Dual attention mechanism module: This module consists of a window multi-head self-attention mechanism in the spatial dimension and a self-attention mechanism in the channel group. It processes problems from orthogonal perspectives and performs self-attention mechanisms in the spatial dimension and the channel dimension respectively to improve the accuracy of the landslide automatic identification network model; Spatial window multi-head self-attention mechanism module: The spatial dimension part of the dual attention mechanism module uses spatial dimension information to improve local features; Channel group self-attention mechanism module: The channel dimension part of the dual attention mechanism module uses channel dimension information to capture global connections and features; PDC module: This module consists of a 1*1 convolution, a PatchEmded module, a dual attention mechanism, a feature transformation operation, and a 1*1 convolution, and performs a skip connection between the input feature map and the final output feature map to adjust the input feature map to adapt to the input of the dual attention mechanism module, and adjusts the output features of the dual attention mechanism module to adapt to the input of the convolutional neural network layer.
7. An automatic landslide recognition device based on a lightweight convolutional neural network and dual attention, characterized in that, it includes: A memory: used to store computer programs; A processor, which is used to implement an automatic landslide recognition method based on a lightweight convolutional neural network and dual attention according to any one of claims 1-5 when executing the computer program.
8. A computer-readable storage medium, characterized in that, it includes: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement an automatic landslide recognition method based on a lightweight convolutional neural network and dual attention according to any one of claims 1-5.