A method for automatically identifying landslides based on an edge-guided attention neural network

Through an attention neural network based on edge guidance, using edge maps and multiple loss functions to train models, the problem of under-learning edge features in landslide recognition is solved, and a higher-precision landslide recognition effect is achieved.

CN115761297BActive Publication Date: 2025-07-08ZHENGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211035893.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-27
Publication Date
2025-07-08
Estimated Expiration
2042-08-27

AI Technical Summary

Technical Problem

The prior art fails to fully learn the edge characteristics of landslides in landslide recognition, resulting in poor recognition effects, especially when using the U-shaped network, landslide details information is lost.

Method used

The attention neural network based on edge guidance is adopted, and the edge diagram is extracted through the Edge Boxes method and combined with the VGG16 network, pyramid pooling module, edge guidance module, adaptive weighted fusion module and attention guidance module are trained using the cross entropy loss function, Generalized Dice loss function and Sobel edge loss function to enhance landslide feature learning.

Benefits of technology

The accuracy of landslide recognition and the accuracy of edge features are improved, and pixel-level landslide recognition is achieved with higher accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115761297B_ABST
    Figure CN115761297B_ABST
Patent Text Reader

Abstract

The present invention provides a method for automatically identifying landslides based on an edge-guided attention neural network. The Edge Boxes method is used to detect the edges of landslide remote sensing images to obtain an edge map. The edge map and the landslide dataset are combined to form a new landslide dataset, which is divided into a training set, a validation set, and a test set. The training set is subjected to data augmentation processing, and the images of the training set, the validation set, and the test set are normalized. The normalized training set is used to train the edge-guided attention neural network, and the normalized validation set is used for validation. The network model with the best performance on the validation set is saved. The normalized test set is used to test on the saved network model to obtain the landslide identification results. This method obtains more image landslide edge feature information that is easily overlooked and improves the accuracy of landslides.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of geological disaster identification, and in particular to a method for automatically identifying landslides based on an edge-guided attention neural network. Background Technique

[0002] Natural disasters of various scales often occur in the world. As one of the most common natural geological disasters, landslides cause huge losses to the economies of various countries. Due to the characteristics of multiple occurrences, suddenness, group occurrence and great harm of landslide geological disasters, they seriously threaten the lives and property safety of the people and hinder the regional economic construction. Therefore, how to quickly and accurately obtain information such as the location of landslides and the affected areas, and rationally and scientifically guide rescue work to reduce the losses caused by landslide disasters is of great significance.

[0003] Obtaining landslides through manual on-site investigation is the most direct and accurate method, but it requires a large amount of manpower, material resources and time. In recent years, with the continuous development of remote sensing technology, the spatial resolution of the obtained remote sensing images has become higher and higher, and richer and more detailed ground object information can be obtained from the remote sensing images, so it has been widely applied to the interpretation of geological disasters. The methods used by researchers to identify landslides from remote sensing images are mainly visual interpretation, pixel-based, object-oriented and machine learning. Visual interpretation is a process in which professionals obtain information on specific target ground objects on remote sensing images by direct observation or with the help of auxiliary interpretation instruments. It has problems such as heavy tasks, low efficiency, high cost and non-intuitive information. The pixel-based landslide identification method uses the spectral and spatial information of individual pixels on remote sensing images to obtain the characteristics of landslides, and uses these characteristics to identify landslides on the image according to a certain method. However, the pixel-based method does not fully utilize the remote sensing image information and is prone to the "salt and pepper phenomenon", resulting in poor identification effects. The object-oriented method regards the adjacent area of pixels as an object and uses the spatial, texture and spectral information of the object to identify the object. However, the object-oriented method selects representative features manually, with low automation. At the same time, the selection of the segmentation scale needs to be determined by the trial-and-error method and is only applicable to specific areas, which limits the accuracy of landslide identification. Researchers use machine learning for landslide identification, which solves some problems existing in pixel-based and object-oriented landslide identification. However, the methods in machine learning are shallow learning networks and are difficult to express complex functions. Facing the PB-level growth of remote sensing image data, machine learning cannot learn useful features, resulting in a reduction in identification accuracy.

[0004] In recent years, deep learning has achieved a series of accomplishments in the field of natural image processing (such as image classification, object detection, image segmentation, etc.). Compared with traditional methods and machine learning methods, deep learning methods input data into a multi-layer neural network, where the inter-layer mapping relationship reduces the size of the data, automatically learns features through convolutional operations, extracts important features of the data, uses hierarchical features instead of manually identified features, and improves the accuracy of recognition and classification. This has led some scholars to apply deep learning methods to the field of landslide disaster detection.

[0005] Convolutional Neural Networks (CNNs), as a representative algorithm of deep learning, have been applied to landslide recognition research and achieved good accuracy. Yu et al. used CNNs and learned the landslide area and outlined the edge contours through discriminant information (region, boundary, and center) without human intervention, achieving high-precision automatic landslide detection. Shi et al. detected landslides based on the images before and after landslide occurrence, using a method that combines CNNs with change detection. The results showed that the accuracy of this method reached 85%. These scholars used deep learning methods to obtain the bounding box positions of landslides from remote sensing images. However, in many landslide-related studies, a detailed landslide inventory is required, such as the area of the landslide and the accurate landslide boundary, in order to carry out rescue and prevention work scientifically. Therefore, some research scholars have applied image semantic segmentation to landslide recognition to achieve pixel-level extraction of landslides. Bragagnolo et al. established a landslide database in Nepal based on Landsat8 images and used the U-Net deep learning model for landslide recognition, with improved accuracy compared to the previous study in the same area. Ji et al. evaluated the effects of pooling strategies, the design of convolutional blocks, the scaling ratio in the attention mechanism, and different positions of the attention mechanism on the model performance, and designed a 3D attention mechanism to enhance landslide information, suppress background information, improve the learning ability of CNNs, and greatly improve the accuracy of landslide recognition.

[0006] Although the above scholars' research has well extracted landslide information and improved the accuracy of landslide recognition, meeting the needs of users to a certain extent, there are still some limitations in the above scholars' research, which can be summarized into the following two aspects: The above scholars used various methods to improve the network model to enhance the accuracy of landslide recognition, but their networks did not fully learn the features of landslides, such as the boundary features of landslides. In traditional landslide recognition methods, the edge features of landslides are one of the important bases for identifying landslides. When using the U-shaped network for landslide recognition, although skip connections are used to restore spatial information, many landslide detail information is lost during upsampling, resulting in poor landslide recognition effects.

[0007] Therefore, a method for automatically identifying landslides based on an edge-guided attention neural network is provided to solve the above problems. Summary of the Invention

[0008] Aiming at the problems existing in the prior art, a method for automatically identifying landslides based on an edge-guided attention neural network provided by the present invention uses an edge map and designs different modules to enable the network to better learn landslide features and obtain a highly accurate landslide identification map.

[0009] The present invention adopts the following technical solutions:

[0010] A method for automatically identifying landslides based on an edge-guided attention network, characterized by comprising the following steps:

[0011] Step S1: Extract an edge map from the landslide remote sensing images in the landslide dataset by using the Edge Boxes method to obtain the edge map input to the network;

[0012] Step S2: Add the edge map extracted in Step S1 to the landslide dataset to form a new landslide dataset, and divide the new landslide dataset into a training set, a validation set, and a test set for network training, validation, and testing;

[0013] Step S3: Perform data augmentation on the training set divided in Step S2 to obtain the training set images after data augmentation, and then normalize the validation set, test set divided in Step S2, and the training set images after data augmentation in Step S3 to prevent the network from overfitting;

[0014] Step S4: Construct an edge-guided attention neural network to obtain a deep learning model for landslide identification;

[0015] Step S5: Use a loss function composed of a cross-entropy loss function, a Generalized Dice loss function, and a Sobel edge loss function to train the edge-guided attention neural network constructed in Step S4 to obtain a preliminarily trained landslide identification model;

[0016] Step S6: Use the training set normalized in Step S3 to train the edge-guided attention neural network trained in Step S5, and use the validation set normalized in Step S3 for validation, and save the network model with good performance on the validation set;

[0017] Step S7: Use the test set normalized in Step S3 to test on the network model saved in Step S6 to obtain the landslide identification result.

[0018] In the said Step S2, the dataset containing landslides is randomly divided into 10 parts and divided into a training set, a validation set, and a test set according to the ratio of 6:1:3.

[0019] The specific operation of step S3 is as follows:

[0020] Step S301: Read the RGB remote sensing image and the edge map of a certain area in the training set at the same time, perform data augmentation on the data such as horizontal and vertical flipping, 90-degree rotation, and optical distortion according to the probability, and add Gaussian or salt-and-pepper noise to complete the data augmentation of the training set;

[0021] Step S302: Use the bilinear interpolation method to normalize the pixel values of each channel of the images in the validation set, the test set, and the training set of the new landslide data set in step S2 to 0-1.

[0022] The specific operation of step S4 is as follows:

[0023] Step S401: Construct an edge-guided attention neural network including a VGG16 network, a pyramid pooling module, an edge guidance module, an adaptive weighted fusion module, and an attention guidance module, and process the landslide remote sensing image and the remote sensing edge map respectively to obtain a high-precision pixel-level landslide identification map;

[0024] Step S402: Construct the VGG16 network and the pyramid pooling module in step S401. The landslide remote sensing image goes through 5 different convolutional modules of the VGG16 network. Each convolutional module has a side output, and after the side output, there is a convolutional layer with a kernel size of , and the number of channels is 128. The pyramid pooling module uses pooling layer sizes of 1, 2, 4, and 8 to obtain feature maps with different resolutions, connects the 4 feature maps with different resolutions through a connection layer, and then passes through a convolutional layer with a kernel size of , and the number of channels is 128 as the input feature map of the encoder;

[0025] Step S403: Construct the edge guidance module in step S401. The edge guidance module obtains the edge feature map from the edge image through two-layer convolutional operations, and adds the extracted edge features to the network to assist the network in identifying landslides; in order to utilize the edge information, edge features are added in the shallow layer and edge features are added in the deep layer; at the same time, convolutional operations with different convolutional kernels and strides are used to obtain edge image feature maps with different receptive fields;

[0026] S404. Construct the adaptive weighted fusion module in step S401. The adaptive weighted fusion module adaptively updates the weights of different features by means of the network's learning to ensure the maximization of the contribution of the fused features. In the module, different fusion weights are assigned to the two feature maps to be fused for fusion. During the network training process, the network uses backpropagation to propagate the error information of the output layer back to all neurons to complete the update of the weights in the network. During the learning process of the network, the fusion weights are corrected and optimized so that they are fused in an appropriate proportion.

[0027] S405. Construct the attention guidance module in step S401. The key of the attention guidance module is to generate an attention feature that focuses on landslide pixels, so that the attention learned by the network is concentrated on landslide recognition. At the same time, low-level features are used to refine high-level features to obtain the next-level refined prediction map.

[0028] In step S5, the cross-entropy loss function is expressed as the following formula:

[0029]

[0030] where represents the label of the sample, and p represents the probability value of the predicted value belonging to the category;

[0031] The Generalized Dice loss function can be expressed as:

[0032]

[0033] where represents the true pixel category of class l at the nth position, while represents the corresponding predicted probability value, represents the weight of each category; The formula of

[0034] The Sobel loss function calculates its error:

[0035]

[0036] where for the th sample, is the true value, is the predicted value, and N is the number of samples; finally, the total loss function of the model is: .

[0037] Advantages of the present invention: Compared with the prior art, the present invention has at least the following advantages: By adding the edge map of remote sensing images to the network and using the edge guidance module and the adaptive weighted fusion module, the network can better learn the edge features of landslides, and the landslides identified by the network have better edges; the attention guidance module in the network can better refine the prediction map, enabling the landslides identified by the network to have better accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is the flowchart of the present invention;

[0039] Figure 2 is the block diagram of the network model of the present invention;

[0040] Figure 3 is the pyramid pooling module in the present invention;

[0041] Figure 4 is the edge guidance module in the present invention;

[0042] Figure 5 is the adaptive weighted fusion module in the present invention;

[0043] Figure 6 is the attention guidance module in the present invention;

[0044] Figure 7 is the remote sensing edge map in the present invention;

[0045] Figure 8 is the landslide recognition result map of different deep learning models of the present invention. 1 is the original image, 2 is the label map, 3 is Deeplab v3plus, 4 is SegNet, 5 is Unet, and 6 is the method of the present invention; the areas circled in white in the figure are the areas with large recognition errors of the model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following will describe the implementation scheme of the present invention in detail in combination with embodiments. However, those skilled in the art will

[0047] understand that the following embodiments are only used to illustrate the present invention and should not be regarded as limiting the scope of the present invention.

[0048] The following will describe the implementation manner of the present invention in detail with reference to the drawings.

[0049] Refer to Figure 1 , a method for automatic recognition of an attention neural network based on edge guidance of the present invention includes the following steps:

[0050] Step S1: Extract the edge map of the landslide remote sensing images in the landslide dataset by using the Edge Boxes method;

[0051] Step S2: Add the edge map to the landslide dataset to form a new landslide dataset, and divide it into a training set, a validation set, and a test set;

[0052] Step S3: Perform data augmentation on the training set divided in Step S2, and then normalize the validation set, test set divided in Step S2, and the training set images after data augmentation in Step S3. Data augmentation includes methods such as random horizontal and vertical flipping, random rotation, and optical distortion; The normalization formula is as follows:

[0053]

[0054] where R represents the RGB image or DEM image after normalization processing, I represents the RGB image before normalization processing, max(I) and min(I) represent the operations of taking the maximum value and the minimum value respectively.

[0055] Step S4: Construct an edge-guided attention neural network;

[0056] Specifically:

[0057] Step S401: Construct an edge-guided attention neural network that includes a VGG16 network, a pyramid pooling module, an edge guidance module, an adaptive weighted fusion module, and an attention guidance module. Process the landslide remote sensing image and the remote sensing edge map respectively to obtain a high-precision pixel-level landslide identification map, such as Figure 6 the remote sensing edge map in;

[0058] S402: As shown in Figure 2 and Figure 3 adopt the first five convolutional blocks of the VGG16 network as the backbone network for extracting landslide image features. Each convolutional module has a side output, followed by a convolutional layer with a kernel size of , and the number of channels is 128. The pyramid pooling module uses pooling layer sizes of 1, 2, 4, and 8 to obtain feature maps of different resolutions. Then, connect the 4 feature maps of different resolutions through a connection layer, and finally pass through a convolutional layer with a kernel size of , and the number of channels is 128 as the input feature map of the encoder;

[0059] Step S403: As shown in Figure 4 , the edge guidance module obtains an edge feature map from the edge image through two convolutional operations. The specific convolutional layer parameters are shown in Table 1. In order to make better use of the edge information, we add edge features not only in the shallow layer but also in the deep layer. There are a total of 5 edge guidance modules between the encoder and decoder of the network.

[0060] Table 1 Specific details of the edge convolution module

[0061] Number of layers Convolution kernel size Stride Number of channels Edge convolution block 1 2 3*3 3*3 1 1 128 Edge convolution block 2 2 3*3 3*3 2 2 128 Edge convolution block 3 2 3*3 3*3 2 2 128 Edge convolution block 4 2 3*3 5*5 2 4 128 Edge convolution block 5 2 5*5 5*5 4 4 128

[0062] Step S404. As shown in Figure 5 , after each side output of VGG16 and the edge guidance module, the adaptive weighted fusion module fuses the features of both. A total of 5 adaptive weighted fusion modules are used. The 1×1 convolutional channels in the adaptive fusion module are 128.

[0063] Step S405. As shown in Figure 6 , after each upsampling in the attention guidance module, the key of this module is to generate an attention feature that focuses on landslide pixels, enabling the attention learned by the network to concentrate on landslide recognition. At the same time, low-level features refine the high-level features to obtain the next-level refined prediction map. The high-level features are convolved by a 1×1 convolution with 128 channels and then multiplied element-wise with the attention map obtained by passing the low-level features through the Sigmoid function and the processed high-level features. After that, a 1×1 convolution operation with the same number of channels as the low-level features is performed to obtain the enhanced feature map . Finally, it is added to the low-level features to prevent gradient disappearance, obtaining the output feature map of the attention guidance module .

[0064] Step S5. Train the edge-guided attention neural network constructed in Step S4 using a loss function composed of a cross-entropy loss function, a Generalized Dice loss function, and a Sobel edge loss function.

[0065] The cross-entropy loss function can be expressed as the following formula:

[0066]

[0067] where y represents the label of the sample, and p represents the probability value of the predicted value belonging to the class.

[0068] The Generalized Dice loss function can be expressed as:

[0069]

[0070] where represents the true pixel class of class l at the nth position, while represents the corresponding predicted probability value, and represents the weight of each class. The formula for

[0071] The Sobel loss function calculates its error:

[0072]

[0073] Among them, for the th sample, is the true value, is the predicted value, and N is the number of samples. Finally, the total loss function of the model is: ;

[0074] Step S6: Use the training set normalized in step S3 to train the edge-guided attention neural network trained in step S5, and use the validation set normalized in step S3 for validation, and save the network model with the best performance on the validation set;

[0075] Step S7: Use the test set normalized in step S3 to test on the network model saved in step S6 to obtain the landslide recognition result.

[0076] Simulation experiment

[0077] The effect of the present invention can be further illustrated by the following specific example:

[0078] 1. The experimental area of the present invention is located in Tianshui City, Gansu Province, China. The data used in this experiment is the landslide dataset of Tianshui City, Gansu Province made by Qi et al., as Figure 8 shown. This dataset mainly includes high-resolution landslide images and corresponding landslide mask shapes. There are 1443 landslide images in total. When training, the dataset is divided into a training set and a test set according to a ratio of 8:2.

[0079] The network model of the present invention is implemented based on the TensorFlow 2.2 deep learning network framework, and the mini-batch stochastic gradient descent is used as the optimizer. The initial learning rate is set to 0.001, and the exponential decay function in TensorFlow is adopted, and other default parameters are used. There are a total of 250 epochs. The minimum batch size is 4, the momentum is 0.9, and the weight decay is 5×10-5.

[0080] 2. Evaluation indicators

[0081] In order to evaluate the performance of the network model, the precision (P) and recall (R) are used to evaluate the model. The precision reflects the probability of being correct among the targets detected as positive samples, and the recall reflects the probability of being correctly identified among all positive samples. The calculation formulas are as follows:

[0082]

[0083] Wherein, TP represents the actual positive sample predicted as a positive sample; FP represents the actual negative sample predicted as a positive sample; FN represents the actual positive sample predicted as a negative sample; TN represents the actual negative sample predicted as a negative sample.

[0084] The F1 score is a weighted average of precision and recall. Its maximum value is 1 and its minimum value is 0. The larger the value, the better the model:

[0085]

[0086] The present invention identifies and segments the shape of landslides. Therefore, to reasonably evaluate the segmentation results, the mean intersection over union (MIoU) index is introduced. MIoU is the average of the ratios of the intersections to the unions between the prediction results of each class and the true masks. The calculation formula is as follows:

[0087]

[0088] Wherein, n is the number of predicted classes; represents the original class predicted as class i; represents the original class i predicted as class j; represents the original class j predicted as class i.

[0089] 3. Landslide identification results

[0090] Table 2 Indexes for evaluating landslide identification

[0091] P (%) R (%) MIoU(%) F1(%) Deeplabv3plus 84.9 84.1 75.7 84.5 SegNet 83.7 76.1 72.8 79.7 UNet 86.5 88.5 80.5 87.5 EGANet <![CDATA 95.7 > <![CDATA 95.8 > <![CDATA 92.1 > <![CDATA 95.7 >

[0092] Table 2 shows the evaluation index results of the EGANet proposed by the present invention and other landslide identification models. It can be seen from it that the EGANet model proposed by the present invention has the best results in terms of accuracy, precision, recall, F1 score and MIoU. The F1 score and MIoU of EGANet are 8.2% and 11.6% higher than those of Unet.

[0093] To have a more intuitive understanding of the performance of the model, Figure 8 shows the result diagram of predicting the selected validation image. Figure 8 In it, the first column is the original image, and the second column is the true landslide image, which provides a reference for the accuracy evaluation of extracting landslides by different methods. Figure 8Columns 3 - 5 are the extraction results of Deeplab v3plus, SegNet, and Unet respectively, and the 6th column is the result graph of the EGANet model of the present invention. It can be seen from the figure that there are many mis - extractions and missed extractions in the landslide recognition results of the compared models, the recognition accuracy is not ideal, and the recognized landslide boundaries are relatively smooth without detecting the complete boundaries. While the EGANet proposed by us has a better effect in recognizing landslides, and the extracted landslide boundaries are also relatively complete.

[0094] When recognizing mountain roads or bare plots similar to landslides, the compared models will have problems of mis - recognition (such as a3 - a5, b3 - b5, e3 - e5 in the attachment Figure 8 . The main reason is that the models do not fully learn the features of landslides. Our model adds PPM after VGG - 16, which enables the network to fully learn the high - level semantic information of landslides. At the same time, when down - sampling, the AGM module refines the landslide information with the help of low - level features from the encoder part, so that the EGANet model can well distinguish roads and buildings similar to landslides (such as b6, e6 in Figure 6 Figure 8), making the extraction results of landslides have a higher accuracy. For some small landslides, due to their small area and few pixels, some compared models fail to completely recognize the landslides ( Figure 6 a3 - a5, b3 - b5, c3 - c5, g3 - g5 in Figure 8), and it can be seen that some small landslides are not recognized. In Figure 6 a6, b6, g6 in Figure 8, there are also small landslides that the model fails to recognize, and the recognition effect of small landslides in c6 is better than that of the compared models. This is because PPM is added to extract landslide features at multiple scales, enabling the encoder to obtain more information. For larger features in the image that are easily confused with landslides, there is a small gap in the recognition of the position and shape of landslides between the compared models and the EGANet model, but there is a large gap in the boundaries. It can be seen from Figure 8 rows b, d, e, f in Figure 8 that the boundary of the landslide recognition result of EGANet has a very small gap with the real boundary. This is because an edge - guiding module, an edge - detection loss function are added in EGANet, and the AGM uses low - level features to refine the landslide features during down - sampling, making the recognized landslides have better shape boundaries.

Claims

1. A method for automatically identifying landslides based on an edge-guided attention network, characterized in that, It includes the following steps: Step S1: Use the Edge Boxes method to extract the edge map from the landslide remote sensing images in the landslide dataset to obtain the edge map for network input; Step S2: Add the edge map extracted in Step S1 to the landslide dataset to form a new landslide dataset, and divide the new landslide dataset into a training set, a validation set, and a test set for network training, validation, and testing; Step S3: Perform data augmentation on the training set divided in Step S2 to obtain the training set images after data augmentation, and then normalize the validation set, test set divided in Step S2, and the training set images after data augmentation in Step S3 to prevent the network from overfitting; Step S4: Construct an edge-guided attention neural network to obtain a deep learning model for landslide recognition; Step S5: Use a loss function composed of a cross-entropy loss function, a Generalized Dice loss function, and a Sobel edge loss function to train the edge-guided attention neural network constructed in Step S4 to obtain a preliminarily trained landslide recognition model; Step S6: Use the training set normalized in Step S3 to train the edge-guided attention neural network trained in Step S5, and use the validation set normalized in Step S3 for validation, and save the network model with good performance on the validation set; Step S7: Use the test set normalized in Step S3 to test on the network model saved in Step S6 to obtain the landslide recognition result.

2. The method for automatically identifying landslides based on an edge-guided attention neural network according to claim 1, characterized in that In Step S2, the dataset containing landslides is randomly divided into 10 parts and divided into a training set, a validation set, and a test set according to the ratio of 6:1:

3.

3. The method for automatically identifying landslides based on an edge-guided attention neural network according to claim 1, characterized in that The specific operation of Step S3 is as follows: Step S301: Read the RGB remote sensing image and the edge map of a certain area in the training set at the same time, perform data augmentation on the data such as horizontal and vertical flipping, 90-degree rotation, and optical distortion according to the probability, and add Gaussian or salt-and-pepper noise to complete the data augmentation of the training set; Step S302: Use the bilinear interpolation method to normalize the pixel values of each channel of the images in the validation set, test set, and the training set of the new landslide dataset in Step S2 to 0-1.

4. The method for automatically identifying landslides based on an edge-guided attention neural network according to claim 1, characterized in that The specific operation of Step S4 is as follows: Step S401: Construct an edge-guided attention neural network including a VGG16 network, a pyramid pooling module, an edge guidance module, an adaptive weighted fusion module, and an attention guidance module, and process the landslide remote sensing image and the remote sensing edge map respectively to obtain a high-precision pixel-level landslide recognition map; Step S402: Construct the VGG16 network and the pyramid pooling module in Step S401. Apply the 5 different convolutional modules of the VGG16 network to the landslide remote sensing image. Each convolutional module has a side output, followed by a convolutional layer with a kernel size of , and a channel number of 128. The pyramid pooling module uses pooling layer sizes of 1, 2, 4, and 8 to obtain feature maps of different resolutions. Connect the 4 feature maps of different resolutions through a connection layer, and then pass through a convolutional layer with a kernel size of , and a channel number of 128 as the input feature map of the encoder; Step S403: Construct the edge guidance module in Step S401. The edge guidance module obtains the edge feature map from the edge image through two-layer convolution operations, and adds the extracted edge features to the network to assist the network in landslide recognition; in order to utilize the edge information, edge features are added in the shallow layer and edge features are added in the deep layer; at the same time, convolution operations with different convolution kernels and strides are used to obtain edge image feature maps with different receptive fields; S404. Construct the adaptive weighted fusion module in step S401. The adaptive weighted fusion module adaptively updates the weights of different features by means of the network's learning to ensure the maximization of the contribution of the fused features. In the module, different fusion weights are assigned to the two feature maps to be fused for fusion. During the network training process, the network uses backpropagation to propagate the error information of the output layer back to all neurons to complete the update of the weights in the network. During the learning process, the network corrects and optimizes the fusion weights to make them fuse in an appropriate proportion. S405. Construct the attention guidance module in step S401. The key of the attention guidance module is to generate an attention feature that focuses on landslide pixels, enabling the attention learned by the network to concentrate on landslide recognition. At the same time, low-level features are used to refine high-level features to obtain the next-level refined prediction map.

5. The landslide identification method based on attention mechanism and multi-modal representation learning according to claim 1, characterized in that In step S5, the cross-entropy loss function is expressed as the following formula: where represents the label of the sample, and p represents the probability value of the category to which the predicted value belongs; The Generalized Dice loss function can be expressed as: where represents the true pixel category of class l at the nth position, and represents the corresponding predicted probability value, represents the weight for each class; The formula for The Sobel loss function calculates its error: Among them, for the th sample, is the true value, is the predicted value, and N is the number of samples; finally, the total loss function of the model is: .

Citation Information

Patent Citations

  • Remote sensing image landslide automatic detection method based on three-dimensional space-channel attention mechanism

    CN111222466A

  • Landslide identification method and device based on multi-model decision level fusion

    CN114463643A