Ancient landslide detection method
Through the coordination of the mask reconstruction task of global context feature extraction and local key feature extraction and the contrast learning task, combined with the self-distillation learning architecture, the visual blur and small sample problems in ancient landslide recognition are solved, and the accuracy of landslide edge segmentation and the reliability of the model are improved.
Patent Information
- Application Number
- CN202510219485.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-27
AI Technical Summary
The high-resolution remote sensing satellite image data of ancient landslides have blurred sample visuals and small sample data sets, which leads to extremely poor accuracy of the existing methods on the edge of the landslide and are prone to missed judgments.
The coordination of mask reconstruction tasks of global context feature extraction, local key feature extraction and comparison learning tasks is adopted, and multi-scale and multi-level semantic features are fused. Through self-distillation learning architecture and multi-task collaboration framework, sample utilization and semantic feature extraction efficiency are improved.
It effectively solves the overfitting problem caused by visual blur and small sample problems in ancient landslide identification, and improves the accuracy of landslide edge segmentation and the reliability of the model.
Smart Images

Figure CN120047843A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of landslide detection, and particularly relates to a method for detecting ancient landslides. Background Art
[0002] Landslides are frequent and highly hazardous natural disasters, and are one of the most serious natural disaster forms in China. Approximately 70% of China's territory is mountainous, with complex geological conditions and frequent geological disasters, which usually result in disasters such as traffic interruption, river channel blockage, farmland damage, factory and mine destruction, village burial, and human and livestock being crushed to death, bringing huge personal and economic losses. Studying the geological and geomorphic characteristics of landslides and their inducing factors, and using information technology to accurately detect and warn of landslides are of great significance for disaster reduction and prevention. Traditional landslide identification is obtained through manual field observation and mapping, which is time-consuming and laborious, and it is difficult to meet the requirements of high efficiency and automation for large-scale general surveys. With the rapid development of remote sensing technology, due to its high precision, diverse sources, and wide coverage, it has been widely used in landslide detection. Therefore, using remote sensing data to achieve landslide detection has become the mainstream technology.
[0003] Due to the limitations of high-resolution remote sensing satellite image data of ancient landslides, the technology for automatically identifying ancient landslides from high-resolution remote sensing satellite images mainly faces two challenges: the problem of sample visual blurriness and the problem of small sample datasets.
[0004] 1. Problem of sample visual blurriness: Ancient landslides have been formed for a long time and have experienced long-term natural environmental evolution and human activities. Their surface morphology is highly similar to the surrounding environment, resulting in blurred morphological features; Remote sensing interpretation experts point out that the height change law of the landslide boundary is the key to identifying landslides. However, high-resolution remote sensing satellite image data lacks height information, and the texture features generated by projecting the height change of the landslide boundary onto a 2D image are extremely similar to the shadows in the background, resulting in blurred semantic features. The visual blurriness caused by both makes the existing methods usually have extremely poor accuracy in landslide edge segmentation and are prone to a large number of misjudgments and missed detections.
[0005] 2. Problem of small sample datasets: It is extremely difficult to construct a dataset of ancient landslides because of the high technical difficulty of accurately identifying landslides and the high time cost required for drawing landslide boundaries. In particular, the accurate annotation of the landslide edge morphology usually requires satellite remote sensing interpretation experts to invest a large amount of time in manual delineation work, so there is a lack of high-quality and large-scale public datasets. When the sample data volume is small and the types are not rich enough, it is difficult to fully train deep learning models, resulting in problems such as underfitting and overfitting of the models, which severely restricts the performance of existing landslide identification methods.
[0006] The disclosed patent CN201510041864.2, a multi-index fusion landslide detection method based on high-resolution remote sensing images, uses high-resolution remote sensing image data for multi-scale image segmentation to obtain ground feature indicators, and obtains terrain feature indicators through a digital elevation model established by a stereo image pair. The terrain and ground features are fused and processed, and each feature indicator is compared with a set rule set to achieve the detection of landslides.
[0007] The disclosed patent CN202111496192.6, a landslide recognition method and system based on attention mechanism and multi-modal representation learning, proposes a method that depends on the attention mechanism and multi-modal strategy to fuse the features of high-resolution remote sensing image data and digital elevation data to achieve the detection of landslides.
[0008] The disclosed patent CN202211022212.0, a landslide recognition method and device based on a satellite data UNet network model, proposes a landslide recognition method and device based on a satellite data UNet network model. It improves the common ordinary convolution operation in the UNet network model to depthwise separable convolution. When constructing the depthwise separable convolution, multiple groups of depthwise separable convolution units connected in series by depth convolution and pointwise convolution are set, and an attention mechanism module CBAM is added between multiple groups of depthwise separable convolution units. At the same time, dilated convolution is introduced to increase the receptive field and extract multi-scale information.
[0009] The landslide detection methods disclosed in the above patents optimize in feature fusion, feature extraction, and network architecture, and improve the performance of the model through multi-modal data, multi-level feature extraction, or attention mechanism. However, the existing methods cannot simultaneously solve the problems of visual blur and small samples faced by the ancient landslide detection task. Summary of the Invention
[0010] Aiming at the deficiencies in the prior art, the present invention provides an ancient landslide detection method. Through the collaboration of global context feature extraction, local key feature extraction mask reconstruction tasks, and contrast learning tasks, it fuses multi-scale and multi-level semantic features, improves the sample utilization rate, and improves the efficiency and reliability of semantic feature extraction, thus effectively solving the overfitting problem caused by visual blur and small samples in ancient landslide recognition.
[0011] The present invention provides the following technical solutions:
[0012] An ancient landslide detection method, which uses a detection model to detect ancient landslides from landslide images. The landslide images are high-resolution remote sensing satellite images. The detection model includes a feature extraction module, a feature fusion module, and a decoder. The feature extraction module includes a semantic feature extraction sub-module and a local feature extraction sub-module;
[0013] The method includes:
[0014] Use the landslide images with annotations to train the detection model, continuously adjust the parameters of the detection model based on the loss function, and stop training when the performance of the detection model no longer improves or reaches the preset number of iterations, so as to obtain the trained detection model for detecting ancient landslides from landslide images;
[0015] Obtain the landslide image pair to be detected as the input image pair, and input the input image pair into the semantic feature extraction sub-module and the local feature extraction sub-module respectively;
[0016] The semantic feature extraction sub-module extracts the global context features of the ancient landslides in each image of the input image pair to obtain a semantic feature image pair;
[0017] The local feature extraction sub-module extracts the local features of the terrain height change at the edge of the ancient landslide in each image of the input image pair through the mask reconstruction task and the contrast learning task to obtain a local feature image pair;
[0018] The feature fusion module fuses the semantic feature image pair and the local feature image pair to obtain a fused feature image pair;
[0019] The decoder restores the resolution of the fused feature image pair to the original image pair resolution to obtain the ancient landslide detection result of the input image pair.
[0020] As a further improvement of the present invention, the local feature extraction sub-module includes an image division unit, a mask unit, a self-distillation teacher-student network, a mask reconstruction unit, a cross-contrast learning unit, and a feature fusion unit; the input image pair consists of a first input image and a second input image, and the first input image and the second input image are respectively collected from the same geographical area at different time points;
[0021] The steps for the local feature extraction sub-module to extract the second feature image pair include:
[0022] The image division unit divides the pixel blocks on the first input image and the second input image respectively to obtain a first divided image and a second divided image; wherein, each pixel block on the first divided image and the second divided image corresponds to a feature point on the feature map;
[0023] The mask unit performs mask processing on the first divided image and the second divided image respectively to obtain a first mask image and a second mask image;
[0024] The teacher network extracts the local features of the first input image and the second input image respectively to obtain a first original feature map and a second original feature map;
[0025] The student network extracts local features of the first masked image and the second masked image respectively to obtain a first masked feature map and a second masked feature map;
[0026] The masked reconstruction unit reconstructs the masked part of the first masked feature map based on the first original feature map and the first masked feature map to obtain a first masked reconstruction feature map; the masked reconstruction unit reconstructs the masked part of the second masked feature map based on the second original feature map and the second masked feature map to obtain a second masked reconstruction feature map;
[0027] The cross-contrast learning unit performs cross-contrast learning on the feature points on the first masked feature map and the second original feature map to obtain a first cross-contrast feature map; performs cross-contrast learning on the feature points on the second masked feature map and the first original feature map to obtain a second cross-contrast feature map;
[0028] The feature fusion unit fuses the first masked reconstruction feature map and the first cross-contrast feature map to obtain a local feature map of the first input image; the feature fusion unit fuses the first masked reconstruction feature map and the first cross-contrast feature map to obtain a local feature map of the second input image; the local feature image pair is composed of the local feature map of the first input image and the local feature of the second input image.
[0029] As a further improvement of the present invention, the image division unit divides the pixel blocks into non-landslide blocks, landslide edge blocks, and landslide internal blocks based on the proportion of landslide pixels contained in each pixel block;
[0030] The division steps include:
[0031] Set a first division ratio and a second division ratio for dividing the pixel blocks;
[0032] Divide the pixel blocks containing fewer landslide pixels than the first division ratio into non-landslide blocks, divide the pixel blocks containing landslide pixels greater than or equal to the first division ratio and less than or equal to the second division ratio into landslide edge blocks, and divide the pixel blocks containing landslide pixels greater than the second division ratio into landslide internal blocks.
[0033] As a further improvement of the present invention, the masking unit randomly selects a certain number of non-landslide blocks and landslide edge blocks on the first divided image and the second divided image for masking processing to obtain a first masked image and a second masked image;
[0034] Record the positions corresponding to the masked non-landslide blocks and landslide edge blocks on the first masked image and the second masked image on the first input image and the second input image respectively.
[0035] As a further improvement of the present invention, the number of non-landslide blocks and landslide edge blocks randomly selected by the mask unit on the first divided image is equal, and the number of non-landslide blocks and landslide edge blocks randomly selected by the mask unit on the second divided image is equal.
[0036] As a further improvement of the present invention, the steps for the mask reconstruction unit to perform mask reconstruction include:
[0037] Based on the positions of the masked non-landslide blocks and landslide edge blocks on the first mask image corresponding to the first input image, the mask reconstruction unit reconstructs the masked feature points on the first mask feature map with the feature points corresponding to other non-landslide blocks and landslide edge blocks on the first original feature map corresponding to the first input image, so as to obtain the first mask reconstructed feature map;
[0038] The steps for the mask reconstruction unit to obtain the second mask reconstructed feature map are the same as those for obtaining the first mask reconstructed feature map.
[0039] As a further improvement of the present invention, when training the model, the mean square error loss function is used to calculate the mask reconstruction loss of the input image pair, including:
[0040] Based on the positions of the masked non-landslide blocks and landslide edge blocks on the mask image corresponding to the corresponding input image, the feature points on the mask reconstructed feature map are obtained to construct the reconstructed feature, and the feature points on the original feature map are obtained to construct the label feature;
[0041] The mean square error loss function is used to calculate the loss between the reconstructed feature and the label feature, so as to obtain the mask reconstruction loss of the mask reconstruction unit. The calculation formula is:
[0042]
[0043] L Reconst =L MSE_1 +L MSE_2
[0044] where N is the number of labeled landslide image pairs participating in the training, f' re 、f' l are the reconstructed feature and the label feature corresponding to each landslide image respectively, L MSE is the loss between the reconstructed feature and the label feature of each landslide image, L Reconst is the mask reconstruction loss of the input image pair, L MSE_1 is the loss between the reconstructed feature and the label feature corresponding to the first input image, L MSE_2 is the loss between the reconstructed feature and the label feature corresponding to the second input image.
[0045] As a further improvement of the present invention, when the model is trained, the cross-contrast learning loss of the input image pair is calculated using a supervised contrastive learning loss function, including:
[0046] Mark the feature points corresponding to all landslide edge blocks on the original feature map and the mask feature map as class 1 to form a class 1 set Mark the feature points corresponding to all non-landslide blocks on the original feature map and the mask feature map as class 0 to form a class 0 set Both the class 1 set and the class 0 set contain M feature points;
[0047] Use the supervised contrastive learning loss function to calculate the loss between the class 1 set and the class 0 set corresponding to the first mask feature map and the second original feature map, and the loss between the class 1 set and the class 0 set corresponding to the second mask feature map and the first original feature map, so as to obtain the cross-contrast loss of the input image pair. The calculation formula is:
[0048]
[0049] L Cons =L cons_1 +L cons_2
[0050] Where L cons is the loss between the class 1 set and the class 0 set; is the loss of the feature vector of the i-th feature point in the class 1 set; I [i≠j] is an indicator function. When i≠j, I [i≠j] =1; x i , x j are the feature vectors of the i-th and j-th feature points in the class 1 set respectively; T is a temperature hyperparameter; is the feature vector of the k-th feature point in the class 0 set; L Cons is the cross-contrast loss of the input image pair; L cons_1 is the loss between the class 1 set and the class 0 set corresponding to the first mask feature map and the second original feature map; L cons_2 is the loss between the class 1 set and the class 0 set corresponding to the second mask feature map and the first original feature map.
[0051] As a further improvement of the present invention, the parameters of the student network are updated using stochastic gradient descent and backpropagation;
[0052] The parameters of the teacher network are constructed by a momentum editor based on the parameters of the student network. The parameter update strategy of the teacher network is:
[0053] θ t =λθ t +(1 - λ)θ s
[0054] Where θt The parameters of the teacher network, θ s The parameters of the student network, and λ is a hyperparameter.
[0055] As a further improvement of the present invention, the decoder stacks two transposed convolutional layers to restore the resolution of the original image pair for the fused feature image pair;
[0056] The decoder further includes: a dropout layer for avoiding overfitting of the model, a batch normalization layer for restricting the data fluctuation range, and a ReLU activation function for increasing the sparsity of the model and avoiding the vanishing of the model gradient.
[0057] Compared with the prior art, the beneficial effects of the present invention are:
[0058] The present invention designs a multi-task collaborative framework based on self-distillation, uses a self-distillation learning architecture, and performs multi-task collaboration of reconstruction learning and contrast learning. Taking two different images as input pairs, the original input pair and the masked input pair are respectively input into the teacher network and the student network; performing a masked reconstruction task on the features of the same image extracted by the teacher and student networks, and using a momentum encoder to update the parameters of the teacher network from the average value of the student network to suppress the interference of the background environment and extract the target general semantic features, thereby suppressing overfitting; performing a cross-contrast learning task on the features of different images extracted by the teacher and student networks, constructing a positive sample space for contrast learning using the landslide boundary features in different pictures, constructing negative samples from the background, and increasing the diversity of contrast samples through three combinations of positive and negative samples, improving the utilization rate of samples, and accelerating the convergence of the model.
[0059] In the masked reconstruction task, targeted masked reconstruction is performed on the landslide edge blocks and non-landslide blocks of the landslide at the feature layer. By masking some landslide edge blocks and background blocks (non-landslide blocks), and using the features of other landslide edges and non-landslide parts in the original image to reconstruct the features of the masked area, the model can learn the features of the landslide edge and the background, thereby enhancing the feature extraction ability of the model.
[0060] Performing supervised contrast learning on the features of the landslide edge and the background. Different from constructing target-level sample pairs in traditional contrast learning, the present invention constructs positive samples from the landslide boundary and negative samples from the environmental background, directly focusing on the differences between similar features, and enhancing the model's ability to distinguish visually blurred features.
[0061] The present invention designs a multi - task collaborative architecture based on dual - branch interaction, including an upper branch for performing semantic segmentation tasks and a lower branch for performing local significant feature enhancement tasks. The upper branch is responsible for extracting the global context features of ancient landslides, and the lower branch is responsible for extracting the edge - key height change features of ancient landslides. By using the collaboration of the dual - branch tasks, the semantic features of landslides are supplemented and enhanced, improving the accuracy of pixel - level classification of landslides in automatic detection. Different from traditional multi - task learning architectures, in the present invention, the parameters of the feature extractors of the semantic feature extraction sub - module and the local feature extraction sub - module are not shared, and the attention mechanism is discarded. An interactive fusion method based on task - goal guidance is designed, and through the interactive fusion enhancement of the features of multiple task branches, the accuracy of pixel - level classification of landslides in automatic detection is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is a flowchart of the method for detecting ancient landslides;
[0063] Figure 2 is a schematic structural diagram of the detection model;
[0064] Figure 3 is a schematic diagram of the image - processing flow of the self - distillation teacher - student network. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0066] The following further describes the present invention in detail with reference to the accompanying drawings:
[0067] This embodiment provides a method for detecting ancient landslides. The method uses a detection model to detect ancient landslides from landslide images. The landslide images are high - resolution remote - sensing satellite images. The detection model includes a feature extraction module, a feature fusion module, and a decoder. The feature extraction module includes a semantic feature extraction sub - module and a local feature extraction sub - module. The structure of the detection model is as Figure 2 shown.
[0068] Please refer to Figure 1 , the method includes:
[0069] Use the landslide images with annotations to train the detection model, continuously adjust the parameters of the detection model based on the loss function, and stop training when the performance of the detection model no longer improves or reaches the preset number of iterations, so as to obtain the trained detection model for detecting ancient landslides from landslide images;
[0070] Obtain the landslide image pairs to be detected as input image pairs, and input the input image pairs into the semantic feature extraction sub-module and the local feature extraction sub-module respectively;
[0071] The semantic feature extraction sub-module extracts the global context features of ancient landslides in each image of the input image pair to obtain a pair of semantic feature images;
[0072] The local feature extraction sub-module extracts the local features of the terrain height change at the edge of the ancient landslide in each image of the input image pair through the mask reconstruction task and the contrast learning task to obtain a pair of local feature images;
[0073] The feature fusion module fuses the semantic feature image pair and the local feature image pair to obtain a pair of fused feature images;
[0074] The decoder restores the resolution of the fused feature image pair to the resolution of the original image pair to obtain the detection result of the ancient landslide in the input image pair.
[0075] The feature extraction module of the detection model includes two branches (i.e., two sub-modules). The semantic feature extraction sub-module (i.e., the segmentation branch extraction) extracts the global context features of the ancient landslide in the input image pair to obtain the semantic features f 1 and f 2 , and the local feature extraction sub-module (i.e., the feature enhancement branch) extracts the local features (i.e., enhanced features) of the terrain height change at the edge of the ancient landslide in the input image pair to obtain the local features f ad_1 and f ad_2 .
[0076] The feature fusion module adds the semantic features f 1 and f 2 , and the local features f ad_1 and f ad_2 point by point to obtain a pair of fused feature images.
[0077] The decoder stacks two transposed convolutional layers to restore the resolution of the fused feature image pair to the resolution of the original image pair and tries to avoid information redundancy. The decoder also includes: a dropout layer for avoiding model overfitting, a batch normalization layer for restricting the data fluctuation range, and a ReLU activation function for increasing the sparsity of the model and avoiding model gradient disappearance. The fused fused feature image pair is input into the decoder to perform the landslide detection task.
[0078] The present invention designs a multi-task collaborative architecture based on dual-branch interaction, including an upper branch for performing semantic segmentation tasks and a lower branch for performing local significant feature enhancement tasks. The upper branch is responsible for extracting the global context features of ancient landslides, and the lower branch is responsible for extracting the key height change features of the edges of ancient landslides. The cooperation of the dual-branch tasks is used to supplement and enhance the semantic features of landslides, improving the accuracy of pixel-level classification of landslides in automatic detection. Different from traditional multi-task learning architectures, in the present invention, the parameters of the feature extractors of the semantic feature extraction sub-module and the local feature extraction sub-module are not shared, and the attention mechanism is discarded. An interactive fusion method guided by task objectives is designed, and through the interactive fusion enhancement of the features of multiple task branches, the accuracy of pixel-level classification of landslides in automatic detection is improved.
[0079] Further, the local feature extraction sub-module includes an image partitioning unit, a masking unit, a self-distillation teacher-student network, a mask reconstruction unit, a cross-contrast learning unit, and a feature fusion unit; the input image pair consists of a first input image and a second input image, and the first input image and the second input image are respectively collected from the same geographical area at different time points.
[0080] The image partitioning unit partitions the pixel blocks on the first input image and the second input image respectively based on the proportion of landslide pixels contained in each pixel block. The pixel blocks are divided into non-landslide blocks, landslide edge blocks, and landslide interior blocks, obtaining a first partitioned image and a second partitioned image; among them, each pixel block on the first partitioned image and the second partitioned image corresponds to a feature point on the feature map.
[0081] The partitioning steps include:
[0082] Set a first partitioning ratio and a second partitioning ratio for partitioning pixel blocks;
[0083] The pixel blocks containing fewer landslide pixels than the first partitioning ratio are divided into non-landslide blocks, the pixel blocks containing landslide pixels greater than or equal to the first partitioning ratio and less than or equal to the second partitioning ratio are divided into landslide edge blocks, and the pixel blocks containing landslide pixels greater than the second partitioning ratio are divided into landslide interior blocks;
[0084] In this embodiment, the first partitioning ratio is 10% and the second partitioning ratio is 90%.
[0085] The masking unit respectively performs masking processing on a certain number of non-landslide blocks and landslide edge blocks randomly selected on the first partitioned image and the second partitioned image, obtaining a first masked image and a second masked image; the number of non-landslide blocks and landslide edge blocks randomly selected by the masking unit on the first partitioned image is equal, and the number of non-landslide blocks and landslide edge blocks randomly selected by the masking unit on the second partitioned image is equal;
[0086] Record the positions corresponding to the masked non-landslide blocks and landslide edge blocks on the first masked image and the second masked image on the first input image and the second input image respectively.
[0087] The teacher network extracts the local features of the first input image and the second input image respectively to obtain the first original feature map and the second original feature map; the student network extracts the local features of the first masked image and the second masked image respectively to obtain the first masked feature map and the second masked feature map; the processing flow of the self-distillation teacher-student network is as Figure 3 shown.
[0088] The teacher network and the student network have the same structure but different network parameters.
[0089] The parameters of the student network are updated using stochastic gradient descent (SGD) and backpropagation; the parameters of the teacher network are constructed by the momentum editor based on the student network parameters;
[0090] The parameter update strategy of the teacher network is:
[0091] θ t = λθ t + (1 - λ)θ s
[0092] where θ t is the parameter of the teacher network, θ s is the parameter of the student network, and λ is a hyperparameter and is set to 0.996.
[0093] The masked reconstruction unit reconstructs the masked part of the first masked feature map based on the first original feature map and the first masked feature map to obtain the first masked reconstruction feature map; the masked reconstruction unit reconstructs the masked part of the second masked feature map based on the second original feature map and the second masked feature map to obtain the second masked reconstruction feature map.
[0094] The cross-contrast learning unit performs cross-contrast learning on the feature points on the first original feature map based on the first original feature map and the second masked feature map to obtain the first cross-contrast feature map; the cross-contrast learning unit performs cross-contrast learning on the feature points on the second original feature map based on the second original feature map and the first masked feature map to obtain the second cross-contrast feature map.
[0095] The feature fusion unit fuses the first masked reconstruction feature map and the first cross - contrast feature map to obtain the local feature map of the first input image; the feature fusion unit fuses the first masked reconstruction feature map and the first cross - contrast feature map to obtain the local feature map of the second input image; the local feature image pair is composed of the local feature map of the first input image and the local feature of the second input image.
[0096] Different from the existing masked reconstruction schemes, the masked reconstruction of the present invention is to perform targeted masked reconstruction on the features of key positions such as the side wall and the back wall of the landslide at the feature layer, which can more efficiently guide the model to focus on the extraction and representation ability of local significant semantic features that contribute the most to landslide detection. By masking some landslide edge blocks and background blocks (i.e., non - landslide blocks) and reconstructing the masked areas using other landslide edges and non - landslide parts in the original image, the model can especially learn the features of the landslide edge and the background, thereby enhancing the model's feature extraction ability. The reconstruction task of feature point recovery on the feature map is because after multiple convolutional operations, the high - level feature map contains more abundant and more abstract semantic features than the RGB image.
[0097] Semantic contrast enhancement is to perform supervised contrast learning on the features of the landslide edge and the background. Different from the traditional contrast learning that constructs target - level sample pairs, the present invention constructs positive samples from the landslide edge and negative samples from the environmental background, directly focusing on the differences between similar features and enhancing the model's ability to distinguish visually blurred features. Semantic feature contrast enhancement performs contrast learning between feature points from different samples, effectively increasing the distance between the landslide edge features and the non - landslide category features in the high - dimensional semantic feature space. Using the supervised contrast loss function, feature points with the same label can form positive sample pairs, better capturing the similarity of features within the same category.
[0098] The specific steps for the local feature extraction sub - module to process the input image pair are as follows:
[0099] Divide the input image with size H×W into blocks, each block is a pixel of size n×n, and the value of n can be determined according to the down - sampling rate of the feature extractor. At this time, each n×n pixel block in the input image exactly corresponds to a feature point on the feature map.
[0100] The masking unit classifies the n×n pixel blocks into three categories: non - landslide blocks containing less than 10% landslide pixels, landslide edge blocks containing greater than or equal to 10% and less than or equal to 90% landslide pixels, and landslide interior blocks containing more than 90% landslide pixels. It should be noted that the landslide interior blocks are highly similar to the surrounding environmental features in the RGB image and contribute less to landslide detection. Therefore, the masking unit will directly discard them.
[0101] For each image, the masking unit randomly selects landslide edge blocks (i.e., class 1) and non-landslide blocks (i.e., class 0) for masking operations to generate a masked image paired with the original input image. The masking process ensures that the number of two types of blocks selected on each image is equal. In addition, the masking unit uses maskList_1 and maskList_0 to record the positions of the masked landslide edge blocks and non-landslide blocks on the original input image respectively.
[0102] Randomly select two different landslide samples to create input image pairs, and perform masking on each sample. The masked samples (masked images) are input into the student network to obtain the feature maps of the masked samples (masked feature maps) The original samples (input images) are input into the teacher network to obtain the feature maps of the original samples (original feature maps) Each feature point in the feature map corresponds to an n×n pixel block in the original sample. Subsequently, the feature maps f mask and f img are flattened in the spatial dimension.
[0103] According to the masking position information recorded in the maskList, the reconstructed feature points f mask_i 1×1×C ∈f mask are collected from the feature map of the masked sample to construct the reconstructed feature and the feature points f img_i 1×1×C ∈f img corresponding to the corresponding positions in the feature map of the original sample are collected to construct the label feature where N represents the number of randomly masked blocks in the original sample.
[0104] During model training, the following mean squared error (MSE) loss function is used to calculate the loss between the reconstructed feature and the label feature:
[0105]
[0106] where N is the number of annotated landslide image pairs participating in training, f' re and f' l represent the reconstructed feature and the label feature corresponding to each sample (landslide image) respectively, and L MSE is the loss between the reconstructed feature and the label feature of each landslide image.
[0107] For the two samples in the input image pair, calculate the loss between the reconstructed feature and the label feature of each sample respectively, and then add the two to obtain the masked reconstruction loss of the input image pair. The calculation formula is as follows:
[0108] L Reconst = L MSE_1 + L MSE_2
[0109] Among them, L Reconst is the masked reconstruction loss of the input image pair, L MSE_1 is the loss between the reconstructed feature and the label feature corresponding to the first input image, L MSE_2 is the loss between the reconstructed feature and the label feature corresponding to the second input image.
[0110] By masking some landslide edge blocks and non-landslide blocks and using other landslide edges and non-landslide parts in the original image to reconstruct the masked area, the model can learn the features of landslide edges and the background, thereby enhancing the model's feature extraction ability. Setting the reconstruction task at the feature layer level to restore the feature points of the masked part is because after multiple convolutional operations, the high-level feature map contains more rich and abstract semantic features than the highly similar image features in the RGB image, which enables the model to obtain more knowledge.
[0111] Normalize and flatten the feature map f mask of the masked sample and the feature map f img of the original sample. Each feature point in the feature map corresponds to the semantic feature of an n×n pixel block in the original sample. According to the position information of the masked pixel block recorded in the mask unit on the original image, extract the reconstructed feature point f 1 from the feature map f mask1 of mask mask1_i 1×1×C ∈ f mask1 , extract the original feature point f 2 from the feature map f img2 of img img2_i 1×1×C ∈ f img2 , and perform cross-contrast enhancement on these extracted feature points; at the same time, extract the reconstructed feature point f 2 from the feature map f mask2 of mask mask2_i 1×1×C ∈ f mask2 , extract the original feature point f 1 from the feature map f img1 of img img1_i 1×1×C ∈ f img1 , and perform cross-contrast enhancement on these extracted feature points.
[0112] In the cross-contrast enhancement of the reconstructed feature points and the original feature points, mark all the feature points corresponding to the landslide edge blocks as class 1 to form a class 1 set All the feature points corresponding to the non-landslide blocks are marked as class 0, forming the class 0 set. Calculate the loss between the class 1 set and the class 0 set using the supervised contrastive learning loss function. The calculation formula is:
[0113]
[0114]
[0115] Among them, L cons is the loss between the class 1 set and the class 0 set; is the loss of the feature vector of the i-th feature point in the class 1 set; I [i≠j] is the indicator function. When i≠j, I [i≠j] =1; x i and x j are the feature vectors of the i-th and j-th feature points in the class 1 set respectively; T is the temperature hyperparameter; is the feature vector of the k-th feature point in the class 0 set;
[0116] Finally, add the losses of the two cross-contrast enhancements to obtain the cross-contrast loss of the input image pair. The calculation formula is as follows:
[0117] L Cons =L cons_1 +L cons_2
[0118] Among them, L Cons is the cross-contrast loss of the input image pair; L cons_1 is the loss between the class 1 set and the class 0 set corresponding to the first masked feature map and the second original feature map; L cons_2 is the loss between the class 1 set and the class 0 set corresponding to the second masked feature map and the first original feature map.
[0119] Although a reconstruction task is set at the feature layer level to learn more semantic information, it is still limited to a single sample. Semantic feature contrast enhancement performs contrastive learning on feature points from different samples, effectively increasing the distance between landslide edge features and non-landslide category features in the high-dimensional semantic feature space. Using the supervised contrastive loss function, feature points with the same label can form positive sample pairs, which can better capture the similarity of features within the same category and enhance the model's ability to distinguish landslide edge features and background features.
[0120] During model training, use the cross-entropy loss function to calculate the landslide segmentation loss of the decoder. The calculation formula is:
[0121]
[0122] Among them, y i is the true label. For the predicted output (i.e., the detection result of ancient landslides).
[0123] During the training process, a loss function Loss = αL Reconst + βL Cons + λL CE is constructed, which is composed of weighted mask reconstruction loss, cross-contrast loss, and landslide segmentation loss, and the value of the loss function is minimized, where α, β, and λ are weighting coefficients.
[0124] The present invention proposes a multi-task learning technique, that is, a dual-branch collaborative architecture and a self-distillation collaborative architecture are designed. Through the collaboration of semantic segmentation tasks for global context feature extraction, mask reconstruction tasks for local key feature enhancement, and contrast learning tasks, not only multi-scale and multi-level semantic features are fused, but also the sample utilization rate is improved, and the efficiency and reliability of semantic feature extraction are enhanced, thus effectively solving the overfitting problem caused by visual blurring and small sample problems in ancient landslide recognition.
[0125] The following uses specific examples to test the effect of the detection method provided by the present invention:
[0126] The dataset used in the performance test experiment is an ancient landslide dataset in the Loess Plateau area of northwest China. The landslides occurred a long time ago, and many landslides have been transformed into farmland or residential areas, and their shape and texture are highly similar to the surrounding environment, making it difficult to identify them from high-resolution remote sensing images.
[0127] The original resolution of the pictures is 2m / pixel. The pictures are cropped to obtain 304 landslide samples of 512*512, and the samples are divided into a training set, a test set, and a validation set according to a ratio of 6:2:2. Data augmentation is used for the training set and the validation set, and slope samples with 1 / 3 of the landslide quantity are added as negative class samples to obtain the final dataset. The sample quantities of the training set, the validation set, and the test set are shown in the following table:
[0128] Table 1
[0129] Training set Validation set Test set 1460 486 80
[0130] The method provided by the present invention and the classic Baseline model DeeplabV3+ are used to conduct performance tests on the experimental dataset, and the test results are shown in Table 2.
[0131] Table 2
[0132] Method PA Precision Recall 1-IoU mIoU F1-score Baseline 0.9445 0.4226 0.6284 0.3381 0.6405 0.5054 The method of this patent 0.9447 0.5347 0.5981 0.3975 0.6680 0.5646
[0133] The following observations can be drawn from Table 2: The method proposed in the present invention performs best among all the comparison models. Specifically, compared with the baseline model, Precision increases from 0.4226 to 0.5347, mIoU (mean intersection over union) increases from 0.6405 to 0.6680, 1-IoU increases from 0.3381 to 0.3934, and F1-socre increases from 0.5054 to 0.5646. These improvements clearly demonstrate the effectiveness of the method proposed in the present invention in ancient landslide detection.
[0134] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An ancient landslide detection method, characterized in that: The method uses a detection model to detect ancient landslides from landslide images, wherein the landslide images are high-resolution remote sensing satellite images, and the detection model includes a feature extraction module, a feature fusion module, and a decoder, wherein the feature extraction module includes a semantic feature extraction submodule and a local feature extraction submodule; The method comprises: The detection model is trained using the labeled landslide images, and the parameters of the detection model are continuously adjusted based on the loss function. When the performance of the detection model is no longer improved or reaches a preset number of iterations, the training is stopped to obtain a trained detection model for detecting ancient landslides from landslide images. Obtaining a pair of landslide images to be detected as an input image pair, and inputting the input image pair into a semantic feature extraction submodule and a local feature extraction submodule respectively; The semantic feature extraction submodule extracts the global context features of the ancient landslide of each image of the input image pair to obtain a semantic feature image pair; The local feature extraction submodule extracts the local features of the terrain height change at the edge of the ancient landslide of each input image pair through mask reconstruction tasks and contrast learning tasks to obtain a local feature image pair; The feature fusion module performs feature fusion on the semantic feature image pair and the local feature image pair to obtain a fused feature image pair; The decoder restores the original image pair resolution on the fused feature image pair to obtain the ancient landslide detection result of the input image pair.
2. The ancient landslide detection method according to claim 1, characterized in that: The local feature extraction submodule includes an image segmentation unit, a mask unit, a self-distillation teacher-student network, a mask reconstruction unit, a cross-contrast learning unit, and a feature fusion unit; the input image pair consists of a first input image and a second input image, and the first input image and the second input image are respectively collected from the same geographical area at different time points; The step of extracting the second feature image pair by the local feature extraction submodule comprises: The image division unit divides the pixel blocks on the first input image and the second input image respectively to obtain a first divided image and a second divided image; wherein each pixel block on the first divided image and the second divided image corresponds to a feature point on the feature map; The mask unit performs mask processing on the first segmented image and the second segmented image respectively to obtain a first mask image and a second mask image; The teacher network extracts local features of the first input image and the second input sub-image respectively to obtain a first original feature map and a second original feature map; The student network extracts local features of the first mask image and the second mask image respectively to obtain a first mask feature map and a second mask feature map; The mask reconstruction unit performs mask reconstruction on the mask part of the first mask feature map based on the first original feature map and the first mask feature map to obtain a first mask reconstructed feature map; the mask reconstruction unit performs mask reconstruction on the mask part of the second mask feature map based on the second original feature map and the second mask feature map to obtain a second mask reconstructed feature map; The cross-contrast learning unit performs cross-contrast learning on the feature points on the first mask feature map and the second original feature map to obtain a first cross-contrast feature map; and performs cross-contrast learning on the feature points on the second mask feature map and the first original feature map to obtain a second cross-contrast feature map; The feature fusion unit fuses the first mask reconstruction feature map and the first cross contrast feature map to obtain a local feature map of the first input image; the feature fusion unit fuses the first mask reconstruction feature map and the first cross contrast feature map to obtain a local feature map of the second input image; the local feature image pair is composed of the local feature map of the first input image and the local features of the second input image.
3. The ancient landslide detection method according to claim 2, characterized in that: The image division unit divides the pixel blocks into non-landslide blocks, landslide edge blocks, and landslide internal blocks based on the proportion of landslide pixels contained in each pixel block; The partitioning steps include: Setting a first division ratio and a second division ratio for dividing the pixel block; The pixel blocks containing landslide pixels less than the first division ratio are divided into non-landslide blocks, the pixel blocks containing landslide pixels greater than or equal to the first division ratio and less than or equal to the second division ratio are divided into landslide edge blocks, and the pixel blocks containing landslide pixels greater than the second division ratio are divided into landslide internal blocks.
4. The ancient landslide detection method according to claim 3 is characterized in that: The mask unit randomly selects a certain number of non-landslide blocks and landslide edge blocks on the first segmented image and the second segmented image for masking to obtain a first mask image and a second mask image; The corresponding positions of the masked non-landslide blocks and landslide edge blocks on the first mask image and the second mask image on the first input image and the second input image are recorded respectively.
5. The ancient landslide detection method according to claim 4, characterized in that: The mask unit randomly selects an equal number of non-landslide blocks and landslide edge blocks on the first segmented image, and randomly selects an equal number of non-landslide blocks and landslide edge blocks on the second segmented image.
6. The ancient landslide detection method according to claim 4, characterized in that: The mask reconstruction unit performs mask reconstruction steps including: The mask reconstruction unit reconstructs the masked feature points on the first mask feature map using feature points corresponding to other non-landslide blocks and landslide edge blocks on the first original feature map corresponding to the first input image based on the corresponding positions of the masked non-landslide blocks and landslide edge blocks on the first mask image on the first input image, so as to obtain a first mask reconstructed feature map; The step of the mask reconstruction unit obtaining the second mask reconstruction feature map is the same as the step of obtaining the first mask reconstruction feature map.
7. The ancient landslide detection method according to claim 6, characterized in that: The mean square error loss function is used to calculate the mask reconstruction loss of the input image pair during model training, including: Based on the corresponding positions of the masked non-landslide blocks and landslide edge blocks on the mask image on the corresponding input image, feature points on the mask reconstructed feature map are obtained to construct reconstructed features, and feature points on the original feature map are obtained to construct label features; The mean square error loss function is used to calculate the loss between the reconstructed features and the label features to obtain the mask reconstruction loss of the mask reconstruction unit. The calculation formula is: L Reconst =L MSE_1 +L MSE_2 Where N is the number of labeled landslide image pairs involved in training, f′ re , f′ l are the reconstruction features and label features corresponding to each landslide image, L MSE is the loss between the reconstructed features and the label features of each landslide image, L Reconst is the mask reconstruction loss for the input image pair, L MSE_1 is the loss between the reconstructed features and the label features corresponding to the first input image, L MSE_2 It is the loss between the reconstructed features and the label features corresponding to the second input image.
8. The ancient landslide detection method according to claim 4, characterized in that: The supervised contrast learning loss function is used to calculate the cross-contrast loss of the input image pairs during model training, including: Mark the feature points corresponding to all landslide edge blocks on the original feature map and the mask feature map as class 1 to form a class 1 set Mark all feature points corresponding to non-landslide blocks on the original feature map and mask feature map as class 0 to form a class 0 set Both class 1 and class 0 contain M feature points; The supervised contrast learning loss function is used to calculate the loss between the class 1 set and the class 0 set corresponding to the first mask feature map and the second original feature map, and the loss between the class 1 set and the class 0 set corresponding to the second mask feature map and the first original feature map to obtain the cross contrast loss of the input image pair. The calculation formula is: L Cons =L cons_1 +L cons_2 Among them, L cons is the loss between class 1 and class 0; is the loss of the feature vector of the i-th feature point in class 1; I [i≠j] is the indicator function. When i≠j, I [i≠j] =1;x i 、x j are the feature vectors of the i-th and j-th feature points in class 1 respectively; T is the temperature hyperparameter; is the feature vector of the kth feature point in class 0; L Cons is the cross-contrast loss of the input image pair; L cons_1 is the loss between the first mask feature map and the second original feature map corresponding to class 1 and class 0; L cons_2 It is the loss between the class 1 set and the class 0 set corresponding to the second mask feature map and the first original feature map.
9. The ancient landslide detection method according to claim 2, characterized in that: The parameters of the student network are updated using stochastic gradient descent and back propagation; The parameters of the teacher network are constructed by the momentum editor based on the parameters of the student network, and the parameter update strategy of the teacher network is: i t =λθ t +(1-λ)θ s Among them, θ t is the parameter of the teacher network, θ s is the parameter of the student network, and λ is a hyperparameter.
10. The ancient landslide detection method according to claim 1, characterized in that: The decoder stacks two transposed convolutional layers to restore the original image pair resolution to the fused feature image pair; The decoder also includes: a drpout layer for avoiding model overfitting, a batch normalization layer for limiting the range of data fluctuations, and a ReLU activation function for increasing model sparsity and avoiding model gradient disappearance.
Citation Information
Patent Citations
High-resolution remote sensing image-based multi-index fusion landslide detection method
CN105989322A
Landslide identification method and system based on attention mechanism and multimodal representation learning
CN114170533B
Landslide identification method and device based on satellite data UNet network model
CN115131684A
Landslide detection method and system based on multi-scale feature fusion
CN116740521A
Method for converting remote sensing image into map by fusing discriminant model and generative model
CN117422787A
Cited By
Natural resource supervision method based on three-dimensional GIS scene and video fusion
CN121053317A