Long-time-sequence landslide identification method and system for various terrains

By combining the SwinTransformer encoder, PSA attention mechanism module and FCUperNet decoder, and using the recognition method of the focus binary cross entropy loss function, the problem of landslide recognition under diverse terrain conditions is solved, and high-precision and high-efficiency landslide recognition is achieved.

CN120088545AActive Publication Date: 2025-06-03WUHAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510138469.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-03
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and classify landslides in remote sensing images under diversified terrain conditions, especially in small-size targets and complex backgrounds, and the problem of sample imbalance is serious.

Method used

The recognition method of fusing SwinTransformer encoder, PSA attention mechanism module and FCUperNet decoder is adopted, and the loss function is changed to a focus binary cross entropy loss function to improve the model's ability to identify a few categories and handle complex backgrounds.

Benefits of technology

It significantly improves the accuracy and efficiency of landslide identification, can adapt to a variety of terrain conditions, handle small-size targets, solve sample imbalance problems, and effectively distinguish complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088545A_ABST
    Figure CN120088545A_ABST
Patent Text Reader

Abstract

The invention discloses a long-time-sequence landslide identification method and system for various terrain environments, and the method comprises the following steps: carrying out the construction of a long-time-sequence landslide data set and the enhancement of the data set for various terrains; constructing a long-time-sequence landslide recognition model based on the long-time-sequence landslide data set; performing landslide extraction and category identification on the input remote sensing image by using a long-time-sequence landslide identification model; and finally, data backflow is carried out, whether labels need to be added and expanded or wrong labels need to be deleted is judged through computer post-processing and manual interaction correction, and the mechanism can adjust and optimize the model according to feedback. The method has the characteristics of excellent multi-temporal long-time-sequence landslide identification capability, good expandability and easy deployment and maintenance, improves the long-time-sequence landslide identification precision and efficiency under various terrains, and is especially suitable for landslide long-term monitoring and disaster prevention and control in special geographical environments such as earthquake disaster areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and remote sensing image processing, specifically to a long-term landslide recognition algorithm, and particularly to a long-term landslide recognition method and system for various terrains. Background Art

[0002] In natural disaster monitoring and geological disaster prevention and control, accurately identifying and monitoring different types and evolution processes of landslides is crucial. Landslides are a common geological disaster that poses a serious threat to human society and the natural environment. Traditional long-term landslide recognition and monitoring methods mainly rely on manual interpretation and on-site surveys, which are time-consuming, laborious, and prone to subjective errors, and are difficult to meet the monitoring and management requirements of large-scale regions. With the development of remote sensing technology, it has become easier to obtain high-quality and high-resolution remote sensing images, which provides new possibilities for landslide monitoring, geological disaster assessment, and environmental protection. However, accurately identifying and classifying landslides in these images, especially under diverse terrain conditions, remains a challenge.

[0003] First of all, the problem of landslide recognition in remote sensing images is complex and diverse. Different types of landslides (such as collapses, debris flows, landslides, etc.) exhibit different characteristics in the images. These characteristics are affected by various factors, including the type of landslide, the covered terrain, seasonal changes, lighting conditions, etc. For example, collapses usually present as steep cliffs and broken rocks, while debris flows may form obvious gullies and sediments due to water erosion. These differences limit the recognition effect based on traditional image processing methods and are difficult to meet the requirements of high precision and high efficiency.

[0004] Secondly, the problem of small target recognition in remote sensing images is also very prominent. In large-scale remote sensing images, some targets such as landslides are relatively small and are often mixed with the surrounding natural environment and other geological structures. This not only increases the difficulty of target recognition but also requires the recognition model to have stronger feature extraction capabilities.

[0005] In addition, the problem of sample imbalance is particularly significant in the long-term landslide recognition of remote sensing images. In practical applications, landslide areas are usually scarcer than other ground objects (such as forests, farmlands, buildings, etc.), resulting in a serious imbalance in the number of landslide samples in the training dataset. This imbalance will affect the training effect of the model and make the model perform poorly in identifying landslide types.

[0006] In addition, the limitations of existing remote sensing image processing methods when dealing with complex backgrounds or similar features are also a major problem. Remote sensing images often contain complex background information such as land, vegetation, buildings, etc. These backgrounds may be visually similar to landslides, increasing the risk of misclassification. At the same time, traditional rule-based methods usually have limited effects when dealing with these complex scenarios and are difficult to adapt to changing environmental conditions. Summary of the Invention

[0007] In view of this, it is necessary to develop a long-term landslide recognition method for multiple terrains, which can adapt to various terrain conditions, process small-size targets, solve the problem of sample imbalance, and effectively distinguish long-term landslides in complex backgrounds.

[0008] The recognition method proposed by the present invention emerges in response to the above problems. It constructs a powerful recognition module by integrating the SwinTransformer encoder, PSA attention mechanism module, and FCUperNet decoder, and changing the loss function to the focal binary cross-entropy loss function, which plays an important role in solving the problem of sample imbalance, improving the model's recognition ability for minority classes, and landslide areas in small regions. Finally, data backflow is performed to determine whether to add and expand labels or delete incorrect labels. This mechanism enables the algorithm to adjust and optimize the model according to feedback during actual application, which is crucial for accurately classifying a small number of but important targets such as landslides in images. The present invention not only improves the recognition accuracy but also significantly enhances the processing efficiency, providing new ideas and methods for solving a series of technical problems in long-term landslide recognition.

[0009] To achieve the above objectives, the technical method of the present invention is as follows: A long-term landslide recognition method for multiple terrains includes the following steps:

[0010] Step 1, dataset construction: Obtain remote sensing images of multiple years before and after an earthquake, and mark the landslide phenomena on the images to obtain marked images;

[0011] Step 2, use data augmentation methods to enhance the images in the dataset;

[0012] Step 3, construction of a long-term landslide recognition model, including an encoder composed of Swin Transformer Blocks for obtaining feature maps of different scales of the input image, a pyramid pooling module for adjusting the scale of the feature maps, a decoder composed of FPN and fully connected in series with the feature maps of corresponding scales in the encoder, an attention mechanism module composed of PSA attention mechanism for fusing the output features of the decoder, and a classification head for finally performing pixel-level classification prediction;

[0013] Step 4: Use the trained long-term landslide recognition model to extract and identify landslides in the remote sensing images.

[0014] Furthermore, the data augmentation methods in Step 2 include contrast perturbation, image rotation, overexposure processing, color perturbation, color separation, contrast enhancement, brightness adjustment, image sharpening, image shearing, weather filters, haze processing, and random chunking.

[0015] Furthermore, the encoder contains k Swin Transformer Blocks with the same structure. First, the input image is linearly embedded into smaller Patch blocks, then processed by Swin Transformer Block1, and then through the Patch Merging technique, the number of Patch blocks and the feature dimensions change gradually. Finally, it is processed by k - 1 Swin Transformer Blocks to obtain the final output.

[0016] Furthermore, the input of the pyramid pooling module comes from the final output of the encoder. First, pooling operations, convolutional operations, and upsampling operations of different scales are performed to capture context information of different scales and obtain multiple feature maps. Subsequently, the multiple feature maps are fused and merged through addition operations to form the final fused feature map.

[0017] Furthermore, the full-order long connection is implemented through the full-order long connection conversion layer. The processing process of the full-order long connection is as follows:

[0018] For the feature maps of different scales output by the encoder, first, a 3x3 convolutional layer and a ReLU activation function are applied to process each scale of feature map to change the number of channels of the feature map while keeping the spatial dimensions unchanged; then, a max pooling layer is used for downsampling. The number of downsampling times is determined by the scale of the feature map. The deepest scale of feature map does not need to be downsampled, and the number of downsampling times for the shallowest scale of feature map is k - 1, where k is the number of Swin Transformer Blocks. The number of downsampling times decreases sequentially as the scale of the feature map increases, and the spatial dimensions are halved after each downsampling operation.

[0019] Furthermore, the decoder integrates the feature maps from all levels of the encoder, the final output of the pyramid pooling module PPM, and the output of the full-order long connection conversion layer;

[0020] The decoder includes k FPN layers, where k is the number of Swin Transformer Blocks; the processing process of each FCN layer is as follows: receive the deepest feature map from the pyramid pooling module, upsample it, and then add it to the feature map of the corresponding scale of the encoder and the feature map of the same dimension after full-order long connection conversion processing; finally, output k feature maps with the same number of channels.

[0021] Further, the processing process of the Pyramid Segmentation Attention (PSA) is as follows:

[0022] The input feature map x is regrouped after spatial convolution processing, each group is processed in its corresponding spatial convolution layer, and the feature map of each group calculates the SE weight through the corresponding SE block to generate the weight of each group; in the SE block, first, global average pooling is performed on the feature map of each channel to generate a channel-level global feature descriptor. After global average pooling, the dimensionality reduction convolutional layer reduces the dimension of the feature through a 1x1 convolution. The feature after ReLU activation is upsampled through a 1x1 convolutional layer, and the Sigmoid activation function limits the importance weight between 0 and 1; then, focusing is performed, that is, the normalized weight is multiplied element-wise with the output of the spatial convolution processing to focus on important spatial features; finally, the final output feature map is recombined to match the dimension of the original input feature map.

[0023] The output feature map of each FCN layer is processed by PSA and finally reshaped and concatenated. The concatenated feature map is used for pixel-level classification prediction through the classification head, and the probability distribution of each pixel belonging to each category is output.

[0024] Further, the focal binary cross-entropy loss function is used to train the long-term landslide recognition model. The focal binary cross-entropy loss function is based on the binary cross-entropy loss and applies the focal loss function for weight adjustment; the calculation formula of the binary cross-entropy loss function BCELoss is as follows:

[0025]

[0026] In the formula, N is the number of samples, y i is the true label of the i-th sample, taking values of 0 or 1, and p i is the probability that the model predicts the i-th sample as class 1;

[0027] The calculation formula of the focal loss function FLoss is as follows:

[0028] FLoss(p t ) = -α t (1 - p t ) γ log(p t )

[0029] where p t is the probability predicted by the model, γ is a tuning parameter used to reduce the weight of easy-to-classify samples, and α t is the class-level weight parameter used to further address the class imbalance problem;

[0030] Combining binary cross-entropy loss and focal loss, the overall formula for the focal binary cross-entropy loss function is obtained:

[0031]

[0032] where FocalBCELoss represents the focal binary cross-entropy loss function, N is the total number of samples, is the predicted probability corresponding to the true label of the i-th sample by the model, represents the class weight of the i-th sample, and according to its true label t i select the weight α of the positive class 1 or the weight α of the negative class 0 to balance the difference in the number of classes, where α 0 and α 1 are hyperparameters.

[0033] Furthermore, it also includes optimizing the confidence map post-processing model through a computer post-processing and human interaction correction mechanism. Specifically: set an accuracy threshold a, automatically evaluate whether the recognition result meets the standard. If the test index does not reach the threshold a, then these test data are fed back for manual re-interpretation and annotation to construct new labels, expand the existing dataset, or delete incorrect labels.

[0034] The present invention also provides a long-term landslide recognition system for multiple terrains, including:

[0035] A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the long-term landslide recognition method for multiple terrains as described in the above technical solution.

[0036] In summary, the present invention has developed a long-term landslide recognition method for various terrains, which can adapt to various terrain conditions, process small-sized targets, solve the problem of sample imbalance, and effectively distinguish complex backgrounds, thus solving the problem of accurately identifying and classifying landslides under various terrain conditions. Moreover, through the mechanism of the data feedback module, the algorithm can adjust and optimize the model according to the feedback during the actual application process, which is crucial for accurately classifying a small number of but important targets such as landslides in images. The present invention not only improves the recognition accuracy but also significantly enhances the processing efficiency, providing new ideas and methods for solving a series of technical problems in landslide recognition, and can better support decision-making in fields such as geological disaster monitoring, land resource planning, and environmental monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative efforts.

[0038] Figure 1 is a schematic flowchart of an embodiment of the present invention;

[0039] Figure 2 is a schematic flowchart of the overall model of an embodiment of the present invention;

[0040] Figure 3 is a schematic structural diagram of the PSA module provided by an embodiment of the present invention;

[0041] Figure 4 is a schematic flowchart of the PPM module provided by an embodiment of the present invention;

[0042] Figure 5 is a landslide recognition result diagram of a partial model of an embodiment of the present invention in the test year;

[0043] Figure 6 is a partial schematic diagram of the landslide recognition result diagram of a partial model of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0045] It should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present invention illustrate the operations implemented according to some embodiments of the present invention. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical context may be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present invention. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor systems and / or microcontroller systems.

[0046] The descriptions such as "first" and "second" involved in the embodiments of the present invention are only for implicit purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" may explicitly or implicitly include at least one of such features. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0047] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0048] The main technical difficulties to be solved by the present invention:

[0049] 1. Construct a large-scale landslide dataset containing various terrains and scenarios. By collecting remote sensing images of multiple years before and after earthquakes and marking the landslide phenomena on the images, marked images are obtained. These images cover landslide situations at different time periods, different seasons, and different weather conditions, ensuring the diversity and representativeness of the dataset;

[0050] 2. A landslide type that can adapt to various terrain conditions, process small-size targets, solve the problem of sample imbalance, and effectively distinguish complex backgrounds; the present invention constructs a powerful recognition module by integrating a Swin Transformer encoder, a PSA attention mechanism module, and an FCUperNet decoder, and changing the loss function to a focal binary cross-entropy loss function;

[0051] 3. The identification of landslide types generally requires professional knowledge training, and the cost of manual identification and review is relatively high. The algorithm proposed by the present invention can automatically identify different types of landslides, reduce the dependence on professional knowledge, and lower the cost of manual identification and review. Through the data feedback mechanism, the algorithm can adjust and optimize the model according to the feedback, further improving the accuracy and efficiency of identification.

[0052] As Figure 1 shown, the technical solution of the present invention is a long-term landslide identification method for multiple terrains, including: constructing a landslide data set, enhancing the landslide data set, constructing a long-term landslide identification algorithm, training and generating a long-term landslide identification model, identifying landslides in long-term remote sensing images, and determining whether the data flows back. The specific process is as follows:

[0053] The first step: constructing a long-term landslide data set

[0054] The second step: enhancing the long-term landslide data set

[0055] The third step: constructing a long-term landslide identification algorithm

[0056] By integrating the Swin Transformer encoder, the PSA attention mechanism module and the FCUperNet decoder, and changing the loss function to the focal binary cross-entropy loss function, a powerful identification module is constructed.

[0057] The fourth step: training and generating a long-term landslide identification model

[0058] The fifth step: identifying the corresponding targets on the remote sensing image by the long-term landslide identification model

[0059] The sixth step: data feedback. Through computer post-processing and manual interaction correction, it is judged whether the returned samples need to separately construct new labels, expand the existing data set or delete incorrect labels. Set a threshold a. According to whether the test index of the predicted long-term landslide identification model reaches the threshold, if not, the test data is fed back for manual re-interpretation and annotation.

[0060] The specific implementation manner of the complete technical solution of the present invention is described as follows:

[0061] 1. Process of constructing the data set: Collect remote sensing images of multiple years before and after the earthquake, and mark the landslide phenomena on the images to obtain marked images. These images should cover landslide conditions in different time periods, different seasons and different weather conditions to ensure the diversity and representativeness of the data set.

[0062] 2. Data set augmentation process: The data set is augmented in different types through image algorithms to simulate the imaging effects at different times and seasons. The specific augmentation methods include contrast perturbation, image rotation, overexposure processing, color perturbation, color separation, contrast enhancement, brightness adjustment, image sharpening, image shear, weather filter, haze processing, and random cropping, etc., to enrich the diversity of the data set and improve the generalization ability and robustness of the model.

[0063] 3. The long-term landslide recognition algorithm for multiple terrains of the present invention mainly consists of three modules as Figure 2 shown. Module 1 is an encoder composed of Swin Transformer, Module 2 is an attention mechanism module composed of PSA attention mechanism, and Module 3 is a decoder composed of FCUperNet. These three are connected in series and the loss function is changed to the focal binary cross-entropy loss function. The explanations of each module are as follows:

[0064] (1) Using Swin-Transformer as the encoder, features are gradually extracted and refined through a hierarchical method. At different network levels, the features are downsampled and transformed to form feature representations of different scales. Swin Transformer as the encoder provides powerful feature extraction capabilities. Swin Transformer includes several stacked SwinTransformer Blocks, and each Swin Transformer block includes an MLP layer. The calculation process of the MLP layer is described as:

[0065] Z 1 = XU + b 1

[0066] A = σ(Z 1 )

[0067] Y = AV + b 2

[0068] In the formula, Z 1 is the first linear transformation formula, X is the input, U is the weight matrix of the first linear mapping, b 1 is the bias term corresponding to U, A is the activation function, σ is the ReLU function, Y is the output, V is the weight matrix of the second linear mapping, and b 2 is the bias term corresponding to V.

[0069] In this embodiment, the Swin Transformer encoder includes four Swin Transformer Blocks with the same structure. First, the input image becomes smaller Patch blocks after linear embedding, and its feature dimension is C. The output resolution processed by Swin Transformer Block1 is [8, 128, 32, 32] (corresponding to [batch size, number of channels, height, width] respectively). As the model processes stage by stage backward, through the Patch Merging technology, the number and feature dimension of the Patch blocks change gradually until the final output dimension reaches [8, 1024, 4, 4].

[0070] (2) Regarding the limitations of the traditional feature pyramid decoder in processing multi-scale feature maps and maintaining rich semantic information, such as Figure 2 shown, the FCUperNet decoder effectively captures context information at different scales and supplements low-dimensional features by integrating the pyramid pooling module (PPM), as Figure 4 shown, and the improved full-stage long connection feature pyramid network (FPN). As Figure 3 shown, the pyramid segmentation attention (PSA) mechanism is further adopted to fuse these features. This design not only solves the problem of detail loss caused by the reduction of the feature map resolution but also overcomes the problems of loss of low-dimensional feature information and reduction of semantic information.

[0071] As a feature extraction component, the pyramid pooling module (PPM) generates a set of feature layers rich in high semantic information through multi-scale pooling and downsampling operations. Subsequently, these feature layers are upsampled to adapt to large-size feature maps for small object detection. The module synthesizes the pooling results of each layer to form a richer feature representation, effectively capturing multi-scale context information, and thus improving the segmentation accuracy of the model.

[0072] In the FCUperNet model, the pooling scale of the PPM part is set to (1,2,3,6), and its input feature map comes from the final output of the encoder, and its dimension is [8,1024,4,4] (corresponding to [batch size, number of channels, height, width]). First, a global pooling operation with a scale of 1 is performed on the feature map, so that the processed feature map size becomes [8,1024,1,1]. Then, the number of channels is adjusted to 512 using a 1x1 convolution (the specific number of channels depends on the implementation), and the resulting feature map dimension remains [8,512,1,1], and then upsampled to the initial input dimension [8,512,4,4]. Next, a pooling operation with a scale of 2 is performed, using a 2x2 size, so that the resulting feature map size is adjusted to [8,1024,2,2]. After the 1x1 convolution, the number of channels becomes 512, and the feature map size is [8,512,2,2], which is then upsampled to the original input dimension [8,512,4,4]. For the pooling operation of scale 3, the 3x3 size is used for pooling, resulting in a feature map size of [8,1024,3,3]. The number of channels is adjusted to 512 using 1x1 convolution, and the feature map dimensions remain [8,512,3,3], and then upsampled to the original input dimension [8,512,4,4]. In the pooling operation of scale 6, although the input size is only 4x4, 6x6 pooling is achieved by padding, resulting in a feature map size of [8,1024,6,6]. After the 1x1 convolution, the number of channels is adjusted to 512, the size remains [8,512,6,6], and then downsampled to match the original input dimension [8,512,4,4]. After the above processing, four feature maps of size [8,512,4,4] are obtained, which capture context information of different scales. These feature maps are then fused and merged through addition operations, with the number of channels remaining unchanged, and the final size is maintained as [8,512,4,4]. The feature of each pixel position is the sum of the features of the same position corresponding to feature maps of different scales, forming the final fused feature map.

[0073] Based on the design concept of U3+, a full-stage connection module is designed to transform all output feature maps of the backbone, i.e., the encoder. For example, for the input [8, 128, 32, 32]: By applying a 3x3 convolutional layer and a ReLU activation function, the number of channels is increased from 128 to 512 while keeping the spatial dimensions unchanged, resulting in [8, 512, 32, 32]. Next, three consecutive downsamplings are performed using a MaxPool2d(kernel_size = 2, stride = 2) max pooling layer. Each operation halves the spatial dimensions, and the results are [8, 512, 16, 16], [8, 512, 8, 8], and [8, 512, 4, 4] respectively. For the input [8, 256, 16, 16]: Applying a 3x3 convolution and a ReLU activation function increases the number of channels from 256 to 512 while keeping the spatial dimensions at [8, 512, 16, 16]. Immediately afterwards, two rounds of max pooling downsampling are performed to sequentially obtain feature maps with dimensions [8, 512, 8, 8] and [8, 512, 4, 4]. For the input [8, 512, 8, 8], the number of channels of this input is 512, and the spatial dimensions are directly halved through one max pooling downsampling, resulting in [8, 512, 4, 4]. For the input [8, 1024, 4, 4], the number of channels is reduced from 1024 to 512 using a 3x3 convolutional layer and a ReLU activation function, and the spatial dimensions are maintained at [8, 512, 4, 4]. Since this input is already at the minimum spatial dimension, no pooling operation is performed.

[0074] The Feature Pyramid Network (FPN) is inspired by the human visual system, which can parse visual information at multiple scales and abstraction levels. The FPN adjusts the number of channels of the bottom-up features through 1x1 convolutions and fuses them with the output of the pyramid pooling module on the top-down path. Further feature extraction and lateral connections are performed through 3x3 convolutions to generate a series of feature maps rich in semantic information. In multi-scale feature fusion, the FPN may reduce the resolution of the feature maps by performing multiple downsamplings and upsamplings on the image. In the upsampling part of the FPN, only the high-level features are fused with the previous layer features by nearest neighbor upsampling and addition, which may lead to the loss of low-dimensional feature information. And the reduction of channels in feature fusion may lead to the loss of semantic information, thus affecting the segmentation accuracy. In response to this, we draw on the full-stage connection strategy of Unet3+ and convert the output feature maps of each layer of the backbone network into each layer of the FPN upsampling through convolution and pooling and then add them together, which helps the model obtain more low-dimensional feature information and improve the segmentation accuracy.

[0075] In this embodiment, the improved Feature Pyramid Network (FPN) integrates the data features of all levels from the backbone, the final output of the PPM, and the output of the full-stage long connection conversion layer. First, FPN receives the deepest feature map [8, 512, 4, 4] from the PPM and upsamples it to [8, 512, 8, 8]. Then, it adds it to the [8, 512, 8, 8] from the backbone and the feature map of the same dimension processed by the full-stage long connection conversion. Next, the fused feature map is further upsampled to [8, 512, 16, 16], and then added to the backbone feature map [8, 256, 16, 16] adjusted by 1x1 convolution and the [8, 512, 16, 16] output by the full-stage long connection conversion layer. Finally, the upsampling process continues to [8, 512, 32, 32], and it is added to the backbone feature map [8, 128, 32, 32] adjusted by 1x1 convolution and the output [8, 512, 32, 32] of the full-stage long connection conversion layer. Finally, the deepest feature map [8, 512, 4, 4] of the PPM is added to the output of the full-stage long connection conversion layer.

[0076] Each fusion generates a feature map with 512 channels. All feature maps have been unified to 512 channels before fusion, so the final output feature map of FPN also maintains 512 channels.

[0077] (3) The feature map output by FPN needs to be fused through the PSA self-attention mechanism before splicing. In PSA self-attention, the input feature map x is first regrouped into the shape b, self.S, c / / self.S, h, w through spatial convolution processing (SPC module). Each group is processed in its corresponding spatial convolution layer, and c / / self.S is the number of channels of each group. The feature map of each group calculates the SE weight through the corresponding SE block to generate the weight of each group. In the SE block, first, global average pooling (GAP) is performed on the feature map of each channel to generate a channel-level global feature descriptor. After global average pooling, the dimensionality reduction convolutional layer reduces the dimension of the feature through 1x1 convolution. The ReLU-activated feature is upsampled through a 1x1 convolutional layer, and the Sigmoid activation function limits the importance weight between 0 and 1. Enter the spatial attention focus (SPA module), multiply the normalized weight (processed by Softmax) element-wise with the output of the SPC module, and focus on the important spatial features. Finally, the PSA output feature map is recombined to make its shape b, -1, h, w to match the dimension of the original input feature map.

[0078] The PSA output feature map is finally resized to [8, 512, 32, 32] and concatenated into [8, 2048, 32, 32]. The concatenated feature map is used for pixel-level classification prediction through the classification head, and the probability distribution of each pixel belonging to various categories is output.

[0079] The Pyramid Split Attention (PSA) module constructs a pyramid-shaped feature map by aggregating the convolution results generated by convolution kernels of different sizes, and applies an attention mechanism on this basis to mine rich feature information. The core of the PSA module lies in its ability to enable each element in sequence data processing to establish connections not only with adjacent elements, but with all other elements within the sequence, and adaptively capture long-range dependencies by evaluating the relative importance between elements. The design highlights of the PSA module include: one is polarization filtering, which maintains a high resolution in both the channel and spatial dimensions to reduce information loss; the other is the enhancement mechanism, which precisely simulates the output distribution of fine-grained regression through a combination of non-linear direct fitting. In the FCUperNet decoder, the PSA module is used to perform attention reconstruction in the channel and spatial dimensions on the outputs of FPN and PPM after full-order long connections to optimize feature fusion and achieve effective integration of multi-dimensional and multi-scale context information.

[0080] (4) Improvement of the loss function. In the original cross-entropy loss function, in an imbalanced dataset, the cross-entropy loss may cause the model to over-focus on the majority class, which will lead to a decline in the recognition performance of the model for the minority class. This paper uses the Focal Binary Cross-Entropy Loss (FocalBCE), which is based on the Binary Cross-Entropy Loss Function (BCELoss) and applies the weight adjustment of the Focal Loss (FL). The Binary Cross-Entropy Loss Function:

[0081]

[0082] where N is the number of samples, y i is the true label of the i-th sample, taking values of 0 or 1, and p i is the probability that the model predicts the i-th sample as class 1.

[0083] The Focal loss function:

[0084] FLoss(p t ) = -α t (1 - p t ) γ log(p t )

[0085] where p t is the probability predicted by the model for the correct class. γ is a tuning parameter used to reduce the weight of easy-to-classify samples. α t is the class-level weight parameter, which addresses the problem of imbalanced data distribution by directly amplifying / shrinking the loss values of different classes. It works together with the γ parameter (sample difficulty weight) to enhance the model's ability to identify minority classes or difficult samples.

[0086] Combining binary cross-entropy loss and focal loss, the overall formula for the focal binary cross-entropy loss function can be obtained:

[0087]

[0088] where N is the total number of samples, is the adjusted predicted probability for the correct class. The γ tuning parameter is used to reduce the weight of easy-to-classify samples. represents the class weight of the i-th sample, and according to its true label t i select the weight α 1 for the positive class or the weight α 0 for the negative class, which is used to balance the difference in the number of classes, where α 0 and α 1 are hyperparameters.

[0089] When combining BCELoss and Focal Loss, the weight adjustment mechanism of Focal Loss is used to modify the basic binary cross-entropy loss of each sample. First, calculate the basic binary cross-entropy loss of each sample. Second, for each sample, calculate the weight factor (1 - p t ) γ and α t according to the Focal Loss formula. This weight factor reduces the loss of easy-to-classify (i.e., p t close to 0 or 1) samples while increasing the loss of difficult-to-classify samples. Finally, apply the calculated weight factor to the binary cross-entropy loss of each sample. In this way, the model is guided during training to pay more attention to those difficult-to-classify samples while giving less attention to easy-to-classify samples. Such a loss function is especially suitable for imbalanced datasets and can help improve the model's recognition ability for minority classes.

[0090] In summary, the long-term landslide recognition algorithm for multiple terrains of the present invention mainly consists of three modules. Module one is an encoder composed of Swin Transformer, module two is an attention mechanism module composed of PSA attention mechanism, and module three is a decoder composed of FCUperNet. These three are connected in series and the loss function is changed to the focal binary cross-entropy loss function.

[0091] 4. The present invention trains a long - time - series landslide recognition model based on the enhanced dataset using a long - time - series landslide recognition algorithm for multiple terrains.

[0092] 5. The present invention uses the trained long - time - series landslide recognition model to infer remote sensing images, obtaining the long - time - series landslide recognition results in images of different years. For example, Figure 5 This is the landslide recognition result map of part of the model in the test year in the embodiments of the present invention. Figure 6 This is a partial schematic diagram of the landslide recognition result map of part of the model in the embodiments of the present invention.

[0093] 6. Through computer post - processing and manual interaction correction, the present invention determines whether the returned samples need to separately construct new labels, expand the existing dataset, or delete incorrect labels. This mechanism can adjust and optimize the model according to the feedback. Manual intervention determines whether the returned samples need to separately construct new labels, expand the existing dataset, or be discarded, and then enters the first step. Specifically, it includes setting a threshold a. According to whether the test index of the predicted long - time - series landslide recognition model reaches the threshold, if not, the test data is fed back for manual re - interpretation and annotation. The workload of manual secondary annotation can be controlled by setting the value of the threshold a.

[0094] 7. Finally, the long - time - series landslide recognition is deployed to the mobile end or the cloud.

[0095] On the other hand, the embodiments of the present invention also provide a long - time - series landslide recognition system for multiple terrains, including:

[0096] A processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the long - time - series landslide recognition method for multiple terrains as described in the above technical solution.

[0097] Sample data is used to train and verify the performance of the deep - learning model, which is an important prerequisite for carrying out landslide recognition. The quantity and quality of sample images play a crucial role in the results of training the neural network model. The model test area of the present invention is a 471 - square - kilometer area between a certain town and another town. Within this range, in addition to the landslide catalog data with a time series of more than 10 years, we also obtained Landsat images (path / row 130 / 38) from 2008 to 2018, as shown in Table 1. Using Landsat TM and Landsat OLI data, we can cut out landslide sample images to train and test the deep - learning model.

[0098] Table 1 Satellite remote sensing data table

[0099]

[0100] At the same time, a comparative test was carried out between the SwinFCUperNet model (the model of the present invention) and typical semantic segmentation models such as UNet, UNet3+, TransUNet, Pspnet, DeepLabv3+, Mask2Fomer, SwinUpernet, and Segformer to verify whether SwinFCUperNet has excellent performance. In the test, the images of a single year (2008) were used as samples to train the model, and the landslide recognition performance of the model in other years (2009 - 2018) was tested. If a model can achieve excellent performance without training in other years, it indicates that the model has good transfer ability and can effectively reduce the workload of our long-term landslide mapping.

[0101] Table 2 Test of the time transfer ability of multiple models

[0102]

[0103]

[0104] The data in the table are the average values of the accuracy indicators for the test years (2009 - 2018). Table 2 shows that the SwinFCUperNet model exhibits excellent performance in almost all accuracy indicators. The Precision of this model reaches 74.33%, Recall is 80.18%, mIoU is 76.49%, and F1Score is 77.14%, and these values are the highest among all the compared models. This proves that the SwinFCUperNet model can not only accurately identify landslide events but also maintain high recognition ability within a relatively long time range, demonstrating excellent time transfer ability.

[0105] Further testing the performance of the model in each test year shows that SwinFCUperNet is excellent in the average performance in multiple years, especially in Precision, Recall, mIoU, and F1Score.

[0106] Table 3 Accuracy performance of some models in specific test years

[0107]

[0108]

[0109] The above method has described the embodiments of the present invention in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these fall within the protection scope of the present invention.

Claims

1. A long-term landslide identification method for various terrains, characterized by: The steps include: Step 1: Dataset construction: obtain remote sensing images of multiple years before and after the earthquake, and mark the landslide phenomena on the images to obtain marked images; Step 2, using data enhancement methods to enhance the images in the data set; Step 3, the construction of a long-term landslide recognition model, including an encoder composed of Swin Transformer Block for obtaining feature maps of different scales of the input image, a pyramid pooling module for adjusting the scale of the feature map, a decoder composed of FPN and fully connected with the feature map of the corresponding scale in the encoder, an attention mechanism module composed of PSA attention mechanism for fusing the output features of the decoder, and finally a classification head for pixel-level classification prediction; Step 4: Use the trained long-term landslide recognition model to extract and identify landslides in remote sensing images.

2. The long-term landslide identification method for various terrains as claimed in claim 1 is characterized in that: The data enhancement methods in step 2 include contrast perturbation, image rotation, overexposure processing, color perturbation, tone separation, contrast enhancement, brightness adjustment, image sharpening, image shearing, weather filter, haze processing and random cutting.

3. The long-term landslide identification method for various terrains as claimed in claim 1 is characterized in that: The encoder contains k Swin Transformer Blocks with the same structure. First, the input image is linearly embedded into smaller patches, then processed by Swin Transformer Block1, and then the number and feature dimensions of patches are gradually changed through Patch Merging technology. Finally, the final output is obtained after processing by k-1 Swin Transformer Blocks.

4. The long-term landslide identification method for various terrains as claimed in claim 1 is characterized in that: The input of the pyramid pooling module comes from the final output of the encoder. First, pooling operations of different scales, convolution operations and upsampling operations are performed to capture contextual information of different scales and obtain multiple feature maps. Subsequently, multiple feature maps are fused and merged through addition operations to form the final fused feature map.

5. The long-term landslide identification method for various terrains as claimed in claim 1 is characterized in that: The full-order long connection is realized through the full-order long connection conversion layer. The processing process of the full-order long connection is as follows: For feature maps of different scales output by the encoder, a 3x3 convolution layer and a ReLU activation function are first applied to process the feature maps of each scale, changing the number of channels of the feature map while keeping the spatial size unchanged; then a maximum pooling layer is used for downsampling. The number of downsampling is determined by the scale of the feature map. The deepest feature map does not need to be downsampled, and the number of downsampling of the shallowest feature map is k-1, where k is the number of Swin Transformer Blocks. The number of downsampling decreases with the increase of the feature map scale, and the spatial size is halved after each downsampling operation.

6. The long-term landslide identification method for various terrains as claimed in claim 5 is characterized in that: The decoder integrates the feature maps from all levels of the encoder, the final output of the pyramid pooling module PPM, and the output of the full-order long-connection transformation layer; The decoder consists of k FPN layers, where k is the number of Swin Transformer Blocks; the processing process of each FCN layer is: receiving the deepest feature map from the pyramid pooling module and upsampling it, then adding it to the feature map of the corresponding scale of the encoder and the feature map of the same dimension processed by the full-order long connection conversion; finally outputting k feature maps with the same number of channels.

7. The long-term landslide identification method for various terrains as claimed in claim 1 is characterized in that: The processing process of pyramid segmentation attention PSA is as follows: The input feature map x is regrouped after being processed by spatial convolution. Each group is processed in its corresponding spatial convolution layer. The feature map of each group is calculated by the corresponding SE block to calculate the SE weight and generate the weight of each group. In the SE block, the feature map of each channel is first pooled by local average pooling to generate a channel-level global feature descriptor. After global average pooling, the dimension reduction convolution layer reduces the dimension of the feature through 1x1 convolution. The feature after ReLU activation is increased in dimension through 1x1 convolution layer. The Sigmoid activation function limits the importance weight to between 0 and 1. Then focus is performed, that is, the normalized weight is multiplied element by element with the output of the spatial convolution process to focus on important spatial features. Finally, the final output feature map is reorganized to match the shape of the original input feature map. Each FCN layer is processed by PSA to output feature maps, which are finally reshaped and spliced. The spliced ​​feature maps are used for pixel-level classification prediction through the classification head, and the probability distribution of each pixel belonging to each category is output.

8. The long-term landslide identification method for various terrains as claimed in claim 1 is characterized in that: The long-term landslide recognition model is trained using the focal binary cross entropy loss function. The focal binary cross entropy loss function is based on the binary cross entropy loss and uses the focal loss function for weight adjustment. The calculation formula of the binary cross entropy loss function BCELoss is as follows: Where N is the number of samples, t i is the true label of the i-th sample, which takes the value of 0 or 1, p i is the probability that the model predicts that the i-th sample is category 1; The calculation formula of the focal loss function FLoss is as follows: FLoss(p t )=-a t (1-p t ) γ log(p t ) In the formula, p t is the probability predicted by the model, γ is a tuning parameter used to reduce the weight of easy-to-classify samples, and α t is the category-level weight parameter, which is used to further solve the category imbalance problem; Combining the binary cross entropy loss and the focal loss, we get the overall formula of the focal binary cross entropy loss function: Where FocalBCELoss represents the focal binary cross entropy loss function, N is the total number of samples, is the predicted probability of the model corresponding to the true label of the i-th sample, Represents the category weight of the i-th sample, according to its true label t i The weight α1 of the positive class or the weight α0 of the negative class is selected to balance the difference in the number of categories, where α0 and α1 are hyperparameters.

9. The long-term landslide identification method for various terrains as claimed in claim 1 is characterized in that: It also includes optimizing the confidence map post-processing model through computer post-processing and manual interactive correction mechanism. Specifically: setting an accuracy threshold a to automatically evaluate whether the recognition result meets the standard. If the test index does not reach the threshold a, the test data will be returned for manual re-interpretation and annotation to build new labels, expand existing data sets or delete erroneous labels.

10. A long-term landslide identification system for various terrains, characterized by: include: A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute the long-term landslide identification method for multiple terrains as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on pyramid segmentation attention module

    CN113807210A

  • Automatic landslide identification method, system and device based on lightweight convolutional neural network and double attention, and medium

    CN116206214A

  • Landslide detection early warning model and early warning method based on multi-model fusion

    CN116543308A

  • Landslide identification method and device based on multi-path feature fusion, and storage medium

    CN118736440A

  • Image segmentation model training method, image segmentation method, and apparatus

    WO2024031219A1