Landslide detection method based on Swin Transform and multi-scale feature fusion

By combining Swin Transformer and multi-scale feature fusion method, the problem of insufficient fusion of local and global features in landslide detection is solved, and high-precision and high-root landslide detection is achieved, especially stable detection in complex terrain and environment, reducing false alarms and missed reports.

CN120372524APending Publication Date: 2025-07-25SHAOXING UNIVERSITY +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510230310.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

There are problems in the existing landslide detection technology that insufficient local and global features integration, low multi-scale feature utilization efficiency, poor adaptability for complex scenarios, and bottlenecks in computing efficiency and generalization, making it difficult to achieve high-precision and high-rootability landslide detection.

Method used

Swin Transformer is used as the backbone network, combining the local information aggregation module (LIAM) and the multi-scale feature fusion transverse connection module (MFFLCM), and the model's ability to capture landslide features through self-attention mechanism and cross-scale feature fusion will improve the model's ability to capture landslide features.

Benefits of technology

It realizes high-precision landslide detection, which can stably detect landslides in complex terrain and environment, reduces false alarms and missed reports, improves the accuracy of landslide prediction, and provides reliable technical support for landslide monitoring and early warning systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372524A_ABST
    Figure CN120372524A_ABST
Patent Text Reader

Abstract

The invention provides a landslide detection method based on Swin Transform and multi-scale feature fusion, and relates to the technical field of landslide image detection, and the method comprises the steps: employing a Swin Transform architecture as a backbone network of a detection system, and effectively capturing local and global context information in an input image through a hierarchical self-attention mechanism; a multi-scale feature fusion lateral connection module is introduced, so that cross-scale feature integration is realized, and the capture capability of the model on landslide feature details and the understanding of the model on wider context information are improved; and the local information aggregation module is adopted to enhance the processing precision of local information. Through the advantages of the Swin Transform architecture, the problems of insufficient local and global feature fusion, low multi-scale feature utilization efficiency, poor complex scene adaptability and the like in landslide detection are effectively solved, and high accuracy and high efficiency of landslide detection are realized; particularly, the method shows excellent performance in the aspects of long-distance global dependence mining and landslide and non-landslide area distinguishing, and shows high robustness to environment and topographic changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of landslide image detection, and particularly to a landslide detection method based on Swin Transformer and multi-scale feature fusion. Background Art

[0002] Landslide detection is a core task in geological disaster prevention and control, and its accuracy is directly related to the efficiency of disaster warning and emergency response.

[0003] Currently, landslide detection mainly relies on optical, radar, and LiDAR remote sensing data, combined with traditional computer interpretation methods and early machine learning techniques. Traditional computer interpretation methods include pixel-based and object-based algorithms. The former identifies landslides through spectral feature classification, but is vulnerable to spectral confusion; the latter divides homogeneous regions through features such as texture and morphology, which can alleviate the problem of edge mixing, but requires a large number of manual parameter tuning and relies on professional experience. Early machine learning techniques such as random forest and support vector machine have improved in feature selection, but have limited ability to model complex terrain and non-linear relationships, and are difficult to adapt to multi-scale landslide morphology. Although deep learning techniques such as CNN, UNet, and Res-UNet have significantly improved the accuracy of landslide detection, there are still key problems such as insufficient fusion of local and global features, low utilization efficiency of multi-scale features, poor adaptability to complex scenarios, and bottlenecks in computational efficiency and generalization. For example, CNN is difficult to model long-range dependencies, while Vision Transformer may lead to blurred boundaries; existing models fuse multi-scale features through skip connections, but do not explicitly optimize cross-scale information interaction, resulting in missed detection of small-scale landslides; in high-resolution images, landslide areas are often confused with shadows, vegetation cover, or exposed rock and soil, and existing models are prone to false detection; deep networks and attention mechanism-based models face problems of gradient disappearance and high computational complexity, and are difficult to be deployed in real-time monitoring systems. In recent years, research has attempted to improve in the directions of multi-scale feature fusion, attention mechanism optimization, and introduction of Transformer architecture, but there are still deficiencies.

[0004] To address these problems, the present invention proposes a landslide detection method that integrates Swin Transformer, Local Information Aggregation Module (LIAM), and Multi-scale Feature Fusion Lateral Connection Module (MFFLCM), aiming to overcome the limitations of traditional methods and existing deep learning models through global-local feature collaboration, cross-scale dynamic fusion, and computational efficiency optimization, and provide a high-precision and high-robustness solution for landslide detection under complex terrain and multi-environmental interference. Summary of the Invention

[0005] To solve the technical problems in the prior art, such as insufficient fusion of local and global features in landslide detection, low utilization efficiency of multi-scale features, poor adaptability to complex scenarios, and bottlenecks in computational efficiency and generalization, the present invention provides a landslide detection method based on Swin Transformer and multi-scale feature fusion.

[0006] The technical solution provided by the present invention is as follows:

[0007] A landslide detection method based on Swin Transformer and multi-scale feature fusion provided by the present invention includes:

[0008] S1. Construct an encoder-decoder network architecture, where the encoder uses Swin Transformer as the backbone network to extract multi-level features of the input remote sensing image;

[0009] S2. Introduce a local information aggregation module LIAM at the end of the encoder to enhance local information of the deep features output by the encoder through a channel attention mechanism;

[0010] S3. Introduce a multi-scale feature fusion lateral connection module MFFLCM in the decoder to perform cross-scale fusion of the features of each stage of the encoder and the features output by LIAM;

[0011] S4. Restore the resolution of the feature map through upsampling and convolution operations to generate the final landslide detection result.

[0012] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0013] (1) In the present invention, by combining the Swin Transformer architecture with the multi-scale feature fusion lateral connection module and the local information aggregation module, the accuracy of landslide detection is effectively improved, and the local and global context information of the input image can be captured simultaneously, thereby enhancing the feature representation ability of the model for landslide features and achieving high-precision landslide detection;

[0014] (2) In the present invention, the introduced multi-scale feature fusion lateral connection module enables the model to integrate features from different scales, capturing both fine detail information and broader context information related to landslide characteristics, which helps the model to stably perform landslide detection under complex and variable terrain and environmental conditions;

[0015] (3) In the present invention, the proposed local information aggregation module further optimizes the model's ability to distinguish between landslide and non-landslide areas. By focusing on the aggregation of local information within the region of interest, the model performs excellently in tasks such as edge detection, reducing false alarms and missed detections, improving the accuracy of landslide prediction, reducing the classification error rate, and providing more reliable technical support for the landslide monitoring and early warning system. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0017] Figure 1 Schematic flow chart of a landslide detection method based on Swin Transformer and multi-scale feature fusion provided by an embodiment of the present invention;

[0018] Figure 2 Schematic overall structure diagram of a landslide detection method based on Swin Transformer and multi-scale feature fusion provided by an embodiment of the present invention;

[0019] Figure 3 Schematic encoder architecture diagram of Swin Transformer of a landslide detection method based on Swin Transformer and multi-scale feature fusion provided by an embodiment of the present invention;

[0020] Figure 4 Schematic architecture diagram of the local information aggregation module LIAM of a landslide detection method based on Swin Transformer and multi-scale feature fusion provided by an embodiment of the present invention;

[0021] Figure 5 Schematic architecture diagram of the multi-scale feature fusion lateral connection module MFFLCM of a landslide detection method based on Swin Transformer and multi-scale feature fusion provided by an embodiment of the present invention;

[0022] Figure 6 Visual comparison diagram of the results of different deep learning models of a landslide detection method based on Swin Transformer and multi-scale feature fusion provided by an embodiment of the present invention;

[0023] Figure 7This is a visualization comparison graph of the results of different deep learning algorithms for a landslide detection method based on Swin Transformer and multi-scale feature fusion provided by an embodiment of the present invention. Detailed implementation manners

[0024] The following describes the technical solutions in the present invention with reference to the accompanying drawings.

[0025] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0026] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when their differences are not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when their differences are not emphasized, the meanings they express are the same.

[0027] In the embodiments of the present invention, sometimes subscripts such as W1 may be miswritten as non-subscript forms such as W1. When their differences are not emphasized, the meanings they express are the same.

[0028] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0029] Refer to the attached specification Figure 1 , which shows a schematic flow chart of a landslide detection method based on Swin Transformer and multi-scale feature fusion provided by an embodiment of the present invention.

[0030] Refer to the attached specification Figure 2 , which shows a schematic overall structure diagram of a landslide detection method based on Swin Transformer and multi-scale feature fusion provided by an embodiment of the present invention.

[0031] The embodiments of the present invention provide a landslide detection method based on Swin Transformer and multi-scale feature fusion. The processing flow may include the following steps:

[0032] S1. Construct an encoder-decoder network architecture, where the encoder uses Swin Transformer as the backbone network to extract multi-level features of the input remote sensing image.

[0033] It should be noted that the core concept of Swin Transformer lies in its way of processing large-size images. Instead of flattening the image pixels into a sequence, it divides the image into overlapping windows, called "shifted windows". This block-based method significantly reduces the computational and memory requirements, making the processing of large-size images more efficient. Within each window, the self-attention mechanism is applied to enable pixels to interact and correlate with each other, thereby capturing local features. The block architecture of each stage of Swin Transformer is as Figure 3 shown.

[0034] S2. Introduce a Local Information Aggregation Module (LIAM) at the end of the encoder to enhance the local information of the deep features output by the encoder through the channel attention mechanism.

[0035] It should be noted that the encoder based on Swin Transformer shows limitations in processing local information, especially in its final output stage. Specifically, the model fails to fully consider the inter-window correlations within each local window, resulting in incomplete local information, which may affect the performance of capturing comprehensive global context information. In the final output stage, the model ignores small pixel-level changes, especially at the image edges and details, limiting its ability to extract fine edge details, especially in high-resolution image processing. To address these limitations, this embodiment proposes LIAM as shown in Figure 4 shown. LIAM aims to aggregate the local information of the input features through the channel attention mechanism. This module calculates the global average pooling result of the features through adaptive average pooling and fully connected layers, and generates channel weights through linear transformation. These weights are then applied to different channels of the original features to highlight important information. Extended to the spatial dimension, the channel weights are multiplied element-wise with the input features to obtain the aggregated features. This method enables LIAM to effectively integrate the information of the local regions around each pixel, significantly improving the accuracy and stability of the detection task. With the introduction of LIAM, the network of this embodiment can better capture local information effectively from different regions of the image, thereby enhancing its ability to perceive local details. This ultimately improves the overall performance of the edge detection task.

[0036] S3. Introduce a Multi-Scale Feature Fusion Lateral Connection Module (MFFLCM) in the decoder to perform cross-scale fusion of the features of each stage of the encoder and the features output by LIAM.

[0037] It should be noted that as Figure 5As shown, to enhance the integration of feature extraction by the encoder and the decoder, this embodiment introduces the MFFLCM module into the method, which plays a key role in the horizontal connection process.

[0038] S4. Restore the resolution of the feature map through upsampling and convolution operations to generate the final landslide detection result.

[0039] In a possible implementation manner, the Swin Transformer encoder is divided into four stages, each stage respectively includes 2, 2, 6, and 2 Swin Transformer blocks, and adopts the alternating mechanism of window multi-head self-attention W-MSA and shifted window multi-head self-attention SW-MSA.

[0040] In a possible implementation manner, the local information aggregation module LIAM specifically includes:

[0041] Perform global average pooling on the input features;

[0042] Generate channel weight coefficients through a fully connected layer;

[0043] Perform element-wise multiplication of the channel weights and the input features, and output locally enhanced features.

[0044] In a possible implementation manner, the multi-scale feature fusion horizontal connection module MFFLCM specifically includes:

[0045] Perform 1×1 convolution channel transformation on the features of each stage of the encoder;

[0046] Perform dimension reconstruction on the output features of LIAM to match the resolutions of each stage of the encoder;

[0047] Perform element-wise addition and ReLU activation on the transformed encoder features and the reconstructed LIAM features to generate fused features.

[0048] In a possible implementation manner, the input remote sensing image is preprocessed and segmented into fixed-size image patches of 16×16 pixels, and sequences are generated through linear embedding and input into the Swin Transformer encoder.

[0049] In a possible implementation manner, the binary cross-entropy loss function BCEWithLogitsLoss is used for model training, and the calculation formula of the loss function is:

[0050] loss = -[y·log(σ(z))+(1 - y)·log(1 - σ(z))]

[0051] where σ is the Sigmoid function, z is the model output, and y is the true label.

[0052] In a possible implementation, the method is applicable to high-resolution satellite images with a spatial resolution of not less than 3 meters, and the input image size is 512×512 pixels.

[0053] In a possible implementation, during the training process, the method adopts a cosine annealing learning rate scheduling and warm-up strategy, and uses the AdamW optimizer for parameter update.

[0054] In a possible implementation, the method realizes pixel-level recognition of landslide areas by fusing local details and global context information, and outputs detection results containing refined boundary annotations.

[0055] In a possible implementation, the method is integrated into a landslide real-time monitoring and early warning system for geological disaster risk assessment and emergency response.

[0056] In the encoder-decoder network structure of this embodiment, the encoder uses Swin Transformer as the backbone network, which is divided into 4 stages (S1 to S4). Each stage contains 2, 2, 6, and 2 Swin Transformer blocks respectively, and the number of feature map channels is 24, 48, 96, and 192 in sequence.

[0057] The decoder gradually restores the spatial resolution through upsampling and integrates the MFFLCM module to achieve cross-scale feature fusion.

[0058] In the input preprocessing of this embodiment, first, the size of the input remote sensing image is 512×512×3, which is divided into fixed-size image patches of 16×16 pixels. The image patches are converted into sequence inputs to the encoder through a linear embedding layer, and the formula is:

[0059] X embed =W p ·X patch +b p

[0060] where W p is a learnable projection matrix, and b p is a bias term.

[0061] In the core module design and algorithm of this embodiment, it includes a Swin Transformer encoder, a local information aggregation module (LIAM), and a multi-scale feature fusion lateral connection module (MFFLCM).

[0062] In the Swin Transformer encoder, each Swin Transformer block adopts an alternating mechanism of window multi-head self-attention (W-MSA) and shifted window multi-head self-attention (SW-MSA), and the calculation formula is:

[0063]

[0064] Among them, Q, K, and V are the query, key, and value matrices respectively, and d k is the dimensionality scaling factor.

[0065] The input of the Local Information Aggregation Module (LIAM) is the feature map at the end of the encoder

[0066] The specific operation process is as follows:

[0067] S201. Global average pooling:

[0068] S202. Channel weight generation: W = σ(τ(G))

[0069] Among them, τ is a fully connected layer, and σ is a Sigmoid activation function;

[0070] S203. Feature enhancement: L out = W⊙S l

[0071] Output the locally enhanced feature

[0072] The input of the Multi-Scale Feature Fusion Lateral Connection Module (MFFLCM) is the features at each stage of the encoder and the output of LIAM.

[0073] The specific operation process is as follows:

[0074] S301. Encoder feature channel transformation: ( 1×1 convolution);

[0075] S302. LIAM feature dimensionality reconstruction: L′ out = μ(L out )(μ: interpolation to adjust the resolution);

[0076] S303. Feature fusion:

[0077] Among them, is element-wise addition, is 1×1 convolution.

[0078] In this embodiment, the loss function uses the binary cross-entropy loss function (BCEWithLogitsLoss):

[0079]

[0080] Among them, N is the total number of pixels, σ is the Sigmoid function, and z i is the model output, and y i ∈{0,1} is the true label.

[0081] In this embodiment, the training parameters are as follows:

[0082] Hardware: NVIDIA RTX Titan GPU (24GB video memory), Intel Xeon Gold 5218 CPU, 64GB memory.

[0083] Optimizer: AdamW, initial learning rate 0.001, cosine annealing scheduling and warm-up strategy.

[0084] Batch size: 8, maximum number of iterations 500 epochs, early stopping strategy (terminate training if the validation loss does not decrease for 10 consecutive rounds).

[0085] In the experimental data of this embodiment, the dataset uses data from the landslide area in the southwestern part of Taiwan Province, China, with a range of 320 km 2 range, Planet Scope satellite images (3-meter resolution). The data is divided into 1670 512×512 images, with 1330 for the training set, 170 for the validation set, and 170 for the test set.

[0086] The evaluation metrics include:

[0087] Precision: The precision rate measures the proportion of samples that are actually positive among the samples predicted as positive by the model. The formula is:

[0088]

[0089] Among them, TP (True Positive) represents the number of samples correctly predicted as positive by the model, and FP (False Positive) represents the number of samples incorrectly predicted as positive by the model.

[0090] Recall: The recall rate measures the proportion of samples that are actually positive and are correctly predicted as positive by the model. The formula is:

[0091]

[0092] Among them, FN (False Negative) represents the number of samples incorrectly predicted as negative by the model.

[0093] mIoU: Mean Intersection over Union, mIoU measures the degree of overlap between the predicted region and the true region. The formula is:

[0094]

[0095] Among them, C is the number of categories, TP c is the True Positive of category c, FP c is the False Positive of category c, FN c is the False Negative of category c.

[0096] F1 Score: The F1 score is the harmonic mean of precision and recall, and the formula is:

[0097]

[0098] Kappa Coefficient: The Kappa coefficient measures the consistency between the classification result and random classification, and the formula is:

[0099]

[0100] Among them, P o is the observed consistency, that is, the proportion of correct classifications, P c is the expected consistency, that is, the expected correct proportion of random classification, and the calculation formula is:

[0101]

[0102] Accuracy: The accuracy measures the proportion of samples correctly predicted by the model in the total samples, and the formula is:

[0103]

[0104] Among them, TN (True Negative) represents the number of samples correctly predicted as negative by the model.

[0105] As Figures 6 - 7 shown, the model provided in this embodiment extracts a robust representation of the landslide by combining local and global context modeling, thereby being able to effectively detect subtle landslides. In addition, the results predicted by the model provided in this embodiment perform excellently in depicting fine-grained boundaries. For example, in the area highlighted by the yellow rectangle, the model provided in this embodiment accurately captures subtle details. In Figure 6 and Figure 7When the landslide scale is larger and more obvious, the model provided by this embodiment can identify most landslides, and the boundary definition is clearer. Compared with the existing CNN-based methods, the model provided by this embodiment adopts the Swin Transformer architecture and combines a customized module for multi-level feature fusion and local information aggregation. This unique combination enables the model provided by this embodiment to effectively capture local and global context information, thus performing excellently in landslide detection performance, especially in overcoming the challenges of long-distance global dependencies and differentiating landslide areas from non-landslide areas. In addition, the model provided by this embodiment shows significantly fewer false positive samples compared with other deep learning algorithms, indicating higher accuracy in landslide prediction and reducing misclassification.

[0106] Metrics of this method: mIoU 84.21%, F1 90.76%, Kappa 82.63%, Precision 89.90%, Recall 91.97%.

[0107] Comparison model: Superior to Deeplabv3+, U-Net, DANet, etc. (see Table 1), especially outstanding in boundary refinement and false detection rate control.

[0108] Table 1 Comparative evaluation metrics

[0109]

[0110] Table 1 shows the quantitative results. Due to the larger landslide scale and more significant nature in the southwestern part of Taiwan, China, the model provided by this embodiment achieved the highest values in all evaluation metrics, exceeding other methods by 84.2%, 90.7%, 82.6%, 89.9%, and 91.9% in terms of mIoU, F1-score, kappa, precision, and recall, respectively. These results indicate that our model significantly reduces false positive samples compared with other state-of-the-art methods.

[0111] This method can accurately identify the landslide boundary in complex terrain areas (such as vegetation-covered areas and bare rock and soil areas), reducing misjudgment of shadows and textures. The method in this embodiment can be integrated into the geological disaster monitoring system to support real-time landslide detection and early warning. It is applicable to the processing of high-resolution satellite images (such as Sentinel-2, GF series) and can be extended to other surface deformation detection tasks.

[0112] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:

[0113] (1) In the present invention, by combining the Swin Transformer architecture with a multi-scale feature fusion lateral connection module and a local information aggregation module, the accuracy of landslide detection is effectively improved. The model can simultaneously capture the local and global context information of the input image, thereby enhancing the model's feature representation ability for landslide features and achieving high-precision landslide detection.

[0114] (2) In the present invention, the introduced multi-scale feature fusion lateral connection module enables the model to integrate features from different scales, capturing both fine detail information and broader context information related to landslide characteristics. This helps the model to stably perform landslide detection under complex and variable terrain and environmental conditions.

[0115] (3) In the present invention, the proposed local information aggregation module further optimizes the model's ability to distinguish between landslide and non-landslide areas. By focusing on local information aggregation within the region of interest, the model performs excellently in tasks such as edge detection, reducing false positives and false negatives, improving the accuracy of landslide prediction, reducing the classification error rate, and providing more reliable technical support for landslide monitoring and early warning systems.

[0116] The above content is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

[0117] The following points need to be explained:

[0118] (1) The attached drawings of the embodiments of the present invention only relate to the structures involved in the embodiments of the present invention, and other structures can refer to the general design.

[0119] (2) For clarity, in the attached drawings used to describe the embodiments of the present invention, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn to actual scale. It can be understood that when an element such as a layer, film, region, or substrate is referred to as being "on" or "under" another element, the element can be "directly" on or under the other element or there can be intermediate elements.

[0120] (3) Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.

[0121] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A landslide detection method based on Swin Transformer and multi-scale feature fusion, characterized in that Including: S1. Construct an encoder-decoder network architecture, where the encoder uses Swin Transformer as the backbone network to extract multi-level features of the input remote sensing image; S2. Introduce a Local Information Aggregation Module (LIAM) at the end of the encoder to enhance local information of the deep features output by the encoder through a channel attention mechanism; S3. Introduce a Multi-Scale Feature Fusion Lateral Connection Module (MFFLCM) in the decoder to perform cross-scale fusion of the features at each stage of the encoder and the features output by LIAM; S4. Restore the resolution of the feature map through upsampling and convolution operations to generate the final landslide detection result.

2. The landslide detection method based on Swin Transformer and multi-scale feature fusion according to claim 1, wherein Including: The Swin Transformer encoder is divided into four stages, each stage contains 2, 2, 6, 2 Swin Transformer blocks respectively, and adopts an alternating mechanism of Window Multi-Head Self-Attention (W-MSA) and Shifted Window Multi-Head Self-Attention (SW-MSA).

3. A landslide detection method based on Swin Transformer and multi-scale feature fusion according to claim 1, characterized in that, The Local Information Aggregation Module (LIAM) specifically includes: Perform global average pooling on the input features; Generate channel weight coefficients through a fully connected layer; Perform element-wise multiplication of the channel weights and the input features, and output locally enhanced features.

4. A landslide detection method based on Swin Transformer and multi-scale feature fusion according to claim 1, characterized in that The Multi-Scale Feature Fusion Lateral Connection Module (MFFLCM) specifically includes: Perform 1×1 convolution channel transformation on the features at each stage of the encoder; Perform dimension reconstruction on the features output by LIAM to match the resolution of each stage of the encoder; Perform element-wise addition and ReLU activation on the transformed encoder features and the reconstructed LIAM features to generate fused features.

5. A landslide detection method based on Swin Transformer and multi-scale feature fusion according to claim 1, characterized in that, Including: The input remote sensing image is preprocessed and segmented into fixed-size image patches of 16×16 pixels, and is input into the Swin Transformer encoder through linear embedding to generate a sequence.

6. A landslide detection method based on Swin Transformer and multi-scale feature fusion according to claim 1, characterized in that, Including: Use the Binary Cross Entropy with Logits Loss function (BCEWithLogitsLoss) for model training, and the calculation formula of the loss function is: loss = -[y·log(σ(z)) + (1 - y)·log(1 - σ(z))] where σ is the Sigmoid function, z is the model output, and y is the true label.

7. A landslide detection method based on Swin Transformer and multi-scale feature fusion according to claim 1, characterized in that Including: The method is applicable to high-resolution satellite images with a spatial resolution of not less than 3 meters and an input image size of 512×512 pixels.

8. A landslide detection method based on Swin Transformer and multi-scale feature fusion according to claim 1, characterized in that, Including: During the training process, the method adopts a cosine annealing learning rate scheduling and warm-up strategy, and uses the AdamW optimizer for parameter update.

9. A landslide detection method based on Swin Transformer and multi-scale feature fusion according to claim 1, characterized in that, Including: The method realizes pixel-level recognition of landslide areas by fusing local details and global context information, and outputs a detection result including refined boundary annotations.

10. A landslide detection method based on Swin Transformer and multi-scale feature fusion according to any one of claims 1-9, characterized in that, Including: The method is integrated into a landslide real-time monitoring and early warning system for geological disaster risk assessment and emergency response.

Citation Information

Cited By

  • Image region-of-interest extraction method and system based on Mama architecture

    CN121353650A

  • Slope supporting structure deformation monitoring method and system based on image recognition

    CN121392530A