Remote sensing landslide and surrounding building detection method and system based on VTAEv2-S model

Through the lightweight ViTAEv2-S model and self-supervised pre-training strategy, the problems of large computing resource consumption and poor multi-task performance in remote sensing landslide detection are solved, and efficient landslide and building detection on resource-constrained equipment is realized, which is suitable for real-time application of disaster monitoring.

CN120388293AActive Publication Date: 2025-07-29CHINA ENENG GRP THIRD ENG BUREAU CO LTD +1

Patent Information

Application Number
CN202510875220.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-29
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The existing remote sensing landslide detection technology consumes a large amount of computing resources, relies on large-scale annotation data, has poor multi-task performance, and is not suitable for equipment with limited computing resources, making it difficult to meet the real-time requirements of disaster monitoring and emergency response.

Method used

The lightweight ViTAEv2-S model is adopted, combining multi-scale hollow convolution and self-attention mechanisms, simplifies the U-Net structure, adds space and channel attention modules, and reduces computing and storage requirements through self-supervised pre-training and multi-task learning strategies, and improves detection efficiency and accuracy.

Benefits of technology

It realizes efficient semantic segmentation of landslides and buildings on resource-constrained equipment, reduces the computational volume and storage requirements, improves the inference speed and detection accuracy, and is suitable for real-time applications of disaster monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388293A_ABST
    Figure CN120388293A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing landslide and surrounding building detection method and system based on a VITAEv2-S model, and belongs to the field of image processing. The method comprises the steps that a detection model is constructed, a lightweight segmentation framework is adopted and comprises an encoder and a decoder, the encoder adopts a VITAEv2-S backbone network, multi-scale cavity convolution and a self-attention mechanism are combined, and the detection model is constructed; extracting multi-scale features; the decoder simplifies a U-Net structure, removes a pyramid pooling module, adds a space attention module to focus landslide and building areas, and adds a channel attention module to enhance feature channels corresponding to the landslide and the building; the detection model is trained; and synchronously segmenting the landslide and the building by using the trained detection model. According to the method, a lightweight improvement scheme is provided for a segmentation framework, semantic segmentation precision is improved, computing resource consumption is remarkably reduced, small sample training ability is improved, synchronous detection of landslides and buildings is realized through multi-task learning, and the method is suitable for disaster real-time monitoring scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of image processing and disaster monitoring, and particularly to a method and system for remote sensing landslide and surrounding building detection based on the ViTAEv2-S model. Background Art

[0002] With the rapid development of remote sensing technology, landslide detection has become a key task in geological disaster monitoring. Compared with other data sources, optical remote sensing has the highest resolution at the sub-meter level, so it can provide more detailed surface information. Research shows that the credibility of detection results highly depends on the quality of the dataset and the considered inducing factors. In addition, high-resolution satellite data can provide updates daily or even hourly, which provides important support for real-time monitoring and response to emergencies such as disasters. Therefore, optical remote sensing images have been widely used in fields such as geological exploration, environmental protection, and earthquake disaster monitoring.

[0003] Currently, semantic segmentation technology is mainly used for remote sensing landslide detection. Semantic segmentation usually uses a convolutional neural network (CNN) or a Transformer architecture for image segmentation. These technologies perform pixel-by-pixel classification through multi-level feature extraction to achieve the detection tasks of landslides and buildings in remote sensing images. These models will use convolutional layers or self-attention mechanisms during the feature extraction process. Convolutional neural networks use convolutional operations to extract low-level and high-level features in images, restore the resolution through pooling and upsampling layers, and finally perform pixel-level classification. Compared with CNN, the Transformer architecture can effectively model long-range context relationships and utilize the self-attention mechanism to process global information in remote sensing images.

[0004] The main methods of semantic segmentation technology include the following steps: 1. Data preprocessing: Perform preprocessing operations such as cropping and normalizing remote sensing images to ensure that the input data is suitable for model training.

[0005] 2. Feature extraction: Use convolutional layers or Transformer layers to extract features of the image and construct semantic representations of landslides and buildings.

[0006] 3. Segmentation decoding: Map the extracted features back to the original image resolution through a decoder to generate a segmented image.

[0007] 4. Loss calculation: Evaluate the difference between the output of the model and the true label through common loss functions (such as cross-entropy loss, Dice coefficient loss, etc.) and optimize the model parameters.

[0008] 5. Inference and post-processing: After training, use the model for inference to generate the final segmentation result, and it can be optimized through post-processing methods such as conditional random field (CRF).

[0009] The existing semantic segmentation technologies have the following defects: 1. High computational resource consumption: Traditional CNN or Transformer models have a large number of parameters, making it difficult to be deployed on low-computing-power devices and not suitable for remote sensing detection scenarios with limited computing resources; 2. Strong data dependence: A large amount of labeled data is required, while the cost of labeling remote sensing data is high; 3. Poor multi-task performance: The training of landslide and building detection tasks is unbalanced, resulting in insufficient model stability; 4. Conflict between efficiency and accuracy: Lightweight models have low accuracy, high-precision models cannot run in real time, the inference speed is not fast enough, and it does not meet the requirements of disaster monitoring and emergency response application scenarios that require quick response. Summary of the Invention

[0010] The purpose of the present invention is to overcome the technical problems existing in the prior art, and provides a method and system for detecting remote sensing landslides and surrounding buildings based on the ViTAEv2-S model. A lightweight improvement scheme is proposed for the segmentation framework, and the semantic segmentation accuracy is improved.

[0011] The purpose of the present invention is achieved through the following technical solutions: In the first aspect, a method for detecting remote sensing landslides and surrounding buildings based on the ViTAEv2-S model is provided, including: Constructing a detection model: Adopting a lightweight segmentation framework, including an encoder and a decoder. The encoder uses the ViTAEv2-S backbone network, combines multi-scale dilated convolution and self-attention mechanism to extract multi-scale features; the decoder simplifies the U-Net structure, removes the UPerNet pyramid pooling module, adds a spatial attention module to focus on the landslide and building areas, and adds a channel attention module to enhance the feature channels corresponding to the landslide and building; Training the detection model; Using the trained detection model to synchronously segment landslides and buildings.

[0012] In some embodiments, the spatial attention module includes a 3×3 convolutional layer and a sigmoid activation function, and the channel attention module includes a global average pooling layer and a fully connected layer.

[0013] In some embodiments, the training of the detection model includes: Training the detection model based on self-supervised pre-training and multi-task learning strategies.

[0014] In some embodiments, the multi-task learning strategy includes: Integrating the landslide and building segmentation tasks into a multi-task model and sharing the backbone network.

[0015] In some embodiments, the loss function adopted by the multitask model includes Dice loss and cross-entropy loss, and a GradNorm loss weight adaptation method and an alternating training strategy are adopted.

[0016] In some embodiments, it further includes: Construct a remote sensing landslide dataset and a surrounding building dataset; Preprocess the images in the remote sensing landslide dataset and the surrounding building dataset respectively.

[0017] In some embodiments, before training the detection model, it further includes: Normalize the input image and convert it into the input format of the detection model.

[0018] In some embodiments, the synchronous segmentation of landslides and buildings using the trained detection model includes: Perform threshold segmentation on the generated probability map, and output a binary mask of landslides or buildings after threshold segmentation.

[0019] In a second aspect, a remote sensing landslide and surrounding building detection system based on the ViTAEv2-S model is provided, including: A detection model construction module for constructing a detection model. The detection model adopts a lightweight segmentation framework, including an encoder and a decoder. The encoder adopts a ViTAEv2-S backbone network, combines multi-scale dilated convolution and self-attention mechanism to extract multi-scale features; the decoder simplifies the U-Net structure, removes the UPerNet pyramid pooling module, adds a spatial attention module to focus on the landslide and building areas, and adds a channel attention module to enhance the feature channels corresponding to landslides and buildings; A training module for training the detection model; A detection module for synchronously segmenting landslides and buildings using the trained detection model.

[0020] It should be further noted that the technical features corresponding to the above embodiments can be combined or replaced with each other without conflict to form a new technical solution.

[0021] Compared with the prior art, the beneficial effects of the present invention are: 1. Model lightweight (1) Reducing the number of parameters and computational complexity: By introducing a simplified U-Net framework and integrating spatial and channel attention mechanisms, the number of parameters and computational complexity of the model are reduced by 52% and 80% respectively. This enables the present invention to run on resource-constrained devices (such as mobile terminals and embedded platforms) without sacrificing segmentation accuracy. (2) Reducing memory occupancy: Compared with the original U-Net segmentation framework, the present invention significantly reduces memory usage, is suitable for embedded platforms with limited memory, and reduces the storage space requirements. It realizes high-efficiency semantic segmentation tasks of landslides and buildings on platforms with limited computing power and storage capacity.

[0022] 2. Higher inference efficiency (1) Improved inference speed: The optimized model not only reduces computational resource consumption but also improves inference speed. This is particularly important for disaster monitoring and emergency response application scenarios that require quick responses. (2) Strong adaptability: In an environment with limited computational resources and storage space, the present invention can operate efficiently, solving the problem of insufficient operation of existing high-precision models on devices with limited computing and storage capabilities.

[0023] 3. Reducing the dependence on a large amount of labeled data and accelerating the training cycle By applying self-supervised pre-training (such as MAE and SparK methods), effective training can be carried out on a small-scale unlabeled dataset. This is particularly important for scenarios where annotation in remote sensing data is difficult and expensive, helping to reduce the cost of data annotation and improve the generalization ability of the model.

[0024] 4. Improving the performance of downstream tasks Experimental results show that for the model based on self-supervised pre-training in landslide and building detection tasks, the F1 scores are improved by 3.15 percentage points (on the landslide dataset) and 2.87 percentage points (on the building dataset) respectively, showing the advantage of the pre-training method in terms of accuracy.

[0025] 5. Improving the multi-task learning ability (1) Synchronous segmentation of landslides and buildings: The present invention integrates the landslide and building segmentation tasks into a multi-task model. By sharing the backbone network, it reduces the computational and storage requirements of the model, while improving the running efficiency and practicality of the model. (2) Solving the overfitting problem: By using the GradNorm method and an alternating training strategy, it alleviates the problem that the landslide task affects the accuracy of the building task due to rapid convergence, ensuring the balance of the segmentation accuracy of landslides and buildings and avoiding interference between tasks during separate training. It realizes the synchronous detection of landslides and buildings and improves the detection efficiency.

[0026] 6. Having real-time performance and applicability The optimization of the present invention enables semantic segmentation tasks to run in real time on platforms with limited computing resources, applicable to actual application scenarios, especially suitable for real-time monitoring and early warning of landslide disasters. This feature enables the technology to be better applied to actual disaster prevention and control work, with important social and economic value. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a flowchart of a remote sensing landslide and surrounding building detection method based on the ViTAEv2-S model of the present invention; Figure 2 It is a schematic diagram of the segmentation framework of the detection model of the present invention; Figure 3 It is the class activation heat map of the present invention; Figure 4 It is a comparison chart of the loss changes of four model structures of the present invention in the pre-training stage. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0029] It should be noted that all the defects existing in the above-mentioned prior art solutions are the results obtained by the inventor after practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the embodiments of the present application below for the above problems should be the contributions made by the inventor to the present application during the invention creation process, and should not be understood as the technical content known to those skilled in the art.

[0030] For the technical problems pointed out in the background art, the embodiments provided by the present invention are as follows: For the technical problems pointed out in the background art, the embodiments provided by the present invention are as follows: Referring to Figure 1 , in an exemplary embodiment, a remote sensing landslide and surrounding building detection method based on the ViTAEv2-S model includes: Constructing a detection model: adopting a lightweight segmentation framework, including an encoder and a decoder. The encoder uses the ViTAEv2-S backbone network, combines multi-scale dilated convolution and self-attention mechanism to extract multi-scale features; the decoder simplifies the U-Net structure, removes the UPerNet pyramid pooling module, adds a spatial attention module to focus on the landslide and building areas, and adds a channel attention module to enhance the feature channels corresponding to the landslide and buildings; Training the detection model; The trained detection model is used to simultaneously segment landslides and buildings.

[0031] Among them, the detection model proposes a lightweight segmentation framework based on the ViTAEv2-S backbone network, and adds spatial attention and channel attention mechanisms on the basis of simplifying the U-Net decoder to suppress spatial interference information and channel interference information in the feature map respectively. Figure 2 As shown, the model inputs a remote sensing image (RGB or multispectral), with an example resolution of 224×224 pixels. The encoder uses the ViTAEv2-S backbone network, with four stages of progressive downsampling (resolution 56×56 → 28×28 → 14×14 → 7×7). Each stage contains a reduction cell (multi-scale dilated convolution) and a normal cell (convolution and self-attention branches in parallel), outputting multi-scale feature maps. The decoder uses a simplified U-Net architecture, consisting of upsampling layers and cross-layer connections. The upsampling layers gradually restore the resolution to the input size, and the cross-layer connections fuse the features of the corresponding encoder stages. A spatial attention module (SAM) is added after the output to suppress spatial noise, and a channel attention module (CAM) is added before the prediction head to enhance the weights of key features.

[0032] This segmentation framework removes the pyramid pooling module from the UPerNet segmentation framework to reduce computational complexity. The spatial attention module includes a 3×3 convolutional layer with a sigmoid activation function, while the channel attention module includes a global average pooling layer and a fully connected layer. SAM (3×3 convolution + sigmoid) focuses on landslides and building areas, while CAM (global average pooling + MLP) suppresses invalid channels. This dual spatial and channel attention mechanism addresses the issues of "ambiguous target region positioning" and "redundant feature channels" in remote sensing imagery. In the lightweight ViTAEv2-S framework, their combined application achieves the following: 1. Spatial dimension: Accurately locate landslide boundaries and building outlines, and suppress background noise; 2. Channel dimension: Screen key spectral features to enhance the distinction between target and background; 3. Efficiency: Improve accuracy at extremely low computing costs, suitable for remote sensing detection scenarios with limited computing resources.

[0033] The datasets used in this invention are as follows: (1) Remote sensing landslide dataset Data source: Remote sensing RGB images of a city, including 770 landslide samples and 2003 negative samples (no landslide), with a resolution of 0.8m.

[0034] Preprocessing: Randomly select 600 negative samples and 770 positive samples to form 1370 samples, and divide them into a training set (961 images) and a validation set (409 images) according to a ratio of 7:3.

[0035] Data augmentation: Random scaling (while maintaining the aspect ratio), random cropping, horizontal flipping, photometric distortion (brightness / contrast adjustment).

[0036] Input size: 224×224 pixels.

[0037] (2)Samples_BuiUrbanVill Building Dataset Data source: Remote sensing RGB images of urban villages in a city in northern China, with a resolution of 0.11m, a total of 2328 samples (512×512 pixels).

[0038] Preprocessing: Eliminate images with large blank areas, randomly select 1370 samples, and divide them into a training set (961 images) and a validation set (409 images) according to a ratio of 7:3.

[0039] Label processing: Convert instance segmentation labels to semantic segmentation labels.

[0040] Data augmentation: Similar to the landslide dataset, but without photometric distortion (to avoid color interference).

[0041] (3)Landslide4Sense Multispectral Landslide Dataset Data source: Multispectral images (Sentinel-2 bands + B13 slope + B14 DEM) from 4 regions around the world, a total of 3799 samples (128×128 pixels).

[0042] Preprocessing: Only use the training set, and divide it into a training set (2661 images) and a validation set (1138 images) according to a ratio of 7:3.

[0043] Input features: 14 bands (including multispectral, slope, DEM), and data augmentation is random vertical flipping (without photometric operations).

[0044] Furthermore, a detection model is trained based on self-supervised pre-training and multi-task learning strategies. Specifically, a pre-training model structure and corresponding training settings are proposed based on the ViTAEv2-S backbone network, experiments are conducted based on small-scale unlabeled remote sensing images, and the SparK self-supervised method is used with the backbone network unchanged.

[0045] For the above dataset, based on the ViTAEv2-S backbone network, the landslide and building segmentation tasks are integrated into a multi-task model, and a multi-task model training strategy is proposed. For landslides, the accuracy of each task in multi-task training is improved mainly by sharing the auxiliary training head and other network layers in the backbone network except the BN layer, and using the Dice loss in the landslide branch. For buildings, the accuracy of each task is further improved by using the GradNorm loss weight adaptation method and the proposed alternating training strategy, and the problem that the accuracy of the building task drops significantly during its later convergence due to the fast convergence of the landslide task is alleviated. Finally, experiments are also conducted on datasets of unequal scales, and the experimental results show that the proposed multi-task training strategy is still effective. The multi-task model can significantly reduce the computational cost, the number of parameters, the memory consumption, and accelerate its running speed by sharing the backbone network, especially the running speed of Pytorch weights on the GPU.

[0046] Before training the detection model, the input images are normalized and converted into the input format of the detection model. The specific training process is as follows: 1. Initialization and parameter settings Pre-trained weights: Use the ViTAEv2-S weights pre-trained on ImageNet-1K (provided officially) and convert them to the MMSegmentation framework.

[0047] Optimizer: AdamW, initial learning rate 0.0006, weight decay 0.01, β1 = 0.9, β2 = 0.999.

[0048] Learning rate strategy: Warm-up stage: Use LinearLR for the first 500 iterations, and the learning rate linearly increases from 0.0006×0.005 to 0.0006.

[0049] Decay stage: Use PolyLR subsequently, and the learning rate decays polynomially to 0 with the decay exponent power = 1.

[0050] Batch size: 48 (GPU: RTX 2080Ti / RTX A4000).

[0051] Loss function: Dice loss + cross-entropy loss (CE).

[0052] 2. Training details and improvement points Efficiency improvement brought by lightweight: Compared with the original UPerNet, the number of parameters is reduced from 48.43M to 23.27M (↓52%), and the computational cost is reduced from 45.60G to 8.92G (↓80%). The GPU training memory consumption is reduced from about 12GB to 8GB, and the inference speed is increased from 32 images / s to 43 images / s.

[0053] Role of the attention mechanism: SAM suppresses background noise in the shallow layer of the encoder, such as non-landslide areas like vegetation and roads. CAM enhances the feature channels corresponding to landslides / buildings in front of the decoder prediction head (e.g., the near-infrared band is sensitive to landslides), improving class discrimination. As Figure 3 shown, the method for generating this heatmap is Grad-CAM, and the target feature layer taken is the SiLU layer in the last convolutional branch of the last stage in the backbone network.

[0054] Multi-scale testing (MS): During testing, the input size is adjusted to 256×256, and 0.5 / 0.75 / 1.0 times scaling and horizontal flipping are performed. By fusing the multi-scale prediction results, the F1 score is increased by 0.37% (remote sensing landslide dataset).

[0055] The following presents a comparison of several pre-trained model structures: As Figure 4 shown, a comparison of the loss changes of four model structures during the pre-training stage is given. Among them, pre-training structure one is the SwinU-Net decoder using the SparK self-supervised method, pre-training structure one (+SE) is pre-training structure one with the SE attention module introduced in its backbone network, pre-training structure two is the classical U-Net decoder using the SparK self-supervised method, and pre-training structure three is a pure Transformer architecture based on the Swin MAE self-supervised method. In Figure 4 , the abscissa is the number of iterations (iters), representing the update rounds of model training, and the total number of iterations is 50,000; the ordinate is the self-supervised loss, measuring the error of the model in reconstructing the features of the unmasked area, and the lower the value, the higher the reconstruction accuracy. It can be seen from the figure that the training loss of pre-training structure one drops the fastest and the final loss is the lowest, indicating that its self-supervised learning efficiency is the highest and its ability to reconstruct the features of remote sensing images is the strongest. The loss of pre-training structure one (+SE) is significantly higher than that of pre-training structure one because the SE module increases the model complexity, resulting in a decrease in training stability. The loss of pre-training structure two drops at a slightly higher speed and the final value is slightly higher than that of structure one because the concatenation operation of the U-Net decoder increases the computational complexity and the convergence efficiency is slightly lower. The loss of pre-training structure three drops slowly and the final value is higher because the pure Transformer architecture has insufficient compatibility with convolutional features and the reconstruction effect of the masked area features is poor.

[0056] Key conclusion: The pre-training structure 1 based on SparK performs optimally in self-supervised pre-training. Adding the SE module actually reduces the performance, verifying the conclusion that "a simple structure is more suitable for few-shot pre-training". Tables 1 - 5 present the pre-training results of each model structure. Among them, IoU represents the intersection over union, which is used to measure the difference between the predicted box and the ground truth box in the segmentation task.

[0057]

[0058]

[0059]

[0060]

[0061]

[0062] It can be seen that the model of the present invention has good effects. It not only greatly reduces the number of parameters and the amount of computation, but also has a higher semantic segmentation detection accuracy than the remaining existing model structures. The pre-training model combining the SparK self-supervised method and the SwinU-Net decoder is optimal. Although the improvement in accuracy on downstream tasks is a bit worse compared to the fully supervised pre-training using the large-scale ImageNet-1K dataset, it is effective and greatly reduces the storage and time costs.

[0063] Furthermore, the trained detection model is used to synchronously segment landslides and buildings. The specific detection process and experimental results are as follows: 1. Inference process Input preprocessing: Normalize the single-band / multi-band image (mean / variance standardization) and convert it to the model input format (such as RGB or multi-spectral tensor).

[0064] Forward propagation: The features extract multi-scale features through the encoder, and the decoder generates pixel-by-pixel classification probabilities through upsampling and the attention module.

[0065] Post-processing: Threshold segmentation of the probability map (such as 0.5) to generate a binary mask (landslide / building is 1, background is 0).

[0066] Optional operation: Morphological filtering (dilation / erosion) to optimize the boundary.

[0067] 2. Experimental results Accuracy comparison (see Tables 6 and 7): The improved model has an F1 score of 86.9% on the remote sensing landslide dataset, which is better than Swin-T-UPerNet (86.72%) and ViTAEv2-S-UPerNet (86.14%). On the building dataset, the F1 score is 92.47% and the IoU reaches 85.99%, surpassing U-Net (87.77% F1) and ConvNeXt-S (90.91% F1). It can be seen that the model of the present invention has good effects, not only greatly reducing the number of parameters and the amount of computation, but also having a higher semantic segmentation detection accuracy than the other existing model structures. The pre-trained model of the SparK self-supervised method combined with the SwinU-Net decoder is optimal. Although the improvement in accuracy on downstream tasks is a bit worse compared to full-supervised pre-training using the large-scale ImageNet-1K dataset, it is effective and greatly reduces the storage and time costs.

[0068]

[0069]

[0070] In another exemplary embodiment, based on the same inventive concept as the method embodiment, a remote sensing landslide and surrounding building detection system based on the ViTAEv2-S model is provided, including: A detection model construction module for constructing a detection model. The detection model adopts a lightweight segmentation framework, including an encoder and a decoder. The encoder adopts the ViTAEv2-S backbone network, combines multi-scale dilated convolution and self-attention mechanism to extract multi-scale features; the decoder simplifies the U-Net structure, removes the UPerNet pyramid pooling module, adds a spatial attention module to focus on the landslide and building areas, and adds a channel attention module to enhance the feature channels corresponding to the landslide and buildings; A training module for training the detection model; A detection module for synchronously segmenting landslides and buildings using the trained detection model.

[0071] The above specific embodiments are detailed descriptions of the present invention. It cannot be determined that the specific embodiments of the present invention are only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the inventive concept of the present invention, several simple deductions and substitutions can be made, which should all be regarded as belonging to the protection scope of the present invention.

Claims

1. A method for remotely sensing landslide and surrounding building detection based on the ViTAEv2-S model, characterized in that, Including: Building a detection model: Using a lightweight segmentation framework, including an encoder and a decoder. The encoder adopts the ViTAEv2-S backbone network, combines multi-scale dilated convolution and self-attention mechanism to extract multi-scale features; the decoder simplifies the U-Net structure, removes the UPerNet pyramid pooling module, adds a spatial attention module to focus on landslide and building areas, and adds a channel attention module to enhance the feature channels corresponding to landslides and buildings; Training the detection model; Using the trained detection model to synchronously segment landslides and buildings.

2. The remote sensing landslide and surrounding building detection method based on the ViTAEv2-S model according to claim 1, wherein The spatial attention module includes a 3×3 convolutional layer and a sigmoid activation function, and the channel attention module includes a global average pooling layer and a fully connected layer.

3. A remote sensing landslide and surrounding building detection method based on the ViTAEv2-S model according to claim 1, characterized in that, The training of the detection model includes: Training the detection model based on self-supervised pre-training and multi-task learning strategies.

4. A method for remote sensing landslide and surrounding building detection based on the ViTAEv2-S model according to claim 3, characterized in that, The multi-task learning strategy includes: Integrating the landslide and building segmentation tasks into a multi-task model and sharing the backbone network.

5. A method for remote sensing landslide and surrounding building detection based on the ViTAEv2-S model according to claim 4, characterized in that, The loss function adopted by the multi-task model includes Dice loss and cross-entropy loss, and adopts the GradNorm loss weight adaptive method and the alternating training strategy.

6. The remote sensing landslide and surrounding building detection method based on the ViTAEv2-S model according to claim 1, characterized in that, Also including: Building a remote sensing landslide dataset and a surrounding building dataset; Preprocessing the images in the remote sensing landslide dataset and the surrounding building dataset respectively.

7. A method for remote sensing landslide and surrounding building detection based on the ViTAEv2-S model according to claim 1, characterized in that, Before training the detection model, it also includes: Normalizing the input image and converting it into the input format of the detection model.

8. A method for remote sensing landslide and surrounding building detection based on the ViTAEv2-S model according to claim 1, characterized in that, The using the trained detection model to synchronously segment landslides and buildings includes: Performing threshold segmentation on the generated probability map and outputting a binary mask of the landslide or building after threshold segmentation.

9. A remote sensing landslide and surrounding building detection system based on the ViTAEv2-S model, characterized in that, Including: A detection model construction module for building a detection model. The detection model uses a lightweight segmentation framework, including an encoder and a decoder. The encoder adopts the ViTAEv2-S backbone network, combines multi-scale dilated convolution and self-attention mechanism to extract multi-scale features; the decoder simplifies the U-Net structure, removes the pyramid pooling module, adds a spatial attention module to focus on landslide and building areas, and adds a channel attention module to enhance the feature channels corresponding to landslides and buildings; A training module for training the detection model; A detection module for using the trained detection model to synchronously segment landslides and buildings.

Citation Information

Patent Citations

  • Defect detection method based on joint optimization and mixed attention feature fusion

    CN115294038A

  • Urban streetscape advertisement image segmentation method

    CN116189180A

  • Pyramid cross-layer fusion decoder based on semantic segmentation

    CN116310324A

  • Remote sensing image road segmentation method fusing multi-scale features and double attention mechanism

    CN117078943A

  • Landslide image segmentation method based on multilayer feature information fusion

    CN118261926A

Cited By

  • Dam landslide dangerous case area identification system and method based on unmanned aerial vehicle and remote sensing image

    CN120932100A

  • Slope rockfall detection method and system

    CN121564410A