Landslide identification method, system, equipment and medium

By constructing landslide data sets and using encoders and decoders of neural network models for feature embedding and multi-objective reconstruction, the problems of low accuracy and long processing time in the prior art are solved, and rapid and accurate landslide identification and efficiency improvement of geological disaster prevention are achieved.

CN120107813AActive Publication Date: 2025-06-06CHANGAN UNIV

Patent Information

Application Number
CN202510160459.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-06
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

In the landslide recognition, the model and data training process requires a lot of processing resources and time, and due to the quality of the training data, the recognition accuracy is low and the early warning cannot be carried out quickly and accurately.

Method used

By obtaining the remote sensing image data and digital elevation data of the target slope segment, a landslide data set is constructed, and the data is divided into n×n pixel tokens, and input into a neural network model based on the encoder and decoder structure. The neural network model is optimized to achieve fast and accurate landslide recognition using the feature embedding of the encoder and decoder and multi-objective reconstruction strategy.

Benefits of technology

Through encoding, decoding and reconstruction loss optimization, the entire process time from data annotation to model training and image interpretation is shortened, pixel-level recognition capabilities are provided, rapid and accurate acquisition of landslide distribution is improved, and the efficiency of geological disaster prevention and emergency response is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107813A_ABST
    Figure CN120107813A_ABST
Patent Text Reader

Abstract

The invention discloses a landslide identification method, system, equipment and medium, and relates to the technical field of disaster detection, and the method comprises the steps: obtaining remote sensing image data and digital elevation data of a target slope section, and taking the data as a landslide data set; the method comprises the following steps: segmenting image data in a landslide data set into tokens with n * n pixels, inputting the token into a neural network model based on an encoder and decoder structure, and processing the token by an encoder and a decoder to obtain a token of a reconstructed image and reconstruction loss between the token and an original image token, and obtaining an optimal neural network model by taking the minimized reconstruction loss as a target, inputting remote sensing image data and digital elevation data of a real-time target slope section into the optimal neural network model for image interpretation, and obtaining a landslide map in the target slope section. According to the method, image interpretation is carried out by optimizing the neural network model, high precision is displayed through the image interpretation result of the region, and the application potential of the method in the field of disaster prevention and reduction is further verified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of disaster detection technology, and in particular to a landslide identification method, system, equipment and medium. Background Art

[0002] Landslide events are often associated with specific triggers, such as earthquakes, severe storms, rapid snow melt, or volcanic eruptions. Such events can involve anything from a single landslide to thousands of them, and, similar to earthquakes and volcanic activity, they can cause severe geological hazards.

[0003] After a disaster occurs, it is crucial to quickly and accurately assess the possible disaster losses and quickly and accurately obtain information such as the location of the landslide and the scope of the disaster for earthquake disaster emergency rescue and resettlement planning. In the existing automatic identification of rapid landslides through deep learning network models, the model and data training processes require a lot of processing, and are limited by the training data. Its accuracy is greatly affected by the quality of sample data. When processing these workflows, it will consume a lot of resources and time, and it is not possible to quickly and accurately issue warnings. Summary of the invention

[0004] The purpose of the present invention is to provide a landslide identification method, system, equipment and medium to solve the problems in the prior art in view of the above-mentioned deficiencies in the prior art.

[0005] The present invention specifically provides the following technical solutions:

[0006] A landslide identification method comprises the following steps:

[0007] Obtain remote sensing image data and digital elevation data of the target slope section and construct a landslide dataset;

[0008] The image data in the landslide dataset is divided into n×n pixel tokens, and the tokens are input into the neural network model based on the encoder and decoder structure. The encoder is used to embed the tokens not covered by the mask, and the relationship between all tokens is captured through the self-attention mechanism to generate the final output features. The final output features are reconstructed using the multi-objective reconstruction strategy of the decoder to obtain the tokens of the reconstructed image and the reconstruction loss between the tokens and the original image. The optimal neural network model is constructed with the goal of minimizing the reconstruction loss.

[0009] The real-time remote sensing image data and digital elevation data of the target slope section are input into the optimal neural network model for image interpretation to obtain the landslide map in the target slope section.

[0010] Preferably, the step of acquiring remote sensing image data and digital elevation data of the target slope section and constructing a landslide data set includes:

[0011] Define the time range before and after the earthquake of the target slope section, filter the remote sensing image data according to the time range, and load the digital elevation data;

[0012] The pre-earthquake images were filtered according to time, region and corner points to obtain a set of images that completely covered the study area, and edge masks were applied to remove edge noise. The image with the least cloud was selected from the filtered images and cropped to the target slope segment.

[0013] The remote sensing images of the target slope section were subjected to single-band image separation and NVID processing to obtain the RGB images before and after the earthquake and the NDVI before and after the earthquake. The DEM-related factors of the digital elevation data were calculated to obtain the hillside shadow and slope. The RGB images before and after the earthquake, the NDVI before and after the earthquake, the hillside shadow and the slope were stacked into a multi-band image, and the multi-band image was used as the landslide dataset.

[0014] Preferably, before segmenting the image data in the landslide dataset into n×n pixel tokens, the method further includes:

[0015] Obtain the NDVI images before and after the earthquake, and add them as additional channels of the RGB images before and after the earthquake as image data to be segmented; the specific expression of the added NDVI image is:

[0016]

[0017] Among them, NIR and RED are the reflectances of the near-infrared and red bands, respectively.

[0018] Preferably, the relationship between all tokens is captured by the self-attention mechanism, and the specific expression is:

[0019] Q i =x i W Q ,K i =x i W K ,V i =x i W V ;

[0020]

[0021] z i =Attention(Q i ,K i ,V i )=S i V i ;

[0022] Where d represents the dimension of embedding, Qi For query, K i is the key, V i is the value, x i is the input, S i is the attention score, z i is the output feature, W Q Q i The corresponding weight matrix is ​​used to convert the input hidden state into the query vector Q i , W K K i The corresponding weight matrix is ​​used to convert the input hidden state into the key vector K i , W V V i The corresponding weight matrix is ​​used to convert the input hidden state into the key vector V i , K i The transpose of .

[0023] Preferably, the multi-objective reconstruction strategy of the decoder is used to reconstruct the final output features to obtain the token of the reconstructed image and the reconstruction loss between the token of the original image, including:

[0024] The final output features are reconstructed using the multi-objective reconstruction strategy of the decoder to obtain the token of the reconstructed image; the specific expression is:

[0025]

[0026] in, is the token of the reconstructed image, g φ is the decoder, M is the mask matrix, specifically a binary matrix used to indicate which tokens are masked and which tokens should be retained, x is the token of the original image, and ⊙ is the element-by-element multiplication or Hadamard product;

[0027] Get the reconstruction loss between the token of the reconstructed image and the token of the original image. The specific expression is:

[0028]

[0029] Among them, L is the overall reconstruction loss, L token is the reconstruction loss from token to token, L spectral is the reconstruction loss from spectral to spectral, and λ represents the loss used to adjust L token and L spectral The combined hyperparameters, m is the masked token (H p ×W p), x is the token of the original image, are the tokens of the reconstructed image, i, j, (r, c, 1) and Both represent indices, n is the number of spectral tokens, and vis is the index of the token not covered by the mask, for example, x vis ={x i |i∈vis} represents the token not covered by the mask, and j corresponds to the summation symbol A summation variable, representing the iteration index in a certain range. Specifically, j starts from 1 and ends at n, indicating that each item in this range is summed. Each spectral token consists of (D is the channel dimension of the input image, k is the channel dimension of the token) standard tokens, and (r,c) represents the token in the rth row and cth column in the spectral data.

[0030] The present invention provides a landslide identification system, comprising:

[0031] The acquisition module is used to obtain remote sensing image data and digital elevation data of the target slope section and construct a landslide data set;

[0032] The model training module is used to segment the image data in the landslide dataset into n×n pixel tokens, and input the tokens into the neural network model based on the encoder and decoder structure. The encoder is used to embed the tokens not covered by the mask, and the relationship between all tokens is captured through the self-attention mechanism to generate the final output features. The final output features are reconstructed using the multi-objective reconstruction strategy of the decoder to obtain the tokens of the reconstructed image and the reconstruction loss between the tokens and the original image tokens, and the optimal neural network model is constructed with the goal of minimizing the reconstruction loss.

[0033] The output module is used to input the real-time remote sensing image data and digital elevation data of the target slope section into the optimal neural network model for image interpretation to obtain the landslide map in the target slope section.

[0034] The present invention provides a computer device, comprising a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of the above-mentioned landslide identification method.

[0035] The present invention provides a storage medium on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the above-mentioned landslide identification method are implemented.

[0036] Compared with the prior art, the present invention has the following significant advantages:

[0037] The present invention divides the image data in the landslide data set into tokens of n×n pixels, inputs the tokens not covered by the mask into the neural network model, uses the encoder and decoder of the neural network model to process the remote sensing image data and digital elevation data of the target slope section, and performs image reconstruction optimization neural network model based on the reconstruction loss optimization of the reconstructed image and the original image. The encoding, decoding and reconstruction loss optimization realizes the shortening of the entire process time from data annotation to model training and then to image interpretation, and provides the ability of pixel-level recognition by processing the n×n pixel tokens, which provides a fast and accurate basis for the subsequent acquisition of landslide distribution. The real-time data is input into the trained model for interpretation, so as to obtain the landslide map of the target slope section, thereby improving the efficiency of geological disaster prevention and emergency response. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 The present invention provides an overall process of a landslide identification method;

[0039] Figure 2 Visualize the results for the data: Figure 2 (a) is the image before and after the earthquake (RGB), Figure 2 (b) is the pre- and post-earthquake image (NDVI). Figure 2 (c) is the hillside shadow, Figure 2 (d) is the slope;

[0040] Figure 3 is the accuracy index curve on the validation set during the SpectralGPT model training process; Figure 3 (a) is the curve under recall condition, Figure 3 (b) is the curve under the Precision condition. Figure 3 (c) is the curve under the F1 Score condition, Figure 3 (d) is the curve under IoU conditions, Figure 3 (e) is the curve under Mean IoU condition;

[0041] Figure 4 To complete the deployment of models on the PIE Engine AI platform (by uploading or online model training);

[0042] Figure 5 Schematic diagram of the workflow of the SpectralGPT model and its migration to downstream tasks;

[0043] Figure 6The network architecture for semantic segmentation (top) and change detection (bottom) for the downstream tasks of the SpectralGPT model.

[0044] Figure 7 Image (left) and DEM (right) data of the study area (Bijie City); red dots are identified landslide locations;

[0045] Figure 8 These are different types of landslide examples in the study area, each with a different sample size. A 40-meter extension is retained as the background for each example. Figure 8 (a1)- Figure 8 (f6) are examples of different types of landslides;

[0046] Fig. 9 Images of some areas before and after the earthquake: Fig. 9 (a) is Hokkaido, Fig. 9 (b) is Jiuzhaigou, China;

[0047] Fig.10 YOLOv5 model structure;

[0048] Fig.11 This is the interface saved in the model library after YOLOv5 is deployed, showing information such as accuracy and network structure;

[0049] Fig.12 To use the PIE-Engine AI image intelligent processing platform for intelligent image interpretation;

[0050] Fig.13 The landslide prediction results in the Hokkaido study area (yellow represents the actual location of the landslide, and red represents the predicted result);

[0051] Fig.14 Landslide prediction results for the Hokkaido study area (partial);

[0052] Fig.15 The landslide prediction results in the Jiuzhaigou study area (yellow represents the actual location of the landslide, and red represents the prediction results);

[0053] Fig.16 This is the landslide prediction result (partial) in the Jiuzhaigou study area. DETAILED DESCRIPTION

[0054] The following is a clear and complete description of the technical solutions of the embodiments of the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0055] The combination of PIE Engine cloud platform and deep learning model landslide identification technology aims to overcome the above shortcomings and provide a more efficient, lower-cost and more timely landslide identification solution. This improves the existing landslide identification workflow and brings better experience and results to users. This paper proposes an online intelligent mapping solution to address the timeliness issue of landslide mapping.

[0056] like Figure 1 As shown, the present invention provides a landslide identification method for description below, which specifically includes the following steps:

[0057] Step S1: Acquire remote sensing image data and digital elevation data of the target slope section and construct a landslide dataset.

[0058] First, the Google Earth Engine (GEE) platform is used to acquire and preprocess Sentinel-2 image data, i.e., remote sensing image data. The specific data processing flow of the GEE code operation process is as follows:

[0059] (1) Define the area and initialize parameters:

[0060] Define the study area: Define a rectangular area based on the bounding rectangle of the study area (such as Jiuzhaigou or Hokkaido) to frame the study area and clip the image. Define the rectangular area by geographic coordinates to ensure that the image data in the study area can be accurately clipped and analyzed.

[0061] Set time parameters: define the time range before and after the earthquake of the target slope section, filter remote sensing image data according to the time range, set a specific time period, and ensure that the filtered image collection can accurately reflect the environmental changes before and after the earthquake. And load digital elevation data.

[0062] (2) Load the dataset:

[0063] Load SRTM DEM data: Load SRTM elevation data.

[0064] Load Sentinel-2 image data: Load the Sentinel-2 image dataset s2Sr and cloud probability dataset s2Clouds.

[0065] Define edge mask function: Define edge mask function to remove edge noise.

[0066] (3) Filtering and preprocessing image data:

[0067] Filter pre- / post-earthquake images: filter pre-earthquake images according to time, region, and corner points to obtain a set of images that completely cover the study area, and apply edge masks to remove edge noise; select the image with the least clouds: select the image with the least clouds from the filtered images and crop it to the target slope section (area).

[0068] (4) Image de-clouding:

[0069] Associate cloud probability data: Associate cloud probability data with Sentinel-2 image data and add cloud masks.

[0070] Apply cloud mask: During the image processing, use the cloud removal function to remove the cloud from the image. This step is optional, and the specific implementation depends on whether the cloud cover of the image with the least cloud in the selected time range affects the use of the image. When the cloud cover of the image with the least cloud in the selected time range does not affect the use of the image, directly select the image with the least cloud. If the image with the least cloud in the selected time range cannot meet the requirements, call the cloud removal function to obtain a composite image of multiple images after cloud removal.

[0071] (5) Calculate and display NDVI:

[0072] The remote sensing images of the target slope section were subjected to single-band image separation and NVID processing to obtain the RGB images before and after the earthquake and the NDVI before and after the earthquake (using the normalizedDifference function of GEE to calculate the NVID before and after the earthquake), and the visualization parameters of the NDVI were set to display it on the map. This step is not only used to calculate the NDVI value, but also to ensure the accuracy of the obtained NDVI data through visual inspection.

[0073] (6) Image normalization and resampling:

[0074] Define the target resolution and projection: Set the target resolution to 10 meters and the target projection to UTM Zone48N / UTMZone 54N (Jiuzhaigou / Hokkaido).

[0075] Check and resample bands: Define resampling and reprojection functions, check the resolution of the image, and resample and reproject. If the resolution of the image is inconsistent with the target resolution, use resample('bicubic') to resample and unify the resolution to 10m*10m for subsequent processing.

[0076] Calculate the actual maximum and minimum values ​​of NDVI: Calculate the actual maximum and minimum values ​​of NDVI for normalization processing of NDVI and visualization of NDVI in GEE.

[0077] Normalize the image: Normalize the image based on the calculated value range (such as NDVI) or the known data value range. This step ensures that data from different bands are compared and analyzed on the same scale, thereby improving the accuracy and consistency of data processing. Training the model on the normalized dataset can accelerate model convergence and reduce training time. The normalized dataset scales the input data to the same value range, making it more stable and efficient during the model training process.

[0078] (7) Stack bands and export images:

[0079] And the DEM related factors of the digital elevation data were calculated to obtain the hillside shadow and slope, and the bands were stacked: the normalized bands were arranged in order, a total of 10 channels, and the order was as follows: the hillside shadow, post-earthquake image RGB, post-earthquake NDVI, pre-earthquake image RGB and slope were stacked into a multi-band image, and the multi-band image was used as the landslide dataset.

[0080] Rename Band: Set a new band name for the stacked image to facilitate subsequent use.

[0081] Export images: Export the stacked images as GeoTIFF files and set parameters such as resolution and projection coordinate system to ensure that the exported data meets the requirements.

[0082] Step S2: Divide the image data in the landslide dataset into tokens of n×n pixels, and input the tokens into a neural network model based on an encoder and decoder structure (the SpectralGPT model in this embodiment), use the linear projection matrix of the encoder to embed the tokens not covered by the mask, capture the relationship between all tokens through the self-attention mechanism, generate the final output features of the encoder, and use the multi-objective reconstruction strategy of the decoder to reconstruct the final output features, obtain the tokens of the reconstructed image and the reconstruction loss between the tokens of the original image, and obtain the optimal neural network model with the goal of minimizing the reconstruction loss.

[0083] (1) Dataset construction and model selection:

[0084] In order to fine-tune a semantic segmentation model specifically for identifying landslides, the present invention refers to the SegMunich dataset structure used for fine-tuning semantic segmentation downstream tasks and constructs a Jiuzhaigou landslide dataset with the same structure. Figure 2 The data visualization results are shown: (a) pre- and post-earthquake images (RGB), (b) pre- and post-earthquake images (NDVI), (c) hillside shadows, and (d) slopes. Tokens are vector representations of images divided into small blocks, which are used to represent local areas of the image.

[0085] In the selection of pre-training models, the present invention adopts the SpectralGPT+ model that performs progressive pre-training on multiple datasets. This model shows better performance on multiple datasets.

[0086] (2) Model fine-tuning and parameter setting

[0087] The present invention mainly utilizes the semantic segmentation function in the downstream task of SpectralGPT to realize the pixel-level recognition of landslides, thereby completing landslide mapping. The SpectralGPT model workflow is as follows Figure 5 As shown, in the pre-training stage, SpectralGPT first trains the model from scratch on a dataset (e.g., fMoW-S2, containing 712,874 images), using random weights initialized based on (3D) tensors. Subsequently, the model is progressively trained on more datasets (e.g., BigEarthNet-S2, containing 354,196 images) with different image sizes, time series information, and geographic regions. SpectralGPT is built following the MAE architecture and combines spectral tensor (3D) masks, where 90% of the labels are masked. For downstream tasks such as classification, segmentation, and change detection, the pre-trained SpectralGPT is connected to the task-specific head network, trained, and fine-tuned. The process of fine-tuning SpectralGPT on the Jiuzhaigou landslide dataset is as follows:

[0088] 1) Data preparation: The present invention collects optical data from the Sentinel-2 satellite and the digital elevation model (DEM) of the Shuttle Radar Topography Mission (SRTM) to create the Jiuzhaigou landslide dataset for fine-tuning a semantic segmentation model for landslide identification. Based on the minimum cloud cover principle, two pairs of Level-1C images (one pair each for Jiuzhaigou and Hokkaido) were selected for the dual-phase data corresponding to each study area. The Normalized Difference Vegetation Index (NDVI), as a key indicator reflecting the growth status of vegetation, is easily affected by the loss of vegetation that may be caused by landslides. Therefore, the NDVI images before and after the event are added as additional channels to overcome the shortcomings of using only RGB spectral data to detect landslides.

[0089] Before segmenting the image data in the landslide dataset into n×n pixel tokens, it also includes:

[0090] Obtain the NDVI images before and after the earthquake, and add them as additional channels of the RGB images before and after the earthquake as image data to be segmented; the specific expressions of the added NDVI images before and after the earthquake are as follows:

[0091]

[0092] Among them, NIR and RED are the reflectances of the near-infrared and red bands, respectively.

[0093] Slope is a key topographic factor that directly affects landslide movement. Hillshade is another topographic factor widely used in landslide mapping. Hillshade map and slope are extracted from DEM, and these topographic factors are added as auxiliary data to enhance model performance. Finally, there are ten input bands of the same resolution (10 meters) for each study area, which are arranged in the following stacking order: hillshade, post-earthquake BGR band, post-earthquake NDVI, pre-earthquake BGR band, pre-earthquake NDVI, slope.

[0094] 2) Data segmentation and enhancement: The image data is segmented into 128×128 pixel tokens. The dataset is divided into a training set and a validation set in a ratio of 7:3, and data enhancement techniques are performed, including random flipping and rotation.

[0095] 3) Fine-tuning parameters: During fine-tuning, a batch size of 96 is used and the base learning rate is set to 5×10 -4 , Warmup-epochs is set to the first 30 epochs to reduce the oscillation and instability that may occur in the early stage of training and reduce the risk of overfitting. Weight-decay is set to 5×10 -3 , to enhance the generalization ability of the model and prevent the model from overfitting the training data.

[0096] 4) Fine-tuning the model: During the fine-tuning process, a specific head network needs to be connected to SpectralGPT, such as Figure 6 As shown in Figure 2. In this architecture, based on the SpectralGPT pre-trained model, the present invention further trains the UperNet head network to support the downstream tasks of semantic segmentation (top) and change detection (bottom). Among them, MLP stands for multi-layer perceptron. During the fine-tuning process, the present invention continuously monitors the accuracy indicators on the validation set, such as Figure 5 to evaluate model performance and make necessary adjustments.

[0097] Among them, the specific model architecture of SpectralGPT mentioned in the "SpectralGPT model fine-tuning" section is as follows:

[0098] 1) Encoder ViT:

[0099] The encoder of SpectralGPT is defined as follows:

[0100] Input processing: All visible tokens (i.e. tokens not covered by the mask) are first passed through a shared linear projection matrix E s Perform feature embedding transformation. These feature embeddings are consistent with the position encoding Epos Combined to form the input to the encoder f θ final expression of .

[0101] Encoder structure: Encoder f θ It consists of multiple stacked self-attention (SA) Transformer blocks. The input embedding z of each SA module is i Generate query Q through linear transformation i 、Key(Key)K i Sum value V i Embed.

[0102] Self-Attention Mechanism: Query Q i and key K i The attention score S between embeddings i The dot product is calculated and scaled by dimension d, and then normalized by the softmax function. The normalized attention score is used for weighted embedding to generate the final output embedding z i .

[0103] The relationship between all tokens is captured through the self-attention mechanism. The specific expression is:

[0104] Q i =x i W Q ,K i =x i W K ,V i =x i W V ,;

[0105]

[0106] z i =Attention(Q i ,K i ,V i )=S i V i ,;

[0107] Where d represents the dimension of embedding, Q i For query, K i is the key, V i is the value, x i is the input, S i is the attention score, z i is the output feature, W Q Q i The corresponding weight matrix is ​​used to convert the input hidden state into the query vector Q i , W KK i The corresponding weight matrix is ​​used to convert the input hidden state into the key vector K i , W V V i The corresponding weight matrix is ​​used to convert the input hidden state into the key vector V i , K i The transpose of .

[0108] Output feature: The final output feature z i With input x i have the same dimensions and can be further processed by subsequent encoders.

[0109] The main function of the encoder is to convert the input visible tokens into feature embeddings and capture the relationship between these tokens through the self-attention mechanism. The encoder consists of multiple stacked self-attention Transformer blocks, each of which processes the input embeddings through linear transformations and self-attention mechanisms to generate the final output features. These output features can be further used for subsequent tasks such as classification, segmentation, etc.

[0110] 2) Decoder MAE and loss function:

[0111] The decoder of SpectralGPT is defined as follows:

[0112] Input and target: Given the features z output by the encoder, train a lightweight decoder g at the same time φ , using a multi-objective reconstruction strategy to restore the token of the original image. The final output features are reconstructed using the decoder's multi-objective reconstruction strategy to obtain the token of the reconstructed image and the reconstruction loss between the token and the original image, including:

[0113] The final output features are reconstructed using the multi-objective reconstruction strategy of the decoder to obtain the token of the reconstructed image; the specific expression is:

[0114]

[0115] in, is the token of the reconstructed image, g φ is the decoder, M is the mask matrix, specifically a binary matrix used to indicate which tokens are masked and which tokens should be retained, x is the token of the original image, and ⊙ is the element-by-element multiplication or Hadamard product.

[0116] Decoder structure: Decoder g φTypically narrower and shallower than the encoder, usually consisting of several Transformer blocks and a linear reconstruction layer.

[0117] Training method: The proposed SpectralGPT model trains the encoder f in an end-to-end manner θ and decoder g φ , to minimize the reconstructed image token The reconstruction loss between the token x of the original image can be obtained. The reconstruction loss between the token of the reconstructed image and the token of the original image is obtained. Reconstruction loss: The reconstruction loss consists of two parts: token-to-token and spectral-to-spectral. This multi-objective reconstruction allows the learned representation to effectively capture the spatial-spectral coupling characteristics and spectral sequence information. The overall loss LL is quantified in the pixel space using the mean square error (MSE), and the specific expression is:

[0118]

[0119]

[0120] Among them, L is the overall reconstruction loss, L token is the reconstruction loss from token to token, L spectral is the reconstruction loss from spectral to spectral, and λ represents the loss used to adjust L token and L spectral The combined hyperparameters, m is the masked token (H p ×W p ), x is the token of the original image, is the token of the reconstructed image, i,j,(r,c,1), etc. all represent indices, n is the number of spectral tokens, and vis is the index of the token not covered by the mask, for example, x vis ={x i |i∈vis} represents the token not covered by the mask, and j corresponds to the summation symbol A summation variable, representing the iteration index in a certain range. Specifically, j starts from 1 and ends at n, indicating that each item in this range is summed. Each spectral token consists of (D is the channel dimension of the input image, k is the channel dimension of the token) standard tokens, and (r,c) represents the token in the rth row and cth column in the spectral data.

[0121] (3) Model deployment:

[0122] After fine-tuning is completed, a Docker image is built based on the model and used as the operating environment. Subsequently, the present invention binds the Docker image to the trained model and uploads it to the PIE EngineAI platform to complete the deployment of the model. Figure 4 shown.

[0123] This deployment model can be used for online prediction and intelligent image interpretation, providing efficient and accurate solutions for application scenarios such as landslide identification.

[0124] Step S3: inputting the real-time remote sensing image data and digital elevation data of the target slope section into the optimal neural network model for image interpretation to obtain the landslide map in the target slope section.

[0125] In another embodiment of the present invention, YOLOv5 model training deployment is adopted, specifically:

[0126] (1) Construction of cloud platform dataset:

[0127] Since YOLOv5 needs to be trained directly on the PIE Engine AI platform, an online training dataset that meets the requirements of the platform must be produced. To this end, the present invention uses the sample collaborative annotation platform of PIE Engine AI to produce a cloud platform Bijie landslide dataset that fully meets the platform training requirements based on the Bijie landslide dataset, providing support for the online training model.

[0128] The Bijie landslide dataset contains TripleSat satellite images taken from May to August 2018, as well as shape files and digital elevation models (DEM) of landslide boundaries. The dataset includes 770 landslide samples, including rock collapses, rock slides, and a few debris landslides, as well as 2003 negative samples with different backgrounds. Figure 7 The red dots in the figure are the actual locations of landslides. The landslide samples vary in size, and a 40-meter buffer is added to the bounding box of each sample as a background. This helps the deep learning network better learn the characteristics of landslides. The RGB image resolution in the dataset is 0.8 meters, and the elevation accuracy of the DEM is 2 meters. The landslide shape vectors were manually drawn using ArcGIS, and the non-landslide samples included mountains, villages, roads, rivers, forests, and farmland. These samples were used to comprehensively evaluate the effectiveness of the landslide detection method. Some landslide instances in the area are shown in Figure 8 middle.

[0129] In addition, when training the model, the label format of the Bijie landslide dataset (binary images in .png format as masks) cannot be directly used for model training, so the corresponding label conversion python code is written to traverse all files in the original label directory and convert them into the format required for subsequent training. Finally, the various labels and images are divided into training sets, validation sets, and test sets, and the dataset that can be used to directly train the model is reorganized according to the directory format required by the dataset.

[0130] (2) Online model training and deployment:

[0131] In the case of landslide mapping without the need for high-precision prediction results, if only the landslide needs to be quickly identified, selecting a target detection model with a small number of parameters and capable of rapid prediction can meet the needs. For this application scenario, the present invention adopts the target detection model YOLOv5 built into the PIE Engine AI platform. Thanks to the built-in model structure and training environment of the PIE cloud platform, the present invention only needs to set parameters in the platform model training interface and select the online data set produced by the present invention to quickly train the required model and achieve one-click deployment, greatly improving efficiency. Utilizing the rapid prediction characteristics of YOLOv5, the present invention can quickly identify the location of the landslide, thereby efficiently completing the task.

[0132] Specifically, YOLOv5 (You One Look Once version 5) mentioned in the “(2) Online Model Training and Deployment” section is a single-stage target detection algorithm that can complete the target detection task in one neural network forward propagation. The YOLOv5 network structure is a standard CSPDarknet+PAFPN+non-decoupled Head. The overall structure of the model is as follows: Fig.10 As shown in the figure, the upper part is an overview of the model, the middle part is the specific composition of each submodule, and the lower part is the specific network structure. The network structure mainly includes three parts: backbone, neck and detection head:

[0133] (1) Model Backbone:

[0134] The overall structure of CSPDarknet is similar to ResNet. The P5 model structure contains 5 layers, including 1 StemLayer and 4 Stage Layers.

[0135] The Stem Layer uses a 6x6 kernel ConvModule, which is more efficient than the Focus module before v6.1. Except for the last Stage Layer, the other Stage Layers are composed of a Con vModule and a CSPLayer. The specific structure is shown in the Details section of the figure above. Among them, ConvModule is composed of Conv2d, BatchNorm and SiLU activation functions. CSPLayer is the C3 module in the YOLO v5 official repository, which consists of 3 ConvModules and n Darknet Bottleneck with residual connections.

[0136] The final Stage Layer adds the SPPF module. The SPPF module works by passing the input serially through multiple 5x5 MaxPool2d layers. It can achieve the same effect as the SPP module while running faster. The model outputs a feature map after Stage Layer 2-4 and enters the Neck structure. Taking a 640x640 input image as an example, its output features are (B, 256, 80, 80), (B, 512, 40, 40), and (B, 1024, 20, 20), and the corresponding strides are 8, 16, and 32, respectively, where B represents the batch size.

[0137] (2) Model neck:

[0138] YOLOv5 uses the PAFPN structure, combined with the CSP2 structure designed by CSPNet and PANET as the neck for feature aggregation. The neck is mainly used to generate feature pyramids, enhance the model's detection ability for objects of different scales, and realize the recognition of objects of different sizes and scales.

[0139] (3) Model detection head:

[0140] The head structure of YOLOv5 is exactly the same as that of YOLOv3, and adopts a non-decoupled head design. The head module contains only three convolutional layers that do not share weights, which are used to transform the input feature map. PAFP N (Path Aggregation Feature Pyramid Network) still outputs three feature maps of different scales, whose shapes are (B, 256, 80, 80), (B, 512, 40, 40) and (B, 1024, 20, 20).

[0141] Since YOLOv5 uses a non-decoupled output method, tasks such as classification and bounding box detection are completed in different channels of the same convolution. Taking the COCO 80-category dataset as an example, when the input resolution is 640×640, the feature map shapes output by the Head module are: (B, 3×(4+1+80), 80, 80), (B, 3×(4+1+80), 40, 40) and (B, 3×(4+1+80), 20, 20). Among them, 3 represents 3 anchors, 4 represents the bounding box (bbox) prediction branch, 1 represents the object existence (obj) prediction branch, and 80 represents the category prediction branch of the COCO dataset.

[0142] (4) Design of loss function:

[0143] YOLOv5 contains a total of 3 Loss, namely:

[0144] Classes loss: BCE loss is used.

[0145] Objectness loss: BCE loss is used.

[0146] Location loss: CIoU loss is used.

[0147] The three losses are summarized according to a certain ratio:

[0148] Loss = λ 1 L cls +λ 2 L obj +λ 3 L loc .

[0149] (5) The model processes the data as it passes through it:

[0150] In the data preprocessing stage of YOLOv5, a variety of data enhancement techniques are used, such as Mosaic, RandomAffine, and MixUp. These techniques increase the diversity of data and improve the generalization ability of the model by randomly transforming the training data. In addition, the image pixel values ​​are normalized to the range of [0,1] to facilitate neural network processing.

[0151] The method further comprises:

[0152] The image data in the landslide dataset is input into the YOLOv5 model. In terms of model structure, YO LOv5 uses CSPDarknet as the feature extraction network (Backbone) to extract multi-scale features of the image. Features of different scales are further fused through PAFPN (Path Aggregation Network) (Neck).

[0153] The detection head uses a non-decoupled output method to perform image data classification and bounding box detection in different channels of the same convolution, and outputs the bounding box coordinates, category probability, and confidence score of each target.

[0154] The image is parsed using the bounding box coordinates, class probability, and confidence score of each target to obtain a landslide map in the target slope segment.

[0155] During training, YOLOv5 uses a composite loss function, including bounding box regression loss, classification loss, and target confidence loss. Adam or SGD optimizers are usually used to update model parameters. During inference, the input image is fed into the model after the preprocessing steps. The backbone network extracts image features, the neck network fuses features, and the head network generates predictions, including bounding boxes and class probabilities. Finally, non-maximum suppression (NMS) is used to remove overlapping prediction boxes and retain the best detection results. Through the above steps, YOLOv5 can quickly and efficiently detect target objects in the input image and output the bounding box coordinates, class probability, and confidence score of each target.

[0156] Finally, the dataset is used to complete the training and deployment of the YOLOv5 model on the platform. The deployed YOLOv5 model can be used for online prediction to achieve intelligent image interpretation. During the training process, 660 rounds of training are performed with reference to the parameter settings of offline model training. After the training is completed, the trained model is saved to the model library of the cloud platform. The model training and deployment results on the cloud platform are shown in Figure 2. Fig.11 shown.

[0157] Intelligent interpretation of images:

[0158] Through the PIE-Engine AI image intelligent processing platform, users can select the deployed model and upload the image to be processed. After completing the intelligent interpretation of the image using the powerful computing power of the cloud platform, users can view the interpretation results online. Fig.12 In addition, users can also interpret the results (such as Figures 13 to 16) is downloaded from the platform to the local computer and imported into ArcGIS and other software for further analysis or other tasks. In order to evaluate the generalization ability of the model, the present invention downloads the prediction results to the local computer and combines the landslide cataloging in the test set area to calculate the accuracy index of the model in the completely unseen area (Hokkaido), as shown in Table 1.

[0159] Technical effect: YOLOv5 deployed on the PIE-Engine AI platform takes only 2 minutes and 24 seconds to interpret an image of size 5888×6313, while SpectralGPT takes only 4 minutes and 19 seconds to interpret an image of size 3350×2867, fully demonstrating the efficiency of deep learning models deployed on the PIE cloud platform. In addition, the system's image interpretation results in completely unseen areas show high accuracy, further verifying its application potential in the field of disaster prevention and mitigation.

[0160] Table 1 Precision indicators of the prediction results of the trained SpectralGPT model in the Hokkaido area

[0161] category IoU Precision Recall background 0.9462 0.9744 0.9703 landslide 0.4008 0.5543 0.5914 mean 0.6735 0.7643 0.7809

[0162] Based on the above method, the present invention provides a landslide identification system, including: a collection module, a model training module and an output module.

[0163] Among them, the acquisition module is used to obtain the remote sensing image data and digital elevation data of the target slope section and construct a landslide data set; the model training module is used to segment the image data in the landslide data set into n×n pixel tokens, and input the tokens into the neural network model based on the encoder and decoder structure, and use the encoder to embed the tokens not covered by the mask, and capture the relationship between all tokens through the self-attention mechanism to generate the final output features, and use the decoder's multi-objective reconstruction strategy to reconstruct the final output features, obtain the tokens of the reconstructed image and the reconstruction loss between the tokens and the original image, and build the optimal neural network model with the goal of minimizing the reconstruction loss; the output module is used to input the real-time remote sensing image data and digital elevation data of the target slope section into the optimal neural network model for image interpretation, and obtain the landslide map in the target slope section.

[0164] The present invention also provides a computer device, comprising a memory and a processor. The memory stores a program, and when the program is executed by the processor, the processor executes the steps of a landslide identification method.

[0165] In accordance with the disclosed embodiments, a computing device may communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth communications, etc.), or with any device (e.g., routers, modems, etc.) that enables a computing device to communicate with one or more other computing devices.

[0166] The present invention also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of a landslide identification method are implemented.

[0167] According to the disclosed embodiments, the storage medium may be a non-volatile computer-readable storage medium, such as but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0168] The above content is a further detailed description of the present invention in combination with a specific preferred embodiment. For technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the protection scope of the present invention.

Claims

1. A landslide identification method, characterized in that: include: Obtain remote sensing image data and digital elevation data of the target slope section and construct a landslide dataset; The image data in the landslide dataset is divided into n×n pixel tokens, and the tokens are input into the neural network model based on the encoder and decoder structure. The encoder is used to embed the tokens not covered by the mask, and the relationship between all tokens is captured through the self-attention mechanism to generate the final output features. The final output features are reconstructed using the multi-objective reconstruction strategy of the decoder to obtain the tokens of the reconstructed image and the reconstruction loss between the tokens and the original image. The optimal neural network model is constructed with the goal of minimizing the reconstruction loss. The real-time remote sensing image data and digital elevation data of the target slope section are input into the optimal neural network model for image interpretation to obtain the landslide map in the target slope section.

2. A landslide identification method according to claim 1, characterized in that: The step of acquiring remote sensing image data and digital elevation data of the target slope section and constructing a landslide data set includes: Define the time range before and after the earthquake of the target slope section, filter the remote sensing image data according to the time range, and load the digital elevation data; The pre-earthquake images were filtered according to time, region and corner points to obtain a set of images that completely covered the study area, and edge masks were applied to remove edge noise. The image with the least cloud was selected from the filtered images and cropped to the target slope segment. The remote sensing images of the target slope section were subjected to single-band image separation and NVID processing to obtain the RGB images before and after the earthquake and the NDVI before and after the earthquake. The DEM-related factors of the digital elevation data were calculated to obtain the hillside shadow and slope. The RGB images before and after the earthquake, the NDVI before and after the earthquake, the hillside shadow and the slope were stacked into a multi-band image, and the multi-band image was used as the landslide dataset.

3. A landslide identification method as claimed in claim 2, characterized in that: Before segmenting the image data in the landslide dataset into n×n pixel tokens, the method further includes: Obtain the NDVI images before and after the earthquake, and add them as additional channels of the RGB images before and after the earthquake as image data to be segmented; the specific expressions of the added NDVI images before and after the earthquake are as follows: Among them, NIR and RED are the reflectances of the near-infrared and red bands, respectively.

4. A landslide identification method according to claim 1, characterized in that: The self-attention mechanism is used to capture the relationship between all tokens. The specific expression is: Q i =x i W Q ,K i =x i W K ,V i =x i W V ; z i =Attention(Q i ,K i ,V i )=S i V i ; Where d represents the dimension of embedding, Q i For query, K i is the key, V i is the value, x i is the input, S i is the attention score, z i is the output feature, W Q Q i The corresponding weight matrix is ​​used to convert the input hidden state into the query vector Q i , W K K i The corresponding weight matrix is ​​used to convert the input hidden state into the key vector K i , W V V i The corresponding weight matrix is ​​used to convert the input hidden state into the key vector V i , K i The transpose of .

5. A landslide identification method according to claim 1, characterized in that: The multi-objective reconstruction strategy of the decoder is used to reconstruct the final output features, obtain the token of the reconstructed image and the reconstruction loss between the token of the original image, including: The final output features are reconstructed using the multi-objective reconstruction strategy of the decoder to obtain the token of the reconstructed image; the specific expression is: in, is the token of the reconstructed image, g φ is the decoder, M is the mask matrix, specifically a binary matrix used to indicate which tokens are masked and which tokens should be retained, x is the token of the original image, and ⊙ is the element-by-element multiplication or Hadamard product; Get the reconstruction loss between the token of the reconstructed image and the token of the original image. The specific expression is: L=L token +λL spectral Among them, L is the overall reconstruction loss, L token is the reconstruction loss from token to token, L spectral is the reconstruction loss from spectral to spectral, and λ represents the loss used to adjust L token and L spectral The combined hyperparameters, m is the masked token (H p ×W p ), x is the token of the original image, are the tokens of the reconstructed image, i, j, (r, c, 1) and Both represent indices, n is the number of spectral tokens, and vis is the index of the token not covered by the mask, for example, x vis =[x i |i∈vis} represents the token not covered by the mask, and j corresponds to the summation symbol A summation variable, representing the iteration index in a certain range. Specifically, j starts from 1 and ends at n, indicating that each item in this range is summed. Each spectral token consists of (D is the channel dimension of the input image, k is the channel dimension of the token) standard tokens, and (r,c) represents the token in the rth row and cth column in the spectral data.

6. A landslide identification system, characterized in that: include: The acquisition module is used to obtain remote sensing image data and digital elevation data of the target slope section and construct a landslide data set; The model training module is used to segment the image data in the landslide dataset into n×n pixel tokens, and input the tokens into the neural network model based on the encoder and decoder structure. The encoder is used to embed the tokens not covered by the mask, and the relationship between all tokens is captured through the self-attention mechanism to generate the final output features. The final output features are reconstructed using the multi-objective reconstruction strategy of the decoder to obtain the tokens of the reconstructed image and the reconstruction loss between the tokens and the original image tokens, and the optimal neural network model is constructed with the goal of minimizing the reconstruction loss. The output module is used to input the real-time remote sensing image data and digital elevation data of the target slope section into the optimal neural network model for image interpretation to obtain the landslide map in the target slope section.

7. A computer device, characterized in that: The method comprises a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of a landslide identification method as claimed in any one of claims 1 to 5.

8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a landslide identification method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Landslide disaster automatic identification method based on lightweight deep learning network

    CN117911866A

  • Sky-sky-ground landslide monitoring method and system based on multi-source data

    CN118171918A

  • Landslide recognition method based on laplacian pyramid remote sensing image fusion

    US11521377B1

  • System and method for joint detection, localization, segmentation and classification of anomalies in images

    WO2024102565A1

Cited By

  • Method, medium and equipment for analyzing landslide influence based on landslide boundary vector diagram

    CN120279427A

  • Landslide identification method and system, computer equipment and medium

    CN120877125A