Landslide intelligent identification model and method based on large model prior knowledge

By adopting a landslide intelligent recognition model based on large-model prior knowledge in landslide recognition, combined with technical means of CLIP model, PKI module, encoder and CFA module, the difficulty of identifying the existing model under the conditions of data scarcity is solved, and high-precision and high-efficiency landslide recognition is achieved.

CN120032269AActive Publication Date: 2025-05-23CHENGDU UNIV

Patent Information

Application Number
CN202510519550.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-23
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing deep learning models are limited by the number and diversity of training data in landslide recognition tasks, making it difficult to fully extract landslide features and accurately identify landslide areas under complex terrain.

Method used

The intelligent landslide identification model based on large-model prior knowledge is adopted, and the local and global information of landslide image data is extracted through the CLIP model, and the prior knowledge of landslide is obtained in combination with the PKI module. The encoder is used to extract multi-scale and multi-level image feature information, and the prior knowledge and feature information are fused through the CFA module. Finally, the precise mask of landslide is obtained through the decoder step by step decoding.

Benefits of technology

It effectively improves the accuracy and efficiency of landslide identification, can improve the adaptability and stability of the model under the conditions of scarcity of data, and provides an important scientific basis for disaster prevention and mitigation of landslide disasters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032269A_ABST
    Figure CN120032269A_ABST
Patent Text Reader

Abstract

The invention relates to an intelligent landslide identification model and method based on large model prior knowledge. The method comprises the following steps: firstly, constructing a landslide identification sample library by adopting a visual interpretation mode according to remote sensing image data and the like, and cutting the landslide identification sample library into a plurality of small images; dividing the landslide identification sample library into a training set, a verification set and a test set, and inputting the training set, the verification set and the test set into the model; secondly, a DS Net model is constructed through the structural design of double encoders and a single decoder, and moving window multi-head attention is added into the first encoder so as to extract multi-scale and multi-level feature information of the image; in the other encoder, a PKI (Public Key Infrastructure) module is introduced, a multi-head attention mechanism is applied, and a Siam attention mechanism is integrated at the same time, so that local details and a global context relationship of the image are obtained; and finally, recovering landslide boundary information step by step by adopting a hierarchical up-sampling decoder, and generating a high-precision landslide mask. Therefore, based on large model prior knowledge, the model can effectively improve the precision and efficiency of landslide identification by using a deep learning technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of geological disaster monitoring and identification, and in particular to a landslide intelligent identification model and method based on large model prior knowledge. Background Art

[0002] Landslides are the most common geological disasters, with wide distribution, high frequency, and complex causes. At the same time, they are affected by natural variations and human activities, and various disasters such as traffic interruption, river blockage, farmland damage, building destruction, village burial, and human and animal casualties have become one of the most threatening geological problems for the safety of life and property of local residents, the healthy and sustainable development of regional economy and society, and the construction and operation safety of major projects. Therefore, it is urgent to study the distribution and activity of landslide disasters, and high-precision identification and extraction of landslides is an important basis for studying landslide disasters, which can provide an important reference for the research on landslide disaster prevention and mitigation.

[0003] Since the occurrence of landslides is often accompanied by the influence of other external environmental or geological tectonic factors, the terrain at the landslide site is fragmented and difficult to access. Traditional landslide identification methods mainly rely on field surveys and visual interpretation. Although they can provide intuitive landslide information, they are limited by high costs, long cycles, and accuracy that depends on manual experience. They cannot accurately determine the specific location, scale, and number of landslide disasters in a timely manner. In recent years, with the significant improvement of remote sensing platforms and sensor performance, high-resolution remote sensing images have been widely used in geological disaster identification and investigation research. Therefore, the use of optical remote sensing images to carry out high-precision landslide identification has become one of the hot research issues in the study of landslide disasters.

[0004] Deep learning has made significant progress in the field of image recognition and has been gradually applied to remote sensing image analysis. Compared with traditional machine learning methods, deep learning has greatly improved the ability of image classification and target recognition through multi-layer feature extraction and automatic learning mechanisms. With the enrichment of remote sensing data, landslide recognition technology based on optical images has developed from traditional manual visual interpretation to the stage of intelligent algorithm automatic recognition. However, the occurrence of landslide events is sporadic and spatially heterogeneous, which makes it very difficult to obtain high-quality training samples, thereby limiting the generalization ability of the model. Existing deep learning models are still limited by the number and diversity of training data in landslide recognition tasks, making it difficult to fully extract landslide features and accurately identify landslide areas under complex terrain. Therefore, how to make full use of the prior knowledge of large models and improve the adaptability and stability of landslide recognition models under data scarcity conditions is one of the key issues in the development of current intelligent landslide recognition technology. Summary of the invention

[0005] The present invention provides a landslide intelligent identification model and method based on large model prior knowledge. At the model and method level, a dual coding segmentation network (DS Net) is used to extract features by integrating prior knowledge, which is composed of a CLIP model, a PKI module, a CFA module, an encoder, etc. The DS Net technology is used to identify the precise boundaries of landslides, delicately capture the unique characteristics of landslides, and accurately distinguish between landslide areas and non-landslide areas. The accuracy and efficiency of landslide identification can be effectively improved, and an important scientific basis for landslide disaster prevention and mitigation can be provided.

[0006] This application is implemented through the following technical solutions:

[0007] A landslide intelligent identification model based on large model prior knowledge, including CLIP model, PKI module, CFA module, encoder and decoder;

[0008] The CLIP model is used to extract local and global information of landslide image data in the landslide identification sample library;

[0009] The PKI module is used to integrate the local and global information extracted by the CLIP model and obtain the prior knowledge of the landslide;

[0010] The encoder is used to extract multi-scale and multi-level image feature information of landslide image data in the landslide identification sample library;

[0011] The CFA module is used to fuse the prior knowledge obtained by the PKI module with the feature information obtained by the encoder to obtain an image feature representation containing landslide information;

[0012] The decoder is used for decoding step by step to obtain an accurate mask of the landslide.

[0013] Based on the above model, this application also provides a landslide intelligent identification method based on large model prior knowledge, which includes the following steps:

[0014] S1. Based on remote sensing image data, regional survey data and landslide spectral characteristics, landslide identification sample data is constructed and landslide boundaries are delineated by visual interpretation to obtain a landslide identification sample library;

[0015] S2, extracting local and global information of landslide image data in the landslide identification sample library through the CLIP model, integrating the local and global information extracted by the CLIP model through the PKI module, and obtaining prior knowledge of landslides;

[0016] S3, extracting multi-scale and multi-level image feature information of landslide image data in the landslide identification sample library through an encoder;

[0017] S4, integrating the prior knowledge with the feature information through the CFA module to obtain the image feature representation containing the landslide information;

[0018] S5. Obtain the accurate mask of the landslide by decoding step by step through the decoder.

[0019] In particular, in S1, the landslide identification sample library includes data sets for model training and testing.

[0020] In particular, the data set includes a training set, a validation set and a test set, and is processed by ArcGIS and Python to generate label data required for the sample library and then input into the CLIP model.

[0021] In particular, in S2, the PKI module is based on a multi-head attention mechanism, with the global information extracted by the CLIP model as the query and the local information extracted by the CLIP model as the key and value, so as to enhance the extraction effect of the landslide prior information.

[0022] In particular, in S3, after the landslide image data enters the decoder, it is first divided into blocks through a block structure, and the blocks are flattened in the channel dimension. Then, these feature maps are sequentially passed through four Swin Transformer modules to extract feature information.

[0023] In particular, in the extraction of feature information, in the first processing, each block will be linearly transformed first and then input into two consecutive Swin Transformer modules. In the subsequent processing, the feature map is downsampled by merging adjacent blocks and then input into the next Swin Transformer module.

[0024] In particular, in S4, the CFA module is based on the reverse multi-head attention mechanism, with landslide prior knowledge as query and image feature information as key and value, to further strengthen the expression of landslide prior information in image features.

[0025] In particular, the landslide image data are randomly divided into a training set, a validation set, and a test set in a ratio of 6:2:2, wherein the training set is used for model training to extract features, the validation set is used to evaluate model performance, and the test set is used to evaluate the model.

[0026] In particular, the evaluation model indicators include OA, PA, F1_score and recall.

[0027] In particular, the CFA module incorporates the Siam attention mechanism, which enhances the model's ability to extract key information and reduces redundant information interference by assigning unique weights to neurons.

[0028] In particular, the dual encoding structure Swin Transformer encoder modifies the original multi-head attention structure into a moving window multi-head attention structure, and transmits feature information in different windows by moving the positions of these windows, thereby ensuring the recognition accuracy of the network.

[0029] The present invention proposes a DS Net deep learning model based on large model prior knowledge to effectively capture the multi-scale, multi-level and global information characteristics of landslide images in landslide identification. The present invention effectively combines the advantages of large model prior knowledge, deep learning feature extraction, multi-head attention mechanism and hierarchical decoding technology. Compared with the prior art, the present invention has the following beneficial effects:

[0030] (1) The dual-coding segmentation network (DS Net) that integrates the features extracted from prior knowledge aims to utilize DS Net technology to effectively improve the accuracy and efficiency of landslide identification in terms of identifying the precise boundaries of landslides, capturing the unique characteristics of landslides in detail, and accurately distinguishing landslide areas from non-landslide areas, thus providing an important scientific basis for landslide disaster prevention and mitigation.

[0031] (2) The Swin Transformer is used as the first encoder, and the moving window multi-head attention structure is used to transfer feature information in different windows by moving the positions of these windows. This strategy greatly improves the recognition accuracy of the network. Compared with ResNet50 and ResNet101 as the backbone network, DS Net with Transformer as the backbone network not only steadily improves the verification accuracy and reaches a higher level, but also has a relatively low training loss.

[0032] (3) The DS Net model has shown a significant improvement over other comparison models in the specific task of landslide identification. Compared with other efficient and powerful semantic segmentation models, DS Net is significantly better than SegFormer, SegNeXt, FeedFormer, and U-MixFormerd in terms of OA, PA, F1_score, and recall in landslide identification, with a lead ranging from 3.5% to 7.3%, because of its deep feature fusion strategy, sophisticated information processing mechanism, and ability to accurately capture key landslide information. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings described herein are used to provide a further understanding of the embodiments of the present application, constitute a part of the present application, and do not constitute a limitation on the embodiments of the present invention.

[0034] Figure 1 It is a flow chart of a landslide intelligent identification method based on large model prior knowledge according to an embodiment of the present invention;

[0035] Figure 2 A schematic diagram of the spatial distribution of landslides selected for the embodiment of the present invention;

[0036] Figure 3 A flow chart for making a sample library according to an embodiment of the present invention;

[0037] Figure 4 This is a flow chart of a Swin Transformer encoder in an embodiment of the present invention;

[0038] Figure 5 FIG. 4 is a flow chart of a DS Net network in an embodiment of the present invention. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0040] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0041] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. It should be noted that the various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same and similar parts between the various embodiments can be referred to each other.

[0042] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0043] In the description of the present invention, it should be noted that the terms "upper", "lower", "inside", "outside", etc. indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, or are the orientations or positional relationships in which the inventive product is conventionally placed when in use, or are the orientations or positional relationships conventionally understood by those skilled in the art. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.

[0044] In this embodiment, the CLIP model, whose full name in English is Contrastive Language-Image Pre-training, is a multimodal model;

[0045] PKI module, the full name of which is Prior Knowledge Integration, is the prior knowledge integration module;

[0046] The CFA module, whose full English name is Cross Fusion Attention, is a cross-fusion attention module.

[0047] Example 1

[0048] A landslide intelligent identification model based on large model prior knowledge, characterized by comprising a CLIP model, a PKI module, a CFA module, an encoder and a decoder;

[0049] The CLIP model is used to extract local and global information of landslide image data in the landslide identification sample library;

[0050] The PKI module is used to integrate the local and global information extracted by the CLIP model and obtain the prior knowledge of the landslide;

[0051] The encoder is used to extract multi-scale and multi-level image feature information of landslide image data in the landslide identification sample library;

[0052] The CFA module is used to fuse the prior knowledge obtained by the PKI module with the feature information obtained by the encoder to obtain an image feature representation containing landslide information;

[0053] The decoder is used for decoding step by step to obtain an accurate mask of the landslide.

[0054] The purpose of this design is to use the dual coding segmentation network (DS Net) that relies on the CLIP model, PKI module, CFA module, encoder, etc. to extract features by integrating prior knowledge. It aims to use DS Net technology to effectively improve the accuracy and efficiency of landslide identification in identifying the precise boundaries of landslides, delicately capturing the unique characteristics of landslides, and accurately distinguishing landslide areas from non-landslide areas, providing an important scientific basis for landslide disaster prevention and mitigation.

[0055] Example 2

[0056] like Figure 1 As shown, a landslide intelligent identification method based on large model prior knowledge includes the following steps: S1, constructing landslide identification sample data and outlining landslide boundaries by visual interpretation according to remote sensing image data, regional survey data and landslide spectral characteristics, and performing data conversion, data cutting and other processes to prepare a landslide identification sample library;

[0057] In this embodiment, Figure 2 As shown, taking the spatial distribution map of landslides in Beichuan County as an example, it is located in the transition zone from the northwest edge of the Sichuan Basin to the western Sichuan Plateau, with a land area of ​​3082.72 square kilometers, and the landslide boundary prediction for the entire region has been completed.

[0058] This landslide identification sample library uses high-resolution remote sensing images with a spatial resolution of 0.5 meters, combined with the spectral characteristics of landslides in the study area, and uses visual interpretation to construct landslide identification sample data. The landslide surface vector data is delineated with the landslide contour as the boundary to generate label data. The sample library production process includes vector editing, field assignment, vector-raster conversion, label generation, data cutting, and sample data set division. The specific process is as follows: Figure 3 shown.

[0059] Based on the existing landslide hazard interpretation signs, the vector editing tool in ArcGIS software was used to edit the TIFF format remote sensing image, and the boundary was manually drawn by visual interpretation. Since the accuracy of boundary delineation will directly affect the quality of the landslide sample library, in order to improve the accuracy of landslide samples and model training, the error of vector delineation needs to be strictly controlled in the process of delineating vectors, and the error range of the delineated vector boundary should be controlled within 1-2 pixels.

[0060] ArcGIS software was used to assign attributes to the landslide interpretation data, and the attribute field "label" was added to the delineated landslide boundary data. The attribute field of the landslide was assigned a value of "1", and the non-landslide background was assigned a value of "0".

[0061] Python is used in combination with the GDAL library to convert vector files into raster files according to the assigned field values, read the corresponding information, form a binary map, and finally form the label data required for vegetation sample library construction, and ensure that the number of rows and columns of the formed raster data is consistent with that of the image data.

[0062] Due to the limited computer memory, the entire remote sensing image has too much information to be used as input for the network model. The remote sensing image and the corresponding landslide label data need to be cut into multiple small pictures to ensure that they can be input into the model for training. In order to make the remote sensing image processing process easier and enable the network model to better extract the landslide details and effectively converge, the data set is split into 256×256 pixels.

[0063] After the sample segmentation operation, a high-resolution remote sensing landslide sample library consisting of 4879 images was obtained. In order to ensure the diversity of samples and make the data set as diverse as possible, the obtained data was randomly divided into training set, validation set and test set in a ratio of 6:2:2, of which 2927 images of the training set were used for model training and feature extraction, 976 images of the validation set were used for evaluating model performance, and 976 images of the test set were used for evaluating the model.

[0064] S2, extracting local and global information of landslide image data in the landslide identification sample library through the CLIP model, integrating the local and global information extracted by the CLIP model through the PKI module, and obtaining prior knowledge of landslides;

[0065] Among them, the CLIP model can mine the global and local information of landslide images. By introducing the PKI module, we effectively integrate the local and global information provided by CLIP to obtain the prior knowledge of landslides. At the same time, the multi-head attention mechanism is used in the process of information integration and fusion. In the integration stage of global and local information, the global information is used as the query (q), and the local information is used as the key (k) and value (v), so as to enhance the extraction effect of landslide prior information.

[0066] In this way, the PKI module is introduced to integrate the local and global information provided by CLIP to obtain the prior knowledge of landslides, and the multi-head attention mechanism is used to enhance the expression of prior information.

[0067] S3, extracting multi-scale and multi-level image feature information of landslide image data in the landslide identification sample library through an encoder;

[0068] The image is first flattened in blocks in the encoder, and then features are extracted step by step through four steps. Each step is processed by linear transformation and Swin Transformer module, and the merged adjacent blocks are used for downsampling in subsequent steps to reduce the size of feature maps.

[0069] Specifically, after entering the encoder, the image will be divided into blocks through a slicing structure and flattened in the channel dimension. Then, these feature maps will be extracted through four steps in sequence. In the first step, each block will be linearly transformed and then input into two consecutive Swin Transformer modules. In the subsequent steps, in order to further extract features and reduce the size of the feature map, the feature map will be downsampled by merging adjacent blocks and then input into the next Swin Transformer module. The specific structure is as follows: Figure 4 .

[0070] S4, integrating the prior knowledge with the feature information through the CFA module to obtain the image feature representation containing the landslide information;

[0071] In the fusion stage of feature information and prior knowledge, the roles are reversed, with landslide prior information as the query (q) and image feature information as the key (k) and value (v), thereby further strengthening the expression of landslide prior information in image features. The specific structure is as follows: Figure 5 .

[0072] In this way, the image features extracted by Swin Transformer are deeply integrated with prior knowledge in the CFA module to optimize the expression of landslide information;

[0073] The Siam attention mechanism is integrated into the CFA module. By assigning unique weights to neurons, the model can more accurately focus on information that is critical to the current task and effectively ignore irrelevant or minor parts.

[0074] S5. Obtain the accurate mask of the landslide by decoding step by step through the decoder.

[0075] In the specific implementation process, in terms of model training parameter setting, 100 training rounds were set for the Swin Transformer and Dual-coded Segmentation Network experiments to fully explore the learning ability of the model and avoid overfitting; the batch size was uniformly set to 50 to balance memory utilization and model convergence speed; the remote sensing image size of the network input was 256×256×3, which meets the format requirements of most deep learning models for input data; the stochastic gradient descent (SGD) optimizer was used to optimize the learning rate to promote the iterative update of the model in the direction of minimizing the loss function.

[0076] The above specific implementation methods further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above are only specific implementation methods of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A landslide intelligent identification model based on large model prior knowledge, characterized in that: Includes CLIP model, PKI module, CFA module, encoder and decoder; The CLIP model is used to extract local and global information of landslide image data in the landslide identification sample library; The PKI module is used to integrate the local and global information extracted by the CLIP model and obtain the prior knowledge of the landslide; The encoder is used to extract multi-scale and multi-level image feature information of landslide image data in the landslide identification sample library; The CFA module is used to fuse the prior knowledge obtained by the PKI module with the feature information obtained by the encoder to obtain an image feature representation containing landslide information; The decoder is used for decoding step by step to obtain an accurate mask of the landslide.

2. A landslide intelligent identification method based on large model prior knowledge, implemented based on the landslide intelligent identification model based on large model prior knowledge as claimed in claim 1, characterized in that: The following steps are involved: S1. Based on remote sensing image data, regional survey data and landslide spectral characteristics, visual interpretation is used to construct landslide identification sample data and outline landslide boundaries to obtain a landslide identification sample library; S2, extracting local and global information of landslide image data in the landslide identification sample library through the CLIP model, integrating the local and global information extracted by the CLIP model through the PKI module, and obtaining prior knowledge of landslides; S3, extracting multi-scale and multi-level image feature information of landslide image data in the landslide identification sample library through an encoder; S4, integrating the prior knowledge with the feature information through the CFA module to obtain the image feature representation containing the landslide information; S5. Obtain the accurate mask of the landslide by decoding step by step through the decoder.

3. The intelligent landslide identification method based on large model prior knowledge according to claim 2 is characterized in that: In S1, the landslide identification sample library includes data sets for model training and testing.

4. The method for intelligent landslide identification based on large model prior knowledge according to claim 3 is characterized in that: The data set includes a training set, a validation set, and a test set, and is processed by ArcGIS and Python to generate the label data required for the sample library and then input into the CLIP model.

5. The method for intelligent landslide identification based on large model prior knowledge according to claim 2 is characterized in that: In S2, the PKI module is based on a multi-head attention mechanism, with the global information extracted by the CLIP model as the query, and the local information extracted by the CLIP model as the key and value.

6. The method for intelligent landslide identification based on large model prior knowledge according to claim 2 is characterized in that: In S3, after the landslide image data enters the decoder, it is first divided into blocks through a block structure, and the blocks are flattened in the channel dimension. Then, these feature maps are sequentially passed through four Swin Transformer modules to extract feature information.

7. The method for intelligent landslide identification based on large model prior knowledge according to claim 6 is characterized in that: In the extraction of feature information, in the first processing, each block will be linearly transformed first, and then input into two consecutive Swin Transformer modules. In the subsequent processing, the feature map is downsampled by merging adjacent blocks, and then input into the next Swin Transformer module.

8. The method for intelligent landslide identification based on large model prior knowledge according to claim 2 is characterized in that: In S4, the CFA module is based on the reverse multi-head attention mechanism, with landslide prior knowledge as the query and image feature information as the key and value.

9. The method for intelligent landslide identification based on large model prior knowledge according to claim 2 is characterized in that: The landslide image data are randomly divided into a training set, a validation set, and a test set in a ratio of 6:2:2, wherein the training set is used for model training to extract features, the validation set is used to evaluate model performance, and the test set is used to evaluate the model.

10. The method for intelligent identification of landslides based on large model prior knowledge according to claim 9, characterized in that: The evaluation model indicators include OA, PA, F1_score and recall.

Citation Information

Patent Citations

  • Land coverage classification method based on dual-attention fusion mode and specific mode combined network

    CN117541918A

  • Semantic segmentation method based on Transform laser radar point cloud and camera image information fusion

    CN117788823A

  • Remote sensing image semantic segmentation method fusing multi-feature channel input and medium

    CN118212406A

  • Transform-based two-stage surface temperature prediction method and device

    CN118259376A

  • Underwater target detection method based on multi-scale feature cross fusion

    CN119091286A

Cited By

  • Neural network model and method for extracting image features of power equipment

    CN121033439A