A landslide intelligent recognition model and method based on the prior knowledge of large models
The DS Net model integrates prior knowledge and advanced deep learning techniques to improve landslide identification precision and efficiency by capturing multi-scale features and distinguishing landslide regions, addressing data scarcity and complexity challenges.
Patent Information
- Application Number
- CN202510519550.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing deep learning models are limited by the number and diversity of training data in landslide recognition tasks, making it difficult to accurately identify landslide areas under complex terrain under data scarcity. The traditional methods are costly, long periods, and accuracy rely on manual experience.
The intelligent landslide recognition model based on the prior knowledge of the big model is adopted, and the local and global information of the landslide image is extracted using the CLIP model, prior knowledge is obtained through the integration of the PKI module, multi-scale and multi-level image features are extracted in combination with the encoder, and prior knowledge and feature information are fused through the CFA module, and the precise mask is finally obtained by decoder decoding.
It significantly improves the accuracy and efficiency of landslide identification, can accurately identify landslide areas under scarcity of data, and provides important scientific basis for disaster prevention and mitigation.
Smart Images

Figure CN120032269B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of geological disaster monitoring and identification, and in particular to a landslide intelligent identification model and method based on prior knowledge of large models. Background Art
[0002] As the top geological disaster, landslides are characterized by wide distribution, high frequency, and complex formation mechanisms. At the same time, affected by both natural variations and human activities, various disaster accidents such as traffic interruption, river channel blockage, farmland damage, building destruction, village burial, and human and livestock casualties have become one of the most threatening geological problems to the life and property safety of local residents, the healthy and sustainable development of regional economy and society, and the safety of major project construction and operation. Therefore, it is urgent to carry out research on the distribution and activity of landslide disasters, and the high-precision identification and extraction of landslides are an important basis for studying landslide disasters, which can provide important reference for the research of disaster prevention and mitigation of landslide disasters.
[0003] Since the occurrence of landslides is often accompanied by the influence of other external environmental or geological structure activity factors, the terrain at the landslide site is fragmented and difficult to access. Traditional landslide identification methods mainly rely on field surveys and visual interpretation. Although they can provide intuitive landslide information, they are limited by problems such as high cost, long cycle, and accuracy dependence on manual experience, and cannot determine the specific location, scale, and quantity of landslide disasters in a timely and accurate manner. In recent years, with the significant improvement of the performance of remote sensing platforms and sensors, high-resolution remote sensing images have been widely used in the identification and investigation of geological disasters. Therefore, using optical remote sensing images to carry out high-precision landslide identification has become one of the hot research issues in the study of landslide disasters.
[0004] Deep learning has made remarkable progress in the field of image recognition and has been gradually applied to remote sensing image analysis. Compared with traditional machine learning methods, deep learning has greatly improved the ability of image classification and target recognition through multi-layer feature extraction and automatic learning mechanisms. With the enrichment of remote sensing data, the landslide identification technology based on optical images has developed from traditional manual visual interpretation to the stage of intelligent algorithm automatic identification. However, the occurrence of landslide events is sporadic and spatially heterogeneous, making it very difficult to obtain high-quality training samples, which in turn limits the generalization ability of the model. Existing deep learning models are still limited by the quantity and diversity of training data in landslide identification tasks, and it is difficult to comprehensively extract landslide features and accurately identify landslide areas under complex terrains. Therefore, how to make full use of the prior knowledge of large models to improve the adaptability and stability of landslide identification models under the condition of scarce data is one of the key issues in the development of current intelligent landslide identification technology. Summary of the Invention
[0005] The present invention provides a landslide intelligent recognition model and method based on prior knowledge of large models. At the model and method levels, it relies on a dual-coding segmentation network (DS Net) that combines prior knowledge and extracts features using a CLIP model, a PKI module, a CFA module, an encoder, etc. The aim is to use DS Net technology to effectively improve the accuracy and efficiency of landslide recognition in terms of identifying the precise boundary of a landslide, delicately capturing the unique features of a landslide, and accurately distinguishing the landslide area from the non-landslide area, providing an important scientific basis for disaster prevention and mitigation of landslide disasters.
[0006] This application is achieved through the following technical solutions:
[0007] A landslide intelligent recognition model based on prior knowledge of large models, comprising a CLIP model, a PKI module, a CFA module, an encoder, and a decoder;
[0008] The CLIP model is used to extract local and global information of landslide image data in the landslide recognition sample library;
[0009] The PKI module is used to integrate the local and global information extracted by the CLIP model and obtain prior knowledge of the landslide;
[0010] The encoder is used to extract multi-scale and multi-level image feature information of landslide image data in the landslide recognition sample library;
[0011] The CFA module is used to fuse the prior knowledge obtained by the PKI module with the feature information obtained by the encoder to obtain an image feature representation containing landslide information;
[0012] The decoder is used to gradually decode to obtain the precise mask of the landslide.
[0013] Based on the above model, this application also provides a landslide intelligent recognition method based on prior knowledge of large models, which includes the following steps:
[0014] S1. According to remote sensing image data, regional survey data, and landslide spectral characteristics, construct landslide recognition sample data by visual interpretation and delineate the landslide boundary to obtain a landslide recognition sample library;
[0015] S2. Extract local and global information of landslide image data in the landslide recognition sample library through the CLIP model, integrate the local and global information extracted by the CLIP model through the PKI module, and obtain prior knowledge of the landslide;
[0016] S3. Extract multi-scale and multi-level image feature information of landslide image data in the landslide recognition sample library through the encoder;
[0017] S4. The prior knowledge and the feature information are fused through the CFA module to obtain an image feature representation containing landslide information;
[0018] S5. The precise mask of the landslide is obtained through hierarchical decoding by the decoder.
[0019] Particularly, in S1, the landslide recognition sample library contains the data sets for model training and testing.
[0020] Particularly, the data set contains a training set, a validation set and a test set. After being processed by ArcGIS and Python to generate the label data required by the sample library, it is input into the CLIP model.
[0021] Particularly, in S2, the PKI module is based on the multi-head attention mechanism. Using the global information extracted by the CLIP model as the query, and the local information extracted by the CLIP model as the key and value, so as to enhance the extraction effect of the landslide prior information.
[0022] Particularly, in S3, after the landslide image data enters the decoder, it is first processed by a chunking structure for chunking, and the chunks are flattened in the channel dimension, and then these feature maps are sequentially passed through four Swin Transformer modules to extract feature information.
[0023] Particularly, in the extraction of feature information, in the first processing, each chunk will first perform a linear transformation and then be input into two consecutive Swin Transformer modules. In subsequent processing, the feature map is downsampled by merging adjacent chunks and then input into the next Swin Transformer module.
[0024] Particularly, in S4, the CFA module is based on the inverse multi-head attention mechanism. Using the landslide prior knowledge as the query, and the image feature information as the key and value, to further strengthen the expression of the landslide prior information in the image features.
[0025] Particularly, the landslide image data is randomly divided into a training set, a validation set and a test set according to the ratio of 6:2:2. The training set is used for model training to extract features, the validation set is used to evaluate the model performance, and the test set is used to evaluate the model.
[0026] Particularly, the evaluation model metrics include OA, PA, F1_score and recall.
[0027] Particularly, the Siam attention mechanism is incorporated into the CFA module. By assigning unique weights to neurons, it enhances the model's ability to extract key information and reduces the interference of redundant information.
[0028] Specifically, in the Swin Transformer encoder with a double-coding structure, the original multi-head attention structure is modified into a shifted window multi-head attention structure. By shifting the positions of these windows, the transfer of feature information in different windows is carried out, ensuring the recognition accuracy of the network.
[0029] The present invention proposes a DS Net deep learning model based on the prior knowledge of large models, which can effectively capture the multi-scale, multi-level, and global information and other features of landslide images in landslide recognition. The present invention effectively combines the advantages of prior knowledge of large models, deep learning feature extraction, multi-head attention mechanism, and hierarchical decoding technology. Compared with the prior art, the present invention has the following beneficial effects:
[0030] (1) A dual-coding segmentation network (DS Net) that fuses features extracted from prior knowledge, aiming to effectively improve the accuracy and efficiency of landslide recognition by using DS Net technology in accurately identifying the precise boundaries of landslides, delicately capturing the unique features of landslides, and accurately distinguishing landslide areas from non-landslide areas, providing an important scientific basis for disaster prevention and mitigation of landslide disasters.
[0031] (2) Using Swin Transformer as the first encoder, and at the same time using the shifted window multi-head attention structure, the transfer of feature information in different windows is carried out by shifting the positions of these windows. This strategy greatly improves the recognition accuracy of the network. Compared with ResNet50 and ResNet101 as the backbone networks, the DS Net with Transformer as the backbone network not only steadily improves in the validation accuracy, reaching a relatively high level, but also has a relatively low training loss.
[0032] (3) The DS Net model demonstrates significantly superior capabilities over other comparison models in the specific task of landslide recognition. Compared with other efficient and powerful semantic segmentation models, the DS Net is significantly superior to comparison models such as SegFormer, SegNeXt, FeedFormer, and U-MixFormerd in the four indicators of OA, PA, F1_score, and recall in landslide recognition, with the specific leading margin ranging from 3.5% to 7.3%, because it has a deep feature fusion strategy, a fine information processing mechanism, and an accurate capture ability for key landslide information. Description of the Drawings
[0033] The drawings described herein are used to provide a further understanding of the embodiments of the present application, form a part of the present application, and do not constitute a limitation on the embodiments of the present invention.
[0034] Figure 1 It is a flowchart of the landslide intelligent recognition method based on the prior knowledge of large models according to the embodiments of the present invention.
[0035] Figure 2 Schematic diagram of the spatial distribution of landslides selected for the embodiments of the present invention;
[0036] Figure 3 Flowchart for making the sample library in the embodiments of the present invention;
[0037] Figure 4 Flowchart of the Swin Transformer encoder in the embodiments of the present invention;
[0038] Figure 5 Flowchart of the DS Net network in the embodiments of the present invention. Detailed implementation manners
[0039] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0040] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0041] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.
[0042] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0043] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the invention product is customarily placed during use, or the orientation or positional relationship commonly understood by those skilled in the art. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.
[0044] In this embodiment, the full English name of the CLIP model is Contrastive Language-Image Pre-training, that is, a multi-modal model;
[0045] The full English name of the PKI module is Prior Knowledge Integration, that is, a prior knowledge integration module;
[0046] The full English name of the CFA module is Cross Fusion Attention, that is, a cross fusion attention module.
[0047] Embodiment 1
[0048] A landslide intelligent recognition model based on the prior knowledge of a large model, characterized in that it includes a CLIP model, a PKI module, a CFA module, an encoder, and a decoder;
[0049] The CLIP model is used to extract the local and global information of the landslide image data in the landslide recognition sample library;
[0050] The PKI module is used to integrate the local and global information extracted by the CLIP model and obtain the prior knowledge of the landslide;
[0051] The encoder is used to extract the multi-scale and multi-level image feature information of the landslide image data in the landslide recognition sample library;
[0052] The CFA module is used to fuse the prior knowledge obtained by the PKI module with the feature information obtained by the encoder to obtain an image feature representation containing landslide information;
[0053] The decoder is used to decode step by step to obtain an accurate mask of the landslide.
[0054] The purpose of such a design is to rely on the dual-coding segmentation network (DS Net) that extracts features by integrating prior knowledge, which consists of the CLIP model, PKI module, CFA module, encoder, etc. It aims to effectively improve the accuracy and efficiency of landslide recognition by using DS Net technology in accurately identifying the precise boundaries of landslides, delicately capturing the unique features of landslides, and accurately distinguishing landslide areas from non-landslide areas, providing an important scientific basis for disaster prevention and mitigation of landslide disasters.
[0055] Example 2
[0056] As Figure 1 shown, a landslide intelligent recognition method based on prior knowledge of large models includes the following steps: S1. According to remote sensing image data, regional survey data, and landslide spectral characteristics, visually interpret to construct landslide recognition sample data and delineate landslide boundaries, and at the same time perform processes such as data conversion and data cutting to produce a landslide recognition sample library;
[0057] In this embodiment, as Figure 2 shown, taking the landslide spatial distribution map of Beichuan County as an example, it is located in the transition area from the northwest edge of the Sichuan Basin to the western Sichuan Plateau, with a land area of 3,082.72 square kilometers, and the prediction of landslide boundaries for the entire region is completed.
[0058] For this landslide recognition sample library, high-resolution remote sensing images with a spatial resolution of 0.5 meters are used, combined with the landslide spectral characteristics of the study area, and visually interpret to construct landslide recognition sample data. Using the landslide contour as the boundary, delineate the landslide surface vector data to generate label data. The process of making the sample library includes vector editing, field assignment, vector-raster conversion, generating labels, data cutting, dividing the sample data set, etc. The specific process is as Figure 3 shown.
[0059] Based on the existing landslide disaster interpretation marks, use the vector editing tool in ArcGIS software to perform vector editing on the TIFF format remote sensing images, and manually delineate the boundaries by visual interpretation. Since the accuracy of boundary delineation will directly affect the quality of the landslide sample library, in order to improve the accuracy of landslide samples and the accuracy of model training, during the process of delineating vectors, it is necessary to strictly control the error of vector delineation and control the range of the delineated vector boundary error within 1 - 2 pixels.
[0060] Use ArcGIS software to perform attribute assignment operations on the landslide interpretation data, and add an attribute field "label" to the delineated landslide boundary data. Assign the attribute field of the landslide as "1", and assign the non-landslide background as "0".
[0061] Use Python in combination with the GDAL library to convert vector files into raster files according to the assigned field values, read the corresponding information therein, form binary images, and finally form the label data required for constructing the vegetation sample library, and ensure that the formed raster data is consistent with the number of rows and columns of the image data.
[0062] Due to limited computer memory, the information volume of the entire remote sensing image is too large to be used as input for the network model. The remote sensing image and the corresponding landslide label data need to be cut into multiple small images to ensure that they can be input into the model for training. In order to make the operation in the remote sensing image processing process more convenient and enable the network model to better extract the detailed features of the landslide and converge effectively, it is selected to divide the dataset into small images of 256×256 pixels in size.
[0063] After the sample segmentation operation, a high-resolution remote sensing landslide sample library consisting of 4,879 pictures is obtained in total. In order to ensure the diversity of the samples and make the dataset as diverse as possible, the obtained data is randomly divided into a training set, a validation set, and a test set according to the ratio of 6:2:2. Among them, 2,927 pictures in the training set are used for the model to train and extract features, 976 pictures in the validation set are used to evaluate the model performance, and 976 pictures in the test set are used to evaluate the model.
[0064] S2. Extract the local and global information of the landslide image data in the landslide recognition sample library through the CLIP model, integrate the local and global information extracted by the CLIP model through the PKI module, and obtain the prior knowledge of the landslide;
[0065] Among them, the CLIP model can mine the global and local information of the landslide image. By introducing the PKI module, we effectively integrate the local and global information provided by the CLIP to obtain the prior knowledge of the landslide. At the same time, the multi-head attention mechanism is used in the process of information integration and fusion. In the integration stage of global and local information, the global information is used as the query (q), and the local information is used as the key (k) and value (v) to enhance the extraction effect of the landslide prior information.
[0066] In this way, the PKI module is introduced to integrate the local and global information provided by the CLIP to obtain the prior knowledge of the landslide, and at the same time, the multi-head attention mechanism is used to strengthen the expression of the prior information.
[0067] S3. Extract the multi-scale and multi-level image feature information of the landslide image data in the landslide recognition sample library through the encoder;
[0068] Among them, the picture is first flattened in blocks in the encoder, and then the features are extracted step by step through four steps. Each step is processed through a linear transformation and a Swin Transformer module, and the adjacent blocks are merged and downsampled in the subsequent steps to reduce the size of the feature map;
[0069] Specifically, after entering the encoder, the image will be divided into blocks through a slicing structure and flattened in the channel dimension. Then, these feature maps will be extracted through four steps in sequence. In the first step, each block will be linearly transformed and then input into two consecutive Swin Transformer modules. In the subsequent steps, in order to further extract features and reduce the size of the feature map, the feature map will be downsampled by merging adjacent blocks and then input into the next Swin Transformer module. The specific structure is as follows: Figure 4 .
[0070] S4, integrating the prior knowledge with the feature information through the CFA module to obtain the image feature representation containing the landslide information;
[0071] In the fusion stage of feature information and prior knowledge, the roles are reversed, with landslide prior information as the query (q) and image feature information as the key (k) and value (v), thereby further strengthening the expression of landslide prior information in image features. The specific structure is as follows: Figure 5 .
[0072] In this way, the image features extracted by Swin Transformer are deeply integrated with prior knowledge in the CFA module to optimize the expression of landslide information;
[0073] The Siam attention mechanism is integrated into the CFA module. By assigning unique weights to neurons, the model can more accurately focus on information that is critical to the current task and effectively ignore irrelevant or minor parts.
[0074] S5. Obtain the accurate mask of the landslide by decoding step by step through the decoder.
[0075] In the specific implementation process, in terms of model training parameter setting, 100 training rounds were set for the Swin Transformer and Dual-coded Segmentation Network experiments to fully explore the learning ability of the model and avoid overfitting; the batch size was uniformly set to 50 to balance memory utilization and model convergence speed; the remote sensing image size of the network input was 256×256×3, which meets the format requirements of most deep learning models for input data; the stochastic gradient descent (SGD) optimizer was used to optimize the learning rate to promote the iterative update of the model in the direction of minimizing the loss function.
[0076] In the above specific embodiments, the purpose, technical solutions, and beneficial effects of the present application have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A landslide intelligent recognition model based on the prior knowledge of large models, characterized in that, It includes a multi-modal model, a prior knowledge fusion module, a cross-fusion attention module, an encoder, and a decoder; The multi-modal model is used to extract the local and global information of landslide image data in the landslide recognition sample library; The prior knowledge fusion module is used to integrate the local and global information extracted by the multi-modal model. Among them, the prior knowledge fusion module is based on the multi-head attention mechanism, uses the global information extracted by the multi-modal model as the query, the local information extracted by the multi-modal model as the key and value, and obtains the prior knowledge of the landslide; The encoder is used to extract multi-scale and multi-level image feature information of landslide image data in the landslide recognition sample library; The cross-fusion attention module is used to fuse the prior knowledge obtained by the prior knowledge fusion module with the feature information obtained by the encoder. Among them, the cross-fusion attention module is based on the inverse multi-head attention mechanism, uses the landslide prior knowledge as the query, and the image feature information as the key and value to obtain an image feature representation containing landslide information; The decoder is used to gradually decode to obtain an accurate mask of the landslide.
2. A landslide intelligent recognition method based on prior knowledge of large models, which is implemented based on the landslide intelligent recognition model described in claim 1, and is characterized in that, It includes the following steps: S1. According to remote sensing image data, regional survey data, and landslide spectral characteristics, visually interpret to construct landslide recognition sample data and draw the landslide boundary to obtain a landslide recognition sample library; S2. Extract the local and global information of landslide image data in the landslide recognition sample library through the multi-modal model, integrate the local and global information extracted by the multi-modal model through the prior knowledge fusion module, and obtain the prior knowledge of the landslide; S3. Extract multi-scale and multi-level image feature information of landslide image data in the landslide recognition sample library through the encoder; S4. Fuse the prior knowledge and feature information through the cross-fusion attention module to obtain an image feature representation containing landslide information; S5. Gradually decode through the decoder to obtain an accurate mask of the landslide.
3. The landslide intelligent recognition method based on the prior knowledge of the large model according to claim 2, wherein In S1, the landslide recognition sample library contains a data set for model training and testing.
4. The landslide intelligent recognition method based on the prior knowledge of the large model according to claim 3, characterized in that, The data set contains a training set, a validation set, and a test set. After being processed by ArcGIS and Python to generate the label data required by the sample library, it is input into the multi-modal model.
5. The landslide intelligent recognition method based on prior knowledge of large models according to claim 2, wherein In S3, after the landslide image data enters the decoder, it is first block-processed through a chunking structure and flattened in the channel dimension, and then these feature maps are sequentially passed through four Swin Transformer modules to extract feature information.
6. The landslide intelligent recognition method based on the prior knowledge of the large model according to claim 5, characterized in that, In the extraction of feature information, in the first processing, each chunk will first perform a linear transformation and then be input into two consecutive Swin Transformer modules. In subsequent processing, the feature map is downsampled by merging adjacent chunks and then input into the next Swin Transformer module.
7. The landslide intelligent recognition method based on the prior knowledge of the large model according to claim 2, characterized in that The landslide image data is randomly divided into a training set, a validation set, and a test set in a ratio of 6:2:
2. The training set is used for model training to extract features, the validation set is used to evaluate the model performance, and the test set is used to evaluate the model.
8. The landslide intelligent recognition method based on the prior knowledge of the large model according to claim 7, characterized in that, The metrics for evaluating the model include OA, PA, F1_score, and recall.