Landslide disaster early warning method and device based on deep learning, medium and product

By constructing a deep learning model of multimodal data set and DenseASPP module, the problems of low landslide recognition efficiency and poor accuracy in traditional methods are solved, and efficient, accurate identification and disaster warning of landslide areas are achieved.

CN120299220AActive Publication Date: 2025-07-11SICHUAN INST OF LAND & SPACE ECOLOGICAL RESTORATION & GEOLOGICAL DISASTER PREVENTION

Patent Information

Application Number
CN202510771624.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

Traditional artificial visual interpretation and pixel-based automation methods have low efficiency and poor accuracy in landslide recognition, making it difficult to distinguish landslides from other land objects, especially in high spatial resolution urban environments.

Method used

Using a deep learning-based method, a multimodal data set is constructed and a deep learning model using the DenseASPP module is used. Combined with image encoder, text encoder and visual sampler, the precise identification of landslide areas is achieved through multi-scale feature extraction and dense connection.

Benefits of technology

It improves the efficiency and accuracy of landslide area identification, can accurately extract landslide characteristics in complex environments, reduce the rate of misjudgment, and supports real-time landslide disaster warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299220A_ABST
    Figure CN120299220A_ABST
Patent Text Reader

Abstract

The invention discloses a landslide disaster early warning method and device based on deep learning, a medium and a product, and relates to the field of geological disaster monitoring. Firstly, a remote sensing landslide data set is obtained and preprocessed, and a multi-modal data set is constructed; each sample in the multi-modal data set comprises a remote sensing image, a text query, a non-text query and a corresponding mask; constructing a deep learning model, wherein the deep learning model comprises an encoder, a DenseASPP module, a decoder and a prediction head; the encoder comprises an image encoder, a text encoder and a visual sampler; dividing the multi-modal data set into a training set, a verification set and a test set which are respectively used for training, verification and testing of a deep learning model, and taking the trained deep learning model as a landslide extraction model; the landslide extraction model is adopted to identify the landslide area in the remote sensing image, and the landslide disaster early warning is carried out according to the landslide area change, so that the efficiency and precision of landslide area identification and landslide disaster early warning can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of geological disaster monitoring, and particularly to a landslide disaster early warning method, device, medium and product based on deep learning. Background Technique

[0002] Landslide is one of the main geological disasters in the world. Due to the complexity of the ground object background and morphology, the automatic and rapid identification of landslides and disaster early warning are major challenges in disaster prevention and mitigation work. At present, traditional manual visual interpretation requires rich experience, and there are problems such as low automation, time-consuming and laborious, and inability to be widely promoted. Although pixel-based automated methods are widely used in landslide identification, they only rely on spectral information and ignore the spatial features of the target, thus greatly limiting the effect of landslide identification.

[0003] With the development of remote sensing and geographic information system technologies, the means of collecting and analyzing spatial information have become gradually rich and efficient, providing favorable conditions for remote sensing landslide identification. Remote sensing technology has irreplaceable advantages in landslide information extraction due to its wide coverage, diverse dimensions and strong real-time performance, providing strong technical support for the monitoring, assessment and management of geological disasters. In this context, with the release of a large number of remotely sensed landslide datasets with fine annotations and the rapid development of artificial intelligence technology from the theoretical knowledge level to practical applications, deep learning technology has also flourished. Given the strong feature extraction ability of deep learning methods in image processing, many researchers in the field of remote sensing have applied them to the task of remotely sensed image landslide extraction. Currently, in the task of extracting high-spatial-resolution urban landslides, the spectral information of artificial ground objects such as landslides, roads, parking lots, and bare open spaces in high-resolution images is very similar. This characteristic of low between-class variance and high within-class variance brings great difficulties to accurately extracting landslides. This similarity makes it difficult for traditional extraction algorithms to distinguish landslides with close or overlapping edges from other ground objects, resulting in reduced extraction accuracy and increased misjudgment rate. Summary of the Invention

[0004] The purpose of the present application is to provide a landslide disaster early warning method, device, medium and product based on deep learning to improve the efficiency and accuracy of landslide area identification and landslide disaster early warning.

[0005] To achieve the above purpose, the present application provides the following solutions.

[0006] In the first aspect, the present application provides a landslide disaster early warning method based on deep learning, including: Obtain a remotely sensed landslide dataset and perform preprocessing to construct a multi-modal dataset; each sample in the multi-modal dataset includes a remotely sensed image, a text query, a non-text query, and a corresponding mask; Build a deep learning model, which includes an encoder, a DenseASPP module, a decoder, and a prediction head; the encoder includes an image encoder, a text encoder, and a visual sampler; the DenseASPP module refers to a densely connected atrous spatial pyramid pooling module; Divide the multi-modal dataset into a training set, a validation set, and a test set, which are used to train, validate, and test the deep learning model respectively, and use the trained deep learning model as a landslide extraction model; Use the landslide extraction model to identify the landslide area in the remote sensing image and issue a landslide disaster warning according to the change of the landslide area.

[0007] Optionally, the acquisition of the remote sensing landslide dataset and its preprocessing, and the construction of the multi-modal dataset specifically include: Acquire a remote sensing landslide dataset, which includes multiple landslide images, multiple non-landslide images, and corresponding label information; both the landslide images and non-landslide images are remote sensing images; Use the landslide images as positive samples and the non-landslide images as negative samples, and clean the remote sensing landslide dataset according to the requirement that the ratio of positive to negative samples is 1:1 to obtain a cleaned dataset; Save the label information of each sample in the cleaned dataset in a structured text form and convert it into a corresponding mask; among them, the mask corresponding to the landslide area is called a positive mask, and the mask corresponding to the background area is called a negative mask; Automatically imitate users to randomly generate text queries and non-text queries corresponding to each sample; the text queries include categories and statements; the non-text queries include points, boxes, scribbles, polygons, and example images; Store the remote sensing images, text queries, non-text queries, and masks of each sample in one-to-one correspondence to construct a multi-modal dataset.

[0008] Optionally, the image encoder is used to perform layer-by-layer convolution and downsampling operations on the input remote sensing image to generate multi-scale feature maps with different resolutions and semantic information levels ; divide the feature maps with higher resolution but less semantic information in the multi-scale feature maps into original feature maps ; divide the feature maps with lower resolution but stronger semantic information in the multi-scale feature maps into to-be-enhanced feature maps ; The visual sampler is used to use the formula to convert all types of non-text queries and multi-scale feature maps into visual cues ; where Operations performed by the visual sampler; A mask for points, boxes, scribbles, polygons, and / or sampling regions from example images; The text encoder is used to convert a text query into a text prompt .

[0009] Optionally, the DenseASPP module is used to perform feature enhancement and fusion on the feature map to be enhanced in the multi-scale feature map to obtain an enhanced feature map ; The enhanced feature map and the original feature map are jointly used as the feature map to be mapped .

[0010] Optionally, the DenseASPP module includes a multi-scale dilated convolutional layer and densely connected channels; The multi-scale dilated convolutional layer contains multiple 3×3 dilated convolutional layers, each with a different dilation rate; a 1×1 convolutional layer is added before each dilated convolutional layer; the dilated convolutional layers are densely connected through densely connected channels; the outputs of the dilated convolutional layers are concatenated with the original input feature map to be enhanced and then passed through a 1×1 convolutional layer to obtain an enhanced feature map .

[0011] Optionally, the feature map to be mapped is mapped to the image-text joint semantic space together with the text prompt , the visual prompt and the memory prompt , and after a scale alignment operation, it is passed to the decoder in a unified form; On the one hand, the decoder generates the memory prompt for the current stage through a masked cross-attention mechanism based on the formula ; where is the memory prompt of the previous stage; is the mask of the previous stage; is the feature map to be mapped of the current stage; ; On the other hand, the decoder outputs a mask embedding and a class embedding through a masked self-attention mechanism based on the formula ; where represents a learnable query.

[0012] Optionally, the prediction head is based on the formula Inferring the mask ; Mask represents the extracted landslide area; Represents the prediction operation for the mask; the prediction head is based on the formula Inferring semantic concepts ; Semantic concept Represents the predicted category or statement; Represents a semantic prediction operation.

[0013] In a second aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the landslide hazard warning method based on deep learning.

[0014] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the landslide disaster warning method based on deep learning.

[0015] In a fourth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the deep learning-based landslide hazard warning method.

[0016] According to the specific embodiments provided in this application, this application discloses the following technical effects.

[0017] In a landslide disaster warning method, device, medium and product based on deep learning provided by the present application, a deep learning model constructed includes an encoder, a DenseASPP module, a decoder and a prediction head; the encoder includes an image encoder, a text encoder and a visual sampler, which can realize multimodal data fusion of remote sensing images, text queries and non-text queries; the deep learning model is based on multimodal data fusion and dense multi-scale perception, which can not only perform fine-grained segmentation of global and local features at the same time, but also extract multi-scale information of images at different scales through void convolution, and enhance the feature expression ability of the model by densely connecting feature maps of different scales, thereby effectively improving the efficiency and accuracy of landslide area identification and landslide disaster warning. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1 A flowchart of a landslide disaster early warning method based on deep learning is provided for this application; Figure 2 Schematic diagram of positive and negative samples and their corresponding label information; Figure 3 A schematic diagram of the overall framework of the deep learning model constructed for this application; Figure 4 It is the structural diagram of ASPP module; Figure 5 Schematic diagram of the multi-scale receptive field of the DenseASPP module; Figure 6 It is a structural diagram of the DenseASPP module; Figure 7 A schematic diagram of the interaction mode of query and prompt; Figure 8 This is a visual comparison of the landslide identification results of the deep learning model and the SEEM model in this application; Figure 9 Schematic diagram of the segmentation capability of the deep learning model in this application in different task scenarios. DETAILED DESCRIPTION

[0020] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0021] The purpose of this application is to propose a landslide disaster warning method, equipment, medium and product based on deep learning, aiming to improve the efficiency and accuracy of landslide area identification and landslide disaster warning.

[0022] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0023] In recent years, deep learning technology has shown significant advantages in the field of image processing. In particular, with the gradual improvement of computing resources, deep learning algorithms have been continuously optimized, and can deeply explore the deep features in remote sensing images (also called images), such as morphology, texture, spectrum, and spatial relationships. This progress enables deep learning to be effectively applied to the accurate identification of landslide areas and landslide disaster warning, improving the accuracy and timeliness of landslide monitoring. On this basis, the present application provides a landslide disaster warning method based on deep learning. In an exemplary embodiment, Figure 1As shown, the landslide disaster early warning method based on deep learning includes the following steps 1 to 4.

[0024] Step 1: Obtain a remote sensing landslide dataset and perform preprocessing to construct a multi-modal dataset; each sample in the multi-modal dataset includes a remote sensing image, a text query, a non-text query, and a corresponding mask.

[0025] The specific steps of Step 1 include the following steps 1.1 to 1.5.

[0026] Step 1.1: Obtain a remote sensing landslide dataset, which includes multiple landslide images, multiple non-landslide images, and corresponding label information.

[0027] Among them, both the landslide images and the non-landslide images are remote sensing images. The landslide images are called positive samples, and the non-landslide images are called negative samples. Examples of some positive and negative sample images and their corresponding label information are shown Figure 2 As shown. The label information refers to the binary mask pixel values (0 / 255) at each pixel position in the image. The pixel values in the landslide area are 255, corresponding to the white area; the pixel values in the background area are 0, corresponding to the black area. The obtained remote sensing landslide dataset can be an existing dataset, such as the open remote sensing landslide dataset named Bijie Landslide Dataset created by Wuhan University, or a dataset created by the user himself.

[0028] Step 1.2: Use the landslide images as positive samples and the non-landslide images as negative samples, and clean the remote sensing landslide dataset according to the requirement that the ratio of positive to negative samples is 1:1 to obtain the cleaned dataset. That is to say, it is necessary to ensure that the ratio of positive to negative samples in the cleaned dataset is 1:1.

[0029] Step 1.3: Save the label information of each sample in the cleaned dataset in a structured text form and convert it into a corresponding mask; the mask corresponding to the landslide area is called the positive mask (pm), and the mask corresponding to the background area is called the negative mask (nm).

[0030] As Figure 2 shown, the label information of each sample in the cleaned dataset itself shows visual content, and it is necessary to save the visual label information in the image in a structured text form. For example, convert the binary mask pixel values (0 / 255) of the label information itself into the 0 / 1 values required for model training through a Python script, and use it as the mask (masks) corresponding to the sample image. Through such processing, the label information in the image is saved in a structured text form, which is convenient for subsequent training, inference, or evaluation work.

[0031] Step 1.4: Automatically imitate the user to randomly generate text queries and non-text queries corresponding to each sample; the text queries include categories and statements; the non-text queries include points, boxes, scribbles, polygons, and example images.

[0032] The overall framework of the deep learning model constructed in this application is as Figure 3 shown, including an encoder, a DenseASPP module, a decoder, and a prediction head, where the encoder further includes an image encoder, a text encoder, and a visual sampler. Since this deep learning model needs to input text queries and non-text queries while inputting remote sensing images, including points, boxes, scribbles, polygons, and their masks. So on the one hand, it is necessary to convert the cleaned dataset into the format of the COCO dataset. Specifically, through Python code, the pixel values of 0 and 255 in the image mask can be converted into the values of 0 and 1 required for training, and a JSON file required for creating the COCO dataset is created, which contains the name of the remote sensing image, the corresponding ID, the input text query, and the mask in text form.

[0033] That is to say, the encoder of the deep learning model in this application needs to combine text queries (such as "landslide affecting the road") and multi-modal inputs such as points, boxes, scribbles, polygons, and masks. Therefore, on the one hand, it is necessary to convert the binary mask pixel values (0 / 255) of the image into the values of 0 / 1 required for training through a Python script, and construct a JSON file in COCO format - this file includes the name and ID of the remote sensing image, the text query, and the masked annotation information structured in text (i.e., the mask in text form), so as to realize the standardized mapping of label data from the pixel space to the semantic space and provide a unified multi-modal data interface for model training, inference, and evaluation.

[0034] On the other hand, it is also necessary to automatically imitate the user to randomly generate text queries and non-text queries corresponding to each sample. Among them, the text queries include categories and statements. For example, categories: "slope", "landslide"; statements: "Identify landslide bodies that may have geological disasters", "Landslide areas affecting roads", etc. The non-text queries include points, boxes, scribbles, polygons, and example images (also called reference images), etc. During the training process, the non-text queries are randomly generated by automatically imitating the user. During the use process after the model training is completed, the non-text queries can be given manually by the user or randomly generated.

[0035] Step 1.5: Store the remote sensing images, text queries, non-text queries, and masks of each sample in one-to-one correspondence to construct a multi-modal dataset.

[0036] In the multimodal dataset constructed in this application, the remote sensing image, text query, and non-text query corresponding to each sample are used as the input of the deep learning model, and the mask and semantic concept corresponding to each sample are used as the output of the deep learning model to train, validate, and test the deep learning model.

[0037] Step 2: Construct a deep learning model, which includes an encoder, a DenseASPP module, a decoder, and a prediction head; the encoder includes an image encoder, a text encoder, and a visual sampler.

[0038] To achieve more efficient and accurate landslide information extraction, this application proposes a new type of deep learning model, as Figure 3 shown. The deep learning model combines the advantages of multimodal data fusion and the DenseASPP module. Based on the advantages in panoramic segmentation and multi-scale feature extraction, the DenseASPP module realizes efficient interaction and information transmission between different-scale features by enhancing the receptive field and densely connecting multi-scale features, thereby providing more refined feature support for the model. This design can significantly improve the recognition accuracy of the model for landslide areas in high-resolution remote sensing images.

[0039] As Figure 3 shown, the overall framework of the deep learning model includes: 2.1) an encoder; 2.2) a DenseASPP module; 2.3) a decoder; and 2.4) a prediction head, which are specifically introduced as follows.

[0040] 2.1) Encoder.

[0041] As Figure 3 shown, the deep learning model of this application generally adopts an encoder-decoder architecture. In the encoder of the deep learning model, it specifically includes: 2.1.1) an image encoder; 2.1.2) a text encoder; and 2.1.3) a visual sampler.

[0042] 2.1.1) Image encoder.

[0043] Referring to Figure 3 , the input of the image encoder is the remote sensing image , and the output is a multi-scale feature map . The image encoder generates multi-scale feature maps with different resolutions and semantic information levels by performing layer-by-layer convolution and downsampling operations on the input remote sensing image . Among them, the feature map with a higher resolution but less semantic information in the multi-scale feature map is called the original feature map ; the multi-scale feature map ; the multi-scale feature map Feature maps with relatively low resolution but strong semantic information are called feature maps to be enhanced. 。

[0044] For example, given a remote sensing image with a size of H × W ×3, where and H and W are the height and width of the remote sensing image respectively. In a specific embodiment, H = W = 512. As Figure 3 shows, the remote sensing image with a size of 512×512×3 is input into the image encoder. The image encoder is usually based on a Convolutional Neural Networks (CNN) or Transformer architecture, and gradually extracts the global features of the remote sensing image through multiple convolutional or Transformer modules to generate multi-scale feature maps . Among them, represent the image features (also called feature maps) of the 2nd, 3rd, 4th, and 5th scales respectively, Figure 3 represents them as in

[0045] The first scale usually refers to the original image features.

[0046] For example, an encoder based on ResNet will gradually extract the semantic information of the image through multiple convolutional layers and residual blocks. During the extraction process, the image encoder will gradually reduce the spatial resolution of the feature map through downsampling operations (such as pooling or convolutions with a stride greater than 1), while increasing the number of channels, so as to extract deeper global semantic information. These feature maps Res2, Res3, Res4, and Res5 respectively correspond to the outputs of different levels in the image encoder. The feature maps at each level are gradually generated through convolutional layers and downsampling operations. Among them, Res2 has the highest spatial resolution and the lowest number of channels, and is suitable for capturing the detailed information of the image, such as edges, textures, etc. The spatial resolution of Res3 is lower than that of Res2 and is obtained through downsampling. While retaining certain details, it begins to capture more abstract semantic information. The spatial resolution of Res4 is further reduced and the number of channels increases. It can capture higher-level semantic information, such as the shape and partial structure of objects. And Res5 has the lowest spatial resolution and the largest number of channels, and is suitable for capturing global semantic information, such as the overall layout and category information of the scene. These multi-scale feature maps do not all directly enter the image-text joint semantic space, but a part of the feature maps are enhanced through the DenseASPP module and then processed subsequently.Specifically, the shallow feature maps Res2 and Res3 contain a lot of redundant information, which is not directly helpful for the final semantic segmentation task. Simply expanding the receptive field will only increase the computational pressure of the model. On the other hand, Res4 and Res5 are rich in a large amount of high-level semantic information and global context. Using the DenseASPP module here can further enhance the representation ability of the feature maps, thereby improving the model accuracy. Therefore, the feature maps Res4 and Res5 are used as the feature maps to be enhanced. , through the DenseASPP module, the recognition ability of the model for landslide areas is improved by expanding the receptive field and dense connection. While the feature maps Res2 and Res3 are used as the original feature maps. They are directly mapped into the image-text joint semantic space.

[0047] 2.1.2) Text Encoder.

[0048] During the landslide recognition process, the text encoder converts text queries such as categories and sentences into text prompts. .

[0049] 2.1.3) Visual Sampler.

[0050] As Figure 3 shown, the remote sensing image first extracts multi-scale feature maps by the image encoder , and then the visual sampler converts the multi-scale feature maps and all types of non-text queries (such as points, boxes, scribbles, polygons, and sampling regions from another example image) into visual prompts , as follows: (1).

[0051] Where is the operation performed by the visual sampler. is the mask corresponding to points, boxes, scribbles, polygons, and / or sampling regions from the example image (also called reference regions). The visual sampler first pools the features of the corresponding regions from the image features by the point sampling method. For all visual prompts (such as points, boxes, scribbles, polygons, etc.), up to 512 point feature vectors are extracted from the sampling regions specified by the prompts in a uniform interpolation manner. When the correct sentence segment cannot be recognized only by the text prompt , these non-text queries help to eliminate the ambiguity of the user's intention.

[0052] In this application, through the joint training of panoramic segmentation and reference segmentation, visual prompts and text prompts Natural alignment, thus significantly improving the segmentation accuracy and the intuitiveness of user interaction. Meanwhile, it has flexibility and scalability, and is suitable for fine-grained segmentation tasks in complex scenarios.

[0053] 2.2) DenseASPP module.

[0054] Before introducing the DenseASPP module adopted in this application, the basic ASPP module is first introduced. ASPP (Atrous Spatial Pyramid Pooling) is a technique widely used in deep learning and computer vision, especially excellent in semantic segmentation tasks. It combines the concepts of atrous convolution and spatial pyramid pooling, aiming to capture multi-scale context information, thereby improving the model's recognition ability for objects of various sizes in images.

[0055] As Figure 4 shown, the ASPP module mainly consists of two parts, namely the atrous convolution layer and the spatial pyramid pooling layer. Among them, atrous convolution is a special type of convolution operation, aiming to expand the network receptive field without increasing the number of parameters or the amount of computation. The traditional convolution operation extracts spatial features by sliding the convolution kernel on the input feature map, where each element of the convolution kernel is only multiplied by the corresponding element of the input feature map. In contrast, atrous convolution introduces an additional parameter - the dilation rate, which represents the spacing between adjacent elements in the convolution kernel. When the dilation rate is 1, atrous convolution is equivalent to standard convolution; when the dilation rate is greater than 1, the convolution kernel is "expanded", and the convolution operation covers a wider input area, but the number of parameters actually involved in the calculation remains unchanged.

[0056] Let the input feature map be , the convolution kernel be , the dilation rate be , then the output of the atrous convolution layer can be calculated by the following formula: (2); Among them, represents the position on the output feature map; represents the relative position in the convolution kernel.

[0057] In traditional CNN models, due to the existence of fully connected layers, the size of the input image must be fixed. This limitation requires that all images be scaled to the same size before being input into the network, which may lead to image distortion or information loss. The Spatial Pyramid Pooling layer is designed to address this issue. It allows the network to receive input images of arbitrary sizes. By adding a Spatial Pyramid Pooling layer after the convolutional layer, it generates a fixed-length output vector, making the size of the image no longer restricted. The key to how the Spatial Pyramid Pooling layer works is that it divides the output feature map of the last convolutional layer into several regions and performs pooling operations (such as max pooling or average pooling) within each region. These regions are divided according to predefined pyramid levels, with different numbers of regions at each level, forming a pyramid structure. In this way, regardless of the size of the input image, the Spatial Pyramid Pooling layer can output a fixed-length feature vector, providing input for subsequent fully connected layers or other processing layers.

[0058] After combining these two, the ASPP module achieves the effect of pyramid pooling by applying different dilation rates on multiple parallel dilated convolutional branches. As Figure 4 shown, among multiple dilated convolutional branches, each branch uses a different dilation rate, usually including dilation rates of 1 (i.e., ordinary convolution), 6, 12, 18, etc. The purpose is to cover different ranges of receptive fields and capture more local information. The global average pooling layer performs global average pooling on the input feature map, then passes through a 1×1 convolutional layer, and finally upsamples it to the same size as the original input feature map. This step captures global context information and helps improve performance. The fusion layer is responsible for fusing the outputs of all dilated convolutional branches and the result of global average pooling in the channel dimension, and then adjusts the number of channels through an additional 1×1 convolutional layer, finally generating a comprehensive feature representation. With such a design, ASPP cleverly combines dilated convolution and spatial pyramid pooling, enhancing the model's ability to understand image details and global structures, and greatly improving the model's segmentation accuracy and boundary clarity by effectively fusing multi-scale context information.

[0059] DenseASPP effectively overcomes the limitations of traditional semantic segmentation networks when dealing with targets of different scales based on ASPP. Traditional convolutional networks are difficult to capture both fine-grained local details and global context information simultaneously. DenseASPP stacks various dilated convolutional layers in a densely connected manner, achieving full fusion of multi-scale features, thereby constructing a richer feature pyramid and significantly expanding the receptive field. Specifically, while maintaining the resolution of the feature map, DenseASPP continuously stacks dilated convolutional layers and utilizes the dense connection mechanism, effectively compensating for the deficiency of single-layer dilated convolution in responding to multi-scale targets.

[0060] In traditional ASPP, the dilated convolutional layers work in parallel, and the four sub-branches do not share any information during the forward propagation process. In contrast, the dilated convolutional layers in DenseASPP share information through dense connections. The layers with smaller and larger dilation rates (i.e., hole rates) are interdependent, where the forward propagation process not only forms a denser feature pyramid but also presents a larger filter to perceive a larger context. The dilated convolutional layers are organized in a cascaded structure, with the dilation rate gradually increasing for each layer: smaller dilation rates are used in the lower layers, and larger dilation rates are used in the deeper layers. The output of each dilation layer is connected not only to the input feature map of the current layer but also cascaded with the outputs of all lower layers to form a richer feature map, which is then fed into the next layer. The feature map finally output by the DenseASPP module is jointly generated by dilated convolutions with multiple rates and scales. The DenseASPP module can build a denser and more powerful feature pyramid with just a few dilated convolutional layers. Compared with the original ASPP, this change mainly brings two benefits: a denser feature pyramid and a larger receptive field.

[0061] "Denser" here not only means that the feature pyramid has better scale diversity but also means that more pixels are involved in the convolution than in ASPP. The DenseASPP module can not only sample the input at different scales but also use dense connections to achieve a diverse integration of layers with different dilation rates. Each integration is equivalent to a kernel of different scales, i.e., different receptive fields. Therefore, a feature map with more scales than in ASPP is obtained. For a dilated convolutional layer with a dilation rate and a kernel size , the equivalent receptive field size is: (3); For example, for a 3×3 convolutional layer, when = 3, the corresponding receptive field size is 7.

[0062] Stacking two convolutional layers together can give a larger receptive field. Suppose there are two convolutional layers with receptive field sizes of and , respectively. The new receptive field is: (4).

[0063] For example, stacking a convolutional layer with a receptive field of 7 and a convolutional layer with a receptive field of 13 will result in a convolutional layer with a receptive field of 19. Taking the dense stacking of dilated convolutions with dilation rates (3, 6, 12, 18) as an example, the simplified structure is as Figure 5 shown. Among them, the Represents the receptive field size of the corresponding combination. The feature pyramid generated by DenseASPP has greater scale diversity (i.e., high resolution on the scale axis) and a larger receptive field. Let denote the maximum receptive field of the feature pyramid, and the function denotes a dilated convolutional layer with kernel size and dilation rate . For ASPP with = {6, 12, 18, 24}, its maximum receptive field is: (5).

[0064] In DenseASPP, when = {6, 12, 18, 24}, its maximum receptive field is: (6).

[0065] According to formula (4), when every two receptive fields are stacked together, the receptive fields are added and then subtract 1. So in formula (6), the receptive fields are the sum of 4, thus subtract 3.

[0066] The structural diagram of the DenseASPP module is as shown in Figure 6 , and mainly includes: 2.2.1) Multi-scale dilated convolutional layer and 2.2.2) Dense connection channels.

[0067] 2.2.1) Multi-scale dilated convolutional layer: The DenseASPP module contains multiple 3×3 dilated convolutional layers, and each convolutional layer uses a different dilation rate = {3, 6, 12, 18, 24} to extract multi-scale long-range spatial information and enhance the feature fusion ability, thereby effectively expanding the receptive field. To avoid the network structure from being too wide, a 1×1 convolutional layer is added before each dilated convolutional layer to perform channel compression and reduce the computational complexity. Essentially, dilated convolution upsamples by inserting pixels with value 0 between the standard convolutional filters, so that the larger the dilation rate , the wider the receptive field, thus enhancing the perception ability for large-scale targets and complex ground objects.

[0068] 2.2.2) Dense connection channels: Introduce a dense connection mechanism between different feature layers to promote feature reuse and improve information flow. This connection method can enhance the expression ability of edge details and improve the accuracy and clarity of the object boundary in the segmentation task.

[0069] Specifically for the Figure 3 shown deep learning model, where the input of the DenseASPP module is the feature map to be enhanced , and by the feature map to be enhanced in the multi-scale feature map Perform feature enhancement and fusion to obtain the enhanced feature map . And use the enhanced feature map and the original feature map together as the feature map to be mapped .

[0070] Therefore, as Figure 6 shown, when the feature map to be enhanced enters the DenseASPP module, it first enters a 1×1 convolutional layer to achieve preliminary channel compression and feature extraction. Subsequently, the feature map is fed into the first 3×3 dilated convolutional layer with a dilation rate = 3. This dilated convolutional layer realizes sparse sampling of dilated convolution by inserting two zeros between adjacent values of the convolutional kernel, thereby significantly expanding the receptive field without increasing the number of parameters and effectively capturing multi-scale spatial information. Then, the output of this layer is concatenated with the original input feature map to form a richer feature representation and provide a more sufficient information basis for subsequent layers. Subsequently, the concatenated feature map will successively enter dilated convolutional layers with larger dilation rates ( = 6, 12, 18, 24). In each layer, the convolutional kernel further expands the receptive field by inserting the corresponding number of zeros to capture a wider range of context information. To prevent the network from being too wide, a 1×1 convolutional layer is added before each dilated convolutional layer to reduce the depth of the feature map to half of the original. And the output of each layer is densely connected to the feature maps of all previous layers, which not only promotes the reuse of features but also enables information at different scales to be complementarily fused at deeper levels. After the outputs of each dilated convolutional layer are concatenated with the original input feature map to be enhanced , and then passed through a 1×1 convolutional layer, the enhanced feature map is finally output.

[0071] 2.3) Decoder.

[0072] As Figure 3 shown, the feature map to be mapped is mapped to the image-text joint semantic space together with the text prompt , the visual prompt and the memory prompt , and after scale alignment operation, it is transmitted to the decoder in a unified form.

[0073] On the one hand, map the feature map to be mapped obtained by the encoder and various prompts to the image-text joint semantic space, and the decoder can use the masked cross-attention mechanism , combined with the memory prompt generated during the previous segmentation , obtain the memory prompt to be used for this segmentation , and the formula is as follows: (7).

[0074] Among them, is the memory prompt for the current stage; is the memory prompt for the previous stage; the memory prompt encodes historical information by using a masked guided cross-attention layer. is the mask for the previous stage; is the feature map to be mapped for the current stage. In this way, the cross-attention takes effect only within the area specified by the previous mask. The updated memory prompt interacts with other prompts in the decoder to convey the historical information of the current round.

[0075] On the other hand, the decoder passes through a self-attention mechanism with a mask , and makes the learnable query interact synergistically with the text, visual, and memory prompts to output the masked embedding and the category embedding , and the formula is as follows: (8).

[0076] In and , the symbols ";" and " " are delimiters used to separate different input parameters. For example in, the symbol ";" is used to separate the main input parameters, namely the memory prompt for the previous stage and the mask for the previous stage . The symbol " " is used to separate the main input parameters from the additional input parameters, such as separating the feature map to be mapped . represents the set of text prompts, visual prompts, and memory prompts.

[0077] Although the architecture of the deep learning model in this application is simple, a complex interaction method is also adopted between the queries and the prompts, as Figure 7 shown. The learnable query is a set of trainable vectors in the decoder, which are usually used to guide the attention mechanism of the model and help the model extract relevant information from the input features. These query vectors are the parameters of the model and will be optimized through backpropagation during the training process. During the training process, the learnable query in the decoder is copied as the object query , the text query and visual queries , each task has the same weight for different segmentation tasks such as general, reference, and interactive segmentation, and the corresponding prompts freely interact with their queries through masked self-attention.

[0078] The deep learning model on which this application is based is a unified multi-modal segmentation architecture that realizes the multi-task unity of general segmentation (semantic / instance / panoramic) and interactive segmentation through an encoder-decoder framework. It constructs a cross-modal prompt encoding system: the image encoder extracts image features of the remote sensing image and outputs a feature map; the visual sampler converts points, boxes, scribbles, polygons, and example images into visual prompts; the text encoder maps semantic queries into text prompts; the three are fused into the image-text joint semantic space through a prompt embedding layer with shared parameters. The decoder based on the cross-attention mechanism and self-attention mechanism dynamically integrates four types of interactive prompts (feature map, text prompt, memory prompt, and visual prompt) to achieve the collaborative optimization of visual-text features, and finally outputs the segmentation result and text information through a unified prediction head. This design breaks through the dependence of traditional segmentation models on a single input modality.

[0079] 2.4) Prediction head.

[0080] The prediction head is based on mask embedding and class embedding to infer the mask and semantic concepts , and the formula is as follows: (9); (10); wherein, represents the prediction operation for the mask; represents the prediction operation for the semantics. The mask represents the extracted landslide area, that is, it indicates the landslide location area in the remote sensing image. The semantic concept represents the predicted category or statement, such as "landslide", "landslide area affecting the road", etc.

[0081] The deep learning model constructed in this application introduces the DenseASPP module into the multi-modal segmentation architecture based on the attention mechanism, giving full play to the advantages of both. The deep learning model can not only perform fine-grained segmentation of global and local features simultaneously, but also extract multi-scale information of the image at different scales through dilated convolution, and enhance the feature expression ability of the network by densely connecting feature maps at different scales, thereby improving the accuracy of landslide area extraction and recognition.

[0082] Step 3: Divide the multi-modal dataset into a training set, a validation set, and a test set, which are used to train, validate, and test the deep learning model respectively. The trained deep learning model is used as the landslide extraction model.

[0083] Wuhan University created an open remote sensing landslide dataset named Bijie Landslide Dataset to support the development of automatic landslide detection methods. The Bijie Landslide Dataset includes satellite optical remote sensing images, shape files of landslide boundaries, and digital elevation models. The remote sensing images in the Bijie Landslide Dataset consist of 770 landslide images and 2003 non-landslide images. All the images are from the TripleSat satellite images taken during the period from May to August 2018 and have been cropped.

[0084] In a specific embodiment, in order to train, validate, and test the deep learning model, 2773 remote sensing images in the dataset are cleaned into 770 landslide images and 770 non-landslide images according to the positive-negative sample ratio of 1:1, totaling 1540 images. The cleaned landslide images and non-landslide images are divided into a training set, a validation set, and a test set according to the ratio of 8:1:1. There are 1232 images in the training set, 154 images in the validation set, and 154 images in the test set. The positive-negative sample ratio in the training set, validation set, and test set is 1:1.

[0085] Based on Figure 3 the new deep learning model architecture shown, during the training and inference process of the landslide for the deep learning model, first, the text encoder converts the text query into a text prompt. At the same time, a remote sensing image with a size of 512×512×3 is input into the image encoder to extract the multi-scale feature maps of the landslide image, namely Res2 (high resolution, rich in detailed information), as well as Res3, Res4, and Res5 (low resolution, strong in semantic information). Among them, the feature maps Res4 and Res5 do not directly enter the image-text joint semantic space but are enhanced by the DenseASPP module and then processed subsequently.

[0086] Then, the shallow feature maps Res2 and Res3 and the enhanced feature maps Res4 and Res5 are, on the one hand, converted into visual prompts through the visual sampler and all types of non-text queries. On the other hand, they are mapped together with the text prompt, visual prompt, and memory prompt into the image-text joint semantic space, and after the scale alignment operation, they are transmitted to the decoder in a unified form. Among them, the memory prompt is the memory prompt of the current stage obtained through cross-attention in the decoder and is used for the next training. Since there is no previous stage during the first calculation, there is no memory prompt.

[0087] Finally, the decoder uses self-attention and cross-attention mechanisms to interact collaboratively based on learnable queries and various prompts, and outputs masked embeddings and class embeddings, thereby generating the segmentation result of the landslide (output in the form of a mask) and semantic concepts.

[0088] The training method proposed in this application is summarized in the following PyTorch-style pseudocode. It should be noted that the parameters / functions in the pseudocode cannot be represented in italics or other forms, so they will be different from the parameter / function forms in the previous text. Refer to the comments after each line of code. The comments of the corresponding code start with the symbol "#".

[0089] The pseudocode for the training process of the deep learning model is as follows: Input: Remote sensing image img: [B, 3, H, W]; positive mask pm and negative mask nm: [B, 1, H, W]; text query: txt [a, b, c......]; #B represents the batch size, H and W represent the height and width. The positive and negative masks can come from remote sensing image labels and non-text queries Variables: Learnable query Qh; self-attention mask msa between query Q and prompt P Functions: Image encoder: Img_Encoder( ), text encoder: Text_Encoder( ), visual sampler: Visual_Sampler( ), DenseASPP module: DenseASPP( ), cross-attention mechanism: feature_attn( ), self-attention mechanism: prompt_attn( ), output: output( ) def init( ): #Initialization operation Qo, Qt, Qv = Qh.copy( ); #Initialize object query Qo, text query Qt, and visual query Qv Fv, Pt = img_Encoder(img), Text_Encoder(txt); #Fv and Pt represent image features (feature maps) and text prompts respectively Fv[4],Fv[5]= DenseASPP(Fv[4]), DenseASPP(Fv[5]); #Fv[4],Fv[5] represent the image features at the 4th and 5th scales output by the image encoder Pv = Visual_Sampler(Fv, pm, nm);#Sample visual prompt Pv from image features, positive and negative masks def Deep_Learning_Model_Decoder(Fv, Qo, Qt, Qv, Pv, Pt, Pm): # Define the decoder of the deep learning model Qo, Qt, Qv = feature_attn(Fv, Qo, Qt, Qv,); # Perform cross-attention calculation between the query and image features Qo, Qt, Qv = prompt_attn(msa, Qo, Qt, Qv, Pv, Pt, Pm); # Perform self-attention calculation between the query and the prompt, where Pm represents the memory prompt Om, Oc, Pm = output(Fv, Qo, Qt, Qv, Pv, Pt, Pm) # Calculate the mask Om and the class (semantic concept) output Oc def forward(img, pm, nm, txt): # Training forward propagation Fv, Qo, Qt, Qv, Pv, Pt = init(); Pm = None; # Initialize variables for i in range(max_iter): # Start iteration Om, Oc, Pm = Deep_Learning_Model_Decoder(Fv, Qo, Qt, Qv, Pv, Pt, Pm); # Decoder operation

[0090] Among them, the positive mask pm refers to the landslide area mask, and the negative mask nm refers to the background area mask. The learnable query Qh is a trainable parameter randomly initialized when building the model and automatically learns the most appropriate representation in a data-driven manner throughout the training process. msa refers to the self-attention mask between the query Q (such as Qo, Qt, Qv) and the prompt P (such as Pv, Pt, Pm). It is a weight matrix converted by calculating the similarity between Q and P. Matching pairs below a certain threshold will be weakened or even masked, thus forming a binary or continuous-valued mask matrix, which in turn suppresses irrelevant or invalid information interaction.

[0091] The construction and training of the deep learning model were carried out under the PyTorch 1.8.1 framework, and GPU acceleration was performed using CUDA 10.2.0. All comparative experiments were uniformly trained for 50 epochs, with a batch size of 4, an initial learning rate of 6e-05, and a polynomial decay strategy was adopted to adjust the learning rate. At the same time, a learning rate warm-up mechanism was introduced. At the beginning of training, the learning rate gradually increased from a lower value to the preset initial value. In the initial stage of training, the learning rate gradually increased from a small value to the set initial learning rate, and the increase of the learning rate during the warm-up process was linear. The minimum value of the learning rate decay was not lower than 1.0000000000000002e-06. The specific parameters during the training process are shown in Table 1 below.

[0092] Table 1 Experimental parameters and training settings

[0093] In the landslide recognition task, the landslide areas often have different spatial scales and morphological changes. The deep learning model can effectively overcome this problem through its unique structure, so as to better capture multi-scale information and significantly improve the recognition accuracy. To verify the effectiveness of the deep learning model proposed in this application, experiments were carried out on the Bijie landslide dataset of Wuhan University. The dataset was divided into a training set, a validation set, and a test set in a ratio of 8:1:1, and the positive and negative sample ratio was ensured to be 1:1. The training set and the validation set were used to train and validate the deep learning model, and the performance of the improved deep learning model was evaluated on the test set. The parameters of the trained deep learning model were encapsulated and used as a landslide extraction model during landslide recognition.

[0094] The trained deep learning model (Deep Learning Model, denoted as DLM) was tested using the test set, and the visualization results obtained on the test set were compared with the vision large model SEEM (Segment Everything Everywhere All at Once) model. The results are as Figure 8 shown. It can be seen from the Figure 8 results shown that the deep learning model constructed in this application is superior to the SEEM model in landslide recognition ability. The newly proposed deep learning model can not only generate more realistic landslide edges, but also significantly improve the recognition accuracy and stability in complex environments.

[0095] Furthermore, the global average accuracy (aAcc), average F-score (mFscore), average precision (mPrecision), average recall (mRecall), average intersection over union (mIoU), and average Dice coefficient (mDice) are used to compare with traditional models such as FCN, DeepLab v3+, SegFormer, and SEEM. Among them, aAcc reflects the ability of the model to correctly classify pixels on the entire image, which is the ratio of the number of all correctly classified pixels to the total number of pixels, and the formula is as follows: (11); Wherein, is the total number of categories, represents the number of the -th category correctly identified as the positive class, represents the number of the -th category correctly identified as the negative class, refers to the total number of actually positive classes, is the total number of actually negative classes. is the calculated global average accuracy (aAcc) value.

[0096] mFscore, mPrecision, and mRecall comprehensively measure the accuracy and integrity of the model in identifying specific categories. mPrecision is the ratio of the number of correctly predicted positive classes (TP) to the number of all predicted positive classes (TP + FP), as shown in formula (12). mRecall is the ratio of the number of correctly predicted positive classes to the number of actually positive classes, as shown in formula (13). mFscore is the harmonic mean of precision and recall, as shown in formula (14).

[0097] (12); (13); (14).

[0098] Wherein, represents the number of the -th category correctly identified as the positive class, represents the number of the -th category correctly identified as the negative class, represents the number of other categories misidentified as the -th category, represents the number of the -th category misidentified as other categories. , and They are the calculated average precision (mPrecision) value, average recall (mRecall) value, and average F-score (mFscore) value respectively.

[0099] mIoU and mDice evaluate the overlap between the predicted region and the ground truth region, which are key metrics in the field of semantic segmentation. mIoU is the ratio of the intersection of the predicted positive class and the actual positive class to their union, as shown in Equation (15). The mDice coefficient is the ratio of twice the intersection of the predicted positive class and the actual positive class to the sum of their respective quantities, which is similar to the F-score, as shown in Equation (16).

[0100] (15); (16); where and are the values of the calculated mean intersection over union (mIoU) and the mean Dice coefficient value respectively.

[0101] These metrics are used to evaluate the performance of the trained deep learning model (landslide extraction model), and by comparing with some visualized results of the test set, it is proved that the deep learning model of this application has the advantages of effectiveness and stability in the high-resolution urban landslide extraction task. The quantitative evaluation results of different models on the test set are compared in Table 2 below.

[0102] Table 2 Quantitative evaluation results of different models on the test set

[0103] It can be seen from the results shown in Table 2 that compared with the traditional FCN, DeepLab v3+, SegFormer, and SEEM models, the deep learning model (DLM) of this application performs better in metrics such as aAcc, mFscore, mPrecision, mRecall, mIoU, and mDice, indicating the effectiveness and stability of the deep learning model of this application in the landslide extraction task.

[0104] To comprehensively evaluate the performance of the landslide extraction model of this application, in addition to the indicators of semantic segmentation, key evaluation indicators such as Panoptic Quality (PQ), Segmentation Quality (SQ), and Recognition Quality (RQ) are also adopted to evaluate the recognition ability of the deep learning model (DLM) and the SEEM model. For interactive segmentation, the number of clicks (NoC) metric is used to evaluate the interactive segmentation performance, which measures the number of clicks required to achieve a certain IoU (Intersection over Union), namely 50%, 85%, and 90%, denoted as NoC50, NoC85, and NoC90 respectively. These indicators not only reflect the accuracy and robustness of the model in the image segmentation task but also can fully measure its comprehensive ability in object recognition and region segmentation. Among them, SQ focuses on the segmentation accuracy of the target area; RQ calculates the recognition accuracy of the target, that is, the indicator to judge whether each target is correctly recognized; and PQ combines SQ and RQ, and the formula is as follows.

[0105] (17); (18); (19); Among them, represents the correctly matched predicted instances, represents the wrongly predicted instances, represents the truly existing instances that are missed. represents the predicted instances and the truly existing instances of the intersection over union. , and represent the calculated RQ, SQ, and PQ values respectively. The comparison results are shown in Table 3 below.

[0106] Table 3 Comparison of Quantitative Evaluation Results of Deep Learning Model and SEEM Model on the Test Set

[0107] In Table 3, the smaller the value of the number of clicks (NoC) indicator, the better. Taking NOC50 as an example, the SEEM model on average requires 13.72 clicks to achieve 50% accuracy, while the deep learning model (DLM) of the present application on average only requires 8.59 clicks. In the interactive segmentation task, the deep learning model of the present application can not only quickly generate the mask and corresponding labels of the landslide area through simple clicks or scribbles, but also supports more efficient user interaction methods, such as single-clicking or stroking on the reference image to intelligently identify and segment areas with similar semantics in the target image. This ability mainly benefits from the visual sampler introduced in the training process of the deep learning model and the dynamic alignment technology of the joint visual-semantic space, enabling the model to unify different types of spatial queries (such as points, boxes, scribbles, polygons, and masks), thereby enhancing the flexibility and interactivity of the segmentation task. A new memory prompt mechanism proposed by the deep learning model can also gradually optimize the segmentation results during multiple interactions, enabling the effective transfer of the previously generated mask knowledge and guiding the optimization of the current training batch.

[0108] Figure 9 Shows the segmentation ability of the deep learning model of the present application in different task scenarios. Figure 9 Part (a) in it shows that after learning from the training set, the deep learning model can automatically identify the landslide body (landslide area) without additional prompts when inputting remote sensing images, and accurately segment the mask boundary. Figure 9 Part (b) in it shows the intention-driven interactive segmentation ability of the deep learning model of the present application. Under the condition of providing click or scribble prompts, the deep learning model can accurately extract the landslide area according to the user's input intention. Figure 9 Part (c) in it reflects the cross-scene reference segmentation ability of the deep learning model of the present application. Experts can mark typical features on historical landslide remote sensing images (reference images), and the deep learning model can identify areas with similar visual semantics in the target image based on the joint space matching method, thereby realizing landslide analogy recognition under zero-sample conditions. This ability significantly reduces the annotation cost of new scenarios and improves the generalization performance and applicability of the model.

[0109] Step 4: Use the landslide extraction model to identify the landslide area in the remote sensing image and issue a landslide disaster warning based on the change of the landslide area.

[0110] When identifying landslide areas, simply input the current remotely sensed image to be identified, along with the corresponding text and non-text queries, into the landslide extraction model, and the corresponding mask and semantic concepts can be output. The mask precisely outlines the contour of the landslide area in the remotely sensed image, and the semantic concepts indicate the category or statement corresponding to the predicted mask. Through time-series image analysis, monitoring the spatial distribution and morphological changes (such as location, shape, area, etc.) of landslides in the same area over a period of time, and combining the dual-index early warning method of deformation speed and deformation area, real-time, fast, and accurate early warning of landslide disasters can be achieved.

[0111] The landslide extraction model trained in this invention based on multi-modal data fusion and multi-scale feature dense connection can not only perform fine-grained segmentation of global and local features simultaneously, but also extract multi-scale information of the image at different scales through dilated convolution, and enhance the feature expression ability of the network by densely connecting feature maps at different scales, thereby improving the landslide recognition accuracy and efficiency. In addition, to verify the model accuracy, a large number of effectiveness verification experiments were conducted on the public landslide dataset of Wuhan University and compared with various advanced extraction models. The experimental results show that the deep learning model proposed in this application is effective. Specifically, the experimental results can be summarized as follows: 1) In terms of quantitative evaluation, the deep learning model proposed in this application has significant improvements in multiple key performance indicators (such as aAcc, mIoU, and mDice) compared to mainstream semantic segmentation models (such as SEEM, FCN, DeepLab v3+, etc.). 2) In terms of visualization results, the deep learning model proposed in this application can effectively identify and extract landslide features in complex environments. Especially when dealing with small-scale landslide features, the deep learning model of this application demonstrates strong feature extraction ability and excellent edge detail recovery performance, further verifying its superior performance in practical applications.

[0112] The reason for this advantage is that the introduction of the ASPP module cleverly combines dilated convolution and spatial pyramid pooling, thereby enhancing the deep learning model's ability to perceive image details and global structure. The use of dilated convolution enables the network to expand the receptive field without increasing the computational cost and capture a larger range of context information; while spatial pyramid pooling further improves the model's feature extraction ability at different scales through multi-scale processing. In this way, the ASPP module not only enhances the sensitivity to details but also effectively fuses multi-scale context information, significantly improving the segmentation accuracy of the landslide area and the clarity of the boundary. Based on this, DenseASPP, by introducing the design concept of DenseNet, makes full use of its efficient information flow and gradient transfer characteristics. By means of dense connections, the utilization efficiency of features is improved, enabling each layer to receive information from all previous layers, thus avoiding information loss and the problem of gradient disappearance. Therefore, the deep learning model proposed by combining the DenseASPP module with multi-modal data fusion can demonstrate excellent performance in the landslide extraction task. Consequently, this application effectively improves the accuracy and efficiency of landslide area recognition and landslide disaster warning through deep learning technology, providing scientific support for geological disaster risk control.

[0113] In an exemplary embodiment, this application also provides a computer device, which may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface, and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for the processor to exchange information with external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. The computer program, when executed by the processor, implements the method for landslide disaster warning based on deep learning.

[0114] In an exemplary embodiment, this application also provides a computer-readable storage medium, on which a computer program is stored. The computer program, when executed by the processor, implements the method for landslide disaster warning based on deep learning.

[0115] In an exemplary embodiment, this application also provides a computer program product, including a computer program. The computer program, when executed by the processor, implements the method for landslide disaster warning based on deep learning.

[0116] Those of ordinary skill in the art can understand that all or part of the processes in the above-described method embodiments can be completed by hardware related to computer program instructions. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the method embodiments as described above. Among them, any reference to a memory or other medium provided in the embodiments of the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0117] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant regulations.

[0118] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0119] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, based on the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A landslide disaster warning method based on deep learning, characterized in that, including: Obtain a remote sensing landslide dataset and perform preprocessing to construct a multimodal dataset; Each sample in the multimodal dataset includes a remote sensing image, a text query, a non-text query, and a corresponding mask; Construct a deep learning model, which includes an encoder, a DenseASPP module, a decoder, and a prediction head; the encoder includes an image encoder, a text encoder, and a visual sampler; the DenseASPP module refers to a densely connected atrous spatial pyramid pooling module; Divide the multimodal dataset into a training set, a validation set, and a test set, which are used to train, validate, and test the deep learning model respectively. Use the trained deep learning model as a landslide extraction model; Use the landslide extraction model to identify the landslide area in the remote sensing image and issue a landslide disaster warning based on the change in the landslide area.

2. The landslide disaster warning method based on deep learning according to claim 1, wherein The obtaining of the remote sensing landslide dataset and performing preprocessing to construct a multimodal dataset specifically includes: Obtain a remote sensing landslide dataset, which includes multiple landslide images, multiple non-landslide images, and corresponding label information; both the landslide images and non-landslide images are remote sensing images; Use the landslide images as positive samples and the non-landslide images as negative samples, and clean the remote sensing landslide dataset according to the requirement that the ratio of positive to negative samples is 1:1 to obtain a cleaned dataset; Save the label information of each sample in the cleaned dataset in a structured text form and convert it into a corresponding mask; the mask corresponding to the landslide area is called a positive mask, and the mask corresponding to the background area is called a negative mask; Automatically imitate users to randomly generate a text query and a non-text query corresponding to each sample; the text query includes a category and a statement; the non-text query includes points, boxes, scribbles, polygons, and example images; Store the remote sensing image, text query, non-text query, and mask of each sample in one-to-one correspondence to construct a multimodal dataset.

3. The landslide disaster warning method based on deep learning according to claim 2, wherein The described image encoder is used to perform layer-by-layer convolution and downsampling operations on the input remote sensing image to generate multi-scale feature maps with different resolutions and semantic information levels. The multi-scale feature maps are divided into original feature maps and feature maps to be enhanced ; The visual sampler is used to adopt the formula to convert all types of non-text queries and multi-scale feature maps into visual cues ; where is the operation performed by the visual sampler; is a mask for points, boxes, scribbles, polygons, and / or sampling regions from example images; The text encoder is used to convert a text query into a text prompt .

4. The landslide disaster warning method based on deep learning according to claim 3, wherein, The DenseASPP module is used to perform feature enhancement and fusion on the feature map to be enhanced in the multi-scale feature map to obtain the enhanced feature map ; The enhanced feature map and the original feature map are jointly used as the feature map to be mapped .

5. The landslide disaster warning method based on deep learning according to claim 4, characterized in that The DenseASPP module includes a multi-scale atrous convolutional layer and a densely connected channel; The multi-scale dilated convolution layer contains multiple 3×3 dilated convolution layers, each with a different dilation rate; a 1×1 convolution layer is added before each dilated convolution layer; the dilated convolution layers are densely connected through dense connection channels; the output of each dilated convolution layer and the original input feature map to be enhanced are concatenated and then passed through a 1×1 convolution layer to obtain the enhanced feature map .

6. The landslide disaster early warning method based on deep learning according to claim 4, wherein The to-be-mapped feature map is mapped to the image-text joint semantic space together with the text prompt , the visual prompt and the memory prompt , and after the scale alignment operation, it is transmitted to the decoder in a unified form; On the one hand, the decoder generates a memory prompt for the current stage through a cross-attention mechanism with a mask , based on the formula ; where is the memory prompt of the previous stage; is the mask of the previous stage; is the feature map to be mapped in the current stage; ; ; On the other hand, the decoder outputs a masked embedding , based on the formula and a class embedding through a masked self-attention mechanism ; where represents a learnable query.

7. The landslide disaster early warning method based on deep learning according to claim 6, characterized in that The prediction head infers a mask based on the formula ; the mask represents the extracted landslide area; represents the prediction operation for the mask; the prediction head infers a semantic concept based on the formula ; the semantic concept represents the predicted class or statement; represents the prediction operation for the semantics.​​ 8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the deep learning-based landslide disaster warning method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the deep learning-based landslide disaster warning method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the deep learning-based landslide disaster warning method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method and device, computer equipment and storage medium

    CN113034506A

  • Landslide and surrounding ground feature identification method and system based on multi-modal feature fusion

    CN115830469A

  • Landslide automatic identification method and system based on visual large model, and computer equipment

    CN117671480A

  • Landslide disaster automatic identification method based on lightweight deep learning network

    CN117911866A

  • Zero sample surface defect segmentation method based on text guidance

    CN119785026A

Cited By

  • Road slope landslide debris flow disaster identification method, device, equipment and medium

    CN121746939A

  • Fault removal agent training method and device and electronic equipment

    CN122241242A

  • Landslide image segmentation method and device based on improved SegFormer, and medium

    CN122335878A