A landslide hazard identification method and system
Patent Information
- Application Number
- CN202611226460.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-13
- Publication Date
- 2026-09-15
AI Technical Summary
[0005]综上所述,现有智能化识别研究大多采用单任务网络分别对蠕滑型滑坡和承灾体进行识别,导致两类对象在建模过程中相互独立,缺乏关联信息的表达,从而难以实现滑坡隐患的高效识别
[0076] This invention constructs a multimodal, multi-task deep learning model, realizing end-to-end pixel-level feature segmentation and object-level feature spatial relationship prediction by combining MT-InSAR deformation rate maps and optical remote sensing data.
Smart Images

Figure CN122761202A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and system for identifying landslide hazards, belonging to the field of geological disaster monitoring and remote sensing information processing technology. Background Technology
[0002] Landslide hazards possess two attributes: a natural attribute, meaning the slope has already deformed or failed; and a social attribute, meaning landslide instability may threaten and damage people or affected structures in the relevant area. In recent years, with the rapid development of remote sensing technology and related algorithms, such as multi-temporal InSAR (MT-InSAR), high-resolution optical remote sensing, and airborne lidar, landslide hazard identification methods have gradually shifted from traditional personnel patrols to manual interpretation based on multi-source remote sensing platform observations. Analysis of the composition of landslide hazards reveals that they essentially contain two key objects: the intrinsic object—creep-type landslides—and the associated object—affected structures. When landslide hazards are in their early development stages, the slope typically exhibits slow deformation and lacks obvious geomorphic damage characteristics. In this situation, relying solely on spectral and textural information from optical images is insufficient for accurate identification. InSAR technology, however, can accurately invert surface deformation information using multi-temporal radar echo signals, thereby effectively distinguishing between stable and unstable slopes. Therefore, for landslide hazard identification at this stage, an identification system has gradually formed, primarily based on InSAR deformation information and supplemented by high-resolution optical remote sensing imagery. On the other hand, when a landslide has undergone or is undergoing significant deformation and damage, the slope surface often exhibits typical optical hazard characteristics, such as steep rear walls, tension cracks, armchair-shaped landforms, and leading-edge deposits. For these types of landslide hazards, research typically constructs an identification system primarily based on multi-temporal high-resolution optical imagery and supplemented by InSAR deformation information. The integrated application of these two technical systems is generally referred to as integrated remote sensing technology. Its core objective is to leverage the advantages of different remote sensing data to comprehensively interpret landslide hazards using deformation information, geomorphic features, and the spatial distribution characteristics of the affected bodies. Currently, this technical system has been successfully applied in multiple regions, and the overall technical approach is gradually maturing, but its identification efficiency remains low.
[0003] Identifying creeping landslides can efficiently acquire candidate areas for potential landslide hazards. With the widespread application of visual models such as Convolutional Neural Networks (CNNs) and Transformers in the field of intelligent recognition of optical remote sensing images, creeping landslide identification methods are gradually evolving from traditional threshold segmentation and mathematical morphology methods to intelligent identification based on deep learning. Currently, numerous research cases based on deep learning have emerged for the identification of landslide hazard entities (creeping landslides), with the main goal of improving the detection efficiency of landslide deformation areas. However, most of these studies can only identify the location and boundaries of landslide deformation areas within the monitoring period, making it difficult to further determine whether the landslide poses a disaster risk. Identification of disaster-bearing bodies provides crucial social attribute information for landslide hazard identification. Disaster-bearing bodies refer to human social entities within a certain area that may be affected by disasters, including humans, buildings, engineering facilities, the environment, and economic and cultural activities. Based on whether they have a physical form, disaster-bearing bodies can be divided into direct disaster-bearing bodies and indirect disaster-bearing bodies. Direct disaster-bearing bodies include buildings, roads, farmland, waterways, infrastructure, and vehicles, while indirect disaster-bearing bodies include economic and cultural activities and the ecological environment. Currently, conducting risk analysis and loss assessment on indirect disaster-bearing bodies is quite challenging; therefore, disaster prevention and mitigation research typically focuses primarily on direct disaster-bearing bodies. Furthermore, direct disaster-bearing bodies can be divided into fixed disaster-bearing bodies and mobile disaster-bearing bodies. Fixed disaster-bearing bodies include buildings, roads, farmland, and waterways, while mobile disaster-bearing bodies include vehicles, ships, and personnel. With the rapid development of sensor technology and artificial intelligence technology, the cost of acquiring remote sensing data has gradually decreased, and spatial and spectral resolution has significantly improved. This has led to a shift in the investigation methods for fixed disaster-bearing bodies from traditional manual remote sensing interpretation combined with on-site verification to computer-based intelligent interpretation, significantly improving identification efficiency. Taking building and road extraction as an example, relevant research methods can be divided into methods based on shallow features and methods based on deep learning. Among these, deep learning methods, with their powerful feature representation capabilities, have become the most mainstream technical means in disaster-bearing body extraction research.
[0004] Looking back at the initial design intent of visual models, their core objective is to construct an algorithm capable of simulating human visual perception and cognitive processes. When people understand images, they often comprehensively consider the spatial and contextual relationships between various elements within the image, combining their own experience and knowledge to understand and identify targets of interest from the overall scene perspective. Similarly, the problem of landslide hazard identification essentially involves the relationships between multiple objects. Specifically, landslide hazards typically include two core objects: the ontological object and the associated object. The ontological object is the creeping landslide, and the associated object is the affected body. The former can be obtained and quantified using temporal InSAR technology to acquire surface deformation information. Based on this, semantic segmentation models can accurately extract the creeping landslide ontological object. The latter can be effectively extracted from the affected body object using high-resolution optical remote sensing imagery combined with semantic segmentation models. However, the key to landslide hazard identification is not merely identifying a single object, but, given the existence of the ontological object, further determining whether the landslide poses a potential threat to the surrounding affected bodies. This process requires a comprehensive analysis of the spatial distribution relationship and relative positional relationship between the landslide ontological object and the affected body. Therefore, designing a model that can characterize the relationship between creeping landslides and the elements of the disaster-bearing body, and understand landslide hazard scenarios from the perspective of spatial relationships, is the key issue for realizing intelligent identification of landslide hazards.
[0005] In summary, most existing intelligent identification studies use single-task networks to identify creeping landslides and disaster-bearing bodies separately. This results in the two types of objects being independent of each other during the modeling process, lacking the expression of related information, thus making it difficult to achieve efficient identification of landslide hazards. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a method and system for identifying landslide hazards.
[0007] To solve the above-mentioned technical problems, the present invention is implemented using the following technical solution.
[0008] In a first aspect, the present invention discloses a method for identifying landslide hazards, comprising:
[0009] Acquire high-resolution optical remote sensing images of the target work area;
[0010] Obtain a low-resolution map of the surface deformation rate of the target work area;
[0011] The low-resolution surface deformation rate map is sampled to the same resolution as the high-resolution optical remote sensing image;
[0012] Optical remote sensing images with consistent resolution and surface deformation rate maps are input into a pre-trained multi-task multimodal deep learning network to identify the segmentation results of significant deformation areas and disaster-bearing bodies, as well as the prediction results of spatial relationships between elements.
[0013] Based on the land cover type of the target work area obtained in advance and multi-source remote sensing data, potential landslide significant deformation areas are screened from the segmentation results of the significant deformation areas and the disaster-bearing bodies. The spatial relationship prediction results between the elements are used to perform constraint analysis on the potential landslide significant deformation areas to determine the landslide hazard identification results.
[0014] Further, acquiring a low-resolution surface deformation rate map of the target working area includes:
[0015] Acquire synthetic aperture radar data, precise orbit data, and digital elevation model data for the target working area;
[0016] Based on synthetic aperture radar data, precise orbit data, and digital elevation model data of the target working area, the time-series surface deformation variables are inverted and the surface deformation rate is estimated using multi-time-series radar interferometry technology to determine the surface deformation rate map.
[0017] Furthermore, the recognition process of the multi-task multimodal deep learning network includes:
[0018] Semantic features of disaster-bearing elements are extracted from optical remote sensing images;
[0019] Deformation characteristics of significant deformation zones are extracted from the surface deformation rate map;
[0020] The semantic features and the deformation features are fused to obtain a multi-scale feature map;
[0021] Obtain the target query vector and target location code, and determine the segmentation results of the significant deformation area and the disaster-bearing body, and the spatial relationship prediction results between elements based on the multi-scale feature map, target query vector and target location code.
[0022] Furthermore, the input of consistent resolution optical remote sensing images and surface deformation rate maps into a pre-trained multi-task multimodal deep learning network to identify the segmentation results of significant deformation zones and disaster-bearing bodies, and the prediction results of spatial relationships between elements, includes:
[0023] The multi-task multimodal deep learning network includes a CNN feature extraction module, a multimodal feature fusion module, and a Transformer module;
[0024] Optical remote sensing images and surface deformation rate maps with consistent resolution are input into the CNN feature extraction module to extract semantic features of disaster-bearing elements and deformation features of significant deformation areas.
[0025] The semantic features of the disaster-bearing body elements and the deformation features of the significant deformation areas are superimposed and then input into the multimodal feature fusion module to generate a fused multi-scale feature map.
[0026] The multi-scale feature map, target query vector, and target location encoding are input into the Transformer module to obtain object element representation vector and relation representation vector. Based on the object element representation vector and relation representation vector, a subject-predicate-object triple relationship is predicted. The spatial relationship prediction result between elements is output according to the subject-predicate-object triple relationship. The target query vector includes object element query vector and relation query vector. The subject and the object represent the objects in the object element query vector, and the predicate represents the relation in the relation query vector.
[0027] By processing multi-scale feature maps and object feature representation vectors using a panoramic segmentation head, the object features are segmented to obtain the precise spatial range of the corresponding object and the segmentation results of the significant deformation region.
[0028] Further, the step of inputting the multi-scale feature map, target query vector, and target location encoding into the Transformer module to obtain object element representation vectors and relation representation vectors, and predicting subject-predicate-object triple relations based on the object element representation vectors and relation representation vectors, includes:
[0029] The j-th object feature representation vector is generated using a feedforward neural network (FFN). Distinguish into the j-th subject representation Representation of the j-th object ;
[0030] ;
[0031] ;
[0032] In the formula, This indicates that the feedforward neural network is used to extract the subject's representation features; This indicates that the feedforward neural network is used to extract object representation features; This indicates that a feedforward neural network is used to extract relational representation features; This represents the i-th relation. This represents the i-th relation representation vector; This represents the collection of subject representations; Represents a collection of object representations; Represents a set of relations;
[0033] The spatial relation matching module using class hints represents the relations respectively. Explicitly model the representational features of candidate object elements, and model the triplet construction task as a fill-in-the-blank question with hints;
[0034] Based on relational representation Given a hint, select the most suitable pair of objects from the candidate object representations and input them into the fill-in-the-blank question. Use cosine similarity calculation to complete the matching to construct a complete set of subject-predicate-object triple relations. ;
[0035] ;
[0036] ;
[0037] ;
[0038] In the formula, Indicates the transpose operation; Indicates the set of features representing the subject The middle corresponds to the relation representation The index value of the highest cosine similarity; This refers to the set of characteristics representing objects. The middle corresponds to the relation representation The index value of the highest cosine similarity; Indicates the first Each object represents an operation that maximizes the objective function; Norm operations are used to represent vectors; express and Vector dot product operation; express and Vector dot product operation; This represents the correspondence to the relational representation. The subject representation with the highest similarity; This represents the correspondence to the relational representation. The object representation with the highest similarity.
[0039] Furthermore, the training process of the multi-task multimodal deep learning network includes:
[0040] Construct a multi-task, multi-modal deep learning network that includes a CNN feature extraction module, a multi-modal feature fusion module, a Transformer encoder, and a Transformer decoder;
[0041] Obtaining a training set, the construction of which includes: delineating significant deformation zones based on historical surface deformation rate maps, delineating disaster-bearing body elements by combining historical high-resolution optical imagery, and labeling the spatial relationship between significant deformation zones and disaster-bearing bodies based on spatial relationship knowledge criteria to construct the training set required for model training; the construction of the spatial relationship knowledge criteria includes: assessing the potential threat of significant deformation zones to disaster-bearing bodies using the principle of spatial proximity to construct spatial relationship knowledge criteria, which include: inclusion, proximity, and connection;
[0042] Construct the overall loss function ;
[0043] Based on the training set and the overall loss function The multi-task multimodal deep learning network is trained to obtain a trained multi-task multimodal deep learning network;
[0044] The overall loss function Represented as:
[0045] ;
[0046] In the formula, The loss function is the Dice similarity coefficient. The cross-entropy loss function; For triple query set; This represents the predicted segmentation result of part k in the i-th predicted triplet. This represents the predicted category of the k-th part in the i-th predicted triple; This represents the true segmentation result of the k-th part in the true triplet of the optimal matching; This represents the true class of the k-part of the true triplet in the optimal matching; Represents the index of the true triple under the optimal mapping; Represents the set of subject representations; Represents a set of object representations; This represents the set of predicate representations.
[0047] Secondly, the present invention also discloses a landslide hazard identification system, comprising:
[0048] The first acquisition module is used to acquire high-resolution optical remote sensing images of the target work area;
[0049] The second acquisition module is used to acquire a low-resolution surface deformation rate map of the target working area; and to sample the low-resolution surface deformation rate map to the same resolution as the high-resolution optical remote sensing image.
[0050] The identification module is used to input optical remote sensing images and surface deformation rate maps with consistent resolution into a pre-trained multi-task multimodal deep learning network to identify the segmentation results of significant deformation areas and disaster-bearing bodies, as well as the prediction results of spatial relationships between elements.
[0051] The determination module is used to screen potential landslide significant deformation zones from the segmentation results of the significant deformation zones and disaster-bearing bodies based on the land cover type and multi-source remote sensing data of the target working area obtained in advance, and to perform constraint analysis on the potential landslide significant deformation zones using the spatial relationship prediction results between the elements to determine the landslide hazard identification results.
[0052] Furthermore, the multi-task multimodal deep learning network includes a CNN feature extraction module, a multimodal feature fusion module, and a Transformer module;
[0053] The CNN feature extraction module is used to extract semantic features of disaster-bearing body elements and deformation features of significant deformation areas based on optical remote sensing images and surface deformation rate maps with consistent resolution.
[0054] The multimodal feature fusion module is used to fuse the semantic features of the superimposed disaster-bearing body elements and the deformation features of the significant deformation area to generate a fused multi-scale feature map.
[0055] The Transformer module is used to generate object element representation vectors and relation representation vectors based on the multi-scale feature map, target query vector, and target location encoding. Based on these vectors, it predicts subject-predicate-object triple relationships and outputs spatial relationship prediction results between elements. The target query vector includes object element query vectors and relation query vectors. The subject and object represent the objects in the object element query vector, and the predicate represents the relation in the relation query vector. The panoramic segmentation head processes the multi-scale feature map and object element representation vectors to segment the object elements, obtaining the precise spatial range of the corresponding objects and the segmentation results of the significant deformation region.
[0056] Furthermore, the Transformer module also includes a triplet construction unit for:
[0057] The j-th object feature representation vector is generated using a feedforward neural network (FFN). Distinguish into the j-th subject representation Representation of the j-th object ;
[0058] ;
[0059] ;
[0060] In the formula, This indicates that the feedforward neural network is used to extract the subject's representation features; This indicates that the feedforward neural network is used to extract object representation features; This indicates that a feedforward neural network is used to extract relational representation features; This represents the i-th relation. This represents the i-th relation representation vector; This represents the collection of subject representations; Represents a collection of object representations; Represents a set of relations;
[0061] The spatial relation matching module using class hints represents the relations respectively. Explicitly model the representational features of candidate object elements, and model the triplet construction task as a fill-in-the-blank question with hints;
[0062] Based on relational representation Given a hint, select the most suitable pair of objects from the candidate object representations and input them into the fill-in-the-blank question. Use cosine similarity calculation to complete the matching to construct a complete set of subject-predicate-object triple relations. ;
[0063] ;
[0064] ;
[0065] ;
[0066] In the formula, Indicates the transpose operation; Indicates the set of features representing the subject The middle corresponds to the relation representation The index value of the highest cosine similarity; This refers to the set of characteristics representing objects. The middle corresponds to the relation representation The index value of the highest cosine similarity; Indicates the first Each object represents an operation that maximizes the objective function; Norm operations are used to represent vectors; express and Vector dot product operation; express and Vector dot product operation; This represents the correspondence to the relational representation. The subject representation with the highest similarity; This represents the correspondence to the relational representation. The object representation with the highest similarity.
[0067] Furthermore, the recognition module further includes: a training unit, used for:
[0068] Construct a multi-task, multi-modal deep learning network that includes a CNN feature extraction module, a multi-modal feature fusion module, a Transformer encoder, and a Transformer decoder;
[0069] Obtaining a training set, the construction of which includes: delineating significant deformation zones based on historical surface deformation rate maps, delineating disaster-bearing body elements by combining historical high-resolution optical imagery, and labeling the spatial relationship between significant deformation zones and disaster-bearing bodies based on spatial relationship knowledge criteria to construct the training set required for model training; the construction of the spatial relationship knowledge criteria includes: assessing the potential threat of significant deformation zones to disaster-bearing bodies using the principle of spatial proximity to construct spatial relationship knowledge criteria, which include: inclusion, proximity, and connection;
[0070] Construct the overall loss function ;
[0071] Based on the training set and the overall loss function The multi-task multimodal deep learning network is trained to obtain a trained multi-task multimodal deep learning network;
[0072] The overall loss function Represented as:
[0073] ;
[0074] In the formula, The loss function is the Dice similarity coefficient. The cross-entropy loss function; For triple query set; This represents the predicted segmentation result of part k in the i-th predicted triplet. This represents the predicted category of the k-th part in the i-th predicted triple; This represents the true segmentation result of the k-th part in the true triplet of the optimal matching; This represents the true class of the k-part of the true triplet in the optimal matching; Represents the index of the true triple under the optimal mapping; Represents the set of subject representations; Represents a set of object representations; This represents the set of predicate representations.
[0075] The beneficial effects achieved by this invention are as follows:
[0076] This invention constructs a multimodal, multi-task deep learning model, realizing end-to-end pixel-level feature segmentation and object-level feature spatial relationship prediction by combining MT-InSAR deformation rate maps and optical remote sensing data.
[0077] This invention fully considers the social attributes of whether a landslide is a disaster, models the spatial relationship between creeping landslides and the disaster-bearing bodies, and realizes the transformation from single landslide deformation identification to comprehensive landslide hazard identification. In this way, while identifying the landslide itself, it can also determine its potential threat to the surrounding disaster-bearing bodies, thereby improving the level of intelligence in landslide hazard identification. Attached Figure Description
[0078] Figure 1 This is a flowchart for identifying potential landslide hazards. Detailed Implementation
[0079] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.
[0080] Example 1: This example introduces a method for identifying landslide hazards, including:
[0081] S1: Acquire synthetic aperture radar (SAR) data, precise orbit data, and digital elevation model data for the working area; use MT-InSAR (Multi-Temporal Interferometric Synthetic Aperture Radar) technology to invert temporal surface deformation and estimate the surface deformation rate.
[0082] S2: Acquire high-resolution optical remote sensing images of the working area and construct knowledge criteria for the spatial relationship between significant deformation areas and disaster-bearing bodies; delineate significant deformation areas based on surface deformation rate results, delineate disaster-bearing body elements (rivers, roads, buildings) in combination with high-resolution optical images, and label the spatial relationship between significant deformation areas and disaster-bearing bodies to construct the sample library required for model training.
[0083] S3: Construct and train a multi-task multi-modal network for joint panoptic segmentation and relation prediction (PSRP-MTMNet) to obtain the segmentation results of salient deformation zones in the work area. Finally, under the constraints of spatial relationship information between salient deformation zones and disaster-bearing bodies, land cover type, slope, elevation, and other information, identify landslide hazards in the work area.
[0084] The technical process for identifying landslide hazards is as follows: Figure 1 As shown, the specific processing steps are described below:
[0085] S1: Acquire multi-temporal SAR images of the working area, precise orbital data for the corresponding dates (for calculating the spatiotemporal baseline of the interferogram), and digital elevation model data (for removing terrain behavior and image registration). After steps such as SAR data import, multi-view processing, image registration, interferometric pair connection based on the short spatiotemporal baseline principle, differential interferometry, phase filtering, spatial baseline simplification, phase unwrapping, tropospheric atmospheric delayed phase removal, terrain residual phase removal, singular value decomposition, and least squares regression deformation rate estimation, the surface deformation rate map of the working area is obtained.
[0086] S2: Acquire high-resolution remote sensing optical images of the work area, and delineate the surrounding disaster-bearing elements, such as buildings, roads, and rivers, based on the location of the deformation zone. Use preprocessing methods such as cropping and resampling to sample the low-resolution surface deformation rate map to the resolution of the high-resolution optical remote sensing image, ensuring the consistency of its multimodal data size. Finally, construct a knowledge criterion for the spatial relationship between salient deformation zones and disaster-bearing elements. Based on this knowledge criterion, label the spatial relationships between salient deformation zone elements and disaster-bearing element elements (roads, buildings, rivers), obtaining the object-level element (salient deformation zone, road, building, river) labels and object-level element triple spatial relationship (e.g., salient deformation zone-adjacent-building) labels required for model training.
[0087] S3: To achieve intelligent identification of landslide hazards, this invention proposes a multi-task, multi-modal deep learning network that combines panoramic segmentation and spatial relationship prediction. This network employs a hybrid architecture combining Convolutional Neural Networks (CNN) and Transformer, consisting of three parts: a CNN feature extraction module, a multi-modal feature fusion module, and a Transformer encoder and decoder module. Finally, based on the segmentation results of salient deformation zones, and combined with multi-source remote sensing data such as land cover type, slope, and elevation, potential salient deformation zones for landslides are extracted through threshold screening. Furthermore, constraint analysis is performed using the spatial relationship between the predicted salient deformation zones and the affected body to further discriminate and screen candidate deformation zones, ultimately identifying the spatial distribution of landslide hazards within the study area. The detailed technical process of the PSRP-MTMNet network is as follows:
[0088] (1) Optical remote sensing images and surface deformation rate maps of the same size were input into the ResNet-50 feature extraction network to extract key features from different modalities. The optical remote sensing images were mainly used to extract semantic features of disaster-bearing elements (such as buildings, roads, and rivers), while the surface deformation rate maps were used to extract deformation features of significant deformation areas. During model training, the officially released ResNet-50 pre-trained weights were loaded to accelerate network convergence and improve feature expression capabilities. Subsequently, the multi-scale feature maps extracted from the two branches were superimposed and input into the multi-modal feature fusion module. This module consists of a convolutional layer (Conv), a batch normalization layer (BN), and a rectified linear unit activation function (ReLU), aiming to achieve deep interaction and information fusion between different modal features, thereby obtaining richer scene semantic expressions. After multimodal feature fusion, the fused multi-scale feature map, along with the query embedding and positional encoding, is input into the Transformer encoder and decoder. The Transformer structure models global scene information, enabling the query vector to learn the latent representation of triple relationships in the scene graph. For each triple query, the network uses three independent feed-forward networks (FFNs) to predict the subject, predicate, and object, respectively, thus obtaining the subject-predicate-object triple relationship (e.g., salient deformation region-containment-building). Simultaneously, the network uses a panoptic segmentation head to segment the subject and object objects, obtaining the precise spatial extent of the corresponding objects.
[0089] (2) Predicting the spatial relationship between significant deformation zones and disaster-bearing bodies is a crucial step in the intelligent identification of landslide hazards. Traditional methods typically rely on object-pair features to implicitly model the relationship during the Transformer decoding stage, which can easily overlook some key relationship information. To address this issue, this invention uses a class-hinted spatial relationship matching module to match the relationship representation vectors. and object element representation vector Explicit modeling is employed to enhance the model's focus on key objects such as salient deformation regions and buildings. The query matching module models the triple construction task as a hinted fill-in-the-blank exercise. Given relational hints, the model must select the most appropriate pair of objects from the candidate object query to construct the complete subject-predicate-object triple. For example, with the relational hint "contains," the model needs to select "salient deformation region" and "building" from the object candidates to generate the triple "salient deformation region-contains-building." Specifically, this is represented by relational characteristics. As a hint, both the subject selector and the object selector must return the most suitable candidate to form a complete triple. This invention uses the cosine similarity index to calculate the object element representation (…). and Representation of a given relation The similarity is used to select the highest similarity result to determine the subject and object candidates. It is important to note that the selection of subjects and objects should be based on the correlation between object element representations and relation representations, rather than simple semantic similarity. Furthermore, the same object element representation may play different roles (subject or object) in different selectors. Therefore, this invention uses two independent feedforward networks (FFNs) to extract specific feature representations for subjects and objects respectively, thereby extracting specific features from the same object element representation vector. To obtain distinguishable subject representations With object representation The entire calculation process can be written as formula (1) - formula (4).
[0090] (1);
[0091] (2);
[0092] (3);
[0093] (4);
[0094] In the formula, This represents the representation vector of the j-th object element (at this point, the representation of the subject and the object are not distinguished). This represents the i-th relation representation vector; This indicates that the feedforward neural network is used to extract the subject's representation features; This indicates that the feedforward neural network is used to extract object representation features; This indicates that a feedforward neural network is used to extract relational representation features; and Let them represent the j-th subject and object representations, respectively; This represents the i-th relation. This represents the collection of features that characterize the subject. Represents the set of characteristics that represent an object; Represents a set of relational characterization features. This represents the characteristic of the j-th subject; This represents the characteristic of the j-th object; Indicates the transpose operation; Indicates the set of features representing the subject The corresponding relational representation features The index value of the highest cosine similarity; This refers to the set of characteristics representing objects. The corresponding relational representation features The index value of the highest cosine similarity; Indicates the first Each object represents an operation that maximizes the objective function. Norm operations are used to represent vectors; express and Vector dot product operation; express and Vector dot product operation.
[0095] Finally, a complete set of triples can be obtained. :
[0096] ;
[0097] In the formula, This represents a subject-predicate-object triple; This represents the i-th relation. This represents the correspondence to the relational representation. The subject representation with the highest similarity; This represents the correspondence to the relational representation. The object representation with the highest similarity.
[0098] (3) During the training phase, after the triplet prediction and object-level feature segmentation are completed, the cross-entropy loss function is applied. Loss calculation for relationship prediction; Dice similarity coefficient / F-1 score loss function The loss function used for object-level segmentation. The overall loss function can be expressed by formula (6).
[0099] ;
[0100] Indicates the first Predicted segmentation results and categories of the k-part of each predicted triple; This represents the true segmentation result of the k-th part in the true triplet of the optimal matching; This represents the true class of the k-part of the true triplet in the optimal matching; This represents the index of the true triple under the optimal mapping.
[0101] Example 2, based on the same inventive concept as Example 1, introduces a landslide hazard identification system, including:
[0102] The first acquisition module is used to acquire high-resolution optical remote sensing images of the target work area;
[0103] The second acquisition module is used to acquire a low-resolution surface deformation rate map of the target working area; and to sample the low-resolution surface deformation rate map to the same resolution as the high-resolution optical remote sensing image.
[0104] The identification module is used to input optical remote sensing images and surface deformation rate maps with consistent resolution into a pre-trained multi-task multimodal deep learning network to identify the segmentation results of significant deformation areas and disaster-bearing bodies, as well as the prediction results of spatial relationships between elements.
[0105] The determination module is used to screen potential landslide significant deformation zones from the segmentation results of the significant deformation zones and disaster-bearing bodies based on the land cover type and multi-source remote sensing data of the target working area obtained in advance, and to perform constraint analysis on the potential landslide significant deformation zones using the spatial relationship prediction results between the elements to determine the landslide hazard identification results.
[0106] In this embodiment, the identification module includes:
[0107] The CNN feature extraction module is used to extract semantic features of disaster-bearing body elements and deformation features of significant deformation areas based on optical remote sensing images and surface deformation rate maps with consistent resolution.
[0108] The multimodal feature fusion module is used to fuse the semantic features of the superimposed disaster-bearing body elements and the deformation features of the significant deformation area to generate a fused multi-scale feature map.
[0109] The Transformer module is used to generate object element representation vectors and relation representation vectors based on the multi-scale feature map, target query vector, and target location encoding. Based on these vectors, it predicts subject-predicate-object triple relationships and outputs spatial relationship prediction results between elements. The target query vector includes object element query vectors and relation query vectors. The subject and object represent the objects in the object element query vector, and the predicate represents the relation in the relation query vector. The panoramic segmentation head processes the multi-scale feature map and object element representation vectors to segment the object elements, obtaining the precise spatial range of the corresponding objects and the segmentation results of the significant deformation region.
[0110] In this embodiment, the Transformer module further includes a triplet construction unit, used for:
[0111] The j-th object feature representation vector is generated using a feedforward neural network (FFN). Distinguish into the j-th subject representation Representation of the j-th object ;
[0112] ;
[0113] ;
[0114] In the formula, This indicates that the feedforward neural network is used to extract the subject's representation features; This indicates that the feedforward neural network is used to extract object representation features; This indicates that a feedforward neural network is used to extract relational representation features; This represents the i-th relation. This represents the i-th relation representation vector; This represents the collection of subject representations; Represents a collection of object representations; Represents a set of relations;
[0115] The spatial relation matching module using class hints represents the relations respectively. Explicitly model the representational features of candidate object elements, and model the triplet construction task as a fill-in-the-blank question with hints;
[0116] Based on relational representation Given a hint, select the most suitable pair of objects from the candidate object representations and input them into the fill-in-the-blank question. Use cosine similarity calculation to complete the matching to construct a complete set of subject-predicate-object triple relations. ;
[0117] ;
[0118] ;
[0119] ;
[0120] In the formula, This represents the j-th subject. Represents the j-th object; Indicates the transpose operation; Indicates the set of features representing the subject The middle corresponds to the relation representation The index value of the highest cosine similarity; This refers to the set of characteristics representing objects. The middle corresponds to the relation representation The index value of the highest cosine similarity; Indicates the first Each object represents an operation that maximizes the objective function; Norm operations are used to represent vectors; express and Vector dot product operation; express and Vector dot product operation; This represents the correspondence to the relational representation. The subject representation with the highest similarity; This represents the correspondence to the relational representation. The object representation with the highest similarity.
[0121] In this embodiment, the recognition module further includes a training unit, used for:
[0122] Construct a multi-task, multi-modal deep learning network that includes a CNN feature extraction module, a multi-modal feature fusion module, a Transformer encoder, and a Transformer decoder;
[0123] Obtaining a training set, the construction of which includes: delineating significant deformation zones based on historical surface deformation rate maps, delineating disaster-bearing body elements by combining historical high-resolution optical imagery, and labeling the spatial relationship between significant deformation zones and disaster-bearing bodies based on spatial relationship knowledge criteria to construct the training set required for model training; the construction of the spatial relationship knowledge criteria includes: assessing the potential threat of significant deformation zones to disaster-bearing bodies using the principle of spatial proximity to construct spatial relationship knowledge criteria, which include: inclusion, proximity, and connection;
[0124] Construct the overall loss function ;
[0125] Based on the training set and the overall loss function The multi-task multimodal deep learning network is trained to obtain a trained multi-task multimodal deep learning network;
[0126] The overall loss function Represented as:
[0127] ;
[0128] In the formula, The loss function is the Dice similarity coefficient / F-1 score. The cross-entropy loss function; For triple query set; This represents the predicted segmentation result of part k in the i-th predicted triplet. This represents the predicted category of the k-th part in the i-th predicted triple; This represents the true segmentation result of the k-th part in the true triplet of the optimal matching; This represents the true class of the k-part of the true triplet in the optimal matching; Represents the index of the true triple under the optimal mapping; Represents the set of subject representations; Represents a set of object representations; This represents the set of predicate representations.
[0129] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0130] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0131] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0132] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0133] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A landslide hazard identification method characterized by, include: Acquire high-resolution optical remote sensing images of the target work area; Obtain a low-resolution map of the surface deformation rate of the target work area; The low-resolution surface deformation rate map is sampled up to the resolution of the high-resolution optical remote sensing image; Optical remote sensing images with consistent resolution and surface deformation rate maps are input into a pre-trained multi-task multimodal deep learning network to identify the segmentation results of significant deformation areas and disaster-bearing bodies, as well as the prediction results of spatial relationships between elements. Based on the land cover type of the target work area obtained in advance and multi-source remote sensing data, potential landslide significant deformation areas are screened from the segmentation results of the significant deformation areas and the disaster-bearing bodies. The spatial relationship prediction results between the elements are used to perform constraint analysis on the potential landslide significant deformation areas to determine the landslide hazard identification results.
2. The landslide hazard identification method according to claim 1, characterized in that, The acquisition of a low-resolution surface deformation rate map of the target working area includes: Acquire synthetic aperture radar data, precise orbit data, and digital elevation model data for the target working area; Based on synthetic aperture radar data, precise orbit data, and digital elevation model data of the target working area, the time-series surface deformation variables are inverted and the surface deformation rate is estimated using multi-time-series radar interferometry technology to determine the surface deformation rate map.
3. The landslide hazard identification method according to claim 1, characterized in that, The recognition process of the multi-task multimodal deep learning network includes: Semantic features of disaster-bearing elements are extracted from optical remote sensing images; Deformation characteristics of significant deformation zones are extracted from the surface deformation rate map; The semantic features and the deformation features are fused to obtain a multi-scale feature map; Obtain the target query vector and target location code, and determine the segmentation results of the significant deformation area and the disaster-bearing body, and the spatial relationship prediction results between elements based on the multi-scale feature map, target query vector and target location code.
4. The landslide hazard identification method according to claim 1, characterized in that, The process involves inputting consistent resolution optical remote sensing images and surface deformation rate maps into a pre-trained multi-task multimodal deep learning network to identify the segmentation results of significant deformation zones and disaster-bearing bodies, as well as the prediction results of spatial relationships between elements, including: The multi-task multimodal deep learning network includes a CNN feature extraction module, a multimodal feature fusion module, and a Transformer module; Optical remote sensing images and surface deformation rate maps with consistent resolution are input into the CNN feature extraction module to extract semantic features of disaster-bearing elements and deformation features of significant deformation areas. The semantic features of the disaster-bearing body elements and the deformation features of the significant deformation areas are superimposed and then input into the multimodal feature fusion module to generate a fused multi-scale feature map. The multi-scale feature map, target query vector, and target location encoding are input into the Transformer module to obtain object element representation vector and relation representation vector. Based on the object element representation vector and relation representation vector, a subject-predicate-object triple relationship is predicted. The spatial relationship prediction result between elements is output according to the subject-predicate-object triple relationship. The target query vector includes object element query vector and relation query vector. The subject and the object represent the objects in the object element query vector, and the predicate represents the relation in the relation query vector. By processing multi-scale feature maps and object feature representation vectors using a panoramic segmentation head, the object features are segmented to obtain the precise spatial range of the corresponding object and the segmentation results of the significant deformation region.
5. The landslide hazard identification method according to claim 4, characterized in that, The process involves inputting the multi-scale feature map, target query vector, and target location encoding into the Transformer module to obtain object element representation vectors and relation representation vectors. Based on these vectors, prediction is performed to obtain subject-predicate-object triple relations, including: The j-th object feature representation vector is generated using a feedforward neural network (FFN). Distinguish into the j-th subject representation Representation of the j-th object ; ; ; In the formula, This indicates that the feedforward neural network is used to extract the subject's representation features; This indicates that the feedforward neural network is used to extract object representation features; This indicates that a feedforward neural network is used to extract relational representation features; This represents the i-th relation. This represents the i-th relation representation vector; This represents the collection of subject representations; Represents a collection of object representations; Represents a set of relations; The spatial relation matching module using class hints represents the relations respectively. Explicitly model the representational features of candidate object elements, and model the triplet construction task as a fill-in-the-blank question with hints; Based on relational representation Given a hint, select the most suitable pair of objects from the candidate object representations and input them into the fill-in-the-blank question. Use cosine similarity calculation to complete the matching to construct a complete set of subject-predicate-object triple relations. ; ; ; ; In the formula, Indicates the transpose operation; Indicates the set of features representing the subject The middle corresponds to the relation representation The index value of the highest cosine similarity; This refers to the set of characteristics representing objects. The middle corresponds to the relation representation The index value of the highest cosine similarity; Indicates the first Each object represents an operation that maximizes the objective function; Norm operations are used to represent vectors; express and Vector dot product operation; express and Vector dot product operation; This represents the correspondence to the relational representation. The subject representation with the highest similarity; This represents the correspondence to the relational representation. The object representation with the highest similarity.
6. The landslide hazard identification method according to claim 1, characterized in that, The training process of the multi-task multimodal deep learning network includes: Construct a multi-task, multi-modal deep learning network that includes a CNN feature extraction module, a multi-modal feature fusion module, a Transformer encoder, and a Transformer decoder; Obtaining a training set, the construction of which includes: delineating significant deformation zones based on historical surface deformation rate maps, delineating disaster-bearing body elements by combining historical high-resolution optical imagery, and labeling the spatial relationship between significant deformation zones and disaster-bearing bodies based on spatial relationship knowledge criteria to construct the training set required for model training; the construction of the spatial relationship knowledge criteria includes: assessing the potential threat of significant deformation zones to disaster-bearing bodies using the principle of spatial proximity to construct spatial relationship knowledge criteria, which include: inclusion, proximity, and connection; Construct the overall loss function ; Based on the training set and the overall loss function The multi-task multimodal deep learning network is trained to obtain a trained multi-task multimodal deep learning network; The overall loss function Represented as: ; In the formula, The loss function is the Dice similarity coefficient. The cross-entropy loss function; For triple query set; This represents the predicted segmentation result of part k in the i-th predicted triplet. This represents the predicted category of the k-th part in the i-th predicted triple; This represents the true segmentation result of the k-th part in the true triplet of the optimal matching; This represents the true class of the k-part of the true triplet in the optimal matching; Represents the index of the true triple under the optimal mapping; Represents the set of subject representations; Represents a set of object representations; This represents the set of predicate representations.
7. A landslide hazard identification system, characterized in that, include: The first acquisition module is used to acquire high-resolution optical remote sensing images of the target work area; The second acquisition module is used to acquire a low-resolution surface deformation rate map of the target working area; The low-resolution surface deformation rate map is sampled to the same resolution as the high-resolution optical remote sensing image; The identification module is used to input optical remote sensing images and surface deformation rate maps with consistent resolution into a pre-trained multi-task multimodal deep learning network to identify the segmentation results of significant deformation areas and disaster-bearing bodies, as well as the prediction results of spatial relationships between elements. The determination module is used to screen potential landslide significant deformation zones from the segmentation results of the significant deformation zones and disaster-bearing bodies based on the land cover type and multi-source remote sensing data of the target working area obtained in advance, and to perform constraint analysis on the potential landslide significant deformation zones using the spatial relationship prediction results between the elements to determine the landslide hazard identification results.
8. The landslide hazard identification system according to claim 7, characterized in that, The multi-task multimodal deep learning network includes a CNN feature extraction module, a multimodal feature fusion module, and a Transformer module; The CNN feature extraction module is used to extract semantic features of disaster-bearing body elements and deformation features of significant deformation areas based on optical remote sensing images and surface deformation rate maps with consistent resolution. The multimodal feature fusion module is used to fuse the semantic features of the superimposed disaster-bearing body elements and the deformation features of the significant deformation area to generate a fused multi-scale feature map. The Transformer module is used to generate object element representation vectors and relation representation vectors based on the multi-scale feature map, target query vector, and target location encoding; predict the subject-predicate-object triple relationship based on the object element representation vectors and relation representation vectors; and output the spatial relationship prediction result between elements based on the subject-predicate-object triple relationship. The target query vector includes an object element query vector and a relation query vector; the subject and the object represent the objects in the object element query vector, and the predicate represents the relation in the relation query vector. By processing multi-scale feature maps and object feature representation vectors using a panoramic segmentation head, the object features are segmented to obtain the precise spatial range of the corresponding object and the segmentation results of the significant deformation region.
9. The landslide hazard identification system according to claim 8, characterized in that, The Transformer module also includes a triplet construction unit, used for: The j-th object feature representation vector is generated using a feedforward neural network (FFN). Distinguish into the j-th subject representation Representation of the j-th object ; ; ; In the formula, This indicates that the feedforward neural network is used to extract the subject's representation features; This indicates that the feedforward neural network is used to extract object representation features; This indicates that a feedforward neural network is used to extract relational representation features; This represents the i-th relation. This represents the i-th relation representation vector; This represents the collection of subject representations; Represents a collection of object representations; Represents a set of relations; The spatial relation matching module using class hints represents the relations respectively. Explicitly model the representational features of candidate object elements, and model the triplet construction task as a fill-in-the-blank question with hints; Based on relational representation Given a hint, select the most suitable pair of objects from the candidate object representations and input them into the fill-in-the-blank question. Use cosine similarity calculation to complete the matching to construct a complete set of subject-predicate-object triple relations. ; ; ; ; In the formula, Indicates the transpose operation; Indicates the set of features representing the subject The middle corresponds to the relation representation The index value of the highest cosine similarity; This refers to the set of characteristics representing objects. The middle corresponds to the relation representation The index value of the highest cosine similarity; Indicates the first Each object represents an operation that maximizes the objective function; Norm operations are used to represent vectors; express and Vector dot product operation; express and Vector dot product operation; This represents the correspondence to the relational representation. The subject representation with the highest similarity; This represents the correspondence to the relational representation. The object representation with the highest similarity.
10. The landslide hazard identification system according to claim 8, characterized in that, The recognition module includes: a training unit, used for: Construct a multi-task, multi-modal deep learning network that includes a CNN feature extraction module, a multi-modal feature fusion module, a Transformer encoder, and a Transformer decoder; Obtaining a training set, the construction of which includes: delineating significant deformation zones based on historical surface deformation rate maps, delineating disaster-bearing body elements by combining historical high-resolution optical imagery, and labeling the spatial relationship between significant deformation zones and disaster-bearing bodies based on spatial relationship knowledge criteria to construct the training set required for model training; the construction of the spatial relationship knowledge criteria includes: assessing the potential threat of significant deformation zones to disaster-bearing bodies using the principle of spatial proximity to construct spatial relationship knowledge criteria, which include: inclusion, proximity, and connection; Construct the overall loss function ; Based on the training set and the overall loss function The multi-task multimodal deep learning network is trained to obtain a trained multi-task multimodal deep learning network; The overall loss function Represented as: ; In the formula, The loss function is the Dice similarity coefficient. The cross-entropy loss function; For triple query set; This represents the predicted segmentation result of part k in the i-th predicted triplet. This represents the predicted category of the k-th part in the i-th predicted triple; This represents the true segmentation result of the k-th part in the true triplet of the optimal matching; This represents the true class of the k-part of the true triplet in the optimal matching; Represents the index of the true triple under the optimal mapping; Represents the set of subject representations; Represents a set of object representations; This represents the set of predicate representations.