Multi-mode intelligent interpretation method and system for collapse remote sensing monitoring
By integrating multi-source remote sensing data and domain knowledge through a multimodal intelligent interpretation method, the problems of data modality incompatibility and rigid interaction methods in the remote sensing interpretation of landslide erosion bodies were solved, achieving high-precision detection and classification of landslide erosion bodies and improving the intelligence and efficiency of the monitoring system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies for remote sensing interpretation of landslides suffer from problems such as incompatible data modes, lack of internalization of domain knowledge, and rigid interaction methods, resulting in insufficient monitoring accuracy and flexibility, making it difficult to meet the needs of multi-source information fusion and professional intent for complex geological disasters.
Employing a multimodal intelligent interpretation method, combining true-color imagery, near-infrared bands, and digital elevation models, and leveraging adaptive feature selection, multi-scale target localization, and knowledge graph empowerment, this approach achieves deep fusion and semantic understanding of multi-source data, supporting natural language interaction and domain knowledge integration.
It significantly improves the accuracy and flexibility of remote sensing monitoring of landslides, enables high-precision detection and classification of landslide erosion bodies, lowers the barrier to entry for using professional tools, and enhances the intelligence and efficiency of the system.
Smart Images

Figure CN121789035A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent interpretation of remote sensing images, geographic information systems and artificial intelligence. It relates to a multimodal intelligent interpretation method and system for remote sensing monitoring of landslides, specifically an intelligent monitoring method and system for landslide geological disasters that integrates multi-source remote sensing data and multimodal interactive information. Background Technology
[0002] Landslides, a severe form of soil erosion, are a typical geological hazard in the red soil regions of southern my country. Their development is accompanied by intense erosion and collapse, posing a serious threat to farmland, roads, and settlements. Therefore, timely and accurate monitoring and mapping of landslides are crucial. With the rapid development of Earth observation technology, high spatial resolution, multispectral remote sensing imagery, and high-precision digital elevation models are becoming increasingly available geospatial data, providing a solid data foundation for the macroscopic and dynamic monitoring of landslides.
[0003] Traditional remote sensing interpretation methods for landslides mainly rely on three technical approaches: 1. Traditional remote sensing and index methods; For example, the invention patent with patent application number 201910740430.X, "A method for judging the activity level of ridge collapse based on vegetation cover", and the invention patent with patent application number 201910553113.7, "A method for extracting ridge collapse information based on Sentinel-2A satellite remote sensing images", rely on manually designed remote sensing indices (such as vegetation index and bare soil index) and texture features, and identify and judge them through thresholds or simple models.
[0004] These methods have the following limitations: (1) They rely heavily on expert experience to select features, making it difficult to capture the complex and abstract nonlinear features of gully collapse, and the accuracy drops sharply in areas with fragmented forms and complex backgrounds. (2) They have a low level of intelligence, and are essentially "pattern matching" rather than "semantic understanding". They cannot understand what "gully collapse" means, let alone make adaptive adjustments based on the user's complex intentions.
[0005] 2. Employ deep learning model methods; For example, the invention patent with patent application number 202210591275.1, "A Method and System for Extracting Collapsed Landslides Based on an Improved U-net Model", and the invention patent with patent application number 202210591266.2, "An Automatic Recognition Method and System for Collapsed Landslides", both improve the ability to extract and fuse features from multi-source data by designing more complex and sophisticated neural network architectures (such as introducing 3D convolution, attention mechanisms, and channel switching modules).
[0006] These methods have the following limitations: (1) Poor flexibility, as these are "static" models. Once trained, their behavior is fixed. Users cannot guide the model to focus on specific areas or attributes in real time during inference using natural language or other means (e.g., "only look for collapses during active periods"). (2) Zero interactivity, as this is a "black box" automated process. When the model makes an incorrect judgment, users have no effective channels for intervention and correction, and can only accept the result or retrain.
[0007] (3) Process-oriented and mechanism-based analysis methods; For example, the invention patent with patent application number 202310024537.0, "Method, Device, Electronic Equipment and Storage Medium for Identifying Landslides", and the invention patent with patent application number 202110072337.3, "A Method and Device for Identifying and Predicting Internal Channels of Landslides", both improve reliability by designing multi-stage processing procedures or by combining geoscientific mechanism models for deeper analysis.
[0008] These methods have the following limitations: (1) Knowledge is not internalized. Whether it is a two-stage process or channel prediction, the rules or mechanisms they rely on are still external and hard-coded. The system itself does not have domain knowledge and cannot generalize. (2) Functions are singular and rigid. CN202310024537.0 focuses on the single task of "recognition", and CN202110072337.3 focuses on the single object of "channel". The whole system is rigid and difficult to extend to a unified platform that supports multiple tasks (such as recognition, classification, and change detection).
[0009] In recent years, computer vision technologies, represented by deep learning, have made groundbreaking progress. In particular, Meta's SegmentAnythingModel, with its powerful general zero-shot segmentation capabilities and flexible cueing mechanism, has brought revolutionary changes to the field of image segmentation. However, directly applying such general-purpose visual models to specialized remote sensing interpretation tasks such as the Benggang earthquake relief project faces severe challenges and incompatibility: 1. Data Modality Uniformity and Lack of Domain-Specific Features: Native models, such as SAM, are primarily designed and trained on RGB three-channel images of natural scenes. However, remote sensing interpretation, especially landslide identification, is a typical multi-source information fusion process. Models require not only visible light information but also multispectral bands sensitive to vegetation cover, such as near-infrared, to assist in identification, as well as key topographic factors like slope, aspect, and undulation provided by DEM data. Existing models lack the ability to effectively encode and fuse multispectral and topographic data, resulting in the waste of a large amount of valuable domain-specific feature information.
[0010] 2. Lack of Domain Knowledge and Insufficient Semantic Understanding: General models lack prior knowledge about the specific geological object of gully collapse. The development of gullies has specific geomorphic features (such as steep slopes), material composition (loose deposits), and evolutionary stages (active and stable phases). Existing models cannot understand this deep-seated domain knowledge, resulting in segmentation results that may only be visually similar regions, failing to ensure accuracy in a geological sense and making it difficult to perform precise, semantic segmentation based on the user's professional intent (such as "extracting all gullies in the active phase in the image").
[0011] 3. Rigid Human-Computer Interaction and Efficiency Bottlenecks: Existing interactive segmentation systems largely rely on mouse-based geometric prompts such as points and boxes, resulting in a limited range of interaction methods. In complex remote sensing interpretation, experts often seek more natural and efficient ways to express their intentions. For example, they might directly describe the data using voice ("circle the large landslide at the mountaintop") or refine it using existing vector boundaries. Current technologies lack integration with semantically rich prompts such as voice and text descriptions, failing to construct a multimodal, integrated intelligent interaction loop. This limits the incorporation of expert experience and further improvements in interpretation efficiency.
[0012] In summary, while existing technologies excel in general image segmentation, they suffer from three core bottlenecks when facing complex remote sensing geoscience interpretation tasks such as Benggang Mountain: data modality incompatibility, lack of domain knowledge internalization, and inflexible interaction methods. Although a few studies have attempted to introduce text prompts, these often remain at a superficial descriptive level, failing to achieve deep end-to-end fusion with multi-source remote sensing data (imagery + DEM), and further failing to systematically integrate prior knowledge such as voice interaction, geometric interaction, and knowledge graphs to jointly guide the interpretation process.
[0013] Therefore, there is an urgent need in this field for an innovative technical solution that can overcome the above limitations and build an intelligent interpretation system that can deeply integrate multi-source remote sensing data, fully understand multimodal interaction intentions, and effectively utilize domain knowledge, thereby comprehensively improving the accuracy, flexibility, and intelligence level of remote sensing monitoring of landslides. Summary of the Invention
[0014] The purpose of this invention is to build a more "intelligent" ecosystem by integrating multimodal interactive prompts (text, voice, geometry) with domain knowledge graphs, thereby achieving a leap from "automation" to "human-machine collaboration" in remote sensing monitoring of landslides, from "data-driven" to "knowledge-driven" and from "single function" to "flexible empowerment".
[0015] The method of this invention employs the following technical means: a multimodal intelligent interpretation method for remote sensing monitoring of landslides, comprising the following steps: Step 1: Acquisition and preprocessing of multi-source remote sensing data; the multi-source remote sensing data includes true-color RGB imagery, near-infrared NIR band data, and digital elevation model (DEM); Step 2: Input the multi-source remote sensing data into the gully erosion monitoring network and output five types of gully erosion bodies at preset scales; The landslide remote sensing monitoring network includes an adaptive feature selection module, an improved HyperACE module, and a multi-scale target localization module; the improved HyperACE module has 5 scales for both input and output. The adaptive feature selection module is a dual-branch, multi-level skip fusion adaptive feature selection encoder; the true-color image RGB and near-infrared band NIR are input into the first branch, and the digital elevation model DEM outputs the second branch; outputs B3, B4, B5, B6, and B7 are processed by the HyperACE improvement module and output as H3, H4, H5, H6, and H7, respectively. The multi-scale target localization module, through bidirectional prior weighting, five-stage sequential upsampling, and a lightweight tower structure with independent detection heads at each stage, simultaneously completes high-precision localization and scale-adaptive classification of five-scale erosion bodies in a single forward pass.
[0016] Preferably, in step 1, the preprocessing first performs geometric fine correction on the original image to ensure geographic coordinate accuracy, and performs radiometric calibration and atmospheric correction to eliminate the effects of illumination and aerosols; then, using the true-color image RGB as a reference, the DEM and NIR are registered, and the registration error is controlled within a preset range; finally, image block processing is performed to maintain spatial alignment.
[0017] Preferably, the adaptive feature selection module includes two parallel branches; each branch consists of a first Conv layer, a first DS_C3k2 layer, a second Conv layer, a second DS_C3k2 layer, a third Conv layer, a third DS_C3k2 layer, a first DSConv layer, a first A2C2f layer, a second DSConv layer, and a second A2C2f layer connected in sequence; the true-color image RGB and near-infrared band NIR are input into the first branch, and the digital elevation model (DEM) outputs into the second branch; the outputs of the first DS_C3k2 layers of the two branches are fused by a first Concat layer and output as B3; the outputs of the second DS_C3k2 layers of the two branches are fused by a second Concat layer and output as B4; the outputs of the third DS_C3k2 layers of the two branches are fused by a third Concat layer and output as B5; the outputs of the first A2C2f layers of the two branches are fused by a fourth Concat layer and output as B6; the outputs of the second A2C2f layers of the two branches are fused by a fifth Concat layer and output as B7. The DS_C3k2 layer comprises a sequentially connected Conv layer, a Split layer, and n DS_C3k layers. The outputs of the Split layer and the DS_C3k layer are fused by a Concat layer and processed by a Conv layer before being output. The DS_C3k layer comprises a sequentially connected Conv layer and n DS-Bottleneck layers. The original input data is processed by the Conv layer and then fused with the output of the DS-Bottleneck layer by a Concat layer, and further processed by the Conv layer before being output. The DS-Bottleneck layer comprises two sequentially connected DSConv layers. The original input data is added pixel-by-pixel to the output of the DSConv layer before being output. The DSConv layer comprises a sequentially connected DWConv layer, a PWConv layer, a BN layer, and a SiLU layer. The A2C2f layer includes a sequentially connected Conv layer, two Split layers, and n Bottleneck layers. The outputs of each of the two Split layers and the n Bottleneck layers are fused by a Concat layer and then further processed by a Conv layer before being output. After being processed by parallel Height Attention and Width Attention layers, the outputs are added pixel by pixel and then processed by a Conv layer before being output.
[0018] Preferably, the HyperACE improvement module first downsamples B3 and B4, upsamples B6 and B7, and then processes them with B5 through the Concat layer, Conv layer and Split layer. After processing by the first C3AH layer, the second C3AH layer and the DS-C3K layer set in parallel respectively, it is output through the Concat layer and the Conv layer. The C3AH layer includes a sequentially connected Conv layer, an Adaptive Hyperedge Generation layer, and a Hypergraph Convolution layer. The original input data is processed by the Conv layer and then merged with the output of the Hypergraph Convolution layer by the Concat layer and processed by the Conv layer before being output. The Adaptive Hyperedge Generation layer's input is processed by the Flatten layer, the parallel MaxPooling layer, and the AvgPooling layer, then fused by the Concat layer, and further processed by the Projection layer, the DynamicOffset layer, and the Global Proto layer. The outputs of the Dynamic Offset layer and the Global Proto layer are added pixel by pixel, and then multiplied pixel by pixel with the output of the Flatten layer after the Projection layer is processed. The Hypergraph Convolution layer includes a sequentially connected hyperedge-to-node aggregation layer, a first Feature Projection layer, a node-to-hyperedge propagation layer, and a second Feature Projection layer. The original data is added pixel by pixel to the output of the second Feature Projection layer before being output.
[0019] Preferably, the multi-scale target localization module includes a first Upsample layer, a sixth Concat layer, a fourth DS_C3k2 layer, a second Upsample layer, a seventh Concat layer, a fifth DS_C3k2 layer, a third Upsample layer, an eighth Concat layer, a sixth DS_C3k2 layer, a fourth Upsample layer, a ninth Concat layer, a seventh DS_C3k2 layer, a fourth Conv layer, a tenth Concat layer, an eighth DS_C3k2 layer, a fifth Conv layer, an eleventh Concat layer, a ninth DS_C3k2 layer, a sixth Conv layer, a twelfth Concat layer, a tenth DS_C3k2 layer, a seventh Conv layer, a thirteenth Concat layer, and an eleventh DS_C3k2 layer, all connected in sequence. B7 and H7 are added pixel by pixel and then input to the first Upsample layer. The following steps are performed: B6 and H6 are added pixel-by-pixel and then input into the sixth Concat layer; B5 and H5 are added pixel-by-pixel and then input into the seventh Concat layer; B4 and H4 are added pixel-by-pixel and then input into the eighth Concat layer; B3 and H3 are added pixel-by-pixel and then input into the ninth Concat layer; the output of the seventh DS_C3k2 layer is added pixel-by-pixel with H3 and then input into the fourth Conv layer; the output of the sixth DS_C3k2 layer is added pixel-by-pixel with H4 and then input into the tenth Concat layer; the output of the fifth DS_C3k2 layer is added pixel-by-pixel with H5 and then input into the eleventh Concat layer; the output of the fourth DS_C3k2 layer is added pixel-by-pixel with H6 and then input into the twelfth Concat layer; B7 and H7 are added pixel-by-pixel and then input into the thirteenth Concat layer; the output of the seventh DS_C3k2 layer is added pixel-by-pixel with H3 and then passed through the first detection head layer (detect). The output of the DS_C3k2 layer is the first type of preset scale erosion body after the _head. The output of the eighth DS_C3k2 layer is added to H4 pixel by pixel and then passed to the second detection head layer detect _head to output the second type of preset scale erosion body. The output of the ninth DS_C3k2 layer is added to H5 pixel by pixel and then passed to the third detection head layer detect _head to output the third type of preset scale erosion body. The output of the tenth DS_C3k2 layer is added to H6 pixel by pixel and then passed to the fourth detection head layer detect _head to output the fourth type of preset scale erosion body. The output of the eleventh DS_C3k2 layer is added to H7 pixel by pixel and then passed to the fifth detection head layer detect _head to output the fifth type of preset scale erosion body.
[0020] As a preferred option, the five types of erosion bodies at preset scales are further processed through a large model segmentation module, a natural language understanding module, and a knowledge graph empowerment module to output accurate classification results. The large model segmentation module uses an image encoder to perform pixel-level segmentation of the five types of eroded landslide bodies at preset scales, and outputs candidate region masks, each of which corresponds to a possible target instance and includes category labels and location information. The natural language understanding module is used to preprocess multimodal input into text, and after word segmentation and normalization, form a word sequence; the Transformer dynamically embeds and generates vectors, which are then fused through context encoding, semantic parsing, and knowledge enhancement to finally output a machine-executable structured object; The knowledge graph empowerment module integrates natural language semantics with candidate masks, links categories to entities, retrieves spectral topography rules to verify the rationality of the mask, reorders to suppress false detections, supplements attributes and generates interpretable descriptions, and finally outputs a structured result containing the optimal mask, corrected category confidence, and knowledge basis.
[0021] Preferably, the large model segmentation module specifically includes: (1) The input data is encoded by Image Encoder, multi-scale feature maps are extracted, and a series of high-level semantic representations are output as shared context information for subsequent decoders; (2) Input prompt information (box, text) into Prompt Encoder, where the bounding box is the normalized coordinates output by the multi-scale target localization module; the text prompt is the semantic vector output by the natural language understanding module; (3) The Prompt Encoder encodes the two prompts into a high-dimensional vector representation and aligns it spatially and semantically with the feature map output by the Image Encoder; (4) The Mask Decoder receives global features from the Image Encoder and prompt embeddings from the Prompt Encoder, and fuses the information from both through a cross-attention mechanism to achieve accurate localization of the target region; (5) Finally, output multiple candidate masks, each corresponding to a possible target instance and containing category label and location information.
[0022] Preferably, the context encoding uses a Transformer Encoder to calculate the dependencies between words through a self-attention mechanism and outputs an enhanced sequence of context vectors. The semantic parsing includes two parallel subtasks: intent recognition and slot filling. Intent recognition is used to obtain the classification results of the sentence, while slot filling is used to obtain the label sequence of each word. The knowledge enhancement fusion involves fusing the current token with the knowledge vector generated by the knowledge graph empowerment module and updating the semantic representation.
[0023] Preferably, the knowledge graph empowerment module takes as input the structured semantics output by the natural language understanding module and multiple candidate masks output by the large model segmentation module, with each mask accompanied by a category label and probability distribution. First, the mask categories are linked to knowledge graph entity nodes. Then, based on the user input type and mask category labels and positions obtained from the natural language understanding module, relevant triples and rules are retrieved. Finally, it is verified whether each mask satisfies the prior constraints. Based on the reasoning results, the knowledge graph empowerment module reorders the mask or suppresses unreasonable options, supplements the attribute information of the finally selected mask, and generates an interpretability description; finally, it outputs the enhanced structured result, including: the optimal mask ID, the corrected category and confidence level, and the knowledge explanation.
[0024] The technical solution adopted by the system of this invention is: a multimodal intelligent interpretation system for remote sensing monitoring of landslides, comprising: One or more processors; A storage device is provided for storing one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the multimodal intelligent interpretation method for remote sensing monitoring of landslides.
[0025] Compared with the prior art, the technical solution provided by the present invention has the following significant advantages and beneficial effects: (1) The natural language understanding module designed in this invention has a high degree of intelligence, can process complex natural language input, supports context awareness and geographic semantic understanding, and enables the system to respond more accurately to the diverse needs of users.
[0026] (2) The adaptive feature selection module designed in this invention achieves accurate feature fusion; by dynamically generating spatially adaptive weight maps, it changes the traditional fusion method of simply stacking multimodal data, and can perform pixel-level accurate calibration of basic visual features, effectively enhancing task-related features and suppressing irrelevant features and noise, thereby significantly improving the model's perception ability and recognition accuracy in complex scenes; at the same time, it innovatively combines data-driven high-level semantic context with knowledge-driven domain prior features to jointly guide the feature calibration process, so that the model decision not only conforms to the data distribution but also follows physical laws, thereby improving the interpretability and robustness of the model.
[0027] (3) The multi-scale target localization module designed in this invention overcomes the problem of detecting targets at extreme scales: by using a "high-frequency detail enhancer" and "global context attention" to asymmetrically enhance the "two poles" of the feature pyramid, it effectively solves the core pain points of traditional methods, such as serious loss of details of small targets and insufficient global perception of huge targets. It achieves accurate coverage of the entire scale spectrum: through specialized multi-scale design, the model can simultaneously and accurately capture the entire series of erosion bodies from small gullies to giant landslides, greatly improving the recall rate and boundary localization accuracy of the detection task.
[0028] (4) The large-scale model segmentation module designed in this invention realizes the professional empowerment of the general model: through the "semantic tutor" mechanism, the knowledge of the professional domain is deeply distilled into the general segmentation large model at the feature encoding level, so that it transforms the "aimless" general segmentation into "targeted" professional segmentation, fundamentally solving the problem of blind and inefficient segmentation of the large model in the professional domain. It takes into account both strong segmentation ability and high professional efficiency: while retaining the powerful zero-sample segmentation ability of the large model, it makes its output highly focused on the task-related target, greatly reducing invalid proposals and improving the efficiency of professional interpretation.
[0029] (5) The knowledge graph empowerment module designed in this invention realizes the efficient injection of symbolic knowledge into neural networks: it provides a lightweight, plug-and-play method to transform structured domain knowledge (triples) into feature maps and integrate them with the model. Human prior knowledge can be integrated into the model without designing complex loss functions, thereby enhancing the model's logical reasoning ability and generalization in data-scarce scenarios.
[0030] (6) The intelligent interaction module designed in this invention creates a natural and efficient human-computer collaboration paradigm: by constructing a “voice-CLIP-SAM” closed loop, it seamlessly connects human language description and high-level semantic understanding with the model’s segmentation ability, creating an interactive experience of “speaking and revising as you go, becoming more and more accurate as you speak”, which greatly reduces the threshold for using professional tools. Attached Figure Description
[0031] The technical solutions of the present invention will be further illustrated below using embodiments and specific implementation methods. In addition, some accompanying drawings are used in the description of the technical solutions. Those skilled in the art can obtain other drawings and the intent of the present invention from these drawings without any creative effort.
[0032] Figure 1 This is a schematic diagram of the method according to an embodiment of the present invention; Figure 2 This is a diagram of the remote sensing network structure for hill collapse monitoring according to an embodiment of the present invention; Figure 3This is a structural diagram of the adaptive feature selection module according to an embodiment of the present invention; Figure 4 This is a structural diagram of the improved HpyerACE module according to an embodiment of the present invention; Figure 5 This is a structural diagram of the multi-scale target localization module according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the large model segmentation module according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the intelligent interaction module in an embodiment of the present invention. Detailed Implementation
[0033] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0034] Please see Figure 1 This embodiment provides a multimodal intelligent interpretation method for remote sensing monitoring of landslides, characterized by the following steps: Step 1: Acquisition and preprocessing of multi-source remote sensing data; the multi-source remote sensing data includes true-color RGB imagery, near-infrared NIR band data, and digital elevation model (DEM); In one embodiment, the preprocessing first performs geometric fine correction on the original image to ensure geographic coordinate accuracy, and performs radiometric calibration and atmospheric correction to eliminate the effects of illumination and aerosols; then, using the true-color image RGB as a reference, the DEM and NIR are registered, and the registration error is controlled within a preset range; finally, image block processing is performed to maintain spatial alignment.
[0035] Step 2: Input the multi-source remote sensing data into the gully erosion monitoring network and output five types of gully erosion bodies at preset scales; Please see Figure 2 The remote sensing monitoring network for landslides includes an adaptive feature selection module, a HyperACE improvement module (Hypergraph adaptive correlation enhancement), and a multi-scale target localization module; the HyperACE improvement module has 5 scales for both input and output. The adaptive feature selection module is a dual-branch, multi-level skip fusion adaptive feature selection encoder; the true-color image RGB and near-infrared band NIR are input into the first branch, and the digital elevation model DEM outputs the second branch; outputs B3, B4, B5, B6, and B7 are processed by the HyperACE improvement module and output as H3, H4, H5, H6, and H7, respectively. The multi-scale target localization module, through bidirectional prior weighting, five-stage sequential upsampling, and a lightweight tower structure with independent detection heads at each stage, simultaneously completes high-precision localization and scale-adaptive classification of five-scale erosion bodies in a single forward pass.
[0036] In one implementation, please see Figure 3 The adaptive feature selection module includes two parallel branches; each branch consists of a first Conv layer, a first DS_C3k2 layer, a second Conv layer, a second DS_C3k2 layer, a third Conv layer, a third DS_C3k2 layer, a first DSConv layer, a first A2C2f layer, a second DSConv layer, and a second A2C2f layer connected in sequence; the true-color image RGB and near-infrared band NIR are input into the first branch, and the digital elevation model (DEM) outputs into the second branch; the outputs of the first DS_C3k2 layers of the two branches are fused by a first Concat layer and output as B3; the outputs of the second DS_C3k2 layers of the two branches are fused by a second Concat layer and output as B4; the outputs of the third DS_C3k2 layers of the two branches are fused by a third Concat layer and output as B5; the outputs of the first A2C2f layers of the two branches are fused by a fourth Concat layer and output as B6; the outputs of the second A2C2f layers of the two branches are fused by a fifth Concat layer and output as B7. The DS_C3k2 layer comprises a sequentially connected Conv layer, a Split layer, and n DS_C3k layers. The outputs of the Split layer and the DS_C3k layer are fused by a Concat layer and processed by a Conv layer before being output. The DS_C3k layer comprises a sequentially connected Conv layer and n DS-Bottleneck layers. The original input data is processed by the Conv layer and fused with the output of the DS-Bottleneck layer by a Concat layer, and then further processed by the Conv layer before being output. The DS-Bottleneck layer comprises two sequentially connected DSConv layers. The original input data is added pixel-by-pixel to the output of the DSConv layer before being output. The DSConv layer comprises a sequentially connected DWConv layer, a PWConv layer, a BN layer, and a SiLU layer. DWConv (Depthwise Convolution) and PWConv (Pointwise Convolution) are existing and widely used standard modules. DWConv performs convolution on each channel of the input feature map separately and does not fuse across channels. PWConv uses 1×1 convolution kernels to linearly combine all channels, achieving channel fusion or dimensionality increase / decrease.
[0037] The A2C2f layer includes a sequentially connected Conv layer, two Split layers, and n Bottleneck layers. The outputs of each of the two Split layers and the n Bottleneck layers are fused by a Concat layer, further processed by the Conv layer, and then output. After processing by parallel Height Attention and Width Attention layers, the outputs are added pixel-by-pixel and then processed by the Conv layer before being output. The Bottleneck layer, through a design of "dimensionality reduction → feature extraction → dimensionality increase + residual connections," maintains expressive power while reducing computational cost. Height Attention captures long-distance dependencies in the vertical direction of the feature map. Width Attention captures long-distance dependencies in the horizontal direction of the feature map.
[0038] In this embodiment, DEM+NIR and RGB image data are input through dual streams. Both are initially extracted through convolutional layers (Conv) and the DS_C3k2 module, and then concatenated at the same level to generate multi-scale feature sequences B3 to B7.
[0039] In one implementation, please see Figure 4 The HyperACE improvement module first downsamples B3 and B4, upsamples B6 and B7, and then processes them with B5 through the Concat layer, Conv layer and Split layer. After processing by the first C3AH layer, the second C3AH layer and the DS-C3K layer set in parallel respectively, it is output through the Concat layer and the Conv layer. The C3AH layer includes a sequentially connected Conv layer, an Adaptive Hyperedge Generation layer, and a Hypergraph Convolution layer. The original input data is processed by the Conv layer and then merged with the output of the Hypergraph Convolution layer by the Concat layer and processed by the Conv layer before being output. The Adaptive Hyperedge Generation layer is used to generate adaptive hyperedges. The input is processed by the Flatten layer, the parallel MaxPooling layer and the AvgPooling layer, and then fused by the Concat layer. It is further processed by the Projection layer, the Dynamic Offset layer and the Global Proto layer. The outputs of the Dynamic Offset layer and the Global Proto layer are added pixel by pixel, and then multiplied pixel by pixel with the output of the Flatten layer after the Projection layer is processed. The Hypergraph Convolution layer is used to perform convolution operations on the hypergraph; it includes a sequentially connected hyperedge-to-node aggregation layer, a first Feature Projection layer, a node-to-hyperedge propagation layer, and a second Feature Projection layer. The original data is added pixel-by-pixel to the output of the second Feature Projection layer before outputting the final data. Specifically, the Projection layer maps the input features to an intermediate space; the Global Proto layer extracts global contextual information; the Dynamic Offset layer generates learnable offsets for aligning or deforming feature maps; and the Flatten layer flattens a multi-dimensional input tensor into a one-dimensional tensor.
[0040] In this embodiment, the HyperACE module receives B3, B4, B5, B6, and B7 as inputs. Using an intermediate level (such as B5) as the baseline resolution, it downsamples the high-resolution features (B3, B4) and upsamples the low-resolution features (B6, B7), unifying them to the same resolution as B5 before concatenating them. The concatenated features are then processed by a 1×1 convolution and divided into two paths: one path is stacked twice by the C3AH module to model high-order semantic information; the other path is processed by the DS-C3k module to extract low-order local details. The two paths are concatenated again and fused by a 1×1 convolution, with the final output feature resolution being consistent with the input B5. At this point, the feature is upsampled or downsampled to the original input resolution to obtain the enhanced feature representations H3, H4, H5, H6, and H7 for use in subsequent tasks.
[0041] In one implementation, please see Figure 5The multi-scale target localization module includes the following layers connected in sequence: a first Upsample layer, a sixth Concat layer, a fourth DS_C3k2 layer, a second Upsample layer, a seventh Concat layer, a fifth DS_C3k2 layer, a third Upsample layer, an eighth Concat layer, a sixth DS_C3k2 layer, a fourth Upsample layer, a ninth Concat layer, a seventh DS_C3k2 layer, a fourth Conv layer, a tenth Concat layer, an eighth DS_C3k2 layer, a fifth Conv layer, an eleventh Concat layer, a ninth DS_C3k2 layer, a sixth Conv layer, a twelfth Concat layer, a tenth DS_C3k2 layer, a seventh Conv layer, a thirteenth Concat layer, and an eleventh DS_C3k2 layer; B7 and H7 are added pixel by pixel and then input to the first Upsample layer. The following inputs are processed pixel-by-pixel: B6 and H6 are added together and then input to the sixth Concat layer; B5 and H5 are added pixel-by-pixel and then input to the seventh Concat layer; B4 and H4 are added pixel-by-pixel and then input to the eighth Concat layer; B3 and H3 are added pixel-by-pixel and then input to the ninth Concat layer; the output of the seventh DS_C3k2 layer is added pixel-by-pixel to H3 and then input to the fourth Conv layer; the output of the sixth DS_C3k2 layer is added pixel-by-pixel to H4 and then input to the tenth Concat layer; the output of the fifth DS_C3k2 layer is added pixel-by-pixel to H5 and then input to the eleventh Concat layer; the output of the fourth DS_C3k2 layer is added pixel-by-pixel to H6 and then input to the twelfth Concat layer; B7 and H7 are added pixel-by-pixel and then input to the thirteenth Concat layer; the output of the seventh DS_C3k2 layer is added pixel-by-pixel to H3 and then passed through the first detection head layer (detect). The output of the DS_C3k2 layer is the first type of preset scale erosion body after the _head. The output of the eighth DS_C3k2 layer is added to H4 pixel by pixel and then passed to the second detection head layer detect _head to output the second type of preset scale erosion body. The output of the ninth DS_C3k2 layer is added to H5 pixel by pixel and then passed to the third detection head layer detect _head to output the third type of preset scale erosion body. The output of the tenth DS_C3k2 layer is added to H6 pixel by pixel and then passed to the fourth detection head layer detect _head to output the fourth type of preset scale erosion body. The output of the eleventh DS_C3k2 layer is added to H7 pixel by pixel and then passed to the fifth detection head layer detect _head to output the fifth type of preset scale erosion body.
[0042] In this embodiment, B3 to B7 are used as input features. First, they are concatenated with the same-level enhanced features H3 to H7 output by the HyperACE module. Then, they are concatenated again with the upsampled high-resolution features of the next stage. The concatenated features are then processed by the DS_C3k2 module to extract local context information.
[0043] In the upsampling fusion path (left-hand network), this process proceeds layer by layer from bottom to top. For each layer i (i=7, 6, 5, 4), the features of the current layer... First, an upsampling operation is used to elevate the sample to the next higher level. It has the same spatial resolution, and is then concatenated with the corresponding hierarchical features from the backbone network to form a cross-scale semantic enhancement path.
[0044] The concatenated features are refined by the DS_C3k2 module and then fused with the output of the corresponding level of the HyperACE improvement module to achieve progressive feature enhancement from high-level semantics to low-level details, ultimately resulting in the enhanced upsampled feature sequences H3, H4, H5, H6, and H7.
[0045] In the downsampling fusion path (intermediate network), this process proceeds layer by layer from top to bottom. For each layer i (i=3, 4, 5, 6), the features of the current layer... First, downsample to the level of the next lower layer through a convolution operation (Conv). The same spatial resolution, then compared with the features output by the next level ( The elements are spliced together to construct a spatial detail propagation path from high to low.
[0046] The downsampled and concatenated features from each layer are further fused with local and global information using the DS_C3k2 module, and output at the same level as the HyperACE improvement module. Joint optimization is performed to generate a multi-scale feature map for object detection, which is then fed into the detection head to output bounding boxes and class predictions.
[0047] In one implementation, the five types of erosion bodies at preset scales are further processed through a large model segmentation module, a natural language understanding module, and a knowledge graph empowerment module to output accurate classification results. The large model segmentation module uses an image encoder to perform pixel-level segmentation of the five types of eroded landslide bodies at preset scales, and outputs candidate region masks, each of which corresponds to a possible target instance and includes category labels and location information. The natural language understanding module is used to preprocess multimodal input into text, and after word segmentation and normalization, form a word sequence; the Transformer dynamically embeds and generates vectors, which are then fused through context encoding, semantic parsing, and knowledge enhancement to finally output a machine-executable structured object; The knowledge graph empowerment module integrates natural language semantics with candidate masks, links categories to entities, retrieves spectral topography rules to verify the rationality of the mask, reorders to suppress false detections, supplements attributes and generates interpretable descriptions, and finally outputs a structured result containing the optimal mask, corrected category confidence, and knowledge basis.
[0048] In one implementation, please see Figure 6 The large model segmentation module specifically includes: (1) The input data is encoded by Image Encoder, multi-scale feature maps are extracted, and a series of high-level semantic representations are output as shared context information for subsequent decoders; (2) Input prompt information (box, text) into Prompt Encoder, where the bounding box is the normalized coordinates output by the multi-scale target localization module; the text prompt is the semantic vector output by the natural language understanding module; (3) The Prompt Encoder encodes the two prompts into a high-dimensional vector representation and aligns it spatially and semantically with the feature map output by the Image Encoder; (4) The Mask Decoder receives global features from the Image Encoder and prompt embeddings from the Prompt Encoder, and fuses the information from both through a cross-attention mechanism to achieve accurate localization of the target region; (5) Finally, output multiple candidate masks, each corresponding to a possible target instance and containing category label and location information.
[0049] In one implementation, the context encoding uses a Transformer Encoder to calculate the dependencies between words through a self-attention mechanism and outputs an enhanced context vector sequence; the semantic parsing includes two parallel subtasks: intent recognition and slot filling, where intent recognition is to obtain the classification result of the sentence, and slot filling is to obtain the label sequence of each word; the knowledge enhancement fusion is to fuse the current token with the knowledge vector generated by the knowledge graph empowerment module and update the semantic representation.
[0050] In one implementation, the knowledge graph empowerment module takes as input the structured semantics output by the natural language understanding module and multiple candidate masks output by the large model segmentation module, each mask being accompanied by a category label and probability distribution; firstly, the mask categories are linked to knowledge graph entity nodes, then relevant triples and rules are retrieved based on the user input type and mask category labels and positions obtained from the natural language understanding module, and finally, it is verified whether each mask satisfies prior constraints such as spectrum and topography; Based on the reasoning results, the knowledge graph empowerment module reorders the mask or suppresses unreasonable options, supplements the attribute information of the finally selected mask, and generates an interpretability description; finally, it outputs the enhanced structured result, including: the optimal mask ID, the corrected category and confidence level, and the knowledge explanation.
[0051] In one implementation, please see Figure 7 It also features an intelligent interaction module that can receive two types of input: text / voice input from the user and structured results output by the knowledge graph empowerment module. If the user's text / voice input matches the follow-up question, i.e., matches the type of existing processing results, it can be directly filtered based on the structured results to output the final segmentation result; otherwise, it returns to the initial input and reprocesses it.
[0052] This invention employs a large deep learning model, which can automatically learn the deep essential features of a collapse from end to end, fundamentally breaking through the upper limit of traditional feature engineering capabilities.
[0053] This invention introduces a "multimodal prompting" mechanism. By using text, voice, and geometric elements, it achieves dynamic, interactive, and guided intelligent interpretation, upgrading the "static model" into a "dialogue-enabled intelligent agent," thus solving the core pain points of flexibility and interactivity.
[0054] This invention introduces a "knowledge graph" to internalize the domain knowledge of landslides (morphology, developmental stages, causes, etc.) into structured information that the model can understand and reason about. This enables the system not only to "see" but also to "understand," possessing the potential for semantic understanding and logical reasoning, thus laying the foundation for realizing a generalized geoscientific interpretation platform.
[0055] This invention comprehensively utilizes topographic information represented by high-resolution remote sensing imagery and digital elevation models, combined with various interactive prompts such as text descriptions, voice commands, and geometric elements (e.g., points, boxes, lines), to achieve accurate, efficient, and interactive intelligent identification and boundary interpretation of landslides, driven by a domain knowledge graph. This system is particularly suitable for professional applications requiring high precision, high flexibility, and support for natural interaction, such as natural resource surveys and geological disaster monitoring.
[0056] It should be understood that the embodiments described above are only some, not all, of the embodiments of the present invention. Furthermore, the technical features of the various embodiments or individual embodiments provided by the present invention can be arbitrarily combined to form feasible technical solutions. Such combinations are not constrained by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.
[0057] It should be understood that the above description of the preferred embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art, under the guidance of this invention, can make substitutions or modifications without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.
Claims
1. A multimodal intelligent interpretation method for remote sensing monitoring of landslides, characterized in that, Includes the following steps: Step 1: Acquisition and preprocessing of multi-source remote sensing data; the multi-source remote sensing data includes true-color RGB imagery, near-infrared NIR band data, and digital elevation model (DEM); Step 2: Input the multi-source remote sensing data into the gully erosion monitoring network and output five types of gully erosion bodies at preset scales; The landslide remote sensing monitoring network includes an adaptive feature selection module, an improved HyperACE module, and a multi-scale target localization module; the improved HyperACE module has 5 scales for both input and output. The adaptive feature selection module is a dual-branch, multi-level skip fusion adaptive feature selection encoder; the true-color image RGB and near-infrared band NIR are input into the first branch, and the digital elevation model DEM outputs the second branch; outputs B3, B4, B5, B6, and B7 are processed by the HyperACE improvement module and output as H3, H4, H5, H6, and H7, respectively. The multi-scale target localization module, through bidirectional prior weighting, five-stage sequential upsampling, and a lightweight tower structure with independent detection heads at each stage, simultaneously completes high-precision localization and scale-adaptive classification of five-scale erosion bodies in a single forward pass.
2. The multimodal intelligent interpretation method for remote sensing monitoring of landslides according to claim 1, characterized in that: In step 1, the preprocessing first performs geometric fine correction on the original image to ensure geographic coordinate accuracy, and performs radiometric calibration and atmospheric correction to eliminate the effects of illumination and aerosols; then, using the true-color image RGB as a reference, the DEM and NIR are registered, and the registration error is controlled within a preset range; finally, image block processing is performed to maintain spatial alignment.
3. The multimodal intelligent interpretation method for remote sensing monitoring of landslides according to claim 1, characterized in that: The adaptive feature selection module includes two parallel branches; each branch consists of a first Conv layer, a first DS_C3k2 layer, a second Conv layer, a second DS_C3k2 layer, a third Conv layer, a third DS_C3k2 layer, a first DSConv layer, a first A2C2f layer, a second DSConv layer, and a second A2C2f layer connected in sequence; the true-color image RGB and near-infrared NIR bands are input into the first branch, and the digital elevation model (DEM) outputs into the second branch; the outputs of the first DS_C3k2 layers of the two branches are fused by a first Concat layer and output as B3; the outputs of the second DS_C3k2 layers of the two branches are fused by a second Concat layer and output as B4; the outputs of the third DS_C3k2 layers of the two branches are fused by a third Concat layer and output as B5; the outputs of the first A2C2f layers of the two branches are fused by a fourth Concat layer and output as B6; the outputs of the second A2C2f layers of the two branches are fused by a fifth Concat layer and output as B7. The DS_C3k2 layer comprises a sequentially connected Conv layer, a Split layer, and n DS_C3k layers. The outputs of the Split layer and the DS_C3k layer are fused by a Concat layer and processed by a Conv layer before being output. The DS_C3k layer comprises a sequentially connected Conv layer and n DS-Bottleneck layers. The original input data is processed by the Conv layer and then fused with the output of the DS-Bottleneck layer by a Concat layer, and further processed by the Conv layer before being output. The DS-Bottleneck layer comprises two sequentially connected DSConv layers. The original input data is added pixel-by-pixel to the output of the DSConv layer before being output. The DSConv layer comprises a sequentially connected DWConv layer, a PWConv layer, a BN layer, and a SiLU layer. The A2C2f layer includes a sequentially connected Conv layer, two Split layers, and n Bottleneck layers. The outputs of each of the two Split layers and the n Bottleneck layers are fused by a Concat layer and then further processed by a Conv layer before being output. After being processed by parallel Height Attention and Width Attention layers, the outputs are added pixel by pixel and then processed by a Conv layer before being output.
4. The multimodal intelligent interpretation method for remote sensing monitoring of landslides according to claim 1, characterized in that: The HyperACE improvement module first downsamples B3 and B4, upsamples B6 and B7, and then processes them with B5 through the Concat layer, Conv layer and Split layer. After processing by the first C3AH layer, the second C3AH layer and the DS-C3K layer set in parallel respectively, it is output through the Concat layer and the Conv layer. The C3AH layer includes a sequentially connected Conv layer, an Adaptive Hyperedge Generation layer, and a Hypergraph Convolution layer. The original input data is processed by the Conv layer and then merged with the output of the Hypergraph Convolution layer by the Concat layer and processed by the Conv layer before being output. The Adaptive Hyperedge Generation layer's input is processed by the Flatten layer, the parallel MaxPooling layer, and the AvgPooling layer, then fused by the Concat layer, and further processed by the Projection layer, the DynamicOffset layer, and the Global Proto layer. The outputs of the Dynamic Offset layer and the Global Proto layer are added pixel by pixel, and then multiplied pixel by pixel with the output of the Flatten layer after the Projection layer is processed. The Hypergraph Convolution layer includes a sequentially connected hyperedge-to-node aggregation layer, a first Feature Projection layer, a node-to-hyperedge propagation layer, and a second Feature Projection layer. The original data is added pixel by pixel to the output of the second Feature Projection layer before being output.
5. The multimodal intelligent interpretation method for remote sensing monitoring of landslides according to claim 1, characterized in that: The multi-scale target localization module includes the following layers connected in sequence: a first Upsample layer, a sixth Concat layer, a fourth DS_C3k2 layer, a second Upsample layer, a seventh Concat layer, a fifth DS_C3k2 layer, a third Upsample layer, an eighth Concat layer, a sixth DS_C3k2 layer, a fourth Upsample layer, a ninth Concat layer, a seventh DS_C3k2 layer, a fourth Conv layer, a tenth Concat layer, an eighth DS_C3k2 layer, a fifth Conv layer, an eleventh Concat layer, a ninth DS_C3k2 layer, a sixth Conv layer, a twelfth Concat layer, a tenth DS_C3k2 layer, a seventh Conv layer, a thirteenth Concat layer, and an eleventh DS_C3k2 layer; B7 and H7 are added pixel by pixel and then input to the first Upsample layer. The following inputs are processed pixel-by-pixel: B6 and H6 are added together and then input to the sixth Concat layer; B5 and H5 are added pixel-by-pixel and then input to the seventh Concat layer; B4 and H4 are added pixel-by-pixel and then input to the eighth Concat layer; B3 and H3 are added pixel-by-pixel and then input to the ninth Concat layer; the output of the seventh DS_C3k2 layer is added pixel-by-pixel to H3 and then input to the fourth Conv layer; the output of the sixth DS_C3k2 layer is added pixel-by-pixel to H4 and then input to the tenth Concat layer; the output of the fifth DS_C3k2 layer is added pixel-by-pixel to H5 and then input to the eleventh Concat layer; the output of the fourth DS_C3k2 layer is added pixel-by-pixel to H6 and then input to the twelfth Concat layer; B7 and H7 are added pixel-by-pixel and then input to the thirteenth Concat layer; the output of the seventh DS_C3k2 layer is added pixel-by-pixel to H3 and then passed through the first detection head layer (detect). The output of the DS_C3k2 layer is the first type of preset scale erosion body after the _head. The output of the eighth DS_C3k2 layer is added to H4 pixel by pixel and then passed to the second detection head layer detect _head to output the second type of preset scale erosion body. The output of the ninth DS_C3k2 layer is added to H5 pixel by pixel and then passed to the third detection head layer detect _head to output the third type of preset scale erosion body. The output of the tenth DS_C3k2 layer is added to H6 pixel by pixel and then passed to the fourth detection head layer detect _head to output the fourth type of preset scale erosion body. The output of the eleventh DS_C3k2 layer is added to H7 pixel by pixel and then passed to the fifth detection head layer detect _head to output the fifth type of preset scale erosion body.
6. The multimodal intelligent interpretation method for remote sensing monitoring of landslides according to any one of claims 1-5, characterized in that: Furthermore, through the large model segmentation module, natural language understanding module, and knowledge graph empowerment module, the five types of erosion bodies at preset scales are processed to output accurate classification results; The large model segmentation module uses an image encoder to perform pixel-level segmentation of the five types of eroded landslide bodies at preset scales, and outputs candidate region masks, each of which corresponds to a possible target instance and includes category labels and location information. The natural language understanding module is used to preprocess multimodal input into text, and then normalize it into a word sequence. Transformer dynamically embeds generated vectors, which are then fused through context encoding, semantic parsing, and knowledge enhancement to finally output a machine-executable structured object; The knowledge graph empowerment module integrates natural language semantics with candidate masks, links categories to entities, retrieves spectral topography rules to verify the rationality of the mask, reorders to suppress false detections, supplements attributes and generates interpretable descriptions, and finally outputs a structured result containing the optimal mask, corrected category confidence, and knowledge basis.
7. The multimodal intelligent interpretation method for remote sensing monitoring of landslides according to claim 6, characterized in that: The large model segmentation module specifically includes: (1) The input data is encoded by Image Encoder, multi-scale feature maps are extracted, and a series of high-level semantic representations are output as shared context information for subsequent decoders; (2) Input prompt information (box, text) into Prompt Encoder, where the bounding box is the normalized coordinates output by the multi-scale target localization module; the text prompt is the semantic vector output by the natural language understanding module; (3) The Prompt Encoder encodes the two prompts into a high-dimensional vector representation and aligns it spatially and semantically with the feature map output by the Image Encoder; (4) The Mask Decoder receives global features from the Image Encoder and prompt embeddings from the Prompt Encoder, and fuses the information from both through a cross-attention mechanism to achieve accurate localization of the target region; (5) Finally, output multiple candidate masks, each corresponding to a possible target instance and containing category label and location information.
8. The multimodal intelligent interpretation method for remote sensing monitoring of landslides according to claim 6, characterized in that: The context encoding uses Transformer Encoder to calculate the dependencies between words through a self-attention mechanism and outputs an enhanced sequence of context vectors. The semantic parsing includes two parallel subtasks: intent recognition and slot filling. Intent recognition is used to obtain the classification results of the sentence, while slot filling is used to obtain the label sequence of each word. The knowledge enhancement fusion involves fusing the current token with the knowledge vector generated by the knowledge graph empowerment module and updating the semantic representation.
9. The multimodal intelligent interpretation method for remote sensing monitoring of landslides according to claim 6, characterized in that: The knowledge graph empowerment module takes as input the structured semantics output by the natural language understanding module and multiple candidate masks output by the large model segmentation module. Each mask is accompanied by a category label and probability distribution. First, the mask categories are linked to the knowledge graph entity nodes. Then, based on the user input type and mask category labels and positions obtained from the natural language understanding module, relevant triples and rules are retrieved. Finally, it verifies whether each mask satisfies the prior constraints. Based on the reasoning results, the knowledge graph empowerment module reorders the mask or suppresses unreasonable options, supplements the attribute information of the finally selected mask, and generates an interpretability description; finally, it outputs the enhanced structured result, including: the optimal mask ID, the corrected category and confidence level, and the knowledge explanation.
10. A multimodal intelligent interpretation system for remote sensing monitoring of landslides, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the multimodal intelligent interpretation method for remote sensing monitoring of landslides as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Method for extracting slope collapse information based on Sentinel-2A satellite remote sensing image
CN110334623A
Vegetation coverage-based collapse activity degree discrimination method
CN110503018A
Method and device for identifying and predicting internal channel of collapse hill
CN113158588A
Improved U-net model-based collapse hill extraction method and system
CN114913424A
Automatic slope collapse identification method and system
CN114972991A