Multi-mode remote sensing intelligent identification method for hidden geological disasters
By introducing pixel-level cross-modal vision converters and related mechanisms in multimodal remote sensing data fusion, the problems of limited effect and high computational complexity in the prior art are solved, and efficient and accurate identification of hidden geological disasters is achieved, and it is suitable for large-scale geological disaster monitoring.
Patent Information
- Application Number
- CN202510600195.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing multimodal data fusion method is difficult to fully explore the complementarity between different mode data, resulting in limited fusion effect, high computational complexity, and difficult to apply to large-scale geological disaster monitoring. It lacks special designs for hidden geological disaster characteristics, and it is impossible to effectively capture the complex relationship between slow deformation and surface features.
A multimodal remote sensing intelligent recognition method for hidden geological disasters is proposed. Through an efficient pixel-level cross-modal fusion mechanism, using the complementary information of RGB optical images, InSAR deformation rate maps and DEM elevation data, a pixel-level cross-modal vision converter is constructed, including a multimodal encoder and a decoder, and multi-scale features are extracted using a multi-stage structure, and precise identification of hidden geological disasters is achieved through pixel-level cross-modal attention module, layer adaptive noise mechanism and relationship discriminator.
It has achieved efficient and accurate identification of hidden geological disasters, significantly improved the identification accuracy, especially in the deformation monitoring capabilities in vegetation-covered areas, and has strong generalization capabilities. It is suitable for multi-source remote sensing data of different terrain conditions and different seasons to meet the actual monitoring application needs.
Smart Images

Figure CN120107808A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of geological disaster monitoring, and in particular to a multi-modal remote sensing intelligent identification method for concealed geological disasters. Background Art
[0002] Hidden geological disasters such as slow landslides and ground subsidence are often difficult to detect in time because of their slow occurrence and unclear manifestations. They may suddenly turn into catastrophic events after a long period of development, posing a serious threat to people's lives and property. Traditional geological disaster investigation methods mainly rely on on-site surveys and expert experience and judgment, which have problems such as limited coverage, low efficiency, and low accuracy.
[0003] With the development of remote sensing technology, multimodal remote sensing data has been widely used in the field of geological disaster monitoring. Synthetic aperture radar differential interferometry (InSAR) can provide millimeter-level deformation information of the surface; optical remote sensing images can reflect the surface coverage characteristics; digital elevation models (DEM) can describe the terrain undulation characteristics. These complementary information are of great value for identifying hidden geological disasters. However, the existing multimodal data fusion methods mainly have the following problems:
[0004] 1. Existing methods use simple splicing or serial processing to fuse information, which makes it difficult to fully explore the complementarity between different modal data, resulting in limited fusion effect;
[0005] 2. The cross-modal fusion method based on cross-attention has high computational complexity when processing high-resolution remote sensing images and is difficult to apply to large-scale geological disaster monitoring;
[0006] 3. The fusion method based on feature replacement may cause key information loss during the replacement process, affecting the final recognition accuracy;
[0007] 4. There is a lack of special design targeting the characteristics of hidden geological hazards, and it is impossible to effectively capture the complex relationship between slow deformation and surface characteristics.
[0008] Therefore, it is urgent to develop an efficient and accurate multimodal fusion method to make full use of the complementary information of different remote sensing data and realize the identification of hidden geological hazards. Summary of the invention
[0009] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, one purpose of the present invention is to propose a multimodal remote sensing intelligent identification method for hidden geological disasters, which makes full use of the complementary information of RGB optical images, InSAR deformation rate images and DEM elevation data, and realizes accurate identification of hidden geological disasters through an efficient pixel-level cross-modal fusion mechanism.
[0010] In order to solve the above problems, the present invention provides a multi-modal remote sensing intelligent identification method for hidden geological hazards, comprising the following steps:
[0011] Normalize multimodal remote sensing data including optical RGB images, InSAR deformation maps and DEM terrain data, and map different modal data into a unified feature space;
[0012] Construct a pixel-level cross-modal visual converter, including a multimodal encoder and decoder; the encoder uses a multi-stage structure to extract multi-scale features, and integrates a pixel-level cross-modal fusion module in each stage; the decoder integrates multimodal encoding features based on a multi-layer perceptron to generate segmentation predictions;
[0013] Construct a pixel-level cross-modal attention module, including a pixel self-attention submodule, a pixel mutual attention submodule, and a layer adaptation hybrid module, and implement an efficient cross-modal attention mechanism with linear complexity through pixel alignment;
[0014] Construct a layer-adaptive noise mechanism to adaptively adjust the weight ratio of intra-modality self-attention and inter-modality mutual attention in different layers to balance the degree of modality fusion at different depth layers;
[0015] Construct a relation discriminator to modulate the attention calculation process by evaluating the difference between the modalities of spatial corresponding points and enhance the ability to extract complementary information;
[0016] The training model uses a multi-task loss function including semantic segmentation loss and boundary perception loss to jointly optimize and generate hidden geological disaster segmentation results.
[0017] Preferably, the step of normalizing the multimodal remote sensing data comprises:
[0018] For RGB optical images, standard normalization is applied: the pixel value of the RGB image is subtracted from the mean and then divided by the standard deviation;
[0019] For the InSAR deformation rate map, robust normalization is used: the InSAR data are mapped to the range of [-1,1], and 95% quantile clipping is applied to avoid the influence of outliers;
[0020] For DEM elevation data, logarithmic transformation is applied and then normalized: first calculate the logarithmic transformation log(1+DEM), and then apply standard normalization.
[0021] Preferably, the encoder of the pixel-level cross-modal visual converter includes four stages, each of which includes:
[0022] Overlapping image partitioning module, which divides the input image into serialized blocks through convolution operation;
[0023] Multiple pixel-level cross-modal Transformer blocks, each containing a pixel-level cross-modal attention module and a feed-forward neural network;
[0024] Downsampling module, which halves the spatial resolution and increases the channel dimension through convolution with a stride of 2;
[0025] The configurations of the four stages include: the first stage has a block size of 7×7, a step size of 4, and an output channel of 64; the second stage has a block size of 3×3, a step size of 2, and an output channel of 128; the third stage has a block size of 3×3, a step size of 2, and an output channel of 320; the fourth stage has a block size of 3×3, a step size of 2, and an output channel of 512.
[0026] Preferably, the calculation formula of the pixel self-attention submodule is:
[0027] ;
[0028] in, , , are query, key, and value matrices, , , is the learnable projection weight, d is the dimension of the attention head, and T is the transposed sign; the pixel mutual attention submodule only establishes connections between pixels that are aligned in spatial position, and its calculation formula is:
[0029] ;
[0030] in, , and represent the query of modality i at position n and the key and value of modality j respectively, is the relation discriminator.
[0031] Preferably, the layer adaptive noise mechanism includes:
[0032] Create unique noisy embeddings for each modality and each network layer: ,in, is the noise embedding vector, EmBed is the embedding function, and according to the modal index and layer index Generate a specific noise vector;
[0033] The noise embedding is added to the key matrix in the self-attention process: ,in, Indexed by modality and layer index The resulting bond matrix, is the key matrix updated after the noise embedding is added;
[0034] Noise injection guides self-attention computation: .
[0035] Preferably, the structure of the relationship discriminator is designed as follows:
[0036] ;
[0037] Among them, [X i ;X j ] represents the feature connection operation, MLP is a multi-layer perceptron, is the Softmax function; the relationship score is used to modulate the key value calculation in the mutual attention: ,in, represents element-wise multiplication, is the key value obtained by modal index j, To calculate the updated key value.
[0038] Preferably, the multi-task loss function includes:
[0039] Focal Loss: , where p t is the probability of predicting category t, α t is the class balance parameter, γ is the focusing parameter;
[0040] Boundary-aware loss: , where CE is the cross entropy loss, w n is the boundary weight: , d n is the distance from the pixel to the nearest segmentation boundary, θ is the parameter that controls the weight decay speed, represents the predicted value, represents the truth value in this case;
[0041] Combined loss: , where λ 1 and λ 2 is the loss weight, set to,λ 1 = 0.7, λ 2 =0.3.
[0042] Preferably, the step of training the model includes:
[0043] Initialize the encoder using a pre-trained visual transformer to speed up training and improve stability;
[0044] The cosine annealing learning rate scheduling is used, and the initial learning rate is set to 2×10 -5 , the minimum learning rate is 1×10 -7 ;
[0045] The batch size is set to 32 and the total number of training epochs is 150.
[0046] Preferably, the method further comprises the steps of reasoning and practical application:
[0047] Acquire multimodal remote sensing data of the target area, including RGB images, InSAR deformation maps and DEM data;
[0048] Pre-process the data, including registration, cropping and normalization;
[0049] Input the processed multimodal data into the trained model for inference;
[0050] Process the segmentation results output by the model to generate a map of hidden geological hazard risk areas;
[0051] Combined with on-site investigation, verify and refine the geological hazard risk assessment results.
[0052] The advantages of the present invention compared with the prior art are:
[0053] 1. This paper proposes a new pixel-level cross-modal attention mechanism. Compared with the global cross-attention method, the computational complexity is reduced from the quadratic level to the linear level, which greatly reduces the computational resource requirements while maintaining the fusion effect;
[0054] 2. The present invention designs a layer-adaptive noise mechanism to automatically adjust the ratio of intra-modality self-attention and inter-modality mutual attention in different network depth layers, optimize feature representation, and improve fusion efficiency;
[0055] 3. The present invention designs a special relationship discriminator to enhance the acquisition of complementary information and avoid information redundancy and homogeneity;
[0056] 4. Based on the complementarity of multimodal data, the present invention significantly improves the recognition accuracy of hidden geological disasters, especially the deformation monitoring capability of vegetation-covered areas;
[0057] 5. The method of the present invention has strong generalization ability and can be applied to multi-source remote sensing data in different terrain conditions and different seasons to meet the actual monitoring application needs.
[0058] The present invention can realize accurate identification of hidden geological disasters by efficiently fusing multimodal remote sensing data, and can be widely used in the fields of early warning of geological disasters, risk assessment, disaster prevention and mitigation, etc., and has important theoretical significance and practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0060] Figure 1 The schematic diagram of the overall framework of the method of the present invention (the complete processing flow from multimodal data input to hidden geological disaster identification);
[0061] Figure 2 The network architecture diagram of the pixel-level cross-modal visual converter (structural design of the model encoder and decoder);
[0062] Figure 3 Schematic diagram of the structure of the pixel-level cross-modal attention module (the connection relationship between the pixel self-attention submodule, the pixel mutual attention submodule and the layer adaptation hybrid module). DETAILED DESCRIPTION
[0063] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.
[0064] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0065] The present invention is further described in detail below in conjunction with the accompanying drawings.
[0066] Reference Figure 1 , Figure 1 Schematic diagram of the overall framework of an embodiment of the present invention.
[0067] like Figure 1As shown, the hidden geological disaster intelligent identification method based on pixel-level cross-modal visual converter proposed in the present invention includes five main steps: multimodal remote sensing data preprocessing, construction of pixel-level cross-modal visual converter, construction of pixel-level cross-modal attention module, construction of layer adaptive noise mechanism, construction of relationship discriminator and model training and application.
[0068] The pixel-level cross-modal visual converter (PixCrossTrans for short) of the present invention adopts an encoder-decoder architecture. Figure 2 As shown in the figure, the encoder has a four-stage hierarchical structure for extracting multi-scale features; the decoder adopts a multi-layer perceptron-based design to fuse the multi-scale features of the encoder to generate the final hidden geological disaster segmentation prediction.
[0069] In terms of multimodal input, the present invention mainly processes three types of remote sensing data:
[0070] 1. RGB optical image: provides surface cover type and visible spectrum characteristics;
[0071] 2. InSAR deformation rate map: provides centimeter to millimeter level surface displacement information;
[0072] 3.DEM elevation data: provides terrain features such as terrain undulation and slope.
[0073] These three modal data have their own characteristics and complement each other: RGB images can intuitively display the surface coverage but are easily affected by weather and light; InSAR deformation data can accurately reflect the slight deformation of the surface but are easily affected by temporal decoherence; DEM data reflects the terrain characteristics but has low timeliness. The PixCrossTrans model of the present invention makes full use of the complementarity of the three modal data and extracts comprehensive features through an efficient pixel-level cross-modal fusion mechanism.
[0074] In this embodiment, the method for intelligently identifying hidden geological hazards using a pixel-level cross-modal visual converter includes the following steps:
[0075] Step S1: Multimodal remote sensing data preprocessing.
[0076] Before model processing, different modal data need to be standardized to facilitate network learning. For each modality, a modality-specific normalization strategy is used:
[0077] For RGB optical images, standard normalization is applied:
[0078] ;
[0079] Among them, μ RGB and σ RGB are the mean and standard deviation of the image respectively.
[0080] For the InSAR deformation rate map, due to the presence of outliers, robust normalization is used:
[0081] ;
[0082] Among them, min(X InSAR ) and max(X InSAR ) Use robust statistical methods such as 95% quantile to avoid the influence of outliers.
[0083] For DEM elevation data, considering the distribution characteristics of elevation data, logarithmic transformation is used and then normalized:
[0084] ;
[0085] Among them, μ log(DEM) and σ log(DEM) are the mean and standard deviation of the DEM data after logarithmic transformation.
[0086] The above normalization process ensures that different modal data are mapped to the appropriate numerical range, which is beneficial to network training and feature fusion between modalities.
[0087] Step S2: Build a pixel-level cross-modal visual transformer.
[0088] Encoder structure design:
[0089] The encoder of PixCrossTrans adopts a hierarchical architecture consisting of four stages, each of which contains overlapping image blocks, multiple pixel-level cross-modal Transformer blocks, and a downsampling module.
[0090] The overlapping image partitioning module divides the input image into serialized blocks.
[0091] ;
[0092] in, Indicates The modality is in the Transformer encoder The feature representation of the layer output, is the convolution projection operation, is the number of modes. For the first stage, is the i-th modal input data after the normalization processing in step S1. The normalization processing is to normalize the original modal data according to the normalization formula in step S1 so that the numerical ranges of different modal data are consistent.
[0093] The pixel-level cross-modal Transformer block processes the input features at each stage:
[0094] ;
[0095] in, Indicates The mode is in The output of the kth Transformer block in layer , is a pixel-level cross-modal attention module, is layer normalization, It is a feed-forward neural network.
[0096] The resolution of the feature map generated at each stage is 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the input, respectively. It should be understood that the above configuration can be adjusted according to specific application scenarios.
[0097] Step S3: Pixel Cross-Modal Attention Module (PCM).
[0098] The pixel-level cross-modal attention module overcomes the quadratic complexity problem of traditional cross-attention and provides an efficient cross-modal fusion mechanism with linear complexity.
[0099] Pixel self-attention submodule construction:
[0100] For each modality, self-attention is calculated to capture contextual information:
[0101] ;
[0102] in, , , are query, key, and value matrices, , , is the learnable projection weight and d is the dimension of the attention head.
[0103] Pixel Cross-Attention (PCA) submodule construction:
[0104] Unlike traditional global cross attention, pixel mutual attention only establishes connections between pixels that are aligned in spatial positions:
[0105] ;
[0106] in, , and represent the query of modality i at position n and the key and value of modality j respectively, is the relation discriminator.
[0107] Layer-Adaptive Mixing Module (LAM) construction
[0108] Combine the outputs of self-attention and mutual attention, and introduce a layer-adaptive noise mechanism:
[0109] ;
[0110] Among them, N i l is the adaptive noise of the lth layer, α l i,j is a learnable mixing weight that controls the contribution of mutual attention of different modalities.
[0111] Through the above design, the computational complexity of the PCM module is reduced from O(N) to 2 ) is reduced to O(N), where N is the number of input tokens. Taking typical remote sensing image processing as an example (512×512 pixels), the traditional cross attention requires about 17G FLOPs, while the PCM module of the present invention only requires 0.14G FLOPs, and the computing efficiency is improved by about 99%.
[0112] Step S4: Layer adaptive noise mechanism construction
[0113] In order to solve the problem that different network depth layers have different requirements for modal fusion, this paper constructs a layer-adaptive noise mechanism (LAN)
[0114] Create unique noisy embeddings for each modality and each network layer:
[0115] ;
[0116] Among them, EmBed is the embedding function, according to the modal index i and the layer index Generates a specific noise vector.
[0117] The noise embedding is added to the key matrix in the self-attention process:
[0118] ;
[0119] Noise injection guides self-attention computation:
[0120] ;
[0121] Layer-adaptive noise has a regulatory effect on the soft attention distribution: at shallow layers, noise guides the model to pay more attention to intra-modal information; at deep layers, noise promotes the model to pay more attention to cross-modal information. This adaptive mechanism can automatically balance the intra-modal feature extraction and inter-modal information fusion of different depth layers.
[0122] Step S5: Construct a relation discriminator
[0123] In order to enhance the extraction of complementary information between modalities, the present invention designs a relation discriminator (RD) to evaluate the differences between modalities at corresponding spatial positions.
[0124] Relation discriminator structure design:
[0125] ;
[0126] Among them, [X i ;X j ] represents the feature connection operation, MLP is a multi-layer perceptron, is the Softmax function.
[0127] The relation score is used to modulate the key-value calculation in mutual attention:
[0128] ;
[0129] in, Represents element-wise multiplication
[0130] Through the relationship discriminator, the model can dynamically adjust the attention allocation according to the degree of complementarity of different modal data at a specific location, avoid redundant information exchange, and enhance meaningful cross-modal interaction. For example, in areas with significant terrain changes but consistent optical image coverage, the relationship discriminator will guide the model to pay more attention to InSAR deformation data; while in areas with small deformation but obvious optical features, more RGB image information will be used.
[0131] Step S6: Model training and application
[0132] Loss function design:
[0133] In order to improve the segmentation accuracy of hidden geological hazards, especially the accuracy of the boundary area, the present invention adopts a combined loss function:
[0134] Focal Loss:
[0135] ;
[0136] Among them, p t is the probability of predicting category t, α t is the category balance parameter and γ is the focus parameter, which is used to solve the foreground-background category imbalance problem.
[0137] Boundary-Aware Loss:
[0138] ;
[0139] Among them, CE is the cross entropy loss, w n is the boundary weight:
[0140] ;
[0141] d n is the distance from the pixel to the nearest segmentation boundary, and θ is a parameter that controls the speed of weight decay.
[0142] Combined loss:
[0143] ;
[0144] Among them, λ 1 and λ 2 is the loss weight, which is set to λ in this embodiment. 1 = 0.7, λ 2 = 0.3. The model training adopts the following strategies to optimize performance:
[0145] Initializing the encoder with a pre-trained visual transformer (such as the Transformer pre-trained on ImageNet) speeds up training and improves stability.
[0146] Cosine annealing learning rate scheduling is adopted, the initial learning rate is set to 2×10^-5, and the minimum learning rate is 1×10^-7.
[0147] The AdamW optimizer with weight decay set to 0.05 and betas parameter in [0.9, 0.999] was used.
[0148] A learning rate warmup of 1×10^-6 is used in the first 3000 iterations.
[0149] The batch size is set to 32 and the total number of training epochs is 150.
[0150] After the model training is completed, it can be applied to actual hidden geological disaster monitoring:
[0151] 1. Obtain multimodal remote sensing data of the target area, including RGB images, InSAR deformation maps and DEM data.
[0152] 2. Pre-process the data, including registration, cropping and normalization.
[0153] 3. Input the processed multimodal data into the trained PixCrossTrans model for inference.
[0154] 4. Process the segmentation results output by the model to generate a map of hidden geological disaster risk areas.
[0155] 5. Combine field investigations to verify and refine the results of geological hazard risk assessments.
[0156] Variants and extensions of the invention
[0157] The PixCrossTrans method of the present invention can have many variations and extensions:
[0158] S01: The number of modalities can be flexibly adjusted. In addition to the three modalities of RGB, InSAR and DEM, other remote sensing data sources such as thermal infrared and SAR intensity maps can also be integrated.
[0159] S02: The encoder can use different backbone networks, such as Swin Transformer or other visual converters, and select the most suitable architecture according to the specific application scenario.
[0160] S03: The design of the layer adaptive noise mechanism can be further optimized, such as using a dynamic parameter generation network to adaptively adjust the noise parameters according to the characteristics of the input data.
[0161] In summary, the intelligent identification method of hidden geological hazards based on pixel-level cross-modal visual converter proposed in this invention realizes the efficient fusion of multimodal remote sensing data through an efficient pixel-level attention mechanism, improves the accuracy and robustness of hidden geological hazards identification while maintaining linear computational complexity, and provides strong support for early warning and risk management of geological hazards.
[0162] The present invention and its embodiments are described above, and such description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if ordinary technicians in the field are inspired by it, without departing from the purpose of the invention, they can design a structure and embodiment similar to the technical solution without creativity, which should belong to the protection scope of the present invention.
Claims
1. A multi-modal remote sensing intelligent identification method for hidden geological hazards, characterized in that: The following steps are involved: Normalize multimodal remote sensing data including optical RGB images, InSAR deformation maps and DEM terrain data, and map different modal data into a unified feature space; Construct a pixel-level cross-modal visual converter, including a multimodal encoder and decoder; the encoder uses a multi-stage structure to extract multi-scale features, and integrates a pixel-level cross-modal fusion module in each stage; The decoder integrates multimodal encoding features based on a multi-layer perceptron to generate segmentation predictions; Construct a pixel-level cross-modal attention module, including a pixel self-attention submodule, a pixel mutual attention submodule, and a layer adaptation hybrid module, and implement an efficient cross-modal attention mechanism with linear complexity through pixel alignment; Construct a layer-adaptive noise mechanism to adaptively adjust the weight ratio of intra-modality self-attention and inter-modality mutual attention in different layers to balance the degree of modality fusion at different depth layers; Construct a relation discriminator to modulate the attention calculation process by evaluating the difference between the modalities of spatial corresponding points and enhance the ability to extract complementary information; The training model uses a multi-task loss function including semantic segmentation loss and boundary perception loss to jointly optimize and generate hidden geological disaster segmentation results.
2. According to claim 1, a multi-modal remote sensing intelligent identification method for hidden geological hazards is characterized by: The step of normalizing the multimodal remote sensing data comprises: For RGB optical images, standard normalization is applied: the pixel value of the RGB image is subtracted from the mean and then divided by the standard deviation; For the InSAR deformation rate map, robust normalization is used: the InSAR data are mapped to the range of [-1,1], and 95% quantile clipping is applied to avoid the influence of outliers; For DEM elevation data, logarithmic transformation is applied and then normalized: first calculate the logarithmic transformation log(1+DEM), and then apply standard normalization.
3. The multi-modal remote sensing intelligent identification method for hidden geological hazards according to claim 1 is characterized by: The encoder of the pixel-level cross-modal visual converter includes four stages, each of which includes: Overlapping image partitioning module, which divides the input image into serialized blocks through convolution operation; Multiple pixel-level cross-modal Transformer blocks, each containing a pixel-level cross-modal attention module and a feed-forward neural network; Downsampling module, which halves the spatial resolution and increases the channel dimension through convolution with a stride of 2; The configurations of the four stages include: the first stage has a block size of 7×7, a step size of 4, and an output channel of 64; the second stage has a block size of 3×3, a step size of 2, and an output channel of 128; the third stage has a block size of 3×3, a step size of 2, and an output channel of 320; the fourth stage has a block size of 3×3, a step size of 2, and an output channel of 512.
4. The multi-modal remote sensing intelligent identification method for hidden geological hazards according to claim 1 is characterized in that: The calculation formula of the pixel self-attention submodule is: ; in, , , are query, key, and value matrices, , , is the learnable projection weight, d is the dimension of the attention head, and T is the transposed sign; the pixel mutual attention submodule only establishes connections between pixels that are aligned in spatial position, and its calculation formula is: ; in, , and represent the query of modality i at position n and the key and value of modality j respectively, is the relation discriminator.
5. The multi-modal remote sensing intelligent identification method for hidden geological hazards according to claim 1 is characterized by: The layer adaptive noise mechanism includes: Create unique noisy embeddings for each modality and each network layer: ,in, is the noise embedding vector, EmBed is the embedding function, and according to the modal index and layer index Generate a specific noise vector; The noise embedding is added to the key matrix in the self-attention process: ,in, Indexed by modality and layer index The resulting bond matrix, is the key matrix updated after the noise embedding is added; Noise injection guides self-attention computation: .
6. The multi-modal remote sensing intelligent identification method for hidden geological hazards according to claim 1 is characterized by: The structure design of the relationship discriminator is as follows: ; Among them, [X i ;X j ] represents the feature connection operation, MLP is a multi-layer perceptron, is the Softmax function; the relationship score is used to modulate the key value calculation in the mutual attention: ,in, represents element-wise multiplication, is the key value obtained by modal index j, To calculate the updated key value.
7. The multi-modal remote sensing intelligent identification method for hidden geological hazards according to claim 1 is characterized by: The multi-task loss function includes: Focal Loss: , where p t is the probability of predicting category t, α t is the class balance parameter, γ is the focusing parameter; Boundary-aware loss: , where CE is the cross entropy loss, w n is the boundary weight: , d n is the distance from the pixel to the nearest segmentation boundary, θ is the parameter that controls the weight decay speed, represents the predicted value, represents the truth value in this case; Combined loss: , where λ1 and λ2 are loss weights, set to λ1=0.7, λ2=0.
3.
8. The multi-modal remote sensing intelligent identification method for hidden geological hazards according to claim 1 is characterized by: The steps of training the model include: Initialize the encoder using a pre-trained visual transformer to speed up training and improve stability; The cosine annealing learning rate scheduling is used, and the initial learning rate is set to 2×10 -5 , the minimum learning rate is 1×10 -7 ; The batch size is set to 32 and the total number of training epochs is 150.
9. The multi-modal remote sensing intelligent identification method for hidden geological hazards according to claim 1 is characterized by: The method also includes the steps of reasoning and practical application: Acquire multimodal remote sensing data of the target area, including RGB images, InSAR deformation maps and DEM data; Pre-process the data, including registration, cropping and normalization; Input the processed multimodal data into the trained model for inference; Process the segmentation results output by the model to generate a map of hidden geological hazard risk areas; Combined with on-site investigation, verify and refine the geological hazard risk assessment results.
Citation Information
Patent Citations
Landslide deformation accumulation area prediction model generation method and landslide deformation accumulation area prediction method
CN111929683A
Geological disaster risk partition automatic optimization method and system based on artificial intelligence
CN118822271A
Point cloud segmentation method and system, medium, equipment and information data processing terminal
CN118967698A
SAR image ship target detection method based on multi-attention fusion
CN119360196A
Transform remote sensing semantic segmentation method based on adaptive dynamic attention mechanism
CN119516197A
Cited By
Land property detection method based on space-time adaptive fusion and geographic features
CN120635657A
Land property detection method based on spatio-temporal adaptive fusion and geographical features
CN120635657B
Landslide edge identification method and system fusing CNN and ViT
CN120976784A
A landslide edge recognition method and system fusing CNN and ViT
CN120976784B
Liquid crystal display screen defect automatic classification method based on deep learning
CN121147202A