Intelligent Identification Method for Creep-Type Landslide Hazards Combining Image Processing and Semantic Understanding

By combining image processing and semantic understanding, the system achieves automated detection and feedback of landslide hazards, solving the problem of dependence on high-quality data and expert knowledge, improving the accuracy and applicability of landslide identification, and enhancing the monitoring and emergency response capabilities for landslide risks.

CN119851278BActive Publication Date: 2026-03-06CHANGAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411902586.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2026-03-06
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing landslide hazard identification technologies rely heavily on high-quality data and expert knowledge, resulting in inaccurate identification results and difficulty in applying them to complex terrain and low-coherence areas. Furthermore, they lack self-interpretation and self-analysis mechanisms.

Method used

This study employs a combined image processing and semantic understanding approach, utilizing geometric modeling, phase filtering, deep learning, and multimodal data processing to achieve automated detection and feedback of landslide hazards. Specific steps include: interferogram phase suppression, deformation signal enhancement, semantic segmentation, image caption generation, and a visual question-answering mechanism, reducing reliance on data quality and expert knowledge.

Benefits of technology

It improves the accuracy and applicability of landslide hazard identification, enhances the monitoring and emergency response capabilities for landslide risks, and provides technical support for the prevention, monitoring, and early warning of landslide disasters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851278B_ABST
    Figure CN119851278B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent identification method for creep-type landslide hazards that combines image processing and semantic understanding. First, high-precision geometric modeling and filtering techniques are used to suppress irrelevant phases in InSAR interferograms while preserving deformation information related to landslides. Next, signal features are enhanced through phase gradient, RGB channel mapping, and generative adversarial networks to improve signal representation capabilities. Subsequently, convolutional operations, self-attention mechanisms, and structural reparameterization are used to optimize the deep learning model, improving its learning ability for landslide hazards. In the feature encoding stage, a multimodal caption generation model is used to obtain visual and linguistic features, which are then mapped to the same embedding space for seamless integration. Finally, an autoregressive language generation model and a multilayer perceptron are used to predict landslide hazards. This approach, by integrating visual, linguistic, and deep learning technologies, significantly improves the efficiency and accuracy of landslide hazard monitoring, providing effective support for disaster response and management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of synthetic aperture radar technology, specifically involving an intelligent identification method for creep-type landslide hazards that combines image processing and semantic understanding. Background Technology

[0002] Landslides are a common and highly destructive geological hazard, referring to the downward and outward sliding movement of rock and soil masses along a specific sliding surface under the influence of gravity. Landslide hazards mainly fall into two categories: one is areas that have not yet experienced a landslide but are undergoing continuous deformation under external forces, and may even become unstable and slide (i.e., potential landslide areas); the other is areas where landslides have occurred in the past, where the landslide body has stabilized, but due to changes in external conditions, there is a potential risk of re-deformation or recurrence. Landslides not only cause significant changes in the landform but also severely impact human society and economic activities. Therefore, accurately identifying and assessing landslide hazards is of great importance. Traditional risk identification methods mainly rely on field surveys and on-site data collection; however, these methods have many limitations. First, the spatial coverage of field surveys is usually limited, making it difficult to comprehensively monitor large-scale landslides. Second, the real-time nature of on-site data collection is insufficient, causing data to lag behind geological changes, thus weakening its effectiveness in disaster prevention and mitigation. Furthermore, field surveys often require significant investment of manpower and resources, resulting in high costs.

[0003] The rapid development of remote sensing technology is profoundly transforming the way landslide hazard identification is conducted. Remote sensing technology enables landslide monitoring to be carried out on a large scale and in multiple dimensions, significantly improving the efficiency and accuracy of data collection. The non-contact nature of remote sensing imagery not only avoids the manpower and time costs of traditional field surveys but also covers a wide area, allowing for geological hazard monitoring in remote or inaccessible regions. However, the effectiveness of optical remote sensing imagery is affected by various factors, such as sunlight conditions, weather changes, cloud cover, and vegetation cover. In contrast, synthetic aperture radar (SAR) data exhibits unique advantages. First, SAR possesses all-weather, all-time imaging capabilities, penetrating clouds and precipitation, thus providing high-quality imagery data even in emergency scenarios such as severe weather. Furthermore, SAR employs active imaging technology, and its strong penetrating power allows it to effectively acquire more accurate information about the Earth's surface.

[0004] Interferometric Synthetic Aperture Radar (InSAR) is a remote sensing technology based on SAR data, which uses phase difference calculations to accurately measure surface deformation and elevation. InSAR technology can capture minute displacement changes on the surface, and its high sensitivity to these changes makes it valuable for landslide deformation monitoring. Differential InSAR (DInSAR) is affected by the quality of the elevation model and suffers from spatiotemporal discorrelation problems; therefore, time-series InSAR technology is more commonly used in research. Currently, various time-series InSAR methods are widely used in landslide activity monitoring. However, due to the diversity of landslide disaster characteristics and complex geomorphological environments, the interpretation of InSAR deformation products still requires visual judgment through human-computer interaction, making the accuracy of the results largely dependent on expert knowledge.

[0005] Deep learning, with its powerful feature extraction and pattern recognition capabilities, is profoundly transforming InSAR monitoring technology and its applications in disaster scenarios such as earthquakes, volcanoes, and glaciers. Convolutional Neural Networks (CNNs) excel in spatial feature extraction, capable of identifying key edges, textures, and shapes from remote sensing images, accurately capturing subtle changes. Transformer-type models, with their long-range dependency processing capabilities, effectively correlate different time points in the temporal dimension to complete sequence prediction, while breaking the receptive field limitation in the spatial dimension, making the network more globally focused. Mamba-type models, as an emerging architecture, are based on the Selective State-Space Model (SSM), achieving efficient inference and real-time response while reducing computational overhead, making them particularly suitable for the time-sensitive needs of disaster monitoring.

[0006] In the research on intelligent landslide hazard identification, the combination of deep learning and InSAR technology has evolved from patch-level or pixel-level regional localization and sensitivity mapping to more fine-grained precise localization and boundary recognition technologies. Currently, research typically uses deformation products as network input, employing semantic segmentation and target detection techniques to accurately delineate landslide area boundaries. Although deformation products have shown significant potential in intelligent landslide identification, several technical challenges remain in practical applications. First, in areas with low coherence and complex terrain, unwrapping errors are prone to occur, and existing automatic unwrapping algorithms struggle to completely resolve these issues, leading to the gradual accumulation of noise phases in the deformation results and increasing the uncertainty of identification. Furthermore, post-unwrapping data processing requires comprehensive analysis of multi-dimensional information such as time, space, and deformation amplitude, and often relies on expert knowledge for parameter optimization. Second, there is a lack of effective self-interpretation and self-analysis mechanisms. Subsequent analysis and interpretation work still requires the integration of rich geological or disaster knowledge to perform type identification, disaster-causing mechanism analysis, and risk assessment of landslide hazard areas identified by the visual model.

[0007] In summary, although current research on intelligent landslide hazard identification has made some progress, significant limitations still exist. First, the high dependence of identification results on data quality remains a key issue. Current research generally uses unwrapped products as model input, but this method is susceptible to phase unwrapping errors and error accumulation, thus reducing identification accuracy. Second, the InSAR processing workflow involves a series of complex and specialized steps, such as parameter optimization and error correction, requiring expert knowledge. Furthermore, in practical applications, users' lack of relevant professional background can lead to biases in data analysis and result interpretation. Therefore, this invention focuses on reducing the dependence on high-quality data and expert knowledge to improve the accuracy and broad applicability of intelligent landslide hazard identification. Summary of the Invention

[0008] To address the technical problems existing in the background art, this invention aims to provide an intelligent identification method for creep-type landslide hazards that combines image processing and semantic understanding, thereby solving the current high dependence on high-quality data and expert knowledge in the landslide hazard identification process. This method will achieve intelligent and automated detection, analysis, and feedback of landslide hazards, not only improving identification accuracy but also enhancing the monitoring and emergency response capabilities for landslide risks, providing innovative technical support for the prevention, monitoring, and early warning of landslide disasters.

[0009] To solve the technical problem, the technical solution of the present invention is as follows:

[0010] A method for intelligent identification of creep-type landslide hazards combining image processing and semantic understanding, the method comprising:

[0011] S1: Input: Wound InSAR interferogram; First, the image is processed through geometric modeling, phase filtering, and the introduction of external data to suppress irrelevant phase caused by the Earth's ellipsoid, topography, atmosphere, and Doppler frequency shift; then, the deformation signal of landslide hazards is highlighted by calculating the deformation phase gradient, signal trigonometric function decomposition, and RGB three-color channel mapping; finally, a deep learning semantic segmentation method combining convolution operation, self-attention mechanism, structure reparameterization, and selective state-space model technology is used to identify landslide hazards in the target area; Output: Raster and vector data of the landslide hazard area identified by the deep learning semantic segmentation method;

[0012] S2: Input the wrapped InSAR interferogram and the landslide hazard area segmentation results generated in step S1; First, generate a text description of the landslide hazard area by defining templates and expert annotations, including landslide type, landslide disaster-causing mechanism, and landslide state characteristics; then, use a deep learning image captioning generation method that combines visual and text processing to describe the landslide hazard generation features; output the text description of the landslide hazard area interpreted by the deep learning image captioning generation method.

[0013] S3: Input the landslide hazard area segmentation results generated in step S1 and the text description generated in step S2; First, generate a question-and-answer dataset for the landslide hazard area by defining templates and expert annotations. The question formats include: "Will the landslide hazard cause a landslide?", "Is it necessary to evacuate personnel?", "What emergency measures are needed?", etc.; Finally, generate an interactive question-and-answer mechanism for the landslide hazard area by using a deep learning visual question-and-answer mechanism that combines visual and text processing; Output the interpretation and judgment of the landslide hazard area and related questions by the deep learning visual question-and-answer mechanism.

[0014] Furthermore, step S1 includes:

[0015] S101: Through high-precision geometric modeling, filtering, and external data correction, irrelevant phase interference in the interferogram is suppressed, pure interferometric phase information related to landslide hazards is preserved, and a filtered and corrected interferogram is generated, providing high-quality input data for subsequent deformation signal enhancement and learning;

[0016] S102: By using phase gradient enhancement, RGB channel mapping and generative adversarial network methods, RGB images or other high-dimensional feature data containing enhanced landslide hazard deformation features are generated from InSAR data, providing a landslide deformation signal representation that is easy to learn and recognize for semantic segmentation networks.

[0017] S103: By integrating convolutional operations, self-attention mechanisms, and structural reparameterization methods, the training and inference structures are optimized. The optimized deep learning model is a segmentation network with the ability to accurately identify landslide hazard areas and perform efficient inference, outputting accurate segmentation results of landslide hazards.

[0018] Furthermore, step S2 includes:

[0019] S201: Create a dataset that matches landslide images (i.e., interferograms and segmentation results) with their corresponding text descriptions to support multimodal learning of the IC model. First, text descriptions are generated by defining description templates and annotating them with geological experts. Second, images and text are matched and diversity checks are performed to form visual-language pairs. Finally, visual and text data are augmented separately to prevent potential overfitting risks during training.

[0020] S202: By constructing a multimodal caption generation model, a visual encoder is used to encode the landslide hazard segmentation results and interference phase into visual feature vectors. Simultaneously, a language encoder is used to map the descriptive text into language feature vectors, thereby generating a landslide feature description, including visual feature vectors from the landslide hazard segmentation results and interference phase, and language feature vectors from the descriptive text, providing basic data for subsequent feature alignment and language generation. By constructing a shared feature space, visual feature vectors and language feature vectors are mapped to the same embedding space to maximize the similarity between matching image and text pairs. An embedding vector is generated by combining a lightweight mapping network, forming a seamless connection from vision to language. By inputting the connection sequence into the autoregressive language generation model, the landslide description is predicted under prefix conditions. The cross-entropy loss function is used to train the generation component to optimize the prediction effect, and finally, an accurate text description of the landslide hazard is output.

[0021] Furthermore, step S3 includes:

[0022] S301: Create a dataset that matches landslide images with their corresponding interpreted text descriptions and question-and-answer text descriptions to support multimodal learning of the VQA model. Specifically, first, text descriptions are generated by defining description templates and annotating them with geological experts; second, images and text are matched and diversity checks are performed to form visual-language pairs; and finally, visual and text data are augmented separately to prevent potential overfitting risks during training.

[0023] S302: First, ViT and BERT models are used to encode features of image and text data respectively, generating form-aligned visual feature vectors and linguistic feature vectors to form a preliminary embedding representation; second, a multimodal fusion method is used to weight and fuse the visual feature vectors and linguistic feature vectors to capture subtle differences in information; finally, the fused embedding is input into a multilayer perceptron to achieve accurate prediction of the problem, and multi-task learning is used to enhance the model's adaptability to different problem types.

[0024] Furthermore, step S1 specifically includes:

[0025] S101: Suppressing interferogram-independent phase:

[0026] To address irrelevant phase interference in the entangled phase image, firstly, the interference phase caused by the reference ellipsoid and terrain is characterized and removed through high-precision geometric modeling. Secondly, for noisy phase, before interferometry, the phase caused by Doppler frequency shift is removed by azimuth filtering, and speckle noise caused by the coherent superposition of backscattered waves from multiple scatterers is removed by multi-view processing. After interferometry, the data is further smoothed by Goldstein frequency domain filtering based on statistical characteristics. For atmospheric delay phase caused by atmospheric phase screen, it is simulated and removed using meteorological and GNSS data.

[0027] In SAR imagery, the complex signal of each pixel can be expressed using Euler's formula: S = A·e jφ , where e jφ =cos(φ) +jsin(φ), where A represents the amplitude. Let R represent the phase, λ represent the propagation path, λ represent the wavelength of the radar signal, and j be the imaginary unit. For two interferometric SAR images, they are represented in complex form as follows:

[0028]

[0029] in, S 1 indicates the first SAR image. S 2 represents the second SAR image, obtained by multiplying the complex conjugates of S1 and S2:

[0030]

[0031] Where * denotes complex conjugate, A1 represents the amplitude of S1, A2 represents the amplitude of S2, φ1 represents the phase of S1, and φ2 represents the phase of S2. After complex multiplication expansion, the phase difference can be expressed as:

[0032]

[0033] Where Δφ represents the phase difference, arg represents the phase angle, and the interference phase can be expressed according to its components as:

[0034]

[0035] in, Indicates the phase of the reference ellipsoid. Indicates terrain phase, Indicates the deformation phase. Indicates atmospheric delay phase, Indicates the phase of random noise.

[0036] Without considering and In this case, as well as This can be represented by a signal interference model. This refers to the interferometric phase corresponding to the equidistant projection points of any point on the ground onto the reference ellipsoid. The observation point P1 is projected onto the reference ellipsoid to P0 with a radius of R1 (distance from the main image) and an axis representing the satellite's velocity vector. The interferometric phase at P0 is the same as the interferometric phase at P1. It can be represented as:

[0037]

[0038] Where R1 and R2 represent the distances from the first and second imaging points to P1, respectively, and r1 and r2 represent the distances from the first and second imaging points to P0, respectively. || B represents the length of the horizontal baseline, θ represents the angle of incidence, and α represents the horizontal angle between the first and second imaging points. This refers to the additional interference phase caused by the Δθ resulting from the ground height h. The slant range difference ΔR between the two images can be expressed as: ΔR = Bsin(θ-α+Δθ) = Bsin(θ-α)cosΔθ + BsinΔθcos(θ-α). And in the state where Δθ→0, we have... Therefore, it can be further simplified to: ΔR = Bsin(θ-α) + BΔθcos(θ-α), where ΔR represents the interference phase, i.e.:

[0039]

[0040] Therefore, the terrain phase can be expressed as: in, B ⊥ This represents the length of the vertical baseline, and because the Earth's curvature is relatively small over a small area... Where H represents the geodetic height of the first imaging point, therefore After substituting this, the terrain phase can be further simplified to:

[0041]

[0042] SAR satellites experience additional propagation paths due to atmospheric phase barriers during signal propagation, resulting in atmospheric phase delay. This delay primarily consists of ionospheric and tropospheric delays. Ionospheric delay, caused by variations in electron density, is inversely correlated with the signal wavelength.

[0043]

[0044] in, The ionospheric delay phase is represented by TEC, and the total electron content is represented by TEC. Sentinel-1A / B radar images are used as the raw data, i.e., the L-band. Therefore, it is assumed that... Right now The tropospheric delay phase is dominated by tropospheric delay, which consists of stratification and turbulence components. The stratification component is related to terrain height and is removed by considering terrain factors. The turbulence component is related to water vapor content and is removed by simulating the introduction of meteorological and GNSS data.

[0045] Depending on the stage of data processing, interferometric phase filtering can be divided into pre-filtering and post-filtering. Pre-filtering is used to suppress noise caused by Doppler decorrelation and speckle effect. Doppler decorrelation manifests as spectral shifts in the azimuth and slant range directions of the main and sub-SAR images. This effect can be reduced by pre-filtering in both directions. The speckle effect refers to the multiplicative noise caused by the coherent superposition of backscattered waves generated by scatterers on the incident signal between scatterers. This noise can be suppressed through multi-look processing.

[0046]

[0047] in, The average intensity value after multi-view processing is represented by I, where I represents the intensity value of the i-th view, and L represents the total number of views. Post-filtering is based on the morphology or statistical characteristics of the interference phase to further remove noise from the interferogram. Goldstein frequency domain filtering will be used to further remove noise from the interferogram.

[0048] H(u,v)=S{Z(u,v)} α Z(u,v) (10)

[0049] Where a∈[0,1] represents the filter parameters; S{·} is the smoothing operator; u and v are the spatial frequencies; Z(u,v) and H(u,v) are the Fourier spectra of the interferogram before and after filtering, respectively;

[0050] The above methods transform the interference result into... The dominant interference phase requires... The relevant enhancement operations are performed; firstly, the sensitivity of the data to deformation perception is enhanced by using phase gradient. Phase gradient refers to the gradient of the phase difference between adjacent pixels in the interferogram, i.e., the rate of change, which represents the intensity and direction of local surface deformation.

[0051]

[0052] in, This represents the gradient of the phase value along the x-direction. φ(x,y) represents the gradient of the phase value along the direction y, φ(x,y) represents the phase value at a pixel position (x,y) in the interferogram, φ(x+1,y)-φ(x,y) and φ(x,y+1)-φ(x,y) represent the phase difference between adjacent pixels in the x and y directions, respectively, and Δx and Δy represent the pixel spacing along the x and y directions, respectively.

[0053] S102: Enhanced deformation signal:

[0054] To address the issue of weak landslide hazard signals, two different research perspectives are used to enhance them. First, from the perspective of InSAR data processing, phase gradient and trigonometric function decomposition are employed to enhance the signal's ability to represent landslide hazards. Second, from the perspective of computer vision, generative adversarial networks are used to perform simulation-level data augmentation on small sample tasks. Additionally, RGB three-color channel mapping can be used to reduce the learning difficulty of the network.

[0055] First, color rendering is used to enhance the data's ability to represent deformation features. After mapping single-channel phase or phase gradient data to an RGB image, different color channels are used to encode the strength and direction information of the phase gradient, thereby enhancing the feature representation capability. Specifically, the interference phase is first normalized to the range [0,1].

[0056]

[0057] Where φ and φ norm Let these represent the phases before and after normalization, respectively. Then, a discrete color list C = {c i} i=1,2,...,n , where each c i =(R i G i B i ), and calculate RGB values ​​through linear interpolation; in semantic segmentation tasks, convolution operations in CNNs extract local features of the image through convolution kernels; in convolution operations, input data With convolution kernel A convolution is performed, where U and V represent the length and width of the input, respectively, C represents the number of input channels, H and W represent the length and width of the convolution kernel, and D represents the number of output channels, along with a bias matrix. R and T represent the length and width of the output, B z This represents the bias of the z-th convolution to generate the output. The value of its z-th channel can be represented as:

[0058]

[0059] Among them, O (:,:,z)M represents the output value of the z-th channel. (:,:,k) This represents the input value of the k-th channel, * indicates the convolution operation, and F... (:,:,k,z) Indicates with O (:,:,z) and M (:,:,k) The corresponding convolution kernel, B (:,:,z) It is a bias B z The matrix formed:

[0060] In the self-attention mechanism, the input sequence is first transformed into three matrices through three linear transformations: the query matrix Q, the key matrix K, and the sum matrix V. Attention is then calculated.

[0061]

[0062] Where Attention(Q,K,V) represents the value of the self-attention mechanism, d k These are the dimensions of Q and K, and softmax is the activation function;

[0063] Step S103: Learn the deformation signal:

[0064] The learning ability of the model is enhanced by optimizing the network training and inference structures. Specifically, in terms of the network training structure, convolutional operations and self-attention mechanisms are integrated to enhance the network's grasp of global information. Secondly, for the inference structure, structural reparameterization and selective state space models are used to optimize the convolutional blocks in the training structure to achieve efficient algorithm transfer. CNNs are deeply integrated with self-attention mechanisms, combining spatial features and global dependencies to enhance the detection capability of landslides in complex terrain areas. In specific implementation, a multi-level feature extraction module combining CNNs and self-attention mechanisms is designed. It captures the local boundary features of the landslide area through convolutional operations and then uses the self-attention mechanism to achieve cross-regional feature information flow. This module extracts and integrates the spatial structural features and deformation information of the landslide area layer by layer through the stacking of multiple deep networks.

[0065] Based on the above convolution calculation formula, it can be inferred that convolution involves only linear operations throughout the entire calculation process; therefore, it is a linear operation. For convolutions with the same structure, they exhibit distribution and additivity:

[0066] Let F'←cF, then c(M*F)=M*F' (15)

[0067] M*F 1 +M*F 2 =M*(F 1 +F 2 Let F'←F 1 +F 2 There is M*F 1 +M*F2 =M*F'(16)

[0068] Where M represents an arbitrary constant, F, F 1 F 2 Different convolution kernels are represented by '*', and convolution operation is represented by '*'. Structural reparameterization is applied to the convolution operation part in the multi-level feature extraction module, so that the segmentation network can contain multiple branches, redundant convolution kernels or non-standard structures during training. During the inference phase, these redundant structures are merged into a simpler equivalent structure, thereby accelerating inference while maintaining the performance during training.

[0069] A state space is the minimum number of variables that can fully describe a system; it is a method of mathematically representing a problem by defining the possible states of the system; a state-space model is used to describe these states and predict the next state of the system based on the input.

[0070]

[0071] Where A, B, C, and D represent coefficient matrices, h(t) and h'(t) represent the current and next states, respectively, x(t) represents the input, and y(t) represents the output. For matrix A, to retain a large historical memory, a high-order polynomial projection HiPPO operator is used for its construction.

[0072]

[0073] Here, n and k are used to define the index of A; in order for the SSM to receive sequence data and generate sequence output, the matrix for calculating the state needs to be replaced with a discrete matrix:

[0074]

[0075] Where Δ represents the discretization time step, The discrete matrix is ​​I, which represents the identity matrix. At the same time, the data is selectively compressed to the state. Even though matrices B, C and Δ depend on the input, the selective SSM divides the entire time series into multiple shorter blocks, each containing multiple consecutive time steps. Within each block, the state updates still maintain the recursive order to ensure the temporal dependencies of the data. In the segmentation model, the selective SSM is combined with the self-attention mechanism.

[0076] Furthermore, step S2 specifically includes:

[0077] Step S201: Constructing a visual-linguistic representation dataset:

[0078] Define a template τ = [T1, T2, ..., T n Each template T iIncludes variable placeholders for inserting landslide feature information; the template format is as follows:

[0079] T i ='Landslide type<category>, landslide characteristics<evaluation>'(20)

[0080] Among them, <Type> is used to record the type of landslide, and <Evaluation> is used to record the landslide's precipitating mechanism and current state characteristics; image features F = [f1, f2, ..., f] are extracted based on the interferogram and segmentation results. n The geological experts then fill in the template with features; subsequently, image-text matching is performed, matching image features F with text template τ to generate a paired dataset P = {(f1,T1),(f2,T2),...,(f...}. n ,T n Then, diversity checks are performed on the dataset, including visual diversity checks based on image feature distribution and linguistic diversity checks based on syntactic structure. Finally, the dataset is augmented: for visual data, geometric transformations, color transformations, and noise injection are used for augmentation; for text data, synonym replacement and variant descriptions are used for augmentation. At this point, the final dataset is output. in, and These represent augmented visual features and textual descriptions, respectively.

[0081] Step S202: Learn the characteristics of landslides:

[0082] Based on a multimodal learning framework, a multimodal caption generation model is constructed using a visual encoder, a language encoder, and an attention mechanism. The visual encoder encodes the input landslide hazard segmentation result and its corresponding interference phase into a visual feature vector V. The language encoder maps the descriptive text of the landslide hazard area to the same semantic space and encodes it into a linguistic feature vector L.

[0083]

[0084] Where, x i y i These represent the values ​​of the i-th index in the (wound) InSAR interferogram and the landslide hazard area segmentation result produced by S1, respectively. i This represents the value at index i in the landslide text description; Encoder vision and Encoder language These represent the visual and speech encoders, respectively.

[0085] By constructing a shared feature space, V and L are mapped to the same embedding space; the training objective is to maximize the cosine similarity of similar image and text pairs while minimizing the similarity of non-matching pairs.

[0086]

[0087] Where ||·|| represents the Euclidean norm, Sim(V,L) represents the similarity between visual and linguistic features, and then a lightweight mapping network is used to map the embedding vectors in the feature space into k embedding vectors:

[0088]

[0089] Wherein, F(CLIP(x) i )) indicates that the input x i The feature vectors generated by the CLIP model are further mapped to functions in the embedding space; each vector It has the same dimensions as the word embedding; and is connected to the caption embedding along with the visual embedding:

[0090]

[0091] By connecting the prefix and subtitle sequence during training The input is fed into a language model, which then predicts the title tag in an autoregressive manner under prefix conditions; the mapping component is trained using cross-entropy loss.

[0092] Furthermore, step S3 specifically includes:

[0093] Step S301: Construct a visual-language question answering dataset:

[0094] The question-and-answer text description includes the question Q and the answer A. First, define the question template set T. Q =[T Q1 ,T Q2 ,...,T Qn ], where each question template T Qi Questions are posed based on specific characteristics of the landslide; subsequently, a set of answer templates T is defined. A =[T A1 ,T A2 ,...,T An ], where each question template T Ai Corresponding features extracted from landslide images are used for feature filling, and question-answer pairs (Q,A) = (T) are formed using questions and answers. Q (F),T A(F)); then, perform diversity checks on the dataset, including visual diversity checks based on image feature distribution and linguistic diversity checks based on syntactic structure; finally, perform data augmentation on the dataset; at this point, output the final dataset. in, and These represent the augmented visual features and the text (question and corresponding answer), respectively.

[0095] Step S302: Learn the characteristics of landslides. Answer:

[0096] Design a landslide hazard VQA model comprising three steps: feature extraction, fusion, and prediction. In the feature extraction part, ViT and BERT models are used to extract features from images and text, respectively. The image input is divided into fixed-size patches by ViT and encoded into a series of embedding representations, defined as:

[0097]

[0098] Where, d v Let K represent the dimension of the embedding, and K represent the number of tokens after segmentation. For a given question text input, the BERT model segments the text into several word embeddings, defined as:

[0099]

[0100] Where, d q The embedding dimension is represented by M, and the number of tokens after block division is represented by M. Context-dependent embeddings are generated by introducing self-attention and cross-attention mechanisms. Given the text representation Q generated by BERT... i Attention weights are calculated using two fully connected layers and the ReLU activation function:

[0101]

[0102] in, This represents the attention weight score of the m-th token. and These are the learned weight matrices, and then the attention weights are applied to the text representation q. m We obtain attention-weighted text embeddings by performing weighted summation:

[0103]

[0104] in, Let V represent the attention weight score of the m-th token after softmax activation. Then, for the image embedding V generated by ViT... i First, it is combined with an attention-weighted text representation q:

[0105]

[0106] in, and It is a linear transformation matrix; subsequently, the combined image is represented as v′. k Calculate attention weights;

[0107] We obtain attention-weighted image embeddings by performing a weighted summation on the image embeddings:

[0108]

[0109] in, Indicates the embedding of the k-th image v k Attention weights are used in the fusion stage, and MUTAN is chosen to combine image and text embeddings. MUTAN uses Tucker decomposition to fuse the tensors. Decomposed into three matrices and a small kernel tensor:

[0110] T=((T c ×1W q )×2W v )×3W o (31)

[0111] in, These are matrices for preparing the fusion process;

[0112] The fused embeddings are fed into a multilayer perceptron, with each dimension corresponding to a possible answer category. At this stage, the model also integrates multi-task learning, enhancing its adaptability by training attention modules and classification weights for each question type separately.

[0113]

[0114] Where y represents the final output, and MLP represents a multilayer perceptron. This represents a multimodal fusion mechanism used to integrate image features. It is fused with text features q.

[0115] Compared with the prior art, the advantages of the present invention are as follows:

[0116] This will enable the detection, analysis, and feedback of landslide hazards through intelligent and automated methods. This will not only improve identification accuracy but also enhance landslide risk monitoring and emergency response capabilities, providing innovative technological support for landslide disaster prevention, monitoring, and early warning. This will address the current heavy reliance on high-quality data and expert knowledge in the landslide hazard identification process. Attached Figure Description

[0117] Figure 1Overall technical roadmap of this invention;

[0118] Figure 2 A technical roadmap for InSAR semantic segmentation models oriented towards wrapped phase processing;

[0119] Figure 3 A technical roadmap for multimodal semantic parsing and image caption generation mechanisms;

[0120] Figure 4 A technical roadmap for intelligent feedback and visual question-answering mechanisms. Detailed Implementation

[0121] The specific implementation of the present invention is described below with reference to embodiments:

[0122] It should be noted that the structures, proportions, sizes, etc. shown in this specification are only used to complement the content disclosed in the specification for those skilled in the art to understand and read, and are not intended to limit the conditions under which the present invention can be implemented. Any modifications to the structure, changes in the proportions, or adjustments to the size, without affecting the effects and objectives that the present invention can produce, should still fall within the scope of the technical content disclosed in the present invention.

[0123] Furthermore, the terms such as "upper," "lower," "left," "right," "middle," and "one" used in this specification are merely for clarity of description and are not intended to limit the scope of the invention. Any changes or adjustments to their relative relationships, without substantially altering the technical content, should also be considered within the scope of the invention.

[0124] Example 1:

[0125] The key technologies focused on in this invention are the automation, intelligentization, and multimodal interpretation of landslide hazard identification, aiming to overcome the heavy reliance of traditional methods on data quality and expert experience, and solve the bottleneck problems in landslide disaster monitoring. The specific key technologies are as follows:

[0126] Landslide hazard identification from interferograms. In landslide hazard identification, the phase unwrapping process of InSAR data often introduces errors and operational complexity. This invention designs a semantic segmentation model directly based on the wrapped phase, reducing the dependence of traditional phase unwrapping and time series processing on data quality, thereby improving the model's adaptability in low-coherence and complex terrain areas. Key technologies include:

[0127] 1) Interferogram-Independent Phase Suppression. To address irrelevant phase interference in entangled phase images, this invention first characterizes and removes the interferometric phase caused by the reference ellipsoid and terrain through high-precision geometric modeling. Secondly, for noisy phase, before interferometry, azimuth (range) filtering is used to remove the phase caused by Doppler frequency shift, and multi-look processing is used to remove speckle noise caused by the coherent superposition of backscattered waves from multiple scatterers. After interferometry, Goldstein frequency domain filtering based on statistical characteristics further smooths the data. Atmospheric delay phase (mainly tropospheric delay) caused by atmospheric phase screens is simulated and removed using meteorological and GNSS data.

[0128] 2) Deformation Signal Enhancement. To address the weak signal of landslide hazards, this invention enhances it through two different research perspectives. First, from the perspective of InSAR data processing, phase gradient and / or trigonometric function decomposition is used to enhance the signal's ability to represent landslide hazards. Second, from the perspective of computer vision, a generative adversarial network is used to perform simulation-level data augmentation on small sample tasks. Simultaneously, RGB three-color channel mapping is employed to reduce the learning difficulty of the network.

[0129] 3) Learning Deformation Signals. Addressing the issue of landslide hazard signals being easy to fit but difficult to learn, this invention enhances the model's learning ability by optimizing the network training and inference structures. Specifically, in the network training structure, convolutional operations and self-attention mechanisms are integrated to enhance the network's grasp of global information. Secondly, for the inference structure, this invention uses structural reparameterization and a selective state-space model to optimize the convolutional blocks in the training structure, achieving efficient algorithm transfer.

[0130] From visual segmentation to semantic understanding. In landslide hazard identification, the results identified by segmentation models (detection models) often require further interpretation of landslide type, precipitating mechanism, and landslide state based on expert knowledge. This presents problems such as high interpretation threshold, reliance on expert knowledge, and lack of self-analysis and self-interpretation mechanisms. Therefore, this invention relates to the collaborative interpretation of different data modalities (images, text). To achieve automated type classification and causal analysis of landslide hazards, this invention introduces a multimodal image captioning generation model, combining InSAR deformation images with landslide segmentation results to generate semantic descriptions through alignment of visual and linguistic features. Key technologies include:

[0131] 1) Constructing a Visual-Language Representation Dataset. A dataset matching landslide images (interferograms and segmentation results) with their corresponding text descriptions is created to support multimodal learning of the image captioning generation model. Specifically, first, text descriptions are generated by defining description templates and annotating them with geological experts. Second, images and text are matched and diversity checks are performed to form visual-language pairs. Finally, the visual and text data are augmented separately to prevent potential overfitting risks during training.

[0132] 2) Learning Landslide Feature Descriptions. Through multimodal model learning, visual segmentation results are transformed into precise textual descriptions, accurately expressing the characteristics and risks of landslide hazards. Specifically, firstly, landslide feature descriptions are generated using CLIP (contrastive language-image pre-trained model) embedding as a prefix. Secondly, a mapping network is used to construct a seamless connection between vision and language. Finally, an automatic regression language generation model is used to generate predicted landslide descriptions.

[0133] From static identification to dynamic interpretation. To achieve closed-loop feedback capability in the landslide hazard identification system, this invention designs an intelligent feedback mechanism based on visual question answering, further enhancing the practicality of the interpretation results through interactive question-and-answer sessions with the user. Key technologies include:

[0134] 1) Constructing a Visual-Language Question Answering Dataset. A dataset is created that matches landslide images (interferograms and / or segmentation results) with their corresponding interpreted text descriptions, as well as question-and-answer text descriptions, to support multimodal learning of the visual question answering model. Specifically, first, text descriptions are generated by defining description templates and annotating them with geological experts. Second, images and text are matched and diversity checks are performed to form visual-language pairs. Finally, the visual and text data are augmented separately to prevent potential overfitting risks during training.

[0135] 2) Learning Landslide Feature Answers. Through multimodal model learning, accurate answer descriptions are generated based on user-submitted questions, precisely representing landslide hazards in readable text. Specifically, firstly, visual and linguistic features are encoded separately to form form-aligned visual (linguistic) feature vectors. Secondly, multimodal fusion is used to weight and fuse these visual (linguistic) feature vectors. Finally, a multilayer perceptron is used to predict the problem.

[0136] Example 2:

[0137] like Figure 1 As shown, this invention provides an intelligent landslide hazard identification framework that combines InSAR technology and deep learning. The complete steps of the method of this invention include:

[0138] S1: Construction of an InSAR Semantic Segmentation Model for Wrapped Phase Processing: Addressing the potential unwrapping errors and operational complexity in InSAR data processing, this invention constructs a semantic segmentation model that directly analyzes interferograms to reduce the dependence of traditional methods on data quality for phase unwrapping and parameter tuning. Specific details include: 1) Effectively suppressing irrelevant phases through geometric modeling, phase filtering, and the introduction of external data; 2) Effectively highlighting landslide hazard deformation signals from the perspective of InSAR data processing or / and computer vision; 3) Enhancing model learning by optimizing the network training and inference structures; (See details...) Figure 2 )

[0139] S2: Construction of Multimodal Semantic Analysis and Image Captioning Generation Mechanism: Addressing the intelligent interpretation needs of landslide disasters, this invention introduces multimodal image captioning generation technology. Landslide area segmentation results and InSAR deformation images are input into a multimodal model. Text analysis results are generated through the alignment of visual and linguistic features, achieving deep semantic analysis of landslide hazards. Specific content includes: 1) Creating a dataset matching landslide images (interferograms, segmentation results) and their corresponding text descriptions to support multimodal learning of the captioning generation model; 2) Through multimodal model learning, visual segmentation results are transformed into precise text descriptions, accurately expressing the characteristics and risks of landslide hazards; (See details...) Figure 3 )

[0140] S3: Intelligent Feedback and Visual Question-Answering Mechanism Construction: To achieve closed-loop feedback capability in the intelligent landslide hazard identification system, this invention designs an automated result feedback mechanism after the identification and interpretation stages. By introducing a visual question-answering mechanism, the identification and analysis results are clearly conveyed to the user in an interactive manner. Specifically, this includes: 1) Creating a dataset that matches landslide images (interferograms, segmentation results) with their corresponding interpreted text descriptions and question-answering text descriptions to support multimodal learning of the visual question-answering model. 2) Through multimodal model learning, generating accurate answer descriptions based on user-submitted questions, accurately representing landslide hazards as readable text. (See details...) Figure 4 )

[0141] InSAR semantic segmentation model construction for wrapped phase processing

[0142] Step S101: Suppress irrelevant phases in the interferogram.

[0143] This step addresses irrelevant phase interference in the entangled phase image. First, the invention characterizes and removes the interferometric phase caused by the reference ellipsoid and terrain through high-precision geometric modeling. Second, for noisy phase, before interferometry, azimuth (range) filtering is used to remove the phase caused by Doppler frequency shift, and multi-look processing is used to remove speckle noise caused by the coherent superposition of backscattered waves from multiple scatterers. After interferometry, the data is further smoothed using Goldstein frequency domain filtering based on statistical characteristics. For atmospheric delay phase (mainly tropospheric delay) caused by atmospheric phase screens, the invention simulates and removes it using meteorological and GNSS data.

[0144] In SAR imagery, the complex signal of each pixel can be expressed using Euler's formula: S = A·e jφ Among them, e jφ =cos(φ)+jsin(φ), where A represents the amplitude (intensity). Let R represent the phase, λ represent the propagation path, λ represent the wavelength of the radar signal, and j be the imaginary unit. Two interferometric SAR images can be represented in complex form:

[0145]

[0146] Where S1 represents the first SAR image and S2 represents the second SAR image. By multiplying the complex conjugates of S1 and S2, we obtain:

[0147]

[0148] Where * denotes complex conjugate, A1 represents the amplitude of S1, A2 represents the amplitude of S2, φ1 represents the phase of S1, and φ2 represents the phase of S2. After complex multiplication expansion, the phase difference can be expressed as:

[0149]

[0150] Where Δφ represents the phase difference, arg represents the phase angle, and the interference phase can be expressed according to its components as:

[0151]

[0152] in, Indicates the phase of the reference ellipsoid. Indicates terrain phase, Indicates the deformation phase. Indicates atmospheric delay phase, This represents the phase of random noise.

[0153] Without considering and In this case, as well as It can be represented by a signal interference model. This refers to the interferometric phase (de-flattening effect) corresponding to the equidistant projection points of any point on the ground onto the reference ellipsoid. The observation point P1 is projected onto the reference ellipsoid to P0 with a radius of distance R1 from the main image and the satellite's velocity vector as the axis. The interferometric phase at P0 is the same as the interferometric phase at P1. It can be represented as:

[0154]

[0155] Where R1 and R2 represent the distances from the first and second imaging points to P1, respectively; r1 and r2 represent the distances from the first and second imaging points to P0, respectively; B represents the length of the horizontal baseline; θ represents the incident angle; and α represents the horizontal angle between the first and second imaging points. This refers to the additional interference phase caused by the Δθ resulting from the ground height h. The slant range difference ΔR between the two images can be expressed as: ΔR = Bsin(θ-α+Δθ) = Bsin(θ-α)cosΔθ + BsinΔθcos(θ-α). And in the state where Δθ→0, we have... Therefore, it can be further simplified to: ΔR = Bsin(θ-α) + BΔθcos(θ-α). Using ΔR to represent the interference phase, that is:

[0156]

[0157] Therefore, the terrain phase can be expressed as: Among them, B ⊥ This indicates the length of the vertical baseline. Also, because the Earth's curvature is relatively small over a small area... Where H represents the geodetic height of the first imaging point. Therefore After substituting this, the terrain phase can be further simplified to:

[0158]

[0159] SAR satellites experience additional propagation paths due to atmospheric phase barriers during signal propagation, resulting in atmospheric phase delay. This primarily includes ionospheric and tropospheric delays. Ionospheric delay, caused by variations in electron density, is inversely correlated with the signal wavelength.

[0160]

[0161] in, The ionospheric delay phase is represented by TEC, and the total electron content is represented by TEC. This invention primarily uses Sentinel-1A / B radar images as raw data (L-band); therefore, it is assumed that... Right now The tropospheric delay phase is dominated by tropospheric delay. This delay phase consists of stratification and turbulence components. The stratification component is related to terrain height and can be removed by considering terrain factors. The turbulence component is related to water vapor content and can be removed through simulation by incorporating meteorological and GNSS data.

[0162] Interferometric phase contains a large amount of noise phase, which is usually suppressed through filtering. Depending on the stage of data processing, filtering can be divided into pre-filtering and post-filtering. Pre-filtering is mainly used to suppress noise caused by Doppler decorrelation and speckle effects. Doppler decorrelation manifests as spectral shifts in the azimuth and slant range directions in both main and sub-SAR images; this effect can be reduced by pre-filtering in each of these directions. The speckle effect refers to the multiplicative noise speckles caused by the coherent superposition of backscattered waves from scatterers on the incident signal between scatterers; this noise can be suppressed through multi-look processing.

[0163]

[0164] in, Let I represent the average intensity value after multi-view processing, and L represent the intensity value of the i-th view. Post-filtering, based on the morphological or statistical characteristics of the interference phase, further removes noise from the interferogram. This invention employs Goldstein frequency domain filtering to further remove noise from the interferogram.

[0165] H(u,v)=S{Z(u,v)} α Z(u,v) (10)

[0166] Where a∈[0,1] represents the filter parameters; S{·} is the smoothing operator; u and v are the spatial frequencies; Z(u,v) and H(u,v) are the Fourier spectra of the interferogram before and after filtering, respectively.

[0167] The above methods can transform the interference result into... The dominant interference phase requires... Further enhancement operations are performed. First, the sensitivity of the data to deformation perception is enhanced using phase gradient methods. Phase gradient refers to the gradient (i.e., rate of change) of the phase difference between adjacent pixels in an interferogram, which can be used to represent the intensity and direction of local surface deformation.

[0168]

[0169] in, Indicates the phase value along the direction x gradient, Let φ(x,y) represent the phase gradient along the direction y, φ(x,y) represent the phase value at a pixel position (x,y) in the interferogram, and φ(x+1,y)-φ(x,y) and φ(x,y+1)-φ(x,y) represent the phase values ​​at... x The phase difference between adjacent pixels in the y-direction, Δx and Δy represent the phase difference along the direction, respectively. x The pixel spacing in the y-direction;

[0170] Step S102: Enhance the deformation signal.

[0171] This step addresses the issue of weak landslide hazard signals. This invention enhances these signals through two different research perspectives. First, from an InSAR data processing perspective, phase gradient and / or trigonometric function decomposition can be used to enhance the signal's ability to represent landslide hazards. Second, from a computer vision perspective, generative adversarial networks can be used for simulation-level data augmentation on small sample tasks. Simultaneously, RGB three-color channel mapping can be employed to reduce the learning difficulty of the network.

[0172] First, color rendering enhances the data's ability to represent deformation features. In single-channel interferograms, the limited information dimensionality fails to fully express the complex spatial changes easily discernible to the human eye. By mapping single-channel phase or phase gradient data to an RGB image, different color channels can be used to encode information such as the strength and direction of the phase gradient, thereby enhancing feature representation. Specifically, the interferometric phase is first normalized to the range [0,1].

[0173]

[0174] Where φ and φ norm These represent the phase before and after normalization, respectively. Then, a discrete color list C = {c i} i=1,2,...,n , where each c i =(R i G i B i The RGB values ​​are calculated using linear interpolation. In semantic segmentation tasks, convolution operations in CNNs extract local features of an image through convolution kernels. During convolution operations, the input data... (U and V represent the length and width of the input, respectively, and C represents the number of input channels) and convolution kernel (H and W represent the length and width of the convolution kernel, respectively, and D represents the number of output channels) Perform convolution and add a bias matrix. (R and T represent the length and width of the output, B) z (representing the bias of the z-th convolution) to generate the output. The value of its z-th channel can be represented as:

[0175]

[0176] Among them, O (:,:,z) M represents the output value of the z-th channel. (:,:,k) This represents the input value of the k-th channel, * indicates the convolution operation, and F... (:,:,k,z) Indicates with O (:,:,z) and M (:,:,k) The corresponding convolution kernel, B (:,:,z) It is a bias B z The matrix formed.

[0177] However, a fixed kernel size limits convolution to a fixed receptive field, making it unable to capture global features. Self-attention mechanisms, on the other hand, incorporate long-distance dependencies into the model by calculating the correlations between different locations in the feature map, enabling effective integration of spatial and temporal features. In a self-attention mechanism, the input sequence first undergoes three linear transformations to obtain three matrices: a query matrix Q, a key matrix K, and a sum matrix V. Then, the attention (d) is calculated. k (These are the dimensions of Q and K):

[0178]

[0179] Where Attention(Q,K,V) represents the value of the self-attention mechanism, d k These are the dimensions of Q and K, and softmax is the activation function;

[0180] Step S103: Learn the deformation signal.

[0181] This step addresses the problem of landslide hazard signals being easy to fit but difficult to learn. This invention enhances the model's learning ability by optimizing the network training and inference structures. Specifically, in the network training structure, convolutional operations and self-attention mechanisms are integrated to enhance the network's grasp of global information. Secondly, for the inference structure, this invention uses structural reparameterization and a selective state-space model to optimize the convolutional blocks in the training structure, achieving high-efficiency algorithm transfer. This invention deeply integrates CNNs with self-attention mechanisms, combining spatial features and global dependencies to enhance the detection capability of landslides in complex terrain areas. In specific implementation, a multi-layered feature extraction module combining CNNs and self-attention mechanisms is designed. Convolutional operations capture local boundary features of the landslide area, and self-attention mechanisms enable cross-regional feature information flow. This module extracts and integrates spatial structural features and deformation information of the landslide area layer by layer through multi-layer deep network stacking.

[0182] Based on the convolution calculation formula above, it can be deduced that convolution involves only linear operations throughout the entire calculation process; therefore, it is a linear operation. For convolutions with the same structure, they exhibit distribution and additivity:

[0183] Let F'←cF, then c(M*F)=M*F' (15)

[0184] M*F 1 +M*F 2 =M*(F 1 +F 2 Let F'←F 1 +F 2 There is M*F 1 +M*F 2 =M*F' (16)

[0186] Where M represents an arbitrary constant, F, F 1 F 2 * indicates different convolution kernels, and * indicates a convolution operation. Structure reparameterization is an optimization technique based on the distribution and additivity of homogeneous convolution operations. Its core lies in employing complex structures during the training phase to enhance feature learning and model expressive power, and then simplifying them into efficient equivalent structures during the inference phase, thereby reducing computational overhead and improving inference efficiency. This technique is mainly applied to the convolution operation part of multi-level feature extraction modules, allowing the segmentation network to contain multi-branch, redundant convolution kernels, or non-standard structures during training to better adapt to data and enhance model robustness. During the inference phase, these redundant structures are merged into a simpler equivalent structure, thereby accelerating inference while maintaining training performance.

[0187] A state space is the minimum number of variables that can fully describe a system. It is a method of mathematically representing a problem by defining the possible states of the system. State-space models are used to describe these states and predict the next state of the system based on the input.

[0188]

[0189] Where A, B, C, and D represent coefficient matrices, h(t) and h'(t) represent the current and next states, respectively, x(t) represents the input, and y(t) represents the output. For matrix A, to retain a large historical memory, a high-order polynomial projection HiPPO operator is used for its construction.

[0190]

[0191] Here, n and k are used to define the indices of A. To enable the SSM to receive (discrete) sequence data and generate sequence output, the matrix used to compute the state must be replaced with a discrete matrix:

[0192]

[0193] Where Δ represents the discretization time step, Let I be the discretized matrix, representing the identity matrix. To selectively compress data to the state, even though matrices B, C, and Δ depend on the input, selective SSM divides the entire time series into multiple shorter blocks, each containing multiple consecutive time steps. Within each block, state updates maintain a recursive order to ensure the temporal dependencies of the data. In the segmentation model, combining selective SSM with a self-attention mechanism avoids the computational bottleneck of self-attention in long sequence processing, improving both the efficiency and accuracy of the segmentation model.

[0194] Constructing a multimodal semantic parsing and image captioning generation mechanism

[0195] Step S201: Construct a visual-linguistic representation dataset

[0196] Define a template τ = [T1, T2, ..., T n Each template T i Includes variable placeholders for inserting landslide feature information. The template format is as follows:

[0197] T i ='Landslide type<category>, landslide characteristics<evaluation>'(20)

[0198] Among them, <Type> records the type of landslide, and <Evaluation> records the landslide's precipitating mechanism and current status, among other characteristic information. Image features F = [f1, f2, ..., f...] are extracted based on the interferogram and segmentation results. n The template is filled in by geological experts based on features. Then, image-text matching is performed, matching image features F with text template τ to generate a paired dataset P = {(f1,T1),(f2,T2),...,(f...}. n ,T n Then, the dataset undergoes diversity checks, including visual diversity checks based on image feature distribution and linguistic diversity checks based on syntactic structure. Finally, the dataset is augmented: for visual data, geometric transformations, color transformations, and noise injection are used for augmentation; for text data, synonym replacement and variant descriptions are used for augmentation. At this point, the final dataset is output. in, and These represent augmented visual features and textual descriptions, respectively.

[0199] Step S202: Learn the characteristics of landslides.

[0200] This step is based on a multimodal learning framework, using a visual encoder, a language encoder, and an attention mechanism to construct a multimodal caption generation model. The visual encoder encodes the input landslide hazard segmentation result and its corresponding interference phase into a visual feature vector V. The language encoder maps the descriptive text of the landslide hazard area to the same semantic space, encoding it into a linguistic feature vector L.

[0201]

[0202] Where, x i y i These represent the values ​​of the i-th index in the (wound) InSAR interferogram and the landslide hazard area segmentation result produced by S1, respectively. i This represents the value at the i-th index in the landslide text description. (Encoder) vision and Encoder language These represent the visual encoder and the speech encoder, respectively.

[0203] This step maps V and L to the same embedding space by constructing a shared feature space. Its training objective is to maximize the cosine similarity of similar image and text pairs while minimizing the similarity of non-matching pairs.

[0204]

[0205] Where ||·|| represents the Euclidean norm, and Sim(V,L) represents the similarity between visual and linguistic features. Then, a lightweight mapping network is used to map the embedding vectors (from CLIP) in the feature space into k embedding vectors:

[0206]

[0207] Wherein, F(CLIP(x) i )) indicates that the input x i The feature vectors generated by the CLIP model are further mapped to functions in the embedding space. Each vector It has the same dimensions as the word embedding and is linked to the caption embedding along with the visual embedding.

[0208]

[0209] This step involves connecting the prefix and subtitle sequence during training. The data is input into a language model. The model is then trained to predict the title's tag in an autoregressive manner, given the prefix. For this purpose, a simple yet effective cross-entropy loss method is used to train the mapping component.

[0210] Building an intelligent feedback and visual question answering mechanism

[0211] Step S301: Construct a visual-language question answering dataset

[0212] The question-and-answer text description includes the question Q and the answer A. First, define the question template set T. Q =[T Q1 ,T Q2 ,...,T Qn ], where each question template T Qi Questions are posed based on specific characteristics of the landslide. Then, a set of answer templates T is defined. A =[T A1 ,T A2 ,...,T An ], where each question template T Ai Corresponding features extracted from landslide images are used for feature filling, and question-answer pairs (Q,A) = (T) are formed using questions and answers. Q (F),T A (F) Then, diversity checks are performed on the dataset, including visual diversity checks based on image feature distribution and linguistic diversity checks based on syntactic structure. Finally, the dataset is augmented. At this point, the final dataset is output. in, and These represent the augmented visual features and the text (question and corresponding answer), respectively.

[0213] Step S302, Learning the characteristics of landslides:

[0214] This step designs a landslide hazard VQA model comprising three steps: feature extraction, fusion, and prediction. In the feature extraction part, ViT and BERT models are mainly used to extract features from images and text, respectively. The image input is divided into fixed-size patches by ViT and encoded into a series of embedding representations, defined as:

[0215]

[0216] Where, d v Let K represent the dimension of the embedding, and K represent the number of tokens after segmentation. For a given question text as input, the BERT model segments the text into several word embeddings, defined as:

[0217]

[0218] Where, dq Let M represent the dimension of the embedding, and M represent the number of tokens after block division. Then, context-sensitive embeddings are generated by introducing self-attention and cross-attention mechanisms. Subsequently, the text representation Q generated by BERT can be used as a reference. i Attention weights are calculated using two fully connected layers and the ReLU activation function:

[0219]

[0220] in, This represents the attention weight score of the m-th token. and These are the learned weight matrices, and then the attention weights are applied to the text representation q. m We obtain attention-weighted text embeddings by performing weighted summation:

[0221]

[0222] in, This represents the attention weight score of the m-th token after softmax activation. Then, for the image embedding V generated by ViT... i First, it is combined with an attention-weighted text representation q:

[0223] v' k =(W v v k )⊙(W q q) (29)

[0224] in, and It is a linear transformation matrix. The combined image v′ can then be represented. k Calculate attention weights.

[0225] This step performs a weighted summation of the image embeddings to obtain attention-weighted image embeddings:

[0226]

[0227] in, Indicates the embedding of the k-th image v k Attention weights are assigned. During the fusion phase, MUTAN (Multimodal Tucker Fusion) can be chosen to combine image and text embeddings. MUTAN uses Tucker decomposition to fuse the tensors. Decomposed into three matrices and a small kernel tensor:

[0228] T=((T c ×1 W q )×2 W v )×3 Wo (31)

[0229] in, These are matrices used to prepare for the fusion process. MUTAN's design effectively captures subtle differences in information from the combination of visual and linguistic data, enabling the model to better handle remote sensing question-answering tasks.

[0230] This step feeds the fused embeddings into a multilayer perceptron, with each dimension corresponding to a possible answer category. The model also integrates multi-task learning at this stage, enhancing its adaptability by training attention modules and classification weights separately for each question type.

[0231]

[0232] Where y represents the final output, and MLP represents a multilayer perceptron. This represents a multimodal fusion mechanism used to integrate image features. It is fused with text features q.

[0233] Through the steps described above, the method of this invention improves the model's question-answering performance in RSVQA tasks while maintaining the integration of multimodal data. The combination of these techniques and formulas enables the model to dynamically understand the relationship between visual data and text queries, thereby providing accurate answers to complex remote sensing visual question answering.

[0234] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

[0235] Many other changes and modifications can be made without departing from the concept and scope of this invention. It should be understood that this invention is not limited to the specific embodiments, and the scope of this invention is defined by the appended claims.

Claims

1. A method for intelligent identification of creeping landslide hazards combining image processing and semantic understanding, characterized in that, The method comprises: S1: input the wrapped InSAR interferogram; first, process the image to suppress irrelevant phases caused by the earth ellipsoid, terrain, atmosphere and Doppler shift through geometric modeling, phase filtering and the introduction of external data; then, highlight the landslide hazard deformation signal by calculating the deformation phase gradient, signal trigonometric function decomposition and RGB three-color channel mapping method; finally, realize the identification of the target area landslide hazard by the deep learning semantic segmentation method of joint convolution operation, self-attention mechanism, structure reparameterization and selective state space model technology; output the raster and vector data of the landslide hazard area identified by the deep learning semantic segmentation method; S2: input the wrapped InSAR interferogram and the landslide hazard area segmentation result produced by step S1; first, generate a text description of the landslide hazard area by defining a template and expert annotation, including landslide type, landslide disaster mechanism and landslide state characteristics, then generate a feature expression of the landslide hazard by the deep learning image caption generation method combining visual and text processing; output the text description of the landslide hazard area interpreted by the deep learning image caption generation method; S3: input the landslide hazard area segmentation result produced by step S1 and the text description produced by step S2; first, generate a question-answer data set of the landslide hazard area by defining a template and expert annotation, and finally generate an interactive question and answer mechanism for the landslide hazard area by the deep learning visual question and answer mechanism method combining visual and text processing; output the interpretation and interpretation of the landslide hazard area and related questions by the deep learning visual question and answer mechanism method. 2.The method according to claim 1, characterized in that, The step S1 comprises: S101: suppress irrelevant phase interference in the interferogram and retain pure interferometric phase information related to landslide hazards by high-precision geometric modeling, filtering and external data correction, to generate filtered and corrected interferograms and provide high-quality input data for subsequent deformation signal enhancement and learning; S102: generate an RGB image or other high-dimensional feature data containing enhanced landslide hazard deformation features from InSAR data through phase gradient enhancement, RGB channel mapping and generative adversarial network method, to provide a landslide deformation signal representation easy to learn and identify for the semantic segmentation network; S103: optimize the training and inference structure by fusing convolution operation, self-attention mechanism and structure reparameterization method, the optimized deep learning model has a segmentation network with precise identification of landslide hazard areas and efficient inference ability, and outputs the accurate segmentation result of the landslide hazard. 3.The method according to claim 1, characterized in that, The step S2 comprises: S201: create a data set matching the landslide image, i.e. the interferogram and the segmentation result, and its corresponding text description to support the multimodal learning of the IC model; first, generate a text description by defining a description template and labeling by a geological expert, second, match the image and the text and perform diversity check to form a visual-linguistic pair, and finally, augment the visual data and text data respectively to prevent potential overfitting risk in the training process; S202: By constructing a multi-modal caption generation model, the landslide hazard segmentation result and the interference phase are encoded into a visual feature vector using a visual encoder, and a language encoder is used to map the description text into a language feature vector, thereby generating a landslide feature description, including a visual feature vector from the landslide hazard segmentation result and the interference phase, and a language feature vector from the description text, providing basic data for subsequent feature alignment and language generation; by constructing a shared feature space, the visual feature vector and the language feature vector are mapped to the same embedding space to maximize the similarity of the image and text pair, and a lightweight mapping network is combined to generate an embedding vector, forming a seamless connection from vision to language; by inputting the connection sequence into the autoregressive language generation model, the landslide description is predicted under the prefix condition, thereby training the generation component using the cross-entropy loss function to optimize the prediction effect, and finally outputting an accurate landslide hazard text description.

4. The intelligent identification method of the creep landslide hazard combining image processing and semantic understanding according to claim 1, characterized in that, The step S3 comprises: S301: Create a data set matching the landslide image and its corresponding interpreted text description, and the question and answer text description, to support multi-modal learning of the VQA model. First, define a description template and label it by a geologist to generate a text description. Second, match the image and text and perform diversity check to form a visual-linguistic pair. Finally, augment the visual data and text data respectively to prevent potential overfitting risk during training; S302: First, use ViT and BERT models to encode the image and text data respectively to generate form-aligned visual feature vectors and language feature vectors to form preliminary embedding representations. Second, use a multi-modal fusion method to weight and fuse the visual feature vectors and language feature vectors to capture subtle differences in information. Finally, input the fused embedding into a multi-layer perceptron to achieve accurate prediction of the question, and enhance the model's adaptability to different question types through multi-task learning.

5. The intelligent identification method of the creep landslide hazard combining image processing with semantic understanding according to claim 3, characterized in that, The step S2 specifically comprises: Step S201: Construct a visual-linguistic expression data set: A set of templates is defined Each template Including variable placeholders for inserting the characteristic information of the landslide, the template form is as follows: (20) wherein, for recording the type of landslide, for recording the disaster-producing mechanism of landslide and the state characteristic information of the landslide; image features are extracted according to the interferogram and the segmentation result , filled in by a geological expert in combination with the characteristics; subsequently, image-text matching is performed, and the image features and the text template are matched to generate a paired data set ; then, the data set is subjected to diversity checking, including visual diversity checking based on the distribution of image features and language diversity checking based on the syntactic structure; finally, the data set is subjected to data augmentation: for visual data, augmentation is performed in the manner of geometric transformation, color transformation and noise injection; for text data, augmentation is performed in the manner of synonym replacement and variant description; at this time, the final data set is output wherein, and represent the augmented visual features and text description, respectively Step S202: Learn landslide feature expression: Based on a multi-modal learning framework, a multi-modal subtitle generation model is constructed using a visual encoder, a language encoder and an attention mechanism; the visual encoder encodes the input landslide hazard segmentation result and its corresponding interference phase into a visual feature vector ; the language encoder maps the description text of the landslide hazard area to the same semantic space and encodes it into a language feature vector : (21) wherein, , respectively represent the wrapped InSAR interferogram and the landslide hazard area segmentation result indexed by S1 produced by the first model, respectively represent the wrapped InSAR interferogram and the landslide hazard area segmentation result indexed by S1 produced by the first model, respectively represent the wrapped InSAR interferogram and the landslide hazard area segmentation result indexed by S1 produced by the first model; respectively represent the wrapped InSAR interferogram and the landslide hazard area segmentation result indexed by S1 produced by the first model; and respectively represent the visual and linguistic encoders; By constructing a shared feature space, we map and to the same embedding space; the training objective is to maximize the cosine similarity of similar image-text pairs while minimizing the similarity of non-matching pairs: (22) wherein, denotes the Euclidean norm, denotes the similarity between visual and linguistic features, after which a light-weight mapping network is employed to map the embedding vectors of the feature space into embedding vectors: (23) wherein, represents inputting The feature vectors generated by the CLIP model are further mapped to a function of the embedding space; each vector has the same dimension as the word embedding; and is connected to the caption embedding with the visual embedding: (24) By connecting the prefix and subtitle sequence during training The input is fed into a language model, which then predicts the title tag in an autoregressive manner under prefix conditions; the mapping component is trained using cross-entropy loss.

Citation Information

Patent Citations

  • Depth representation learning and fusion method based on multi-modal trajectory

    CN116956224A

  • Landslide automatic identification method and system based on visual large model, and computer equipment

    CN117671480A