Landslide disaster monitoring method and device and computer readable storage medium

By acquiring infrared and visible light image sequences and combining them with registration and deep prediction models, the problems of small coverage and low accuracy in landslide monitoring have been solved. This has enabled quantifiable characterization and risk warning of landslide movement processes, improving the accuracy and coverage of the monitoring system.

CN121921608APending Publication Date: 2026-04-24SOUTH SURVEYING & MAPPING INSTR
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTH SURVEYING & MAPPING INSTR
Filing Date
2025-12-30
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing landslide disaster monitoring methods suffer from problems such as sparse monitoring points, high costs, expensive equipment, great susceptibility to weather conditions, data heterogeneity, and unstable identification, resulting in small monitoring coverage and low accuracy.

Method used

Infrared and visible light image sequences are acquired, spatially aligned using a pre-defined registration algorithm, and combined with a feature extraction network and a deep prediction model to generate a probability map of the landslide area. The physical displacement and risk level are then calculated through threshold segmentation and consistency verification, achieving multi-source imaging fusion and dynamic identification.

Benefits of technology

It improves the accuracy and coverage of landslide monitoring, enables quantifiable characterization of landslide movement processes and timely risk warnings, and has value for engineering decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921608A_ABST
    Figure CN121921608A_ABST
Patent Text Reader

Abstract

The invention discloses a landslide disaster monitoring method and device and a computer readable storage medium. The method comprises the following steps: acquiring an infrared image sequence and a visible light image sequence; respectively inputting the infrared image sequence and the visible light image sequence into a preset feature extraction network to obtain infrared features and visible light features; fusing the infrared features and the visible light features to obtain fused features; generating a landslide region probability graph of fused features based on a preset deep prediction model; processing the landslide area probability graph based on a preset threshold segmentation method and a consistency verification method to obtain a target mask sequence; calculating a physical displacement based on the target mask sequence, and calculating a landslide rate, an acceleration and an area change rate based on the physical displacement; and determining a risk level based on the landslide rate, the acceleration, the area change rate and a preset risk index model, and performing risk early warning based on the risk level to complete landslide disaster monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a landslide disaster monitoring method, device, and computer-readable storage medium. Background Technology

[0002] Landslides are common and severe natural disasters in mountainous geological environments, characterized by their suddenness, insidiousness, and irreversibility. Many landslides rapidly destabilize and slide down at high speeds without any obvious warning signs, often causing numerous casualties and property losses in a short period. Furthermore, hidden factors such as underground weak surfaces, saturated layers, or fissures increase the difficulty of prevention, and once landslides occur, the resulting topographical and ecological damage is difficult to restore. Landslides often trigger secondary disasters such as barrier lakes, debris flows, traffic disruptions, and power and water supply outages, expanding the disaster area and significantly increasing rescue and recovery costs. Their long-term impact on the regional economy, infrastructure, and ecological environment cannot be ignored. Through multi-source monitoring, threshold setting, and prediction and early warning systems combined with numerical / statistical models, valuable time can be gained before a disaster occurs for evacuation and engineering protection, significantly reducing casualties and property losses.

[0003] Existing landslide monitoring methods, based on ground sensors, can provide high-precision displacement data, but monitoring points are typically sparse, deployment and maintenance costs are high, and it is difficult to cover large areas of landslides. Point cloud monitoring based on lidar or drones can accurately reflect terrain details, but equipment and operating costs are expensive, power consumption is high, and it is significantly affected by weather and visibility, making it unsuitable for long-term continuous deployment. Image methods based on visible light cameras are low-cost and flexible in deployment, but they are easily affected by external conditions such as changes in lighting, shadows, and rain and fog, leading to unstable landslide identification, blurred boundaries, and increased false alarms. Near-infrared imaging has good imaging stability under low light and haze conditions, but its weak texture features, low contrast, and insufficient structural details limit the accuracy of accurate target segmentation and image-based displacement calculation. In addition, cross-sensor data heterogeneity, insufficient real-time processing and automated identification capabilities, and the control of occlusion and scale effects in complex terrain all contribute to low landslide detection accuracy and small coverage. Summary of the Invention

[0004] This invention provides a landslide disaster monitoring method, device, and computer-readable storage medium to improve the accuracy of landslide monitoring while expanding the monitoring coverage.

[0005] To address the aforementioned technical problems, this invention provides a landslide disaster monitoring method, comprising: Infrared image sequences and visible light image sequences are acquired, and the infrared image sequences and visible light image sequences are spatially aligned based on a preset registration algorithm to obtain the target infrared sequence and the target visible light sequence; The target infrared sequence and the target visible light sequence are respectively input into a preset feature extraction network to obtain infrared features and visible light features; and the infrared features and visible light features are fused to obtain fused features; A landslide area probability map based on the fused features is generated based on a preset deep prediction model; and the landslide area probability map is processed based on a preset threshold segmentation method and a consistency verification method to obtain a target mask sequence. The physical displacement is calculated based on the target mask sequence, and the landslide rate, acceleration, and area change rate are calculated based on the physical displacement. The risk level is determined based on the landslide rate, acceleration, area change rate and the preset risk index model, and a risk warning is issued based on the risk level to complete the landslide disaster monitoring.

[0006] This invention acquires infrared and visible light image sequences and uses a pre-defined registration algorithm to achieve spatial alignment, ensuring geometric consistency of the two types of imaging at the pixel level. This effectively overcomes the limitations of single-modal imaging under conditions of illumination, haze, or weakened texture, laying a reliable foundation for subsequent feature fusion. Secondly, by inputting the target infrared and visible light sequences into a feature extraction network, deep features with modal difference perception capabilities can be obtained. Furthermore, a feature fusion strategy combines the stable imaging advantages of infrared with the detail representation capabilities of visible light, significantly improving the accuracy and robustness of landslide area identification. After generating a landslide area probability map based on the fused features using a deep prediction model, threshold segmentation and consistency verification are performed to effectively reduce noise response, false boundaries, and temporal jitter, resulting in a more coherent and realistic landslide target mask. Then, by calculating the physical displacement based on the target mask sequence and further obtaining the landslide rate, acceleration, and area change rate, the monitoring results are no longer limited to static image-level identification but achieve a quantifiable and interpretable representation of the landslide movement process, which is beneficial for capturing the accelerated evolution or abrupt change trends of landslides. Finally, based on the aforementioned dynamic indicators, a risk index model is constructed and risk levels are determined, enabling timely and graded early warning of landslide hazards. This gives the monitoring system engineering decision-making value and field application capabilities. It achieves a high degree of integration of multi-source imaging fusion, dynamic evolution identification, and physical-driven early warning, improving the accuracy of landslide monitoring while expanding its coverage.

[0007] Furthermore, the acquisition of infrared image sequences and visible light image sequences, and the spatial alignment of the infrared image sequences and visible light image sequences based on a preset registration algorithm to obtain target infrared sequences and target visible light sequences, includes: Infrared image sequences and visible light image sequences are acquired based on a preset synchronization signal; and grayscale normalization and noise filtering are performed on the infrared image sequences and visible light image sequences to obtain a set of target infrared and visible light images. Feature points are extracted from each image in the target infrared image and visible light image set based on a preset feature point detection algorithm, and the homography matrix is ​​calculated based on the feature points; Based on the homography matrix, the infrared image sequence and the visible light image sequence are spatially aligned to obtain the target infrared sequence and the target visible light sequence.

[0008] This invention implements synchronous triggering, grayscale normalization, and noise filtering in the image acquisition stage, and performs registration based on the homography matrix extracted from feature points. This ensures strict temporal and spatial consistency between infrared and visible light data. Temporal synchronization avoids motion distortion and registration errors caused by camera acquisition delays; preprocessing improves the comparability and matching rate of cross-modal features; furthermore, solving the homography matrix based on feature points lays the geometric foundation for subsequent pixel-by-pixel fusion, thereby improving the fusion effect and the accuracy and stability of downstream segmentation and displacement estimation.

[0009] Furthermore, the step of spatially aligning the infrared image sequence and the visible light image sequence based on the homography matrix to obtain the target infrared sequence and the target visible light sequence includes: The homography matrix is ​​validated for consistency using the RANSAC algorithm to obtain the target registration model. Based on the target registration model, the infrared image sequence and the visible light image sequence are spatially aligned to obtain the target infrared sequence and the target visible light sequence; wherein each frame of the target infrared sequence and the target visible light image sequence is spatially consistent.

[0010] This invention introduces the RANSAC algorithm during the registration process to verify the consistency of the homography matrix, effectively eliminating erroneous matches and outliers, resulting in a robust target registration model. This significantly reduces the risk of registration distortion or local misalignment caused by erroneous matches, ensuring a reliable pixel-level correspondence between the generated target infrared sequence and the target visible light sequence. This provides a solid geometric guarantee for subsequent pixel- or feature-based multimodal fusion, mask segmentation, and accurate displacement calculation, enhancing the anti-interference capability and reliability of the entire monitoring link.

[0011] Furthermore, the step of inputting the target infrared sequence and the target visible light sequence into a preset feature extraction network to obtain infrared features and visible light features respectively; and fusing the infrared features and visible light features to obtain fused features, includes: The target infrared sequence is input into a first feature extraction network to obtain infrared features; the first feature extraction network is trained using ResNet18 as the backbone network. The second feature extraction network of the target visible light sequence is used to obtain visible light features; the second feature extraction network is trained using ResNet50 as the backbone network. The infrared and visible light features are normalized to obtain the target infrared and visible light features. The infrared and visible light features of the target are fused based on a preset cross-modal alignment network to obtain fused features.

[0012] This invention employs deep backbones adapted to modal characteristics for feature extraction and achieves feature fusion through channel normalization and cross-modal alignment networks, balancing feature expressiveness and computational efficiency. It utilizes a lighter network to process low-contrast near-infrared signals to reduce computational overhead, while leveraging a deeper visible light backbone to extract rich semantic information. Furthermore, normalization and alignment mechanisms effectively bridge modal differences, ensuring that fused features retain their respective advantages while maintaining comparability. This improves segmentation accuracy and the reliability of temporal displacement tracking, while also facilitating real-time or near-real-time inference on edge devices.

[0013] Furthermore, the landslide area probability map is generated based on the preset deep prediction model, and the landslide area probability map is processed based on a preset threshold segmentation method and a consistency check method to obtain a target mask sequence, including: A probability map of the landslide area based on the fused features is generated based on a preset deep prediction model. A high-confidence mask is constructed based on a preset high threshold and the probability map of the landslide area; and a candidate mask is constructed based on a preset low threshold and the probability map of the landslide area. A preliminary mask is generated based on all connected regions of the high-confidence mask and the candidate mask; The initial mask is filtered and consistency checked to obtain the target mask sequence.

[0014] This invention generates a probability map based on a deep prediction model and employs a high- and low-threshold hysteresis strategy to construct high-confidence masks and candidate masks, which can suppress false alarms while ensuring detection sensitivity. By retaining low-confidence neighborhoods connected to high-confidence regions through dual-threshold hysteresis, it helps maintain target coherence and reduce false breakage detections. Subsequently, targeted filtering and consistency verification of the initial mask yields a mask sequence with more complete structure and higher confidence, providing more robust input for subsequent displacement estimation and risk calculation, thus improving the overall accuracy and stability of the system.

[0015] Furthermore, the step of filtering and consistency verification of the preliminary mask to obtain the target mask sequence includes: The preliminary mask is filtered and corrected based on a preset morphological algorithm to obtain the first mask; The second mask is obtained by performing least-squares polynomial fitting on the set of boundary points of each connected domain of the first mask; Perform regional consistency verification and temporal consistency verification on the second mask to obtain the target mask sequence.

[0016] This invention performs step-by-step optimization on the initial mask, including morphological filtering, boundary least-squares polynomial fitting, and region / temporal consistency verification. This effectively removes small-area noise, corrects pseudo-aliased boundaries, and suppresses temporal jitter. Morphological operations and connected component filtering improve the spatial integrity of the mask; furthermore, region and temporal consistency verification enhances the temporal stability of the detection, improving the reliability of the mask in spatial geometric description and its continuity in the time series. This provides high-quality basic data for accurate displacement calculation and dynamic analysis.

[0017] Furthermore, the step of calculating physical displacement based on the target mask sequence, and calculating landslide velocity, acceleration, and area change rate based on the physical displacement, includes: Feature extraction is performed on the target mask sequence to obtain feature points of each consecutive frame in the target mask sequence; Calculate the single-point displacement of feature points in any two consecutive frames, calculate the physical displacement based on the single-point displacement, and calculate the physical displacement in combination with preset calibration parameters and preset ground measurement scale factor. The landslide rate, acceleration, and area change rate are calculated based on the physical displacement.

[0018] This invention extracts feature points and calculates single-point displacement based on an optimized mask sequence. By combining camera calibration parameters and ground scaling factors, it converts these displacements into physical displacements. Based on this, it calculates velocity, acceleration, and area change rate, transforming image-level detection results into quantitative parameters required for engineering and decision-making. This quantification process not only enables risk assessment based on physical quantities for more objective judgment but also provides comparable data and historical traceability for tiered early warning, engineering response, and long-term trend analysis. Furthermore, by employing calibration and measured factor correction, measurement errors can be further reduced, enhancing the application value and credibility of monitoring results in practical engineering.

[0019] In a second aspect, the present invention provides a communication device including a module for performing the method.

[0020] Thirdly, the present invention provides a communication device, including a processor and an interface circuit, wherein the interface circuit is used to receive signals from other communication devices and transmit them to the processor or to send signals from the processor to other communication devices, and the processor is used to implement the method through logic circuits or execution code instructions.

[0021] Fourthly, the present invention provides a computer-readable storage medium storing a computer program or instructions that, when executed by a communication device, implement the method described above. Attached Figure Description

[0022] Figure 1 A schematic flowchart of a landslide disaster monitoring method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a multimodal deep fusion network provided in an embodiment of the present invention. Detailed Implementation

[0023] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0024] The terms "first" and "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.

[0025] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0026] Example 1 See Figure 1 , Figure 1 This is a schematic flowchart illustrating a landslide disaster monitoring method provided by an embodiment of the present invention. The embodiment of the present invention provides a landslide disaster monitoring method, including steps 101 to 105, as detailed below: Step 101: Acquire infrared image sequences and visible light image sequences, and spatially align the infrared image sequences and visible light image sequences based on a preset registration algorithm to obtain the target infrared sequence and the target visible light sequence; In this embodiment, the acquisition of infrared image sequences and visible light image sequences, and the spatial alignment of the infrared image sequences and visible light image sequences based on a preset registration algorithm to obtain target infrared sequences and target visible light sequences, includes: Infrared image sequences and visible light image sequences are acquired based on a preset synchronization signal; and grayscale normalization and noise filtering are performed on the infrared image sequences and visible light image sequences to obtain a set of target infrared and visible light images. Feature points are extracted from each image in the target infrared image and visible light image set based on a preset feature point detection algorithm, and the homography matrix is ​​calculated based on the feature points; Based on the homography matrix, the infrared image sequence and the visible light image sequence are spatially aligned to obtain the target infrared sequence and the target visible light sequence.

[0027] In this embodiment, a binocular monitoring unit consisting of a near-infrared camera and a visible light camera is installed in a landslide-prone area. The two cameras are mounted on a fixed bracket so that their fields of view overlap. The system completes the geometric calibration of the two cameras using a calibration plate to obtain the intrinsic parameter matrix. and extrinsic parameter matrix This ensures the accuracy of subsequent image registration.

[0028] In this embodiment, a sequence of frames is simultaneously acquired by a near-infrared camera and a visible light camera, and the resulting original sequences are denoted as follows: and ,in, Represents near-infrared images, Represents a visible light image. This represents a time frame. A precise timestamp is written to each frame on the edge computing nodes to ensure timing consistency.

[0029] In this embodiment, each frame of image is then first normalized to grayscale and subjected to noise suppression and adaptive histogram equalization to enhance the texture features of the near-infrared low-contrast region, thereby obtaining the target infrared image and visible light image set.

[0030] In this embodiment, key points are extracted and descriptors are calculated from the target infrared image and the visible light image set using a preset feature point detection and descriptor algorithm (such as SIFT, SURF, or ORB). A matcher is used to perform preliminary matching of floating-point descriptors using BFMatcher with L2 distance or FLANN, and preliminary matching of binary descriptors using BFMatcher with Hamming distance. Lowe's ratio test is applied to eliminate low-quality matches. The homography matrix between the near-infrared and visible light images is calculated based on the candidate matching point set. homography matrix This describes the spatial transformation relationship from pixel coordinates in a visible light image to pixel coordinates in a near-infrared image: (1) in, These are the pixel coordinates of the visible light image. These are the corresponding coordinates in the near-infrared image.

[0031] In this embodiment, the step of spatially aligning the infrared image sequence and the visible light image sequence based on the homography matrix to obtain the target infrared sequence and the target visible light sequence includes: The homography matrix is ​​validated for consistency using the RANSAC algorithm to obtain the target registration model. Based on the target registration model, the infrared image sequence and the visible light image sequence are spatially aligned to obtain the target infrared sequence and the target visible light sequence; wherein each frame of the target infrared sequence and the target visible light image sequence is spatially consistent.

[0032] In this embodiment, RANSAC is used to robustly estimate the homography matrix H, resulting in a robust registration model. Mismatched points are then eliminated to obtain the target registration model. (2) in, It is the original visible light image; It was through After transformation, the visible light image, which is spatially aligned with the near-infrared image, becomes the target visible light image sequence. The two can then undergo pixel-by-pixel feature fusion.

[0033] In this embodiment, the RANSAC algorithm is introduced during the registration process to verify the consistency of the homography matrix. This effectively eliminates erroneous matching points and outliers, resulting in a robust target registration model. This significantly reduces the risk of registration distortion or local misalignment caused by erroneous matching, thereby ensuring a reliable pixel-level correspondence between the generated target infrared sequence and the target visible light sequence. This provides a solid geometric guarantee for subsequent pixel- or feature-based multimodal fusion, mask segmentation, and accurate displacement calculation, enhancing the anti-interference capability and reliability of the entire monitoring link.

[0034] In this embodiment, synchronous triggering, grayscale normalization, and noise filtering are implemented in the image acquisition stage. Registration is then performed by solving the homography matrix based on feature point extraction, ensuring strict temporal and spatial consistency between the infrared and visible light data streams. Temporal synchronization avoids motion distortion and registration errors caused by camera acquisition delays; preprocessing improves the comparability and matching rate of cross-modal features; furthermore, solving the homography matrix based on feature points lays the geometric foundation for subsequent pixel-by-pixel fusion, thereby improving the fusion effect and the accuracy and stability of downstream segmentation and displacement estimation.

[0035] Step 102: Input the target infrared sequence and the target visible light sequence into a preset feature extraction network to obtain infrared features and visible light features; and fuse the infrared features and visible light features to obtain fused features; In this embodiment, the step of inputting the target infrared sequence and the target visible light sequence into a preset feature extraction network to obtain infrared features and visible light features, and then fusing the infrared features and visible light features to obtain fused features, includes: The target infrared sequence is input into a first feature extraction network to obtain infrared features; the first feature extraction network is trained using ResNet18 as the backbone network. The second feature extraction network of the target visible light sequence is used to obtain visible light features; the second feature extraction network is trained using ResNet50 as the backbone network. The infrared and visible light features are normalized to obtain the target infrared and visible light features. The infrared and visible light features of the target are fused based on a preset cross-modal alignment network to obtain fused features.

[0036] Please refer to Figure 2 , Figure 2 This is a schematic diagram of a multimodal deep fusion network provided in an embodiment of the present invention.

[0037] In this embodiment, the registered target visible light sequence is first... With the target infrared sequence The two independent feature extraction networks, VisibleNet and NIRNet, are input frame by frame into the multimodal deep fusion network. The first feature extraction network, VisibleNet, uses a convolutional network with ResNet-50 as its backbone, while the second feature extraction network, NIRNet, uses a convolutional network with ResNet-18 as its backbone, to obtain a multi-scale feature set. and .

[0038] In this embodiment, the first feature extraction network is VisibleNet (visible light feature branch): visible light images mainly provide texture, color, and detail features of the land surface, which are crucial for identifying soil, vegetation, and other surface elements in landslide areas. To efficiently extract this information, this invention selects ResNet50 as the backbone network. This network can learn features more deeply through its residual structure, avoiding the gradient vanishing problem that may occur in deep networks. The input size is... The image is a color visible light image (RGB). Multiple convolutional layers are used to extract features from the image, outputting multi-layer feature maps. ,in This represents different convolutional layers; each convolutional layer extracts feature information at different scales, such as edges, textures, and shapes, resulting in a set of extracted feature maps. .

[0039] In this embodiment, the second feature extraction network, NIRNet (Near-Infrared Feature Branch), utilizes near-infrared images to provide information related to surface materials and soil reflectance, which is particularly important for identifying key soil and rock types in landslide monitoring. Because near-infrared images have weaker detail and contrast, this invention employs the ResNet18 network, which has fewer parameters than ResNet50 and can efficiently extract low-contrast features. The input size is... Near-infrared single-channel images, similar to VisibleNet, use the ResNet18 network for convolutional feature extraction, outputting feature maps. The feature map set extracted has the same dimension as the feature map of the visible light branch. .

[0040] In this embodiment, to eliminate the channel dimension differences between modalities, a 1×1 convolution is first applied at each scale. and Projecting onto the same channel C1, then performing batch normalization (BatchNorm) and ReLU activation on both to stabilize the numerical distribution and suppress numerical shifts due to modal differences. Cross-modal alignment employs a channel-level adaptive weighting mechanism: global average pooling (GAP) is first performed on each scale feature after projection to obtain the channel description vector. (3) (4) In this embodiment, the pooled vector is mapped through a multilayer perceptron (MLP) to generate the modality weight vector. and This indicates the degree of contribution of each component to the final output: (5) Based on the calculated weights and Weighted fusion of feature maps from the two modalities: (6) In this embodiment, to ensure information flow and gradient propagation, residual skipping and attention recalibration are introduced. Enhancement is performed, and finally, the fused features at each scale are fed into the temporal enhancement module ConvLSTM or the deep decoder through cross-layer skip connections to generate subsequent probability maps.

[0041] In this embodiment, in landslide monitoring, images not only contain spatial features, such as boundaries and shapes, but also temporal dynamic features, such as displacement and movement trajectories. To capture these spatiotemporal features, ConvLSTM is used to enhance temporal information, especially between consecutive frames, to capture the dynamic evolution of landslides. ConvLSTM is a network module that combines convolutional operations and Long Short-Term Memory (LSTM), which can effectively capture the spatiotemporal dependencies in image sequences. In this module, the fused feature maps... The data is fed into a ConvLSTM, which utilizes its memory properties to handle temporal slippage variations, by inputting feature maps from multimodal fusion. Then, output the temporally enhanced feature map. It contains information about the dynamic changes between consecutive frames.

[0042] (7) In this embodiment, feature extraction is performed using deep backbones adapted to modal characteristics, and feature fusion is achieved through channel normalization and cross-modal alignment networks, balancing feature expressiveness and computational efficiency. It utilizes a lighter network to process low-contrast near-infrared signals to reduce computational overhead, while leveraging a deeper visible light backbone to extract rich semantic information. Furthermore, normalization and alignment mechanisms effectively bridge modal differences, ensuring that fused features retain their respective advantages while maintaining comparability. This improves segmentation accuracy and the reliability of temporal displacement tracking, while also facilitating real-time or near-real-time inference on edge devices.

[0043] Step 103: Generate a landslide area probability map based on the fused features using a preset deep prediction model; and process the landslide area probability map using a preset threshold segmentation method and a consistency verification method to obtain a target mask sequence; In this embodiment, the landslide area probability map based on the fused features is generated using a preset deep prediction model; and the landslide area probability map is processed using a preset threshold segmentation method and a consistency check method to obtain a target mask sequence, including: A probability map of the landslide area based on the fused features is generated based on a preset deep prediction model. A high-confidence mask is constructed based on a preset high threshold and the probability map of the landslide area; and a candidate mask is constructed based on a preset low threshold and the probability map of the landslide area. A preliminary mask is generated based on all connected regions of the high-confidence mask and the candidate mask; The initial mask is filtered and consistency checked to obtain the target mask sequence.

[0044] In this embodiment, the fusion feature or First, the data is fed into a pre-defined deep prediction model, which is a pixel-level segmentation network using the U-Net++ architecture. After upsampling and skip connections, the network outputs a landslide probability map for each pixel through Sigmoid activation. To generate a stable sequence of target masks from the probability map.

[0045] In this embodiment, the deep prediction model uses the U-Net++ architecture, where each layer receives information from previous encoding layers (including low-level and high-level features) to capture more spatial details; the input feature map is upsampled to progressively restore the spatial resolution of the image; skip connections are used to concatenate the feature maps of lower layers with those of the current layer to preserve edge information and local details; the decoded feature map is then transformed into a binary landslide region prediction map using a sigmoid activation function. The value of each pixel represents the probability of whether that location belongs to a landslide region.

[0046] (8) The final output is a probability map of the landslide area. It can be used in subsequent landslide risk analysis and early warning systems.

[0047] In this embodiment, a dual threshold hysteresis strategy is adopted: a high threshold is set. With low threshold , exemplary =0.8, =0, construct a high-confidence mask based on a preset high threshold. and candidate masks And use a lag join rule to retain all connections with Connected The area is used to form a preliminary mask. .

[0048] In this embodiment, the step of filtering and consistency verification of the preliminary mask to obtain the target mask sequence includes: The preliminary mask is filtered and corrected based on a preset morphological algorithm to obtain the first mask; The second mask is obtained by performing least-squares polynomial fitting on the set of boundary points of each connected domain of the first mask; Perform regional consistency verification and temporal consistency verification on the second mask to obtain the target mask sequence.

[0049] In this embodiment, the following is performed on Morphological filtering is implemented, first opening operation for noise reduction, then closing operation for hole filling. The structuring element can be a circle or square with radius r and connected component analysis: pseudo-detection regions with an area less than the threshold Amin are removed by connecting component annotation, and hole filling and minimum bounding polygon correction are performed on the retained regions to enhance region integrity.

[0050] In this embodiment, to further improve the quality of spatial boundaries, least-squares polynomial or spline curve fitting is performed on the boundary points of each connected domain to obtain smooth boundaries, and the fitted curve is used to replace the original contour to reduce curvature abrupt changes. For the time dimension, temporal consistency verification is introduced: the mask of the previous frame is mapped to the current frame through the estimated motion field, and the IoU with the current mask is calculated. The motion field is obtained through dense optical flow or feature tracking.

[0051] In this embodiment, if the IoU is less than a set threshold θ, a short-time window median filter or a confidence low-pass filter is applied to the connected component to suppress noise; simultaneously, a median time filter with a width of w is performed on the mask sequence to smooth out instantaneous jitter. Finally, after the above adaptive thresholding, morphological filtering, connected component constraint, boundary smoothing, and temporal consistency verification processes, the output target mask sequence meets the requirements of spatial integrity, boundary continuity, and temporal stability. It also records the mask confidence and quality indicators for each frame, such as average probability, reprojection error, and connected component stability, for subsequent displacement estimation and risk assessment modules to perform weighted fusion and decision support.

[0052] In this embodiment, the initial mask undergoes progressive optimization processes, including morphological filtering, boundary least-squares polynomial fitting, and region / temporal consistency verification. This effectively removes small-area noise, corrects pseudo-aliased boundaries, and suppresses temporal jitter. Morphological operations and connected component filtering enhance the spatial integrity of the mask; furthermore, region and temporal consistency verification improves the temporal stability of the detection, increasing the reliability of the mask in spatial geometric description and its continuity in the time series. This provides high-quality foundational data for accurate displacement calculation and dynamic analysis.

[0053] In this embodiment, a probability map is generated based on a deep prediction model, and a high-confidence mask and candidate masks are constructed using a high- and low-threshold hysteresis strategy. This approach can suppress false alarms while ensuring detection sensitivity. By retaining the low-confidence neighborhood connected to the high-confidence region through dual-threshold hysteresis, the continuity of the target is maintained and the false breakage is reduced. Subsequently, targeted filtering and consistency verification of the initial mask yields a mask sequence with a more complete structure and higher confidence, providing a more robust input for subsequent displacement estimation and risk calculation, thereby improving the overall accuracy and stability of the system.

[0054] Step 104: Calculate the physical displacement based on the target mask sequence, and calculate the landslide rate, acceleration, and area change rate based on the physical displacement; In this embodiment, the step of calculating physical displacement based on the target mask sequence, and calculating landslide velocity, acceleration, and area change rate based on the physical displacement, includes: Feature extraction is performed on the target mask sequence to obtain feature points of each consecutive frame in the target mask sequence; Calculate the single-point displacement of feature points in any two consecutive frames, calculate the physical displacement based on the single-point displacement, and calculate the physical displacement in combination with preset calibration parameters and preset ground measurement scale factor. The landslide rate, acceleration, and area change rate are calculated based on the physical displacement.

[0055] In this embodiment, a landslide area mask is used for continuous time intervals. The motion trajectory of feature points within a region is calculated by optical flow tracing. Let the feature points in two consecutive frames be... and Then the displacement of a single point is calculated as follows: (9) The physical displacement is calculated using the regional average displacement: (10) Where N is the total number of single-point displacements.

[0056] In this embodiment, camera calibration parameters are combined with ground-measured scale factors. Convert pixel displacement into physical displacement: This enables displacement monitoring of the landslide area in the actual physical space.

[0057] In this embodiment, based on the displacement sequence Calculate the landslide's velocity and acceleration: This invention extracts feature points and calculates single-point displacement based on an optimized mask sequence. By combining camera calibration parameters and ground scaling factors, it converts these displacements into physical displacements. Based on this, it calculates velocity, acceleration, and area change rate, transforming image-level detection results into quantitative parameters required for engineering and decision-making. This quantification process not only enables risk assessment based on physical quantities for more objective judgment but also provides comparable data and historical traceability for tiered early warning, engineering response, and long-term trend analysis. Furthermore, by employing calibration and measured factor correction, measurement errors can be further reduced, enhancing the application value and credibility of monitoring results in practical engineering.

[0058] Step 105: Determine the risk level based on the landslide rate, acceleration, area change rate and the preset risk index model, and conduct risk warning based on the risk level to complete landslide disaster monitoring.

[0059] In this embodiment, based on the landslide rate, acceleration, and area change rate of the landslide area calculated in the previous steps, the system uses a preset risk index model to quantitatively assess the landslide risk.

[0060] First, by calculating the rate of change of displacement and the rate of change of acceleration over continuous time steps, and combining this with the rate of change of area of ​​the landslide region, the dynamic trend of landslide change is determined. Specifically, the landslide velocity v(t) and acceleration a(t) are used to reflect the rate of change of landslide activity, while the area change rate A characterizes the expansion of the landslide activity range.

[0061] In this embodiment, a landslide risk index model is established based on landslide rate, acceleration, and area change rate: in, This represents the rate of change in the landslide area.

[0062] In this embodiment, a preset risk index model R is used to determine the risk level and output a risk index value R(t). This value can be mapped to multiple risk levels, such as low risk, medium risk, and high risk. The risk index model can be trained and optimized based on historical data and actual landslide disaster cases through machine learning or statistical methods.

[0063] After determining the risk level, the system uses the set threshold. The system automatically determines the current landslide risk level and triggers corresponding early warning mechanisms based on that level. For example, when the risk index R(t) exceeds a set high-risk threshold... When the risk index is less than or equal to the high-risk threshold, the system will automatically trigger a high-risk warning and send the warning information to the monitoring center or mobile terminal in real time via the communication module; when the risk index is less than or equal to the high-risk threshold... Greater than the low risk threshold When the risk index is in the medium range, a medium-risk warning is triggered, reminding relevant personnel to be aware of potential disasters; when the risk index is below or equal to the set low-risk threshold... At this time, the system can operate normally or enter observation mode.

[0064] In this embodiment, not only can the risk level of landslide disasters be assessed in real time, but early warning information can also be issued in a timely manner based on the risk level, providing data support and decision-making basis for the prevention, emergency response and rescue work of landslide disasters, and effectively improving the real-time and accuracy of landslide disaster monitoring.

[0065] In this embodiment, by acquiring infrared and visible light image sequences and using a preset registration algorithm to achieve spatial alignment, the geometric consistency of the two types of imaging at the pixel level is ensured. This effectively overcomes the limitations of single-modal imaging under conditions of illumination, haze, or weakened texture, laying a reliable foundation for subsequent feature fusion. Secondly, by inputting the target infrared and visible light sequences into a feature extraction network, deep features with modal difference perception capabilities can be obtained. Furthermore, a feature fusion strategy combines the stable imaging advantages of infrared with the detail representation capabilities of visible light, significantly improving the accuracy and robustness of landslide area identification. After generating a landslide area probability map based on the fused features using a deep prediction model, threshold segmentation and consistency verification are performed to effectively reduce noise response, false boundaries, and temporal jitter, resulting in a more coherent and realistic landslide target mask. Then, by calculating the physical displacement based on the target mask sequence and further obtaining the landslide rate, acceleration, and area change rate, the monitoring results are no longer limited to static image-level identification but achieve a quantifiable and interpretable representation of the landslide movement process, which is beneficial for capturing the accelerated evolution or abrupt change trends of landslides. Finally, based on the aforementioned dynamic indicators, a risk index model is constructed and risk levels are determined, enabling timely and graded early warning of landslide hazards. This gives the monitoring system engineering decision-making value and field application capabilities. It achieves a high degree of integration of multi-source imaging fusion, dynamic evolution identification, and physical-driven early warning, improving the accuracy of landslide monitoring while expanding its coverage.

[0066] This invention also provides a communication device, including a module for performing the method.

[0067] This invention also provides a communication device, including a processor and an interface circuit. The interface circuit is used to receive signals from other communication devices and transmit them to the processor, or to send signals from the processor to other communication devices. The processor is used to implement the method through logic circuits or execution code instructions.

[0068] In this embodiment of the invention, a terminal device is also provided, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the above-described landslide disaster monitoring method.

[0069] In this embodiment of the invention, a computer-readable storage medium is also provided, which includes a stored computer program, wherein the computer program controls the device where the computer-readable storage medium is located to execute the above-described landslide disaster monitoring method when it is running.

[0070] For example, a computer program can be divided into one or more modules, one or more of which are stored in memory and executed by a processor to perform the present invention. The one or more modules can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in a terminal device.

[0071] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor, memory, and display. Those skilled in the art will understand that the above components are merely examples of terminal devices and do not constitute a limitation on the terminal device. It may include more or fewer components, or combinations of certain components, or different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.

[0072] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device through various interfaces and lines.

[0073] Memory can be used to store computer programs and / or modules. The processor implements various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, application programs required for at least one function (such as sound playback, text conversion, etc.), etc.; the data storage area can store data created based on the use of the mobile phone (such as audio data, text message data, etc.). In addition, memory can include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital cards (SD cards), flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0074] In this invention, if the landslide disaster monitoring module is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. Those skilled in the art can understand and implement this invention without any inventive effort.

[0075] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A method for monitoring landslide disasters, characterized in that, include: Infrared image sequences and visible light image sequences are acquired, and the infrared image sequences and visible light image sequences are spatially aligned based on a preset registration algorithm to obtain the target infrared sequence and the target visible light sequence; The target infrared sequence and the target visible light sequence are respectively input into a preset feature extraction network to obtain infrared features and visible light features; The infrared and visible light features are then fused to obtain a fused feature. A probability map of the landslide area based on the fused features is generated based on a preset deep prediction model. The probability map of the landslide area is processed based on a preset threshold segmentation method and a consistency check method to obtain a target mask sequence; The physical displacement is calculated based on the target mask sequence, and the landslide rate, acceleration, and area change rate are calculated based on the physical displacement. The risk level is determined based on the landslide rate, acceleration, area change rate and the preset risk index model, and a risk warning is issued based on the risk level to complete the landslide disaster monitoring.

2. The landslide disaster monitoring method as described in claim 1, characterized in that, The acquisition of infrared and visible light image sequences, and spatial alignment of the infrared and visible light image sequences based on a preset registration algorithm to obtain target infrared and target visible light sequences, includes: Infrared image sequences and visible light image sequences are acquired based on a preset synchronization signal; and grayscale normalization and noise filtering are performed on the infrared image sequences and visible light image sequences to obtain a set of target infrared and visible light images. Feature points are extracted from each image in the target infrared image and visible light image set based on a preset feature point detection algorithm, and the homography matrix is ​​calculated based on the feature points; Based on the homography matrix, the infrared image sequence and the visible light image sequence are spatially aligned to obtain the target infrared sequence and the target visible light sequence.

3. The landslide disaster monitoring method as described in claim 2, characterized in that, The step of spatially aligning the infrared image sequence and the visible light image sequence based on the homography matrix to obtain the target infrared sequence and the target visible light sequence includes: The homography matrix is ​​validated for consistency using the RANSAC algorithm to obtain the target registration model. Based on the target registration model, the infrared image sequence and the visible light image sequence are spatially aligned to obtain the target infrared sequence and the target visible light sequence; wherein each frame of the target infrared sequence and the target visible light image sequence is spatially consistent.

4. The landslide disaster monitoring method as described in claim 3, characterized in that, The target infrared sequence and the target visible light sequence are respectively input into a preset feature extraction network to obtain infrared features and visible light features; The infrared and visible light features are then fused to obtain fused features, including: The target infrared sequence is input into a first feature extraction network to obtain infrared features; the first feature extraction network is trained using ResNet18 as the backbone network. The second feature extraction network of the target visible light sequence is used to obtain visible light features; the second feature extraction network is trained using ResNet50 as the backbone network. The infrared and visible light features are normalized to obtain the target infrared and visible light features. The infrared and visible light features of the target are fused based on a preset cross-modal alignment network to obtain fused features.

5. A landslide disaster monitoring method as described in claim 4, characterized in that, The landslide area probability map based on the preset deep prediction model is generated by the fused features. The probability map of the landslide area is processed based on a preset threshold segmentation method and a consistency check method to obtain a target mask sequence, including: A probability map of the landslide area based on the fused features is generated based on a preset deep prediction model. A high-confidence mask is constructed based on a preset high threshold and the probability map of the landslide area; and a candidate mask is constructed based on a preset low threshold and the probability map of the landslide area. A preliminary mask is generated based on all connected regions of the high-confidence mask and the candidate mask; The initial mask is filtered and consistency checked to obtain the target mask sequence.

6. The landslide disaster monitoring method as described in claim 5, characterized in that, The step of filtering and consistency verification of the preliminary mask to obtain the target mask sequence includes: The preliminary mask is filtered and corrected based on a preset morphological algorithm to obtain the first mask; The second mask is obtained by performing least-squares polynomial fitting on the set of boundary points of each connected domain of the first mask; Perform regional consistency verification and temporal consistency verification on the second mask to obtain the target mask sequence.

7. A landslide disaster monitoring method as described in claim 6, characterized in that, The step of calculating physical displacement based on the target mask sequence, and calculating landslide velocity, acceleration, and area change rate based on the physical displacement, includes: Feature extraction is performed on the target mask sequence to obtain feature points of each consecutive frame in the target mask sequence; Calculate the single-point displacement of feature points in any two consecutive frames, calculate the physical displacement based on the single-point displacement, and calculate the physical displacement in combination with preset calibration parameters and preset ground measurement scale factor. The landslide rate, acceleration, and area change rate are calculated based on the physical displacement.

8. A communication device, characterized in that, Includes a module for performing the method as described in any one of claims 1 to 7.

9. A communication device, characterized in that, The device includes a processor and an interface circuit, wherein the interface circuit is used to receive signals from other communication devices and transmit them to the processor or to send signals from the processor to other communication devices, and the processor is used to implement the method as described in any one of claims 1 to 7 through logic circuits or execution code instructions.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program or instructions, which, when executed by a communication device, implement the method as described in any one of claims 1 to 7.