Method and system for dynamic monitoring and evaluation of soil body collapse loss in a bank collapse test

By using a deep convolutional neural network-based bank collapse monitoring model, combined with the physical scale features of image sequences and scale perturbation perception enhancement, the accuracy and quantification issues of bank collapse experimental monitoring in existing technologies have been solved, achieving high-precision collapse loss assessment.

CN121904485BActive Publication Date: 2026-05-29NANJING HYDRAULIC RES INST

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING HYDRAULIC RES INST
Filing Date
2026-03-19
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing bank collapse test monitoring technologies are unable to achieve full-process, high-precision dynamic monitoring, especially in complex scenarios where they cannot effectively identify minute cracks and locate ambiguous boundaries, and lack quantitative means for three-dimensional volume loss.

Method used

A bank collapse monitoring model based on deep convolutional neural networks is adopted. The physical scale features of the image sequence are acquired to enhance the perception of scale perturbation. Multi-scale bank collapse feature maps are extracted using a feature extraction network and a dual-branch decoupling head. The semantic probability map of the region and the boundary contour probability map are predicted in a coordinated manner, and the collapse loss assessment index is calculated.

Benefits of technology

It improves the accuracy and full-process perception capability of bank collapse monitoring, solves the problems of difficulty in identifying micro-cracks, inaccurate positioning of fuzzy boundaries and lack of volume quantification, and achieves high-precision assessment of collapse loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904485B_ABST
    Figure CN121904485B_ABST
Patent Text Reader

Abstract

The application discloses a kind of bank collapse test soil body collapse loss dynamic monitoring evaluation method and system, the method includes: obtaining bank collapse test image sequence, calculate physical scale feature to perform adaptive scale disturbance perception enhancement;Enhanced data is input into the improved network embedded with C2f boundary attention unit, high-frequency edge feature extraction is strengthened using Sobel operator;Through the double-branch decoupling head, the regional semantic probability graph and the boundary contour probability graph are cooperatively predicted;Based on the double-branch output, the collapse area and volume index are calculated.The application solves the problems of difficult micro-fissure recognition, inaccurate fuzzy boundary positioning and missing volume quantification through physical scale perception enhancement and gradient consistency constraint soft labeling training, and improves the precision and whole-process perception ability of bank collapse monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of geological disaster monitoring technology, and in particular to a method and system for dynamic monitoring and evaluation of soil collapse loss in bank collapse tests. Background Technology

[0002] Bank collapse is a common natural disaster in the evolution of rivers in alluvial plains. It not only causes soil erosion on bank slopes, threatening the safety of dikes, but also alters river boundary conditions and affects river stability. Conducting generalized model experiments is an important means of studying bank collapse mechanisms. By simulating the soil failure process under different hydrodynamic conditions in a flume, theoretical support can be provided for bank collapse early warning. During the experiment, real-time and accurate acquisition of the time, location, morphological evolution, and soil loss of bank collapse is crucial for revealing the mechanical mechanisms of bank collapse.

[0003] Currently, monitoring methods for bank collapse tests mainly rely on manual observation, contact sensors (such as displacement gauges), and conventional machine vision technology. Manual observation is inefficient and struggles to capture transient collapses; contact sensors can only monitor single-point deformation and are insufficient to reflect area damage; early visual monitoring methods were mostly based on background subtraction or traditional edge detection algorithms, or directly applied general object detection models (such as the standard YOLO series) to identify the bank collapse area. For example, some studies have used temporal difference methods to extract the bank collapse contour, or used UAV aerial imagery combined with semantic segmentation networks to assess bank slope stability.

[0004] However, existing monitoring technologies still have significant limitations when facing complex bank collapse test scenarios. Bank collapse processes span a large scale, from initial micro-cracks (millimeter-level) to later overall landslides (meter-level). Conventional models easily lose high-frequency features of micro-cracks during downsampling, resulting in low early recognition rates. Bank collapse boundaries often exhibit non-rigid, fuzzy transition characteristics (such as soil-water mixing zones), and existing hard-annotation training methods struggle to adapt to this uncertainty, leading to coarse boundary localization. Furthermore, most vision-based solutions can only output two-dimensional areas, lacking effective means to estimate three-dimensional volume loss, making it difficult to meet the depth requirements for quantitative assessment of collapse. These problems prevent existing methods from achieving full-process, high-precision dynamic monitoring of bank collapses. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for dynamic monitoring and evaluation of soil collapse loss in bank collapse tests, so as to solve the above-mentioned problems existing in the prior art.

[0006] According to one aspect of this application, a method for dynamic monitoring and evaluation of soil collapse loss in bank collapse tests includes:

[0007] Acquire image sequences of the bank collapse test process and calculate the physical scale characteristics of the bank collapse area in the image sequences;

[0008] Based on physical scale characteristics, scale perturbation sensing enhancement processing is performed on the image sequence to obtain enhanced image data;

[0009] The enhanced image data is input into a pre-trained bank collapse monitoring model. The feature extraction network in the bank collapse monitoring model is used to extract multi-scale bank collapse feature maps. During the feature extraction process, the high-frequency boundary information in the multi-scale bank collapse feature maps is weighted and enhanced.

[0010] Using the dual-branch decoupling head in the bank collapse monitoring model, based on the weighted and enhanced multi-scale bank collapse feature map, we collaboratively predict and output a regional semantic probability map reflecting the extent of soil collapse, as well as a boundary contour probability map reflecting cracks and erosion edges.

[0011] Based on the regional semantic probability map and the boundary contour probability map, the collapse loss assessment index is calculated.

[0012] According to another aspect of this application, a dynamic monitoring and evaluation system for soil collapse loss in bank collapse tests includes:

[0013] The data acquisition module is used to acquire image sequences of the bank collapse test process and calculate the physical scale characteristics of the bank collapse area in the image sequence;

[0014] The enhancement processing module is used to perform scale perturbation-aware enhancement processing on image sequences based on physical scale features;

[0015] The model processing module is used to input the enhanced image data into the pre-trained bank collapse monitoring model, and use the feature extraction network and the two-branch decoupling head in the model to jointly predict and output the regional semantic probability map and the boundary contour probability map.

[0016] The assessment module is used to calculate collapse loss assessment indicators based on the regional semantic probability map and the boundary contour probability map.

[0017] Beneficial effects: This invention solves the problems of difficulty in identifying micro-cracks, inaccurate positioning of fuzzy boundaries, and lack of volume quantization by using soft annotation training with enhanced physical scale perception and gradient consistency constraints, thereby improving the accuracy and full-process perception capability of bank collapse monitoring. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the overall process of the dynamic monitoring and evaluation method for soil collapse loss in bank collapse tests provided in the embodiments of this application.

[0019] Figure 2 This is a schematic diagram illustrating the physical scale characteristics of a bank collapse area in a computational image sequence provided in this application embodiment.

[0020] Figure 3This is a schematic diagram of the process for performing scale perturbation perception enhancement processing on image sequences based on physical scale features, provided in an embodiment of this application.

[0021] Figure 4 This is a schematic diagram of the process of dynamically generating boundary uncertainty soft weights based on the physical geometric properties of the collapsed bank boundary, using isotropic soft annotation based on geometric complexity provided in the embodiments of this application.

[0022] Figure 5 This is a graph showing the model pass rate provided in the embodiments of this application. Detailed Implementation

[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0024] For ease of understanding, the relationships between some terms used in this invention are explained below:

[0025] C2f module and Cross Stage Partial Darknet module: In this invention, the two have the same meaning. C2f is an abbreviation for Cross Stage Partial Darknet 2-stage FPN, which is an improved cross stage partial darknet structure.

[0026] Backbone and Feature Extraction Network: In this invention, the two have the same meaning, referring to the backbone network in a deep learning model that is responsible for extracting features from the input image;

[0027] FPN / PANet and Shared Feature Layer: FPN (Feature Pyramid Network) and PANet (Path Aggregation Network) are specific network structures for implementing multi-scale feature fusion, and are used as shared feature layers in the dual-branch decoupling head of this invention.

[0028] Example 1 provides an overall flowchart for a method of dynamic monitoring and evaluation of soil collapse loss in a bank collapse test, as follows: Figure 1 As shown, this addresses the technical problems of low efficiency of manual observation in bank collapse experiments, insufficient accuracy of conventional visual models in identifying micro-cracks, and difficulty in quantifying collapse volume.

[0029] In this embodiment, the method relies on a generalized bank collapse model test system, which typically includes a model flume, a water level control subsystem, a multi-angle high-definition camera acquisition subsystem, and a high-performance computing server. The water level control subsystem simulates the rise and fall of river water levels to induce bank collapse; the multi-angle high-definition camera acquisition subsystem records real-time changes in bank morphology during the test; and the high-performance computing server deploys the bank collapse monitoring model of this embodiment for real-time or offline processing of the acquired image data.

[0030] Step 101: Obtain the image sequence of the bank collapse test process and calculate the physical scale characteristics of the bank collapse area in the image sequence.

[0031] In this embodiment, acquiring the image sequence of the bank collapse test process specifically refers to continuously acquiring RGB image frames of the bank slope soil at preset time intervals using image acquisition equipment positioned on the side or above the model flume. The acquisition equipment can specifically employ a high-resolution industrial camera or a high-definition surveillance camera, and the acquisition frequency can be set to 25 frames per second or higher to ensure that the transient process of the collapse is captured. The acquired image sequence can reflect the gradual failure process of the bank slope under the action of water level circulation, including stages such as crack generation, development, penetration, and soil sliding.

[0032] Building upon this, calculating the physical scale features of the landslide area in the image sequence involves performing preliminary geometric morphological analysis on each input frame or keyframe image to quantify the spatial scale of the current landslide event. Specifically, physical scale features are quantitative indicators used to describe the proportion of the landslide area in the image and the complexity of its edges. For example, traditional image processing methods such as thresholding or background subtraction can be used to roughly locate the landslide area and calculate the proportion of pixels in the landslide area to the total number of pixels in the image, as well as the ratio of the length of the landslide shoreline to the image's side length. These two ratios constitute the physical scale features required for subsequent steps. Calculating these features helps to determine whether the current landslide event is a large-scale overall collapse or a small-scale local erosion, providing adaptive guidance for subsequent data augmentation.

[0033] Step 102: Based on physical scale characteristics, scale perturbation sensing enhancement processing is performed on the image sequence to obtain enhanced image data.

[0034] In this embodiment, this step addresses the problem of limited sample data and uneven scale distribution of collapsed bank landslides. Because real-world collapsed bank landslide experiments are costly and time-consuming, the number of effective samples obtained is often limited, and different types of collapsed bank landslides (such as large-area collapses and the development of micro-cracks) differ significantly in visual scale. The scale perturbation perception enhancement processing does not involve blindly performing random scaling, but rather, based on the physical scale characteristics calculated in step 101, it performs targeted geometric transformations on the image.

[0035] Specifically, if the physical scale characteristics indicate that the current situation mainly involves the development of small cracks (small-scale features), the system will automatically perform a magnification operation, using bilinear interpolation or bicubic interpolation algorithms to enlarge the crack region, thereby increasing the model's learning weight for subtle features. Conversely, if a large-scale collapse occurs (large-scale features), a reduction operation is used to ensure that the model can capture the complete collapse morphology within a single field of view. Furthermore, this enhancement process can also include perspective transformation, simulating the imaging effects at different camera heights and tilt angles to expand the diversity of the data. Through targeted enhancement, the robustness of the model to bank collapse events of different scales can be significantly improved, and the enhanced image data can be used as input for training or inference.

[0036] Step 103: Input the enhanced image data into the pre-trained bank collapse monitoring model, use the feature extraction network in the bank collapse monitoring model to extract multi-scale bank collapse feature maps, and in the feature extraction process, perform weighted enhancement on the high-frequency boundary information in the multi-scale bank collapse feature maps.

[0037] In this embodiment, the bank collapse monitoring model is built upon a deep convolutional neural network, preferably using the improved YOLOv8 (You Only Look Once version 8) architecture as its base framework. The feature extraction network (Backbone) is responsible for extracting multi-level features from the input image, ranging from shallow to deep layers. Shallow features typically contain rich texture and edge information, while deep features contain rich semantic information.

[0038] To address the issue of minute cracks being easily lost during multi-layer convolution in bank collapse monitoring, this embodiment introduces a targeted mechanism to enhance high-frequency information during feature extraction. High-frequency boundary information in the multi-scale bank collapse feature map is weighted and enhanced, specifically by embedding boundary attention units in key nodes of the feature extraction network (e.g., within the C2f module). This unit explicitly extracts gradient information, i.e., high-frequency components, from the feature map using edge detection operators (such as the Sobel operator) and generates an attention weight map based on the gradient intensity. This weight map is then multiplied element-wise with the original feature map, resulting in a numerical amplification of high-frequency features such as cracks and fracture edges, while suppressing smooth background areas. This process ensures that subsequent prediction heads receive multi-scale bank collapse feature maps with clear edges and well-preserved details.

[0039] Step 104: Using the dual-branch decoupling head in the bank collapse monitoring model, based on the weighted and enhanced multi-scale bank collapse feature map, collaboratively predict and output a regional semantic probability map reflecting the extent of soil collapse, as well as a boundary contour probability map reflecting cracks and erosion edges.

[0040] In this embodiment, to simultaneously achieve accurate segmentation of the collapsed bank area and precise localization of the fracture boundaries, the model adopts a dual-task parallel output structure, namely a dual-branch decoupled head. This structure abandons the traditional single-task network approach of outputting only a single mask, and instead decouples the tasks into two branches: a region segmentation branch and a boundary detection branch.

[0041] Specifically, the region segmentation branch focuses on identifying which pixels belong to the collapsed bank soil, and its output is a region semantic probability map. The value of each pixel in the map represents the confidence level (between 0 and 1) that the point belongs to the collapsed bank region. The boundary detection branch focuses on identifying crack lines and collapse leading edges, and its output is a boundary contour probability map, where highlighted areas correspond to high-frequency edges in the image. Collaborative prediction refers to the two branches sharing the feature maps extracted in the preceding steps and mutually constraining each other during training through a joint loss function. This design not only improves the model's ability to analyze complex bank collapse morphologies, but the output boundary contour probability map can also provide intuitive evidence for subsequent analysis of bank collapse mechanisms (such as tensile or shear failure).

[0042] Step 105: Calculate the collapse loss assessment index based on the regional semantic probability map and the boundary contour probability map.

[0043] In this embodiment, this step transforms the model's probability output into specific physical quantification indicators to support the assessment of bank collapse disasters. The main indicators for assessing collapse loss include collapse area and collapse volume. For the collapse area, the system first binarizes the regional semantic probability map to generate a defined bank collapse area mask, counts the number of pixels within the mask, and combines this with camera calibration parameters—that is, the ratio of pixels to actual physical length—to calculate the actual projected area of ​​the collapsed bank.

[0044] For the collapse volume, since monocular vision cannot directly obtain depth information, this embodiment utilizes the geometric constraints provided by the boundary contour probability map, combined with a preset collapse depth estimation model, for estimation. For example, based on the identified collapse bank length and collapse area width, the soil loss volume can be calculated using empirical formulas or prism-like models. The final output evaluation index can be presented in the form of a time series curve, intuitively reflecting the soil loss rate and cumulative loss during the collapse process.

[0045] Based on the methods provided in the above embodiments, this embodiment also provides a dynamic monitoring and evaluation system for soil collapse loss in bank collapse tests based on an improved deep learning model. Logically, this system includes a data acquisition module, an enhancement processing module, a model processing module, and an evaluation module.

[0046] The data acquisition module is used to perform the image acquisition and initial feature calculation functions in step 101 above;

[0047] The enhancement processing module is used to perform the scale perturbation-aware enhancement function in step 102, and is typically integrated into the image preprocessing pipeline;

[0048] The model processing module is used to load the deep learning model with pre-trained weights and perform the feature extraction and bi-branch prediction functions in steps 103 and 104. This module runs on a server with GPU acceleration.

[0049] The evaluation module is used to perform the post-processing and index calculation functions in step 105, and output the final results to the visualization interface or database.

[0050] Through the collaborative work of various modules, fully automated monitoring from capturing experimental phenomena to outputting quantitative data has been achieved.

[0051] As a preferred implementation, the dynamic monitoring and evaluation system for soil collapse loss in bank collapse tests also includes an anomaly handling mechanism, as detailed below:

[0052] Image acquisition anomaly handling: When N consecutive frames (e.g., N=5) of image acquisition failure or abnormal image quality (e.g., overexposure, underexposure, blur, etc.) are detected, the system automatically triggers an alarm and uses the most recent valid image for interpolation filling, while marking the reliability level of the data for that period in the output results.

[0053] Illumination change adaptation: When a drastic change in the overall brightness of the image is detected (e.g., the average gray value changes by more than 30%), the system automatically enables histogram equalization or adaptive gamma correction for preprocessing to ensure the stability of feature extraction.

[0054] Model inference failure handling: When the model inference time exceeds the preset threshold (e.g., the single frame inference time exceeds 5 seconds) or an exception such as memory overflow occurs, the system automatically restarts the inference process and marks the current frame as pending processing. Supplementary inference will be performed after the system recovers.

[0055] No-bank collapse area identification: When the maximum probability value of the semantic probability map of a region is lower than a preset threshold (e.g., 0.3), the system determines that there is no obvious bank collapse phenomenon in the current frame, outputs that the collapse area and volume are both zero, and records this state in the log.

[0056] Example 2 describes the calculation method of physical scale characteristics in bank collapse monitoring and the scale disturbance sensing enhancement strategy based on these characteristics, such as... Figure 2 , Figure 3As shown, conventional random data augmentation methods are often inefficient and may even disrupt the topology of minute cracks, given the large scale span (from centimeter-level tensile cracks to meter-level overall collapse) exhibiting in the spatiotemporal evolution of bank collapse phenomena. This embodiment introduces physical geometric priors to construct an adaptive perception-decision-augmentation closed-loop mechanism, ensuring that the model can learn bank collapse features at different scales in a balanced manner during the training phase, thereby improving its ability to capture subtle precursor features.

[0057] Step 201: Extract the mask of the collapsed bank area and the mask of the water body area from the image sequence.

[0058] In this embodiment, this step is fundamental to scale analysis. The image sequence typically consists of consecutive RGB image frames, each frame I having a size of H*W (H being the image height and W being the image width). To quantify the scale of the bank collapse, the image is first roughly segmented into three main semantic regions: the bank collapse region (i.e., the exposed soil or the area that is collapsing), the water region (the river surface), and the background region (the undeformed bank or other environment).

[0059] Specifically, extracting the mask M of the collapsed bank area c With water area mask M w Various image segmentation techniques can be employed. As a basic implementation method, threshold segmentation based on color space can be used. For example, since the newly exposed soil after a bank collapse usually presents a different color (such as dark brown or reddish-brown) than the surrounding vegetation or topsoil, while the water usually appears grayish-green or turbid yellow, the image can be converted to the HSV (Hue-Saturation-Value) color space, and binary masks of the soil and water can be extracted separately using preset hue and saturation threshold ranges.

[0060] As a preferred implementation, to improve the robustness of extraction, a lightweight semantic segmentation network (such as UNet or a simplified version of DeepLabV3+) can be pre-trained on a small number of samples. This pre-model can then be used to quickly infer the image sequence and generate the corresponding binary mask. The mask M... c (x,y)=1 indicates that pixel (x,y) belongs to the collapsed shore region, otherwise it is 0; similarly, M w (x,y)=1 indicates that the pixel belongs to the water body region. The extracted mask does not need to achieve perfect pixel-level accuracy; it only needs to reflect the macroscopic distribution of the region to meet the requirements of subsequent scale calculations.

[0061] Step 202: Count the number of pixels in the mask of the collapsed bank area, calculate the ratio of the number of pixels in the mask of the collapsed bank area to the total number of pixels in the image, and obtain the area ratio of the collapsed bank area.

[0062] In this embodiment, the area of ​​the collapsed bank accounts for R. a It is a dimensionless scalar used to visually represent the areal scale of a bank collapse event. The calculation process is based on the mask M generated in step 201. c Specifically, the system iterates through the mask matrix, counts the total number of pixels with a value of 1, and divides it by the total number of pixels in the image.

[0063] The calculation formula is as follows: R a =∑(M c ) / (H*W);

[0064] Among them, R a The percentage of the bank collapse area is ∑(M) c ) represents the total number of pixels with a value of 1 in the mask of the collapsed shore area, H is the image height, and W is the image width.

[0065] For example, for an image with a resolution of 1920*1080, if the collapsed bank area is found to contain 207360 pixels, then R a =207360 / (1920*1080)=0.1. The larger this value, the wider the area of ​​bank collapse damage. Typically, R... a The value of R is in the range of [0,1], but in actual bank collapse tests, due to the limitation of the camera's field of view, R... a It typically fluctuates between 0.01 and 0.6.

[0066] Step 203: Extract the boundary contour line between the mask of the collapsed bank area and the mask of the water body area, calculate the ratio of the pixel length of the boundary contour line to the edge length of the image, and obtain the relative length ratio of the shoreline.

[0067] In this embodiment, the relative length of the shoreline is R l It is a linear index used to describe the morphological complexity of the leading edge of a bank collapse. When a bank collapse occurs, the soil entering the water causes significant morphological changes in the boundary between land and water (shoreline). Severe collapses are often accompanied by the tortuosity, sawing, or elongation of the shoreline.

[0068] Specifically, extracting the boundary contour line can be done by analyzing M... c and M w Implemented by performing logical operations or edge detection. The spatially adjacent regions of two masks can be found through dilation operations, or directly by adjusting M. c Perform Canny edge detection and filter out those edges that are also located in M. w The set of pixels near the edge is denoted as set E. border Through statistical set E borderThe pixel length L of the boundary contour line can be obtained by counting the number of pixels or by calculating its geometric length using a chain code algorithm. c To eliminate the influence of image resolution, it needs to be normalized. The normalization reference is usually chosen as the diagonal length or the length of the longer side L of the image. side =max(H,W).

[0069] The calculation formula is as follows: R l =L c / L side ;

[0070] Among them, R l L represents the ratio of the relative lengths of the shorelines. c L is the pixel length of the outline of the boundary between the collapsed bank area and the water body area. side The length of the longer side of the image is in pixels.

[0071] Step 204: Use the ratio of the area of ​​the collapsed bank to the relative length of the shoreline as a physical scale feature.

[0072] In this embodiment, after the above calculations, R a and R l Together, they constitute the physical scale feature vector V describing the current state of the collapsed bank. scale =[R a ,R l This feature vector contains not only area information but also implicit shape information. For example, when R... a Smaller but R l When R is large, it often corresponds to narrow, elongated fissure collapses or shoreline erosion; when R a Larger and R l When the value is smaller, it may correspond to the overall slippage of large soil masses (with regular block shapes). These two features provide a multi-dimensional basis for decision-making regarding subsequent enhancement strategies.

[0073] Step 205: Weighted summation of the area ratio of the collapsed bank area and the relative length ratio of the shoreline to obtain the comprehensive scale index.

[0074] To transform multidimensional physical characteristics into a single decision variable, it is necessary to calculate the comprehensive scale index λ. Considering that in bank collapse monitoring, area proportion usually reflects the disaster level more directly, it is often given a higher weight in the weighting process.

[0075] Specifically, a weighted sum is calculated by taking the area ratio of the collapsed bank area as a percentage of the relative length of the shoreline. In particular, a weighting coefficient is assigned to the area ratio of the collapsed bank area as a percentage of the relative length of the shoreline to calculate the comprehensive scale index.

[0076] The calculation formula is as follows: λ=α*R a +β*R l ;

[0077] Among them, λ is the comprehensive scale index, α is the weight coefficient of the proportion of the bank collapse area, β is the weight coefficient of the relative shoreline length ratio, and α + β = 1 is satisfied. Usually, α = 0.6 to 0.7 and β = 0.3 to 0.4 are preferably selected.

[0078] For example, set α = 0.7 and β = 0.3. If the current frame R a = 0.1, R l = 0.2, then λ = 0.7 * 0.1 + 0.3 * 0.2 = 0.13. This index λ comprehensively reflects the visual significance of bank collapse.

[0079] Step 206: Based on the preset scale threshold interval, determine the bank collapse scale level to which the comprehensive scale index belongs.

[0080] In this embodiment, the system presets a threshold interval to divide the bank collapse scale level. Usually, two thresholds T1 and T2 (T1 < T2) are set, and the scale is divided into three levels: small scale stage, medium scale stage, and large scale stage.

[0081] If λ < T1, it is determined as the small scale stage. At this time, it usually corresponds to the initial stage of bank collapse, manifested as the萌生 of细微裂隙 or local soil particle spalling.

[0082] If T1 ≤ λ ≤ T2, it is determined as the medium scale stage. At this time, it corresponds to the typical collapse development period.

[0083] If λ > T2, it is determined as the large scale stage. At this time, it corresponds to a large-scale landslide event.

[0084] As an example, the thresholds can be set as T1 = 0.15 and T2 = 0.4.

[0085] Step **********: If the bank collapse scale level is determined as the small scale stage, calculate a scaling ratio greater than 1 to perform zoom-in enhancement on the image sequence to strengthen the crack features.

[0086] In this embodiment, for small scale samples, such as micro cracks, the main problem is that the features account for too small proportion in the image and are likely to disappear during the downsampling process of the convolutional network. The zoom-in enhancement strategy is adopted.

[0087] Specifically, the scaling ratio S greater than 1 can be obtained through the following inverse proportional function or linear mapping:

[0088] S = 1 + k1 * (T1 - λ);

[0089] Among them, S is the scaling ratio, k1 is the zoom-in gain coefficient (for example, take 2.0), T1 is the small scale threshold, and λ is the current comprehensive scale index.

[0090] For example, if T1=0.15, λ=0.05, and k1=2.0, then S=1+2.0*(0.15-0.05)=1.2. This means the image is magnified by a factor of 1.2. The magnification operation typically involves cropping and scaling centered on the centroid of the collapsed area, allowing the tiny cracks to occupy more pixels in the input tensor, thus enhancing their features.

[0091] Step 208: If the scale level of the bank collapse is determined to be large-scale, a scaling factor less than 1 is calculated to perform downscaling enhancement on the image sequence to maintain the integrity of the collapse outline. If the scale level of the bank collapse is determined to be medium-scale, the original scale of the image sequence is maintained or a small-amplitude random scaling enhancement is performed.

[0092] In this embodiment, for large-scale samples, the main problem is that the collapsed bank area may exceed the camera's field of view, resulting in truncated contours. A zoom-out strategy is adopted.

[0093] Specifically, the scaling factor S, which is less than 1, can be obtained using the following formula:

[0094] S = 1 - k2 * (λ - T2);

[0095] Where S is the scaling ratio, k2 is the shrinkage gain coefficient (e.g., 0.5), T2 is the large-scale threshold, and λ is the current overall scale exponent. To prevent the image from becoming too small, a lower limit for S is usually set, such as 0.5.

[0096] For example, if T2=0.4, λ=0.6, k2=0.5, then S=1-0.5*(0.6-0.4)=0.9. That is, the image is reduced by a factor of 0.9, and gray or black pixels are filled at the edges to simulate the effect of observation from a greater distance while preserving the complete collapsed outline.

[0097] As a supplement and preferred implementation of steps 207 and 208, the scale perturbation perception enhancement processing also includes perspective transformation enhancement. Specifically, perspective transformation enhancement is performed by constructing a homography matrix based on the preset observation height and tilt angle range of the bank collapse monitoring scene, and then performing random four-corner point offset transformation on the image sequence to simulate the differences in bank collapse imaging from different perspectives.

[0098] Specifically, in bank collapse experiments, the camera positions are usually fixed, resulting in a lack of perspective diversity in the dataset. To improve the model's generalization ability and enable it to adapt to monitoring scenarios with different installation heights and angles, this embodiment introduces simulated perspective transformation.

[0099] This process involves using the four corner points (x, y) of the original image. i ,y i Random offset (Δx) is superimposed on top of ) i ,Δy iThis is achieved through a virtual observation mechanism. The range of the offset is controlled by preset virtual observation parameters. For example, when simulating changes in camera pitch angle, the horizontal offset of the two upper corner points of the image is mainly adjusted; when simulating changes in side view, the vertical offset of the left or right corner points is mainly adjusted.

[0100] Based on the coordinates of the four new corner points after offset, the homography matrix H is solved using the least squares method. matrix The matrix is ​​then used to project and transform the original image, generating a new image that appears to be taken from a different perspective. This enhancement allows the model to learn spatially invariant bank collapse features, ensuring high accuracy even with minor adjustments to the camera angle during actual deployment.

[0101] Example 3 describes the improved feature extraction network structure in the bank collapse monitoring model, particularly the embedded C2f edge attention mechanism. Existing convolutional neural networks often downsample images at high magnification when extracting features. While this expands the receptive field, it also leads to a significant loss of high-frequency detail information (such as tiny hairline cracks in the early stages of bank collapse). To address this issue, this example designs a plug-and-play edge-attention unit on the critical path of the feature extraction network. Through explicit gradient calculation and channel weighting, it forces the network to focus on edge regions with significant grayscale transitions.

[0102] Step 301: The feature extraction network includes a cross-stage local network module, and a boundary attention unit is embedded in the cross-stage local network module.

[0103] In this embodiment, the feature extraction network (Backbone) adopts an improved version of the CSP (Cross Stage Partial) network concept, specifically the C2f module. The C2f module achieves a richer gradient flow through a split-merge strategy. A standard C2f module typically includes an input layer, two branch paths (main path and residual path), multiple stacked bottleneck layers, and a final concatenation layer.

[0104] This embodiment does not directly use the standard C2f structure, but instead embeds a boundary attention unit within its internal structure. Specifically, this embedding position is designed to be located after the bottleneck layer stacked output and before the final concatenation operation. This layout allows the attention mechanism to directly act on the high-dimensional features after deep transformation, performing weighted filtering. The specific distribution of the cross-stage local network module (C2f) in the network is typically located at layers P3, P4, and P5 of the backbone, corresponding to 8x, 16x, and 32x downsampled feature maps, respectively. The shallower the layer (such as P3), the more significant the effect of embedding this unit is on preserving small gap features.

[0105] Step 302: In the feature extraction process, the high-frequency boundary information in the multi-scale collapse feature map is weighted and enhanced. Specifically, this is achieved by configuring the boundary attention unit between the bottleneck layer output and the splicing operation layer input in the cross-stage local network module.

[0106] This embodiment describes the data flow of the structure in detail. Assume the input feature map of the C2f module is X. i n, after the splitting operation, is divided into two parts: X res (Residual path) and X main (Main path). X main After processing through N bottleneck layers, the intermediate feature F is obtained. mid .

[0107] At this point, the traditional C2f will F mid Directly with X res The parts are then assembled. In this embodiment, F... mid First, the input is fed into the boundary attention unit, and after processing, the enhanced feature F is output. enhanced The system performs a concatenation operation: Concat([X...) res ,F enhanced The final features are then fused through a 1x1 convolutional layer.

[0108] This configuration ensures that only the main path features that have undergone deep nonlinear transformation are given attention weights, while the residual paths remain unchanged. This enhances edge information while avoiding gradient vanishing or exploding, thus maintaining training stability.

[0109] Step 303: Use the Sobel operator to perform convolution calculation on the intermediate feature maps in the cross-stage local network module to obtain spatial gradient features containing horizontal and vertical gradient components.

[0110] In this embodiment, this is the first step in the boundary attention unit processing flow. The input is the intermediate feature map F. midIts dimensions are C*H*W. To capture edges in the image, i.e. regions with drastic grayscale changes, the classic Sobel operator is used as a convolution kernel with fixed weights.

[0111] Specifically, the horizontal direction kernel K is applied respectively. x and vertical direction kernel K y For F mid Convolution is performed on each channel.

[0112] The specific numerical definition of the Sobel operator kernel is as follows:

[0113] K x =[[-1,0,1],[-2,0,2],[-1,0,1]];

[0114] K y =[[-1,-2,-1],[0,0,0],[1,2,1]];

[0115] Among them, K x Used to detect vertical edges (responding to gradient changes in the horizontal direction), K y Used to detect horizontal edges (in response to gradient changes in the vertical direction).

[0116] After the convolution operation, the horizontal gradient components G are obtained respectively. x and vertical gradient component G y Calculate the gradient magnitude G as a spatial gradient feature:

[0117] G=sqrt(G x 2 +G y 2 );

[0118] Where G is the gradient magnitude feature map, and its dimension is the same as the input feature. Figure 1 The characteristic map G exhibits a high numerical response at the crack edge, while the response is close to zero at flat soil surfaces or water surfaces.

[0119] Step 304: Perform max pooling and average pooling operations in parallel on the spatial gradient features along the channel dimension, and then concatenate and fused the pooling results to generate a single-channel spatial attention weight map.

[0120] In this embodiment, in order to reduce the computational load and extract the most representative edge responses, it is necessary to compress the multi-channel gradient features G.

[0121] Perform global max pooling and global average pooling on the channel dimension respectively:

[0122] Max pooling A max=MaxPool(G): Extracts the most prominent gradient response across all channels at each spatial location, which helps capture the strongest edge signals.

[0123] Average pooling A avg =AvgPool(G): Extracts the average gradient response of all channels at each spatial location, which helps to preserve the overall texture structure.

[0124] Both have an output dimension of 1*H*W.

[0125] The two feature maps are concatenated along the channel dimension to obtain feature map A with dimensions 2*H*W. cat .

[0126] Using a 7x7 or 3x3 convolutional layer for A cat The attention map is then fused and compressed back to a single channel to obtain the unnormalized attention map A. raw Using large convolutional kernels (such as 7*7) can expand the receptive field, giving edge detection a certain degree of local context awareness.

[0127] Step 305: The spatial attention weight map is mapped to the normalized interval using the Sigmoid activation function, and then element-wise multiplication is performed between it and the intermediate feature map to enhance high-frequency boundary information.

[0128] In this embodiment, A is first placed raw Inputting the Sigmoid activation function generates the final spatial attention weight map M. The Sigmoid function maps values ​​to the (0,1) interval, so that each pixel value in M ​​represents the probability weight that the location is an edge.

[0129] The calculation formula is as follows: M(x,y)=1 / (1+exp(-A) raw (x,y)));

[0130] Where M(x,y) is the attention weight at position (x,y), and its value ranges from (0,1).

[0131] Perform feature reweighting. Combine the generated weight map M with the intermediate feature map F of the original input. mid Perform element-wise multiplication (also known as Hadamard product).

[0132] Its weighting formula is as follows: F enhanced (c,x,y)=F mid (c,x,y)*M(x,y);

[0133] Among them, Fenhanced For the enhanced feature map, F mid is the original intermediate feature map, and c is the channel index.

[0134] Through this operation, originally in F mid Weak fracture features with low median values, if their gradient response is high (i.e., corresponding to a high value in M), will be preserved or even enhanced; while background noise with high original values, such as reflected light, will be suppressed if its gradient response is low (i.e., corresponding to a low value in M). This achieves targeted enhancement of high-frequency boundary information, allowing the network to focus on the physical damage interface during the bank collapse process.

[0135] The following is a simplified numerical example to illustrate the calculation process of the boundary attention unit:

[0136] Assume intermediate feature map F mid The size is 256×80×80 (256 channels, 80×80 spatial resolution). In a certain local area (e.g., a 3×3 window corresponding to the edge of a crack in an image), F mid The value on a certain channel is:

[0137] [[0.2,0.3,0.8],[0.2,0.4,0.9],[0.3,0.5,1.0]];

[0138] After convolution using the Sobel operator, the horizontal gradient G x and vertical gradient G y At the center point P c The calculation results are as follows:

[0139] G x (P c =(-1×0.2+1×0.8)+(-2×0.2+2×0.9)+(-1×0.3+1×1.0)=2.7;

[0140] G y (P c =(-1×0.2-2×0.3-1×0.8)+(1×0.3+2×0.5+1×1.0)=0.7;

[0141] Gradient magnitude G(P) c =sqrt(2.7) 2 +0.7 2 )≈2.79;

[0142] After channel pooling, fusion, and sigmoid activation, the attention weight M(P) at this location is... c = sigmoid(2.79)≈0.94, close to 1.0, indicating that the position is identified as a high-probability edge region.

[0143] Connect M and F mid After multiplication, the eigenvalue at that position changes from 0.4 to 0.4 × 0.94 = 0.376 (because the weights are close to 1 after sigmoid, the edge features are basically preserved); while in flat regions (assuming G(P) c With λ ≈ 0.1 and M ≈ 0.52, the eigenvalues ​​will be moderately attenuated, achieving relative enhancement of the edges.

[0144] Example 4 describes the specific architecture of the dual-branch decoupled head in the landslide monitoring model and the method for quantitatively assessing landslide loss based on the output of this architecture. Traditional semantic segmentation networks typically output category masks directly, leading to feature conflicts between different tasks (region identification and boundary localization). The dual-branch structure designed in this example focuses on surface features and line features respectively through physical decoupling and logical synergy, and calculates high-precision landslide area and volume indices accordingly.

[0145] Step 401: The dual-branch decoupling head includes a shared feature layer, a region segmentation branch, and a boundary detection branch.

[0146] In this embodiment, the dual-branch decoupling header is connected after the feature extraction network (Backbone). To reduce the number of parameters and facilitate information exchange, the two branches are not completely independent, but start from a shared feature layer. This shared feature layer is typically composed of a Feature Pyramid Network (FPN) or a Path Aggregation Network (PANet), responsible for fusing feature maps from different scales (e.g., P3, P4, P5) and outputting a unified feature tensor with a fixed number of channels (e.g., 256 channels).

[0147] Step 402: Input the multi-scale bank collapse feature map into the shared feature layer to obtain the prototype features.

[0148] In this embodiment, the multi-scale bank collapse feature map refers to the set of features extracted and enhanced by the Backbone. These features are input into the shared feature layer of the FPN / PANet structure, and after upsampling and feature fusion operations, a set of high-resolution prototype features is output. The prototype features are a set of features with dimension C. proto *H out *W out tensor (C proto H represents the number of channels (semantic dimension) of the prototype feature. out W is the height (in pixels) of the output feature map. out H is the width (in pixels) of the output feature map. out W out Typically, it is 1 / 4 or 1 / 8 of the original image size. It acts as a universal base, containing all the basic elements that constitute the semantics and contours of an image, waiting for subsequent branches to perform specific combinations and decoding.

[0149] Step 403: Use the region segmentation branch to perform semantic convolutional encoding on the prototype features and output a region semantic probability map representing the probability that a pixel belongs to the collapsed shore region.

[0150] In this embodiment, the task of the region segmentation branch is to decode the prototype features into a mask for a specific category. This branch typically consists of two parts:

[0151] Convolutional layer group: contains multiple 3*3 convolutional layers, used to extract semantic context information and adjust the number of channels.

[0152] Prediction layer: A 1x1 convolutional layer that compresses the number of channels to N. class (In this example, it is 1, i.e., the collapse-type).

[0153] Activation function: This is followed by a Sigmoid activation function.

[0154] Its output is the region semantic probability map P. seg Its size is consistent with the input image size (or restored to consistent size through interpolation). The value P of each pixel (x, y) in the image... seg (x,y) belongs to [0,1], which physically represents the probability that the point belongs to a collapsed bank. For example, P seg (x,y)>0.5 is usually identified as a bank collapse area.

[0155] Step 404: Use the boundary detection branch to extract edge features and perform fine-grained convolution on the prototype features, and output a boundary contour probability map that represents the probability that a pixel belongs to a bank collapse fissure or erosion boundary.

[0156] In this embodiment, the boundary detection branch is designed with greater finesse to capture high-frequency details. Unlike the segmentation branch, this branch typically does not include significant downsampling operations to maintain spatial resolution.

[0157] The specific structure includes:

[0158] Fine-grained convolution: Employs dilated convolution or small kernel convolution to extract subtle texture features without losing resolution.

[0159] Gating mechanism (optional): A gating unit can be introduced to allow only features with larger gradient values ​​to pass through.

[0160] Prediction layer and activation: The output is also activated by 1*1 convolution and Sigmoid.

[0161] Its output is a boundary contour probability map P. edge The value P of each pixel (x, y) in the image. edgeThe value (x,y) belonging to [0,1] represents the probability that the point belongs to the boundary of a collapsed bank (crack line or contour line). In visualization, this is typically presented as a thin, highlighted line. The output of this branch is crucial for distinguishing between global slip and local collapse, as their boundary morphologies are drastically different.

[0162] Step 405: Binarize the regional semantic probability map, count the number of pixels in the landslide area and convert it into physical area to obtain the landslide loss area index.

[0163] In this embodiment, this is the first step in assessing the severity of a bank collapse hazard.

[0164] For P seg Binarization is performed to obtain the mask M. seg :

[0165] If P seg (x,y)≥T seg Then M seg (x,y)=1; otherwise, 0. Where the threshold T... seg The value is usually 0.5.

[0166] The total number of pixels N with a value of 1 in the statistical mask. pixels .

[0167] Based on the camera's calibration parameters (pixel equivalent), the number of pixels is converted into physical area.

[0168] The calculation formula is as follows: S loss =N pixels *Area pixel ;

[0169] Among them, S loss N represents the area of ​​landslide loss (unit: square meters). pixels Area represents the total number of pixels in the collapsed bank area. pixel Area represents the actual physical area corresponding to a single pixel (unit: square meters / pixel). pixel Pre-determination can be achieved by setting up a calibration board (such as a checkerboard of known size) in the test scenario.

[0170] Step 406: Extract the boundary geometric information of the bank collapse area based on the boundary contour probability map, and calculate the bank collapse depth characteristics based on the boundary geometric information and the preset depth extrapolation model; combine the bank collapse depth characteristics with the collapse loss area index to calculate the collapse loss volume index.

[0171] This step uses a two-dimensional image to infer a three-dimensional volume.

[0172] This embodiment provides a preferred method for extrapolating depth based on boundary geometric features. The assumption is that the depth of a collapsed bank often has a positive correlation with the width of its surface opening (or crack width), which can be obtained through experimental fitting under specific soil materials.

[0173] The specific calculation process is as follows:

[0174] Extracting the width of the collapsed bank: Based on the boundary contour probability map P edge Extract the left and right boundaries of the collapsed bank region and calculate the width distribution W(y) of the collapsed bank region in each row (or along the normal direction).

[0175] Depth distribution estimation: The depth of each point is estimated using the preset depth estimation model h(y)=f(W(y)).

[0176] The deep simulation model can take one of the following specific forms:

[0177] Linear depth model: h(y)=k depth W(y), where k depth This is a depth coefficient, ranging from 0.3 to 0.8. The specific value is determined through testing based on the soil material properties and the type of bank collapse. For example, for silty sandy soil, k... depth A typical value is 0.5; for cohesive soils, k depth The typical value is 0.4.

[0178] Semi-elliptical depth model: Assuming the cross-section of the collapsed body is approximately a semi-ellipse, the depth calculation formula is: h(y) = (b / a)sqrt(a) 2 -(W(y) / 2) 2 ), where a is the major semi-axis of the ellipse (equal to W(y) / 2), b is the minor semi-axis of the ellipse, and b / a is the major-minor axis ratio coefficient, typically ranging from 0.5 to 1.0.

[0179] Exponential decay depth model: h(y) = h max (1-exp(-γW(y))), where h max The maximum bank collapse depth is defined as γ (which can be set based on the channel depth or historical data), and γ is the attenuation coefficient, typically ranging from 0.5 to 2.0. This model is applicable when the bank collapse depth tends to saturate as the width increases.

[0180] Volume integral: Integrate the estimated depth over the area of ​​the collapsed bank (or sum the discrete values).

[0181] The calculation formula is as follows: V loss =∑(Area pixel *h(x,y));

[0182] Or simplified to: V loss=S loss *h avg ;

[0183] Among them, V loss h represents the volumetric collapse loss (unit: cubic meters), and h(x,y) represents the estimated depth at pixel (x,y). avg This represents the average estimated depth of the collapse area.

[0184] As an alternative implementation, if experimental conditions permit, three-dimensional laser scanning can be performed before and after the bank collapse to establish a realistic digital elevation model (DEM). The accurate volume can be obtained by subtracting the DEMs before and after the collapse, and this can be used as the ground truth to train the aforementioned depth estimation model f(W(y)). This can also achieve high volume estimation accuracy when relying solely on camera monitoring in the future.

[0185] Example 5 describes the training process of the bank collapse monitoring model, particularly the soft label generation mechanism. Traditional deep learning segmentation tasks typically use hard labels (0 or 1), ignoring the inherent ambiguity and uncertainty of bank collapse boundaries, such as the mud transition zone at the soil-water interface and the gradual edges of micro-cracks. Forcing the model to learn hard boundaries that are either 0 or 1 can easily lead to overfitting or jagged predicted edges. By introducing soft weights based on physical geometric properties (such as geometric complexity and gradient direction), a supervision signal that conforms to the physical characteristics of bank collapse is constructed, improving the model's adaptability to complex boundaries.

[0186] Step 501: The bank collapse monitoring model is obtained in advance through a training process based on a multi-task collaborative loss function.

[0187] In this embodiment, model training is an end-to-end optimization process. The input consists of pairs of training data {I,GT}, where I is the enhanced image and GT is the manually annotated ground truth mask.

[0188] To simultaneously optimize both region segmentation and boundary detection tasks, this embodiment constructs a total loss function L. total This function consists of three parts:

[0189] L total =λ1*L seg +λ2*L edge +λ3*L cons ;

[0190] Among them, L seg The region segmentation loss is used to supervise the Seg branch (region segmentation branch); L edge The boundary detection loss is used to supervise the Edge branch (boundary detection branch); L consThe gradient consistency constraint loss is shown in Example 6; λ1, λ2, and λ3 are the hyperparameter weights that balance each loss term (e.g., 1.0, 1.0, and 0.5).

[0191] As a preferred implementation, the following hyperparameter configuration is used during the training process of the bank collapse monitoring model:

[0192] Optimizer: The Adam W optimizer is used, with the initial learning rate set between 1e-4 and 5e-4, and the weight decay set between 0.01 and 0.05.

[0193] Learning rate scheduling: Cosine annealing strategy is adopted, and the training cycle is 100 to 300 epochs (rounds), of which the first 5 to 10 epochs are used for warm-up.

[0194] Batch size: Set between 8 and 32 based on GPU memory; a typical value is 16 for single-card training.

[0195] Input image size: During the training phase, the image is scaled to 640×640 or 1280×1280 pixels; during the inference phase, it can be adjusted according to actual needs.

[0196] Data augmentation: In addition to the aforementioned scale perturbation perception enhancement, it may also include conventional enhancement methods such as random horizontal flipping (probability 0.5), color jitter (brightness, contrast, and saturation ±20% each), and Gaussian blur (kernel size 3×3 to 7×7);

[0197] Loss function weights: Typical values ​​for λ1, λ2, and λ3 are 1.0, 1.0, and 0.5, respectively. In practical applications, these values ​​can be fine-tuned based on the performance of the validation set. The value of λ3 ranges from 0.3 to 1.0. A smaller λ3 focuses on the accuracy of region segmentation, while a larger λ3 focuses on the consistency of boundaries.

[0198] Step 502, the multi-task collaborative loss function includes the region segmentation loss for optimizing the region segmentation branch, the boundary detection loss for optimizing the boundary detection branch, and the gradient consistency constraint loss for constraining the consistency of the outputs of the two branches.

[0199] In this embodiment, the basic form of each loss term is defined in detail.

[0200] For the region segmentation loss L seg Typically, binary cross-entropy loss (BCE Loss) or Dice loss is used. In this embodiment, a weighted improvement is applied (see step 503 for details).

[0201] For boundary detection loss L edgeSince boundary pixels account for a small proportion of the overall image (class imbalance), weighted BCE Loss is typically used.

[0202] L edge =-∑(β'*GT edge *log(P edge )+(1-β')*(1-GT edge )*log(1-P edge ));

[0203] Among them, GT edge For the true boundary label, P edge To predict the boundary probability, β' is a positive sample weighting coefficient (e.g., the total number of pixels in the image / the number of boundary pixels) to balance the influence of positive and negative samples.

[0204] Step 503: The region segmentation loss is calculated by introducing soft weights for boundary uncertainty.

[0205] Traditional BCE Loss assigns the same weight to all pixel locations, but in collapsed bank images, the pixel categories inside are definite, while the pixel categories near the boundary tend to be ambiguous.

[0206] This embodiment will use L seg Modified to weighted form:

[0207] L seg =-∑(W soft (x,y)*(GT(x,y)*log(P seg (x,y))+(1-GT(x,y))*log(1-P seg (x,y))));

[0208] Among them, W soft (x,y) represents the soft weights for boundary uncertainty, GT(x,y) represents the ground truth mask, and P seg (x,y) represents the semantic probability value of the region. The design principle of this weight is: in regions far from the boundary (high confidence), the weight is higher (or kept at 1); in regions close to the boundary (low confidence, blurry), the weight is lower, reducing the penalty for the model to force classification of these blurry pixels and allowing for a certain smooth transition in the prediction results.

[0209] Step 504: When calculating the region segmentation loss, boundary uncertainty soft weights are introduced. The numerical distribution of the boundary uncertainty soft weights is dynamically generated based on the physical geometric properties of the collapsing bank boundary. The distance from the collapsing bank boundary is negatively correlated with its weight value. That is, the closer the region is to the collapsing bank boundary, the lower its weight value. Moreover, the decay range of the weight is controlled by the physical geometric properties.

[0210] This embodiment provides two specific generation strategies, which can be used individually or in combination.

[0211] One type is isotropic soft annotation based on geometric complexity, such as Figure 4 As shown.

[0212] This approach posits that the more tortuous and complex the boundary, the higher the uncertainty in its positioning, requiring a wider softening zone.

[0213] Step 504a: Obtain the true annotation mask of the collapsed shore region corresponding to the training sample, extract the true boundary of the collapsed shore region in the true annotation mask, calculate the ratio of the contour length of the true boundary of the collapsed shore region to the contour length of the convex hull, and obtain the boundary complexity index.

[0214] Extract the boundary contour B of the labeled mask GT true Calculate its length L contour .

[0215] Calculate the convex hull of the contour to obtain the convex hull perimeter L. hull .

[0216] Calculate the boundary complexity exponent C b :C b =L contour / L hull ;

[0217] Obviously, C b >=1. When the edge of the collapsed bank is smooth and rounded, C b Approximately 1; when the edges are serrated or dendritic, C b It is significantly greater than 1.

[0218] Step 504b: Based on the boundary complexity index and image resolution, dynamically calculate the boundary band width of the current sample. The boundary band width is positively correlated with the boundary complexity index, that is, the larger the boundary complexity index, the wider the boundary band width.

[0219] Establish mapping relationship f(C) b The width ω (in pixels) of the softening band is determined using a linear mapping method.

[0220] ω=ω base *C b ;

[0221] Where, ω base The base width is (e.g., 5 pixels). If a sample C... b =2.0, then ω=10 pixels. This means that for complex boundaries, soft weights are applied within a 10-pixel radius around them.

[0222] Step 504c: For pixels located within the width of the boundary band, calculate a linearly decaying weight value based on their Euclidean distance to the true boundary, as a soft weight for boundary uncertainty.

[0223] For each pixel (x, y), calculate its distance to the boundary B. true The shortest Euclidean distance d(x,y).

[0224] If d(x,y)>=ω, then W soft (x,y)=1 (hard-labeled area).

[0225] If d(x,y)<ω, then the weights decay linearly:

[0226] W soft (x,y)=d(x,y) / ω;

[0227] For example, with ω=10, a point only 1 pixel away from the boundary has a weight of only 0.1. This means that the contribution of the model's prediction bias at that point to the total loss is significantly reduced, and the model is not forced to fit this edge point that may have labeling errors.

[0228] Another type is anisotropic soft labeling based on gradient direction.

[0229] This approach is a preferred alternative to the above scheme, further considering the directionality of the boundary. It assumes that the uncertainty of the boundary mainly occurs in the direction perpendicular to the boundary (normal), while the uncertainty is lower in the direction along the boundary (tangential).

[0230] Step 504d: Extract the bank collapse boundary pixels based on the real annotation mask, and calculate the normal and tangent directions of the bank collapse boundary pixels.

[0231] For each point p on the boundary, use the Sobel operator or the structure tensor to compute its normal vector n and tangent vector t.

[0232] Step 504e: Construct an anisotropic decay function based on the normal direction, so that the boundary uncertainty soft weight exhibits a rapid decay characteristic along the boundary normal direction, while maintaining a smooth transition along the boundary tangent direction.

[0233] Construct a Gaussian distribution function centered at the boundary points and aligned in direction as the weight decay model.

[0234] The specific formula is as follows: W soft (x,y)=1-exp(-(d n 2 / (2*σ n 2 ))-(d t 2 / (2*σt 2 )));

[0235] Where, d n d is the projected distance from the pixel to the boundary point along the normal direction. t σ is the projected distance in the tangential direction; n and σ t These are the Gaussian standard deviations in the normal and tangential directions, respectively.

[0236] Set σ n <σ t For example, let σ n =2,σ t =10.

[0237] This indicates that the weights change rapidly in the normal direction (the direction that crosses the boundary), defining a clear edge band; while they change slowly in the tangent direction (the direction along the boundary), ensuring the continuity and smoothness of the boundary.

[0238] Final soft weight map W soft It exhibits a tubular distribution characteristic that flows along the boundary, guiding the model to learn a more continuous, smooth, and physically consistent bank collapse boundary.

[0239] Example 6 describes the gradient consistency loss function used to constrain the synergy of the two-branch output in the bank collapse monitoring model.

[0240] In existing multi-task learning frameworks, semantic segmentation and edge detection branches are often independent of each other, implicitly related only by sharing low-level features. This easily leads to a common logical paradox: the region boundaries predicted by the segmentation branch and the crack lines predicted by the edge detection branch do not coincide spatially (misalignment). This embodiment introduces explicit gradient constraints, utilizing mathematical duality (the gradient of the region is the boundary), to force the outputs of the two branches to maintain strict consistency in a physical sense, improving the geometric accuracy and logical consistency of the prediction results.

[0241] Step 601: The gradient consistency constraint loss is used to constrain the gradient distribution of the region segmentation branch output to be consistent with the prediction result of the boundary detection branch.

[0242] In this embodiment, the physical principle of the loss function is first clarified.

[0243] Mathematically, the spatial gradient ▽(M) of a region mask M should numerically correspond to the edge contour E of that region. In other words, if the probability map P output by the region segmentation branch... seg Taking the derivative, the result should be very close in shape to the boundary contour probability map P output by the boundary detection branch. edgeAnd the true boundaries of GT edge .

[0244] This embodiment constructs a monitoring mechanism that not only monitors P seg The pixel value (this is the region segmentation loss L) seg (task), and also supervise P seg The rate of change (i.e., gradient). This supervision is not uniform across the entire graph, but rather selective.

[0245] Step 602: Using the boundary contour probability map output by the boundary detection branch as spatial attention weights, supervised constraints are applied to the spatial gradient field of the region semantic probability map output by the region segmentation branch, so that in the high-confidence boundary region, the gradient direction and intensity of the region semantic probability map converge to the true boundary gradient.

[0246] This embodiment elaborates on the implementation logic of this constraint.

[0247] Constrained object: Spatial gradient field of the region semantic probability map |▽(P) seg )|. This can be achieved by analyzing P seg It is obtained by performing Sobel convolution calculation.

[0248] Supervision target: the gradient field |▽(GT)| of the true labeled mask, i.e. the true boundary location, usually a binary boundary map.

[0249] Spatial weights: Boundary contour probability map P edge .

[0250] The logic is as follows: only focus on the region that the model considers to be the boundary, i.e., P. edge The region with higher values. If P edge The gradient is very high, indicating that the Edge branch considers this point to be the boundary. Therefore, the gradient of the Seg branch at this point is |▽(P). seg The probability threshold (P) must also be very large (manifested as a sharp probability jump) and must be close to the true value. Conversely, in a flat region (P... edge (Approximately equal to 0), the constraint on the gradient can be relaxed.

[0251] This mechanism essentially uses the prediction results of the Edge branch as a form of spatial attention to guide the Seg branch in optimizing its edge morphology.

[0252] Step 603, the formula for calculating the gradient consistency constraint loss is:

[0253] L cons =∑(P edge (x,y)*L bce (|▽(P seg (x,y))|,|▽(GT(x,y))|));

[0254] Among them, L cons Let P represent the gradient consistency constraint loss, ∑ represent the summation over all pixel positions (x, y) in the image, and P represent the summation over all pixel positions (x, y) in the image. edge (x,y) represents the boundary profile probability value at position (x,y) (from the Edge branch), P seg (x,y) represents the semantic probability value of the region (from the Seg branch), and ▽ represents the spatial gradient operator (which can be implemented using the Sobel operator, |▽(f)|=sqrt(G)). x 2 +G y 2 ), G x and G y These are approximations of the first-order partial derivatives of the image in the x and y directions, respectively. GT represents the ground truth mask, and L... bce This represents the binary cross-entropy loss function.

[0255] In this embodiment, for ease of engineering implementation and gradient backpropagation, the above formula can be expanded into a more specific discrete form.

[0256] Calculate P seg Gradient magnitude map G seg :

[0257] G seg =|▽(P seg )|≈|Sobel x (P seg Sobel y (P seg )|;

[0258] To ensure differentiability, we usually use the sum of absolute values ​​or the square root of the sum of squares as an approximation.

[0259] Calculate the gradient magnitude plot G of the true value GT gt (In fact, it is the boundary truth value GT) edge ):

[0260] G gt =|▽(GT)|≈GT edge ;

[0261] Calculate the difference between the two (using BCE Loss or L1 Loss), and use P edge Weighting:

[0262] L cons =(1 / N)*∑(P edge (x,y)*(-G gt (x,y)*log(G seg (x,y)+ε)-(1-Ggt (x,y))*log(1-G seg (x,y)+ε)));

[0263] Where N is the total number of pixels, and ε is the minimum value of log(0) (e.g., 1e-8).

[0264] Suppose we are at a certain position (x0, y0) in the image:

[0265] Scenario A (ideal consistency): Edge branch predicts P edge =0.9 (the boundary value), Seg branch predicts P seg The jump from 0.1 to 0.9 causes G to... seg Approximately 0.8. Truth value G gt =1. In this case, the loss is smaller.

[0266] Case B (Inconsistent): Edge branch predicts P edge =0.9 (is the boundary), but the Seg branch prediction is ambiguous, P seg A smooth transition leads to G seg It is approximately equal to 0.1. At this time, P edge This difference is amplified by the weights, resulting in a huge loss, which forces the Seg branch to sharpen the prediction at that point, making its gradient larger, so that the segmentation edge coincides with the detection edge.

[0267] Case C (Background Area): Edge branch predicts P edge =0.01. At this point, regardless of the gradient of the Seg branch (which may be noise), its contribution to the total loss is small due to the suppression effect of the weights.

[0268] Through the above methods, L cons Like a bridge, it tightly couples the two branches, realizing a virtuous cycle of detection guiding segmentation and segmentation verifying detection.

[0269] According to one aspect of this application, a method for dynamic monitoring and evaluation of soil collapse loss in bank collapse tests based on an improved YOLOV8 model is provided, comprising the following steps:

[0270] S1, perform bank collapse data augmentation processing to introduce scale and perspective uncertainties.

[0271] Conduct tests on typical water flow scour dynamics within a semi-circular collapse pit and bank collapse under different slope ratios. Deploy a high-frequency video acquisition system at the test site to continuously capture images of the entire bank collapse process.

[0272] In the data construction phase of the bank collapse process, to address the issues of changes in observation scale and uncertain shooting perspective caused by continuous shoreline retreat, expansion of collapse range, and changes in camera deployment conditions in the bank collapse test images, a data augmentation strategy of scale perturbation perception and perspective perturbation was introduced.

[0273] Specifically, based on the relative scale characteristics of the collapsed bank area in the image, the input image is processed in a hierarchical manner: when the collapsed bank area is in the early stage of crack expansion or local erosion, the small cracks and local boundary features are enhanced by enlarging the scale; when the collapsed bank area enters the stage of overall collapse or large-scale collapse, the integrity of the overall collapse outline and contextual structure information is maintained by reducing the scale.

[0274] Meanwhile, by applying perspective transformation to the original images, the imaging differences under different observation heights, pitch angles, and tilt angles are simulated. This allows the model to learn the morphological characteristics of landslide banks under multiple scales and perspectives during the training phase, thereby improving the model's adaptability to complex experimental setup conditions. Combined with time-series indexing, a continuous dynamic sample set of landslides is constructed as the training input for the YOLOv8n recognition model.

[0275] S2 performs soft annotation based on boundary uncertainty.

[0276] In the image annotation stage of bank collapse, a soft annotation mechanism for areas with uncertain boundaries is introduced to address the issues of the bank collapse boundary evolving over time, becoming blurred, and developing irregular shapes during the collapse process.

[0277] In practice, a boundary band region with a preset width is formed by extending outwards from the manually labeled bank collapse boundary. The width of the boundary band is set according to the image spatial resolution and the local complexity of the bank collapse boundary. Inside the boundary band, pixels are assigned continuously varying labeling weights, so that the labeling results smoothly transition from the boundary center to the inner and outer regions, thereby explicitly expressing the uncertainty of the boundary position.

[0278] During model training, the aforementioned weight information is incorporated into the loss calculation process. A weaker constraint intensity is applied to the boundary zone than to the interior of the zone to reduce the impact of manual annotation errors on model parameter updates and improve the continuity and stability of the landslide boundary prediction results over time. The soft annotation information is stored in a one-to-one correspondence with the original image and is associated with the collapse stage labels and time-series information for subsequent model training and dynamic analysis.

[0279] S3 implements the YOLOv8n-seg model (an instance segmentation model based on the YOLOv8 nano architecture) that introduces boundary attention and dual-task collaboration.

[0280] In terms of model structure design, the model is improved based on the YOLOv8n-seg network. In the feature extraction stage, a boundary-aware enhancement mechanism is introduced into the C2f module. By explicitly modeling the spatially significant regions in the intermediate feature map, the model can distinguish between the extended areas of bank collapse fissures and the smooth background regions.

[0281] Based on the extracted boundary response information, boundary attention weights are generated and introduced into the feature fusion process. Spatially weighting of the backbone features strengthens the feature representation related to the irregular boundary and crack structure of the collapsed bank in the multi-scale feature fusion stage, while suppressing the interference of non-critical region features on the segmentation results.

[0282] In the segmentation head structure, a dual-task collaborative mechanism of region segmentation and boundary detection is introduced to decouple the semantic segmentation task of the collapsed bank loss region from the collapsed bank boundary localization task. The region segmentation branch focuses on maintaining the overall continuity of the predicted results for the collapsed bank region, while the boundary detection branch focuses on characterizing fine-grained structural features such as crack edges, erosion boundaries, and collapse contours. During model training, a collaborative relationship is established between the two branches at the loss constraint level, allowing the region segmentation results to be guided by the boundary detection results within the boundary neighborhood, achieving synergistic optimization of region integrity and boundary refinement.

[0283] To verify the model's detection performance and practical application effectiveness, an independent test set was used for model evaluation. An optimized YOLOv8n-seg model was established, capable of real-time detection, dynamic segmentation, and continuous tracking of collapse areas during bank collapse experiments, providing accurate visual recognition support for subsequent volume conversion and risk analysis. The model's recognition performance was quantitatively analyzed using three metrics: precision, recall, and F1-score. Figure 5 As shown, to more intuitively reflect the model's accuracy in recognizing boundary shapes, the Mean Contour Matching Error (MBD) metric is introduced, which is defined as:

[0284] ;

[0285] In the formula, and , where are the predicted and true boundary lengths, respectively, and N is the number of test samples. Model validation results are as follows: Figure 5 As shown, mAP@0.5 represents the average precision calculated when the threshold is set to 0.5, that is, a correct match is judged when the intersection-union ratio (IoU) between the predicted box and the ground truth box is not less than 0.5; mAP@0.5:0.95 represents the index obtained by averaging the AP (average precision) calculated under the standard of IoU threshold (intersection-union ratio threshold) from 0.5 to 0.95 with a step size of 0.05.

[0286] This application employs a physical scale-aware enhancement strategy. By calculating the ratio of the area of ​​the collapsed bank region to the length of the shoreline, it adaptively performs magnification or reduction processing on the image, ensuring the model's ability to capture minute cracks. Simultaneously, a C2f boundary attention unit is embedded in the feature extraction network, and the Sobel operator is used to explicitly extract and weight high-frequency gradient information, effectively preventing the loss of subtle features in deep networks. This solves the problems of difficult identification of minute cracks and large scale spans.

[0287] This application introduces a physically-driven soft labeling mechanism and gradient consistency loss. By calculating the geometric complexity and normal direction of the boundary, dynamically decaying soft weight labels are generated, enabling the model to adapt to the non-rigid transition at the water-soil interface. A dual-branch decoupling architecture, combined with gradient consistency constraints, forces the edges of the segmented region to maintain spatial gradient consistency with the detected crack lines, improving the geometric accuracy of boundary localization. This solves the problems of blurred boundaries and coarse localization.

[0288] This application utilizes the boundary contour probability map output by the dual-branch model as a geometric constraint to construct a depth extrapolation model based on crack width, transforming two-dimensional monitoring data into three-dimensional volume indicators, thus achieving a full-dimensional assessment of collapse loss. This solves the problem of missing volume quantification.

[0289] It should be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present invention will not describe the various possible combinations separately.

Claims

1. A method for dynamic monitoring and evaluation of soil collapse loss in bank collapse tests, characterized in that, include: Acquire image sequences of the bank collapse test process and calculate the physical scale characteristics of the collapsed areas; Based on physical scale characteristics, scale perturbation sensing enhancement processing is performed on the image sequence to obtain enhanced image data; The enhanced image data is input into a pre-trained bank collapse monitoring model. The feature extraction network in the bank collapse monitoring model is used to extract multi-scale bank collapse feature maps. During the feature extraction process, the high-frequency boundary information in the multi-scale bank collapse feature maps is weighted and enhanced. Using the dual-branch decoupling head in the bank collapse monitoring model, based on the weighted and enhanced multi-scale bank collapse feature map, we collaboratively predict and output a regional semantic probability map reflecting the extent of soil collapse, as well as a boundary contour probability map reflecting cracks and erosion edges. Calculate collapse loss assessment indicators based on the regional semantic probability map and the boundary contour probability map; Calculate the physical scale characteristics of the bank collapse area in the image sequence, including: Extract the masking of the bank collapse area and the masking of the water body area from the image sequence; The number of pixels in the mask of the collapsed bank area is counted, and the ratio of the number of pixels to the total number of pixels in the image is calculated to obtain the area ratio of the collapsed bank area. Extract the boundary contour line between the mask of the collapsed bank area and the mask of the water body area, calculate the ratio of its pixel length to the image side length, and obtain the relative length ratio of the bank line. The ratio of the area of ​​the collapsed bank to the relative length of the shoreline is used as a physical scale feature. The feature extraction network includes cross-stage local network modules, in which boundary attention units are embedded; During feature extraction, high-frequency boundary information in the multi-scale collapsed bank feature map is weighted and enhanced by configuring the boundary attention unit between the bottleneck layer output and the splicing operation layer input in the cross-stage local network module. The boundary attention unit performs the following processing steps: The Sobel operator is used to perform convolution calculation on the intermediate feature maps in the cross-stage local network module to obtain spatial gradient features containing horizontal and vertical gradient components. Max pooling and average pooling operations are performed in parallel on the spatial gradient features along the channel dimension, and the pooling results are concatenated and fused by convolution to generate a single-channel spatial attention weight map. The spatial attention weight map is mapped to a normalized interval using the Sigmoid activation function, and then element-wise multiplication is performed between it and the intermediate feature map to enhance high-frequency boundary information.

2. The method according to claim 1, characterized in that, Based on physical scale characteristics, scale perturbation-aware enhancement processing is performed on image sequences, including: The comprehensive scale index is obtained by weighted summation of the area ratio of the collapsed bank area and the relative length of the shoreline. Based on the preset scale threshold range, determine the bank collapse scale level to which the comprehensive scale index belongs; If the scale of the bank collapse is determined to be small-scale, a scaling factor greater than 1 is calculated to perform magnification enhancement on the image sequence to strengthen the crack features. If the scale level of the landslide is determined to be large-scale, a scaling factor of less than 1 is calculated to perform downscaling enhancement on the image sequence in order to maintain the integrity of the landslide outline.

3. The method according to claim 1, characterized in that, The dual-branch decoupling head includes a shared feature layer, a region segmentation branch, and a boundary detection branch; it collaboratively predicts and outputs, including: The multi-scale bank collapse feature map is input into the shared feature layer to obtain the prototype feature; The prototype features are semantically convolutionally encoded using a region segmentation branch, and the output is a region semantic probability map representing the probability that a pixel belongs to a landslide region. The boundary detection branch is used to extract edge features and perform fine-grained convolution on the prototype features, and outputs a boundary contour probability map that represents the probability that a pixel belongs to a bank collapse fissure or erosion boundary.

4. The method according to claim 3, characterized in that, Based on the regional semantic probability map and the boundary contour probability map, collapse loss assessment indicators are calculated, including: The semantic probability map of the region is binarized, the number of pixels in the landslide area is counted and converted into physical area to obtain the landslide loss area index. The boundary geometric information of the bank collapse area is extracted based on the boundary contour probability map, and the bank collapse depth features are calculated based on the boundary geometric information and the preset depth inference model. By combining the characteristics of the bank collapse depth with the collapse loss area index, the collapse loss volume index is calculated.

5. The method according to claim 3, characterized in that, The bank collapse monitoring model is pre-obtained through a training process based on a multi-task collaborative loss function; the multi-task collaborative loss function includes: The region segmentation loss is used to optimize the region segmentation branch; This is used to optimize the boundary detection loss of the boundary detection branch. And a gradient consistency constraint loss used to constrain the consistency of outputs between the region segmentation branch and the boundary detection branch; The region segmentation loss incorporates soft weights to account for boundary uncertainties during computation. The numerical distribution of the boundary uncertainty soft weights is dynamically generated based on the physical geometric properties of the collapsing bank boundary. The closer the region is to the collapsing bank boundary, the lower its weight value, and the decay range of the weight is controlled by the physical geometric properties.

6. The method according to claim 5, characterized in that, Dynamically generate soft weights for boundary uncertainties based on the physical and geometric properties of the collapsed bank boundary, including: Obtain the true labeled mask of the collapsed shore region corresponding to the training sample, extract the true boundary of the collapsed shore region in the true labeled mask, calculate the ratio of the contour length of the true boundary of the collapsed shore region to the contour length of the convex hull, and obtain the boundary complexity index. The boundary band width of the current sample is dynamically calculated based on the boundary complexity index and image resolution. The larger the boundary complexity index, the wider the boundary band width. For pixels located within the width of the boundary band, a linearly decaying weight value is calculated based on their Euclidean distance to the true boundary, serving as a soft weight for boundary uncertainty.

7. The method according to claim 5, characterized in that, The gradient consistency constraint loss is used to ensure that the gradient distribution of the region segmentation branch output is consistent with the prediction result of the boundary detection branch; specifically, it manifests as follows: By using the boundary contour probability map output by the boundary detection branch as spatial attention weight, a supervisory constraint is applied to the spatial gradient field of the region semantic probability map output by the region segmentation branch, so that in the high-confidence boundary region, the gradient direction and intensity of the region semantic probability map converge to the true boundary gradient. The formula for calculating the gradient consistency constraint loss is: L cons =∑(P edge (x,y)*L bce (|▽(P seg (x,y))|,|▽(GT(x,y))|)); Among them, L cons P represents the gradient consistency constraint loss. edge (x,y) represents the probability value of the boundary profile at position (x,y), P seg (x,y) represents the semantic probability value of the region, ▽ represents the spatial gradient operator, GT represents the ground truth mask, and L bce This represents the binary cross-entropy loss function.

8. A dynamic monitoring and evaluation system for soil collapse loss in bank collapse tests, characterized in that, include: The data acquisition module is used to acquire image sequences of the bank collapse test process and calculate the physical scale characteristics of the bank collapse area in the image sequence; The enhancement processing module is used to perform scale perturbation-aware enhancement processing on image sequences based on physical scale features; The model processing module is used to input the enhanced image data into the pre-trained bank collapse monitoring model, and use the feature extraction network and the two-branch decoupling head in the model to jointly predict and output the regional semantic probability map and the boundary contour probability map. The assessment module is used to calculate collapse loss assessment indicators based on the regional semantic probability map and the boundary contour probability map. The system operates by implementing the method as described in any one of claims 1-7.