A Debris Flow Identification Method and System Based on Background Decoupling and Flow Attention Masking

CN122574746BActive Publication Date: 2026-09-18SICHUAN HUADIAN MULIHE HYDROPOWER DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611064244.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-17
Publication Date
2026-09-18
Estimated Expiration
2046-07-17

AI Technical Summary

Technical Problem

[0004]针对现有技术中的上述不足,本发明提供的基于背景解耦与流态注意力掩膜的泥石流识别方法及系统,解决了现有泥石流监测技术因静态相似背景干扰且缺乏时域运动约束导致监测误报率高、且跨场景泛化能力不足的技术问题

Benefits of technology

1、在图像特征层面抑制静态相似背景干扰

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122574746B_ABST
    Figure CN122574746B_ABST
Patent Text Reader

Abstract

This invention provides a debris flow identification method and system based on background decoupling and flow-state attention masking, belonging to the field of debris flow monitoring. The method includes: acquiring historical videos without debris flow events and debris flow monitoring videos, and performing background decoupling on the image sequence to obtain a spatial change mask for the current frame; obtaining a temporal flow-state feature map through a dense optical flow network based on the current frame image and the previous frame image; inputting the current frame image into a feature extraction network, and obtaining a fused feature map through solid-phase sensing and liquid-phase sensing branches; scaling and stitching the spatial change mask and the temporal flow-state feature map, then inputting them into an attention generation network to output a flow-state attention mask, which is then fused with the fused feature map and the deep feature map to obtain a reconstructed feature map; finally, outputting the debris flow bounding box and confidence score via a detection head. This invention solves the problems of high false alarm rates and insufficient cross-scene generalization ability in existing debris flow monitoring technologies due to static similar background interference and lack of temporal motion constraints.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of debris flow monitoring, and particularly relates to a debris flow identification method and system based on background decoupling and flow state attention masking. Background Technology

[0002] Debris flow disasters are characterized by their suddenness, destructive power, and high monitoring difficulty. Intelligent monitoring based on image vision is the mainstream technical means for debris flow disaster early warning. Common technical solutions mainly include deep learning detection methods based on single-frame images, that is, using models such as convolutional neural networks or Vision Transformer to extract and classify single-frame features from real-time video frames captured by monitoring cameras. However, such methods only rely on spatial appearance information and lack the constraint of temporal motion information. They are prone to misinterpreting static similar backgrounds such as static deposits remaining in the gully, moist riverbeds, and exposed rock and soil slopes as target categories, resulting in a high false alarm rate.

[0003] Another common approach is based on inter-frame differencing and background modeling. For example, Gaussian mixture models, inter-frame differencing, or the ViBe algorithm are used to model the static background in video sequences. The moving foreground region is extracted by differencing the current frame with the dynamically updated background model, and then this region is identified. While these methods introduce temporal information, they mostly extract the foreground in the preprocessing stage, failing to deeply integrate motion priors into the image feature extraction stage. Furthermore, traditional background modeling is sensitive to interference from lighting changes and shadows. In addition, there are post-processing methods based on temporal information, which introduce target tracking algorithms such as Kalman filtering and Hungarian tracking based on the single-frame detection results. These methods decouple temporal information from feature extraction, performing only logical filtering in the backend, and cannot suppress responses from statically similar backgrounds at the feature level. Meanwhile, existing attention mechanisms are mostly based on general spatial features, failing to consider the special motion patterns of solid-liquid two-phase mixtures in debris flows, and failing to utilize the continuity prior of fluid motion, resulting in insufficient cross-scene generalization ability. Summary of the Invention

[0004] To address the aforementioned shortcomings in existing technologies, the present invention provides a debris flow identification method and system based on background decoupling and flow state attention masking, which solves the technical problems of high false alarm rate and insufficient cross-scene generalization ability of existing debris flow monitoring technologies due to static similar background interference and lack of temporal motion constraints.

[0005] To achieve the above objectives, the technical solution adopted in this invention is: a debris flow identification method based on background decoupling and flow state attention masking, comprising the following steps: Acquire historical videos of debris flow events and debris flow monitoring videos, and obtain a debris flow spatial variation mask in the current frame image of the debris flow monitoring video by decoupling the image sequence background. Based on the current frame and the previous frame of the debris flow monitoring video, the optical flow field is calculated through a dense optical flow network, the optical flow amplitude is extracted and normalized, and a time-domain flow characteristic map is obtained. The current frame image of the debris flow monitoring video is input into the feature extraction network, which outputs a deep feature map. The features of debris flow rocks and muddy water are extracted through parallel solid phase sensing branches and liquid phase sensing branches, respectively, to obtain a fused feature map. The debris flow spatial variation mask and temporal flow feature map are scaled to the same spatial size as the fused feature map, and then stitched together along the channel before being input into the attention generation network to output the flow attention mask; Based on the fluid attention mask, attention weighting is performed on the fused feature map, and then residual fusion is performed on the deep feature map to obtain the reconstructed feature map, which is then input into the detection head to output the bounding box and confidence score of the debris flow target.

[0006] In existing debris flow image monitoring technologies, detection methods based on single-frame deep learning rely solely on spatial appearance information, which easily leads to misidentification of statically similar backgrounds such as moist riverbeds and exposed soil as debris flows, resulting in a high false alarm rate. This invention generates a spatial variation mask by decoupling the background of image sequences and combines it with temporal flow feature maps extracted by a dense optical flow network to generate a flow attention mask at the feature level. This dynamically enhances and statically suppresses deep feature maps, reducing the misidentification of statically similar backgrounds, achieving motion-static decoupling at the feature level, and improving the accuracy of debris flow identification. Existing methods based on inter-frame difference or background modeling only extract the foreground in the preprocessing stage, failing to deeply integrate motion priors in the feature extraction stage, and are sensitive to changes in illumination. This invention scales and stitches a spatially varying mask and a temporal flow feature map and inputs them into an attention generation network to achieve the fusion of temporal information and feature extraction, thereby enhancing robustness to interferences such as illumination. Existing attention mechanisms are mostly based on the spatial characteristics of general debris flow mountains, without considering the special motion form of the solid-liquid two-phase mixture of debris flow, resulting in insufficient cross-scene generalization ability. This invention sets up parallel solid phase sensing branches and liquid phase sensing branches and fuses them, which improves the ability to represent the edges and textures of rocks and reduces the fluid edge positioning error.

[0007] Further: the debris flow spatial variation mask obtained from the current frame image of the debris flow monitoring video specifically includes: Acquire historical videos of events without mudslides, select several consecutive frames, and construct a pure background baseline image by calculating the median value element by element. Acquire the current frame image of the debris flow monitoring video and calculate the difference image with the pure background baseline image; An adaptive threshold is dynamically calculated based on the local neighborhood grayscale standard deviation, and the differential image is binarized and segmented to obtain a debris flow spatial variation mask in the current frame image of the debris flow monitoring video.

[0008] Furthermore, the expression for the spatial variation mask is as follows:

[0009]

[0010]

[0011]

[0012] in, For spatially varying masks, The x-coordinate is in pixels. The vertical coordinate is in pixels. For difference images, For adaptive threshold, For the current frame image, For the first frame, A pure background baseline image. From frame 1 to frame 2 At pixel position in frame The median of all pixel values. The total number of selected video frames without debris flow events. This is the scaling factor. In pixel coordinates Center, size Local window difference image The corresponding gray standard deviation, This represents the size of the sliding window in the local neighborhood.

[0013] The further beneficial effects mentioned above are as follows: This invention selects historical videos without debris flow events and calculates the median pixel by pixel in consecutive frames to construct a pure background baseline image, suppressing interference from temporary moving targets; it dynamically calculates an adaptive threshold based on the local neighborhood grayscale standard deviation, enabling differential segmentation to adapt to differences in illumination and texture changes in different regions; it solves the problem of poor robustness of fixed thresholds in complex scenes, realizes accurate identification of debris flow change areas by spatially varying masks, reduces false detections caused by factors such as illumination, and improves the recall rate of debris flow candidate regions.

[0014] Furthermore, the expression for the time-domain flow characteristic map is as follows:

[0015]

[0016]

[0017] in, For the temporal flow state feature map at pixel location The value at that location, For the Euclidean norm, For the optical flow field calculated by the dense optical flow network in The vector at that location, This represents the maximum optical flow amplitude of all pixels in the entire image. It is a constant. This represents the horizontal pixel displacement between the current frame and the previous frame. This represents the pixel displacement in the vertical direction between the current frame and the previous frame. For optical flow field, The horizontal optical flow component. This represents the optical flow component in the vertical direction.

[0018] The further beneficial effects mentioned above are as follows: This invention calculates the optical flow vector between the current frame image and the previous frame image through a dense optical flow network, captures the horizontal and vertical displacement of each pixel, extracts the optical flow amplitude for global normalization, maps the motion intensity to the [0,1] interval, and obtains a temporal flow feature map; it solves the interference of motion amplitude differences on attention weights under different debris flow scenarios, and keeps the temporal flow feature map and the actual physical motion intensity of debris flow in a consistent scale.

[0019] Further: obtaining the fused feature map specifically includes: The current frame image of the debris flow monitoring video is input into the feature extraction network, and the deep feature map is output through the backbone network. In the neck feature layer of the feature extraction network, parallel solid-phase sensing branches and liquid-phase sensing branches are set up. Based on the deep feature map, feature extraction is performed using 7×7 large kernel convolution through the solid phase sensing branch to obtain the features of debris flow rocks. Based on the deep feature map, feature extraction is performed using 3×3 small kernel convolution through the liquid phase sensing branch to obtain the mud water body features; Based on the characteristics of debris flow rocks and mud and water, spatial adaptive fusion weights are obtained through channel splicing and convolution. Based on spatial adaptive fusion weights, the features of debris flow rocks and mud and water bodies are weighted and fused to obtain a fused feature map.

[0020] Furthermore, the expression for the fused feature map is as follows:

[0021]

[0022] in, To fuse feature maps, For spatial adaptive fusion weights, Characteristics of debris flow rocks. Characteristics of muddy water bodies It is the Sigmoid activation function. It is a 1×1 convolutional layer. This is for channel splicing operations.

[0023] The further beneficial effects mentioned above are as follows: By setting up parallel solid-phase sensing branches and liquid-phase sensing branches, the present invention extracts the texture of rocks and the features of mud and water respectively, and adopts channel splicing and convolution to generate spatial adaptive fusion weights. It can select the solid-phase or liquid-phase features to focus on according to the image region, which enhances the feature expression ability of debris flow edges and textures, reduces fluid edge positioning error, and works in synergy with the flow state attention mask to improve the overall detection accuracy of debris flow.

[0024] Furthermore, the expression for the fluid attention mask is as follows:

[0025] in, For fluid attention masking, It is the Sigmoid activation function. For convolution operations, For splicing operations, For space size adjustment operations, For spatially varying masks, This is a time-domain flow characteristic diagram.

[0026] The further beneficial effects mentioned above are: the present invention uses a spatial variation mask... With time-domain flow characteristics The images are scaled and stitched together, and then mapped via a convolutional network to generate a fluid attention mask. This approach makes attention weights positively correlated with regions that exhibit significant apparent changes and motion, and negatively correlated with regions that exhibit only apparent changes and remain static. It solves the problem of misjudging static backgrounds caused by the lack of motion constraints in traditional attention methods, and achieves selective enhancement of dynamic regions of debris flows and suppression of static similar backgrounds, thereby improving the interpretability and cross-scene generalization performance of the model.

[0027] Furthermore, the expression for the reconstructed feature map is as follows:

[0028]

[0029] in, To reconstruct the feature map, For deep feature maps, These are learnable residual weight parameters. To enhance the feature map, To fuse feature maps, For element-wise multiplication, For fluid attention masking.

[0030] The further beneficial effects mentioned above are: the present invention preserves deep features through residual connections. Furthermore, a learnable weight adaptive control is introduced to enhance the feature injection intensity, which solves the problem that a fixed fusion ratio is difficult to adapt to different debris flow scenarios. This ensures that the original effective information is not weakened. At the same time, the static background area is suppressed by the attention mask, and the debris flow movement area is enhanced, making the reconstructed feature map semantically complete and the response accurate, thereby improving the accuracy of debris flow recognition.

[0031] This invention also provides a debris flow identification system based on background decoupling and flow state attention masking, comprising: The multi-frame input module is used to acquire video frame images of historical debris flow events and video frame images of debris flow monitoring videos. The background decoupling module is used to obtain the spatial variation mask in the current frame image by decoupling the background of the image sequence. The temporal flow feature extraction module is used to calculate the optical flow field based on the current frame image and the previous frame image through a differentiable dense optical flow network, extract the optical flow amplitude and normalize it to obtain the temporal flow feature map; The feature extraction and fusion module is used to input the current frame image into the feature extraction network to output a deep feature map, and to extract features of debris flow rocks and muddy water through parallel solid phase sensing branches and liquid phase sensing branches respectively, to obtain a fused feature map; The attention mask generation module is used to scale the spatially varied mask and the temporal fluid dynamic feature map to the same spatial size as the fused feature map, stitch them along the channel and input them into the attention generation network, and output the fluid attention mask. The feature reconstruction module is used to perform attention weighting based on the fluid attention mask and the fused feature map, and then perform residual fusion with the deep feature map to obtain the reconstructed feature map; The detection output module is used to input the reconstructed feature map into the detection head and output the bounding box and confidence score of the debris flow target.

[0032] The beneficial effects of this invention are as follows: 1. Suppress static similar background interference at the image feature level This invention generates an attention mask at the feature fusion layer by decoupling the background and constraining the temporal flow state, thereby reducing the response weight of static similar backgrounds at the feature level and thus reducing false positive interference in image processing. Compared with the baseline YOLOv8, the false positive rate of static backgrounds is significantly reduced, the recall rate of debris flow targets is significantly improved, and the attention heatmap intuitively shows that the response of static background regions is effectively suppressed and the features of debris flow regions are enhanced.

[0033] 2. Improve the cross-scene generalization of feature representations This method guides the network to focus on image regions with dynamic flow properties, reduces the influence of static appearance features of specific scenes on feature extraction, and enables the extracted features to better represent the motion attributes of the target, thereby improving the robustness and generalization ability of the algorithm to unknown scenes.

[0034] 3. Fusion of multimodal priors and deep networks By transforming the low-dimensional spatiotemporal physical priors obtained from traditional image processing algorithms into attention weights to guide the high-level feature reconstruction of deep networks, the feature reconstruction process is made physically interpretable and end-to-end optimized. With a slight increase in the number of parameters, the inference speed is significantly reduced, making it suitable for deployment in real-time monitoring devices at the edge.

[0035] 4. Enhancement effect of solid-liquid two-phase bibranching This invention designs a dual-branch structure to address the different visual characteristics of solid phases such as gravel and boulders in debris flows and liquid phases such as mud and water, thereby enhancing the edge and texture features of debris flows and significantly reducing fluid edge positioning errors. In synergy with subsequent attention masks, it improves the overall accuracy of debris flow identification and detection. Attached Figure Description

[0036] Figure 1 This is a schematic diagram of the debris flow identification method based on background decoupling and flow state attention mask. Detailed Implementation

[0037] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0038] Example 1 like Figure 1 The diagram shows a flowchart of a debris flow identification method based on background decoupling and flow state attention masking. This invention provides a debris flow identification method based on background decoupling and flow state attention masking, comprising the following steps: Acquire historical videos of debris flow events and debris flow monitoring videos, and obtain a debris flow spatial variation mask in the current frame image of the debris flow monitoring video by decoupling the image sequence background. Based on the current frame and the previous frame of the debris flow monitoring video, the optical flow field is calculated through a dense optical flow network, the optical flow amplitude is extracted and normalized, and a time-domain flow characteristic map is obtained. The current frame image of the debris flow monitoring video is input into the feature extraction network, which outputs a deep feature map. The features of debris flow rocks and muddy water are extracted through parallel solid phase sensing branches and liquid phase sensing branches, respectively, to obtain a fused feature map. The debris flow spatial variation mask and temporal flow feature map are scaled to the same spatial size as the fused feature map, and then stitched together along the channel before being input into the attention generation network to output the flow attention mask; Based on the fluid attention mask, attention weighting is performed on the fused feature map, and then residual fusion is performed on the deep feature map to obtain the reconstructed feature map, which is then input into the detection head to output the bounding box and confidence score of the debris flow target.

[0039] In a specific embodiment of the present invention, existing background modeling methods based on inter-frame difference or Gaussian mixture models often use a globally fixed threshold to segment the difference image. This is difficult to adapt to interferences such as changes in lighting, shadows, and swaying leaves in different mountain debris flow monitoring scenarios. Furthermore, mountain gullies often contain residual debris flow deposits and moist gully beds, which have high similarity to debris flow targets in the image color space. These static and dynamic interferences easily lead to misjudging areas of sudden lighting changes as potential debris flow areas, or missing debris flow areas with little difference from the static background. The present invention uses image sequence background decoupling to solve this technical problem. It acquires historical videos without debris flow events and debris flow monitoring videos, and through image sequence background decoupling, obtains a debris flow spatial change mask in the current frame image of the debris flow monitoring video, specifically including: Obtain historical videos of mudslide events and select several consecutive frames. To better display the continuous flow of mudslides, generally select more than 30 frames, up to 50 frames. Construct a pure background baseline image by calculating the median element-wise. The expression is as follows:

[0040] in, A pure background baseline image. The x-coordinate is in pixels. The vertical coordinate is in pixels. From frame 1 to frame 2 At pixel position in frame The median of all pixel values. In this embodiment, the total number of selected video frames without debris flow events is [number]. By constructing a pure background baseline image through element-wise median calculation, temporary noise such as birds and swaying leaves can be effectively suppressed, resulting in a stable background. Acquire the current frame image of the debris flow monitoring video. And calculate the baseline image with pure background. difference image Its expression is as follows:

[0041] in, This is a difference image, representing the current frame image. Baseline image with pure background At pixel position Color differences at the location; Adaptive threshold is dynamically calculated based on the local neighborhood gray-level standard deviation. and the difference image Binarization segmentation is performed to obtain the spatial variation mask of debris flow in the current frame image of the debris flow monitoring video. The expression of the spatial variation mask is as follows:

[0042]

[0043]

[0044]

[0045] in, For spatially varying masks, when ,but , indicating that the pixel An apparent change has occurred, indicating it belongs to a debris flow candidate region; otherwise, the value is 0. This is achieved through a spatial variation mask. It can initially exclude static backgrounds such as continuously stationary exposed rock masses and riverbeds, while avoiding missed or false detections caused by fixed thresholds, and reducing false detections of debris flows caused by sudden changes in lighting and local shadows. The x-coordinate is in pixels. The vertical coordinate is in pixels. For difference images, For adaptive threshold, For the current frame image, For the first frame, A pure background baseline image. From frame 1 to frame 2 At pixel position in frame The median of all pixel values. The total number of selected video frames without debris flow events. This is the scaling factor, with a value range of 1.5. 2.5, the present invention is acceptable , In pixel coordinates Center, size Local window difference image The corresponding gray standard deviation, For the local neighborhood sliding window size, an odd-sized window can be used, such as a 7×7 window.

[0046] In a specific embodiment of the present invention, existing debris flow monitoring methods that incorporate temporal information mostly perform logical filtering, such as tracking filtering, at the back end of the detection results. These methods only extract the foreground during the preprocessing stage and fail to reduce overfitting and false responses of the network to static, similar backgrounds during the image feature extraction stage. They also cannot perform dynamic enhancement and static suppression at the image feature level. Therefore, the present invention addresses this issue as follows: Acquire the current frame image of the debris flow monitoring video. And the previous frame image Two frames of images are input into a differentiable dense optical flow network, which calculates the optical flow of each pixel in the image from... Frame to Horizontal displacement of the frame and vertical displacement The optical flow vector is obtained. The entire optical flow field is denoted as In one embodiment of the present invention, the dense optical flow network consists of multiple convolutions and transposed convolutions, and can be trained end-to-end:

[0047] in, For all trainable parameters, including convolution kernel, bias, BN parameters, etc.; the optical flow field is differentiable and can participate end-to-end in the backpropagation of the subsequent loss function.

[0048] Extracting optical flow amplitude And find the maximum value of the optical flow amplitude in the entire image. Normalizing each pixel yields the following expression for the temporal flow feature map:

[0049] in, For the temporal flow state feature map at pixel location The value at that location, Let be a constant, set to a very small positive number to prevent the denominator from being zero and to ensure numerical stability. A value of can be taken as... After normalization, the range of values ​​for the time-domain flow characteristic map is: , characterization The motion intensity at a pixel location is close to 0 in stationary areas and close to 1 in areas of rapid motion.

[0050] The differentiable dense optical flow network of this invention can participate in the backpropagation of the subsequent loss function end-to-end, enabling optical flow features to be jointly optimized with the detection task. Therefore, the temporal flow feature map can characterize the pixel-level dense and scale-uniform motion intensity distribution, providing accurate motion priors for the flow attention mask, solving the problem of incomparable motion amplitudes in different debris flow scenarios, and improving the ability to distinguish between real debris flow motion areas and static interference areas.

[0051] In specific embodiments of the present invention, existing attention mechanisms or multi-branch feature extraction methods are mostly based on general spatial feature design, such as calculating global or local importance weights through channel attention or spatial attention, but do not consider the special motion characteristics of the solid-liquid two-phase mixture in debris flows, such as the rolling of solid rocks and the surging of liquid mud. The solid phases such as rocks and gravel in debris flows exhibit rough textures and obvious edges, requiring a large receptive field to effectively capture their contour and shape information; while the liquid phases such as mud and water exhibit smooth areas and reflective properties, requiring small kernel convolution to retain detailed textures. In response, the present invention sets up a solid phase perception branch and a liquid phase perception branch in the neck feature layer of the feature extraction network. The current frame image of the debris flow monitoring video is input into the feature extraction network, and a deep feature map is output. The features of debris flow rocks and mud and water are extracted by the parallel solid phase perception branch and liquid phase perception branch respectively to obtain a fused feature map, specifically including: The current frame image of the debris flow monitoring video Input to the feature extraction network, output deep feature maps via the backbone network. In this embodiment, CSPDarknet can be used as the backbone network, and the spatial size of the output deep feature map is 1 / 16 of the original image, with 256 channels. In the neck feature layer of the feature extraction network, parallel solid-phase sensing and liquid-phase sensing branches are set up. The solid-phase sensing branch uses a 7×7 large kernel convolution with a stride of 1 and 256 output channels to extract coarse-grained debris flow rock features. The focus is on the edge and shape information of gravel and boulders; the liquid phase sensing branch uses 3×3 small kernel convolution with a stride of 1 and 256 output channels to extract fine-grained mud-water features. Pay special attention to the smooth areas and reflective details of the mud and water; Based on the characteristics of debris flow rocks Characteristics of muddy water bodies Spatial adaptive fusion weights are obtained through channel concatenation and convolution. ; Based on spatial adaptive fusion weights, the features of debris flow rocks and muddy water are weighted and fused to obtain a fused feature map. The expression of the fused feature map is as follows:

[0052]

[0053] in, To fuse feature maps, For spatial adaptive fusion weights, Characteristics of debris flow rocks. Characteristics of muddy water bodies It is the Sigmoid activation function. It is a 1×1 convolutional layer. This is for channel splicing operations.

[0054] This invention achieves pixel-level adaptive fusion of solid and liquid features by setting up parallel solid-phase sensing branches and liquid-phase sensing branches. Through adaptive fusion weights, it can favor the solid-phase sensing branch in solid rock regions and the liquid-phase sensing branch in muddy water regions, thereby enhancing the ability to represent the edge and texture features of debris flows, reducing fluid edge positioning errors, and providing richer physical semantic features for subsequent flow attention masks.

[0055] In a specific embodiment of the present invention, in actual debris flow monitoring scenarios, there are often static backgrounds that are highly similar to debris flow targets in terms of appearance, color, and texture, i.e., static similar backgrounds, such as static deposits remaining in gullies, moist riverbeds, and exposed rock and soil slopes. Traditional single-frame image feature extraction networks rely only on spatial appearance information and lack the constraint of temporal motion information, making it easy to mistakenly identify such static similar backgrounds as debris flow targets, leading to an increased false alarm rate. At the same time, during feature extraction, deep networks tend to overfit the static appearance features of specific scenes in the training set, such as the rock texture of a specific gully or the slope reflection under specific lighting conditions, while ignoring the common dynamic motion features of the debris flow itself, resulting in a decrease in the model's feature representation ability in unseen scenes. To address this, the present invention uses a debris flow spatial variation mask. and time-domain flow characteristics Scale to fused feature map Using the same spatial dimensions and spliced ​​along the channels, an attention-based network is generated. Attention Generative Network The structure consists of a series of 3×3 convolutions, LeakyReLU, 1×1 convolutions, and a Sigmoid function, outputting a fluid attention mask. The expression for the fluid attention mask is as follows:

[0056] in, The fluid attention mask is a binary image of size H×W, used to identify candidate regions of apparent change. It is the Sigmoid activation function. For convolution operations, the input has 2 channels and the output has 1 channel. For the stitching operation, the two scaled single-channel images are stitched together to form a 2-channel feature map. For spatial dimension adjustment operations, bilinear interpolation is typically used to mask the spatial changes. and time-domain flow characteristics Scale to deep feature map Same size, For spatially varying masks, This is a time-domain flow characteristic diagram; In fluidic attention masks, for and For larger areas, i.e., actual debris flow areas, the output is close to 1; for but For regions close to 0, i.e., static similar backgrounds, the output is close to 0; for The region, i.e., the background region, has an output close to 0. This invention explicitly injects spatiotemporal physical priors into the feature reconstruction process through a fluid attention mask, combining physical interpretability with end-to-end optimization capabilities.

[0057] In a specific embodiment of the present invention, after obtaining the flow-state attention mask, the flow-state attention mask and the fused feature map are attention-weighted, and then residually fused with the deep feature map to obtain the reconstructed feature map, which is then input into the detection head to output the bounding box and confidence score of the debris flow target; the expression of the reconstructed feature map is as follows:

[0058]

[0059] in, To reconstruct the feature map, For deep feature maps, These are learnable residual weight parameters, which can be initialized to 0 to ensure that the original feature distribution is maintained during the initial training phase. To enhance the feature map, To fuse feature maps, For element-wise multiplication, For fluid attention masking.

[0060] In one embodiment of this invention, the detection head can be a decoupled detection head from the YOLO series, including a classification branch and a regression branch, which output the target category confidence and bounding box coordinate offset corresponding to each anchor point or grid, respectively. After non-maximum suppression, the final bounding box and confidence of the debris flow target are output. This invention reconstructs feature maps by constructing a flow-state attention mask, achieving dynamic enhancement and static suppression at the feature level while maintaining the integrity of the original semantic information, thus improving the accuracy of debris flow recognition and detection and reducing the false alarm rate of statically similar backgrounds. Example 2 Based on Embodiment 1, the present invention also provides a debris flow identification system based on background decoupling and flow state attention masking, which can execute the debris flow identification method based on background decoupling and flow state attention masking of Embodiment 1, including: The multi-frame input module is used to acquire video frame images of historical debris flow events and video frame images of debris flow monitoring videos. The background decoupling module is used to obtain the spatial variation mask in the current frame image by decoupling the background of the image sequence. The temporal flow feature extraction module is used to calculate the optical flow field based on the current frame image and the previous frame image through a differentiable dense optical flow network, extract the optical flow amplitude and normalize it to obtain the temporal flow feature map; The feature extraction and fusion module is used to input the current frame image into the feature extraction network to output a deep feature map, and to extract features of debris flow rocks and muddy water through parallel solid phase sensing branches and liquid phase sensing branches respectively, to obtain a fused feature map; The attention mask generation module is used to scale the spatially varied mask and the temporal fluid dynamic feature map to the same spatial size as the fused feature map, stitch them along the channel and input them into the attention generation network, and output the fluid attention mask. The feature reconstruction module is used to perform attention weighting based on the fluid attention mask and the fused feature map, and then perform residual fusion with the deep feature map to obtain the reconstructed feature map; The detection output module is used to input the reconstructed feature map into the detection head and output the bounding box and confidence score of the debris flow target.

[0061] The beneficial effects of this invention are as follows: The debris flow recognition system based on background decoupling and flow-state attention mask of this invention generates a flow-state attention mask at the image feature level, deeply fusing the spatial variation mask with the temporal flow-state feature map. This effectively suppresses the response of static similar backgrounds at the feature level. Compared with baseline detection systems that rely solely on spatial appearance information, it reduces the false alarm rate of static backgrounds such as wet riverbeds and exposed soil, while improving the recall rate of debris flow targets. The system guides the feature extraction network to focus on regions with dynamic flow-state attributes, weakening the influence of static appearance features in specific scenes, and enabling the extracted features to better represent the common motion of targets, thus improving generalization ability. In addition, the solid-liquid two-phase dual-branch design adaptively fuses the different visual characteristics of debris flow rocks and muddy water, enhancing the recognition of debris flow edge and texture features and reducing the error in fluid edge localization. The overall computational load of this invention increases only slightly, the number of model parameters increases only slightly, and the inference speed meets real-time requirements, making it suitable for deployment in edge monitoring equipment.

Claims

1. A debris flow identification method based on background decoupling and flow state attention mask, characterized in that, Includes the following steps: Acquire historical videos of debris flow events and debris flow monitoring videos, and obtain a debris flow spatial variation mask in the current frame image of the debris flow monitoring video by decoupling the image sequence background. Based on the current frame and the previous frame of the debris flow monitoring video, the optical flow field is calculated through a dense optical flow network, the optical flow amplitude is extracted and normalized, and a time-domain flow characteristic map is obtained. The current frame image of the debris flow monitoring video is input into the feature extraction network, which outputs a deep feature map. The features of debris flow rocks and muddy water are extracted through parallel solid phase sensing branches and liquid phase sensing branches, respectively, to obtain a fused feature map. The debris flow spatial variation mask and temporal flow feature map are scaled to the same spatial size as the fused feature map, and then stitched together along the channel before being input into the attention generation network to output the flow attention mask; Based on the fluid attention mask, attention weighting is performed with the fused feature map, and then residual fusion is performed with the deep feature map to obtain the reconstructed feature map, which is then input into the detection head to output the bounding box and confidence of the debris flow target; The obtained fused feature map specifically includes: The current frame image of the debris flow monitoring video is input into the feature extraction network, and the deep feature map is output through the backbone network. In the neck feature layer of the feature extraction network, parallel solid-phase sensing branches and liquid-phase sensing branches are set up. Based on the deep feature map, feature extraction is performed using 7×7 large kernel convolution through the solid phase sensing branch to obtain the features of debris flow rocks. Based on the deep feature map, feature extraction is performed using 3×3 small kernel convolution through the liquid phase sensing branch to obtain the mud water body features; Based on the characteristics of debris flow rocks and mud and water, spatial adaptive fusion weights are obtained through channel splicing and convolution. Based on spatial adaptive fusion weights, the features of debris flow rocks and mud and water bodies are weighted and fused to obtain a fused feature map. The expression for the fluid attention mask is as follows: in, For fluid attention masking, It is the Sigmoid activation function. For convolution operations, For splicing operations, For space size adjustment operations, For spatially varying masks, This is a time-domain flow characteristic diagram.

2. The debris flow identification method based on background decoupling and flow state attention mask according to claim 1, characterized in that, The debris flow spatial variation mask obtained from the current frame image of the debris flow monitoring video specifically includes: Acquire historical videos of events without mudslides, select several consecutive frames, and construct a pure background baseline image by calculating the median value element by element. Acquire the current frame image of the debris flow monitoring video and calculate the difference image with the pure background baseline image; An adaptive threshold is dynamically calculated based on the local neighborhood grayscale standard deviation, and the differential image is binarized and segmented to obtain a debris flow spatial variation mask in the current frame image of the debris flow monitoring video.

3. The debris flow identification method based on background decoupling and flow state attention mask according to claim 2, characterized in that, The expression for the spatial variation mask is as follows: in, For spatially varying masks, The x-coordinate is in pixels. The vertical coordinate is in pixels. For difference images, For adaptive threshold, For the current frame image, For the first frame, A pure background baseline image. From frame 1 to frame 2 At pixel position in frame The median of all pixel values. The total number of selected video frames without debris flow events. This is the scaling factor. In pixel coordinates Center, size Local window difference image The corresponding gray standard deviation, This represents the size of the sliding window in the local neighborhood.

4. The debris flow identification method based on background decoupling and flow state attention mask according to claim 1, characterized in that, The expression for the time-domain flow state feature map is as follows: in, For the temporal flow state feature map at pixel location The value at that location, For the Euclidean norm, For the optical flow field calculated by the dense optical flow network in The vector at that location, This represents the maximum optical flow amplitude of all pixels in the entire image. It is a constant. This represents the horizontal pixel displacement between the current frame and the previous frame. This represents the pixel displacement in the vertical direction between the current frame and the previous frame. For optical flow field, The horizontal optical flow component. This represents the optical flow component in the vertical direction.

5. The debris flow identification method based on background decoupling and flow state attention mask according to claim 1, characterized in that, The expression for the fused feature map is as follows: in, To fuse feature maps, For spatial adaptive fusion weights, Characteristics of debris flow rocks. Characteristics of muddy water bodies It is the Sigmoid activation function. It is a 1×1 convolutional layer. This is for channel splicing operations.

6. The debris flow identification method based on background decoupling and flow state attention mask according to claim 1, characterized in that, The expression for the reconstructed feature map is as follows: in, To reconstruct the feature map, For deep feature maps, These are learnable residual weight parameters. To enhance the feature map, To fuse feature maps, For element-wise multiplication, For fluid attention masking.

7. A debris flow identification system based on background decoupling and flow state attention masking, used to execute the debris flow identification method based on background decoupling and flow state attention masking as described in any one of claims 1-6, characterized in that, include: The multi-frame input module is used to acquire video frame images of historical debris flow events and video frame images of debris flow monitoring videos. The background decoupling module is used to obtain the spatial variation mask in the current frame image by decoupling the background of the image sequence. The temporal flow feature extraction module is used to calculate the optical flow field based on the current frame image and the previous frame image through a differentiable dense optical flow network, extract the optical flow amplitude and normalize it to obtain the temporal flow feature map; The feature extraction and fusion module is used to input the current frame image into the feature extraction network to output a deep feature map, and to extract features of debris flow rocks and muddy water through parallel solid phase sensing branches and liquid phase sensing branches respectively, to obtain a fused feature map; The attention mask generation module is used to scale the spatially varied mask and the temporal fluid dynamic feature map to the same spatial size as the fused feature map, stitch them along the channel and input them into the attention generation network, and output the fluid attention mask. The feature reconstruction module is used to perform attention weighting based on the fluid attention mask and the fused feature map, and then perform residual fusion with the deep feature map to obtain the reconstructed feature map; The detection output module is used to input the reconstructed feature map into the detection head and output the bounding box and confidence score of the debris flow target.

Citation Information

Patent Citations

  • Optical flow estimation method, device and equipment

    CN114677412A

  • Debris flow disaster early warning method and emergency response system based on big data analysis

    CN119479242A