A method and device for flood inundation monitoring that integrates optical and SAR technologies across multiple levels and modes.

By employing cross-modal adaptive interactive fusion, frequency domain adaptive dual-stream fusion, and hierarchical multi-interactive fusion modules, the problems of insufficient features and blurred boundaries in existing flood inundation monitoring methods under complex scenarios are solved, achieving high-precision and robust flood inundation monitoring.

CN121482612BActive Publication Date: 2026-04-17WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN UNIV
Filing Date
2026-01-07
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing flood inundation monitoring methods suffer from insufficient single-modal features, blurred boundaries, false detections, and missed detections in complex scenarios, making it difficult to achieve high-precision and robust flood monitoring.

Method used

By employing cross-modal adaptive interactive fusion, frequency domain adaptive dual-stream fusion, and hierarchical multi-interactive fusion modules, and through the complementary and modeling of depth features of optical and SAR images, the accuracy of global consistency and local details is enhanced.

Benefits of technology

It significantly improves the accuracy and robustness of flood monitoring, enhances extraction performance in complex scenarios, and provides reliable data support for flood inundation monitoring and emergency response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482612B_ABST
    Figure CN121482612B_ABST
Patent Text Reader

Abstract

This invention relates to a flood inundation monitoring method and apparatus based on multi-level cross-modal fusion of optical and SAR images. The method first acquires and preprocesses pre-disaster optical images and post-disaster SAR images to construct training samples. Through three core modules—cross-modal adaptive interactive fusion, frequency-domain adaptive dual-stream fusion, and hierarchical multi-interactive fusion—it achieves complementary deep features and modeling of optical and SAR images, improving feature robustness and boundary accuracy. The cross-modal adaptive interactive fusion module enhances inter-modal complementarity, the frequency-domain adaptive dual-stream fusion module balances consistency and boundary accuracy through high- and low-frequency modeling, and the hierarchical multi-interactive fusion module considers both global semantics and local details, further improving the model's adaptability to different scenarios. This invention can effectively improve the accuracy of flood monitoring, enhance extraction performance in complex scenarios, and provide reliable methodological support for flood inundation monitoring, emergency response, and disaster assessment, with broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent interpretation and multimodal fusion technology of remote sensing images, specifically relating to a flood inundation monitoring method and device based on multimodal remote sensing image fusion. Background Technology

[0002] Rapidly and accurately determining the extent of flood inundation after a flood is crucial for disaster assessment, emergency response, and disaster prevention and mitigation. With the rapid development of remote sensing technology, its advantages of wide coverage and high spatiotemporal resolution have enabled it to play an irreplaceable role in natural disaster research such as landslide evolution, earthquake monitoring, and flood monitoring, and it is gradually becoming an important technical means for flood monitoring.

[0003] Optical remote sensing imagery provides detailed texture and spectral information, offering advantages in characterizing land cover types and water body boundaries. However, floods are often accompanied by heavy rainfall and thick cloud cover, making post-disaster optical imagery susceptible to obstruction and limiting its applicability. Synthetic Aperture Radar (SAR) employs an active microwave transmission and reception mechanism, independent of sunlight, and can penetrate clouds and rain curtains. This allows it to stably acquire surface echo signals and characterize flood inundation areas even under complex weather conditions, providing all-weather, all-time imaging capabilities. However, it lacks the fine-grained texture and semantic features provided by optical imagery and is prone to misjudgment when water bodies are confused with dark-colored features. A single modality is insufficient to comprehensively and accurately characterize flood features in complex flood scenarios, highlighting the importance of fusing optical and SAR technologies.

[0004] In recent years, deep learning technology has become a core tool in remote sensing image analysis, achieving breakthroughs in tasks such as image fusion, semantic segmentation, and change detection. Multimodal fusion methods based on convolutional neural networks (CNNs) achieve joint modeling of local spatial features through feature concatenation, weighting, or sharing of convolutional kernels. Meanwhile, the Transformer architecture, relying on its self-attention mechanism, can capture long-range dependencies, demonstrating superior global modeling capabilities in cross-modal interactions. Simultaneously, frequency domain and multi-scale feature fusion have also gained increasing attention, effectively improving boundary characterization and overall consistency through frequency decomposition and hierarchical complementarity. Existing work has made some progress in improving the accuracy of flood monitoring.

[0005] However, while single-modal methods are widely used in flood monitoring, optical imagery is easily obscured by clouds and rain after a disaster, and SAR imagery, although possessing all-weather advantages, lacks fine-grained texture and semantic information. This often leads to problems such as blurred boundaries, false detections, or missed detections in complex flood scenarios. Although multimodal methods improve the accuracy and robustness of flood monitoring, they still lack effective redundancy suppression and complementary modeling mechanisms, especially in joint modeling at the frequency, scale, and semantic levels. Cross-level feature interaction is also limited, making it difficult to balance global consistency and local boundary accuracy, thus restricting the model's generalization ability.

[0006] These shortcomings highlight the necessity of developing more efficient cross-modal fusion strategies. Therefore, it is necessary to propose a novel optical and SAR multi-level cross-modal fusion network to achieve flood inundation monitoring with high accuracy, robustness and lightweight, in order to solve problems such as insufficient single-modal features, blurred boundaries, false detection and missed detection. Summary of the Invention

[0007] This invention addresses the problems of insufficient single-modal features, blurred boundaries, false detections, and missed detections in existing flood inundation monitoring methods under complex scenarios. It provides a flood inundation monitoring method and device based on multi-level cross-modal fusion of optical and SAR images. The method constructs training samples by acquiring and preprocessing pre-disaster optical images and post-disaster SAR images. It employs cross-modal adaptive interactive fusion, frequency domain adaptive dual-stream fusion, and hierarchical multi-interactive fusion modules to achieve complementary deep features and modeling of optical and SAR images, improving feature robustness and boundary accuracy. Each module enhances global consistency and the accuracy of local details through adaptive fusion, collaborative modeling, and cross-level interaction. This invention effectively improves the accuracy and robustness of flood monitoring, enhances extraction performance in complex scenarios, provides reliable data support for flood inundation monitoring and emergency response, and has broad application prospects.

[0008] The technical solution of this invention is: a flood inundation monitoring method that integrates optical and SAR multi-level cross-modal fusion, comprising the following steps:

[0009] Acquire optical images before the flood and SAR images after the flood;

[0010] The optical image and SAR image are encoded separately using the backbone network to obtain multi-scale optical features and SAR features.

[0011] A cross-modal adaptive interactive fusion module is constructed, and multi-scale dynamic spatial downsampling is introduced to obtain optical and SAR features after cross-modal interaction. Modal adaptive weighted fusion is introduced to output the final control features.

[0012] A frequency-domain adaptive dual-stream fusion module is constructed. First, the control features are divided into high-resolution features and low-resolution features, and cross-scale information alignment is achieved in a unified channel space. Then, filtering and corresponding resampling operations are performed to obtain high-level features and low-level features.

[0013] A hierarchical multi-interaction fusion module is constructed, which interacts with high-level features and low-level features through various complementary forms to obtain multi-level features that fuse global semantic information and local details. Through these multi-level features, a high-precision flood inundation map is generated.

[0014] Furthermore, a pyramid visual transformer is used as the backbone network.

[0015] Furthermore, obtaining the optical and SAR features after cross-modal interaction specifically includes:

[0016] First, a bidirectional cross-attention mechanism is employed to achieve information interaction and fusion between optical and SAR features, fully capturing complex nonlinear information from different modalities; the optical branch is obtained through a linear layer. After multi-scale dynamic spatial downsampling, the following was obtained. and SAR branch is obtained through linear layers After multi-scale dynamic spatial downsampling, the following was obtained. and Then, the query vector of the optical branch. Key-value pairs with SAR branches Matching is performed to capture complementary information about the optical branch of SAR features; conversely, the query vector of the SAR branch is used to... Key-value pairs with optical branches Interact with the system to introduce fine-grained representations of optical features;

[0017]

[0018]

[0019] in, This represents the dimension of the feature vector in each attention head. and These represent the optical and SAR features after cross-modal interaction, respectively. Indicates scale.

[0020] Furthermore, modal adaptive weighted fusion specifically includes:

[0021] The final modulation feature is obtained by dynamically assigning weights to different modal features at each position in space using a learnable gating function. :

[0022]

[0023]

[0024] in, and These represent the optical and SAR features after cross-modal interaction, respectively. Indicates scale. For learnable gating functions, This indicates a splicing operation. Represents the normalization function. This indicates element-wise multiplication. Indicates adaptive weights, and Let represent the adaptive weights of optical features and SAR features at each position, respectively, and satisfy . .

[0025] Furthermore, the modulation features obtained from the cross-modal adaptive interactive fusion module are categorized into high-resolution features based on their relative resolution. With low-resolution features First, high-resolution features With low-resolution features The features are compressed into a uniform channel space using 1×1 convolution; residual information is extracted using global average pooling (GAP); and position-adaptive weights are generated using the sigmoid function to finally obtain aligned high-resolution features. With low-resolution features :

[0026]

[0027]

[0028] in Indicates Sigmoid activation. ⊙ indicates global average pooling, and ⊙ indicates position-wise weighted pooling. and These represent the compressed high-resolution and low-resolution features, respectively.

[0029] Furthermore, the specific implementation methods for obtaining high-level and low-level features are as follows:

[0030] Adaptive low-pass and adaptive high-pass filters are used to generate position-adaptive low-pass and high-pass filter kernels respectively, thereby achieving joint modeling of low-frequency global consistency and high-frequency boundary details;

[0031]

[0032]

[0033] in, and These represent an adaptive low-pass filter and an adaptive high-pass filter, respectively. The upsampling operation maps low-resolution features to a space consistent with the high-resolution one. This indicates a normalization operation, ensuring the numerical stability of the convolution kernel;

[0034] Finally, low-resolution feature intermediate variables Content-aware upsampling is used to obtain low-frequency consistent low-level features. High-resolution feature intermediate variables The low-frequency components are separated to obtain high-frequency residuals, and these residuals are then back-injected to form high-level features that enhance details. :

[0035]

[0036]

[0037] in, This indicates a resampling operation under the constraints of an adaptive low-pass or high-pass filter, used to generate low-frequency consistency and high-frequency residual information.

[0038] Furthermore, the specific processing procedure of the multi-level interactive fusion module is as follows:

[0039] High-level features obtained from the Frequency Domain Adaptive Dual Stream Fusion Module (FADF) With low-level features Mapped to a unified channel dimension, three complementary interaction forms are then introduced to obtain interactive features, where addition is used to model joint information to capture the commonalities of features across levels, differencing is used to emphasize hierarchical differences to enhance complementarity, and multiplication is used to characterize the correlation between features to strengthen consistent expression.

[0040]

[0041]

[0042]

[0043] in, and These represent the high-level and low-level features of the unified channel dimension, respectively. This represents position-by-position weighted operations. , and These represent the results of adding high-level features and low-level features, the results of absolute difference, and the results of multiplication, respectively.

[0044] The multi-way interaction features are element-wise accumulated, multiplied, and concatenated to form a comprehensive feature. A gating mechanism is introduced to adaptively control different interaction branches, resulting in multi-level features that integrate global semantic information and local details.

[0045]

[0046]

[0047] in, and These represent element-wise addition, multiplication, and concatenation operations, respectively. Represented as a non-linear activation function, Represented as a gating weight function, it consists of a 1×1 convolution, batch normalization, and sigmoid activation. It represents a comprehensive characteristic of element-wise addition, multiplication, or concatenation. This represents the multi-level features after empowerment.

[0048] Furthermore, a high-precision flood inundation map is generated using residual blocks, as detailed below:

[0049]

[0050] in, Indicates the first j High-level features at various scales This indicates a depth-separable residual block. This represents a flooded area.

[0051] Secondly, the present invention also provides a flood inundation monitoring device that integrates optical and SAR multi-level cross-modal fusion, including a processor and a memory. The memory is used to store program instructions, and the processor is used to call the program instructions in the memory to execute the flood inundation monitoring method that integrates optical and SAR multi-level cross-modal fusion as described in the above technical solution.

[0052] Thirdly, the present invention also provides a computer-readable storage medium, including a readable storage medium on which a computer program is stored, wherein when the computer program is executed, it implements the flood inundation monitoring method of optical and SAR multi-level cross-modal fusion as described in the above technical solution.

[0053] The beneficial effects of the technical solution provided by this invention are as follows:

[0054] (1) Existing flood inundation monitoring methods often suffer from insufficient single-modal features and blurred boundaries when dealing with complex scenarios, making it difficult to effectively cope with diverse flood features. This invention constructs a cross-modal adaptive interactive fusion module to achieve complementary deep features and modeling of optical images and SAR images, effectively enhancing feature robustness and boundary accuracy, and enabling more accurate capture of details and boundary information in complex flood scenarios.

[0055] (2) Existing flood monitoring methods often neglect the imbalance between frequency domains after modal fusion. The frequency domain adaptive dual-stream fusion module proposed in this invention achieves high- and low-frequency synergistic enhancement of cross-scale features through dual-stream modeling with low-pass and high-pass filtering. At the same time, it uses global residual information to dynamically guide the weight allocation at each location, enabling the model to adaptively balance the overall consistency of low-frequency features with the detailed representation of high-frequency features.

[0056] (3) To address the problem of multi-scale flood feature fusion, this invention designs a hierarchical multi-interaction fusion module. By using addition, difference and multiplication to model the jointness, difference and correlation of cross-level features, it can take into account both global semantics and local details, improve the model's cross-scene adaptability, enhance the fine characterization of boundary areas, and significantly improve the recognition effect of flood-inundated areas. Attached Figure Description

[0057] Figure 1 This is a diagram illustrating the overall network structure of an embodiment of the present invention.

[0058] Figure 2 This is a schematic diagram of the cross-modal adaptive interactive fusion module (MAIF) in an embodiment of the present invention;

[0059] Figure 3 This is a schematic diagram of the Frequency Domain Adaptive Dual-Stream Fusion Module (FADF) in an embodiment of the present invention;

[0060] Figure 4 This is a schematic diagram of the Hierarchical Multi-Interaction Fusion Module (HMIF) in an embodiment of the present invention;

[0061] Figure 5 This is a visualization of flood prediction results in an embodiment of the present invention. Detailed Implementation

[0062] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings.

[0063] This invention addresses the problems of boundary ambiguity, false detection, and missed detection in existing flood inundation monitoring methods under complex scenarios. It proposes a multi-level cross-modal fusion method for flood inundation monitoring based on optical and SAR imagery, comprising:

[0064] Data acquisition and preprocessing of remote sensing images before and after the flood were carried out, and the images were cropped, normalized and labeled to form a training sample dataset.

[0065] The optical image and SAR image are encoded separately using the backbone network to obtain multi-scale optical features and SAR features.

[0066] This invention constructs a cross-modal adaptive interactive fusion module, introduces multi-scale dynamic spatial downsampling to obtain optical and SAR features after cross-modal interaction, and introduces modal adaptive weighted fusion to output the final control features. This invention achieves deep complementarity and modeling of optical and SAR image features through a multi-level cross-modal fusion framework, enhancing the robustness and accuracy of the model under different flood scenarios.

[0067] A frequency-domain adaptive dual-stream fusion module is constructed. First, the control features are divided into high-resolution features and low-resolution features, and cross-scale information alignment is achieved in a unified channel space. Then, filtering and corresponding resampling operations are performed to obtain high-level features and low-level features. This invention uses a frequency-domain adaptive dual-stream fusion module to perform high- and low-frequency collaborative modeling and residual control, balancing overall consistency and local boundary accuracy, and further optimizing the recognition effect in flood areas.

[0068] A hierarchical multi-interaction fusion module is constructed to interact with high-level and low-level features through various complementary methods to obtain multi-level features that integrate global semantic information and local details. These multi-level features are then used to generate high-precision flood inundation maps. This invention designs a hierarchical multi-interaction fusion module that models the jointness, differences, and correlations of cross-level features using addition, difference, and multiplication methods, taking into account both global semantics and local details, thereby further improving adaptability to different flood scenarios. Ultimately, this module obtains multi-level features that integrate global semantic information and local details. Using these features, the model can accurately identify flood change areas and generate high-precision flood inundation maps, thus providing reliable data support for emergency response and disaster assessment.

[0069] like Figure 1 As shown, the specific steps of this embodiment of the invention are as follows:

[0070] We propose an end-to-end, multi-level cross-modal fusion framework, MFHNet, to address the information redundancy and insufficient complementarity issues arising from the fusion of optical and SAR images in flood inundation monitoring. This framework introduces three key modules: a Modality-Adaptive Interaction Fusion Module (MAIF), a Frequency-Adaptive Dual-stream Fusion Module (FADF), and a Hierarchical Multi-Interaction Fusion Module (HMIF). These modules achieve multi-level complementary modeling at the modality, frequency domain, and feature levels, significantly improving the accuracy and robustness of flood inundation monitoring. Given the significant differences between optical and SAR remote sensing images in their imaging mechanisms and acquisition timelines, this invention employs a Pyramid Vision Transformer (PVT) backbone network to effectively characterize their complementary properties. This network, with its self-attention mechanism, can model global dependencies and capture long-range interactions between cross-modal features, thereby enhancing the model's ability to jointly represent global semantic consistency and local detailed structure.

[0071] (1)

[0072] (2)

[0073] in, Represents feature encoding operation. and These represent encoded optical and SAR remote sensing images, respectively.

[0074] To address the issues of insufficient cross-modal feature alignment and rigid weight allocation, this invention designs a cross-modal adaptive interactive fusion module. This module, by introducing a cross-modal attention mechanism and a modal adaptive weighting strategy, achieves deep interaction and dynamic fusion of optical and SAR features, effectively alleviating the problem of insufficient inter-modal complementarity.

[0075] (3)

[0076] However, the frequency domain distribution of optical and SAR fused features still suffers from an imbalance between low-frequency ambiguity and high-frequency noise. To address this, we propose a Frequency Domain Adaptive Dual-Stream Fusion Module (FADF), which combines low-pass and high-pass filtering for dual-stream modeling and introduces an adaptive residual control mechanism (ARM) to dynamically balance the overall consistency of low frequencies with the boundary details of high frequencies, thereby improving the completeness and discriminativeness of feature representation. Furthermore, considering the differences in semantic and detail representation among features at different fusion levels, directly stacking features from adjacent levels often leads to an imbalance between local texture and global semantics. To solve this problem, this invention further proposes a Hierarchical Multi-Interaction Fusion Module (HMIF). This module explicitly models the jointness, differences, and correlations of features from adjacent levels through three complementary interaction methods: addition, difference, and multiplication. Guided by a gating mechanism, it achieves adaptive fusion, thus balancing global consistency and local fine detail in cross-level feature integration, improving the accuracy and stability of flood inundation range extraction.

[0077] (4)

[0078] (5)

[0079] Finally, in the decoding stage, we employ a lightweight multi-level decoder. Multi-level features are fused through linear projection and alignment mechanisms, and the flood inundation range is output via a prediction head, achieving a unified characterization of global semantics and local boundaries.

[0080] Step a, construct a cross-modal adaptive interaction fusion module, such as Figure 2 As shown;

[0081] This invention designs a cross-modal adaptive interactive fusion module (MAIF), which achieves deep fusion of optical and SAR features by combining cross-modal attention and modal adaptive weighting with fine-grained information from pre-disaster optical data and robust observations from post-disaster SAR data, thereby enabling high-precision extraction of flood inundation range.

[0082] In the MAIF module, based on the multi-scale features extracted by the Pyramid Vision Transformer (PVT) backbone network, optical and SAR modal features are first mapped to queries and keys / values, respectively, and information interaction is achieved through a bidirectional cross-attention mechanism. To fully utilize the complementary characteristics of optical and SAR remote sensing images, we introduce Multi-Scale Dynamic Reduction (MSDR) and Modality-Adaptive Weighting (MAW) on the basis of cross-modal interaction, improving fusion efficiency while ensuring robustness of information representation. Specifically, we first employ a bidirectional cross-attention mechanism to achieve information interaction and fusion between optical and SAR features, fully capturing complex nonlinear information from different modalities. The optical branch is obtained through a linear layer. After multi-scale dynamic spatial downsampling, the following was obtained. and SAR branch is obtained through linear layers After multi-scale dynamic spatial downsampling, the following was obtained. and Then, the query vector of the optical branch. Key-value pairs with SAR branches Matching is performed to capture complementary information about the optical branch of SAR features. Conversely, the query vector of the SAR branch... Key-value pairs with optical branches Interact with the system to introduce fine-grained expressions of optical features.

[0083] (6)

[0084] (7)

[0085] Among them, when adopting a multi-attention strategy, This represents the dimension of the feature vector in each attention head. This process allows optical features to be combined with SAR structured features, while SAR features can supplement the optical spectral and texture priors, thereby achieving bidirectional information enhancement. and These represent the optical and SAR features after cross-modal interaction, respectively.

[0086] However, in high-resolution remote sensing imagery, the computational complexity of cross-modal attention typically increases quadratically with sequence length, i.e. ,in This represents spatial dimension. Such complexity is unacceptable in large-scale remote sensing scenarios, thus requiring a reduction in computational overhead while maintaining information representation capabilities. To address this, this invention introduces the MSDR strategy in the MAIF module to compactly model key-value branches. Unlike the single-scale convolutional downsampling commonly used in existing methods, MSDR combines multi-scale receptive fields through 1×1, 3×3, and 5×5 depthwise convolutions, resulting in features with stronger spatial context modeling capabilities.

[0087] After completing the cross-modal feature interaction, in order to further achieve adaptive control of the contribution of optical and SAR features, this invention uses the learnable gating function of MAW to dynamically assign weights to different modal features at each position in space. This enables the model to automatically balance the fine-grained texture priors provided by the optical modality and the robust water body response of the SAR modality based on scene semantic features and observation conditions, thereby improving the discriminative power of flood-inundated area extraction.

[0088] (8)

[0089] (9)

[0090] in, For learnable gating functions, This indicates a splicing operation. This indicates element-wise multiplication. and Let represent the adaptive weights of optical and SAR features at each position, respectively, and satisfy . .

[0091] Step b, construct a frequency-domain adaptive dual-stream fusion module, such as Figure 3 As shown;

[0092] This invention proposes a Frequency-Adaptive Dual-stream Fusion (FADF) module, which achieves high- and low-frequency synergistic enhancement of cross-scale features through dual-stream modeling using low-pass and high-pass filtering. Simultaneously, an Adaptive Residual Modulation (ARM) mechanism is introduced during the feature compression stage, dynamically guiding position-by-position weight allocation using global residual information, enabling the model to adaptively balance overall low-frequency consistency with high-frequency detail representation.

[0093] In the FADF module, in order to achieve cross-scale information alignment in a unified channel space, the information obtained from the MAIF module is... Based on relative resolution size, it is divided into high-resolution features. With low-resolution features First, high-resolution features With low-resolution features By compressing the data into a unified channel space using 1×1 convolution, and considering that in the fusion of optical and SAR features, features at different scales often contribute unevenly in complex scenes, simply adopting a uniform fusion method can easily weaken the effective utilization of key information. Therefore, this invention employs ARM (Average Array Pooling), extracting residual information through Global Average Pooling (GAP) and using the Sigmoid function to generate location-adaptive weights. This dynamically highlights discriminative features and suppresses redundant interference, enabling the model to better adapt to the complex feature distribution in flood scenarios.

[0094] (10)

[0095] (11)

[0096] in Indicates Sigmoid activation. This indicates global adaptive pooling, and ⊙ indicates position-wise weighted pooling. and These represent the compressed high-resolution and low-resolution features, respectively.

[0097] To simultaneously characterize the overall low-frequency structure and high-frequency detail boundaries, we use an Adaptive Low-Pass Filter (ALPF) and an Adaptive High-Pass Filter (AHPF) to generate position-adaptive low-pass and high-pass filter kernels, respectively. This enables joint modeling of low-frequency global consistency and high-frequency boundary details. Stable kernel weights are obtained through normalization, thereby improving the ability to collaboratively represent global and local information.

[0098] (12)

[0099] (13)

[0100] in, and These represent the adaptive low-pass filter generation module and the adaptive high-pass filter generation module, respectively. The upsampling operation maps low-resolution features to a space consistent with the high-resolution one. This indicates a normalization operation, ensuring the numerical stability of the convolution kernel.

[0101] Finally, content-aware upsampling is performed on low-resolution features to obtain low-frequency consistency. High-resolution features separate low-frequency components to obtain high-frequency residuals, and then use residual backinjection to enhance details. .

[0102] (14)

[0103] (15)

[0104] in, This represents a resampling operation under adaptive low / high-pass filter constraints, used to generate low-frequency consistency and high-frequency residual information.

[0105] By using cross-scale modeling that complements high and low frequencies, the fused features are enhanced at both the global and local levels. This not only ensures the integrity of the flood extent representation but also improves the smoothness and delicacy of boundary transitions, thus providing more discriminative feature support for flood detection.

[0106] Step c, construct a hierarchical multi-interaction integration module, such as Figure 4 As shown;

[0107] This invention designs a Hierarchical Multi-InteractionFusion Module (HMIF), which explicitly models the jointness, difference, and correlation of features at adjacent levels through three complementary interaction methods: addition, difference, and multiplication. Under the guidance of a gating mechanism, it achieves adaptive fusion, thereby taking into account both global semantics and local details in cross-level feature integration, and improving the precision and stability of flood inundation range extraction.

[0108] In the HMIF module, the high-level features obtained from FADF are first processed. With low-level features Mapping to a unified channel dimension mitigates fusion bias caused by dimensional inconsistencies. Three complementary interaction methods are then introduced: addition models joint information to capture commonalities across layers; differencing emphasizes layer differences to enhance complementarity; and multiplication characterizes correlations between features to strengthen consistent representation.

[0109] (16)

[0110] (17)

[0111] (18)

[0112] in, and These represent the high-level and low-level features of the unified channel dimension, respectively. This represents position-by-position weighted operations. , and These represent the results of adding high-level features and low-level features, the results of absolute difference, and the results of multiplication, respectively.

[0113] To unify the modeling of information from different interaction perspectives, the multi-way interaction features are element-wise accumulated, multiplied, and concatenated to form a comprehensive feature, providing a more discriminative feature foundation for subsequent adaptive fusion. Based on this, a gating mechanism is introduced to adaptively regulate different interaction branches, thereby suppressing redundant information and highlighting salient regions. Finally, the regulation results are aggregated and integrated through residual convolutional blocks, achieving a dynamic balance between local boundaries and global consistency.

[0114] (19)

[0115] (20)

[0116] (twenty one)

[0117] in, and These represent element-wise addition, multiplication, and concatenation operations, respectively. Represented as a non-linear activation function, Represented as a gating weight function, it consists of a 1×1 convolution, batch normalization, and sigmoid activation. This indicates a depth-separable residual block. It represents the characteristics of element-wise addition, multiplication, or concatenation. Represents high-level features at scale j. and These represent the weighted features and the flood inundation map, respectively.

[0118] The environment used in this invention was trained on an NVIDIA GeForce RTX3090 GPU with 24GB of memory within the PyTorch framework, using the Adam optimizer, with an initial learning rate set to 2e-4, a batch size of 32, and a total of 100 training periods. Using pre-flood optical images and post-flood SAR images as input, the BCEDice loss function was employed during training to optimize the network, balancing accuracy and structural feature learning, and outputting the extent of flooding.

[0119] pass Figure 5The results demonstrate that this invention has significant advantages and positive effects: by introducing a multi-level cross-modal fusion network and combining the deep complementarity and redundancy modeling of optical and SAR images, the accuracy and robustness of flood inundation monitoring are effectively improved, especially in complex flood scenarios where it exhibits stronger adaptability; at the same time, the frequency domain adaptive dual-stream fusion module enhances the model's sensitivity to complex boundaries, significantly improving the identification accuracy and boundary characterization ability of flood areas; furthermore, the design of a hierarchical multi-interaction fusion module further strengthens the jointness and differences of features at different levels, enabling the model to exhibit stable performance under multi-scale and cross-modal conditions.

[0120] On the other hand, embodiments of the present invention also provide a flood inundation monitoring device that integrates optical and SAR multi-level cross-modal fusion, including a processor and a memory. The memory is used to store program instructions, and the processor is used to call the program instructions in the memory to execute the flood inundation monitoring method that integrates optical and SAR multi-level cross-modal fusion as described in the above technical solution.

[0121] Meanwhile, embodiments of the present invention also provide a computer-readable storage medium, including a readable storage medium on which a computer program is stored. When the computer program is executed, it implements the flood inundation monitoring method of optical and SAR multi-level cross-modal fusion as described in the above technical solution.

[0122] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments, or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.

Claims

1. A flood inundation monitoring method that integrates optical and SAR multi-level cross-modal fusion, characterized in that, Includes the following steps: Acquire optical images before the flood and SAR images after the flood; The optical image and SAR image are encoded separately using the backbone network to obtain multi-scale optical features and SAR features. A cross-modal adaptive interactive fusion module is constructed, and multi-scale dynamic spatial downsampling is introduced to obtain optical and SAR features after cross-modal interaction. Modal adaptive weighted fusion is introduced to output the final control features. A frequency-domain adaptive dual-stream fusion module is constructed. First, the control features are divided into high-resolution features and low-resolution features, and cross-scale information alignment is achieved in a unified channel space. Then, filtering and corresponding resampling operations are performed to obtain high-level features and low-level features. The specific implementation methods for obtaining high-level and low-level features are as follows: Adaptive low-pass and adaptive high-pass filters are used to generate position-adaptive low-pass and high-pass filter kernels respectively, thereby achieving joint modeling of low-frequency global consistency and high-frequency boundary details; in, and These represent high-resolution features and low-resolution features, respectively. and These represent an adaptive low-pass filter and an adaptive high-pass filter, respectively. The upsampling operation maps low-resolution features to a space consistent with the high-resolution one. This indicates a normalization operation, ensuring the numerical stability of the convolution kernel; Finally, low-resolution feature intermediate variables Content-aware upsampling is used to obtain low-frequency consistent low-level features. High-resolution feature intermediate variables The low-frequency components are separated to obtain high-frequency residuals, and these residuals are then back-injected to form high-level features that enhance details. : in, This indicates a resampling operation under the constraints of an adaptive low-pass or high-pass filter, used to generate low-frequency consistency and high-frequency residual information; A hierarchical multi-interaction fusion module is constructed, which interacts with high-level features and low-level features through various complementary forms to obtain multi-level features that fuse global semantic information and local details. Through these multi-level features, a high-precision flood inundation map is generated.

2. The flood inundation monitoring method based on multi-level cross-modal fusion of optical and SAR technologies as described in claim 1, characterized in that: A pyramid visual transformer is used as the backbone network.

3. The flood inundation monitoring method based on multi-level cross-modal fusion of optical and SAR technologies as described in claim 1, characterized in that: The specific optical and SAR features obtained after cross-modal interaction include: First, a bidirectional cross-attention mechanism is employed to achieve information interaction and fusion between optical and SAR features, fully capturing complex nonlinear information from different modalities; the optical branch is obtained through a linear layer. After multi-scale dynamic spatial downsampling, the following was obtained. and SAR branch is obtained through linear layers After multi-scale dynamic spatial downsampling, the following was obtained. and Then, the query vector of the optical branch. Key-value pairs with SAR branches Matching is performed to capture complementary information about the optical branch of SAR features; conversely, the query vector of the SAR branch is used to... Key-value pairs with optical branches Interact with the system to introduce fine-grained representations of optical features; in, This represents the dimension of the feature vector in each attention head. and These represent the optical and SAR features after cross-modal interaction, respectively. Indicates scale.

4. The flood inundation monitoring method based on multi-level cross-modal fusion of optical and SAR technologies as described in claim 1, characterized in that: Modal adaptive weighted fusion specifically includes: The final modulation feature is obtained by dynamically assigning weights to different modal features at each position in space using a learnable gating function. : in, and These represent the optical and SAR features after cross-modal interaction, respectively. Indicates scale. For learnable gating functions, This indicates a splicing operation. Represents the normalization function. This indicates element-wise multiplication. Indicates adaptive weights, and Let represent the adaptive weights of optical features and SAR features at each position, respectively, and satisfy . .

5. The flood inundation monitoring method based on multi-level cross-modal fusion of optical and SAR technologies as described in claim 1, characterized in that: The modulation features obtained from the cross-modal adaptive interactive fusion module are categorized into high-resolution features based on their relative resolution. With low-resolution features First, high-resolution features With low-resolution features Compressed to a uniform channel space by 1×1 convolution; Residual information is extracted using global average pooling (GAP), and position-adaptive weights are generated using the sigmoid function to finally obtain aligned high-resolution features. With low-resolution features : in Indicates Sigmoid activation. ⊙ indicates global average pooling, and ⊙ indicates position-wise weighted pooling. and These represent the compressed high-resolution and low-resolution features, respectively.

6. The flood inundation monitoring method based on multi-level cross-modal fusion of optical and SAR technologies as described in claim 1, characterized in that: The specific processing procedure of the hierarchical multi-interaction fusion module is as follows: High-level features obtained from the Frequency Domain Adaptive Dual Stream Fusion Module (FADF) With low-level features Mapped to a unified channel dimension, three complementary interaction forms are then introduced to obtain interactive features, where addition is used to model joint information to capture the commonalities of features across levels, differencing is used to emphasize hierarchical differences to enhance complementarity, and multiplication is used to characterize the correlation between features to strengthen consistent expression. in, and These represent the high-level and low-level features of the unified channel dimension, respectively. This represents position-by-position weighted operations. , and These represent the results of adding high-level features and low-level features, the results of absolute difference, and the results of multiplication, respectively. The multi-way interaction features are element-wise accumulated, multiplied, and concatenated to form a comprehensive feature. A gating mechanism is introduced to adaptively control different interaction branches, resulting in multi-level features that integrate global semantic information and local details. in, and These represent element-wise addition, multiplication, and concatenation operations, respectively. Represented as a non-linear activation function, Represented as a gating weight function, it consists of a 1×1 convolution, batch normalization, and sigmoid activation. It represents a comprehensive characteristic of element-wise addition, multiplication, or concatenation. This represents the multi-level features after empowerment.

7. The flood inundation monitoring method based on multi-level cross-modal fusion of optical and SAR technologies as described in claim 6, characterized in that: A high-precision flood inundation map is generated using residual blocks, as detailed below: in, Indicates the first j High-level features at various scales This indicates a depth-separable residual block. This represents a flooded area.

8. A flood inundation monitoring device integrating optical and SAR multi-level cross-modal fusion, characterized in that: It includes a processor and a memory, the memory being used to store program instructions, and the processor being used to call the program instructions in the memory to execute the flood inundation monitoring method of optical and SAR multi-level cross-modal fusion as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, It includes a readable storage medium on which a computer program is stored, and when the computer program is executed, it implements the flood inundation monitoring method of optical and SAR multi-level cross-modal fusion as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Flood detection method and device based on fusion of optics and SAR (Synthetic Aperture Radar)

    CN116682020A

  • Visible light and infrared image fusion method based on cross-modal dynamic collaboration

    CN120525735A