Method and device for extracting flood depth levels by fusing optical, sar and dem data

The flood depth level extraction method based on multimodal remote sensing data fusion solves the problems of unstable flood depth information extraction and insufficient boundary recognition accuracy in complex scenarios, and achieves efficient and accurate flood depth level discrimination, supporting flood monitoring and emergency response in multiple scenarios.

CN121746875BActive Publication Date: 2026-05-29WUHAN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN UNIV
Filing Date
2026-02-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing flood monitoring methods are inadequate for characterizing water depth differences within inundated areas in complex scenarios, have low end-to-end inference efficiency, and are highly complex in multimodal modeling, making it difficult to meet the needs of flood monitoring and emergency management in multiple scenarios and on a large scale.

Method used

By employing a fusion-priority modal sensing feature module, a depth-constrained local-global guidance module, and a multi-scale structural sensing feature refinement module, a network for constructing flood depth level extraction is built, enabling complementary deep feature modeling and efficient fusion of multimodal remote sensing data.

Benefits of technology

It achieves end-to-end extraction of flood depth levels, improves the accuracy and robustness of the discrimination results, enhances stability and spatial consistency in complex scenarios, and supports large-scale, multi-scenario flood monitoring and emergency response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746875B_ABST
    Figure CN121746875B_ABST
Patent Text Reader

Abstract

The application provides a flood depth level extraction method and device based on optical, SAR and DEM data fusion. Through the fusion of a priority modal perception feature module, a depth-constrained local-global guidance module and a multi-scale structure perception feature refining module, efficient fusion of multi-modal information and depth feature modeling are realized, and the stability and spatial consistency of flood depth level discrimination are improved. The priority modal perception feature module enhances the complementarity between different modal information, the depth-constrained local-global guidance module takes into account local details and global semantic information, and the multi-scale structure perception feature refining module represents the multi-scale structure features of complex flood patterns, thereby improving the adaptability and robustness of the model under complex scenes and multi-terrain conditions. The application realizes end-to-end extraction of flood depth levels, effectively improves flood monitoring effects in complex scenes, provides technical support for flood monitoring, emergency response and disaster assessment, and has good application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent interpretation of remote sensing images and multimodal information fusion technology, specifically relating to a method and apparatus for extracting flood depth levels based on the fusion of optical, SAR and DEM data. Background Technology

[0002] With the increasing frequency of extreme hydrological and climatic events, the risks and uncertainties of flood disasters have intensified. How to quickly and accurately obtain spatial distribution information of floods has become a key issue in improving disaster monitoring and emergency management capabilities. Remote sensing technology, with its advantages of wide coverage and fast acquisition speed, has been widely used in flood disaster monitoring. Existing research mainly focuses on the extraction of flood inundation areas. Early methods were mostly based on the spectral characteristics of water bodies in optical remote sensing images, but these are easily affected by cloud cover and complex terrain backgrounds. Subsequently, synthetic aperture radar (SAR) imagery, due to its all-weather and all-day imaging capabilities, has been widely used for flood change detection, effectively improving the reliability of flood extent identification. In recent years, with the development of multi-source remote sensing data and deep learning methods, the multimodal fusion of optical and SAR imagery has further enhanced the robustness of flood extent extraction in complex scenarios.

[0003] However, most existing flood monitoring studies focus on determining whether a flood has occurred and its spatial extent, failing to characterize water depth differences within the inundated area and thus failing to meet the practical needs of refined post-disaster risk assessment and emergency decision-making. To obtain more detailed flood information, some studies have begun to explore flood depth estimation methods. Traditional flood depth acquisition methods mainly rely on field surveys and hydrological station monitoring, which, while highly accurate, are limited by high costs and limited spatial coverage, making them difficult to apply to large-scale flood monitoring. Flood depth estimation methods based on remote sensing data are gradually emerging, with most of these methods relying on the physical relationship between topographic data and water surface changes, obtaining flood depth distribution through water surface elevation inference and spatial interpolation.

[0004] While the aforementioned methods have achieved flood depth estimation to some extent, they still have shortcomings when applied to multi-scenario and large-scale applications. On the one hand, some methods are heavily reliant on data-intensive urban scenarios, limiting their applicability in areas with complex terrain or sparse data. On the other hand, existing flood depth estimation workflows typically involve multiple independent modeling steps, making it difficult to achieve efficient end-to-end inference from multi-source remote sensing data to flood depth results. Furthermore, the direct overlay of multimodal data often leads to a significant increase in computational complexity and resource consumption, hindering rapid deployment and application in real-world business scenarios.

[0005] Therefore, there is an urgent need for a technical solution that can make full use of multimodal remote sensing data and achieve efficient and stable extraction of flood depth information while ensuring computational efficiency, so as to meet the application needs of flood monitoring and emergency management in multiple scenarios and on a large scale. Summary of the Invention

[0006] This invention addresses the shortcomings of existing flood monitoring methods, such as difficulty in characterizing water depth differences within inundated areas under complex scenarios, low end-to-end inference efficiency, and high complexity of multimodal modeling. It provides a flood depth level extraction method and apparatus based on multimodal remote sensing image fusion. The method constructs flood depth level training samples by acquiring and preprocessing optical, SAR, and DEM data. The method employs a fusion-priority modal sensing feature module, a depth-constrained local-global guidance module, and a multi-scale structural sensing feature refinement module to achieve complementary modeling and efficient fusion of depth features from multimodal remote sensing data, improving the stability, spatial consistency, and boundary sensitivity of flood depth level discrimination. Each module effectively characterizes the changing characteristics of complex flood morphologies under different scenarios and terrain conditions through modal sensing fusion, cross-scale collaborative modeling, and multi-scale structural enhancement. This invention enables end-to-end extraction of flood depth levels, improving the accuracy and robustness of discrimination results while maintaining computational efficiency. It provides reliable technical support for large-scale, multi-scenario flood monitoring, emergency response, and post-disaster assessment, and has broad application prospects.

[0007] This invention provides a method for extracting flood depth levels by fusing optical, SAR, and DEM data, including:

[0008] Collect and preprocess multi-source remote sensing data;

[0009] A flood depth classification extraction network is constructed to extract flood depth classifications, specifically including the following steps:

[0010] (1) Lightweight multi-scale feature modeling and fusion of pre-disaster optical images, post-disaster SAR images and DEM data are performed by using the fusion priority modal sensing feature module;

[0011] (2) Input the fused features into the encoder to obtain low-level local detail features at n scales;

[0012] (3) Input the low-level local detail features of the highest scale, i.e. the nth scale, into the nth multi-scale structure perception feature refining module to obtain the high-level global semantic features of the nth scale. Input the low-level local detail features of the n-1th scale and the high-level global semantic features of the nth scale into the local-global guidance module of the n-1th depth constraint to obtain the corresponding results.

[0013] (4) The nth high-level global semantic feature and the output of the (n-1)th deep-constrained local-global guidance module are summed and input into the (n-1)th multi-scale structure-aware feature refining module. The processing of step (3) is repeated until the output of the first multi-scale structure-aware feature refining module is obtained as the final refined feature.

[0014] (5) Input the final refined features into the fully connected layer for feature classification to obtain the flood depth level results.

[0015] Furthermore, the processing procedure for fusing the priority modality-aware feature module in step (1) is as follows:

[0016] First, a modality-independent multi-scale feature encoding is employed, constructing a lightweight multi-scale feature modality encoding branch for each input modality, and introducing a learnable modality bias into each encoding modality branch. Adaptive correction is performed on multi-scale features. Each coding mode is extracted through convolutional layers only. Grouped convolution is introduced to reduce the number of parameters and computational complexity, thereby obtaining flood-related features that effectively represent different spatial scales.

[0017] Then, flood-related features are fused, and the fused features are guided by spatial saliency, specifically expressed by the following formula:

[0018]

[0019]

[0020] in These represent the optical, SAR, and DEM features after multi-scale feature encoding, respectively. This indicates the characteristics after the channels are connected. This represents the final result after fusing the priority modality-aware feature module. Represents global average pooling. Represents global max pooling. Represents average pooling. Represents the Sigmoid function. This represents a channel connection operation.

[0021] Furthermore, flood-related characteristics are obtained using the following formula:

[0022]

[0023]

[0024] k

[0025] in, These are respectively abbreviations for pre-disaster optical images, post-disaster SAR images, and DEM data. These represent the input pre-disaster optical images, post-disaster SAR images, and DEM data, respectively. for The set, Indicates modal bias. and These represent the features after ordinary convolution and grouped convolution, respectively. and These represent 1×1 convolution and 3×3 convolution, respectively. This represents a 3×3 grouped convolution with dilation rates of 1, 2, and 4.

[0026] Furthermore, the encoder mentioned in step (2) is a PVT-2 model.

[0027] Furthermore, in steps (3) and (4), the multi-scale structure perception feature refinement module uses multiple sets of depthwise separable convolutions with different kernel sizes to extract multi-scale features in parallel, and performs global statistical aggregation on each scale feature to generate a scale description vector. Then, it combines all scale description vectors to generate adaptive weights. Finally, by weighted summation of the multi-scale features and combining residual connections, it achieves adaptive fusion of features at different spatial scales.

[0028] Furthermore, the extracted multi-scale features Represented as:

[0029]

[0030] in, For activation function, For layer normalization, Indicates the kernel size as The depthwise separable convolution operator, where j is the scale, taking values ​​from 1 to n. Represents the high-level global semantic features at the j-th scale;

[0031] The adaptive weights are calculated as follows:

[0032]

[0033] ([ ])

[0034] in These represent the feature's channels, height, and length, respectively. This represents the weight of the j-th scale feature. This represents a mapping consisting of fully connected layers and nonlinear functions. For normalized exponential functions, These are the adaptive weights corresponding to each scale;

[0035] Finally, the output of the j-th multi-scale structure-aware feature refinement module for:

[0036] .

[0037] Furthermore, in steps (3) and (4), the processing of the local-global guidance module for depth constraints includes:

[0038] First, high-level global semantic features and low-level local detail features are mapped to a unified channel dimension, and additive features and differential features are constructed respectively. The additive features are used to characterize the consistency of local-global semantics, while the differential features are used to highlight the change areas between features at different scales.

[0039] Further, a deep continuity constraint is introduced on the above additive and differential features to obtain additive and differential features with continuity constraints.

[0040] The depth continuity constraint is expressed as:

[0041]

[0042]

[0043] in This represents the average operation over the channel dimension. This represents a depth continuity modeling function composed of convolution and nonlinearity; and These are respectively represented as additive features and differential features subject to continuity constraints;

[0044] By combining additive and differential features constrained by continuity, a gating mechanism is used to modulate low-level local detail features pixel by pixel. The modulation result is then injected into the low-level local detail features in the form of residuals after fusion, thereby achieving synergistic enhancement of low-level local detail features and high-level global semantic features.

[0045] Furthermore, additive and differential features are represented as follows:

[0046]

[0047]

[0048] in, and These represent the high-level global semantic features at the (t+1)th scale and the low-level local detail features at the tth scale, respectively, after unifying the channel dimensions. and These represent additive features and differential features, respectively.

[0049] The collaborative enhancement of low-level local detail features and high-level global semantic features is represented as follows:

[0050]

[0051]

[0052]

[0053] in, This represents the original low-level local detail features at the t-th scale. Represented as a non-linear activation function, Represented as a gating weight function, it consists of a 1×1 convolution, batch normalization, and sigmoid activation. and This represents the intermediate feature result after addition and concatenation operations. This is the final result of the local-global bootstrapping module for the t-th depth constraint.

[0054] The present invention also provides a flood depth level extraction device based on optical, SAR and DEM data fusion, including a processor and a memory. The memory is used to store program instructions, and the processor is used to call the program instructions in the memory to execute the flood depth level extraction method based on optical, SAR and DEM data fusion as described in the above technical solution.

[0055] The present invention also provides a computer-readable storage medium, including a readable storage medium on which a computer program is stored, wherein when the computer program is executed, it implements the flood depth level extraction method by fusion of optical, SAR and DEM data as described in the above technical solution.

[0056] The beneficial effects of the technical solution provided by this invention are as follows:

[0057] (1) Existing methods for extracting flood depth or inundation information often rely on single modality or simple multimodal superposition in complex scenarios, which makes it difficult to give full play to the complementary advantages between different remote sensing data, resulting in unstable feature representation and insufficient discrimination ability. This invention introduces a modal perception mechanism into a single backbone structure by setting a fusion priority modal perception feature module, realizing complementary modeling of depth features of optical, SAR and DEM data, reducing feature redundancy while enhancing the collaborative expression ability of multimodal information, thereby improving the stability and reliability of flood depth level discrimination.

[0058] (2) In view of the problem that it is difficult to stably characterize the flood structure features under complex terrain and diverse flood morphology conditions, this invention sets up a multi-scale structure perception feature refinement module to enhance the fused features at multiple scales. This can effectively characterize the changes in flood morphology at different scales and directions, improve the model's cross-scene adaptability, and enhance the robustness and fine expression of flood depth level discrimination in complex scenarios.

[0059] (3) In view of the problem that existing methods are prone to insufficient spatial continuity and blurred boundaries of different depth levels during multi-scale feature fusion, this invention introduces a local-global guidance module with depth constraints. By strengthening the guidance and interaction between cross-scale features, it takes into account both local detailed information and global semantic structure, effectively improves the spatial consistency of flood depth level discrimination results, and enhances the sensitivity to boundary areas of different depth levels. Attached Figure Description

[0060] Figure 1 This is a schematic diagram of the overall structure of the flood depth level extraction network in an embodiment of the present invention.

[0061] Figure 2 This is a schematic diagram of the structure of the Fusion Priority Modality Awareness Feature Module (FMFM) in an embodiment of the present invention;

[0062] Figure 3 This is a schematic diagram of the structure of the Multi-Scale Structure-Aware Feature Refinement Module (MSFRM) in an embodiment of the present invention;

[0063] Figure 4 This is a schematic diagram of the structure of the depth-constrained local-global guidance module (DLGM) in an embodiment of the present invention;

[0064] Figure 5 This is a visualization of the flood depth prediction results in an embodiment of the present invention. Detailed Implementation

[0065] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0066] This invention proposes a method and apparatus for extracting flood depth levels based on multimodal remote sensing image fusion, addressing the problems of unstable flood depth information extraction and insufficient boundary recognition accuracy in complex scenarios. In this embodiment, optical, SAR, and DEM data of the target area are first acquired. The multimodal remote sensing data are then registered, cropped, normalized, and labeled to construct flood depth level training samples. Subsequently, the preprocessed multimodal data is input into the backbone network PVT-2 to extract multi-scale features. During the feature extraction stage, a fusion-priority modal sensing feature module is used to jointly model different modal features, achieving efficient fusion and complementary expression of multimodal features. Based on this, a multi-scale structural sensing feature refinement module is used to structurally enhance the fused features, improving the model's stability and robustness under complex terrain and multi-scenario conditions. Subsequently, a depth-constrained local-global guidance module is introduced to strengthen the interaction between cross-scale features, improving the spatial consistency and boundary sensitivity of flood depth level discrimination results. Finally, the flood depth level distribution results of the target area are generated based on the fused features, achieving automatic mapping of flood depth levels and providing data support for flood monitoring, emergency response, and post-disaster assessment.

[0067] The environment used in this invention was all on an NVIDIA GeForce RTX 3090 GPU with 24GB of memory running the PyTorch framework. The specific steps of this invention's embodiments are as follows:

[0068] The specific steps of this embodiment of the invention are as follows:

[0069] like Figure 1 As shown, to address the problems of large parameter scale, high computational complexity, and limited inference efficiency caused by the multi-branch structure in traditional multimodal flood depth level extraction methods, this invention first inputs pre-disaster optical images, post-disaster SAR images, and DEM data into the Fusion-first modality-aware feature module (FMFM). Through lightweight, modality-aware feature encoding and early fusion, it significantly reduces computational overhead while achieving effective interaction of multimodal information and generating a unified and efficient fused feature representation. Subsequently, the fused features are fed into the encoder for deep semantic encoding, utilizing its hierarchical structure to extract multi-scale features step by step, enhancing semantic expressiveness while avoiding information redundancy caused by multiple branches.

[0070] (1)

[0071] (2)

[0072] in, This is a collection of preprocessed pre-disaster optical images, post-disaster SAR images, and DEM data as input. These are the encoded features.

[0073] To adapt to the diversity of flood spatial morphology in complex flood scenarios in terms of scale and direction, a multi-scale structure-aware feature refinement module (MSFRM) is proposed. This module refines and enhances features from multiple perspectives, including channel statistics, spatial orientation structure, and multi-scale receptive fields. Furthermore, addressing the issues of insufficient interaction between multi-scale features and the difficulty in explicitly modeling the spatial continuity of flood depth, this invention proposes a depth-constrained local-global guidance module (DLGM). Through local-global guidance and depth continuity constraints, this module enhances the spatial consistency and boundary sensitivity of flood depth classification.

[0074] (3)

[0075] (4)

[0076] (5)

[0077] in, The result is the refined feature set using MSFRM. To obtain the final flood depth result, These are the feature results after DLGM.

[0078] The following is a detailed explanation of several key modules:

[0079] (1) Fusion priority modality sensing feature module

[0080] This invention proposes a Fusion Priority Modality Awareness Feature Module (FMFM). This module is placed before the backbone network to perform lightweight multi-scale feature modeling and fusion of pre-disaster optical, post-disaster SAR, and DEM data.

[0081] like Figure 2 As shown, FMFM employs modality-independent multi-scale feature encoding, constructing a lightweight multi-scale feature encoding branch for each input modality. Simultaneously, a learnable modality bias is introduced into each modality branch. Adaptive correction is applied to modal features, enabling the network to automatically balance the relative contributions of different modalities. Each modality is feature-extracted using only a small number of convolutional layers, and grouped convolutions are introduced to reduce the number of parameters and computational complexity. At the same time, convolutions with different dilation rates are combined to model the spatial context at multiple scales, effectively representing flood-related features at different spatial scales.

[0082] (6)

[0083] (7)

[0084] k (8)

[0085] in This term is used to refer to pre-disaster optical imagery, post-disaster SAR imagery, and DEM data. These represent the input pre-disaster optical images, post-disaster SAR images, and DEM data, respectively. for The set, Indicates modal bias. and These represent the features after ordinary convolution and grouped convolution, respectively. and These represent 1×1 convolution and 3×3 convolution, respectively. This represents a 3×3 grouped convolution with dilation rates of 1, 2, and 4.

[0086] Further considering that the spatial distribution of floods is constrained by both topographic conditions and water cover, a spatial saliency guidance (SSG) mechanism is implemented for the fusion characteristics. This mechanism constructs a spatial weight distribution based on the average and maximum convergence of the channel dimension, thereby effectively suppressing interference from highlands and non-inundated areas, highlighting low-lying and potentially inundated areas, and providing reliable spatial support for flood depth classification.

[0087] (9)

[0088] (10)

[0089] in These represent the optical, SAR, and DEM features after lightweight multi-scale feature encoding, respectively. This indicates the characteristics after connecting their channels. This indicates the final result after processing by the FMFM module. Represents global average pooling. Represents global max pooling. Represents average pooling. Represents the Sigmoid function. This represents a channel connection operation.

[0090] FMFM achieves a unified representation of multimodal information in the early stages of feature extraction. While maintaining discriminative ability, it effectively reduces the model parameter scale and computational complexity, providing an efficient unified feature representation for subsequent flood depth level extraction.

[0091] (2) Multi-scale structural perception feature refinement module

[0092] This invention designs a multi-scale structure-aware feature refinement module (MSFRM), which uses an adaptive multi-scale depth separable convolution module to solve the problem that a fixed-scale receptive field is difficult to adapt to complex flood morphologies. By using scale-adaptive weighted fusion of multi-scale features, it enhances the ability to represent flood structures at different spatial scales.

[0093] like Figure 3 As shown, first, the result obtained in (1) Obtained via encoder PVT-2 Features. Then, MSFRM employs multiple sets of depthwise separable convolutions with different kernel sizes to extract multi-scale features in parallel. Different scale convolution kernels correspond to different receptive field ranges, used to capture flood-related features ranging from local details to global structure. Using depthwise separable convolutions can significantly reduce the number of parameters and computational complexity while maintaining multi-scale feature modeling capabilities.

[0094] (11)

[0095] (12)

[0096] (13)

[0097] in, Indicates the kernel size as Depth-separable convolution operators.

[0098] To avoid redundancy and scale bias caused by simply stacking features at different scales, MSFRM further introduces a scale-adaptive weighting mechanism. Specifically, for each scale feature... Perform global statistical aggregation to generate scale description vectors, and combine all scale description vectors to generate weight results.

[0099] (14)

[0100] ([ ]) (15)

[0101] in These represent the feature's channels, height, and length, respectively. This represents a mapping consisting of fully connected layers and nonlinear functions. These are the adaptive weights corresponding to each scale. Ultimately, MSFRM achieves adaptive fusion of features at different spatial scales by weighted summation of multi-scale features and combining residual connections.

[0102] (16)

[0103] Through the above design, MSFRM can dynamically adjust the contribution of different spatial scales based on feature content, enhancing the perception of local flood boundaries and detailed structures while also taking into account the overall modeling of large-scale flood expansion patterns. This module significantly improves the adaptability of features to complex flood spatial structures while maintaining efficient computation, providing a more stable and discriminative multi-scale feature representation for subsequent flood depth classification.

[0104] (3) Depth-constrained local-global guidance module

[0105] This invention proposes a depth-constrained local-global guided module (DLGM) to enhance cross-scale feature interaction and improve the spatial consistency and boundary sensitivity of flood depth classification.

[0106] like Figure 4 As shown, firstly The result obtained in (2) As a high-level global semantic feature, Local detail features at lower levels are entered into the DLGM together, and so on for others, and mapped to a unified channel dimension to mitigate fusion bias caused by dimensional inconsistencies. Subsequently, additive and differential relationships of features are constructed separately. Additive features are used to characterize the consistency of local-global semantics, while differential features are used to highlight the regions of change between features at different scales, which is particularly helpful in characterizing the boundaries and transition regions corresponding to changes in flood depth levels.

[0107] (17)

[0108] (18)

[0109] in, and These represent the high and low layer features of the same channel dimension, respectively. and These represent additive features and differential features, respectively.

[0110] Considering that flood depth typically exhibits continuous spatial variation, DLGM further introduces a depth continuity constraint (DCC) on top of the aforementioned additive and differential features. A single-channel response is generated by statistically converging the features along the channel dimension, and then smoothed using lightweight convolutions to obtain constraint weights reflecting depth continuity. Subsequently, the features are continuously modulated.

[0111] (19)

[0112] (20)

[0113] in This represents the average operation over the channel dimension. This represents a deep continuity modeling function consisting of convolution and nonlinearity. and These are respectively represented as additive features and differential features subject to continuity constraints.

[0114] Building upon this foundation, DLGM constructs adaptive guiding weights based on additive and differential features, respectively, and modulates local features pixel-by-pixel through a gating mechanism. The two types of guiding results are then fused and injected into the local features in residual form, thereby achieving synergistic enhancement of local details and global semantics.

[0115] (twenty one)

[0116] (twenty two)

[0117] (twenty three)

[0118] in, Represented as a non-linear activation function, Represented as a gating weight function, it consists of a 1×1 convolution, batch normalization, and sigmoid activation. and This represents the intermediate feature result after addition and concatenation operations. This is the final result of this module.

[0119] DLGM can explicitly model the spatial continuity and hierarchical differences of flood depth during multi-scale feature interactions. This makes the network more sensitive to spatial transition regions between different depth levels while maintaining overall semantic consistency, thereby improving stability and discriminative ability in boundary regions and complex terrain conditions. Ultimately, this will refine the features. The input is fed into a fully connected layer for feature classification to obtain the flood depth level extraction result. In this embodiment, the cross-entropy loss function CELows is used when training the above network.

[0120] pass Figure 5 The results show that while FuseNet and DeepLab can identify the main inundation areas on a global scale, they often exhibit confusion and local classification errors at the boundaries of different flood depth levels. SegFormer lacks the ability to collaboratively model multimodal remote sensing information, resulting in insufficient connectivity and fragmented structure in flood boundaries and depth transition areas. In contrast, this invention achieves excellent results in both the overall flood boundary and local areas with different flood depth variations, and is more consistent with the ground truth in detail.

[0121] Compared with existing technologies, this invention has the following significant advantages and positive effects: By introducing a flood depth level extraction network that integrates multimodal remote sensing data, this invention collaboratively models optical, SAR, and DEM data, fully utilizing the complementary information between different modalities, effectively improving the stability and robustness of flood depth level discrimination results, especially demonstrating stronger adaptability under complex terrain, multiple underlying surfaces, and multiple scene conditions. Simultaneously, by fusing a priority modality perception feature module, the multimodal feature expression capability is enhanced while reducing feature redundancy and model complexity; by using a depth-constrained local-global guidance module, cross-scale feature interaction is strengthened, significantly improving the spatial consistency and boundary sensitivity of flood depth level discrimination results; and by using a multi-scale structure perception feature refinement module, multi-scale structural enhancement is performed on complex flood morphologies, enabling the model to maintain stable and reliable discrimination performance under multi-scale and multi-scene conditions.

[0122] On the other hand, embodiments of the present invention also provide a flood depth level extraction device for optical, SAR and DEM data fusion, including a processor and a memory. The memory is used to store program instructions, and the processor is used to call the program instructions in the memory to execute the flood depth level extraction method for optical, SAR and DEM data fusion as described in the above technical solution.

[0123] Thirdly, embodiments of the present invention also provide a computer-readable storage medium, including a readable storage medium on which a computer program is stored, wherein when the computer program is executed, it implements the flood depth level extraction method by fusion of optical, SAR and DEM data as described in the above technical solution.

[0124] The above description is merely a specific embodiment of the present invention, used to illustrate the technical solution of the present invention and not to limit it. Those skilled in the art can make various modifications, equivalent substitutions, or additions to the above embodiments without departing from the spirit and scope of the claims, all of which should fall within the protection scope of the present invention.

Claims

1. A method for extracting flood depth levels by fusing optical, SAR, and DEM data, characterized in that, include: Collect and preprocess multi-source remote sensing data; A flood depth classification extraction network is constructed to extract flood depth classifications, specifically including the following steps: (1) Lightweight multi-scale feature modeling and fusion of pre-disaster optical images, post-disaster SAR images and DEM data are performed by using the fusion priority modal sensing feature module; The processing procedure for fusing the priority modality sensing feature module in step (1) is as follows: First, a modality-independent multi-scale feature encoding is employed, constructing a lightweight multi-scale feature modality encoding branch for each input modality, and introducing a learnable modality bias into each encoding modality branch. Adaptive correction is performed on multi-scale features. Each coding mode is extracted through convolutional layers only. Grouped convolution is introduced to reduce the number of parameters and computational complexity, thereby obtaining flood-related features that effectively represent different spatial scales. Then, flood-related features are fused, and spatial saliency guidance is applied to the fused features; (2) Input the fused features into the encoder to obtain low-level local detail features at n scales; (3) Input the low-level local detail features of the highest scale, i.e. the nth scale, into the nth multi-scale structure perception feature refining module to obtain the high-level global semantic features of the nth scale. Input the low-level local detail features of the n-1th scale and the high-level global semantic features of the nth scale into the local-global guidance module of the n-1th depth constraint to obtain the corresponding results. (4) The nth high-level global semantic feature and the output of the (n-1)th deep-constrained local-global guidance module are summed and input into the (n-1)th multi-scale structure-aware feature refining module. The processing of step (3) is repeated until the output of the first multi-scale structure-aware feature refining module is obtained as the final refined feature. In steps (3) and (4), the multi-scale structure perception feature refinement module uses multiple sets of depthwise separable convolutions with different kernel sizes to extract multi-scale features in parallel, and performs global statistical aggregation on each scale feature to generate a scale description vector. Then, it combines all scale description vectors to generate adaptive weights. Finally, by weighted summation of the multi-scale features and combined with residual connections, it achieves adaptive fusion of features at different spatial scales. In steps (3) and (4), the processing of the local-global guidance module for depth constraints includes: First, high-level global semantic features and low-level local detail features are mapped to a unified channel dimension, and additive features and differential features are constructed respectively. The additive features are used to characterize the consistency of local-global semantics, while the differential features are used to highlight the change areas between features at different scales. Further, a deep continuity constraint is introduced on the above additive and differential features to obtain additive and differential features with continuity constraints. By combining additive and differential features after continuity constraints, a gating mechanism is used to modulate low-level local detail features pixel by pixel. The modulation result is further injected into the low-level local detail features in the form of residuals after fusion, thereby achieving synergistic enhancement of low-level local detail features and high-level global semantic features. (5) Input the final refined features into the fully connected layer for feature classification to obtain the flood depth level results.

2. The flood depth level extraction method based on optical, SAR, and DEM data fusion as described in claim 1, characterized in that: Flood-related features are fused, and the spatial saliency of the fused features is guided, specifically expressed by the following formula: in These represent the optical, SAR, and DEM features after multi-scale feature encoding, respectively. This indicates the characteristics after the channels are connected. This represents the final result after fusing the priority modality-aware feature module. Represents global average pooling. Represents global max pooling. Represents average pooling. Represents the Sigmoid function. This represents a channel connection operation.

3. The flood depth level extraction method based on optical, SAR, and DEM data fusion as described in claim 2, characterized in that: Flood-related characteristics are represented by the following formula: k in, These are abbreviations for pre-disaster optical images, post-disaster SAR images, and DEM data, respectively. These represent the input pre-disaster optical imagery, post-disaster SAR imagery, and DEM data, respectively. for The set, Indicates modal bias. and These represent the features after ordinary convolution and grouped convolution, respectively. and These represent 1×1 convolution and 3×3 convolution, respectively. This represents a 3×3 grouped convolution with dilation rates of 1, 2, and 4.

4. The flood depth level extraction method based on optical, SAR, and DEM data fusion as described in claim 1, characterized in that: The encoder mentioned in step (2) is a PVT-2 model.

5. The flood depth level extraction method based on optical, SAR, and DEM data fusion as described in claim 1, characterized in that: The extracted multi-scale features are represented as follows: in, For activation function, For layer normalization, Indicates the kernel size as The depthwise separable convolution operator, where j is the scale, taking values ​​from 1 to n. Represents the high-level global semantic features at the j-th scale; The adaptive weights are calculated as follows: ([ ]) in These represent the feature's channels, height, and length, respectively. This represents a mapping consisting of fully connected layers and nonlinear functions. For normalized exponential functions, These are the adaptive weights corresponding to each scale; Finally, the output of the multi-scale structure-aware feature refinement module is: 。 6. The flood depth level extraction method based on optical, SAR, and DEM data fusion as described in claim 1, characterized in that: The depth continuity constraint is expressed as: in This represents the average operation over the channel dimension. This represents a depth continuity modeling function composed of convolution and nonlinearity; and These are respectively represented as additive features and differential features subject to continuity constraints.

7. The flood depth level extraction method based on optical, SAR, and DEM data fusion as described in claim 6, characterized in that: Additive and differential features are represented as follows: in, and These represent the high-level global semantic features at the (t+1)th scale and the low-level local detail features at the tth scale, respectively, after unifying the channel dimensions. and These represent additive features and differential features, respectively. The collaborative enhancement of low-level local detail features and high-level global semantic features is represented as follows: in, This represents the original low-level local detail features at the t-th scale. Represented as a non-linear activation function, Represented as a gating weight function, it consists of a 1×1 convolution, batch normalization, and sigmoid activation. and This represents the intermediate feature result after addition and concatenation operations. This is the final result of the local-global bootstrapping module for the t-th depth constraint.

8. A flood depth level extraction device based on the fusion of optical, SAR, and DEM data, characterized in that: It includes a processor and a memory, the memory being used to store program instructions, and the processor being used to call the program instructions in the memory to execute the flood depth level extraction method of optical, SAR and DEM data fusion as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The method includes a readable storage medium on which a computer program is stored, which, when executed, implements the flood depth level extraction method based on the fusion of optical, SAR, and DEM data as described in any one of claims 1-7.