A Method for Constructing a Large-Scale Model for Flood Change Detection Based on Prior Topographic Knowledge
By introducing prior knowledge of terrain and a dual attention mechanism for cross-modal feature alignment and multi-task water feature enhancement, the semantic ambiguity and noise interference problems in flood change detection under complex terrain are solved, and high-precision flood inundation range identification is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHENGZHOU UNIV
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-31
AI Technical Summary
Existing cross-modal deep learning schemes struggle to effectively decouple semantic ambiguities caused by terrain elevation differences or shadows in complex terrains. Furthermore, the high-frequency speckle noise interference in radar images leads to missed detections, false detections, and boundary fragmentation in the results of flood inundation range extraction.
It integrates prior knowledge of elevation and terrain with a cross-modal deep network, adopts a pseudo-twin dual-branch architecture, introduces a lightweight DEM feature extraction module to explicitly encode terrain structure features, and uses a dual attention mechanism to perform cross-modal feature alignment and multi-task water feature enhancement, combined with gradient constraints to ensure the continuity of water body boundaries.
It achieves high-precision and robust flood inundation range identification in complex terrain, significantly reducing missed detections, false detections, and boundary fragmentation problems, and improving the accuracy and stability of flood change detection.
Smart Images

Figure CN122493311A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing image classification technology, specifically to a method for constructing a large-scale model for flood change detection based on prior terrain knowledge. Background Technology
[0002] The earliest method for obtaining information on the extent of flooding was a combination of records from surface hydrological stations and on-site measurements by manual patrols. Its advantage lies in the direct acquisition of data and the reliability of the results. However, its limitations are due to constraints on manpower, resources, the number of stations, and transportation conditions. Especially in mountainous or complex disaster-stricken areas, on-site investigations are often difficult to conduct in a timely manner, resulting in a significant lag in obtaining information on large-scale flooding and making it difficult to achieve full regional coverage (Alizadeh B, Behzadan A H. Scalableflood inundation mapping using deep convolutional networks and trafficsignage[J]. Computational Urban Science, 2023, 3(1): 17.).
[0003] With the development of remote sensing and Earth observation technologies, the main flood inundation range identification models currently cover traditional feature extraction models (image difference method / log ratio method), dual-temporal optical image models, dual-temporal radar (SAR) image models, and cross-modal image fusion models.
[0004] Traditional feature extraction models are characterized by their relatively simple implementation process and the ability to quickly obtain results in small-scale experiments. However, their disadvantages lie in the fact that feature extraction mainly relies on human experience, is extremely sensitive to threshold selection, and often requires repeated parameter tuning in different scenarios. Furthermore, their recognition accuracy and stability are insufficient under complex surface conditions (Silwal A, Subedi A, Tamrakar R, et al. A Comprehensive Review of Machine Learning and DeepLearning Methods for Flood Inundation Mapping[J]. Earth, 2026, 7(2): 44.).
[0005] The advantage of dual-temporal optical imagery is that it can provide rich spectral information of ground features, but its limitation is that cloud and fog obscure the image during disasters, which greatly restricts the feasibility of obtaining complete post-disaster images (Gebrehiwot A, Hasemi-Beni L. 3d inundation mapping: A comparison between deep learning image classification and geomorphic flood index approaches[J]. Frontiers in Remote Sensing, 2022, 3: 868104.). The dual-temporal radar imagery approach relies on the all-weather imaging capability of SAR and is widely used, but its limitation is that the inherent speckle noise of SAR and its single information type restrict the ability to accurately extract flood change information (Destefanis T, Guliyeva S, Boccardo P, et al. Advancing flood detection and mapping: a review of earth observation services, 3D data integration, and AI-Based techniques[J]. Remote Sensing, 2025, 17(17):2943.). Cross-modal image fusion schemes (pre-disaster optical + post-disaster SAR) utilize high-quality benchmarks and post-disaster real-time data to achieve information complementarity, with a high theoretical upper limit. However, they face limitations in mining deep spatiotemporal features in complex scenarios and require complex adjustments and optimizations when dealing with different types of data.
[0006] With the rapid advancements in artificial intelligence technology, deep learning techniques, represented by convolutional neural networks, are rapidly rising to become the mainstream methodology for flood change detection, replacing traditional algorithms, thanks to their powerful nonlinear feature extraction and end-to-end optimization capabilities. Deep learning methods can automatically capture data features and distributions, automatically extract deep information from multi-source remote sensing data, and establish a mapping relationship between image features and inundation extent. In contrast, traditional remote sensing extraction methods require excessive manual intervention and threshold settings under certain conditions, while deep learning models can theoretically adapt to complex environmental conditions. However, existing cross-modal deep learning schemes still face significant shortcomings in practical applications: on the one hand, optical images, limited by their imaging mechanisms, lack the ability to represent terrain and spatial structures, making it difficult to reflect the intrinsic relationship between water distribution and topographic relief, and easily leading to semantic ambiguity in complex terrain (such as misjudging mountain shadows as water bodies); on the other hand, high-frequency speckle noise in radar images can severely interfere with the stable modeling of water features during disasters.
[0007] In summary, flood disasters are characterized by their suddenness and destructive power. Timely and accurate acquisition of flood inundation ranges is of paramount theoretical and practical significance for dynamic disaster assessment, emergency rescue decision-making, and post-disaster reconstruction. However, in complex disaster environments, on-site surveys are resource-intensive and time-consuming. Deep learning, as a mainstream emerging technology for remote sensing interpretation, can automatically extract deep-level feature information from multi-source spatiotemporal data and establish mapping relationships. However, existing cross-modal flood detection deep learning networks generally lack the ability to perceive topography and geomorphology. In complex terrain, they are prone to semantic ambiguity due to differences in mountain elevation or shadows. Furthermore, single change detection tasks cannot effectively decouple the inherent high-frequency speckle noise of SAR images, resulting in a large number of missed detections, false detections, and boundary fragmentation in the final inundation range extraction results.
[0008] Therefore, a method for constructing a large-scale flood change detection model based on prior terrain knowledge is provided to solve the above-mentioned technical problems. Summary of the Invention
[0009] The purpose of this invention is to provide a method for constructing a large-scale flood change detection model based on terrain prior knowledge. This invention integrates elevation and terrain prior knowledge with a cross-modal deep network. First, it employs a pseudo-twin dual-branch architecture and introduces a lightweight DEM feature extraction module to explicitly encode elevation information into terrain structure features, providing clear terrain prior knowledge for cross-modal flood detection and effectively reducing spurious change responses caused by terrain undulations. Building upon this, a cross-modal feature alignment module guided by a dual attention mechanism enables interaction of different modal features at multiple scales. Furthermore, a water body semantic segmentation branch is introduced to construct a multi-task feature enhancement module, providing additional explicit semantic constraints for feature learning. This allows the network to focus on stable water body regions when extracting change information, thereby effectively suppressing speckle noise. Finally, gradient constraints are combined to ensure the continuity of water body boundaries, enabling effective extraction of globally dependent water body features and accurate reconstruction of subtle water body edges. Ultimately, this achieves high-precision and robust identification of flood inundation ranges.
[0010] The objective of this invention is achieved as follows: A large-scale flood change detection model guided by terrain prior knowledge is proposed, comprising a terrain prior feature encoding module, a dual attention-guided cross-modal feature alignment module, and a multi-task water feature enhancement module. It extracts features from pre-disaster optical and post-disaster SAR images through a pseudo-twin dual-branch architecture and introduces a lightweight DEM feature extraction module to explicitly encode elevation information into a terrain prior feature encoding module, providing the network with clear terrain prior constraints. The cross-modal feature alignment module facilitates multi-scale interaction, bridging the heterogeneous semantic gap between optical and SAR images. Subsequently, a water semantic segmentation branch is introduced to construct a multi-task water feature enhancement module, and gradient constraints are applied to force the network to decouple speckle noise and focus on stable water bodies at the feature level. Finally, a dual attention decoding module for both channel and spatial aspects is used to decode the data, outputting a high-precision flood change mask.
[0011] A method for constructing a large-scale model for flood change detection based on prior terrain knowledge includes the following steps: Step S1: Construct a terrain prior feature encoding enhancement module based on lightweight DEM feature extraction. The lightweight DEM feature extraction module is introduced to extract DEM terrain auxiliary information, and the spectral features of the two-dimensional optical image are constrained by three-dimensional geometric features to effectively suppress pseudo-change responses in the optical image that contradict the terrain logic. The construction of the terrain prior feature encoding enhancement module based on lightweight DEM extraction includes the following operation steps: feature standardization calibration, lightweight terrain encoding, and layer-by-layer geometric feature injection. Step S2: Construct a cross-modal feature alignment module based on dual attention guidance. Using the optical and SAR features output by the terrain prior feature encoding module in step S1, high-value semantic information in the features affected by different modalities is filtered through cross-modal channel attention. Combined with cross-modal spatial attention, the SAR features are reconstructed and denoised with optical features as guidance, thereby providing high-purity feature representation for the final mask prediction. Step S3: Construct a multi-task water feature enhancement module. Receive the attention map output from step S2 and input the attention map as the deepest semantic feature into the multi-task decoding network. Perform the main task of change detection and the auxiliary task of water segmentation respectively. Through the semantic constraints of backpropagation in the auxiliary branch, further force the preceding attention module and encoder to decouple noise from the real water body; thus obtaining a large-scale flood change detection model guided by terrain prior knowledge. Step S4, loss function and evaluation metric: The Adam optimizer is used to train the model, and the training loss value between the actual flood inundation range and the model prediction value is calculated by a hybrid loss function consisting of Cross-Entropy Loss and Dice loss, so as to effectively alleviate the problem of severe imbalance between the proportion of water bodies and background pixels in remote sensing images.
[0012] The specific operation of step S1 is as follows: S1.1 Standardization and Calibration of Optical Features: Based on the optical feature maps output from different stages of the backbone network (such as ResNet50), the number of channels increases continuously with the network depth (from 256 to 2048), and the feature distribution varies significantly. Direct cross-modal fusion would lead to a mismatch in computational dimensions and optimization difficulties. Therefore, before feature injection, a Residual Convolution Block (RCB) is introduced to preprocess the optical features. The RCB first uses 1×1 convolutions to uniformly project features from different levels to a fixed low-dimensional feature space, achieving channel unification. Subsequently, a convolutional block containing residual connections (consisting of two 3×3 convolutional layers) is connected for local feature extraction. The update process of the RCB can be described as follows:
[0013] In the formula, These are low-dimensional features after 1×1 convolutional projection. Represents a 3×3 convolutional nonlinear transformation and ReLU activation operation; after this calibration operation, the large-scale flood change detection model guided by terrain prior knowledge obtains a set of standardized optical features with consistent channel dimensions and optimized by local context, thus preparing the numerical data for incorporating terrain features. S1.2 Lightweight Multi-Scale Terrain Feature Encoding: To achieve strict alignment of terrain features with the optical backbone network in terms of spatial resolution, an independent lightweight multi-scale terrain feature encoder (DEM Encoder) is constructed. Unlike networks loaded with heavy pre-trained weights, the lightweight DEM feature encoder consists of five cascaded convolutional stages, specifically designed for hierarchical extraction of terrain semantics from single-channel DEM data. Each encoding stage includes a convolutional layer, batch normalization, and ReLU activation function. By setting specific convolutional stride and pooling operations, five sets of terrain feature maps with progressively decreasing spatial resolution (gradually reducing dimensionality from 1 / 2 to 1 / 32 of the input size) and a fixed number of channels are generated. , This layered mechanism ensures that topographic features can comprehensively cover multi-scale spatial structure information, from local micro-topography to macro-topography. S1.3 Layer-by-layer geometric feature injection and fusion: After acquiring multi-scale terrain features and standardized optical features, a layer-by-layer injection strategy is adopted to enhance the spectral features with geometric priors; for the first... For each level of features, bilinear interpolation is first used to force the terrain features to spatially align to a resolution consistent with the current level's optical features; subsequently, a feature concat operation is performed along the channel dimension.
[0014] In the formula, For the first Standardized optical characteristics of the layer The aligned topographic structure features of the i-th layer; the stitched feature tensor The system simultaneously incorporates both spectral texture and topographic geometry information of land features. Finally, a fusion layer composed of 3×3 convolutions is used to perform deep computation on the mixed features, enabling the network to automatically learn the spatial mapping relationship between "topographic elevation" and "water distribution". Through the above steps, this invention successfully and explicitly injects elevation priors into the network, ultimately outputting enhanced optical features with terrain-aware capabilities. This not only filters out spurious changes caused by complex terrain from a physical and logical perspective, but also provides high-quality, unambiguous baseline feature input for the next step of "cross-modal feature alignment and interaction" with SAR imagery.
[0015] The specific operation of step S2 is as follows: S2.1, Cross-modal Spatial Attention Feature Enhancement: The purpose of the spatial attention mechanism is to overcome the limitations of the local receptive field of traditional convolution. By explicitly modeling the dependencies between arbitrary spatial locations in the feature map, it enhances the model's ability to perceive key areas of global structure and flood changes. This is achieved for the deepest features output by the optical and SAR branches. and The spatial attention module first generates the corresponding query Q, key K, and value V feature matrices through 1×1 convolutions, and then flattens them into C×N vector representations (where N = H×W) in the spatial dimension. Subsequently, by calculating the matrix multiplication between Q and K and applying Softmax normalization, single-modal spatial attention matrices are generated. ;
[0016] To achieve efficient interaction of cross-modal information, the spatial attention results of the two modalities are fused to construct a cross-modal spatial cross-attention matrix. This simultaneously encodes the synergistic spatial variation relationship between optical and SAR features; subsequently, the cross matrix is applied to the value feature V and fused with the original feature via residual connection:
[0017] Finally, the spatial attention features from the optical and SAR branches are added element-wise to form the final cross-modal spatial attention feature representation: ; S2.2 Channel attention machines characterize the semantic differences in the importance of different feature channels. In multimodal scenarios, the semantic channels corresponding to different sensors (optical reflectivity and radar backscattering) differ significantly. Introducing cross-modal channel attention is key to improving feature fusion quality and aligning heterogeneous semantics. Similar to spatial attention, the channel module also... and It generates corresponding Q, K, and V features for the input; the difference lies in that it focuses on the similarity between channels, by calculating Q and The channel attention matrix is obtained from the correlation at the channel dimension:
[0018] Subsequently, the channel attention matrices of the two methods are fused to construct a cross-modal channel attention representation. This is then used to weight the value feature V, and channel self-attention features are obtained through residual connections: ; Finally, the channel features of the two modalities are added element-wise to obtain the cross-modal channel attention feature representation: ; S2.3 Dual Attention Feature Fusion and Alignment Output: After acquiring cross-modal enhancement features at the spatial and semantic levels respectively, this step generates channel attention feature maps. Spatial attention feature map Concatenation or addition can be performed along the channel dimension:
[0019] Through weighting and reconstruction using a dual attention mechanism, the model simultaneously enhances cross-modal feature representation at both the location (spatial) and semantic (channel) levels. Discrete noise in SAR images is effectively filtered out under the "guidance and suppression" of optical texture features, resulting in a high-purity combined attention map output. The deepest semantic features are input into the final decoder, thereby significantly improving the recognition accuracy and recovery capability of complex water body edges.
[0020] The specific operation of step S3 is as follows: S3.1 Construction of a Multi-Task Water Feature Enhancement Module: To effectively suppress the inherent speckle noise interference of SAR images at the source of feature extraction, this study constructs a water segmentation task branch in addition to the change detection backbone network. Considering the real-time requirements of flood monitoring and to avoid computational redundancy, the water segmentation task branch is built on the deepest enhanced features of the shared encoder. Although the deepest features have lower spatial resolution, they have the strongest semantic abstraction ability and the largest receptive field, and best reflect the overall properties of the water body. The water segmentation branch consists of only a 3×3 convolution (for feature smoothing), a 1×1 convolution (for channel compression), and a sigmoid activation function, outputting a single-channel water probability map. Its core innovation lies in constructing a short-path gradient flow from the semantic label directly to the encoder. During the training phase, when the encoder produces incorrect activations due to overfitting to SAR speckle noise, the water segmentation branch generates a huge penalty gradient; gradient backpropagation acts as an explicit feature regularizer, forcing the encoder parameters to be updated in the direction of "suppressing discrete noise and responding to homogeneous water bodies." S3.2, Multi-task Hybrid Loss Function Design and Joint Optimization: To alleviate the severe imbalance between positive and negative samples in flood detection and to guide the stable convergence of the dual-branch network, this step designs a multi-task hybrid loss function training framework. For the main task of change detection, a hybrid loss function combining Weighted Binary Cross-Entropy (BCE) and Dice Loss is adopted. Weighted binary cross-entropy assigns higher weights to the minority class (variable flood areas), strengthening the network's focus on sparse samples; Dice Loss optimizes the set similarity between predicted results and ground truth from a global perspective, improving the internal integrity of flood patches. Its calculation formula is as follows:
[0021] In the formula, N is the total number of pixels. For the first The true label is a pixel, where 1 represents a positive sample with variation and 0 represents the background. This represents the probability value predicted by the network. These are the weighting coefficients for positive samples. To prevent the denominator from being zero and to smooth the gradient from a minimum; For the auxiliary water body segmentation branch, the standard binary cross-entropy loss function is used. As a supervised objective, since the water segmentation branch does not pursue extreme boundary segmentation but aims to provide smooth and continuous gradient feedback, the standard BCE can robustly guide the model to decouple water features from the background in a low-resolution, high-dimensional semantic space, avoiding gradient oscillations in the early stages of training caused by using boundary-sensitive losses (such as Dice). The formula is as follows:
[0022] Finally, the total loss function for multi-task joint optimization is obtained by weighted summation of the main task loss and the auxiliary task loss:
[0023] In the formula, For multi-task balancing hyperparameters; experimental verification shows that when Setting it to 0.3 can minimize overall loss while ensuring the accuracy of change detection tasks; Through the above multi-task collaboration and hybrid loss constraints, the features of the main branch of inflow change detection are "pre-cleaned" at the source, and non-water body features and SAR high-frequency noise are effectively removed, significantly improving the signal-to-noise ratio of subsequent difference map generation; the main and auxiliary tasks are dynamically optimized under a unified loss framework, and finally output a high-precision flood change mask with continuous boundaries and no voids.
[0024] The specific operation of step S4 is as follows: The accuracy verification criteria are divided into two dimensions: pixel-level classification accuracy and spatial morphology accuracy. At the pixel-level classification level, the overall accuracy (OA), precision (Precision), recall (Recall), F1 score (F1-Score), and intersection-over-union (IoU) are used to evaluate the model's ability to capture flood-affected areas. Generally speaking, the higher the values of these five metrics, the better the model's pixel-level segmentation performance.
[0025] To address the spatial physical characteristic that flood-inundated areas typically exhibit as continuous patches rather than highly discrete fragments, the Perimeter-Area Ratio (PAR) and the Shape Fragmentation Index (SFI), based on the isoperimetric inequality, are introduced to provide a constrained evaluation of the model output from a topological perspective. PAR characterizes boundary complexity and internal voids, while SFI measures the deviation of the target shape from an ideal compact structure. Generally, smaller values for these two morphological indices indicate a more compact water patch structure, smoother boundaries, and a stronger ability to suppress noise fragmentation. The calculation formulas for the main evaluation indices are as follows:
[0026] Wherein, TP (True Positive) is the number of pixels that are both true floods and predicted as floods; FP (False Positive) is the number of pixels that are true non-floods but predicted as floods; FN (False Negative) is the number of pixels that are true floods but predicted as non-floods; in the morphological indices, A is the total pixel area of water bodies in the predicted results; L is the total boundary length of all water body patches (including the outer contour and the edge of the internal voids); SFI is a normalized dimensionless quantity, with a value range of [0, 1].
[0027] The beneficial effects of this invention are: 1. This invention uses a lightweight DEM feature extraction module to explicitly encode elevation information, guiding the deep network to overcome terrain pseudo-changes and enhancing the geographic interpretability of the deep learning model in complex areas; at the same time, it introduces a water body semantic segmentation branch and gradient constraint mechanism to decouple stable water bodies from discrete noise at the feature level, and supplements it with a dual attention mechanism of channel and space to accurately restore the water body edge.
[0028] 2. This invention constructs a change detection framework with dual constraints of physical mechanism (topography prior) and semantic mechanism (multi-task learning), which fully reveals the intrinsic dependency between multimodal image features and the actual flood distribution. On this basis, it completely solves the problems of missed detection, false detection and fragmentation of floods under complex topography and strong noise interference, significantly improves the expression quality of water body information during disasters, and shows greater advantages in flood inundation identification in complex environments. Attached Figure Description
[0029] Figure 1 This is a diagram of the overall network architecture of the TGFE-CDNet of this invention.
[0030] Figure 2 This is a schematic diagram illustrating the structure of the dataset used in the experiments of this invention.
[0031] Figure 3 This is a visual comparison of the experimental results of the present invention and the comparative model. Detailed Implementation
[0032] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0033] like Figure 1 As shown, a large-scale flood change detection model based on terrain prior knowledge includes a terrain prior feature encoding module, a dual attention-guided cross-modal feature alignment module, and a multi-task water feature enhancement module. It extracts pre-disaster optical and post-disaster SAR image features through a pseudo-twin dual-branch architecture and introduces a lightweight DEM feature extraction module to encode elevation information, providing the network with explicit terrain prior information. Then, it obtains attention maps through spatial and channel dual attention modules, using optical features to correct SAR features, achieving cross-modal feature alignment and multi-scale image interaction, bridging the heterogeneous semantic gap between optical and SAR. Subsequently, a water semantic segmentation branch is introduced to construct a multi-task water feature enhancement module, and gradient constraints are applied, forcing the network to decouple speckle noise at the feature level and focus on stable water bodies. Finally, it outputs a high-precision flood change mask.
[0034] A method for constructing a large-scale model for flood change detection based on prior terrain knowledge includes the following steps: Step S1: Construct a terrain prior feature encoding enhancement module based on lightweight DEM feature extraction. The lightweight DEM feature extraction module is introduced to extract elevation features as network auxiliary information, and the two-dimensional spectral features are constrained by three-dimensional geometric features to effectively suppress pseudo-change responses in optical images that contradict the terrain logic. The construction of the terrain prior feature encoding enhancement module based on lightweight DEM extraction includes the following operation steps: feature standardization calibration, lightweight terrain encoding, and layer-by-layer geometric feature injection.
[0035] The specific operation of step S1 is as follows: S1.1 Standardization and Calibration of Optical Features: Based on the optical feature maps output from different stages of the backbone network (such as ResNet50), the number of channels increases continuously with the network depth (from 256 to 2048), and the feature distribution varies significantly. Direct cross-modal fusion would lead to a mismatch in computational dimensions and optimization difficulties. Therefore, before feature injection, a Residual Convolution Block (RCB) is introduced to preprocess the optical features. The RCB first uses 1×1 convolutions to uniformly project features from different levels to a fixed low-dimensional feature space, achieving channel unification. Subsequently, a convolutional block containing residual connections (consisting of two 3×3 convolutional layers) is connected for local feature extraction. The update process of the RCB can be described as follows:
[0036] In the formula, These are low-dimensional features after 1×1 convolutional projection. Represents a 3×3 convolutional nonlinear transformation and ReLU activation operation; through this calibration, the large-scale flood change detection model guided by terrain prior knowledge obtains a set of standardized optical features with consistent channel dimensions and optimized by local context, thus preparing it numerically for incorporating terrain features. S1.2 Lightweight Multi-Scale Terrain Feature Encoding: To achieve strict alignment of terrain features with the optical backbone network in terms of spatial resolution, an independent lightweight multi-scale terrain feature encoder (DEM Encoder) is constructed. Unlike networks loaded with heavy pre-trained weights, the terrain feature encoder consists of five cascaded convolutional stages, specifically designed for hierarchical extraction of terrain semantics from single-channel DEM data. Each encoding stage includes a convolutional layer, batch normalization, and ReLU activation function. By setting specific convolutional stride and pooling operations, the terrain feature encoder generates five sets of terrain feature maps with progressively decreasing spatial resolution (gradually reducing the dimensionality from 1 / 2 of the input size to 1 / 32) and a fixed number of channels. , This layered mechanism ensures that topographic features can comprehensively cover multi-scale spatial structure information, from local micro-topography to macro-topography. S1.3 Layer-by-layer geometric feature injection and fusion: After acquiring multi-scale terrain features and standardized optical features, a layer-by-layer injection strategy is adopted to enhance the spectral features with prior terrain features; for the first... For each level of features, bilinear interpolation is first used to force the terrain features to spatially align to a resolution consistent with the current level's optical features; subsequently, a feature concat operation is performed along the channel dimension.
[0037] In the formula, For the first Standardized optical characteristics of the layer The aligned topographic structure features of the i-th layer; the stitched feature tensor The system simultaneously incorporates both spectral texture and topographic geometry information of land features. Finally, a fusion layer composed of 3×3 convolutions is used to perform deep computation on the mixed features, enabling the network to automatically learn the spatial mapping relationship between "topographic elevation" and "water distribution". Through the above steps, this invention successfully and explicitly injects elevation priors into the network, ultimately outputting enhanced optical features with terrain-aware capabilities. This not only filters out spurious changes caused by complex terrain from a physical and logical perspective, but also provides high-quality, unambiguous baseline feature input for the next step of "cross-modal feature alignment and interaction" with SAR imagery.
[0038] Step S2: Construct a cross-modal feature alignment module based on dual attention guidance. The optical image features containing terrain features output by the terrain prior feature encoding module in step S1 are aligned with SAR image features. High-value semantic dimensions are selected through cross-modal channel attention, and combined with cross-modal spatial attention, SAR features are reconstructed and denoised with optical features as guidance, thereby providing high-purity feature representation for the final mask prediction. The specific operation of step S2 is as follows: S2.1, Spatial Attention Feature Enhancement: The purpose of the spatial attention mechanism is to overcome the limitations of the local receptive field in traditional convolution. By explicitly modeling the dependencies between arbitrary spatial locations in the feature map, it enhances the model's ability to perceive key areas of global structure and flood changes. This is achieved by enhancing the deepest features output by the optical and SAR branches. and The spatial attention module first generates the corresponding query Q, key K, and value V feature matrices through 1×1 convolutions, and then flattens them into C×N vector representations (where N = H×W) in the spatial dimension. Subsequently, by calculating the matrix multiplication between Q and K and applying Softmax normalization, single-modal spatial attention matrices are generated.
[0039] To achieve efficient interaction of cross-modal information, the spatial attention results of the two modalities are fused to construct a cross-modal spatial cross-attention matrix. This simultaneously encodes the synergistic spatial variation relationship between optical and SAR features; subsequently, the cross matrix is applied to the value feature V and fused with the original feature via residual connection:
[0040] Finally, the spatial attention features from the optical and SAR branches are added element-wise to form the final cross-modal spatial attention feature representation: ; S2.2 Channel attention machines characterize the semantic differences in the importance of different feature channels. In multimodal scenarios, the semantic channels corresponding to different sensors (optical reflectivity and radar backscattering) differ significantly. Introducing cross-modal channel attention is key to improving feature fusion quality and aligning heterogeneous semantics. Similar to spatial attention, the channel module also... and It generates corresponding Q, K, and V features for the input; the difference lies in that it focuses on the similarity between channels, by calculating Q and The channel attention matrix is obtained from the correlation at the channel dimension:
[0041] Subsequently, the channel attention matrices of the two methods are fused to construct a cross-modal channel attention representation. This is then used to weight the value feature V, and channel self-attention features are obtained through residual connections: ; Finally, the channel features of the two modalities are added element-wise to obtain the cross-modal channel attention feature representation: ; S2.3 Dual Attention Feature Fusion and Alignment Output: After acquiring cross-modal enhancement features at the spatial and semantic levels respectively, this step generates channel attention feature maps. Spatial attention feature map Concatenation or addition can be performed along the channel dimension:
[0042] Through weighting and reconstruction using a dual attention mechanism, the model simultaneously enhances cross-modal feature representation at both the location (spatial) and semantic (channel) levels. Discrete noise in SAR images is effectively filtered out under the "guidance and suppression" of optical texture features, resulting in a high-purity combined attention map output. The deepest semantic features are input into the final decoder, thereby significantly improving the recognition accuracy and recovery capability of complex water body edges.
[0043] Step S3: Construct a multi-task water feature enhancement module, receive the attention map output in step S2, and input the attention map as the deepest semantic feature into the multi-task decoding network. Perform the main task change detection and the auxiliary task water segmentation respectively. Through the semantic constraints of backpropagation of the auxiliary branch, the preceding attention module and encoder are further forced to decouple noise from the real water body; thus, the distribution map of water cover change caused by surface flooding is obtained.
[0044] The specific operation of step S3 is as follows: S3.1 Construction of a Multi-Task Water Feature Enhancement Module: To effectively suppress the inherent speckle noise interference of SAR images at the source of feature extraction, this study constructs a lightweight auxiliary water segmentation branch in addition to the change detection backbone network. Considering the real-time requirements of flood monitoring and to avoid computational redundancy, the multi-task water feature enhancement module is built on the deepest enhanced features of the shared encoder. Although the deepest features of the shared encoder have lower spatial resolution, they have the strongest semantic abstraction ability and the largest receptive field, and best reflect the overall properties of the water body. The water segmentation task branch consists of only a 3×3 convolution (for feature smoothing), a 1×1 convolution (for channel compression), and a sigmoid activation function, outputting a single-channel water probability map. Its core innovation lies in constructing a short-path gradient flow from the semantic label directly to the encoder. During the training phase, when the encoder produces incorrect activations due to overfitting to SAR speckle noise, the water segmentation task branch generates a huge penalty gradient; gradient backpropagation acts as an explicit feature regularizer, forcing the encoder parameters to be updated in the direction of "suppressing discrete noise and responding to homogeneous water bodies". S3.2, Multi-task Hybrid Loss Function Design and Joint Optimization: To alleviate the severe imbalance between positive and negative samples in flood detection and to guide the stable convergence of the dual-branch network, this step designs a multi-task hybrid loss function training framework. For the main task of change detection, a hybrid loss function combining Weighted Binary Cross-Entropy (BCE) and Dice Loss is adopted. Weighted binary cross-entropy assigns higher weights to the minority class (variable flood areas), strengthening the network's focus on sparse samples; Dice Loss optimizes the set similarity between predicted results and ground truth from a global perspective, improving the internal integrity of flood patches. Its calculation formula is as follows:
[0045] In the formula, N is the total number of pixels. For the first The true label is a pixel, where 1 represents a positive sample with variation and 0 represents the background. This represents the probability value predicted by the network. These are the weighting coefficients for positive samples. To prevent the denominator from being zero and to smooth the gradient from a minimum; For the auxiliary water body segmentation branch, the standard binary cross-entropy loss function is used. As a supervised objective, since the water segmentation branch does not pursue extreme boundary segmentation but aims to provide smooth and continuous gradient feedback, the standard BCE can robustly guide the model to decouple water features from the background in a low-resolution, high-dimensional semantic space, avoiding gradient oscillations in the early stages of training caused by using boundary-sensitive losses (such as Dice). The formula is as follows:
[0046] Finally, the total loss function for multi-task joint optimization is obtained by weighted summation of the main task loss and the auxiliary task loss:
[0047] In the formula, For multi-task balancing hyperparameters; experimental verification shows that when Setting it to 0.3 can minimize overall loss while ensuring the accuracy of change detection tasks; Through the above multi-task collaboration and hybrid loss constraints, the features of the main branch of inflow change detection are "pre-cleaned" at the source, and non-water body features and SAR high-frequency noise are effectively removed, significantly improving the signal-to-noise ratio of subsequent difference map generation; the main and auxiliary tasks are dynamically optimized under a unified loss framework, and finally output a high-precision flood change mask with continuous boundaries and no voids.
[0048] Step S4, loss function and evaluation metrics: The Adam optimizer of adaptive moment estimation is used to train a flood change detection model guided by terrain prior knowledge. The training loss value between the actual flood inundation range and the model prediction value is calculated by a hybrid loss function consisting of cross-entropy loss and Dice loss, so as to effectively alleviate the problem of severe imbalance between the proportion of water body and background pixels in remote sensing images.
[0049] The specific operation of step S4 is as follows: The accuracy verification criteria are divided into two dimensions: pixel-level classification accuracy and spatial morphology accuracy. At the pixel-level classification level, the overall accuracy (OA), precision (Precision), recall (Recall), F1 score (F1-Score), and intersection-over-union (IoU) are used to evaluate the model's ability to capture flood-affected areas. Generally speaking, the higher the values of these five metrics, the better the model's pixel-level segmentation performance.
[0050] To address the spatial physical characteristic that flood-inundated areas typically exhibit as continuous patches rather than highly discrete fragments, the Perimeter-Area Ratio (PAR) and the Shape Fragmentation Index (SFI), based on the isoperimetric inequality, are introduced to provide a constrained evaluation of the model output from a topological perspective. PAR characterizes boundary complexity and internal voids, while SFI measures the deviation of the target shape from an ideal compact structure. Generally, smaller values for these two morphological indices indicate a more compact water patch structure, smoother boundaries, and a stronger ability to suppress noise fragmentation. The calculation formulas for the main evaluation indices are as follows:
[0051] Wherein, TP (True Positive) is the number of pixels that are both true floods and predicted as floods; FP (False Positive) is the number of pixels that are true non-floods but predicted as floods; FN (False Negative) is the number of pixels that are true floods but predicted as non-floods; in the morphological indices, A is the total pixel area of water bodies in the predicted results; L is the total boundary length of all water body patches (including the outer contour and the edge of the internal voids); SFI is a normalized dimensionless quantity, with a value range of [0, 1].
[0052] Model Structure Comparison and Analysis: To verify the superiority of the proposed cross-modal flood change detection model (the complete method of this invention) that integrates topographic priors and multi-task water feature enhancement, it is compared with several mainstream and widely used deep learning models for change detection. All models use the same self-made cross-modal flood change detection dataset (CMFlood), with a total of 2,680 sample pairs, the composition of which is described in [link to dataset]. Figure 2 The dataset includes preflood optical imagery, afterflood imagery, topographic elevation map (DEM), and change label data. The dataset is randomly divided into training, testing, and validation sets in a 6:2:2 ratio, and the input images are standardized to a resolution of 256×256 pixels.
[0053] This experiment was conducted on a Windows Server 2019 system environment, using one Intel Xeon Silver 4316 processor (2.3GHz) and one Nvidia GeForce RTX 4090 GPU with 24GB of video memory. Based on this, the PyTorch (v1.12.1) deep learning framework was used for model design and construction, and the timm library was used to call the pre-trained backbone network. During training, the Adam optimizer was used to learn the model weights, with an initial learning rate set to 2e-5, a batch size of 8, and a total of 100 training epochs. To address the severe imbalance between water bodies and background pixels in remote sensing imagery, a hybrid loss function consisting of cross-entropy loss and Dice loss was constructed, and the best-performing model parameters during training were saved.
[0054] The selected comparison models cover classic change detection models (such as BIT) and recent mainstream deep networks (such as FC-Diff, CDNeXt, SNUNet, CMCD-Net, etc.). Accuracy validation criteria use overall accuracy (OA), precision, recall, F1 score, and intersection-over-union ratio (IoU) to evaluate the model's ability to capture flood change areas and its boundary fit.
[0055] Table 1 Comparison of experimental results
[0056] Table 2 Morphological indices of different change detection models
[0057] To ensure fairness in the comparison, all detection models were trained and tested in the same hardware and software environment.
[0058] As shown in Table 1, TGFE-CDNet exhibits the best overall performance. In terms of key metrics F1 score and IoU, our proposed method achieves 91.09% and 84.03%, respectively. Compared to the second-best model, MTCDN, it surpasses TGFE-CDNet by 1.22% in F1 score and SNUNet, the best in IoU among our comparative models, by 3.37%. Although SNUNet has a 0.97% higher recall rate than our proposed method, its precision (83.84%) is significantly lower than TGFE-CDNet's 92.95%. In terms of overall accuracy (OA), it is 0.97% lower than the MTCDN model. These results demonstrate that the proposed TGFE-CDNet significantly reduces the false alarm rate while maintaining a low false negative rate, enabling more accurate elimination of pseudo-variable regions.
[0059] Table 2 shows the morphological indices of the prediction results of each model. It is worth noting that the morphological indices of each model in this experiment are generally at a high level. This is mainly attributed to the natural topological features of floodwater bodies, which are often elongated or dendritic, and the jagged edges caused by speckle noise in SAR images. The high index level of the ground truth (Label) objectively reflects the geometric complexity of the scene; therefore, the key to evaluating model performance lies in measuring the degree of deviation of its results from the ground truth.
[0060] The ground truth (label) model exhibits the lowest perimeter-to-area ratio and shape fragmentation index, quantitatively characterizing the smooth and dense topological properties of natural water body edges. In the comparative models, the perimeter-to-area ratios of MTCDN and CDNeXt rise to 0.8761 and 0.8701, respectively, significantly higher than the ground truth, reflecting severe edge roughness and jaggedness in the prediction results. Meanwhile, the shape fragmentation indices of ELGC-Net and HFA-Net both exceed 0.75, revealing certain spatial discretization issues in their results. The proposed TGFE-CDNet achieves significant improvements in key morphological features. Its perimeter-to-area ratio decreases to 0.8273, demonstrating the model's effective suppression of edge noise; simultaneously, its shape fragmentation index is controlled at 0.7328, effectively avoiding the high degree of discretization observed in models such as ELGC-Net. Overall, TGFE-CDNet outperforms mainstream comparative methods in suppressing edge jaggedness and maintaining spatial connectivity of ground features.
[0061] Figure 3 To compare with the visualization results of the experiment, these results further corroborate the above quantitative analysis. As can be seen from the figure, methods such as ELGC-Net and STANet tend to generate more salt-and-pepper noise and boundary adhesion when facing complex backgrounds. In contrast, the detection map generated by TGFE-CDNet is closest to the label, effectively suppressing not only the speckle noise interference commonly found in radar images but also maintaining the internal continuity and boundary integrity of the changing areas. Furthermore, the false positive and false negative patch information in the figure also shows that TGFE-CDNet, after acquiring richer and purer semantic information about the water body, effectively solves the problems of false positives and false negatives.
[0062] Analysis of the morphological indices in Table 1 (comparison results) and Table 2 (model experiment results) Figure 3 Visual features reveal the following: 1) Excellent overall performance with progressive advantages: The proposed complete model demonstrates absolute superiority in all core metrics. Experiments show that simply introducing terrain priors on top of the baseline model significantly improves precision and effectively reduces false positives caused by terrain elevation differences. Furthermore, the complete method of this invention, after further introducing multi-task water feature enhancement and gradient constraints, achieves an overall F1 score of 91.09%, a significant improvement over the best result of the comparison model. This verifies that the present invention can focus more on stable water body region features while extracting change information.
[0063] 2) High convergence efficiency and robustness: During the training phase, this invention not only exhibits extremely fast convergence speed and can efficiently lock the effective gradient direction, but also maintains a highly smooth curve without significant oscillations during long training epochs. This proves that the combination of the multi-branch architecture and hybrid loss function in this invention not only has high feature fitting efficiency, but also possesses excellent training robustness in complex scenarios.
[0064] 3) Verification of dual capabilities in terrain perception and noise suppression: Analysis of morphological visualization results shows that in complex areas with low surface reflectivity or dramatic terrain undulations, other contrast models are prone to large-area missed detections or spurious change responses. However, the method of this invention maintains extremely high detection integrity by relying on the guidance of terrain structure features. At the same time, under the interference of high-frequency speckle noise in SAR images, contrast methods often result in fragmented detection results. This invention effectively decouples noise interference by relying on multi-task semantic constraints and dual attention mechanisms, greatly enhancing the continuity and integrity of water body boundaries.
[0065] By integrating terrain mechanisms with multi-task water feature enhancement strategies, this invention demonstrates significantly higher accuracy and adaptability than single-module or traditional deep learning methods in cross-modal flood spatiotemporal detection tasks, exhibiting clear geographic interpretability and strong noise suppression capabilities. This invention constructs a change detection framework constrained by both physical mechanisms (terrain prior) and semantic mechanisms (multi-task learning), comprehensively revealing the intrinsic dependency between multimodal image features and the actual inundation distribution. Based on this, it thoroughly solves the problems of missed detections, false detections, and fragmentation of floods under complex terrain and strong noise interference.
Claims
1. A method for constructing a large-scale model for flood change detection based on prior terrain knowledge, characterized in that, The large-scale flood change detection model guided by terrain prior knowledge includes a terrain prior feature encoding module, a dual-attention-guided cross-modal feature alignment module, and a multi-task water body feature enhancement module. It extracts features from pre-disaster optical and post-disaster SAR images through a pseudo-twin dual-branch architecture and introduces a lightweight DEM feature extraction module to explicitly encode elevation information as terrain prior features, providing clear terrain prior constraints for the network. The channel and spatial dual-attention-guided cross-modal feature alignment module facilitates multi-scale feature interaction, bridging the heterogeneous semantic gap between optical and SAR. Subsequently, a water body semantic segmentation branch is introduced to construct a multi-task water body feature enhancement module, and gradient constraints are applied to force the network to decouple speckle noise and focus on stable water bodies at the feature level. Finally, multiple F1 metrics are used to calculate the accuracy of the prediction results, and a high-precision flood change mask is output. A method for constructing a large-scale model for flood change detection based on prior terrain knowledge includes the following steps: Step S1: Construct a terrain prior feature encoding module based on lightweight DEM feature extraction. Introduce DEM features extracted by the lightweight DEM feature extraction module as auxiliary information. Constrain two-dimensional spectral features through three-dimensional geometric features to effectively suppress pseudo-change responses in optical images that contradict the terrain logic. The construction of the terrain prior feature encoding enhancement module based on lightweight DEM extraction includes the following operation steps: feature normalization calibration, lightweight terrain encoding, and layer-by-layer geometric feature injection. Step S2: Construct a cross-modal feature alignment module based on dual attention guidance. Using the terrain prior feature encoding module of step S1, the optical features and SAR features output by S1 are input into the cross-modal channel attention to filter high-value semantic dimensions. Combined with cross-modal spatial attention, the SAR features are reconstructed and denoised with optical features as guidance, thereby providing high-purity feature representation for the final mask prediction. Step S3: Construct a multi-task water feature enhancement module. Receive the attention map output from step S2 and input the attention map as the deepest semantic feature into the multi-task decoding network. Execute the main task of change detection and the auxiliary task of water segmentation respectively. Through semantic constraints propagated back through the auxiliary branch, the preceding attention module and encoder are further forced to decouple noise from the real water body. This enables the model to extract features of the complete water body, thereby improving the accuracy of flood change detection in identifying changed areas. Step S4, loss function and evaluation metric: The Adam optimizer is used to train the model, and the training loss value between the actual flood inundation range and the model prediction value is calculated by a hybrid loss function consisting of Cross-Entropy Loss and Dice loss, so as to effectively alleviate the problem of severe imbalance between the proportion of water bodies and background pixels in remote sensing images.
2. The method for constructing a large-scale flood change detection model based on terrain prior knowledge as described in claim 1, characterized in that, The specific operation of step S1 is as follows: S1.1 Before feature injection, a Residual Convolution Block (RCB) is introduced to preprocess the optical features. The RCB first uses 1×1 convolution to project features from different levels to a fixed low-dimensional feature space, achieving channel unification. Then, it connects to a convolution block containing residual connections for local feature extraction. The update process of the RCB can be described as follows: In the formula, These are low-dimensional features after 1×1 convolutional projection. This represents a 3×3 convolutional nonlinear transformation and the ReLU activation operation; Through this calibration, the large-scale flood change detection model guided by prior terrain knowledge obtained a set of standardized optical features with consistent channel dimensions and local context optimization, thus preparing it numerically for incorporating terrain features. S1.2 Lightweight Multi-Scale Terrain Feature Encoding: To achieve strict alignment of terrain features with the optical backbone network in terms of spatial resolution, an independent lightweight multi-scale terrain feature encoder (DEM Encoder) is constructed. Unlike networks loaded with heavy pre-trained weights, the lightweight DEM feature extraction module consists of five cascaded convolutional stages used to extract terrain semantics hierarchically from single-channel DEM data. Each encoding stage includes a convolutional layer, batch normalization, and ReLU activation function. By setting specific convolutional stride and pooling operations, the encoder generates five sets of terrain feature maps with sequentially decreasing spatial resolution and a fixed number of channels. , This layered mechanism ensures that topographic features can comprehensively cover multi-scale spatial structure information, from local micro-topography to macro-topography. S1.3 Layer-by-layer geometric feature injection and fusion After acquiring multi-scale terrain features and standardized optical features, a layer-by-layer injection strategy is adopted to inject DEM features into multi-level optical features, thereby enhancing the spectral features with geometric priors. For the For each level of features, bilinear interpolation is first used to force the terrain features to spatially align to a resolution consistent with the current level's optical features; subsequently, a feature concat operation is performed along the channel dimension. In the formula, For the first Standardized optical characteristics of the layer The i-th layer of terrain structure features after alignment; The concatenated feature tensor The system simultaneously incorporates both spectral texture and topographic geometry information of land features. Finally, a fusion layer composed of 3×3 convolutions is used to perform deep computation on the mixed features, enabling the network to automatically learn the spatial mapping relationship between "topographic elevation" and "water distribution".
3. The method for constructing a large-scale flood change detection model based on terrain prior knowledge as described in claim 1, characterized in that, The specific operation of step S2 is as follows: S2.1, Spatial Attention Feature Enhancement: The purpose of the spatial attention mechanism is to overcome the limitations of the local receptive field in traditional convolution. By explicitly modeling the dependencies between arbitrary spatial locations in the feature map, it enhances the model's ability to perceive key areas of global structure and flood changes. This is achieved by enhancing the deepest features output by the optical and SAR branches. and The spatial attention module first generates the corresponding query Q, key K and value V feature matrices through 1×1 convolution, and then flattens them into a C×N vector representation in the spatial dimension. Subsequently, by calculating the matrix multiplication between Q and K and applying Softmax normalization, the spatial attention matrices for each unimodal mode are generated: ; To achieve efficient interaction of cross-modal information, the spatial attention results of the two modalities are fused to construct a cross-modal spatial cross-attention matrix. This simultaneously encodes the synergistic spatial variation relationship between optical and SAR features; subsequently, the cross matrix is applied to the value feature V and fused with the original feature via residual connection: ; Finally, the spatial attention features from the optical and SAR branches are added element-wise to form the final cross-modal spatial attention feature representation: ; S2.2 Channel attention machine characterizes the differences in the importance of different feature channels at the semantic level; in multimodal scenarios, the semantic channels corresponding to different sensors have significant differences. Introducing cross-modal channel attention is the key to improving the quality of feature fusion and aligning heterogeneous semantics. Similar to spatial attention, the channel module also uses... and Generate corresponding Q, K, and V features for the input; different The key point is that it focuses on the similarity between channels, by calculating Q and The channel attention matrix is obtained from the correlation at the channel dimension: ; Subsequently, the channel attention matrices of the two methods are fused to construct a cross-modal channel attention representation. This is then used to weight the value feature V, and channel self-attention features are obtained through residual connections: ; Finally, the channel features of the two modalities are added element-wise to obtain the cross-modal channel attention feature representation: ; S2.3 Dual Attention Feature Fusion and Alignment Output: After acquiring cross-modal enhancement features at the spatial and semantic levels respectively, this step generates channel attention feature maps. Spatial attention feature map Concatenation or addition can be performed along the channel dimension: ; Through weighted and reconstructed processing using a dual attention mechanism, the model simultaneously enhances cross-modal feature representation at both the location and semantic levels. Discrete noise in SAR images is effectively filtered out by the "guidance and suppression" of optical texture features, resulting in a high-purity combined attention map output. The deepest semantic features are input into the final decoder, thereby improving the recognition accuracy and recovery capability of complex water body edges.
4. The method for constructing a large-scale flood change detection model based on terrain prior knowledge as described in claim 1, characterized in that, The specific operation of step S3 is as follows: S3.1 Multi-task Feature Enhancement Module Construction: The multi-task feature enhancement module introduces a water segmentation task branch based on the change detection task. The water segmentation branch is built on the deepest features of the encoder and is constructed through gradient backpropagation. Although the deepest features have lower spatial resolution, they have the strongest semantic abstraction ability and the largest receptive field, and best reflect the overall properties of the water body. The water segmentation branch consists of only a 3×3 convolution, a 1×1 convolution, and a sigmoid activation function, outputting a single-channel water probability map. During the training phase, when the encoder produces incorrect activations due to overfitting SAR speckle noise, the water segmentation task branch generates a large penalty gradient. Gradient backpropagation acts as an explicit feature regularizer, forcing the encoder parameters to be updated in the direction of "suppressing discrete noise and responding to homogeneous water bodies". S3.2 Design and Joint Optimization of Multi-Task Hybrid Loss Function: For the main task of change detection, a hybrid loss function combining Weighted Binary Cross-Entropy (BCE) and Dice Loss is adopted. Weighted binary cross-entropy strengthens the network's focus on sparse samples by assigning higher weights to the minority class; Dice Loss optimizes the set similarity between predicted results and ground truth from a global perspective, improving the internal integrity of flood patches. Its calculation formula is as follows: ; In the formula, N is the total number of pixels. For the first The true label is a pixel, where 1 represents a positive sample with variation and 0 represents the background. This represents the probability value predicted by the network. These are the weighting coefficients for positive samples. To prevent the denominator from being zero and to smooth the gradient from a minimum; For the auxiliary water body segmentation branch, the standard binary cross-entropy loss function is used. As a supervised objective; since the water segmentation branch does not pursue extreme boundary segmentation, but aims to provide smooth and continuous gradient feedback, the standard BCE can robustly guide the model to decouple water features from the background in a low-resolution high-dimensional semantic space, avoiding gradient oscillations in the early stages of training caused by using boundary-sensitive loss; the formula is as follows: ; Finally, the total loss function for multi-task joint optimization is obtained by weighted summation of the main task loss and the auxiliary task loss: ; In the formula, To balance hyperparameters for multi-task operation; Experiments have verified that when Setting it to 0.3 can minimize overall loss while ensuring the accuracy of change detection tasks; Through the above multi-task collaboration and hybrid loss constraints, the features of the main branch of inflow change detection are "pre-cleaned" at the source, and non-water body features and SAR high-frequency noise are effectively removed, significantly improving the signal-to-noise ratio of subsequent difference map generation; the main and auxiliary tasks are dynamically optimized under a unified loss framework, and finally output a high-precision flood change mask with continuous boundaries and no voids.
5. The method for constructing a large-scale flood change detection model based on prior terrain knowledge as described in claim 1, characterized in that, The specific operation of step S4 is as follows: The accuracy verification criteria are divided into two dimensions: pixel-level classification accuracy and spatial morphology accuracy. At the pixel-level classification level, the overall accuracy (OA), precision (Precision), recall (Recall), F1 score (F1-Score), and intersection-over-union (IoU) are used to evaluate the model’s ability to capture flood-affected areas. To address the spatial physical characteristic that flood-inundated areas typically exhibit as continuous patches rather than highly discrete fragments, we introduce the perimeter-to-area ratio (PAR) and the shape fragmentation index (SFI) based on the isoperimetric inequality to constrain the model output from a topological perspective. PAR characterizes boundary complexity and internal voids; SFI measures the deviation of the target shape from an ideal compact structure. Smaller values for these two morphological indices indicate a more compact water feature structure, smoother boundaries, and a stronger ability to suppress noise fragmentation. The calculation formulas for the main evaluation indices are as follows: ; Where TP is the number of pixels that are actually floods and predicted as floods; FP is the number of pixels that are actually not floods but predicted as floods; FN is the number of pixels that are actually floods but predicted as not floods; in the morphological indices, A is the total pixel area of water targets in the prediction results; L is the total boundary length of all water body patches; SFI is a normalized dimensionless quantity, whose value ranges from [0, 1).