Floodwater-oriented optical and sar collaborative domain adaptive segmentation method
By using an optical and SAR collaborative domain adaptive segmentation method, the problem of cross-modal complementarity between optical and SAR images in flood disaster remote sensing monitoring is solved, and the accuracy of water accumulation range and boundary extraction is improved in complex backgrounds. This method is suitable for flood emergency mapping and urban waterlogging monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UESTC (SHENZHEN) ADVANCED RES INST
- Filing Date
- 2026-04-28
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies for remote sensing monitoring of flood disasters, optical remote sensing images are easily affected by cloud and fog obscuration and complex surface backgrounds, while SAR images are easily affected by speckle noise and strong scattering, resulting in inaccurate extraction of the water accumulation range and boundaries. Furthermore, the generalization ability across regions, seasons, and sensor conditions is insufficient.
An optical and SAR collaborative domain adaptive segmentation method is constructed. Through bi-branch feature extraction, cross-modal attention enhancement, teacher-student domain adaptive training, pseudo-label screening, and dynamic weight control, multi-scale cross-modal attention alignment and boundary enhancement of optical and SAR images are achieved, thereby improving the extraction accuracy of water accumulation range and boundary.
Under conditions where no manual annotation is required in the target disaster area, it significantly reduces the decline in model accuracy, minimizes interference from clouds, shadows, noise, and strong scattering, improves the accuracy of water accumulation extraction and the integrity of boundaries, and is suitable for rapid inference and mosaicking of large-format remote sensing images.
Smart Images

Figure CN122115873A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing monitoring technology for flood disasters, and in particular to an adaptive segmentation method for optical and SAR co-domains for flood water accumulation. Background Technology
[0002] In remote sensing monitoring of flood disasters, rapidly and accurately obtaining information on the extent and boundaries of water accumulation in disaster areas is crucial for emergency mapping, disaster assessment, and subsequent rescue and dispatch. Because floods are often accompanied by continuous rainfall, cloud cover, and complex surface backgrounds, relying solely on optical remote sensing imagery is prone to errors or omissions due to shadows, clouds, dark features, and turbid water. While relying solely on SAR (Synthetic Aperture Radar) imagery (SAR imagery is active microwave remote sensing imagery acquired by a synthetic aperture radar system) offers all-weather, day-and-night observation capabilities, it is susceptible to speckle noise, strong scattering from buildings, mountain shadows, and cross-regional imaging differences. Therefore, the "optical + SAR" approach to flood water accumulation extraction has become an important technical route in remote sensing disaster monitoring.
[0003] From the perspective of existing technical approaches, the relevant solutions can be broadly divided into three categories: The first category is the method of extracting water bodies by combining active and passive remote sensing data, which focuses on solving the limitations of single optical or single SAR; the second category is the deep learning water body detection method based on multi-source remote sensing images, which focuses on improving the segmentation accuracy under noise label conditions; and the third category is the remote sensing image classification or recognition method based on multimodal attention fusion, which focuses on exploring the correlation and complementarity between different modalities.
[0004] However, the existing technologies still have the following shortcomings: First, many "optical + SAR" water extraction schemes mainly rely on threshold rules, coarse-fine combination or post-processing constraints, and lack a deep cross-modal feature alignment mechanism that can adapt to complex disaster area backgrounds; Second, although existing multi-source water detection schemes use deep learning, they usually assume that the training and testing scenarios are relatively consistent, and do not adequately consider the generalization ability under cross-regional, cross-seasonal and cross-sensor conditions; Third, most existing multimodal attention methods are geared towards classification tasks and have not been further combined with pixel-level flood boundary segmentation.
[0005] For example, the invention patent CN114202698A, entitled "A Method for Precise Water Body Extraction Using Combined Active and Passive Remote Sensing Data," improves the spatiotemporal accuracy and precision of large-scale water body extraction by combining Sentinel-1 active radar data with Sentinel-2 passive optical remote sensing data. This invention is closest to the present invention in terms of application scenario, already embodying the basic idea of "optical + SAR / active and passive remote sensing collaborative water body extraction," which can reduce errors caused by urban shadows, rooftops, and turbid water. However, this method still leans towards a rule-based and phased processing approach, focusing on coarse extraction followed by fine extraction. It lacks a deep cross-modal semantic alignment mechanism for flood and waterlogging boundaries and does not address the domain offset problem under unlabeled cross-regional target domain conditions.
[0006] The invention patent with publication number CN111666849B, entitled "Multi-view depth network iterative evolution method for water body detection in multi-source remote sensing images," discloses the following technical solution: (1) Divide the original dataset into multiple non-overlapping subsets; (2) Train a deep semantic segmentation network with multiple perspectives using different subset datasets; (3) Utilize multi-view network to collaboratively update labels and retrain the network; (4) During the testing phase, the results of the multi-view network are voted on, and the final water body detection results are output.
[0007] This patent emphasizes "collaborative label updates" and "achieving a good deep semantic segmentation network after multiple iterations," aiming to solve the problems of low resolution and high noise in water body labels during training data. This invention has transitioned from traditional rule-based methods to a deep semantic segmentation approach, focusing on improving detection accuracy under conditions of multi-source remote sensing imagery and noisy labels. However, it primarily addresses the issues of training label quality and multi-model collaboration, without designing explicit cross-modal attention alignment between optical and SAR technologies, nor constructing a source-target domain adaptive training framework for cross-disaster area migration scenarios.
[0008] Therefore, an optical-SAR collaborative domain adaptive segmentation method is needed for fine extraction of floodwater across regions. This method combines the texture / spectral information of optical images with the all-weather scattering information of SAR images. Under the condition that the target disaster area is unlabeled, the method improves the extraction accuracy of the water accumulation range and boundary through cross-modal attention enhancement, teacher-student pseudo-label learning and boundary perception constraints. Summary of the Invention
[0009] The technical problem to be solved by the embodiments of the present invention is to provide an optical and SAR collaborative domain adaptive segmentation method for flood water accumulation, so as to improve the extraction accuracy of water accumulation range and boundary.
[0010] To address the aforementioned technical problems, this invention proposes an optical and SAR co-domain adaptive segmentation method for flood-prone waterlogging, comprising: Step 1: Construct source and target domain sample sets. The source domain samples consist of optical images with pixel-level annotations of flood water accumulation, SAR images corresponding to them in time and space, and manually annotated masks. The target domain samples consist of optical images and SAR images of the area to be interpreted. The two modal samples in the sample sets are preprocessed to ensure that the same pixel location corresponds to the same geographical area. Step 2: Using the two modal samples after preprocessing in the sample set as collaborative input, perform bi-branch feature extraction, and enhance the features of the two modalities through a cross-modal attention mechanism to obtain multi-scale joint features and segmentation / boundary prediction results; Step 3: Construct student and teacher models with identical structures. Supervised training of the student model is performed using source domain samples, and consistent training of the student model is performed using target domain samples. The teacher model does not directly participate in gradient backpropagation; its parameters are updated by the student model parameters through exponential moving average. Step 4: Filter target domain pseudo-tags and perform boundary enhancement; Step 5: Construct the loss function, perform dynamic weight adjustment, backpropagate to update the student model parameters, and obtain the trained student model; Step 6: Input the optical image and SAR image of the area to be tested into the trained student model to obtain the corresponding flood probability map, binary mask map and boundary map.
[0011] The beneficial effects of this invention are as follows: 1. In the absence of manual pixel-level annotation in the target disaster area, this invention achieves cross-regional transfer through teacher-student domain adaptive training, which can significantly reduce the accuracy degradation problem when the model is directly transferred.
[0012] 2. By aligning the multi-scale cross-modal attention of optical and SAR images, this invention can simultaneously leverage the advantages of optical texture and the all-weather advantages of SAR to reduce the interference of clouds, shadows, low illumination, speckle noise, and strong scattering on flood extraction results.
[0013] 3. By jointly screening target domain pseudo-labels through cross-modal consistency constraints and boundary uncertainties, the propagation of erroneous pseudo-labels in boundary regions and complex terrain areas can be effectively reduced, thereby improving training stability and target domain segmentation robustness.
[0014] 4. Through the design of boundary reinforcement heads and boundary loss, this invention can improve the boundary integrity and fine depiction capabilities of building edges, road edges, embankment edges, and localized small water accumulation areas.
[0015] 5. Through a dynamic weight control mechanism, the influence of auxiliary tasks can be adaptively adjusted according to the training stage and the credibility of pseudo-labels, avoiding under-constraint or over-constraint problems caused by fixed hyperparameters in different regions and with different sample quality.
[0016] 6. This invention employs non-adversarial domain adaptive training, which does not rely on an additional domain discriminator. The training process is relatively stable and facilitates engineering deployment and parameter tuning.
[0017] 7. This invention is applicable to sliding window inference and mosaic output of large-format remote sensing images, facilitating rapid implementation in flood emergency mapping, urban waterlogging monitoring, and post-disaster assessment. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating the optical and SAR cooperative domain adaptive segmentation method for flood-prone waterlogging according to an embodiment of the present invention.
[0019] Figure 2 This is a schematic diagram of dual-branch feature extraction according to an embodiment of the present invention.
[0020] Figure 3 This is a schematic diagram of teacher-student domain adaptive training according to an embodiment of the present invention. Detailed Implementation
[0021] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] In this embodiment of the invention, directional indicators (such as up, down, left, right, front, back, etc.) are only used to explain the relative positional relationship and movement of each component in a specific posture (as shown in the figure). If the specific posture changes, the directional indicator will also change accordingly.
[0023] Furthermore, in this invention, descriptions involving "first," "second," etc., are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features.
[0024] Please refer to Figure 1The optical and SAR collaborative domain adaptive segmentation method for flood-prone areas includes steps 1 to 6. Specifically, the standardized source / target domain multimodal samples output in step 1 serve as the dual-branch input for step 2; the multi-scale joint features and segmentation / boundary prediction results output in step 2 are used for teacher-student training, pseudo-label screening, and loss construction in steps 3 to 5, respectively; the stable model obtained after parameter updating in step 5 is used for inference output in step 6, thus forming a complete computational processing chain that connects the preceding and following steps.
[0025] This invention does not involve simple stitching or fusion of optical and SAR images. Instead, it addresses the specific challenges of fine-grained extraction of flood-affected areas across regions. It tackles the practical problems of "lack of labeled target domains, susceptibility of optical images to cloud and fog interference, susceptibility of SAR images to speckle noise and strong scattering, and difficulty in accurately extracting flood boundaries" in flood scenarios. The invention constructs a complete technical solution comprising bi-branch feature extraction, cross-modal attention enhancement, teacher-student domain adaptive training, pseudo-label selection, boundary enhancement, and dynamic weight control. This allows for stable extraction of the flood-affected area's extent and boundaries, even when the source domain is labeled but the target domain is unlabeled. Utilizing the texture details of optical images and the all-weather observation capabilities of SAR images, the invention achieves stable extraction of the flood-affected area's extent and boundaries through cross-modal attention enhancement, teacher-student domain adaptive training, pseudo-label selection, boundary enhancement, and dynamic weight adjustment.
[0026] Step 1, Training Data Construction and Preprocessing: First, construct the source and target domain sample sets. The source domain samples consist of optical remote sensing images with pixel-level annotations of flood water accumulation, corresponding spatiotemporal SAR images, and manually annotated masks; the target domain samples consist of optical remote sensing images and SAR images of the area to be interpreted, without the requirement for manual annotation. The source and target domains can come from different regions, seasons, sensors, or imaging conditions, thus forming a typical cross-domain scenario.
[0027] Subsequently, optical and SAR images are preprocessed (including spatial registration, radiometric normalization, SAR logarithmic transformation, tile sampling, and data augmentation) to ensure that the same pixel location corresponds to the same geographic area. For large-format SAR images, fixed windows or overlapping windows can be used to divide them into training sample blocks. Standardization is then performed on both modalities separately. , ; in, Represents an optical image sample. Represents SAR image samples. and These represent the mean and standard deviation of the corresponding modes, respectively. To prevent smoothing terms with a denominator of zero, preprocessing effectively reduces the impact of differences in the radiation range of different sensors and the scale of SAR amplitude on training stability. After standardization, the source domain samples are denoted as... The target domain sample is denoted as .in, Represents the source domain. Indicates the target domain. Indicates optical modes, Indicates SAR mode, This represents the sample label. The standardized output above is directly used as the input to the dual-branch encoder in step 2.
[0028] Step 2, Bi-branch Feature Extraction and Cross-modal Attention Enhancement: Using preprocessed samples from two modalities in the sample set as collaborative input, bi-branch feature extraction is performed. A cross-modal attention mechanism is then used to interactively enhance the features of the two modalities, resulting in multi-scale joint features and segmentation / boundary prediction results. This invention uses optical image branches and SAR image branches as collaborative inputs, and a cross-modal attention mechanism is used to interactively enhance the features of the two modalities, enabling mutual compensation between the two modalities in difficult regions.
[0029] The preprocessed optical image and SAR image are input into the optical encoder and SAR encoder respectively to extract feature maps at multiple scales, as shown in the attached figure. Figure 2 As shown. Let the optical features and SAR features at the i-th scale be respectively... and To highlight waterlogged areas against a complex terrain background, this invention introduces cross-modal attention modules at multiple scales, enabling effective information from one modality to guide and enhance another.
[0030] Taking SAR features guiding optical features as an example, the query, key, and value matrix are first obtained through linear mapping. , and Then calculate cross-modal attention: ; Further enhanced optical features: ; Similarly, SAR features can be derived by reverse-engineering optical features to obtain enhanced SAR features. .in, Indicates the feature channel dimension. This represents the cross-modal enhancement coefficient when SAR guides the optical branch. Through this process, the sensitivity of SAR to the scattering characteristics of water-filled areas can be used to correct the shadow false detection problem in optical images, while the boundary texture information of optical images can be used to suppress false judgments caused by SAR speckle noise.
[0031] The enhancement results from both directions are further fused into joint features: ; in, Indicates feature splicing, This represents convolutional mapping and nonlinear activation operations. It combines features from different scales. After being fed into the shared decoder, it can output pixel-level segmentation probability maps and boundary probability maps. That is, for any sample domain... The network outputs segmentation probabilities. With boundary prediction values The above outputs will serve as direct inputs for teacher-student training in step 3, pseudo-label filtering in step 4, and loss calculation in step 5, respectively.
[0032] Step 3, Construction of the Teacher-Student Domain Adaptive Training Framework: Construct student and teacher models with identical structures, where the student model participates in backpropagation updates, while the teacher model does not directly participate in gradient updates, and its parameters are updated by the student model parameters through an exponential moving average; the teacher model receives weakly augmented samples from the target domain to generate pseudo-labels, and the student model receives strongly augmented samples from the target domain for consistency training.
[0033] During the network training phase, as shown in the attached document... Figure 3 As shown, student models with the same structure are set up. Teacher Model The student model directly participates in backpropagation updates, and its input includes samples from the source domain and samples from the target domain. The teacher model is only used to generate pseudo-labels in the target domain and does not directly participate in gradient updates. The teacher model parameters are updated by the student model parameters using an exponential moving average (EMA). ; in, and Let represent the parameters of the teacher model and the student model respectively at the Kth iteration. This is a smoothing coefficient, and its value range is typically [value range missing]. This update method avoids drastic changes in the teacher model with single gradient fluctuations, resulting in a more stable target domain supervision signal.
[0034] Source domain samples are used for supervised segmentation training, while target domain samples are predicted by the teacher model to obtain target domain probability maps and boundary probability maps. To enhance the stability of the model in cross-region transfer, the teacher model receives target domain samples with weak perturbations, while the student model receives target domain samples with strong perturbations, thus forming a consistent training mechanism.
[0035] More specifically, for the source domain input The student model outputs the segmentation probability. With boundary prediction Weakly enhanced input to the target domain Teacher model output and ; Enhance the input to the target domain Student model output and In this process, the output of the teacher model serves as the direct input for constructing pseudo-labels and boundary confidence in step 4, while the output of the student model serves as the direct object for backpropagation optimization in step 5.
[0036] Step 4: Filter target domain pseudo-labels and perform boundary enhancement.
[0037] For the target domain samples, the teacher model first outputs the pixel-level class probabilities of the fusion branch. And based on this, generate initial pseudo-tags: , ; in, This represents the pseudo-label category of the i-th pixel. This indicates the corresponding confidence level. To prevent erroneous labels from accumulating and propagating in complex regions, this invention further introduces cross-modal consistency and boundary uncertainty measures. The cross-modal consistency measure directly utilizes the auxiliary classification outputs of the optical and SAR branches in step 2, while the boundary uncertainty measure directly utilizes the boundary prediction results of the teacher model in step 3. This ensures a one-to-one correspondence between the pseudo-label selection and the aforementioned feature extraction and teacher prediction processes.
[0038] set up and Let represent the class probabilities of the optical branch and the SAR branch for the i-th pixel, respectively. Then, the cross-modal bifurcation can be expressed as: ; Simultaneously, the boundary uncertainty of the i-th pixel is obtained from the boundary branch output. Boundary uncertainty can be defined as... This ensures that the uncertainty is low when the boundary prediction is close to 0 or 1, and high when the boundary prediction is close to 0.5. Considering confidence level, cross-modal divergence, and boundary uncertainty, the target domain pseudo-label weight is defined as: ; in, Indicates the confidence threshold. Indicates the cross-modal divergence threshold. This is the indicator function. For pixels with high confidence, low divergence, and low boundary uncertainty, For values close to 1, the pseudo-labels will be retained and participate in target domain supervision; for regions with ambiguous boundaries or significant modal divergence, The smaller the value, the less the negative impact of these pseudo-labels on training. This results in... The target domain weighted pseudo-label set is constructed and directly fed into the target domain pseudo-label loss calculation in step 5.
[0039] Meanwhile, the boundary branch utilizes source domain labeled boundaries and target domain high-confidence boundary pseudo-labels to enhance the characterization of areas such as embankment edges, road edges, and building edges. Source domain boundary labels can be generated from the source domain ground truth mask through morphological gradients or edge extraction operators. Taking morphological gradients as an example, the ground truth boundary values... It can be represented as: ; Where r represents the radius of the structuring element. and These represent the dilation and erosion operations, respectively. The origin domain boundary truth value... Together with the high-confidence boundary pseudo-labels of the target domain, they constitute the supervision signal for the boundary enhancement loss in step 5.
[0040] Step 5: Construct the loss function, perform dynamic weight adjustment, backpropagate to update the student model parameters, and obtain the trained student model.
[0041] The total loss of this invention consists of source domain supervised segmentation loss, target domain pseudo-label loss, boundary enhancement loss, and cross-modal alignment loss. Specifically, the source domain supervised segmentation loss corresponds to the student model's output on the source domain samples in step 3; the target domain pseudo-label loss corresponds to the weighted pseudo-labels of the target domain selected in step 4; the boundary enhancement loss corresponds to the boundary supervision signal generated in step 4; and the cross-modal alignment loss directly constrains the bimodal enhancement features obtained in step 2.
[0042] In the source domain, cross-entropy loss is used to constrain the consistency between the fused output and the manually labeled output: ; in, and These represent the height and width of the sample block, respectively. This represents the true label of the i-th pixel in the source domain in category c. This represents the student model's prediction probability for the source domain sample.
[0043] In the target domain, a target domain loss is constructed using weighted pseudo-labels: ; in, Represents the set of pixels in the target domain. This represents the student model's prediction probability for pixels in the target domain.
[0044] For the boundary enhancement branch, binary cross-entropy loss or Dice loss can be used to constrain the boundary graph. Taking binary cross-entropy as an example, it can be written as: ; in, The source domain boundary truth value generated in step 4 Or, a pseudo-label with high credibility at the target domain boundary. This is the corresponding student model boundary prediction value.
[0045] For cross-modal alignment modules, consistency constraints can be set for modal features or branch outputs before and after enhancement. Mean squared error loss at the feature level can be used. ; in, and Let L represent the projection function that maps the two modal features to a unified embedding space, and let L represent the number of scale layers involved in the alignment.
[0046] Therefore, the total loss can be written as: ; in, , and These represent the dynamic weights of the target domain pseudo-label loss, boundary loss, and alignment loss at the Kth iteration. They can be adaptively updated based on the average confidence level of the current batch of pseudo-labels and the average boundary uncertainty. ; ; Cross-modal alignment weights can also be gradually increased with each training epoch: ; in, This indicates the weight warm-up round. Therefore, in the early stages of training, high-confidence regions are prioritized to stabilize network parameters, while in the later stages of training, constraints on boundary recovery and cross-modal alignment are gradually strengthened, thus balancing training stability and target domain adaptability. In the Kth iteration, the teacher model from step 3 first generates the target domain prediction results, then step 4 completes pseudo-label selection and weight calculation, and subsequently, the calculation follows the same steps. And only update the student model parameters. Finally, update the teacher model parameters according to the EMA formula in step 3. This completes the closed-loop training process of "prediction-screening-weighting-optimization-updating teachers".
[0047] Step 6, Model Inference and Result Output: Input the optical image and SAR image of the area to be tested into the trained student model to obtain the corresponding flood probability map, binary mask map and boundary map.
[0048] After training, during the inference phase, only the optical and SAR images of the area to be tested need to be input into the trained stable student model. Alternatively, a corresponding stable teacher weight model can be used to obtain the flood probability map, binary mask map, and boundary map. For large-format optical imagery and SAR imagery pairs, sliding window inference can be used to perform weighted fusion of overlapping areas. ; in, This represents the prediction result of the nth window for pixel x. This represents the blending weight of the window at pixel x. Pixels located in the center of the window are given higher weights, while pixels closer to the window edges are given lower weights to reduce window stitching errors. It comes directly from the water accumulation probability output by the decoder segmentation head after completing the training and convergence in step 5, which is derived from step 2.
[0049] As one implementation method, it can be based on a threshold. Generate a binary water accumulation mask The data is then converted into vectorized water accumulation boundaries, statistical values of water accumulation area, and thematic maps of the disaster situation, which are used for flood emergency mapping, urban waterlogging monitoring, post-disaster assessment, and business system calls.
[0050] Existing floodwater extraction methods mostly focus on single-modality remote sensing imagery or simply perform conventional fusion processing of optical and SAR images, making it difficult to simultaneously achieve all-weather observation capabilities, target recognition capabilities in complex backgrounds, and cross-regional migration capabilities. Furthermore, while existing adaptive methods in the remote sensing domain can utilize source domain annotation information to assist target domain training to some extent, they typically do not address the complementary relationship between optical and SAR modalities in flood scenarios, nor do they fully consider the problem of false label propagation most prevalent in floodwater boundary areas within the target domain. Therefore, existing solutions are prone to widespread false detections, missed detections, and boundary blurring in flood scenarios spanning different regions, seasons, sensors, or imaging conditions.
[0051] This invention introduces a cross-modal attention enhancement mechanism combining optics and SAR into the floodwater extraction task, enabling the two modalities to compensate for each other in challenging regions. Simultaneously, a pseudo-label selection strategy within a teacher-student framework incorporates modal consistency and boundary uncertainty into the credibility evaluation process of the target domain supervision signal, reducing the interference of low-quality pseudo-labels on model training. Furthermore, a dynamic weight control mechanism assigns different importance to target domain pseudo-label supervision, boundary enhancement, and cross-modal alignment at different training stages, thereby improving the model's training stability and final segmentation accuracy in unlabeled target domains. Due to the synergistic effect of these techniques, this invention effectively addresses the key challenges in cross-regional migration extraction of floodwater, namely, the difficulty in utilizing modal complementarity, unstable target domain supervision, and susceptibility to errors in boundary regions.
[0052] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for adaptive segmentation of the optical and SAR co-domains for flood-prone areas, characterized in that, include: Step 1: Construct source and target domain sample sets. The source domain samples consist of optical images with pixel-level annotations of flood water accumulation, SAR images corresponding to them in time and space, and manually annotated masks. The target domain samples consist of optical images and SAR images of the area to be interpreted. The two modal samples in the sample sets are preprocessed to ensure that the same pixel location corresponds to the same geographical area. Step 2: Using the two modal samples after preprocessing in the sample set as collaborative input, perform bi-branch feature extraction, and enhance the features of the two modalities through a cross-modal attention mechanism to obtain multi-scale joint features and segmentation / boundary prediction results; Step 3: Construct student and teacher models with identical structures. Supervised training of the student model is performed using source domain samples, and consistent training of the student model is performed using target domain samples. The teacher model does not directly participate in gradient backpropagation; its parameters are updated by the student model parameters through exponential moving average. Step 4: Filter target domain pseudo-tags and perform boundary enhancement; Step 5: Construct the loss function, perform dynamic weight adjustment, backpropagate to update the student model parameters, and obtain the trained student model; Step 6: Input the optical image and SAR image of the area to be tested into the trained student model to obtain the corresponding flood probability map, binary mask map and boundary map.
2. The optical and SAR co-domain adaptive segmentation method for flood-prone areas as described in claim 1, characterized in that, The preprocessing includes spatial registration, radiometric normalization, SAR logarithmic transformation, slice sampling, and data augmentation. In step 1, normalization is performed on the two modes respectively: ; in, Represents an optical image sample. Represents SAR image samples. and These represent the mean and standard deviation of the corresponding modes, respectively. To prevent smooth terms with a denominator of zero; After standardization, the source domain samples are denoted as The target domain sample is denoted as ,in, Represents the source domain. Indicates the target domain. Indicates optical modes, Indicates SAR mode, Indicates the sample label.
3. The optical and SAR co-domain adaptive segmentation method for flood-prone areas as described in claim 1, characterized in that, In step 2, the enhanced optical features are obtained by guiding the optical features with SAR features: First, obtain the query, key, and value matrix through linear mapping. , and Then calculate cross-modal attention: ; Further enhanced optical features : ; in, Indicates the feature channel dimension. This represents the cross-modal enhancement coefficient when SAR guides optical branching; Similarly, SAR features are guided in reverse from optical features to obtain enhanced SAR features. ; The enhancement results from both directions are further fused into joint features: ; in, Indicates feature splicing, This represents convolution mapping and non-linear activation operations; Combine features across scales The data is fed into a shared decoder, which outputs a pixel-level segmentation probability map and a boundary probability map.
4. The optical and SAR co-domain adaptive segmentation method for flood-prone areas as described in claim 3, characterized in that, In step 3, the teacher model parameters are updated from the student model parameters using an exponential moving average: ; in, and Let represent the parameters of the teacher model and the student model respectively at the Kth iteration. For smoothing coefficients; Source domain samples are used for supervised segmentation training, while target domain samples are predicted by the teacher model to obtain target domain probability maps and boundary probability maps, which are then used to input the source domain. The student model outputs the segmentation probability. With boundary prediction Weakly enhanced input to the target domain Teacher model output and ; Enhance the input to the target domain Student model output and .
5. The optical and SAR cooperative domain adaptive segmentation method for flood-prone waterlogging as described in claim 4, characterized in that, In step 4, for the target domain samples, the teacher model first outputs the pixel-level class probabilities of the fusion branch. And based on this, generate initial pseudo-tags: , ; in, This represents the pseudo-label category of the i-th pixel. Indicates the corresponding confidence level; set up and Let represent the class probabilities of the optical image branch and the SAR image branch for the i-th pixel, respectively. Then, the cross-modal divergence can be expressed as: ; Define the target domain pseudo-label weight as: ; in, Indicates the confidence threshold. Indicates the cross-modal divergence threshold. For indicator functions, Let be the boundary uncertainty of the i-th pixel; By utilizing source domain labeled boundaries and target domain high-confidence boundary pseudo-labels, the characterization of edge regions is enhanced.
6. The optical and SAR cooperative domain adaptive segmentation method for flood-prone waterlogging as described in claim 5, characterized in that, Boundary truth Represented as: ; Where r represents the radius of the structuring element. and These represent expansion and corrosion operations, respectively.
7. The optical and SAR cooperative domain adaptive segmentation method for flood-prone waterlogging as described in claim 5, characterized in that, In step 5, the total loss is composed of the source domain supervision segmentation loss, the target domain pseudo-label loss, the boundary enhancement loss, and the cross-modal alignment loss.
8. The optical and SAR cooperative domain adaptive segmentation method for flood-prone waterlogging as described in claim 7, characterized in that, In step 5, within the source domain, cross-entropy loss is used to constrain the consistency between the fused output and the manually labeled data. ; in, and These represent the height and width of the sample block, respectively. This represents the true label of the i-th pixel in the source domain in category c. This represents the student model's prediction probability for the source domain samples; In the target domain, a target domain loss is constructed using weighted pseudo-labels: ; in, Represents the set of pixels in the target domain. This represents the student model's prediction probability for a pixel in the target domain; Total loss function for: ; in, , and These represent the dynamic weights of the target domain pseudo-label loss, boundary enhancement loss, and cross-modal alignment loss at the Kth iteration, respectively. , To separate the boundary enhancement loss and cross-modal alignment loss; ; in, This refers to the ground truth value of the source domain boundary or the pseudo-label of the highly reliable boundary of the target domain. This corresponds to the student model boundary prediction value; ; in, and Let L represent the projection function that maps the features of the two modalities to a unified embedding space, and let L represent the number of scale layers involved in the alignment.
9. The optical and SAR cooperative domain adaptive segmentation method for flood-prone waterlogging as described in claim 8, characterized in that, In step 5, the target domain pseudo-label loss weights and boundary augmentation loss weights are adaptively updated based on the average confidence level of the current batch of pseudo-labels and the average boundary uncertainty: ; ; Cross-modal alignment weights gradually increase with each training epoch: ; in, This indicates the weighted warm-up round.
10. The optical and SAR co-domain adaptive segmentation method for flood-prone areas as described in claim 1, characterized in that, In step 6, for the large-format optical image and SAR image pair, sliding window inference is used and weighted fusion is performed on the overlapping areas: ; in, This represents the prediction result of the nth window for pixel x. This indicates the fusion weight of the window at pixel x.