A multimodal flood area recognition method based on dual-correlation cross-attention

By using a multimodal flood area recognition method based on dual-correlation cross-attention and a dual-branch encoder and feature fusion module, the problem of insufficient utilization of complementary information of SAR and optical images in the existing technology is solved, and higher-precision flood area recognition is achieved.

CN119693804BActive Publication Date: 2025-09-19YUNNAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510024085.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-09-19
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

When using SAR and optical images to identify flooded areas, existing technologies fail to fully explore and utilize the complementary advantages of the two modalities, resulting in low recognition accuracy and difficulty in achieving high-precision identification of flooded areas under severe weather conditions.

Method used

A multimodal flood area recognition method based on dual-correlation cross-attention is adopted. The multi-scale features of SAR and optical images are extracted respectively through a dual-branch encoder. Combining the dual-correlation cross-attention mechanism and the self-attention mechanism, a feature fusion module is constructed to realize the complementary information fusion of SAR and optical images. Finally, the flood coverage area is generated through a U-Net-like decoder.

Benefits of technology

It significantly improves the accuracy of flood area identification, fully utilizes the complementary information of SAR and optical images, overcomes the limitations of single-modal images, and achieves more accurate flood area identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693804B_ABST
    Figure CN119693804B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal flood area recognition method based on dual-correlation cross-attention, which relates to the technical fields of remote sensing image processing and remote sensing image recognition, and aims to solve the problem of low accuracy in existing flood inundation recognition. The method comprises performing image segmentation processing on SAR images and optical images respectively to obtain a first non-overlapping image block and a second non-overlapping image block; using a dual-branch encoder to encode the first non-overlapping image block and the second non-overlapping image block respectively to obtain multi-layer target SAR features and multi-layer target optical features; determining the dual-correlation features between the target SAR features and the target optical features of each layer based on the dual-correlation cross-attention mechanism; fusing the dual-correlation features, and determining the flood coverage area in the flood area to be identified based on the fused features. The multimodal flood area recognition method based on dual-correlation cross-attention provided by the present invention is used to improve the accuracy of flood inundation recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing and remote sensing image recognition, and in particular to a multimodal flood area recognition method based on dual-correlation cross attention. Background Art

[0002] Floods are one of the most common and destructive natural disasters worldwide and one of the most frequent meteorological disasters in my country, affecting a large number of people. Against the backdrop of climate change, the scale and impact of floods are expanding due to increased extreme precipitation, posing serious challenges to socioeconomic development and the ecological environment. The development of accurate flood forecasting and real-time monitoring technologies is an urgent requirement for disaster prevention and mitigation, with significant social and economic significance. Remote sensing technology, capable of acquiring large-scale surface information in real time, quickly, and accurately, is widely used to identify and locate flood-inundated areas. Using remote sensing technology to obtain real-time and accurate data on flood-inundated areas has become fundamental for relevant departments to fully understand the disaster situation, rapidly and efficiently launch emergency responses, and conduct flood emergency rescue management.

[0003] Synthetic aperture radar (SAR) offers significant advantages in remote sensing-based flood inundation area identification. It operates around the clock and in all weather conditions, unaffected by adverse weather conditions. Furthermore, SAR signals are highly sensitive to water bodies and possess strong penetration, providing information on inundation status. Therefore, they are the preferred choice for flood inundation area identification and monitoring. However, due to limitations in their imaging mechanism, SAR imagery typically has low spatial resolution and is significantly affected by noise, terrain, and shadows, which, to a certain extent, limits its accuracy in flood inundation area identification. In contrast, optical imagery offers rich spectral information and relatively high spatial resolution, offering distinct advantages in ground object identification. However, it is susceptible to cloud cover and struggles to provide information on inundation status. Given the complementary nature of SAR and optical data, some studies have utilized both optical and SAR imagery for flood-covered area identification. However, these methods, often based on traditional machine learning approaches, lack sufficient feature extraction and characterization for both optical and SAR imagery. Furthermore, simply overlaying SAR and optical imagery fails to achieve a deep interaction between the features of the different modalities, making it difficult to fully leverage the complementary advantages of SAR and optical data in flood-covered area identification, resulting in low flood inundation identification accuracy. Summary of the Invention

[0004] The purpose of the present invention is to provide a multimodal flood area recognition method based on dual-correlation cross-attention, which is used to fully explore and utilize the complementary advantages of SAR and optical imaging modalities and significantly improve the accuracy of remote sensing flood recognition.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] The present invention provides a multimodal flood area recognition method based on dual-correlation cross attention, comprising:

[0007] Obtain SAR images and optical images of the flood area to be identified;

[0008] performing image segmentation processing on the SAR image and the optical image respectively to obtain a first non-overlapping image block corresponding to the SAR image and a second non-overlapping image block corresponding to the optical image;

[0009] Using a dual-branch encoder to encode the first non-overlapping image block and the second non-overlapping image block respectively to obtain multi-layer target SAR features and multi-layer target optical features;

[0010] Determine the dual-correlation features between the target SAR features and the target optical features at each layer based on the dual-correlation cross-attention mechanism;

[0011] Fusing the dual correlation features to obtain target fusion features;

[0012] A flood coverage area in the to-be-identified flood area is determined based on the target fusion features.

[0013] Technical effect: Compared with the existing technology, the present invention provides a multimodal flood area recognition method based on dual-correlation cross-attention, which performs flood area recognition based on SAR images and optical images, overcomes the limitations of single-modal images in flood area recognition, and achieves more accurate recognition performance. By segmenting the SAR and optical images, and then using a dual-branch encoder to encode the first non-overlapping image block and the second non-overlapping image block respectively, multi-layer target SAR features and multi-layer target optical features are obtained; the dual-correlation features between the SAR features and optical features of each layer are determined through the dual-correlation cross-attention mechanism, which can not only pay attention to the common information between the modalities, but also pay attention to the unique information between the modalities, and fully retain the complementary information between the multimodal data, thereby improving the discrimination ability and expression effect of the multimodal features, and fusing the dual-correlation features can further enhance the feature expression ability, thereby improving the accuracy of flood area recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0015] Figure 1A flow chart of a multimodal flood area recognition method based on dual-correlation cross-attention provided by the present invention;

[0016] Figure 2 This is a general framework diagram of a multimodal flood area recognition method based on dual-correlation cross-attention provided by the present invention;

[0017] Figure 3 Schematic diagram of the calculation process of the feature fusion module composed of DCCMF and SA provided by the present invention. DETAILED DESCRIPTION

[0018] To facilitate a clear description of the technical solutions of the embodiments of the present invention, the words "first" and "second" are used in the embodiments of the present invention to distinguish between identical or similar items with substantially the same functions and effects. For example, the first threshold and the second threshold are merely used to distinguish between different thresholds and do not limit their order. Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or execution order, and the words "first" and "second" do not necessarily mean different.

[0019] It should be noted that, in the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present invention should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0020] In the present invention, "at least one" means one or more, and "more" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, the combination of a and b, the combination of a and c, the combination of b and c, or the combination of a, b and c, where a, b, c can be single or plural.

[0021] Before introducing the embodiments of the present invention, the following definitions are given for the relevant terms involved in the embodiments of the present invention:

[0022] Swin-Transformer: A visual Transformer model that has achieved remarkable results in multiple fields, including image classification, object detection, and semantic segmentation. Swin-Transformer improves upon the Vision Transformer concept, addressing Vision Transformer's issues with low-resolution feature mapping and quadratic growth in complexity with image size. By introducing a sliding window mechanism and a hierarchical structure, Swin-Transformer achieves efficient image processing and multi-scale feature extraction.

[0023] U-Net-like decoder: The U-Net-like decoder is a network structure designed and improved based on the decoder part of the U-Net network architecture. U-Net is a deep learning architecture designed specifically for image segmentation tasks. Its decoder part plays a vital role in image segmentation, responsible for gradually restoring the spatial resolution and details of the image. The main function of the U-Net decoder is to gradually restore the spatial resolution and details of the image through an upsampling process. This is usually achieved through a series of deconvolution layers and upsampling layers. After each upsampling step, the decoder merges the feature map with the feature map of the corresponding layer of the encoder to restore the lost spatial information. This skip connection mechanism is a key feature of the U-Net architecture, which helps to restore detail information during the upsampling process and allows the network to learn more accurate outputs.

[0024] In recent years, deep learning has developed rapidly, and a series of new deep learning models have emerged, demonstrating powerful feature learning and representation capabilities. These models can automatically extract multi-level, high-level feature representations from complex data, overcoming the limitations of hand-crafted features in traditional machine learning methods. Therefore, based on the significant complementary characteristics between SAR and optical data, and the effective extraction and representation capabilities of deep learning for multimodal features, research on multimodal flood coverage area identification methods based on deep learning-based synergy between optical and SAR imagery is being conducted. This can fully explore and utilize the complementary advantages between the two modalities and significantly improve the effectiveness of remote sensing flood monitoring. Existing remote sensing-based flood inundation area identification methods do not adequately consider the feature interaction between SAR and optical imagery, cannot fully utilize the complementary advantages between SAR and optical imagery, and have low flood inundation identification accuracy.

[0025] To address the above issues, this application proposes a multimodal flood area recognition method based on dual-correlation cross-attention. First, an asymmetric dual-branch encoder is used to extract multi-scale features of SAR images and optical images respectively. Then, a feature fusion module containing a dual-correlation cross-attention mechanism and a self-attention mechanism is constructed. The dual-correlation cross-attention mechanism is used to extract high-correlation features and low-correlation features from optical images to SAR images and from SAR images to optical images respectively, and the complement of the attention weights is used to capture the unique information between the modalities. For high-correlation and low-correlation features, self-attention mechanisms are applied respectively to further enhance the feature expression ability, highlight important features, suppress redundant information, and achieve effective extraction and fusion of optical features and SAR features. Finally, the decoder constructed adopts a U-Net-like structure, and fuses features at different levels through a multi-level step-by-step feature fusion module. At the same time, skip connections are introduced in the decoder to alleviate the gradient vanishing problem in deep networks and promote information flow and gradient propagation. By combining an asymmetric dual-branch encoder, a cross-modal cross-attention feature fusion module, and a U-Net-like decoder, this paper constructs an efficient and robust multimodal flood area recognition framework. This framework fully leverages the complementary information of optical and SAR imagery to achieve accurate flood area extraction. The present invention can be automated using computer software. This is described below with reference to the accompanying figures.

[0026] See also Figure 1 ,This application provides a multimodal flood area recognition method based on ,double-correlation cross attention, such as Figure 1 As shown, the method includes the following steps:

[0027] Step 101: Acquire SAR images and optical images of the flood area to be identified;

[0028] Specifically, obtain SAR images of the same phase after the flood and optical imaging ;in, l is the number of bands of SAR image, h is the height of the SAR image, w is the width of the SAR image; L is the number of bands of the optical image, H For the height of optical image, W is the width of the optical image; , , ; R is the size of the feature.

[0029] Perform image preprocessing operations on the SAR image and the optical image to obtain preprocessed SAR image and preprocessed optical image. The preprocessing may be operations such as spatial registration and normalization.

[0030] Upsample the preprocessed SAR image to obtain the upsampled SAR image , the spatial resolution of the upsampled SAR image is the same as that of the preprocessed optical image.

[0031] Step 102: performing image segmentation processing on the SAR image and the optical image respectively to obtain a first non-overlapping image block corresponding to the SAR image and a second non-overlapping image block corresponding to the optical image;

[0032] The SAR image in this step is the SAR image after upsampling in step 101 The optical image is the pre-processed optical image. Figure 2 As shown, the optical feature extraction network extracts features from the optical image, and the SAR feature extraction network extracts features from the SAR image. Before feature extraction, the image is segmented by the image segmentation module, and then encoded by the encoder. Specifically, step 102 includes: segmenting the SAR image according to a preset block size to obtain a plurality of first non-overlapping image blocks;

[0033] The optical image is segmented according to a preset block size to obtain a plurality of second non-overlapping image blocks.

[0034] Exemplarily, the preset block size may be 4×4.

[0035] Step 103: using a dual-branch encoder to encode the first non-overlapping image block and the second non-overlapping image block respectively to obtain multi-layer target SAR features and multi-layer target optical features;

[0036] like Figure 2 As shown, the dual-branch encoder includes a SAR encoder and an optical encoder. In the optical feature extraction network, the optical encoder is used to extract the multi-scale features of the optical image, and in the SAR feature extraction network, the SAR encoder is used to extract the multi-scale features of the SAR image; wherein, the SAR encoder and the optical encoder both include 4 coding layers, wherein the first coding layer is composed of a linear embedding layer and a Swin-Transformer module, and any layer from the second to the fourth coding layer is composed of an image block merging layer and a Swin-Transformer module. The linear embedding layer is used to embed the segmented image blocks into a high-dimensional vector space, and the embedding dimension size is CIt can be 96. The image block merging layer is used to spatially downsample the feature map, reduce the spatial resolution, and increase the number of channels. The downsampling rate can be 2 each time, and the number of channels can be increased by a multiple of 2. The Swin-Transformer module is used for self-attention calculation and feature transformation. In the optical encoder, the depths of the Swin-Transformer modules used from the first to the fourth layers are 2, 2, 6, and 2 layers, respectively. In particular, the depth of the Swin-Transformer module used in the third encoding layer of the SAR encoder is smaller than the depth of the Swin-Transformer module used in the third encoding layer of the optical encoder. The depths of the Swin-Transformer modules used from the first to the fourth layers of the SAR encoder are 2, 2, 2, and 2 layers, respectively.

[0037] Specifically, a dual-branch encoder is used to encode the first non-overlapping image block and the second non-overlapping image block respectively, and the multi-layer target SAR features and the multi-layer target optical features are obtained, including:

[0038] For a first coding layer of the SAR encoder: mapping the first non-overlapping image blocks into an initial embedding space to obtain a first image block feature map;

[0039] Performing self-attention calculation and feature transformation on the feature map of the first image block to obtain the SAR output feature of the current layer;

[0040] For any coding layer from the second to the fourth coding layer of the SAR encoder: spatially downsample the SAR output features of the previous layer and increase the number of channels to obtain processed SAR output features;

[0041] Perform self-attention calculation and feature transformation on the processed SAR output features to obtain the SAR output features of the current layer;

[0042] Complete the encoding processing of the four-layer SAR coding layer and obtain the four-layer SAR output features;

[0043] Reshape each layer of SAR output features into a four-dimensional tensor form to obtain four layers of target SAR features;

[0044] Four-layer target SAR features Indicates. i Indicates the sequence number of the coding layer, , B is the batch size, C i For the i The channel dimension size of the features generated by the layer encoding layer, H i For the iThe size of the spatial width dimension of the features generated by the layer encoding layer, W i For the i The layer encodes the size of the high-dimensional space of features generated by the layer.

[0045] For a first encoding layer of the optical encoder: mapping the second non-overlapping image block to the initial embedding space to obtain a second image block feature map;

[0046] Performing self-attention calculation and feature transformation on the feature map of the second image block to obtain the optical output feature of the current layer;

[0047] For any coding layer from the second coding layer to the fourth coding layer of the optical encoder: spatially downsampling the optical output features of the previous layer and increasing the number of channels to obtain processed optical output features;

[0048] Perform self-attention calculation and feature transformation on the processed optical output features to obtain the optical output features of the current layer;

[0049] Complete the encoding processing of the four optical coding layers to obtain the four-layer optical output features;

[0050] The optical output features of each layer are reshaped into a four-dimensional tensor form to obtain four layers of target optical features.

[0051] Four-layer target optical features express.

[0052] like Figure 2 As shown in the figure, the target SAR features and target optical features generated in each coding layer will be input into the cross-modal feature fusion module for fusion. The cross-modal feature fusion module is a feature fusion module that includes dual-correlation cross-attention fusion DCCMF and self-attention SA mechanism. The cross-modal feature fusion module first performs cross-attention calculation on the two input modalities to capture the high-correlation features and low-correlation features between the two modalities of SAR image and optical image. Figure 3 The calculation process of the feature fusion module composed of DCCMF and SA is shown, and is described in detail below in conjunction with step 104 and step 105. The calculation process of cross attention is as shown in step 104.

[0053] Step 104: determining dual-correlation features between the target SAR features and the target optical features at each layer based on a dual-correlation cross-attention mechanism;

[0054] The dual correlation features include high correlation features of target SAR features to target optical features, low correlation features of target SAR features to target optical features, high correlation features of target optical features to target SAR features, and low correlation features of target optical features to target SAR features. The high correlation features refer to features with a high degree of correlation, which can be understood as common information between modalities, and the low correlation features refer to features with a low degree of correlation, which can be understood as unique information between modalities.

[0055] The specific steps of step 104 are as follows:

[0056] The dual correlation features between the target SAR features and the target optical features of each layer determined based on the cross attention mechanism include:

[0057] According to formula (1):

[0058] (1)

[0059] Calculate the query vector, key vector and value vector of the target optical feature to the target SAR feature;

[0060] in, For the i Layer target optical characteristics, For the i Layer target SAR characteristics, For the i The size of the features generated by the layer encoding layer, B is the size of the batch, C i For the i The channel dimension size of the features generated by the layer encoding layer, H i For the i The size of the spatial width dimension of the features generated by the layer encoding layer, W i For the i The size of the high-dimensional space of the features generated by the layer encoding layer, is the query vector of the target optical feature to the target SAR feature, h is the number of multi-head attention layers, For the i The influence of optical characteristics of layer targets on SAR characteristics of targets h The query vector of the attention head; is the key vector of the target SAR feature, is the value vector of the target SAR feature, For the i SAR characteristics of layer targets h The key vector of the attention head, For the iSAR characteristics of layer targets h The value vector of the attention head, , and For different change matrices;

[0061] According to formula (2):

[0062] (2)

[0063] Calculate the query vector, key vector and value vector of the target SAR feature to the target optical feature;

[0064] in, is the query vector of the target SAR feature to the target optical feature, For the i The influence of SAR characteristics of layered targets on target optical characteristics h The query vector of the attention head; is the bond vector of the target optical feature, is the value vector of the target optical characteristics, For the i Layer target optical characteristics h The key vector of the attention head, For the i Layer target optical characteristics h The value vector of the attention head;

[0065] Substitute the query vector of the target SAR feature to the target optical feature and the key vector of the target optical feature into formula (3):

[0066] (3)

[0067] Obtain the attention weight of the target SAR feature to the target optical feature;

[0068] Substitute the query vector of the target optical feature to the target SAR feature and the key vector of the target SAR feature into formula (4):

[0069] (4)

[0070] Obtain the attention weight of the target optical features to the target SAR features;

[0071] in, For Flatten and transpose it, For Flatten and transpose it, For Flatten and transpose it, For Flatten and transpose it, is the attention weight of the target SAR feature to the target optical feature, is the attention weight of the target optical feature to the target SAR feature;

[0072] Substitute the attention weight of the target optical feature to the target SAR feature into formula (5):

[0073] (5)

[0074] Obtain high correlation characteristics between target optical characteristics and target SAR characteristics;

[0075] Substitute the attention weight of the target SAR feature to the target optical feature into formula (6):

[0076] (6)

[0077] Obtain high correlation characteristics between target SAR characteristics and target optical characteristics;

[0078] in, For Flatten and transpose it, For Flatten and transpose it, It is a high correlation feature between the target optical characteristics and the target SAR characteristics. It is a feature with high correlation between target SAR characteristics and target optical characteristics;

[0079] Substitute the attention weight of the target optical feature to the target SAR feature into formula (7):

[0080] (7)

[0081] Obtaining low correlation characteristics of target optical characteristics to target SAR characteristics;

[0082] Substitute the attention weight of the target SAR feature to the target optical feature into formula (8):

[0083] (8)

[0084] Obtaining low correlation characteristics between target SAR characteristics and target optical characteristics;

[0085] in, is the low correlation characteristic of the target optical characteristics to the target SAR characteristics, It is a low correlation feature between the target SAR feature and the target optical feature.

[0086] Step 105: Fusing the dual correlation features to obtain target fusion features;

[0087] Specifically, step S1 includes cascading the high correlation features of the target SAR features to the target optical features and the high correlation features of the target optical features to the target SAR features in each layer to obtain the cascaded high correlation features. , , Indicates a cascade operation.

[0088] Step S2: Cascade the low correlation features of the target SAR features to the target optical features and the low correlation features of the target optical features to the target SAR features in each layer to obtain the cascaded low correlation features , .

[0089] Step S3: using a self-attention mechanism to perform weighted processing on the concatenated high-correlation features and the concatenated low-correlation features to obtain weighted high-correlation features and weighted low-correlation features;

[0090] According to formula (9):

[0091] (9)

[0092] Calculate the weighted high correlation features; is the weighted high correlation feature. It is the self-attention mechanism;

[0093] According to formula (10):

[0094] (10)

[0095] The weighted low-correlation features are calculated. is the weighted low correlation feature.

[0096] The calculation process of the self-attention mechanism SA is:

[0097] First, given the input feature map X, since the steps for weighting the high-correlation features after cascading are the same as those for weighting the low-correlation features after cascading, X is used here to represent both. The query vector, key vector, and value vector of the feature map X are calculated according to formula (11):

[0098] (11)

[0099] in, Q is the query vector of feature map X, K is the key vector of the feature graph X, Vis the value vector of the feature map X; it can be understood that the input feature map X is the feature map corresponding to the high correlation features after cascading and the low correlation features after cascading.

[0100] Then according to formula (12):

[0101] (12)

[0102] Calculating attention weights A ;

[0103] in, For After flattening and transposing, For After flattening and transposing, d is the scaling factor.

[0104] Then according to formula (13):

[0105] (13)

[0106] Calculating weighted features .

[0107] in, For Flatten and transpose.

[0108] Finally, according to formula (14), the residual connection calculation is performed with the original feature map to obtain the weighted features, as shown in formula (14):

[0109] (14)

[0110] in, is the original feature map, is the weighted feature of the output, It is understood that the original feature map is the feature map corresponding to the high correlation feature after cascading and the low correlation feature after cascading.

[0111] Step S4: fusing the weighted high-correlation features and the weighted low-correlation features to obtain fused features;

[0112] Step S5: Perform convolution operation on the fused features to obtain a convolution feature map; reduce the dimension of the target fused features to the target output channel number by convolution. .

[0113] Step S6: performing a batch normalization operation on the convolution feature map to obtain a normalized feature map;

[0114] Step S7: Perform nonlinear mapping on the normalized feature map through the ReLU activation function to obtain the target fusion feature.

[0115] Specifically, according to formula (15), batch normalization and ReLU activation function are performed on the convolution feature map for nonlinear mapping, as shown in formula (15):

[0116] (15)

[0117] in, For a 1×1 convolutional layer, For the i Features output by the cross-modal feature fusion module.

[0118] Step 106: Determine the flood coverage area in the to-be-identified flood area based on the target fusion features.

[0119] The specific steps include:

[0120] Step S10: using a multi-layer decoder to gradually upsample and decode the target fusion features to obtain a flood coverage area feature map of the flood area to be identified;

[0121] The decoder can be a U-Net-like decoder. The main task of the decoder is to gradually upsample and decode the target fusion features obtained by the feature fusion module to generate the final segmentation output. Figure 2 As shown in the figure, the first three layers of the decoder are composed of an image expansion layer and a Swin-Transformer module. The image expansion layer is used to spatially upsample the feature map, improve the spatial resolution, and reduce the number of channels of the feature map. The upsampling rate is 2 each time, and the number of channels is reduced to 1 / 2 of the input each time. The Swin-Transformer module is mainly used for feature decoding. j The decoding layer receives two inputs. j =1, receives the features output by the feature fusion module of the corresponding level And the skip connection features output from the previous fusion module ,when When , the decoding layer receives the output features of the decoding layer above it and the features output from the feature fusion module of the previous layer The decoding layer first performs skip connection features Perform average pooling and adjust the spatial size to obtain , and then and Cascading in the channel dimension, we get , then adjust the features The dimension in the channel dimension and the linear projection layer Linear is used to reduce the number of channels of the feature map to obtain , and finally the feature map Enter to j The Swin-Transformer module in the layer decoding layer performs feature decoding and extraction to obtain the decoded features .

[0122] like Figure 2 As shown in the figure, the target fusion features output by the feature fusion module are, from top to bottom: the first layer fusion features, the second layer fusion features, the third layer fusion features and the fourth layer fusion features; the decoder is, from bottom to top, the first layer decoding layer, the second layer decoding layer, the third layer decoding layer and the fourth layer decoding layer.

[0123] The specific step S10 includes the following steps:

[0124] Step S101: For the first decoding layer: use average pooling AvgPool The third layer fusion features The spatial dimension is halved, and the number of channels is doubled using 1×1 convolution to obtain the first increased dimensional feature ;

[0125] The first dimension feature is increased Fusion features with the fourth layer Cascade to obtain the first cascade feature ; The cascade splicing is performed in the channel dimension.

[0126] The first cascade feature Perform upsampling and reduce the number of channels to obtain the first reduced dimensionality feature ; Upsampling is performed through the image expansion layer, and the number of channels is reduced through the linear projection layer Linear. The upsampling multiple is 2, and the number of channels after reduction is 1 / 2 of the original.

[0127] The first reduced dimensionality feature is processed using the Swin-Transformer module in the first decoding layer Perform feature decoding and extraction to obtain the output features of the first decoding layer ;

[0128] Step S102: For the second decoding layer, use average pooling to fusion features of the second layer The spatial dimension of the corresponding feature map is halved, and the number of channels of the feature map corresponding to the second layer of fusion features is increased to twice the original number, and the second increased dimension feature is obtained. ;

[0129] The second dimension feature is increased The output features of the first decoding layer Cascade on the channel dimension to obtain the second cascade feature ;

[0130] Reduce the second cascade features The number of channels corresponding to the feature map is used to obtain the second reduced dimensional feature Specifically, the number of channels of the feature map corresponding to the second cascade feature is reduced to half of the current one through the linear projection layer;

[0131] The second reduced dimensional feature is input into the Swin-Transformer module in the second decoding layer for feature decoding and extraction, and the output feature of the second decoding layer is obtained. .

[0132] Step S103: For the third decoding layer, use average pooling to fusion features of the first layer The spatial dimension of the feature map is halved, and the number of channels of the feature map corresponding to the first layer of fusion features is increased to 2 times of the original, and the third increased dimension feature map is obtained. ;

[0133] Add the third dimension feature The output features of the second decoding layer Cascade on the channel dimension to obtain the third cascade feature ;

[0134] Reduce third cascade features The number of channels corresponding to the feature map is used to obtain the third reduced dimensional feature Specifically, the number of channels of the feature map corresponding to the third cascade feature is reduced to half of the current one through the linear projection layer;

[0135] The third reduced dimensional feature is input into the Swin-Transformer module in the third decoding layer for feature decoding and extraction, and the output feature of the third decoding layer is obtained. .

[0136] The output features of the third decoding layer are expanded by increasing the number of channels and spatial dimensions of the feature map to generate a flood coverage feature map for the flood area to be identified. Specifically, the linear mapping layer first expands the number of channels to 16 times the original number. Then, a rearrangement operation expands the spatial dimension to 4 times the original number, while the number of channels is correspondingly reduced to the original size. The resulting output flood coverage feature map for the flood area to be identified has the same spatial resolution as the input data. It is understood that the input data is the upsampled SAR image and the preprocessed optical image.

[0137] Step S20: Segment the flood coverage area feature map of the flood area to be identified to obtain the flood coverage area in the flood area to be identified.

[0138] Specifically, a 1×1 convolutional layer is used to reduce the number of channels of the flood coverage area feature map of the flood area to be identified to the number of categories, which is 1 in this task, namely the flood coverage area. The feature map with the number of channels reduced to 1 is mapped to the probability space through the Sigmoid nonlinear activation function to obtain the flood coverage area in the flood area to be identified.

[0139] The beneficial effects of the present invention are:

[0140] 1. The present invention uses SAR images and optical images of the same flood-affected area at the same time phase to carry out research on a multimodal flood-inundated area identification method. Compared with the existing mainstream single-modality: flood coverage area identification method using only SAR or optical images, the present invention can fully tap the complementary advantages of SAR images and optical images, overcome the limitations of single-modality images in flood-inundated area identification, and achieve more accurate identification performance.

[0141] 2. The method of the present invention is a coding and decoding structure, and both the encoder and the decoder use Swin-Transformer as the basic network. In view of the difference in information richness contained in optical images and SAR images, an asymmetric dual-branch encoder is designed. Among them, considering that optical images usually contain richer texture and detail information than SAR images, a deeper network is required to extract high-level features. Therefore, 6 Swin-Transformer modules are used in the third encoding layer of the optical encoder, while the SAR encoder only uses 2 Swin-Transformer modules at all levels. The deeper encoder structure can extract richer texture and semantic features of the optical image, while the SAR encoder uses a relatively shallow encoder structure to capture the geometric structure and scattering characteristics of the data structure, which ensures effective feature extraction while reducing the number of parameters of the SAR encoder, improving the generalization ability of the model and preventing overfitting.

[0142] 3. The present invention proposes a new feature fusion module, which aims to effectively fuse the features of optical images and SAR images to solve the problem of extracting complementary information in multimodal data. In particular, the feature fusion module includes a dual correlation cross attention fusion (DCCAF) module and a self-attention (SA) module. DCCAF is used to extract complementary information between optical images and SAR images. Unlike the current cross-attention method, which can only focus on the highly correlated features between multimodal data, but ignore the unique features of each modality. The present invention constructs a new dual correlation cross attention mechanism (DCCAF), which includes two branches. One branch uses the cross-self-attention mechanism to extract features with high correlation between modalities, while the other branch uses the complement cross-attention mechanism to extract unique information with low correlation between modalities. In this way, the model can not only pay attention to the common information between modalities, but also the unique information between modalities, fully retaining the complementary information between multimodal data, thereby improving the ability to distinguish and express multimodal features.

[0143] 4. The present invention uses the self-attention mechanism to fuse the high-correlation and low-correlation features between modalities obtained by the dual-correlation cross-attention fusion module. The self-attention mechanism is applied to the high-correlation and low-correlation features respectively to further enhance the feature expression capability. The introduction of the self-attention mechanism enables the model to automatically focus on important features during the fusion process, suppress the redundant and unimportant information between the high-correlation and low-correlation features of multimodal data, and help improve the model's discrimination ability and generalization performance. On the other hand, by performing self-attention enhancement on high-correlation and low-correlation features respectively, the model can better capture the complex dependencies in spatial and channel dimensions, improve the feature expression capability, and thus improve the accuracy of flood coverage area identification.

[0144] 5. The decoder developed in this paper adopts a U-Net-like architecture, fusing features at different levels through a multi-level, step-by-step feature fusion module. Shallow, high-resolution features provide fine spatial details, while deeper, low-resolution features contain global semantic information. This cross-scale feature fusion enables the decoder to simultaneously learn the precise boundaries and overall structural features of flooded areas, resulting in high-quality recognition results.

[0145] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit may implement several functions listed in the claims. The fact that certain measures are recorded in mutually different dependent claims does not mean that these measures cannot be combined to produce good results.

[0146] Although the present invention has been described with reference to specific features and embodiments thereof, it will be apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the invention. Accordingly, this specification and drawings are merely illustrative of the invention as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the invention. It will be apparent that various modifications and variations may be made to the present invention by those skilled in the art without departing from the spirit and scope of the invention. Thus, the present invention is intended to include such modifications and variations as fall within the scope of the claims of the present invention and their equivalents.

Claims

1. A multimodal flood area recognition method based on dual-correlation cross attention, characterized by: include: Obtain SAR images and optical images of the flood area to be identified; performing image segmentation processing on the SAR image and the optical image respectively to obtain a first non-overlapping image block corresponding to the SAR image and a second non-overlapping image block corresponding to the optical image; Using a dual-branch encoder to encode the first non-overlapping image block and the second non-overlapping image block respectively to obtain multi-layer target SAR features and multi-layer target optical features; Determine the dual-correlation features between the target SAR features and the target optical features at each layer based on the dual-correlation cross-attention mechanism; The dual correlation features include a high correlation feature of the target SAR feature to the target optical feature, a low correlation feature of the target SAR feature to the target optical feature, a high correlation feature of the target optical feature to the target SAR feature, and a low correlation feature of the target optical feature to the target SAR feature; Fusing the dual correlation features to obtain target fusion features; determining a flood coverage area in the to-be-identified flood area based on the target fusion features; The fusing of the dual correlation features to obtain a target fusion feature comprises: Cascading the high correlation features of the target SAR features to the target optical features and the high correlation features of the target optical features to the target SAR features in each layer to obtain the cascaded high correlation features; Cascading the low correlation features of the target SAR features to the target optical features and the low correlation features of the target optical features to the target SAR features in each layer to obtain the cascaded low correlation features; A self-attention mechanism is used to perform weighted processing on the concatenated high-correlation features and the concatenated low-correlation features to obtain weighted high-correlation features and weighted low-correlation features; Fusing the weighted high-correlation features and the weighted low-correlation features to obtain fused features; Performing a convolution operation on the fused features to obtain a convolution feature map; Performing a batch normalization operation on the convolutional feature map to obtain a normalized feature map; The normalized feature map is nonlinearly mapped through the ReLU activation function to obtain the target fusion feature.

2. The multimodal flood area recognition method based on dual-correlation cross attention according to claim 1 is characterized in that: The performing image segmentation processing on the SAR image and the optical image respectively to obtain a first non-overlapping image block corresponding to the SAR image and a second non-overlapping image block corresponding to the optical image comprises: Segmenting the SAR image according to a preset block size to obtain a plurality of first non-overlapping image blocks; The optical image is segmented according to a preset block size to obtain a plurality of second non-overlapping image blocks.

3. The multimodal flood area recognition method based on dual-correlation cross attention according to claim 1 is characterized in that: The determining of the flood coverage area in the to-be-identified flood area based on the target fusion feature includes: A multi-layer decoder is used to gradually upsample and decode the target fusion features to obtain a flood coverage area feature map of the flood area to be identified; The flood coverage area feature map of the flood area to be identified is segmented to obtain the flood coverage area in the flood area to be identified.

4. The multimodal flood area recognition method based on dual-correlation cross attention according to claim 3 is characterized in that: The target fusion features include first-layer fusion features, second-layer fusion features, third-layer fusion features, and fourth-layer fusion features; The multi-layer decoder is used to gradually upsample and decode the target fusion features to obtain a flood coverage area feature map of the flood area to be identified, including: For the first decoding layer: the spatial dimension of the third layer fusion feature is halved and the number of channels is increased to obtain a first increased dimension feature; Cascading the first dimension-increased feature and the fourth-layer fusion feature to obtain a first cascade feature; Performing upsampling and channel reduction processing on the first cascaded features to obtain first reduced-dimensionality features; Performing feature decoding and extraction processing on the first reduced-dimensional features to obtain output features of a first decoding layer; For the second decoding layer: the spatial dimension of the second layer fusion feature is halved and the number of channels is increased to obtain the second increased dimension feature; Cascading the second dimension-increased feature with the output feature of the first decoding layer in the channel dimension to obtain a second cascade feature; Adjusting the dimension of the second concatenated feature in the channel dimension and reducing the number of channels to obtain a second reduced-dimensionality feature; Decoding and extracting the second reduced-dimensional features to obtain output features of a second decoding layer; For the third decoding layer: the spatial dimension of the first layer fusion feature is halved and the number of channels is increased to obtain the third increased dimensional feature; Cascading the third dimension-increased feature with the output feature of the second decoding layer in the channel dimension to obtain a third cascade feature; Adjusting the dimension of the third concatenated feature in the channel dimension and reducing the number of channels to obtain a third reduced-dimensional feature; Performing feature decoding and extraction on the third reduced-dimensionality features to obtain output features of a third decoding layer; The number of channels and the spatial dimension of the output features of the third decoding layer are expanded to obtain a flood coverage area feature map of the flood area to be identified.

5. The multimodal flood area recognition method based on dual-correlation cross attention according to claim 3 is characterized in that: The flood coverage area characteristic map of the flood area to be identified is segmented to obtain the flood coverage area in the flood area to be identified, including: The number of channels of the flood coverage area feature map of the flood area to be identified is reduced to 1, and the feature map with the channel number reduced to 1 is mapped to the probability space through a nonlinear activation function to obtain the flood coverage area in the flood area to be identified.

6. The multimodal flood area recognition method based on dual-correlation cross attention according to claim 1 is characterized in that: The dual-branch encoder includes a SAR encoder and an optical encoder, and each branch encoder includes four coding layers; the dual-branch encoder is used to encode the first non-overlapping image block and the second non-overlapping image block respectively to obtain multi-layer target SAR features and multi-layer target optical features, including: For a first coding layer of the SAR encoder: mapping the first non-overlapping image blocks into an initial embedding space to obtain a first image block feature map; Performing self-attention calculation and feature transformation on the feature map of the first image block to obtain the SAR output feature of the current layer; For any coding layer from the second to the fourth coding layer of the SAR encoder: spatially downsample the SAR output features of the previous layer and increase the number of channels to obtain processed SAR output features; Perform self-attention calculation and feature transformation on the processed SAR output features to obtain the SAR output features of the current layer; Complete the encoding processing of the four-layer SAR coding layer and obtain the four-layer SAR output features; Reshape each layer of SAR output features into a four-dimensional tensor form to obtain four layers of target SAR features; For a first encoding layer of the optical encoder: mapping the second non-overlapping image block into the initial embedding space to obtain a second image block feature map; Performing self-attention calculation and feature transformation on the feature map of the second image block to obtain the optical output feature of the current layer; For any coding layer from the second coding layer to the fourth coding layer of the optical encoder: spatially downsampling the optical output features of the previous layer and increasing the number of channels to obtain processed optical output features; Perform self-attention calculation and feature transformation on the processed optical output features to obtain the optical output features of the current layer; Complete the encoding processing of the four optical coding layers to obtain the four-layer optical output features; The optical output features of each layer are reshaped into a four-dimensional tensor form to obtain four layers of target optical features.

7. The multimodal flood area recognition method based on dual-correlation cross attention according to claim 1 is characterized in that: The dual correlation features between the target SAR features and the target optical features of each layer determined based on the dual correlation cross attention mechanism include: According to the formula: Calculate the query vector, key vector and value vector of the target optical feature to the target SAR feature; in, is the target optical feature of the i-th layer, is the SAR feature of the target at layer i, is the size of the feature generated by the i-th encoding layer, B is the batch size, C i The channel dimension size of the feature generated by the i-th encoding layer, H i The size of the spatial width dimension of the feature generated by the i-th encoding layer, W i The size of the high-dimensional space for generating features for the i-th encoding layer, is the query vector of the target optical feature to the target SAR feature, h is the number of multi-head attention layers, is the query vector of h attention heads of target optical features on target SAR features at the i-th layer; is the key vector of the target SAR feature, is the value vector of the target SAR feature, is the key vector of the h attention heads of the SAR features of the i-th target, is the value vector of h attention heads of the SAR feature of the target in the i-th layer, W Q , W K and W V For different change matrices; According to the formula: Calculate the query vector, key vector and value vector of the target SAR feature to the target optical feature; in, is the query vector of the target SAR feature to the target optical feature, is the query vector of h attention heads of target SAR features on target optical features at the i-th layer; is the bond vector of the target optical feature, is the value vector of the target optical characteristics, is the key vector of the h attention heads of the target optical features at the i-th layer, is the value vector of h attention heads of the target optical features in the i-th layer; Substitute the query vector of the target SAR feature to the target optical feature and the key vector of the target optical feature into: Obtain the attention weight of the target SAR feature to the target optical feature; Substitute the query vector of the target optical feature to the target SAR feature and the key vector of the target SAR feature into: Obtain the attention weight of the target optical features to the target SAR features; in, For Flatten and transpose it, For Flatten and transpose it, For Flatten and transpose it, For A is obtained by flattening and transposing s→o is the attention weight of the target SAR feature to the target optical feature, A o→s is the attention weight of the target optical feature to the target SAR feature; Substitute the attention weight of the target optical feature to the target SAR feature into: Obtain high correlation characteristics between target optical characteristics and target SAR characteristics; Substitute the attention weight of the target SAR feature to the target optical feature into: Obtain high correlation characteristics between target SAR characteristics and target optical characteristics; in, For Flatten and transpose it, For Flatten and transpose it, It is a high correlation feature between the target optical characteristics and the target SAR characteristics. It is a feature with high correlation between target SAR characteristics and target optical characteristics; Substitute the attention weight of the target optical feature to the target SAR feature into: Obtaining low correlation characteristics of target optical characteristics to target SAR characteristics; Substitute the attention weight of the target SAR feature to the target optical feature into: Obtaining low correlation characteristics between target SAR characteristics and target optical characteristics; in, is the low correlation characteristic of the target optical characteristics to the target SAR characteristics, It is a low correlation feature between the target SAR feature and the target optical feature.

8. The multimodal flood area recognition method based on dual-correlation cross attention according to claim 6 is characterized in that: The depth of the Swin-Transformer module used in the third coding layer of the SAR encoder is smaller than the depth of the Swin-Transformer module used in the third coding layer of the optical encoder.

9. The multimodal flood area recognition method based on dual-correlation cross attention according to claim 1 is characterized in that: The obtaining of SAR images and optical images in the flood area to be identified includes: Obtain SAR images and optical images at the same time after the flood occurs; The SAR image is upsampled to obtain an upsampled SAR image, and the spatial resolution of the upsampled SAR image is the same as that of the optical image.