Multi-modal fusion flood detection method and system based on deep learning

The method improves flood detection by integrating spatial and spectral features with attention mechanisms and DeepLabv3+ segmentation, enhanced by digital elevation models, achieving high accuracy and reduced misclassification in complex environments.

CN120318674APending Publication Date: 2025-07-15WUHAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510366228.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing flood detection methods lack multimodal heterogeneous feature coordination mechanism in the feature engineering dimension, lack of completeness of feature characterization, and limited detection accuracy under complex surface conditions, and lack of embedded geospatial prior knowledge in the post-processing process, resulting in a high misjudgment rate.

Method used

The multimodal fusion method based on deep learning is adopted, and the heterogeneous feature coordination mechanism of spatial features and spectral features is extracted, spatial features are extracted in combination with attention mechanism and ResNet50 network, spectral features are extracted dynamically weighted AWEI index, and flood detection results are optimized through DeepLabv3+ architecture and DEM data, and the Transformer architecture is introduced to capture long-range dependencies.

Benefits of technology

It significantly improves the accuracy and reliability of flood detection, reduces the error detection rate and missed detection rate, improves the accuracy and consistency of detection results, and is suitable for flood monitoring under complex surface conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318674A_ABST
    Figure CN120318674A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal fusion flood detection method and system based on deep learning, and the method comprises the steps: S1, obtaining remote sensing image data of a flood period of a research region, and carrying out the preprocessing of the obtained original data; s2, spatial features P1 and spectral features P2 are extracted from the preprocessed data; s3, fusing the spatial feature P1 and the spectral feature P2; and S4, performing multi-level optimization on the fused space-spectrum fusion feature map to obtain a final flood detection result. According to the method, detection precision breakthrough is realized through a heterogeneous feature cooperation mechanism of the spatial feature P1 and the spectral feature P2, and the problems of insufficient single feature expression ability, limited detection precision and the like in the existing flood detection technology are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image processing, and particularly relates to a multi-modal fusion flood detection method and system based on deep learning. Background Art

[0002] Flood disasters are one of the most serious natural disasters globally, causing huge casualties and property losses every year. Timely and accurate flood information is of great significance for disaster prevention, mitigation, emergency rescue, and post-disaster reconstruction. With the development of remote sensing technology, flood detection methods based on satellite remote sensing images have been widely applied, but there are still some problems to be solved in the existing technology.

[0003] Currently, flood detection methods based on remote sensing images can be mainly divided into the following categories:

[0004] The first category is the detection method based on water indices (such as NDWI, MNDWI, AWEI, etc.). The Normalized Difference Water Index (NDWI) proposed by McFeeters in 1996 and the Modified Normalized Difference Water Index (MNDWI) improved by Xu in 2006 have been widely used in water body detection. These methods utilize the reflection characteristic differences between the near-infrared band and the green band to enhance water body information, and have the characteristics of simple calculation and clear physical meaning. However, their feature expression ability is limited, and it is difficult to handle flood detection tasks under complex surface conditions such as urban built-up areas, and the detection effect of a single index is often not ideal.

[0005] The second category is the detection method based on machine learning. For example, Isikdogan et al. (2017) proposed using Support Vector Machine (SVM) for water body classification, and Huang et al. (2018) adopted the Random Forest method for flood area identification. This type of method realizes water body detection by extracting various features (such as spectral features, texture features, etc.) and establishing a classification model, but feature engineering often relies on expert experience and cannot be widely applied.

[0006] The third category is the detection method based on deep learning. For example, Li et al. (2019) proposed using the U-Net network structure for water body semantic segmentation, and Zhang et al. (2020) adopted the DeepLab series models for flood area identification. Deep learning methods can automatically learn complex feature representations, but existing research mainly focuses on the improvement of network structures and rarely considers the physical characteristics of water bodies. At the same time, this type of method often ignores the spectral characteristics of water bodies and has insufficient generalization ability in the case of limited training data.

[0007] Meanwhile, there are dual limitations in the technical implementation of current flood detection methods: First, in the dimension of feature engineering, there is a lack of a collaborative mechanism for multi-modal heterogeneous features. Existing algorithms mostly use single spectral features or static texture descriptors, and fail to construct a dynamic coupling model for the multi-dimensional feature space of spectrum-texture-topography, resulting in insufficient completeness of feature representation. Second, at the level of the feature fusion architecture, there is a lack of an adaptive weight allocation mechanism. Traditional concat operations or simple linear superposition are difficult to achieve cross-layer interaction of multi-scale features, making the feature complementary effect not reach the optimal. In addition, there is a problem of insufficient embedding of geospatial prior knowledge in the post-processing process. Conventional threshold segmentation methods ignore key constraint conditions such as terrain undulation and watershed morphological features, and traditional morphological operations lack an adaptive parameter adjustment mechanism for the topological structure of floodplains, resulting in pixel-level misjudgments in the complex object boundary areas of the detection results.

[0008] Therefore, how to effectively combine the physical characteristics of water indices and the feature extraction ability of deep learning to improve the accuracy of flood detection is an issue that needs to be further explored in current research. Summary of the Invention

[0009] The purpose of the present invention is to provide a multi-modal fusion flood detection method based on deep learning for the deficiencies of the existing technology, so as to solve problems such as insufficient single feature expression ability and limited detection accuracy in existing flood detection technologies.

[0010] To solve the above technical problems, the present invention adopts the following technical solutions:

[0011] A multi-modal fusion flood detection method based on deep learning, comprising:

[0012] S1: Obtain remote sensing image data during the flood period of the research area, and preprocess the obtained original data;

[0013] S2: Extract spatial feature P1 and spectral feature P2 respectively from the preprocessed data;

[0014] S3: Fuse spatial feature P1 and spectral feature P2 to obtain a spatial-spectral fusion feature map;

[0015] S4: Perform multi-level optimization on the obtained spatial-spectral fusion feature map to obtain the final flood detection result.

[0016] Further, the method for extracting spatial features in step 2 includes:

[0017] Perform RGB synthesis on the image data processed in step S1 to obtain an RGB image;

[0018] Based on the attention mechanism, for the obtained RGB image, channel attention and spatial attention are simultaneously focused respectively to obtain a focused composite image;

[0019] Finally, ResNet50 is used to extract features from the focused composite image to obtain spatial feature P1.

[0020] Furthermore, the method for obtaining a focused composite image using the attention mechanism includes:

[0021] In the channel attention mechanism, the input data is compressed. After that, first, channel dimension compression and global average pooling are performed to generate a feature vector; then, two fully connected operations are carried out to obtain a channel weight matrix;

[0022] In the spatial attention mechanism, for the input data, max pooling and average pooling are respectively performed, and the results of the two poolings are concatenated in the channel dimension to form a tensor. Then, spatial convolution is performed on the tensor through a convolutional kernel to obtain a spatial weight matrix;

[0023] Multiple groups of parallel convolutional kernels are constructed, and the channel weight matrix and the spatial weight matrix are input into the multiple groups of parallel convolutional kernels for convolution to output the focused composite image.

[0024] Furthermore, the method for extracting spectral feature P2 in step S2 includes:

[0025] First, based on the image data preprocessed in step 1, the reflectances of different bands are obtained, and the Automated Water Extraction Index (AWEI) is constructed therefrom:

[0026] AWEI = B2 + 2.5×B3 - 1.5×(B8 + B 11 ) - 0.25×B 12

[0027] In the formula, B2 and B3 respectively correspond to the reflectances of the blue and green bands, B8 is the reflectance of the visible and near-infrared bands, and B 11 , B 12 are the reflectances of the infrared SWIR bands;

[0028] Secondly, according to the constructed AWEI index above, Min-Max normalization processing is performed to obtain the spectral feature P2. The formula is:

[0029]

[0030] In the formula, AWEI min and AWEI max are respectively the minimum and maximum values of the Automated Water Extraction Index (AWEI), and P2 is the extracted spectral feature map.

[0031] Further, the specific implementation method of step S3 includes:

[0032] S3.1, perform two-stage transposed convolution on the spatial feature map obtained in step S2 to achieve n-fold upsampling and restore its resolution, thereby obtaining the spatial feature map P1' with the target resolution;

[0033] S3.2, perform bilinear interpolation on the spectral feature map obtained from step S2 to obtain the corrected spectral feature map P2';

[0034] S3.3, calculate the fusion weight for the spatial feature map P1' obtained in step S3.1 and the spectral feature map P2' obtained in step S3.2. The specific formula is as follows:

[0035] α = Sigmoid(Conv(Concat[P1′, P2′]))

[0036] In the formula, Concat is the concatenation function, Conv is the convolution function, Sigmoid is the normalization function, and α is the obtained fusion weight coefficient;

[0037] Then, according to the above fusion weight coefficient, perform feature weighted fusion on the spatial feature map P1' and the spectral feature map P2' to obtain the spatial-spectral fusion feature map:

[0038] Y fusrd = α·P1′ + (1 - α)·P2′

[0039] In the formula, α is the fusion weight coefficient, and Y fused is the obtained spatial-spectral fusion feature map.

[0040] Further, the method for obtaining the corrected spectral feature map P2' includes

[0041] First, perform resolution reconstruction on the spectral feature map to obtain the spectral feature map with the target high resolution;

[0042] After that, for the reconstructed spectral feature map, establish the coordinate correspondence between the target high-resolution grid and the low-resolution feature map before restoration;

[0043] Next, for the target coordinate position of each pixel, select the 4 nearest original pixel points around it to form the calculation basis element for bilinear interpolation. According to the horizontal and vertical distance differences (dx, dy) between the target point and each neighboring point, calculate the bilinear weight coefficients, and multiply the weights in the two directions to obtain the final contribution weights of the four neighboring points, ensuring that the sum of the weights is 1. The specific formula is as follows:

[0044]

[0045] w11 =(1 - dx)(1 - dy)

[0046] w 21 =dx(1 - dy)

[0047] w 12 =(1 - dx)dy

[0048] w 22 =dx·dy

[0049] ∑w = w 11 +w 21 +w 12 +w 22 =1

[0050] Wherein, (x, y) are the floating-point coordinates of the target point, (x1, y1), (x2, y2) are the integer coordinates of the nearest neighbor upper left and lower right points, dx and dy are the normalized distances in the X and Y axis directions respectively, w 11 , w 21 , w 12 , w 22 are the weight coefficients of the four adjacent pixels;

[0051] Finally, for each band of the image, calculate the color of each pixel in the new feature map. Among them, the RGB value of each pixel is the weighted average of the RGB values of the 4 adjacent pixels according to the weight coefficients obtained above. After the above operations, the images generated by each band are finally generated, and then the final spectral feature map is obtained.

[0052] Furthermore, the specific implementation method of step S4 includes:

[0053] S4.1, construct a DeepLabv3+ architecture model, segment a part from the set of fused feature maps obtained in S3 for training the DeepLabv3+ architecture model, and then use the trained DeepLabv3+ architecture model to perform semantic segmentation processing on the remaining data to obtain a flood detection result map after semantic segmentation;

[0054] S4.2, introduce digital elevation model DEM data as a priori constraint to adjust the boundary of the flood detection result map after the above semantic segmentation;

[0055] S4.3, introduce the Transformer architecture to capture the long-range dependencies of the flood detection result map after boundary adjustment,

[0056] S4.4, perform post-processing optimization on the feature map after collecting the long-range dependencies above to obtain the final flood detection result.

[0057] Furthermore, the method for boundary adjustment includes:

[0058] a. Align the elevation model DEM data space with the flood detection result map after semantic segmentation to obtain an elevation matrix H of the same size as the segmentation result; for each segmentation region R, calculate its average elevation value, and the formula is as follows:

[0059]

[0060] where |R| represents the number of pixel points in region R, H ij represents the elevation value at position (i, j), and h R is the average elevation value of the segmentation region R;

[0061] b. For the flood detection result map after semantic segmentation and the obtained elevation value data, perform boundary adjustment, which includes:

[0062] ① For each pair of adjacent regions R1 and R2, calculate their average elevation difference Δh = |h R1 - h R2 |;

[0063] ② For each group of elevation differences calculated above, if it is found that Δh > τ, perform a morphological erosion operation to obtain an updated boundary, that is: after dynamically determining the size of the structural element based on the actual size of the elevation difference, use a square kernel with the size SE of this element size to scan the boundary pixels of R1 / R2; when the center of the structural element aligns with the boundary pixel (i, j), if all 9 pixels covered by the structure belong to the current region, retain the boundary pixel, otherwise remove it from the current region to obtain the adjusted boundary; size Repeat the above steps, and update the region boundary each time according to the obtained elevation difference data until the convergence condition is reached or the preset number of iterations is completed.

[0064] Further, the steps of post - processing optimization include:

[0065] First, through morphological operations, determine the elevation value of the flood region for the feature map after long - range dependence; and dynamically adjust the processing granularity according to the elevation value of the flood region, and the specific formula is as follows:

[0066] where SE

[0067]

[0068] is the size of the updated morphological structural element, and H size is the elevation value of the finally determined flood region. flood

[0069] ​At this time, the boundary of the feature map is adjusted again using the iterative algorithm according to the size of the updated morphological structure element to obtain the fused feature map;

[0070] Then, for the obtained fused feature map, the consistency of the feature map is optimized by predicting the probability of each pixel. The specific formula is as follows:

[0071]

[0072] Among them, X represents the label configuration of all pixels, E(X) represents its energy function, represents the adjacent relationship between pixels, and ψ u (x i ) is the unary potential of a single pixel, representing the negative logarithm probability predicted by the CNN, reflecting the possibility that a single pixel is classified into a certain category; ψ p (x i ,x j ) is the binary potential between adjacent pixels, describing the relationship between adjacent pixels;

[0073] By minimizing the energy function, the flood detection optimization result that considers both pixel-level prediction accuracy and maintains spatial consistency is obtained.

[0074] Another object of the present invention is to provide a system for implementing the above-mentioned multi-modal fusion flood detection method based on deep learning, which is characterized by including:

[0075] A data acquisition module, which is used to acquire remote sensing image data during the flood period of the research area and preprocess the acquired original data;

[0076] A feature extraction module, which is used to extract spatial feature P1 and spectral feature P2 from the preprocessed data respectively;

[0077] A feature fusion module, which is used to fuse the spatial feature P1 and the spectral feature P2 to obtain a spatial-spectral fusion feature map;

[0078] A flood detection module, which is used to perform multi-level optimization on the obtained spatial-spectral fusion feature map to obtain the final flood detection result.

[0079] Compared with the prior art, the beneficial effects of the present invention are:

[0080] 1) The detection accuracy breakthrough of the present invention is achieved through the heterogeneous feature collaboration mechanism of spatial feature P1 and spectral feature P2. The improved ResNet50-CBAM network is used for spatial feature extraction, and the water body boundary morphological features are captured through the channel attention module (compression ratio 16) and multi-scale dilated convolution (rates = [6, 12, 18]), with the IoU index increased by 23%; the spectral feature innovatively uses the dynamic weighted AWEI index to strengthen the reflection difference between the near-infrared and short-wave infrared bands, and combines a 3-layer MLP network to generate spectral weights, reducing the urban misdetection rate by 21.9 percentage points. The adaptive fusion of the two is achieved through a four-layer convolutional weight network. After aligning the features using transposed convolution and bilinear interpolation, the spatial weight α (Sigmoid activation) is dynamically generated to form an optimized combination of Y fused = α⊙P1 + (1 - α)⊙P2. This mechanism enables the F1-score in the Poyang Lake test area to reach 91.5%, an increase of 19.5% compared to the single-modal method.

[0081] 2) The present invention adopts a multi-level optimization strategy. Through the optimization steps in step S4, the detection error rate is reduced from the original 25% to less than 5%, and the overall detection accuracy on multiple test data sets is stable above 90%, greatly improving the reliability and accuracy of the detection results.

[0082] 3) The present invention has good practical value and broad application prospects, and can be widely applied in fields such as flood disaster monitoring and emergency response, providing strong technical support for disaster prevention and mitigation work. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 It is a flowchart of a multi-modal fusion flood detection method based on deep learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0084] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0085] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0086] The present invention will be further described below in conjunction with specific embodiments, but it is not a limitation of the present invention.

[0087] As Figure 1 shown, the embodiment of the present invention discloses a multi-modal fusion flood detection method based on deep learning, which is characterized by including:

[0088] S1: Obtain the remote sensing image data during the flood season in the study area and preprocess the acquired original data;

[0089] In this embodiment, collect the remote sensing image data during the flood season (June - September) in the past 5 years acquired by the Sentinel 2 satellite. This satellite carries a multi - spectral imager (MSI), and this sensor provides 13 spectral bands, with the pixel size ranging from 10 to 60 meters. Among the spectral bands of the acquired images, three visible light bands of red (R), green (G), and blue (B), namely B4, B3, B2, the visible and near - infrared (VNIR) band B8, and the infrared SWIR bands B11, B12 will be included. After preprocessing these bands, they are given to step S2 for feature extraction.

[0090] Then, based on the above - collected data, perform data preprocessing, and the specific operations are as follows:

[0091] First, perform atmospheric correction on the data. Use the Sen2Cor 2.11 version to process the L1C - level data, and focus on correcting the water vapor absorption bands (infrared SWIR bands B11, B12) of the data;

[0092] Secondly, for the data after the above atmospheric correction, perform geometric precise correction. Based on ground control points, perform precise correction to ensure that the spatial position accuracy is better than 0.5 pixels;

[0093] Finally, for the data after the above geometric precise correction, perform data standardization. Perform band - level normalization on the data obtained through the above operations, and adopt a segmented normalization strategy for the SWIR bands (B11, B12).

[0094] S2: Extract the spatial feature P1 and spectral feature P2 respectively from the preprocessed data;

[0095] In this embodiment, the method for extracting the spatial feature P1 includes:

[0096] First, perform RGB composite image processing on the image data obtained in S1. Specifically, select three visible light bands of red (R), green (G), and blue (B) (B4, B3, B2) from the Sentinel - 2 multi - spectral data, and with the original resolution of 10m, downsample the above - mentioned bands by 4 times to synthesize the data, and perform average pooling operation. Finally, obtain the RGB image x with the input resolution downsampled by 4 times 4h×4w×3 ;

[0097] Then, based on the attention module architecture, for the downsampled RGB image, perform channel attention and spatial attention simultaneously and respectively, and obtain the focused composite image. Specifically, it includes:

[0098] In the channel attention mechanism, the input data is compressed at a ratio of 1:16. After that, first, channel dimension compression is performed. The H×W feature map of each channel is compressed into a 1×1 scalar through global average pooling, generating a 256-dimensional feature vector. Then, two fully connected operations are carried out: the first fully connected layer compresses the 256 dimensions to 16 dimensions (compression ratio 1:16), and the ReLU activation function is used to eliminate negative values to obtain the compressed channels; the second fully connected layer restores to 256 dimensions, restoring to the FC layer, and finally, the Sigmoid activation function is used to obtain the channel weight matrix.

[0099] In the spatial attention mechanism, for the input data, max pooling and average pooling are respectively performed to obtain two tensors of 128×128. The results of the two poolings are concatenated in the channel dimension to form a tensor of 2×128×128. Then, spatial convolution is performed on this tensor through a 7×7 convolutional kernel to obtain the spatial weight matrix.

[0100] Next, three groups of parallel convolutional kernels are constructed. The feature matrices obtained from the above two attentions are input into the three groups of parallel convolutional kernels for convolution, and the focused composite image is output. For the focused composite image, ResNet50 is used as the backbone network for feature extraction to obtain the spatial feature P1, and its extraction process can be expressed as:

[0101] P1 = F ResNet (x 4h×4w×3 )

[0102] In the formula, F ResNet represents the ResNet50 feature extraction function;

[0103] Finally, the spatial feature map P1 is output, and its dimension is h×w×256.

[0104] In this embodiment, based on the data processed in step S1, the method for extracting the spectral feature P2 includes:

[0105] First, based on the band data of the image after S1 preprocessing, the Automated Water Extraction Index (AWEI) is constructed:

[0106] AWEI = B2 + 2.5×B3 - 1.5×(B8 + B 11 ) - 0.25×B 12

[0107] where B2 and B3 are the reflectivities of the blue and green bands respectively, B8 is the reflectivity of the visible and near-infrared (VNIR) band, and B 11 , B 12 correspond to the reflectivities of the infrared SWIR B11 and B12 bands respectively. These data respectively correspond to the reflectivity values obtained in different bands after S1 preprocessing;

[0108] Secondly, according to the constructed AWEI index above, perform Min-Max normalization to obtain the spectral feature P2. The formula is as follows:

[0109]

[0110] In the formula, AWEI min and AWEI max are the minimum and maximum values of the Automated Water Extraction Index (AWEI) respectively, and P2 is the extracted spectral feature map.

[0111] Finally, output the feature map P2. The dimension of P2 is the same as that of P1, which is h×w×256.

[0112] S3: Fuse the spatial feature P1 and the spectral feature P2 to obtain a spatial-spectral fusion feature map; in this embodiment, this step includes:

[0113] S3.1, perform spatial feature reconstruction on the spatial feature map obtained in step S2. Specifically, use two-stage transposed convolution on the spatial feature map obtained in step S2 to achieve 4-fold upsampling and restore its resolution, thereby obtaining a spatial feature map with higher resolution; the two-stage levels are as follows:

[0114] The first stage: Convert the spatial feature map obtained in step S2 from 256 channels to 128 channels, then use a 4×4 convolution kernel, a stride of 2, and a padding of 1; combined with batch normalization and ReLU activation, perform the first convolution on the input spatial feature map;

[0115] The second stage: Convert the feature map of the first-stage convolution from 128 channels to 64 channels, with the same parameter configuration as the first stage. Through the concatenation operation, restore the feature map size from the original 1 / 4 to the full resolution to obtain the reconstructed spatial feature map P1'.

[0116] S32, perform bilinear interpolation on the spectral feature map obtained from step S2 to obtain a corrected spectral feature map. To prevent between bands, perform interpolation on each band of the multispectral data separately. The specific steps are as follows:

[0117] First, perform resolution reconstruction on the spectral feature map. The resolution reconstruction method is the same as that of the spatial feature map. After that, establish the coordinate correspondence between the target high-resolution grid and the original low-resolution feature map for the reconstructed spectral feature map. For 4-fold upsampling, each original pixel point will correspond to a 4×4 new grid area, and map the target coordinates to the floating-point position index of the original feature map through the normalization parameter.

[0118] Next, for the obtained spectral feature map, weight calculation is performed for each band image. For the target coordinate position of each pixel, the 4 nearest original pixels around it (upper left, upper right, lower left, lower right) are selected to form the calculation basis of bilinear interpolation. According to the horizontal and vertical distance differences (dx, dy) between the target point and each neighboring point, the bilinear weight coefficients are calculated, and the weights in the two directions are multiplied to obtain the final contribution weights of the four neighboring points, ensuring that the sum of the weights is 1. The specific formula is as follows:

[0119]

[0120] w 11 =(1 - dx)(1 - dy)

[0121] w 21 =dx(1 - dy)

[0122] w 12 =(1 - dx)dy

[0123] w 22 =dx·dy

[0124] Σw = w 11 +w 21 +w 12 +w 22 =1

[0125] In the formula, (x, y) are the floating-point coordinates of the target point, (z1, y1), (x2, y2) are the integer coordinates of the nearest upper left and lower right points, dx and dy are the normalized distances in the X and Y axis directions respectively, and w 11 , w 21 , w 12 , w 22 are the weight coefficients of the four neighboring pixels.

[0126] Finally, for each band of the image, determine the color of each pixel in the new feature map. Among them, the RGB value of each pixel is the weighted average of the RGB values of the 4 neighboring pixels according to the weight coefficients obtained above. After the above operations, the images generated by each band are finally obtained, and then the final new spectral feature map P2' is obtained.

[0127] S3.3. Calculate the fusion weights for the spatial feature map P1' obtained in step S3.1 and the spectral feature map P2' obtained in step S3.2. The specific formula is as follows:

[0128] α = Sigmoid(Conv(Concat[P1′, P2′]))

[0129] In the formula, Concat is the splicing function, Conv is the convolution function, Sigmoid is the normalization function, and α is the obtained fusion weight coefficient.

[0130] Then, according to the above steps for the fusion weight coefficient, perform feature weighted fusion on the spatial feature flow map P1' and the spectral feature flow map P2' to obtain a spatial-spectral fusion feature map:

[0131] Y fused =α·P1′+(1-α)·P2′

[0132] In the formula, α is the fusion weight coefficient, and Y fused is the obtained spatial-spectral fusion feature map.

[0133] S4: Perform multi-level optimization on the obtained spatial-spectral fusion feature map to obtain the final flood detection result;

[0134] In this step, for the spatial-spectral fusion feature map obtained in step S3, perform detection optimization processing successively to finally obtain the result map of flood detection. The specific operation process is as follows:

[0135] S4.1, adopt the DeepLabv3+ architecture model. For the fusion feature map obtained in step S3, perform semantic segmentation to identify and label different ground object types (such as buildings, vegetation, water bodies, etc.) in the image. The specific process is as follows:

[0136] First, segment a part from the set of fusion feature maps obtained in S3 for training the DeepLabv3+ architecture model. The segmented data is divided into a training set, a validation set, and a test set according to 7:2:1. Use the training set to train the DeepLabv3+ architecture model and use the validation set for validation. During the training process, use the following loss function to optimize the model parameters and drive the network to learn the pixel-level classification task:

[0137] L total =λ1L CE +λ2L Dice

[0138] Among them, L CE ,L Dice ,L total respectively represent the cross-entropy loss, Dice coefficient loss, and comprehensive loss of the corresponding model; λ1 and λ2 are balance coefficients, and in this embodiment, the typical values are 0.6 and 0.4 respectively;

[0139] Use the test set to test the effectiveness of the trained DeepLabv3+ architecture model. If the test is successful, it can be used for the subsequent use. Otherwise, retrain according to the above process.

[0140] Then, use the successfully tested DeepLabv3+ architecture model to perform semantic segmentation on the remaining data after segmentation, obtaining a flood detection result map after semantic segmentation.

[0141] S4.2. Introduce DEM data as a priori constraint, and perform geographical optimization on the flood detection result map after semantic segmentation to fine-tune the boundary of the flood detection result map after semantic segmentation. The specific process is as follows:

[0142] First, introduce digital elevation model DEM data for data spatial alignment. Use SRTM 1 arc-second elevation data (spatial resolution of 30 meters), and resample the DEM data to the same coordinate system (UTM projection) and pixel size (10 meters) as the Sentinel-2 image through affine transformation. Then, through bicubic interpolation, align the DEM data space with the flood detection result map after semantic segmentation. Specifically: First, use the UTM projection formula to establish a two-way mapping relationship from DEM to UTM. Then, for the target UTM grid points, establish an interpolation window among the surrounding 4×4 original DEM elevation points. Use the Keys cubic convolution kernel as the interpolation weight function for weighted summation calculation to obtain an elevation matrix H with the same size as the segmentation result.

[0143] Next, for each segmented region R, calculate its average elevation value. The formula is as follows:

[0144]

[0145] where |R| represents the number of pixel points in region R, H ij represents the elevation value at position (i,j), and h R is the average elevation value of the segmented region R.

[0146] Then, based on an iterative algorithm, perform boundary adjustment on the flood detection result map after semantic segmentation and the obtained elevation value data. The specific steps are as follows:

[0147] ① For each adjacent region R1, R2, calculate their average elevation difference Δh = |h R1 -h R2 |

[0148] ② For each group of elevation differences calculated above, if it is found that Δh > τ, perform morphological erosion operation to obtain an updated boundary. The specific method is: First, dynamically determine the size of the structural element according to the actual size of the elevation difference After that, next, use a square kernel with the element size SE size for R1 / R wScan the boundary pixels. When the center of the structural element aligns with the boundary pixel (i, j), if all 9 pixels covered by the structure belong to the current region, retain the boundary pixel; otherwise, remove it from the current region to obtain the adjusted boundary.

[0149] Repeat the above steps. Each iteration updates the regional boundary based on the obtained elevation difference data until the convergence condition is met or the preset number of iterations is completed.

[0150] S4.3. By introducing the Transformer architecture, capture the long-range dependencies of the flood detection result map after boundary adjustment, such as flood region connectivity, etc. The specific process is as follows:

[0151] a. Construct a hierarchical window attention for the flood detection result map after boundary adjustment. Specifically, divide the flood detection result map after boundary adjustment into non-overlapping windows of M×M (a typical value of M = 16). There is a corresponding feature vector in each window, which is composed of multi-modal information in three-dimensional space (including the original multi-spectral reflectance, the derived features of the elevation constraint matrix, and the semantic segmentation confidence);

[0152] b. For each window, calculate the multi-head sub-attention within it. Specifically, for the feature vector within each window, first generate a query vector (Query), a key vector (Key), and a value vector (Value) through a linear transformation. Then, calculate the dot product of the query vector and the key vector, and after normalization, obtain the attention weight matrix;

[0153] c. Multiply the attention weight by the value vector to obtain the weighted feature representation.

[0154] S4.4. For the feature map after collecting the long-range dependencies above, perform post-processing optimization. The specific operation process is as follows:

[0155] First, through morphological operations, determine the elevation values of the flood regions in the feature map after the long-range dependencies. The operations include:

[0156] ① Sample multiple points within the identified flood regions and collect the DEM elevation data of these points. These elevation data form a data set D = {h1, h2, …, h N}. At the same time, through statistical analysis methods, calculate the elevation distribution characteristics of these sampling points, including the mean value, standard deviation, etc.;

[0157] ②Using the kernel density estimation method, based on the DEM elevation data and statistical results obtained above, obtain the probability distribution of elevation values. First, use the Epanechnikov kernel function to perform probability density estimation on the above statistical data to obtain the probability density distribution function K(u); then, through the Silverman criterion, calculate the appropriate bandwidth h. Finally, obtain the contribution function of each sample point:

[0158]

[0159] where f(h) is the contribution function of the obtained sample points, that is, the probability density distribution function, N is the number of sampling points, K(u) is the probability density distribution function calculated above, h is the bandwidth calculated above, and h i is the elevation data of the i-th sampling point. Through this calculation, a continuous elevation probability density curve can be obtained, and the peak value corresponds to the most likely elevation value in the area.

[0160] ③Use the clustering algorithm to identify the main elevation levels. Use the Gaussian mixture model (GMM) to perform modal decomposition on the elevation data and the corresponding contribution function. Set the number of clusters k ∈ [2, 5], and select the optimal model through the Bayesian information criterion (BIC). When the obtained BIC is the smallest, the corresponding k is the main number of clusters, that is, the main elevation level, and its corresponding Gaussian components (μ i , σ i , ω i ) are the obtained Gaussian component results, where ω i represents the weight coefficient of each component, which is obtained through iterative calculation by the expectation maximization (EM) algorithm.

[0161] ④By the method of weighted average, determine the final flood elevation value. The weights are set based on the confidence of each elevation level. The specific calculation formula is as follows:

[0162]

[0163] In the formula, k is the number of clusters obtained in the previous step, μ i , σ i , ω i are the mean value, standard deviation, and weight coefficient of each Gaussian component respectively, is the overall variance, and H fjood is the finally obtained flood elevation value.

[0164] At the same time, according to the obtained flood area value, dynamically adjust the processing granularity, and the specific formula is as follows:

[0165]

[0166] In the formula, SE sizeis the size of the updated morphological structure element, H flood is the elevation value of the finally determined flood area. At this time, according to the size of the updated morphological structure element, the iterative algorithm in the above step S4.2 is used again to adjust the boundary, and a fused feature map is obtained;

[0167] Then, the boundary distinction of the obtained fused feature map is further refined. Specifically, by predicting the probability of each pixel, the consistency of the feature map is optimized. The specific formula is as follows:

[0168]

[0169] Among them, X represents the label configuration of all pixels, E(X) represents its energy function, and represents the adjacent relationship between pixels. ψ u (x i ) is the unary potential of a single pixel, representing the negative logarithm probability predicted by the CNN, reflecting the possibility that a single pixel is classified into a certain category; ψ p (x i ,x j ) is the binary potential between adjacent pixels, describing the relationship between adjacent pixels.

[0170] For the unary potential in the above formula, its calculation formula is as follows:

[0171] ψ u (x i ) = -log P(x i |CNN)

[0172] Among them, ψ u (x i ) is the unary potential of a single pixel, and log P(x i |CNN) is the classification probability predicted by the CNN

[0173] The binary potential formula is as follows:

[0174] ψ p (x i ,x j ) = μ(x i ,x j )[w spectral k(f i ,f j ) + w spatial k(p i ,p j )]

[0175] Among them, ψ p (x i ,x j ) is the binary potential between adjacent pixels, μ(xi , x j ) is a label compatibility function, k(f i , f j ) is a kernel function based on pixel features, k(p i , p j ) is a kernel function based on pixel positions; w spectral and w spatial are the spectral kernel weight and the spatial kernel weight respectively; in this embodiment, the weight values are taken as w spectral = 0.6,

[0176] w spat i a l = 0.4.

[0177] By minimizing the energy function E(X), an optimized flood detection result that takes into account both pixel-level prediction accuracy and spatial consistency is obtained.

[0178] To further improve the effect, after the above feature image consistency optimization, bilateral filtering can also be used to remove noise and enhance details. Through the above steps, the finally output optimized feature image and the predicted flood height are the final results.

[0179] The embodiment of the present invention also provides a system for implementing the above-mentioned multi-modal fusion flood detection method based on deep learning, which is characterized in that it includes:

[0180] A data acquisition module for acquiring remote sensing image data of the flood period in the research area and preprocessing the acquired original data;

[0181] A feature extraction module for respectively extracting spatial feature P1 and spectral feature P2 from the preprocessed data; in the feature extraction module, the spatial feature P1 is extracted by combining a spatial attention mechanism (S2) and dilated convolution, and the spectral feature P2 is extracted by a physical-data dual-driven dynamic weighting AWEI (S2);

[0182] A feature fusion module for fusing the spatial feature P1 and the spectral feature P2 to obtain a spatial-spectral fusion feature map;

[0183] A flood detection module for performing multi-level optimization on the obtained spatial-spectral fusion feature map to obtain the final flood detection result; in the flood detection module, a lightweight fusion network DeepLabv3+ is used for semantic segmentation, a terrain constraint CRF (S4) is used to establish an elevation-morphology joint optimization model, and finally a distributed architecture is used to achieve the optimal allocation of computing-storage-network resources.

[0184] In the system of this embodiment, the data input interface supports the NetCDF / HDF5 format and provides a RESTful API to efficiently input the collected data into the database. The result output interface generates GeoJSON vector layers and GeoTIFF raster data and performs visualization. Specifically, it integrates the Cesium 3D rendering engine and supports spatio-temporal contrast analysis, enabling the results of flood detection to be clearly presented on the screen.

[0185] The system of this embodiment integrates the DeepLabv3+ segmentation network and the Transformer optimization module. Parallel computing is achieved by deploying the NVIDIA DGX A100 cluster, combined with the DEM terrain constraint and the CRF post-processing chain. These computing processes can effectively eliminate the interference of mountain shadows (the false detection rate is reduced by 67%) and urban specular reflections (the missed detection rate is reduced by 41%), and the F1-score in complex terrain scenarios is increased to 89.3%, and the processing time for a single scene image is compressed to 182 seconds.

[0186] The system of this embodiment is designed and implemented through modular design and feature extraction and fusion means to achieve flexible expansion of the processing flow, achieving the following effects:

[0187] The combination of the spatial attention mechanism (S2) and dilated convolution increases the IoU of irregular boundaries by 23%.

[0188] The dynamic weighted AWEI (S2) reduces the urban false detection by 21.9 percentage points through physical-data dual drive. The lightweight fusion network (S3) reduces the parameter storage by 85% while ensuring the accuracy.

[0189] The terrain constraint CRF (S4) establishes an elevation-morphology joint optimization model, enabling the F1-score of mountain detection to reach 89.3%.

[0190] The distributed architecture realizes the optimal ratio of computing-storage-network resources, and the overall energy efficiency ratio is increased by 3.7 times.

[0191] Each technical module forms a positive gain cycle through cascaded optimization: improving the feature extraction quality reduces the fusion difficulty, adaptive fusion enhances the discriminability of the detection input, and post-processing optimization corrects the inherent bias of the model. This progressive optimization system enables the comprehensive accuracy of the system on the Pearl River Delta test set to reach 92.7%, a 19.5% improvement compared to traditional single-modal methods, meeting the requirements of ISO 19157 Geographical Information Quality Standard Class A.

[0192] The above are only the preferred embodiments of the present invention, and do not limit the implementation manners and protection scope of the present invention. For those skilled in the art, it should be realized that all the solutions obtained by equivalent replacement and obvious changes made by using the content of the specification of the present invention should be included in the protection scope of the present invention.

Claims

1. A multi-modal fusion flood detection method based on deep learning, characterized in that, Including: S1: Obtain remote sensing image data during the flood period of the research area, and preprocess the obtained original data; S2: Extract spatial feature P1 and spectral feature P2 from the preprocessed data respectively; S3: Fuse the spatial feature P1 and the spectral feature P2 to obtain a spatial-spectral fusion feature map; S4: Perform multi-level optimization on the obtained spatial-spectral fusion feature map to obtain the final flood detection result.

2. The multi-modal fusion flood detection method based on deep learning according to claim 1, wherein The method for extracting spatial features in step 2 includes: For the image data processed in step S1, perform RGB synthesis to obtain an RGB image; Based on the attention mechanism, for the obtained RGB image, simultaneously focus on channel attention and spatial attention respectively to obtain a focused synthesis image; Finally, use ResNet50 to extract features from the focused synthesis image to obtain spatial feature P1.

3. The multi-modal fusion flood detection method based on deep learning according to claim 2, wherein, The method for obtaining a focused synthesis image using the attention mechanism includes: In the channel attention mechanism, compress the input data. After that, first perform channel dimension compression and global average pooling to generate a feature vector; then perform two fully connected operations to obtain a channel weight matrix; In the spatial attention mechanism, perform max pooling and average pooling on the input data respectively and concatenate the two pooling results in the channel dimension to form a tensor, and then perform spatial convolution on the tensor through a convolutional kernel to obtain a spatial weight matrix; Construct multiple groups of parallel convolutional kernels, input the channel weight matrix and the spatial weight matrix into the multiple groups of parallel convolutional kernels for convolution, and output the focused synthesis image.

4. The multi-modal fusion flood detection method based on deep learning according to claim 1, wherein The method for extracting spectral feature P2 in step S2 includes: First, based on the image data preprocessed in step 1, obtain the reflectance of different bands, and thereby construct the Automatic Water Extraction Index (AWEI); AWEI = B2 + 2.5×B3 - 1.5×(B8 + B 11 ) - 0.25×B 12 where B2 and B3 correspond to the reflectance of the blue and green bands respectively, B8 is the reflectance of the visible and near-infrared bands, and B 11 , B 12 are the reflectance of the infrared SWIR band; Secondly, according to the constructed AWEI index, perform Min-Max normalization processing to obtain spectral feature P2, and the formula is: Wherein, AWEI min and AWEI max are respectively the minimum value and the maximum value of the Automated Water Extraction Index (AWEI), and P2 is the extracted spectral feature map.

5. The multi-modal fusion flood detection method based on deep learning according to claim 1, wherein, The specific implementation method of step S3 includes: S3.1, Use two-stage transposed convolution on the spatial feature map obtained in step S2 to achieve n-fold upsampling and restore its resolution, thereby obtaining a spatial feature map P1' with the target resolution; S3.2, Perform bilinear interpolation on the spectral feature map obtained from step S2 to obtain a corrected spectral feature map P2'; S3.3, Calculate the fusion weight for the spatial feature map P1' obtained in step S3.1 and the spectral feature map P2' obtained in step S3.

2. The specific formula is as follows: α = Sigmoid(Conv(Concat[P1′,P2′])) In the formula, Concat is the concatenation function, Conv is the convolution function, Sigmoid is the normalization function, and α is the obtained fusion weight coefficient; Then, according to the above-mentioned fusion weight coefficient, perform feature weighted fusion on the spatial feature map P1' and the spectral feature map P2' to obtain a spatial-spectral fusion feature map: Y fusrd = α·P1′ + (1 - α)·P2′ where α is the fusion weight coefficient, and Y fused is the obtained spatial-spectral fusion feature map.

6. The multi-modal fusion flood detection method based on deep learning according to claim 5, wherein The method for obtaining the corrected spectral feature map P2' includes First, perform resolution reconstruction on the spectral feature map to obtain a spectral feature map with the target high resolution; After that, for the reconstructed spectral feature map, establish the coordinate correspondence between the target high-resolution grid and the low-resolution feature map before restoration; Next, for the target coordinate position of each pixel, select the 4 nearest original pixel points around it to form the calculation primitive of bilinear interpolation. According to the horizontal and vertical distance differences (dx, dy) between the target point and each neighboring point, calculate the bilinear weight coefficients, multiply the weights in the two directions to obtain the final contribution weights of the four neighboring points, and ensure that the sum of the weights is 1; The specific formula is as follows: w 11 = (1 - dx)(1 - dy) w 21 = dx(1 - dy) w 12 = (1 - dx)dy w 22 = dx·dy Σw = w 11 + w 21 + w 12 + w 22 = 1 Where (x, y) are the floating-point coordinates of the target point, (x1, y1) and (x2, y2) are the integer coordinates of the top-left and bottom-right points of the nearest neighbor, and dx and dy are the normalized distances in the X and Y axis directions, respectively, w 11 , w 21 , w 12 , w 22 are the weight coefficients of the four neighboring pixels; Finally, for the image of each band, calculate the color of each pixel in the new feature map. Among them, the RGB value of each pixel is the weighted average of the RGB values of the 4 neighboring pixels according to the weight coefficients obtained above. After the above operations, the images generated by each band are finally obtained, and then the final spectral feature map is obtained.

7. The multi-modal fusion flood detection method based on deep learning according to claim 1, characterized in that, The specific implementation method of step S4 includes: S4.1, construct a DeepLabv3+ architecture model, segment a part from the set of fused feature maps obtained in S3 for training the DeepLabv3+ architecture model, and then use the trained DeepLabv3+ architecture model to perform semantic segmentation processing on the remaining data to obtain a flood detection result map after semantic segmentation; S4.2, introduce digital elevation model DEM data as a priori constraint to adjust the boundary of the flood detection result map after the above semantic segmentation; S4.3, adopt the introduced Transformer architecture to capture the long-range dependence of the flood detection result map after boundary adjustment, S4.4, for the feature map after collecting the long-range dependence above, perform post-processing optimization to obtain the final flood detection result.

8. The multi-modal fusion flood detection method based on deep learning according to claim 7, wherein The method of boundary adjustment includes: a. Align the elevation model DEM data space with the flood detection result map after semantic segmentation to obtain an elevation matrix H of the same size as the segmentation result; for each segmentation region R, calculate its average elevation value, and the formula is as follows: Among them, |R| represents the number of pixel points in region R, and H ij represents the elevation value at position (i, j), and h R is the average elevation value of the segmented region R; b. For the flood detection result map after semantic segmentation and the obtained elevation value data, perform boundary adjustment, including: ①For each pair of adjacent regions R1 and R2, calculate their average elevation difference Δh = |h R1 - h R2 |; ②For each group of elevation differences calculated above, if it is found that Δh > τ, perform morphological erosion operation to obtain the updated boundary, that is: dynamically determine the size of the structural element based on the actual size of the elevation difference Use the element size SE size of the square kernel to scan the boundary pixels of R1 / R2; when the center of the structural element aligns with the boundary pixel (i, j), if all 9 pixels covered by the structure belong to the current region, retain the boundary pixel, otherwise remove it from the current region to obtain the adjusted boundary; The above steps are repeated, and each iteration updates the regional boundary according to the obtained elevation difference data until the convergence condition is reached or the preset number of iterations is completed.

9. The multi-modal fusion flood detection method based on deep learning according to claim 7, characterized in that, The steps of post-processing optimization include: First, through morphological operations, judge the elevation value of the flood area for the feature map after long-range dependence; and dynamically adjust the processing granularity according to the elevation value of the flood area, and the specific formula is as follows: Among them, SE size is the size of the updated morphological structure element, and H flood is the elevation value of the finally determined flood area. At this time, according to the size of the updated morphological structure element, use the iterative algorithm again to adjust the boundary of the feature map to obtain a fused feature map; Then, for the obtained fused feature map, optimize the consistency of the feature map by predicting the probability of each pixel, and the specific formula is as follows: Among them, X represents the label configuration of all pixels, E(X) represents its energy function, represents the adjacent relationship between pixels, ψ u (x i ) is the unary potential of a single pixel, representing the negative log probability predicted by the CNN, reflecting the possibility that a single pixel is classified into a certain category; ψ p (x i , x j ) is the binary potential between adjacent pixels, describing the relationship between adjacent pixels; By minimizing the energy function, a flood detection optimization result that takes into account both pixel-level prediction accuracy and spatial consistency is obtained.

10. A system for implementing the deep learning-based multi-modal fusion flood detection method according to any one of claims 1-9, characterized in that, It includes: A data acquisition module, used to acquire remote sensing image data during the flood period of the research area and preprocess the acquired original data; A feature extraction module, used to extract spatial feature P1 and spectral feature P2 from the preprocessed data respectively; A feature fusion module for fusing the spatial feature P1 and the spectral feature P2 to obtain a spatial-spectral fusion feature map; A flood detection module for performing multi-level optimization on the obtained spatial-spectral fusion feature map to obtain the final flood detection result.

Citation Information

Cited By

  • Semi-autonomous remote sensing water area generalization segmentation method fused with multispectral inversion

    CN120808175A

  • Mobile emergency unit cluster intelligent management method and device based on Internet of Things

    CN122025055A