Methods, devices, storage media and electronic equipment for monitoring flood disasters

CN121415258BActive Publication Date: 2026-09-01NORTH CHINA ELECTRICAL POWER RES INST +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511787119.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-09-01
Estimated Expiration
2045-12-01

AI Technical Summary

Technical Problem

两者均具有未充分利用SAR图像的双极化(VV/VH)数据特性,且模型仅依赖空间维度的特征注意力,未融合时序维度的动态关联,导致对洪涝灾害动态监测场景适配性弱、精度不足、动态变化显著区域的误判率较高

Benefits of technology

[0018]本申请并非采用传统固定卷积核尺寸的多尺度结构,而是通过自适应极化融合多尺度金字塔结构提取多尺度特征图,既有效解决了小尺度与全局特征失衡的问题,又弥补了现有技术对SAR双极化数据利用不足的缺陷。该结构通过计算双极化通道的极化散射矩阵特征值与纹理一致性协同系数,动态调整卷积核权重,同时根据前一层级特征响应强度迭代更新卷积核尺寸;配合基于特征稀疏度的动态池化层,既能精准捕捉零星积水、水体边缘等小尺度细节,又能保留大范围淹没区域等全局语义信息,还能强化双极化特征一致的洪涝区域响应。相比传统单一尺度或固定多尺度方法,这种多尺度特征提取方式,克服了小尺度特征丢失问题,确保对复杂地形中多样洪涝形态的识别能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121415258B_ABST
    Figure CN121415258B_ABST
Patent Text Reader

Abstract

This application discloses a method, device, storage medium, and electronic device for monitoring flood disasters. The method includes: acquiring multi-temporal synthetic aperture radar (SAR) image data of a target area; preprocessing the multi-temporal SAR image data to obtain SAR images that retain flood boundary features; extracting multi-scale feature maps from the preprocessed SAR images using an adaptive polarization fusion multi-scale pyramid structure; dynamically weighting the multi-scale feature maps using a spatiotemporal joint attention weight map; performing adaptive noise suppression and reconstruction on the dynamically weighted multi-scale feature maps to obtain low-dimensional feature representations; performing feature calibration on the low-dimensional feature representations based on historical flood data of the target area; performing multi-scale feature fusion and attention weighting adjustment on the calibrated low-dimensional feature representations; and outputting flood classification results for the target area by combining fully connected mapping and probability calculation; and generating flood disaster distribution information based on the flood classification results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, storage medium and electronic device for monitoring flood disasters. Background Technology

[0002] Existing flood disaster monitoring methods based on Synthetic Aperture Radar (SAR) images can be divided into two categories: traditional threshold analysis methods and deep learning-driven methods that have emerged in recent years. However, both types of methods have technical bottlenecks and are difficult to meet the monitoring requirements of high precision and high adaptability in complex scenarios.

[0003] One early approach was the backscattering coefficient analysis method based on a traditional fixed threshold. The process involves: first, preprocessing the original SAR image, such as radiometric correction, to convert pixel values ​​into physically meaningful backscattering coefficients; then, leveraging the characteristic that water bodies typically exhibit low backscattering in SAR images, a fixed backscattering coefficient threshold is artificially set; finally, through pixel-by-pixel comparison, areas with backscattering coefficients less than the threshold are identified as flooded areas, while areas with backscattering coefficients greater than or equal to the threshold are identified as non-flooded areas, thus completing the identification and monitoring of flooded areas. However, this method suffers from significant drawbacks, including weak anti-interference capabilities, making it unsuitable for monitoring complex scenarios. The core cause of this deficiency lies in the diversity and dynamism of the scattering characteristics of ground objects in the natural environment. On the one hand, the backscattering coefficients of some non-floodable ground objects (such as bare soil moistened after rain and densely growing low vegetation) may be close to or even lower than the set threshold, leading to their misidentification as flooded areas. On the other hand, some special flooded areas (such as water bodies with a large amount of floating debris and riverbeds with rough bottoms in shallow water areas) may have their actual backscattering coefficients higher than the threshold due to changes in the scattering mechanism, thus being incorrectly excluded from flooded areas. At the same time, factors such as surface roughness in different regions and the incident angle during SAR imaging can significantly affect the backscattering characteristics of ground objects. Fixed thresholds cannot adapt to these variables, resulting in large differences in monitoring accuracy under different scenarios, making it difficult to meet the high-precision monitoring requirements in complex flood scenarios.

[0004] To overcome the limitations of traditional methods, deep learning-based monitoring schemes have emerged in recent years, with representative techniques including convolutional network methods and multi-scale large convolutional kernel feature extraction methods. Both of these methods fail to fully utilize the dual-polarization (VV / VH) data characteristics of SAR images, and the models rely solely on spatial dimension feature attention without incorporating dynamic correlations in the temporal dimension. This results in poor adaptability to dynamic monitoring scenarios of flood disasters, insufficient accuracy, and a high misclassification rate in areas with significant dynamic changes.

[0005] In summary, existing SAR image-based flood monitoring technologies generally suffer from four common problems: 1) Insufficient utilization of SAR dual-polarization data, failing to fully leverage the synergistic and complementary effects of the two channels; 2) Rigid multi-scale feature extraction structures, unable to adapt to the differences in feature distribution across different regions; 3) Lack of a spatiotemporal joint feature attention mechanism, making it difficult to accurately capture the dynamic evolution characteristics of flood disasters; 4) Failure to construct prior constraints based on historical flood data, resulting in weak adaptability and generalization ability of the models to different regions and scenarios. These problems make it difficult for existing technologies to meet the practical needs of "high-precision identification, full-cycle monitoring, and wide-scenario adaptation" in complex flood scenarios, necessitating an improved monitoring method that can overcome these bottlenecks. Summary of the Invention

[0006] In view of the above problems, this application provides a method, device, storage medium and electronic equipment for monitoring flood disasters.

[0007] To solve the above-mentioned technical problems, this application proposes the following solution:

[0008] Firstly, this application provides a method for monitoring flood disaster conditions. The method includes: acquiring multi-temporal synthetic aperture radar (SAR) image data of a target area, wherein the target area is a flood-sensitive area, and the multi-temporal SAR image data contains dual polarization data of VV and VH polarization channels; preprocessing the multi-temporal SAR image data to obtain a SAR image that retains flood boundary features; and extracting multi-scale feature maps from the preprocessed SAR image using an adaptive polarization fusion multi-scale pyramid structure. The adaptive polarization fusion multi-scale pyramid structure is implemented through alternating operations of multiple layers of adaptive polarization fusion convolutional units and dynamic pooling layers. The adaptive polarization fusion convolutional unit has a built-in VV-VH channel interaction module, which calculates the eigenvalues ​​of the polarization scattering matrices of the VV and VH channels in conjunction with the texture consistency parameters. The system dynamically adjusts the convolutional kernel weights based on coefficients and iteratively updates them according to the feature response intensity of the previous layer. The dynamic pooling layer adaptively adjusts the pooling window size and stride based on the feature sparsity of the corresponding convolutional layer, achieving feature dimensionality reduction and key information preservation. A spatiotemporal joint attention weight map is used to dynamically weight multi-scale feature maps to improve the spatiotemporal correlation feature weights of flood-sensitive areas. Adaptive noise suppression and reconstruction are performed on the dynamically weighted multi-scale feature maps to obtain a low-dimensional feature representation that retains key information about flood-sensitive areas. Based on historical flood data of the target area, the low-dimensional feature representation is calibrated, and multi-scale feature fusion and attention weighting are applied to the calibrated low-dimensional feature representation. The flood classification results for the target area are output by combining fully connected mapping and probability calculation. Flood disaster distribution information is generated based on the flood classification results to monitor the flood disaster situation in the target area.

[0009] Secondly, this application provides a flood disaster monitoring device, which includes:

[0010] The acquisition module is used to acquire multi-temporal synthetic aperture radar image data of the target area, which is a flood-sensitive area. The multi-temporal synthetic aperture radar image data includes dual polarization data of VV polarization channel and VH polarization channel.

[0011] The preprocessing module is used to preprocess multi-temporal synthetic aperture radar image data to obtain synthetic aperture radar images that retain flood boundary features;

[0012] The feature extraction module extracts multi-scale feature maps from preprocessed synthetic aperture radar images using an adaptive polarization fusion multi-scale pyramid structure. This structure is achieved through alternating operations of multiple adaptive polarization fusion convolutional units and dynamic pooling layers. Each convolutional unit incorporates a VV-VH channel interaction module, which dynamically adjusts the convolutional kernel weights by calculating the eigenvalues ​​of the polarization scattering matrices and the texture consistency parameter of the VV and VH channels. The kernel weights are iteratively updated based on the feature response intensity of the previous layer. The dynamic pooling layer adaptively adjusts the pooling window size and stride based on the feature sparsity of the corresponding convolutional layer, achieving feature dimensionality reduction and key information preservation. A spatiotemporal joint attention weight map is used to dynamically weight the multi-scale feature maps, enhancing the spatiotemporal correlation feature weights for flood-sensitive areas. Adaptive noise suppression and reconstruction are then applied to the dynamically weighted multi-scale feature maps to obtain a low-dimensional feature representation that retains key information about flood-sensitive areas.

[0013] The classification module is used to perform feature calibration on low-dimensional feature representations based on historical flood data of the target area, perform multi-scale feature fusion and attention weighting adjustment on the calibrated low-dimensional feature representations, and output the flood classification results of the target area by combining fully connected mapping and probability calculation.

[0014] The generation module is used to generate flood disaster distribution information based on flood classification results in order to monitor the flood disaster situation in the target area.

[0015] To achieve the above objectives, according to a third aspect of this application, a storage medium is provided, the storage medium including a stored program, wherein, when the program is executed, the device on which the storage medium is located is controlled to perform the flood disaster monitoring method of the first aspect.

[0016] To achieve the above objectives, according to a fourth aspect of this application, an electronic device is provided, the device including at least one processor, and at least one memory and bus connected to the processor; wherein the processor and memory communicate with each other through the bus; the processor is used to call program instructions in the memory to execute the flood disaster monitoring method of the first aspect described above.

[0017] By employing the above-described technical solution, the technical solution provided in this application has at least the following advantages:

[0018] This application does not employ the traditional multi-scale structure with fixed convolutional kernel size. Instead, it extracts multi-scale feature maps through an adaptive polarization fusion multi-scale pyramid structure. This effectively solves the problem of imbalance between small-scale and global features and compensates for the shortcomings of existing technologies in utilizing SAR dual-polarization data. This structure dynamically adjusts the convolutional kernel weights by calculating the eigenvalues ​​of the polarization scattering matrix of the dual-polarization channels and the texture consistency co-coefficient, while iteratively updating the convolutional kernel size based on the feature response intensity of the previous layer. Combined with a dynamic pooling layer based on feature sparsity, it can accurately capture small-scale details such as sporadic water accumulation and water body edges, while preserving global semantic information such as large-scale flooded areas, and enhancing the consistent flood area response of dual-polarization features. Compared to traditional single-scale or fixed multi-scale methods, this multi-scale feature extraction approach overcomes the problem of small-scale feature loss, ensuring the ability to identify diverse flood patterns in complex terrain.

[0019] By employing a spatiotemporal joint attention mechanism to achieve dynamic weighting, this approach focuses on the static features of flood-sensitive areas such as low-lying areas and riverbanks in the spatial dimension, while capturing the temporal correlation of features from adjacent time phases in the temporal dimension. This dual-dimensional approach synergistically enhances the feature weights of flood-sensitive areas, automatically strengthening the feature responses of key regions while suppressing interfering features from non-flood areas such as moist soil and dense vegetation. This adaptive weighting mechanism overcomes the limitations of fixed threshold methods in misjudging complex features and compensates for the shortcomings of existing technologies that rely solely on spatiotemporal joint attention and cannot capture dynamic flood features, thus reducing false detections caused by low-scattering features outside of water bodies.

[0020] By adaptively suppressing and reconstructing noise in dynamically weighted multi-scale feature maps, feature robustness is further enhanced. It effectively filters inherent speckle noise in SAR images while preserving key information such as boundary features, flooding range features, and temporal variation features. Compared to traditional methods that directly rely on threshold judgments of the original image, feature-level denoising improves the model's tolerance to noise, preventing noise from being misclassified as small areas of water accumulation, and providing a cleaner, more focused low-dimensional feature representation for subsequent classification.

[0021] Multi-scale feature fusion, attention-weighted adjustment, and prior feature calibration are performed on low-dimensional feature representations. Flood classification results are then output through fully connected mapping and probability calculation, generating flood disaster distribution information and achieving precise monitoring and dynamic feedback. Multi-scale feature fusion integrates features from different levels and combines them with temporal variation features, ensuring both detailed accuracy and global consistency in the classification results. Probability calculation achieves precise conversion from features to classification results. The output classification results accurately distinguish between water bodies and non-water bodies while reflecting the dynamic changes of the disaster, solving the problem that traditional methods cannot balance accuracy and dynamic monitoring. Ultimately, this achieves efficient and accurate monitoring of flood disaster conditions in the target area.

[0022] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0023] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0024] Figure 1 A flowchart illustrating a flood disaster monitoring method provided in an embodiment of this application is shown.

[0025] Figure 2 This paper shows a schematic diagram of the structure of a flood disaster monitoring device provided in an embodiment of this application;

[0026] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0027] Exemplary embodiments of the present application will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thorough and complete, and will fully convey the scope of the present application to those skilled in the art.

[0028] In the embodiments of this application, the terms "first," "second," etc., do not have a logical or temporal dependency, nor do they limit the quantity or execution order. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. These terms are merely used to distinguish one element from another.

[0029] In this application, the term "at least one" means one or more, and the term "multiple" means two or more.

[0030] It should also be understood that the term “if” can be interpreted as “when” or “upon”, or “in response to determination” or “in response to detection”. Similarly, depending on the context, the phrase “if determination…” or “if detection [the stated condition or event]” can be interpreted as “when determination…” or “in response to determination…” or “when detection [the stated condition or event]” or “in response to detection [the stated condition or event]”.

[0031] In traditional methods of flood monitoring using SAR images, a common approach is to analyze backscattering coefficients based on a fixed threshold. The steps are roughly as follows: First, preprocessing operations such as radiometric correction are performed on the original SAR image to convert the pixel values ​​into backscattering coefficients with practical physical meaning. Next, given that water bodies often exhibit low backscattering in SAR images, a fixed backscattering coefficient threshold is manually set. Finally, each pixel in the image is compared, and areas with backscattering coefficients less than the threshold are identified as flooded areas, while areas with backscattering coefficients greater than or equal to the threshold are considered non-flooded areas. This method is used to identify and monitor the extent of flooding.

[0032] However, this method has poor anti-interference capabilities and is difficult to meet the monitoring needs in complex scenarios. This is mainly because the scattering characteristics of ground objects in the natural environment are diverse and dynamic. For example, some ground objects that are not flooded, such as bare soil moistened after rain or vigorously growing low vegetation, may have backscattering coefficients close to or even lower than the set threshold, which can easily lead to these ground objects being misclassified as flooded areas. In addition, some special flooded areas, such as water bodies with a large amount of floating debris or riverbeds with rough bottoms in shallow water areas, have altered scattering mechanisms, and their actual backscattering coefficients may be higher than the set threshold, thus being incorrectly excluded from flooded areas. Moreover, factors such as surface roughness in different regions and the incident angle during SAR imaging can significantly affect the backscattering characteristics of ground objects. Fixed thresholds cannot adapt to these changes, which makes the monitoring accuracy of this method fluctuate greatly in different scenarios, making it difficult to meet the high monitoring accuracy requirements in complex flood scenarios.

[0033] Based on this, this application provides a method for monitoring flood disasters, which will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating a flood disaster monitoring method provided in this application. It specifically includes the following steps:

[0034] Step 110: Acquire multi-temporal synthetic aperture radar image data of the target area.

[0035] Regarding the definition of the target area, this application targets flood-sensitive areas, i.e., areas prone to flooding disasters. The delineation of such areas requires comprehensive consideration of factors such as geographical environment, hydrological characteristics, and historical disaster records. Examples include low-lying plains, riverbanks and lakeside areas, urban built-up areas with weak drainage systems, and areas that have frequently experienced flooding in the past. Due to their inherent natural or anthropogenic conditions, these areas have a higher risk of flooding during heavy rainfall, snowmelt, or riverbank breaches, thus becoming the key monitoring targets of this application. By focusing on flood-sensitive areas, the targeting and efficiency of monitoring can be improved, unnecessary full-area data processing can be avoided, and effective coverage of areas with potential disaster risks can be ensured.

[0036] Acquiring multi-temporal SAR image data requires a suitable satellite remote sensing platform. Considering the timeliness and data quality of monitoring, priority should be given to SAR satellites with high spatial resolution and short revisit periods, such as the European Space Agency's Sentinel-1 satellite. This satellite can provide C-band SAR image data with a spatial resolution of several meters to tens of meters, sufficient to capture detailed features of flooded areas. Simultaneously, its short revisit period (e.g., 12 days or 6 days, depending on the observation mode) can meet the dynamic monitoring needs of flood disaster occurrence and development, ensuring the acquisition of multiple sets of image data before, during, and after the disaster. In practice, SAR image data of the target area can be obtained through a satellite data distribution platform (such as the ESA Data Center). When acquiring data, the geographical scope of the target area must be clearly specified (usually defined by latitude and longitude coordinates), and image data covering the area and occurring in different temporal phases should be selected to ensure that the temporal span fully reflects the evolution of the flood disaster. For example, it should include at least a baseline image before the disaster, dynamic images during the disaster, and post-disaster assessment images.

[0037] The acquired multi-temporal SAR image data must include dual-polarization data in both the VV polarization channel (vertical transmission-vertical reception) and the VH polarization channel (vertical transmission-horizontal reception). This combination of polarization channels forms the core data foundation for fully exploring the scattering characteristics of ground objects and improving the accuracy of flood identification. SAR images with different polarization modes can provide complementary ground object scattering information. The VV polarization channel is more sensitive to surface roughness and structural features, exhibiting a significant difference in low scattering response when distinguishing water bodies from rigid ground objects (such as buildings and paved surfaces). The VH polarization channel is more discriminative of vegetation cover and humidity changes, effectively distinguishing flooded water bodies from areas covered by moist soil, reducing misjudgments caused by soil moisture. The collaborative analysis of dual-polarization data can overcome the limitations of single-polarization data, accurately identifying flooded areas in complex scenarios by capturing the scattering differences and collaborative features of the two channels, providing rich raw data support for subsequent adaptive polarization fusion, feature extraction, and other steps.

[0038] Furthermore, the acquired SAR image data must meet certain technical parameter requirements. For imaging modes, interferometric wide swath (IW) or stripe mode (SM) can be selected to balance spatial resolution and coverage. Simultaneously, it is necessary to ensure good consistency in radiometric and geometric aspects of the acquired image data. For example, images from the same satellite and under the same imaging mode should be selected whenever possible to reduce data deviations caused by differences in sensors or imaging conditions, laying the foundation for subsequent preprocessing and feature extraction. For dual-polarization data, the data acquisition parameters (such as imaging incident angle, resolution, and radiometric accuracy) of the VV and VH channels must be consistent to avoid significant differences between the two channels affecting the subsequent polarization feature fusion effect. This ensures that the dual-polarization data can fully exert its synergistic and complementary effects, providing reliable data support for accurate monitoring of flood disasters.

[0039] Step 120: Preprocess the multi-temporal synthetic aperture radar image data to obtain a synthetic aperture radar image that retains the characteristics of flood boundaries.

[0040] Preprocessing multi-temporal SAR image data is a key step to ensure the accuracy of subsequent feature extraction and classification. Its core objective is to remove image noise, correct data bias, and preserve the boundary and detailed features of flooded areas to the greatest extent possible through a series of technical operations, so as to provide high-quality image input for subsequent processing.

[0041] Specifically, the preprocessing mainly includes three key steps: radiometric correction, geometric correction, and speckle noise removal. Radiometric correction is the first step, aiming to convert the raw echo intensity of the SAR image into physically meaningful backscattering coefficients, eliminating radiometric biases caused by factors such as sensor gain, incident angle differences, and atmospheric attenuation. Radiometric correction ensures that SAR images acquired at different times and under different imaging conditions are comparable in intensity values, guaranteeing that subsequent analysis of flood-affected areas is based on a unified physical benchmark. For example, using radiometric calibration parameters provided by the satellite, absolute radiometric correction is performed on the image, converting pixel values ​​into normalized backscattering coefficients, avoiding the impact of sensor-specific characteristics on the extraction of ground features.

[0042] Geometric correction is another crucial preprocessing step, aiming to eliminate geometric distortions in the image and ensure that SAR images accurately correspond to the real geographic coordinate system. This process includes interior orientation correction, exterior orientation correction, and topographic correction: interior orientation correction primarily corrects errors in the sensor's own geometric parameters; exterior orientation correction aligns the image with a reference map using ground control points to ensure the image's geographic location accuracy; and topographic correction, for mountainous and other terrain-complex areas, eliminates image point displacement caused by terrain undulations, ensuring that the positions of ground features in the image match their actual geographic locations. Through geometric correction, multi-temporal SAR images can be precisely matched spatially, guaranteeing the positional correspondence of flood-prone areas in images from different time periods, providing a reliable spatial reference for subsequent analysis of the spatial distribution changes of floods. For example, using professional remote sensing processing software such as ISCE, combined with a digital elevation model (DEM), topographic correction of images ensures the accurate geometric location of flood-sensitive areas such as low-lying areas along rivers.

[0043] Despeculiarization is a crucial preprocessing operation tailored to the characteristics of SAR images. Due to the coherent imaging mechanism, SAR images inevitably contain speckle noise, manifested as randomly distributed bright and dark spots that interfere with the identification of flooded area boundaries and details. Therefore, this application employs an adaptive filtering algorithm for denoising, the core of which is to dynamically adjust the filtering parameters based on the texture complexity and scale characteristics of local image regions. For regions with simple textures and significant noise (such as open water), a stronger filtering intensity is used to effectively suppress noise. For regions with complex textures and containing important boundary information (such as the boundary between flooded and land areas), the filtering intensity is appropriately reduced to avoid blurring boundary features. For example, by combining Lee filtering with edge detection or adaptive Kuan filtering, the region characteristics are determined by calculating statistical measures such as the mean and variance within the local window, and the filtering coefficients are dynamically adjusted. While removing speckle noise, this process preserves the edge contours of flooded areas, subtle water body branches, and other key boundary features to the maximum extent possible, laying the foundation for subsequent multi-scale feature extraction to capture flood boundary information.

[0044] Furthermore, the preprocessing process requires consistency optimization based on the characteristics of multi-temporal data. Since SAR images from different time phases may differ in imaging angle, atmospheric conditions, etc., radiometric normalization is necessary to further reduce radiometric differences between multi-temporal data, ensuring consistent representation of the same ground feature across different time-phase images. For example, stable areas unaffected by flooding (such as buildings and bare land) are selected as references, and the radiometric intensity of images from different time phases is normalized. This ensures that changes in flood-sensitive areas across the multi-temporal sequence are solely due to actual disaster conditions, rather than biases in the data itself. Through these preprocessing steps, the final SAR image not only removes various noises and biases but, more importantly, fully preserves the boundary features of flooded areas and the details of the water-land interface. This provides high-quality input data for subsequent feature extraction and noise suppression using an improved variational autoencoder, ensuring the model can accurately capture the key features of flood-sensitive areas.

[0045] To achieve accurate monitoring of flood-sensitive areas in SAR images, this application adopts a network architecture that integrates an improved variational autoencoder core mechanism. The encoder and decoder work together to enhance feature processing accuracy. The encoder is implemented using a multi-scale pyramid structure, comprising multiple convolutional and pooling layers. Through alternating operations of multiple convolutional and pooling layers with different kernel sizes, local features of the SAR image are extracted layer by layer, and dimensionality reduction is achieved. Simultaneously, an improved multi-layer convolutional neural network structure is employed, leveraging parameter sharing to reduce the number of model parameters and improve computational efficiency. This allows for more efficient capture of complex features in SAR images, ranging from subtle boundaries to global distributions, ultimately mapping the input SAR image to a low-dimensional feature space.

[0046] The decoder uses alternating deconvolution and upsampling layers that match the multi-scale pyramid structure of the encoder. It performs layer-by-layer dimensionality upscaling and feature restoration based on low-dimensional features. During the reconstruction process, nonlinear transformation and regularization are performed simultaneously. Based on the spatial distribution characteristics of the reconstructed features, noise and redundant information are suppressed, the feature response of flood-sensitive areas is enhanced, and the consistency of the reconstructed features with the original image in terms of structure and content is ensured.

[0047] The following steps 130-150 provide a detailed explanation of the collaborative feature processing process between the encoder and decoder in the improved variational autoencoder network.

[0048] Step 130: Extract multi-scale feature maps from the preprocessed synthetic aperture radar image using an adaptive polarization fusion multi-scale pyramid structure.

[0049] After preprocessing the SAR images, including radiometric correction, geometric correction, and adaptive despeccary noise removal, high-quality images with clear flood boundary features are obtained, laying the foundation for subsequent multi-scale feature extraction. This implementation method uses an adaptive polarization fusion multi-scale pyramid structure for feature extraction. The core of this method is to rely on the alternating collaboration of multiple layers of adaptive polarization fusion convolutional units and dynamic pooling layers to fully explore the synergistic information of dual-polarization data and adapt to flood features at different scales.

[0050] The hierarchical architecture of the adaptive polarization fusion multi-scale pyramid structure matches the feature distribution characteristics of the preprocessed SAR image. The overall structure employs a progressively layered network, with each layer consisting of an adaptive polarization fusion convolutional unit and a dynamic pooling layer. These layers are sequentially connected in the order of feature extraction. The output of the dynamic pooling layer in the previous layer directly serves as the input to the adaptive polarization fusion convolutional unit in the next layer, forming a feature extraction chain from low to high dimensions and from local to global. In the initial layer, the input to the adaptive polarization fusion convolutional unit is the preprocessed SAR dual-polarization image. The initial kernel size is preset based on the typical feature scale of flood-sensitive areas. For example, for small- to medium-scale flood boundary features, the initial kernel size can be set to 3×3 to ensure accurate capture of subtle boundary textures.

[0051] The core functionality of the adaptive polarization fusion convolutional unit is implemented through a built-in VV-VH channel interaction module. The core of this module is to dynamically optimize the convolutional kernel weights by quantifying the collaborative characteristics of the dual polarization channels. Specifically, the input VV and VH channel image data are first divided into local regions, and the entire image is traversed using a 3×3 or 5×5 sliding window. For the pixel data within each window, the eigenvalues ​​of the polarization scattering matrices corresponding to the VV and VH channels are calculated. By constructing the scattering matrices of the dual polarization channels, the eigenvalues ​​of the matrices are solved to quantify the difference in scattering intensity and collaborative distribution characteristics of the two channels in this local region. Simultaneously, the texture features of the two channels within each window are extracted. Texture parameters such as contrast, energy, and entropy can be calculated using the gray-level co-occurrence matrix. Then, by calculating the cosine similarity of the corresponding texture parameters of the two channels, a texture consistency parameter is obtained. This parameter measures the degree of matching between the two channels in terms of texture patterns. Subsequently, the eigenvalues ​​of the polarization scattering matrix and the texture consistency parameter are weighted and combined to obtain the synergy coefficient. The weights can be determined adaptively through training data. For example, higher weights are assigned to the texture consistency parameter in flood-prone areas, and higher weights are assigned to the eigenvalues ​​of the polarization scattering matrix in large-scale flooded areas. The magnitude of the synergy coefficient directly reflects the information complementarity value of the dual polarization channels in that local region. Based on this synergy coefficient, the weight allocation of each convolutional kernel in the current convolutional unit is dynamically adjusted. For regions with high synergy coefficients and strong information complementarity, the weight of the corresponding convolutional kernel is increased to enhance feature extraction; for regions with low synergy coefficients and potential noise interference, the weight of the convolutional kernel is decreased to suppress invalid information.

[0052] The iterative update mechanism of the convolutional kernel weights is closely related to the feature response intensity of the previous layer, ensuring the dynamic adaptability of feature extraction. After feature extraction is completed at each level, the response entropy value of the multi-scale feature map output by that level is calculated. The response entropy value is used to quantify the uniformity of the feature response distribution in the feature map. If the response entropy value is greater than a preset threshold, it indicates that the feature distribution extracted at the current level is scattered, and there are many key features that have not been fully captured. The kernel size of the next level's adaptive polarization fusion convolutional unit is increased by a preset stride (e.g., 1 or 2) to expand the feature capture range. If the response entropy value is less than or equal to the preset threshold, it indicates that the feature extraction at the current level is relatively sufficient. The kernel size of the next level remains unchanged or decreases by a preset stride to avoid redundancy caused by over-extraction. For example, if the initial layer convolution kernel size is 3×3, and the response entropy value of its output feature map is higher than the threshold, the size of the next layer convolution kernel is adjusted to 5×5. By gradually increasing the size of the convolution kernel, the gradual capture of the fine texture of the flood boundary to the global features of the large-scale flooded area is achieved, forming a precise fit with the multi-scale flood features retained after the preprocessing in the previous text.

[0053] After each adaptive polarization fusion convolutional unit completes feature extraction, the corresponding dynamic pooling layer immediately starts processing. Its core principle is to adaptively adjust the pooling window size and stride based on the sparsity of the convolutional layer's output feature map, achieving a balance between feature dimensionality reduction and key information preservation. The sparsity of the feature map is calculated by statistically analyzing the proportion of non-zero elements; a lower proportion of non-zero elements indicates more concentrated features and higher sparsity. The window size and stride of the dynamic pooling layer are calculated using preset formulas: window size = round(α × sparsity of the current convolutional layer's output feature map + β), stride = round(γ × window size), where α∈[0.5,1.5], β∈[1,3], and γ∈[0.5,1.0] are preset adjustment coefficients determined through training with a large amount of SAR flood data. For example, when the sparsity of the output feature map of the convolutional layer is 0.3, α is 1.0, β is 2, and γ is 0.8, the window size is round(1.0×0.3+2)=2, and the stride is round(0.8×2)=2. Through this dynamic adjustment mechanism, a smaller window and stride are used in feature-dense areas such as flood boundaries to preserve boundary details to the maximum extent; while a larger window and stride are used in feature-sparse areas such as large-scale flooded areas, which efficiently achieves feature dimensionality reduction and provides multi-scale feature maps with moderate dimensionality and focused information for subsequent dynamic weighting through spatiotemporal joint attention weight maps.

[0054] The entire feature extraction process is closely integrated with the pre- and post-processing steps: the flood boundary features retained after preprocessing provide high-quality input for the adaptive polarization fusion convolutional unit; the VV-VH channel interaction module's in-depth mining of dual-polarization data compensates for the shortcomings of traditional methods in utilizing polarization information; the adaptive adjustment of the dynamic pooling layer ensures the effective extraction of multi-scale features; and the final output multi-scale feature map contains both small-scale flood boundary details and large-scale inundation range features, providing a solid feature foundation for subsequent steps such as spatiotemporal joint attention dynamic weighting, adaptive noise suppression, and reconstruction.

[0055] Step 140: Dynamically weight the multi-scale feature map using a spatiotemporal joint attention weight map to improve the spatiotemporal correlation feature weight of flood-sensitive areas.

[0056] Global average pooling and global max pooling are applied to multi-scale feature maps to extract globally representative information, laying the foundation for the subsequent generation of spatiotemporal joint attention weight maps. Global average pooling calculates the average value of all pixels in each channel of the multi-scale feature map, resulting in a global average pooled feature map. This map reflects the overall intensity level of features in each channel, weakens the interference of local noise, and highlights relatively stable feature information within the region. For example, the average backscattering intensity of the entire flooded area helps capture the macroscopic features of contiguous inundated areas. Global max pooling, on the other hand, generates a global max pooled map by taking the maximum value of all pixels in each channel. It focuses more on the most significant features within a channel, such as strong gradient changes at flood boundaries and small, high-response inundated areas, preserving the extreme value information of key local features. The combination of these two pooling operations allows for the extraction of global features from different perspectives, encompassing both the overall distribution trend of features and not overlooking prominent local features, providing comprehensive input for subsequent global feature fusion.

[0057] By concatenating the global average pooling feature map and the global max pooling feature map, a global feature map is formed by merging the channel dimensions. This operation integrates the complementary information captured by the two pooling methods, allowing the global feature map to simultaneously contain average trends and extreme value features. For example, the concatenated global feature map includes both the overall intensity level information of the flooded area and significant change features at the boundaries, thus more comprehensively reflecting the global flood-related attributes in the multi-scale feature map and providing richer feature basis for the subsequent generation of the weight map.

[0058] Dimensionality reduction and nonlinear transformations of the global feature map are performed to simplify feature complexity and enhance feature expressiveness. Dimensionality reduction is typically achieved through 1×1 convolutional layers, reducing the number of feature channels while fusing information from different channels, thus lowering computational costs. Nonlinear transformations (such as using activation functions like ReLU) introduce nonlinear factors, enabling the model to learn more complex feature mapping relationships and capture hidden nonlinear correlations in global features, such as the complex interaction between flood intensity and boundary features. After dimensionality reduction and nonlinear transformations, further mapping is performed using activation functions (such as the Sigmoid function) to compress feature values ​​to the range of 0 to 1, generating a spatiotemporal joint attention weight map with the same size as the multi-scale feature map. Each pixel value in this weight map represents the importance of the corresponding location in the multi-scale feature map; the closer the value is to 1, the more critical the feature at that location is to flood monitoring. For example, the weight values ​​of areas such as the water body edge and the core of the inundation zone will be significantly higher than those of non-sensitive areas.

[0059] The core operation for achieving dynamic weighting is to multiply the spatiotemporal joint attention weight map pixel-by-pixel with the multi-scale feature map. Through pixel-by-pixel multiplication, features in high-weight regions (such as flood-sensitive areas) of the multi-scale feature map are enhanced, while features in low-weight regions (such as non-submerged buildings and vegetation areas) are suppressed. For example, features at flood boundaries show increased signal strength after multiplication, making them easier to capture in subsequent processing steps, while interfering features in non-target areas are weakened. This allows the multi-scale feature map to focus on spatial regions crucial for flood monitoring, improving the accuracy and efficiency of subsequent feature reconstruction, classification, and other steps. This dynamic weighting process fully utilizes the characteristics of the spatiotemporal joint attention mechanism, enabling the model to adaptively focus on key regions. This aligns perfectly with the design goal of strengthening the feature response of flood-sensitive areas in this application, providing strong support for the accurate extraction of flood-related features.

[0060] Step 150: Perform adaptive noise suppression and reconstruction on the dynamically weighted multi-scale feature map to obtain a low-dimensional feature representation that retains key information of flood-sensitive areas.

[0061] Before performing adaptive noise suppression and reconstruction on the dynamically weighted multi-scale feature map, it is necessary to first deeply mine and fuse the temporal features of multi-temporal data in order to make full use of the flood evolution information contained in SAR images of different time phases. The core of this step is to extract the dynamic change features of flood disaster through temporal correlation analysis and organically combine them with spatial features to form a more comprehensive feature representation.

[0062] Specifically, the multi-scale feature maps, dynamically weighted across different time phases, are arranged chronologically to form a multi-scale feature map sequence. This operation forms the basis for temporal series modeling of multi-temporal data. Here, "time phase" corresponds to SAR images acquired at different points in time, such as the baseline time phase before a flood, multiple monitoring time phases during the disaster, and the assessment time phase after the disaster. By arranging them chronologically, the feature map sequence can intuitively reflect the changes in the target area over time. The multi-scale feature maps of each time phase have been enhanced with features of flood-sensitive areas through a spatiotemporal joint attention mechanism. Therefore, this sequence contains both spatial features at different scales and implicitly contains evolutionary trends in the temporal dimension, providing a structured data foundation for subsequent temporal correlation capture.

[0063] Iterative calculations on multi-scale feature map sequences, performed phase by phase, dynamically adjust the focus on features from different time phases using preset update and reset parameters to accurately capture the temporal correlations between multi-temporal SAR image data. The update parameter controls the proportion of historical temporal feature maps retained in the current iteration. For example, for recent historical features strongly correlated with the current time (such as the inundation area expansion features of the previous time phase), a higher update parameter value is set to retain more information. The reset parameter controls the proportion of historical temporal feature maps forgotten. For historical features with longer time intervals and weaker correlation with the current time (such as features from earlier time phases before the disaster), an appropriate reset parameter value is set to reduce their impact on the current calculation, avoiding redundant information interference. This dynamic allocation of focus allows the model to adaptively focus on temporal information meaningful for flood evolution. For example, during flooding, it prioritizes retaining and utilizing the inundation range change features between adjacent time phases, thereby more accurately capturing the development trend of the disaster, such as the temporal correlations of the expansion speed and direction of the inundated area.

[0064] Based on the aforementioned temporal correlation, a phase-by-phase calculation is performed on the feature map sequences before, during, and after flooding to extract the dynamic changes of flood disasters over time. This process, by comparing the feature differences at different stages, uncovers information directly related to the evolution of the disaster, such as the initial formation characteristics of the inundated area before and during the flood, the continuous expansion characteristics during the flood, and the changes in the inundated area from the flood to the stable period. Specifically, by calculating the differences, ratios, or other statistics between adjacent temporal feature maps, dynamic indicators such as the area change of the inundated area and the speed of boundary movement can be quantified. These indicators can effectively reflect the development stage and severity of flood disasters, providing a temporal basis for subsequent classification and identification.

[0065] Finally, the extracted dynamic change features are fused with the multi-scale feature map obtained through spatiotemporal joint attention processing, channel by channel, to form a multi-scale feature map that integrates spatiotemporal features. Channel-by-channel fusion refers to merging the channels of dynamic change features and spatial features along the feature dimension, so that the fused feature map simultaneously contains multi-scale details in the spatial dimension (such as the boundaries and textures of the flooded area) and dynamic change information in the temporal dimension (such as the expansion trend of the flooded area). This fusion method can fully leverage the complementarity of spatiotemporal features. For example, spatial features accurately locate the spatial distribution of the flooded area, while dynamic change features reveal its evolutionary patterns. After the two are combined, the feature map can more comprehensively describe the characteristics of flood-sensitive areas, providing richer and more discriminative input for subsequent adaptive noise suppression and reconstruction steps, thereby improving the adaptability and recognition accuracy of the entire monitoring method to complex flood scenarios.

[0066] After constructing a multi-scale feature map that integrates spatiotemporal features, feature reconstruction is required to accurately reconstruct information from these integrated features that matches the original image feature structure and focuses on flood-sensitive areas. This process needs to effectively complement the multi-scale feature extraction described earlier. Therefore, alternating operations of deconvolutional layers and upsampling layers, corresponding to the multi-scale pyramid structure, are used to progressively upscale and restore the features of the multi-scale feature map that integrates spatiotemporal features. The core of this approach lies in achieving accurate reconstruction of multi-scale features through a network structure symmetrical to the feature extraction stage.

[0067] The multi-scale pyramid structure extracts features at different scales layer by layer from high to low resolution through alternating operations of convolutional and pooling layers during feature extraction. Corresponding deconvolutional and upsampling layers operate along the opposite path, forming the decoder part of the improved variational autoencoder (VAE) network. The decoder reconstructs the image based on the low-dimensional features output by the encoder, containing multiple deconvolutional and upsampling layers as described above. This ensures the reconstructed image maintains structural and content consistency with the original image. The deconvolutional layer increases the number of channels in the feature map to achieve dimensionality upscaling, restoring the semantic richness of the features. The upsampling layer expands the spatial size of the feature map through interpolation (such as bilinear interpolation), achieving restoration from low to high resolution. This alternating operation strictly matches the scale levels of the feature extraction stage. For example, in the feature extraction stage, three convolutional and pooling layers are used to obtain three feature maps of different scales. In the reconstruction stage, three deconvolutional and upsampling layers are used to restore the feature maps respectively, ensuring that the feature maps reconstructed by each layer are consistent with the corresponding layers in the extraction stage in terms of spatial resolution and semantic level. This achieves accurate reconstruction of the feature structure, enabling the multi-scale feature maps that integrate spatiotemporal features to accurately reflect the spatial distribution and temporal evolution of flood-sensitive areas in the original image after reconstruction.

[0068] Simultaneous nonlinear transformation and regularization during feature reconstruction aim to improve feature quality while restoring the feature structure. Nonlinear transformations (such as activation functions like ReLU and LeakyReLU) introduce nonlinear mapping relationships, enhancing the model's ability to express complex features and making the reconstructed features more closely reflect the nonlinear distribution characteristics of flooded areas, such as distinguishing subtle differences between water bodies and wet ground. Regularization, by standardizing the feature distribution, reduces the risk of overfitting and improves the model's generalization ability. The combination of these two methods, based on the spatial distribution characteristics of the reconstructed features (such as the continuous distribution of flooded areas and gradient changes at boundaries), effectively suppresses noise and redundant information introduced during feature reconstruction (such as random disturbances in non-flooded areas and duplicated invalid features), while strengthening the feature responses of flood-sensitive areas. For example, nonlinear transformations amplify the feature signals of water bodies, while regularization weakens interference from non-target areas, making the reconstructed feature map more clearly highlight key information such as flood boundaries and inundation ranges.

[0069] The reconstructed features after noise suppression are compressed and integrated to generate a low-dimensional feature representation that is both compact and discriminative. Compression and integration are typically achieved through operations such as global pooling and 1×1 convolution, reducing the feature dimensionality to below that of the preprocessed synthetic aperture radar (SAR) image features while preserving key information. This reduces subsequent computational burden and highlights core features. The key information here is strictly focused on flood-sensitive areas, including at least boundary features (such as the contour of the water-land boundary and edge gradient changes), inundation extent features (such as the shape and area proportion of contiguous inundated areas), and temporal variation features (such as the expansion or contraction trend of inundated areas across different time phases). Through compression and integration, multi-scale and multi-temporal feature information is condensed into a low-dimensional vector. This removes redundant non-critical information while ensuring that the retained features comprehensively reflect the spatial distribution and temporal evolution of flood disasters, providing high-quality and targeted input for subsequent classification and identification, ultimately improving the efficiency and accuracy of flood disaster monitoring. This process aligns with the feature extraction objective of improved variational autoencoders, enabling efficient encoding of key information in flood-sensitive areas through low-dimensional feature representations while maintaining the discriminative power and robustness of the features.

[0070] After performing layer-by-layer dimensionality upscaling, feature restoration, noise suppression, and compression integration on multi-scale feature maps that fuse spatiotemporal features to obtain low-dimensional feature representations that retain key information about flood-sensitive areas, fully utilizing the polarization characteristics of SAR images becomes an important optimization direction to further improve the accuracy of feature extraction and reconstruction. Therefore, in the process of feature extraction through a multi-scale pyramid structure and feature reconstruction through deconvolution and upsampling, introducing polarization feature constraints on the VV and VH channels of the SAR image, i.e., designing a polarization consistency constraint term as a component of the improved variational autoencoder loss function, is the key optimization method for improving the accuracy of flood-sensitive area identification in this application. Its core lies in using the complementary information of the dual polarization channels to enhance the discriminativeness of features, while optimizing the loss function to ensure that the model extracts features from flood-sensitive areas more stably and accurately.

[0071] Specifically, the sum of squares of the intensity differences between corresponding pixels in the VV and VH channels of a SAR image is calculated to obtain an intensity difference parameter. This operation aims to quantify the difference in backscattering intensity between the two polarization channels in the same region. The VV (vertical transmit-vertical receive) and VH (vertical transmit-horizontal receive) channels, as common dual polarization modes, exhibit different responses to the scattering characteristics of different ground features. For example, water bodies typically show low backscattering intensity in the VV channel, while the response in the VH channel is more complex, but the intensity variation trends of both should have a certain consistency in flood-prone areas. By calculating the sum of squares of the intensity differences between corresponding pixels, the degree of deviation in intensity between the two channels can be captured, constituting the intensity loss component of the polarization consistency constraint. If a region is flood-prone, the intensity difference between its VV and VH channels should be within a reasonable range; excessive differences may indicate noise interference or feature extraction bias. This parameter provides a quantitative basis for subsequent loss optimization, ensuring that the model can make the intensity characteristics of the two channels more consistent in flood-sensitive areas during feature extraction and reconstruction.

[0072] Extracting texture features from the VV and VH channels and calculating the cosine similarity between them to obtain a texture consistency parameter further constrains the feature consistency of the dual-polarization channels from the texture dimension. Texture features (such as contrast, entropy, and correlation extracted through the Gray-Level Co-occurrence Matrix (GLCM)) reflect the spatial distribution pattern of pixels in an image, which is crucial for distinguishing between flooded areas (typically with relatively uniform texture) and non-flooded areas (such as the rough texture of buildings or the complex texture of vegetation). Calculating the cosine similarity between the texture features of two channels measures the degree of similarity in their texture patterns. Higher similarity indicates more consistent texture features between the two channels in that region, resulting in stronger reliability of feature extraction; conversely, lower similarity may indicate noise or misclassification. For example, the VV and VH channels in a flooded area should exhibit smooth and uniform textures, with a high cosine similarity. Conversely, if the textures of the two channels differ significantly in a certain area, it may be a non-flooded area or an area affected by noise. This parameter, together with the intensity difference parameter, constitutes a comprehensive constraint on the dual-polarization features.

[0073] The intensity difference parameter and texture consistency parameter are weighted and combined to form a polarization consistency constraint term, which is then incorporated into the overall loss calculation of the VAE feature extraction and reconstruction process (forming the loss function together with reconstruction error and KL divergence). The weights of the weighted combination can be dynamically adjusted according to the importance of intensity and texture features in the actual scene. For example, in flood boundary areas, the weight of texture consistency can be appropriately increased to strengthen the constraint on boundary details. Through this structured design, the model not only needs to minimize the feature reconstruction error and KL divergence during training, but also needs to minimize the intensity difference between the two channels and maximize texture consistency through the polarization consistency constraint term. This proactively corrects the bias of the dual polarization channels during feature extraction and reconstruction, improving the sensitivity to features in flooded areas.

[0074] The backpropagation mechanism, which collaboratively optimizes the parameters of convolutional, pooling, deconvolutional, and upsampling layers involved in feature extraction and reconstruction, is a specific means of achieving the aforementioned polarization consistency constraint. Backpropagation calculates the gradient of the loss function containing the polarization consistency constraint term with respect to the parameters of each layer, adjusting the weights of convolutional kernels, pooling windows, and deconvolutional layers layer by layer. This allows the model to continuously reduce the intensity difference parameter (narrowing the intensity difference between the VV and VH channels) and increase the texture consistency parameter (enhancing the consistency of texture features between the two channels) during iterative training. This collaborative optimization ensures that the multi-scale pyramid structure can fully utilize the complementary information of the dual polarization channels when extracting features, reducing misjudgments caused by single-channel noise or characteristic differences. Simultaneously, the deconvolutional and upsampling layers can generate high-quality reconstruction results based on more consistent polarization features when reconstructing features. Ultimately, this allows the extracted low-dimensional feature representation to more accurately focus on flood-sensitive areas, providing a more reliable feature foundation for subsequent classification and recognition. This process aligns closely with the design goals of the improved variational autoencoder in this application, significantly enhancing the model's accuracy and anti-interference capabilities in identifying flood-prone areas through joint optimization of polarization features.

[0075] To ensure that the improved variational autoencoder network can accurately extract key features of flood-sensitive areas in SAR images, it is necessary to optimize the network parameters through scientific training methods. Based on the encoder-decoder architecture design described above, this application adopts an unsupervised learning method during the training phase of the network. By minimizing reconstruction error and KL divergence, the network parameters are optimized, enabling the network to learn the latent distribution of the input SAR image. Combining the characteristics of the encoder-decoder architecture described above, a spatiotemporal joint attention module is implicitly introduced during training. By dynamically weighting the feature importance of different regions (corresponding to the dynamic weighting logic of the spatiotemporal joint attention weight map in step 140) and coordinating with reconstruction error optimization, the suppression effect on speckle noise is further enhanced during feature extraction, while accurately preserving key feature information of flood-sensitive areas (such as boundary features, inundation range features, etc.).

[0076] This unsupervised learning approach does not rely on a large amount of labeled data, which reduces the cost of data preparation and improves the accuracy and stability of feature extraction through feature learning in the latent space. After training, the SAR image preprocessed in step 120 is input into the network, which can efficiently extract the low-dimensional feature representation described in step 150, while simultaneously removing noise. This provides high-quality feature input for the subsequent multi-scale feature fusion and classification tasks in step 160, forming a complete technical chain from feature extraction to classification and monitoring.

[0077] Step 160: Based on historical flood data of the target area, perform feature calibration on the low-dimensional feature representation, perform multi-scale feature fusion and attention weighting adjustment on the calibrated low-dimensional feature representation, and output the flood classification result of the target area by combining fully connected mapping and probability calculation.

[0078] After completing the feature extraction and optimization in step 150, an efficient classification mechanism is needed to transform the processed features into clear flood identification results. Therefore, a convolutional neural network (CNN) is constructed to classify the features extracted in step 150. The CNN constructed in this application, through multi-scale feature fusion and attention mechanisms, can capture complex multi-scale features in SAR images and automatically focus on key parts of flooded areas, thereby improving the accuracy of classification and identification. Specifically, the CNN constructed in this application extracts and fuses features at different scales, enabling the model to capture more layers of flooded area information; simultaneously, the attention mechanism allows the model to automatically adjust its focus on features in different areas, suppressing interference from non-target areas. This network structure design is compatible with the multi-scale low-dimensional features extracted in step 150, fully utilizing the various information contained in the features, from subtle boundaries to macroscopic distributions, preparing for subsequent accurate classification.

[0079] To further improve the fit between features and flood characteristics of the target area, after the model fuses multi-scale features and generates a preliminary low-dimensional feature representation, feature calibration is performed on the low-dimensional feature representation based on historical flood data of the target area. The specific implementation is as follows: First, systematically acquire multi-temporal synthetic aperture radar image data, field survey data, and hydrological monitoring data of historical floods in the target area in recent years. Perform multi-dimensional analysis on these data to accurately extract typical features of historical flood events in terms of polarization scattering, texture structure, and spatiotemporal evolution, thereby constructing a priori feature library covering various flood scenarios. This feature library serves as an important reference for subsequent calibration. Next, use a feature matching algorithm to calculate the matching degree between the low-dimensional feature representation and similar scene features in the priori feature library. The calculation of the matching degree needs to comprehensively consider factors such as polarization feature consistency, texture similarity, and spatiotemporal distribution fit. Finally, according to the formula... Calibrate the low-dimensional feature representation, where F calibrated For the calibrated low-dimensional feature representation, F original For low-dimensional feature representation, F prior θ represents the prior feature with the highest matching degree in the prior feature library, and θ is the weight coefficient after the matching degree is normalized. This formula realizes the adaptive fusion and calibration of the original low-dimensional features and prior features, so that the feature representation is more in line with the historical flood characteristics of the target area, and provides more targeted feature support for the accurate classification of subsequent convolutional neural networks.

[0080] Specifically, the convolutional neural network constructed in this application processes low-dimensional feature representations through multiple feature extraction stages. This aims to further extract flood-related information at different scales from the compressed and integrated key features, laying the foundation for subsequent accurate classification. Each feature extraction stage employs a combination of convolutional and pooling layers of varying depths. Here, "depth" refers to the number of stacked convolutional layers; different depths imply different levels of feature abstraction. Shallower convolutional layer combinations retain more detailed information from low-dimensional features, making them suitable for extracting small-scale features, such as subtle boundaries of flooded areas and scattered small inundated areas. Deeper convolutional layer combinations, through more convolutional operations, progressively abstract features, making them better at capturing large-scale features, such as the overall distribution of contiguous inundated areas and their macroscopic correlation with surrounding terrain. Through this multi-stage design, the feature maps output by each stage correspond to flood features at different scales, forming a multi-scale feature map sequence. This ensures that the feature coverage of flood-sensitive areas includes both local details and global distribution.

[0081] By fusing multi-scale feature maps from different stages through lateral connections and upsampling operations, the aim is to break down the isolation between features at different scales and achieve information complementarity and enhancement. Lateral connections stitch feature maps from different stages together along the channel dimension, allowing detailed information (such as boundary textures) in shallow features to be fused with semantic information (such as flood zone categories) in deep features. Upsampling operations scale up deep feature maps (e.g., through bilinear interpolation or deconvolution) to match their spatial resolution with that of shallow feature maps, ensuring that the fused feature maps maintain spatial consistency. This fusion method can simultaneously capture subtle flood boundary features and macroscopic flood zone distribution features. For example, combining boundary gradient information from shallow features with regional distribution information from deep features can accurately locate the edges of flooded areas and clarify their overall extent, providing more comprehensive input for subsequent feature adjustments.

[0082] Calculating the global mean or global maximum value of each channel in the fused multi-scale feature map to obtain a measure of the importance of the feature channels is to quantify the contribution of different feature channels to flood identification. Each feature channel corresponds to a specific feature pattern (e.g., one channel may focus on the texture features of water bodies, while another channel may reflect the intensity features of inundated areas). The global mean reflects the overall activation level of the features in that channel, while the global maximum highlights the most significant feature response in that channel. These two statistical measures can be used to determine the importance of each channel in distinguishing between flooded and non-flooded areas. For example, channels sensitive to water body features typically have higher global means or peak values ​​in flooded areas, resulting in higher importance measures, while channels sensitive to noise or non-target features have lower measures.

[0083] Weights are assigned to different feature channels based on importance metrics. Then, convolutional operations are used to extract local correlation information from the fused multi-scale feature map to generate spatial response weights. This process achieves precise feature adjustment through a combination of channel attention and spatiotemporal joint attention. Assigning higher weights to high-importance channels strengthens their influence on classification results while suppressing interference from low-importance channels. Convolutional operations capture pixel correlations within local regions (such as pixel gradient changes at flood boundaries) to generate spatial response weights, enabling the model to focus on key spatial locations for flood identification (such as water body edges and the core of inundated areas). Through the synergistic effect of these two weights, the features of flood-sensitive areas are significantly enhanced in the fused multi-scale feature map, while features of non-target areas (such as buildings and vegetation) are effectively suppressed, further improving feature discriminative power.

[0084] After completing the aforementioned feature adjustments, the constructed convolutional neural network needs to be trained to better learn and distinguish the feature patterns of flooded and non-flooded areas. This is done by training the convolutional neural network with a small amount of labeled SAR image data. This labeled data is preprocessed and includes samples from both flooded and non-flooded areas, providing a clear learning target for the convolutional neural network.

[0085] To improve training efficiency and model performance, this application employs the cross-entropy loss function for network training. This loss function effectively measures the difference between the model's predictions and the actual annotations, guiding the network parameters towards optimization in the correct direction. Simultaneously, data augmentation techniques (such as rotation, translation, and scaling) are used to expand the labeled dataset. Considering the scarcity of labeled data often encountered in flood disaster monitoring, this application adopts a hybrid strategy combining traditional data augmentation with physical model-based synthetic data augmentation.

[0086] Traditional data augmentation expands the diversity of datasets by performing geometric transformations such as rotation, translation, scaling, and mirroring on existing labeled samples, simulating flood scene features from different observation angles and scales. These operations allow the model to access a more diverse distribution of features.

[0087] Physical model-based synthetic data augmentation further overcomes the limitations of the number of real samples. It generates virtual flood scenarios with varying inundation levels and terrain conditions through fluid dynamics simulations. These are then combined with SAR imaging physical models (such as radar backscattering characteristic models) to generate corresponding synthetic SAR images. Furthermore, speckle noise distributions (e.g., conforming to Gamma or K distributions) statistically obtained from real SAR data are overlaid to construct a large-scale labeled pre-training sample. For example, based on the digital elevation model (DEM) and hydrological parameters of the target area, the accumulation of floodwater in low-lying areas and its flow in river channels are simulated to generate inundation range data at different time phases. Then, based on the radar scattering characteristics of different land features (water bodies, vegetation, buildings), the inundation range data is converted into SAR image features with realistic radiation characteristics, making the synthetic data closer to the actual observation scenario.

[0088] This hybrid data augmentation strategy not only preserves the feature distribution of the actual scene using limited real-world labeled data, but also significantly expands the sample size through synthetic data, covering more extreme flood conditions and terrain combinations. This effectively improves the generalization ability of the convolutional neural network and its efficiency in utilizing scarce labeled data, enabling it to maintain good classification performance even when faced with complex real-world SAR images. The optimized network, after training, can process feature-adjusted input information more accurately.

[0089] The adjusted features are then input into a fully connected layer for integration and mapping, efficiently compressing high-dimensional features into low-dimensional vectors directly relevant to the classification task. The fully connected layer linearly combines and transforms the features using a weight matrix, extracting the most crucial information for classification. The integrated features are mapped to feature vectors corresponding to the classification categories, which at least include flooded and non-flooded areas, ensuring a clear distinction between the disaster status of the target area.

[0090] Furthermore, the low-dimensional features extracted in step 150 are input into the trained convolutional neural network for classification and recognition. At this point, the input low-dimensional features have undergone key information extraction and mapping through fully connected layers, achieving a precise match with the input dimension of the trained convolutional neural network. The trained convolutional neural network, relying on the previous optimization of multi-scale feature fusion and attention mechanisms, can quickly locate feature patterns in flood-sensitive areas from low-dimensional features. It identifies both small-scale edge details and captures large-scale regional distributions, deriving classification logic through hierarchical feature matching, and finally outputting preliminary classification results, providing structured input for subsequent probability calculations.

[0091] Specifically, by using an activation function (such as the Softmax function) to calculate the probability values ​​of each category and form a probability distribution, the feature vector is transformed into an interpretable probability output. The probability value of each category reflects the confidence that the feature belongs to the corresponding category. Based on this probability distribution, the category of each pixel or image patch in the preprocessed SAR image is determined as the category with the highest probability value, thereby generating a classification result that includes the division between flooded and non-flooded areas. This process considers both the global statistical characteristics of the features and local details, and the final output classification result can accurately reflect the actual disaster situation in flood-sensitive areas, providing a direct basis for subsequent generation of flood disaster distribution maps and statistical disaster information, which is highly consistent with the goal of this application to quickly and accurately monitor flood disasters.

[0092] In the process of multi-scale feature fusion and attention-weighted adjustment of low-dimensional feature representations, and outputting flood classification results for the target region through fully connected mapping and probability calculation, it is necessary to specifically optimize the loss calculation method to further improve classification accuracy, especially to enhance the recognition effect of the key area of ​​flood boundaries. The premise of achieving this optimization is to accurately locate the boundary between flooded and non-flooded areas. Therefore, identifying the boundary pixels between flooded and non-flooded areas in the preprocessed synthetic aperture radar (SAR) image through edge detection algorithm is to accurately locate the intersection of the two types of areas, providing a basis for the subsequent differentiated allocation of loss weights.

[0093] The preprocessed SAR images retain clear flood boundary features. Edge detection algorithms (such as Canny edge detection and the Sobel operator) can identify gradient changes at flood boundaries by recognizing abrupt changes in pixel grayscale values, thereby determining the location of boundary pixels. These boundary pixels are crucial for distinguishing flooded and non-flooded areas, and their accuracy directly affects the determination of the inundation range. For example, inundation boundaries along riverbanks and the edges of water accumulation in low-lying urban areas must be accurately extracted using edge detection algorithms to ensure that subsequent weight allocation focuses on these regions that are critical to the classification results.

[0094] The core idea is to assign loss weights to pixels in different regions based on the recognition results, thereby enhancing the model's focus on learning about flood boundary areas through differentiated weight settings. Since pixels in flood boundary areas are more prone to misclassification (e.g., misclassifying shallow water areas at the edge as non-flooded areas, or vice versa), pixels in these areas are assigned higher loss weights than pixels in non-boundary areas. Specifically, the loss weight for boundary pixels can be set several times (e.g., 2-5 times) that of non-boundary pixels, making the model more sensitive to classification errors of boundary pixels during training. This weight allocation method guides the model to focus more on learning the feature differences of boundary areas, such as distinguishing the differences in backscattering characteristics between water bodies and damp ground at the boundary, thus reducing classification bias caused by boundary ambiguity.

[0095] Integrating this dynamic weight into the loss calculation during feature fusion and classification processes allows for targeted optimization of boundary regions through the design of the loss function. In the feature fusion stage, the dynamic weight influences the fusion weights of features at different scales, giving boundary features greater attention during fusion. In the classification stage, the dynamic weight amplifies the loss caused by classification errors in boundary pixels by adjusting the calculation of classification losses such as cross-entropy loss. Through backpropagation, this loss signal containing the dynamic weight is passed layer by layer to parameters (such as convolutional kernel weights, attention weight coefficients, and fully connected layer weights) in the multi-scale feature fusion, attention weighting adjustment, and fully connected layer mapping processes, driving targeted optimization of these parameters. For example, adjusting the convolutional kernel enhances the ability to capture boundary textures, optimizing the attention mechanism makes the feature responses in boundary regions more significant, or correcting the mapping relationship of fully connected layers to improve the classification accuracy of boundary pixels.

[0096] Ultimately, through this series of operations, the classification accuracy of flood boundary areas was effectively improved, enabling a more accurate definition of the spatial boundaries of inundation areas and avoiding errors in flood area estimation due to boundary misjudgment. This process echoes the overall design of this application, which strengthens the characteristic response of flood-sensitive areas. By focusing on the fine optimization of boundary areas, the accuracy and reliability of flood disaster monitoring results were further improved, providing crucial support for the subsequent generation of high-precision flood disaster distribution maps and statistical analysis of inundation parameters.

[0097] Step 170: Generate flood disaster distribution information based on flood classification results to monitor flood disaster conditions in the target area.

[0098] Generating flood disaster distribution information based on flood classification results to monitor flood disaster conditions in target areas is the final output of the entire methodology. Its core lies in transforming the abstract features obtained from classification and identification into intuitive and usable disaster information, providing decision support for disaster emergency response and disaster reduction and relief.

[0099] Specifically, the flood classification results clearly distinguish between flooded and non-flooded areas within the target region (flood-sensitive area). The process of generating flood disaster distribution information begins with visualizing these classification results. Using specialized SAR image processing software, flooded and non-flooded areas in the classification results are marked with different colors or symbols. For example, a blue gradient is used to represent flooded areas, with varying shades corresponding to differences in inundation depth. Non-flooded areas are marked with gray or other neutral colors, allowing the distribution information to intuitively reflect the spatial location and extent of the inundated areas. Simultaneously, the generated distribution map needs to be georegistered. Combined with the geometric correction results completed in the preprocessing stage, this ensures the distribution map is accurately aligned with the real geographic coordinate system, including latitude, longitude, scale, and other geographic information, so that relevant departments can accurately locate disaster-stricken areas, such as specific towns and river sections.

[0100] Based on the generated distribution map, statistical analysis of key parameters of flood disasters is also required. These parameters include, but are not limited to, the extent, area, and inundation depth of flooded regions. Area calculation is performed by converting the number of pixels representing flooded areas in the statistical classification results into the actual geographical area, combined with the spatial resolution of the image (such as the meter-level resolution of the Sentinel-1 satellite), ensuring data accuracy. Inundation depth estimation can be achieved by combining the backscatter intensity characteristics of SAR images, historical hydrological data, or topographic data (such as digital elevation models, DEMs), establishing a mapping relationship between intensity values ​​and inundation depth. For example, areas with lower backscatter intensity typically correspond to deeper inundation depths. This statistical information, presented in tabular or text form, can quantify the severity of the disaster, providing data support for assessing disaster losses and allocating relief resources.

[0101] The flood disaster distribution information generated through the above process can comprehensively reflect the flood disaster situation in the target area: the distribution map intuitively shows the spatial distribution and diffusion trend of the inundated area, while the statistical parameters quantify the scope and degree of the disaster's impact. Based on this information, dynamic monitoring of flood-sensitive areas can be achieved. For example, by comparing distribution information generated at different time phases, the expansion or contraction of the inundated area can be analyzed to determine the development stage of the disaster (such as continuous aggravation, stabilization, or gradual receding); at the same time, combined with the geographical characteristics of the target area (such as population density and infrastructure distribution), the potential risks posed by the disaster can be assessed, such as the degree of threat to power facilities and residential areas. This monitoring not only covers areas that have already experienced flooding but also provides early warning of potential risks in flood-sensitive areas through historical data and trend analysis, achieving full-process support from disaster identification to risk assessment, and ultimately providing a scientific and accurate basis for disaster emergency response decisions, fully leveraging the technical advantages of this application in rapid and efficient flood disaster monitoring.

[0102] It is understood that, in order to achieve the functions in the above embodiments, the computer device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and method steps described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.

[0103] Furthermore, as a response to the above Figure 1 The implementation of the method embodiment shown in this application provides a flood disaster monitoring device. This device embodiment corresponds to the foregoing method embodiments. For ease of reading, this embodiment will not repeat the details of the foregoing method embodiments one by one, but it should be clear that the device in this embodiment can correspondingly implement all the contents of the foregoing method embodiments. Specifically, as shown... Figure 2 As shown, the flood disaster monitoring device 200 includes:

[0104] The acquisition module 210 is used to acquire multi-temporal synthetic aperture radar image data of the target area, which is a flood-sensitive area. The multi-temporal synthetic aperture radar image data includes dual polarization data of VV polarization channel and VH polarization channel.

[0105] Preprocessing module 220 is used to preprocess multi-temporal synthetic aperture radar image data to obtain synthetic aperture radar images that retain flood boundary features;

[0106] The feature extraction module 230 is used to extract multi-scale feature maps from the preprocessed synthetic aperture radar image through an adaptive polarization fusion multi-scale pyramid structure. The adaptive polarization fusion multi-scale pyramid structure is implemented through the alternating operation of multiple layers of adaptive polarization fusion convolutional units and dynamic pooling layers. The adaptive polarization fusion convolutional unit has a built-in VV-VH channel interaction module. The VV-VH channel interaction module dynamically adjusts the convolution kernel weights by calculating the eigenvalues ​​of the polarization scattering matrices of the VV and VH channels and the co-correlation coefficients of the texture consistency parameters, and iteratively updates the convolution kernel weights according to the feature response intensity of the previous layer. The dynamic pooling layer adaptively adjusts the pooling window size and stride according to the feature sparsity of the corresponding convolutional layer to achieve feature dimensionality reduction and key information preservation. The multi-scale feature map is dynamically weighted through a spatiotemporal joint attention weight map to improve the spatiotemporal correlation feature weights of flood-sensitive areas. The dynamically weighted multi-scale feature map is then subjected to adaptive noise suppression and reconstruction to obtain a low-dimensional feature representation that retains key information of flood-sensitive areas.

[0107] The classification module 240 is used to perform feature calibration on the low-dimensional feature representation based on historical flood data of the target area, perform multi-scale feature fusion and attention weighting adjustment on the calibrated low-dimensional feature representation, and output the flood classification result of the target area by combining fully connected mapping and probability calculation.

[0108] The generation module 250 is used to generate flood disaster distribution information based on flood classification results in order to monitor the flood disaster situation in the target area.

[0109] Optionally, the flood disaster monitoring device may be an electronic device with data processing capabilities, or a functional module within such electronic device; there is no limitation on this.

[0110] For example, the electronic device can be a server, which can be a single server or a server cluster consisting of multiple servers. As another example, the electronic device can be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, as well as cellular phone, personal digital assistant (PDA), augmented reality (AR), virtual reality (VR) device, and other terminal devices. Furthermore, the electronic device can also be a recording device, video surveillance device, etc. This application does not impose any special limitations on the specific form of the electronic device.

[0111] The following example uses electronic devices for monitoring flood disasters, such as... Figure 3 As shown, Figure 3 The hardware structure of an electronic device 300 provided in this application.

[0112] like Figure 3 As shown, the electronic device 300 includes a processor 310, a communication line 320, and a communication interface 330.

[0113] Optionally, the electronic device 300 may also include a memory 340. The processor 310, memory 340, and communication interface 330 can be connected via a communication line 320.

[0114] The processor 310 can be a central processing unit (CPU), a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 310 can also be any other device with processing capabilities, such as a circuit, device, or software module, without limitation.

[0115] In one example, processor 310 may include one or more CPUs, for example Figure 3 CPU0 and CPU1 in the CPU.

[0116] As an optional implementation, the electronic device 300 may include multiple processors, for example, in addition to processor 310, it may also include processor 370. A communication line 320 is used to transmit information between the components included in the electronic device 300.

[0117] Communication interface 330 is used for communication with other devices or other communication networks. These other communication networks can be Ethernet, Radio Access Network (RAN), Wireless Local Area Networks (WLAN), etc. Communication interface 330 can be a module, circuit, transceiver, or any device capable of enabling communication.

[0118] The memory 340 is used to store instructions. These instructions can be computer programs.

[0119] The memory 340 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and / or instructions; it may also be a random access memory (RAM) or other type of dynamic storage device capable of storing information and / or instructions; it may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, etc., without limitation.

[0120] It should be noted that the memory 340 can exist independently of the processor 310, or it can be integrated with the processor 310. The memory 340 can be used to store instructions, program code, or some data, etc. The memory 340 can be located inside or outside the electronic device 300, without restriction.

[0121] The processor 310 is configured to execute instructions stored in the memory 340 to implement the communication method provided in the following embodiments of this application. For example, when the electronic device 300 is a terminal or a chip in a terminal, the processor 310 can execute instructions stored in the memory 340 to implement the steps performed by the sending end in the following embodiments of this application.

[0122] As an optional implementation, the electronic device 300 also includes an output device 350 and an input device 360. The output device 350 can be a display screen, speaker, or other device capable of outputting data from the electronic device 300 to the user. The input device 360 ​​can be a keyboard, mouse, microphone, joystick, or other device capable of inputting data into the electronic device 300.

[0123] It should be pointed out that, Figure 3 The structure shown does not constitute a limitation on the electronic device, except... Figure 3 In addition to the components shown, the electronic device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0124] The flood disaster monitoring device and application scenarios described in this application are for the purpose of more clearly illustrating the technical solutions of this application, and do not constitute a limitation on the technical solutions provided in this application. As those skilled in the art will know, with the evolution of flood disaster monitoring devices and the emergence of new business scenarios, the technical solutions provided in this application are also applicable to similar technical problems.

[0125] This application provides a storage medium storing a program that, when executed by a processor, implements the flood disaster monitoring method.

[0126] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0127] In a typical configuration, the device includes one or more processors (CPUs), memory, and a bus. The device may also include input / output interfaces, network interfaces, etc.

[0128] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM, and memory includes at least one memory chip. Memory is an example of computer-readable media.

[0129] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0130] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0131] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0132] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for monitoring flood disaster conditions, characterized in that, The method includes: Acquire multi-temporal synthetic aperture radar image data of a target area, wherein the target area is a flood-sensitive area, and the multi-temporal synthetic aperture radar image data includes dual polarization data of VV polarization channel and VH polarization channel; The multi-temporal synthetic aperture radar image data is preprocessed to obtain synthetic aperture radar images that retain flood boundary features; Multi-scale feature map extraction is performed on preprocessed synthetic aperture radar images using an adaptive polarization fusion multi-scale pyramid structure. This structure is achieved through alternating operations of multiple adaptive polarization fusion convolutional units and dynamic pooling layers. Each adaptive polarization fusion convolutional unit incorporates a VV-VH channel interaction module. This module dynamically adjusts the convolutional kernel weights by calculating the eigenvalues ​​of the polarization scattering matrices of the VV and VH channels and the co-operation coefficients of the texture consistency parameters. Furthermore, it iteratively updates the kernel weights based on the feature response intensity of the previous layer. The dynamic pooling layer adaptively adjusts the pooling window size and stride according to the feature sparsity of the corresponding convolutional layer, achieving feature dimensionality reduction and key information preservation. The multi-scale feature map is dynamically weighted by a spatiotemporal joint attention weight map to improve the spatiotemporal correlation feature weight of flood-sensitive areas; Adaptive noise suppression and reconstruction are performed on the dynamically weighted multi-scale feature map to obtain a low-dimensional feature representation that retains key information of flood-sensitive areas. Based on historical flood data of the target area, the low-dimensional feature representation is calibrated, and the calibrated low-dimensional feature representation is subjected to multi-scale feature fusion and attention weighting adjustment. The flood classification result of the target area is output by combining fully connected mapping and probability calculation. Based on the flood classification results, flood disaster distribution information is generated to monitor the flood disaster situation in the target area; The adaptive polarization fusion multi-scale pyramid structure includes multiple adaptive polarization fusion convolutional units. Each adaptive polarization fusion convolutional unit is followed by a dynamic pooling layer. Each adaptive polarization fusion convolutional unit and its corresponding dynamic pooling layer form a hierarchical unit. The hierarchical units are connected sequentially according to the order of feature extraction. The output of the dynamic pooling layer of the previous hierarchical unit is used as the input of the convolutional layer of the next hierarchical unit. The initial size of the convolutional kernels within the same level unit is the same, and iterative updates are achieved in the following way: calculate the response entropy value of the multi-scale feature map of the current level. If the response entropy value is greater than a preset threshold, the size of the convolutional kernel of the next level increases by a preset step size. If the response entropy value is less than or equal to the preset threshold, the size of the convolutional kernel of the next level remains unchanged or decreases by a preset step size. The window size and stride of the dynamic pooling layer satisfy the following: window size = round(α × sparsity of the current convolutional layer output feature map + β), stride = round(γ × window size), where α, β, and γ are preset adjustment coefficients, and round is the rounding function.

2. The method according to claim 1, characterized in that, The multi-scale feature map is dynamically weighted using a spatiotemporal joint attention weight map, including: Global average pooling and global max pooling are performed on the multi-scale feature map to obtain a global average pooling feature map and a global max pooling feature map. The temporal correlation coefficient between adjacent temporal multi-scale feature maps is calculated by using the cosine similarity and Euclidean distance of the two temporal multi-scale feature maps. The global average pooling feature map, the global max pooling feature map, and the temporal correlation coefficient are fused through a channel-by-channel weighted fusion to obtain a spatiotemporal joint global feature map. The spatiotemporal joint global feature map is subjected to dimensionality reduction and nonlinear transformation, and then mapped by an activation function to generate a spatiotemporal joint attention weight map with the same size as the multi-scale feature map. The spatiotemporal joint attention weight map is multiplied pixel by pixel with the multi-scale feature map to achieve dynamic weighting of the multi-scale feature map.

3. The method according to claim 2, characterized in that, Before performing adaptive noise suppression and reconstruction on the dynamically weighted multi-scale feature map, the method further includes: The multi-scale feature maps, dynamically weighted for each time phase, are arranged in chronological order to form a multi-scale feature map sequence. The multi-scale feature map sequence is subjected to time-phase iterative calculation. The retention ratio of historical time-phase feature maps is adjusted by preset update parameters, and the forgetting ratio of historical time-phase feature maps is controlled by preset reset parameters. In this way, the attention to different time-phase feature maps is dynamically allocated, and the temporal correlation between multi-time-phase synthetic aperture radar image data is captured. Based on the aforementioned temporal correlation, the feature map sequences before, during, and after the flood are calculated on a time-by-time basis to extract the dynamic change features of the flood disaster in the time dimension. The extracted dynamic change features are then stitched together channel by channel with the multi-scale feature map obtained through spatiotemporal joint attention processing to form a multi-scale feature map that integrates spatiotemporal features.

4. The method according to claim 3, characterized in that, Adaptive noise suppression and reconstruction are performed on the dynamically weighted multi-scale feature maps to obtain low-dimensional feature representations that retain key information of flood-sensitive areas, including: By alternating operations of deconvolutional layers and upsampling layers corresponding to the multi-scale pyramid structure, the multi-scale feature map fused with spatiotemporal features is upgraded and its features are restored layer by layer. The feature structure is reconstructed by matching the scale level of the feature extraction stage. Nonlinear transformation and regularization are performed simultaneously during feature restoration. Based on the spatial distribution characteristics of the reconstructed features, noise and redundant information are suppressed, and the feature response of flood-sensitive areas is enhanced. The reconstructed features after noise suppression are compressed and integrated to form a low-dimensional feature representation with a dimension lower than that of the preprocessed synthetic aperture radar image features and focusing on flood-sensitive areas.

5. The method according to claim 4, characterized in that, The method further includes: In the process of feature extraction through adaptive polarization fusion multi-scale pyramid structure and feature reconstruction through adaptive deconvolution and upsampling, the polarization scattering matrix feature values ​​of the corresponding regions of the synthetic aperture radar image VV channel and VH channel are calculated to obtain polarization feature parameters. The texture features of VV channel and VH channel are extracted, and the cosine similarity between the texture features of the two channels is calculated to obtain texture consistency parameters. The polarization feature parameters and the texture consistency parameters are weighted and combined, and then incorporated into the loss calculation of the feature extraction and reconstruction process. The parameters of the convolutional layers, dynamic pooling layers, adaptive deconvolutional layers, and upsampling layers involved in the feature extraction and reconstruction process are optimized in a coordinated manner through the backpropagation mechanism to enhance the coordination and consistency of the dual-polarization channel features.

6. The method according to claim 5, characterized in that, Feature calibration is performed on the low-dimensional feature representation based on historical flood data of the target area, including: Acquire multi-temporal synthetic aperture radar image data, field survey data, and hydrological monitoring data of historical floods in the target area in recent years, extract typical features of historical flood events, and construct a priori feature library; Calculate the matching degree between the low-dimensional feature representation and the features of the same type of scene in the prior feature library; according to The low-dimensional feature representation is calibrated, where F calibrated For the calibrated low-dimensional feature representation, F original For the low-dimensional feature representation, F prior θ represents the prior feature with the highest matching degree in the prior feature library, and θ is the weight coefficient after the matching degree is normalized.

7. The method according to claim 6, characterized in that, Multi-scale feature fusion and attention-weighted adjustment are performed on the calibrated low-dimensional feature representation. Combined with fully connected mapping and probability calculation, the flood classification results for the target region are output, including: The calibrated low-dimensional feature representation is processed through three feature extraction stages. The first stage uses a combination of shallow convolutional layers and dynamic pooling layers to extract small-scale flood boundary features. The second stage uses a combination of medium-level convolutional layers and dynamic pooling layers to extract medium-scale inundation range features. The third stage uses a combination of deep convolutional layers and dynamic pooling layers to extract large-scale flood distribution features. The small-scale flood boundary features, the medium-scale inundation range features, and the large-scale flood distribution features are fused through lateral connection and upsampling operations; The importance measure of the feature channels is obtained by calculating the weighted sum of the global mean and global maximum of each channel in the fused multi-scale feature map. Weights are assigned to different feature channels based on the importance metric, and local correlation information of the fused multi-scale feature map is extracted through convolution operation to generate spatial response weights. The spatial response weights are then used to adjust the features at different spatial locations in the fused multi-scale feature map to enhance the features of flood-sensitive areas. The adjusted features are input into a fully connected layer for integration and mapping. The integrated features output by the fully connected layer are mapped into feature vectors corresponding to the classification categories, which include at least flooded areas and non-flooded areas. The probability values ​​of each category are calculated by the activation function and a probability distribution is formed. Based on the probability distribution, the category of each pixel or image patch in the preprocessed synthetic aperture radar image is determined, and a classification result containing the division between flooded and non-flooded areas is generated.

8. The method according to claim 7, characterized in that, The method further includes: The boundary pixels between flooded and non-flooded areas in the preprocessed synthetic aperture radar image are identified using an edge detection algorithm. Based on the recognition results, loss weights are assigned to pixels in different regions, and pixels in flood boundary regions are given higher loss weights than pixels in non-boundary regions. The dynamic weight is incorporated into the loss calculation of the feature fusion and classification process. The parameters in the multi-scale feature fusion, attention weighting adjustment and fully connected mapping process are optimized through backpropagation to improve the classification accuracy of flood boundary areas.

9. A flood disaster monitoring device, characterized in that, The device includes: The acquisition module is used to acquire multi-temporal synthetic aperture radar image data of a target area, wherein the target area is a flood-sensitive area, and the multi-temporal synthetic aperture radar image data includes dual polarization data of VV polarization channel and VH polarization channel; The preprocessing module is used to preprocess the multi-temporal synthetic aperture radar image data to obtain a synthetic aperture radar image that retains the flood boundary features; The feature extraction module is used to extract multi-scale feature maps from preprocessed synthetic aperture radar images using an adaptive polarization fusion multi-scale pyramid structure. This structure is implemented through alternating operations of multiple adaptive polarization fusion convolutional units and dynamic pooling layers. Each convolutional unit incorporates a VV-VH channel interaction module, which dynamically adjusts the convolutional kernel weights by calculating the eigenvalues ​​of the polarization scattering matrices and the texture consistency parameter of the VV and VH channels. The kernel weights are iteratively updated based on the feature response intensity of the previous layer. The dynamic pooling layer adaptively adjusts the pooling window size and stride according to the feature sparsity of the corresponding convolutional layer, achieving feature dimensionality reduction and key information preservation. The multi-scale feature maps are dynamically weighted using a spatiotemporal joint attention weight map to enhance the spatiotemporal correlation feature weights of flood-sensitive areas. Adaptive noise suppression and reconstruction are then applied to the dynamically weighted multi-scale feature maps to obtain a low-dimensional feature representation that retains key information about flood-sensitive areas. The classification module is used to perform feature calibration on the low-dimensional feature representation based on historical flood data of the target area, perform multi-scale feature fusion and attention weighting adjustment on the calibrated low-dimensional feature representation, and output the flood classification result of the target area by combining fully connected mapping and probability calculation. A generation module is used to generate flood disaster distribution information based on the flood classification results, so as to monitor the flood disaster situation in the target area; The adaptive polarization fusion multi-scale pyramid structure includes multiple adaptive polarization fusion convolutional units. Each adaptive polarization fusion convolutional unit is followed by a dynamic pooling layer. Each adaptive polarization fusion convolutional unit and its corresponding dynamic pooling layer form a hierarchical unit. The hierarchical units are connected sequentially according to the order of feature extraction. The output of the dynamic pooling layer of the previous hierarchical unit is used as the input of the convolutional layer of the next hierarchical unit. The initial size of the convolutional kernels within the same level unit is the same, and iterative updates are achieved in the following way: calculate the response entropy value of the multi-scale feature map of the current level. If the response entropy value is greater than a preset threshold, the size of the convolutional kernel of the next level increases by a preset step size. If the response entropy value is less than or equal to the preset threshold, the size of the convolutional kernel of the next level remains unchanged or decreases by a preset step size. The window size and stride of the dynamic pooling layer satisfy the following: window size = round(α × sparsity of the current convolutional layer output feature map + β), stride = round(γ × window size), where α, β, and γ are preset adjustment coefficients, and round is the rounding function.

10. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, it controls the device containing the storage medium to perform the flood disaster monitoring method as described in any one of claims 1-8.

11. An electronic device, characterized in that, The device includes at least one processor, at least one memory connected to the processor, and a bus; wherein the processor and the memory communicate with each other through the bus; the processor is used to call program instructions in the memory to execute the flood disaster monitoring method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Deep learning and SAR image-based flood water body dynamic monitoring and evidence storage method

    CN118865244A

  • Multi-scale synthetic aperture radar flood detection method and device

    CN120997563A