Flood disaster range monitoring method and system based on satellite remote sensing image
By preprocessing high-resolution SAR satellite remote sensing images and using a multimodal flood segmentation model, combined with real-time inference and dynamic monitoring technologies, the problems of insufficient resolution and real-time performance in flood disaster monitoring have been solved, enabling high-precision and rapid identification and early warning of flood disaster extent.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- STATE GRID SIJI DIGITAL TECH (BEIJING) CO LTD
- Filing Date
- 2025-11-26
- Publication Date
- 2026-04-24
AI Technical Summary
Existing flood disaster monitoring technologies suffer from insufficient spatial resolution, susceptibility to topography and buildings, low real-time performance and automation levels due to reliance on multi-source data fusion, and a lack of dynamic monitoring capabilities.
High-resolution SAR satellite remote sensing images are preprocessed to construct a multimodal flood segmentation model. Combining the focal loss function and the dice loss function, the TensorRT inference engine is used to monitor the extent of flood disasters in real time, and the ConvLSTM network is used for dynamic monitoring and early warning.
It achieves high-precision and rapid monitoring of flood disaster range, improves the accuracy of identifying inundation boundaries in small water bodies and complex scenarios, and provides near real-time dynamic monitoring and early warning capabilities, making it suitable for large-scale and long-term disaster monitoring.
Smart Images

Figure CN121921668A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural disaster monitoring and artificial intelligence, and in particular to a method and system for monitoring the extent of flood disasters based on satellite remote sensing imagery. Background Technology
[0002] Floods are among the most frequent and destructive natural disasters globally, affecting a wide area and severely threatening people's lives and property while also causing long-term negative impacts on the ecological environment and socio-economic development. Statistics show that floods cause billions of dollars in direct economic losses worldwide each year, often accompanied by secondary disasters such as soil erosion, water quality deterioration, and the spread of infectious diseases. Therefore, achieving high-precision and efficient monitoring of flood inundation areas is crucial for early disaster warning, emergency command, disaster assessment, and water resource management.
[0003] Traditional flood monitoring relies primarily on ground-based observation methods, such as water level stations, current meters, and manual inspections. While these methods are reliable, they suffer from high deployment costs, limited coverage, and low efficiency, making them particularly unsuitable for large-scale flood-prone areas or remote regions. With the development of remote sensing technology, satellite imagery has gradually become an important data source for flood monitoring; however, its practical application still faces many challenges.
[0004] In satellite remote sensing monitoring, early methods primarily employed optical remote sensing imagery (such as Landsat and Sentinel-2), utilizing the Normalized Differential Water Index (NDWI) and the Modified Normalized Differential Water Index (MNDWI) for water body identification. While optical imagery is intuitive and easy to interpret, it is susceptible to interference from weather conditions such as clouds, rain, and fog. During severe weather events like floods, timely acquisition of effective data is often impossible, leading to insufficient monitoring timeliness. Furthermore, low-to-medium resolution optical imagery (such as Sentinel-2 with a spatial resolution of 10 meters) struggles to accurately capture the inundation boundaries of narrow river channels or urban areas, and shadows and building obstructions in complex terrain further impact identification accuracy.
[0005] To overcome the limitations of optical remote sensing, Synthetic Aperture Radar (SAR), with its advantages of all-day, all-weather imaging, has gradually become the main technical means for flood monitoring. SAR actively transmits and receives microwave signals, enabling it to penetrate clouds and some vegetation for stable Earth observation. C-band SAR satellites, represented by Sentinel-1, have been widely used for flood identification, often employing threshold segmentation methods (such as the OTSU algorithm) or machine learning models (such as support vector machines) to distinguish between water bodies and land. However, existing SAR technology still has significant shortcomings when monitoring narrow inland water bodies: on the one hand, C-band SAR has a low spatial resolution (e.g., 10 meters for Sentinel-1), limiting its ability to identify small tributaries and inundation details; on the other hand, its longer wavelength (approximately 5.6 cm) is less sensitive to minor water surface movements and is easily affected by wind and waves, leading to noise and errors in the extracted results.
[0006] In recent years, the development of high-resolution SAR satellites has brought new opportunities for flood monitoring. For example, the GF-3 satellite has improved the accuracy of water body identification through algorithm optimization (such as the Q-OTSU thresholding method), but in complex scenarios such as urban and mountainous areas, it is still difficult to achieve fine-grained monitoring due to overlapping and shadowing effects. Further technological breakthroughs come from Ka-band SAR, which has a shorter wavelength (about 8 mm) and higher spatial resolution and the ability to depict ground features in detail. Launched in 2023, Luojia2-01, as the world's first high-resolution Ka-band SAR satellite, has a resolution of up to 1 meter, bringing significant progress to inland water monitoring. This satellite can clearly capture water body boundaries and texture features under adverse weather conditions, significantly improving the ability to identify flood inundation areas. Research shows that extracting texture features based on Luojia2 images (such as the Gray-Level Co-occurrence Matrix (GLCM)) and combining it with a Support Vector Machine (SVM) classification model can effectively improve recognition accuracy. With the aid of optical image correction, local details can also be optimized. Nevertheless, existing methods still have the following problems: First, Ka-band SAR is susceptible to geometric distortion in steep mountainous areas and urban areas, leading to misidentification; second, texture feature extraction algorithms have high computational complexity, low efficiency in large-scale data processing, and rely on synchronous optical images for correction, making it difficult to meet the timeliness requirements of emergency response to sudden disasters; in addition, existing technologies mostly focus on static range extraction and lack the ability to monitor the dynamic evolution of flood processes over time.
[0007] In summary, although the current flood disaster range monitoring technology continues to develop, the following key issues still need to be addressed: (1) Insufficient spatial resolution of SAR data leads to the omission of small water bodies or blurred boundaries; (2) The identification algorithm has weak anti-interference ability and is easily affected by terrain, buildings and noise; (3) It relies heavily on multi-source data fusion, and the real-time performance and automation level are not high; (4) There is a lack of dynamic monitoring schemes that can take into account high accuracy.
[0008] Therefore, how to provide a method and system for monitoring the extent of flood disasters based on satellite remote sensing imagery, so as to improve the accuracy and timeliness of flood disaster monitoring, has become an urgent technical problem to be solved. Summary of the Invention
[0009] The technical problem to be solved by the present invention is to provide a method and system for monitoring the extent of flood disasters based on satellite remote sensing imagery, so as to improve the accuracy and timeliness of monitoring the extent of flood disasters.
[0010] In a first aspect, the present invention provides a method for monitoring the extent of flood disasters based on satellite remote sensing imagery, comprising the following steps: Step S1: The server acquires a large number of historical SAR satellite remote sensing images with spatial resolution higher than a preset resolution threshold. The historical SAR satellite remote sensing images are preprocessed, including at least radiometric calibration, filtering and removal, georegistration and terrain correction. The preprocessed historical SAR satellite remote sensing images are labeled to construct a dataset. Step S2: The server creates a flood segmentation model based on a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer, and sets the joint loss function of the flood segmentation model; the joint loss function combines the focus loss function and the dice loss function; Step S3: The server trains the flood segmentation model using the dataset. During the training process, the model parameters of the flood segmentation model are optimized using the joint loss function. The trained flood segmentation model is then deployed. Step S4: The server acquires target SAR satellite remote sensing images of the target area before and during the flood event, with a resolution higher than the resolution threshold. After preprocessing each target SAR satellite remote sensing image, it inputs the deployed flood segmentation model for inference to obtain a water distribution map. Step S5: The server performs a difference operation on the adjacent water body distribution maps to obtain the newly added flood disaster area; Step S6: The server generates a time-series flood inundation range map of the target area based on each of the flood disaster ranges, and calculates flood characteristic parameters including inundation area, inundation depth and inundation volume. Based on the time-series flood inundation range map and flood characteristic parameters, the server performs dynamic monitoring and early warning of the flood disaster range.
[0011] Furthermore, step S1 specifically includes: The server acquires a large number of historical SAR satellite remote sensing images with a spatial resolution higher than a preset resolution threshold; the resolution threshold is 1 meter; the historical SAR satellite remote sensing images carry at least backscatter intensity, metadata, polarization information, spatial and geometric information; the metadata includes at least imaging geometric parameters, time information, location information, data processing level, and polarization mode; the spatial and geometric information includes at least spatial resolution and geographic coordinates; Each of the aforementioned historical SAR satellite remote sensing images is subjected to preprocessing including at least radiometric calibration, filtering and removal, georegistration, and terrain correction. The radiometric calibration specifically involves: extracting the pixel amplitude values DN from historical SAR satellite remote sensing images; calculating the backscattering coefficient σ° based on the amplitude values DN and the calibration constant K; converting the unit of the backscattering coefficient σ° to decibels to obtain the decibel value σ°_dB of the backscattering coefficient σ°, thus completing the radiometric calibration. σ° = DN² / K; σ°_dB = 10 * log10(σ°); The filtering and removal process employs the Lee Filter algorithm. The georegistment specifically involves georegistring historical SAR satellite remote sensing images using orbital ephemeris data and digital elevation models. The terrain correction specifically involves: reconstructing the three-dimensional geometric relationship of historical SAR satellite remote sensing images using orbital ephemeris data and digital elevation models, identifying and correcting distortions to perform geometric terrain correction, calculating the local incident angle of each pixel in the geometrically corrected historical SAR satellite remote sensing images, and normalizing the brightness of the local incident angle using the cosine correction method to perform radiation terrain correction. The preprocessed historical SAR satellite remote sensing images are labeled with at least water body pixels, non-water body pixels, and easily confused regions to construct a dataset; the water body pixels are permanent water body pixels or flooded water body pixels; the easily confused regions include at least wind-blown water surfaces, vegetation-covered water areas, and flat non-water surfaces.
[0012] Furthermore, step S2 specifically includes: The server creates a flood segmentation model based on a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer. The joint loss function of the flood segmentation model is set, which combines the focus loss function and the dice loss function. The multimodal feature extraction layer is constructed based on the SAR image feature extraction module, the metadata encoding module, and the polarization feature enhancement module; The SAR image feature extraction module is used to extract multi-scale features from SAR satellite remote sensing images through residual networks and dilated convolutions, and integrate the multi-scale features through a feature pyramid to obtain multi-scale image features; the metadata encoding module is used to map metadata in SAR satellite remote sensing images into high-dimensional features through a multilayer perceptron, and capture the dependencies between the high-dimensional features through a self-attention mechanism to obtain metadata features; the polarization feature enhancement module is used to perform channel compression and enhancement of polarization information in SAR satellite remote sensing images through 1x1 convolutions, and focus on water-sensitive polarization channels through a polarization attention mechanism to obtain polarization enhancement features; The modality fusion layer is constructed based on a spatial alignment module and a dynamic weight fusion module; The spatial alignment module performs spatial alignment operations on multi-scale image features, metadata features, and polarization enhancement features through deformable convolution to correct residual geometric distortions and expands the metadata features to the spatial dimension through a broadcast operation. The dynamic weight fusion module calculates the correlation weights between the multi-scale image features, metadata features, and polarization enhancement features output by the spatial alignment module through a cross-attention mechanism, and dynamically fuses the multi-scale image features, metadata features, and polarization enhancement features according to the correlation weights through a gated recurrent unit to obtain fused features. The feature enhancement layer is constructed based on a self-supervised contrastive learning module and an enhanced terrain attention module; The self-supervised contrastive learning module is used to construct positive and negative sample pairs through data augmentation, and to enhance the water body discrimination of the fused features using a contrastive loss function to obtain first-level enhanced features; the enhanced terrain attention module is used to calculate attention weights through geometric information, and to integrate geometric information into the first-level enhanced features based on the attention weights through conditional batch normalization to obtain second-level enhanced features. The segmentation prediction layer is constructed based on a multi-scale segmentation decoding module and an uncertainty estimation module; The multi-scale segmentation and decoding module is used to upsample and skip connections the secondary enhancement features through the U-Net decoder structure, and output a high-resolution water distribution map by combining the multi-scale image features. The uncertainty estimation module is used to calculate the prediction variance of the water distribution map as uncertainty by sampling multiple times during inference through Monte Carlo dropout. After filtering the uncertainty region in the water distribution map by the uncertainty and the preset uncertainty threshold, the module outputs a water distribution map carrying the uncertainty. The formula for the joint loss function is: ; ; ; in, Indicates the value of the joint loss function; This represents the value of the focus loss function; This represents the value of the dice loss function; This represents the total number of pixels processed in one training iteration; i represents the pixel index, ranging from 1 to N, used to iterate through each pixel; This represents the true category label of the i-th pixel, with a value of 0 or 1; Indicate category Weighting coefficients; This indicates that the i-th pixel predicted by the flood segmentation model belongs to... The probability; γ represents the focusing parameter, a real hyperparameter greater than or equal to 0, used to adjust the weights of easy and difficult samples; This represents the true category label of the i-th pixel; ϵ represents the predicted probability of the i-th pixel; ϵ represents the smoothing constant. This represents twice the area of the intersection between the actual water body pixels and the predicted water body pixels, used to measure the degree of overlap. This represents the sum of the actual total number of water body pixels and the predicted probability of water body pixels, used for normalization.
[0013] Furthermore, step S3 specifically includes: The server divides the dataset into a training set, a validation set, and a test set based on a spatiotemporal partitioning method, such that the time of the validation set is newer than the time of the training set, the time of the test set is newer than the time of the validation set, and the training set, validation set, and test set come from geographically non-overlapping regions. Before training begins, a tree-structured Parzen estimator is used to search for hyperparameters on the training and validation sets to obtain the optimal combination of hyperparameters; the hyperparameters include at least the learning rate, batch size, network depth, network width, Dropout ratio, decay step size, and decay rate. The flood segmentation model is trained on the training set using the hyperparameter combination. During training, the training loss value is calculated using the joint loss function, and the gradient of the training loss value with respect to the model parameters of the flood segmentation model is calculated using the backpropagation algorithm. The model parameters are updated based on the gradient, and early stopping is monitored using the validation set until a preset early stopping condition is triggered to complete the training. The minimum model parameters include: the kernel weights and biases of the convolutional layers in the residual network of the SAR image feature extraction module, and the parameters of the dilated convolution kernels of the dilated convolution; the weights and biases of the fully connected layers in the multilayer perceptron of the metadata encoding module, and the weights of the query matrix, key matrix, and value matrix in the self-attention mechanism; the kernel weights and biases of the 1x1 convolutional layers in the polarization feature enhancement module, and the attention weight parameters in the polarization attention mechanism; the kernel weights and biases of the deformable convolution in the spatial alignment module, and the parameters of the convolutional layers that learn the offset; the projection matrix in the cross-attention mechanism of the dynamic weight fusion module, and the weight matrix and bias vector in the gated recurrent unit; the parameters of the attention weights used to calculate geometric information in the enhanced terrain attention module, and the scaling and offset parameters in the conditional batch normalization; and the kernel weights and biases of all upsampled convolution kernels in the U-Net decoder of the multi-scale segmentation decoding module. Precision, recall, F1-Score, intersection-over-union ratio, overall accuracy, and inference speed are calculated using the test set to test the trained flood segmentation model. The tested flood segmentation model is then deployed using the TensorRT inference engine.
[0014] Furthermore, step S4 specifically includes: Based on the received flood disaster monitoring instructions, the server obtains target SAR satellite remote sensing images of the target area with a resolution higher than the resolution threshold from the remote sensing image database before and during the flood event. After performing preprocessing on each target SAR satellite remote sensing image in sequence through the streaming computing engine, including at least radiometric calibration, filtering and removal, georegistration and terrain correction, the server inputs each target SAR satellite remote sensing image into the deployed flood segmentation model and performs real-time inference through the TensorRT inference engine to obtain the water distribution map. Step S5 specifically involves: The server performs differential operations on the time-adjacent water body distribution maps, combines morphological opening operations for noise reduction, and combines morphological closing operations for edge refinement to obtain the newly added flood disaster range. Step S6 specifically involves: The server generates a time-series flood inundation range map of the target area based on the flood disaster ranges, and calculates flood characteristic parameters including inundation area, inundation depth, and inundation volume; the inundation area is calculated based on the time-series flood inundation range map and spatial resolution; the inundation depth is calculated based on the time-series flood inundation range map and digital elevation model; the inundation volume is calculated based on the inundation area and inundation depth. The server inputs the time-series flood inundation range map into a pre-trained ConvLSTM network to obtain a predicted flood inundation range map. The time-series flood inundation range map, the predicted flood inundation range map, and flood characteristic parameters are displayed in real time on a large screen. The server monitors the time-series flood inundation range map, the predicted flood inundation range map, and the flood characteristic parameters through a preset monitoring threshold group. When the monitoring threshold group is triggered, a flood warning notification is immediately pushed to the management terminal, thereby dynamically monitoring and warning of the flood disaster range.
[0015] Secondly, the present invention provides a flood disaster range monitoring system based on satellite remote sensing imagery, comprising the following modules: The dataset construction module is used to acquire a large number of historical SAR satellite remote sensing images with spatial resolution higher than a preset resolution threshold. The module performs preprocessing on each of the historical SAR satellite remote sensing images, including at least radiometric calibration, filtering and removal, georegistration and terrain correction. The preprocessed historical SAR satellite remote sensing images are then labeled to construct the dataset. The flood segmentation model creation module is used by the server to create a flood segmentation model based on a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer, and to set the joint loss function of the flood segmentation model; the joint loss function combines the focus loss function and the dice loss function; The flood segmentation model training module is used by the server to train the flood segmentation model using the dataset. During the training process, the model parameters of the flood segmentation model are optimized through the joint loss function, and the trained flood segmentation model is deployed. The flood segmentation model inference module is used by the server to acquire target SAR satellite remote sensing images of the target area before and during a flood event, with a resolution higher than the resolution threshold. After preprocessing each target SAR satellite remote sensing image, the image is input into the deployed flood segmentation model for inference to obtain a water distribution map. The differential operation module is used by the server to perform differential operations on the distribution maps of adjacent water bodies to obtain the newly added flood disaster range; The flood disaster range monitoring module is used by the server to generate a time-series flood inundation range map of the target area based on each of the flood disaster ranges, and to calculate flood characteristic parameters including inundation area, inundation depth and inundation volume. Based on the time-series flood inundation range map and flood characteristic parameters, dynamic monitoring and early warning of the flood disaster range are carried out.
[0016] Furthermore, the dataset construction module is specifically used for: The server acquires a large number of historical SAR satellite remote sensing images with a spatial resolution higher than a preset resolution threshold; the resolution threshold is 1 meter; the historical SAR satellite remote sensing images carry at least backscatter intensity, metadata, polarization information, spatial and geometric information; the metadata includes at least imaging geometric parameters, time information, location information, data processing level, and polarization mode; the spatial and geometric information includes at least spatial resolution and geographic coordinates; Each of the aforementioned historical SAR satellite remote sensing images is subjected to preprocessing including at least radiometric calibration, filtering and removal, georegistration, and terrain correction. The radiometric calibration specifically involves: extracting the pixel amplitude values DN from historical SAR satellite remote sensing images; calculating the backscattering coefficient σ° based on the amplitude values DN and the calibration constant K; converting the unit of the backscattering coefficient σ° to decibels to obtain the decibel value σ°_dB of the backscattering coefficient σ°, thus completing the radiometric calibration. σ° = DN² / K; σ°_dB = 10 * log10(σ°); The filtering and removal process employs the Lee Filter algorithm. The georegistment specifically involves georegistring historical SAR satellite remote sensing images using orbital ephemeris data and digital elevation models. The terrain correction specifically involves: reconstructing the three-dimensional geometric relationship of historical SAR satellite remote sensing images using orbital ephemeris data and digital elevation models, identifying and correcting distortions to perform geometric terrain correction, calculating the local incident angle of each pixel in the geometrically corrected historical SAR satellite remote sensing images, and normalizing the brightness of the local incident angle using the cosine correction method to perform radiation terrain correction. The preprocessed historical SAR satellite remote sensing images are labeled with at least water body pixels, non-water body pixels, and easily confused regions to construct a dataset; the water body pixels are permanent water body pixels or flooded water body pixels; the easily confused regions include at least wind-blown water surfaces, vegetation-covered water areas, and flat non-water surfaces.
[0017] Furthermore, the flood segmentation model creation module is specifically used for: The server creates a flood segmentation model based on a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer. The joint loss function of the flood segmentation model is set, which combines the focus loss function and the dice loss function. The multimodal feature extraction layer is constructed based on the SAR image feature extraction module, the metadata encoding module, and the polarization feature enhancement module; The SAR image feature extraction module is used to extract multi-scale features from SAR satellite remote sensing images through residual networks and dilated convolutions, and integrate the multi-scale features through a feature pyramid to obtain multi-scale image features; the metadata encoding module is used to map metadata in SAR satellite remote sensing images into high-dimensional features through a multilayer perceptron, and capture the dependencies between the high-dimensional features through a self-attention mechanism to obtain metadata features; the polarization feature enhancement module is used to perform channel compression and enhancement of polarization information in SAR satellite remote sensing images through 1x1 convolutions, and focus on water-sensitive polarization channels through a polarization attention mechanism to obtain polarization enhancement features; The modality fusion layer is constructed based on a spatial alignment module and a dynamic weight fusion module; The spatial alignment module performs spatial alignment operations on multi-scale image features, metadata features, and polarization enhancement features through deformable convolution to correct residual geometric distortions and expands the metadata features to the spatial dimension through a broadcast operation. The dynamic weight fusion module calculates the correlation weights between the multi-scale image features, metadata features, and polarization enhancement features output by the spatial alignment module through a cross-attention mechanism, and dynamically fuses the multi-scale image features, metadata features, and polarization enhancement features according to the correlation weights through a gated recurrent unit to obtain fused features. The feature enhancement layer is constructed based on a self-supervised contrastive learning module and an enhanced terrain attention module; The self-supervised contrastive learning module is used to construct positive and negative sample pairs through data augmentation, and to enhance the water body discrimination of the fused features using a contrastive loss function to obtain first-level enhanced features; the enhanced terrain attention module is used to calculate attention weights through geometric information, and to integrate geometric information into the first-level enhanced features based on the attention weights through conditional batch normalization to obtain second-level enhanced features. The segmentation prediction layer is constructed based on a multi-scale segmentation decoding module and an uncertainty estimation module; The multi-scale segmentation and decoding module is used to upsample and skip connections the secondary enhancement features through the U-Net decoder structure, and output a high-resolution water distribution map by combining the multi-scale image features. The uncertainty estimation module is used to calculate the prediction variance of the water distribution map as uncertainty by sampling multiple times during inference through Monte Carlo dropout. After filtering the uncertainty region in the water distribution map by the uncertainty and the preset uncertainty threshold, the module outputs a water distribution map carrying the uncertainty. The formula for the joint loss function is: ; ; ; in, Indicates the value of the joint loss function; This represents the value of the focus loss function; This represents the value of the dice loss function; This represents the total number of pixels processed in one training iteration; i represents the pixel index, ranging from 1 to N, used to iterate through each pixel; This represents the true category label of the i-th pixel, with a value of 0 or 1; Indicate category Weighting coefficients; This indicates that the i-th pixel predicted by the flood segmentation model belongs to... The probability; γ represents the focusing parameter, a real hyperparameter greater than or equal to 0, used to adjust the weights of easy and difficult samples; This represents the true category label of the i-th pixel; ϵ represents the predicted probability of the i-th pixel; ϵ represents the smoothing constant. This represents twice the area of the intersection between the actual water body pixels and the predicted water body pixels, used to measure the degree of overlap. This represents the sum of the actual total number of water body pixels and the predicted probability of water body pixels, used for normalization.
[0018] Furthermore, the flood segmentation model training module is specifically used for: The server divides the dataset into a training set, a validation set, and a test set based on a spatiotemporal partitioning method, such that the time of the validation set is newer than the time of the training set, the time of the test set is newer than the time of the validation set, and the training set, validation set, and test set come from geographically non-overlapping regions. Before training begins, a tree-structured Parzen estimator is used to search for hyperparameters on the training and validation sets to obtain the optimal combination of hyperparameters; the hyperparameters include at least the learning rate, batch size, network depth, network width, Dropout ratio, decay step size, and decay rate. The flood segmentation model is trained on the training set using the hyperparameter combination. During training, the training loss value is calculated using the joint loss function, and the gradient of the training loss value with respect to the model parameters of the flood segmentation model is calculated using the backpropagation algorithm. The model parameters are updated based on the gradient, and early stopping is monitored using the validation set until a preset early stopping condition is triggered to complete the training. The minimum model parameters include: the kernel weights and biases of the convolutional layers in the residual network of the SAR image feature extraction module, and the parameters of the dilated convolution kernels of the dilated convolution; the weights and biases of the fully connected layers in the multilayer perceptron of the metadata encoding module, and the weights of the query matrix, key matrix, and value matrix in the self-attention mechanism; the kernel weights and biases of the 1x1 convolutional layers in the polarization feature enhancement module, and the attention weight parameters in the polarization attention mechanism; the kernel weights and biases of the deformable convolution in the spatial alignment module, and the parameters of the convolutional layers that learn the offset; the projection matrix in the cross-attention mechanism of the dynamic weight fusion module, and the weight matrix and bias vector in the gated recurrent unit; the parameters of the attention weights used to calculate geometric information in the enhanced terrain attention module, and the scaling and offset parameters in the conditional batch normalization; and the kernel weights and biases of all upsampled convolution kernels in the U-Net decoder of the multi-scale segmentation decoding module. Precision, recall, F1-Score, intersection-over-union ratio, overall accuracy, and inference speed are calculated using the test set to test the trained flood segmentation model. The tested flood segmentation model is then deployed using the TensorRT inference engine.
[0019] Furthermore, the flood segmentation model inference module is specifically used for: Based on the received flood disaster monitoring instructions, the server obtains target SAR satellite remote sensing images of the target area with a resolution higher than the resolution threshold from the remote sensing image database before and during the flood event. After performing preprocessing on each target SAR satellite remote sensing image in sequence through the streaming computing engine, including at least radiometric calibration, filtering and removal, georegistration and terrain correction, the server inputs each target SAR satellite remote sensing image into the deployed flood segmentation model and performs real-time inference through the TensorRT inference engine to obtain the water distribution map. The difference operation module is specifically used for: The server performs differential operations on the time-adjacent water body distribution maps, combines morphological opening operations for noise reduction, and combines morphological closing operations for edge refinement to obtain the newly added flood disaster range. The flood disaster range monitoring module is specifically used for: The server generates a time-series flood inundation range map of the target area based on the flood disaster ranges, and calculates flood characteristic parameters including inundation area, inundation depth, and inundation volume; the inundation area is calculated based on the time-series flood inundation range map and spatial resolution; the inundation depth is calculated based on the time-series flood inundation range map and digital elevation model; the inundation volume is calculated based on the inundation area and inundation depth. The server inputs the time-series flood inundation range map into a pre-trained ConvLSTM network to obtain a predicted flood inundation range map. The time-series flood inundation range map, the predicted flood inundation range map, and flood characteristic parameters are displayed in real time on a large screen. The server monitors the time-series flood inundation range map, the predicted flood inundation range map, and the flood characteristic parameters through a preset monitoring threshold group. When the monitoring threshold group is triggered, a flood warning notification is immediately pushed to the management terminal, thereby dynamically monitoring and warning of the flood disaster range.
[0020] The advantages of this invention are: 1. Acquire a large number of historical SAR satellite remote sensing images with spatial resolution higher than a preset resolution threshold through a server. Perform preprocessing on each historical SAR satellite remote sensing image, including at least radiometric calibration, filtering and removal, georegistration, and terrain correction. Label the preprocessed historical SAR satellite remote sensing images to construct a dataset. Then, create a flood segmentation model based on a multimodal feature extraction layer, modality fusion layer, feature enhancement layer, and segmentation prediction layer. Set the joint loss function of the flood segmentation model. Train the flood segmentation model using the dataset. During the training process, optimize the model parameters of the flood segmentation model through the joint loss function. Deploy the trained flood segmentation model. Next, acquire target SAR satellite remote sensing images of the target area with resolution higher than the resolution threshold before and during flood events. After preprocessing each target SAR satellite remote sensing image, input them into the flood segmentation model for inference to obtain... The system generates a water body distribution map, performs differential operations on adjacent water body distribution maps to obtain newly added flood disaster areas, and generates a time-series flood inundation range map of the target area based on each flood disaster range. It also calculates flood characteristic parameters, including inundation area, inundation depth, and inundation volume. Based on the time-series flood inundation range map and flood characteristic parameters, dynamic monitoring and early warning of flood disaster areas are conducted. In terms of accuracy, leveraging the detail advantages of high-resolution SAR satellite remote sensing imagery and the powerful multimodal feature extraction and anti-interference capabilities of the flood segmentation model, it can accurately identify inundation boundaries in small water bodies and complex scenarios. In terms of timeliness, by constructing a fully automated analysis process based on a single SAR satellite remote sensing image, it eliminates the reliance on multi-source data fusion, achieving rapid processing from data input to result output. Furthermore, by utilizing differential operations and time-series analysis of pre- and post-disaster images, it ultimately greatly improves the accuracy and timeliness of flood disaster area monitoring.
[0021] 2. By using SAR satellite remote sensing imagery with a spatial resolution higher than 1 meter as the data source, and combining preprocessing methods such as radiometric calibration, filtering and removal, georegistration, and terrain correction, the accuracy and quality of the data are improved from the source. On this basis, an innovative multimodal flood segmentation model is constructed. This flood segmentation model enhances the ability to distinguish features of water bodies and complex terrain scenes by fusing SAR image features, metadata, and polarization information, and by using attention mechanisms and self-supervised contrastive learning. Combined with a joint loss function optimized for boundaries and an uncertainty estimation module, the accuracy and anti-interference ability of flood range identification are significantly improved, and the omission of small water bodies and boundary ambiguity are effectively reduced. In terms of timeliness, the preprocessing and model inference are accelerated by the TensorRT inference engine and streaming computing engine, achieving near real-time processing. The newly added inundation range is quickly extracted by temporal difference analysis, and prediction is performed using a ConvLSTM network. Finally, an automated process from real-time monitoring to dynamic early warning is formed, thus achieving a dual improvement in the accuracy and timeliness of flood disaster range monitoring.
[0022] 3. Utilizing SAR (Synthetic Aperture Radar) satellite remote sensing imagery for flood monitoring offers advantages such as all-weather, all-time imaging, unaffected by cloud cover or weather conditions, ensuring the continuity and reliability of data acquisition. Preprocessing (such as radiometric calibration, filtering, georegistration, and terrain correction) effectively improves the quality of image data, reduces noise and geometric distortion, and provides a high-quality dataset for subsequent model training. This systematic data processing workflow not only improves efficiency but also reduces the need for manual intervention, making it suitable for large-scale, long-term disaster monitoring applications.
[0023] 4. The flood segmentation model integrates a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer. This multi-layered structure can capture complex features in SAR satellite remote sensing images and improve segmentation accuracy. At the same time, the joint loss function combines the focus loss function and the dice loss function, which effectively solves the common class imbalance problem in flood monitoring (such as the imbalance of the ratio of water body to non-water body pixels), and enhances the robustness and generalization ability of the model. This design not only improves monitoring accuracy but also reduces the risk of overfitting, making the model more adaptable to actual disaster scenarios.
[0024] 5. The flood segmentation model is trained using a dataset, and the model parameters are optimized using a joint loss function, achieving an end-to-end automated learning process. The trained flood segmentation model can be deployed directly without frequent manual adjustments, greatly improving operational efficiency. This automated process lowers the professional threshold, making it easy for non-expert users to apply. At the same time, it supports continuous iterative optimization of the model, adapting to different geographical regions and climatic conditions, enhancing the scalability and practicality of the technology.
[0025] 6. By acquiring SAR satellite remote sensing images of the target area before and after the flood event, and using a trained flood segmentation model for inference and differential calculation, the newly added flood disaster area can be directly identified. This method avoids the subjectivity of relying on manual interpretation in traditional monitoring and improves the accuracy and speed of change detection. Differential calculation can accurately distinguish between permanent water bodies and temporary floods, reduce false alarms, and provide timely and reliable data support for disaster emergency response.
[0026] 7. By generating time-series flood inundation range maps and calculating flood characteristic parameters such as inundation area, inundation depth, and inundation volume, dynamic monitoring and trend analysis of flood disasters are realized. This time-series analysis capability allows users to track the evolution of disasters and supports early warning and disaster assessment. Combined with parameterized output, it not only provides spatial range information but also quantifies disaster intensity, providing a scientific basis for disaster prevention and mitigation decisions (such as evacuation planning and resource allocation), and improving the comprehensiveness and foresight of overall disaster management.
[0027] 8. By acquiring historical SAR satellite remote sensing images with a spatial resolution higher than 1 meter, this high-resolution data can capture more subtle surface changes and significantly improve the accuracy of flood disaster range identification. Compared with traditional low-resolution images, it can more accurately distinguish water body boundaries and minute features, reducing the risk of missed or false detections, thus providing more reliable data support in disaster emergency response and loss assessment, reflecting the cutting-edge nature and practicality of the technology.
[0028] 9. Historical SAR satellite remote sensing images carry multiple data dimensions such as backscatter intensity, metadata, polarization information, spatial and geometric information. This comprehensive information allows for cross-validation and in-depth analysis. For example, polarization information helps distinguish water body types, and the temporal information in the metadata supports time series analysis, thereby improving the robustness and adaptability of flood monitoring, reducing the uncertainty that may be caused by a single data source, and enhancing the comprehensiveness of the method and its industrial application value.
[0029] 10. By specifying in detail the preprocessing steps such as radiometric calibration, filtering removal, georegistration, and terrain correction, a standardized data processing workflow has been formed. This systematic processing can effectively eliminate noise, geometric distortion, and terrain effects in SAR satellite remote sensing images, ensure data consistency and comparability, lay a high-quality foundation for subsequent monitoring, and improve the repeatability and efficiency of the method.
[0030] 11. The specific algorithm selections, such as using the Lee Filter algorithm for filtering and removing backscattering, and using mathematical formulas to calculate the backscattering coefficient and convert it to decibel values in radiometric calibration, reflect technical optimization. The Lee Filter can effectively suppress speckle noise without losing details, while the decibel conversion facilitates data standardization. These optimizations not only improve processing speed but also enhance the reliability of the results.
[0031] 12. Geographic registration and terrain correction using orbital ephemeris data and digital elevation models, including 3D geometric reconstruction and local incident angle correction, can effectively address monitoring challenges in mountainous or undulating terrain. This meticulous terrain correction mechanism reduces errors such as terrain shading and perspective shrinkage, enabling the method to maintain high accuracy in complex geographical environments, highlighting its adaptability and innovation, and helping to expand its application scope.
[0032] 13. By labeling preprocessed SAR satellite remote sensing images with water body pixels, non-water body pixels, and easily confused areas (such as wind-blown water surfaces and vegetation-covered water areas), a dataset is constructed. This targeted labeling helps train a more intelligent classification model, reduces common confusion problems in flood monitoring, and improves the accuracy and generalization ability of automated monitoring. It reflects the forward-looking and practical value of the technology in artificial intelligence-assisted disaster management.
[0033] 14. By integrating the SAR image feature extraction module, metadata encoding module, and polarization feature enhancement module through a multimodal feature extraction layer, the system fully leverages the diversity of SAR satellite remote sensing data. The SAR image feature extraction module employs residual networks and dilated convolutions to extract multi-scale features, which, combined with feature pyramid integration, can capture terrain details at different scales. The metadata encoding module uses a multilayer perceptron and self-attention mechanism to map auxiliary data (such as time and location) into high-dimensional features and capture dependencies. The polarization feature enhancement module focuses on water-sensitive polarization channels to enhance information specificity. This multimodal design avoids the limitations of a single data source, improves feature richness and model robustness, and can more accurately identify flood extent and reduce misjudgment rate, especially in complex environments (such as cloud cover or terrain changes).
[0034] 15. The modality fusion layer, based on a spatial alignment module and a dynamic weight fusion module, achieves intelligent feature integration. The spatial alignment module uses deformable convolution to correct residual geometric distortion and extends metadata features to the spatial dimension, ensuring spatial consistency of multimodal features. The dynamic weight fusion module calculates relevance weights through a cross-attention mechanism and dynamically fuses features using gated recurrent units, adaptively adjusting the contribution of each modality. This fusion method avoids the rigidity of traditional methods (such as simple stitching) and can optimize feature combinations according to specific scenarios, thereby improving segmentation accuracy. For example, in flood monitoring, it can effectively combine the texture information of images and the contextual information of metadata to improve the detection capability of weak water body signals.
[0035] 16. The feature enhancement layer introduces a self-supervised contrastive learning module and an enhanced terrain attention module to further optimize feature quality. The self-supervised contrastive learning module constructs positive and negative sample pairs through data augmentation and uses a contrastive loss function to enhance the water body discrimination of the fused features, enabling the model to better distinguish flooded areas from similar backgrounds (such as shadows or vegetation). The enhanced terrain attention module calculates attention weights through geometric information and incorporates terrain information into the features using conditional batch normalization, improving adaptability to terrain changes. This enhancement mechanism not only improves the model's discriminative ability but also enhances its generalization ability, enabling it to maintain stable performance under different geographical regions and seasonal conditions, making it suitable for large-scale flood disaster monitoring.
[0036] 17. The segmentation prediction layer employs a multi-scale segmentation decoding module and an uncertainty estimation module to achieve high-precision segmentation and reliability assessment. The multi-scale segmentation decoding module is based on the U-Net decoder structure, combining upsampling and skip connections to output a high-resolution water distribution map, ensuring detail preservation. The uncertainty estimation module uses Monte Carlo dropout to sample multiple times during inference, calculates the prediction variance as uncertainty, and filters uncertain regions. This provides a reliability index for the segmentation results, allowing users to adjust the output according to the uncertainty threshold, reducing decision-making risks. For example, in emergency response, this design can output a more reliable flood extent map, supporting precise allocation of rescue resources and enhancing the practical value of the technical solution.
[0037] 18. The joint loss function combines the focus loss function and the dice loss function, effectively solving the class imbalance and optimization challenges in flood segmentation. The focus loss emphasizes the learning of difficult samples (such as a small number of water pixels) through weight coefficients and focus parameters, reducing the background dominance problem. The dice loss directly optimizes the segmentation overlap, improving the consistency between prediction and true label. This joint design balances accuracy and robustness, enabling faster convergence and improved model performance during training. Especially in flood disaster scenarios with sparse water pixels, it can significantly reduce false positive and false negative rates, enhancing the accuracy of monitoring results.
[0038] 19. By employing a spatiotemporal partitioning method, the dataset is divided into a training set, a validation set, and a test set. The validation set is newer than the training set, and the test set is newer than the validation set. Furthermore, all three sets originate from geographically non-overlapping regions. This partitioning method ensures that the model training, validation, and testing processes are independent in both time and space, effectively avoiding data leakage and overfitting issues. It also enhances the model's generalization ability to unknown spatiotemporal scenarios, making flood disaster monitoring results more reliable. This method is particularly suitable for dynamically changing natural disaster scenarios, thus enhancing its practicality and robustness.
[0039] 20. The Tree-structured Parzen Estimator (TPE) performs hyperparameter search on the training and validation sets to automatically obtain the optimal combination of hyperparameters (such as learning rate, batch size, etc.). This is an efficient method based on Bayesian optimization, which significantly reduces the cost and time of manual hyperparameter tuning. By intelligently exploring the hyperparameter space, it quickly converges to the optimal configuration, thereby improving model performance (such as segmentation accuracy and training efficiency).
[0040] 21. During training, a joint loss function is used to calculate the training loss value, and early stopping is monitored through backpropagation and validation sets until a preset condition is triggered. This effectively prevents model overfitting and ensures that the training process terminates when the validation set performance is optimal, thereby improving training efficiency and model stability.
[0041] 22. The model is evaluated using multiple metrics such as precision and recall calculated on the test set, and deployed using the TensorRT inference engine, ensuring that the model has been rigorously validated and has high reliability. At the same time, the TensorRT engine optimizes the inference speed, making it suitable for real-time or near real-time flood disaster monitoring applications.
[0042] 23. By integrating the streaming computing engine and the TensorRT inference engine, real-time preprocessing and model inference of high-resolution SAR satellite remote sensing images are realized. The streaming computing engine supports continuous data stream processing, avoiding the latency of traditional batch processing, while TensorRT optimizes the inference speed of deep learning models, significantly improving the response efficiency of flood monitoring. This design enables the system to quickly generate water distribution maps when disasters occur, providing timely support for emergency decision-making, which is superior to traditional methods that rely on offline processing.
[0043] 24. By combining differential operations with morphological opening and closing operations, intelligent analysis of time-series water distribution maps is achieved. Differential operations can accurately capture changes in flood range, while morphological operations improve the robustness of the results through noise reduction and edge refinement. This time-series analysis method can not only identify newly added disaster ranges, but also dynamically track the evolution process, providing continuous data support for disaster assessment, which is superior to static monitoring methods.
[0044] 25. By setting up monitoring threshold groups and automatically pushing early warning notifications, the entire process of automated monitoring is achieved; the large display screen shows the time sequence and prediction results in real time, improving the accessibility of information, while the threshold triggering mechanism ensures that the management terminal is notified immediately when the disaster exceeds the limit. This integrated design reduces manual intervention, improves the level of automation in disaster response, is suitable for large-scale monitoring scenarios, and reduces operating costs.
[0045] 26. The flood segmentation model innovatively integrates a multimodal feature extraction layer, a modality fusion layer, and a feature enhancement layer. By utilizing techniques such as residual networks, self-attention mechanisms, and dynamic weight fusion, it effectively combines different data sources (such as image features, metadata, and polarization information). The joint loss function combines focus loss and dice loss to solve the class imbalance problem and improve the model's ability to handle difficult samples (such as easily confused regions). Through hyperparameter search and early stopping monitoring, the model training process optimizes parameters, ensuring efficient convergence and generalization performance, thereby improving overall monitoring efficiency.
[0046] 27. By integrating high-resolution SAR satellite remote sensing image preprocessing, multimodal feature fusion-based flood segmentation models, and real-time inference mechanisms, high-precision and reliable flood disaster monitoring was achieved. Its advantages include improving data quality through radiometric calibration and terrain correction, optimizing model performance through multimodal feature extraction and dynamic weight fusion, ensuring the accuracy and robustness of water body segmentation; simultaneously, it supports real-time differential computation and morphological processing, enabling dynamic identification of newly added flood areas and providing prediction and early warning through ConvLSTM networks, combined with comprehensive parameter analysis such as inundation area and water depth, providing multidimensional disaster assessment; furthermore, the model's generalization ability is enhanced through spatiotemporal data partitioning and hyperparameter optimization, while modular design and TensorRT deployment ensure scalability and ease of use, significantly improving the efficiency and scientific rigor of disaster response overall. Attached Figure Description
[0047] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0048] Figure 1 This is a flowchart of a method for monitoring the extent of flood disasters based on satellite remote sensing imagery, according to the present invention.
[0049] Figure 2 This is a schematic diagram of the structure of a flood disaster range monitoring system based on satellite remote sensing imagery according to the present invention.
[0050] Figure 3 This is a schematic diagram illustrating the training and deployment process of the flood segmentation model of this invention.
[0051] Figure 4This is a schematic diagram of the process for processing SAR satellite remote sensing images of the target of this invention.
[0052] Figure 5 This is a flowchart illustrating the dynamic monitoring and early warning process for flood disasters according to the present invention.
[0053] Figure 6 This is the network structure diagram of the flood diversion model of the present invention. Detailed Implementation
[0054] The overall approach of the technical solution in this application is as follows: In terms of accuracy, by leveraging the detail advantages of high-resolution SAR satellite remote sensing images and the powerful multimodal feature extraction and anti-interference capabilities of the flood segmentation model, it is possible to accurately identify small water bodies and inundation boundaries in complex scenarios; in terms of timeliness, by constructing a fully automated analysis process based on a single SAR satellite remote sensing image, the reliance on multi-source data fusion is eliminated, enabling rapid processing from data input to result output, and by utilizing differential operations and time-series analysis of pre- and post-disaster images, the accuracy and timeliness of flood disaster range monitoring are improved.
[0055] Please refer to Figures 1 to 6 As shown, a preferred embodiment of the present invention, a method for monitoring the extent of flood disasters based on satellite remote sensing imagery, includes the following steps: Step S1: The server acquires a large number of historical SAR satellite remote sensing images with spatial resolution higher than a preset resolution threshold. The server performs preprocessing on each of the historical SAR satellite remote sensing images, including at least radiometric calibration, filtering and removal, georegistration and terrain correction. The server then labels the preprocessed historical SAR satellite remote sensing images to construct a dataset. Step S2: The server creates a flood segmentation model based on a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer, and sets the joint loss function of the flood segmentation model; the joint loss function combines the focus loss function and the dice loss function; Step S3: The server trains the flood segmentation model using the dataset. During the training process, the model parameters of the flood segmentation model are optimized using the joint loss function. The trained flood segmentation model is then deployed. Step S4: The server acquires target SAR satellite remote sensing images of the target area before and during the flood event, with a resolution higher than the resolution threshold. After preprocessing each target SAR satellite remote sensing image, it inputs the deployed flood segmentation model for inference to obtain a water distribution map. Step S5: The server performs a difference operation on the adjacent water body distribution maps to obtain the newly added flood disaster area; Step S6: The server generates a time-series flood inundation range map of the target area based on each of the flood disaster ranges, and calculates flood characteristic parameters including inundation area, inundation depth and inundation volume. Based on the time-series flood inundation range map and flood characteristic parameters, the server performs dynamic monitoring and early warning of the flood disaster range.
[0056] Step S1 specifically involves: The server acquires a large number of historical SAR satellite remote sensing images with a spatial resolution higher than a preset resolution threshold; the resolution threshold is 1 meter; the historical SAR satellite remote sensing images carry at least backscatter intensity (this is the most basic and commonly used information layer of SAR satellite remote sensing images, reflecting the strength of the ground target's ability to reflect radar signals, usually expressed as power value (dB) or amplitude value); bright areas (high backscatter): indicate areas that reflect strong signals, such as rough surfaces (rocks, buildings), metal structures, and slopes perpendicular to the radar beam direction (such as uphill surfaces); dark areas (low backscatter): indicate areas that reflect strong signals. Areas that reflect very weak signals or scatter signals in other directions; calm water is a typical example, reflecting radar pulses away from satellites like a mirror, resulting in very dark blacks in images. This is the core physical principle of using SAR to monitor floods. Metadata (this is the data describing the data itself, crucial, containing all the parameter information of the imaging; metadata is usually provided with the image data in the form of a header file), polarization information (radar waves have specific polarization directions (horizontal H or vertical V), SAR systems can transmit and receive signals in specific polarization patterns, which constitute different polarization channels; co-polarization (HH, VV: The transmitting and receiving polarizations are the same; Cross-polarization (HV, VH): The transmitting and receiving polarizations are different. Different ground features interact differently with microwaves of different polarizations; for example, water bodies show extremely low scattering in images with the same polarization (HH or VV), while the signal is weaker in cross-polarization (HV). Vegetation has a relatively strong response in cross-polarization. Multi-polarization data greatly enhances ground feature classification and capabilities, and can better distinguish between floods, vegetated water bodies, and bare land. The metadata includes at least imaging geometric parameters (satellite orbital altitude, imaging mode (strip, scan, etc.), incident angle, azimuth angle, etc.), time information (image acquisition date and exact time (UTC)), location information (satellite latitude and longitude coordinates, platform attitude (pitch, roll, yaw)), data processing level (indicating whether it is raw data (L0), single-view complex data (SLC), or geocoded ground distance detection product (GRD)) and polarization mode (recording the polarization channel of the image (e.g., HH, VV). The spatial and geometric information includes at least spatial resolution (including distance resolution and azimuth resolution, which are key indicators for measuring the ability to capture image details) and geographic coordinates (so that they can be overlaid with maps and other GIS data layers). Each of the aforementioned historical SAR satellite remote sensing images is subjected to preprocessing including at least radiometric calibration, filtering and removal, georegistration, and terrain correction. The radiometric calibration specifically involves: extracting the pixel amplitude values DN from historical SAR satellite remote sensing images; calculating the backscattering coefficient σ° based on the amplitude values DN and the calibration constant K; converting the unit of the backscattering coefficient σ° to decibels to obtain the decibel value σ°_dB of the backscattering coefficient σ°, thus completing the radiometric calibration. σ° = DN² / K; σ°_dB = 10 * log10(σ°); The fundamental purpose of radiometric calibration is to convert the raw grayscale value (digital value, DN) of each pixel in SAR satellite remote sensing imagery into a backscattering coefficient with a clear physical meaning. The backscattering coefficient quantitatively describes the ground's ability to reflect radar waves. Only radiometrically calibrated data can be used for this purpose. Quantitative comparison of images from different time phases: for example, comparing backscatter values of water bodies before and during floods; combined use of images from different sensors and models: eliminating inconsistencies in radiation levels caused by differences in satellite systems; accurate threshold segmentation or model inference: providing stable and reliable input features for subsequent flood segmentation models.
[0057] After radiometric calibration, calm water surfaces (specular reflection) will exhibit extremely low negative decibel values (e.g., around -20 dB), while rough terrain features such as urban buildings and vegetation will exhibit higher decibel values (e.g., -5 dB to over 0 dB). This strong contrast is the physical basis for using SAR to identify water bodies.
[0058] The backscattering coefficient is a standardized, absolute backscattering intensity, while backscattering intensity usually refers to the original, relative intensity value. They are related but not entirely the same concept. For example, consider the following analogy: Backscattering intensity is like the raw brightness value of a photograph taken with different cameras and settings (such as ISO, aperture, and shutter speed). The brightness of this photograph can only tell you which object in the image is brighter or darker than another object, but you cannot directly compare the brightness value of this photograph with a photograph taken with another camera because the shooting parameters are completely different.
[0059] The backscattering coefficient is like the absolute physical brightness value of each pixel obtained after professional calibration (e.g., candela per square meter). Now you can use this absolute value to make precise quantitative comparisons with any other calibrated photo.
[0060] The filtering and removal process employs the Lee Filter algorithm. The inherent coherent imaging principle of SAR satellite remote sensing imagery results in a large amount of granular "speckle noise" in the images. This is not "noise" in the traditional sense, but rather a random scattering phenomenon caused by the interference of radar waves reflected from multiple scatterers within ground features. Speckle noise can: severely interfere with visual interpretation and automatic computer classification; obscure the true texture and boundaries of ground features; reduce classification accuracy; and cause brightness fluctuations within homogeneous ground features (such as large bodies of water). The purpose of filtering is to achieve the optimal balance between suppressing noise and preserving edge / detail information.
[0061] The Lee Filter algorithm is a classic adaptive filtering algorithm that assumes that the radar backscattering coefficient is constant within a small window and that the observed variation is caused by speckle noise. It adjusts the filtering strength according to the statistical characteristics (such as variance) of the pixels within the window, with a large filtering strength in uniform regions (such as water) and a small filtering strength in edge regions (such as the water-land boundary) to protect the edges.
[0062] The georeferencing specifically involves georeferencing historical SAR satellite remote sensing images using orbital ephemeris data and digital elevation models. Specifically, based on the radar range-Doppler positioning model, combined with the satellite's position, velocity, imaging time, and elevation information provided by the DEM, the geodetic coordinates corresponding to each pixel are calculated, and resampling is performed using bilinear interpolation. The purpose of georegistration is to assign precise geographic coordinates (latitude and longitude) to each pixel on a SAR satellite remote sensing image, enabling it to be spatially aligned precisely with maps, other remote sensing images (such as optical images), or GIS layers (such as administrative boundaries, digital elevation models, DEMs). This is crucial for flood monitoring because it requires: accurately locating the location of floods and performing precise differential calculations with multi-temporal images.
[0063] The terrain correction specifically involves: reconstructing the three-dimensional geometric relationships of historical SAR satellite remote sensing images using orbital ephemeris data and digital elevation models; identifying and correcting distortions (such as overlay and perspective shrinkage caused by terrain) to perform geometric terrain correction (correctly projecting each pixel onto the surface of the Earth's ellipsoid); calculating the local incident angle of each pixel in the geometrically terrain-corrected historical SAR satellite remote sensing image; and normalizing the brightness of the local incident angle using the cosine correction method to perform radiation terrain correction (eliminating brightness differences caused by terrain orientation, ensuring that the backscattering coefficient only reflects the scattering characteristics of the ground object itself, rather than the terrain slope). SAR is a side-looking imaging system. Topographical undulations can cause severe geometric distortions and radiometric distortions, mainly including: Overlapping: High points such as the mountaintop are illuminated by radar waves before the foot of the mountain, causing the image of the mountaintop to be "compressed" or even "overlapped" onto the foot of the mountain.
[0064] Shadows: Back slope areas that cannot be reached by radar waves appear as black (no signal) on the image.
[0065] Perspective shrinkage: A slope facing the radar appears shorter in the image than it actually is.
[0066] The purpose of terrain correction is to eliminate the influence of terrain undulations on the geometry and pixel brightness (backscattering intensity) of SAR satellite remote sensing images, generating an "orthophoto image" that looks like it was taken vertically from directly above.
[0067] The preprocessed historical SAR satellite remote sensing images are labeled with at least water body pixels, non-water body pixels, and easily confused regions to construct a dataset; the water body pixels are permanent water body pixels or flooded water body pixels; the easily confused regions include at least wind-blown water surfaces, vegetation-covered water areas, and flat non-water surfaces.
[0068] Step S2 specifically involves: The server creates a flood segmentation model based on a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer. A joint loss function is set for the flood segmentation model, which combines a focus loss function and a dice loss function. To address the extreme imbalance between flood water body pixels and background pixels and to optimize the segmentation boundary accuracy, a joint loss function combining the focus loss function and the dice loss function is used to optimize the model parameters of the flood segmentation model. The multimodal feature extraction layer is constructed based on the SAR image feature extraction module, the metadata encoding module, and the polarization feature enhancement module; The SAR image feature extraction module is used to extract multi-scale features from SAR satellite remote sensing images through ResNet and dilated convolution, and integrate the multi-scale features through Feature Pyramid (FPN) to obtain multi-scale image features; the metadata encoding module is used to map metadata in SAR satellite remote sensing images into high-dimensional features through a multilayer perceptron, and capture the dependencies between the high-dimensional features through a self-attention mechanism to obtain metadata features; the polarization feature enhancement module is used to compress and enhance the polarization information in SAR satellite remote sensing images through 1x1 convolution, and focus on the water-sensitive polarization channels through a polarization attention mechanism to obtain polarization enhancement features; By employing a high-resolution encoder (residual network combined with dilated convolution) and a feature pyramid, the receptive field can be expanded while maintaining high spatial resolution. This is directly optimized for high-resolution SAR satellite remote sensing imagery, effectively capturing small water bodies and clear boundaries. Furthermore, the multi-scale segmentation decoding module further refines the output through upsampling and skip connections, avoiding blurred boundaries. Therefore, the flood segmentation model significantly improves the accuracy of identifying small water bodies.
[0069] The modality fusion layer is constructed based on a spatial alignment module and a dynamic weight fusion module; The spatial alignment module performs spatial alignment operations on multi-scale image features, metadata features, and polarization enhancement features through deformable convolution to correct residual geometric distortions and expands the metadata features to the spatial dimension through a broadcast operation. The dynamic weight fusion module calculates the correlation weights between the multi-scale image features, metadata features, and polarization enhancement features output by the spatial alignment module through a cross-attention mechanism, and dynamically fuses the multi-scale image features, metadata features, and polarization enhancement features according to the correlation weights through a gated recurrent unit (GRU) to obtain fused features. The feature enhancement layer is constructed based on a self-supervised contrastive learning module and an enhanced terrain attention module; The self-supervised contrastive learning module is used to construct positive sample pairs (such as rotated and scaled features) and negative sample pairs (such as different land cover features) through data augmentation, and uses a contrastive loss function to enhance the water body discrimination of the fused features to obtain the first-level augmented features; the terrain-enhancing attention module is used to calculate attention weights through geometric information (such as local incident angles), and integrates geometric information into the first-level augmented features based on the attention weights through conditional batch normalization to enhance robustness to terrain changes to obtain the second-level augmented features; The robustness of features is enhanced through a self-supervised contrastive learning module and an enhanced terrain attention module. The self-supervised contrastive learning module reduces noise interference by constructing positive and negative sample pairs. The enhanced terrain attention module uses geometric information to correct terrain distortion and reduce the impact of building overlap and shadow effects. The polarization feature enhancement module focuses on water-sensitive polarization channels to reduce micro-motion interference such as wind and waves. The spatial alignment module uses deformable convolution to correct geometric distortion, further improving the anti-interference capability.
[0070] The segmentation prediction layer is constructed based on a multi-scale segmentation decoding module and an uncertainty estimation module; The multi-scale segmentation and decoding module is used to upsample and skip connections the secondary enhancement features through the U-Net decoder structure, and output a high-resolution water distribution map by combining the multi-scale image features. The uncertainty estimation module is used to calculate the prediction variance of the water distribution map as uncertainty by sampling multiple times during inference through Monte Carlo dropout. After filtering the uncertainty region in the water distribution map by the uncertainty and the preset uncertainty threshold, the module outputs a water distribution map carrying the uncertainty. Because it relies solely on SAR satellite remote sensing imagery (including backscatter intensity, metadata, polarization information, etc.), it eliminates the need to fuse optical imagery or other data sources, reducing multi-source dependence. The modality fusion layer efficiently integrates multimodal features through dynamic weight fusion, improving automation. The flood segmentation model adopts an encoder-decoder architecture and introduces an uncertainty estimation module for rapid inference, meeting the real-time response requirements of sudden disasters. Its computational complexity is lower than traditional texture feature extraction (such as GLCM), improving the efficiency of large-scale data processing.
[0071] The formula for the joint loss function is: ; ; ; in, Indicates the value of the joint loss function; This represents the value of the focus loss function; This represents the value of the dice loss function; This represents the total number of pixels processed in one training iteration; i represents the pixel index, ranging from 1 to N, used to iterate through each pixel; This represents the true category label of the i-th pixel, with a value of 0 or 1 (1 represents water (positive class), and 0 represents non-water (negative class)). Indicate category Weighting coefficients; This indicates that the i-th pixel predicted by the flood segmentation model belongs to... The probability; γ represents the focusing parameter, which is a real number hyperparameter greater than or equal to 0, used to adjust the weight of easy and difficult samples (the larger the value, the more the model focuses on difficult samples (i.e., samples with predicted probabilities close to 0 or 1)). This represents the true category label of the i-th pixel; ϵ represents the predicted probability of the i-th pixel (representing the model's confidence that the pixel belongs to water); ϵ represents the smoothing constant (a very small positive number added to the numerator and denominator to prevent numerical instability when the denominator is zero). This represents twice the area of the intersection between the actual water body pixels and the predicted water body pixels, used to measure the degree of overlap. This represents the sum of the actual total number of water body pixels and the predicted probability of water body pixels, used for normalization.
[0072] In the focus loss function and the dice loss function The two terms have the same meaning; both refer to the true class labels used to supervise model training. In the focus loss function, Used to calculate the model's predicted probability; in the dice loss function Similarity is used to calculate and predict probabilities, measuring segmentation accuracy; different notations are used due to standard notation conventions for loss functions, with focus loss typically using... Emphasis is placed on category indexes, while dice loss is used. Emphasize the true value.
[0073] Step S3 specifically involves: The server divides the dataset into a training set, a validation set, and a test set based on a spatiotemporal partitioning method, such that the time of the validation set is newer than the time of the training set, the time of the test set is newer than the time of the validation set, and the training set, validation set, and test set come from geographically non-overlapping regions. Before training begins, a tree-structured Parzen estimator (TPE) is used to search for hyperparameters on the training and validation sets to obtain the optimal combination of hyperparameters (this is a process of multiple iterations); the hyperparameters include at least the learning rate, batch size, network depth, network width, Dropout ratio, decay step size, and decay rate. The flood segmentation model is trained on the training set using the hyperparameter combination. During training, the training loss value is calculated using the joint loss function, and the gradient of the training loss value with respect to the model parameters of the flood segmentation model is calculated using the backpropagation algorithm. The model parameters are updated based on the gradient, and early stopping is monitored using the validation set until a preset early stopping condition is triggered (such as when the validation loss value no longer decreases within several consecutive periods) to complete the training. The at least the model parameters include: the kernel weights and biases of the convolutional layers in the residual network of the SAR image feature extraction module, and the dilated kernel parameters of the dilated convolution; the weights and biases of the fully connected layers in the multilayer perceptron of the metadata encoding module, and the weights of the query matrix, key matrix, and value matrix in the self-attention mechanism; the kernel weights and biases of the 1x1 convolutional layers in the polarization feature enhancement module, and the attention weight parameters in the polarization attention mechanism; the kernel weights and biases of the deformable convolution in the spatial alignment module, and the convolutional layer parameters for learning offsets; the projection matrix used to calculate the relevance weights in the cross-attention mechanism of the dynamic weight fusion module, and the weight matrix and bias vector in the gated recurrent unit; the parameters for calculating the attention weights of geometric information in the enhanced terrain attention module, and the scaling and offset parameters in conditional batch normalization; and the kernel weights and biases of all upsampled convolutional kernels in the U-Net decoder of the multi-scale segmentation decoding module. Precision, recall, F1-Score, intersection-over-union ratio, overall accuracy, and inference speed are calculated using the test set to test the trained flood segmentation model. The tested flood segmentation model is then deployed using the TensorRT inference engine.
[0074] Step S4 specifically involves: Based on the received flood disaster monitoring instructions, the server obtains target SAR satellite remote sensing images of the target area with a resolution higher than the resolution threshold from the remote sensing image database before and during the flood event. After performing preprocessing on each target SAR satellite remote sensing image in sequence through the streaming computing engine, including at least radiometric calibration, filtering and removal, georegistration and terrain correction, the server inputs each target SAR satellite remote sensing image into the deployed flood segmentation model and performs real-time inference through the TensorRT inference engine to obtain the water distribution map. Step S5 specifically involves: The server performs differential operations on the time-adjacent water body distribution maps, combines morphological opening operations for noise reduction, and combines morphological closing operations for edge refinement to obtain the newly added flood disaster range. Step S6 specifically involves: The server generates a time-series flood inundation range map of the target area based on the flood disaster ranges, and calculates flood characteristic parameters including inundation area, inundation depth, and inundation volume. The inundation area is calculated based on the time-series flood inundation range map and spatial resolution (pixel count × single pixel area). The inundation depth is calculated based on the time-series flood inundation range map and digital elevation model (flood water surface elevation - ground elevation; flood water surface elevation is estimated using a water surface fitting algorithm based on the time-series flood inundation range map and digital elevation model). The inundation volume is calculated based on the inundation area and inundation depth (through spatial integration, i.e., calculating the volume of each pixel separately and then summing them). The server inputs the time-series flood inundation range map into a pre-trained ConvLSTM network to obtain a predicted flood inundation range map. The time-series flood inundation range map, the predicted flood inundation range map, and flood characteristic parameters are displayed in real time on a large screen. The server monitors the time-series flood inundation range map, the predicted flood inundation range map, and the flood characteristic parameters through a preset monitoring threshold group. When the monitoring threshold group is triggered, a flood warning notification is immediately pushed to the management terminal, thereby dynamically monitoring and warning of the flood disaster range. A preferred embodiment of the flood disaster range monitoring system based on satellite remote sensing imagery of the present invention includes the following modules: The dataset construction module is used to acquire a large number of historical SAR satellite remote sensing images with spatial resolution higher than a preset resolution threshold. The module performs preprocessing on each of the historical SAR satellite remote sensing images, including at least radiometric calibration, filtering and removal, georegistration and terrain correction. The preprocessed historical SAR satellite remote sensing images are then labeled to construct the dataset. The flood segmentation model creation module is used by the server to create a flood segmentation model based on a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer, and to set the joint loss function of the flood segmentation model; the joint loss function combines the focus loss function and the dice loss function; The flood segmentation model training module is used by the server to train the flood segmentation model using the dataset. During the training process, the model parameters of the flood segmentation model are optimized through the joint loss function, and the trained flood segmentation model is deployed. The flood segmentation model inference module is used by the server to acquire target SAR satellite remote sensing images of the target area before and during a flood event, with a resolution higher than the resolution threshold. After preprocessing each target SAR satellite remote sensing image, the image is input into the deployed flood segmentation model for inference to obtain a water distribution map. The differential operation module is used by the server to perform differential operations on the distribution maps of adjacent water bodies to obtain the newly added flood disaster range; The flood disaster range monitoring module is used by the server to generate a time-series flood inundation range map of the target area based on each of the flood disaster ranges, and to calculate flood characteristic parameters including inundation area, inundation depth and inundation volume. Based on the time-series flood inundation range map and flood characteristic parameters, dynamic monitoring and early warning of the flood disaster range are carried out.
[0075] The dataset construction module is specifically used for: The server acquires a large number of historical SAR satellite remote sensing images with a spatial resolution higher than a preset resolution threshold; the resolution threshold is 1 meter; the historical SAR satellite remote sensing images carry at least backscatter intensity (this is the most basic and commonly used information layer of SAR satellite remote sensing images, reflecting the strength of the ground target's ability to reflect radar signals, usually expressed as power value (dB) or amplitude value); bright areas (high backscatter): indicate areas that reflect strong signals, such as rough surfaces (rocks, buildings), metal structures, and slopes perpendicular to the radar beam direction (such as uphill surfaces); dark areas (low backscatter): indicate areas that reflect strong signals. Areas that reflect very weak signals or scatter signals in other directions; calm water is a typical example, reflecting radar pulses away from satellites like a mirror, resulting in very dark blacks in images. This is the core physical principle of using SAR to monitor floods. Metadata (this is the data describing the data itself, crucial, containing all the parameter information of the imaging; metadata is usually provided with the image data in the form of a header file), polarization information (radar waves have specific polarization directions (horizontal H or vertical V), SAR systems can transmit and receive signals in specific polarization patterns, which constitute different polarization channels; co-polarization (HH, VV: The transmitting and receiving polarizations are the same; Cross-polarization (HV, VH): The transmitting and receiving polarizations are different. Different ground features interact differently with microwaves of different polarizations; for example, water bodies show extremely low scattering in images with the same polarization (HH or VV), while the signal is weaker in cross-polarization (HV). Vegetation has a relatively strong response in cross-polarization. Multi-polarization data greatly enhances ground feature classification and capabilities, and can better distinguish between floods, vegetated water bodies, and bare land. The metadata includes at least imaging geometric parameters (satellite orbital altitude, imaging mode (strip, scan, etc.), incident angle, azimuth angle, etc.), time information (image acquisition date and exact time (UTC)), location information (satellite latitude and longitude coordinates, platform attitude (pitch, roll, yaw)), data processing level (indicating whether it is raw data (L0), single-view complex data (SLC), or geocoded ground distance detection product (GRD)) and polarization mode (recording the polarization channel of the image (e.g., HH, VV). The spatial and geometric information includes at least spatial resolution (including distance resolution and azimuth resolution, which are key indicators for measuring the ability to capture image details) and geographic coordinates (so that they can be overlaid with maps and other GIS data layers). Each of the aforementioned historical SAR satellite remote sensing images is subjected to preprocessing including at least radiometric calibration, filtering and removal, georegistration, and terrain correction. The radiometric calibration specifically involves: extracting the pixel amplitude values DN from historical SAR satellite remote sensing images; calculating the backscattering coefficient σ° based on the amplitude values DN and the calibration constant K; converting the unit of the backscattering coefficient σ° to decibels to obtain the decibel value σ°_dB of the backscattering coefficient σ°, thus completing the radiometric calibration. σ° = DN² / K; σ°_dB = 10 * log10(σ°); The fundamental purpose of radiometric calibration is to convert the raw grayscale value (digital value, DN) of each pixel in SAR satellite remote sensing imagery into a backscattering coefficient with a clear physical meaning. The backscattering coefficient quantitatively describes the ground's ability to reflect radar waves. Only radiometrically calibrated data can be used for this purpose. Quantitative comparison of images from different time phases: for example, comparing backscatter values of water bodies before and during floods; combined use of images from different sensors and models: eliminating inconsistencies in radiation levels caused by differences in satellite systems; accurate threshold segmentation or model inference: providing stable and reliable input features for subsequent flood segmentation models.
[0076] After radiometric calibration, calm water surfaces (specular reflection) will exhibit extremely low negative decibel values (e.g., around -20 dB), while rough terrain features such as urban buildings and vegetation will exhibit higher decibel values (e.g., -5 dB to over 0 dB). This strong contrast is the physical basis for using SAR to identify water bodies.
[0077] The backscattering coefficient is a standardized, absolute backscattering intensity, while backscattering intensity usually refers to the original, relative intensity value. They are related but not entirely the same concept. For example, consider the following analogy: Backscattering intensity is like the raw brightness value of a photograph taken with different cameras and settings (such as ISO, aperture, and shutter speed). The brightness of this photograph can only tell you which object in the image is brighter or darker than another object, but you cannot directly compare the brightness value of this photograph with a photograph taken with another camera because the shooting parameters are completely different.
[0078] The backscattering coefficient is like the absolute physical brightness value of each pixel obtained after professional calibration (e.g., candela per square meter). Now you can use this absolute value to make precise quantitative comparisons with any other calibrated photo.
[0079] The filtering and removal process employs the Lee Filter algorithm. The inherent coherent imaging principle of SAR satellite remote sensing imagery results in a large amount of granular "speckle noise" in the images. This is not "noise" in the traditional sense, but rather a random scattering phenomenon caused by the interference of radar waves reflected from multiple scatterers within ground features. Speckle noise can: severely interfere with visual interpretation and automatic computer classification; obscure the true texture and boundaries of ground features; reduce classification accuracy; and cause brightness fluctuations within homogeneous ground features (such as large bodies of water). The purpose of filtering is to achieve the optimal balance between suppressing noise and preserving edge / detail information.
[0080] The Lee Filter algorithm is a classic adaptive filtering algorithm that assumes that the radar backscattering coefficient is constant within a small window and that the observed variation is caused by speckle noise. It adjusts the filtering strength according to the statistical characteristics (such as variance) of the pixels within the window, with a large filtering strength in uniform regions (such as water) and a small filtering strength in edge regions (such as the water-land boundary) to protect the edges.
[0081] The georeferencing specifically involves georeferencing historical SAR satellite remote sensing images using orbital ephemeris data and digital elevation models. Specifically, based on the radar range-Doppler positioning model, combined with the satellite's position, velocity, imaging time, and elevation information provided by the DEM, the geodetic coordinates corresponding to each pixel are calculated, and resampling is performed using bilinear interpolation. The purpose of georegistration is to assign precise geographic coordinates (latitude and longitude) to each pixel on a SAR satellite remote sensing image, enabling it to be spatially aligned precisely with maps, other remote sensing images (such as optical images), or GIS layers (such as administrative boundaries, digital elevation models, DEMs). This is crucial for flood monitoring because it requires: accurately locating the location of floods and performing precise differential calculations with multi-temporal images.
[0082] The terrain correction specifically involves: reconstructing the three-dimensional geometric relationships of historical SAR satellite remote sensing images using orbital ephemeris data and digital elevation models; identifying and correcting distortions (such as overlay and perspective shrinkage caused by terrain) to perform geometric terrain correction (correctly projecting each pixel onto the surface of the Earth's ellipsoid); calculating the local incident angle of each pixel in the geometrically terrain-corrected historical SAR satellite remote sensing image; and normalizing the brightness of the local incident angle using the cosine correction method to perform radiation terrain correction (eliminating brightness differences caused by terrain orientation, ensuring that the backscattering coefficient only reflects the scattering characteristics of the ground object itself, rather than the terrain slope). SAR is a side-looking imaging system. Topographical undulations can cause severe geometric distortions and radiometric distortions, mainly including: Overlapping: High points such as the mountaintop are illuminated by radar waves before the foot of the mountain, causing the image of the mountaintop to be "compressed" or even "overlapped" onto the foot of the mountain.
[0083] Shadows: Back slope areas that cannot be reached by radar waves appear as black (no signal) on the image.
[0084] Perspective shrinkage: A slope facing the radar appears shorter in the image than it actually is.
[0085] The purpose of terrain correction is to eliminate the influence of terrain undulations on the geometry and pixel brightness (backscattering intensity) of SAR satellite remote sensing images, generating an "orthophoto image" that looks like it was taken vertically from directly above.
[0086] The preprocessed historical SAR satellite remote sensing images are labeled with at least water body pixels, non-water body pixels, and easily confused regions to construct a dataset; the water body pixels are permanent water body pixels or flooded water body pixels; the easily confused regions include at least wind-blown water surfaces, vegetation-covered water areas, and flat non-water surfaces.
[0087] The flood diversion model creation module is specifically used for: The server creates a flood segmentation model based on a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer. A joint loss function is set for the flood segmentation model, which combines a focus loss function and a dice loss function. To address the extreme imbalance between flood water body pixels and background pixels and to optimize the segmentation boundary accuracy, a joint loss function combining the focus loss function and the dice loss function is used to optimize the model parameters of the flood segmentation model. The multimodal feature extraction layer is constructed based on the SAR image feature extraction module, the metadata encoding module, and the polarization feature enhancement module; The SAR image feature extraction module is used to extract multi-scale features from SAR satellite remote sensing images through ResNet and dilated convolution, and integrate the multi-scale features through Feature Pyramid (FPN) to obtain multi-scale image features; the metadata encoding module is used to map metadata in SAR satellite remote sensing images into high-dimensional features through a multilayer perceptron, and capture the dependencies between the high-dimensional features through a self-attention mechanism to obtain metadata features; the polarization feature enhancement module is used to compress and enhance the polarization information in SAR satellite remote sensing images through 1x1 convolution, and focus on the water-sensitive polarization channels through a polarization attention mechanism to obtain polarization enhancement features; By employing a high-resolution encoder (residual network combined with dilated convolution) and a feature pyramid, the receptive field can be expanded while maintaining high spatial resolution. This is directly optimized for high-resolution SAR satellite remote sensing imagery, effectively capturing small water bodies and clear boundaries. Furthermore, the multi-scale segmentation decoding module further refines the output through upsampling and skip connections, avoiding blurred boundaries. Therefore, the flood segmentation model significantly improves the accuracy of identifying small water bodies.
[0088] The modality fusion layer is constructed based on a spatial alignment module and a dynamic weight fusion module; The spatial alignment module performs spatial alignment operations on multi-scale image features, metadata features, and polarization enhancement features through deformable convolution to correct residual geometric distortions and expands the metadata features to the spatial dimension through a broadcast operation. The dynamic weight fusion module calculates the correlation weights between the multi-scale image features, metadata features, and polarization enhancement features output by the spatial alignment module through a cross-attention mechanism, and dynamically fuses the multi-scale image features, metadata features, and polarization enhancement features according to the correlation weights through a gated recurrent unit (GRU) to obtain fused features. The feature enhancement layer is constructed based on a self-supervised contrastive learning module and an enhanced terrain attention module; The self-supervised contrastive learning module is used to construct positive sample pairs (such as rotated and scaled features) and negative sample pairs (such as different land cover features) through data augmentation, and uses a contrastive loss function to enhance the water body discrimination of the fused features to obtain the first-level augmented features; the terrain-enhancing attention module is used to calculate attention weights through geometric information (such as local incident angles), and integrates geometric information into the first-level augmented features based on the attention weights through conditional batch normalization to enhance robustness to terrain changes to obtain the second-level augmented features; The robustness of features is enhanced through a self-supervised contrastive learning module and an enhanced terrain attention module. The self-supervised contrastive learning module reduces noise interference by constructing positive and negative sample pairs. The enhanced terrain attention module uses geometric information to correct terrain distortion and reduce the impact of building overlap and shadow effects. The polarization feature enhancement module focuses on water-sensitive polarization channels to reduce micro-motion interference such as wind and waves. The spatial alignment module uses deformable convolution to correct geometric distortion, further improving the anti-interference capability.
[0089] The segmentation prediction layer is constructed based on a multi-scale segmentation decoding module and an uncertainty estimation module; The multi-scale segmentation and decoding module is used to upsample and skip connections the secondary enhancement features through the U-Net decoder structure, and output a high-resolution water distribution map by combining the multi-scale image features. The uncertainty estimation module is used to calculate the prediction variance of the water distribution map as uncertainty by sampling multiple times during inference through Monte Carlo dropout. After filtering the uncertainty region in the water distribution map by the uncertainty and the preset uncertainty threshold, the module outputs a water distribution map carrying the uncertainty. Because it relies solely on SAR satellite remote sensing imagery (including backscatter intensity, metadata, polarization information, etc.), it eliminates the need to fuse optical imagery or other data sources, reducing multi-source dependence. The modality fusion layer efficiently integrates multimodal features through dynamic weight fusion, improving automation. The flood segmentation model adopts an encoder-decoder architecture and introduces an uncertainty estimation module for rapid inference, meeting the real-time response requirements of sudden disasters. Its computational complexity is lower than traditional texture feature extraction (such as GLCM), improving the efficiency of large-scale data processing.
[0090] The formula for the joint loss function is: ; ; ; in, Indicates the value of the joint loss function; This represents the value of the focus loss function; This represents the value of the dice loss function; This represents the total number of pixels processed in one training iteration; i represents the pixel index, ranging from 1 to N, used to iterate through each pixel; This represents the true category label of the i-th pixel, with a value of 0 or 1 (1 represents water (positive class), and 0 represents non-water (negative class)). Indicate category Weighting coefficients; This indicates that the i-th pixel predicted by the flood segmentation model belongs to... The probability; γ represents the focusing parameter, which is a real number hyperparameter greater than or equal to 0, used to adjust the weight of easy and difficult samples (the larger the value, the more the model focuses on difficult samples (i.e., samples with predicted probabilities close to 0 or 1)). This represents the true category label of the i-th pixel; ϵ represents the predicted probability of the i-th pixel (representing the model's confidence that the pixel belongs to water); ϵ represents the smoothing constant (a very small positive number added to the numerator and denominator to prevent numerical instability when the denominator is zero). This represents twice the area of the intersection between the actual water body pixels and the predicted water body pixels, used to measure the degree of overlap. This represents the sum of the actual total number of water body pixels and the predicted probability of water body pixels, used for normalization.
[0091] In the focus loss function and the dice loss function The two terms have the same meaning; both refer to the true class labels used to supervise model training. In the focus loss function, Used to calculate the model's predicted probability; in the dice loss function Similarity is used to calculate and predict probabilities, measuring segmentation accuracy; different notations are used due to standard notation conventions for loss functions, with focus loss typically using... Emphasis is placed on category indexes, while dice loss is used. Emphasize the true value.
[0092] The flood segmentation model training module is specifically used for: The server divides the dataset into a training set, a validation set, and a test set based on a spatiotemporal partitioning method, such that the time of the validation set is newer than the time of the training set, the time of the test set is newer than the time of the validation set, and the training set, validation set, and test set come from geographically non-overlapping regions. Before training begins, a tree-structured Parzen estimator (TPE) is used to search for hyperparameters on the training and validation sets to obtain the optimal combination of hyperparameters (this is a process of multiple iterations); the hyperparameters include at least the learning rate, batch size, network depth, network width, Dropout ratio, decay step size, and decay rate. The flood segmentation model is trained on the training set using the hyperparameter combination. During training, the training loss value is calculated using the joint loss function, and the gradient of the training loss value with respect to the model parameters of the flood segmentation model is calculated using the backpropagation algorithm. The model parameters are updated based on the gradient, and early stopping is monitored using the validation set until a preset early stopping condition is triggered (such as when the validation loss value no longer decreases within several consecutive periods) to complete the training. The at least the model parameters include: the kernel weights and biases of the convolutional layers in the residual network of the SAR image feature extraction module, and the dilated kernel parameters of the dilated convolution; the weights and biases of the fully connected layers in the multilayer perceptron of the metadata encoding module, and the weights of the query matrix, key matrix, and value matrix in the self-attention mechanism; the kernel weights and biases of the 1x1 convolutional layers in the polarization feature enhancement module, and the attention weight parameters in the polarization attention mechanism; the kernel weights and biases of the deformable convolution in the spatial alignment module, and the convolutional layer parameters for learning offsets; the projection matrix used to calculate the relevance weights in the cross-attention mechanism of the dynamic weight fusion module, and the weight matrix and bias vector in the gated recurrent unit; the parameters for calculating the attention weights of geometric information in the enhanced terrain attention module, and the scaling and offset parameters in conditional batch normalization; and the kernel weights and biases of all upsampled convolutional kernels in the U-Net decoder of the multi-scale segmentation decoding module. Precision, recall, F1-Score, intersection-over-union ratio, overall accuracy, and inference speed are calculated using the test set to test the trained flood segmentation model. The tested flood segmentation model is then deployed using the TensorRT inference engine.
[0093] The flood segmentation model inference module is specifically used for: Based on the received flood disaster monitoring instructions, the server obtains target SAR satellite remote sensing images of the target area with a resolution higher than the resolution threshold from the remote sensing image database before and during the flood event. After performing preprocessing on each target SAR satellite remote sensing image in sequence through the streaming computing engine, including at least radiometric calibration, filtering and removal, georegistration and terrain correction, the server inputs each target SAR satellite remote sensing image into the deployed flood segmentation model and performs real-time inference through the TensorRT inference engine to obtain the water distribution map. The difference operation module is specifically used for: The server performs differential operations on the time-adjacent water body distribution maps, combines morphological opening operations for noise reduction, and combines morphological closing operations for edge refinement to obtain the newly added flood disaster range. The flood disaster range monitoring module is specifically used for: The server generates a time-series flood inundation range map of the target area based on the flood disaster ranges, and calculates flood characteristic parameters including inundation area, inundation depth, and inundation volume. The inundation area is calculated based on the time-series flood inundation range map and spatial resolution (pixel count × single pixel area). The inundation depth is calculated based on the time-series flood inundation range map and digital elevation model (flood water surface elevation - ground elevation; flood water surface elevation is estimated using a water surface fitting algorithm based on the time-series flood inundation range map and digital elevation model). The inundation volume is calculated based on the inundation area and inundation depth (through spatial integration, i.e., calculating the volume of each pixel separately and then summing them). The server inputs the time-series flood inundation range map into a pre-trained ConvLSTM network to obtain a predicted flood inundation range map. The time-series flood inundation range map, the predicted flood inundation range map, and flood characteristic parameters are displayed in real time on a large screen. The server monitors the time-series flood inundation range map, the predicted flood inundation range map, and the flood characteristic parameters through a preset monitoring threshold group. When the monitoring threshold group is triggered, a flood warning notification is immediately pushed to the management terminal, thereby dynamically monitoring and warning of the flood disaster range.
[0094] In summary, the advantages of this invention are as follows: 1. Acquire a large number of historical SAR satellite remote sensing images with spatial resolution higher than a preset resolution threshold through a server. Perform preprocessing on each historical SAR satellite remote sensing image, including at least radiometric calibration, filtering and removal, georegistration, and terrain correction. Label the preprocessed historical SAR satellite remote sensing images to construct a dataset. Then, create a flood segmentation model based on a multimodal feature extraction layer, modality fusion layer, feature enhancement layer, and segmentation prediction layer. Set the joint loss function of the flood segmentation model. Train the flood segmentation model using the dataset. During the training process, optimize the model parameters of the flood segmentation model through the joint loss function. Deploy the trained flood segmentation model. Next, acquire target SAR satellite remote sensing images of the target area with resolution higher than the resolution threshold before and during flood events. After preprocessing each target SAR satellite remote sensing image, input them into the flood segmentation model for inference to obtain... The system generates a water body distribution map, performs differential operations on adjacent water body distribution maps to obtain newly added flood disaster areas, and generates a time-series flood inundation range map of the target area based on each flood disaster range. It also calculates flood characteristic parameters, including inundation area, inundation depth, and inundation volume. Based on the time-series flood inundation range map and flood characteristic parameters, dynamic monitoring and early warning of flood disaster areas are conducted. In terms of accuracy, leveraging the detail advantages of high-resolution SAR satellite remote sensing imagery and the powerful multimodal feature extraction and anti-interference capabilities of the flood segmentation model, it can accurately identify inundation boundaries in small water bodies and complex scenarios. In terms of timeliness, by constructing a fully automated analysis process based on a single SAR satellite remote sensing image, it eliminates the reliance on multi-source data fusion, achieving rapid processing from data input to result output. Furthermore, by utilizing differential operations and time-series analysis of pre- and post-disaster images, it ultimately greatly improves the accuracy and timeliness of flood disaster area monitoring.
[0095] 2. By using SAR satellite remote sensing imagery with a spatial resolution higher than 1 meter as the data source, and combining preprocessing methods such as radiometric calibration, filtering and removal, georegistration, and terrain correction, the accuracy and quality of the data are improved from the source. On this basis, an innovative multimodal flood segmentation model is constructed. This flood segmentation model enhances the ability to distinguish features of water bodies and complex terrain scenes by fusing SAR image features, metadata, and polarization information, and by using attention mechanisms and self-supervised contrastive learning. Combined with a joint loss function optimized for boundaries and an uncertainty estimation module, the accuracy and anti-interference ability of flood range identification are significantly improved, and the omission of small water bodies and boundary ambiguity are effectively reduced. In terms of timeliness, the preprocessing and model inference are accelerated by the TensorRT inference engine and streaming computing engine, achieving near real-time processing. The newly added inundation range is quickly extracted by temporal difference analysis, and prediction is performed using a ConvLSTM network. Finally, an automated process from real-time monitoring to dynamic early warning is formed, thus achieving a dual improvement in the accuracy and timeliness of flood disaster range monitoring.
[0096] 3. Utilizing SAR (Synthetic Aperture Radar) satellite remote sensing imagery for flood monitoring offers advantages such as all-weather, all-time imaging, unaffected by cloud cover or weather conditions, ensuring the continuity and reliability of data acquisition. Preprocessing (such as radiometric calibration, filtering, georegistration, and terrain correction) effectively improves the quality of image data, reduces noise and geometric distortion, and provides a high-quality dataset for subsequent model training. This systematic data processing workflow not only improves efficiency but also reduces the need for manual intervention, making it suitable for large-scale, long-term disaster monitoring applications.
[0097] 4. The flood segmentation model integrates a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer. This multi-layered structure can capture complex features in SAR satellite remote sensing images and improve segmentation accuracy. At the same time, the joint loss function combines the focus loss function and the dice loss function, which effectively solves the common class imbalance problem in flood monitoring (such as the imbalance of the ratio of water body to non-water body pixels), and enhances the robustness and generalization ability of the model. This design not only improves monitoring accuracy but also reduces the risk of overfitting, making the model more adaptable to actual disaster scenarios.
[0098] 5. The flood segmentation model is trained using a dataset, and the model parameters are optimized using a joint loss function, achieving an end-to-end automated learning process. The trained flood segmentation model can be deployed directly without frequent manual adjustments, greatly improving operational efficiency. This automated process lowers the professional threshold, making it easy for non-expert users to apply. At the same time, it supports continuous iterative optimization of the model, adapting to different geographical regions and climatic conditions, enhancing the scalability and practicality of the technology.
[0099] 6. By acquiring SAR satellite remote sensing images of the target area before and after the flood event, and using a trained flood segmentation model for inference and differential calculation, the newly added flood disaster area can be directly identified. This method avoids the subjectivity of relying on manual interpretation in traditional monitoring and improves the accuracy and speed of change detection. Differential calculation can accurately distinguish between permanent water bodies and temporary floods, reduce false alarms, and provide timely and reliable data support for disaster emergency response.
[0100] 7. By generating time-series flood inundation range maps and calculating flood characteristic parameters such as inundation area, inundation depth, and inundation volume, dynamic monitoring and trend analysis of flood disasters are realized. This time-series analysis capability allows users to track the evolution of disasters and supports early warning and disaster assessment. Combined with parameterized output, it not only provides spatial range information but also quantifies disaster intensity, providing a scientific basis for disaster prevention and mitigation decisions (such as evacuation planning and resource allocation), and improving the comprehensiveness and foresight of overall disaster management.
[0101] 8. By acquiring historical SAR satellite remote sensing images with a spatial resolution higher than 1 meter, this high-resolution data can capture more subtle surface changes and significantly improve the accuracy of flood disaster range identification. Compared with traditional low-resolution images, it can more accurately distinguish water body boundaries and minute features, reducing the risk of missed or false detections, thus providing more reliable data support in disaster emergency response and loss assessment, reflecting the cutting-edge nature and practicality of the technology.
[0102] 9. Historical SAR satellite remote sensing images carry multiple data dimensions such as backscatter intensity, metadata, polarization information, spatial and geometric information. This comprehensive information allows for cross-validation and in-depth analysis. For example, polarization information helps distinguish water body types, and the temporal information in the metadata supports time series analysis, thereby improving the robustness and adaptability of flood monitoring, reducing the uncertainty that may be caused by a single data source, and enhancing the comprehensiveness of the method and its industrial application value.
[0103] 10. By specifying in detail the preprocessing steps such as radiometric calibration, filtering removal, georegistration, and terrain correction, a standardized data processing workflow has been formed. This systematic processing can effectively eliminate noise, geometric distortion, and terrain effects in SAR satellite remote sensing images, ensure data consistency and comparability, lay a high-quality foundation for subsequent monitoring, and improve the repeatability and efficiency of the method.
[0104] 11. The specific algorithm selections, such as using the Lee Filter algorithm for filtering and removing backscattering, and using mathematical formulas to calculate the backscattering coefficient and convert it to decibel values in radiometric calibration, reflect technical optimization. The Lee Filter can effectively suppress speckle noise without losing details, while the decibel conversion facilitates data standardization. These optimizations not only improve processing speed but also enhance the reliability of the results.
[0105] 12. Geographic registration and terrain correction using orbital ephemeris data and digital elevation models, including 3D geometric reconstruction and local incident angle correction, can effectively address monitoring challenges in mountainous or undulating terrain. This meticulous terrain correction mechanism reduces errors such as terrain shading and perspective shrinkage, enabling the method to maintain high accuracy in complex geographical environments, highlighting its adaptability and innovation, and helping to expand its application scope.
[0106] 13. By labeling preprocessed SAR satellite remote sensing images with water body pixels, non-water body pixels, and easily confused areas (such as wind-blown water surfaces and vegetation-covered water areas), a dataset is constructed. This targeted labeling helps train a more intelligent classification model, reduces common confusion problems in flood monitoring, and improves the accuracy and generalization ability of automated monitoring. It reflects the forward-looking and practical value of the technology in artificial intelligence-assisted disaster management.
[0107] 14. By integrating the SAR image feature extraction module, metadata encoding module, and polarization feature enhancement module through a multimodal feature extraction layer, the system fully leverages the diversity of SAR satellite remote sensing data. The SAR image feature extraction module employs residual networks and dilated convolutions to extract multi-scale features, which, combined with feature pyramid integration, can capture terrain details at different scales. The metadata encoding module uses a multilayer perceptron and self-attention mechanism to map auxiliary data (such as time and location) into high-dimensional features and capture dependencies. The polarization feature enhancement module focuses on water-sensitive polarization channels to enhance information specificity. This multimodal design avoids the limitations of a single data source, improves feature richness and model robustness, and can more accurately identify flood extent and reduce misjudgment rate, especially in complex environments (such as cloud cover or terrain changes).
[0108] 15. The modality fusion layer, based on a spatial alignment module and a dynamic weight fusion module, achieves intelligent feature integration. The spatial alignment module uses deformable convolution to correct residual geometric distortion and extends metadata features to the spatial dimension, ensuring spatial consistency of multimodal features. The dynamic weight fusion module calculates relevance weights through a cross-attention mechanism and dynamically fuses features using gated recurrent units, adaptively adjusting the contribution of each modality. This fusion method avoids the rigidity of traditional methods (such as simple stitching) and can optimize feature combinations according to specific scenarios, thereby improving segmentation accuracy. For example, in flood monitoring, it can effectively combine the texture information of images and the contextual information of metadata to improve the detection capability of weak water body signals.
[0109] 16. The feature enhancement layer introduces a self-supervised contrastive learning module and an enhanced terrain attention module to further optimize feature quality. The self-supervised contrastive learning module constructs positive and negative sample pairs through data augmentation and uses a contrastive loss function to enhance the water body discrimination of the fused features, enabling the model to better distinguish flooded areas from similar backgrounds (such as shadows or vegetation). The enhanced terrain attention module calculates attention weights through geometric information and incorporates terrain information into the features using conditional batch normalization, improving adaptability to terrain changes. This enhancement mechanism not only improves the model's discriminative ability but also enhances its generalization ability, enabling it to maintain stable performance under different geographical regions and seasonal conditions, making it suitable for large-scale flood disaster monitoring.
[0110] 17. The segmentation prediction layer employs a multi-scale segmentation decoding module and an uncertainty estimation module to achieve high-precision segmentation and reliability assessment. The multi-scale segmentation decoding module is based on the U-Net decoder structure, combining upsampling and skip connections to output a high-resolution water distribution map, ensuring detail preservation. The uncertainty estimation module uses Monte Carlo dropout to sample multiple times during inference, calculates the prediction variance as uncertainty, and filters uncertain regions. This provides a reliability index for the segmentation results, allowing users to adjust the output according to the uncertainty threshold, reducing decision-making risks. For example, in emergency response, this design can output a more reliable flood extent map, supporting precise allocation of rescue resources and enhancing the practical value of the technical solution.
[0111] 18. The joint loss function combines the focus loss function and the dice loss function, effectively solving the class imbalance and optimization challenges in flood segmentation. The focus loss emphasizes the learning of difficult samples (such as a small number of water pixels) through weight coefficients and focus parameters, reducing the background dominance problem. The dice loss directly optimizes the segmentation overlap, improving the consistency between prediction and true label. This joint design balances accuracy and robustness, enabling faster convergence and improved model performance during training. Especially in flood disaster scenarios with sparse water pixels, it can significantly reduce false positive and false negative rates, enhancing the accuracy of monitoring results.
[0112] 19. By employing a spatiotemporal partitioning method, the dataset is divided into a training set, a validation set, and a test set. The validation set is newer than the training set, and the test set is newer than the validation set. Furthermore, all three sets originate from geographically non-overlapping regions. This partitioning method ensures that the model training, validation, and testing processes are independent in both time and space, effectively avoiding data leakage and overfitting issues. It also enhances the model's generalization ability to unknown spatiotemporal scenarios, making flood disaster monitoring results more reliable. This method is particularly suitable for dynamically changing natural disaster scenarios, thus enhancing its practicality and robustness.
[0113] 20. The Tree-structured Parzen Estimator (TPE) performs hyperparameter search on the training and validation sets to automatically obtain the optimal combination of hyperparameters (such as learning rate, batch size, etc.). This is an efficient method based on Bayesian optimization, which significantly reduces the cost and time of manual hyperparameter tuning. By intelligently exploring the hyperparameter space, it quickly converges to the optimal configuration, thereby improving model performance (such as segmentation accuracy and training efficiency).
[0114] 21. During training, a joint loss function is used to calculate the training loss value, and early stopping is monitored through backpropagation and validation sets until a preset condition is triggered. This effectively prevents model overfitting and ensures that the training process terminates when the validation set performance is optimal, thereby improving training efficiency and model stability.
[0115] 22. The model is evaluated using multiple metrics such as precision and recall calculated on the test set, and deployed using the TensorRT inference engine, ensuring that the model has been rigorously validated and has high reliability. At the same time, the TensorRT engine optimizes the inference speed, making it suitable for real-time or near real-time flood disaster monitoring applications.
[0116] 23. By integrating the streaming computing engine and the TensorRT inference engine, real-time preprocessing and model inference of high-resolution SAR satellite remote sensing images are realized. The streaming computing engine supports continuous data stream processing, avoiding the latency of traditional batch processing, while TensorRT optimizes the inference speed of deep learning models, significantly improving the response efficiency of flood monitoring. This design enables the system to quickly generate water distribution maps when disasters occur, providing timely support for emergency decision-making, which is superior to traditional methods that rely on offline processing.
[0117] 24. By combining differential operations with morphological opening and closing operations, intelligent analysis of time-series water distribution maps is achieved. Differential operations can accurately capture changes in flood range, while morphological operations improve the robustness of the results through noise reduction and edge refinement. This time-series analysis method can not only identify newly added disaster ranges, but also dynamically track the evolution process, providing continuous data support for disaster assessment, which is superior to static monitoring methods.
[0118] 25. By setting up monitoring threshold groups and automatically pushing early warning notifications, the entire process of automated monitoring is achieved; the large display screen shows the time sequence and prediction results in real time, improving the accessibility of information, while the threshold triggering mechanism ensures that the management terminal is notified immediately when the disaster exceeds the limit. This integrated design reduces manual intervention, improves the level of automation in disaster response, is suitable for large-scale monitoring scenarios, and reduces operating costs.
[0119] 26. The flood segmentation model innovatively integrates a multimodal feature extraction layer, a modality fusion layer, and a feature enhancement layer. By utilizing techniques such as residual networks, self-attention mechanisms, and dynamic weight fusion, it effectively combines different data sources (such as image features, metadata, and polarization information). The joint loss function combines focus loss and dice loss to solve the class imbalance problem and improve the model's ability to handle difficult samples (such as easily confused regions). Through hyperparameter search and early stopping monitoring, the model training process optimizes parameters, ensuring efficient convergence and generalization performance, thereby improving overall monitoring efficiency.
[0120] 27. By integrating high-resolution SAR satellite remote sensing image preprocessing, multimodal feature fusion-based flood segmentation models, and real-time inference mechanisms, high-precision and reliable flood disaster monitoring was achieved. Its advantages include improving data quality through radiometric calibration and terrain correction, optimizing model performance through multimodal feature extraction and dynamic weight fusion, ensuring the accuracy and robustness of water body segmentation; simultaneously, it supports real-time differential computation and morphological processing, enabling dynamic identification of newly added flood areas and providing prediction and early warning through ConvLSTM networks, combined with comprehensive parameter analysis such as inundation area and water depth, providing multidimensional disaster assessment; furthermore, the model's generalization ability is enhanced through spatiotemporal data partitioning and hyperparameter optimization, while modular design and TensorRT deployment ensure scalability and ease of use, significantly improving the efficiency and scientific rigor of disaster response overall.
[0121] While specific embodiments of the present invention have been described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and not intended to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for monitoring the extent of flood disasters based on satellite remote sensing imagery, characterized in that: Includes the following steps: Step S1: The server acquires a large number of historical SAR satellite remote sensing images with spatial resolution higher than a preset resolution threshold. The historical SAR satellite remote sensing images are preprocessed, including at least radiometric calibration, filtering and removal, georegistration and terrain correction. The preprocessed historical SAR satellite remote sensing images are labeled to construct a dataset. Step S2: The server creates a flood segmentation model based on a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer, and sets the joint loss function of the flood segmentation model; the joint loss function combines the focus loss function and the dice loss function; Step S3: The server trains the flood segmentation model using the dataset. During the training process, the model parameters of the flood segmentation model are optimized using the joint loss function. The trained flood segmentation model is then deployed. Step S4: The server acquires target SAR satellite remote sensing images of the target area before and during the flood event, with a resolution higher than the resolution threshold. After preprocessing each target SAR satellite remote sensing image, it inputs the deployed flood segmentation model for inference to obtain a water distribution map. Step S5: The server performs a difference operation on the adjacent water body distribution maps to obtain the newly added flood disaster area; Step S6: The server generates a time-series flood inundation range map of the target area based on each of the flood disaster ranges, and calculates flood characteristic parameters including inundation area, inundation depth and inundation volume. Based on the time-series flood inundation range map and flood characteristic parameters, the server performs dynamic monitoring and early warning of the flood disaster range.
2. The method for monitoring the extent of flood disasters based on satellite remote sensing imagery as described in claim 1, characterized in that: Step S1 specifically involves: The server acquires a large number of historical SAR satellite remote sensing images with a spatial resolution higher than a preset resolution threshold; the resolution threshold is 1 meter; the historical SAR satellite remote sensing images carry at least backscatter intensity, metadata, polarization information, spatial and geometric information; the metadata includes at least imaging geometric parameters, time information, location information, data processing level, and polarization mode; the spatial and geometric information includes at least spatial resolution and geographic coordinates; Each of the aforementioned historical SAR satellite remote sensing images is subjected to preprocessing including at least radiometric calibration, filtering and removal, georegistration, and terrain correction. The radiometric calibration specifically involves: extracting the pixel amplitude values DN from historical SAR satellite remote sensing images; calculating the backscattering coefficient σ° based on the amplitude values DN and the calibration constant K; converting the unit of the backscattering coefficient σ° to decibels to obtain the decibel value σ°_dB of the backscattering coefficient σ°, thus completing the radiometric calibration. σ° = DN² / K; σ°_dB = 10 * log10(σ°); The filtering and removal process employs the Lee Filter algorithm. The georegistration specifically involves georegistering historical SAR satellite remote sensing images using orbital ephemeris data and digital elevation models. The terrain correction specifically involves: reconstructing the three-dimensional geometric relationship of historical SAR satellite remote sensing images using orbital ephemeris data and digital elevation models, identifying and correcting distortions to perform geometric terrain correction, calculating the local incident angle of each pixel in the geometrically corrected historical SAR satellite remote sensing images, and normalizing the brightness of the local incident angle using the cosine correction method to perform radiation terrain correction. The preprocessed historical SAR satellite remote sensing images are labeled with at least water body pixels, non-water body pixels, and easily confused regions to construct a dataset; the water body pixels are permanent water body pixels or flooded water body pixels; the easily confused regions include at least wind-blown water surfaces, vegetation-covered water areas, and flat non-water surfaces.
3. The method for monitoring the extent of flood disasters based on satellite remote sensing imagery as described in claim 1, characterized in that: Step S2 specifically involves: The server creates a flood segmentation model based on a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer. The joint loss function of the flood segmentation model is set, which combines the focus loss function and the dice loss function. The multimodal feature extraction layer is constructed based on the SAR image feature extraction module, the metadata encoding module, and the polarization feature enhancement module; The SAR image feature extraction module is used to extract multi-scale features from SAR satellite remote sensing images through residual networks and dilated convolutions, and integrate the multi-scale features through a feature pyramid to obtain multi-scale image features; the metadata encoding module is used to map metadata in SAR satellite remote sensing images into high-dimensional features through a multilayer perceptron, and capture the dependencies between the high-dimensional features through a self-attention mechanism to obtain metadata features; the polarization feature enhancement module is used to perform channel compression and enhancement of polarization information in SAR satellite remote sensing images through 1x1 convolutions, and focus on water-sensitive polarization channels through a polarization attention mechanism to obtain polarization enhancement features; The modality fusion layer is constructed based on a spatial alignment module and a dynamic weight fusion module; The spatial alignment module is used to perform spatial alignment operations on multi-scale image features, metadata features, and polarization enhancement features through deformable convolution to correct residual geometric distortions and extend the metadata features to the spatial dimension through a broadcast operation. The dynamic weight fusion module is used to calculate the correlation weights between the multi-scale image features, metadata features, and polarization enhancement features output by the spatial alignment module through a cross-attention mechanism. The gated loop unit dynamically fuses the multi-scale image features, metadata features, and polarization enhancement features according to the correlation weights to obtain the fused features. The feature enhancement layer is constructed based on a self-supervised contrastive learning module and an enhanced terrain attention module; The self-supervised contrastive learning module is used to construct positive and negative sample pairs through data augmentation, and to enhance the water body discrimination of the fused features using a contrastive loss function to obtain first-level enhanced features. The enhanced terrain attention module is used to calculate attention weights through geometric information, and then integrate the geometric information into the first-level enhanced features based on the attention weights through conditional batch normalization to obtain the second-level enhanced features. The segmentation prediction layer is constructed based on a multi-scale segmentation decoding module and an uncertainty estimation module; The multi-scale segmentation and decoding module is used to upsample and skip connections the secondary enhancement features through the U-Net decoder structure, and output a high-resolution water distribution map by combining the multi-scale image features. The uncertainty estimation module is used to calculate the prediction variance of the water distribution map as uncertainty by sampling multiple times during inference through Monte Carlo dropout. After filtering the uncertainty region in the water distribution map by the uncertainty and the preset uncertainty threshold, the module outputs a water distribution map carrying the uncertainty. The formula for the joint loss function is: ; ; ; in, Indicates the value of the joint loss function; This represents the value of the focus loss function; This represents the value of the dice loss function; This represents the total number of pixels processed in one training iteration; i represents the pixel index, which ranges from 1 to N, and is used to traverse each pixel; This represents the true category label of the i-th pixel, with a value of 0 or 1; Indicate category Weighting coefficients; This indicates that the i-th pixel predicted by the flood segmentation model belongs to... The probability; γ represents the focusing parameter, a real hyperparameter greater than or equal to 0, used to adjust the weights of easy and difficult samples; This represents the true category label of the i-th pixel; ϵ represents the predicted probability of the i-th pixel; ϵ represents the smoothing constant. This represents twice the area of the intersection between the actual water body pixels and the predicted water body pixels, used to measure the degree of overlap. This represents the sum of the actual total number of water body pixels and the predicted probability of water body pixels, used for normalization.
4. The method for monitoring the extent of flood disasters based on satellite remote sensing imagery as described in claim 1, characterized in that: Step S3 specifically involves: The server divides the dataset into a training set, a validation set, and a test set based on a spatiotemporal partitioning method, such that the time of the validation set is newer than the time of the training set, the time of the test set is newer than the time of the validation set, and the training set, validation set, and test set come from geographically non-overlapping regions. Before training begins, a tree-structured Parzen estimator is used to search for hyperparameters on the training and validation sets to obtain the optimal combination of hyperparameters; the hyperparameters include at least the learning rate, batch size, network depth, network width, Dropout ratio, decay step size, and decay rate. The flood segmentation model is trained on the training set using the hyperparameter combination. During training, the training loss value is calculated using the joint loss function, and the gradient of the training loss value with respect to the model parameters of the flood segmentation model is calculated using the backpropagation algorithm. The model parameters are updated based on the gradient, and early stopping is monitored using the validation set until a preset early stopping condition is triggered to complete the training. The minimum model parameters include: the kernel weights and biases of the convolutional layers in the residual network of the SAR image feature extraction module, and the parameters of the dilated convolution kernels of the dilated convolution; the weights and biases of the fully connected layers in the multilayer perceptron of the metadata encoding module, and the weights of the query matrix, key matrix, and value matrix in the self-attention mechanism; the kernel weights and biases of the 1x1 convolutional layers in the polarization feature enhancement module, and the attention weight parameters in the polarization attention mechanism; the kernel weights and biases of the deformable convolution in the spatial alignment module, and the parameters of the convolutional layers that learn the offset; the projection matrix in the cross-attention mechanism of the dynamic weight fusion module, and the weight matrix and bias vector in the gated recurrent unit; the parameters of the attention weights used to calculate geometric information in the enhanced terrain attention module, and the scaling and offset parameters in the conditional batch normalization; and the kernel weights and biases of all upsampled convolution kernels in the U-Net decoder of the multi-scale segmentation decoding module. Precision, recall, F1-Score, intersection-over-union ratio, overall accuracy, and inference speed are calculated using the test set to test the trained flood segmentation model. The tested flood segmentation model is then deployed using the TensorRT inference engine.
5. The method for monitoring the extent of flood disasters based on satellite remote sensing imagery as described in claim 1, characterized in that: Step S4 specifically involves: Based on the received flood disaster monitoring instructions, the server obtains target SAR satellite remote sensing images of the target area with a resolution higher than the resolution threshold from the remote sensing image database before and during the flood event. After performing preprocessing on each target SAR satellite remote sensing image in sequence through the streaming computing engine, including at least radiometric calibration, filtering and removal, georegistration and terrain correction, the server inputs each target SAR satellite remote sensing image into the deployed flood segmentation model and performs real-time inference through the TensorRT inference engine to obtain the water distribution map. Step S5 specifically involves: The server performs differential operations on the time-adjacent water body distribution maps, combines morphological opening operations for noise reduction, and combines morphological closing operations for edge refinement to obtain the newly added flood disaster range. Step S6 specifically involves: The server generates a time-series flood inundation range map of the target area based on the flood disaster ranges, and calculates flood characteristic parameters including inundation area, inundation depth, and inundation volume; the inundation area is calculated based on the time-series flood inundation range map and spatial resolution; the inundation depth is calculated based on the time-series flood inundation range map and digital elevation model; the inundation volume is calculated based on the inundation area and inundation depth. The server inputs the time-series flood inundation range map into a pre-trained ConvLSTM network to obtain a predicted flood inundation range map. The time-series flood inundation range map, the predicted flood inundation range map, and flood characteristic parameters are displayed in real time on a large screen. The server monitors the time-series flood inundation range map, the predicted flood inundation range map, and the flood characteristic parameters through a preset monitoring threshold group. When the monitoring threshold group is triggered, a flood warning notification is immediately pushed to the management terminal, thereby dynamically monitoring and warning of the flood disaster range.
6. A flood disaster range monitoring system based on satellite remote sensing imagery, characterized in that: Includes the following modules: The dataset construction module is used to acquire a large number of historical SAR satellite remote sensing images with spatial resolution higher than a preset resolution threshold. The module performs preprocessing on each of the historical SAR satellite remote sensing images, including at least radiometric calibration, filtering and removal, georegistration and terrain correction. The preprocessed historical SAR satellite remote sensing images are then labeled to construct the dataset. The flood segmentation model creation module is used by the server to create a flood segmentation model based on a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer, and to set the joint loss function of the flood segmentation model; the joint loss function combines the focus loss function and the dice loss function; The flood segmentation model training module is used by the server to train the flood segmentation model using the dataset. During the training process, the model parameters of the flood segmentation model are optimized through the joint loss function, and the trained flood segmentation model is deployed. The flood segmentation model inference module is used by the server to acquire target SAR satellite remote sensing images of the target area before and during a flood event, with a resolution higher than the resolution threshold. After preprocessing each target SAR satellite remote sensing image, the image is input into the deployed flood segmentation model for inference to obtain a water distribution map. The differential operation module is used by the server to perform differential operations on the distribution maps of adjacent water bodies to obtain the newly added flood disaster range; The flood disaster range monitoring module is used by the server to generate a time-series flood inundation range map of the target area based on each of the flood disaster ranges, and to calculate flood characteristic parameters including inundation area, inundation depth and inundation volume, and to perform dynamic monitoring and early warning of the flood disaster range based on the time-series flood inundation range map and flood characteristic parameters.
7. A flood disaster range monitoring system based on satellite remote sensing imagery as described in claim 6, characterized in that: The dataset construction module is specifically used for: The server acquires a large number of historical SAR satellite remote sensing images with a spatial resolution higher than a preset resolution threshold; the resolution threshold is 1 meter; the historical SAR satellite remote sensing images carry at least backscatter intensity, metadata, polarization information, spatial and geometric information; the metadata includes at least imaging geometric parameters, time information, location information, data processing level, and polarization mode; the spatial and geometric information includes at least spatial resolution and geographic coordinates; Each of the aforementioned historical SAR satellite remote sensing images is subjected to preprocessing including at least radiometric calibration, filtering and removal, georegistration, and terrain correction. The radiometric calibration specifically involves: extracting the pixel amplitude values DN from historical SAR satellite remote sensing images; calculating the backscattering coefficient σ° based on the amplitude values DN and the calibration constant K; converting the unit of the backscattering coefficient σ° to decibels to obtain the decibel value σ°_dB of the backscattering coefficient σ°, thus completing the radiometric calibration. σ° = DN² / K; σ°_dB = 10 * log10(σ°); The filtering and removal process employs the Lee Filter algorithm. The georegistration specifically involves georegistering historical SAR satellite remote sensing images using orbital ephemeris data and digital elevation models. The terrain correction specifically involves: reconstructing the three-dimensional geometric relationship of historical SAR satellite remote sensing images using orbital ephemeris data and digital elevation models, identifying and correcting distortions to perform geometric terrain correction, calculating the local incident angle of each pixel in the geometrically corrected historical SAR satellite remote sensing images, and normalizing the brightness of the local incident angle using the cosine correction method to perform radiation terrain correction. The preprocessed historical SAR satellite remote sensing images are labeled with at least water body pixels, non-water body pixels, and easily confused regions to construct a dataset; the water body pixels are permanent water body pixels or flooded water body pixels; the easily confused regions include at least wind-blown water surfaces, vegetation-covered water areas, and flat non-water surfaces.
8. A flood disaster range monitoring system based on satellite remote sensing imagery as described in claim 6, characterized in that: The flood diversion model creation module is specifically used for: The server creates a flood segmentation model based on a multimodal feature extraction layer, a modality fusion layer, a feature enhancement layer, and a segmentation prediction layer. The joint loss function of the flood segmentation model is set, which combines the focus loss function and the dice loss function. The multimodal feature extraction layer is constructed based on the SAR image feature extraction module, the metadata encoding module, and the polarization feature enhancement module; The SAR image feature extraction module is used to extract multi-scale features from SAR satellite remote sensing images through residual networks and dilated convolutions, and integrate the multi-scale features through a feature pyramid to obtain multi-scale image features; the metadata encoding module is used to map metadata in SAR satellite remote sensing images into high-dimensional features through a multilayer perceptron, and capture the dependencies between the high-dimensional features through a self-attention mechanism to obtain metadata features; the polarization feature enhancement module is used to perform channel compression and enhancement of polarization information in SAR satellite remote sensing images through 1x1 convolutions, and focus on water-sensitive polarization channels through a polarization attention mechanism to obtain polarization enhancement features; The modality fusion layer is constructed based on a spatial alignment module and a dynamic weight fusion module; The spatial alignment module is used to perform spatial alignment operations on multi-scale image features, metadata features, and polarization enhancement features through deformable convolution to correct residual geometric distortions and extend the metadata features to the spatial dimension through a broadcast operation. The dynamic weight fusion module is used to calculate the correlation weights between the multi-scale image features, metadata features, and polarization enhancement features output by the spatial alignment module through a cross-attention mechanism. The gated loop unit dynamically fuses the multi-scale image features, metadata features, and polarization enhancement features according to the correlation weights to obtain the fused features. The feature enhancement layer is constructed based on a self-supervised contrastive learning module and an enhanced terrain attention module; The self-supervised contrastive learning module is used to construct positive and negative sample pairs through data augmentation, and to enhance the water body discrimination of the fused features using a contrastive loss function to obtain first-level enhanced features. The enhanced terrain attention module is used to calculate attention weights through geometric information, and then integrate the geometric information into the first-level enhanced features based on the attention weights through conditional batch normalization to obtain the second-level enhanced features. The segmentation prediction layer is constructed based on a multi-scale segmentation decoding module and an uncertainty estimation module; The multi-scale segmentation and decoding module is used to upsample and skip connections the secondary enhancement features through the U-Net decoder structure, and output a high-resolution water distribution map by combining the multi-scale image features. The uncertainty estimation module is used to calculate the prediction variance of the water distribution map as uncertainty by sampling multiple times during inference through Monte Carlo dropout. After filtering the uncertainty region in the water distribution map by the uncertainty and the preset uncertainty threshold, the module outputs a water distribution map carrying the uncertainty. The formula for the joint loss function is: ; ; ; in, Indicates the value of the joint loss function; This represents the value of the focus loss function; This represents the value of the dice loss function; This represents the total number of pixels processed in one training iteration; i represents the pixel index, which ranges from 1 to N, and is used to traverse each pixel; This represents the true category label of the i-th pixel, with a value of 0 or 1; Indicate category Weighting coefficients; This indicates that the i-th pixel predicted by the flood segmentation model belongs to... The probability; γ represents the focusing parameter, a real hyperparameter greater than or equal to 0, used to adjust the weights of easy and difficult samples; This represents the true category label of the i-th pixel; ϵ represents the predicted probability of the i-th pixel; ϵ represents the smoothing constant. This represents twice the area of the intersection between the actual water body pixels and the predicted water body pixels, used to measure the degree of overlap. This represents the sum of the actual total number of water body pixels and the predicted probability of water body pixels, used for normalization.
9. A flood disaster range monitoring system based on satellite remote sensing imagery as described in claim 6, characterized in that: The flood segmentation model training module is specifically used for: The server divides the dataset into a training set, a validation set, and a test set based on a spatiotemporal partitioning method, such that the time of the validation set is newer than the time of the training set, the time of the test set is newer than the time of the validation set, and the training set, validation set, and test set come from geographically non-overlapping regions. Before training begins, a tree-structured Parzen estimator is used to search for hyperparameters on the training and validation sets to obtain the optimal combination of hyperparameters; the hyperparameters include at least the learning rate, batch size, network depth, network width, Dropout ratio, decay step size, and decay rate. The flood segmentation model is trained on the training set using the hyperparameter combination. During training, the training loss value is calculated using the joint loss function, and the gradient of the training loss value with respect to the model parameters of the flood segmentation model is calculated using the backpropagation algorithm. The model parameters are updated based on the gradient, and early stopping is monitored using the validation set until a preset early stopping condition is triggered to complete the training. The minimum model parameters include: the kernel weights and biases of the convolutional layers in the residual network of the SAR image feature extraction module, and the parameters of the dilated convolution kernels of the dilated convolution; the weights and biases of the fully connected layers in the multilayer perceptron of the metadata encoding module, and the weights of the query matrix, key matrix, and value matrix in the self-attention mechanism; the kernel weights and biases of the 1x1 convolutional layers in the polarization feature enhancement module, and the attention weight parameters in the polarization attention mechanism; the kernel weights and biases of the deformable convolution in the spatial alignment module, and the parameters of the convolutional layers that learn the offset; the projection matrix in the cross-attention mechanism of the dynamic weight fusion module, and the weight matrix and bias vector in the gated recurrent unit; the parameters of the attention weights used to calculate geometric information in the enhanced terrain attention module, and the scaling and offset parameters in the conditional batch normalization; and the kernel weights and biases of all upsampled convolution kernels in the U-Net decoder of the multi-scale segmentation decoding module. Precision, recall, F1-Score, intersection-over-union ratio, overall accuracy, and inference speed are calculated using the test set to test the trained flood segmentation model. The tested flood segmentation model is then deployed using the TensorRT inference engine.
10. A flood disaster range monitoring system based on satellite remote sensing imagery as described in claim 6, characterized in that: The flood segmentation model inference module is specifically used for: Based on the received flood disaster monitoring instructions, the server obtains target SAR satellite remote sensing images of the target area with a resolution higher than the resolution threshold from the remote sensing image database before and during the flood event. After performing preprocessing on each target SAR satellite remote sensing image in sequence through the streaming computing engine, including at least radiometric calibration, filtering and removal, georegistration and terrain correction, the server inputs each target SAR satellite remote sensing image into the deployed flood segmentation model and performs real-time inference through the TensorRT inference engine to obtain the water distribution map. The difference operation module is specifically used for: The server performs differential operations on the time-adjacent water body distribution maps, combines morphological opening operations for noise reduction, and combines morphological closing operations for edge refinement to obtain the newly added flood disaster range. The flood disaster range monitoring module is specifically used for: The server generates a time-series flood inundation range map of the target area based on the flood disaster ranges, and calculates flood characteristic parameters including inundation area, inundation depth, and inundation volume; the inundation area is calculated based on the time-series flood inundation range map and spatial resolution; the inundation depth is calculated based on the time-series flood inundation range map and digital elevation model; the inundation volume is calculated based on the inundation area and inundation depth. The server inputs the time-series flood inundation range map into a pre-trained ConvLSTM network to obtain a predicted flood inundation range map. The time-series flood inundation range map, the predicted flood inundation range map, and flood characteristic parameters are displayed in real time on a large screen. The server monitors the time-series flood inundation range map, the predicted flood inundation range map, and the flood characteristic parameters through a preset monitoring threshold group. When the monitoring threshold group is triggered, a flood warning notification is immediately pushed to the management terminal, thereby dynamically monitoring and warning of the flood disaster range.
Citation Information
Cited By
Method and device for monitoring flood and waterlogging in a basin based on high-orbit SAR satellite
CN122239021A