Wind cloud 3F satellite image natural water area extraction and abnormal expansion monitoring method
By employing intelligent water body extraction algorithms and efficient cloud-based pixel interpolation methods, combined with a multi-dimensional accuracy assessment system, the accuracy and efficiency issues of water body extraction in complex surface environments have been resolved. This enables real-time monitoring and early warning of water body changes, supporting water resource management and ecological environmental protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF DESERT METEOROLOGY CMA URUMQI
- Filing Date
- 2025-11-24
- Publication Date
- 2026-04-28
AI Technical Summary
Existing water extraction methods lack accuracy and reliability when dealing with complex surface environments. In particular, cloud and shadow interference in medium-resolution remote sensing images leads to a large number of errors in water extraction results. Furthermore, data processing efficiency is low, making it difficult to meet the needs of real-time or near-real-time monitoring. There is also a lack of a multi-dimensional accuracy evaluation system and a real-time monitoring and early warning system.
An intelligent water body extraction algorithm is adopted, combined with an efficient cloud-based pixel interpolation method and a multi-dimensional accuracy evaluation system. Water bodies are classified using the XGBoost algorithm, and data processing is performed using GPU parallel acceleration technology. This enables the generation of daily cloudless multi-channel reflectivity data and real-time monitoring of abnormal water body expansion.
It significantly improves the accuracy and efficiency of water body extraction, enables real-time monitoring and early warning of abnormal water body expansion events, provides high-precision data support, and provides timely and accurate information support for water resource management and ecological environment protection.
Smart Images

Figure CN121937889A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of remote sensing image data processing technology, specifically relating to a method for extracting natural water bodies and monitoring anomalies in Fengyun-3F satellite images. Background Technology
[0002] With the intensification of global climate change and human activities, the rational utilization and protection of water resources has become a global focus. Changes in natural water bodies have a significant impact on regional ecological environment and socio-economic development. Satellite remote sensing provides a key technological means for large-scale, long-term continuous monitoring of natural water bodies. Among them, the Fengyun-3F satellite, as my country's new generation polar-orbiting meteorological satellite, carries multiple sensors that can provide high-resolution, high-precision remote sensing data, providing technical support for the monitoring of natural water bodies. This patent aims to develop an intelligent technology for the extraction, accuracy assessment, and anomaly expansion monitoring of natural water bodies using Fengyun-3F satellite imagery (FY-3F / MERSI-III standard products) to meet the needs of water resource management and ecological environment protection.
[0003] Threshold-based water index methods: These methods utilize water indices such as the Normalized Difference Water Index (NDWI) to identify water bodies by setting thresholds. While simple and fast, these methods have low accuracy in complex surface environments (such as cloud cover and mixed land cover). For example, Klein (2024) et al. developed the Global Water Pack product using MODIS 250m daily surface reflectance data. This automatically generates daily, monthly, and annual masks and water frequency maps of inland water bodies worldwide since 2000, providing a high-temporal-resolution dataset for characterizing the spatiotemporal evolution of global surface water.
[0004] Classification methods based on machine learning: These methods utilize machine learning algorithms such as Support Vector Machine (SVM) and Random Forest (RF) to extract water body features from training samples, thereby achieving water body classification. This approach improves the accuracy of water body extraction to some extent, but it requires high-quality and high-quantity training samples. For example, Yue (2023) et al. proposed a fully automated, high-precision water body mapping framework based on GEE + RF, which automatically constructs training samples and distinguishes between permanent and seasonal water bodies.
[0005] Deep learning-based segmentation methods utilize deep learning algorithms such as convolutional neural networks (CNNs) to extract water bodies through pixel-level segmentation. This method demonstrates high accuracy when handling complex surface environments, but it consumes significant computational resources. For example, Chen et al. (2018) used convolutional neural networks to extract urban water bodies from high-resolution remote sensing images, significantly improving segmentation accuracy in complex backgrounds. Mullen et al. (2023) used 3 m Planet high-resolution imagery and deep learning methods to monitor seasonal area changes in small lakes and ponds in Alaska, highlighting the advantages of deep learning in the dynamic monitoring of small water bodies.
[0006] Under-cloud pixel interpolation methods: Traditional under-cloud pixel interpolation methods are mainly based on linear interpolation of time series or simple interpolation of spatial neighborhood. These methods are not effective when dealing with large-scale cloud cover. For example, Bai (2022) et al. proposed a method for filling gaps in time series images of surface water based on spatiotemporal characteristics. By simultaneously utilizing temporal variations and spatial neighborhood information, it significantly improves the accuracy of under-cloud water body restoration. Chen (2016) proposed a linear interpolation method based on time series for filling data in cloud-covered areas.
[0007] Accuracy assessment methods: Existing accuracy assessment methods mainly rely on a single data source or simple validation metrics, such as overall accuracy (OA) and Kappa coefficient. These methods have limitations in assessing the accuracy and reliability of water body extraction results. For example, Olofsson et al. (2014) proposed a probability sampling-based validation method for assessing the accuracy of land cover change.
[0008] Challenges of High-Precision Water Body Extraction: Existing water body extraction methods suffer from insufficient accuracy and reliability when dealing with complex surface environments (such as cloud cover and mixed land cover). This is especially true in medium-resolution remote sensing imagery, where cloud cover and shadows cause significant errors in water body extraction results. Traditional water body extraction algorithms (such as threshold-based NDWI methods) are poorly adaptable to different seasons, lighting conditions, and land cover types, making it difficult to achieve high-precision water body extraction.
[0009] Improving data processing efficiency: Existing data processing methods are inefficient when handling large-scale remote sensing data, making it difficult to meet the needs of real-time or near-real-time monitoring. This is especially true when processing large-scale, high temporal resolution data from the Fengyun-3F satellite, where the data volume is enormous and processing time is long. The lack of efficient under-cloud pixel interpolation methods leads to data loss in cloud-covered areas, affecting the completeness and continuity of water area extraction.
[0010] Comprehensiveness of accuracy assessment: Existing accuracy assessment methods often rely on a single data source or simple validation metrics, making it difficult to comprehensively evaluate the accuracy and reliability of water body extraction results. There is a lack of a multi-dimensional, multi-level comprehensive accuracy assessment system. Furthermore, there is a lack of systematic validation methods for water body extraction results across different time scales (such as seasonal variations) and different spatial scales (such as small and large water bodies).
[0011] Real-time monitoring of abnormal water body expansion: Existing monitoring technologies lag behind in identifying and issuing early warnings of abnormal water body expansion, making it difficult to promptly detect and respond to sudden water body expansion events such as floods and lake expansion. The lack of a real-time monitoring and early warning system based on high-frequency remote sensing data hinders the provision of timely and accurate information support for water resource management and emergency response. Summary of the Invention
[0012] To overcome the above technical problems, the present invention aims to provide a method for natural water area extraction and anomaly expansion monitoring from Fengyun-3F satellite imagery. This method significantly improves the accuracy and efficiency of natural water area monitoring through intelligent water area extraction algorithms, efficient cloud-based pixel interpolation methods, multi-dimensional accuracy evaluation systems, and real-time monitoring and early warning systems, thus meeting the practical needs of water resource management and ecological environment protection.
[0013] The technical solution adopted in this invention is: The method for extracting natural water bodies and monitoring anomalous expansion from Fengyun-3F satellite imagery includes the following steps; Step 1: Under-cloud pixel interpolation: To acquire FY-3F / MERSI-III L1 level daily multi-channel reflectance data at 250m or 1km, a strategy of labeling, interpolation, and iterative extrapolation is adopted to solve the problem of invalid values in FY-3F optical data caused by cloud cover, and output spatiotemporally continuous daily cloud-free multi-channel reflectance data (at two resolutions: 250m and 1km). Step 2: Intelligent water area extraction: Adaptive water body extraction was performed on the multi-channel cloudless reflectance data output in step 1 to generate binary classification results (1-water, 0-non-water) and optimized vector files at two resolutions: 250m and 1km. Step 3: Multi-dimensional accuracy assessment: Perform independent verification on the output of step 2 to ensure that it meets the requirements of the monitoring application; Step 4: Monitoring of abnormal water body expansion: Based on the verified reliable water body time series classification product in step 3, conduct detection of abnormal expansion of natural water bodies.
[0014] Step 1 specifically involves: (1) Data reading: Acquire FY-3F / MERSI-III L1 level 250m or 1km daily scale multichannel reflectivity data and CLM cloud mask products; (2) Cloud pixel marking: Use the CLM cloud mask of Fengyun 3F satellite to mark invalid pixels, output the location of the pixels marked as invalid (cloud covered), and proceed to step (3). (3) Spatiotemporal cube construction: Take one day before and after in the time dimension, and the initial spatial dimension is a 7×7 pixel window to construct a spatiotemporal cube; output a three-dimensional data block containing the reflectance values of the target pixel and its spatiotemporal neighboring pixels, and proceed to step (4). (4) IDW first round of filling: Search for valid observation pixels within the spacetime cube and reconstruct missing data using the IDW interpolation method.
[0015] The interpolation is given by the following formula: ; In the formula, y i Let d be the value of the i-th effective neighboring pixel. i Let be the distance between the i-th effective neighboring pixel and the target pixel. α This is the distance decay exponent (usually taken as 2). N Set the lower limit of the number of effective neighboring pixels (if insufficient, mark it as waiting for the next iteration); output the interpolated reflectance value or the mark to be iterated, and proceed to step (4). (5) Multi-round iterative extrapolation and convergence conditions: For large-scale continuous data-free areas, a multi-round iterative extrapolation mechanism is adopted. After each round of interpolation, the effective pixel set is updated, and the spatiotemporal search radius is gradually expanded (maximum expansion to a 15×15 pixel window and 15 days before and after), until the filling rate is <0.1% for 3 consecutive rounds or the maximum number of iterations is reached (30 rounds for 250m resolution, 20 rounds for 1km resolution); output multi-channel reflectivity data with complete cloud pixel filling.
[0016] High efficiency: Based on GPU parallel acceleration, it significantly improves the data filling rate, with a stable monthly data filling rate exceeding 98%. Detail preservation: While preserving spatial details, it effectively suppresses cloud pollution, improving data integrity and continuity.
[0017] Step 2 specifically involves: water body extraction, from reflectance to a binary water body map, including the following steps: (1) Feature extraction: Using the multi-channel reflectance data of FY-3F250m and 1000m under cloud interpolation or under the original clear sky conditions without clouds, the data are read in block by block to extract features related to water bodies, including four bands: red light (R), green light (G), blue light (B) and near infrared (NIR). The R and G bands directly reflect the coverage and health status of surface vegetation, the B band is sensitive to atmospheric scattering, and the NIR band is highly sensitive to vegetation biomass and water content. In the feature calculation stage, the two basic spectra of R and G are retained from the original bands, and the normalized water index (NDWI) is further calculated to highlight water bodies and suppress vegetation and soil noise. At the same time, the brightness feature is synthesized by multi-band weighting to characterize the overall surface reflectance intensity. Finally, the five features of R, G, NIR, NDWI and Brightness, which have both physical meaning and discriminative ability, are used as the original input features of XGBoost.
[0018] The specific steps for feature extraction are as follows: Normalized Difference Water Index (NDWI): ; Brightness Index: ; Function: To distinguish highly reflective features (such as bare soil and buildings), while water bodies typically have lower reflectivity (<0.2).
[0019] (2) Sample and model preparation: First, relying on global free satellite resources, Landsat8, Landsat9 and Sentinel2 optical images were selected, and Sentinel1SAR was introduced to compensate for the influence of clouds and rain, and radiometric, atmospheric, orthorectification and projection unification were completed; then, the water body contour was manually vectorized in QGIS, and non-water body samples in easily confused areas were added. The initial mask was generated by combining NDWI, MNDWI (improved normalized difference water index), AWEI (automatic water extraction index) and Otsu threshold (Otsu method), and then manually corrected; the label adopts binary raster, water body 1, non-water body 0, and mixed pixels are marked as ignored values and masked during training; ; Function: MNDWI enhances water body information and suppresses non-water features (especially building shadows) by using the normalized ratio of the green band to the shortwave infrared band. The index ranges from [-1, 1], with a larger positive value indicating a higher probability of water bodies and a negative value usually corresponding to non-water features.
[0020] AWEI contains AWEI nsh (Shadowless version) and AWEI sh (Shadowed version), suitable for different scenarios: AWEI nsh = 4×(Green−SWIR1) −(0.25×NIR+2.75×SWIR2); AWEI sh= Blue+2.5×Green−1.5 × (NIR+SWIR1) − 0.25 × SWIR 2; In the formula, Blue is the blue light band; Green is the green light band; NIR is the near-infrared band; SWIR1 is the short-wave infrared band 1 (such as Landsat B6 / Sentinel-2 B11); SWIR2 is the short-wave infrared band 2 (such as Landsat B7 / Sentinel-2 B12). Function: AWEI leverages the characteristic that water has higher reflectivity in the green / blue light bands than in the NIR / SWIR bands. It enhances the positive values of water pixels through linear combination, while making non-water bodies (especially shadows and dark surfaces) appear with larger negative values. nsh Suitable for general scenarios, AWEI sh The accuracy of water body recognition in shaded areas has been specifically optimized.
[0021] Post-processing removes false water bodies through small patch filtering, morphological opening and closing operations, and shape and area constraints. The sample library archives image patches, labels, metadata, and version records according to a unified standard, forming a high-quality, traceable dataset. Samples are divided into three categories: non-water bodies, water bodies, and others. They are input into XGBoost to construct a gradient boosting tree model, balancing the class ratio with a positive-to-negative sample weight ratio (scale_pos_weight), and using five-fold cross-validation to suppress overfitting. No retraining is required during the inference stage; if the accuracy is not up to standard, a new model is resampled and trained in step 2. (3) XGBoost water body classification: The XGBoost algorithm is used to construct an intelligent water body extraction model. The XGBoost algorithm uses second-order gradient information to accelerate model convergence, and introduces a regularization term to prevent overfitting. Finally, three products are output simultaneously: water body binary raster (.tif), water body vector file (.shp), and true color overlay boundary map (.jpg). Raster and vector data can be used as input data for step (3); the key process is as follows: 1. For the 5-dimensional features of step (1) above, speed up the process to obtain 3 types of probabilities. 2. Take category ID=2 as the initial water body mask (bool array). Output the block water body mask and proceed to step (4) vectorization optimization step; (4) Optimization of water body distribution vectorization: The morphological opening operation of 3×3 structuring elements is used to smooth the water body boundary and remove small noise. Then, connected component analysis is performed to set the small patches with an area smaller than MIN_PATCH_PIXELS (50 pixels) to zero, convert the raster to vector, merge adjacent patches and simplify the topology with pixel resolution as the threshold, and simultaneously write the LZW compressed binary raster _water.tif, the water body vector *_water.shp with the same CRS as the original image, and the cartograph-grade true color overlay map *_shp_overlay.jpg with the red vector line superimposed on the RGB base map. After completion, proceed to step (3) quality inspection.
[0022] Algorithm advantages: High precision: Combining machine learning algorithms significantly improves the precision of water body extraction and adapts to different imaging conditions.
[0023] Adaptability: The model can automatically learn and adapt to different surface environments, reducing human intervention.
[0024] Step 3 specifically involves multi-dimensional accuracy assessment, which includes the following steps: (1) Pixel-level accuracy evaluation (based on raster): After registering the binary raster of water body with the reference water body raster of the same high-resolution image (Gaofen or Sentinel-2) manually interpreted, the pixels are compared one by one, the confusion matrix is calculated, and indicators such as OA, Kappa, IoU, and F1-score are calculated. (2) Validation based on vector water body data: Using Xinjiang regional water body vector data as the reference true value, spatial overlap analysis is carried out to calculate the intersection-over-union ratio (IoU), area difference (ΔA) and water body boundary deviation, etc., to evaluate the accuracy and consistency of the dataset in spatial structure; or the water body boundary vector output in step 2 is directly spatially superimposed with the above-mentioned reference water body vector to calculate the IoU, boundary overlap degree and average distance deviation, etc., to quantitatively analyze the degree of agreement between the water body range and the boundary position; (3) Ground verification based on measured water body sample points: Independent verification sample points (water / non-water bodies) obtained from field GPS measurements or high-resolution aerial image interpretation, or collected measured water body sample point data in Xinjiang region, are spatially overlaid with classification rasters to construct confusion matrix and calculate accuracy indicators such as overall accuracy (OA), Kappa coefficient, user accuracy (Precision) and cartographer accuracy (Recall) to verify the authenticity of the dataset from the ground observation level; (4) Validation based on long-term water body probability map: The probability map of water body occurrence is generated using MODIS satellite multi-temporal images. The overall accuracy (OA), precision, recall, F1 score, Kappa coefficient, intersection-over-union ratio (IoU) and other accuracy indicators are calculated to evaluate the stability and robustness of the dataset under long-term series conditions.
[0025] The formula for calculating the index in step (4) is as follows: Overall Accuracy (OA) reflects the proportion of all pixels that are judged as correct; OA = (TP + TN) / (TP + TN + FP + FN); Precision (or accuracy) is the percentage of pixels that the model identifies as water bodies that are actually water bodies. Precision = TP / (TP + FP). Recall (also known as recall rate or sensitivity) is the proportion of "real water" pixels that are successfully retrieved by the model; Recall = TP / (TP + FN). The F1 score is the harmonic mean of Precision and Recall, which balances accuracy and completeness; F1 = 2·Precision·Recall / (Precision + Recall); The Kappa coefficient (κ) is a consistency index after removing the "random guessing" part. The closer it is to 1, the more the classification result matches the real scene. First, calculate the "random consistency rate" Pe: Pe=[(TP+FN)(TP+FP)+(TN+FP)(TN+FN)] / (TP+TN+FP+FN)². Then calculate Kappa: κ=(OA–Pe) / (1–Pe); The IoU is the ratio of the overlapping area of the "predicted water body" and the "true water body" to the combined area of the two. The closer it is to 1, the more accurately the model captures the boundary of the water body; IoU=|A∩B| / |A∪B|=TP / (TP+FP+FN); In the formula, TP (True Positive) represents the number of pixels that are actually water bodies and correctly identified as water bodies by the model. TN (True Negative) represents the number of pixels that are actually non-water bodies and correctly identified as non-water bodies by the model. FP (False Positive) represents the number of pixels that are actually non-water bodies but misidentified as water bodies by the model ("false alarms"). FN (False Negative) represents the number of pixels that are actually water bodies but misidentified as non-water bodies by the model ("false misses"). A: The set of "water body" pixels in the reference ground truth; B: The set of "water body" pixels predicted by the model; Intersection |A∩B|=TP (the number of pixels correctly identified as water bodies by the model); Union |A∪B|=TP+FP+FN (the sum of all pixels "judged as water bodies" or "actually water bodies", i.e., "multiple detections + false misses + correct detections").
[0026] Evaluation Advantages: Comprehensiveness: The performance of the dataset is systematically verified from three dimensions: temporal stability, spatial geometric consistency, and ground realism, constructing a multi-dimensional and multi-level comprehensive evaluation system. Reliability: Multiple verification methods are combined to ensure the representativeness and stability of the evaluation results.
[0027] Step 4 specifically involves: monitoring abnormal water body expansion. This is achieved through an intelligent water body extraction algorithm, which monitors changes in the water body in real time and identifies abnormal water body expansion events, such as floods and lake expansion. This includes the following steps: (1) Input data that meets the multi-dimensional accuracy thresholds (OA≥90%, IoU≥0.75, F1≥0.85) of step 3, and daily 250m and 1km binary water body mask (GeoTIFF). (2) Data preprocessing: Data standardization: uniformly resampled to WGS-84 / UTM area projection, set pixel depth to 8 bits, water pixels to 1, land pixels to 0, background fill value to 255; Time series construction: sorted by "year-day order", missing dates were filled by linear interpolation, and a three-dimensional Boolean cube of T×H×W was generated (T: number of days; H, W: number of rows and columns); (3) Adaptive benchmark setting for benchmark water bodies: For example, take the 75th quantile of the frequency cube of water bodies during the cloudless period from 2019 to 2023 to generate a probability layer (Pbase) for "normal water bodies". Seasonal benchmark: If the date information is complete, the quantile is calculated again by sliding the window by "ten-day period" to obtain 36 seasonal benchmarks (Pseason) to eliminate seasonal fluctuations such as the ice-covered period and the irrigation period; (4) Flood detection (multi-mode joint): Absolute increment detection: ΔW=Pcurrent−Pbase, trigger threshold 0.05. Relative growth detection: R=ΔW / max(Pbase,0.1), trigger threshold 10%; New water body detection: Pbase<0.1 and Pcurrent>0.3. High water level detection: Pcurrent>0.8 and lasts for ≥5 days; Multi-scale intensity: Perform 3×3, 5×5, 7×7 neighborhood morphological opening and closing reconstruction on the trigger pixel, and calculate the intensity index I=Σ(ΔW·A), where A is the area of the connected domain (km²); (5) Event identification: Area decline criteria: If the flooded area on the day is less than 30% of the peak area, it is marked as "entering decay"; Continuous low water level: If the flooded area is less than 5km² for 3 consecutive days, it is judged as the end of the event; Event segmentation: If the interval between two adjacent peaks is ≥7 days and the area decline is >70%, it is automatically split into independent events; (6) Administrative matching: Overlay county-level administrative division vectors and output the code and name of the corresponding district / county; The monitoring results of abnormal water body expansion include detailed information on flood events, such as event ID, start date, peak date, end date, duration, peak inundation area, latitude and longitude of the flood center, and severity. The dynamic threshold is the mean area over time ±2σ. When the water body area observed in a single observation exceeds this range, an abnormal event is triggered, and it is further distinguished into two categories: "flood / lake expansion" and "drought / shrinkage". Visual charts: flood time series charts, flood event statistics charts; Report documents: Text reports and Excel spreadsheets containing event statistics and accuracy metrics.
[0028] System advantages: Real-time performance: Enables real-time processing and analysis of high-resolution remote sensing data, allowing for timely detection and response to abnormal water body expansion events. Accuracy: Combines intelligent water area extraction algorithms and a multi-dimensional accuracy assessment system to ensure high precision and reliability of monitoring results.
[0029] The beneficial effects of this invention are: Improve the accuracy of water extraction; High-precision extraction: Addressing the core bottlenecks such as severe cloud pixel contamination in optical images, the reliance on manual scene-by-scene adjustment of water extraction thresholds which is labor-intensive, and significant differences in results under different seasons, lighting conditions, and terrain features, this invention achieves adaptive and accurate identification in complex environments by deeply integrating intelligent water extraction algorithms with machine learning technology, significantly improving the accuracy, efficiency, and robustness of water extraction.
[0030] Under-cloud pixel interpolation: Efficient under-cloud pixel interpolation methods effectively fill in data gaps in cloud-covered areas, improve data integrity and continuity, and ensure the comprehensiveness and accuracy of water area extraction.
[0031] Improve data processing efficiency: High-efficiency processing: The GPU-based parallel acceleration cloud-based pixel interpolation algorithm significantly improves data processing efficiency, with a monthly data filling rate consistently exceeding 98%, which can meet the needs of real-time or near-real-time monitoring.
[0032] Automated processing: Intelligent water extraction algorithms reduce human intervention, improve the automation of data processing, and reduce labor and time costs.
[0033] Enhance the comprehensiveness and reliability of accuracy assessment: Multi-dimensional evaluation: A comprehensive accuracy evaluation system with multiple dimensions and levels was constructed to systematically verify the performance of the dataset from three dimensions: temporal stability, spatial geometric consistency, and ground realism, so as to ensure the comprehensiveness and reliability of the evaluation results.
[0034] High reliability: The verification method, which combines vector water body data and measured water body sample points, ensures the representativeness and stability of the assessment results, providing reliable data support for water resource management and ecological environment protection.
[0035] Achieving daily high-frequency monitoring and early warning of abnormal water body expansion: Daily high-frequency monitoring and early warning: By processing Fengyun 3F satellite data in real time and combining it with intelligent water area extraction algorithms, it can promptly detect and respond to abnormal water body expansion events, such as floods and lake expansion.
[0036] It provides important support for water resource management and ecological environment protection: Data Support: The high-precision water dataset provided by this patent offers important basic data support for research on water resource management and ecological environmental protection, and helps to scientifically formulate water resource utilization and protection strategies.
[0037] Decision support: Real-time monitoring and early warning systems provide timely and accurate decision support for government departments and relevant agencies, helping to improve their ability to respond to natural disasters and ecological and environmental problems.
[0038] In summary, this invention significantly improves the accuracy and efficiency of natural water body monitoring by employing intelligent water body extraction algorithms, efficient cloud-based pixel interpolation methods, a multi-dimensional accuracy assessment system, and a real-time monitoring and early warning system. It also enhances the comprehensiveness and reliability of accuracy assessment, enabling daily high-frequency monitoring and early warning of abnormal water body expansion. Attached Figure Description Figure 1 A schematic diagram of reconstructing a spatiotemporal cube from pixels under the cloud (Step 1, core algorithm).
[0039] Figure 2 This is a composite image of three channels 250m before and after pixel interpolation under the cloud.
[0040] Figure 3 This is a water body identification technology route based on the XGBoost algorithm.
[0041] Figure 4 shows the water body identification results of Fengyun 3FXGBoost on different dates (20250412, 20250502, 20250713).
[0042] Figure 5 This is a schematic diagram of the vector distribution of water bodies.
[0043] Figure 6 This is a schematic diagram for boundary accuracy evaluation (identifying the spatial superposition of the water body boundary and the reference water body boundary).
[0044] Figure 7 shows the statistics of the extracted water area across the entire region.
[0045] Figure 8. Monitoring results of abnormal water body expansion.
[0046] Figure 9 This is a schematic diagram of the process of the present invention. Detailed Implementation The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0047] The method for extracting natural water bodies and monitoring anomalous expansion from Fengyun-3F satellite imagery includes the following steps; Step 1: Under-cloud pixel interpolation Acquire FY-3F / MERSI-III L1 level daily multi-channel reflectance data at 250m or 1km. A strategy of labeling, interpolation, and iterative extrapolation is adopted to solve the problem of invalid values caused by cloud cover in FY-3F optical data, and output spatiotemporally continuous daily cloud-free multi-channel reflectance data (at two resolutions: 250m and 1km). Key points of the method: ① Radiometric calibration: Numerical registration between data is performed using the calibration parameters inherent in the data. The calibration parameters for the visible light band provide polynomial coefficients for the quadratic, linear, and constant values of the original values, respectively. For the infrared band, the inverse Planck transform is performed using given parameters; ② Coordinate standardization and reshaping: Given that the latitude and longitude grid of the original data is not regular, after selecting the latitude and longitude range according to the research objective, the data is projected pixel by pixel onto the standard latitude and longitude grid using the latitude and longitude information inherent in the data, and missing pixels are marked. When multiple data exist within a given range, their reshaping results are merged, and the values of duplicate pixels are averaged to finally obtain daily standardized reflectance data that still contains a small amount of invalid data (due to gaps in the scanned data / missing original data / damaged original data, etc.); ③ Cloud pixel marking: Using the FY-3F official cloud detection product ( CLM further marks the pixels obtained at the second point, and marks all cloud pixels as invalid data; ④ Single-round interpolation: The processing result of the third point is regarded as a "spatiotemporal cube" of T*H*W*C, where T is the time span, H and W are the number of rows and columns of data, and C is the number of bands. Then, inverse distance weighting (IDW) combined with adaptive normalization is used to interpolate and fill in all invalid pixels with spatiotemporal adjacency. In practice, the kernel size of IDW is set to 3 (3*3 pixels, ±1d); ⑤ Multi-round iterative extrapolation: Based on the single-round interpolation at the fourth point, based on the case that a small number of pixels do not have any spatiotemporal adjacency, it is necessary to update the set of valid pixels in multiple rounds of iteration, gradually reduce the range of invalid pixel areas, and finally fill in all pixels. Furthermore, during multiple iterations, the IDW convolution kernel size will increase to 5 or 7 every ten iterations after 10 iterations (maximum 7, i.e., 7*7 pixels, ±3d) to accelerate the iteration speed; ⑥ Iteration termination condition: Iteration terminates when the number of missing pixels is zero, or when the missing percentage first falls below 0.01% after reaching the maximum number of iterations (30 iterations), to maximize data integrity; ⑦ Post-processing: The interpolation process uses the PyTorch framework for GPU-accelerated computation, and the output file is in pt format, which does not meet the requirements of subsequent processing. The GDAL library is needed to convert the data into tif format with geographic coordinates. The final output is a radiometrically calibrated and corrected FY-3F multi-channel reflectance dataset with daily cloudless 250m and 1km resolutions, which can directly proceed to step 2.
[0048] Step 2: Intelligent water area extraction Adaptive water body extraction was performed on the multi-channel cloudless reflectance data output from step 1, generating binary classification results (1-water, 0-non-water) and optimized vector files for water bodies at two resolutions: 250m and 1km. Key points of the method: ① Feature extraction: Using FY-3F multi-band imagery (blue (B), green (G), red (R), and near-infrared (NIR) bands), (G−NIR) / (G+NIR) was calculated pixel by pixel to highlight water bodies and suppress vegetation and soil noise. Simultaneously, Brightness=(R+G+NIR) / 3 was generated to characterize the overall reflectance intensity. The final five-dimensional features fed into the model are [R,G,NIR,NDWI,Brightness]. ② Samples and Model: During the training phase, externally labeled water body samples are used, trained according to 3 categories (non-water / water / other). During inference, water bodies are extracted using WATER_CLASS_ID=2; ③ XGBoost (eXtremeGradientBoosting) Water Classification: A gradient boosting tree model is constructed, balancing the water / non-water body ratio with a positive / negative sample weight ratio (scale_pos_weight), and 5-fold cross-validation is used to suppress overfitting; ④ Water Distribution Vectorization Optimization: Morphological opening operations (3×3 kernels) are used to remove minor noise and smooth boundaries in the water body recognition raster results. A minimum patch area threshold (MIN_PATCH_SIZE, in pixels) is set to remove fragmented connected components. The binary water body mask raster is converted to vector, adjacent patches are automatically merged, and the topology is simplified (tolerance at the pixel resolution level), ultimately generating a high-precision vector output. The processing results directly generate vector output files that conform to cartographic specifications. The final output of this step is: a binary water raster (.tif), a water vector file (.shp), and a true-color overlay boundary map (.jpg). If the accuracy in step 3 is not up to standard, return to this step to retrain the XGBoost model or adjust the post-processing parameters. Once the data meets the requirements, you can proceed to step 3.
[0049] Step 3: Multi-dimensional accuracy assessment Independent verification of the output results in step 2 was conducted to ensure that they met the requirements for monitoring applications. Key methodological points: ① Pixel-level accuracy evaluation (raster-based): After registering the binary water body raster with the baseline water body raster manually interpreted from the concurrent high-resolution imagery (Gaofen or Sentinel-2), pixel-by-pixel comparison was performed. The confusion matrix was calculated, and indices such as OA, Kappa, IoU, and F1-score were calculated. ② Validation based on vector water body data: Using Xinjiang regional water body vector data as the reference ground truth, spatial overlap analysis was conducted. Indices such as intersection-to-union ratio (IoU), area difference (ΔA), and water body boundary deviation were calculated to evaluate the accuracy and consistency of the dataset in spatial structure. Alternatively, the water body boundary vector output in step 2 can be directly spatially superimposed with the aforementioned baseline water body vector to calculate indicators such as IoU, boundary overlap, and average distance deviation, quantitatively analyzing the degree of agreement between the water body range and the boundary location; ③ Ground validation based on measured water body sample points: Independent validation sample points (water / non-water bodies) obtained from field GPS measurements or high-resolution aerial image interpretation, or collected measured water body sample point data within the Xinjiang region, are spatially superimposed with the classification raster to construct a confusion matrix and calculate accuracy indicators such as overall accuracy (OA), Kappa coefficient, user accuracy (Precision), and cartographer accuracy (Recall), validating the authenticity of the dataset from a ground observation perspective. Key requirements: Validation samples must be spatially independent and ≥100 per class; results are only reliable and proceed to step 4 when OA ≥90%, IoU ≥0.75, F1 ≥0.85, and user accuracy ≥80% for each typical land cover type. If the accuracy does not meet the requirements, return to step 2 to adjust model parameters or post-processing thresholds.
[0050] Step 4: Monitoring of abnormal water body expansion Based on the water body time-series classification product verified in step 3, anomaly expansion detection of natural water bodies is carried out. Key methodological points: ① Baseline water body establishment: The historical maximum water body extent is extracted from cloudless data over the past 5 years as a baseline; ② Change detection: The COLD algorithm and LandTrendr algorithm are used to detect abrupt changes in water body expansion; ③ Anomaly determination: Combining hydrological and meteorological data (precipitation, temperature), when the water body expansion rate exceeds twice the standard deviation of the historical average for the same period and persists for more than 3 months, it is marked as an abnormal expansion area; ④ Output early warning product: Generate a thematic map and vector file of abnormal water body expansion containing expansion area, expansion rate, and confidence level.
[0051] Example 1: Detailed Description of Under-Cloud Pixel Interpolation Process FY-3F / MERSI Under-Cloud Pixel Interpolation: In the MERSI visible-near-infrared band, under the influence of convective cloud systems in summer, cloud pixels generally account for more than 35%, and in some areas even exceed 60%, leading to a large number of missing values in water body index sequences such as NDWI. This patent designs an "under-cloud pixel interpolation" algorithm. First, invalid pixels are identified using the FY-3F CLM cloud mask. Then, within a constructed spatiotemporal cube (taking one day before and after in the time dimension, and an initial 7×7 pixel window in the spatial dimension), valid observed pixels are searched, and the missing data is reconstructed using an inverse distance-weighted interpolation method. Specifically, for any missing pixel p, within the constructed spatiotemporal cube (taking T adjacent days before and after in the time dimension, and a 7×7 pixel window in the spatial dimension), a set of valid observed pixels is searched. Calculate the spatiotemporal distance from it to the target pixel P; ; In the formula, λ The time weighting coefficient (usually taken as...) λ =0.1~0.5 to balance the differences in spatiotemporal scales.
[0052] The interpolation is given by the following formula: ; In the formula, α =2 is the distance decay exponent. N ≥3 is the lower limit of the number of effective neighboring pixels (if less than 3, it is marked as pending the next iteration).
[0053] For large, continuous areas without data, a multi-round iterative extrapolation mechanism is employed. After each round of interpolation, the effective cell set is updated, gradually expanding the spatiotemporal search radius (maximum expansion to a 15×15 cell window plus 15 days before and after), until the fill rate is <0.1% for three consecutive rounds or the maximum number of iterations is reached (30 rounds for 250m resolution, 20 rounds for 1km resolution). The entire algorithm is accelerated in parallel using GPUs, significantly suppressing cloud contamination while preserving spatial details, ensuring a stable monthly data fill rate exceeding 98%.
[0054] Example 2: XGBoost Model Training and Water Extraction Process XGBoost (eXtreme Gradient Boosting) water classification algorithm: XGBoost is a highly efficient gradient boosting algorithm belonging to the ensemble learning method. It builds a powerful predictive model by combining multiple weak learners, typically Classification and Regression Trees (CART) decision trees. In water extraction tasks, the core of the XGBoost algorithm is the iterative construction of CART decision trees. Each new tree is learned based on the residuals of the previous tree (i.e., the difference between the actual value and the model's prediction), thus progressively reducing the model's prediction error. Compared to traditional gradient boosting methods, a significant advantage of the XGBoost algorithm is its use of second-order gradient information to optimize the objective function. It considers not only the first-order gradient (the first derivative of the loss function with respect to the predicted value) but also the second-order gradient (the second derivative of the loss function with respect to the predicted value). This use of second-order gradient information helps accelerate the model's convergence speed, enabling the model to find the optimal solution more quickly. Simultaneously, the XGBoost algorithm introduces a regularization term Ω(fk), which plays a role in controlling the model's complexity within the objective function. Regularization terms can prevent the model from becoming too complex and overfitting, thereby improving the model's generalization ability on new data.
[0055] Core objective function: ;
[0056] In the formula, l is the loss function and Ω(fk) is the regularization term.
[0057] Through this optimization method, the XGBoost algorithm can achieve efficient and accurate extraction of water bodies from FY-3F satellite remote sensing images. Under different image conditions, such as different lighting, atmospheric conditions, and land cover types, the XGBoost algorithm demonstrates strong adaptability and high extraction accuracy, showing significant advantages over traditional water body extraction methods. Implementation method: Data Acquisition Acquire the MERSI L1 sensor, data, and products from the Fengyun-3F satellite.
[0058] FY-3F Medium Resolution Spectral Imager L1 Data (1km) FY-3F Medium Resolution Spectral Imager L1 Data (GEO / 1km) FY-3F Medium Resolution Spectral Imager L1 Data (250M) FY-3F Medium Resolution Spectral Imager L1 Data (GEO / 250M) MERSI-III cloud detection segment products Data preprocessing Input data: Raw data obtained from the data acquisition stage.
[0059] Radiometric calibration and geometric correction are performed on remote sensing data to ensure data accuracy and consistency.
[0060] Output data: Radiometrically calibrated and geometrically corrected remote sensing data.
[0061] Under-cloud pixel interpolation Input data: Preprocessed remote sensing data.
[0062] Invalid pixels are identified using the CLM cloud mask built into the Fengyun-3F satellite.
[0063] We construct a spatiotemporal cube by taking one day before and after the time dimension and an initial 7×7 pixel window in the spatial dimension.
[0064] The missing data is reconstructed by searching for valid observed pixels within the spacetime cube and using inverse distance weighted (IDW) interpolation.
[0065] For large, continuous areas without data, a multi-round iterative extrapolation mechanism is adopted. After each round of interpolation, the effective cell set is updated, and the spatiotemporal search radius is gradually expanded until the filling rate is less than 0.1% for three consecutive rounds or the maximum number of iterations is reached.
[0066] Output data: Daily cloudless 250m and 1km multi-channel composite data.
[0067] Feature extraction Input data: Daily 250m and 1km multi-channel composite data after cloud interpolation Blue (B), green (G), red (R), and near-infrared (NIR) bands were extracted, and the normalized water index (NDWI) was further calculated to highlight water bodies and suppress vegetation and soil noise. At the same time, the overall surface reflectance intensity was characterized by multi-band weighted composite brightness features.
[0068] Output data: Extracted water feature data (NDWI, Brightness, B, G, R, NIR).
[0069] Intelligent water extraction Input data: Extracted water body feature data.
[0070] An intelligent water area extraction model is constructed using the XGBoost algorithm, achieving XGBoost block prediction + post-processing + vectorization. Block prediction: Divide the image into blocks to improve computational efficiency.
[0071] Post-processing: Smooth the prediction results and remove noise.
[0072] Vectorization: Converts binary water mask (GeoTIFF) to vector format.
[0073] The model was trained and validated using long-term water body data and measured water body sample points from Xinjiang to ensure its adaptability and accuracy under different seasons and different land cover types.
[0074] Output data: Daily 250m and 1km binary water body masks (GeoTIFF), vector (SHP), and overlay map (JPG). Accuracy assessment Input data: Daily 250m and 1km binary water body masks (GeoTIFF) and vector (SHP) Pixel-level accuracy evaluation (based on raster): After registering the binary raster of water body with the reference water body raster of manual interpretation of the high-resolution image (Gaofen or Sentinel-2) of the same period, the pixels are compared one by one, the confusion matrix is calculated, and indicators such as OA, Kappa, IoU, and F1-score are calculated. Validation based on vector water body data: Using Xinjiang regional water body vector data as the reference ground truth, spatial overlap analysis was conducted to calculate indices such as intersection-to-union ratio (IoU), area difference (ΔA), and water body boundary deviation, evaluating the accuracy and consistency of the dataset in terms of spatial structure. Alternatively, the water body boundary vector output in step 2 can be directly spatially superimposed with the aforementioned reference water body vector to calculate indices such as IoU, boundary overlap, and average distance deviation, quantitatively analyzing the degree of agreement between the water body range and the boundary location. Ground-based validation based on measured water body sample points: Independent validation sample points (water / non-water bodies) obtained from field GPS measurements or high-resolution aerial image interpretation, or collected measured water body sample point data within the Xinjiang region, are spatially overlaid with classification rasters to construct a confusion matrix and calculate accuracy indicators such as overall accuracy (OA), Kappa coefficient, user accuracy (Precision), and cartographer accuracy (Recall) to verify the authenticity of the dataset from a ground observation perspective.
[0075] Output data: Precision evaluation results (OA, Precision, Recall, F1, Kappa, IoU, etc.).
[0076] Monitoring of abnormal water body expansion Input data: Daily 250m and 1km binary water mask (GeoTIFF) data that meet the accuracy assessment requirements.
[0077] Data preprocessing Data standardization: Standardize all input data to a uniform format and resolution to ensure data consistency, and convert the data into a binary image.
[0078] Time series construction: Arrange all standardized water masks in chronological order to construct time series data.
[0079] Reference water body setting Adaptive reference water body: The reference water body is calculated based on historical data. It is obtained by calculating the 75th quantile of historical water body data and is used to represent the normal water body distribution.
[0080] Seasonal baseline: If date information is provided, calculate the seasonal baseline water body to account for water body variations in different seasons.
[0081] Flood monitoring Multi-mode joint detection: This involves using multiple detection modes to jointly detect flooding events. The detection modes include: Absolute increment detection: Calculate the absolute increment between the current water body and the reference water body, and detect areas where the increment exceeds the threshold (0.05).
[0082] Relative growth detection: Calculates the relative growth of the current water body to the reference water body and detects areas where the growth exceeds the threshold (10%).
[0083] New water body detection: Areas that are non-water bodies (<0.1) in the benchmark water body, but are currently water bodies (>0.3).
[0084] High water level detection: Detects areas where the current water level exceeds the high water level threshold (0.8).
[0085] Multi-scale intensity calculation: Perform multi-scale intensity calculations on the detected flooded areas to assess the severity of the flooding.
[0086] Event recognition Event recognition logic: Identifying flood events. The improved event recognition logic includes: Area reduction assessment: If the current flooded area is less than 30% of the peak area, the event is considered to be over.
[0087] Continuous low water level assessment: If the flooded area is less than 5 km² for 3 consecutive days, the event is considered over.
[0088] Event splitting logic: If the flooded area decreases by more than 70%, the current event will be split into two independent events.
[0089] Event matching: If administrative division data is provided, the identified flood events will be matched to the corresponding administrative divisions.
[0090] Multi-window analysis Multi-window analysis: Perform 3 / 5 / 7-day multi-window analysis on each flood event to assess the changes in the event within different time windows.
[0091] Visualization and Report Generation Visualization: Generate flood time series plots and flood event statistics charts to intuitively display the distribution and changes of flood events.
[0092] Report generation: Generates detailed flood analysis reports, including event statistics, accuracy indicators, and visualization charts.
[0093] Output data Abnormal water body expansion monitoring results include detailed information on flood events, such as event ID, start date, peak date, end date, duration in days, peak inundation area, latitude and longitude of the flood center, and severity.
[0094] Visual charts: flood time series charts, flood event statistics charts, etc.
[0095] Report documents: Text reports and Excel spreadsheets containing event statistics and accuracy metrics.
[0096] like Figure 1 The diagram shows a spatiotemporal cube for reconstructing pixels under the cloud (step 1, core algorithm). Figure 1The 3D grid visually illustrates the reconstruction strategy of "centered on the target pixel, extending forward and backward along the time axis, and expanding into multiple neighborhoods along the spatial axis," serving as a visual explanation of the core algorithm in step 1: "labeling first, then interpolating, and iterative extrapolation." Here, 1 represents the spatial axis: the horizontal direction represents the image row and column numbers (represented by i, j in the diagram), and square grids represent pixels. The central pixel is the "target pixel" to be reconstructed, located at the geometric center of the 5×5 window (black square). The first-order neighborhood refers to the 8 pixels directly adjacent to the central pixel (top, bottom, left, right + four diagonals, gray squares). The higher-order neighborhood refers to the outermost 16 pixels (light gray squares), used to expand the search range and improve interpolation stability. 2. Time Axis: The vertical direction is "Day". Centered on the target pixel at time Dayt, two days are taken before and after, for a total of 5 observation phases (t-2, t-1, t, t+1, t+2). Each phase stores the multi-channel reflectance values of the same 5×5 window, forming a 5×5×5 "spatiotemporal cube" (T-block). 3. Valid / Invalid Pixel Marking: If a pixel at a certain time phase and location is marked as "cloudy" or "missing" by the FY-3FCLM cloud mask, it is considered invalid (represented by hollow squares in the figure). During interpolation, inverse distance weighting (IDW) is performed only on valid pixels (solid squares), and the cloudless reflectance value of the center pixel at Dayt is estimated by spatiotemporal distance weighting. 4. Iteration strategy: In the first round, IDW is performed using a 5×5×5 cube; if there are fewer than 3 effective pixels, proceed to the next round and expand the spatial window to 7×7 and the time window to ±3 days until the convergence condition is met (see steps 1-④~⑥ in the instruction manual).
[0097] like Figure 2 The image shown is a three-channel composite image of 250m before and after pixel interpolation under the cloud (a schematic diagram of the output data of step 1), which intuitively demonstrates the cloud pollution repair effect of the algorithm in step 1 on the FY-3FMERSI-III 250m data. Figure 2 The "same region, same phase, same stretching" method is used to intuitively demonstrate that the "spatiotemporal cube + iterative IDW" strategy in step 1 can effectively restore the true reflectivity of pixels under clouds while maintaining spatial texture, providing a continuous, cloud-free data foundation for subsequent high-precision water body extraction. like Figure 3 This is the water body identification technology route based on the XGBoost algorithm, corresponding to sections (1) to (4) of step 2 in the instruction manual. Figure 3Using the engineering approach of "first dividing into blocks and then merging", the entire process of "feature calculation → machine learning prediction → morphological post-processing → vectorization" is connected into a closed loop, which solves the memory bottleneck of large images and ensures the output of a raster-vector integrated product with mapping-level precision. This corresponds one-to-one with the subsections (1) to (4) of step 2 in the instruction manual.
[0098] Figure 4 shows the water body identification results from Fengyun 3F data on different dates. Figure 4 uses three 250m false-color images side-by-side to display the automatic water body identification results for the same area on different dates, visually verifying the robustness of the XGBoost model under various phenological and meteorological conditions. Figure 4 uses three typical time phases—spring, summer, and autumn—to demonstrate that even under three different water color states (ice, clear water, and high turbidity) and various background interferences such as bare soil, vegetation, and cloud shadows, this invention can still stably and accurately extract water body boundaries, meeting the accuracy thresholds of OA ≥ 90% and IoU ≥ 0.75 proposed in step 3 of the specification.
[0099] like Figure 5 This is a schematic diagram of the vector distribution of water bodies. Figure 5 The “Water Body Vector Distribution Diagram” is a 250m automatic water body extraction result map of a typical time phase (2025-07-04) in western China (Xinjiang and surrounding areas) for operational display. Figure 5 Using a method of "red vector overlaying false-color imagery," this invention demonstrates its mapping-grade automated water body extraction capabilities in large-scale, complex terrain, and cloudy environments: high boundary geometric accuracy, coverage of water bodies of varying sizes, and no human intervention. It can be directly used for subsequent step 4, "anomaly expansion monitoring," and operational water resource statistics. like Figure 6 The diagram illustrates the boundary accuracy evaluation, where the water body boundary vector (red line) output in step 2 is spatially superimposed with the reference water body vector (green line) (e.g., Ulungur Lake). Figure 6 The diagram for "Boundary Accuracy Evaluation" shows the spatial overlay results of the water body boundary vector (red line) automatically extracted in step 2 and the high-precision reference water body vector (green line) in the same area (Ulungur Lake-Jili Lake Group, 2025-07-04), used to visually evaluate geometric accuracy. Figure 6 The "red-green double-line overlay" method provides a visual demonstration that, at a medium resolution scale of 250m, the geometric position of the lake shoreline extracted by this invention closely matches the true value at 10m, with boundary deviation controlled at the sub-pixel level, meeting the accuracy requirements for subsequent anomaly expansion monitoring. Figure 7 is a schematic diagram of the invention process.
[0100] Figure 7 shows the extracted area statistics of the entire water body. Remote sensing monitoring of 50 major lakes in Xinjiang from April to August 2025 shows: ① The overall water area shows a slight expansion trend (slope 0.000018, R²=0.58), with an average area of 52,288.6 km², but 98.6% are occasional water bodies, and only 1% are permanently stable, indicating a prominent fragile characteristic of "increased total area but changing individual lakes". ② Only 8 lakes can be continuously tracked by intelligent matching (matching rate 16%), among which Lake Ayagkumkuli, Lake Aksaichin, and Lake Whale have the highest reliability (score 0.90–0.98); Lake Ayakkul has the largest average area (974 km²), but its annual fluctuation is as high as 3 times; the vector areas of Lake Ebinur and Lake Sayram are much higher than the actual average, indicating a long-term shrinkage background. ③ The extremely high proportion of occasional water bodies and the large standard deviation indicate that the region is driven by short-term snowmelt or extreme precipitation, resulting in "pulsating" expansion of lakes, and the risk of sustainability remains.
[0101] Figure 8. Monitoring results of abnormal water body expansion. During the monitoring period from April to August 2025, a total of 8 "single-day-short" flood events were identified, all concentrated in the area along the Tianshan Mountains in northern Xinjiang (center longitude 82°±1°, latitude 44°±1°). Although the peak inundation area was 40,000–48,000 km² and the multiple was 1.04–1.18×, the events only lasted for 1–6 days before receding. The abnormality rate was only 12.4%. Overall, the events showed the characteristics of "large scope, short duration, and rapid receding" in spring, which was sudden and repeated in summer, and did not form a regional long-term disaster.
[0102] Figure 9 The complete technical link of this invention, from raw satellite data to anomaly expansion early warning, is condensed into 7 key nodes in the form of an end-to-end flowchart, forming a "daily-scale 250m / 1km cloudless water body monitoring production line" that can be operated in an operational manner.
Claims
1. A method for extracting natural water bodies and monitoring anomalous expansion from Fengyun-3F satellite imagery, characterized in that, Includes the following steps; Step 1: Acquire multi-channel reflectance data, and output spatiotemporally continuous daily cloudless multi-channel reflectance data through labeling, interpolation, and iterative extrapolation strategies; Step 2: Perform adaptive water body extraction on the multi-channel cloudless reflectance data output from Step 1 to generate water body binary classification results and optimized vector files at two resolutions: 250m and 1km. Step 3: Conduct independent verification of the output results of Step 2 to ensure that they meet the monitoring application requirements; Step 4: Based on the verified and reliable water body time series classification products in Step 3, conduct detection of abnormal expansion of natural water bodies.
2. The method for extracting natural water areas and monitoring anomalies in Fengyun-3F satellite imagery according to claim 1, characterized in that, Step 1 specifically involves: (1) Data reading: Acquire FY-3F / MERSI-III L1 level 250m or 1km daily scale multichannel reflectance data and CLM cloud mask products; (2) Cloud pixel marking: Use the CLM cloud mask of Fengyun 3F satellite to mark invalid pixels, output the position of the marked invalid pixels, and proceed to step (3). (3) Spatiotemporal cube construction: Take one day before and after in the time dimension, and the initial spatial dimension is a 7×7 pixel window to construct a spatiotemporal cube; output a three-dimensional data block containing the reflectance values of the target pixel and its spatiotemporal neighboring pixels, and proceed to step (4). (4) IDW first round of filling: Search for valid observed pixels within the spatiotemporal cube and reconstruct missing data using the IDW interpolation method; The interpolation formula is: ; In the formula, y i Let d be the value of the i-th effective neighboring pixel. i Let be the distance between the i-th effective neighboring pixel and the target pixel. α The distance decay exponent, N The number of all valid neighboring pixels; Output the interpolated reflectance value or the tag to be iterated, and proceed to step (4). (5) Multi-round iterative extrapolation and convergence conditions: For large-scale continuous data-free areas, a multi-round iterative extrapolation mechanism is adopted. After each round of interpolation, the effective cell set is updated, and the spatiotemporal search radius is gradually expanded until the filling is completed, or the continuous filling rate is <0.1%, or the maximum number of iterations is reached. Output multi-channel reflectance data that fully fills in cloud pixels.
3. The method for extracting natural water areas and monitoring anomalies in Fengyun-3F satellite imagery according to claim 2, characterized in that, Step 2 specifically involves: water body extraction, from reflectance to a binary water body map, including the following steps: (1) Feature extraction: Using the multi-channel reflectance data of FY-3F 250m and 1000m under cloud interpolation or under the original clear sky conditions without clouds, the data are read in block by block to extract features related to water bodies, including four bands: red light (R), green light (G), blue light (B) and near infrared (NIR). The R and G bands directly reflect the coverage and health status of surface vegetation, the B band is sensitive to atmospheric scattering, and the NIR band is highly sensitive to vegetation biomass and water content. In the feature calculation stage, the two basic spectra of R and G are retained in the original bands, and the normalized water index NDWI is further calculated. At the same time, the brightness feature Brightness is synthesized through multi-band weighting to characterize the overall surface reflectance intensity. Finally, the five features of R, G, NIR, NDWI and Brightness, which have both physical meaning and discriminative ability, are used as the original input features of XGBoost. (2) Sample and model preparation: First, relying on global free satellite resources, Landsat8, Landsat9 and Sentinel2 optical images were selected, and Sentinel1 SAR was introduced to compensate for the influence of clouds and rain, and radiometric, atmospheric, orthorectification and projection unification were completed; then, the water body contour was manually vectorized in QGIS, non-water body samples in easily confused areas were added, and an initial mask was generated by combining NDWI, MNDWI, AWEI and Otsu thresholds, and then manually corrected; the labels used a binary raster, water body 1, non-water body 0, and mixed pixels were marked as ignored values and masked during training; ; AWEI contains AWEI nsh (Shadowless version) and AWEI sh (Shadowed version), suitable for different scenarios: AWEI nsh = 4×(Green−SWIR1) −(0.25×NIR+2.75×SWIR2); AWEI sh = Blue+2.5×Green−1.5 × (NIR+SWIR1) − 0.25 × SWIR 2; In the formula, Blue is the blue light band; Green is the green light band; NIR is the near-infrared band; SWIR1 is the short-wave infrared band 1; SWIR2 is the short-wave infrared band 2; Post-processing removes false water bodies through small patch filtering, morphological opening and closing operations, and shape and area constraints. The sample library archives image patches, labels, metadata, and version records according to a unified standard, forming a high-quality, traceable dataset. Samples are divided into three categories: non-water bodies, water bodies, and others. They are input into XGBoost to construct a gradient boosting tree model, balancing the class ratio with a positive-to-negative sample weight ratio (scale_pos_weight), and using five-fold cross-validation to suppress overfitting. No retraining is required during the inference stage; if the accuracy is not up to standard, a new model is resampled and trained in step 2. (3) XGBoost water body classification: The XGBoost algorithm is used to construct an intelligent water body extraction model; the XGBoost algorithm uses second-order gradient information to accelerate model convergence, and at the same time introduces a regularization term to prevent overfitting; finally, three products are output simultaneously: water body binary raster (.tif), water body vector file (.shp) and true color overlay boundary map (.jpg). Raster and vector data can be used as input data for step (3); output the segmented water body mask and proceed to step (4) vectorization optimization step; (4) Optimization of water body distribution vectorization: The morphological opening operation of 3×3 structuring elements is used to smooth the water body boundary and remove small noise. Then, connected component analysis is performed to set the small fragmented patches with an area smaller than MIN_PATCH_PIXELS to zero, convert the raster to vector, merge adjacent patches and simplify the topology with pixel resolution as the threshold, and simultaneously write the LZW compressed binary raster _water.tif with nodata=0, the water body vector *_water.shp with the same CRS as the original image, and the cartograph-grade true color overlay map *_shp_overlay.jpg with red vector lines superimposed on the RGB base map. After completion, proceed to step (3) quality inspection.
4. The method for extracting natural water areas and monitoring anomalies in Fengyun-3F satellite imagery according to claim 3, characterized in that, The specific steps of feature extraction in (1) are as follows: Normalized Difference Water Index (NDWI): ; Brightness Index: ; Function: To distinguish highly reflective features (such as bare soil and buildings), while water bodies typically have lower reflectivity (<0.2).
5. The method for extracting natural water areas and monitoring anomalies in Fengyun-3F satellite imagery according to claim 4, characterized in that, Step 3 specifically involves multi-dimensional accuracy assessment, which includes the following steps: (1) Pixel-level accuracy evaluation: After registering the binary water body raster with the reference water body raster of the same high-resolution image, the pixels are compared one by one, the confusion matrix is calculated, and the OA, Kappa, IoU and F1-score indices are calculated. (2) Validation based on vector water body data: Using Xinjiang regional water body vector data as the reference true value, spatial overlap analysis is carried out to calculate the intersection-over-union ratio (IoU), area difference (ΔA) and water body boundary deviation, etc., to evaluate the accuracy and consistency of the dataset in spatial structure; or the water body boundary vector output in step 2 is directly spatially superimposed with the above-mentioned reference water body vector to calculate the IoU, boundary overlap degree and average distance deviation, etc., to quantitatively analyze the degree of agreement between the water body range and the boundary position; (3) Ground verification based on measured water body sample points: Independent verification sample points (water / non-water bodies) obtained from field GPS measurements or high-resolution aerial image interpretation, or collected measured water body sample point data in Xinjiang region, are spatially overlaid with classification rasters to construct confusion matrix and calculate accuracy indicators such as overall accuracy (OA), Kappa coefficient, user accuracy (Precision) and cartographer accuracy (Recall) to verify the authenticity of the dataset from the ground observation level; (4) Validation based on long-term water body probability map: The probability map of water body occurrence is generated using MODIS satellite multi-temporal images. The overall accuracy (OA), precision, recall, F1 score, Kappa coefficient, intersection-over-union ratio (IoU) and other accuracy indicators are calculated to evaluate the stability and robustness of the dataset under long-term series conditions.
6. The method for extracting natural water areas and monitoring anomalies in Fengyun-3F satellite imagery according to claim 5, characterized in that, The formula for calculating the index in step (4) is as follows: Overall Accuracy (OA) reflects the proportion of all pixels that are judged as correct; OA = (TP + TN) / (TP + TN + FP + FN); Precision (accuracy) is the percentage of pixels that the model identifies as water bodies that are actually water bodies. Precision = TP / (TP + FP); Recall (also known as recall rate or sensitivity) is the proportion of "real water" pixels that are successfully retrieved by the model; Recall = TP / (TP + FN). The harmonic mean of the F1 score, Precision and Recall, is used to comprehensively evaluate whether the number of tests is excessive and the accuracy of the tests. F1 = 2·Precision·Recall / (Precision + Recall). The Kappa coefficient (κ) is a consistency index after removing the "random guessing" part. The closer it is to 1, the more the classification result matches the real scene. First, calculate the "random consistency rate" Pe: Pe=[(TP+FN)(TP+FP)+(TN+FP)(TN+FN)] / (TP+TN+FP+FN)². Then calculate Kappa: κ=(OA–Pe) / (1–Pe); The IoU is the ratio of the overlapping area of the "predicted water body" and the "true water body" to the combined area of the two. The closer it is to 1, the more accurately the model captures the boundary of the water body; IoU=|A∩B| / |A∪B|=TP / (TP+FP+FN); In the formula, TP (True Positive) represents the number of pixels that are actually water bodies and correctly identified as water bodies by the model. TN (True Negative) represents the number of pixels that are actually non-water bodies and correctly identified as non-water bodies by the model. FP (False Positive) represents the number of pixels that are actually non-water bodies but misidentified as water bodies by the model; FN (False Negative) represents the number of pixels that are actually water bodies but misidentified as non-water bodies by the model; A: the set of "water body" pixels in the ground truth; B: the set of "water body" pixels predicted by the model; intersection |A∩B|=TP; union |A∪B|=TP+FP+FN.
7. The method for extracting natural water areas and monitoring anomalies in Fengyun-3F satellite imagery according to claim 6, characterized in that, Step 4 specifically involves: monitoring abnormal water body expansion. This is achieved through an intelligent water area extraction algorithm, which monitors changes in the water area in real time and identifies abnormal water body expansion events. This includes the following steps: (1) Input data, daily 250m and 1km resolution binary water mask raster data (GeoTIFF) that meet the multi-dimensional accuracy threshold of step 3. (2) Data preprocessing: Data standardization: uniformly resampled to WGS-84 / UTM area projection, set pixel depth to 8 bits, water pixels to 1, land pixels to 0, background fill value to 255; Time series construction: sorted by "year-day order", missing dates were filled by linear interpolation, and a T×H×W three-dimensional Boolean cube was generated. (3) Adaptive benchmark setting for benchmark water bodies: take the quantiles of the frequency cube of water bodies during the cloudless period to generate the probability layer (Pbase) of "normal water bodies"; seasonal benchmark: if the date information is complete, calculate the quantiles again by sliding the window according to "ten-day period" to obtain the seasonal benchmark (Pseason), which is used to eliminate seasonal fluctuations during the ice-covered period and irrigation period; (4) Flood detection: Absolute increment detection: ΔW=Pcurrent−Pbase, relative growth detection: R=ΔW / max(Pbase,0.1), trigger threshold 10%; New water body detection: when Pbase<0.1 and Pcurrent>0.3, it is determined as a new water body; High water level detection: when Pcurrent>0.8 and lasts for ≥5 days, it is determined as a continuous high water level; Multi-scale intensity: perform neighborhood morphological opening and closing reconstruction on the trigger pixel, and calculate the intensity index I=Σ(ΔW·A), where A is the area of the connected domain (km²); (5) Event identification: Area decline criteria: If the flood area on the day is less than 30% of the peak area, it is marked as "entering decay"; Continuous low water level: If the flood area is less than 5km² for 3 consecutive days, it is judged as the end of the event; Event segmentation: If the interval between two adjacent peaks is ≥7 days and the area decline is >70%, it is automatically split into independent events; (6) Administrative matching: Overlay county-level administrative division vectors and output the code and name of the corresponding district / county; The monitoring results of abnormal water body expansion include detailed information on flood events. The dynamic threshold is the mean area over time ±2σ. When the water body area observed in a single observation exceeds this range, an abnormal event is triggered. The results are further distinguished into two categories: "flood / lake expansion" and "drought / shrinkage". Visual charts: flood time series charts, flood event statistics charts; Report documents: Text reports and Excel spreadsheets containing event statistics and accuracy metrics.