A method for precisely extracting small water bodies by combining remote sensing and machine learning

CN122657741APending Publication Date: 2026-08-28ANHUI AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610785015.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0004]基于以上不足之处,本发明提供一种结合遥感与机器学习的细小水体精准提取方法,本发明能够提升高分辨率遥感影像对隐蔽水体的识别精度,同时解决了边缘断裂与阴影误判的问题

Benefits of technology

[0025] The beneficial effects and advantages of this invention are as follows: This invention establishes a multi-dimensional collaborative feature enhancement strategy by integrating a multi-scale edge-guided attention module, a high- and low-frequency decomposition and fusion module, and a convolutional block attention module. This strategy organically integrates edge constraints, frequency domain analysis, and attention mechanisms, effectively suppressing complex background noise such as building shadows while significantly enhancing the model's ability to perceive weak and fragmented water body signals. It exhibits good generalization ability and practical application value in complex scenarios. Furthermore, this invention can improve the accuracy of identifying hidden water bodies in high-resolution remote sensing images, while solving the problems of edge breakage and shadow misjudgment. Its fine-grained segmentation advantage provides a high-precision, low-cost technical solution for detecting small water bodies in large-scale areas. The MEF-TransUNet model designed in this invention successfully achieves high-confidence identification of core water body targets by discarding a very small proportion of blurred edge redundancy. This feature selection mechanism, which prioritizes result reliability, establishes the model's adaptability to the high-purity, low-noise extraction requirements in refined remote sensing mapping tasks, providing a more cost-effective technical solution for automated small water body detection. Furthermore, the MEF-TransUNet model employs a more robust identification strategy than traditional machine learning models. This multi-mechanism collaborative strategy ensures high recall while also improving the accuracy and purity of the extracted results, demonstrating strong robustness and lower manual correction costs in complex surface environments. The extraction results from small water bodies using this technology not only provide high-confidence technical support for refined water resource surveys but also maintain extremely high purity, significantly reducing the workload of manual correction and post-processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122657741A_ABST
    Figure CN122657741A_ABST
Patent Text Reader

Abstract

The application discloses a kind of small water body precision extraction method combined with remote sensing and machine learning, it is related to remote sensing image feature extraction technical field.The method first selects typical river basin to construct heterogeneity sample area;Then fuse 3-meter resolution PlanetScope optical image and C-band radar image, construct seven-channel dataset containing R, G, B, NIR, NDVI, NDWI, VV after preprocessing;MEF-TransUNet model is constructed, encoder fuses ResNet50 and ViT, decoder discards shallow noise features and integrates MEGA, HLFD, CBAM module;BCE-Dice hybrid loss function is used to train and optimize model;Finally, the remote sensing image of the region to be detected is input to realize small water body extraction.The application improves the small water body recognition precision, solves the edge fracture and shadow misjudgment problem, and is suitable for ecological protection and fine management of water resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing image feature extraction, specifically relating to a method for accurately extracting small water bodies by combining remote sensing and machine learning. Background Technology

[0002] Water bodies are a vital component of regional ecosystems, playing a fundamental role in regulating hydrological cycles, maintaining biodiversity, sustaining wetland functions, supporting socio-economic development, and ensuring ecological security. Small water bodies, including ponds, small lakes, low-lying streams, ditches, and springs, are closely related to daily human life and impact the living environment. Compared to large lakes and rivers, the shrinkage of small water bodies weakens their water storage capacity and exacerbates the risk of localized flooding. Therefore, achieving high-precision identification and dynamic change detection of small water bodies is of crucial scientific significance and practical value for regional ecological protection, refined water resource management, and urban planning decisions.

[0003] With the rapid development of remote sensing technology, water body mapping and change detection based on optical remote sensing images have become mainstream. Compared with traditional ground survey methods, remote sensing imagery has advantages such as wide coverage, strong temporal continuity, and high acquisition efficiency. Existing research shows that multi-temporal remote sensing data can accurately extract lake water information and its spatiotemporal evolution characteristics, and reveal its relationship with climate factors and human activities. Among them, multispectral satellite remote sensing images such as Landsat and Sentinel-2 have been widely used in water resource monitoring, flood detection and assessment, wetland mapping, and lake change detection. Surface water extraction algorithms based on optical remote sensing images can be mainly divided into two categories: spectral analysis and image classification. Specifically, spectral analysis-based methods for extracting surface water do not rely on training samples but are achieved by thresholding water body index images, thereby enhancing the contrast between water bodies and other land features. However, the threshold setting for water body indices usually varies from case to case, exhibiting significant uncertainty. Supervised image classification based on machine learning algorithms is an effective and direct method for extracting water targets from optical images. These supervised classification methods rely on the quality and quantity of training samples, and can obtain satisfactory water extraction results given sufficiently representative samples. However, facing increasingly complex remote sensing surface environments and the demand for high-resolution interpretation, traditional machine learning for extracting small water bodies is only applicable to specific scenarios and problems, and it is difficult to meet the dual requirements of high accuracy and high topological integrity for small water bodies. The water extraction results are also subject to resolution loss due to downsampling, resulting in pixel loss. In addition, although radar remote sensing images with vertical polarization are very sensitive to water information, they are rarely used for water extraction. Summary of the Invention

[0004] To address the above shortcomings, this invention provides a method for accurately extracting small water bodies by combining remote sensing and machine learning. This invention can improve the accuracy of identifying hidden water bodies from high-resolution remote sensing images, while solving the problems of edge breakage and shadow misjudgment.

[0005] The technical solution adopted in this invention is as follows: A method for precise extraction of small water bodies combining remote sensing and machine learning, comprising the following steps:

[0006] Step 1, Study Area Selection and Sample Distribution Design: Select a typical river basin as the study area, and select typical small water body distribution areas with spatial heterogeneity within the river basin, covering different geographical zones, land use types and small water body types;

[0007] Step 2, Multi-source remote sensing data preprocessing and dataset construction: For the typical area, download and preprocess optical remote sensing images and radar remote sensing images, calculate spectral indices, and construct a seven-channel input dataset and a binary label dataset containing red band images, green band images, blue band images, near-infrared band images, vegetation index images, water index images, and resampled vertical polarization radar remote sensing images from optical remote sensing images.

[0008] Step 3: Construct the MEF-TransUNet network model: Based on the improved TransUNet, construct the MEF-TransUNet deep learning segmentation model. The encoder uses a combination of ResNet-50 and ViT, and the decoder discards shallow noise features of C1 and restores resolution through cascaded upsampling. At the end of the decoder, a multi-scale edge-guided attention module, a high-low frequency decomposition and fusion module, and a convolutional block attention module are integrated in sequence, and a selective feature fusion strategy is adopted.

[0009] Step 4, Model Training and Parameter Optimization: Train, validate, and test the MEF-TransUNet model using a hybrid loss function, and save the optimal model weights;

[0010] Step 5: Water body extraction in the area to be detected: Download the same type of remote sensing image for the area to be detected and preprocess it, then input it into the trained MEF-TransUNet model to achieve accurate extraction of small water bodies.

[0011] Furthermore, in step 2, a 3-meter resolution PlanetScope multispectral optical image is acquired, including red, green, blue, and near-infrared bands, and preprocessed by radiometric calibration, atmospheric correction, geometric fine correction, cloud and fog removal, and normalization. The radar remote sensing image is a C-band vertical polarization image, which is preprocessed by multi-view, filtering, geocoding, and resampling to match the resolution of the optical image. The normalized vegetation index and normalized water index of the optical image are calculated.

[0012] Furthermore, in step 2, seven-channel input data are constructed, including red band images, green band images, blue band images, near-infrared band images, vegetation index images, water index images, and resampled vertical polarization radar remote sensing images; binary classification labels are created, dividing the dataset into 70% training set, 15% validation set, and 15% test set.

[0013] Furthermore, in step 3, the encoder changes the first layer convolution of ResNet-50 to a seven-channel input, and the deep C4 features are fed into the 12-layer ViT to capture long-range dependencies; the decoder adopts cascaded upsampling, only splicing the features of the C4, C3, and C2 layers, and discarding the shallow C1 features.

[0014] Furthermore, in step 3, the multi-scale edge-guided attention module takes the feature map, the predicted probability map, and the original image as input, extracts edge information through the Laplacian pyramid, constructs a three-way parallel attention mechanism, and outputs it through residual connection.

[0015] Furthermore, in step 3, the high-low frequency decomposition and fusion module decouples the features into low-frequency and high-frequency components. The low-frequency components are enhanced by discrete wavelet transform and Transformer, while the high-frequency components are enhanced by densely connected convolutional blocks. Finally, the features are adaptively fused through channel attention.

[0016] Furthermore, in step 3, the convolutional block attention module performs channel attention and spatial attention calibration sequentially. The channel attention uses a combination of global average pooling and global max pooling in a multilayer perceptron, while the spatial attention uses large kernel convolution to generate a spatial weight map.

[0017] Furthermore, in step 4, the hybrid loss function is a weighted fusion of the binary cross-entropy (BCE) loss and the Dice loss, with the weight coefficient λ set to 0.5; the Adam optimizer is used, and the initial learning rate is set to... Weight decay Cosine annealing to minimum learning rate is set to The batch size is 8, and the total number of epochs is 80.

[0018] Furthermore, in step 4, the optimal weights are loaded during model testing, and a water mask is generated with a binarization threshold of 0.65. The accuracy is evaluated through IoU, precision, and recall.

[0019] This invention also provides a system for precise extraction of small water bodies by combining remote sensing and machine learning to achieve the method for precise extraction of small water bodies by combining remote sensing and machine learning as described above, including: a sample design module, a data processing and dataset construction module, a MEF-TransUNet network model construction module, a model training and optimization module, and a water body extraction execution module;

[0020] The sample design module is used to select typical river basins as study areas, and within the river basins, to select typical small water bodies with spatial heterogeneity, covering different geographical zones, land use types, and small water body types.

[0021] The data processing and dataset construction module is used to download and preprocess optical and radar remote sensing images for the typical region, calculate spectral indices, and construct a seven-channel input dataset and a binary classification label dataset containing multi-band and remote sensing indices. Specifically, it acquires 3-meter resolution PlanetScope multispectral optical images, including red, green, blue, and near-infrared bands, and performs radiometric calibration, atmospheric correction, geometric fine correction, cloud and fog removal, and normalization. The radar remote sensing images are C-band vertical polarization images, processed by... After multi-viewing, filtering, geocoding, and resampling, the data is matched with the resolution of the optical imagery. The normalized vegetation index and normalized water index of the optical imagery are calculated to construct a seven-channel input data set containing red band, green band, blue band, near-infrared band, vegetation index, and water index images of the optical remote sensing imagery, as well as the resampled vertical polarization radar remote sensing imagery. The dataset is then cropped into 512×512 image patches, binary classification labels are created, and the dataset is divided into training, validation, and test sets in a ratio of 7:1.5:1.5.

[0022] The MEF-TransUNet network model building module is used to construct a MEF-TransUNet deep learning segmentation model based on an improved TransUNet. The encoder uses a combination of ResNet-50 and ViT, changing the first convolution of ResNet-50 to a seven-channel input, and feeding the deep C4 features into a 12-layer ViT to capture long-range dependencies. The decoder discards shallow noise features of the C1 layer, uses cascaded upsampling to restore resolution, and only concatenates features from the C4, C3, and C2 layers. At the end of the decoder, a multi-scale edge-guided attention module, a high-low frequency decomposition and fusion module, and a convolutional block attention module are integrated in sequence, and a selective feature fusion strategy is adopted. The multi-scale edge-guided attention module takes feature maps, predicted probability maps, and original images as inputs, extracts edge information through Laplacian pyramids, constructs a three-way parallel attention mechanism, and outputs it through residual connections. The high-low frequency decomposition and fusion module decouples features into low-frequency and high-frequency components. The low-frequency components are enhanced by discrete wavelet transform and Transformer, and the high-frequency components are enhanced by densely connected convolutional blocks. Then, they are adaptively fused through channel attention. The convolutional block attention module performs channel attention and spatial attention calibration in sequence. Channel attention uses a combination of global average pooling and global max pooling with a multilayer perceptron, and spatial attention uses large-kernel convolution to generate a spatial weight map.

[0023] The model training and optimization module is used to train, validate, and test the MEF-TransUNet model using a hybrid loss function, and to save the optimal model weights. The hybrid loss function is a weighted fusion of binary cross-entropy (BCE) loss and Dice loss, with a weight coefficient λ of 0.5. The Adam optimizer is used, and the initial learning rate is set to... Weight decay is set to Cosine annealing to minimum learning rate The batch size is 8, and the total number of epochs is 80. During model testing, the optimal weights are loaded, and a water mask is generated with a binarization threshold of 0.65. The accuracy is evaluated by IoU, precision, and recall.

[0024] The water body extraction execution module is used to download and preprocess the same type of remote sensing images of the area to be detected, and input them into the trained MEF-TransUNet model to achieve accurate extraction of small water bodies.

[0025] The beneficial effects and advantages of this invention are as follows: This invention establishes a multi-dimensional collaborative feature enhancement strategy by integrating a multi-scale edge-guided attention module, a high- and low-frequency decomposition and fusion module, and a convolutional block attention module. This strategy organically integrates edge constraints, frequency domain analysis, and attention mechanisms, effectively suppressing complex background noise such as building shadows while significantly enhancing the model's ability to perceive weak and fragmented water body signals. It exhibits good generalization ability and practical application value in complex scenarios. Furthermore, this invention can improve the accuracy of identifying hidden water bodies in high-resolution remote sensing images, while solving the problems of edge breakage and shadow misjudgment. Its fine-grained segmentation advantage provides a high-precision, low-cost technical solution for detecting small water bodies in large-scale areas. The MEF-TransUNet model designed in this invention successfully achieves high-confidence identification of core water body targets by discarding a very small proportion of blurred edge redundancy. This feature selection mechanism, which prioritizes result reliability, establishes the model's adaptability to the high-purity, low-noise extraction requirements in refined remote sensing mapping tasks, providing a more cost-effective technical solution for automated small water body detection. Furthermore, the MEF-TransUNet model employs a more robust identification strategy than traditional machine learning models. This multi-mechanism collaborative strategy ensures high recall while also improving the accuracy and purity of the extracted results, demonstrating strong robustness and lower manual correction costs in complex surface environments. The extraction results from small water bodies using this technology not only provide high-confidence technical support for refined water resource surveys but also maintain extremely high purity, significantly reducing the workload of manual correction and post-processing. Attached Figure Description

[0026] Figure 1 This is a flowchart of the detection method of the present invention;

[0027] Figure 2 Here is a diagram of the MEF-TransUNet model structure;

[0028] Figure 3 This is a diagram of the MEGA module structure.

[0029] Figure 4 Here is a structural diagram of the HLFD module;

[0030] Figure 5 Here is a structural diagram of the CBAM module;

[0031] Figure 6 Visualize the experimental results for comparison;

[0032] Figure 7 This is a visualization of the ablation experiment results. Specific Implementation

[0033] The present invention will be further described in detail below with reference to specific embodiments.

[0034] Example 1

[0035] To balance the advantages of multiple features and address the challenge of achieving both denoising and fidelity, this invention constructs a method for accurately extracting small water bodies based on an optimized TransUNet network, combining remote sensing and machine learning. This aims to overcome the technical bottlenecks of background confusion and target under-detection in small water body extraction. Specifically, this invention innovatively integrates two spectral index bands—vegetation and water—on top of traditional optical bands, while also incorporating vertically polarized radar remote sensing data sensitive to water body information. Figure 1 As shown, a method for accurately extracting small water bodies by combining remote sensing and machine learning includes the following steps:

[0036] Step 1: Study Area Selection and Sample Distribution Design: Specifically, the ten rivers with the largest drainage areas in China are selected, including the Yangtze River, Heilongjiang River, Yellow River, Pearl River, Tarim River, Haihe River, Yarlung Tsangpo River, Liaohe River, Huaihe River, and Lancang River. These rivers are distributed across different geographical regions of China, with significant differences in soil, topography, landforms, climate, and vegetation. Therefore, the small water bodies within each basin exhibit obvious spatial heterogeneity. Next, several typical areas are randomly selected within the basin of each river to ensure that each selected typical random distribution area within the basin includes different types of small water bodies such as lakes, ponds, and wetlands. Furthermore, the distribution of each type of small water body within the area also covers underlying land use types such as vegetation, forests, topography, landforms, buildings, and bare soil. This provides a solid foundation for the generalization ability of the machine learning network model proposed in this invention during the construction, training, and validation stages in different environments.

[0037] Step 2: Multi-source remote sensing data preprocessing and dataset construction: For the 10 rivers with the largest drainage areas distributed across different geographical regions of China in Step 1—the Yangtze River, Heilongjiang River, Yellow River, Pearl River, Tarim River, Haihe River, Yarlung Tsangpo River, Liaohe River, Huaihe River, and Lancang River—based on the latitude and longitude range of typical randomly distributed areas of small water bodies within each river basin, Level-3B analytical-grade 3-meter spatial resolution multispectral optical remote sensing image data products with surface reflectance correction were selected and downloaded from the Planet Labs website (https: / / www.planet.com / ). The remote sensing image data products include red, green, blue, and near-infrared bands. For this high spatial resolution PlanetScope multispectral optical remote sensing image data product, this invention implements a rigorous preprocessing and fine-grained annotation process. First, radiometric calibration and atmospheric correction are performed on the original optical remote sensing images sequentially. The digital quantization values ​​of each band are converted into physical reflectance, and the influence of atmospheric scattering is eliminated to enhance the spectral discernibility of small water bodies under different lighting conditions. Subsequently, precise geometric correction, spatial cropping, multi-temporal registration, and cloud and fog shadow removal are performed to ensure the spatial geometric accuracy and purity of the images. Band normalization is then used to eliminate numerical scale differences. For the 10 rivers with the largest drainage areas distributed across different geographical regions of China in Step 1, namely the Yangtze River, Heilongjiang River, Yellow River, Pearl River, Tarim River, Haihe River, Yarlung Tsangpo River, Liaohe River, Huaihe River, and Lancang River, the latitude and longitude range of typical random distribution areas of small water bodies selected within the drainage basin of each river was used. Sentinel-1 C-band radar remote sensing images were downloaded to the GEE remote sensing platform. After performing complex data conversion, multi-view processing, filtering, and geocoding on the original radar images, the radar image data were synthesized into backscattering coefficient images (unit: dB), and a vertical polarization band synthesis operation was performed. On this basis, the vertical polarization radar band images were resampled to make their spatial resolution consistent with that of PlanetScope optical remote sensing images. To further improve the water body identification performance of the model designed in this invention under complex backgrounds, this invention utilizes optical remote sensing imagery to calculate two typical spectral indices—Normalized Difference Vegetation Index (NDVI) and Normalized Difference Water Index (NDWI)—on PlanetScope optical remote sensing images of typical areas within different watersheds, and generates remote sensing index images from the calculation results. NDVI can effectively distinguish between vegetation and water bodies, avoiding misclassification of areas with high vegetation cover as water bodies. NDWI can highlight the spectral characteristics of water bodies and enhance their edge information. The specific calculation process for these two spectral indices is as follows:

[0038] Vegetation index image

[0039] Water index image

[0040] In the formula, NIR is the reflectance of the near-infrared band of PlanetScope optical remote sensing image; Red is the reflectance of the red band of PlanetScope optical remote sensing image; and Green is the reflectance of the green band of PlanetScope optical remote sensing image.

[0041] The red band image (R), green band image (G), blue band image (B), near-infrared (NIR) image, NDVI vegetation index image, NDWI water index image, and resampled vertical polarization radar remote sensing image (VV) were used as input bands. Based on the spectral characteristics of these input bands, pixel-by-pixel manual masks for typical small water bodies were created. This process accurately defines the visible boundaries of different types of small water bodies such as lakes, ponds, and wetlands, while effectively eliminating interference from similar features such as vegetation shadows and dark bare land. Finally, a standardized binary classification label dataset was constructed, providing reliable data support for the training and accuracy verification of the small water body extraction model designed in this invention. To ensure strict spatial consistency between images and labels during supervised learning, the label data and their corresponding image data were cropped into 512 × 512 pixel image blocks using the same cropping rules. During the dataset partitioning stage, image samples and their corresponding labels were randomly partitioned according to a ratio of 70% training set, 15% validation set, and 15% test set. The training set is used for learning model parameters, the validation set is used for performance evaluation and parameter tuning during model training, and the test set is used for accuracy verification and generalization ability evaluation of the final model.

[0042] Step 3: Construction of the MEF-TransUNet network model: This includes designing five parts: the model encoder, algorithm module, decoder, convolution module, and model output. Specifically, this invention addresses the limitations of existing machine learning models in spectral discrimination and high-frequency detail recovery when dealing with small water bodies, as well as the problems of fragmentation and shadow confusion in the extraction results of small water bodies. Based on the traditional TransUNet model framework, a novel segmentation network, MEF-TransUNet, designed for accurate extraction of small water bodies, is presented. A seven-channel input is constructed to enhance spectral saliency, improving target discrimination from the source. The overall structure is as follows: Figure 2 As shown.

[0043] Specifically, in the encoder stage, this invention uses ResNet-50 as the base model for feature extraction, adjusting the first convolutional layer to support seven-channel input (R, G, B, NIR, NDWI, NDVI, VV). The feature extraction process covers multiple feature levels of ResNet. The feature maps from the deep feature C4 stage are flattened and serialized, then fed into the ViT module containing a 12-layer structure, and a self-attention mechanism is used to capture long-range dependencies. Before the high-dimensional feature maps output by the decoder are mapped to the final segmentation head, this invention also sequentially introduces three complementary enhancement modules, including a multi-scale edge-guided attention module (MEGA), a high-low frequency decomposition fusion module (HLFD), and a convolutional block attention module (CBAM), aiming to enhance the model's ability to capture the boundaries and detailed structures of small water bodies. Among them, to address the limitation that a single feature source is difficult to accurately define the topological boundaries of small water bodies, the multi-scale edge-guided attention module (MEGA) constructs a three-way parallel attention mechanism that integrates the prior information of the original image and the intermediate prediction information of the network, such as... Figure 3 As shown. This module first receives the current feature map F. in A coarsely predicted probability map P and the original image I are used as multi-source inputs. The Laplacian pyramid algorithm is used to extract the high-frequency texture information E of the first layer of the original image. org Upsampling and alignment are performed to compensate for the original details lost during downsampling. Simultaneously, inverse operations and Laplacian edge detection are performed on the coarse prediction map after Sigmoid activation to generate a background mask M that suppresses non-water regions. bg Boundary mask M that enhances the target contour edge The above three types of information are weighted and fused with the input features respectively. The generation of feature components can be formalized as follows:

[0044]

[0045]

[0046]

[0047] In the formula, F represents the Sigmoid activation function; bg Indicates background features; F edge Indicates boundary features; F org Indicates edge enhancement features; L Represents the Laplace operator; 1 represents element-wise multiplication; UP represents upsampling.

[0048] Subsequently, these three features are concatenated along the channel dimension and passed through a convolutional layer to generate a unified fused attention map A. This attention map is then element-wise multiplied with the fused convolutional features to achieve dynamic weighting of the features. After residual connections are established with the original input features through element-wise addition, the output features are finally refined by the MEGA module. The calculation formula is as follows:

[0049]

[0050] In the formula, F MEGA This indicates the final output of the module, where Conv represents the convolution operation.

[0051] Given that high-frequency details are easily lost during deep network downsampling of small water bodies, the High-Low Frequency Decomposition and Fusion (HLFD) module involved in this invention introduces a frequency domain analysis perspective and employs a decoupling-independent enhancement-adaptive fusion strategy to reconstruct features, such as... Figure 4 As shown, the input features are first processed by an average pooling operation (AvgPool) to extract low-frequency overview components. High-frequency detail components containing edge textures are separated by interpolation subtraction. This achieves initial decoupling in the frequency domain, and the decomposition process is defined as follows:

[0052]

[0053]

[0054] Subsequently, the low-frequency components are fed into a sub-network embedding two-dimensional discrete wavelet transform (DWT) and Transformer, leveraging the long-range modeling capabilities of the Transformer to enhance the global semantic structure and macroscopic topological relationships in the low-frequency domain, resulting in... High-frequency components are processed through convolutional blocks containing dense connections to enhance gradient propagation and feature reuse capabilities for minute textures, resulting in... In the fusion stage, the enhanced high- and low-frequency features are concatenated along the channel dimension and processed by a 1×1 convolutional dimensionality reduction and a channel attention layer (CALayer). This layer adaptively calculates the importance weights of each channel, thereby dynamically adjusting the contribution ratio of high- and low-frequency information in the final output. The final feature reconstruction formula is:

[0055]

[0056] In the formula, F HLFD Indicates the output characteristics of the module; This indicates the low-frequency characteristics after processing; denoted by , represents the enhanced high-frequency features; f represents the convolution operation. This equation shows that while recovering macroscopic semantics, the model effectively preserves accurate boundary details through the residual path.

[0057] To further enhance the sensitivity of feature maps to water targets and suppress background noise response, the Convolutional Block Attention Module (CBAM) involved in this invention, such as... Figure 5 As shown, this module follows a serial processing logic, aiming to achieve adaptive calibration of the feature map. First, the channel attention submodule extracts the spatial context descriptor of the feature map F through parallel global average pooling (AvgPool) and global max pooling (MaxPool), and feeds it into a shared multilayer perceptron (MLP) to capture the dependencies between channels, generating channel weight vectors. Used to recalibrate the channel dimension of the input features and obtain :

[0058]

[0059]

[0060] Subsequently, the feature map was calibrated by the channel. Average pooling and max pooling are performed on the channel dimension to generate a two-dimensional spatial feature descriptor, and local spatial information is aggregated through a large-kernel convolutional layer to generate a spatial attention map. The spatial weight matrix is ​​ultimately applied to the feature map to complete the module's output feature F. CBAM Final selection and refinement:

[0061]

[0062]

[0063] Through the above steps, the model ultimately emphasizes the target area of ​​the water body with pixel-level accuracy and can effectively suppress irrelevant background interference.

[0064] In the decoder stage, a cascaded upsampling strategy is used to gradually restore the spatial resolution of the feature maps. During this process, a selective feature fusion strategy is implemented: the decoder only concatenates the upsampled deep semantic features with stages C4, C3, and C2 in the encoder (corresponding to 1 / 16, 1 / 8, and 1 / 4 resolutions, respectively), and discards the C1 shallow features (1 / 2 resolution) which are rich in high-frequency noise. This is because although the decoding stage corresponding to the C1 shallow features rich in high-frequency noise has high spatial resolution, due to insufficient network representation depth, it is filled with a large amount of high-frequency non-water body noise such as unfiltered building shadows and dark bare ground, directly introducing background response overload and computational redundancy at the decoder, leading to shadow misjudgment. Therefore, this invention blocks the C1 shallow features rich in high-frequency noise to... To eliminate shallow clutter and restore water body boundaries without introducing shallow noise, a precise mechanism is employed by the subsequent MEGA and HLFD modules after the decoder completes cascaded upsampling and before the final result is output by the segmentation head. Specifically, the MEGA module directly introduces the high-frequency texture prior from the original image after Laplacian transform, achieving direct restoration of clean edge information while bypassing shallow noise contamination. Combined with the HLFD module, this enables independent enhancement of minute textures in the frequency domain. This feature selection strategy at the decoder end effectively suppresses complex background noise while ensuring the integrity of the topological boundaries of small water bodies, achieving multi-dimensional synergy between global denoising and local fidelity preservation.

[0065] Step 4, Model Training and Parameter Optimization: This mainly involves training and testing the MEF-TransUNet model designed in this invention using different module combinations from Step 3, thereby verifying the effectiveness and specific contributions of the overall network model architecture described in this invention. Supervised learning is used during the model training phase. Considering the imbalance between water and non-water categories in remote sensing scenes, this invention applies a hybrid loss function composed of binary cross-entropy loss (BCE) and Dice loss to balance pixel-level classification accuracy with overall region overlap. The expression for the BCE loss function is:

[0066]

[0067] In the formula, and Let and represent the true label and predicted probability of the i-th pixel, respectively.

[0068] Dice loss is used to measure the degree of overlap between the predicted area and the actual water body area, and its form is as follows:

[0069]

[0070] In the formula, To prevent the stability constant from having a denominator of zero.

[0071] BCE and Dice are used together in a weighted manner during training, and their combined loss is:

[0072]

[0073] In the formula, λ is a hyperparameter, which is usually taken as 0.5 to balance the contributions of both.

[0074] During training, this invention uses the Adam optimizer to set the initial learning rate to... Weight decay is set to In conjunction with the cosine annealing strategy, the minimum learning rate is set to To ensure rapid convergence in the early stages of training and stable optimization in later stages, a training batch size of 8 and a total of 80 training epochs were implemented. To improve the model's generalization ability and suppress overfitting, a rigorous data augmentation strategy was implemented during the data preprocessing stage. Specifically, geometric augmentation included random 90° rotations and horizontal and vertical flips, while color augmentation randomly applied perturbations to hue, saturation, brightness, and contrast on the RGB channels to simulate differences in lighting and sensor imaging conditions. IoU, Precision, and Recall were calculated on the validation set for each training epoch to comprehensively evaluate model performance. Model weights were saved based on the weights with the minimum validation loss after each validation epoch; after epochs ≥ 60, the optimal weights were additionally saved to balance global optimum and later convergence, ensuring the final model maintains overall performance while also considering later optimization potential. Training was performed entirely in a GPU parallel environment to improve the efficiency of large-scale high-resolution remote sensing image processing.

[0075] To verify the model's generalization performance, this invention performed rigorous testing and evaluation after model training. During the testing phase, the optimal weights with the minimum loss on the validation set were first loaded, and the network was switched to evaluation mode. In this mode, all Batch Normalization layers in the network used the moving average and variance statistically obtained during training to ensure the determinism of the inference results. The preprocessing of the test data remained consistent with the training phase, but random augmentation operations such as rotation or flipping were no longer performed to maintain the spatial distribution characteristics of the original images. The inference process was performed in a non-gradient environment, reducing memory usage while improving computational efficiency. Furthermore, after sigmoid activation, the model output was binarized with a threshold of 0.65 to generate the final segmentation mask. Through threshold iteration experiments on the validation set, 0.65 was found to be the optimal threshold balancing precision and recall. Finally, by comparing with the ground truth labels, the average index of the entire test set was calculated, thereby comprehensively quantifying the model's overall accuracy and internal consistency in the small water body extraction task.

[0076] Step 5: Automatic Extraction of Small Water Bodies in the Detection Area: Based on the geographical latitude and longitude range of the area to be detected for small water bodies, select and download Level-3B analytical-grade 3-meter spatial resolution multispectral optical remote sensing image data products with surface reflectance correction from the Planet Labs website (https: / / www.planet.com / ). Radiometric calibration and atmospheric correction are then performed on the high spatial resolution PlanetScope multispectral optical remote sensing image, converting digital quantization values ​​into physical reflectance and eliminating atmospheric scattering effects to enhance the spectral discernibility of small water bodies under different lighting conditions. Subsequently, through precise geometric correction, spatial cropping, multi-temporal registration, and cloud / shadow removal, the spatial geometric accuracy and purity of the image are ensured, and band normalization is used to eliminate numerical scale differences. Simultaneously, based on the geographical latitude and longitude range of the area to be detected in small water bodies, Sentinel-1 C-band radar remote sensing images were downloaded to the GEE remote sensing platform. After performing complex data conversion, multi-view processing, filtering, and geocoding on the original radar images, the radar image data were synthesized into a DB image with backscattering coefficients, and a vertical polarization band synthesis operation was performed. Furthermore, the vertical polarization radar band image was resampled to ensure its spatial resolution was consistent with the PlanetScope optical remote sensing image. Further, this invention utilizes the PlanetScope optical remote sensing image to extract the normalized vegetation index and normalized water index for the area to be detected in small water bodies, and generates a remote sensing index image from the calculation results. Subsequently, the red band image (R), green band image (G), blue band image (B), near-infrared band image (NIR), vegetation index image (NDVI), water index image (NDWI), and the resampled vertical polarization radar remote sensing image (VV) of the optical remote sensing image within the area to be detected in small water bodies were used as the seven input variables of the model. The seven input variable images are input into the MEF-TransUNet model established in step 4 to accurately extract small water bodies within the detection area.

[0077] To verify the superiority of the method of this invention in the task of accurately extracting small water bodies, this invention also randomly selected nine images of small water bodies covering different water bodies and land use types, and conducted comparative verification experiments on the extraction effect of small water bodies. First, a comparative experiment was conducted with other classic models. Several mainstream models were introduced, including variants of U-Net (U-Net, AttenUNet, UNet++), DeepLabV3+, and SegFormer based on the Transformer architecture. The performance of the model proposed in this invention in the task of identifying small water bodies was comprehensively evaluated. The specific evaluation visualization results are shown below. Figure 6As shown, in terms of preserving the morphology of small water bodies, compared with the traditional machine model mentioned above, the model framework proposed in this invention exhibits the best extraction performance in terms of topological coherence, boundary detail delineation, and overall stability. Cross-sectional testing confirms the superior performance of this method in cross-model architecture comparisons.

[0078] To verify the specific effectiveness of the MEGA, HLFD, and CBAM modules, this invention also randomly selected nine other small water body images covering different water bodies and land use types to design controlled variable ablation experiments. This was done to isolate and quantify the performance contribution of each core module. All ablation experiments were performed under identical training conditions and parameter settings to ensure the fairness and comparability of the experimental results. The ablation experiments of this invention used TransUNet as the baseline network and constructed four progressive schemes: by independently embedding the MEGA, HLFD, and CBAM modules, the performance of edge guidance, frequency domain denoising, and attention enhancement mechanisms in solving bottleneck problems such as water body fragmentation, background obfuscation, and weak signals was tested. Subsequently, the three modules were organically integrated to construct a complete MEF-TransUNet model, comprehensively evaluating the overall improvement effect of the multi-dimensional feature collaboration strategy on fine-grained segmentation performance. The visualization results of the ablation experiments are shown below. Figure 7 As shown, compared to the TransUNet baseline network which suffers from missed detections and false positives, and the scheme that introduces each module separately, the MEF-TransUNet model, which integrates three modules, achieves the best accuracy in improving the sensitivity to small targets and suppressing background interference.

Claims

1. A method for precise extraction of small water bodies combining remote sensing and machine learning, characterized in that, Includes the following steps: Step 1, Study Area Selection and Sample Distribution Design: Select a typical river basin as the study area, and select typical small water body distribution areas with spatial heterogeneity within the river basin, covering different geographical zones, land use types and small water body types; Step 2, Multi-source remote sensing data preprocessing and dataset construction: For the typical area, download and preprocess optical remote sensing images and radar remote sensing images, calculate spectral indices, and construct a seven-channel input dataset and a binary label dataset containing red band images, green band images, blue band images, near-infrared band images, vegetation index images, water index images, and resampled vertical polarization radar remote sensing images from optical remote sensing images. Step 3: Construct the MEF-TransUNet network model: Based on the improved TransUNet, construct the MEF-TransUNet deep learning segmentation model. The encoder uses a combination of ResNet-50 and ViT, and the decoder discards shallow noise features of C1 and restores resolution through cascaded upsampling. At the end of the decoder, a multi-scale edge-guided attention module, a high-low frequency decomposition and fusion module, and a convolutional block attention module are integrated in sequence, and a selective feature fusion strategy is adopted. Step 4, Model Training and Parameter Optimization: Train, validate, and test the MEF-TransUNet model using a hybrid loss function, and save the optimal model weights; Step 5: Water body extraction in the area to be detected: Download the same type of remote sensing image for the area to be detected and preprocess it, then input it into the trained MEF-TransUNet model to achieve accurate extraction of small water bodies.

2. The method for precise extraction of small water bodies combining remote sensing and machine learning according to claim 1, characterized in that, In step 2, a 3-meter resolution PlanetScope multispectral optical image is acquired, including red, green, blue, and near-infrared bands, and preprocessed by radiometric calibration, atmospheric correction, geometric fine correction, cloud and fog removal, and normalization. The radar remote sensing image is a C-band vertical polarization image, which is preprocessed by multi-view, filtering, geocoding, and resampling to match the resolution of the optical image. The normalized vegetation index and normalized water index of the optical image are calculated.

3. The method for precise extraction of small water bodies combining remote sensing and machine learning according to claim 1, characterized in that, In step 2, binary labels are created, dividing the dataset into a 70% training set, a 15% validation set, and a 15% test set.

4. The method for precise extraction of small water bodies combining remote sensing and machine learning according to claim 1, characterized in that, In step 3, the encoder changes the first layer convolution of ResNet-50 to a seven-channel input, and the deep C4 features are fed into the 12-layer ViT to capture long-range dependencies; the decoder adopts cascaded upsampling, only splicing the features of the C4, C3, and C2 layers, and discarding the shallow C1 features.

5. The method for precise extraction of small water bodies combining remote sensing and machine learning according to claim 1, characterized in that, In step 3, the multi-scale edge-guided attention module takes the feature map, the predicted probability map, and the original image as input, extracts edge information through the Laplacian pyramid, constructs a three-way parallel attention mechanism, and outputs it through residual connection.

6. The method for precise extraction of small water bodies combining remote sensing and machine learning according to claim 1, characterized in that, In step 3, the high-low frequency decomposition and fusion module decouples the features into low-frequency and high-frequency components. The low-frequency components are enhanced by discrete wavelet transform and Transformer, while the high-frequency components are enhanced by densely connected convolutional blocks. Finally, the features are adaptively fused through channel attention.

7. The method for precise extraction of small water bodies combining remote sensing and machine learning according to claim 1, characterized in that, In step 3, the convolutional block attention module performs channel attention and spatial attention calibration sequentially. Channel attention uses a combination of global average pooling and global max pooling in a multilayer perceptron, while spatial attention uses large kernel convolution to generate a spatial weight map.

8. The method for precise extraction of small water bodies combining remote sensing and machine learning according to claim 1, characterized in that, In step 4, the hybrid loss function is a weighted fusion of the binary cross-entropy (BCE) loss and the Dice loss, with a weight coefficient λ of 0.5; the Adam optimizer is used, and the initial learning rate is... Weight decay Cosine annealing to minimum learning rate The batch size is 8, and the total number of epochs is 80.

9. The method for precise extraction of small water bodies combining remote sensing and machine learning according to claim 8, characterized in that, In step 4, the optimal weights are loaded during model testing, and a water mask is generated with a binarization threshold of 0.

65. The accuracy is evaluated by IoU, precision, and recall.

10. A system for precise extraction of small water bodies combining remote sensing and machine learning, used to implement the method for precise extraction of small water bodies combining remote sensing and machine learning as described in any one of claims 1-9, characterized in that, include: The module includes a sample design module, a data processing and dataset construction module, a MEF-TransUNet network model construction module, a model training and optimization module, and a water extraction execution module. The sample design module is used to select typical river basins as study areas, and within the river basins, to select typical small water bodies with spatial heterogeneity, covering different geographical zones, land use types, and small water body types. The data processing and dataset construction module is used to download and preprocess optical and radar remote sensing images for the typical region, calculate spectral indices, and construct a seven-channel input dataset and a binary classification label dataset containing multi-band and remote sensing indices. Specifically, it acquires 3-meter resolution PlanetScope multispectral optical images, including red, green, blue, and near-infrared bands, and performs radiometric calibration, atmospheric correction, geometric fine correction, cloud and fog removal, and normalization. The radar remote sensing images are C-band vertical polarization images, processed by... After multi-viewing, filtering, geocoding, and resampling, the data is matched with the resolution of the optical imagery. The normalized vegetation index and normalized water index of the optical imagery are calculated to construct a seven-channel input data set containing red band, green band, blue band, near-infrared band, vegetation index, and water index images of the optical remote sensing imagery, as well as the resampled vertical polarization radar remote sensing imagery. The dataset is then cropped into 512×512 image patches, binary classification labels are created, and the dataset is divided into training, validation, and test sets in a ratio of 7:1.5:1.

5. The MEF-TransUNet network model building module is used to construct a MEF-TransUNet deep learning segmentation model based on an improved TransUNet. The encoder uses a combination of ResNet-50 and ViT, changing the first convolution of ResNet-50 to a seven-channel input, and feeding the deep C4 features into a 12-layer ViT to capture long-range dependencies. The decoder discards shallow noise features of the C1 layer, uses cascaded upsampling to restore resolution, and only concatenates features from the C4, C3, and C2 layers. At the end of the decoder, a multi-scale edge-guided attention module, a high-low frequency decomposition and fusion module, and a convolutional block attention module are integrated in sequence, and a selective feature fusion strategy is adopted. The multi-scale edge-guided attention module takes feature maps, predicted probability maps, and original images as inputs, extracts edge information through Laplacian pyramids, constructs a three-way parallel attention mechanism, and outputs it through residual connections. The high-low frequency decomposition and fusion module decouples features into low-frequency and high-frequency components. The low-frequency components are enhanced by discrete wavelet transform and Transformer, and the high-frequency components are enhanced by densely connected convolutional blocks. Then, they are adaptively fused through channel attention. The convolutional block attention module performs channel attention and spatial attention calibration in sequence. Channel attention uses a combination of global average pooling and global max pooling with a multilayer perceptron, and spatial attention uses large-kernel convolution to generate a spatial weight map. The model training and optimization module is used to train, validate, and test the MEF-TransUNet model using a hybrid loss function, and to save the optimal model weights. The hybrid loss function is a weighted fusion of binary cross-entropy (BCE) loss and Dice loss, with a weight coefficient λ of 0.

5. The Adam optimizer is used, and the initial learning rate is set to... Weight decay Cosine annealing to minimum learning rate The batch size is 8, and the total number of epochs is 80. During model testing, the optimal weights are loaded, and a water mask is generated with a binarization threshold of 0.

65. The accuracy is evaluated by IoU, precision, and recall. The water body extraction execution module is used to download and preprocess the same type of remote sensing images of the area to be detected, and input them into the trained MEF-TransUNet model to achieve accurate extraction of small water bodies.