A method for extracting water body information from remote sensing images by combining weak supervision and sample migration
Patent Information
- Application Number
- CN202610041338.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-13
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-01-13
AI Technical Summary
[0006]本发明的目的在于克服传统技术中存在的上述问题,提供一种联合弱监督与样本迁移的遥感影像水体信息提取方法,采用跨尺度样本迁移、NDWI约束的类激活映射优化以及多区域联合训练框架,有效解决了现有方法在水体提取过程中的标注依赖、精度和鲁棒性问题
[0025]本发明实现了在高分辨率遥感影像中,无需人工标注数据的条件下,自动、准确地提取水体信息。本发明方法在水体提取精度、处理效率和跨区域适用性上均具有显著优势,尤其适用于洪涝灾害应急监测和大尺度水体动态监测等实时应用场景。
Smart Images

Figure CN121837935B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of remote sensing geoscience application technology, specifically relating to a method for extracting water body information from remote sensing images by combining weak supervision and sample migration. Background Technology
[0002] High-resolution remote sensing imagery for water body extraction is of great value in fields such as flood monitoring and water resource management. Existing water body extraction methods can be mainly divided into three categories: spectral index-based methods, machine learning-based methods, and deep learning-based methods.
[0003] Spectral index-based methods identify water bodies by constructing mathematical combinations between different spectral bands. While computationally simple and fast, these methods are susceptible to interference from shadows and buildings in complex environments, and the lack of key spectral bands required for water body indices in high-resolution imagery limits their application. Machine learning-based methods can learn multidimensional features of water bodies, but they suffer from limitations such as computational complexity, strong dependence on training data, susceptibility to overfitting, and insufficient spatial transfer capabilities.
[0004] Deep learning-based methods can automatically learn multi-level feature representations of water bodies and perform well in complex scenarios. However, deep learning methods rely heavily on a large number of manually labeled samples. The labeling process is time-consuming and costly, and the differences in water body characteristics in different regions make it difficult to transfer labeled data across regions, which seriously restricts their rapid deployment in practical scenarios such as emergency response.
[0005] Therefore, how to quickly and accurately extract water bodies from high-resolution remote sensing images without manual annotation has become a pressing technical challenge. Summary of the Invention
[0006] The purpose of this invention is to overcome the aforementioned problems in traditional technologies and provide a method for extracting water information from remote sensing images by combining weak supervision and sample transfer. This method employs cross-scale sample transfer, NDWI-constrained class activation mapping optimization, and a multi-region joint training framework, which effectively solves the problems of annotation dependence, accuracy, and robustness in the water extraction process of existing methods.
[0007] To achieve the above-mentioned technical objectives and effects, the present invention is implemented through the following technical solution:
[0008] This invention provides a method for extracting water body information from remote sensing images by combining weak supervision and sample migration, comprising the following steps:
[0009] S1. Based on spatial matching of medium-resolution NDWI imagery and high-resolution RGB imagery, a sliding window strategy and threshold decision mechanism are used to generate scene-level water / non-water body samples.
[0010] S2. Train a weakly supervised classification model using scene-level samples, generate class activation maps, and perform logical AND fusion with NDWI images to optimize the generation of pixel-level pseudo-labels.
[0011] S3. A training set is constructed based on pixel-level pseudo-labels of multiple cities. The DeepLabV3+ semantic segmentation model is used for cross-regional joint training. A composite loss function and a balanced sampling strategy are used to achieve high-precision water body extraction.
[0012] Furthermore, in step S1, the sliding window size is 128×128 pixels, and the step size is 128 pixels.
[0013] Furthermore, in step S1, the segmentation threshold of NDWI is preset. When the proportion of pixels with NDWI > 0 reaches 30% or more, it is marked as a water scene and assigned a value of 1; when the proportion of pixels with NDWI ≤ 0 reaches 95% or more, it is marked as a non-water scene and assigned a value of 0; scenes that do not meet the above conditions are classified as uncertain samples, assigned a value of 2, and are removed.
[0014] Furthermore, in step S2, the weakly supervised classification model uses ResNet-101 as the backbone network and ImageNet pre-trained weights.
[0015] The fusion method of class activation graph and NDWI is as follows:
[0016] ;
[0017] In the formula, This represents the optimized activation graph. This represents the original class activation graph, where I is the indicator function.
[0018] Furthermore, in step S3, the semantic segmentation model adopts the DeepLabV3+ architecture, with the backbone network being ResNet-101; a composite loss function is used, which is a weighted combination of Focal Loss and Dice Loss, expressed as:
[0019] ;
[0020] In the formula, The recommended value for the balance coefficient is 0.5.
[0021] Furthermore, in step S3, a cross-regional joint training strategy is adopted.
[0022] Furthermore, in step S3, a cross-regional balanced sampling strategy is adopted, and the sampling weights are allocated according to the square root of the total number of pixels in the training city images.
[0023] Furthermore, the method is applicable to GF-2 and Sentinel-2 multi-source remote sensing images, supporting water body extraction tasks across sensors and resolutions; the method achieves water body extraction without manual annotation, and is suitable for flood disaster emergency response and dynamic water resource monitoring application scenarios.
[0024] The beneficial effects of this invention are:
[0025] This invention enables the automatic and accurate extraction of water body information from high-resolution remote sensing images without the need for manual data annotation. The method of this invention has significant advantages in water body extraction accuracy, processing efficiency, and cross-regional applicability, and is particularly suitable for real-time applications such as flood emergency monitoring and large-scale dynamic monitoring of water bodies.
[0026] Of course, any product implementing this invention does not necessarily need to achieve all of the above advantages at the same time. Attached Figure Description
[0027] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is a flowchart of the method of the present invention;
[0029] Figure 2 Example images of scene-level samples automatically generated for Wuhan, Zhengzhou, and Guangzhou;
[0030] Figure 3 This is a schematic diagram of the cross-scale sample migration process from medium-resolution NDWI to high-resolution scene labels based on a determined threshold.
[0031] Figure 4 This is a schematic diagram of CAM generation and optimization.
[0032] Figure 5 Comparison of the final water body segmentation results for each study area;
[0033] Among them, (a) and (b) are magnified views of parts of Nanjing, and (c), (d), and (e) are magnified views of parts of Tianjin.
[0034] Figure 6 Heatmaps of F1 scores for class activation maps under different NDWI threshold combinations were generated.
[0035] Figure 7 A comparison chart showing whether CAM and NDWI intersect and merge;
[0036] Among them, (a) and (b) are magnified views of parts of Wuhan, (c) and (d) are magnified views of parts of Zhengzhou, and (e) and (f) are magnified views of parts of Guangzhou. Detailed Implementation
[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] This invention discloses a method for extracting water body information from remote sensing images by combining weak supervision and sample migration, comprising the following steps:
[0039] Step S1: Scene-level sample generation for cross-scale knowledge transfer
[0040] The Normalized Differential Water Index (NDWI) was calculated based on medium-resolution Sentinel-2 remote sensing imagery. Preliminary water body and non-water body discrimination results were generated at 10 m resolution using a threshold segmentation method, and spatial matching was performed with the corresponding 0.5 m high-resolution RGB imagery. A sliding window strategy was employed to traverse the high-resolution imagery, and scene-level labels were assigned based on the pixel statistical features within the corresponding NDWI window, achieving effective migration from medium-resolution pixel-level water body index to high-resolution scene-level labels.
[0041] Step S2: Generation of pseudo-labels for class activation mapping of NDWI constraints
[0042] A weakly supervised classification model is constructed using scene-level training samples, and a spatial activation map of water bodies is generated through a Class Activation Map (CAM) mechanism. To address the issue that traditional CAM only activates the most salient parts of water bodies and suffers from ambiguous boundary localization, a logical AND operation is performed between the generated CAM and spatially registered medium-resolution NDWI imagery. This systematically removes potentially misclassified pixels that are not water bodies as indicated by NDWI, generating complete and accurate pixel-level pseudo-labels for water bodies.
[0043] Step S3: Construction of a Multi-Region Robust Semantic Segmentation Framework
[0044] A semantic segmentation model is constructed using the DeepLabV3+ architecture based on pixel-level pseudo-labels. To ensure the model's cross-regional generalization ability, a multi-city joint training strategy is adopted, and a balanced sampling strategy based on the square root of the region area and a composite loss function of FocalLoss and Dice Loss are introduced. This not only ensures accurate classification of the main water body region, but also enhances the ability to accurately characterize boundary details.
[0045] Specific embodiments of the present invention are as follows:
[0046] like Figure 1 As shown in the figure, this embodiment provides an automatic method for extracting water body information from high-resolution remote sensing images by combining weak supervision and sample migration, including the following steps:
[0047] Step S1: Scene-level sample generation for cross-scale knowledge transfer
[0048] Acquire Sentinel-2 and high-resolution RGB images covering the same area. Calculate the Normalized Difference Water Index (NDWI):
[0049] ;
[0050] in, and These represent the surface reflectance of the Sentinel-2 image in the green band (B3) and near-infrared band (B8), respectively.
[0051] like Figure 2 and Figure 3 As shown, this method employs spatial registration technology to establish a precise spatial correspondence between a 10 m resolution NDWI image and a 0.5 m resolution high-resolution RGB image. A 128×128 pixel sliding window is used to traverse the high-resolution image with a step size of 128 pixels, simultaneously extracting the corresponding medium-resolution NDWI window at the spatial location.
[0052] Suppose that within a certain scene window, the NDWI is greater than the segmentation threshold. The pixel ratio is NDWI is less than or equal to The pixel ratio is The sample classification rule is: when When the value is ≥0.3, the corresponding high-resolution scene is labeled as a water sample; when When the value is ≥0.95, it is marked as a non-water sample; scenes that do not meet the criteria are removed. This decision-making strategy effectively avoids classification uncertainty caused by mixed pixels in the boundary region.
[0053] Step S2: Generation of pseudo-labels for class activation mapping of NDWI constraints
[0054] A pre-trained ResNet-101 was used as the feature extraction backbone network. Global average pooling layers and fully connected layers were integrated at the ends of the backbone network to construct a scene-level classification model. During the training phase, the network parameters were optimized by minimizing the cross-entropy loss function.
[0055] like Figure 4As shown, CAM generates a spatial attention map by calculating a weighted linear combination of the feature maps from the last convolutional layer of ResNet-101. For water class c, the class activation value at location (x,y) in the image is... The calculation is as follows:
[0056] ;
[0057] in, K Indicates the number of feature channels. The first fully connected layer k Each channel is a category c The weighting coefficients, For the first k Each convolutional feature map at location (x,y) The activation response value at that location.
[0058] Traditional CAM typically activates only the most prominent parts of water bodies rather than the entire target, and activation values in edge regions exhibit gradient diffusion. This embodiment proposes a CAM optimization mechanism based on NDWI spectral constraints. A logical AND operation is performed between the generated preliminary water body activation map and a spatially registered medium-resolution NDWI image to remove potentially misclassified pixels that are not water bodies according to NDWI. The optimized activation map is represented as follows:
[0059] ;
[0060] in, This represents the optimized activation graph. Primitive class activation graph, I This is the indicator function. The optimization strategy integrates deep learning features with water spectral response characteristics to generate high-precision pixel-level pseudo-labels for water bodies.
[0061] Step S3: Construction of a Multi-Region Robust Semantic Segmentation Framework
[0062] The DeepLabV3+ architecture is adopted as the core segmentation framework, with ResNet-101 as the feature extraction backbone network. The dilated spatial pyramid pooling module of DeepLabV3+ consists of four parallel dilated convolutional layers, which effectively expand the receptive field and capture multi-scale contextual information by using different dilation rates.
[0063] To ensure the model's cross-regional generalization ability, joint training was conducted in three cities—Wuhan, Zhengzhou, and Guangzhou—with significant hydrological differences. A balanced sampling strategy based on the square root of the region's area was adopted for each city. Its sampling weight Total number of pixels in the city image The square root of the expression is proportional to and normalized. This strategy significantly improves the sampling opportunities in small regions while ensuring sufficient training in large areas.
[0064] ;
[0065] The loss function employs a combination of Focal Loss and Dice Loss. Focal Loss introduces a modulation factor to suppress the loss contribution of easily classified samples, allowing the model to focus on difficult-to-classify samples. Dice Loss optimizes the regional overlap between the predicted results and the true labels from a global perspective, exhibiting higher sensitivity to boundary pixels.
[0066] ;
[0067] in, This is a balancing coefficient used to coordinate global classification accuracy with the preservation of boundary details. This composite loss function ensures accurate classification of the main water body region while enhancing the precise characterization of boundary details.
[0068] Step S4: Model Inference and Water Extraction
[0069] After training, the model is applied to new high-resolution imagery for water body extraction. During the inference phase, a sliding window strategy (256×256 pixels, 128-pixel stride) is employed to perform a weighted average of the segmentation results for overlapping regions. This enables rapid and accurate extraction of water body information from new high-resolution remote sensing imagery without any manual annotation, ultimately outputting a complete water body segmentation mask.
[0070] This embodiment describes the construction process of an automated water extraction method based on combined weak supervision and sample migration, as follows:
[0071] 1) Selection of study area and construction of dataset
[0072] To systematically verify the generalization and robustness of this method, five representative cities with significant geographical differences within China were selected as the study area: Guangzhou (Guangdong Province), Nanjing (Jiangsu Province), Tianjin, Wuhan (Hubei Province), and Zhengzhou (Henan Province). Each region encompasses various water body types, mainly including rivers, lakes, ponds, and reservoirs, as well as complex environmental features such as building shadows, mountain shadows, small water bodies, and turbid water. The study used 0.5-meter high-resolution Google Earth RGB imagery and 10-meter resolution Sentinel-2 imagery as data sources, and pixel-level ground truth labels were generated through manual annotation for verification. All images underwent preprocessing including radiometric calibration, atmospheric correction, cloud detection, and median synthesis, and geometric registration was implemented to ensure spatial consistency.
[0073] 2) Scene-level sample generation
[0074] Normalized Differential Water Index (NDWI) was calculated based on Sentinel-2 imagery, and a reasonable global optimal threshold was set. A sliding window strategy (128×128 pixels, 128-pixel step size) was used to generate scene-level water and non-water body labels for high-resolution RGB images. Samples were filtered based on pixel statistical characteristics within the NDWI window: scenes with NDWI greater than 0 were labeled as water scenes (over 30%), and scenes with NDWI less than or equal to 0 were labeled as non-water scenes (over 95%). Scenes that did not meet these criteria were discarded.
[0075] 3) Pixel-level pseudo-tag generation
[0076] A ResNet-101 classification model was trained using scene-level labels from three cities: Wuhan, Zhengzhou, and Guangzhou. The model was trained for 30 epochs using the AdamW optimizer (learning rate 0.0002) and Focal Loss. Subsequently, class activation maps were generated, and pixel-level pseudo-labels were generated using an NDWI constraint optimization strategy. This optimization strategy performs a logical AND operation between CAM and NDWI images, effectively eliminating misclassified pixels that are not water bodies.
[0077] 4) Semantic segmentation model training and optimization
[0078] A DeepLabV3+ segmentation model (ResNet-101 backbone network) was constructed with an input size of 256×256 pixels. The Adam optimizer (learning rate 0.001) was used, and the loss function was a weighted combination of Focal Loss (α=[3.0, 1.0], γ=2.0) and Dice Loss (weights...). =0.5). Joint training was conducted on the optimized pseudo-labels of Wuhan, Zhengzhou, and Guangzhou, for a total of 50 rounds.
[0079] 5) Accuracy assessment and comparative analysis
[0080] To comprehensively verify the overall performance of this invention, this embodiment conducts a systematic evaluation from three aspects: quantitative accuracy, qualitative quality, and the effectiveness of core components. Experimental results show that this invention achieves water extraction performance comparable to fully supervised methods without requiring any manual labeling, and each core component has been proven to play a crucial role in improving the final performance.
[0081] In terms of quantitative accuracy, under completely unannotated conditions, the overall accuracy of this method on the Nanjing and Tianjin test sets reached 96.67%, the F1 score was 87.27%, and the intersection-union ratio was 77.54%. Compared with fully supervised deep learning methods that rely on a large amount of manual annotation, the differences in all indicators are controlled within a very small range, proving that the technical framework constructed in this embodiment can achieve high-precision automated extraction of water body information without relying on costly manual annotation.
[0082] In terms of qualitative quality, the visualization results further confirm the effectiveness of this method. For example... Figure 5 As shown, this method exhibits significant advantages in the Nanjing and Tianjin test areas. Regarding anti-interference capabilities, this method effectively suppresses misidentification of noise points such as building shadows and reflections generated by the NDWI method. In terms of boundary refinement, the water body boundaries generated by this method are smoother and more coherent, avoiding the jagged and fragmented phenomena of NDWI results. Compared with fully supervised methods, this method performs comparably in most scenarios.
[0083] In terms of evaluating core algorithm components, such as Figure 6 As shown, NDWI threshold sensitivity analysis indicates that when At that time, the model performed best in all study areas: the F1 score was 87.08% in Wuhan, 78.67% in Guangzhou, and 90.78% in Zhengzhou, with an average F1 score of 85.51%. Figure 7 As shown, the effectiveness analysis of the intersection fusion strategy of CAM and NDWI shows that by utilizing the strong semantic prior of CAM to filter out noise caused by shadows in NDWI, and by using the fine spectral information of NDWI to make up for the shortcomings of CAM in boundary details, the best water extraction effect can be achieved in complex urban scenes.
[0084] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to specific implementations. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for extracting water body information from remote sensing images by combining weak supervision and sample migration, characterized in that, Includes the following steps: S1. Based on spatial matching of medium-resolution NDWI imagery and high-resolution RGB imagery, a sliding window strategy and threshold decision mechanism are used to generate scene-level water / non-water body samples. Preset NDWI segmentation threshold When the proportion of pixels with NDWI > 0 reaches 30% or more, it is marked as a water scene and assigned a value of 1; when the proportion of pixels with NDWI ≤ 0 reaches 95% or more, it is marked as a non-water scene and assigned a value of 0; scenes that do not meet the above conditions are classified as uncertain samples, assigned a value of 2, and are removed. S2. Train a weakly supervised classification model using scene-level samples, generate class activation maps, and perform logical AND fusion with NDWI images to optimize the generation of pixel-level pseudo-labels. The weakly supervised classification model uses ResNet-101 as the backbone network and ImageNet pre-trained weights. The fusion method of class activation graph and NDWI is as follows: ; In the formula, This represents the optimized activation graph. This represents the primitive class activation graph, where I is the indicator function; S3. A training set is constructed based on pixel-level pseudo-labels of multiple cities. The DeepLabV3+ semantic segmentation model is used for cross-regional joint training. A composite loss function and a balanced sampling strategy are used to achieve high-precision water body extraction. The semantic segmentation model adopts the DeepLabV3+ architecture, with ResNet-101 as the backbone network; it uses a composite loss function, which is a weighted combination of Focal Loss and Dice Loss, expressed as follows: ; In the formula, The recommended value for the balance coefficient is 0.
5.
2. The method for extracting water body information from remote sensing images by combining weak supervision and sample migration as described in claim 1, characterized in that, In step S1, the sliding window size is 128×128 pixels, and the step size is 128 pixels.
3. The method for extracting water body information from remote sensing images by combining weak supervision and sample migration as described in claim 1, characterized in that, In step S3, a cross-regional joint training strategy is adopted.
4. The method for extracting water body information from remote sensing images by combining weak supervision and sample migration as described in claim 1, characterized in that, In step S3, a cross-regional balanced sampling strategy is adopted, and the sampling weights are allocated according to the square root of the total number of pixels in the training city images.
5. The method for extracting water body information from remote sensing images by combining weak supervision and sample migration as described in claim 1, characterized in that, The method is applicable to GF-2 and Sentinel-2 multi-source remote sensing images and supports water body extraction tasks across sensors and resolutions. The method achieves water body extraction without manual annotation and is suitable for application scenarios such as flood disaster emergency response and dynamic water resource monitoring.
Citation Information
Patent Citations
Satellite remote sensing image water extraction method
CN113642663A