A method, system, and device for water body extraction from remote sensing images based on UNetMFormer network

By combining global and local feature extraction structures with the UNetMFormer network, the problem of insufficient generalization ability in water body extraction from remote sensing images is solved, achieving higher accuracy and faster processing speed.

CN117152607BActive Publication Date: 2025-10-31FUJIAN SATELLITE DATA DEV CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311068122.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-23
Publication Date
2025-10-31
Estimated Expiration
2043-08-23

AI Technical Summary

Technical Problem

Existing methods for extracting water bodies from remote sensing images are insufficient in terms of generalization ability and accuracy, especially in complex scenes where it is difficult to accurately segment water bodies. Traditional methods require a large amount of manual intervention and are inefficient.

Method used

We employ a deep learning approach based on the UNetMFormer network, combining global and local feature extraction structures. By optimizing downsampling and multi-output upsampling structures, and utilizing the Multiscale_Transformer module to balance global and local information, we improve semantic segmentation accuracy.

Benefits of technology

It improves the accuracy and efficiency of water extraction, increasing IoU by 2.23% compared to traditional methods and shortening training time by 3.25 hours, effectively solving the problem of water extraction in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117152607B_ABST
    Figure CN117152607B_ABST
Patent Text Reader

Abstract

This invention discloses a method, system, and device for water body extraction from remote sensing images based on the UNetMFormer network. First, remote sensing images of the target watershed are acquired and water bodies in the images are labeled. Then, the labeled remote sensing images are matched using histograms. Finally, the matched remote sensing images are input into the UNetMFormer network to predict the water bodies in the remote sensing images. This invention improves the accuracy of deep learning applied to water body extraction from remote sensing images, meeting the model's accuracy requirements for water body extraction from high-resolution remote sensing images in both local and global aspects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of remote sensing technology and computer vision technology, and relates to a method, system and device for extracting water bodies from remote sensing images, specifically a method, system and device for extracting water bodies from remote sensing images based on the UNetMFormer network. Background Technology

[0002] With technological advancements, the increasing size of sensors, and the widespread application of drones, remote sensing technology is further developing and expanding (Reference 1). Many sensors, especially high-resolution optical satellites and drones, have become important platforms for Earth observation. NASA launched the first satellite in 1972 to measure and study the Earth's surface, marking the beginning of monitoring land cover from space (Reference 2). Subsequently, numerous satellites, such as Sentinel, SPOT, and Gaofen, were launched for their respective purposes. Based on the different imaging bands, remote sensing sensors can be divided into optical remote sensing sensors and synthetic aperture radar sensors. Optical remote sensing sensors can be further divided into coarse resolution, medium resolution, and high resolution based on their spatial resolution (Reference 2). Furthermore, drones, as a supplement to traditional remote sensing technologies, are becoming a new generation of sensors. Through these sensors, researchers obtain high-resolution optical remote sensing images, which are widely applied in specific remote sensing tasks (Reference 3).

[0003] In remote sensing images, rivers and lakes are important landmarks, playing a crucial role in water resource surveys, regional water resource management, flood monitoring, and water resource protection planning (Reference 4). Research on river and lake detection is receiving increasing attention. Accurate water body segmentation is a vital step in lake and river research. Traditional segmentation methods include threshold-based segmentation, edge-based segmentation, active region-based segmentation, and support vector machine-based segmentation (References 5-7). Zhu et al. used filtering and morphological methods combined with a region growing algorithm to detect changes in river regions (Reference 8). However, this algorithm is an iterative method with high time and space overhead and lacks versatility. Sun proposed a novel synthetic aperture radar (SAR) image river detection algorithm (Reference 9), which extracts edges in the wavelet domain and merges water bodies by ridge tracking. While this algorithm improves river edge detection to some extent, its parameter settings are highly susceptible to human intervention. McFeeters proposed the Normalized Differential Water Index (NDWI) method (Reference 10), which utilizes near-infrared and green light to enhance image features and then performs accurate segmentation. However, this method is highly susceptible to environmental influences. In summary, the above methods suffer from poor generalization ability, require significant human intervention, and suffer from inaccurate information.

[0004] The emergence of deep learning technology has brought new hope to water classification. When deep learning was first developed, it was mainly used for image-level classification. The main method uses continuous convolution and pooling to obtain feature information of the input image and finally obtain the probability of each category (Reference 11). However, these models have difficulty accurately segmenting details. To solve these problems, some scholars have extracted image pixel-level networks, which can obtain more detailed feature information. In 2014, Long et al. proposed the FCN segmentation network model, which can achieve image pixel-level classification (Reference 12), bringing a qualitative leap to image classification. However, this network is not accurate enough because it ignores the relationship between pixels, and it is not sensitive enough to detailed and global feature information. SegNet stores position indices to preserve image details through the pooling process (Reference 13), but the effect is low. The Unet semantic segmentation model (Reference 14) obtains rich feature information through upsampling and achieves accurate segmentation, but it is difficult to recover rich image features during the upsampling process, which easily leads to information redundancy. PSPNet obtains rich feature information by combining local and global cues through the context of different regions (Reference 15). In summary, compared to traditional methods, deep learning-based methods are better able to adapt to water bodies of different scales, shapes, and textures, while improving classification accuracy and efficiency (Reference 16). Current deep learning-based water classification methods have high practical value, but they also face many challenges, such as how to improve the model's generalization ability and how to handle large-scale, complex remote sensing images. To better address these issues, further research and development of deep learning-based water classification technology are needed, along with the development of more effective solutions tailored to specific scenarios and requirements.

[0005] References:

[0006] [Document 1] MCCABE MF, RODELL M, ALSDORF DE, et al. The future of Earthobservation in hydrology [J]. Hydrology and earth system sciences, 2017, 21(7): 3879-914.

[0007] [Reference 2] HUANG L, LIU L, JIANG L, et al. Automatic mapping of thermokarst landforms from remote sensing images using deep learning: A case study in the Northeastern Tibetan Plateau[J]. Remote Sensing, 2018, 10(12):2067.

[0008] [Reference 3] T, J, PáDUA L, et al. Hyperspectral imaging: A review on UAV-based sensors, data processing and applications for agriculture and forestry[J]. Remote sensing, 2017, 9(11):1110.

[0009] [Reference 4] VERMA U, CHAUHAN A, MM M P, et al. DeepRivWidth: Deep learning based semantic segmentation approach for river identification and width measurement in SAR images of Coastal Karnataka[J]. Computers&Geosciences, 2021, 154(104805.

[0010] [Reference 5] LI W, YANG M, LIANG Z, et al. Assessment for surface water quality in Lake Taihu Tiaoxi River Basin China based on support vector machine[J]. Stochastic environmental research and risk assessment, 2013, 27(1861-70.

[0011] [Reference 6] SUN X, LI L, ZHANG B, et al. Soft urban water cover extraction using mixed training samples and support vector machines[J]. International Journal of Remote Sensing, 2015, 36(13): 3331-44.

[0012] [Reference 7] LI C, XU C, GUI C, et al. Distance regularized level set evolution and its application to image segmentation[J]. IEEE transactions on image processing, 2010, 19(12): 3243-54.

[0013] [Reference 8] ZHU L, ZHANG J, PA L. River change detection based on remote sensing image and vector;proceedings of the First International Multi-Symposiums on Computer and Computational Sciences(IMSCCS'06), F, 2006[C]. IEEE.

[0014] [Reference 9] SUN J, MAO S. River detection algorithm in SAR images based on edge extraction and ridge tracing techniques[J]. International journal of remote sensing, 2011, 32(12): 3485-94.

[0015] [Literature 10] MCFEETERS S K. The use of the Normalized Difference Water Index(NDWI) in the delineation of open water features[J]. International journal of remote sensing, 1996, 17(7): 1425-32.

[0016] [Literature 11] HINTON G E, SALAKHUTDINOV R R. Reducing the dimensionality of data with neural networks[J]. science, 2006, 313(5786): 504-7.

[0017] [Literature 12] LONG J, SHELHAMER E, DARRELL T. Fully convolutional networks for semantic segmentation;proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, F, 2015[C].

[0018] [Literature 13] BADRINARAYANAN V, KENDALL A, CIPOLLAR. Segnet: A deep convolutional encoder-decoder architecture for image segmentation[J]. IEEE transactions on pattern analysis and machine intelligence, 2017, 39(12): 2481-95.

[0019] [Document 14] RONNEBERGER O, FISCHER P, BROX TU-net: Convolutional networks for biomedical image segmentation; procedures of the Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015:18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18,F,2015[C].Springer.

[0020] [Document 15] ZHAO H, SHI J, QI X, et al. Pyramid scene parsing network; proceedings of the Proceedings of the IEEE conference on computer vision and pattern recognition, F, 2017[C].

[0021] [Document 16] GUO H, HE G, JIANG W, et al. A multi-scale water extraction convolutional neural network (MWEN) method for GaoFen-1remote sensing images [J]. ISPRS International Journal of Geo-Information, 2020, 9(4): 189. Summary of the Invention

[0022] To address the aforementioned technical problems, this invention provides a method, system, and device for water body extraction from deep learning remote sensing images based on the UNetMFormer network, aiming to improve the accuracy and efficiency of semantic segmentation.

[0023] The technical solution adopted by the method of the present invention is: a method for extracting water bodies from remote sensing images based on UNetMFormer network, comprising the following steps:

[0024] Step 1: Acquire remote sensing images of the target watershed and label the water bodies in the images;

[0025] Step 2: Match the labeled remote sensing images using histograms;

[0026] Step 3: Input the matched remote sensing image into the UNetMFormer network to predict the water bodies in the remote sensing image.

[0027] Preferably, in step 3, the UNetMFormer network includes an optimized downsampling structure, a global and local feature extraction structure, and an optimized upsampling structure that combines multiple outputs.

[0028] The optimized downsampling structure includes a five-level downsampling module. Each downsampling module downsamples the feature map with a scaling factor of 2. Each downsampling module is followed by a normalization operation layer and an activation layer.

[0029] The global and local feature extraction structure includes a first global and local feature extraction module, a first weighted sum operation module, a second global and local feature extraction module, a second weighted sum operation module, a third global and local feature extraction module, a third weighted sum operation module, and a convolutional layer. The output of the fifth-level downsampling module is input to the first global and local feature extraction module. The outputs of the first global and local feature extraction module and the fourth-level downsampling module are processed by the first weighted sum operation module and then input to the second global and local feature extraction module. The outputs of the second global and local feature extraction module and the third-level downsampling module are processed by the second weighted sum operation module and then input to the third global and local feature extraction module. The outputs of the third global and local feature extraction module and the second-level downsampling module are processed by the third weighted sum operation module and then input to the convolutional layer.

[0030] The optimized upsampling structure combining multiple outputs includes a first upsampling module, a second upsampling module, a third upsampling module, a fourth upsampling module, a first fusion module, and a second fusion module. The outputs of the first global and local feature extraction module, the second global and local feature extraction module, and the third global and local feature extraction module are respectively processed by the first upsampling module, the second upsampling module, and the third upsampling module, and then fused by the first fusion module to output the final entity extraction result. The outputs of the convolutional layers are respectively processed by the fourth upsampling module, and then fused with the output of the first fusion module by the second fusion module to output the final entity extraction result.

[0031] Preferably, in step 3, the semantic features generated by each downsampling module are aggregated with the features generated by the global and local feature extraction modules of the decoder using a weighted sum operation. The weighted sum operation selectively weights the features based on their contribution to the segmentation accuracy; the weighting calculation formula is as follows:

[0032] F = φ.R + (1-φ).MT;

[0033] Where F represents the fused feature, R represents the feature generated by each downsampling module, MT represents the feature generated by the global and local feature extraction modules of the decoder, and φ represents the weight ratio.

[0034] Preferably, the first global and local feature extraction module is a global-local multi-scale context structure, including a global branch, a local branch, a fusion layer, a convolutional layer, a batch normalization operation layer, and a convolutional layer. The outputs of the global branch and the local branch are fused by the fusion layer and then sequentially passed through the convolutional layer, the batch normalization operation layer, and the convolutional layer before being output.

[0035] The global branch includes a window segmentation layer, a first dot product layer, a normalized exponential function layer, a second dot product layer, and a cross-shaped window context interaction module, outputting global context features. The window segmentation layer consists of a sequentially connected convolution module, a window segmentation module, a window deformation module, and an allocation module, used to segment the output into a query Q, a key K, and a value V carrier. The query Q and key K, after passing through the first dot product layer and the normalized exponential function layer, and the value V carrier, after passing through the second dot product layer, are input into the cross-shaped window context interaction module. The cross-shaped window context interaction module consists of parallel horizontal average pooling layers and vertical average pooling layers, and a fusion layer. The fusion layer merges the two feature maps generated by the horizontal average pooling layer and the vertical average pooling layer, outputting global features.

[0036] The local branch includes five convolutional layers arranged in parallel. Each convolutional layer is followed by a batch normalization function layer. The output of the batch normalization function layer is fused by a fusion layer to output local features.

[0037] Preferably, the window segmentation layer first uses a standard 1*1 convolution to segment the input 2D image ∈ R. B×C×H×W The channel dimension is increased to 3 times; then window partitioning is applied to the 1D sequence. It is split into query Q, key K, and value V carriers; where B represents the batch size, H and W represent the length and width of the feature map, respectively, C is the channel dimension, w is the window size, and h is the number of heads.

[0038] Preferably, the UNetMFormer network is a pre-trained network; the training process includes the following steps:

[0039] (1) Acquire several high-resolution remote sensing images of target watersheds as a sample set; and label the water bodies in the high-resolution remote sensing images as a tag set;

[0040] (2) Obtain the preprocessed remote sensing image set by matching the labeled remote sensing images with histograms;

[0041] (3) Use the preprocessed remote sensing image set and label set to train the UNetMFormer network and obtain the trained UNetMFormer network;

[0042] During training, Dice loss is used as the loss function, Adam is used as the optimizer, and accuracy is used as the evaluation metric.

[0043] The expression for the loss function, Dice loss, is as follows:

[0044]

[0045] Where N is the total number of pixels, g i To determine whether the i-th pixel in the standard reference result belongs to a body of water, if it does, then g... i =1, otherwise g i =0, p i To predict the probability that the i-th pixel in the image is a building.

[0046] The technical solution adopted by the system of the present invention is: a remote sensing image water body extraction system based on UNetMFormer network, comprising the following modules:

[0047] The remote sensing image annotation acquisition and annotation module is used to acquire remote sensing images of the target watershed and annotate the water bodies in the remote sensing images;

[0048] The remote sensing image preprocessing module is used to match labeled remote sensing images using histograms;

[0049] The remote sensing image water body prediction module is used to input the matched remote sensing images into the UNetMFormer network to predict the water bodies in the remote sensing images.

[0050] The technical solution adopted by the device of the present invention is: a remote sensing image water body extraction device based on UNetMFormer network, comprising:

[0051] One or more processors;

[0052] A storage device for storing one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the remote sensing image water body extraction method based on the UNetMFormer network.

[0053] The advantages of this invention are:

[0054] 1. In the upsampling process, the present invention incorporates the Multiscale_Transformer (Mtransformer) module, which takes into account both global and local information, enriches multi-scale spatial details, enhances semantic segmentation, and improves segmentation accuracy.

[0055] 2. The water body optimization effect of this invention is more targeted. Due to the varying complexity of high-resolution image backgrounds, some water bodies are located in complex scenes, with significant differences in size and shape. Previous methods can extract most water bodies, but some small tributaries are missed, and edge extraction is not effective. This invention incorporates a global and local feature extraction structure, which can effectively take into account both local and global feature extraction, effectively improving and alleviating the problems of poor edge extraction for small rivers, and improving extraction accuracy.

[0056] 3. While traditional CNN deep learning models can extract both global and local features, their extraction efficiency and accuracy are relatively low. This invention improves IoU by 2.23% compared to the control method, enhancing water body optimization efficiency and reducing training time by 3.25 hours. Attached Figure Description

[0057] The technical solutions described herein are further illustrated below using examples and specific implementation methods. Additionally, accompanying drawings are used in the description of the technical solutions. Those skilled in the art can, without any creative effort, obtain other drawings and the intent of the present invention based on these drawings.

[0058] Figure 1 This is a flowchart of a method according to an embodiment of the present invention.

[0059] Figure 2 This is a diagram of the overall network architecture of UNetMFormer according to an embodiment of the present invention.

[0060] Figure 3 This is a schematic diagram of the global and local feature extraction structure according to an embodiment of the present invention.

[0061] Figure 4 This is the result of water body extraction from the Minjiang dataset. Detailed Implementation

[0062] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0063] Please see Figure 1 This embodiment provides a method for extracting water bodies from remote sensing images based on the UNetMFormer network, which includes the following steps:

[0064] Step 1: Acquire remote sensing images of the target watershed and label the water bodies in the images;

[0065] Step 2: Match the labeled remote sensing images using histograms;

[0066] Step 3: Input the matched remote sensing image into the UNetMFormer network to predict the water bodies in the remote sensing image.

[0067] Please see Figure 2 The UNetMFormer network includes an optimized downsampling structure, a global and local feature extraction structure, and an optimized upsampling structure that combines multiple outputs.

[0068] The optimized downsampling structure includes a five-level downsampling module. Each downsampling module downsamples the feature map with a scaling factor of 2. Each downsampling module is followed by a normalization operation layer and an activation layer.

[0069] The global and local feature extraction structure includes a first global and local feature extraction module, a first weighted sum operation module, a second global and local feature extraction module, a second weighted sum operation module, a third global and local feature extraction module, a third weighted sum operation module, and a convolutional layer. The output of the fifth-level downsampling module is input to the first global and local feature extraction module. The outputs of the first global and local feature extraction module and the fourth-level downsampling module are processed by the first weighted sum operation module and then input to the second global and local feature extraction module. The outputs of the second global and local feature extraction module and the third-level downsampling module are processed by the second weighted sum operation module and then input to the third global and local feature extraction module. The outputs of the third global and local feature extraction module and the second-level downsampling module are processed by the third weighted sum operation module and then input to the convolutional layer.

[0070] The optimized upsampling structure combining multiple outputs includes a first upsampling module, a second upsampling module, a third upsampling module, a fourth upsampling module, a first fusion module, and a second fusion module. The outputs of the first global and local feature extraction module, the second global and local feature extraction module, and the third global and local feature extraction module are respectively processed by the first upsampling module, the second upsampling module, and the third upsampling module, and then fused by the first fusion module to output the final entity extraction result. The outputs of the convolutional layers are respectively processed by the fourth upsampling module, and then fused with the output of the first fusion module by the second fusion module to output the final entity extraction result.

[0071] In one implementation, the semantic features generated by each downsampling module are aggregated with the features generated by the global and local feature extraction modules of the decoder using a weighted sum operation. The weighted sum operation selectively weights the features based on their contribution to the segmentation accuracy; the weighting calculation formula is as follows:

[0072] F = φ.R + (1-φ).MT;

[0073] Where F represents the fused feature, R represents the feature generated by each downsampling module, MT represents the feature generated by the global and local feature extraction modules of the decoder, and φ represents the weight ratio.

[0074] Please see Figure 3 In one embodiment, the first global and local feature extraction module is a global-local multi-scale context structure, including a global branch, a local branch, a fusion layer, a convolutional layer, a batch normalization operation layer, and a convolutional layer. The outputs of the global branch and the local branch are fused by the fusion layer and then sequentially passed through the convolutional layer, the batch normalization operation layer, and the convolutional layer before being output.

[0075] The global branch includes a window segmentation layer, a first dot product layer, a normalized exponential function layer, a second dot product layer, and a cross-shaped window context interaction module, outputting global context features. The window segmentation layer consists of a sequentially connected convolution module, a window segmentation module, a window deformation module, and an allocation module, used to segment the output into a query Q, a key K, and a value V carrier. The query Q and key K, after passing through the first dot product layer and the normalized exponential function layer, and the value V carrier, after passing through the second dot product layer, are input into the cross-shaped window context interaction module. The cross-shaped window context interaction module consists of parallel horizontal average pooling layers and vertical average pooling layers, and a fusion layer. The fusion layer merges the two feature maps generated by the horizontal average pooling layer and the vertical average pooling layer, outputting global features.

[0076] The local branch includes convolutional layers with convolutional kernels of 1, 3, 5, 7, and 9 arranged in parallel. Each convolutional layer is followed by a batch normalization function layer. The output of the batch normalization function layer is fused by a fusion layer to output local features.

[0077] The window segmentation layer first uses a standard 1*1 convolution to segment the input 2D image ∈ R. B×C×H×W The channel dimension is increased to 3 times; then window partitioning is applied to the 1D sequence. The algorithm is split into query Q, key K, and value V carriers; where B represents the batch size, H and W represent the length and width of the feature map, respectively, C is the channel dimension, w is the window size, and h is the number of heads. Finally, a cross-shaped window context interaction module is used to fuse the two feature maps generated by the horizontal average pooling layer and the vertical average pooling layer, thereby capturing the global context. In this embodiment, the channel dimension C is 64, the window size w, and the number of heads h are both set to 8.

[0078] The cross-shaped window context interaction module described above fuses two feature maps generated by the horizontal average pooling layer and the vertical average pooling layer to capture the global context. Specifically, the horizontal average pooling layer establishes the horizontal relationship between windows, such as Win1 = H (Win2). For any point P1 in window 1... (m,n) Its relationship with the window The correspondence is as follows:

[0079]

[0080] P1 (m+i,n) =D i (P1 (m,n) );

[0081]

[0082]

[0083] Where w represents the size of the window, D represents self-attention computation, which models the dependency of pixel pairs in a local window; m, n, i, and j represent points within it.

[0084] In one implementation, the optimized upsampling structure combining multiple outputs consists of three output layers, D1, D2, and D3, which enhance the network's refinement capability by acquiring low-level feature information. The feature map sizes of the three output layers are 16×16, 32×32, and 64×64, respectively. Then, an upsampling operation transforms the feature map sizes of the three output layers to the same 64×64. Subsequently, the features of the three output layers are fused and upsampled by a factor of 8 to restore the original input size, resulting in an F1 score.

[0085] In one implementation, the UNetMFormer network is a pre-trained network; the training process includes the following steps:

[0086] (1) Acquire several high-resolution remote sensing images of target watersheds as a sample set; and label the water bodies in the high-resolution remote sensing images as a tag set;

[0087] In one implementation, remote sensing images of the Minjiang River basin are acquired, and water bodies in the high-resolution remote sensing images are labeled using ArcGIS software to form a sample set.

[0088] In one implementation, water body annotation results are obtained based on manual annotation, and the pixel values ​​of water bodies and non-water bodies in the image are different.

[0089] (2) Obtain the preprocessed remote sensing image set by matching the labeled remote sensing images with histograms;

[0090] In one implementation, ENVI is used to perform histogram matching on the image, and image histogram transformation is performed with reference to the histogram of a specified image to enhance the display of contrasting images.

[0091] (3) Use the preprocessed remote sensing image set and label set to train the UNetMFormer network and obtain the trained UNetMFormer network;

[0092] In this embodiment, the 31 optimized remote sensing images of the Minjiang River Basin are first divided into 25 images as the training and validation set and 6 images as the test set, according to a 6:2:2 ratio. Due to insufficient computer video memory, the original large images after histogram matching are cropped to a size of 512*512 to obtain the sample set.

[0093] Then, the training samples are input into the UNetMFormer network, whose network architecture and training parameters have been set, for training.

[0094] Regarding the parameter settings for network training, since the amount of water and non-water bodies in remote sensing images is uneven, and Dice loss is beneficial for solving the data balance problem, the UNetMFormer network uses Dice loss as the loss function during training. The adam optimizer has the advantage of fast convergence speed, so adam is used as the optimizer. Accuracy is used as the evaluation metric, batch size = 2, and iterations = 50.

[0095] Dice loss is a widely used loss function in image segmentation, used to measure image classification accuracy. Its expression is as follows:

[0096]

[0097] N is the total number of pixels, g i To determine whether the i-th pixel in the standard reference result belongs to a body of water, if it does, then g... i =1, otherwise g i =0, p iTo predict the probability that the i-th pixel in the image is a building.

[0098] This embodiment also provides a remote sensing image water body extraction system based on the UNetMFormer network, including the following modules:

[0099] The remote sensing image annotation acquisition and annotation module is used to acquire remote sensing images of the target watershed and annotate the water bodies in the remote sensing images;

[0100] The remote sensing image preprocessing module is used to match labeled remote sensing images using histograms;

[0101] The remote sensing image water body prediction module is used to input the matched remote sensing images into the UNetMFormer network to predict the water bodies in the remote sensing images.

[0102] This embodiment also provides a remote sensing image water body extraction device based on the UNetMFormer network, including:

[0103] One or more processors;

[0104] A storage device for storing one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the remote sensing image water body extraction method based on the UNetMFormer network.

[0105] Please see Figure 4 This image shows the water body extraction results based on the UNetMFormer network on the Minjiang River dataset. The first row shows the overall comparison image, and the second to fifth rows show the detail comparison images. Figure 4 It can be seen that the water extraction effect of the method of the present invention is the best. For example... Figure 4 As shown in the second and third rows, traditional methods are prone to producing voids when extracting water from large lakes. For example... Figure 4 As shown in the fourth row, the Unet model is prone to incomplete extraction when extracting small rivers. As shown in the fifth row, spectral information can be disrupted by sediment at river edges, leading to missed detections. To address these issues, this invention utilizes global and local feature extraction modules to enhance the extraction of global and local features of water bodies. Furthermore, for rivers and lakes of different scales, the global and local feature extraction modules fuse multiple scales, strengthening the extraction of water bodies at different scales and improving extraction accuracy.

[0106] As shown in Table 1, the present invention uses the UNetMFormer network, which improves the IoU by 2.23% compared with the control method, thus improving the water body optimization efficiency and shortening the training time by 3.25 hours.

[0107]

[0108] This invention considers the water morphology and spectral characteristics of remote sensing images and proposes a method and system for water body extraction from remote sensing images based on a globally fused local feature network. This addresses the problems of high difficulty, low accuracy, and low efficiency in water body extraction in the field of remote sensing technology. This invention proposes a method for water body extraction from remote sensing images based on a globally fused local feature network. This method fuses traditional CNNs and transformers, obtaining local features while also considering a better global feature design. This network has good water body extraction capabilities, achieving high accuracy, good results, and high efficiency in water body extraction from remote sensing images.

[0109] It should be understood that the above description of the preferred embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art, under the guidance of this invention, can make substitutions or modifications without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.

Claims

1. A method for extracting water bodies from remote sensing images based on UNetMFormer network, characterized in that, Includes the following steps: Step 1: Acquire remote sensing images of the target watershed and label the water bodies in the images; Step 2: Match the labeled remote sensing images using histograms; Step 3: Input the matched remote sensing image into the UNetMFormer network to predict the water bodies in the remote sensing image; The UNetMFormer network includes an optimized downsampling structure, a global and local feature extraction structure, and an optimized upsampling structure that combines multiple outputs. The optimized downsampling structure includes a five-level downsampling module. Each downsampling module downsamples the feature map with a scaling factor of 2. Each downsampling module is followed by a normalization operation layer and an activation layer. The global and local feature extraction structure includes a first global and local feature extraction module, a first weighted sum operation module, a second global and local feature extraction module, a second weighted sum operation module, a third global and local feature extraction module, a third weighted sum operation module, and a convolutional layer. The output of the fifth-level downsampling module is input to the first global and local feature extraction module. The outputs of the first global and local feature extraction module and the fourth-level downsampling module are processed by the first weighted sum operation module and then input to the second global and local feature extraction module. The outputs of the second global and local feature extraction module and the third-level downsampling module are processed by the second weighted sum operation module and then input to the third global and local feature extraction module. The outputs of the third global and local feature extraction module and the second-level downsampling module are processed by the third weighted sum operation module and then input to the convolutional layer. The optimized upsampling structure combining multiple outputs includes a first upsampling module, a second upsampling module, a third upsampling module, a fourth upsampling module, a first fusion module, and a second fusion module. The outputs of the first global and local feature extraction module, the second global and local feature extraction module, and the third global and local feature extraction module are respectively processed by the first upsampling module, the second upsampling module, and the third upsampling module, and then fused by the first fusion module to output the final entity extraction result. The outputs of the convolutional layers are respectively processed by the fourth upsampling module, and then fused with the output of the first fusion module by the second fusion module to output the final entity extraction result.

2. The method for extracting water bodies from remote sensing images based on UNetMFormer network according to claim 1, characterized in that: The semantic features generated by each downsampling module are aggregated with the features generated by the global and local feature extraction modules of the decoder using a weighted sum operation. The weighted sum operation selectively weights the features based on their contribution to the segmentation accuracy. The weighted calculation formula is: ; Where F represents the fused feature, R represents the feature generated by each downsampling module, and MT represents the feature generated by the global and local feature extraction modules of the decoder. This represents the weighting ratio.

3. The method for extracting water bodies from remote sensing images based on UNetMFormer network according to claim 1, characterized in that: The first global and local feature extraction module is a global-local multi-scale context structure, including a global branch, a local branch, a fusion layer, a convolutional layer, a batch normalization operation layer, and a convolutional layer. The outputs of the global branch and the local branch are fused by the fusion layer and then sequentially passed through the convolutional layer, the batch normalization operation layer, and the convolutional layer before being output. The global branch includes a window segmentation layer, a first dot product layer, a normalized exponential function layer, a second dot product layer, and a cross-shaped window context interaction module, which outputs global context features. The window segmentation layer consists of a sequentially connected convolution module, window segmentation module, window deformation module, and allocation module, used to segment the output into a query Q, a key K, and a value V carrier. The query Q and key K are processed through the first dot product layer and the normalized exponential function layer, and then, together with the value V carrier, are processed through the second dot product layer before being input into the cross-shaped window context interaction module. The cross-shaped window context interaction module consists of parallel horizontal average pooling layers and vertical average pooling layers, and a fusion layer. The fusion layer merges the two feature maps generated by the horizontal average pooling layer and the vertical average pooling layer to output global features. The local branch includes five convolutional layers arranged in parallel. Each convolutional layer is followed by a batch normalization function layer. The output of the batch normalization function layer is fused by a fusion layer to output local features.

4. The method for extracting water bodies from remote sensing images based on UNetMFormer network according to claim 3, characterized in that: The window segmentation layer first uses convolution to process the input 2D image. The channel dimension is expanded to a predetermined multiple; then window partitioning is applied to the 1D sequence. It is split into query Q, key K, and value V carriers; where B represents the batch size, H and W represent the length and width of the feature map, respectively, and C is the channel dimension. For window size, h The number of heads.

5. The method for extracting water bodies from remote sensing images based on UNetMFormer network according to any one of claims 1-4, characterized in that: The UNetMFormer network is a pre-trained network; The training process includes the following steps: (1) Acquire several high-resolution remote sensing images of the target watershed as a sample set; and label the water bodies in the high-resolution remote sensing images as a tag set; (2) Obtain the preprocessed remote sensing image set by matching the labeled remote sensing images with histograms; (3) Use the preprocessed remote sensing image set and label set to train the UNetMFormer network and obtain the trained UNetMFormer network; During training, Dice loss is used as the loss function, Adam is used as the optimizer, and accuracy is used as the evaluation metric. The expression for the loss function, Dice loss, is as follows: ; in, This represents the total number of pixels. The first in the standard reference results Does each pixel belong to a body of water? If it does, then... ,otherwise , To predict the first in the graph The probability that a pixel represents a building.

6. A water body extraction system based on UNetMFormer network in remote sensing images, characterized in that, Includes the following modules: The remote sensing image annotation acquisition and annotation module is used to acquire remote sensing images of the target watershed and annotate the water bodies in the remote sensing images; The remote sensing image preprocessing module is used to match labeled remote sensing images using histograms; The remote sensing image water body prediction module is used to input the matched remote sensing images into the UNetMFormer network to predict the water bodies in the remote sensing images. The UNetMFormer network includes an optimized downsampling structure, a global and local feature extraction structure, and an optimized upsampling structure that combines multiple outputs. The optimized downsampling structure includes a five-level downsampling module. Each downsampling module downsamples the feature map with a scaling factor of 2. Each downsampling module is followed by a normalization operation layer and an activation layer. The global and local feature extraction structure includes a first global and local feature extraction module, a first weighted sum operation module, a second global and local feature extraction module, a second weighted sum operation module, a third global and local feature extraction module, a third weighted sum operation module, and a convolutional layer. The output of the fifth-level downsampling module is input to the first global and local feature extraction module. The outputs of the first global and local feature extraction module and the fourth-level downsampling module are processed by the first weighted sum operation module and then input to the second global and local feature extraction module. The outputs of the second global and local feature extraction module and the third-level downsampling module are processed by the second weighted sum operation module and then input to the third global and local feature extraction module. The outputs of the third global and local feature extraction module and the second-level downsampling module are processed by the third weighted sum operation module and then input to the convolutional layer. The optimized upsampling structure combining multiple outputs includes a first upsampling module, a second upsampling module, a third upsampling module, a fourth upsampling module, a first fusion module, and a second fusion module. The outputs of the first global and local feature extraction module, the second global and local feature extraction module, and the third global and local feature extraction module are respectively processed by the first upsampling module, the second upsampling module, and the third upsampling module, and then fused by the first fusion module to output the final entity extraction result. The outputs of the convolutional layers are respectively processed by the fourth upsampling module, and then fused with the output of the first fusion module by the second fusion module to output the final entity extraction result.

7. A water body extraction device based on UNetMFormer network in remote sensing images, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the remote sensing image water body extraction method based on the UNetMFormer network as described in any one of claims 1 to 5.