Vision Transform-based high-resolution remote sensing image water body extraction method and device
By introducing a space detail information retention module and a multi-scale water body feature extraction module into the Vision Transformer network, combining the improved multi-head self-attention mechanism and global multi-layer perceptron, a multi-scale fusion ViT model (MSMViT) is designed, which solves the problem of insufficient spatial detail feature extraction capabilities in water body extraction in the existing ViT network, achieving more efficient water body extraction effect and lower computational complexity.
Patent Information
- Application Number
- CN202510149295.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The existing Vision Transformer (ViT) network has the problem of strong global context information capture capabilities but insufficient spatial detail feature extraction capabilities in water extraction, and the computational complexity is high, resulting in huge memory requirements and affecting application potential.
A multi-scale fusion ViT model (MSMViT) is designed to dynamically balance local spatial detail information and global context information by introducing a spatial detail information retention module and a multi-scale water feature extraction module into the encoder, combining the improved multi-head self-attention mechanism and a convolution-based global multi-layer perceptron.
The accuracy of water extraction of high-score remote sensing images is improved, taking into account the extraction of global features and detailed features, reducing the computational complexity and memory requirements, and enhancing the application potential of the model.
Smart Images

Figure CN120107812A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for extracting water from high-resolution remote sensing images. Aiming at the shortcomings of the existing Vision Transformer (ViT) network in water extraction, a method and device for extracting water from high-resolution remote sensing images based on a multi-scale fusion ViT model (MSMViT) are proposed. Background Art
[0002] Water is an important resource that plays a key role in human production and life. Therefore, understanding the distribution of water resources is of great guiding significance for the rational use and protection of water resources. Remote sensing technology is macroscopic, timely and authentic, and can obtain spectral image information on a large range of the earth's surface. In particular, the data obtained by various remote sensing methods represented by satellite remote sensing have greatly improved spatial resolution and spectral resolution. Using satellite remote sensing data to extract surface water information has become one of the important contents of remote sensing applications.
[0003] Water body extraction from high-resolution remote sensing images mainly includes single-band threshold method, multi-band spectral relationship method and water body index method. With the development of deep learning methods, great progress has been made in water body information extraction from remote sensing images. Classification networks such as VGG and ResNet and segmentation networks such as FCN have achieved good results in water body information extraction. Vision Transformer can segment remote sensing images into multiple patches. By modeling the relationship between patches through the Transformer encoder, ViT can capture the global context information in remote sensing images, thereby improving the recognition ability of large-area water bodies. However, ViT requires a large amount of data for pre-training and has high computational complexity, resulting in huge memory requirements for ViT, which seriously affects its application potential. The improved Swin Transformer adopts a hierarchical structure and designs a multi-head self-attention mechanism to improve efficiency, but its complexity will still increase quadratically with the increase of windows. ViT mainly focuses on capturing global context information and ignores spatial detail features. Therefore, how to maintain ViT's global information extraction ability and improve its spatial detail feature extraction ability without increasing the computational complexity of the model is very important for water body extraction from high-resolution remote sensing images. Summary of the invention
[0004] Based on the shortcomings of the prior art, the present invention designs a method and device for water body extraction from high-resolution remote sensing images based on Vision Transformer, which can take into account both the global features and detail features of high-resolution remote sensing images and achieve better water body extraction effect.
[0005] The method for extracting water from high-resolution remote sensing images based on Vision Transformer designed by the present invention comprises the following steps:
[0006] S1, obtain high-resolution remote sensing images and perform preprocessing;
[0007] S2, visually interpret the water bodies of the pre-processed high-resolution remote sensing images and delineate the water body boundaries in the images to produce a dataset for deep learning model training;
[0008] S3, constructing an MSMViT network model, which includes an encoder and a decoder; wherein the encoder includes a detail information retention module and a water body feature extraction module; there are multiple multi-scale water body feature extraction modules, and the jump connection part thereof includes a dynamic information fusion module;
[0009] The detail information retaining module first obtains an image feature map, and then processes the image feature map to obtain spatial detail information;
[0010] The multi-scale water feature extraction modules all include a mixW-Block module, which is composed of LMSMHSA, a convolution-based global multi-layer perceptron CGMLP, a normalization layer, and a residual connection. The LMSMHSA is improved from the multi-head self-attention mechanism MHSA of the ordinary ViT. The specific improvements include using convolution to replace the full-link layer in the ordinary ViT, and using the Taylor formula to replace the Softmax function in the ordinary ViT.
[0011] S4, using the data set in step S2 to train the constructed MSMViT network model, and using the trained MSMViT network model for water body extraction.
[0012] Preferably, the preprocessing in S1 includes performing radiation correction, geometric correction, fusion, mosaicking and cropping on the high-resolution remote sensing image.
[0013] Preferably, in S2, the water bodies on the multispectral image are visually interpreted, and the boundaries of the water bodies in the image are outlined using Labelme software and converted into true value label data of the water bodies.
[0014] Preferably, in S2, after reading the high-resolution remote sensing image data and the corresponding water body label data, the sliding window method is used to cut them into sample data of the same size, and the entire data set is divided into a training set, a validation set and a test set according to the proportion.
[0015] Preferably, the spatial detail information retention module includes two parts A and B, wherein part A provides a feature map for the main structure, which is mainly composed of two downsampling modules and a fusion module, and part B is a spatial detail information extraction module.
[0016] Furthermore, the spatial detail information acquisition module is implemented by multi-layer modules. The processing process of each layer of modules includes convolution, batch normalization and activation function. Each layer obtains detailed spatial information by stepwise downsampling.
[0017] Preferably, the encoder also includes a patch merging module for adjusting the size of the feature map. The specific process includes: using convolution and pooling operations to downsample the feature map to half of the original feature map, and expanding the dimension of the feature map to twice the original, and then splicing the feature maps obtained by the two operations.
[0018] Preferably, there are four multi-scale water body feature extraction modules.
[0019] Preferably, the decoder uses a basic self-attention mechanism to restore the resolution of the feature map by upsampling layer by layer.
[0020] Based on the same inventive concept, the present invention also provides an electronic device, including:
[0021] one or more processors;
[0022] A storage device for storing one or more programs;
[0023] When one or more programs are executed by the one or more processors, the one or more processors implement a high-resolution remote sensing image water body extraction method based on Vision Transformer.
[0024] Based on the same inventive concept, the present invention also designs a computer-readable medium on which a computer program is stored. When the program is executed by a processor, a method for extracting water from high-resolution remote sensing images based on Vision Transformer is implemented.
[0025] The special features of the MSMViT network model proposed in the present invention include the following points:
[0026] 1) In the MSMViT network model, a spatial detail information preservation module is designed. This module not only provides feature maps for the main structure, but also provides more detailed spatial detail information for the network model, enhancing the model's ability to extract detail features.
[0027] 2) In the MSMViT network model, a multi-scale module mixW-Block is designed to extract water feature information in the image. The improvement of this module lies in the linear multi-scale multi-head self-attention module LMSMHSA and the convolution-based global multi-layer perceptron CGMLP. LMSMHSA can obtain richer spatial information, avoid the lack of feature diversity caused by the single feature head of the self-attention mechanism, and optimize the calculation method of the multi-head attention mechanism, effectively reducing the amount of self-attention calculation, which is more suitable for large-scale high-resolution remote sensing image processing; CGMLP strengthens cross-window interaction and further enhances the model's ability to capture global context information.
[0028] 3) The MSMViT network model can retain spatial detail information while taking into account the spatial position relationship of water bodies in remote sensing images, dynamically balancing the local spatial detail information and global context information during the segmentation process, and improving the accuracy of water body segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is the overall structure diagram of the MSMViT network in the present invention.
[0030] Figure 2 This is a structural diagram of the spatial detail information retention module in the present invention.
[0031] Figure 3 This is a structural diagram of the mixW-Block module in the present invention.
[0032] Figure 4 This is a structural diagram of the global multi-layer perceptron in the present invention.
[0033] Figure 5 The following table shows the image data, corresponding labels and water extraction results of each method. DETAILED DESCRIPTION
[0034] In order to make the technical means, creative features, workflow, usage method, purpose and effect of the present invention easy to understand, the present invention is further explained below with reference to the accompanying drawings.
[0035] Embodiment 1
[0036] A method for extracting water from high-resolution remote sensing images based on Vision Transformer includes the following steps:
[0037] S1, obtain high-resolution remote sensing images of the study area, and perform pre-processing such as radiation correction, geometric correction, image fusion, mosaicking, and cropping on the high-resolution remote sensing images;
[0038] S2, perform visual interpretation of water bodies on the preprocessed high-resolution remote sensing images, use Labelme software to outline the water body mask, and convert the mask data into water body true value label data. After aligning the image data and the corresponding water body label data, use the sliding window method to crop the image and label data into a size of 512*512, and obtain many paired image and label data blocks. The entire data set is divided into training set, validation set and test set in a ratio of 7:2:1. The training set is used to train the model, that is, the model learns the characteristics and rules in the data through the data of the training set; the validation set is used to adjust the model hyperparameters and evaluate the model performance to avoid model overfitting; the test set is used to finally evaluate the performance of the model and simulate the performance of the model in real scenes.
[0039] S3, construct the MSMViT network model, its overall structure is as follows Figure 1 As shown in Figure 1, the main body still follows the network structure of the encoder and decoder. The encoder part consists of a spatial detail information retention module mix-stem and four multi-scale water feature extraction modules w-mix stage 1、2、3、4 Composition; the middle jump connection part contains three dynamic information fusion modules WFstage 1、2、3 ; The decoder part consists of 4 water feature recovery modules w-mix stage 5、6、7、8 And a fusion module merge; finally, the water body extraction result is output through the head;
[0040] The encoder is used to extract image features and capture context information. In order to avoid losing spatial detail information during network downsampling, the present invention designs a spatial detail information preservation module mix-stem at the front end of the encoder. The specific structure is as follows Figure 2 As shown in the figure, the spatial detail information retention module consists of two parts, A and B. A mainly provides feature maps for the main structure through downsampling. It consists of two downsampling modules and a fusion module, which is composed of a series combination of a convolution kernel of 4*4 with a step length of 4, a convolution kernel of 1*1 with a step length of 1, and an activation layer; B mainly provides more detailed spatial detail information for the network model. This module consists of 6 parts, each of which consists of convolution, batch normalization and activation function. Each layer obtains detailed spatial information through step-by-step downsampling. There are 4 multi-scale water feature extraction modules w-mix stage after the spatial detail information retention module. The encoder specifically includes:
[0041] 1) Patch merging, its main function is to adjust the size of the feature map during downsampling to complete the downsampling task. The entire encoder includes three patch merging layers. 2、3、4To complete the three downsampling tasks. First, the feature map is downsampled to half of the original feature map by using convolution and pooling operations respectively, and the dimension of the feature map is expanded to twice the original. Then the feature maps obtained by the two operations are spliced. This design is conducive to enhancing the foreground information. Finally, the standard deep convolution and residual connection operations are used to strengthen the fused adjustment information.
[0042] 2) The mixW-Block module is the core component of the encoder for extracting features. It obtains high-level semantic information at each level and performs the w-mix stage. 1、2、3、4 The mixW-Block module is included in the mixW-Block. The mixW-Block contains LMSMHSA, convolution-based global multilayer perceptron CGMLP, normalization layer and residual connection settings. The specific structure is as follows Figure 3 As shown in the figure, after the spatial detail information is input into mixW-Block, LMSMHSA first calculates the weight information of the feature information, improves the network's information perception ability of the global water body characteristics, and outputs a detailed water body feature map. Compared with the multi-head self-attention mechanism MHSA of ordinary ViT, LMMHSA has certain advantages in terms of calculation method and amount of calculation. Ordinary ViT directly flattens the input feature map into Q (Query), K (Key), and V (Value), calculates the attention score through QKV, and inputs it after weighted summation. LMMHSA uses convolution to replace the fully connected layer in ordinary ViT. Convolution is conducive to preserving the spatial relationship between objects in the image. At the same time, convolution layers of different sizes are constructed to obtain feature information of different sizes, avoiding the problem of information loss caused by a single scale, enhancing the diversity of information, and being more suitable for extracting cross-scale water body features. In addition, the Taylor formula is used to replace the Softmax function in ordinary ViT, reducing the computational complexity of LMSMHSA from N 2 It is reduced to N. It can be seen that the efficiency and effect of feature extraction can be improved by improving the QKV generation module. CGMLP uses multi-scale convolution operations to enhance local information fusion. The specific structure is as follows Figure 4 As shown in the figure. First, the 1*1 convolution combination module and the 3*3 convolution combination module are connected in parallel and the dimension of the feature map is increased by 2 times respectively. After superposition, the dimension of the feature map becomes 4 times the dimension of the input feature map. Then a 3*3 convolution combination module is performed, and finally the dimension of the feature map is restored to the input dimension through a 1*1 convolution combination module, and then a drop layer is passed to alleviate the overfitting problem. Therefore, the CGMLP module directly outputs the feature map after strengthening the information interaction and spatial detail information of the water body feature map. The role of the normalization layer is to stabilize the training process and accelerate the convergence of the model; the main role of the residual connection is to alleviate the gradient disappearance and reduce the loss of feature information.
[0043] The function of the decoder part is to restore the resolution and fuse the features to generate accurate output. The present invention uses the basic self-attention mechanism to restore the resolution of the feature map by upsampling layer by layer. The entire decoder uses three basic self-attention mechanisms, and makes full use of the global features and spatial detail features in the decoding process. The jump connection between the encoder and the decoder uses different connection methods according to the characteristics of the feature map. When the global information features and the upsampled features are connected to each other, an additive connection method is adopted; when the spatial detail features and the upsampled features are connected to each other, the present invention designs a special feature fusion method, which controls the addition of the two feature maps according to a dynamic training weight, so that the fusion of different types of feature maps is more scientific, and finally the feature map is restored to its original size through the segmentation head to complete the segmentation task.
[0044] S4, using the data set in step S2 to train the constructed MSMViT network model, and using the trained MSMViT network model for water body extraction.
[0045] Experimental setup
[0046] To ensure the repeatability of this experiment and the fairness of the comparative experiments, all comparative experiments were conducted under the same implementation environment, with the hardware configuration as follows: NVIDIA GeForce RTX 3090 GPU was used for display, and the memory size was 64GB; the software configuration was as follows: the network model was built based on the pytorch1.10 framework under python3.9, and the gradient descent optimization method used was the AdamW optimizer; the loss function adopted the binary cross entropy loss function; the enhancement methods for the training set data during the training process included flipping, random cropping, random cropping combination, and mirroring.
[0047] Case Showcase
[0048] The data used in the experiment are self-made data sets and open source GID data sets. The test area selected for the self-made water body data set is located in Chongqing and Yueyang, China, including urban and rural areas. The images used are Gaofen-2 satellite image data with an image resolution of 1m. The water body vector data was obtained by visual interpretation and converted into binary image labels. A total of 4910 data sets with a size of 512*512 were obtained through image cropping. The data were also divided into training set, validation set, and test set according to the ratio of 7:2:1. The GID data set is a high-resolution remote sensing image provided by Gaofen-2. After image fusion, the image resolution reaches 1m. The image includes a total of 150 remote sensing images. The original data set is a 5-category land feature classification data. In this experiment, only water body information is retained and set as classification information, and other categories are set as background information. At the same time, the missing water bodies in the labels are supplemented accordingly.
[0049] This study mainly uses five evaluation indicators, namely overall accuracy (OA), precision (P), recall (R), F1-score (F1), and intersection over union (IoU), to evaluate the extraction accuracy of the network. These five evaluation indicators can effectively evaluate the comprehensive extraction ability of the semantic segmentation network. Among them, the overall accuracy (OA) represents the ratio of the correct water pixel value extracted by the network model to the total pixel value; precision (P) represents the ratio of the correct water pixel value extracted by the network model to the pixel value of all water extraction; recall (R) represents the ratio of the water pixel value extracted by the network model to the water pixel value in the true label; F1-score (F1) represents the harmonic mean of precision and recall; intersection over union (IoU) represents the ratio of the intersection of the true label water pixel and the predicted water pixel to their union. In order to verify the effectiveness and superiority of the proposed water extraction method, this paper selects several commonly used and effective water extraction methods and deep learning methods for comparative experiments. The selected comparative methods include Unet, PSPNet, SegNet, MWEN, and ViT. Table 1 is a comparison table of the progress of six water extraction methods. Figure 5 The following table shows the image data, corresponding labels and water extraction results of each method.
[0050] Table 1. Accuracy comparison of MSMViT and other methods
[0051]
[0052] Embodiment 2
[0053] Based on the same inventive concept, the present invention also provides an electronic device, including one or more processors; a storage device for storing one or more programs; when one or more programs are executed by the one or more processors, the one or more processors implement the method described in Example 1.
[0054] Since the device introduced in the second embodiment of the present invention is an electronic device used to implement the high-resolution remote sensing image water extraction method based on VisionTransformer in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, the technical personnel in the field can understand the specific structure and deformation of the electronic device, so it is not repeated here. All electronic devices used in a method of the embodiment of the present invention belong to the scope of protection of the present invention.
[0055] Embodiment 3
[0056] Based on the same inventive concept, the present invention further provides a computer-readable medium on which a computer program is stored, and when the program is executed by a processor, the method described in the first embodiment is implemented.
[0057] Since the device introduced in the third embodiment of the present invention is a computer-readable medium used to implement the high-resolution remote sensing image water extraction method based on VisionTransformer in the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, a person skilled in the art can understand the specific structure and deformation of the electronic device, so it is not repeated here. All electronic devices used in a method of the embodiment of the present invention belong to the scope of protection of the present invention.
[0058] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for extracting water from high-resolution remote sensing images based on Vision Transformer, characterized in that: The following steps are involved: S1, obtain high-resolution remote sensing images and perform preprocessing; S2, visually interpret the water bodies of the pre-processed high-resolution remote sensing images and delineate the water body boundaries in the images to produce a dataset for deep learning model training; S3, constructing the MSMViT network model, which includes an encoder and a decoder; the encoder includes a detail information retention module and a water body feature extraction module; there are four multi-scale water body feature extraction modules, and the jump connection part thereof includes a dynamic information fusion module; The detail information retaining module first obtains an image feature map, and then processes the image feature map to obtain spatial detail information; The multi-scale water feature extraction modules all include a mixW-Block module, which is composed of LMSMHSA, a global multi-layer perceptron based on convolution, a normalization layer, and a residual connection. The LMSMHSA is improved from the multi-head self-attention mechanism MHSA of the ordinary ViT. The specific improvements include using convolution to replace the full connection layer in the ordinary ViT, and using the Taylor formula to replace the Softmax function in the ordinary ViT. S4, using the data set in step S2 to train the constructed MSMViT network model, and using the trained MSMViT network model for water body extraction.
2. The method for extracting water from high-resolution remote sensing images based on Vision Transformer according to claim 1, characterized in that: The preprocessing in S1 includes performing radiation correction, geometric correction, image fusion, mosaicking, and cropping on the high-resolution remote sensing image.
3. The method for extracting water from high-resolution remote sensing images based on Vision Transformer according to claim 1, characterized in that: In the S2, the high-resolution image water body is visually interpreted, and the image water body boundary is outlined using Labelme software and converted into water body true value label data.
4. The method for extracting water from high-resolution remote sensing images based on Vision Transformer according to claim 1, characterized in that: In S2, after reading the high-resolution remote sensing image data and the corresponding water body label data, the sliding window method is used to cut them into sample data of the same size, and the entire data set is divided into a training set, a validation set and a test set according to the proportion.
5. The method for extracting water from high-resolution remote sensing images based on Vision Transformer according to claim 1, characterized in that: The spatial detail information retention module consists of two parts, A and B. Part A provides a feature map for the main structure, which is mainly composed of two downsampling modules and a fusion module. Part B is a spatial detail information extraction module.
6. The method for extracting water from high-resolution remote sensing images based on Vision Transformer according to claim 5, characterized in that: The spatial detail information acquisition module is implemented by multiple layers of modules. The processing process of each layer of modules includes convolution, batch normalization and activation function. Each layer obtains detailed spatial information by stepwise downsampling.
7. The method for extracting water from high-resolution remote sensing images based on Vision Transformer according to claim 1, characterized in that: The encoder also includes a patch merging module, which is used by the water body feature extraction module to adjust the size of the feature map. The specific process includes: using convolution and pooling operations to downsample the feature map to half of the original feature map, and expanding the dimension of the feature map to twice the original, and then splicing the feature maps obtained by the two operations.
8. The method for extracting water from high-resolution remote sensing images based on Vision Transformer according to claim 1, characterized in that: The decoder uses a basic self-attention mechanism to restore the resolution of the feature map by upsampling layer by layer.
9. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the high-resolution remote sensing image water body extraction method as described in any one of claims 1-8.
10. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for extracting water from high-resolution remote sensing images as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Polarization three-dimensional reconstruction method and system based on multiple receptive field fusion network
CN117036613A
Target region segmentation method and system of ultrasonic image and electronic equipment
CN117934824A
Remote sensing image semantic segmentation method based on Mama and Transform architecture fusion
CN119206227A
Automatic seismic facies identification method based on combination of self-attention mechanism and u-shaped structure
WO2024000709A1