A method, system, device and medium for extracting surface water bodies based on an improved Segformer network and remote sensing images
The optimal band combination is obtained through the improved Segformer network model and Bayesian optimization method, which solves the problem of band combination neglecting in surface water extraction of remote sensing images, and achieves efficient and accurate water extraction, which is suitable for large-scale and complex environments.
Patent Information
- Application Number
- CN202310617074.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-29
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2043-05-29
AI Technical Summary
The prior art ignores the rich band combination of images in the surface water body extraction of remote sensing images, resulting in poor water body extraction effect, poor stability, poor robustness, and low efficiency and accuracy of traditional convolutional neural network models.
The improved Segformer network model is adopted to obtain the optimal band combination through Bayesian optimization method, and the water body extraction performance is evaluated based on the determination coefficient and root mean square error, data enhancement and feature fusion are carried out to build a water body extraction network.
It improves the accuracy and robustness of water body extraction, reduces time and labor costs, is suitable for large-scale, long-term, and highly complex water body extraction tasks, and provides efficient water body extraction evaluation standards.
Smart Images

Figure CN116612387B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of surface water body extraction, and in particular relates to a surface water body extraction method, system, equipment and medium based on an improved Segformer network and remote sensing images. Background Art
[0002] Surface water bodies are a vital component of the Earth's hydrosphere, and their extraction is a crucial prerequisite for water resource management and conservation. Influenced by climate, seasons, and human activities, surface water bodies are constantly changing and exhibit highly dynamic characteristics in both space and time. Therefore, accurate and rapid monitoring of water bodies is crucial for environmental research and the management of terrestrial ecosystems. Currently, methods for extracting surface water bodies are generally categorized into three types: thresholding methods based on single-band imagery, identification methods based on spectral indices, and image classification methods. However, because the spectral reflectance characteristics of water bodies are affected by numerous natural conditions, single-band images only capture limited surface features, resulting in poor performance of single-band thresholding methods. Spectral index-based water body identification methods can better reflect image features and utilize different indices to identify water image information in different bands. However, determining thresholds can be difficult in some cases. Previous studies on remote sensing image segmentation have mostly considered image segmentation using combinations of visible light bands, ignoring the richer range of band combinations available in remote sensing imagery.
[0003] Remote sensing provides us with the advantages of macroscopic, dynamic, continuous and low-cost monitoring of land objects, which can help us better understand spatiotemporal changes. Due to its characteristics of large observation range, fast update time and rich information, remote sensing data has been widely used in water resource change monitoring, water quality assessment, flood disaster monitoring and loss assessment. At present, the most effective and efficient method for extracting surface water is to use satellite remote sensing images to systematically extract surface water bodies in a certain area. However, different data sources and different technical solutions will have a huge impact on the efficiency, accuracy and extraction results of surface water bodies based on remote sensing images. Among the existing technical solutions, most of the current research mainly uses three types of technologies to achieve the purpose of surface water body extraction based on remote sensing images: threshold method based on single-band image, recognition method based on spectral index and image classification method. However, these methods do not pay attention to the rich band combinations of remote sensing images, and the efficiency and accuracy of the model need to be improved. Therefore, the existing technology has the following shortcomings:
[0004] 1. Since the reflectance characteristics of water spectra are affected by many natural conditions, the ground features reflected by single-band images are limited, resulting in poor results of the threshold segmentation method based on single band.
[0005] 2. The water body recognition method based on spectral index needs to obtain water body image information in different bands according to different design indices, but in some cases the threshold is difficult to determine, and an imperfect threshold will have a great impact on the water body extraction results.
[0006] 3. Most water extraction techniques based on remote sensing image segmentation only consider imagery with visible light band combinations for ground feature segmentation, ignoring the rich band combinations of remote sensing images. This results in poor robustness and accuracy in water extraction results. In summary, the aforementioned methods suffer from poor performance, stability, robustness, and accuracy in water extraction tasks based on remote sensing images.
[0007] The patent application with application number [CN202010806778.7] discloses a deep learning remote sensing image water body extraction system, which uses ResNet50 as the extraction network core and uses the coarse-grained semantic map and fine-grained semantic map of the image to be segmented at different positions in the neural network to obtain a binary result map of water body extraction in a certain area. The patent application with application number [CN201910059294.8)] discloses a remote sensing image water body extraction method and system based on deep learning, which combines U-net and Densene neural networks to construct a water body extraction network CNNs, thereby achieving the water body extraction task of remote sensing images. Both of the above-mentioned existing technologies use variants of traditional convolutional neural networks, and both ignore the rich band combination information of the remote sensing image itself, which may lead to poor water body extraction results. Patent application number [CN201510272030.2] discloses an automatic water extraction method based on Landsat OLI multispectral remote sensing imagery. This method employs a water index and then uses the maximum between-class variance (Otsd) method to automatically select a threshold for water extraction. Essentially, this is a threshold-based segmentation method suitable for extracting water bodies at small scales and within short timeframes. However, in the case of long time series, large scales, and highly complex water bodies, this method is difficult to determine the thresholds for water and other ground features, significantly impacting the extraction results and resulting in poor model robustness and stability. Summary of the Invention
[0008] In order to overcome the shortcomings of the above-mentioned prior art, the purpose of the present invention is to provide a surface water body extraction method, system, equipment and medium based on an improved Segformer network and remote sensing images, by comparing different band combinations of original images and evaluating the data source to obtain a reliable band combination, and using a combination of the coefficient of determination (R 2) and root mean square error (RMSE) are used to evaluate the performance of the water body extraction network model in extracting water body width; the present invention has the advantages of good effect, high efficiency, high stability and high robustness in the task of water body extraction from remote sensing images, and can evaluate the performance of the model in extracting river width in specific scenarios, greatly saving manpower, material resources and time costs.
[0009] In order to achieve the above object, the technical solution adopted by the present invention is:
[0010] A surface water body extraction method based on an improved Segformer network and remote sensing images includes the following steps:
[0011] Step 1: Obtain remote sensing image data of surface water bodies and create a one-to-one correspondence between the original water body image dataset and the labeled image dataset;
[0012] Step 2: Establish an improved Segformer network model;
[0013] Step 3: Perform data enhancement and band analysis on the cropped water body original image dataset and the labeled image dataset obtained in step 1 to obtain the training set, test set, and validation set image data sources with the optimal band combination;
[0014] Step 4: Use the Bayesian optimization method to obtain the optimal hyperparameters and train the improved Segformer network model established in step 2 to obtain the trained improved Segformer water extraction network model;
[0015] Step 5: Use the training set, test set, and test set of the validation set image data source of the optimal band combination obtained in step 3 to test the improved Segformer water body extraction network model trained in step 4. Use precision, recall, F1-score, and mean intersection over union (mIoU) as indicators to evaluate the performance of the improved Segformer water body extraction network model trained in step 4.
[0016] Step 6: Combine the coefficient of determination R 2 The performance of the trained improved Segformer water body extraction network model obtained in step 4 for water body width extraction is evaluated by the root mean square error (RMSE).
[0017] The specific method of step 1 is:
[0018] Step 1.1: Obtain remote sensing image data of surface water bodies and use LabelMe software to label the water bodies in the remote sensing image data using a visual interpretation method to obtain binary classification labels corresponding to water bodies and non-water bodies;
[0019] Step 1.2: Crop the remote sensing image obtained in step 1.1 and the binary classification labels corresponding to water bodies and non-water bodies marked using the Labelme software to m × n pixels, and obtain the cropped water body original image dataset and label image dataset that correspond one to one.
[0020] The specific method of step 2 is:
[0021] The water body extraction network backbone is constructed by connecting Segformerblock and CNN block in parallel. CNN block and Segformerblock send the extracted different texture features into two decoders with the same structure before performing feature fusion. Each CNN Block consists of two 3*3 convolutional layers, each of which is followed by a linear unit (ReLU) and a 2*2 maximum pooling layer with a stride of 2. Each TransformerBlock consists of an Efficient Self-Attention layer, a Mix-FNN layer, and an Overlap PatchMerging layer. The Mix-FNN layer is composed of a linear connection of an MLP layer, a 3*3 convolutional layer, a GELU layer, and another MLP layer.
[0022] The specific method of step 3 is:
[0023] Step 3.1: Expand the cropped, one-to-one corresponding water body original image dataset and label image dataset obtained in step 1 by rotating, mirroring, blurring, and adding noise to obtain expanded training data;
[0024] Step 3.2: Divide the expanded training data obtained in step 3.1 into n groups according to different band combinations to obtain n groups of data with different band combinations. Perform Gaussian stretching on the n groups of data with different band combinations to obtain n groups of data with different band combinations after Gaussian stretching. Divide the n groups and the n groups of data with different band combinations after Gaussian stretching into training sets, validation sets, and test sets in the same proportion to obtain data with the same amount of data and the same division ratio but different band combinations, or data with the same band combination but after Gaussian stretching.
[0025] Step 3.3: Use the improved Segformer network model obtained in step 2 to conduct experimental analysis on the data obtained in step 3.2 with the same data volume, the same division ratio but different band combinations, or the same band combination but after Gaussian stretching. By comparing the loss iteration curves during the experiment of the improved Segformer network model, select the optimal band combination image and obtain the training set, test set, and validation set image data sources of the optimal band combination.
[0026] The specific method of step 4 is:
[0027] Step 4.1: Use the Bayesian optimization method to obtain the optimal hyperparameters of the improved Segformer network model established in step 2, and obtain the improved Segformer water extraction network model with the optimal hyperparameters;
[0028] Step 4.2: Use the training set, test set, and validation set image data source of the optimal band combination obtained in step 3.3 to train the improved Segformer water body extraction network model with optimal hyperparameters obtained in step 4.1. During the training process, use the training set, test set, and validation set image data source of the optimal band combination obtained in step 3.3 to evaluate the current performance of the improved Segformer water body extraction network model with optimal hyperparameters obtained in step 4.1. After the training is completed, the trained improved Segformer water body extraction network model is obtained.
[0029] The calculation formulas for the precision, recall, and F1-score evaluation indicators in step 5 are as follows:
[0030]
[0031]
[0032]
[0033]
[0034]
[0035] Among them, the meanings of the parameters in Formula 1, Formula 2, Formula 3, and Formula 4 are as follows:
[0036] TP (True Positive): The prediction is correct, the prediction result is positive, and the actual result is positive; FP (False Positive): The prediction is wrong, the prediction result is positive, and the actual result is negative; FN (False Negative): The prediction is wrong, the prediction result is negative, and the actual result is positive; TN (True Negative): The prediction is correct, the prediction result is negative, and the actual result is negative;
[0037] In Formula 5, P refers to Precision, R refers to Recall, and F1 refers to the balanced F score.
[0038] The expression for evaluating the water body width extraction performance in step 6 is:
[0039] N i =GETNUM(Latitude[1]:Latitude[n]), i∈(Longitude[1],Longitude[m]) Equation 6
[0040] Width i =N i *resolution Formula 7
[0041]
[0042]
[0043] In Equation 6, n and m represent the number of pixels on the longitude and latitude lines in the study area, respectively. GETNUM represents the operation of counting the total number of water body points. N i Represents the number of longitude river pixels at longitude i;
[0044] In formula 7, resolution is the resolution of the remote sensing image, Width i is the actual river width obtained by calculation;
[0045] In formula 8 and formula 9, PreWidth i and Width i They represent the predicted river width and the calculated actual river width at longitude i, MeanWidth represents the average river width, Longitude[i] represents the i-th longitude, and R 2 stands for coefficient of determination, and RMSE stands for mean square error.
[0046] The present invention also provides a surface water body extraction system based on an improved Segformer network and remote sensing images, comprising:
[0047] Data processing module: used to crop a complete remote sensing image data to be segmented into m×n pixels to produce a segmentation data set;
[0048] Model building module: used to build an improved Segformer network model;
[0049] Data analysis module: used to analyze the optimal band combination of remote sensing image data to be segmented;
[0050] Water body segmentation module: used to input the remote sensing image to be segmented with the optimal band combination into the improved Segformer network model for water body extraction;
[0051] Result evaluation module: used to evaluate the performance of the trained improved Segformer water extraction network model using precision, recall, F1-score, and mean intersection over Union (mIoU) as indicators; and used to combine the coefficient of determination (R 2 ) and root mean square error (RMSE) are used to evaluate the performance of the trained improved Segformer water body extraction network model for water body width extraction.
[0052] The present invention also provides a surface water body extraction device based on an improved Segformer network and remote sensing images, comprising:
[0053] A memory for storing a computer program for implementing the surface water body extraction method based on an improved Segformer network and remote sensing images;
[0054] A processor is used to implement the surface water body extraction method based on the improved Segformer network and remote sensing images when executing the computer program.
[0055] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement a surface water body extraction method based on an improved Segformer network and remote sensing images.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] 1. The present invention uses the trained improved Segformer model to extract surface water bodies from remote sensing images to obtain accurate water body extraction results, which not only saves time costs, labor costs and financial costs, but also has high accuracy and strong technical robustness in water body extraction.
[0058] 2. The present invention analyzes remote sensing data of different band combinations and selects the band combination with the best water body extraction effect for extraction operation, which not only enhances the reliability of the sample, but also improves the accuracy of the model in extracting water bodies.
[0059] 3. The present invention designs a new method combining the coefficient of determination (R 2 ) and root mean square error (RMSE), which provides an evaluation criterion for water body extraction tasks in specific scenarios.
[0060] 4. Since the present invention adopts deep learning technology and uses the improved Segformer network model to extract water bodies in the study area, it achieves an end-to-end water body extraction effect from the image to be segmented to the water body extraction result. In addition, the performance of the improved Segformer network in this paper is better than that of many semantic segmentation networks.
[0061] 5. The present invention adopts the Segformer network with superior performance as the extraction network core and improves its decoder. The present invention fully considers the rich band combination characteristics of the remote sensing image itself, compares the results of different band combinations, and selects the optimal band for segmentation tasks, which greatly improves the model segmentation performance and efficiency.
[0062] 6. The water body extraction method based on the improved deep neural network designed in the present invention can be applied to large-scale, long-time series, and highly complex water body extraction tasks, and does not require too much manual work.
[0063] In summary, the improved Segformer network model proposed in the present invention has superior performance in surface water extraction tasks. Both accuracy and robustness are far superior to traditional water extraction methods and water extraction methods based on classic segmentation networks. In addition, the improved Segformer network model proposed in the present invention can be applied to large-scale, long-time series, and highly complex water extraction tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 Flow chart of the method of the present invention.
[0065] Figure 2 This is the extraction network structure diagram of the present invention.
[0066] Figure 3 This is the Landsat 8-OLI image band description used in the embodiment of the present invention, including the band name, bandwidth, and resolution.
[0067] Figure 4 This is the band combination setting performed in the embodiment of the present invention.
[0068] Figure 5 The results verify the model of the present invention under different band combinations.
[0069] Figure 6 This is the extraction result of an embodiment of the present invention in a certain research area.
[0070] Figure 7 This is the performance evaluation of the embodiment of the present invention in the study area.
[0071] Figure 8This is an evaluation of various indicators in a certain research area according to an embodiment of the present invention. DETAILED DESCRIPTION
[0072] The following is an example of a surface water body extraction method based on an improved Segformer network and remote sensing images, and the technical solution of the present invention is further explained and illustrated in detail with reference to the accompanying drawings.
[0073] like Figure 1 As shown in Figure 1, step 1: obtain Landsat 8-OLI remote sensing image data and create a one-to-one corresponding water body original image dataset and label image dataset.
[0074] Step 1.1: Obtain Landsat 8-OLI remote sensing image data of the Wei River Basin in 2016. Use LabelMe software to visually interpret and label the water bodies in the Landsat 8-OLI remote sensing image of the Wei River Basin in 2016 to obtain binary classification labels corresponding to water bodies and non-water bodies.
[0075] Step 1.2: The original Landsat8-OLI remote sensing image data of the Wei River Basin in 2016 obtained in Step 1.1 and the corresponding binary classification labels of water bodies and non-water bodies marked with Labelme software are cropped to 256×256 pixels, where water bodies are marked as 1 and non-water bodies are marked as 0. This gives the cropped original water body image dataset and labeled image dataset of the Wei River Basin, which correspond one to one.
[0076] like Figure 2 As shown, step 2: establish an improved Segformer network model; the specific method of step 2 is:
[0077] Referring to the idea of multi-feature fusion, a robust water body extraction network backbone is constructed by connecting Segformerblock and CNN block in parallel. CNNblock and Segformer block send the extracted different texture features into two decoders with the same structure before performing feature fusion. Each CNN Block consists of two 3*3 convolutional layers, each of which is followed by a linear unit (ReLU) and a 2*2 maximum pooling layer with a stride of 2. Each TransformerBlock consists of an Efficient Self-Attention layer, a Mix-FNN layer, and an Overlap Patch Merging layer. The Mix-FNN layer is composed of a linear connection of an MLP layer, a 3*3 convolutional layer, a GELU layer, and another MLP layer.
[0078] like Figure 3 As shown, step 3: In order to improve the robustness of the model and prevent the model from overfitting, data enhancement operations and band analysis operations are performed on the one-to-one corresponding original water body image dataset and labeled image dataset of the Weihe River Basin obtained in step 1.2.
[0079] Step 3.1: Expand the one-to-one correspondence between the original water body image dataset and the labeled image dataset of the Weihe River Basin obtained in step 1.2 by rotating, mirroring, blurring, and adding noise to obtain the expanded training data.
[0080] like Figure 4 As shown, step 3.2: group the expanded training data obtained in step 3.1 according to different band combinations, and combine (Band4, Band3, Band2) and (Band4, Band3, Band2) after Gaussian stretching, (Band5, Band6, Band4) and (Band5, Band6, Band4) after Gaussian stretching, and divide these four groups of data into training set, validation set and test set according to 6:3:1, so as to obtain four groups of data with the same data amount and the same division ratio but different band combinations.
[0081] like Figure 5 As shown, step 3.3: Use the improved Segformer network model obtained in step 2 to conduct experimental analysis on the four groups of data with the same data volume and division ratio but different band combinations obtained in step 3.2. By comparing the loss iteration curves during the experiment of the improved Segformer network model, the (Band5, Band6, Band4) data after Gaussian stretching is selected as the optimal band combination image, and the training set, test set and validation set image data sources of the optimal band combination are obtained.
[0082] Step 4: Select the improved Segformer network model as the core model of the present invention, use the Bayesian optimization method to obtain the optimal hyperparameters, train the improved Segformer network model established in step 2, and obtain an improved segformer water extraction network model with superior performance after training.
[0083] Step 4.1: Use the Bayesian optimization method to obtain the optimal hyperparameters of the improved Segformer network model established in step 2, and obtain an improved Segformer water body extraction network model with optimal hyperparameters.
[0084] Step 4.2: Use the training set, test set, and validation set image data source of the optimal band combination obtained in step 3.3 to train the improved Segformer water body extraction network model with optimal hyperparameters obtained in step 4.1. During the training process, use the training set, test set, and validation set image data source of the optimal band combination obtained in step 3.3 to evaluate the current performance of the improved Segformer water body extraction network model with optimal hyperparameters obtained in step 4.1. After the training is completed, a trained improved Segformer water body extraction network model is obtained.
[0085] like Figure 6 As shown, step 5: use the training set, test set and test set in the validation set image data source of the optimal band combination obtained in step 3.3 to test the trained improved Segformer water body extraction network model obtained in step 4.2, and use precision, recall, F1 score (F1-score), and Mean Intersection over Union (mIoU) as indicators to evaluate the performance of the trained improved Segformer water body extraction network model obtained in step 4.2 in water body extraction in the Weihe River Basin.
[0086] The calculation formulas for the precision, recall and F1 score evaluation indicators are as follows:
[0087]
[0088]
[0089]
[0090]
[0091]
[0092] Among them, the meanings of the parameters in Formula 1, Formula 2, Formula 3, and Formula 4 are as follows:
[0093] TP (True Positive): The prediction is correct, the prediction result is positive, and the actual is positive. FP (False Positive): The prediction is wrong, the prediction result is positive, and the actual is negative. FN (False Negative): The prediction is wrong, the prediction result is negative, and the actual is positive. TN (True Negative): The prediction is correct, the prediction result is negative, and the actual is negative.
[0094] In Formula 5, P refers to Precision, R refers to Recall, and F1 refers to the balanced F score.
[0095] In addition, the present invention proposes a new method for evaluating the performance of water body width extraction, namely, a new method for evaluating the performance of water body extraction network model for water body width extraction by combining the decision system (R2) and the root mean square error (RMSE).
[0096] like Figure 7 As shown, step 6: Combined determination coefficient (R 2 ) and root mean square error (RMSE) are used to evaluate the performance of the trained improved Segformer water body extraction network model obtained in step 4.2 for water body width extraction.
[0097] The expression for evaluating the water body width extraction performance is:
[0098] N i =GETNUM(Latitude[1]:Latitude[n]), i∈(Longitude[1],Longitude[m]) Equation 6
[0099] Width i =N i *resolution Formula 7
[0100]
[0101]
[0102] In Equation 6, n and m represent the number of pixels on the longitude and latitude of the study area, respectively. GETNUM represents the operation of counting the total number of water body points. Ni represents the number of longitude river pixels when the longitude is i.
[0103] In formula 7, resolution is the resolution of the remote sensing image, Width i is the actual river width obtained by calculation;
[0104] In formula 8 and formula 9, PreWidth i and Width i They represent the predicted river width and the calculated actual river width at longitude i, MeanWidth represents the average river width, Longitude[i] represents the i-th longitude, and R 2 stands for coefficient of determination, and RMSE stands for mean square error.
[0105] The present invention also provides a surface water body extraction system based on an improved Segformer network and remote sensing images, comprising:
[0106] Data processing module: used to implement the step 1 of cropping a complete Landsat8-OLI remote sensing image data to be segmented into m×n pixels to produce a segmentation dataset;
[0107] Model building module: used to implement the improved Segformer network model established in step 2;
[0108] Data analysis module: used to implement the optimal band combination of the Landsat 8-OLI remote sensing image data to be segmented in step 3, so as to facilitate the subsequent segmentation task;
[0109] Water body segmentation module: used to implement the step 4 of inputting the Landsat 8-OLI remote sensing image to be segmented with the optimal band combination selected into the improved Segformer network model proposed in the present invention to extract water bodies;
[0110] Result evaluation module: used to implement the performance evaluation of the trained improved Segformer water extraction network model obtained in step 4 using precision, recall, F1-score, and mean intersection over Union (mIoU) as indicators in step 5; and used to implement the combined determination coefficient (R 2 ) and root mean square error (RMSE) are used to evaluate the performance of the improved Segformer water body extraction network model trained in step 4.2 for water body width extraction.
[0111] The present invention also provides a surface water body extraction device based on an improved Segformer network and remote sensing images, comprising:
[0112] A memory for storing a computer program for implementing the surface water body extraction method based on an improved Segformer network and remote sensing images;
[0113] A processor is used to implement the surface water body extraction method based on the improved Segformer network and remote sensing images when executing the computer program.
[0114] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement a surface water body extraction method based on an improved Segformer network and remote sensing images.
[0115] Figure 1The figure is a flowchart of the method of the present invention. After image cropping and band analysis, the original image is divided into training data and test data. The training data, after data enhancement, is fed into the improved Segformer water body extraction network proposed in the present invention for network training. After model training is complete, the test data is fed into the model, and the extraction results are compared with the labels. The performance of the water body extraction network model for water body width extraction is evaluated using precision, recall, F1-score, mean intersection over union (mIoU), and a new method proposed in the present invention that combines the decision square (R2) and root mean square error (RMSE) to evaluate the performance of the water body extraction network model.
[0116] Attachment Figure 2 The improved Segformer network structure proposed in this invention consists of four CNN blocks and four Transformer blocks in parallel to form the network backbone. Each CNN block consists of two 3*3 convolutional layers, each of which is followed by a linear unit (ReLU) and a 2*2 maximum pooling layer with a stride of 2. Each Transformer block consists of an Efficient Self-Attention layer, a Mix-FNN layer, and an Overlap Patch Merging layer. The Mix-FNN layer is composed of an MLP layer, a 3*3 convolutional layer, a GELU layer, and another MLP layer linearly connected.
[0117] Figure 3 This is information on different bands and resolutions of the Landsat 8-OLI remote sensing image used in the embodiment of the present invention.
[0118] Figure 4 For the band combination experiment designed for the present invention, four groups of experiments are set up, namely, (Band4, Band3, Band2), (Band4, Band3, Band2) after Gaussian stretching, (Band5, Band6, Band4) and (Band5, Band6, Band4) after Gaussian stretching.
[0119] Figure 5 The error iteration graph of the four band combination experiments on the validation set as the number of model training rounds increases. From the analysis of the graph, it can be seen that the band combination (Band5, Band6, Band4) after Gaussian stretching of the data set is the best band combination.
[0120] Figure 6The following is a comparison of the water body extraction results of the present invention in several typical scenes in the Weihe River Basin with those of U-Net and Seg-Net. The first column is the test image scene, the second column is the label of the test image, the third column is the extraction results of the model proposed by the present invention in these scenes, the fourth column is the extraction results of U-Net in these scenes, and the fifth column is the extraction results of Seg-Net in these scenes. Among them, the yellow dotted box is an undetected water body, and the red dotted box is a misclassified water body. The comparison of the results shows that the water body extraction model proposed in the patent of this invention not only outperforms other methods in thick water bodies, small water bodies, reservoirs and tiny water bodies, but also has very few misclassified water bodies, so it has the best performance.
[0121] Figure 7 In order to evaluate the overall water extraction results of the water extraction model proposed in the present invention in the Weihe River Basin, precision, recall, F1-score, and mean intersection over union (mIoU) were used as indicators to evaluate the performance of the water extraction model proposed in the present invention in the Weihe River Basin.
[0122] Figure 8 The water body extraction model proposed in the present invention is used to extract water bodies in the Xi'an-Xianyang section of the Weihe River. Then, a new method proposed in the present invention that combines the determination coefficient R2 and the root mean square error (RMSE) to evaluate the performance of the water body extraction network model in extracting water body width is used to evaluate the extraction results. The correlation coefficient can reach 0.954, and the root mean square error is 41.66m (one pixel is 30m).
[0123] like Figure 8 As shown, compared with the existing technology, the improved Segformer water body extraction network model proposed in the present invention is far superior to classic network models such as UNet and SegNet in terms of accuracy and robustness for water body extraction. The improved Segformer water body extraction network model proposed in the present invention can pay good attention to the edge information of the water body, and far exceeds classic network models such as UNet and SegNet in the extraction performance of small water bodies and discontinuous water bodies, and the misclassification rate is very low. In addition, the improved Segformer water body extraction network model proposed in the present invention performs well in the task of water body extraction in the Weihe River Basin. After evaluating the extraction results by combining the determination coefficient R2 and the root mean square error (RMSE), the extracted water body width is very close to the actual width of the water body.
[0124] The application prospects of the present invention are as follows: surface water includes water bodies such as rivers, lakes, and swamps, which are very important for agriculture, industry, aquaculture, and aquatic and terrestrial ecosystems. Changes in the area of water bodies have a huge impact on ecosystems, biogeochemical cycles, and other environmental changes. Therefore, quickly and accurately obtaining spatial distribution information of surface water is of great significance to the monitoring and management of water resources. In addition, affected by climate, seasons, and human activities, surface water is constantly changing and exhibits highly dynamic characteristics in space and time. Therefore, accurate and rapid monitoring of water bodies is crucial for environmental research and the management of terrestrial ecosystems. However, monitoring and understanding the spatiotemporal dynamics of large-scale water bodies remains a major challenge. Currently, water body extraction technology based on remote sensing images has many problems such as low accuracy, poor robustness, low efficiency, and high resource consumption.
[0125] This paper proposes a surface water extraction technology based on a deep learning algorithm for remote sensing imagery. The technology considers and studies the impact of the rich band combinations of remote sensing imagery on water extraction results, selects the optimal band combination and extraction model, and designs an evaluation metric for evaluating water body width extraction, providing an evaluation reference for specific water body extraction tasks. This technology has broad applications in digital twins, water resource monitoring and management, disaster warning, post-disaster analysis, and ecological protection, and provides researchers with a high-precision model and research approach for surface water and surface feature extraction tasks.
Claims
1. A surface water extraction method based on an improved Segformer network and remote sensing images, characterized by: The following steps are involved: Step 1: Obtain remote sensing image data of surface water bodies and create a one-to-one correspondence between the original water body image dataset and the labeled image dataset; Step 2: Establish an improved Segformer network model; The specific method of step 2 is: The backbone of the water body extraction network is constructed by connecting the Segformer block and the CNN block in parallel. The CNN block and the Segformer block send the extracted different texture features to two decoders with the same structure, and then perform feature fusion. Each CNN block consists of two 3*3 convolutional layers, each of which is followed by a linear unit (ReLU) and a 2*2 maximum pooling layer with a stride of 2. Each Transformer Block consists of an Efficient Self-Attention layer, a Mix-FNN layer, and an Overlap Patch Merging layer. The Mix-FNN layer is composed of an MLP layer, a 3*3 convolutional layer, a GELU layer, and another MLP layer connected linearly. Step 3: Perform data enhancement and band analysis on the cropped water body original image dataset and the labeled image dataset obtained in step 1 to obtain the training set, test set, and validation set image data sources with the optimal band combination; Step 4: Use the Bayesian optimization method to obtain the optimal hyperparameters and train the improved Segformer network model established in step 2 to obtain the trained improved Segformer water extraction network model; Step 5: Use the training set, test set, and test set of the validation set image data source of the optimal band combination obtained in step 3 to test the improved Segformer water body extraction network model trained in step 4. Use precision, recall, F1-score, and mean intersection over union (mIoU) as indicators to evaluate the performance of the improved Segformer water body extraction network model trained in step 4. Step 6: Combine the coefficient of determination R 2 The performance of the trained improved Segformer water body extraction network model obtained in step 4 for water body width extraction is evaluated by the root mean square error (RMSE).
2. The surface water extraction method based on the improved Segformer network and remote sensing imagery according to claim 1, characterized in that: The specific method of step 1 is: Step 1.1: Obtain remote sensing image data of surface water bodies and use LabelMe software to label the water bodies in the remote sensing image data using a visual interpretation method to obtain binary classification labels corresponding to water bodies and non-water bodies; Step 1.2: Crop the remote sensing image obtained in step 1.1 and the binary classification labels corresponding to water bodies and non-water bodies marked using the Labelme software to m × n pixels, and obtain the cropped water body original image dataset and label image dataset that correspond one to one.
3. The surface water extraction method based on the improved Segformer network and remote sensing imagery according to claim 1, characterized in that: The specific method of step 3 is: Step 3.1: Expand the cropped, one-to-one corresponding water body original image dataset and label image dataset obtained in step 1 by rotating, mirroring, blurring, and adding noise to obtain expanded training data; Step 3.2: Divide the expanded training data obtained in step 3.1 into n groups according to different band combinations to obtain n groups of data with different band combinations. Perform Gaussian stretching on the n groups of data with different band combinations to obtain n groups of data with different band combinations after Gaussian stretching. Divide the n groups and the n groups of data with different band combinations after Gaussian stretching into training sets, validation sets, and test sets in the same proportion to obtain data with the same amount of data and the same division ratio but different band combinations, or data with the same band combination but after Gaussian stretching. Step 3.3: Use the improved Segformer network model obtained in step 2 to conduct experimental analysis on the data obtained in step 3.2 with the same data volume, the same division ratio but different band combinations, or the same band combination but after Gaussian stretching. By comparing the loss iteration curves during the experiment of the improved Segformer network model, select the optimal band combination image and obtain the training set, test set, and validation set image data sources of the optimal band combination.
4. The surface water extraction method based on the improved Segformer network and remote sensing imagery according to claim 1, characterized in that: The specific method of step 4 is: Step 4.1: Use the Bayesian optimization method to obtain the optimal hyperparameters of the improved Segformer network model established in step 2, and obtain the improved Segformer water extraction network model with the optimal hyperparameters; Step 4.2: Use the training set, test set, and validation set image data source of the optimal band combination obtained in step 3.3 to train the improved Segformer water body extraction network model with optimal hyperparameters obtained in step 4.
1. During the training process, use the training set, test set, and validation set image data source of the optimal band combination obtained in step 3.3 to evaluate the current performance of the improved Segformer water body extraction network model with optimal hyperparameters obtained in step 4.
1. After the training is completed, the trained improved Segformer water body extraction network model is obtained.
5. The surface water extraction method based on the improved Segformer network and remote sensing imagery according to claim 1, characterized in that: The calculation formulas for the precision, recall, and F1-score evaluation indicators in step 5 are as follows: Among them, the meanings of the parameters in Formula 1, Formula 2, Formula 3, and Formula 4 are as follows: TP (True Positive): The prediction is correct, the prediction result is positive, and the actual result is positive; FP (False Positive): The prediction is wrong, the prediction result is positive, and the actual result is negative; FN (False Negative): The prediction is wrong, the prediction result is negative, and the actual result is positive; TN (True Negative): The prediction is correct, the prediction result is negative, and the actual result is negative; In Formula 5, P refers to Precision, R refers to Recall, and F1 refers to the balanced F score.
6. The surface water extraction method based on an improved Segformer network and remote sensing images according to claim 1, characterized in that: The expression for evaluating the water body width extraction performance in step 6 is: N i =GETNUM(Latitude[1]:Latitude[n]), i∈(Longitude[1],Longitude[m]) Equation 6 Width i = N i *resolution Equation 7 In Equation 6, n and m represent the number of pixels on the longitude and latitude lines in the study area, respectively. GETNUM represents the operation of counting the total number of water body points. N i Represents the number of longitude river pixels at longitude i; In formula 7, resolution is the resolution of the remote sensing image, Width i is the actual river width obtained by calculation; In formula 8 and formula 9, PreWidth i and Width i They represent the predicted river width and the calculated actual river width at longitude i, MeanWidth represents the average river width, Longitude[i] represents the i-th longitude, and R 2 stands for coefficient of determination, and RMSE stands for mean square error.
7. A surface water extraction system based on an improved Segformer network and remote sensing images based on the method of claim 1, characterized in that: include: Data processing module: used to crop a complete remote sensing image data to be segmented into m×n pixels to produce a segmentation data set; Model building module: used to build an improved Segformer network model; Data analysis module: used to analyze the optimal band combination of remote sensing image data to be segmented; Water body segmentation module: used to input the remote sensing image to be segmented with the optimal band combination into the improved Segformer network model for water body extraction; Result evaluation module: used to evaluate the performance of the trained improved Segformer water extraction network model using precision, recall, F1-score, and mean intersection over Union (mIoU) as indicators; and used to combine the coefficient of determination (R 2 ) and root mean square error (RMSE) are used to evaluate the performance of the trained improved Segformer water body extraction network model for water body width extraction.
8. A surface water extraction device based on an improved Segformer network and remote sensing images, characterized by: include: A memory for storing a computer program for implementing the surface water body extraction method based on an improved Segformer network and remote sensing images as described in any one of claims 1 to 6; A processor is used to implement the surface water body extraction method based on an improved Segformer network and remote sensing images as described in any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement the surface water body extraction method based on an improved Segformer network and remote sensing images as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Automatic extraction method of water bodies based on landsat OLI multispectral remote sensing images
CN104915954B
A remote sensing image water body extraction method and system based on deep learning
CN109934095A
Remote sensing image water body extraction system based on deep learning
CN111985372A