A Remote Sensing Image Land-Sea Segmentation Method Based on Pyramid Pooling U-Net
Through the pyramid-pooled U-shaped network, learning multi-scale remote sensing image features in U-net, the problem of coastline missegment in traditional methods is solved, and high-precision sea-land segmentation and coastline information extraction are achieved.
Patent Information
- Application Number
- CN202111083141.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-09-15
AI Technical Summary
Traditional remote sensing image sea-land segmentation methods tend to ignore the semantic relationships up and down of the coastline, resulting in missegment and difficulty in obtaining coastline information accurately.
Using a pyramid-pooled U-shaped network, the segmentation accuracy is improved by learning the ocean and terrestrial features of multi-scale remote sensing images in the jump connection of U-net.
It realizes highly consistent sea and land segmentation with expert manual segmentation, improving the segmentation accuracy of high-resolution remote sensing images and the accuracy of coastline information extraction.
Smart Images

Figure CN114266795B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image segmentation, and particularly to a method for remote sensing image land-sea segmentation based on a pyramid pooling U-shaped network. Background Art
[0002] China is a major maritime country, and the coastline, as one of the very important landmarks, is the boundary line between the ocean and the land. With the rapid development of the ocean economy, the southern coastal areas of China have gradually become the main areas of human activities due to their geographical advantages. The coastline will change correspondingly due to external and human factors. For example, sea water erosion, silt discharge, beach reclamation, and sea sand collection will all lead to the expansion and contraction of the coastline. In recent years, China's remote sensing technology has made progress with the rapid development of the remote sensing satellite industry. The advantage of remote sensing technology is that it is not affected by surface changes, weather differences, and geographical environments, so it has been widely applied in the ocean development industry. High-resolution remote sensing images are beneficial for people to obtain image information, extract image features, and interpret images because of their high clarity. Among them, image semantic segmentation plays a key role in the application of remote sensing images. Especially the segmentation of the ocean and the land can accurately obtain coastline information, which is of great significance for the dynamic changes of the coast and the extraction of important information.
[0003] With the continuous development of artificial intelligence technology, machine learning methods have been widely applied in various fields and have become the research focus and hot issue of image semantic segmentation. The convolutional neural network (CNN) has achieved remarkable results in the field of remote sensing image processing with its great advantages. High-resolution remote sensing images have good imaging quality and high clarity, which are of great significance for detecting coastline changes and the macroscopic change trend of the shore beach. Extracting coastline information from remote sensing images is of great significance for the development around the ocean, and usually, the coastline is extracted by segmenting ocean and land images. However, traditional methods are prone to ignoring the semantic relationship above and below the coastline when performing remote sensing image land-sea segmentation, thus obtaining a wrong feature discrimination mechanism, resulting in difficulty in distinguishing seawater with a high sediment concentration, other coastal waters, aquaculture ponds, etc. The existing methods for high-resolution remote sensing image segmentation mainly include threshold segmentation method, edge detection method, wavelet transform method, region growing method, and machine learning algorithm. Most traditional algorithms perform image segmentation based on the principle of pixel value difference in remote sensing images, but it is easy to have missegmentation based only on the pixel theory. Traditional machine learning algorithms distinguish the ocean and the land in the form of features, but it is also very difficult to obtain more accurate coastline information for remote sensing images with unclear upper and lower semantic features. Summary of the Invention
[0004] The object of the present invention is to provide a method for remote sensing image sea-land segmentation based on a pyramid pooling U-shaped network. By adding pyramid pooling to the skip connections of the U-net to learn the features of the sea and land in multi-scale remote sensing images, the problem of blurred boundaries is solved, and the sea-land segmentation accuracy of high-resolution remote sensing images is improved, so as to solve the problems raised in the above-mentioned background technology.
[0005] The present invention is realized through the following technical solutions: The present invention discloses a method for remote sensing image sea-land segmentation based on a pyramid pooling U-shaped network, adding pyramid pooling to the skip connections of the U-net to learn the features of the sea and land in multi-scale remote sensing images. The method includes the following steps:
[0006] Obtain high-resolution remote sensing images, crop the high-resolution remote sensing images, and draw the corresponding true sea-land segmentation map.
[0007] Block and perform image rigid transformation on the cropped high-resolution remote sensing images in sequence, and divide the training set and the test set based on the transformation results.
[0008] Build a pyramid U-shaped convolutional neural network, input the data in the training set into the pyramid U-shaped convolutional neural network for learning and training to obtain a high-resolution remote sensing image sea-land segmentation model.
[0009] Input the data in the test set into the pyramid U-shaped convolutional neural network to obtain the sea-land segmentation result of the remote sensing image.
[0010] Optionally, crop the high-resolution remote sensing images, and the cropped images contain the area near the coastline and all information of the land.
[0011] Optionally, when drawing the corresponding true sea-land segmentation map, the process includes: Based on the ArcGIS10.2 tool, manually draw the sea and land areas in the cropped high-resolution remote sensing image to obtain a vector file in shp format composed of points, lines and surfaces as the true map.
[0012] Optionally, when performing block and image rigid transformation on the cropped high-resolution remote sensing images in sequence, the process includes:
[0013] Perform block processing on the cropped high-resolution remote sensing images, and the block size is N×N, where N is a natural number not exceeding 256.
[0014] Flip the blocked images up and down, left and right, and rotate by a certain angle to expand the sample size.
[0015] Optionally, when inputting the data in the training set into the pyramid U-shaped convolutional neural network for learning and training to obtain a high-resolution remote sensing image sea-land segmentation probability map, the process includes:
[0016] Set \(A = \{A_1, A_2, \ldots, A i \}\) contains all high - resolution remote sensing image training datasets where \(d m , d n represents the size of sample \(A i \);
[0017] Input the training set \(A i into the first layer of the pyramid U - shaped convolutional neural network for convolution to obtain feature \(E_1\), and input the feature \(E_1\) into the pooling layer of the pyramid U - shaped convolutional neural network for downsampling to obtain feature \(F_1\);
[0018] Input the feature \(E_1\) into the pyramid pooling module of the pyramid U - shaped convolutional neural network to obtain feature \(P_1\);
[0019] Convolve the feature \(F_1\) to obtain feature \(E_2\), and at the same time input the feature \(E_2\) into the pooling layer of the pyramid U - shaped convolutional neural network for downsampling to obtain feature \(F_2\);
[0020] Input the feature \(E_2\) into the pyramid pooling module of the pyramid U - shaped convolutional neural network to obtain feature \(P_2\);
[0021] Convolve the feature \(F_2\) to obtain feature \(E_3\), and at the same time input the feature \(E_3\) into the pooling layer of the pyramid U - shaped convolutional neural network for downsampling to obtain feature \(F_3\);
[0022] Input the feature \(F_3\) into the decoder of the pyramid U - shaped convolutional neural network for bilinear upsampling to obtain feature \(D_3\);
[0023] Cascade - fuse the feature \(E_3\), feature \(P_2\), and feature \(D_3\) to obtain feature \(C_1\), and convolve the feature \(C_1\), and input the convolution result into the decoder for bilinear upsampling to obtain feature \(D_2\);
[0024] Cascade - fuse the feature \(E_2\), feature \(P_1\), and feature \(D_2\) to obtain feature \(C_2\), and convolve the feature \(C_2\), and input the convolution result into the decoder for bilinear upsampling to obtain feature \(D_1\);
[0025] Cascade - fuse the feature \(E_1\) and the feature \(D_1\) to obtain feature \(C_3\), and convolve the feature \(C_3\) to obtain a high - resolution remote sensing image land - sea segmentation model.
[0026] Optionally, the pyramid pooling module includes four levels, where
[0027] The first level uses global pooling to generate a single output;
[0028] At the second level, the feature map is divided into 2×2 sub-regions, pooling is performed on each sub-region, and finally the outputs containing position information are combined;
[0029] At the third level, the feature map is divided into 3×3 sub-regions, pooling is performed on each sub-region, and finally the outputs containing position information are combined;
[0030] At the fourth level, the feature map is divided into 6×6 sub-regions, pooling is performed on each sub-region, and finally the outputs containing position information are combined.
[0031] Optionally, the method further includes:
[0032] Comparing the land-sea segmentation probability map of the high-resolution remote sensing image with the ground truth map. If the similarity is high, it indicates that the land-sea segmentation probability map of the high-resolution remote sensing image is correct and the training of the pyramid U-shaped convolutional neural network is completed. Otherwise, the data in the training set is re-input into the pyramid U-shaped convolutional neural network for learning and training.
[0033] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0034] A method for land-sea segmentation of ocean remote sensing images based on a pyramid pooling U-shaped network provided by the present invention can achieve a high degree of consistency with manual segmentation by experts in land-sea segmentation of high-resolution remote sensing images through the pyramid U-shaped network; embedding the pyramid pooling structure into the U-shaped network, performing two pooling operations and combining feature maps of different scales, improving the segmentation accuracy of different scales of high-resolution remote sensing images; adding a deep supervision function in the decoder stage to aggregate hierarchical representations learned in the feature map and improve the accuracy of coastline information extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only the preferred embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0036] Figure 1 It is a flowchart of a method for land-sea segmentation of remote sensing images based on a pyramid pooling U-shaped network provided by the present invention;
[0037] Figure 2 It is a structure diagram of a pyramid U-shaped convolutional neural network provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] To make the objectives, technical solutions and advantages of the present invention more apparent, exemplary embodiments according to the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein. Based on the embodiments of the present invention described herein, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0039] In the following description, numerous specific details are given to provide a more thorough understanding of the present invention. However, it will be apparent to one of ordinary skill in the art that the present invention may be practiced without one or more of these details. In other instances, some well-known technical features are not described in order to avoid obscuring the present invention.
[0040] It should be understood that the present invention can be implemented in different forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to make the disclosure thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.
[0041] The purpose of the terms used herein is only to describe specific embodiments and is not a limitation of the present invention. When used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the terms "comprising" and / or "including", when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups. When used herein, the term "and / or" includes any and all combinations of the related listed items.
[0042] To thoroughly understand the present invention, detailed structures will be presented in the following description to illustrate the technical solutions proposed by the present invention. The optional embodiments of the present invention are described in detail below. However, in addition to these detailed descriptions, the present invention may also have other embodiments.
[0043] The objective of the present invention is to propose a method for sea-land segmentation of ocean remote sensing images using a pyramid pooling U-shaped network, that is, adding pyramid pooling to the skip connections of U-net to learn the features of the ocean and land in multi-scale remote sensing images, solve the problem of blurred boundaries, and improve the sea-land segmentation accuracy of high-resolution remote sensing images. See Figures 1 to 2 , which includes the following steps:
[0044] S1. Obtain high-resolution remote sensing images, crop the high-resolution remote sensing images, and draw the corresponding sea-land segmentation ground truth map;
[0045] S2. Sequentially perform blocking and image rigid transformation on the cropped high-resolution remote sensing image, and divide the training set and test set based on the transformation results;
[0046] S3. Establish a pyramid U-shaped convolutional neural network, and input the data in the training set into the pyramid U-shaped convolutional neural network for learning and training to obtain a high-resolution remote sensing image land-sea segmentation model;
[0047] S4. Input the data in the test set into the pyramid U-shaped convolutional neural network to obtain the land-sea segmentation result of the remote sensing image.
[0048] In this embodiment, the collected high-resolution remote sensing image is a false-color image of the coastal area within the South China Sea taken by the GF-1 satellite. In step S1, since the shape of the high-resolution remote sensing image taken by the satellite is irregular, which increases the difficulty of coastline extraction, the collected remote sensing image is cropped. The cropped image contains all the information of the area near the coastline and the land.
[0049] Furthermore, based on the ArcGIS 10.2 tool, manually draw the ocean and land areas in the cropped high-resolution remote sensing image to obtain a vector file in shp format composed of points, lines, and surfaces as the ground truth map.
[0050] In step S2, due to the high resolution of the high-resolution remote sensing image and the large image size, the cropped high-resolution remote sensing image is block-processed. The block size is N×N, where N is a natural number not exceeding 256. For example, N = 256;
[0051] Deep learning requires a large number of training samples. Therefore, the block-processed images are flipped up and down, left and right, and rotated by a certain angle to expand the sample size;
[0052] Finally, the block-processed and expanded sample data is divided into a training set and a test set according to a certain ratio. The ratio of the training set to the test set is 4:1.
[0053] In step S3, the present invention further discloses a training method for inputting the data in the training set into the pyramid U-shaped convolutional neural network for learning and training to obtain a high-resolution remote sensing image land-sea segmentation probability map. The process includes:
[0054] S301. Set A = {A1, A2,..., A i} to include all high-resolution remote sensing image training data sets where d m , d n represents the size of sample A i ;
[0055] S302. Input the training set A iThe first layer of the input pyramid U-shaped convolutional neural network performs m×m convolution operations twice and follows a ReLU function with a decay rate of 0.85, where m is 3, to obtain feature E1. The feature E1 is input into the pooling layer of the pyramid U-shaped convolutional neural network for downsampling to obtain feature F1;
[0056] S303. Input the feature E1 into the pyramid pooling module of the pyramid U-shaped convolutional neural network to obtain feature P1;
[0057] S304. Perform m×m convolution twice on the feature F1 and follow a ReLU function with a decay rate of 0.85, where m is 3, to obtain feature E2. At the same time, input the feature E2 into the pooling layer of the pyramid U-shaped convolutional neural network for downsampling to obtain feature F2;
[0058] S305. Input the feature E2 into the pyramid pooling module of the pyramid U-shaped convolutional neural network to obtain feature P2;
[0059] S306. Perform m×m convolution twice on the feature F2 and follow a ReLU function with a decay rate of 0.85, where m is 3, to obtain feature E3. At the same time, input the feature E3 into the pooling layer of the pyramid U-shaped convolutional neural network for downsampling to obtain feature F3;
[0060] S307. Input the feature F3 into the decoder of the pyramid U-shaped convolutional neural network for bilinear upsampling, and perform deep supervision at the last layer of the decoder to obtain feature D3;
[0061] In S307, deep supervision is to feed the last layer of the decoder into an m×m convolutional layer. The cross-entropy loss calculated in terms of pixels at this stage is: l1 = ∑ x∈Ω ω(x)logp(x), where m is 3;
[0062] S308. Concatenate and fuse the feature E3, feature P2, and feature D3 to obtain feature C1, and perform m×m convolution twice on the feature C1 and follow a ReLU function with a decay rate of 0.85, where m is 3. Finally, input the convolution result into the decoder for bilinear upsampling, and perform deep supervision at the last layer of the decoder to obtain feature D2;
[0063] In S308, deep supervision is to feed the last layer of the decoder into an m×m convolutional layer. The cross-entropy loss calculated in terms of pixels at this stage is: l2 = ∑ x∈Ω ω(x)logp(x), where m is 3;
[0064] S309. Cascade and fuse the feature E2, feature P1, and feature D2 to obtain feature C2. Then perform a convolution of m×m on the feature C2, followed by a ReLU function with a decay rate of 0.85. Here, m is 3. Finally, input the convolution result into the decoder for bilinear upsampling, and perform deep supervision at the last layer of the decoder to obtain feature D1;
[0065] In S309, the deep supervision is to feed the last layer of the decoder into an m×m convolutional layer. The cross-entropy loss calculated in terms of pixels at this stage is: l3 = ∑ x∈Ω ω(x)logp(x), where m is 3;
[0066] S310. Cascade and fuse the feature E1 and the feature D1 to obtain feature C3. Perform two convolutions of m×m on the feature C3, followed by a ReLU function with a decay rate of 0.85, to obtain the probability map for land-sea segmentation of the high-resolution remote sensing image.
[0067] Furthermore, in this embodiment, the pyramid pooling module includes four hierarchical features of different scales. Among them,
[0068] For the first level, use global pooling to generate a single output;
[0069] For the second level, divide the feature map into 2×2 sub-regions, perform pooling on each sub-region, and finally combine the output containing position information;
[0070] For the third level, divide the feature map into 3×3 sub-regions, perform pooling on each sub-region, and finally combine the output containing position information;
[0071] For the fourth level, divide the feature map into 6×6 sub-regions, perform pooling on each sub-region, and finally combine the output containing position information.
[0072] Furthermore, in this embodiment, the Dice function is used as part of the loss function to solve the problem of class imbalance, and combined with the three loss functions in the deep supervision process. The final loss function is expressed as: Loss = L(Θ) + DL(X) = αl1 + βl2 + γl3 + DL(X), where L(Θ) = αl1 + βl2 + γl3,
[0073] Optionally, the method further includes:
[0074] S311. Compare the probability map for land-sea segmentation of the high-resolution remote sensing image with the ground truth map. If the similarity is high, it indicates that the probability map for land-sea segmentation of the high-resolution remote sensing image is correct, and the training of the pyramid U-shaped convolutional neural network is completed. Otherwise, re-input the data in the training set into the pyramid U-shaped convolutional neural network for learning and training.
[0075] In step S4, the data in the test set is input into the trained pyramid U-shaped convolutional neural network to obtain the final high-resolution remote sensing image land-sea segmentation probability map.
[0076] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for remote sensing image land-sea segmentation based on a pyramid pooling U-shaped network, characterized in that, Add pyramid pooling to the skip connections of U-net to learn the features of ocean and land in multi-scale remote sensing images. The method includes the following steps: Obtain high-resolution remote sensing images, crop the high-resolution remote sensing images, and draw the corresponding true value map of sea-land segmentation; Block and perform image rigid transformation on the cropped high-resolution remote sensing images in sequence, and divide the training set and test set based on the transformation results; Build a pyramid U-shaped convolutional neural network, input the data in the training set into the pyramid U-shaped convolutional neural network for learning and training to obtain a high-resolution remote sensing image sea-land segmentation model; Input the data in the test set into the pyramid U-shaped convolutional neural network to obtain the sea-land segmentation result of the remote sensing image; Crop the high-resolution remote sensing image, and the cropped image includes the area near the coastline and all information of the land; When drawing the corresponding true value map of sea-land segmentation, the process includes: based on the ArcGIS 10.2 tool, manually draw the ocean and land areas in the cropped high-resolution remote sensing image to obtain a vector file in shp format composed of points, lines and surfaces as the true value map; When performing block and image rigid transformation on the cropped high-resolution remote sensing image in sequence, the process includes: Perform block processing on the cropped high-resolution remote sensing image, and the block size is N×N, where N is a natural number not exceeding 256; Flip the blocked image up and down, left and right, and rotate it by a certain angle to expand the sample size; When inputting the data in the training set into the pyramid U-shaped convolutional neural network for learning and training to obtain a high-resolution remote sensing image sea-land segmentation model, the process includes: Set \(A = \{A_1, A_2, \ldots, A\) i \} contains all high - resolution remote sensing image training datasets where \(d\) m , \(d\) n represents the size of sample \(A\) i ; Input the training set A i into the first layer of the Pyramid U-shaped convolutional neural network for convolution to obtain the feature E1, and input the feature E1 into the pooling layer of the Pyramid U-shaped convolutional neural network for downsampling to obtain the feature F1; Input the feature E1 into the pyramid pooling module of the pyramid U-shaped convolutional neural network to obtain the feature P1; Convolve the feature F1 to obtain the feature E2, and at the same time input the feature E2 into the pooling layer of the pyramid U-shaped convolutional neural network for downsampling to obtain the feature F2; Input the feature E2 into the pyramid pooling module of the pyramid U-shaped convolutional neural network to obtain the feature P2; Convolve the feature F2 to obtain the feature E3, and at the same time input the feature E3 into the pooling layer of the pyramid U-shaped convolutional neural network for downsampling to obtain the feature F3; Input the feature F3 into the decoder of the pyramid U-shaped convolutional neural network for bilinear upsampling to obtain the feature D3; Cascade and fuse the feature E3, feature P2, and feature D3 to obtain the feature C1, convolve the feature C1, and input the convolution result into the decoder for bilinear upsampling to obtain the feature D2; Cascade and fuse the feature E2, feature P1, and feature D2 to obtain the feature C2, convolve the feature C2, and input the convolution result into the decoder for bilinear upsampling to obtain the feature D1; Cascade and fuse the feature E1 and the feature D1 to obtain the feature C3, convolve the feature C3 to obtain a high-resolution remote sensing image sea-land segmentation model.
2. The method for remote sensing image land-sea segmentation based on a pyramid pooling U-shaped network according to claim 1, wherein The pyramid pooling module includes four levels, where, The first level uses global pooling to generate a single output; At the second level, divide the feature map into 2×2 sub-regions, perform pooling on each sub-region, and finally combine the output containing location information; At the third level, divide the feature map into 3×3 sub-regions, perform pooling on each sub-region, and finally combine the output containing location information; At the fourth level, divide the feature map into 6×6 sub-regions, perform pooling on each sub-region, and finally combine the output containing location information.
3. The method for remote sensing image land-sea segmentation based on a pyramid pooling U-shaped network according to claim 2, characterized in that, The method further includes: Compare the land-sea segmentation probability map of the high-resolution remote sensing image with the ground truth map. If the similarity is high, it indicates that the land-sea segmentation probability map of the high-resolution remote sensing image is correct and the training of the Pyramid U-shaped convolutional neural network is completed. Otherwise, re-enter the data in the training set into the Pyramid U-shaped convolutional neural network for learning and training.
Citation Information
Patent Citations
Remote sensing image coastline extraction method based on deep semantic segmentation network
CN113139550A
Remote sensing image water body region extraction method based on neural network model
CN113343861A