Deep learning method and system for extracting polar sea ice using multispectral remote sensing images
By introducing feature map reconstruction module in multispectral remote sensing image processing, the problem of accuracy and efficiency limitations in sea ice extraction of CNN models is solved, and more efficient sea ice monitoring is achieved.
Patent Information
- Application Number
- CN202411907401.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-12-24
AI Technical Summary
The existing multispectral remote sensing image processing methods have accuracy and efficiency limitations in polar sea ice extraction. In particular, the local receptive field of the CNN model is limited, which cannot effectively capture long-distance spatial relationships, and requires a large number of labeled data sets for training, which is time-consuming and labor-intensive for data acquisition and labeling.
Deep learning method based on feature map reconstruction is adopted, and the depth features extracted by CNN are graphically reconstructed through the graph reconstruction module, capturing the long-distance dependence between pixels and supplementing global context information, and improving the sea ice extraction accuracy. At the same time, the semi-automatic labeling of sea ice is completed through simple image processing technology, reducing the difficulty and time of obtaining label data.
It improves the accuracy and efficiency of sea ice extraction, reduces model deployment parameters, simplifies the acquisition process of sea ice labeled data, and achieves faster and more efficient sea ice monitoring.
Smart Images

Figure CN119339250B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multispectral remote sensing image processing, and specifically relates to a deep learning method and system for extracting polar sea ice using multispectral remote sensing images based on feature map reconstruction. Background Art
[0002] It is of great significance to monitor polar sea ice and study its distribution and dynamic changes. The amount of data and spatial range obtained by traditional monitoring methods are very limited, and it is impossible to monitor the changes of sea ice in real time and on a large scale. Remote sensing technology obtains information on the earth's surface through sensors on satellites or aircraft. It has the advantages of wide range, long time series and fast acquisition. It has become the main means of sea ice monitoring. Among them, multispectral remote sensing records the reflection or radiation information of surface objects in different bands through sensors, and uses information from multiple bands to synthesize the required images, which can clearly reflect the distribution of polar sea ice. At the same time, compared with other remote sensing methods, multispectral remote sensing usually provides higher spatial resolution, lower cost and shorter revisit time. Therefore, the use of multispectral remote sensing images to extract sea ice has great research value.
[0003] At present, the mainstream methods for extracting sea ice based on multispectral remote sensing images can be divided into two categories: model-driven and data-driven. Model-driven mainly uses image processing technology to segment and extract sea ice based on the image features of sea ice itself, such as superpixel segmentation based on image texture, snake segmentation based on gradient vector flow (GVF), watershed transform segmentation based on Sobel filter, etc. Although these methods have achieved good extraction effects in certain specific scenarios, they usually require the design of complex image features, which have a great impact on the final extraction results.
[0004] At present, some multispectral satellites are able to provide massive amounts of remote sensing data, which makes data-driven sea ice extraction methods mainstream. Among them, the most typical is the use of convolutional neural networks (CNNs) to segment and extract sea ice in multispectral remote sensing images. However, the local receptive field of CNN is limited, and it cannot effectively capture long-distance spatial relationships, represent geographic objects and their topological relationships, thereby limiting the accuracy and efficiency of the extraction process. Some researchers have made improvements by increasing the number of network layers or adding improved modules, such as multi-scale pooling, hole convolution, and introducing attention mechanisms, but this inevitably introduces a huge number of model parameters, which seriously affects the reasoning speed of the model. At the same time, another important issue is that deep learning requires a large number of annotated data sets for training. There are currently few public data on sea ice multispectral remote sensing images available, and collecting and annotating sea ice remote sensing images is a time-consuming and labor-intensive task. Summary of the invention
[0005] In order to solve the problems existing in the above background, the present invention proposes a deep learning method and system for extracting polar sea ice using multispectral remote sensing images based on feature graph reconstruction. Taking Sentinel-2 data as an example, the main steps of the present invention include: preprocessing the original Sentinel-2 data to obtain TCI, and completing the sea ice annotation of the TCI synthesized by multispectral remote sensing through simple image processing technology. The present invention proposes to use graph reconstruction to reconstruct the deep features extracted by CNN, so as to capture the long-distance dependencies between pixels and supplement the global context information, so as to improve the extraction accuracy of sea ice after the final upsampling.
[0006] The specific scheme of the present invention is as follows:
[0007] A deep learning method for extracting polar sea ice using multispectral remote sensing images. The main steps are:
[0008] S1. Obtain image data of polar sea ice areas;
[0009] S2. Preprocess and annotate the image data obtained in step S1 to obtain a true color image TCI of sea ice and annotated images to form a data set;
[0010] S3. Use the deep learning model based on feature map reconstruction to train and test the data set obtained in step S2;
[0011] S4. The optimal model parameters are obtained through tuning, and the optimal model parameters are loaded into the deep learning model. The deep learning model inputs the true color image TCI of sea ice to extract the sea ice and obtain the output result.
[0012] Preferably, in step S1, the image data is Sentinel-2 data.
[0013] Preferably, in step S2, the original image (such as Sentinel-2) is synthesized by the 2nd (Blue), 3rd (Green), and 4th (Red) bands, 2% linear stretching operation, and non-overlap cropping to obtain a TCI sample of 512 pixels × 512 pixels; then a simple image processing technique is used to complete the semi-automatic annotation of sea ice in the TCI: first, the TCI is converted from the RGB space to the LAB space and the L channel is taken as the reference image, and then the image is enhanced by adaptive equalization, mean filtering, morphological operations, etc., and finally the image is labeled using threshold segmentation. For some complex scenes, manual correction can be performed.
[0014] Preferably, in step S2, the reference image is selected by converting the TCI from the RGB space to the LAB space and taking the L channel as the reference image, specifically as follows: For a bit RGB image, the value range of the three channels R, G, and B is , convert the image from RGB space to XYZ space, and then convert from XYZ space to LAB space. The specific calculation process is as follows:
[0015] ;
[0016] ;
[0017] in, Indicates the color depth, Represents the value of the L channel of the image in the LAB space, Represents the value of the Y channel of the image in the XYZ space. is a nonlinear mapping function, and the specific calculation formula is:
[0018] .
[0019] Preferably, in step S3, the deep learning model is mainly composed of an encoder module, a graph reconstruction module and a decoder module. The encoder module is composed of convolution blocks and residual blocks, and is responsible for extracting multi-layer features of the input image TCI; the graph reconstruction module is responsible for reconstructing the features extracted by the encoder module to capture the contextual information in the features; the decoder module is responsible for upsampling the reconstructed features layer by layer and outputting the final extraction results.
[0020] The graph reconstruction module reconstructs the deep features extracted by the encoder module. It consists of two stages: graph construction stage and feature reconstruction stage, as follows:
[0021] In the graph construction phase, it is assumed that the input features ,in, represents a real number, is the number of channels of the feature, Represents the height and width of the feature, Flatten the node features ,in , and generate the corresponding adjacency matrix When constructing the adjacency matrix, the global average adaptive similarity (GAAST) threshold is used to determine the adjacency relationship between each node. Specifically, for node features First, normalize each feature of the node. The inner product of the similarity matrix , and then construct a threshold based on the global average similarity to determine the adjacency relationship between each node. It can be expressed as:
[0022] ;
[0023] ;
[0024] in, ( ) indicates the Nodes and The adjacency relationship between nodes, It is a learnable weight parameter with an initial value of 1. It automatically finds the best adjacency matrix construction solution through data training optimization.
[0025] After obtaining the adjacency matrix, perform the following operations:
[0026] ;
[0027] ;
[0028] in, The adjacency matrix representing its own data dependency, express The elements in the degree matrix of express Middle Row, No. The elements of the column, , is the identity matrix.
[0029] In the feature reconstruction stage, GCN is used to update the features in the node, specifically:
[0030] ;
[0031] in, Indicates layer, represents the learnable weight parameters, represents the activation function, for degree matrix; get the updated graph node features , for the above formula represents the graph node features before updating, represents the updated graph node features, that is, = , and restore it to the original input feature size to obtain the reconstructed feature .
[0032] In the feature reconstruction stage, the updated node features are restored to the same shape as the input features to obtain the reconstructed features. Since the number of nodes depends on the size of the input features, when the number of nodes is too large, it will take up more computing resources. Therefore, for larger feature maps, a parameter-free adaptive average pooling operation is used to compress the feature map to limit its size. At the same time, a parameter-free bilinear interpolation upsampling operation is used to restore it to its original size after feature reconstruction.
[0033] The present invention also discloses a deep learning system for extracting polar sea ice using multispectral remote sensing images, which is used to execute the above method and includes the following modules:
[0034] Data acquisition module: obtain image data of polar sea ice areas;
[0035] Dataset formation module: preprocesses and annotates the acquired image data to obtain the true color image TCI of sea ice and annotated images to form a data set;
[0036] Training module: Use the deep learning model based on feature map reconstruction to train and test the formed data set;
[0037] Result output module: The optimal model parameters are obtained through tuning, and the optimal model parameters are loaded into the deep learning model, and TCI is input to extract sea ice to obtain output results.
[0038] Compared with the existing methods, the present invention uses simple image processing technology to complete the semi-automatic annotation of sea ice in TCI. Compared with manual annotation, it effectively reduces the difficulty of obtaining sea ice annotation data and reduces the annotation time. At the same time, the present invention proposes a graph reconstruction method, which can effectively solve the limitations of traditional models applied to remote sensing images, and has the advantage of fewer deployment parameters while improving the extraction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flow chart of a deep learning method for extracting polar sea ice using multispectral remote sensing images according to a preferred embodiment of the present invention;
[0040] Figure 2 It is a specific flow chart of the preferred embodiment of the present invention in the pre-processing stage;
[0041] Figure 3 It is a structural diagram of a preferred embodiment model of the preferred embodiment of the present invention;
[0042] Figure 4 is a specific structural diagram of an encoder module and a decoder module in a preferred embodiment of the present invention;
[0043] Figure 5It is a specific structural diagram of the reconstruction module of the preferred embodiment of the present invention;
[0044] Figure 6 is a detailed step diagram of extracting sea ice from TCI using a model loaded with optimal parameters in an embodiment of the present invention;
[0045] Figure 7 It is a block diagram of a deep learning system for extracting polar sea ice using multispectral remote sensing images according to a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0046] The implementation process of the present invention is described below through preferred embodiments. Those skilled in the art can easily understand the advantages of the present invention through the contents disclosed in this specification.
[0047] The following article introduces in detail the specific model for improvement and compares it with some current mainstream deep learning models to further demonstrate the superiority of the present invention. In terms of data, the preferred embodiment of the present invention uses Sentinel-2 L1C data near the Ross Sea in Antarctica for detailed description. At the same time, pixel accuracy (PA), intersection over union (IoU), and F1 score (F1-Score) are used as evaluation indicators to evaluate the model. At the same time, in order to illustrate the practicality of the present invention, this method is applied to the production of a specific ice map.
[0048] This embodiment provides a deep learning method for extracting polar sea ice using Sentinel-2 data based on feature graph reconstruction. This method uses simple image processing technology to annotate the synthesized sea ice TCI to form a data set. At the same time, it is improved based on U-net and introduces a graph reconstruction module to achieve the best extraction results while maintaining fewer learning parameters. The specific process is as follows Figure 1 The detailed operation steps are as follows:
[0049] S1. Get Sentinel-2 L1C data near the Ross Sea in Antarctica for free through the ESA website.
[0050] S2. Figure 2As shown in the figure, the acquired Sentinel-2 data is synthesized into 2 (Blue), 3 (Green), and 4 (Red) bands, linearly stretched by 2%, and cropped without overlap to obtain a TCI sample of 512 pixels × 512 pixels. The above operations can be completed using ENVI software. Furthermore, based on the OpenCV library in Python, simple image processing technology is used to complete the semi-automatic annotation of sea ice in TCI: First, the TCI is converted from RGB space to LAB space and the L channel is taken as the reference image. The specific formula is:
[0051] ;
[0052] ;
[0053] in Indicates the color depth, Represents the value of the L channel of the image in the LAB space, Represents the value of the Y channel of the image in the XYZ space. is a nonlinear mapping function, and the specific calculation formula is:
[0054] .
[0055] Then, the basic image is adaptively equalized, mean filtered, and morphological operations are performed to enhance the image. The specific parameters need to be manually fine-tuned. Finally, the image is segmented by threshold segmentation to separate the sea ice. For some complex scenes, manual correction can be performed. The TCI and the labeled samples are divided into training data sets and test data sets in a ratio of 7:3.
[0056] S3. Use the deep learning model based on feature map reconstruction to train and test the data. Specifically, the classic U-net model in deep learning is improved. The specific model structure diagram after improvement is as follows: Figure 3 As shown in Figure 2, the model mainly consists of three modules: encoder module, graph reconstruction module and decoder module.
[0057] The encoder module is mainly composed of convolution blocks and residual blocks, which are responsible for downsampling the image to obtain multi-scale features. In the convolution operation, k represents the convolution kernel size, p represents the edge padding size, and s represents the convolution kernel moving step. The decoder is mainly composed of parameter-free bilinear interpolation and convolution operations, which are responsible for upsampling the features, gradually restoring the feature size and outputting the final result. The specific structure of the encoder and decoder is as follows: Figure 4 shown.
[0058] In order to capture the long-distance dependencies between pixels in the deep features and supplement the global context information, a graph reconstruction module is introduced between the encoder module and the decoder module to reconstruct the deep features extracted by the encoder module. The specific structure is as follows: Figure 5 As shown in Figure 1, it includes two stages: graph construction stage and feature reconstruction stage. The details are as follows:
[0059] In the graph construction phase, it is assumed that the input features ,in, is the number of channels of the feature, Represents the height and width of the feature, Flatten the node features ,in , and generate the corresponding adjacency matrix When constructing the adjacency matrix, the global average adaptive similarity (GAAST) threshold is used to determine the adjacency relationship between each node. The specific refinement steps are as follows: First, normalize each feature of the node. The inner product of the similarity matrix , and then construct a threshold based on the global average similarity to determine the adjacency relationship between each node. It can be expressed as:
[0060] ;
[0061] ;
[0062] in represents the normalization operation, Indicates Nodes and The similarity between nodes, represents the operation of finding the average value, Indicates Nodes and The adjacency relationship between nodes, , It is a learnable weight parameter with an initial value of 1. It automatically finds the best adjacency matrix construction solution through data training optimization.
[0063] In getting the adjacency matrix Then do the following:
[0064] ;
[0065] ;
[0066] in, The adjacency matrix representing its own data dependency, express The elements in the degree matrix of express Middle Row, No. The elements of the column, , is the identity matrix.
[0067] In the feature reconstruction stage, GCN is used to update the features in the node, specifically:
[0068] ;
[0069] in, Indicates layer, represents the learnable weight parameters, represents the activation function, for degree matrix; get the updated graph node features , for the above formula represents the graph node features before updating, represents the updated graph node features, i.e. = , and restore it to the original input feature size to obtain the reconstructed feature .
[0070] The reconstructed features The input is sent to the decoder module for upsampling, and finally a convolution operation with a convolution kernel size of 1 is used to output the final extraction result.
[0071] Input the data of the training data set and the test data set in step S2 into the deep learning model to obtain the output result. Using Focal loss As the loss function during model training, the specific definition is:
[0072] ;
[0073] Among them, the coefficient Used to improve the problem of unbalanced sample size, the coefficient It is used to adjust the weight of the difference between the predicted result and the true value on the loss size. The larger the value, the greater the adjustment.
[0074] S4. The best model parameters are obtained through tuning, and the best model parameters are loaded into the deep learning model. Evaluation indicators such as pixel accuracy (PA), intersection over union (IoU), and F1-score (F1-Score) of the annotated image and the output image are calculated, and the above indicators are optimized by manually adjusting the learning rate, model parameter initialization and other hyperparameters. The parameters with the best evaluation indicators obtained during model testing are taken as the best model parameters, and the parameters are loaded into the model, and TCI is input for sea ice extraction to obtain the final results.
[0075] This embodiment realizes the extraction of sea ice using Sentinel-2 data and deep learning network based on feature map reconstruction. In the data preprocessing stage, simple image processing technology is used to complete the TCI labeling work, which effectively reduces the workload and working time compared with manual labeling. At the same time, a graph reconstruction module is proposed to reconstruct the deep features extracted by the encoder module, which effectively overcomes the limitations of CNN.
[0076] In summary, Table 1 is a comparison of the extraction results of the model of this embodiment with other existing mainstream models, and Table 2 is a comparison of the model parameters of the model of this embodiment with other existing mainstream models:
[0077] Table 1 Comparison of the results of the embodiment model and the mainstream model
[0078]
[0079] Table 2 Parameter comparison table of the embodiment model and the mainstream model
[0080]
[0081] According to the results shown in the above table, it can be concluded that compared with increasing the number of network layers and other improved modules, the graph reconstruction module proposed in the present invention is more effective and lighter.
[0082] After obtaining the optimal model parameters, the parameters can be loaded into the model to complete the extraction of other TCI sea ice, such as Figure 6 As shown, the specific steps are:
[0083] 1) Crop the entire TCI into sub-images with non-overlapping sliding windows of size 512 × 512 pixels;
[0084] 2) Input each sub-image into the model to complete the segmentation;
[0085] 3) Merge the segmented images in the order of cropping to obtain the extraction result of the entire TCI.
[0086] like Figure 7 As shown, this embodiment discloses a deep learning system for extracting polar sea ice using multispectral remote sensing images, which is used to execute the above method and includes the following modules:
[0087] Data acquisition module: obtain image data of polar sea ice areas;
[0088] Dataset formation module: preprocesses and annotates the acquired image data to obtain the true color image (TCI) of sea ice and annotated images to form a data set;
[0089] Training module: Use the deep learning model based on feature map reconstruction to train and test the obtained data set;
[0090] Result output module: The optimal model parameters are obtained through tuning, and the optimal model parameters are loaded into the deep learning model, and TCI is input to extract sea ice to obtain output results.
[0091] For other contents of this embodiment, please refer to the above method embodiment.
[0092] Although all the above embodiments effectively describe the present invention, those skilled in the art will appreciate that there are many variations of the present invention. For example, other embodiments similar to the spirit of the present invention can be easily proposed as long as multispectral data is involved, and they also fall within the protection scope of the present invention.
Claims
1. A deep learning method for extracting polar sea ice using multispectral remote sensing images, characterized in that: The specific steps include: S1. Obtain image data of polar sea ice areas; S2. Preprocess and annotate the image data obtained in step S1 to obtain a true color image TCI of sea ice and annotated images to form a data set; S3. Use the deep learning model based on feature map reconstruction to train and test the data set obtained in step S2; S4. Obtain optimal model parameters through tuning, and load the optimal model parameters into the deep learning model, and extract the sea ice by inputting the true color image TCI of the sea ice into the deep learning model to obtain an output result; In step S3, the deep learning model includes an encoder module, a graph reconstruction module and a decoder module, wherein: The encoder module consists of convolutional blocks and residual blocks to extract features from the input TCI; The graph reconstruction module is responsible for reconstructing the features extracted by the encoder module and capturing the contextual information in the features; The decoder module is responsible for upsampling the reconstructed features layer by layer and outputting the final extraction results; The processing of the graph reconstruction module is divided into the graph construction stage and the feature reconstruction stage. Specifically: In the construction stage, let the input feature X∈R C×H×W , where R represents a real number, C is the number of channels of the feature, H and W represent the height and width of the feature respectively, and X is flattened to obtain the node feature V∈R N×C , where N = H × W, and generate the corresponding adjacency matrix A, and perform the following operations on A: in, The adjacency matrix representing its own data dependency, express The elements in the degree matrix of express The element in the i-th row and j-th column, i,j∈1,2,3…N, I is the unit matrix; In the feature reconstruction stage, the graph convolutional network is used to update the features in the nodes, specifically: Where l represents the lth layer, w represents the learnable weight parameter, σ(·) represents the activation function, for The degree matrix, V (l) represents the graph node features before updating, V (l+1) represents the updated graph node features, and V (l+1) =V′; finally, the updated graph node feature V′ is obtained, and then restored to the original input feature size to obtain the reconstructed feature X′; In the graph construction phase, the adjacency matrix is constructed as follows: for node feature V∈R N×C First, normalize each feature of the node and obtain the similarity matrix Z∈R through the inner product of the node feature V N×N , and then construct a threshold based on the global average similarity to determine the adjacency relationship between each node; the formula is expressed as: Z=Norm(V)·Norm(V T ) Among them, Norm represents the normalization operation, z ij represents the similarity between the i-th node and the j-th node, mean represents the operation of finding the average value, A ij represents the adjacency relationship between the i-th node and the j-th node, i≠j, i,j∈1,2,3…N, α is a learnable weight parameter; In the feature reconstruction stage, the updated node features are restored to the same shape as the input features to obtain the reconstructed features.
2. The deep learning method for extracting polar sea ice using multispectral remote sensing images according to claim 1, characterized in that: In step S2, the preprocessing is as follows: a TCI sample with a size of 512 pixels×512 pixels is obtained by synthesizing the second, third, and fourth bands, performing a 2% linear stretching operation, and performing non-overlap cropping on the original image.
3. The deep learning method for extracting polar sea ice using multispectral remote sensing images according to claim 1 or 2, characterized in that: In step S2, the sea ice in TCI is labeled: first, the TCI is converted from RGB space to LAB space and the L channel is taken as the reference image, then image enhancement is performed, and finally threshold segmentation is used to complete the image labeling.
4. The deep learning method for extracting polar sea ice using multispectral remote sensing images according to claim 3, characterized in that: In step S2, the TCI is converted from the RGB space to the LAB space, and the L channel is taken as the reference image, as follows: For an RGB image with a color depth of n bits, the value range of the three channels R, G, and B is [0,2 n ], convert the image from RGB space to XYZ space, and then convert from XYZ space to LAB space. The specific calculation process is as follows: Where n represents the color depth, L LAB represents the value of the L channel of the image in the LAB space, Y represents the value of the Y channel of the image in the XYZ space, and f(·) is a nonlinear mapping function. The specific calculation formula is:
5. A deep learning system for extracting polar sea ice using multispectral remote sensing images, used to execute the method according to any one of claims 1 to 4, characterized in that: Includes the following modules: Data acquisition module: obtain image data of polar sea ice areas; Dataset formation module: preprocesses and annotates the acquired image data to obtain the true color image TCI of sea ice and annotated images to form a data set; Training module: Use the deep learning model based on feature map reconstruction to train and test the obtained data set; Result output module: The optimal model parameters are obtained through tuning, and the optimal model parameters are loaded into the deep learning model, and TCI is input to extract sea ice to obtain output results.
Citation Information
Patent Citations
Remote sensing image sea ice identification method based on depth U-Net model
CN112102324A
Fusion filtering multi-scale high-resolution remote sensing glacier extraction method
CN118135239A