Glacier disintegration leading edge extraction method based on edge and semantic feature fusion
By adopting a dual-branch supervision network based on the fusion of edge and semantic features in the frontier extraction of glacier disintegration, the problem of extraction result deviation in complex glacier scenarios is solved, and higher extraction accuracy and automation are achieved.
Patent Information
- Application Number
- CN202510091402.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art has deviations in extracting the frontier of glacier disintegration, especially in complex glacier scenarios, where the boundaries of geographic distinctions are blurred, resulting in a large deviation in the extraction results.
A dual-branch supervision network based on the fusion of edge and semantic features is adopted, and edge and semantic features are extracted through the backbone network, combined with edge enhancement module, semantic enhancement module and feature fusion module, feature enhancement and fusion are performed, and ultimately constrain the errors of edge prediction and region segmentation through joint supervision of the loss function.
It improves the accuracy of the front-line extraction of glacier disintegration, effectively alleviates the deviation of extraction results in complex glacier scenarios, and achieves higher automation, efficiency and accuracy.
Smart Images

Figure CN120032264A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image segmentation, and more specifically to a glacier collapse front extraction method based on edge and semantic feature fusion. Background Art
[0002] Glaciers are sensitive indicators of global climate change. Timely monitoring of glacier collapse and growth can help promote research on the laws of global climate change and has important practical significance for addressing global warming. The glacier collapse front refers to the boundary between an ice shelf that is still fixed to the ice and its surrounding open environment (ocean, lake or sea ice). It is an important parameter in ice sheet dynamics, and its location has irreplaceable value for polar research in oceanography, glaciology, and terrestrial or marine ecology. Therefore, accurately obtaining the location of the glacier collapse front is of great significance for global warming response and polar scientific research.
[0003] Due to the limitations of early computer technology, the extraction of glacier calving fronts was usually done manually, that is, glaciologists manually drew the location of the glacier calving front on remote sensing images; however, due to the complex topography of polar regions, manual annotation may lead to certain subjective deviations. In order to overcome the shortcomings of manual annotation of glacier calving fronts, semi-automatic methods based on traditional image processing technologies such as threshold segmentation and edge detection operators have gradually replaced manual methods, greatly reducing manual work; however, semi-automatic methods still require a lot of manual post-processing, and professional glaciological knowledge is required to annotate the calving front of difficult glaciers. In recent years, the introduction of a large number of remote sensing task algorithms based on deep learning has promoted the research on automated methods for extracting glacier calving fronts. Compared with manual annotation and semi-automatic methods, the automation of glacier calving front extraction has the advantages of low cost, high efficiency and high accuracy. It can also quickly obtain accurate detection results for large-scale glacier scenes with long time series, providing greater possibilities for timely response to global warming issues and in-depth development of polar scientific research.
[0004] The extraction of the glacier calving front based on deep learning methods is usually regarded as a classification task. By segmenting the glacier from the surrounding open environment into different categories of glacier and non-glacier, and further vectorizing the boundaries between the ground objects, the position of the glacier calving front is obtained. At present, deep learning-based image segmentation algorithms have achieved excellent performance in the task of extracting the glacier calving front; however, due to the complexity of the glacier scene in the polar region, there are often floating ice, sea ice mixtures, etc. near many glacier calving areas all year round, which increases the difficulty of detecting the glacier calving front. Most of the existing automatic methods for extracting the glacier calving front only regard the extraction of the glacier calving front as an ordinary image segmentation task, without considering the characteristics of the glacier scene. Therefore, it is restricted by the unclear expression feature differences between some ground objects in the remote sensing image, resulting in blurred boundaries for distinguishing ground objects, and further causing large deviations in the extraction results of the calving front in complex glacier scenes. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for extracting the glacier calving front based on the fusion of edge and semantic features, which can improve the accuracy of the network for automatically extracting the glacier calving front.
[0006] The present invention provides a method for extracting the glacier calving front based on the fusion of edge and semantic features, including the following steps: Step 1: Preprocess the remote sensing image and the corresponding regional label to obtain a training set and a test set; Step 2: Construct a two-branch supervised network and a loss function, and use the training set and the loss function to train the two-branch supervised network to obtain a trained two-branch supervised network; Step 3: Use the trained two-branch supervised network to predict the image in the test set to obtain a regional segmentation result; Post-process the regional segmentation result to obtain the position of the glacier calving front.
[0007] Further, the above two-branch supervised network includes a backbone network, an edge enhancement module, a semantic enhancement module, and a feature fusion module; the backbone network is used to obtain edge features and semantic features according to the input image; the edge enhancement module is used to enhance the edge features to obtain enhanced edge features and predicted edge results; the semantic enhancement module is used to enhance the semantic features to obtain enhanced semantic features; the feature fusion module is used to perform feature fusion and upsampling on the enhanced edge features and enhanced semantic features to obtain a regional segmentation result.
[0008] Furthermore, the above-mentioned backbone network is obtained by removing the last fully connected layer based on the VGG-16 network; the backbone network includes Block1 module, Block2 module, Block3 module, Block4 module, and Block5 module connected in series in sequence; Block1 module is used to generate a first hierarchical feature, Block2 module is used to generate a second hierarchical feature, Block3 module is used to generate a third hierarchical feature, Block4 module is used to generate a fourth hierarchical feature, and Block5 module is used to generate a fifth hierarchical feature.
[0009] Furthermore, the above-mentioned edge features include first hierarchical features and second hierarchical features; and the above-mentioned semantic features include fourth hierarchical features and fifth hierarchical features.
[0010] Furthermore, the Block1 module and the Block2 module of the above-mentioned backbone network are connected to the input end of the edge enhancement module; the Block4 module and the Block5 module of the backbone network are connected to the input end of the semantic enhancement module; the output end of the edge enhancement module and the output end of the semantic enhancement module are connected to the input end of the feature fusion module.
[0011] Furthermore, the edge enhancement module is specifically configured as follows: the first hierarchical feature and the second hierarchical feature are respectively subjected to a 1×1 convolution and a 3×3 convolution, and each is upsampled to the same size as the input image, and then spliced according to the channel dimension to obtain a first spliced edge feature; the first spliced edge feature is downsampled to the same size as the second hierarchical feature, and the number of channels is adjusted to 128 through a 1×1 convolution to obtain an enhanced edge feature; the first spliced edge feature is subjected to a 1×1 convolution to adjust the number of channels to 1, and The layer obtains the predicted edge results.
[0012] Furthermore, the feature fusion module is specifically configured as follows: upsampling the enhanced semantic features to the same size as the enhanced edge features, and splicing them with the enhanced edge features according to the channel dimension, successively undergoing a 1×1 convolution and a 3×3 convolution to change the number of channels to 128 and performing preliminary feature fusion to obtain preliminary fusion features; performing global average pooling on the preliminary fusion features, then undergoing a 1×1 convolution to perform feature compression and dimensionality reduction to change the number of channels to 32, then using the ReLU activation function for nonlinear transformation, then undergoing a 1×1 convolution to perform feature recovery and dimensionality increase to change the number of channels to 128, and then using the Sigmoid activation function to compress the output value to between 0 and 1 to implement a gating mechanism, and obtaining attention to the channel dimension of the preliminary fusion features; multiplying the attention to the channel dimension of the preliminary fusion features with the preliminary fusion features and then adding them to obtain fused features; upsampling the fused features to the size of the input image to obtain a region segmentation result.
[0013] Furthermore, the above loss function is as follows: , , , in, is the loss function, and Represent the weights of edge loss and regional loss respectively; is the marginal loss, is the predicted probability of the marginal outcome, and There are two hyperparameters in the edge loss, which are used to adjust the sample weight and the attention of the difficult and easy samples respectively; is the area loss, is the Dice loss, is the cross entropy loss, and are two hyperparameters in the region loss.
[0014] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned glacier collapse front extraction method based on the fusion of edge and semantic features.
[0015] The implementation of the glacier collapse front extraction method based on edge and semantic feature fusion provided by the present invention has the following beneficial effects: The present invention makes full use of image features by fusing edge features and semantic features, and constrains the errors of object edge prediction and regional segmentation through joint supervision of edge loss and semantic loss, so as to reduce the deviation of difficult glacier collapse front extraction. In the process of downsampling and learning image features by the deep learning network, since the shallow features represent more spatial and texture information, and the deep features represent more structural information, the present invention extracts the first two layers and the last two layers of features obtained by the model feature as the edge features and semantic features extracted by the network respectively; the present invention integrates the prior knowledge of glaciology into the deep learning model to carry out the glacier collapse front in complex glacier scenes. The boundary between glaciers and non-glaciers reflects the limit of the area where different objects are located. On the basis of the deep learning model of semantic segmentation, the edge information in the remote sensing image is extracted and the edge extraction result is constrained by using edge loss supervision, which can alleviate the problem of blurred edges in object segmentation and effectively improve the accuracy of glacier collapse front extraction in complex glacier scenes.
[0016] Specifically, the present invention uses the first two layers and the last two layers of features in the network downsampling process as the edge information and semantic information of the image respectively. After enhancing the edge features and semantic features respectively, the feature fusion module is used to enable the network to learn the effective information contained in the image in multiple scales and aspects, fully integrating the edge and semantic features of the image, and improving the accuracy of the network in automatically extracting the glacier collapse front; the present invention combines the characteristics of the glacier scene, integrates the boundary information between different landforms into the model as prior information, and constrains the deviation of edge prediction and regional segmentation from the true label through the joint supervision of edge loss and semantic loss, effectively alleviating the problem of blurred boundaries of image segmentation results; the method provided by the present invention is applied to remote sensing images covering the Antarctic Peninsula and Greenland. Even in the face of glacier scenes with floating ice all year round, the method can show excellent glacier collapse front extraction effect, indicating that the present invention can effectively alleviate the problem of difficulty in collapse front extraction in complex glacier scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which: Figure 1 It is a flow chart of a glacier collapse front extraction method based on edge and semantic feature fusion provided by the present invention; Figure 2 It is a dual-branch supervision network model diagram of edge and semantic feature fusion provided by the present invention; Figure 3 is an input image example provided by the present invention; Figure 4 is an output image example provided by the present invention. DETAILED DESCRIPTION
[0018] In order to have a clearer understanding of the technical features, purposes and effects of the present invention, specific embodiments of the present invention are now described in detail with reference to the accompanying drawings.
[0019] Figure 1 A schematic diagram of a glacier calving front extraction method based on edge and semantic feature fusion in this embodiment is shown. In this embodiment, the glacier calving front extraction method based on edge and semantic feature fusion includes the following steps: Step 1: Preprocess the remote sensing images and corresponding region labels to obtain training sets and test sets; As an exemplary embodiment, step 1 specifically includes: Step 1.1: Perform data enhancement on remote sensing images, including brightness adjustment, image deformation, Gaussian blur, rotation (90°, 180° and 270°) and flipping, with the probabilities of these operations being 0.1, 0.1, 0.5, 0.5 and 0.2 respectively; perform random data enhancement operations on the input data during the data loading phase to increase the diversity of samples; Step 1.2: Divide the remote sensing image into blocks; As an exemplary embodiment, step 1.2 specifically includes: Step 1.2.1: Determine whether the image height and width are integer multiples of the block height and width. If not, fill the lower right corner of the image with a pixel value of 0; Step 1.2.2: Use the sliding window strategy to divide the image into blocks, and set the block size to 256×256; the window based on the block size uses a certain step size to slide the window from the upper left corner of the image, and the image area covered by the window after each sliding is the current block; for the training set, the step size is 256, that is, there is no overlap between blocks; for the test set, the step size is 128, and there is half of the overlap between blocks to alleviate the problem of discontinuous prediction results after block splicing; Step 1.2.3: For each block, add the row and column number corresponding to its upper left corner in the original image to the original file name of the image to generate the file name corresponding to each block; Step 1.3: Process the region labels to generate corresponding edge labels; region labels usually include glaciers, oceans, rocks and other regions, so the corresponding edge labels generated are the boundaries between each type of region; As an exemplary embodiment, step 1.3 specifically includes: Step 1.3.1: Read the region label image , fill different numbers of rows and columns around it to generate different offset images that cause the region labels to shift relatively; specifically, in the region label Fill one row or column above, below, left and right to generate an unshifted image ; Fill one column on the left and right of the region label and two rows below to generate an upward-shifted image ; Similarly, generate images that are offset downward, left, and right , and ; Step 1.3.2: Detect the edges between each class using a logical formula. The formula is as follows:
[0020] The meaning of this logical formula is The pixel value in the image is not equal to the value of at least one pixel in the four directions around it, that is, there are other area categories around the pixel, so this pixel is regarded as the edge between the categories, and finally we get for the edges between each category; Step 1.3.3: Cropping Remove the padding area by one row above and one column below and one column on the left and right to get edge labels of the same size as the original area labels ; Step 2: construct a two-branch supervision network and a loss function, and train the two-branch supervision network using the training set and the loss function to obtain a trained two-branch supervision network; In an exemplary embodiment, the dual-branch supervision network includes a backbone network, an edge enhancement module, a semantic enhancement module and a feature fusion module; the backbone network is used to obtain edge features and semantic features according to an input image; the edge enhancement module is used to perform feature enhancement on the edge features to obtain enhanced edge features and predicted edge results; the semantic enhancement module is used to perform feature enhancement on the semantic features to obtain enhanced semantic features; the feature fusion module is used to perform feature fusion and up-sampling on the enhanced edge features and enhanced semantic features to obtain a region segmentation result; In an exemplary embodiment, the backbone network is obtained by removing the last fully connected layer based on the VGG-16 network; the backbone network includes a Block1 module, a Block2 module, a Block3 module, a Block4 module, and a Block5 module connected in series in sequence; the Block1 module is used to generate a first hierarchical feature, the Block2 module is used to generate a second hierarchical feature, the Block3 module is used to generate a third hierarchical feature, the Block4 module is used to generate a fourth hierarchical feature, and the Block5 module is used to generate a fifth hierarchical feature; In an exemplary embodiment, the edge feature includes a first hierarchical feature and a second hierarchical feature; the semantic feature includes a fourth hierarchical feature and a fifth hierarchical feature; In an exemplary embodiment, the Block1 module and the Block2 module of the backbone network are connected to the input end of the edge enhancement module; the Block4 module and the Block5 module of the backbone network are connected to the input end of the semantic enhancement module; the output end of the edge enhancement module and the output end of the semantic enhancement module are connected to the input end of the feature fusion module; In an exemplary embodiment, the edge enhancement module is specifically configured as follows: the first hierarchical feature and the second hierarchical feature are respectively subjected to a 1×1 convolution and a 3×3 convolution, and each is upsampled to the same size as the input image, and then spliced according to the channel dimension to obtain a first spliced edge feature; the first spliced edge feature is downsampled to the same size as the second hierarchical feature, and the number of channels is adjusted to 128 through a 1×1 convolution to obtain an enhanced edge feature; the first spliced edge feature is subjected to a 1×1 convolution to adjust the number of channels to 1, and then subjected to The layer obtains the predicted edge result; In an exemplary embodiment, the feature fusion module is specifically configured as follows: upsampling the enhanced semantic features to the same size as the enhanced edge features, and splicing them with the enhanced edge features according to the channel dimension, successively undergoing a 1×1 convolution and a 3×3 convolution to change the number of channels to 128 and performing preliminary feature fusion to obtain preliminary fusion features; performing global average pooling on the preliminary fusion features, then undergoing a 1×1 convolution to perform feature compression and dimensionality reduction to change the number of channels to 32, then using a ReLU activation function to perform nonlinear transformation, then undergoing a 1×1 convolution to perform feature recovery and dimensionality increase to change the number of channels to 128, and then using a Sigmoid activation function to compress the output value to between 0 and 1 to implement a gating mechanism, and obtaining attention to the channel dimension of the preliminary fusion features; multiplying the attention to the channel dimension of the preliminary fusion features with the preliminary fusion features and then adding them to obtain fused features; upsampling the fused features to the size of the input image to obtain a region segmentation result; In an exemplary embodiment, the loss function is as follows: , , , in, is the loss function, and Represent the weights of edge loss and regional loss respectively; is the marginal loss, is the predicted probability of the marginal outcome, and are two hyperparameters in the edge loss, which are used to adjust the sample weights and the attention to easy and difficult samples respectively; is the region loss, is the Dice loss, is the cross-entropy loss, and are two hyperparameters in the region loss; Step 3: Use the trained dual-branch supervised network to predict the images in the test set to obtain the region segmentation result; post-process the region segmentation result to obtain the position of the glacier calving front; In an exemplary embodiment, the post-processing of the region segmentation result includes block merging, extracting the boundary line between the glacier and the ocean as the glacier calving front, and obtaining the position of the glacier calving front corresponding to the remote sensing image.
[0021] In some embodiments, the above method for extracting the glacier calving front based on the fusion of edge and semantic features can also be implemented in the following manner. In this embodiment, the method for extracting the glacier calving front based on the fusion of edge and semantic features includes: Step S1: Perform data preprocessing on the remote sensing image and the corresponding region label, mainly including data augmentation, block division of the remote sensing image, and processing the region label to generate the corresponding edge label; Step S2: Design a dual-branch supervised network for fusing edge and semantic features, perform layer-by-layer abstraction and representation learning on the input image, and finally obtain the edge prediction and region segmentation result with the same size as the input image through decoding and restoration; Step S3: Automatically post-process the region segmentation result output by the network, including block merging, extracting the boundary line between the glacier and the ocean as the glacier calving front, and obtaining the position of the glacier calving front corresponding to the remote sensing image; Step S4: Use the trained network to automatically extract the glacier calving front of the remote sensing image of the glaciers in the polar region; Further, the specific implementation of Step S1 includes the following sub-steps: Step S1.1: Perform data augmentation on the remote sensing image, including brightness adjustment, image deformation, Gaussian blur, rotation (90°, 180°, and 270°), and flipping, and set different probabilities for these data augmentation operations respectively. Perform random data augmentation operations on the input data during the data loading stage to increase the diversity of samples; Step S1.2: Divide the remote sensing image into blocks; the specific steps are as follows: Step S1.2.1: Judge whether the height and width of the image are integer multiples of the block height and width. If not, fill the lower right corner of the image with pixels with a value of 0; Step S1.2.2: Use a sliding window strategy to divide the image into blocks, that is, a window based on the block size uses a certain step size to slide the window from the upper left corner of the image, and the image area covered by the window after each sliding is the current block; Step S1.2.3 For each block, append the row and column numbers corresponding to the upper left corner of the block in the original image to the original file name of the image to generate a file name corresponding to each block; Step S1.3: Process the region labels to generate corresponding edge labels; region labels usually include glaciers, oceans, rocks and other regions, so the corresponding edge labels generated are the boundaries between each type of region; the specific steps are as follows: Step S1.3.1 Read region label image , fill different numbers of rows and columns around it to generate different offset images that cause the region labels to shift relatively; specifically, in the region label Fill one row or column above, below, left and right to generate an unshifted image ; Fill one column on the left and right of the region label and two rows below to generate an upward-shifted image ; Similarly, generate images that are offset downward, left, and right , and ; Step S1.3.2 detects the edges between each class using a logical formula, the formula is as follows:
[0022] The meaning of this logical formula is The pixel value in the image is not equal to the value of at least one pixel in the four directions around it, that is, there are other area categories around the pixel, so this pixel is regarded as the edge between the categories, and finally we get for the edges between each category; Step S1.3.3 Cropping Remove the padding area by one row above and one column below and one column on the left and right to get edge labels of the same size as the original area labels ; Furthermore, the specific implementation of step S2 includes the following sub-steps: Step S2.1 Based on VGG-16 as the backbone network, the input image Perform feature extraction and generate hierarchical features in sequence , and remember for As edge features, for As a semantic feature; Step S2.2 Use edge enhancement module to enhance edge features Perform feature enhancement to generate enhanced edge features , and output the predicted edge results ; Step S2.3 Use semantic enhancement module to enhance semantic features Perform feature enhancement to generate enhanced semantic features ; Step S2.4 Use feature fusion module to fusion edge features and semantic features Perform feature fusion to obtain the fused features , then Upsample to input image The size of the region segmentation result is output ; Step S2.5: Edge prediction results And the region segmentation results , and calculate both and edge labels before each back propagation and area labels The loss is , and the output of the network is optimized by continuously updating the parameters; Furthermore, the specific implementation of step S2.2 includes the following sub-steps: Step S2.2.1 For edge features and A 1×1 convolution and a 3×3 convolution are performed successively, and each is upsampled to the same size as the input image. The same size, then spliced according to the channel dimension to obtain ; Step S2.2.2 Downsample to edge features The same size, and then adjust the number of channels through a 1×1 convolution to obtain enhanced edge features ; Step S2.2.3: After a 1×1 convolution, the number of channels is adjusted to 1, and after The layer obtains the predicted probability and outputs the predicted edge result ; Furthermore, the specific implementation of step S2.3 includes the following sub-steps: Step S2.3.1 Count Sketch method as a low-dimensional approximation of bilinear pooling for semantic features and Perform feature fusion to obtain a semantic feature whose channel number is much smaller than the product of the channel numbers of the two semantic features. ; Approximate bilinear pooling fusion based on Count Sketch method and The specific steps are as follows: Step S2.3.1.1 Random generation Fixed hash index , and random symbols ;in, , corresponding to two semantic features and ; Representation characteristics The number of channels; Indicates the number of channels after feature fusion; Step S2.3.1.2 Initialize output vector is a zero vector, where ; Step S2.3.1.3 Use the Count Sketch method to convert semantic features and Project to Dimension, the projection formula is as follows:
[0023] in, is the element index; Representation characteristics No. The value of the element; and Respectively represent indexes Random symbols and hash index The value of Then the output vector No. The value of the element; Step S2.3.1.4: For the two output vectors and Perform fast Fourier transform to convert them from time domain to frequency domain. After vector inner product in frequency domain, inverse fast Fourier transform is used to transform the result back to time domain to obtain the fused features. , the number of channels is The calculation formula is as follows:
[0024] in, and represent fast Fourier transform and inverse fast Fourier transform, respectively. It means element-wise multiplication; Step S2.3.2: After a 1×1 convolution, the number of channels is adjusted to obtain enhanced semantic features. ; Furthermore, the specific implementation of step S2.4 includes the following sub-steps: Step S2.4.1: Upsampled to The same size and According to the channel dimension, the concatenation is performed, and a 1×1 convolution and a 3×3 convolution are performed to change the number of channels and perform preliminary feature fusion to obtain the feature ; Step S2.4.2 First Perform global average pooling, then perform a 1×1 convolution to compress and reduce the dimension of the feature, then use the ReLU activation function for nonlinear transformation, then perform a 1×1 convolution to restore and increase the dimension of the feature, and then use the Sigmoid activation function to compress the output value to between 0 and 1 to implement the gating mechanism, and finally generate the feature Attention in the channel dimension ; Step S2.4.3 Focus on With features Multiply and add to obtain the fused features The calculation formula is as follows:
[0025] Step S2.4.4: Upsample to input image The size of the region segmentation result is output ; Step S2.5 is the loss function The calculation of is the marginal loss and regional losses The weighted sum of ; Among them, the edge loss uses the Focal loss function to alleviate the problem of insufficient learning of edge features caused by class imbalance; the regional loss is a combination of Dice loss and cross entropy loss; the calculation formula of the loss function is as follows:
[0026]
[0027]
[0028] in, and Respectively represent the weights of edge loss and regional loss. In view of the parallel double-branch structure designed in the present invention, it is set and All are 0.5; is the predicted probability of the marginal outcome, and There are two hyperparameters in the edge loss, which are used to adjust the sample weight and the attention of the difficult and easy samples respectively. They are set to 0.25 and 2 respectively according to experience; and These are two hyperparameters in the regional loss, and are set to 0.5 based on research experience; Furthermore, the specific implementation of step S3 includes the following sub-steps: Step S3.1: Merge the region segmentation results output by the network into blocks and restore them to the same size as the original corresponding image. The specific steps are as follows: Step S3.1.1: determine whether there is overlap between the blocks according to the block size of each image and the row and column number in the upper left corner of the block file name; Step S3.1.2 If there is no overlap between the blocks, they are directly spliced according to the row and column numbers in the block file names; Step S3.1.3 If there is overlap between blocks, calculate the Gaussian importance weighting of the blocks, by element-wise multiplying the prediction results of each block with a Gaussian kernel of similar size, giving higher weights to pixels close to the center of the block, and then taking the average value of the overlapping part as the final prediction result value, and splicing according to the row and column numbers in the block file name; Step S3.1.4 align the stitched image with the original corresponding image according to the upper left corner, and crop the stitched image according to the size of the original image to obtain the final region segmentation image; Step S3.2 extracts the boundary between the glacier and the ocean as the predicted glacier collapse front. The process is similar to step S1.3 in generating edge labels. Specifically, different numbers of rows and columns are filled around the final region segmentation image to generate unshifted images. , and images offset up, down, left, and right , , and , according to the pixel values of ocean and glacier in the regional segmentation results and , the glacier collapse front is obtained by the following logical formula :
[0029]
[0030] The meaning of this logical formula is The pixel value in the image is And at least one of the adjacent pixels has a value of ,get The boundary between the ocean and the glacier in the regional label is the glacier calving front. The filled area is removed in one row or one column on the top, bottom, left and right sides to obtain the glacier calving front extraction result with the same size as the original corresponding image.
[0031] In some embodiments, the above-mentioned glacier calving front extraction method based on edge and semantic feature fusion can also be implemented in the following manner. In this embodiment, the glacier calving front extraction method based on edge and semantic feature fusion includes: Step A1: preprocessing the remote sensing images of the training set and the test set and the corresponding region labels, mainly including data enhancement and segmentation of the remote sensing images, and processing the region labels to generate corresponding edge labels; Step A2 Design a dual-branch supervision network that fuses edge and semantic features, such as Figure 2 As shown, the data of the input training set is abstracted and represented layer by layer, and finally the edge prediction and region segmentation results with the same size as the input image are obtained through decoding and restoration; Step A3: automatically post-processing the regional segmentation results output by the network, including merging blocks, extracting the boundary between the glacier and the ocean as the glacier calving front, and obtaining the glacier calving front position corresponding to the remote sensing image; Step A4: Use the trained network to automatically extract the glacier collapse front from the remote sensing images in the test set. Input the remote sensing images and corresponding labels as follows: Figure 3 As shown, Figure 3 (a) is a grayscale image of an optical remote sensing image in the test set, taken by the Landsat-8 optical satellite on December 27, 2014, of the Fleming Glacier on the Antarctic Peninsula; Figure 3 (b) is the region label corresponding to the image, where white, light gray, and black represent ocean, glacier, and rock categories, respectively; Figure 3 (c) is the corresponding edge label generated based on the region label; Output example of region segmentation, glacier calving front extraction and visualization results Figure 4 As shown; among them, Figure 4 (a) is Figure 3 The region segmentation result of the corresponding image; Figure 4 (b) is the glacier collapse front extraction result of the image; Figure 4 (c) is the visualization result of the glacier calving front extraction result and the real glacier calving front on the true color image. The dark and light lines represent the real and predicted glacier calving front positions respectively. Furthermore, the specific implementation of step A1 includes the following sub-steps: Step A1.1 Perform data enhancement on the remote sensing images in the training set, including brightness adjustment, image deformation, Gaussian blur, rotation (90°, 180° and 270°) and flipping, with the probabilities of these operations being 0.1, 0.1, 0.5, 0.5 and 0.2 respectively; perform random data enhancement operations on the input data during the data loading phase to increase the diversity of samples; Step A1.2: Divide the remote sensing image into blocks; the specific steps are as follows: Step A1.2.1 determines whether the image height and width are integer multiples of the block height and width. If not, fill the lower right corner of the image with a pixel value of 0; Step A1.2.2 Use the sliding window strategy to divide the image into blocks, and set the block size to 256×256; the window based on the block size uses a certain step size to slide the window from the upper left corner of the image, and the image area covered by the window after each sliding is the current block; for the training set, the step size is 256, that is, there is no overlap between blocks; for the test set, the step size is 128, and there is half an overlap between blocks to alleviate the problem of discontinuous prediction results after block splicing; Step A1.2.3 For each block, append the row and column numbers corresponding to the upper left corner of the block in the original image to the original file name of the image to generate a file name corresponding to each block; Step A1.3: Process the region labels to generate corresponding edge labels; the region labels include three regions: glaciers, oceans, and rocks, and the corresponding edge labels generated are the boundaries between each type of region; the specific steps are as follows: Step A1.3.1 Read the region label image , fill different numbers of rows and columns around it to generate different offset images that cause the region labels to shift relatively; specifically, in the region label Fill one row or column above, below, left and right to generate an unshifted image ; Fill one column on the left and right of the region label and two rows below to generate an upward-shifted image ; Similarly, generate images that are offset downward, left, and right , and ; Step A1.3.2 Detect the edges between each class using a logical formula, the formula is as follows:
[0032] The meaning of this logical formula is The pixel value in the image is not equal to the value of at least one pixel in the four directions around it, that is, there are other area categories around the pixel, so this pixel is regarded as the edge between the categories, and finally we get for the edges between each category; Step A1.3.3 Cropping Remove the padding area by one row above and one column below and one column on the left and right to get edge labels of the same size as the original area labels ; Furthermore, the specific implementation of step A2 includes the following sub-steps: Step A2.1 Based on VGG-16 as the backbone network, the input image Perform feature extraction and generate hierarchical features in sequence ; According to the form of [batch size, number of channels, height, width], the shapes of each layer of features are [16, 64, 128, 128], [16, 128, 64, 64], [16, 256, 32, 32], [16, 512, 16, 16] and [16, 512, 16, 16]; and, record for As edge features, for As a semantic feature; Step A2.2 Use edge enhancement module to enhance edge features Perform feature enhancement to generate enhanced edge features , and output the predicted edge results ; Step A2.3 Use semantic enhancement module to enhance semantic features Perform feature enhancement to generate enhanced semantic features ; Step A2.4 Use the feature fusion module to fusion edge features and semantic features Perform feature fusion to obtain the fused features , then Upsample to input image The size of the region segmentation result is output ; Step A2.5: For edge prediction results And the region segmentation results , and calculate both and edge labels before each back propagation and area labels The loss is , and the output of the network is optimized by continuously updating the parameters; Furthermore, the specific implementation of step A2.2 includes the following sub-steps: Step A2.2.1 For edge features and A 1×1 convolution and a 3×3 convolution are performed respectively, and each is upsampled to the same size as the input image. The same size, then spliced according to the channel dimension to obtain ; Step A2.2.2 Downsample to edge features The same size, and then adjust the number of channels to 128 through a 1×1 convolution to obtain enhanced edge features ; Step A2.2.3: After a 1×1 convolution, the number of channels is adjusted to 1, and after The layer obtains the predicted probability and outputs the predicted edge result ; Furthermore, the specific implementation of step A2.3 includes the following sub-steps: Step A2.3.1 Count Sketch method as a low-dimensional approximation of bilinear pooling for semantic features and Perform feature fusion and set the number of output channels to 8000 (much smaller than and The product of their respective dimensions is 512×512), and the fused semantic features are obtained. ; Approximate bilinear pooling fusion based on Count Sketch method and The specific steps are as follows: Step A2.3.1.1 Random Generation Fixed hash index , and random symbols ;in, , corresponding to two semantic features and ; Representation characteristics The number of channels; Indicates the number of channels after feature fusion; Step A2.3.1.2 Initialize output vector is a zero vector, where ; Step A2.3.1.3 Use the Count Sketch method to convert semantic features and Project to Dimension, the projection formula is as follows:
[0033] in, is the element index; Representation characteristics No. The value of the element; and Respectively represent indexes Random symbols and hash index The value of Then the output vector No. The value of the element; Step A2.3.1.4 For the two output vectors and Perform fast Fourier transform to convert them from time domain to frequency domain. After vector inner product in frequency domain, inverse fast Fourier transform is used to transform the result back to time domain to obtain the fused features. , the number of channels is The calculation formula is as follows:
[0034] in, and represent fast Fourier transform and inverse fast Fourier transform, respectively. It means element-wise multiplication; Step A2.3.2 After a 1×1 convolution, the number of channels is adjusted to obtain enhanced semantic features. ; Furthermore, the specific implementation of step 2.4 includes the following sub-steps: Step A2.4.1 Upsampled to The same size and According to the channel dimension, the concatenation is performed, and a 1×1 convolution and a 3×3 convolution are performed to change the number of channels to 128 and perform preliminary feature fusion to obtain the feature ; Step A2.4.2 First Perform global average pooling, then perform a 1×1 convolution to compress features and reduce the dimension to change the number of channels to 32, then use the ReLU activation function for nonlinear transformation, then perform a 1×1 convolution to restore features and increase the dimension to change the number of channels to 128, and then use the Sigmoid activation function to compress the output value to between 0 and 1 to implement the gating mechanism, and finally generate the feature Attention in the channel dimension ; Step A2.4.3 Focus on With features Multiply and add to obtain the fused features The calculation formula is as follows:
[0035] Step A2.4.4: Upsample to input image The size of the region segmentation result is output ; Step A2.5 is the loss function The calculation of is the marginal loss and regional losses The weighted sum of ; Among them, the edge loss uses the Focal loss function to alleviate the problem of insufficient learning of edge features caused by class imbalance; the regional loss is a combination of Dice loss and cross entropy loss; the calculation formula of the loss function is as follows:
[0036]
[0037]
[0038] in, and Respectively represent the weights of edge loss and regional loss. In view of the parallel double-branch structure designed in the present invention, it is set and All are 0.5; is the predicted probability of the marginal outcome, and There are two hyperparameters in the edge loss, which are used to adjust the sample weight and the attention of the difficult and easy samples respectively. They are set to 0.25 and 2 respectively according to experience; and These are two hyperparameters in the regional loss, and are set to 0.5 based on research experience; Furthermore, the specific implementation of step A3 includes the following sub-steps: Step A3.1: Merge the region segmentation results output by the network into blocks and restore the prediction results to the same size as the original corresponding image. The specific steps are as follows: Step A3.1.1: Determine whether there is overlap between blocks based on the block size of each image and the row and column number in the upper left corner of the block file name; Step A3.1.2 For the training set, there is no overlap between the blocks, and they are directly spliced according to the row and column numbers in the block file name; Step A3.1.3 For the test set, there is overlap between blocks. Calculate the Gaussian importance weighting of the blocks. By element-wise multiplying the prediction results of each block with a Gaussian kernel of a similar size, higher weights are assigned to the pixels closer to the center of the block. Then, take the average value in the overlapping part as the final prediction result value, and splice according to the row and column numbers in the block file name; Step A3.1.4 Align the spliced image with the original corresponding image according to the upper left corner, and crop the spliced image according to the size of the original image to obtain the final region segmentation image; Step A3.2 Extract the glacier-ocean boundary line as the predicted glacier calving front. The process is similar to that of generating the edge label in Step A1.3; specifically, perform padding with different numbers of rows and columns around the final region segmentation image to generate the non-offset image , as well as the images offset upward, downward, leftward, and rightward , , and . According to the pixel values of the ocean and glacier in the region segmentation result and , obtain the glacier calving front through the following logical formula :
[0039]
[0040] The meaning of this logical formula is In the image, the pixel value is and at least one adjacent pixel value is , obtain as the boundary between the ocean and the glacier in the region label, that is, the glacier calving front; finally, remove the padding area by cropping one row or one column from the top, bottom, left, and right to obtain the glacier calving front extraction result with the same size as the original corresponding image.
[0041] This embodiment provides a computer program product, including a computer program, which when executed by a processor implements the steps of the above-mentioned method for extracting the glacier calving front based on the fusion of edge and semantic features.
[0042] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose and scope protected by the claims of the present invention. These all fall within the protection scope of the present invention.
Claims
1. A glacier collapse front extraction method based on edge and semantic feature fusion, characterized in that: The following steps are involved: Step 1: Preprocess the remote sensing images and corresponding region labels to obtain training sets and test sets; Step 2: construct a two-branch supervision network and a loss function, and train the two-branch supervision network using the training set and the loss function to obtain a trained two-branch supervision network; Step 3: Use the trained dual-branch supervised network to predict the images in the test set to obtain regional segmentation results; post-process the regional segmentation results to obtain the position of the glacier collapse front.
2. The method for extracting glacier collapse front based on edge and semantic feature fusion according to claim 1, characterized in that: The dual-branch supervision network includes a backbone network, an edge enhancement module, a semantic enhancement module and a feature fusion module; the backbone network is used to obtain edge features and semantic features according to the input image; The edge enhancement module is used to perform feature enhancement on the edge features to obtain enhanced edge features and predicted edge results; the semantic enhancement module is used to perform feature enhancement on the semantic features to obtain enhanced semantic features; the feature fusion module is used to perform feature fusion and up-sampling on the enhanced edge features and enhanced semantic features to obtain region segmentation results.
3. The method for extracting glacier collapse front based on edge and semantic feature fusion according to claim 2 is characterized in that: The backbone network is obtained by removing the last fully connected layer from the VGG-16 network; the backbone network includes a Block1 module, a Block2 module, a Block3 module, a Block4 module, and a Block5 module connected in series in sequence; the Block1 module is used to generate a first hierarchical feature, the Block2 module is used to generate a second hierarchical feature, the Block3 module is used to generate a third hierarchical feature, the Block4 module is used to generate a fourth hierarchical feature, and the Block5 module is used to generate a fifth hierarchical feature.
4. The method for extracting glacier collapse front based on edge and semantic feature fusion according to claim 3 is characterized in that: The edge features include first hierarchical features and second hierarchical features; the semantic features include fourth hierarchical features and fifth hierarchical features.
5. The method for extracting glacier collapse front based on edge and semantic feature fusion according to claim 3, characterized in that: The Block1 module and the Block2 module of the backbone network are connected to the input end of the edge enhancement module; the Block4 module and the Block5 module of the backbone network are connected to the input end of the semantic enhancement module; the output end of the edge enhancement module and the output end of the semantic enhancement module are connected to the input end of the feature fusion module.
6. The method for extracting glacier collapse front based on edge and semantic feature fusion according to claim 3, characterized in that: The edge enhancement module is specifically configured as follows: performing a 1×1 convolution and a 3×3 convolution on the first hierarchical feature and the second hierarchical feature respectively, and upsampling each to the same size as the input image, and then splicing them according to the channel dimension to obtain a first spliced edge feature; The first spliced edge feature is downsampled to the same size as the second hierarchical feature, and then the number of channels is adjusted to 128 through a 1×1 convolution to obtain an enhanced edge feature; the first spliced edge feature is adjusted to 1 through a 1×1 convolution, and the number of channels is adjusted to 1 through The layer obtains the predicted edge results.
7. The method for extracting glacier collapse front based on edge and semantic feature fusion according to claim 3, characterized in that: The feature fusion module is specifically configured as follows: upsampling the enhanced semantic features to the same size as the enhanced edge features, and concatenating them with the enhanced edge features according to the channel dimension, successively undergoing a 1×1 convolution and a 3×3 convolution to change the number of channels to 128 and performing preliminary feature fusion to obtain preliminary fused features; The preliminary fusion features are globally averaged pooled, and then a 1×1 convolution is performed to compress the features and reduce the dimension to change the number of channels to 32, and then a ReLU activation function is used for nonlinear transformation, and then a 1×1 convolution is performed to restore the features and increase the dimension to change the number of channels to 128, and then a Sigmoid activation function is used to compress the output value to between 0 and 1 to implement a gating mechanism, and obtain the attention to the channel dimension of the preliminary fusion features; the attention to the channel dimension of the preliminary fusion features is multiplied and added with the preliminary fusion features to obtain the fused features; The fused features are upsampled to the size of the input image to obtain a region segmentation result.
8. The method for extracting glacier collapse front based on edge and semantic feature fusion according to claim 1, characterized in that: The loss function is as follows: , , , in, is the loss function, and Represent the weights of edge loss and regional loss respectively; is the marginal loss, is the predicted probability of the marginal outcome, and There are two hyperparameters in the edge loss, which are used to adjust the sample weight and the attention of the difficult and easy samples respectively; is the area loss, is the Dice loss, is the cross entropy loss, and are two hyperparameters in the region loss.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for extracting the glacier collapse front based on the fusion of edge and semantic features as described in any one of claims 1 to 8 are implemented.