A Semantic Segmentation Method for High-Resolution Remote Sensing Images Based on Encoding-Decoding Indexed Edge Representation

By adopting codec indexed edge characterization technology in the semantic segmentation model of high-resolution remote sensing image, the problem of poor segmentation effect of remote sensing images at the edge of objects is solved, the recognition accuracy of small-size objects and complex boundary information is improved, and high-precision semantic segmentation of remote sensing object edges is achieved.

CN116524189BActive Publication Date: 2025-06-17DALIAN MARITIME UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310496605.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-05
Publication Date
2025-06-17
Estimated Expiration
2043-05-05

AI Technical Summary

Technical Problem

The semantic segmentation model of high-resolution remote sensing images has poor segmentation effect at the edge of the object, especially the recognition accuracy of small-volume objects and complex boundary information is not high.

Method used

The semantic segmentation method of high-resolution remote sensing image based on edge characterization based on codec index is adopted. The multi-scale semantic features of the image are extracted through a multi-scale feature encoder, and the spatial context information is captured in parallel by separable pyramid units, and the segmentation effect of edge information is enhanced through codec index.

Benefits of technology

The feature extraction and processing capability of remote sensing object edge information is improved, the recognition accuracy of small-size objects and complex boundary information in remote sensing images is improved, and the accurate semantic segmentation of remote sensing object edges is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116524189B_ABST
    Figure CN116524189B_ABST
Patent Text Reader

Abstract

The present invention discloses a semantic segmentation method for high-resolution remote sensing images based on codec-indexed edge representation, including: acquiring and augmenting a remote sensing image set, normalizing the remote sensing images, constructing a remote sensing image semantic segmentation model based on codec-indexed edge representation, obtaining the predicted labels of each pixel in the remote sensing image after training the segmentation model according to a training set, calculating the loss according to the ground truth labels and the predicted labels, determining whether the loss value meets a threshold, if it does not meet the threshold, updating the parameters of the segmentation model, if it meets the threshold, obtaining the trained segmentation model, acquiring the processed remote sensing images and inputting them into the trained segmentation model, and outputting the semantic segmentation result map of the remote sensing images. It improves the feature extraction and processing of ground object edge information, enhances the recognition accuracy of small-size objects and complex boundary information in remote sensing images, and realizes the precise semantic segmentation of remote sensing ground object edges.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image processing, and in particular to a high-resolution remote sensing image semantic segmentation method based on codec-indexed edge representation. Background Art

[0002] High-spatial-resolution remote sensing images (high-resolution remote sensing images) are an important part of modern remote sensing images, with characteristics such as high spatial resolution, high definition, high timeliness, and large amounts of information. Through them, rich ground object detail information and the relationships between adjacent ground objects can be clearly and intuitively presented. Currently, semantic segmentation of images is a research hotspot in the field of computer vision. Its task essence is the category recognition of image regions, that is, assigning category labels to each pixel in the image. As an important part of the semantic segmentation direction, high-resolution remote sensing image semantic segmentation can automatically extract surface features in remote sensing images and assign semantic categories to ground object targets. High-resolution remote sensing image semantic segmentation has wide applications in fields such as disaster assessment and prediction, environmental protection, urban planning, traffic navigation, and military security.

[0003] In recent years, deep learning, especially deep convolutional neural network technology, has developed and been applied rapidly. It shows amazing feature extraction capabilities in tasks such as image classification, object detection, and semantic segmentation, and can adaptively extract shallow and deep features in images, especially having good understanding capabilities for complex scenes. Therefore, applying deep learning technology to the semantic segmentation of high-resolution remote sensing images has important practical significance and will bring new development opportunities for the processing of remote sensing images. However, high-resolution remote sensing images usually consist of large and complex scenes and heterogeneous objects, and occlusion and shadow problems caused by lighting conditions and imaging angles during image acquisition lead to poor segmentation effects at the edges of objects in existing deep remote sensing segmentation models. In addition, objects with small volumes have a higher proportion of edge pixels in the overall pixels of the object, and if the segmentation at the edges is not ideal, it will also lead to poor segmentation effects for the entire object. Summary of the Invention

[0004] The present invention provides a high-resolution remote sensing image semantic segmentation method based on codec-indexed edge representation to overcome the above technical problems.

[0005] A high-resolution remote sensing image semantic segmentation method based on codec-indexed edge representation includes:

[0006] Step 1: Obtain a remote sensing image set, augment the remote sensing image set. The augmentation is to rotate the remote sensing images at any angle and store them in the remote sensing image set, perform normalization processing on the remote sensing images respectively, and divide the remote sensing image set into a training set and a test set.

[0007] Step 2: Construct a remote sensing image semantic segmentation model based on encoded-decoded indexed edge representation. After training the remote sensing image semantic segmentation model based on encoded-decoded indexed edge representation using the training set, obtain the predicted label of each pixel in the remote sensing image, calculate the loss based on the true label and predicted label of each pixel, and determine whether the value of the loss satisfies the threshold. If it does not satisfy the threshold, optimize the parameters of the remote sensing image semantic segmentation model based on encoded-decoded indexed edge representation according to the difference between the value of the loss and the threshold. If it satisfies the threshold, obtain the trained remote sensing image semantic segmentation model based on encoded-decoded indexed edge representation.

[0008] The remote sensing image semantic segmentation model based on encoded-decoded indexed edge representation includes a multi-scale feature encoder, a separable pyramid unit, an encoded-decoded indexed edge representation unit, and an upsampling decoder.

[0009] The multi-scale feature encoder is used to generate four initial feature matrices according to the size h of the remote sensing image, and the sizes of the four initial feature matrices are h / 2, h / 4, h / 8, and h / 16 respectively.

[0010] The separable pyramid unit is used to obtain four context feature matrices according to the four initial feature matrices of the remote sensing image, and the sizes of the four context feature matrices are h / 2, h / 4, h / 8, and h / 16 respectively.

[0011] The encoded-decoded indexed edge representation unit is used to obtain a first encoding index and a first decoding index according to the context feature matrix with a size of h / 2, obtain a second encoding index and a second decoding index according to the context feature matrix with a size of h / 4, fuse the first encoding index with the context feature matrix with a size of h / 2, fuse the second encoding index with the context feature matrix with a size of h / 4, and obtain the fused context feature matrix with a size of h / 2 and the context feature matrix with a size of h / 4.

[0012] The upsampling decoder is used to perform decoded upsampling on the four context feature matrices in ascending order of size to obtain the semantic segmentation result map of the remote sensing image, and the semantic segmentation result map includes the predicted label of each pixel in the remote sensing image.

[0013] Step 3: Obtain the processed remote sensing image and input it into the trained remote sensing image semantic segmentation model based on encoded-decoded indexed edge representation, and output the semantic segmentation result map of the remote sensing image.

[0014] Preferably, the upsampling decoder is used to perform decoded upsampling on the context feature matrix with a size of h / 16 to obtain an output feature matrix x with a size of h / 8 d1 , and x d1Concatenate the context feature matrix with size h / 8 along dimension 1 using the torch.cat function to form a new feature matrix x m1 ;

[0015] Perform decoding upsampling on x m1 to obtain an output feature matrix x with size h / 4 d2 , concatenate x d2 with the context feature matrix with size h / 4 along dimension 1 using the torch.cat function to form a new feature matrix x m2 , perform matrix multiplication on the second decoding index and the feature matrix x m2 and then output the feature matrix x n2 ;

[0016] Perform decoding upsampling on x n2 to obtain an output feature matrix with size h / 2, then concatenate it with the context feature matrix with size h / 2 along dimension 1 using the torch.cat function to form a new feature matrix x m3 , perform matrix multiplication on the first decoding index and the feature matrix x m3 and output the feature matrix x n3 ;

[0017] Perform decoding upsampling on x n3 to obtain a feature matrix x with size h d4 , perform one convolution on x d4 and then input it into the softmax activation function to obtain the semantic segmentation result map.

[0018] Preferably, calculating the loss according to the ground truth label and the predicted label of each pixel includes calculating the loss according to formula (1),

[0019] Loss focal = -(1 - p t ) γ log(p t ) (1)

[0020] where p t is the predicted probability of the ground truth label, the predicted probability is obtained according to the ground truth label and the predicted label, γ is a hyperparameter, and Loss focal represents the focal loss function.

[0021] Preferably, the multi-scale context feature encoder includes a spatial feature extraction branch, a self-attention feature extraction branch, and a fusion branch. The spatial feature extraction branch is used to extract local feature information of the remote sensing image, the self-attention feature extraction branch is used to extract global feature information of the remote sensing image, and the fusion branch is used to fuse the local feature information and the global feature information according to formula (2).

[0022] x = concatnate(Conv2d(x ci ), Conv2d(x si ))

[0023] y = sigmoid(Conv2d(ReLU(Conv2d(AdaptiveAvgPool2d(x)))))

[0024] x fi = x * reshape(y)(2)

[0025] Where x si represents the feature matrix of the i-th stage of the self-attention feature extraction branch, x ci represents the feature matrix of the i-th stage of the spatial feature extraction branch, x fi represents the fused feature, Conv2d(*) represents 2D convolution, AdaptiveAvgPool2d(*) represents the adaptive pooling function, sigmoid(*) represents the sigmoid activation function, ReLU(*) represents the ReLU activation function, concatnate(*) represents concatenating two matrices along dimension 1, and reshape(*) represents the shape change function.

[0026] Preferably, the obtaining of the first encoding index and the first decoding index according to the context feature matrix of size h / 2 includes

[0027] S11. Represent the context feature matrix of size h / 2 as x i , and obtain the shape parameters of x i , where the shape parameters include the batch size value batch size, the number of channels c, the height h, and the width w.

[0028] S12. Input x i into the Conv2d function to obtain x i1 , input x i1 into the BatchNorm2d function to obtain x i2 , input x i2 into the BatchNorm2d function to obtain x i3 , input x i3 into the BatchNorm2d function to obtain x i4 .

[0029] S13. Perform max pooling operations on x i1 , x i2 , x i3 , x i4 respectively to obtain four initial indices x1, x2 , x 3 , x 4 ,

[0030] S14. Concatenate the initial indices x1, x2, x3, x4 along dimension 1 into a new matrix through the torch.cat function and pass it to the sigmoid activation function to obtain the initial decoding index y; S15. The initial decoding index y passes through the softmax function to obtain the initial encoding index z. Use the view function to adjust the shape parameters of the initial decoding index y and the initial encoding index z. The adjustment is to adjust the shape parameters to batch size, c×4, h / 2, w / 2, and obtain the adjusted initial decoding index y and initial encoding index z.

[0031] S15. Use the pixel_shuffle function to reorganize the adjusted initial decoding index y and initial encoding index z into the size before adjustment to obtain the first encoding index and the first decoding index.

[0032] The present invention provides a high-resolution remote sensing image semantic segmentation method based on encoded-decoded indexized edge representation. The multi-scale feature encoder in the remote sensing image semantic segmentation model based on encoded-decoded indexized edge representation extracts the multi-scale semantic features of the image. The separable pyramid units can capture the spatial context information in parallel. By extracting the encoded-decoded indices containing edge information, the segmentation effect of the remote sensing ground object edge information is strengthened, the feature extraction and processing of the remote sensing ground object edge information are improved, the recognition accuracy of small-size objects and complex boundary information in the remote sensing image is enhanced, and the accurate semantic segmentation of the remote sensing ground object edge is realized. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0034] Figure 1 is the flowchart of the method of the present invention;

[0035] Figure 2 is the schematic diagram of the process of the remote sensing image semantic segmentation model based on encoded-decoded indexized edge representation of the present invention;

[0036] Figure 3 is the schematic diagram of the structure of the multi-scale feature encoder of the present invention;

[0037] Figure 4 is the schematic diagram of the structure of the separable pyramid unit of the present invention;

[0038] Figure 5 This is a schematic diagram of the index structure generated by the present invention. Specific implementation manners

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0040] Figure 1 This is a flowchart of the method of the present invention. As Figure 1 shown, the method of this embodiment may include:

[0041] Step 1: Obtain a remote sensing image set, perform augmentation on the remote sensing image set. The augmentation is to rotate the remote sensing images at any angle and then store them in the remote sensing image set, perform normalization processing on the remote sensing images respectively, and divide the remote sensing image set into a training set and a test set.

[0042] Step 2: Construct a remote sensing image semantic segmentation model based on an encoder-decoder indexed edge representation. After training the remote sensing image semantic segmentation model based on the encoder-decoder indexed edge representation according to the training set, obtain the predicted label of each pixel in the remote sensing image, calculate the loss according to the true label and the predicted label of each pixel, and determine whether the value of the loss satisfies a threshold. If it does not satisfy the threshold, optimize the parameters of the remote sensing image semantic segmentation model based on the encoder-decoder indexed edge representation according to the difference between the value of the loss and the threshold. If it satisfies the threshold, obtain the trained remote sensing image semantic segmentation model based on the encoder-decoder indexed edge representation.

[0043] The remote sensing image semantic segmentation model based on the encoder-decoder indexed edge representation includes a multi-scale feature encoder, a separable pyramid unit, an encoder-decoder indexed edge representation unit, and an upsampling decoder.

[0044] The multi-scale feature encoder is used to generate four initial feature matrices according to the size h of the remote sensing image. The sizes of the four initial feature matrices are h / 2, h / 4, h / 8, and h / 16 respectively.

[0045] The separable pyramid unit is used to obtain four context feature matrices according to the four initial feature matrices of the remote sensing image. The sizes of the four context feature matrices are h / 2, h / 4, h / 8, and h / 16 respectively.

[0046] The encoding and decoding indexed edge representation unit is used to obtain a first encoding index and a first decoding index according to the context feature matrix with a size of h / 2, obtain a second encoding index and a second decoding index according to the context feature matrix with a size of h / 4, fuse the first encoding index with the context feature matrix with a size of h / 2, fuse the second encoding index with the context feature matrix with a size of h / 4, and obtain the fused context feature matrix with a size of h / 2 and the context feature matrix with a size of h / 4.

[0047] The upsampling decoder is used to decode and upsample the context feature matrix in ascending order of size to obtain the semantic segmentation result map of the remote sensing image. The semantic segmentation result map includes the predicted labels of each pixel in the remote sensing image.

[0048] Step 3: Obtain the processed remote sensing image and input it into the trained remote sensing image semantic segmentation model based on the encoding and decoding indexed edge representation to output the semantic segmentation result map of the remote sensing image.

[0049] Based on the above solution, the multi-scale feature encoder in the remote sensing image semantic segmentation model based on the encoding and decoding indexed edge representation extracts the multi-scale semantic features of the remote sensing image. The separable pyramid unit captures the spatial context information in parallel. By extracting the encoding and decoding indexes containing edge information, the segmentation effect of the remote sensing ground object edge information is strengthened, the feature extraction and processing of the remote sensing ground object edge information are improved, the recognition accuracy of small-size objects and complex boundary information in the remote sensing image is enhanced, and the accurate semantic segmentation of the remote sensing ground object edge is realized.

[0050] Step 1: Obtain a remote sensing image set, perform augmentation on the remote sensing image set. The augmentation is to rotate the remote sensing image at an arbitrary angle and store it in the remote sensing image set. Data augmentation is to prevent the model from overfitting and improve the robustness of the model. Generally, it includes operations such as vertical and horizontal flipping, and rotating 90° of the remote sensing image, and perform normalization processing on the remote sensing image respectively. Normalization is a data preprocessing operation that maps the original remote sensing image data to the range of 0 to 1, and divide the remote sensing image set into a training set and a test set.

[0051] Step 2: Construct a remote sensing image semantic segmentation model based on the encoding and decoding indexed edge representation, as Figure 2As shown, after training a remote sensing image semantic segmentation model based on an encoder-decoder indexed edge representation using a training set, the predicted label of each pixel in the remote sensing image is obtained. The loss is calculated based on the ground truth label and the predicted label of each pixel. It is determined whether the value of the loss meets a threshold. If the threshold is not met, the parameters of the remote sensing image semantic segmentation model based on the encoder-decoder indexed edge representation are updated according to the difference between the loss and the threshold. If the threshold is met, the trained remote sensing image semantic segmentation model based on the encoder-decoder indexed edge representation is obtained. The calculation of the loss based on the ground truth label and the predicted label of each pixel includes calculating the loss according to formula (1).

[0052]

[0053] where p t is the predicted probability of the ground truth label, and the predicted probability is obtained based on the ground truth label and the predicted label. γ is a hyperparameter, and Loss focal represents the focal loss function.

[0054] The remote sensing image semantic segmentation model based on the encoder-decoder indexed edge representation includes a multi-scale feature encoder, a separable pyramid unit, an encoder-decoder indexed edge representation unit, and an upsampling decoder.

[0055] The multi-scale feature encoder is used to generate four initial feature matrices according to the size h of the remote sensing image. The sizes of the four initial feature matrices are h / 2, h / 4, h / 8, and h / 16 respectively.

[0056] The multi-scale context feature encoder includes a spatial feature extraction branch, a self-attention feature extraction branch, and a fusion branch. As Figure 3 shown, the spatial feature extraction branch is used to extract local feature information of the remote sensing image. The self-attention feature extraction branch is used to extract global feature information of the remote sensing image. The fusion branch is used to fuse the local feature information and the global feature information according to formula (2).

[0057]

[0058] where x si represents the feature matrix of the i-th stage of the self-attention feature extraction branch, x ci represents the feature matrix of the i-th stage of the spatial feature extraction branch, x fi represents the fused feature. Conv2d(*) represents a 2d convolution, AdaptiveAvgPool2d(*) represents an adaptive pooling function, sigmoid(*) represents a sigmoid activation function, ReLU(*) represents a ReLU activation function, concatnate(*) represents concatenating two matrices along dimension 1, and reshape(*) represents a shape change function.

[0059] Specifically, when using a multi-scale context feature encoder to extract global and local feature information of remote sensing images, the following method is adopted:

[0060] Use the first five stages of the lightweight feature extraction model EfficientNet-B5 as the spatial feature extraction branch to extract local feature information of remote sensing images. EfficientNetB-5 is a type of convolutional neural network structure and belongs to the EfficientNet series of models. The design goal of the EfficientNet series of models is to provide better performance while keeping the computational cost low. EfficientNetB-5 is the fifth model in the EfficientNet series;

[0061] Use Swin Transformer as the self-attention feature extraction branch to extract global feature information of remote sensing images. Swin Transformer is a new type of neural network architecture that uses a Transformer-based method to solve computer vision tasks. Its innovation lies in using a method called "Shifted Windows" to process images, which enables it to perform calculations more efficiently when dealing with large-scale image data;

[0062] The separable pyramid unit is used to obtain four context feature matrices based on four initial feature matrices of the remote sensing image. The sizes of the four context feature matrices are h / 2, h / 4, h / 8, and h / 16 respectively. The structure of the separable pyramid unit Figure 4 is shown. The function of the separable pyramid unit is to use separable dilated convolutions with different dilation rates to parallelly capture the spatial context information of multiple scales of each initial feature map and output context feature matrices of the same size. The dilation rates are 0, 1, 6, and 12;

[0063] The encoding and decoding indexation edge representation unit is used to obtain the first encoding index and the first decoding index based on the context feature matrix with a size of h / 2. The structure for generating the index is as Figure 5 shown. Specifically, obtaining the first encoding index and the first decoding index based on the context feature matrix with a size of h / 2 includes

[0064] S11. Represent the context feature matrix with a size of h / 2 as x i , and obtain the shape parameters of x i . The shape parameters include the batch size value batch size, the number of channels c, the height h, and the width w,

[0065] S12. Input x i into the Conv2d function to get xi1 , input x i1 into the BatchNorm2d function to obtain x i2 , input x i2 into the BatchNorm2d function to obtain x i3 , input x i3 into the BatchNorm2d function to obtain x i4 ,

[0066] S13. Perform max pooling operations on x i1 , x i2 , x i3 , x i4 respectively to obtain four initial indices x1, x 2 , x 3 , x 4 ,

[0067] S14. Concatenate the initial indices x1, x2, x3, and x4 along dimension 1 using the torch.cat function and pass them to the sigmoid activation function to obtain the initial decoding index y; S15. The initial decoding index y passes through the softmax function to obtain the initial encoding index z. Use the view function to adjust the shape parameters of the initial decoding index y and the initial encoding index z. The adjustment is to adjust the shape parameters to batch size, c×4, h / 2, w / 2, and obtain the adjusted initial decoding index y and initial encoding index z.

[0068] S16. Use the pixel_shuffle function to reorganize the adjusted initial decoding index y and initial encoding index z into the size before adjustment to obtain the first encoding index and the first decoding index.

[0069] Obtain the second encoding index and the second decoding index according to the context feature matrix with size h / 4. For the context feature matrix with size h / 4, it can also be processed according to steps S11 - S16 to obtain the second encoding index and the second decoding index.

[0070] Fuse the first encoding index with the context feature matrix with size h / 2, and fuse the second encoding index with the context feature matrix with size h / 4 to obtain the fused context feature matrix with size h / 2 and the context feature matrix with size h / 4.

[0071] Specifically, when using the separable pyramid unit to parallelly capture the spatial context information of multiple scales of the initial feature map, the following method is adopted:

[0072] The separable pyramid unit replaces all 3×3 convolutions in the dilated pyramid module with depthwise separable convolutions, and constructs and applies the separable pyramid unit based on four feature matrices of different sizes. Among them, the dilated pyramid module is a commonly used convolutional neural network module in deep learning for extracting features of different scales in images. Its structure is as follows:

[0073] Input layer: Accepts the input data from the previous layer.

[0074] Convolutional layer: Performs convolution operations on the input data using convolutional kernels of different sizes to extract features of different scales.

[0075] Dilation layer: Performs dilation operations on the feature maps output by the convolutional layer to obtain a larger receptive field.

[0076] Fusion layer: Fuses the feature maps of different scales to obtain a more abundant feature representation.

[0077] The upsampling decoder is used to decode and upsample the context feature matrix in ascending order of size to obtain the semantic segmentation result map of the remote sensing image. The semantic segmentation result map includes the predicted labels of each pixel in the remote sensing image.

[0078] Specifically, the upsampling decoder is used to decode and upsample the context feature matrix of size h / 16 to obtain the output feature matrix x of size h / 8 d1 , and concatenate x d1 with the context feature matrix of size h / 8 along dimension 1 using the torch.cat function to form a new feature matrix x m1 ;

[0079] Decode and upsample x m1 to obtain the output feature matrix x of size h / 4 d2 , concatenate x d2 with the context feature matrix of size h / 4 along dimension 1 using the torch.cat function to form a new feature matrix x m2 , perform matrix multiplication on the second decoding index and the feature matrix x m2 and output the feature matrix x n2 ;

[0080] Decode and upsample x n2 to obtain the output feature matrix of size h / 2 and then concatenate it with the context feature matrix of size h / 2 along dimension 1 using the torch.cat function to form a new feature matrix x m3 , perform matrix multiplication on the first decoding index and the feature matrix x m3 and output the feature matrix x n3 ;

[0081] Decode and upsample x n3 to obtain a feature matrix x of size h d4 , and then d4 perform a convolution on x and input it into the softmax activation function to obtain the semantic segmentation result map.

[0082] Based on the above technical solution, a remote sensing image semantic segmentation method based on codec-indexed edge representation provided by the present invention uses a deep learning classification framework based on images in the segmentation model, and inputs the remote sensing image into the model for training and prediction. First, the multi-scale feature encoder is used to extract and fuse features of the remote sensing image. Second, the separable pyramid unit with dilated convolutions of different dilation rates is used to capture the spatial context information of multiple scales of the initial feature map in parallel. Third, the two largest-sized feature maps are input into the index generation module to extract the encoded index and decoded index containing edge feature information, and the encoded index is incorporated into the feature matrix in the form of a matrix product. Fourth, the smallest-sized feature map is decoded and upsampled four times. The output matrix of each decoding and upsampling is fused with the corresponding-sized feature map in a skip connection manner, and the two decoded indexes are incorporated into the results of the second and third decoding and upsampling in the form of a matrix product. Finally, the result of the fourth decoding and upsampling is classified using a transposed convolution and a softmax activation function once to output the final semantic segmentation result map. Using the remote sensing image semantic segmentation model based on codec-indexed edge representation improves the feature extraction and processing of the edge information of remote sensing ground objects, and enhances the recognition accuracy of small-sized objects and complex boundary information in remote sensing images.

[0083] In this embodiment, real remote sensing image data is used for experiments, and two publicly available real remote sensing image datasets are used to test and illustrate a remote sensing image semantic segmentation model based on codec-indexed edge representation provided by the present invention, as well as to analyze and evaluate the application effect.

[0084] 1. Dataset and parameter setting

[0085] In this embodiment, two publicly available high-resolution remote sensing image datasets (Potsdam dataset and Vaihingen dataset) in the ISPRS 2D semantic annotation competition are used for experiments and analysis. The datasets use digital surface models (DSM) generated by high-resolution orthophotos and corresponding dense image matching techniques.

[0086] The ISPRS Vaihingen dataset consists of a total of 33 images of varying sizes, with an average size of 2494×2064. The spatial resolution of the images is 9 cm, and the images contain three bands: near-infrared (NIR), red (R), and green (G). The label categories include impervious surfaces, buildings, low vegetation, trees, cars, and others, a total of 6 categories.

[0087] The ISPRS Potsdam dataset contains 38 images with a pixel size of 6000×6000. The spatial resolution of the images is 5 cm, and it uses three bands: red (R), green (G), and blue (B). The label categories and quantities are the same as those in the Vaihingen dataset.

[0088] When training the model, the batch size is set to 16, and each training runs for 300 epochs. The learning rate is dynamically adjusted using cosine annealing. The initial learning rate is set to 1e-3, the learning rate decay coefficient is 0.2, and the learning rate decay interval is 5. The AdamW optimizer is used to optimize the parameters.

[0089] 2. Experimental evaluation metrics

[0090] Overall Accuracy (OA) is a performance metric used to evaluate classification models. In the image semantic segmentation task, it refers to the ratio of the number of correctly classified pixels to the total number of pixels. Its calculation formula is:

[0091]

[0092] F1 score and mF1 score are metrics for measuring the performance of classification models and are commonly used to evaluate the accuracy of binary or multi-classification models. Intersection over Union (IoU) and Mean Intersection over Union (mIoU) are commonly used metrics for measuring the performance of object detection and semantic segmentation models.

[0093] 3. Analysis and evaluation of experimental results

[0094] The results of the remote sensing image semantic segmentation model based on codec-indexed edge representation provided in this embodiment in experiments using two sets of remote sensing image data are shown in Tables 1 and 2.

[0095] Table 1 Comparative experimental results table for Vaihingen dataset

[0096]

[0097] Table 2 Comparative experimental results table for Potsdam dataset

[0098]

[0099] The experiment introduced the DCNN-based models UNet and SegNet, the improved Transformer-based model TransUNet, and the CapsUNet model using the same dataset. According to the classification results, the following conclusions can be analyzed:

[0100] As can be seen from the table, the segmentation model can achieve the best results. Due to the limited amount of remote sensing image data, the experimental results of the TransUNet model are poor. Transformer is used as the encoder in the TransUNet model to present the modeled long-range dependencies and add low-level detail information to the feature maps in the decoder through skip connections. However, since the Transformer model requires a large amount of data for training, TransUNet is not as good as the remote sensing image semantic segmentation model based on codec-indexed edge representation proposed in this invention in this experiment. Comparing the experimental results of the improved DCNN-based models (Unet, SegNet, CapsUNet) and the improved Transformer-based model TransUNet, the model proposed in this invention has achieved better performance.

[0101] Overall beneficial effects:

[0102] This invention provides a high-resolution remote sensing image semantic segmentation method based on codec-indexed edge representation. The multi-scale feature encoder in the remote sensing image semantic segmentation model based on codec-indexed edge representation extracts the multi-scale semantic features of the image. The separable pyramid unit captures the spatial context information in parallel. By extracting the codec index containing edge information, the segmentation effect of the remote sensing ground object edge information is strengthened, the feature extraction and processing of the remote sensing ground object edge information are improved, the recognition accuracy of small-size objects and complex boundary information in the remote sensing image is enhanced, and the accurate semantic segmentation of the remote sensing ground object edge is realized.

[0103] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A semantic segmentation method for high - resolution remote sensing images based on codec - indexed edge representation, characterized in that, including, Step 1: Obtain a remote sensing image set, augment the remote sensing image set, where the augmentation is to rotate the remote sensing images at arbitrary angles and store them in the remote sensing image set, perform normalization processing on the remote sensing images respectively, and divide the remote sensing image set into a training set and a test set. Step 2: Construct a remote sensing image semantic segmentation model based on an encoder-decoder indexed edge representation. After training the remote sensing image semantic segmentation model based on the encoder-decoder indexed edge representation according to the training set, obtain the predicted label of each pixel in the remote sensing image, calculate the loss based on the true label and the predicted label of each pixel, and determine whether the value of the loss satisfies the threshold. If it does not satisfy the threshold, optimize the parameters of the remote sensing image semantic segmentation model based on the encoder-decoder indexed edge representation according to the difference between the value of the loss and the threshold. If it satisfies the threshold, obtain the trained remote sensing image semantic segmentation model based on the encoder-decoder indexed edge representation. The remote sensing image semantic segmentation model based on the encoder-decoder indexed edge representation includes a multi-scale feature encoder, a separable pyramid unit, an encoder-decoder indexed edge representation unit, and an upsampling decoder. The multi-scale feature encoder is used to generate four initial feature matrices according to the size h of the remote sensing image, and the sizes of the four initial feature matrices are h / 2, h / 4, h / 8, and h / 16 respectively. The separable pyramid unit is used to obtain four context feature matrices according to the four initial feature matrices of the remote sensing image, and the sizes of the four context feature matrices are h / 2, h / 4, h / 8, and h / 16 respectively. The encoder-decoder indexed edge representation unit is used to obtain a first encoding index and a first decoding index according to the context feature matrix with a size of h / 2, obtain a second encoding index and a second decoding index according to the context feature matrix with a size of h / 4, fuse the first encoding index with the context feature matrix with a size of h / 2, fuse the second encoding index with the context feature matrix with a size of h / 4, and obtain the fused context feature matrix with a size of h / 2 and the context feature matrix with a size of h / 4. The upsampling decoder is used to decode and upsample the four context feature matrices in ascending order of size to obtain the semantic segmentation result map of the remote sensing image, and the semantic segmentation result map includes the predicted label of each pixel in the remote sensing image. Step 3: Obtain the processed remote sensing image and input it into the trained remote sensing image semantic segmentation model based on the encoder-decoder indexed edge representation, and output the semantic segmentation result map of the remote sensing image.

2. The semantic segmentation method for high - resolution remote sensing images based on codec - indexed edge representation according to claim 1, characterized in that, The upsampling decoder is used to decode and upsample the context feature matrix with a size of h / 16 to obtain an output feature matrix x with a size of h / 8 d1 , and x d1 and the context feature matrix with a size of h / 8 are concatenated along dimension 1 using the torch.cat function to form a new feature matrix x m1 ; Decode and upsample x m1 to obtain an output feature matrix x with a size of h / 4 d2 , and concatenate x d2 with a context feature matrix of size h / 4 along dimension 1 using the torch.cat function to form a new feature matrix x m2 , perform a matrix multiplication operation on the second decoding index and the feature matrix x m2 and then output the feature matrix x n2 ; For x n2 Perform decoded upsampling. After obtaining the output feature matrix with a size of h / 2, concatenate it with the context feature matrix of size h / 2 along dimension 1 using the torch.cat function to form a new feature matrix x m3 , perform matrix multiplication on the first decoding index and the feature matrix x m3 to output the feature matrix x n3 ; Decode and upsample x n3 to obtain a feature matrix x of size h d4 , perform a convolution on x d4 and input it into the softmax activation function to obtain the semantic segmentation result map.

3. The semantic segmentation method for high - resolution remote sensing images based on codec - indexed edge representation according to claim 1, characterized in that, Calculating the loss based on the true label and the predicted label of each pixel includes calculating the loss according to formula (1). Loss focal = -(1 - p t ) γ log(p t ) (1) Among them, p t is the predicted probability of the true label, and the predicted probability is obtained according to the true label and the predicted label. γ is a hyperparameter, and Loss focal represents the focal loss function.

4. The semantic segmentation method for high - resolution remote sensing images based on codec - indexed edge representation according to claim 1, characterized in that, The multi-scale context feature encoder includes a spatial feature extraction branch, a self-attention feature extraction branch, and a fusion branch. The spatial feature extraction branch is used to extract local feature information of the remote sensing image, the self-attention feature extraction branch is used to extract global feature information of the remote sensing image, and the fusion branch is used to fuse the local feature information and the global feature information according to formula (2). x = concatnate(Conv2d(x ci ), Conv2d(x si )) y = sigmoid(Conv2d(ReLU(Conv2d(AdaptiveAvgPool2d(x))))) x fi = x × reshape(y) (2) where x si represents the feature matrix of the i-th stage of the self-attention feature extraction branch, and x ci represents the feature matrix of the i-th stage of the spatial feature extraction branch, and x fi represents the fused feature, Conv2d(*) represents 2D convolution, AdaptiveAvgPool2d(*) represents the adaptive pooling function, sigmoid(*) represents the sigmoid activation function, ReLU(*) represents the ReLU activation function, concatnate(*) represents concatenating two matrices along dimension 1, and reshape(*) represents the shape change function.

5. The semantic segmentation method for high - resolution remote sensing images based on codec - indexed edge representation according to claim 1, characterized in that, The obtaining of the first encoding index and the first decoding index according to the context feature matrix of size h / 2 includes S11. Represent the context feature matrix with a size of h / 2 as x i , and obtain x i 's shape parameters, where the shape parameters include the batch size value batchsize, the number of channels c, the height h, and the width w S12. Input x i into the Conv2d function to obtain x i1 . Input x i1 into the BatchNorm2d function to obtain x i2 . Input x i2 into the BatchNorm2d function to obtain x i3 . Input x i3 into the BatchNorm2d function to obtain x i4 . S13. For x i1 x i2 x i3 x i4 perform max pooling operations respectively to obtain four initial indices x1, x 2 x 3 x 4 , S14. Concatenate the initial indices x1, x2, x3, and x4 along dimension 1 into a new matrix through the torch.cat function and pass it to the sigmoid activation function to obtain the initial decoding index y; S15. The initial decoding index y passes through the softmax function to obtain the initial encoding index z, and the view function is used to adjust the shape parameters of the initial decoding index y and the initial encoding index z. The adjustment is to adjust the shape parameters to batch size, c×4, h / 2, w / 2, and obtain the adjusted initial decoding index y and initial encoding index z S15. Use the pixel_shuffle function to reorganize the adjusted initial decoding index y and initial encoding index z into the size before adjustment to obtain the first encoding index and the first decoding index

Citation Information

Patent Citations

  • Remote sensing image road segmentation method based on contextual information and multi-scale feature fusion

    CN113850825A

  • Method and system of extraction of impervious surface of remote sensing image

    US20200026953A1