A method for classifying ground objects in optical remote sensing images guided by boundaries

Through edge extraction and cross-scale fusion modules, the edge guidance method of Transformer architecture is used to solve the problems of spatial information loss and long-distance information not being used in convolutional neural networks, and the geographic classification accuracy of optical remote sensing images is improved.

CN116486158BActive Publication Date: 2025-07-04BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310457467.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-24
Publication Date
2025-07-04
Estimated Expiration
2043-04-24

AI Technical Summary

Technical Problem

The existing land object classification method based on convolutional neural networks has problems such as loss of spatial information and insufficient utilization of long-distance and global context information in optical remote sensing images, resulting in insufficient classification accuracy.

Method used

The edge extraction module is used to fuse low-level local edge information and high-level global position information, fuse with the backbone features through the edge guide module, and feature aggregation is performed based on the Transformer architecture by using the cross-scale fusion module to enhance feature representation and utilize long-distance and global context information.

Benefits of technology

It improves the accuracy of geographic classification of optical remote sensing images, makes up for the loss of spatial information, enhances the feature representation ability, and achieves a better segmentation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116486158B_ABST
    Figure CN116486158B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for classifying ground objects in optical remote sensing images guided by boundaries. The classification of ground objects in optical remote sensing images is a hot research issue in the interpretation of remote sensing images, which plays an important role in both civilian and defense fields. There is a practical need to accurately and timely obtain ground object information from remote sensing images. In the process of learning features layer by layer, the usual ground object classification method based on convolutional neural networks will lose spatial information and pay less attention to long-distance information and global context information in the images. The purpose of the present invention is to better utilize long-distance information and global context information based on the Transformer architecture with the help of the attention mechanism to achieve better feature fusion, thereby improving the segmentation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of optical remote sensing image imaging, and particularly relates to a method for classifying ground objects in optical remote sensing images based on edge guidance. Background Art

[0002] The development of optical remote sensing image imaging technology has promoted the explosive growth of remote sensing image data, showing a trend of higher resolution and larger swath width. The classification of ground objects in optical remote sensing images is a hot research issue in remote sensing image interpretation, and it plays an important role in civilian and even national defense fields such as land resource investigation, economic analysis, and ecological environment monitoring. Ground object classification requires classifying each pixel in the remote sensing image and dividing the image into regions with different ground object semantic labels, which has important applications in the analysis and interpretation of remote sensing images.

[0003] Traditional ground object classification methods mainly include three aspects: remote sensing image feature extraction, remote sensing image feature selection, and classification algorithms. The feature extraction ability and generalization ability of traditional ground object classification methods are insufficient to solve the problem of multi-ground object classification of high-resolution optical remote sensing images. With the development of deep learning technology, ground object classification methods based on deep learning technology have been greatly developed. The ground object classification method based on convolutional neural network trains a neural network classifier by learning features layer by layer from optical remote sensing images and classifies at the pixel level of the image. During the process of learning features layer by layer, the resolution of the optical remote sensing image will gradually decrease, resulting in the loss of spatial information. At the same time, restricted by the receptive field of the convolutional layer, the long-range information and global context information in the image are lost. The good fusion of the spatial information of the shallow feature map and the semantic information of the deep feature map helps to improve the segmentation result.

[0004] Therefore, it has important practical significance and application value to study how to make up for the loss of spatial information, make full use of the long-range information and global context information of the image, and perform good feature fusion to improve the accuracy of ground object classification. Summary of the Invention

[0005] The object of the present invention is to solve the problems of loss of spatial information and insufficient utilization of long-range and even global context information during the process of ground object classification based on convolutional neural network for optical remote sensing images, so as to obtain good ground object classification results.

[0006] To achieve the above object, the present invention is realized through the following technical solutions:

[0007] For the input image, four feature maps at different levels are extracted, and they are sorted in order from shallow to deep as , , , .

[0008] Input and into the edge extraction module. The edge extraction module first fuses the two feature maps and extracts edge information from the fused features. This module aims to fuse low-level local edge information and high-level global position information, and extract edge information related to the target boundary under explicit boundary supervision.

[0009] Then, the extracted edge information , is successively passed through the edge guidance module and , , , for fusion to obtain , , , . This module aims to fuse edge information with backbone features at all levels to guide feature learning.

[0010] Finally, , , , are input into the cross-scale fusion module to obtain the final segmentation result. This module aims to focus on the global context information of the feature map and perform cross-scale aggregation on the feature maps at four levels based on the Transformer architecture to generate stronger and more effective features.

[0011] Train the constructed land cover classification network, and after training, select the model with the best performance metrics for the land cover classification task of optical remote sensing images.

[0012] According to the above technical solution, the specific steps of this method are as follows:

[0013] Step 1: Extract multi-level features from the input image, namely , , , .

[0014] Step 2: Apply the edge extraction module to extract edge information related to the target from the low-level features containing local edge detail information and the high-level features containing global position information under the supervision of the target boundary .

[0015] Good edge priors are helpful for localization and segmentation in land cover classification. Low-level features contain rich edge details but lack high-level semantics. Therefore, this module intends to combine low-level features and high-level features to model and extract edge information.

[0016] Step 3: Use multiple edge guidance modules to process the result obtained in Step 2 and the backbone features at each level , , , for aggregation to obtain , , , , so as to guide feature learning, thereby enhancing the boundary representation and making up for the loss of spatial information.

[0017] This module aims to introduce boundary-related edge information into representation learning to enhance the feature representation with the semantic of the target structure. As we all know, different feature channels usually contain different semantics. Therefore, in order to achieve good fusion and obtain a powerful representation, this module introduces a local channel attention mechanism to explore the interaction between channels.

[0018] Step 4: Adopt a cross-scale fusion module to perform cross-scale aggregation on , , , and predict the segmentation result map.

[0019] This module aims to utilize global semantic information and long-range semantic information, and adopts a Transformer architecture to obtain high-resolution and rich semantic representations, which is crucial for subsequent segmentation. This module mainly uses a cross-attention mechanism to achieve the aggregation of feature maps with different resolutions from coarse to fine, and thus can make full use of long-range and global context information. Brief Description of the Drawings

[0020] Figure 1 This is a structural diagram of a ground object classification network for boundary-guided optical remote sensing images provided by an embodiment of the present invention.

[0021] Figure 2 is an edge extraction module.

[0022] Figure 3 is an edge guidance module.

[0023] Figure 4 is a cross-scale fusion module. Detailed Embodiment

[0024] The present invention will be described in detail below with reference to the accompanying drawings and by way of examples.

[0025] The present invention provides a method for classifying ground objects in boundary-guided optical remote sensing images, constructs a ground object classification network, and is used to perform the following steps:

[0026] Step 1: For the input image, Res2Net-50 is used as the backbone network to extract features at four different levels, which are arranged in the order from shallow to deep as , , , .

[0027] Step 2: Apply the edge extraction module, whose structure is as shown in Figure 2 , to mine edge information related to the target from the low-level features containing local edge details ( ) and the high-level features containing global position information under target boundary supervision ( ). .

[0028] Specifically, first use two 1×1 convolutional blocks to compress the number of channels of and to 64 and 256 respectively, obtaining features and . Then, upsample to get , where has the same size as , and concatenate and fuse and along the channel dimension through the concat operation. After that, obtain the edge information through two 3×3 convolutional blocks, one 1×1 convolutional block, and a Sigmoid function.

[0029] Step 3: Utilize multiple edge guidance modules, whose structure is as shown in Figure 3 , to aggregate with the backbone features at all levels to guide feature learning, thereby enhancing the boundary representation.

[0030] Specifically, given the input feature and the edge information , first use element-wise multiplication with a residual connection and a 3×3 convolution to obtain the initial fusion feature, which can be expressed as:

[0031] (1)

[0032] where D represents downsampling, is a 3×3 convolution. represents element-wise multiplication, represents element-wise addition. Subsequently, use channel global average pooling (GAP) to aggregate the convolutional feature Then, the corresponding channel attention weights are obtained through one-dimensional convolution and the Sigmoid function. Then, the channel attention weights are multiplied by the initial fusion features, and the number of channels is reduced through 1×1 convolution to obtain the final fusion features:

[0033] (2)

[0034] where is a 1×1 convolution, is a one-dimensional convolution with a kernel size of k, represents the Sigmoid function. The kernel size k can be adaptively set to where represents the nearest odd number, and C is the number of channels. The kernel size is proportional to the channel size. Obviously, this attention strategy can highlight the key channels, suppress redundant channels or noise, thereby enhancing the semantic representation.

[0035] Step four, adopt a cross-scale fusion module, the structure of which is as shown in Figure 4 to aggregate , , , and predict to obtain the segmentation result map.

[0036] Specifically, given the feature maps , , , , the cross-scale module first segments the feature map to obtain several patches, and inputs these patches into Mix-FFN to construct the query . The output of Mix-FFN can be defined as:

[0037] (3)

[0038] where represents several patches obtained by segmenting the feature map , MLP represents a multi-layer perceptron, GELU represents the Gelu activation function, is a 3×3 convolution.

[0039] Then, , , are fused through element-wise addition. Specifically, is upsampled to the size and added element-wise to to obtain . Then, is upsampled to the size and added to Element-wise addition is performed to obtain .

[0040] Subsequently, is segmented to obtain a number of patches, and key-value 、 is passed to the cross-attention module and performs cross-scale information aggregation. Then it is passed to Mix-FFN to obtain the fused information. The output of the cross-scale module is defined as:

[0041] (4)

[0042] Wherein, is the query obtained from the feature map , is the mixed feed-forward neural network. and are the key and value obtained from the feature .

[0043] In summary, the above is only a preferred embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for classifying ground objects in optical remote sensing images guided by boundaries, characterized in that The steps of the method include: Step 1: Extract multi-level features from the input image, namely , , , ; Step 2: Apply the edge extraction module to extract edge information related to the target from the low-level features containing local edge detail information and the high-level features containing global position information under the supervision of the target boundary ; ;​ Step 3: Use multiple edge guidance modules to process the result obtained in Step 2 and the backbone features at each level , , , for aggregation to obtain , , , to guide feature learning, thereby enhancing the boundary representation and compensating for the loss of spatial information; Step 4: Use a cross-scale fusion module to perform cross-scale aggregation on , , , and predict the segmentation result map.

2. The method for classifying ground objects in an optical remote sensing image guided by a boundary according to claim 1, wherein In Step 1, for the input image, Res2Net-50 is used as the backbone network to extract features at four different levels, which are arranged in the order from shallow to deep as , , , .

3. The method for classifying ground objects in an optical remote sensing image guided by a boundary according to claim 1, wherein In step two, first use two 1×1 convolution blocks to compress the number of channels of and to 64 and 256 respectively, obtaining feature maps and . Then, upsample to obtain , where has the same size as , and concatenate and fuse and along the channel dimension through a concat operation. After that, obtain the edge information through two 3×3 convolution blocks, one 1×1 convolution block, and a Sigmoid function.

4. A method for classifying ground objects in an optical remote sensing image guided by a boundary, characterized in that In step three, the final fused feature is expressed as: Among them, is a 1×1 convolution, is a one-dimensional convolution with a kernel size of k, represents the Sigmoid function; the kernel size k can be adaptively set to , where represents the nearest odd number, C is the number of channels; the kernel size is proportional to the channel size.

5. A method for classifying ground objects in a boundary-guided optical remote sensing image according to claim 1, characterized in that In step four, the output of Mix-FFN is defined as: Among them, represents a number of patches obtained by segmenting the feature map , MLP represents a multi-layer perceptron, and GELU represents the Gelu activation function, is a 3×3 convolution.

6. The method for classifying ground objects in an optical remote sensing image guided by a boundary according to claim 1, wherein In step four, the output of the cross-scale module is defined as: Among them, is a query obtained from the feature map and is a hybrid feedforward neural network; and are the key and value obtained from the feature respectively.

Citation Information

Patent Citations

  • Rapid saliency detection method based on multi-scale feature attention mechanism

    CN110929735A

  • River and lake remote sensing image segmentation method based on deformable convolution and self-attention model

    CN115601549A