Remote sensing image urban landscape classification method and system based on self-attention network

By segmenting the remote sensing image into image blocks and building a self-attention network, the problems of low generalization performance and difficult to obtain global context dependencies in traditional methods are solved, and efficient remote sensing image urban landscape classification is achieved.

CN120298759APending Publication Date: 2025-07-11SUN YAT SEN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510335272.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Among the existing remote sensing image urban landscape classification methods, the traditional method based on manual features has low generalization performance, while the deep network architecture based on CNN is difficult to obtain the dependencies of the global context of the image.

Method used

The remote sensing image urban landscape classification method based on self-attention network is adopted. By segmenting the remote sensing image into image blocks, a self-attention network of the visual Transformer model is constructed, including an embedding layer, an encoder and a task head. The multi-head self-attention mechanism and data enhancement strategy are used to capture the similarity between landscape elements and generate classification labels.

Benefits of technology

The generalization performance of urban landscape classification in remote sensing images is improved, the misjudgment rate is reduced, and the global dependence relationship of images is captured through self-attention networks is enhanced, and the accuracy of classification is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298759A_ABST
    Figure CN120298759A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image urban landscape classification method and system based on a self-attention network, and relates to the technical field of remote sensing image recognition and classification, and the method comprises the steps: obtaining a remote sensing image and a classification mark, and carrying out the data enhancement; segmenting the remote sensing image in the data set to obtain an image block sequence; constructing a self-attention network based on a visual Transform model, wherein the self-attention network comprises an embedded layer, an encoder and a task head; inputting an image block sequence and a classification mark to the self-attention network, obtaining an embedded representation in an embedding layer, then connecting the embedded representation with the classification mark, and obtaining a result embedded image block sequence with a mark; embedding the marked result into an image block sequence, inputting the image block sequence into an encoder, and generating a feature representation and an attention score; and inputting the feature representation into a task head, generating a classification label according to the attention score, and completing remote sensing image urban landscape classification. The method is high in generalization performance, classification is performed based on the global context dependency relationship of the image, and the accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image recognition and classification, and more specifically, to a method and system for classifying urban landscapes in remote sensing images based on a self-attention network. Background Art

[0002] Remote sensing is a science that can collect information about objects without direct physical contact, usually through satellites, airplanes or drones. Examples of remote sensing applications include geological surveys, land management, urban planning, etc. With the development of satellite sensors, the spatial and temporal resolutions of remote sensing images have been improved, providing opportunities to analyze fine details on the Earth's surface. The classification of urban landscapes in remote sensing images is an active research field in recent years. Its task is to capture images and correctly label them into predefined categories, and it faces the challenging problem of effectively interpreting remote sensing images.

[0003] Early works on the classification of urban landscapes in remote sensing images were based on manually extracted handcrafted features, including local binary patterns, histograms of oriented gradients, and scale-invariant feature transform. Traditional interpretation relied on encoding handcrafted features using different models, such as bag-of-words, Fisher vectors, or locally aggregated descriptor vectors. With the rapid development of computer science and technology, deep learning methods have achieved great success in many application fields such as the interpretation of remote sensing image scenes. In particular, convolutional neural networks (CNNs) have outperformed the above traditional methods in many applications. The main advantage of these methods is to provide an end-to-end solution with minimal feature engineering. In addition to CNNs, a new type of deep learning model called Transformer has also been proposed and is highly favored in the field of computer vision. Transformer relies on a simple but powerful process called the self-attention mechanism, which focuses attention on certain parts of the input to obtain more effective results. Currently, they are considered the most advanced models in sequence data, especially in natural language processing methods such as machine translation, language modeling, and speech recognition.

[0004] Inspired by the success of Transformer in the field of natural language processing, new research attempts to apply Transformer directly to images. However, this is a challenging task because in self-attention applications, each pixel needs to attend to all other pixels. For images, the computational cost is very high because images contain a large number of pixels. Researchers have tried various methods to apply Transformer to images. Some work combines the CNN architecture with self-attention. For example, by replacing some convolutional layers with self-attention layers to enhance CNN, the image classification effect has been improved. However, this method faces high computational costs because the large size of the image leads to a significant increase in the time complexity of self-attention. Therefore, how to block-process the input remote sensing image and add corresponding position encoding information to interpret urban landscape information from the remote sensing image is a challenging task.

[0005] The existing methods for classifying urban landscapes in remote sensing images still have the following drawbacks:

[0006] (1) Traditional remote sensing urban landscape classification methods based on handcrafted features are limited by expert domain knowledge. The models have the advantages of being simple and easy to apply, but their performance has the disadvantage of low generalization.

[0007] (2) The deep network architecture based on CNN can capture the deep semantic features of urban landscapes in remote sensing images. CNN performs excellently in local feature induction, but it is difficult to obtain the dependency relationships of the global context of the image. Summary of the Invention

[0008] The object of the present invention is to overcome the defects that the traditional remote sensing urban landscape classification method based on handcrafted features in the prior art has low generalization performance, and the deep network architecture based on CNN is difficult to obtain the dependency relationships of the global context of the image, and to provide a method and system for classifying urban landscapes in remote sensing images based on a self-attention network.

[0009] To achieve the above object, the technical solution of the present invention is as follows:

[0010] In the first aspect, the present invention provides a method for classifying urban landscapes in remote sensing images based on a self-attention network, including the following steps:

[0011] Obtain a remote sensing image containing an urban landscape and classification labels and perform data augmentation to obtain a remote sensing image dataset;

[0012] Segment the remote sensing images in the remote sensing image dataset into a plurality of image patches to obtain an image patch sequence;

[0013] Construct a self-attention network based on a vision Transformer model, where the self-attention network includes an embedding layer, an encoder, and a task head;

[0014] Input the sequence of image patches and the classification token into the self-attention network. After obtaining the embedded representation of the sequence of image patches through the embedding layer, concatenate it with the classification token to obtain a result embedded image patch sequence with tokens.

[0015] Input the result embedded image patch sequence with tokens into the encoder to generate a feature representation and attention scores.

[0016] Input the feature representation into the task head and generate classification labels according to the attention scores to complete the classification of urban landscapes in remote sensing images.

[0017] Preferably, the method for data augmentation includes:

[0018] Randomly draw two samples (X i , y i ) and (X j , y j ) from the dataset each time, and use the following equation to insert the two samples X i and X j and their related label encodings y i and y j , and repeat several times to expand the dataset:

[0019]

[0020] where λ ∈ [0, 1].

[0021] Preferably, the method for data augmentation further includes cropping, stitching, and color transformation of the remote sensing images in the dataset.

[0022] Preferably, the method for splitting the remote sensing images in the remote sensing image dataset into several image patches to obtain an image patch sequence is specifically as follows:

[0023] For a remote sensing image X of size h×w×c, where h is the height, w is the width, and c represents the number of channels, extract each image patch of size p×p×c from the image X, where p is the side length of the image patch, to form an image patch sequence of length n (x1, x2,..., x n ), where hw / p 2 = 1.

[0024] Preferably, the process of obtaining the result embedded image patch sequence with tokens is as follows:

[0025] In the embedding layer, first, the learned embedding matrix E is used to linearly project the sequence of image patches into a vector of dimension d to obtain the embedded representation of the sequence of image patches. Then, the embedded representation of the sequence of image patches is concatenated with the corresponding classification token. Since the Vision Transformer model treats the embedded sequence of image patches as a set of image features and does not consider the positional order of these image patches, in order to keep the spatial arrangement of the image patches the same as the original remote sensing image, positional information is encoded and appended to the embedded representation of the sequence of image patches. Therefore, the resulting embedded image patch sequence z0 with tokens is defined as follows:

[0026]

[0027] where, v class represents the classification token, E pos represents the positional information, n is the length of the sequence of image patches, p is the side length of each image patch in the sequence of image patches, and c is the number of channels of each image patch in the sequence of image patches.

[0028] Preferably, the encoder includes L identical layers, each layer including a multi-head self-attention block and a feed-forward dense block. A normalization layer LN is connected to the inputs of both the multi-head self-attention block and the feed-forward dense block. The inputs and outputs of the multi-head self-attention block and the feed-forward dense block are connected through residual skip connections.

[0029] Preferably, the multi-head self-attention module includes an input linear layer, a multi-head self-attention layer, a connection layer, and an output linear layer. The input linear layer includes a number of linear layers, and each linear layer is used to generate a query matrix, a key matrix, and a value matrix by performing different linear transformations on the input sequence of patch embeddings. The multi-head self-attention layer includes the same number of self-attention layers as the number of linear layers in the input linear layer. Each self-attention layer is used to calculate the attention result based on the query matrix, the key matrix, and the value matrix. All the attention results are concatenated together in the connection layer and finally projected to the required dimension through a feed-forward layer with learnable weights W, i.e., the output linear layer. This process is represented by the following equation:

[0030]

[0031] where, MSA represents the result of the multi-head self-attention layer, Z represents the input resulting embedded image patch sequence with tokens, Concat() represents the concatenation operation, h represents the number of heads in the multi-head self-attention layer, i.e., the number of self-attention layers, and D K represents the dimension of the key vectors in the key matrix;

[0032] The feed-forward dense block includes two fully-connected layers, and there is a GeLU activation function layer between the two fully-connected layers.

[0033] Preferably, the self-attention layer includes matrix multiplication operation, scaling operation, masking operation, and SoftMax function weighting operation. The query matrix and the key matrix first undergo matrix multiplication operation to obtain the dot product of the query matrix and the key matrix. Then, after the dot product undergoes scaling operation and masking operation, it is weighted by the SoftMax function to obtain weights, and finally, matrix multiplication operation is performed with the value matrix.

[0034] Preferably, at the end of the encoder, the feature representation of the first element in the result of the encoder is taken and passed to the task head to generate classification labels, completing the urban landscape classification of remote sensing images.

[0035] In a second aspect, the present invention provides a remote sensing image urban landscape classification system based on a self-attention network for applying the remote sensing image urban landscape classification method described in the above technical solution. The system includes:

[0036] A data acquisition module for acquiring remote sensing images containing urban landscapes and classification labels and performing data augmentation to obtain a remote sensing image data set;

[0037] An image segmentation module for segmenting the remote sensing images in the remote sensing image data set into a plurality of image patches to obtain an image patch sequence;

[0038] A self-attention network construction module for constructing a self-attention network based on a Vision Transformer model. The self-attention network includes an embedding layer, an encoder, and a classifier;

[0039] A self-attention network classification module for inputting the image patch sequence and the classification label into the self-attention network. After obtaining the embedded representation of the image patch sequence through the embedding layer, it is concatenated with the classification label to obtain a labeled result embedded image patch sequence; inputting the labeled result embedded image patch sequence into the encoder to generate feature representations and attention scores; inputting the feature representations into the classifier and generating classification labels according to the attention scores, completing the urban landscape classification of remote sensing images.

[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0041] The present invention uses a Transformer network architecture with a multi-head self-attention mechanism to globally represent the dependency relationships of the input remote sensing images, thereby capturing the similarities between landscape elements, and provides sufficient learning samples for the training of the self-attention network through a data augmentation strategy, with high generalization performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1Schematic diagram of the steps of the remote sensing image urban landscape classification method based on the self-attention network in Embodiment 1 of the present application;

[0043] Figure 2 Schematic diagram of the structure of the self-attention network in Embodiment 1 of the present application;

[0044] Figure 3 Attention heat map generated by the self-attention network in Embodiment 1 of the present application according to the input image;

[0045] Figure 4 Schematic diagram of the classification accuracy of the self-attention network in Embodiment 1 of the present application under different-layer Transformer architectures. Detailed implementation manners

[0046] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments may be combined with each other.

[0047] Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.

[0048] Embodiment 1

[0049] Please refer to Figure 1 , Embodiment 1 of the present application provides a remote sensing image urban landscape classification method based on the self-attention network, including the following steps S1 - S6:

[0050] S1: Obtain remote sensing images containing urban landscapes and classification labels and perform data augmentation to obtain a remote sensing image data set.

[0051] A simple but effective data augmentation strategy is provided in Embodiment 1 of the present application, which can increase the size and diversity of the urban landscape training data set. (X i , y i ) and (X j , y j ) are two samples randomly drawn from the training data. By using the following equation to insert two samples X i and X j and their related unique label encodings y i and y j to expand the data set, each time two samples (X i , y i ) and (X j , y j) Insert two samples X using the following equation i and X j and their associated label encodings y i and y j , repeating several times to expand the dataset:

[0052]

[0053] where λ ∈ [0, 1].

[0054] The method for data augmentation in this Example 1 further includes cropping, stitching, and color transformation of the remote sensing images in the dataset to generate more datasets.

[0055] S2: Segment the remote sensing images in the remote sensing image dataset into several image patches to obtain an image patch sequence.

[0056] Let represent a set of r remote sensing images, where X i is the image, and y i is the corresponding classification label y i ∈ {1, 2, 3,..., m}, where m is the number of defined classes in this set. First, the remote sensing image X from the dataset (for simplicity, the image index i is omitted here) is subdivided into non - overlapping patches. For an image X of size h×w×c (where h is the height, w is the width, and c represents the number of channels), image patches are extracted, each with a size of p×p×c (where p is the side length of the image patch), forming an image patch sequence (x1, x2,..., x n ), where h w / p 2 = 1. Usually, the image patch size p×p is selected to be 16×16 or 32×32.

[0057] S3: Construct a self - attention network based on the Vision Transformer model, and the self - attention network includes an embedding layer, an encoder, and a classifier (task head).

[0058] The goal of the Vision Transformer model is to learn the mapping from the image patch sequence to the corresponding semantic labels.

[0059] Please refer to Figure 2 , Figure 2 which is the structural schematic diagram of the self - attention network of this Application's Example 1. In this Application's Example 1, the Vision Transformer model, i.e., the self - attention network, consists of a linear embedding layer ( Figure 2 Linear Embeddung (E) in Figure 2 ), an encoder ( Figure 2It consists of an encoder and a classifier. The encoder includes L identical layers, each layer including a multi-head self-attention block (MSA) and a feed-forward dense block (MLP). A normalization layer (Layer Norm, LN) is connected to the inputs of both the multi-head self-attention block and the feed-forward dense block. The inputs and outputs of the multi-head self-attention block and the feed-forward dense block are each connected through a residual skip connection.

[0060] The multi-head self-attention module includes an input linear layer, a multi-head self-attention layer, a connection layer, and an output linear layer. The input linear layer includes several linear layers. The multi-head self-attention layer includes the same number of self-attention layers (i.e., self-attention heads, Figure 2 in Scaled Dot Product Attention) as the number of linear layers in the input linear layer. Each self-attention layer includes a matrix multiplication operation (MatMul), a scaling operation (Scale), a masking operation (Mask), and a SoftMax function weighting operation.

[0061] The feed-forward dense block includes two fully connected layers (FC), and there is a layer of GeLU activation function layer in the middle of the two fully connected layers.

[0062] S4: Input the image patch sequence and the classification token into the self-attention network. After obtaining the embedded representation of the image patch sequence through the embedding layer, concatenate it with the classification token to obtain a labeled result embedded image patch sequence.

[0063] The image patch sequence of length n (x1, x2,..., x n ) After the image patch sequence is flattened, it is input into the linear embedding layer ( Figure 2 in Linear Embeddung (E)). In this linear embedding layer, first use the learned embedding matrix E to linearly project the image patch sequence into a vector of dimension d to obtain the embedded representation of the image patch sequence. Then concatenate the embedded representation of the image patch sequence with the learnable classification token v required for performing the classification task input. cllass Since the original Transformer treats the embedded image patches as a set of image features and does not consider the positional order of these image patches. To keep the spatial arrangement of the image patches the same as the original image, the positional information E pos is encoded and appended to the embedded representation of the image patches. Therefore, the labeled result embedded image patch sequence z0 is defined as follows:

[0064]

[0065] where, v classRepresents a classification label, E pos Represents position information, where n is the length of the image patch sequence, p is the side length of each image patch in the image patch sequence, and c is the number of channels of each image patch in the image patch sequence.

[0066] S5: Input the marked result into the encoder by embedding it into the image patch sequence to generate a feature representation and an attention score.

[0067] The processing flow of this process is as follows:

[0068] z l ' = MSA(LN(z l-1 )) + z l-1 , l = 1, 2, 3,..., L

[0069] z l = MLP(LN(z l ')) + z l ', l = 1, 2, 3,..., L

[0070] At the end of the encoder, the attention scores generated by the encoder and the feature representations of each image patch element are passed to the task head (classifier).

[0071] S6: Input the feature representation into the classifier and generate a classification label according to the attention score to complete the urban landscape classification of remote sensing images.

[0072] The task head (classifier) generates a classification label according to the feature representation of the element ranked first in the attention scores in step S5 to complete the urban landscape classification of remote sensing images. In a specific embodiment, the classifier includes a pooling layer and a fully connected layer.

[0073] Embodiment 2

[0074] Embodiment 2 of this application is based on Embodiment 1 and further illustrates the multi-head self-attention block (MSA) in Embodiment 1 as follows:

[0075] The MSA block in the encoder is the core component of the Transformer self-attention network. Its role is to determine the relative importance of a single image patch embedding relative to other embeddings in the sequence. This module has four layers: an input linear layer, a self-attention layer, a connection layer (connecting the outputs of multiple self-attention heads), and a final output linear layer. From a high level, attention can be represented by attention weights, which are calculated by taking the weighted sum of all values in the sequence z. The self-attention (SA) head learns the attention weights by calculating the query-key-value scaled dot product. For each element in the input sequence, by multiplying the image patch features with three learned matrices UQKV Multiply them to obtain the Q (query), K (key), and V (value) matrices, as specifically shown below:

[0076]

[0077] To determine the relevance of an element to other elements in the sequence, it is necessary to calculate the dot product of the Q vector of this element and the K vectors of other elements. The calculation result can determine the relative importance of each image patch in the sequence. The self-attention layer (self-attention head) includes matrix multiplication operations, scaling operations, masking operations, and SoftMax function weighting operations. First, the query matrix and the key matrix undergo matrix multiplication operations to obtain the dot product of the query matrix and the key matrix. Then, after the dot product undergoes scaling operations and masking operations, it is weighted through the SoftMax function to obtain weights. Finally, matrix multiplication operations are performed with the value matrix.

[0078] The scaled dot product operation performed by the self-attention layer is similar to the standard dot product, but it uses the dimension D of the key matrix K as the scaling factor. Finally, multiply the value matrix V of the embedding vector of each image patch by the output A of the softmax function operation to obtain the attention score of each image patch. The complete operation is given by the following equation:

[0079]

[0080] SA(z) = A · V

[0081] Please refer to Figure 3 and Figure 4 . The input linear layer includes several linear layers. Each linear layer is used to generate the query matrix, key matrix, and value matrix by performing different linear transformations on the input block embedding sequence. The multi-head self-attention layer includes the same number of self-attention layers as the number of linear layers in the input linear layer. Each self-attention layer is used to calculate the attention result based on the query matrix, key matrix, and value matrix. The attention results of all self-attention layers are concatenated together in the concatenation layer, and finally projected to the required dimension through a feed-forward layer with learnable weights W, that is, the output linear layer. This process is represented by the following equation:

[0082]

[0083] where MSA represents the result of the multi-head self-attention layer, Z represents the input sequence of labeled result embedding image patches, Concat() represents the concatenation operation, h represents the number of heads in the multi-head self-attention layer, that is, the number of self-attention layers, and D K represents the dimension of the key vector in the key matrix.

[0084] Other technical details and steps of this Embodiment 2 are the same as those of Embodiment 1 and will not be elaborated here.

[0085] Example 3

[0086] Based on Example 1 and Example 2, Example 3 of this application provides a remote sensing image urban landscape classification system based on a self-attention network, which is used to apply the remote sensing image urban landscape classification method based on the self-attention network described in Example 1. The system includes:

[0087] A data acquisition module, which is used to acquire remote sensing images containing urban landscapes and perform data augmentation to obtain a remote sensing image dataset;

[0088] An image segmentation module, which is used to segment the remote sensing images in the remote sensing image dataset into several image patches to obtain an image patch sequence;

[0089] A self-attention network construction module, which is used to construct a self-attention network based on a Vision Transformer model. The self-attention network includes an embedding layer, an encoder, and a task head;

[0090] A self-attention network classification module, which is used to input the image patch sequence and classification tokens into the self-attention network. After obtaining the embedded representation of the image patch sequence through the embedding layer, it is concatenated with the classification tokens to obtain a marked result embedded image patch sequence; the marked result embedded image patch sequence is input into the encoder to generate feature representations and attention scores; the feature representations are input into the task head, and classification labels are generated according to the attention scores to complete the remote sensing image urban landscape classification.

[0091] Other technical details and steps of Example 3 are the same as those of Example 1 or Example 2, and will not be elaborated here.

[0092] In summary of the above embodiments, the present invention adopts a data augmentation strategy to generate additional training data, providing sufficient training samples for network training and learning. Secondly, the original input remote sensing image is segmented into multiple image patches and converted into a sequence through technical means such as flattening and positional embedding. Then, the generated sequence is input into several multi-head attention layers to generate the final feature representation. Finally, the first token sequence passing through the deep network is input into the Softmax classification layer to obtain the urban landscape label. The main advantages of the present invention are in two aspects: an advanced Transformer deep network architecture is adopted to capture the global context dependencies of remote sensing image urban landscape elements, thereby reducing the misjudgment rate of urban landscapes; in order to overcome the need for a large number of training samples in the learning and training stages of the deep network, the present invention designs a remote sensing image urban landscape data augmentation strategy, thus providing sufficient learning samples for deep network training.

Claims

1. A method for classifying urban landscapes in remote sensing images based on a self-attention network, characterized in that, It includes the following steps: Obtain a remote sensing image containing urban landscapes and classification labels and perform data augmentation to obtain a remote sensing image dataset; Segment the remote sensing images in the remote sensing image dataset into a number of image patches to obtain an image patch sequence; Construct a self-attention network based on a Vision Transformer model, where the self-attention network includes an embedding layer, an encoder, and a classifier; Input the image patch sequence and classification labels into the self-attention network. After obtaining the embedded representation of the image patch sequence through the embedding layer, concatenate it with the classification labels to obtain a labeled result embedded image patch sequence; Input the labeled result embedded image patch sequence into the encoder to generate feature representations and attention scores; Input the feature representations into the classifier and generate classification labels according to the attention scores to complete the classification of urban landscapes in remote sensing images.

2. The method for classifying urban landscapes in remote sensing images based on the self-attention network according to claim 1, wherein, The method of data augmentation includes: Randomly draw two samples (X i , y i ) and (X j , y j ) from the dataset each time, and use the following equation to insert the two samples X i and X j and their related label encodings y i and y j , and repeat several times to expand the dataset: where λ ∈ [0, 1].

3. The method for classifying urban landscapes in remote sensing images based on the self-attention network according to claim 2, characterized in that, The method of data augmentation also includes cropping, splicing, and color transformation of the remote sensing images in the dataset.

4. The method for classifying urban landscapes in remote sensing images based on a self-attention network according to claim 1, wherein, The method of segmenting the remote sensing images in the remote sensing image dataset into a number of image patches to obtain an image patch sequence is specifically as follows: For a remote sensing image X of size h×w×c, where h is the height, w is the width, and c represents the number of channels, image patches of size p×p×c are extracted from the image X, where p is the side length of the image patch, to form a sequence of image patches of length n (x1, x2,..., x n ), where hw / p 2 = 1.

5. The remote sensing image urban landscape classification method based on the self-attention network according to claim 1, wherein The process of obtaining the labeled result embedded image patch sequence is as follows: In the embedding layer, first linearly project the image patch sequence into a vector of dimension d using the learned embedding matrix E to obtain the embedded representation of the image patch sequence. Then, concatenate the embedded representation of the image patch sequence with the corresponding classification labels. Since the Vision Transformer model treats the embedded image patch sequence as a set of image features and does not consider the positional order of these image patches, in order to keep the spatial arrangement of the image patches the same as the original remote sensing image, positional information is encoded and appended to the embedded representation of the image patch sequence. Therefore, the labeled result embedded image patch sequence z0 is defined as follows: Among them, v class represents a classification mark, and E pos represents position information, n is the length of the image patch sequence, p is the side length of each image patch in the image patch sequence, and c is the number of channels of each image patch in the image patch sequence.

6. The method for classifying urban landscapes of remote sensing images based on a self-attention network according to claim 1, wherein The encoder includes L identical layers, each layer including a multi-head self-attention block and a feed-forward dense block. A normalization layer LN is connected to the inputs of both the multi-head self-attention block and the feed-forward dense block. The inputs and outputs of the multi-head self-attention block and the feed-forward dense block are connected through residual skip connections.

7. The method for classifying urban landscapes in remote sensing images based on the self-attention network according to claim 6, characterized in that, The multi-head self-attention module includes an input linear layer, a multi-head self-attention layer, a connection layer, and an output linear layer. The input linear layer includes a number of linear layers, and each linear layer is used to generate a query matrix, a key matrix, and a value matrix by linearly transforming the input block embedding sequence through different linear transformations. The multi-head self-attention layer includes the same number of self-attention layers as the number of linear layers in the input linear layer. Each self-attention layer is used to calculate the attention result according to the query matrix, the key matrix, and the value matrix. The attention results are connected together in the connection layer and finally projected to the required dimension through a feed-forward layer with learnable weights W, that is, the output linear layer. This process is represented by the following equation: Among them, MSA represents the result of the multi-head self-attention layer, Z represents the input sequence of labeled result embedding image patches, Concat() represents the concatenation operation, h represents the number of heads in the multi-head self-attention layer, that is, the number of self-attention layers, and D K represents the dimension of the key vector in the key matrix; The feed-forward dense block includes two fully connected layers, and there is a GeLU activation function layer between the two fully connected layers.

8. The method for classifying urban landscapes in remote sensing images based on the self-attention network according to claim 7, wherein, The self-attention layer includes matrix multiplication operation, scaling operation, masking operation, and SoftMax function weighting operation. The query matrix and the key matrix first undergo matrix multiplication operation to obtain the dot product of the query matrix and the key matrix. Then, after the dot product undergoes scaling operation and masking operation, it is weighted by the SoftMax function to obtain weights. Finally, matrix multiplication operation is performed with the value matrix.

9. The method for classifying urban landscapes in remote sensing images based on the self-attention network according to any one of claims 6-8, characterized in that, At the end of the encoder, the feature representation of the first element in the result of the encoder is taken and passed to the task head to generate classification labels, completing the classification of urban landscapes in remote sensing images.

10. A remote sensing image urban landscape classification system based on a self-attention network, for applying the remote sensing image urban landscape classification method based on a self-attention network according to any one of claims 1-9, characterized in that, The system includes: A data acquisition module for acquiring remote sensing images containing urban landscapes and classification labels and performing data augmentation to obtain a remote sensing image dataset; An image segmentation module for segmenting the remote sensing images in the remote sensing image dataset into a number of image patches to obtain an image patch sequence; A self-attention network construction module for constructing a self-attention network based on the Vision Transformer model, where the self-attention network includes an embedding layer, an encoder, and a task head; A self-attention network classification module for inputting the image patch sequence and the classification labels into the self-attention network. After obtaining the embedded representation of the image patch sequence through the embedding layer, it is concatenated with the classification labels to obtain a labeled result embedded image patch sequence; inputting the labeled result embedded image patch sequence into the encoder to generate feature representations and attention scores; inputting the feature representations into the task head and generating classification labels according to the attention scores to complete the classification of urban landscapes in remote sensing images.

Citation Information

Cited By

  • Online identification method for unbalanced remote sensing image under cluster framework

    CN121033674A

  • Remote sensing image segmentation method and system based on fine screening double-domain attention mechanism

    CN121582565A