Remote sensing image classification method based on improved self-attention

By combining the multi-scale feature extraction and fusion module of self-attention and local attention module, the problem of insufficient global and local feature capture in remote sensing image classification is solved, and the accuracy and robustness of remote sensing image classification is improved.

CN120259729APending Publication Date: 2025-07-04NORTHWEST UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510273444.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art is difficult to effectively capture the global and local features of images in remote sensing image classification, especially when dealing with the problems of feature sparsity and scale differences, the classification performance is degraded.

Method used

A dual-branch multi-scale attention feature extraction fusion module is adopted, combining self-attention and local attention modules, and the attention selection and adjustment strategy is used to enhance attention to important features, making up for the shortcomings of the self-attention mechanism in local information extraction.

Benefits of technology

The accuracy and robustness of remote sensing image classification are significantly improved, especially in scenarios with large-scale changes, the classification accuracy is improved by 7.3-2.5%, enhancing the recognition effect of target objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259729A_ABST
    Figure CN120259729A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image classification method based on improved self-attention, and the method comprises the steps: employing a dual-branch multi-scale attention feature extraction and fusion module in an encoder assembly based on a ViT network frame, combining a local attention module based on a convolutional neural network with a self-attention module, and carrying out the recognition of a remote sensing image through the fusion module. And a multi-scale remote sensing image classification system with global and local feature extraction capability is formed. The local attention module makes up for the deficiency of a self-attention mechanism in the aspect of local information extraction by effectively capturing the local features of the image, so that a remote sensing scene with large-scale change can be better processed. By adopting the technical scheme of the invention, the distinction degree of attention scores can be improved, and the attention capability of a self-attention mechanism on important features during remote sensing image processing is enhanced, so that the classification performance is improved, and meanwhile, the method can better adapt to remote sensing image classification tasks with large target object scale differences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image processing and computer vision, and particularly relates to a remote sensing image classification method based on improved self-attention. Background Art

[0002] Remote sensing image classification is a key task in the field of remote sensing interpretation. By classifying remote sensing images, scene information and target features can be effectively extracted. This technology is of great significance in applications such as land use monitoring, disaster assessment, and resource development. However, with the rapid accumulation of remote sensing data, traditional methods relying on expert manual interpretation can no longer meet the current needs. To address this challenge, deep learning technologies, especially Convolutional Neural Network (CNN) and Self-Attention mechanism, have made significant progress in remote sensing image classification in recent years.

[0003] Convolutional Neural Network (CNN) extracts features from images through adaptive convolutional kernels, especially outstanding in local information extraction. Techniques such as transfer learning and ResNet combined with multi-scale learning methods have effectively improved the classification accuracy of remote sensing images. However, CNN is limited by a fixed receptive field and cannot comprehensively capture global features in remote sensing images.

[0004] The Self-Attention mechanism has unique advantages in capturing the global context of remote sensing images through global weighted processing of feature sequences. Some studies have attempted to apply it to remote sensing image classification tasks and optimize classification performance through methods such as multi-modal fusion and compressed attention. However, when directly applying the Self-Attention mechanism to remote sensing image classification, especially in dealing with the problems of feature sparsity and scale difference unique to remote sensing images, there are still the following deficiencies: First, the sparse feature information leads to a decline in classification performance. Due to the scarcity of semantic information and the large amount of redundant airspace information in remote sensing images, when the traditional Self-Attention mechanism processes these images, its normalization operation may make the attention score distribution too smooth, thus reducing the weight of important features. As a result, the model cannot fully focus on key objects or regions in the image, leading to a decline in classification accuracy. Second, it is difficult to handle the scale difference problem in remote sensing images. There are significant scale differences in the objects in remote sensing images. Although the traditional Self-Attention mechanism has advantages in global feature extraction, its ability to capture local features is relatively weak. Although existing local feature extraction methods based on convolutional neural networks can process some local information, the limitation of their receptive fields leads to insufficient ability to extract global features, making it difficult to meet the precise capture requirements of multi-scale information for remote sensing image classification tasks. Summary of the Invention

[0005] The purpose of the present invention is to provide a remote sensing image classification method based on improved self-attention, which can improve the discrimination of attention scores, enhance the ability of the self-attention mechanism to focus on important features when processing remote sensing images, thereby improving the classification performance, and can better adapt to the remote sensing image classification task with large differences in the scales of target objects.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A remote sensing image classification method based on improved self-attention includes the following steps:

[0008] Step 1: Build a multi-scale remote sensing image classification system, where the multi-scale remote sensing image classification system includes an embedding projection component, an encoder component, and a classification head component;

[0009] The encoder component is composed of a series of identical dual-branch multi-scale attention feature extraction and fusion modules. During the calculation process of the encoder component, the feature information of each image patch is continuously introduced into the classification patch in the sequence data, so that the classification patch contains sufficient information for image classification. After the calculation of the last multi-scale attention feature extraction and fusion module in the series is completed, the information of the image patch will be discarded, and only the content of the classification patch will be retained as the output;

[0010] Step 2: Perform standard preprocessing on the remote sensing image data and convert it into sequence data with a length of L to adapt to the input requirements of the self-attention mechanism;

[0011] Step 3: Perform linear projection on the image patch sequence processed in Step 2, and add learnable spatial position information and classification patch information to form a sequence with a length of L + 1;

[0012] Step 4: Use the sequence with a length of L + 1 as the input of the encoder component, and perform global and local feature extraction on the image patch sequence with spatial position information through a series of multi-layer dual-branch multi-scale attention feature extraction and fusion modules, and use a fusion strategy to fuse the global and local features to obtain a fused feature matrix;

[0013] Step 5: Extract the classification patch from the fused feature matrix, and use the classification head component to perform feature extraction and conversion on it to obtain a category probability distribution vector Category probability distribution vector The category corresponding to the largest value in the is the final classification result of the image.

[0014] Furthermore, the multi-scale attention feature extraction and fusion module described in step 1 consists of two branches. The first branch is a self-attention module with an attention selection and adjustment strategy, and the second branch is mainly composed of a local attention module. Among them, the local attention module uses the common CBAM module, and the self-attention module with an attention selection and adjustment strategy is modified based on the self-attention module of the original ViT model. The specific steps are as follows:

[0015] Input: Sequence matrix X ∈ R L×C (L is the sequence length, C is the vector dimension of the sequence elements), and select the ratio η;

[0016] The specific algorithm steps are as follows:

[0017] Q, K, V := Linear(X)

[0018] Q, K := LayerNorm(Q), LayerNorm(K)

[0019]

[0020] A := MatMul(Q, K T )

[0021] B := ZerosLike(A)

[0022] FOR EACH row IN A:

[0023] Let p be the value of the th element after sorting row from large to small

[0024] Set the values in row that are lower than p to 0

[0025] Replace row with the corresponding row in B

[0026] END FOR

[0027] A := A + B

[0028] S := SoftMax(A)

[0029] Y := Linear(MatMul(S, V))

[0030] Output: Sequence matrix Y ∈ R L×C .

[0031] Furthermore, the specific process of the preprocessing described in step 2 is as follows:

[0032] (1) For the input remote sensing image data X ∈ R H×W×CPerform standard preprocessing so that the values of each channel of each pixel are within [0,1], where R is the real number field, H is the height of the image, W is the width of the image, and C is the number of channels of the image. For common RGB images, the number of channels C is 3;

[0033] (2) Divide the image into several image patches of size p×p, and the dimension of each image patch is p×p×C; subsequently, convert the two-dimensional image data into sequence data to adapt to the input requirements of the self-attention mechanism. Each image patch is flattened into a one-dimensional vector, and the final dimension of the data matrix X is L×(p 2 ×C), where p 2 ×C is the dimension of each sequence element, and L is the sequence length, that is, how many blocks the image is divided into.

[0034] Furthermore, the multi-scale attention feature extraction described in step four includes the calculation of the self-attention branch and the calculation of the local attention branch. Among them, the self-attention branch is calculated using the self-attention module that is good at capturing the overall global information of the image, and the local attention branch is calculated using the local attention block. And before the calculation of the local attention block, it is necessary to first perform an inverse serialization conversion to restore the input sequence data to a two-dimensional feature image, and then perform a serialization conversion after the calculation to restore the two-dimensional feature map to sequence data.

[0035] Furthermore, the steps of a single head (Head) in the calculation of the attention module include:

[0036] S1. The input sequence is presented as a data matrix , and first, three learnable linear transformations and are used to generate the query (Query) matrix Q∈R L×d , the key (Key) matrix K∈R L×d , and the value (Value) matrix V∈R L×d , where d = p 2 ×C is the dimension of the attention; subsequently, Q and K are normalized, and their dot product is calculated to obtain the attention score matrix A before normalization. The formula is as follows:

[0037] K T is the transpose of the key matrix K;

[0038] Next, Q and K are normalized, and their dot product is calculated to obtain the attention score matrix;

[0039] S2. Sort each row of the attention score matrix A in descending order, select the top η important attention scores, and set the other elements to zero to form the adjusted attention matrix, where η is a hyperparameter, and a typical value is 0.3;

[0040] S3. Add the original attention matrix A and the adjusted attention matrix B element by element to obtain a new attention matrix A' = A + B;

[0041] S4. Perform a SoftMax operation on A' to obtain the final normalized attention matrix S, and generate the output feature matrix Y ∈ R through matrix multiplication Y = S × V L×d .

[0042] Furthermore, the specific operation of the deserialization conversion is as follows: Strip out the classification blocks (i.e., the 1s in L + 1) in the data matrix X, and rearrange the information of the remaining L blocks along the L dimension according to the image segmentation method in step two into two dimensions (width, height) that are the same as before segmentation in step two. Finally, restore the feature vectors of the L image blocks to a feature map with the shape of H × W × d', where d' is the last dimension.

[0043] Furthermore, the steps of the serialization conversion are the reverse operations of the deserialization conversion, which are essentially the same as the image segmentation method in step two, but after restoring the two-dimensional feature map to sequence data, the classification block data stripped out during the deserialization process needs to be added back to the beginning of the sequence.

[0044] Furthermore, the specific process of the classification head component in step five includes:

[0045] (1) Layer normalization: Normalize the input features to enhance the stability of the model. The specific calculation method is as follows:

[0046]

[0047] where μ ∈ R and σ ∈ R are the mean and standard deviation of h cls respectively, and γ ∈ R d′ and β ∈ R d′ are learnable scaling and offset parameters;

[0048] (2) Linear projection: Map the normalized features to the category space. The specific calculation method is as follows:

[0049] z = Wh cls '+ b

[0050] where is the weight matrix, is the bias vector, and the output represents the unnormalized scores of each category;

[0051] (3) Probability conversion: Obtain the probability distribution through the Softmax function

[0052]

[0053] Finally, a category probability distribution vector is obtained. where C class is the number of classification categories.

[0054] Compared with the prior art, the present invention has the following beneficial effects:

[0055] By designing an attention selection and adjustment strategy, the present invention significantly improves the discrimination ability of the self-attention mechanism when processing remotely sensed images with sparse feature information. By sorting and screening the attention scores, the attention weights of important features are strengthened, while the influence of redundant information is suppressed, thereby improving the recognition effect of the model on the target object. In some sparse feature scenarios (such as detecting bridge targets in a large-area water background), the classification accuracy of the present invention is improved by 7.3% compared with the standard self-attention mechanism.

[0056] The present invention adopts a dual-branch multi-scale attention feature extraction and fusion module. By combining a local attention module based on a convolutional neural network with a self-attention module, a multi-scale remotely sensed image classification system with global and local feature extraction capabilities is formed. The local attention module effectively captures the local features of the image, making up for the deficiency of the self-attention mechanism in local information extraction, and thus can better process remotely sensed scenes with large-scale changes. Compared with traditional classification models implemented separately based on self-attention or convolutional neural network (CNN), the present invention shows higher robustness and adaptability in multi-scale remotely sensed image classification tasks, especially in scenarios with significant scale changes (such as scenes with both small buildings and large-area vegetation), and the classification accuracy is improved by 2.1%. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is a schematic structural diagram of the original ViT model;

[0058] Figure 2 is a schematic structural diagram of the multi-scale attention feature extraction and fusion module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0060] A method for classifying remotely sensed images based on improved self-attention according to this embodiment includes the following steps:

[0061] Step 1: Build a multi-scale remotely sensed image classification system

[0062] Such as Figure 1As shown, based on the framework of the ViT (Vision Transformer) network, the multi-scale remote sensing image classification system includes an embedding projection component, an encoder component, and a classification head component.

[0063] The embedding projection component includes a simple single-layer FFN network, which realizes the mapping of the original optical information of image patches to the feature space and adds position information to each image patch. Finally, for the purpose of classification, a data block called the classification block is artificially added and concatenated with the data of other image patches to form sequence data.

[0064] The encoder component is composed of several identical double-branch multi-scale attention feature extraction and fusion modules connected in series. As Figure 2 shown, the multi-scale attention feature extraction and fusion module consists of two branches, and each branch independently (logically in parallel) calculates the input of the current module. The first branch of the multi-scale attention feature extraction and fusion module is a self-attention module with an attention selection and adjustment strategy, and the second branch is mainly composed of a local attention module. To enable the local attention module to work in coordination with the self-attention module, an anti-serialization process is added before the local attention module, and a serialization process is added after the local attention module, so that the sequence data is first converted into the data form required by the local attention module, and then converted back into sequence data after being processed by the local attention module. Finally, the calculation results of the two branches are fused through a hybrid strategy. During the calculation process of the encoder component, the feature information of each image patch will be continuously imported into the classification block in the sequence data, so that the classification block contains sufficient information for image classification. After the calculation of the last multi-scale attention feature extraction and fusion module connected in series is completed, the information of the image patch will be discarded, and only the content of the classification block will be retained as the output.

[0065] In the encoder component, the local attention module uses a common CBAM module. The self-attention module with an attention selection and adjustment strategy is modified based on the self-attention module of the original ViT model. The specific algorithm steps are as follows (where steps 5-11 are the steps newly added by the attention selection and adjustment strategy in the calculation process of the original self-attention):

[0066]

[0067]

[0068] Steps 5-11 make the numerical values of the attention matrix A more distinguishable in terms of row values compared to the original attention matrix A after sorting and zeroing, thus alleviating the problem of overly smooth attention.

[0069] The classification head component mainly receives the classification head information output by the encoder component and uses a two-layer FFN network.

[0070] Step 2: Perform standardized preprocessing on the remote sensing image data and convert it into sequence data of length L to adapt to the input requirements of the self-attention mechanism

[0071] First, for the input remote sensing image data X ∈ R H×W×C perform standardized preprocessing so that the values of each channel of each pixel are within [0, 1], where R is the real number field, H is the height of the image, W is the width of the image, and C is the number of channels of the image. For common RGB images, the number of channels C is 3. Then, the image is sliced into several image patches (patches) with a length and width of p × p, and the dimension of each image patch is p × p × C. Subsequently, the two-dimensional image data is converted into sequence data to adapt to the input requirements of the self-attention mechanism. Each image patch is flattened into a one-dimensional vector, and the final dimension of the data matrix X is L × (p 2 × C), where p 2 × C is the dimension of each sequence element, and L is the sequence length, that is, how many blocks the image is sliced into. For example, if the input image size is 256 × 256 and it is sliced into 16 × 16 blocks, then

[0072] The input data is stored in the form of a two-dimensional image. After standardized processing, it is sliced into several image patches of the same size. Each image patch is flattened into a one-dimensional vector to form a serialized input data matrix X. After converting the two-dimensional image data into a one-dimensional sequence, ensure that the data dimension adapts to the input requirements of the self-attention mechanism.

[0073] Step 3: Perform linear projection on the sequence of image patches processed in Step 2 and add learnable spatial position information and classification block information to form a sequence of length L + 1

[0074] Use the linear projection component to perform linear projection on the L image patches in the sequence obtained after Step 2 and add learnable spatial position information parameters. At the same time, add a learnable classification block at the beginning of the sequence to form a sequence of length L + 1.

[0075] The L image patches obtained by cutting in Step 2 will first be flattened through a fully connected layer, and then mapped to the latent space using a learnable projection matrix to form L image feature vectors with a dimension of d. To preserve the spatial structure information of the image, each position index (from 1 to L) will be assigned a learnable spatial position encoding vector. These parameters are randomly initialized and dynamically optimized during training, and finally fused into the corresponding projection vector in an element-wise addition manner. At the same time, to establish a global semantic representation, a learnable classification block with the same dimension d is inserted at the starting position of the sequence. This classification block gradually aggregates the global feature information of the entire image through self-attention interaction with subsequent image patches. After the above processing, the original sequence with a length of L is extended to a composite representation sequence with a dimension of (L + 1)×D, where the first classification block will ultimately serve as the output feature vector for the image classification task, and the subsequent L positions carry both local visual features and spatial position information.

[0076] Step 4: Use the sequence with a length of L + 1 as the input of the encoder component, perform multi-scale attention feature extraction on the sequence of image patches with spatial position information through a cascaded multi-layer double-branch multi-scale attention feature extraction and fusion module, and use a fusion strategy to fuse the global and local features to obtain a fused feature matrix.

[0077] Use the sequence with a length of L + 1 obtained in Step 3 as the input of the encoder component and perform calculations using the multi-scale attention feature extraction and fusion module for several layers. As Figure 2 shown, each layer of the multi-scale attention feature extraction and fusion module is independently calculated by two branches, namely the self-attention branch and the local attention branch, and then the calculation results of the two branches are synthesized using a fusion strategy and output as the result. The specific content is as follows:

[0078] (1) Calculation of the self-attention branch

[0079] This branch uses a self-attention module that is good at capturing the overall global information of the image to calculate, so as to achieve feature extraction in a large-scale range. The specific steps are as follows:

[0080] S1: The input sequence is presented as a data matrix , and first passes through three learnable linear transformations and to generate a query (Query) matrix Q∈R L×d , a key (Key) matrix K∈R L×d , and a value (Value) matrix V∈R L×d respectively, where d = p 2 ×C is the dimension of the attention. Subsequently, Q and K are normalized, and their dot product is calculated to obtain the attention score matrix A before normalization. The formula is as follows:

[0081] K T is the transpose of the key matrix K;

[0082] Next, Q and K are normalized, and their dot product is calculated to obtain the attention score matrix.

[0083] S2. Sort each row of the attention score matrix A in descending order, select the top η important attention scores, and set the other elements to zero to form the adjusted attention matrix, where η is a hyperparameter, and a typical value is 0.3.

[0084] S3. Add the original attention matrix A and the adjusted attention matrix B element by element to obtain the new attention matrix A' = A + B.

[0085] S4. Perform the SoftMax operation on A' to obtain the normalized final attention matrix S, and generate the output feature matrix Y ∈ R through matrix multiplication Y = S × V L×d .

[0086] It should be noted that: the process of a single head in the attention calculation is described here. In fact, the processes of S1 - S4 will be calculated multiple times in the system (where the linear transformations and are different for each head), so as to obtain the attention score matrix of each head. This repeated process is the multi - head attention mechanism. By calculating multiple independent attention heads in parallel, the system can simultaneously focus on different subspace information of the input sequence, thereby enhancing the model's ability to capture diverse global features.

[0087] (2) Calculation of the local attention branch

[0088] To further enhance the ability to capture small - scale features in remote sensing images, this branch uses local attention blocks for calculation. Since the input data is one - dimensional sequence data, while local attention modules usually require the input to be a two - dimensional image, it is necessary to restore the input sequence data to a two - dimensional feature image, and this process is called deserialization. And in order to fuse the calculation results of the local attention module with those of the self - attention module, it is necessary to serialize it back into one - dimensional sequence data again. Here, one - dimensional / two - dimensional does not include the feature dimension, that is, the last dimension.

[0089] The one - dimensional sequence data X ∈ R (L+1)×d loses the relative position relationship (row and column numbers) of each image patch. The purpose of deserialization is to restore the feature matrix X to be similar to the input image X ∈ R H×W×CShape, restore this relative positional relationship, so as to facilitate the fusion with the output of the local attention template. The deserialization process will rearrange the sequentially flattened image patches in step two into a two-dimensional grid structure, restoring the spatial information of the original image. The specific operations are as follows:

[0090] Strip out the classification patches (i.e., 1 in L + 1) in the data matrix X, and rearrange the information of the remaining L patches along the L dimension according to the image segmentation method in step two into two dimensions (width, height) consistent with those before segmentation in step two. Finally, restore the feature vectors of the L image patches to a feature map of shape H×W×d′ according to their original positions, where d′ is the last dimension.

[0091] Subsequently, send the restored feature map into the local feature module for calculation. The local feature module uses the CBAM (Convolutional Block Attention Module) module. This module enhances the system's attention to key regions through channel attention and spatial attention mechanisms. The output dimension of the CBAM module remains H×W×d′.

[0092] Subsequently, perform a serialization conversion. This step is the reverse operation of deserialization. Substantially, it is consistent with the image segmentation method in step two, but after restoring the two-dimensional feature map to sequence data, the classification patch data stripped out during the deserialization process needs to be added back to the beginning of the sequence again.

[0093] (3) Feature fusion

[0094] Align the output of the local attention branch with the output of the self-attention branch, that is, ensure that they have the same feature dimension (i.e., the last feature dimension d′). Subsequently, adopt a pointwise summation strategy to fuse the self-attention and local attention features to obtain the fused feature matrix Y fusion ∈R H×W×d′ .

[0095] Step five: Extract the classification patches in the fused feature matrix, and use the classification head component to perform feature extraction and transformation on them to obtain the class probability distribution vector Class probability distribution vector The class corresponding to the largest value in the class probability distribution vector is the final classification result of the image

[0096] This step follows the operation of the classification head component of the ViT model, that is, it extracts the information of the first block (classification block) in the final result obtained in step four, and uses the classification head component to further extract and transform its features to achieve the mapping from features to the final classification label. This component first standardizes the features through layer normalization to enhance the generalization ability of the system; then it performs a non-linear transformation through a fully connected layer, where activation functions such as GeLU are used in the intermediate layer to introduce non-linear expression ability, and the final output layer compresses the dimension to the total number of categories through a fully connected layer.

[0097] Extract the information of the first block (classification block) in the final result obtained in step four (i.e., the feature vector h corresponding to the classification token cls ∈R d′ ), and use the classification head component to further extract and transform its features. The classification head usually consists of layer normalization and a fully connected layer, and the specific process is as follows:

[0098] (1) Layer normalization (Layer Norm): Normalize the input features to enhance the stability of the model. The specific calculation method is as follows:

[0099]

[0100] where μ∈R and σ∈R are the mean and standard deviation of h cls respectively, and γ∈R d′ and β∈R d′ are learnable scaling and offset parameters.

[0101] (2) Linear projection: Map the normalized features to the category space. The specific calculation method is as follows:

[0102] z = Wh cls ′ + b

[0103] where is the weight matrix, is the bias vector, and the output represents the unnormalized scores (logits) of each category.

[0104] (3) Probability conversion: Obtain the probability distribution through the Softmax function

[0105]

[0106] Finally, obtain the category probability distribution vector where C class is the number of classification categories. The category corresponding to the largest value in the category probability distribution vector is the final classification result of the image.

[0107] The remote sensing image sample data set (including NWPU-RESISC45, UCMerced_LandUse, WHU-RS19, RSSCN7 and AID) is divided into a training set, a validation set and a test set according to a preset ratio to train the system of the present invention.

[0108] Each dataset is divided into five equal parts, and one is randomly selected as the test set (accounting for 20%); the remaining data is divided into a training set (64%) and a validation set (16%) in a ratio of 4:1. RandAugment (random rotation, cropping, color jitter) and Mixup are used to enhance the training dataset during the training phase.

[0109] Weight initialization: Truncated Normal Initialization is used, and the weight distribution satisfies W i,j ~N(0,0.02),|W i,j |<2σ, where σ=0.02.

[0110] Optimizer configuration: Use AdamW optimizer (weight decay coefficient λ = 0.3), basic learning rate η = 3 × 10 -3 , the learning rate scheduling uses linear warm-up (300 steps) and cosine annealing strategy:

[0111]

[0112] Loss function: Use the standard cross-entropy loss function (Cross-Entropy Loss):

[0113]

[0114] where y C is the one-hot encoding of the true label, Predict probabilities for the system.

[0115] During the entire training process, the cross entropy loss on the training set is calculated in each training round and the weights are updated by backpropagation; the model performance is evaluated on the validation set every other epoch and the best weights are saved; if the validation set accuracy does not improve for five consecutive epochs, the training is terminated early or the predetermined maximum number of training rounds (100 rounds) is reached.

[0116] Input a remote sensing image with a resolution of 400×400, sample it to a resolution of 224×224 using the bicubic algorithm, and divide the image into blocks of size 16×16 to obtain 14×14 = 196 blocks (patches). Each block is flattened into a vector of length 196×3 = 588 (3 represents the three color channels in an RGB color image), forming a serialized input data matrix X. Verify the accuracy of the attention selection and adjustment strategy and the multi-scale feature fusion module of the present invention in the remote sensing image classification task on the NWPU-RESISC45, UCMerced_LandUse, WHU-RS19, RSSCN7, and AID datasets. The results are shown in Table 1.

[0117] Table 1

[0118]

[0119] It can be seen from the experimental data in Table 1 that the attention selection and adjustment strategy and the multi-scale feature fusion module proposed in the present invention show significant technical advantages in the remote sensing image classification task. In the feature sparse scenario, taking the WHU-RS19 dataset as an example, the multi-scale attention feature extraction and fusion module increases the classification accuracy from 68.8% to 76.1%, with a relative increase of 7.3 percentage points. This performance leap stems from the strengthening effect of the attention score screening mechanism on key features. For example, when detecting bridge targets in a water background, this strategy can effectively suppress the redundant response of the water surface texture and focus on the key edge features of the bridge structure. On the RSSCN7 dataset, the accuracy increases from 85.5% to 89.3% after introducing this module. When dealing with complex scenarios that simultaneously contain small buildings (local features) and large areas of vegetation (global features), the local attention mechanism of the convolutional branch captures details such as window panes and roofs, while the self-attention branch maintains the overall semantic association of the vegetation distribution. The synergistic effect of the local attention module and the self-attention module results in a cumulative increase of 2.1 percentage points in the AID dataset, with the accuracy increasing from 77.6% to 79.7%. In addition, the synergistic effect is fully verified on the NWPU-RESISC45 dataset, with an accuracy increase of up to 2.5%, indicating that when dealing with remote sensing scenarios with significant scale differences (such as long and thin airport runway targets and large apron areas), the dual improvement scheme can optimize both the feature focusing ability and the multi-scale representation ability. This technical combination not only makes up for the inherent defect of traditional ViT in local feature extraction, but also enables the system to maintain stable target recognition performance under complex background interference through the attention weight dynamic adjustment mechanism, providing a more powerful feature learning framework for remote sensing image interpretation.

Claims

1. A remote sensing image classification method based on improved self-attention, characterized in that, Including the following steps: Step 1: Build a multi-scale remote sensing image classification system, where the multi-scale remote sensing image classification system includes an embedding projection component, an encoder component, and a classification head component; Among them, the encoder component is composed of a series of identical double-branch multi-scale attention feature extraction and fusion modules. During the calculation of the encoder component, the feature information of each image patch is continuously imported into the classification patch in the sequence data, so that the classification patch contains sufficient information for image classification. After the calculation of the last multi-scale attention feature extraction and fusion module in the series is completed, the information of the image patch will be discarded, and only the content of the classification patch will be retained as the output; Step 2: Perform standardized preprocessing on the remote sensing image data and convert it into sequence data with a length of L to adapt to the input requirements of the self-attention mechanism; Step 3: Perform linear projection on the image patch sequence processed in Step 2, and add learnable spatial position information and classification patch information to form a sequence with a length of L + 1; Step 4: Use the sequence with a length of L + 1 as the input of the encoder component, and perform global and local feature extraction on the image patch sequence with spatial position information through a series of multi-layer double-branch multi-scale attention feature extraction and fusion modules, and use a fusion strategy to fuse the global and local features to obtain a fused feature matrix; Step 5: Extract the classification blocks from the fused feature matrix, and use the classification head component to perform feature extraction and transformation on them to obtain a class probability distribution vector Class probability distribution vector The class corresponding to the largest value in the is the final classification result of the image.

2. The remote sensing image classification method based on improved self-attention according to claim 1, wherein The multi-scale attention feature extraction and fusion module described in Step 1 is composed of two branches. The first branch is a self-attention module with an attention selection and adjustment strategy, and the second branch is mainly composed of a local attention module; among them, the local attention module uses a common CBAM module, and the self-attention module with an attention selection and adjustment strategy is modified based on the self-attention module of the original ViT model. The specific steps are as follows: Input: Sequence matrix X ∈ R L×C (where L is the sequence length and C is the vector dimension of sequence elements), select the ratio η; The specific algorithm steps are as follows: Q, K, V := Linear(X) Q, K := LayerNorm(Q), LayerNorm(K) A := MatMul(Q, K T ) B := ZerosLike(A) FOR EACH row IN A: Let p be the value of the element after sorting row from large to small Set the values in row that are lower than p to 0 Replace row with the corresponding row in B END FOR A := A + B S := SoftMax(A) Y := Linear(MatMul(S, V)) Output: Sequence matrix Y ∈ R L×C .

3. The remote sensing image classification method based on improved self-attention according to claim 2, wherein, The specific process of the preprocessing described in Step 2 is: (1)Preprocess the input remote sensing image data \(X\in R\) H×W×C by standardization so that the values of each channel of each pixel are within \([0, 1]\), where \(R\) is the real number field, \(H\) is the height of the image, \(W\) is the width of the image, and \(C\) is the number of channels of the image. For common RGB images, the number of channels \(C\) is 3; (2) Cut the image into several image patches (patches) of size p×p, and the dimension of each image patch is p×p×C; Subsequently, the two-dimensional image data is converted into sequence data to meet the input requirements of the self-attention mechanism. Each image patch is flattened into a one-dimensional vector, and the final dimension of the data matrix X is L×(p 2 ×C), where p 2 ×C is the dimension of each sequence element, and L is the sequence length, that is, the number of patches into which the image is divided.

4. A remote sensing image classification method based on improved self-attention according to claim 3, characterized in that, The multi-scale attention feature extraction described in Step 4 includes the calculation of the self-attention branch and the calculation of the local attention branch. Among them, the self-attention branch is calculated using a self-attention module that is good at capturing the overall global information of the image, and the local attention branch is calculated using a local attention block. And before the calculation of the local attention block, it is necessary to perform deserialization conversion first to restore the input sequence data to a two-dimensional feature image, and then perform serialization conversion after the calculation to restore the two-dimensional feature map to sequence data.

5. A remote sensing image classification method based on improved self-attention according to claim 4, characterized in that The steps for a single head (Head) in the calculation of the attention module include: S1. The input sequence is presented as a data matrix . First, three learnable linear transformations and are used to generate a query matrix Q ∈ R L×d , a key matrix K ∈ R L×d and a value matrix V ∈ R L×d respectively, where d = p 2 × C is the dimension of the attention; subsequently, Q and K are normalized, and their dot product is calculated to obtain the attention score matrix A before normalization, and the formula is as follows: K T is the transpose of the key matrix K; Next, normalize Q and K, and calculate their dot product to obtain the attention score matrix; S2. Sort each row of the attention score matrix A in descending order, select the top η important attention scores, and set the other elements to zero to form the adjusted attention matrix, where η is a hyperparameter, and a typical value is 0.3; S3. Add the original attention matrix A and the adjusted attention matrix B element by element to obtain a new attention matrix A' = A + B; S4. Perform a SoftMax operation on A' to obtain the final normalized attention matrix S, and generate the output feature matrix Y ∈ R through matrix multiplication Y = S × V L×d .

6. The remote sensing image classification method based on improved self-attention according to claim 4, wherein The specific operation of the deserialization conversion is as follows: Strip out the classification blocks (i.e., 1 in L + 1) in the data matrix X, and rearrange the information of the remaining L blocks along the L dimension according to the image segmentation method in step two into two dimensions (width, height) that are the same as before the segmentation in step two. Finally, restore the feature vectors of the L image blocks to a feature map with the shape of H×W×d′ according to their original positions, where d′ is the last dimension.

7. A remote sensing image classification method based on improved self-attention according to claim 4, characterized in that The steps of the serialization conversion are the reverse operations of the deserialization conversion, which are essentially the same as the image segmentation method in step two, but after restoring the two-dimensional feature map to sequence data, the classification block data stripped out during the deserialization process needs to be added back to the beginning of the sequence.

8. A remote sensing image classification method based on improved self-attention according to claim 1, characterized in that The specific process of the classification head component described in step five includes: (1) Layer normalization: Normalize the input features to enhance the stability of the model. The specific calculation method is as follows: where μ ∈ R and σ ∈ R are the mean and standard deviation of h cls respectively, γ ∈ R d′ and β ∈ R d′ are learnable scaling and offset parameters; (2) Linear projection: Map the normalized features to the category space. The specific calculation method is as follows: Among them is the weight matrix, is the bias vector, and the output represents the unnormalized scores for each category; (3) Probability conversion: Obtain the probability distribution through the Softmax function Finally, a category probability distribution vector is obtained. Among them, C class is the number of classification categories.

Citation Information

Cited By

  • Automatic channel system identification method and system based on artificial intelligence

    CN120673263A

  • Transform-based remote sensing image scene classification method

    CN121982500A

  • A remote sensing image scene classification method based on a transformer

    CN121982500B