Remote sensing image classification method based on ResNet18-Transform

By introducing the ResNet18-Transformer model, combining data enhancement and self-attention mechanism, the problems of low computing efficiency and insufficient information integration in remote sensing image classification are solved, and high-precision and robust remote sensing image classification are achieved, which is suitable for applications in complex scenarios.

CN120259765APending Publication Date: 2025-07-04YUNNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510372754.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When existing remote sensing image classification methods deal with high-resolution, multi-spectral complex scenarios, there are problems such as low computing efficiency, insufficient classification accuracy and limited generalization capabilities. Especially under complex backgrounds and scale changes, it is difficult to effectively integrate multi-scale information.

Method used

The ResNet18-Transformer model is adopted, through data enhancement technology and self-attention mechanism, combined with ResNet18's local feature extraction and Transformer's global context information capture, dynamic fusion of features is achieved and the classification performance of the model is improved.

Benefits of technology

It significantly improves the accuracy, accuracy and F1 score of remote sensing image classification, enhances the robustness and generalization capabilities of the model, especially the classification performance in complex scenarios, and is suitable for land use classification, environmental monitoring and disaster assessment and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259765A_ABST
    Figure CN120259765A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image classification method based on ResNet18-Transform. The remote sensing image classification method comprises the following steps: acquiring remote sensing image data; transforms are introduced to carry out data enhancement, and the data enhancement comprises random cutting, overturning, rotating, color shaking and normalization processing; constructing a ResNet18 model, introducing a Transform structure into the model, and respectively capturing local features and global context information of the image; carrying out weighted aggregation on the image features through a self-attention mechanism of Transform; and fusing the extracted features with the context information, and finally performing classification prediction through a full connection layer. According to the method, a Transform structure is introduced into the ResNet18 model, the capturing capability of the model for the complex spatial relationship in the remote sensing image is remarkably improved, multiple indexes are improved to a certain extent, particularly, the classification accuracy and the F1 score are remarkably improved, and the classification accuracy, the generalization performance and the robustness of the model are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image processing, and specifically relates to a remote sensing image classification method based on ResNet18-Transformer. Background Art

[0002] In the field of remote sensing image processing, traditional image classification methods (such as support vector machines (SVM) and random forests (RF)) often face significant technical bottlenecks when processing high-resolution, multi-spectral complex scenes. These methods are highly dependent on manually designed feature engineering, and it is difficult to fully tap the rich spatial, spectral and semantic information in remote sensing images, resulting in limited classification accuracy and generalization capabilities. Especially when faced with large-scale remote sensing data, the computational efficiency of traditional methods is low, and it is difficult to meet the needs of efficient and accurate classification in practical applications. In addition, the scene complexity of remote sensing images (such as fuzzy boundaries of objects, large intra-class differences, and high inter-class similarities) further exacerbates the limitations of traditional methods, causing them to be gradually marginalized in dealing with modern remote sensing tasks.

[0003] In recent years, the rapid development of deep learning technology has brought revolutionary breakthroughs in remote sensing image classification. With its powerful feature extraction capabilities, convolutional neural networks (CNNs) have become the mainstream method for remote sensing image processing. However, traditional CNN models still have some inherent defects when processing remote sensing images: for example, the local receptive field limits its ability to model long-distance dependencies, while complex backgrounds and scale changes also pose severe challenges to its classification performance. Although improved CNN architectures (such as ResNet, DenseNet, etc.) have alleviated the gradient vanishing problem and improved feature reuse capabilities through residual connections and dense connections, their ability to model global contextual information is still limited, making it difficult to fully cope with complex scenes in remote sensing images.

[0004] With the wide application of deep learning technology in remote sensing image processing, researchers have proposed various improved models to enhance classification performance. The AlexNet model is fully described in "ImageNet Classification with Deep Convolutional Neural Networks" and achieved an accuracy of 90.44% in subsequent tests. However, due to its low generalization ability, it is prone to overfitting during data augmentation, has a large number of parameters, and low training efficiency. "Land-Cover Classification Using Deep Learning with High-Resolution Remote-Sensing Imagery" proposed a land cover classification model Pro-Res Net based on deep learning. By optimizing the network structure and high-resolution remote sensing image feature extraction, it achieved a relatively high classification accuracy (91%). However, the adaptability of its model to complex scenarios is limited, and the generalization ability across datasets has not been fully verified, restricting the robustness of actual deployment. "Performance Comparison of CNN Based Hybrid Systems Using UC Merced Land-Use Dataset" designed a ResNet18&SVM hybrid system, combining the classification advantages of convolutional neural network (CNN) and support vector machine (SVM), and achieved an accuracy of 93.17% on the UCMerced dataset. However, this model relies on manual feature fusion strategies, has high computational complexity, low accuracy and recognition ability, and does not effectively solve the classification bias problem of small sample categories. These works have significantly promoted the development of remote sensing image classification technology, but there are still problems such as limited model generalization, insufficient computational efficiency, and inadequate integration of multi-scale information. We have innovated by integrating other models accordingly.

[0005] Based on the limitations of existing remote sensing image classification methods, there is an urgent need to propose an innovative model based on ResNet18 and Transformer - the ResNet18-Transformer model, aiming to further improve the performance of remote sensing image classification. Summary of the Invention

[0006] To overcome the problems in the background technology, the present invention provides a remote sensing image classification method based on ResNet18-Transformer.

[0007] To achieve the above object, the present invention is realized by the following technical solutions:

[0008] A remote sensing image classification method based on ResNet18-Transformer, comprising the following steps:

[0009] Step 1: Obtain remote sensing image data and perform data augmentation, including random horizontal flipping, random rotation, color jitter, brightness change, contrast change, and saturation change, and normalize the images.

[0010] Step 2: Construct a ResNet18 network as the backbone network, which contains a total of 8 BasicBlock structures in four stages, and the output feature map dimensions of each stage are 64, 128, 256, and 512 respectively.

[0011] Step 3: Embed a Transformer structure in the ResNet18 model. The Transformer contains 8 attention heads, the feature dimension is 512, the feed-forward network magnifies the feature dimension from 512 to 2048 and then restores it to 512, and aggregates global context information through self-attention mechanism weighted aggregation.

[0012] Step 4: Fuse the local features extracted by ResNet18 and the global features output by the Transformer, and output classification probabilities through global average pooling and fully connected layers.

[0013] Furthermore, in the data augmentation operation, the normalization process scales the image pixel values to the range of [-1, 1], and the training set and test set are divided in a ratio of 8:2.

[0014] Furthermore, the residual stages of the ResNet18 model perform spatial downsampling in sequence, and the output feature map dimensions are gradually increased from 64 to 512, and the input and output features are fused through residual connections in each BasicBlock.

[0015] Furthermore, the self-attention mechanism of the Transformer structure includes: generating query, key, and value vectors through linear transformation, calculating attention weights and then performing weighted summation on the value vectors, and stabilizing the training process through residual connections and layer normalization.

[0016] Furthermore, the model is trained using a cross-entropy loss function and an Adam optimizer, with an initial learning rate of 0.001, a batch size of 8, the learning rate is dynamically adjusted according to the cosine annealing strategy, and the number of training epochs is 100.

[0017] Furthermore, the method is implemented for classification on the UCMerced_LandUse dataset, which contains 21 categories, with 100 remote sensing images of 256×256 pixels for each category.

[0018] Further, the classification results of the model are evaluated through a confusion matrix and APRF metrics (accuracy, precision, recall, F1-score), where the accuracy reaches 95.45% and the F1-score reaches 95.87%.

[0019] Further, the output features of the Transformer structure are adjusted in dimension through 1×1 convolution and then converted into a sequence input, and after processing, they are restored to the form of a two-dimensional feature map.

[0020] Further, during the model training process, training logs are dynamically recorded, including loss values, learning rates, accuracy, and KAPPA coefficients, where the KAPPA coefficient is increased to 0.92.

[0021] Further, the model significantly improves the accuracy and robustness of remote sensing image classification through data augmentation, local feature extraction of ResNet18, and global context information capture of Transformer, especially showing excellent classification performance in complex scenarios.

[0022] Compared with the prior art, the present invention has the following significant advantages:

[0023] (1) Model performance improvement: The present invention improves on the basis of ResNet18, and significantly improves the model performance by introducing the Transformer structure and Transforms data augmentation technology.

[0024] (2) Significantly improved accuracy: The accuracy of the traditional ResNet18 is 91.5%, while that of ResNet18-Transformer is increased to 95.45%.

[0025] (3) Greatly improved precision: In terms of the accuracy of positive class prediction, the precision of the traditional ResNet18 is 91.3%, while that of ResNet18-Transformer is increased to 95.95%.

[0026] (4) The recall rate of the traditional ResNet18 is 90.81%, while that of ResNet18-Transformer is increased to 94.58%, indicating that the present invention has a stronger ability to identify positive class samples.

[0027] (5) The F1-score of the traditional ResNet18 is 90.76%, while that of ResNet18-Transformer is increased to 95.87%, indicating that the present invention achieves a better balance between precision and recall.

[0028] (6) The Kappa coefficient of the traditional ResNet18 is 0.85, while that of ResNet18-Transformer is improved to 0.9, indicating that the consistency between the classification results of the model of the present invention and the true labels is significantly improved.

[0029] (7) At the same time, the present invention has broad application value. The present invention can efficiently process image classification tasks in complex remote sensing scenarios, providing reliable technical support for fields such as land use classification, environmental monitoring, and disaster assessment, helping to improve the application value of remote sensing data, reduce the cost of manual annotation, and provide a scientific basis for scientific research and decision-making in related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0031] Figure 1 is the flowchart of the model usage of the present invention;

[0032] Figure 2 is the confusion matrix diagram of the classification results of the present invention;

[0033] Figure 3 is the structure diagram of the Basic Block (left) and Bottlenec (right) of the present invention;

[0034] Figure 4 is the overall structure diagram of the model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0036] Embodiment 1

[0037] The present invention provides a remote sensing image classification method based on ResNet18 and Transformer models. The overall structure of the model is as Figure 4 shown. The method includes:

[0038] S1: Dataset construction and preprocessing. The UCMerced_LandUse public dataset is adopted. This dataset contains 21 types of land use (including agricultural areas, airports, baseball fields, etc.), and 100 RGB remote sensing images of 256×256 pixels are provided for each type. The data storage directory follows the requirements of the ImageFolder format, and the images of each category are stored in independent subfolders.

[0039] S2: Training environment configuration.

[0040] The specific configuration is as follows:

[0041] S2a: Hardware environment: Configure the CUDA 11.1 parallel computing platform and the NVIDIA RTX 3090 GPU acceleration device, and set the CUDA_LAUNCH_BLOCKING = 1 parameter to ensure the synchronous execution of the computing kernel function.

[0042] S2b: Software parameters: Define the model hyperparameters through configuration classes: the initial learning rate is 0.001, the batch size is 8, the number of training epochs is 100, the image input size is 224×224 pixels, and the AdamW optimizer is used with a weight decay coefficient of 0.01.

[0043] S2c: Data preprocessing: Construct the ApplyTransform data pipeline class, inherit from the PyTorch Dataset base class, and rewrite the __getitem__ method to implement dynamic data augmentation, including:

[0044] Training set: Random horizontal flipping, random rotation, color jittering.

[0045] Validation set: Only perform normalization processing.

[0046] The data is divided using the stratified sampling strategy, and the training set (1680 images) and the validation set (420 images) are split in an 8:2 ratio, maintaining the consistency of the class distribution.

[0047] S3: Model component construction. The core modules include:

[0048] S3a: Transformer module: Implement the multi-head self-attention mechanism. The specific calculation process is as follows:

[0049] The input features are linearly transformed to generate the query matrix Q, the key matrix K, and the value matrix V.

[0050] The calculation process of the multi-head attention mechanism:

[0051] Q&=XW Q ,&K = XW K ,&V = XW V

[0052] Calculate attention weights:

[0053]

[0054] Among them, X is the input tensor, and W Q , W K , W V are the weight matrices of the query, key, and value respectively, and d k is the dimension of the key. It can simultaneously focus on different positions in the input sequence, thereby capturing long-range dependencies in the sequence. The multi-head attention mechanism decomposes the input feature vectors into multiple heads. Each head independently calculates the attention weights and then concatenates their outputs to enhance the model's expression ability for features in different subspaces.

[0055] The module contains layer normalization and residual connections. The feed-forward network uses two fully connected layers (dimension expansion ratio 4:1) and the GELU activation function.

[0056] S3b: Residual module: Construct a BasicBlock structure, which contains two 3×3 convolutional layers, batch normalization, and ReLU activation. When the input and output dimensions do not match, dimension alignment is performed through a 1×1 convolutional shortcut connection. The structure of the residual network is as Figure 3 shown, where the residual connection formula:

[0057] y = BN(x) + F(BN(x))

[0058] In the formula, x is the input tensor, BN represents batch normalization, and F represents the residual function (such as convolution, activation, etc.).

[0059] S4: Model architecture integration. The overall architecture is as Figure 4 shown, and the specific implementation steps:

[0060] S4a: Feature extraction network: Use the improved ResNet18 as the backbone network, including:

[0061] Initial convolutional layer: 3×3 convolutional kernel, output channels 64, stride 2;

[0062] Four-stage residual layer: The number of channels is 64 / 128 / 256 / 512 in sequence, and each stage contains 2 BasicBlocks;

[0063] Feature dimension conversion: Map the final 512-channel feature to the Transformer input dimension 512 through a 1×1 convolution.

[0064] S4b: Feature Fusion Module: Reshape the two-dimensional feature map output by ResNet into a sequence form, input it into the encoder layer containing 4 Transformer Blocks, and restore the output features to a two-dimensional structure through inverse transformation.

[0065] S4c: Classification Decision Layer: Use global average pooling (GAP) to compress the spatial dimension, and connect to a fully connected layer (input dimension 512, output dimension 21) to generate a class probability distribution.

[0066] S5: Model Training and Optimization.

[0067] Adopt a three-stage training strategy:

[0068] S5a: Parameter Initialization: Load the ImageNet pre-trained weights for the ResNet part, and initialize the Transformer part using the Xavier normal distribution;

[0069] S5b: Optimization Strategy: Use the cross-entropy loss function, adopt AdamW as the optimizer, and adjust the learning rate according to the cosine annealing strategy;

[0070] S5c: Regularization Measures: Set the Dropout rate to 0.2, the weight decay coefficient to 0.01, and adopt the early stopping mechanism.

[0071] S6: Model Validation and Deployment.

[0072] Evaluate the model performance on the validation set:

[0073] Evaluation Metrics: Accuracy, Precision, Recall, F1-Score, Kappa Coefficient;

[0074] Performance: The accuracy reaches 95.45% ± 0.32%, and the Kappa coefficient is 0.92 ± 0.02, which is better than the comparison model. See Table 1 below;

[0075] Deployment Plan: Convert the trained model into a portable format through TorchScript, supporting ONNX Runtime inference;

[0076] Table 1 Model Performance Comparison (UCMerced Dataset)

[0077]

[0078]

[0079] S7: Verification of Technical Effects. Such as Figure 2As shown in the confusion matrix, in the classification tasks of "dense residential areas" and "commercial areas" with high inter-class similarity, the accuracy rates of the present invention reach 94.7% and 96.2% respectively, significantly superior to the traditional ResNet18 (89.3% and 91.5%).

[0080] Example 2

[0081] In this embodiment, an optional implementation manner of the present invention is mainly described in detail, but the content of the present invention is not limited to the described scope.

[0082] In this embodiment, aiming at the problem that the traditional image classification method in the field of remote sensing image processing faces significant technical bottlenecks when dealing with complex scenes of high resolution and multi-spectral, an innovative model ResNet18-Transformer model that combines ResNet18 and Transformer based on deep learning is proposed to further improve the performance of remote sensing image classification. First, this method conducts remote sensing image classification based on the ResNet18-Transformer architecture. The specific steps include: obtaining remote sensing image data and performing data augmentation, while normalizing the images, and dividing the training set and the test set. Then, a ResNet18 model is constructed to extract local features through the residual stage. Next, a Transformer structure is embedded in the model to aggregate global context information using the self-attention mechanism. Subsequently, the local features and the global features are fused, and the classification probability is output through global average pooling and fully connected layers. The model training uses the cross-entropy loss function and the Adam optimizer, and the learning rate is adjusted dynamically. This method realizes classification on the publicly available remote sensing image dataset, and evaluates the performance through the confusion matrix and APRF metrics. In addition, the output features of the Transformer are restored to the form of a two-dimensional feature map after dimension adjustment. Training logs are recorded during the model training process, including indicators such as loss values and accuracy rates. This method shows excellent classification performance for specific classes in complex scenes.

[0083] This embodiment proposes a fusion model based on ResNet18-Transformer. Its core innovation lies in solving the limitations of existing methods through data augmentation strategies, a dual-branch feature extraction architecture, and a dynamic feature fusion mechanism. First, after inputting the remote sensing image, data augmentation operations such as random cropping, flipping, rotation, and color jitter are adopted, combined with normalization processing, which significantly improves the model's adaptability to illumination and perspective changes, and at the same time suppresses the risk of overfitting. Subsequently, the ResNet18 model is used to extract the local detailed features of the image (such as texture and edges), and the global context dependencies (such as spatial layout and semantic associations) are captured by embedding the Transformer module. The multi-head self-attention mechanism of the Transformer dynamically assigns weights to the features, further strengthening the ability to model long-distance information. Finally, the local features and global features are adaptively fused. After being compressed by global average pooling, the classification probability is output by the fully connected layer to achieve end-to-end efficient training.

[0084] To verify the advancement of this method, a comparative experiment is conducted with mainstream models on a public remote sensing dataset. The results show that traditional methods are prone to misclassification due to background interference because of the lack of global semantic modeling ability; while the pure Transformer model is insufficient in capturing local details and is difficult to handle small-scale ground objects. This method achieves significant improvements in both classification accuracy and F1 score through the complementary advantages of the dual-branch architecture, and its robustness in complex scenarios such as noise and occlusion is superior to the existing technology. In addition, the lightweight design of ResNet18 reduces the computational complexity while retaining the global modeling ability of the Transformer, achieving a balance between efficiency and accuracy.

[0085] This method is applicable to remote sensing image classification tasks in fields such as agricultural monitoring and urban planning, and it performs particularly well in scenarios where the boundaries of ground objects are blurred and multiple categories are mixed (such as dense building areas and mixed vegetation coverage areas). The limitations of the existing technology are reflected in single feature extraction strategies, single data augmentation, and inefficient feature fusion, resulting in limited classification accuracy and generalization ability. This embodiment breaks through the bottleneck of traditional methods through the dynamic feature expression that fuses local details and global semantics, combined with data augmentation, providing a reliable technical solution for high-precision remote sensing image analysis and showing higher practical value in complex scene classification.

[0086] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A remote sensing image classification method based on ResNet18-Transformer, characterized in that, It includes the following steps: Step 1: Obtain remote sensing image data and perform data augmentation, including random horizontal flipping, random rotation, color jittering, brightness change, contrast change, and saturation change, and normalize the images; Step 2: Construct a ResNet18 network as the backbone network, which contains a total of 8 BasicBlock structures in four stages, and the output feature map dimensions of each stage are 64, 128, 256, and 512 respectively; Step 3: Embed a Transformer structure in the ResNet18 model. The Transformer contains 8 attention heads, the feature dimension is 512, the feed-forward network amplifies the feature dimension from 512 to 2048 and then restores it to 512, and aggregates global context information through self-attention mechanism weighting; Step 4: Fuse the local features extracted by ResNet18 and the global features output by the Transformer, and output classification probabilities through global average pooling and fully connected layers.

2. The method according to claim 1, wherein In the data augmentation operation, the normalization process scales the image pixel values to the range of [-1,1], and the training set and test set are divided in the ratio of 8:

2.

3. The method according to claim 1, wherein The residual stages of the ResNet18 model perform spatial downsampling in sequence, and the output feature map dimensions are gradually increased from 64 to 512, and the input and output features are fused through residual connections in each BasicBlock.

4. The method according to claim 1, wherein The self-attention mechanism of the Transformer structure includes: generating query, key, and value vectors through linear transformation, calculating attention weights and then performing weighted summation on the value vectors, and stabilizing the training process through residual connection and layer normalization.

5. The method according to claim 1, characterized in that, The training of the model uses the cross-entropy loss function and the Adam optimizer, the initial learning rate is 0.001, the batch size is 8, the learning rate is dynamically adjusted according to the cosine annealing strategy, and the number of training epochs is 100.

6. The method according to claim 1, wherein The method realizes classification on the UCMerced_LandUse dataset, which contains 21 categories, and each category has 100 remote sensing images of 256×256 pixels.

7. The method according to claim 1, wherein The classification results of the model are evaluated through a confusion matrix and APRF metrics (accuracy, precision, recall, F1-score), where the accuracy reaches 95.45% and the F1-score reaches 95.87%.

8. The method according to claim 1, wherein The output features of the Transformer structure are adjusted in dimension through 1×1 convolution and then converted into sequence input, and then restored to the form of a two-dimensional feature map after processing.

9. The method according to claim 1, characterized in that, During the training process of the model, training logs are dynamically recorded, including loss values, learning rates, accuracies, and KAPPA coefficients, and the KAPPA coefficient is increased to 0.

92.

10. The method according to claim 1, characterized in that, Through data augmentation, local feature extraction of ResNet18, and global context information capture of the Transformer, the model significantly improves the accuracy and robustness of remote sensing image classification, especially showing excellent classification performance in complex scenarios.