Remote sensing scene classification method and system based on small sample learning
Through the improved MobileViT network and small sample learning framework, the problems of high annotation cost and insufficient model generalization capabilities in remote sensing image scene classification are solved, and the efficient classification of remote sensing images under small sample conditions is realized and the accuracy and stability of remote sensing image scene classification are improved.
Patent Information
- Application Number
- CN202510765279.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The existing remote sensing image scene classification technology has problems such as high labeling cost, insufficient model generalization ability and difficult to deal with feature complexity under small sample conditions, especially in cross-sensor and cross-regional scenarios, resulting in insufficient classification accuracy and untimely response.
Using the improved MobileViT network, the channel attention mechanism is introduced in the MobileNetV2 module, and local features and global features are integrated in the MobileViT module, combined with a small sample learning framework, a lightweight remote sensing scene classification model is built, and multi-scale feature fusion and dynamic attention weighting are used for classification.
It significantly reduces the cost of remote sensing image annotation, improves the classification accuracy and generalization ability of the model under small sample conditions, and can quickly adapt to new categories, especially maintain stable classification performance in cross-sensors and cross-regional scenarios.
Smart Images

Figure CN120279430A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine vision, and particularly relates to a remote sensing scene classification method and system based on few-shot learning. Background Art
[0002] Remote sensing image scene classification is one of the core technologies in the fields of smart agriculture, disaster monitoring, environmental management, etc. Its purpose is to accurately classify different scenes (such as ports, beaches, forests, airports, etc.) in remote sensing images through automated algorithms. However, current remote sensing image scene classification technologies face many challenges, mainly concentrated in aspects such as sample learning requirements, feature complexity, and model generalization ability. First of all, the annotation of remote sensing images requires professional knowledge, with high annotation costs and time-consuming. For example, hyperspectral images are vulnerable to atmospheric interference and cloud occlusion, resulting in scarce available annotated data. Traditional deep learning models rely on large-scale annotated data for training, but are prone to overfitting under few-shot conditions and difficult to generalize to new categories. Secondly, remote sensing images have significant intra-class differences (such as spectral differences caused by different lighting conditions and shooting angles) and inter-class similarities (such as similar spectra of different vegetation), which makes it difficult for traditional convolutional neural networks (CNNs) to fully capture multi-modal features, resulting in insufficient classification accuracy. Finally, existing few-shot learning methods have performance degradation due to insufficient feature decoupling in cross-sensor and cross-region scenarios. For example, in disaster emergency monitoring, it is necessary to quickly identify sudden disaster areas, but traditional methods cannot respond in a timely manner due to their dependence on a large amount of annotated data. In summary, existing remote sensing image scene classification technologies have problems such as high annotation costs, insufficient model generalization ability, and difficulty in dealing with feature complexity under few-shot conditions, which limit their wide promotion in practical applications. Therefore, there is an urgent need for a remote sensing image scene classification method that can efficiently learn and quickly generalize to new categories under few-shot conditions. Summary of the Invention
[0003] The present invention proposes a remote sensing scene classification method and system based on few-shot learning to solve the problems existing in the above-mentioned prior art.
[0004] To achieve the above object, the present invention provides a remote sensing scene classification method based on few-shot learning, including the following steps:
[0005] Obtain a remote sensing image data set and divide it into a meta-training set, a meta-validation set, and a meta-test set;
[0006] Construct an initial few-shot classification model based on the improved MobileViT network. The improvements include adding a feature fusion module to the main backbone structure of the MobileViT network, introducing a channel attention mechanism into the MobileNetV2 module of the MobileViT network, and fusing local features and global features in the MobileViT module of the MobileViT network.
[0007] Train the initial few-shot classification model using a meta-training set and a meta-validation set, and test it using a meta-test set. After passing the test, obtain the few-shot classification model, and use the few-shot classification model to classify images to obtain image classification results.
[0008] Optionally, the improved MobileViT network includes:
[0009] Extract multi-scale features through a hierarchical structure including MobileNetV2 modules and MobileViT modules;
[0010] Upsample the multi-scale features at different levels to a unified resolution and then concatenate them to obtain fused features;
[0011] Adaptive weight the fused features through a channel attention mechanism;
[0012] Calculate class prototype vectors based on support set samples and achieve classification through distance metrics.
[0013] Optionally, the channel attention mechanism introduced in the MobileNetV2 module includes:
[0014] Perform global average pooling on the input feature map to obtain channel statistics;
[0015] Generate channel weights through a fully connected layer and an activation function;
[0016] Multiply the channel weights by the original feature map to achieve feature recalibration.
[0017] Optionally, the MobileViT module includes: connecting local features and global features along the channel direction through a shortcut branch to fuse local details and global context information.
[0018] Optionally, the model training adopts the N-way K-shot task training method. Each task includes a support set and a query set. Extract features from the support set and calculate class prototypes, and classify based on the similarity between the query set and the class prototypes to optimize the model parameters.
[0019] Optionally, the calculation method of the class prototype includes: taking the mean of all sample features of each class in the support set to obtain the prototype vector of that class.
[0020] The present invention also provides a remote sensing scene classification system based on few-shot learning, including:
[0021] A dataset partitioning module, configured to obtain a remote sensing image dataset and partition it into a meta-training set, a meta-validation set, and a meta-test set;
[0022] A model construction module, configured to construct an initial few-shot classification model according to the improved MobileViT network;
[0023] A training module, configured to train the initial few-shot classification model through the meta-training set and the meta-validation set;
[0024] A testing module, configured to test the trained model through the meta-test set to obtain a few-shot classification model;
[0025] A classification module, configured to classify an image through the few-shot classification model to obtain an image classification result.
[0026] Optionally, the improved MobileViT network includes:
[0027] A feature extraction layer, configured to extract multi-scale features through a hierarchical structure including MobileNetV2 modules and MobileViT modules;
[0028] A multi-scale feature fusion module, configured to upsample features at different levels to a unified resolution and then splice them to obtain fused features;
[0029] A dynamic attention weighting module, configured to adaptively weight the fused features through a channel attention mechanism;
[0030] A prototype classification head, configured to calculate class prototype vectors based on support set samples and achieve classification through distance measurement.
[0031] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the method.
[0032] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method are implemented.
[0033] Compared with the prior art, the present invention has the following advantages and technical effects:
[0034] First, through the few-shot learning framework, the present invention breaks through the dependence of traditional deep learning on large-scale labeled data. Only 1 to 5 labeled samples are required to accurately identify new ground object categories, significantly reducing the manual labeling cost and solving the problems of high remote sensing image labeling cost and scarce samples. Second, the improved MobileViT network structure enhances the model's collaborative modeling ability for local details and global context, effectively improving the classification accuracy, especially when dealing with remote sensing images with significant intra-class differences and inter-class similarities. In addition, the introduced cross-layer connection and feature fusion mechanism optimize the model's generalization ability, enabling it to maintain stable classification performance in multi-scale and multi-resolution remote sensing images across sensors and regions. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0036] Figure 1 is the structural diagram of the improved MobileViT network of the embodiment of the present invention;
[0037] Figure 2 is the flowchart of few-shot learning of the embodiment of the present invention;
[0038] Figure 3 is the structural diagram of the improved MobileNetV2 module of the embodiment of the present invention;
[0039] Figure 4 is the structural diagram of the channel attention mechanism of the embodiment of the present invention;
[0040] Figure 5 is the structural diagram of the improved MobileViT module of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0042] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0043] The related technologies are introduced as follows:
[0044] MobileViT is a lightweight hybrid network architecture that cleverly combines the advantages of convolutional neural networks (CNNs) and Transformers, taking into account both local detail perception and global context modeling capabilities. Its core design is divided into three parts:
[0045] Local feature extraction: Use the deep separable convolutional layer of MobileNetV2 to efficiently extract local features such as image edges and textures, reducing the amount of calculation;
[0046] Global relationship modeling: Through the self-attention mechanism in the lightweight Transformer block, long-range dependencies between pixels are established to capture the spatial layout of objects (such as the arrangement of buildings and the direction of rivers);
[0047] Cross-layer feature fusion: The alternating stacked CNN-Transformer structure realizes the dynamic interaction between local and global features, avoiding the high computational overhead of pure Transformer.
[0048] Compared with traditional CNN (such as ResNet), MobileViT has more than 70% fewer parameters. In remote sensing image classification tasks, it can not only identify local details (such as vehicle contours) but also associate global scenes (such as road network distribution). It is particularly suitable for resource-constrained small sample learning scenarios, and takes into account both high precision and edge device deployment efficiency.
[0049] The original MobileViT network structure:
[0050] MobileViT is a lightweight hybrid network architecture that aims to integrate the local perception capabilities of convolutional neural networks (CNNs) with the global modeling advantages of Transformers. It is suitable for mobile terminals and resource-constrained scenarios. MobileViT adopts a hierarchical design with staged stacking, which includes multiple stages. Each stage gradually expands the receptive field and compresses the spatial resolution while increasing the channel dimension. The typical structure is as follows:
[0051] Initial convolution layer: Conv3*3 in the figure, standard 3×3 convolution, quickly extracts underlying features (such as edges and corners).
[0052] MobileNetV2 module: MV2 in the figure, composed of multiple inverted residual blocks, extracts local detail features through depthwise separable convolution. It extracts local texture and edge features in a lightweight way, and reduces the number of parameters by more than 80% compared with standard convolution.
[0053] MobileViT Module: The MobileViT block in the figure alternately uses convolutional and lightweight Transformer blocks to achieve local-to-global feature fusion. It captures long-range context through self-attention (such as road networks and farmland distributions in remote sensing images).
[0054] Final Convolutional Layer: Conv1*1 in the figure adjusts the channel dimension and feature fusion.
[0055] Global Pooling Linear Layer: Global Pool Linear in the figure performs global average pooling (GAP) to extract the global statistics (mean) of each channel, making the model robust to spatial translations of the input image (e.g., the position of a building in the image does not affect the classification result). The linear layer (Linear) generates the final classification scores (Logits), which are used to obtain the probability distribution after Softmax normalization.
[0056] Classification Head: Global average pooling (GAP) is followed by a fully connected layer to output Logits (unnormalized classification scores).
[0057] The application of the present invention in smart agriculture is as follows:
[0058] The remote sensing scene classification method based on few-shot learning of the present invention has important application value in the field of smart agriculture. Smart agriculture requires real-time monitoring and precise management of various agricultural scenes such as farmland, orchards, and pastures to improve agricultural production efficiency, optimize resource utilization, and reduce environmental impact. However, traditional remote sensing image classification methods face many challenges in agricultural scenes, such as diverse farmland types, significant seasonal changes, and scarce labeled data, resulting in insufficient classification accuracy and difficulty in quickly adapting to new agricultural scenes.
[0059] Through the few-shot learning framework, the present invention can quickly identify new agricultural scene categories with only a very small number of labeled samples (such as 1 - 5 images), significantly reducing the labeling cost and improving the generalization ability of the model at the same time. The improved MobileViT network structure can efficiently capture the local texture features of farmland (such as crop planting patterns and irrigation ditches) and the global spatial layout (such as farmland distribution and topography), and enhance the discriminative ability for complex agricultural scenes through the feature fusion mechanism. For example, when monitoring the growth status of crops, the model can quickly identify crops at different growth stages and timely detect pest and disease areas; in land use planning, it can accurately distinguish different types of farmland, forest land, and water areas, providing data support for the rational allocation of agricultural resources.
[0060] In addition, the high efficiency and flexibility of the present invention enable it to respond in real time to the dynamic changes in agricultural production, such as sudden natural disasters or pest and disease outbreaks. By quickly classifying and analyzing remote sensing images, agricultural managers can take timely measures to reduce losses and optimize the production process. The experimental results show that the present invention performs excellently in the classification tasks of agricultural-related scenarios, can significantly improve the classification accuracy and work efficiency, and provides strong technical support for the precise management and sustainable development of smart agriculture.
[0061] Example 1
[0062] As Figure 1-5 shown, in this embodiment, a remote sensing scene classification method based on few-shot learning is provided, including the following steps:
[0063] Obtain a remote sensing image dataset and divide it into a meta-training set, a meta-validation set, and a meta-test set;
[0064] Construct an initial few-shot classification model based on the improved MobileViT network. The improvement includes adding a feature fusion module to the main backbone structure of the MobileViT network, introducing a channel attention mechanism into the MobileNetV2 module of the MobileViT network, and fusing local features and global features in the MobileViT module of the MobileViT network;
[0065] Train the initial few-shot classification model through the meta-training set and the meta-validation set, and test it through the meta-test set. After passing the test, obtain the few-shot classification model, and classify the image through the few-shot classification model to obtain the image classification result.
[0066] This embodiment uses a few-shot remote sensing image scene classification method based on the improved MobileViT.
[0067] Flow method (taking the NWPU-RESISC45 dataset as an example):
[0068] 1. Dataset division;
[0069] Total dataset: 45 categories (such as airports, forests, etc.), 700 images for each category.
[0070] Division ratio (example):
[0071] Meta-training set: 30 categories (for training the meta-learning model).
[0072] Meta-validation set: 5 categories (for hyperparameter tuning and early stopping).
[0073] Meta-test set: 10 categories (for finally testing the model generalization ability).
[0074] Key principle: The three categories are strictly mutually exclusive, and the meta-test category is completely invisible during training and validation.
[0075] 2. Meta-Train stage;
[0076] (1) Task generation:
[0077] Generate a 5-way 1-shot task in each iteration: Select categories: Randomly select 5 categories from the 30 categories in the meta-training set.
[0078] Sample the support set: Select 1 image for each category (5 images in total).
[0079] Query set: Select 15 images for each category (75 images in total), which do not overlap with the support set.
[0080] Task format: The input is a pair of images of the support set and the query set, and the output is the classification result of the query set.
[0081] (2) Model training:
[0082] The model inputs the support set (5 images) + the query set (75 images).
[0083] Learning objective: Learn task-specific features based on the support set. Calculate the loss on the query set to optimize the generalization ability of the model.
[0084] Furthermore, the model training adopts the N-way K-shot task training method. Each task includes a support set and a query set. Extract features from the support set and calculate the class prototypes. Classify based on the similarity between the query set and the class prototypes to optimize the model parameters.
[0085] N-way K-shot: way represents the number of categories, and shot represents the number of images in each category. 5-way 1-shot means there are 5 categories and 1 image in each category. In the 5-way 1-shot task, two sets are defined, namely the support set and the query set. The goal of the few-shot learning task is to determine which category in the support set the image in the query set belongs to.
[0086] The goal of few-shot classification is to enable the model to quickly adapt to new tasks with only a very small number of samples (e.g., 5 classes × 1 sample). The essential difference from traditional supervised learning is that traditional supervised learning inputs all classes and samples at once, and the model directly learns the classification boundary (such as ResNet classification). The training process is decomposed into a large number of subtasks (Episodes), and each task simulates a "few-shot scenario". The model learns how to generalize knowledge from a small number of samples. Although only a small amount of data is used for each task, all meta-training classes and samples will be covered multiple times through repeated random sampling. The optimization goal of meta-learning is the generalization ability at the task level, rather than the classification accuracy of a single task. Through the training of tens of thousands of tasks, the model learns how to quickly extract features from any 5 classes, rather than memorizing the details of all classes.
[0087] The innovation and advantages of few-shot learning methods in remote sensing image classification are mainly reflected in that they fundamentally change the dependence of traditional deep learning on large-scale labeled data. By simulating the ability of humans to quickly learn and generalize from a small number of samples, a knowledge transfer paradigm closer to the actual application scenario is constructed.
[0088] The overall workflow of using MobileViT for few-shot learning is as Figure 2 shown.
[0089] 3. Meta-Validation (Meta-Val) stage;
[0090] (1) Task generation:
[0091] Generate validation tasks from 5 classes in the meta-validation set, using the same method as Meta-Train: generate 5-way 1-shot tasks each time.
[0092] (2) Validation process:
[0093] Fix the model parameters: Use the trained model without any parameter updates.
[0094] Evaluation metric: Calculate the classification accuracy of the query set.
[0095] Purpose: Adjust hyperparameters (such as learning rate, number of tasks).
[0096] 4. Meta-Test stage;
[0097] (1) Task generation:
[0098] Generate test tasks from 10 classes in the meta-test set: Generate multiple groups of 5-way 1-shot tasks (set 1000 tasks) to ensure coverage of all class combinations.
[0099] (2) Test process:
[0100] Strict zero-shot: The model has never seen the categories and samples of the meta-test set.
[0101] Evaluation metric: Report the average classification accuracy.
[0102] Example result: The model achieves 70% accuracy on the 5-way 1-shot task, indicating its ability to quickly adapt to new categories.
[0103] Furthermore, in the MobileViT network structure, the features extracted by three MobileViT modules are fused. The improved network structure is as Figure 1 shown:
[0104] Feature extraction layer: The input remote sensing image (256×256×3) is preprocessed by a lightweight convolutional layer (Conv3x3) and then passes through five levels (Layer1~5) in sequence. Among them, Layer1~2 are stacked with pure MV2 (MobileNetV2 inverted residual module), and Layer3~5 are composed of cascaded MV2 modules and MobileViT modules.
[0105] Multi-scale feature fusion module: Extract the MobileViT output features of Layer3 (32×32×256), Layer4 (16×16×384), and Layer5 (8×8×512), upsample them to a unified resolution (32×32) by bilinear interpolation, align the number of channels through 1×1 convolution, and then concatenate them along the channel dimension to form fused features.
[0106] Dynamic attention weighting: After fusing the features, a channel attention module is introduced. The channel weights are generated through global average pooling and fully connected layers to achieve adaptive weighting of multi-scale features.
[0107] Prototype classification head: The fused features are input into the prototype head after global average pooling. The class prototype vectors are calculated based on the support set samples, and the query image classification is achieved through negative Euclidean distance measurement.
[0108] Compared with the original MobileViT network, the improved network shows the following advantages in the remote sensing few-shot scene classification task:
[0109] Enhanced multi-scale feature complementarity: The original MobileViT only relies on the features of a single level (usually Layer5), making it difficult to capture both high-resolution details (such as building textures) and low-resolution semantics (such as the overall layout of the airport) simultaneously. The improved network constructs a hybrid representation containing local details and global context by fusing multi-level features of Layer3~5, significantly improving the discriminative ability for complex scenes.
[0110] Generalization Optimization for Small Samples: Remote sensing images have significant intra-class differences (such as farmland in different seasons) and inter-class similarities (such as ports and docks). The original network is prone to overfitting due to insufficient feature diversity, while the fused features suppress noise interference (such as cloud occlusion) through the attention mechanism and strengthen discriminative regions (such as the shape of airport runways).
[0111] Furthermore, the MV2 module is improved by adding an attention mechanism;
[0112] In the MobileViT architecture, the MV2 module is used as a building block to provide efficient feature extraction capabilities. This module adopts an "inverted bottleneck" structure. During the convolution process, the features go through two steps of dimensionality increase and decrease. These techniques help reduce the computational complexity of the model while maintaining good performance. The improved algorithm of the present invention adds a channel attention mechanism module based on the original residual result, enhancing the representation ability of the image channels. The structure of the improved MV2 module is as Figure 3 shown.
[0113] The attention mechanism is used to establish associations between two different input sequences. It can simultaneously focus on relevant elements in both sequences and learn the interaction relationships between them. The role of the channel attention mechanism is to enable the network to automatically identify and assign different importance to different channels, thereby making more efficient use of feature information. This mechanism is achieved by assigning different weights to the features of each channel, enabling the model to focus on processing more useful information while ignoring relatively unimportant parts. This not only helps improve the expressiveness of the model but also enhances the model's ability to understand complex data structures. The structure of the channel attention mechanism is as Figure 4 shown.
[0114] For the input feature map, the channel attention mechanism first uses global pooling to obtain channel statistics of 1×1×N. After the channel statistics are reduced by m times and then increased by m times, the weight coefficients of each channel are normalized between 0 and 1 through activation functions (sigmoid function and ReLU activation function). In this way, the weight coefficient of each channel can be interpreted as a probability, representing the importance of that channel. Finally, these weight coefficients are multiplied by each channel of the original input feature map. This step is feature recalibration, that is, adaptively adjusting the contribution of each channel through the learned channel weights. In this way, the model can emphasize the feature channels that are more important for the current task and suppress those unimportant channels.
[0115] Furthermore, the improved MobileViT module is as Figure 5 shown:
[0116] MobileViT is a lightweight hybrid network architecture that cleverly combines the local perception capabilities of convolutional neural networks (CNNs) with the global modeling advantages of Transformers. It is designed specifically for mobile terminals and resource-constrained scenarios. Its core module, the MobileViT Block, achieves local-to-global feature fusion through alternating stacked convolutions and Transformer operations. The structure of the MobileViT module is shown in the figure, where the red line shows the improvement of the original MobileViT module in this invention: before extracting global feature information, we introduce a shortcut branch that directly connects the obtained global feature information with the local feature information along the channel direction. This improvement makes greater use of local features, allowing the network to more fully extract feature information of remote sensing scene images, especially in capturing details.
[0117] After the image features are extracted by the MobileViT main network, small sample classification is achieved through metric learning. The core idea is that in the feature space, each category can be represented by the feature mean (prototype) of its support set samples, and classification is completed by calculating the similarity between the query sample and all prototypes.
[0118] For N-way K-shot tasks, support set features ,D is the feature dimension, the prototype of each category Calculated as:
[0119] ;
[0120] in Represents the characteristics of the k-th support sample of the n-th class.
[0121] Query sample features , calculate its negative Euclidean distance to all prototypes:
[0122] ;
[0123] The final classification probability is obtained by softmax normalization:
[0124] .
[0125] The following experiments were also conducted in this embodiment:
[0126] NWPU-RESISC45 is a dataset for remote sensing scene classification established by Northwestern Polytechnical University. The dataset contains 31,500 samples in 45 categories, each of which contains 700 RGB images with a sample size of 256×256. The dataset is divided into 25, 10, and 10 categories for training, validation, and testing.
[0127] The UC Merced is a classic remote sensing scene classification dataset proposed by the UC Merced Computer Vision Laboratory in 2010 for classifying land use scenes in urban areas. The spatial resolution of this dataset is approximately 0.3m, the image scale is 256×256, it contains 21 types of scenes, with 100 RGB images for each type, totaling 2,100 images. 10 categories in this dataset are used for training, 6 categories for validation, and 5 categories for testing.
[0128] Experimental results:
[0129] The classification accuracy of 5-way 1-shot on the NWPU-RESISC45 dataset is shown in Table 1:
[0130] Table 1
[0131] Method Accuracy (%) MAML 45.78 Prototypical Network 46.21 MatchingNet 48.62 Relation Network 56.49 DeepEMD 62.75 ResNet50 57.93 Ours 74.67
[0132] The classification accuracy of 5-way 1-shot on the UC Merced dataset is shown in Table 2:
[0133] Table 2
[0134] Method Accuracy (%) MAML 47.84 Prototypical Network 46.87 MatchingNet 51.23 Relation Network 51.87 DeepEMD 59.75 ResNet50 52.17 Ours 75.16
[0135] From the experimental results, it can be seen that the present invention has achieved the following effects:
[0136] 1. Break through the dependence of traditional deep learning on large-scale labeled data: Aiming at the pain points of high cost of remote sensing image annotation and frequent dynamic addition of new categories, through the few-shot learning framework, it can accurately identify new ground object categories with only 1 to 5 labeled samples, significantly reducing the manual annotation cost.
[0137] 2. Improve the collaborative modeling ability of local features and global context: Aiming at the misclassification problems (such as confusing morphologically similar farmland and grassland) caused by the limited field of view of local convolution in existing convolutional neural networks (such as ResNet) in remote sensing images, a MobileViT improved structure that integrates lightweight convolution and self-attention mechanism is designed to capture the correlation between the detailed texture of ground objects and the spatial layout at the same time.
[0138] 3. Optimize the cross-task adaptation efficiency of lightweight models: Aiming at the performance degradation caused by insufficient feature decoupling in existing few-shot methods in cross-sensor and cross-region scenarios, by introducing cross-layer connections and feature fusion mechanisms, the generalization ability of the model to multi-scale and multi-resolution remote sensing images is improved.
[0139] This embodiment also proposes a remote sensing scene classification system based on few-shot learning, including:
[0140] The dataset partitioning module is used to obtain a remote sensing image dataset and partition it into a meta-training set, a meta-validation set, and a meta-test set;
[0141] The model construction module is used to construct an initial few-shot classification model based on the improved MobileViT network;
[0142] The training module is used to train the initial few-shot classification model through the meta-training set and the meta-validation set;
[0143] The testing module is used to test the trained model through the meta-test set to obtain a few-shot classification model;
[0144] The classification module is used to classify an image through the few-shot classification model to obtain an image classification result.
[0145] Furthermore, the improved MobileViT network includes:
[0146] The feature extraction layer is used to extract multi-scale features through a hierarchical structure including MobileNetV2 modules and MobileViT modules;
[0147] The multi-scale feature fusion module is used to upsample the features of different levels to a unified resolution and then splice them to obtain fused features;
[0148] The dynamic attention weighting module is used to adaptively weight the fused features through a channel attention mechanism;
[0149] The prototype classification head is used to calculate class prototype vectors based on support set samples and achieve classification through distance metrics.
[0150] This embodiment also proposes a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method.
[0151] This embodiment also proposes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method are implemented.
[0152] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A remote sensing scene classification method based on few-shot learning, characterized in that, Including the following steps: Obtain a remote sensing image dataset and divide it into a meta-training set, a meta-validation set, and a meta-test set; Construct an initial few-shot classification model based on the improved MobileViT network. The improvement includes adding a feature fusion module to the main backbone structure of the MobileViT network, introducing a channel attention mechanism into the MobileNetV2 module of the MobileViT network, and fusing local features and global features in the MobileViT module of the MobileViT network; Train the initial few-shot classification model through the meta-training set and the meta-validation set, and test it through the meta-test set. After passing the test, obtain the few-shot classification model, and classify the image through the few-shot classification model to obtain the image classification result.
2. The method according to claim 1, wherein The improved MobileViT network includes: Extract multi-scale features through a hierarchical structure including a MobileNetV2 module and a MobileViT module; Upsample the multi-scale features of different levels to a unified resolution and then splice them to obtain fused features; Adaptive weight the fused features through a channel attention mechanism; Calculate the class prototype vector based on the support set samples and achieve classification through distance measurement.
3. The method according to claim 2, wherein The channel attention mechanism introduced in the MobileNetV2 module includes: Perform global average pooling on the input feature map to obtain channel statistics; Generate channel weights through a fully connected layer and an activation function; Multiply the channel weights by the original feature map to achieve feature recalibration.
4. The method according to claim 2, characterized in that The MobileViT module includes: Connect local features and global features along the channel direction through a shortcut branch to fuse local details and global context information.
5. The method according to claim 1, wherein The model training adopts the N-way K-shot task training method. Each task includes a support set and a query set. Extract features from the support set and calculate the class prototype, and classify based on the similarity between the query set and the class prototype to optimize the model parameters.
6. The method according to claim 5, wherein The calculation method of the class prototype includes: Take the mean of all sample features of each class in the support set to obtain the prototype vector of this class.
7. A remote sensing scene classification system based on few-shot learning, characterized in that, Including: A dataset division module for obtaining a remote sensing image dataset and dividing it into a meta-training set, a meta-validation set, and a meta-test set; A model construction module for constructing an initial few-shot classification model according to the improved MobileViT network; A training module for training the initial few-shot classification model through the meta-training set and the meta-validation set; A testing module for testing the trained model through the meta-test set to obtain the few-shot classification model; A classification module for classifying an image through the few-shot classification model to obtain the image classification result.
8. The system according to claim 7, wherein, The improved MobileViT network includes: A feature extraction layer for extracting multi-scale features through a hierarchical structure including a MobileNetV2 module and a MobileViT module; A multi-scale feature fusion module for upsampling the features of different levels to a unified resolution and then splicing them to obtain fused features; A dynamic attention weighting module, which is used to adaptively weight the fusion features through a channel attention mechanism; A prototype classification head, which is used to calculate class prototype vectors based on support set samples and achieve classification through distance measurement.
9. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1-6 are implemented.
Citation Information
Patent Citations
Small sample image classification method and device based on attention mechanism
CN117542075A
Infrared human body posture estimation method based on lightweight ViT and attention mechanism
CN117877122A
Remote sensing image classification method, equipment and device based on small samples
CN118015322A
Image classification method based on improved MobileViT model, electronic equipment and readable storage medium
CN118840608A