A multi-scale token selection and cross-bone token distillation remote sensing scene classification method
Patent Information
- Application Number
- CN202611017381.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-29
AI Technical Summary
现有知识蒸馏方法通常仅通过输出Logits进行监督,难以充分传递教师模型内部的结构化语义信息与多尺度特征融合能力,导致轻量化模型性能受限
(1)本发明通过多尺度Token选择Transformer模块,以深层语义Token作为查询中心,对浅层Token进行选择性稀疏路由,有效增强了模型对关键局部结构信息与全局语义信息的联合建模能力,降低背景冗余信息干扰,提高遥感场景分类精度。
Smart Images

Figure CN122841955A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent interpretation of remote sensing images and computer vision, specifically involving a remote sensing scene classification method with multi-scale token selection and cross-backbone token distillation. Background Technology
[0002] Remote sensing scene classification is an important research direction in remote sensing image understanding. Its goal is to assign corresponding semantic category labels to remote sensing scenes based on the distribution of ground features, spatial structure, and semantic information contained in remote sensing images. Remote sensing scene classification is widely used in land use surveys, urban planning, environmental monitoring, disaster assessment, crop monitoring, and military reconnaissance, and has significant research and application value.
[0003] With the rapid development of high-resolution remote sensing imaging technology, remote sensing images contain richer spatial texture details and more complex scene structures. While this enhances scene representation capabilities, it also increases the difficulty of remote sensing scene classification tasks. Compared to natural images, remote sensing images typically exhibit problems such as large intra-class variability, small inter-class variability, significant target scale variations, severe background interference, and complex spatial layouts, making it difficult for traditional classification methods to obtain stable and accurate classification results.
[0004] In recent years, deep learning-based remote sensing scene classification methods have made significant progress. Convolutional neural networks (CNNs) can learn local texture features and high-level semantic features in remote sensing images through hierarchical feature extraction, and have gradually become the mainstream method in the field of remote sensing scene classification. However, most existing methods only utilize deep features for classification, which can easily lead to the loss of shallow detail information. Although some methods employ multi-scale feature fusion strategies, they usually use direct concatenation or weighted fusion methods, lacking effective modeling of the importance of features at different levels, which can easily introduce redundant background information and reduce the model's feature selection ability.
[0005] Furthermore, the success of the Transformer structure in visual tasks has further promoted the development of remote sensing scene classification. However, traditional Transformers usually adopt a global dense attention mechanism, which has high computational complexity and is easily affected by repeated texture regions and complex background regions in remote sensing images, making it difficult to achieve efficient and effective cross-level feature interaction.
[0006] On the other hand, in practical remote sensing deployments, models not only need high classification accuracy but also require lightweight design and low computational overhead. Existing knowledge distillation methods typically rely solely on output logits for supervision, which fails to fully convey the structured semantic information and multi-scale feature fusion capabilities within the teacher model, thus limiting the performance of lightweight models.
[0007] Therefore, how to effectively utilize multi-level feature information in the process of remote sensing scene classification, achieve selective fusion of key tokens, and improve the learning ability of student models to multi-scale semantic structures while ensuring model lightweightness has become an important technical problem that urgently needs to be solved in the field of remote sensing scene classification. Summary of the Invention
[0008] The purpose of this invention is to provide a remote sensing scene classification method based on multi-scale token selection and cross-backbone token distillation. By constructing a multi-scale token selection Transformer (MSTS) module and a cross-backbone level token distillation (CHTD) mechanism, it achieves efficient modeling and lightweight knowledge transfer of multi-level semantic information in remote sensing images, thereby improving the accuracy and robustness of remote sensing scene classification.
[0009] The technical solution of the present invention includes the following steps: Step S1: Input remote sensing scene images and construct a remote sensing scene classification dataset, dividing the dataset into a training set and a test set; Step S2: Construct a remote sensing scene classification model based on Multi-Scale Token Selection Transformer (MSTS); Step S3: Train the remote sensing scene classification model from Step S2 using the training set, and use the cross-backbone level token distillation (CHTD) mechanism to realize knowledge transfer from the teacher model to the student model; Step S4: Use the trained remote sensing scene classification model to classify and predict the remote sensing scene images in the test set to obtain the remote sensing scene classification results. The main features of the remote sensing scene classification model in step S2 are: S2.1: Construct teacher and student models. The teacher model uses ResNet50 as the backbone network, and the student model uses ResNet18 as the backbone network. S2.2: Both the teacher model and the student model include a backbone feature extraction module, a multi-scale token selection Transformer (MSTS) module, a feature enhancement module (GCE), and a classification module; S2.3: The backbone feature extraction module is used to extract multi-level feature representations of the input remote sensing image, where shallow features contain rich texture and edge information, and deep features contain high-level semantic information; S2.4: The MSTS module includes a channel alignment unit, a tokenization unit, a token routing unit, a gated aggregation unit, and a fusion and reconstruction unit; S2.4.1: Channel alignment units are used to map features from different levels to a unified channel dimension; S2.4.2: The tokenization unit is used to convert multi-level features into a token sequence representation; S2.4.3: The Token routing unit uses deep tokens as query tokens and shallow tokens as candidate tokens, and filters key tokens through the Top-k routing strategy; S2.4.4: The gating aggregation unit uses gating weights to perform weighted fusion of the filtered key tokens; S2.4.5: The fusion reconstruction unit is used to restore the fused tokens into a spatial feature map and generate the final fused feature representation; S2.5: The feature enhancement module is used to further enhance the contextual expressiveness of fused features; S2.6: The classification module is used to output the final remote sensing scene category prediction results; The main feature of the model training process in step S3 is that: S3.1: The total loss function used in training the remote sensing scene classification model consists of classification loss, Logits distillation loss, deep feature distillation loss, and fusion token distillation loss. The total loss function is defined by formula (1): Ltotal = LCE + αkdLKD + αfitLFitNet + αtokLToken; S3.2: Classification loss LCE is used to measure the difference between the predicted class and the true class; S3.3: Logits distillation loss (LKD) is used to constrain the difference between the output distributions of the teacher model and the student model; S3.4: Deep Feature Distillation Loss LFitNet is used to align the deep semantic features of the teacher model and the student model; S3.5: The fusion token distillation loss LToken is used to constrain the consistency of the fusion token representation between the teacher model and the student model; S3.6: The cross-backbone level token distillation mechanism includes three levels: Logits distillation, deep feature distillation, and fused token distillation.
[0010] Compared with existing methods, the present invention has the following advantages: (1) This invention selects the Transformer module through multi-scale Tokens, uses deep semantic Tokens as the query center, and performs selective sparse routing on shallow Tokens, which effectively enhances the model’s ability to jointly model key local structural information and global semantic information, reduces background redundant information interference, and improves the accuracy of remote sensing scene classification.
[0011] (2) This invention uses the Top-k Token routing mechanism to retain only candidate tokens that are highly relevant to the current query token, thereby achieving sparse interaction of cross-level features, reducing the computational complexity of the traditional dense attention mechanism, and improving the model's running efficiency.
[0012] (3) This invention achieves dynamic fusion of multi-scale token information through a gating aggregation mechanism, which enhances the model’s adaptability to changes in target scale and complex spatial structures and improves the model’s robustness.
[0013] (4) This invention proposes a cross-backbone level token distillation mechanism, which realizes the structured knowledge transfer from the teacher model to the lightweight student model by jointly constraining the output Logits, deep semantic features and fused token representation, effectively improving the classification performance of the lightweight model.
[0014] (5) Under the condition of ensuring a low number of model parameters and computational complexity, the present invention improves the classification accuracy and generalization ability of the remote sensing scene classification model in complex remote sensing scenes, and is suitable for remote sensing intelligent interpretation tasks under resource-constrained conditions. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the overall process of the present invention; Figure 2 This is a schematic diagram of the overall structure of the remote sensing scene classification model based on MSTS and CHTD of the present invention; Figure 3 This is a schematic diagram of the Multi-Scale Token Selection Transformer (MSTS) module structure of the present invention. Detailed Implementation
[0016] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, but the scope of protection of the present invention is not limited to the following embodiments.
[0017] like Figure 1 As shown, this invention provides a remote sensing scene classification method based on multi-scale token selection and cross-backbone level token distillation, comprising the following steps: Step S1: Obtain the remote sensing scene image dataset and divide the dataset into a training set and a test set.
[0018] Specifically, this embodiment uses the UCM, AID, and NWPU-RESISC45 public remote sensing scene classification datasets as experimental datasets. The UCM dataset contains 21 scene categories, the AID dataset contains 30 scene categories, and the NWPU-RESISC45 dataset contains 45 scene categories.
[0019] Step S2: Construct a remote sensing scene classification model based on Multi-Scale Token Selection Transformer (MSTS).
[0020] Specifically, such as Figure 2 As shown, the model consists of two parts: a teacher network and a student network. The teacher network uses ResNet50 as the backbone network, and the student network uses ResNet18 as the backbone network. Both include a backbone feature extraction module, an MSTS module, a GCE feature enhancement module, and a classification module.
[0021] Furthermore, such as Figure 3 As shown, the MSTS module includes a channel alignment unit, a tokenization unit, a top-k token routing unit, a gated aggregation unit, and a fusion reconstruction unit.
[0022] The system includes a channel alignment unit for unifying the channel dimensions of features at different levels; a tokenization unit for converting spatial features into token sequences; a token routing unit for using deep tokens as query tokens and selecting candidate tokens with the highest relevance from shallow tokens; a gating aggregation unit for dynamically fusing key tokens using a gating mechanism; and a fusion reconstruction unit for remapping the fused tokens into a spatial feature map to obtain the final fused feature representation.
[0023] Step S3: Train the model using the training set.
[0024] Specifically, a cross-backbone level token distillation mechanism is used to jointly train the teacher model and the student model. The teacher model provides classification decision knowledge, multi-scale fusion knowledge, and deep semantic knowledge, while the student model learns the knowledge representation of the teacher model through Logits distillation, deep feature distillation, and fused token distillation.
[0025] Step S4: Use the trained remote sensing scene classification model to classify and predict the remote sensing scene images in the test set, and output the final remote sensing scene category.
[0026] The effects of this invention can be further illustrated by the following experiments.
[0027] Experimental setup Specifically, for the benchmark dataset, this invention uses five different datasets: Kvasir-SEG, CVC-ClinicDB, CVC-ColonDB, CVC-300, and ETIS. The experiments in this invention follow the same principles as the other methods, using images randomly selected from Kvasir and CVC-ClinicDB for training, totaling 1450 images. Specific information is shown in Table 1.
[0028] Table 1
[0029] Experimental results We compared our proposed method with representative Remote Sensing Scene Classification (RSSC) methods, including CrossViT-15-D, Swin-T, MF2CNet, TCNN, EMTCAL, SKAL, AGOS, MSGNet, SAGN, EMSCNet, EFCOMFF-Net, MGML, HFAM, CGINet, UPetu, and MSCN. Quantitative results are shown in Tables 2, 3, and 4.
[0030]
[0031] Table 1 presents the experimental results of each method on the UCM dataset. Although the UCM dataset is close to performance saturation, the proposed MSTS method achieves the best classification accuracy under both training ratios. This indicates that even in remote sensing scene classification tasks with small sample sizes and relatively low classification difficulty, the MSTS module can still provide additional discriminative power, thereby further improving model performance.
[0032]
[0033] Table 2 presents the experimental results of the proposed method on the AID dataset. The results show that the MSTS method maintains strong competitiveness under both training ratios. Although its performance improvement compared to some state-of-the-art methods is relatively limited, MSTS achieves stable and excellent classification performance under both small and large training sample ratios, demonstrating good generalization ability and robustness.
[0034]
[0035] For the larger and more challenging NWPU-RESISC45 dataset, the results shown in Table 3 further validate the effectiveness of MSTS. Compared to the currently superior MSCN method, MSTS achieves higher classification accuracy at both 10% and 20% training ratios. Specifically, at the 20% training ratio setting, MSTS achieves an overall classification accuracy of 95.12%, significantly outperforming existing comparative methods. Experimental results demonstrate that the proposed selective multi-scale token routing mechanism can effectively mine discriminative information from cross-level features, showing significant advantages for complex remote sensing scene classification tasks with stronger intra-class differences and inter-class similarities.
[0036] Overall, MSTS achieved competitive or even superior performance on three public remote sensing scene classification benchmark datasets: UCM, AID, and NWPU-RESISC45, fully demonstrating the effectiveness and superiority of multi-scale token selection and sparse routing mechanisms in remote sensing scene classification tasks.
[0037] Experimental results show that the present invention can effectively improve the adaptability of remote sensing scene classification models to complex scene structures, multi-scale targets and background interference, and improve classification accuracy and generalization performance while ensuring the model is lightweight.
[0038] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A remote sensing scene classification method based on multi-scale token selection and cross-backbone token distillation, characterized mainly by: Specifically, the following steps are included: Step S1: Obtain the remote sensing scene image dataset and divide the remote sensing scene image dataset into a training set and a test set; Step S2: Construct a remote sensing scene classification model based on Multi-Scale Token Selection (MSTS); Step S3: Use the training set from Step S1 to train the remote sensing scene classification model from Step S2 to obtain the remote sensing scene classification model. Step S4: Use the remote sensing scene classification model trained in step S3 to classify and predict the remote sensing scene images in the test set to obtain the remote sensing scene classification results. The main feature of the remote sensing scene classification model based on multi-scale token selection (MSTS) in step S2 is: S2.1: Multi-level feature extraction is performed on the input remote sensing scene image through a convolutional neural network backbone network to obtain feature map representations at different levels. The feature maps include shallow local structural features and deep semantic features. S2.2: The channel alignment module performs unified channel dimension mapping on feature maps of different levels, and converts the mapped feature maps into token sequence representations; S2.3: Using the deepest token sequence as the query token and the shallowest token sequence as the key token and value token, construct cross-level token association relationships through query-key similarity calculation; S2.4: Use the Top-k Token routing mechanism to perform sparse filtering on the cross-level token association relationship and retain the candidate token with the highest relevance to the currently queried token; S2.5: The filtered tokens are weighted and fused through a gating aggregation mechanism to generate a multi-scale fused token representation; S2.6: The multi-scale fused token representation is restored to a spatial feature map through the fusion reconstruction module, and the final fused feature representation is obtained by using the channel recalibration mechanism; S2.7: Use the feature enhancement module to perform context enhancement on the final fused feature representation, and output the remote sensing scene classification result through the classifier; The main feature of the model training process in step S3 is that: S3.1: The total loss function used during the training of the remote sensing scene classification model consists of classification loss, Logits distillation loss, deep feature distillation loss, and fusion token distillation loss. The total loss function is defined by formula (1): ; Where LCE represents classification loss, LKD represents Logits distillation loss, LFitNet represents deep feature distillation loss, LToken represents fused token distillation loss, and αkd, αfit, and αtok represent the weight coefficients of the corresponding loss terms. S3.2: The classification loss LCE is used to measure the difference between the model's predicted results and the true class labels. Its classification loss is defined by formula (2): ; Where yc represents the true class label, pc represents the model predicted probability, and Ccls represents the total number of classes; S3.3: The Logits distillation loss (LKD) is used to constrain the difference between the output probability distributions of the teacher model and the student model. Its distillation loss is defined by formula (3): ; Where zs represents the student model output Logits, zt represents the teacher model output Logits, T represents the distillation temperature coefficient, and KL represents the KL divergence function; S3.4: The deep feature distillation loss LFitNet is used to align the deep semantic features of the teacher model and the student model. Its feature distillation loss is defined by formula (4): ; Where ŝf4 represents the deep features of the student model after feature adaptation, tf4 represents the deep features of the teacher model, and MSE represents the mean squared error function. S3.5: The fusion token distillation loss LToken is used to constrain the consistency of the multi-scale fusion token representation between the teacher model and the student model. Its token distillation loss is defined by formula (5): ; Where ŝtok represents the student model fusion token representation after projection mapping, and ttok represents the teacher model fusion token representation; S3.6: Construct a cross-backbone level Token Distillation (CHTD) mechanism to transfer knowledge from the teacher model to the student model; S3.7: The teacher model adopts the MSTS remote sensing scene classification model based on ResNet50, and the student model adopts the MSTS remote sensing scene classification model based on ResNet18. S3.8: By jointly optimizing the classification loss, Logits distillation loss, deep feature distillation loss, and fused token distillation loss, the final remote sensing scene classification model is obtained.