Ground feature recognition method based on inverse residual cross-header knowledge distillation neural network
Through the inverted residual cross-head knowledge distillation neural network, combined with the inverted residual attention module, multi-scale spatial attention module and cross-head knowledge distillation structure, the problem that the existing land object recognition methods cannot effectively extract image features is solved, and higher recognition accuracy and better generalization ability are achieved.
Patent Information
- Application Number
- CN202510123684.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-27
AI Technical Summary
The existing land object recognition methods cannot effectively extract image features, resulting in low recognition accuracy.
A land object recognition method based on inverted residual cross-head knowledge distillation neural network is proposed. Through the inverted residual attention module, multi-scale spatial attention module and cross-head knowledge distillation structure, combined with convolution and self-attention mechanism, the multi-scale features of the image are extracted and knowledge distilled.
Effectively extracting image features improves the accuracy of land object recognition, overcomes the shortcomings of the traditional distillation method, and shows better performance on multiple remote sensing scene classification data sets.
Smart Images

Figure CN120047828A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for ground object recognition. Background Art
[0002] Remote Sensing (RS) scene image classification, as a key application field of remote sensing technology, is of irreplaceable importance in aspects such as interpreting Earth's surface information, urban planning, and agricultural monitoring [1]-[4]. With the rapid development of deep learning, especially the rise of Convolutional Neural Networks (CNNs) and Transformers, the field of image classification faces new opportunities and challenges.
[0003] Before the rise of deep learning, traditional remote sensing image classification methods mainly relied on manually designed features (such as texture features [5][6], spectral features [7][8], color features [9]
[10] , and shape features
[11]
[12] ) and machine learning classifiers (Support Vector Machines
[13] and Decision Trees
[14] ). These methods usually involve a large amount of domain expert knowledge and often struggle to capture rich feature information in complex remote sensing scenes. Although such methods have solved the remote sensing image classification problem to a certain extent, with the increase in data scale and the improvement of task complexity, their performance has gradually been limited.
[0004] With the rise of deep learning, Convolutional Neural Networks have achieved great success in the field of image processing. The introduction of convolutional layers enables the network to effectively capture local features of images, while reducing the number of parameters through weight sharing and improving the computational efficiency of the model. Convolutional Neural Networks have performed excellently in remote sensing image processing
[15] -
[17] and have been successfully applied in fields such as agricultural monitoring and urban planning.
[0005] However, traditional Convolutional Neural Networks have certain limitations in processing global information and sequential data, ignoring long-range context information. As can be seen from Figure 2 (a) below, this scene contains many different land covers, such as "airplane", "runway", "parking lot", and "tree". If the model only focuses on local structural features, the "airport" scene may be misclassified as other scenes. Therefore, the expected classification model should consider local land covers and their context relationships, especially long-range connections. As Figure 2In (b), when the relationships among "airplane", "runway", and "building" are well explained, this scene can be correctly classified as "airport". At this time, Vision Transformer
[18] , as a powerful sequence modeling tool, has attracted wide attention. The Transformer achieves global modeling of sequences through the self-attention mechanism, leading to great success in fields such as natural language processing. In remote sensing image classification, the introduction of the Transformer provides a new idea for dealing with the problems of global and long-range dependencies. The model can better understand the context information of the entire scene, which helps to improve the recognition accuracy of complex ground objects.
[0006] Due to the multi-scale characteristics of ground objects in remote sensing scene images, Dilated Convolution
[19] , as an important convolutional variant, has emerged. Dilated Convolution introduces an adjustable dilation rate in the convolutional kernel, enabling the network to more widely perceive context information while maintaining local information. This characteristic provides an effective solution to the problem that traditional convolution is prone to information loss when dealing with multi-scale features.
[0007] In order to fully utilize the respective advantages of the Transformer and convolution, researchers have begun to explore their fusion methods. The model that combines the Transformer and convolution can not only effectively process local and global information but also has better generalization ability and adaptability.
[0008] As a method of model compression, distillation learning brings new possibilities to remote sensing scene image classification. In 2015, the concept of knowledge distillation
[20] was proposed by Hinton et al. and has been widely applied in various fields and tasks. By transferring the knowledge of a complex model to a relatively simple model, usually, the complex model is called the "teacher model", and the simplified model is called the "student model". By using the output of the teacher model as the training target of the student model, the student model obtains dark knowledge from the teacher model. However, sometimes there will be a conflict between the true label and the distillation target. As Figure 3 shown, the true label is "dense residential area", while the predicted probability of "medium-sized residence" is the highest in the teacher's prediction output. The student's prediction imitates both the ground truth label and the teacher's prediction. This will affect the learning of the student model to a certain extent. Summary of the Invention
[0009] The purpose of the present invention is to solve the problem that the existing ground object recognition methods cannot effectively extract the features of images, resulting in low accuracy of ground object recognition, and to propose a ground object recognition method based on an inverted residual cross-head knowledge distillation neural network.
[0010] The specific process of the ground object recognition method based on the inverted residual cross-head knowledge distillation neural network is as follows:
[0011] I. Construct an inverted residual cross-head knowledge distillation model; the specific process is as follows:
[0012] The inverted residual cross-head knowledge distillation model includes a classifier network and a cross-head knowledge distillation structure CHKD;
[0013] The classifier network includes a convolutional block, Block1, Block2, the first inverted residual attention module IRAM with self-attention fusion, the second inverted residual attention module IRAM with self-attention fusion, a multi-scale spatial attention module MSSA, and a first linear layer Linear;
[0014] The cross-head knowledge distillation structure includes feature aggregation and cross-head distillation;
[0015] The feature aggregation includes the first Inception former, the second Inception former, the third Inception former, the fourth Inception former, downsampling Down, a first Fusion module, and a second Fusion module;
[0016] The cross-head distillation part includes a second linear layer Linear;
[0017] The classifier network is also called the student network;
[0018] The feature aggregation is also called the teacher network;
[0019] The specific working process of the inverted residual cross-head knowledge distillation model is as follows:
[0020] The feature map is sequentially input into the convolutional block and Block1 in the classifier network, and Block1 outputs the feature α;
[0021] The feature α is input into Block2, and Block2 outputs the feature α′;
[0022] The feature α′ is input into the first inverted residual attention module with self-attention fusion, and the first inverted residual attention module with self-attention fusion outputs the feature α″′;
[0023] The output feature α″′ of the first inverted residual attention module with self-attention fusion is input into the second inverted residual attention module with self-attention fusion, and the second inverted residual attention module with self-attention fusion outputs the feature
[0024] Feature The input is a multi-scale spatial attention module, and the multi-scale spatial attention module outputs the feature β;
[0025] The feature α is input into the first Inception former, and the output of the first Inception former is input into the downsampling Down, and the output of the downsampling Down is the feature D;
[0026] The output feature α′ of Block2 is input into the second Inception former, and the output of the second Inception former is the feature E;
[0027] The output feature D of the downsampling Down and the output feature E of the second Inception former are input into the first Fusion module, and the first Fusion module outputs the feature F;
[0028] The output feature α″′ of the first self-attention fusion inverted residual attention module is input into the third Inception former, and the output of the third Inception former is the feature G;
[0029] The output feature G of the third Inception former and the output feature F of the first Fusion module are input into the second Fusion module, and the second Fusion module outputs the feature H;
[0030] The output feature of the second self-attention fusion inverted residual attention module is input into the fourth Inception former, and the output of the fourth Inception former is the feature I;
[0031] The output feature I of the fourth Inception former and the output feature H of the second Fusion module are added element-wise to obtain the feature J;
[0032] The output feature β of the multi-scale spatial attention module is input into the first linear layer Linear, and the first linear layer Linear outputs the feature P S ;
[0033] The output feature β of the multi-scale spatial attention module is input into the second linear layer Linear, and the first linear layer Linear outputs the feature
[0034] The feature J is input into the second linear layer Linear, and the first linear layer Linear outputs the feature P t ;
[0035] Based on the feature P S calculate the student network loss function L between the output of the student network and the true label CE ;
[0036] Based on features and P t Calculate the loss function L of knowledge distillation between the output of the student network and the true label CKD ;
[0037] Based on the loss function L of the student network CE and the loss function L of knowledge distillation CKD Calculate the total supervised loss Loss;
[0038] Obtain a trained inverted residual cross-head knowledge distillation model based on the total supervised loss Loss;
[0039] II. Perform ground object recognition on the to-be-tested feature map based on the classifier network in the trained inverted residual cross-head knowledge distillation model.
[0040] The beneficial effects of the present invention are as follows:
[0041] In order to extract the features of the image more effectively, achieve as high an accuracy as possible, and overcome the shortcomings of traditional distillation, in the present invention, an inverted residual cross-head knowledge distillation neural network model is proposed.
[0042] 1. The present invention proposes an inverted residual attention module, which fully considers the advantages of convolution and self-attention and effectively combines convolution and self-attention. First, use inverted residual multi-head self-attention to construct the long-range dependence relationship of the image while increasing the number of channels of the features. Then use convolution to further extract features and play the role of the local inductive bias of convolution.
[0043] 2. A multi-scale spatial attention module is proposed. Use convolutional kernels with different receptive fields to fuse the context information in different ranges in the image. Effectively capture the multi-scale features of the target object and enable the model to more comprehensively understand and process the input image.
[0044] 3. A cross-head knowledge distillation structure is proposed. By fusing the features at all levels to form a teacher branch and passing the intermediate features of the student detection head to the teacher detection head, the resulting cross-head prediction mimics the prediction of the teacher. This makes the head of the student no longer affected by the contradictory supervision signals between the true label and the teacher's prediction, which effectively improves the performance of the student network.
[0045] The present invention proposes a new remote sensing scene classification method, named IRCHKD. This method mainly consists of four key components: inverted residual convolutional block, inverted residual attention module, multi-scale spatial attention module, and cross-head knowledge distillation structure. First, two inverted residual convolutional blocks are introduced to learn the shallow features of the image. Next, an inverted residual attention module is designed. First, the self-attention mechanism is used to learn long-range dependency information and increase the number of feature channels, and then features are further extracted through convolution. Then, a multi-scale spatial attention module is adopted. By using convolutional kernels with different receptive field sizes, the fusion of context information in different ranges of the image is realized, so as to comprehensively capture the multi-scale features of the target object. Finally, a cross-head knowledge distillation structure is introduced. This structure enables the prediction of the student model to mimic the prediction of the teacher by passing the intermediate features of the student detection head to the teacher detection head. This effectively reduces the influence of the student network being affected by the contradictory supervision signals of the true label and the teacher's prediction, thus improving the performance of the student network. It is verified on three widely used remote sensing scene classification datasets. The experimental results show that, compared with some advanced methods, the proposed IRCHKD method shows better performance in remote sensing scene image classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is the overall flowchart of the inverted residual cross-head knowledge distillation model; Figure 2 is the information diagram within the scene selected from the "airport" category, (a) local information; (b) remote context information; Figure 3 is the difference diagram between the hard label and the teacher's prediction;
[0047] Figure 4 is the inverted residual structure diagram adopted by Block1 and Block2 in the model; Figure 5 is the schematic diagram of the dilated convolution. The dilation rates of the dilated convolution from left to right are 1, 2, and 3 respectively, where the orange box is the size of the equivalent convolutional kernel; Figure 6 is the spatial attention structure diagram; Figure 7 are the pictures randomly selected from the four datasets;
[0048] Figure 8Confusion matrix graph obtained at 80% training ratio on the UC-Merced dataset. Predicted Label: Predicted label, True Label: True label, agricultural: Farmland, airplane: Airplane, baseballdiamond: Baseball diamond, beach: Beach, buildings: Buildings, chaparral: Chaparral, denseresidential: Dense residential area, forest: Forest, freeway: Freeway, golfcourse: Golf course, harbor: Harbor, intersection: Intersection, mediumresidential: Medium residential area, mobilehomepark: Mobile home park, overpass: Overpass, parkinglot: Parking lot, river: River, sparseresidential: Sparse residential area, storagetanks: Storage tanks, tenniscourt: Tennis court;
[0049] Figure 9 Confusion matrix graph obtained at 50% training ratio on AID. Predicted Label: Predicted label, True Label: True label, Airport: Airport, BareLand: Bare land, BaseballField: Baseball field, Beach: Beach, Bridge: Bridge, Center: Center, Church: Church, Commercial: Commercial area, DenseResidential: Dense residential area, Desert: Desert, Farmland: Farmland, Forest: Forest, Industrial: Industrial area, Meadow: Meadow, MediumResidential: Medium residential area, Mountain: Mountain, Park: Park, Parking: Parking lot, Playground: Playground, Pond: Pond, Port: Port, RailwayStation: Railway station, Resort: Resort area, River: River, School: School, SparseResidential: Sparse residential area, Square: Square, Stadium: Stadium, StorageTanks: Storage tanks, Viaduct: Viaduct;
[0050] Figure 10The confusion matrix obtained under the 20% training ratio of the NWPU dataset. Predicted Label: Predicted label, True Label: True label, airplane: Airplane, airport: Airport, baseball_diamond: Baseball diamond, basketball_court: Basketball court, beach: Beach, bridge: Bridge, chaparral: Chaparral, church: Church, circular_farmland: Circular farmland, cloud: Cloud, commercial_area: Commercial area, dense_residential: Dense residential area, desert: Desert, forest: Forest, freeway: Freeway, golf_course: Golf course, ground_track_field: Ground track field, harbor: Harbor, industrial_area: Industrial area, intersection: Intersection, island: Island, lake: Lake, meadow: Meadow, medium_residential: Medium residential area, mobile_home_park: Mobile home park, mountain: Mountain, overpass: Overpass, palace: Palace, parking_lot: Parking lot, railway: Railway, railway_station: Railway station, rectangular_farmland: Rectangular farmland, river: River, sea_ice: Sea ice, ship: Ship, snowberg: Snowberg, sparse_residential: Sparse residential area, stadium: Stadium, storage_tank: Storage tank, tennis_court: Tennis court, terrace: Terrace, thermal_power_station: Thermal power station, wetland: Wetland. Detailed implementation manner
[0051] Detailed implementation manner one: Combining Figure 3 To illustrate this implementation manner, the specific process of the feature recognition method based on the inverted residual cross-head knowledge distillation neural network is as follows:
[0052] I. Construct an inverted residual cross-head knowledge distillation model; the specific process is as follows:
[0053] The inverted residual cross-head knowledge distillation model includes a classifier network and a cross-head knowledge distillation structure CHKD;
[0054] The classifier network includes a convolutional block, Block1 for extracting shallow features, Block2, an inverted residual attention module IRAM for the first self-attention fusion, an inverted residual attention module IRAM for the second self-attention fusion, a multi-scale spatial attention module MSSA, and a first linear layer Linear;
[0055] The cross-head knowledge distillation structure includes feature aggregation and cross-head distillation;
[0056] The feature aggregation includes a first Inception former, a second Inception former, a third Inception former, a fourth Inception former, downsampling Down, a first Fusion module, and a second Fusion module;
[0057] The cross-head distillation part includes a second linear layer Linear for generating cross-head predictions;
[0058] The classifier network is also called the student network;
[0059] Feature aggregation is also called the teacher network (TeacherNetwork, TN);
[0060] The specific working process of the inverted residual cross-head knowledge distillation model is as follows:
[0061] The feature map is sequentially input into the convolutional block, Block1 in the classifier network, and Block1 outputs feature α;
[0062] Feature α is input into Block2, and Block2 outputs feature α';
[0063] Feature α' is input into the inverted residual attention module for the first self-attention fusion, and the inverted residual attention module for the first self-attention fusion outputs feature α''';
[0064] The output feature α''' of the inverted residual attention module for the first self-attention fusion is input into the inverted residual attention module for the second self-attention fusion, and the inverted residual attention module for the second self-attention fusion outputs feature
[0065] Feature is input into the multi-scale spatial attention module, and the multi-scale spatial attention module outputs feature β;
[0066] Feature α is input into the first Inception former, the output feature of the first Inception former is input into downsampling Down, and downsampling Down outputs feature D;
[0067] The output feature α′ of Block2 is input into the second Inception former, and the second Inception former outputs feature E;
[0068] The output feature D of the downsampling Down and the output feature E of the second Inception former are input into the first Fusion module, and the first Fusion module outputs feature F;
[0069] The output feature α″′ of the first self-attention fusion inverted residual attention module is input into the third Inception former, and the third Inception former outputs feature G;
[0070] The output feature G of the third Inception former and the output feature F of the first Fusion module are input into the second Fusion module, and the second Fusion module outputs feature H;
[0071] The output feature of the second self-attention fusion inverted residual attention module is input into the fourth Inception former, and the fourth Inception former outputs feature I;
[0072] The output feature I of the fourth Inception former and the output feature H of the second Fusion module are added element-wise to obtain feature J;
[0073] The output feature β of the multi-scale spatial attention module is input into the first linear layer Linear, and the first linear layer Linear outputs feature P S ;
[0074] The output feature β of the multi-scale spatial attention module is input into the second linear layer Linear, and the first linear layer Linear outputs the feature
[0075] Feature J is input into the second linear layer Linear, and the first linear layer Linear outputs feature P t ;
[0076] Based on feature P S Calculate the student network loss function L between the student network output and the true label CE ;
[0077] Based on the feature and P t Calculate the loss function L of knowledge distillation between the student network output and the true label CKD ;
[0078] Based on the student network loss function L CE; and the loss function L of knowledge distillation CKD Calculate the total supervision loss Loss;
[0079] Obtain a trained inverted residual cross-head knowledge distillation model based on the total supervision loss Loss;
[0080] II. Perform ground object recognition on the to-be-tested feature map based on the classifier network in the trained inverted residual cross-head knowledge distillation model.
[0081] Specific Embodiment 2: Different from Specific Embodiment 1, the convolutional block sequentially includes a first 3×3 convolutional layer, BN, and SiLU activation function;
[0082] The Block1 includes a first 1×1 convolutional layer, a first 3×3 depthwise separable convolutional layer, and a second 1×1 convolutional layer;
[0083] The Block2 includes a third 1×1 convolutional layer, a second 3×3 depthwise separable convolutional layer, and a fourth 1×1 convolutional layer;
[0084] The working process of the Block1 is as follows:
[0085] The feature x is sequentially input into the first 1×1 convolutional layer, the first 3×3 depthwise separable convolutional layer, and the second 1×1 convolutional layer. The output feature of the second 1×1 convolutional layer is added element-wise to the feature x to obtain the feature α, and the feature α is used as the output feature of the Block1;
[0086] The working process of the Block2 is as follows:
[0087] The output feature α of the Block1 is sequentially input into the third 1×1 convolutional layer, the second 3×3 depthwise separable convolutional layer, and the fourth 1×1 convolutional layer. The output feature of the fourth 1×1 convolutional layer is added element-wise to the feature α to obtain the feature α′, and the feature α′ is used as the output feature of the Block2.
[0088] Other steps and parameters are the same as those in Specific Embodiment 1.
[0089] Specific Embodiment 3: Different from Specific Embodiment 1 or 2, the first self-attention fusion inverted residual attention module IRAM includes self-attention, relative position bias (RPB), Softmax, a third 3×3 depthwise separable convolutional layer, and a fifth 1×1 convolutional layer;
[0090] The working process of the first self-attention fusion inverted residual attention module IRAM is as follows:
[0091] 1) The feature maps pass through the third linear layer (Linear), the fourth linear layer (Linear), and the fifth linear layer (Linear) respectively. The third linear layer (Linear) outputs the query matrix Q, the fourth linear layer (Linear) outputs the key matrix K, and the fifth linear layer (Linear) outputs the value matrix V.
[0092] 2) Based on the query matrix Q, the key matrix K, the value matrix V, and the relative position bias B, self-attention Attention(Q, K, V) is obtained. The calculation process is
[0093]
[0094] where Q, K, and V are the query matrix, the key matrix, and the value matrix respectively. d represents the dimension of Q / K, and B represents the relative position bias. M 2 represents the number of pixel blocks in the window. represents a real number; the superscript T represents taking the transpose.
[0095] 3) The self-attention Attention(Q, K, V) is input into the third 3×3 depthwise separable convolutional layer. The output feature of the third 3×3 depthwise separable convolutional layer is element-wise added to the self-attention Attention(Q, K, V) to obtain the feature α″.
[0096] The feature α″ is input into the fifth 1×1 convolutional layer. The fifth 1×1 convolutional layer outputs the feature α″′, and the feature α″′ is used as the output of the first self-attention fusion inverted residual attention module (IRAM).
[0097] Other steps and parameters are the same as those in the first or second specific implementation.
[0098] Specific implementation method four: The difference between this implementation method and one of the first to third specific implementation methods is that the second self-attention fusion inverted residual attention module IRAM includes self-attention, relative position bias (relativepositionbias, RPB), Softmax, the fourth 3×3 depthwise separable convolutional layer, and the sixth 1×1 convolutional layer.
[0099] The working process of the second self-attention fusion inverted residual attention module IRAM is as follows:
[0100] 1) The output features of the first self-attention fusion inverted residual attention module IRAM pass through the sixth linear layer (Linear), the seventh linear layer (Linear), and the eighth linear layer (Linear) respectively. The sixth linear layer (Linear) outputs the query matrix Q, the seventh linear layer (Linear) outputs the key matrix K, and the eighth linear layer (Linear) outputs the value matrix V.
[0101] 2), obtain the self-attention Attention(Q, K, V) based on the query matrix Q, the key matrix K, the value matrix V, and the relative position bias B; the calculation process is
[0102]
[0103] where Q, K, and V are the query matrix, the key matrix, and the value matrix respectively, d represents the dimension of Q / K, and B represents the relative position bias, M 2 represents the number of pixel blocks in the window, represents a real number;
[0104] 3), the self-attention Attention(Q, K, V) is input into the fourth 3×3 depthwise separable convolutional layer, and the output features of the fourth 3×3 depthwise separable convolutional layer are element-wise added to the self-attention Attention(Q, K, V) to obtain the features
[0105] features are input into the sixth 1×1 convolutional layer, and the output features of the sixth 1×1 convolutional layer features are used as the output of the inverted residual attention module IRAM for the second self-attention fusion;
[0106] Other steps and parameters are the same as those in any one of the specific embodiments one to three.
[0107] Specific embodiment five: The difference between this embodiment and any one of the specific embodiments one to four is that the multi-scale spatial attention module MSSA includes four branches and a 1×1 convolutional layer;
[0108] The first branch sequentially includes a 1×1 convolutional layer and spatial attention;
[0109] The second branch sequentially includes a 1×1 convolutional layer, a 3×3 convolutional layer, a batch normalization layer BN, a SiLU activation function, and spatial attention;
[0110] The third branch sequentially includes a 1×1 convolutional layer, a 3×3 convolutional layer, a batch normalization layer BN, a SiLU activation function, a 3×3 convolutional layer with a dilation rate of 2, a batch normalization layer BN, a SiLU activation function, and spatial attention;
[0111] The fourth branch sequentially includes a 1×1 convolutional layer, a 3×3 convolutional layer, a batch normalization layer BN, a SiLU activation function, a 3×3 convolutional layer with a dilation rate of 2, a batch normalization layer BN, a SiLU activation function, a 3×3 convolutional layer with a dilation rate of 3, a batch normalization layer, a SiLU activation function, and spatial attention;
[0112] The working process of the multi-scale spatial attention module MSSA is as follows:
[0113] The output features of the inverted residual attention module (IRAM) with second self-attention fusion are respectively input into four branches. The features obtained from the four branches are concatenated, and the concatenated features are input into a 1×1 convolutional layer. The output features of the 1×1 convolutional layer are used as the output features of the multi-scale spatial attention module MSSA.
[0114] Other steps and parameters are the same as those in any one of the first to fourth specific embodiments.
[0115] Specific embodiment six: The difference between this embodiment and any one of the first to fifth specific embodiments is that the spatial attention sequentially includes max pooling, average pooling, a 3×3 convolutional layer, and a Softmax layer;
[0116] The working process of the spatial attention is as follows:
[0117] The feature maps are respectively input into max pooling and average pooling. The feature maps after max pooling and average pooling are concatenated by channels. The output features of the concatenated feature maps after passing through a 3×3 convolutional layer pass through a Softmax layer to obtain an attention map;
[0118] It is expressed as
[0119] Ms(F) = Softmax(f 3 × 3 ([AvgPool(F); MaxPool(F)]))
[0120] where F represents the feature map, and f 3×3 represents the 3×3 convolutional layer, and M s (F) represents the attention map output after the feature map F passes through the spatial attention;
[0121] AvgPool(F) represents average pooling of the feature map F;
[0122] MaxPool(F) represents max pooling of the feature map F.
[0123] Other steps and parameters are the same as those in any one of the first to fifth specific embodiments.
[0124] Embodiment 7: The difference between this embodiment and any one of Embodiments 1 to 6 is that each Inceptionformer among the first Inceptionformer, the second Inceptionformer, the third Inceptionformer, and the fourth Inceptionformer includes a first branch, a second branch, a ninth linear layer, GELU, and a tenth linear layer;
[0125] The first branch sequentially includes a 1×1 convolutional layer, batch normalization BN, and SiLU activation function;
[0126] The second branch sequentially includes a 3×3 convolutional layer, batch normalization BN, and SiLU activation function;
[0127] The working process of each Inceptionformer among the first Inceptionformer, the second Inceptionformer, the third Inceptionformer, and the fourth Inceptionformer is as follows:
[0128] The feature maps are respectively input into the first branch and the second branch, the output features of the first branch and the second branch are concatenated, and the concatenated features are sequentially input into the ninth linear layer, GELU, and tenth linear layer. The output features of the tenth linear layer are added element-wise to the concatenated features to obtain the output features of the Inceptionformer;
[0129] Other steps and parameters are the same as any one of Embodiments 1 to 6.
[0130] Embodiment 8: The difference between this embodiment and any one of Embodiments 1 to 7 is that the downsampling Down includes a first branch, a second branch, and a 1×1 convolutional layer;
[0131] The first branch sequentially includes two-dimensional average pooling, 1×1 convolution, BN, and SiLU activation function;
[0132] The second branch sequentially includes 3×3 convolution, BN, and SiLU activation function;
[0133] The working process of the downsampling Down is as follows:
[0134] The feature maps are respectively input into the first branch and the second branch, the output feature maps of the first branch and the second branch are connected, and a 1×1 convolutional layer is used to perform channel mixing on the connected feature maps to obtain the output features of the downsampling Down; thus, one downsampling of the feature maps is completed.
[0135] Other steps and parameters are the same as any one of Embodiments 1 to 7.
[0136] Embodiment 9: The difference between this embodiment and any one of Embodiments 1 to 8 is that each Fusion module in the first Fusion module and the second Fusion module includes a first branch, a second branch, a 1×1 convolutional layer, BN, and a SiLU activation function;
[0137] The first branch sequentially includes average pooling, a 1×1 convolutional layer, BN, and a SiLU activation function;
[0138] The second branch sequentially includes a 3×3 convolutional layer, BN, and a SiLU activation function;
[0139] The working process of each Fusion module in the first Fusion module and the second Fusion module is as follows:
[0140] The two input feature maps are concatenated along the channel dimension, the channels of the concatenated feature map are shuffled, the shuffled feature maps are respectively input into the first branch and the second branch, the feature maps obtained by the two branches are concatenated in the channel dimension, and the concatenated feature map is sequentially input into a 1×1 convolutional layer, BN, and a SiLU activation function, and the SiLU activation function outputs a feature map, thus completing the second downsampling of the feature map.
[0141] Other steps and parameters are the same as any one of Embodiments 1 to 8.
[0142] Embodiment 10: The difference between this embodiment and any one of Embodiments 1 to 9 is that the student network loss function L between the output of the student network and the true label is calculated based on the feature P S ; CE ;
[0143] Based on the feature and P t the loss function L of knowledge distillation between the output of the student network and the true label is calculated; CKD ;
[0144] Based on the student network loss function L CE and the loss function L of knowledge distillation CKD the total supervision loss Loss is calculated;
[0145] The expression is:
[0146] The loss function L of knowledge distillation CKD is expressed as:
[0147]
[0148] where T represents the temperature hyperparameter, and D KL represents P t and The KL distance between two distributions.
[0149] Generally speaking, the supervision loss of the model consists of two parts. The first part is the cross-head distillation loss between the teacher output distribution and the student cross-head output distribution. The second part is the cross-entropy loss between the student network output and the true label. In addition to the distillation loss, the student network also uses the cross-entropy loss to learn the true label.
[0150] The loss function L of the student network CE Uses the cross-entropy loss to learn the true label, and the cross-entropy loss is expressed as:
[0151]
[0152] where N represents the training samples, x represents the input tensor, θ s represents the parameters in the student network, y i represents the true label, represents the output of the student network after passing through the fully connected layer;
[0153] The total supervision loss Loss is expressed as:
[0154]
[0155] Train the inverted residual cross-head knowledge distillation model based on the total supervision loss Loss to obtain a trained inverted residual cross-head knowledge distillation model.
[0156] As Figure 3 shown, we observe that directly imitating the teacher's prediction will face the problem of target conflict. To alleviate this problem, we propose a cross-head knowledge distillation (CHKD) structure. Like ordinary knowledge distillation, CHKD performs prediction simulation. The difference is that CHKD transfers the intermediate features of the student network to the detection head of the teacher network and generates cross-predictions for distillation. Denote the outputs of the teacher and student networks as P t and P s . In addition to the original inputs of the teacher and the student, CHKD also inputs the feature map output by the student network into the detection head of the teacher network to generate cross-head predictions We use the knowledge distillation loss between the cross-head predictions and the teacher prediction P t as the objective of CHKD.
[0157] In recent years, with the rapid development of deep learning technology, remote sensing scene image classification has made great progress. Compared with natural images, remote sensing scene images have more complex objects, with high inter-class similarity and large intra-class differences, which makes it difficult to effectively extract image features. Convolutional neural networks are widely used in remote sensing scene image classification tasks, in which convolution pays more attention to the high-frequency information of the image. Unlike convolution, Transformer can model long-range feature dependencies and mine contextual information in remote sensing scene images. In addition, Transformer is considered to be a low-pass filter that can complement convolutional neural networks. Considering the complementary characteristics of convolution and Transformer, the present invention combines the two for feature extraction. At the same time, traditional one-hot labels cannot accurately describe the image and cannot provide sufficient information for the feature learning of the supervised network. In order to enable the model to obtain sufficient supervised information, we use knowledge distillation to provide additional supervised information. In this study, an inverted residual cross-head knowledge distillation neural network (IRCHKD) is proposed to generate more robust features. First, two inverted residual convolution blocks are used to extract shallow features. Then, the inverted residual attention module (IRAM) is used to extract local and global information. Then, multi-scale spatial attention (MSSA) is used to extract features of different scales in the image. Finally, cross-head knowledge distillation (CHKD) is used to generate additional supervision signals. The experimental results are verified on three widely used remote sensing scene image datasets, and the effectiveness of the proposed method for remote sensing scene image classification is demonstrated.
[0158] The overall framework diagram of the method proposed in the present invention is as follows Figure 1 As shown, the classifier network in the pink area and the cross-head knowledge distillation structure in the gray area. The classifier network mainly includes Block1 and Block2 for extracting shallow features, the inverted residual attention module (IRAM) and the multi-scale spatial attention module (MSSA) that fuse convolution and self-attention. The cross-head knowledge distillation (CHKD) structure mainly includes the Inception former, Down and Fusion modules. The classifier network is used for remote sensing scene image classification, and the cross-head knowledge distillation structure is used to provide supervision information for the classifier network.
[0159] 2.1 CNN-based RS scene classification method
[0160] Convolutional neural networks (CNNs) are outstanding in feature extraction and are widely used in remote sensing scene image classification. For example, Shi et al.
[21] proposed a branched feature fusion convolutional neural network to fuse the feature information extracted from two branches and simultaneously combined depthwise separable convolution to reduce the model complexity. Zhao et al.
[22] proposed a structure that combines global texture features, local structural features, and spectral features. Li
[23] et al. combined transfer learning and convolutional neural networks to generate robust visual features, improving the classification accuracy in the case of limited data labeling. Singh et al.
[24] proposed a heterogeneous convolution method that introduced two types of convolutional kernels, a k×k convolutional kernel for some channels and a 1×1 convolutional kernel for the remaining channels. The ratio between channels was adjusted by the hyperparameter p. The purpose of this method was to improve the performance of image processing by optimizing the convolutional structure. In the dynamic convolution proposed by Chen et al.
[25] , an attention mechanism was used to dynamically aggregate multiple small-sized convolutional kernels to enhance feature representation and computational efficiency in a non-linear manner. This method could improve feature representation and computational efficiency without increasing the network depth and width. The self-calibrated convolution proposed by Liu et al.
[26] was a method to adaptively establish the long-distance spatial and channel dependencies at each spatial position. More discriminative features were generated through self-calibration operations. From the frequency perspective, Chen et al.
[27] proposed octave convolution. This method divided the input features into high-frequency and low-frequency features and controlled the ratio between the two by the hyperparameter α to reduce the parameters and computational amount and effectively improve the feature expression ability. In dealing with redundant information, Han et al.
[28] improved traditional convolution and proposed ghost convolution. Ghost convolution extracted feature information through traditional convolution and used linear transformation to generate redundant information, thus reducing the model computational complexity. The convolutional parameters in traditional convolution were shared by all samples, while Yang et al.
[29] proposed conditional parameterized convolution to obtain customized convolutional kernels for each input sample in each batch. This method could improve the model capacity and maintain an efficient running speed. Cao et al.
[30] combined depthwise separable convolution with traditional convolution and proposed depthwise separable over-parameterized convolution. This new method provided new ideas for research and applications in the field of image processing.
[0161] 2.2 Attention-based RS Scene Classification Methods
[0162] In remote sensing scene classification, CNN-based methods have achieved remarkable results. However, in the face of complex remote sensing scene images, how to effectively extract key information and capture context dependencies remains an urgent problem to be solved. To address these challenges, visual attention techniques have been introduced into the field of remote sensing image processing. Shi et al.
[31] proposed a multi-branch feature fusion method based on the attention mechanism, combining channel attention and spatial attention, and fully extracting the deep features of images through multi-convolution collaboration. Tang et al.
[32] proposed an attention consistency network (ACNet) for remote sensing scene classification. This network combines spatial attention and channel attention, aiming to extract basic features from remote sensing scenes. In addition, an attention consistency model was designed to unify significant regions in order to extract discriminative features more accurately. Chen et al.
[33] designed a multi-branch local attention network and proposed a convolutional local attention module for effectively obtaining channel and spatial attention weights. This method can automatically emphasize or suppress important or redundant information in remote sensing scenes.
[0163] Dosovitskiy et al.
[18] proposed the Vision Transformer (ViT), which converts natural images into a sequence of image patches and uses self-attention to build long-range dependencies. Subsequently, the Transformer has been widely applied in the field of natural image processing
[34] -
[36] . These positive results demonstrate the potential of the Transformer in image processing. Based on ViT, a multi-scale vision transformer was proposed
[37] , which is a two-branch network for extracting multi-scale features for image classification. In addition, an efficient feature fusion module was designed to integrate the contributions of the two branches to learn richer and more robust features. To reduce the time cost caused by the self-attention mechanism, Liu et al.
[38] proposed a Swin Transformer for various vision tasks. By calculating self-attention within a local window, the time cost can be significantly reduced. In the remote sensing community, the number of Transformer-based models is small. SpectralFormer
[39] is a new backbone network for hyperspectral image (HSI) classification tasks. SpectralFormer applies the original Transformer to extract context information from HSI and proposes a grouped spectral embedding module and a cross-layer adaptive fusion module to learn local and cross-layer information. Bazi et al.
[40] introduced the Transformer to handle remote sensing scene classification. It uses the original Transformer model to capture the context information within the remote sensing scene and achieved excellent classification results. Ma et al.
[41] proposed a homogeneous transformer learning framework (HHTL). By simultaneously mining the homogeneous and heterogeneous information in the remote sensing scene, a comprehensive and discriminative feature representation for remote sensing scene classification was constructed. Tang et al.
[42] proposed a multi-scale Transformer to mine the context information of the content of the RS scene and model the internal relationship at a relatively small time cost. Chen et al.
[43] proposed a model combining convolution and Transformer to classify smoke scenes, which combines local features and global features and achieves a high classification accuracy.
[0164] 2.3 Knowledge Distillation
[0165] The idea of Knowledge Distillation is to guide the training of the student model through the predictions of the teacher model. The student model will be trained according to the prediction results of the teacher model to learn how to generate prediction results similar to those of the teacher model. During the training process, the student model will continuously optimize its own parameters.
[0166] To effectively transfer knowledge, an appropriate loss function is needed to measure the difference between the output of the student model and the output of the teacher model. KL divergence
[44] is usually used to measure the difference between two distributions. Its calculation formula can be expressed as follows:
[0167]
[0168] where i represents the input tensor, P(i) represents the distribution predicted by the teacher network, Q(i) represents the distribution predicted by the student network, and * represents element-wise multiplication;
[0169] Generally speaking, the teacher model should be complex and accurate enough. The commonly used teacher model is a pre-trained deep neural network model. The student model is usually more lightweight to be deployed on resource-constrained devices. With the in-depth research, researchers have designed many improved knowledge distillation methods. For example, Adriana Romero et al.
[45] aligned the intermediate layer outputs of the teacher model and the student model, introducing the concept of intermediate layer alignment. Wonpyo Park et al.
[46] used relational modeling to optimize the effect of knowledge distillation, making knowledge transfer more accurate and effective. Byeongho Heo et al.
[47] designed a new distillation function, paying special attention to the features before the ReLU function and keeping the negative values of these features during the distillation process. Sungsoo Ahn et al.
[48] proposed a variational information distillation framework that can transfer the knowledge learned by the convolutional network to the multi-layer perceptron, maximizing the mutual information between the two neural networks through the maximum variational lower bound.
[0170] Since it is very difficult to select a suitable teacher network and train a large teacher network, some research has turned to self-distillation algorithms. Zhang et al.
[49] proposed a self-distillation framework using ResNet as the backbone network, downsampling the deep features through a bottleneck structure, and finally outputting soft labels through a fully connected layer. The distribution used to supervise the shallow network enables the shallow layer of the network to learn the deep features. Ji et al.
[50] proposed a self-distillation framework for feature refinement, enhancing the feature map through lateral convolution for the purpose of self-knowledge distillation. Hu et al.
[51] proposed a hierarchical self-distillation feature learning framework, supervising the distribution generated by the shallow network with the distribution generated by the deep network. They also introduced a gradient separation and fusion module to ensure that the gradient of the final classification output does not backpropagate to the backbone network. These studies provide new ideas and methods for the field of knowledge distillation, helping to improve the learning efficiency and accuracy of the model.
[0171] Knowledge distillation methods have been applied to the field of remote sensing image analysis. The MSH-Net proposed by Wei et al.
[52] reconstructs complete modality-shared features from incomplete inference modalities to assist the model inference of missing modalities. The Joint Adaptation Distillation (JAD) method among them guides the model to learn modality-shared knowledge from multi-modal models by matching the joint probability distribution between representations and ground truths. Hu et al.
[53] proposed the variational self-distillation method, which hierarchically distills deep features and shallow features through Variational Knowledge Transfer (VKT), and uses the prediction vector of class-entangled information as supplementary class information. Li et al.
[54] proposed dual knowledge distillation, designed dual attention and spatial structures, as well as two effective loss functions, which can effectively transfer the knowledge learned by the teacher network to the student network. Liu et al.
[55] proposed a cross-model knowledge distillation method, using the model pre-trained on RGB images as the teacher model to guide multi-spectral scene classification.
[0172] The overall framework diagram of the proposed method is as Figure 1 shown, including a classifier network in the pink area and a cross-head knowledge distillation structure in the gray area. The classifier network mainly includes Block1 and Block2 for extracting shallow features, an Inverted Residual Attention Module (IRAM) for convolutional and self-attention fusion, and a Multi-Scale Spatial Attention Module (MSSA). The cross-head knowledge distillation (CHKD) structure mainly includes Inception former, Down, and Fusion modules. The classifier network is used for remote sensing scene image classification, and the cross-head knowledge distillation structure is used to provide supervision information for the classifier network.
[0173] 3.1 Inverted Residual Attention Module (IRAM)
[0174] Depthwise separable convolution consists of two layers. The first layer is called depth convolution, which performs filtering by applying a single filter to each input channel. The second layer is a 1×1 convolution, also known as point convolution, whose role is to linearly combine features between channels. The input of standard convolution is h i ×w i ×d i , and the convolution kernel used is to generate an output tensor of h i ×w i ×d j . The computational cost of standard convolution is h i ·w i ·d i ·d j ·k·k. Depthwise separable convolution can directly replace standard convolution, producing an effect comparable to that of standard convolution, while the computational cost is only h i ·w i ·d i (k2 +d j ) effectively reduces the computational complexity.
[0175] As Figure 4 shown, we adopt the inverted residual structure in the first two blocks of the model. First, use a 1×1 point convolution to increase the number of channels of the features, then use depthwise separable convolution to further extract features, and then add the residual structure. This process can be expressed as
[0176] Conv = Residual(Conv dsc (Conv 1×1 (x)))
[0177] where Conv 1×1 represents a 1×1 convolution, Conv dsc represents depthwise separable convolution, and Residual represents the residual connection.
[0178] To combine the advantages of self-attention in long-range modeling, we improved the first 1×1 convolution in Figure 4 . Specifically, as in the IRAM module in Figure 1 , this process can be expressed as
[0179] F(·)Conv(IRSA(·))
[0180] where IRSA represents inverted residual multi-head self-attention, and Conv represents a combination of 3×3 depthwise separable convolution and 1×1 convolution. This design method combines the local feature modeling efficiency of CNN and the modeling ability of Transformer to learn long-range interactions. First, map the feature map to obtain three tensors Q, K, and V. Among them, the number of channels of V is doubled to increase the dimension of the feature map and provide a richer feature representation. Use Q and K to obtain the self-attention map, and then combine the relative position bias (RPB) to obtain the global attention map. After multiplying the attention map by V, the attention interaction is realized. Add depth convolution and residual structure after self-attention. The self-attention mechanism can capture global information, and the depth convolution layer can focus on local feature extraction to better capture details. Combining the two can make the network more comprehensive in understanding images while not losing the fineness of local information. Finally, use a 1×1 convolution to perform information interaction in the channel dimension and combine the features between different channels to generate new features.
[0181] For effective modeling, the remote sensing image is divided into some non-overlapping windows, and the self-attention calculation is completed within the windows, which can reduce the computational complexity. Suppose each window contains M×M pixel blocks. On an h×w image, the computational complexities of global self-attention and window-based self-attention are respectively:
[0182] Ω(SA) = 4hwC 2 + 2(hw) 2 CΩ(WSA) = 4hwC 2 + 2M 2 hwC
[0183] Among them, Ω(SA) represents the computational complexity of global self-attention, Ω(WSA) represents the computational complexity of window-based self-attention, M represents the size of the divided window, C represents the number of channels of the feature map, and hw represents the size of the input feature map. The former is quadratic with respect to the feature map size, and the latter is linear when the value of M (set to 7) is fixed. It can be seen that the latter can significantly reduce the computational complexity.
[0184] 3.2 Multi-Scale Spatial Attention (MSSA)
[0185] In recent years, CNN has been widely used in the field of deep learning due to its powerful feature extraction ability. However, due to the limitations of traditional convolution, many different convolution methods have been derived. Dilated convolution is one of these methods. Dilated convolution can increase the receptive field (Receptive Field, RF) of the convolution kernel while keeping the number of parameters unchanged, enabling each convolution to contain a larger range of information while ensuring that the size of the output feature map remains unchanged. The process of dilated convolution is as Figure 5 shown, and the RF increases with the increase of the dilation rate. In the case of a convolution kernel of size r×r and a dilation rate of d, the equivalent convolution kernel size is r′×r′.
[0186] r′ = r + (r - 1)(d - 1)
[0187] When the dilation rate is 1, the receptive field of dilated convolution is the same as that of ordinary convolution. When the dilation rate is equal to 2, the RF of a 3×3 convolution kernel of dilated convolution is the same as that of a 5×5 convolution kernel of standard convolution. When the dilation rate is equal to 3, the RF of a 3×3 convolution kernel of dilated convolution is the same as that of a 7×7 convolution kernel of standard convolution. The calculation process of RF is
[0188] R i+1 = R i + (r′ - 1)S i
[0189] where i represents the total number of layers of dilated convolution used, r′ represents the size of the equivalent convolution kernel of the (i + 1)-th layer, R i represents the RF of the i-th layer, R i+1 represents the RF of the (i + 1)-th layer, stride represents the stride adopted by each layer of dilated convolution, and S i represents the product of all strides of the first i layers.
[0190] The multi-scale spatial attention module includes four branches: The first branch first uses a 1×1 convolution and then performs spatial attention; the second branch first uses a 1×1 convolution, then uses a 3×3 convolution and a batch normalization layer, and then performs spatial attention; the third branch first uses a 1×1 convolution, then uses a 3×3 convolution and a batch normalization layer, then uses a 3×3 convolution with a dilation rate of 2 and a batch normalization layer, and then performs spatial attention; the fourth branch first uses a 1×1 convolution, then uses a 3×3 convolution and a batch normalization layer, then uses a 3×3 convolution with a dilation rate of 2 and a batch normalization layer, and then uses a 3×3 convolution with a dilation rate of 3, and finally performs spatial attention. Finally, the features obtained from the four branches are concatenated. Through multiple parallel dilated convolution branches, features at different scales can be captured, enabling the model to understand the image content more comprehensively, including objects or structures at different scales. By using convolution kernels with different receptive field sizes, the context information in different ranges of the image can be effectively fused. Feature extraction can be targeted at the characteristics of complex backgrounds and multi-scale ground objects in remote sensing scene images.
[0191] The spatial attention mechanism can further enhance the model's representation ability. The spatial attention mechanism focuses on important regions in the image and emphasizes or suppresses certain features by assigning different weights to different regions. It helps improve the model's perception ability of the image, enabling it to better understand the image content. Especially for objects or targets of different sizes and recognition tasks in complex backgrounds, the spatial attention mechanism weights the features, allowing the model to ignore unimportant information when processing important regions. The spatial attention we use is as Figure 6 shown. Max-pooling and average-pooling are performed on the feature map along the channel dimension, and the processed feature maps are concatenated by channel. Then, a 3×3 convolution is used to mix the feature maps, and after passing through the Softmax layer, an attention map is obtained.
[0192] 3.3 Cross-Head Knowledge Distillation (CHKD)
[0193] The cross-head knowledge distillation we proposed is Figure 1 the grey block part in Figure 1As can be seen, the cross-head knowledge distillation structure mainly consists of two parts: the feature aggregation part and the cross-head distillation part. The feature aggregation part mainly includes three components: Inception former, Down, and Fusion. Additionally, the feature aggregation part is also called the Teacher Network (TN). Features at different layers contain different spatial structure information in the remote sensing scene. Shallow neural networks are more inclined to learn low-level features of data, such as edges and textures. Deeper neural networks, on the other hand, are capable of learning more advanced features and are better at capturing global and abstract associations. For these features at different layers, we adopt a progressive aggregation strategy to gradually fuse the low-level feature maps with the high-level feature maps.
[0194] Four Inception formers are used to further extract features from the feature maps of four different stages of the student network. The specific structure of the Inception former is shown in Figure 1 . In the Inception former structure, the feature map is input into two branches. The first branch contains a 1×1 convolution, batch normalization (Batch Normalization, BN), and the SiLU activation function. The second branch contains a 3×3 convolution, BN, and the SiLU activation function. Then, the features obtained from the two branches are concatenated. The concatenated features are then input into a multi-layer perceptron (MLP) to perform information interaction in the channel dimension of the feature map. We use two parallel convolutions to mimic the self-attention structure in Transformer and retain the MLP in Transformer. The MLP contains two linear layers and a GELU activation function. Using this structure can not only extract spatial features but also fuse the feature maps in the channel dimension. The choice of convolution to replace self-attention mainly takes into account the computational efficiency of the model.
[0195] First, the features obtained from the first Inception former are input into the downsampling structure. The downsampling structure is shown in Figure 1 in the 'Down' part. The feature map is input into two parallel branches. The first branch contains a two-dimensional average pooling, a 1×1 convolution, BN, and the SiLU activation function. The second branch contains a 3×3 convolution, BN, and the SiLU activation function. Then, the feature maps obtained from the two branches are concatenated, and a 1×1 convolution layer is used to perform channel mixing on the concatenated feature map. This completes one downsampling of the feature map. Next, the feature map obtained from the Down structure and the feature map obtained from the second Inception former are simultaneously input into the first Fusion structure. The Fusion structure is shown in Figure 1。In the Fusion structure, first, the two input feature maps are concatenated along the channel dimension. Then, channel shuffle is used to shuffle the channels, and the shuffled feature maps are input into two branches. Next, the feature maps obtained from the two branches are concatenated in the channel dimension. The concatenated feature maps are sequentially input into a 1×1 convolutional layer, BN, and the SiLU activation function, thus completing the second downsampling of the feature maps. Then, the feature map obtained from the first Fusion and the feature map obtained from the third Inception former are simultaneously input into the second Fusion structure. Similar to the first Fusion structure, after passing through the second Fusion structure, the feature maps are downsampled again. Finally, the feature map output by the second Fusion structure and the feature map output by the fourth Inception former structure are added together to obtain the feature map after multi-level feature aggregation. The obtained feature map is flattened and then input into a linear layer to obtain the output of the teacher network.
[0196] The following embodiments are used to verify the beneficial effects of the present invention:
[0197] Embodiment 1:
[0198] To evaluate the effectiveness of our proposed IRCHKD, we conducted various experiments on three public and challenging datasets and compared them with some advanced methods proposed in recent years. The three datasets are the UC-Merced dataset
[56] , AID
[57] , and NWPU-RESISC45 dataset
[58] . The experimental results demonstrate the effectiveness of our proposed method.
[0199] Datasets: In this part, the three datasets used in the experiments are introduced. Some sample examples selected from these datasets are shown as Figure 7 shown.
[0200] The training and test sample ratios of UCM, AID, and NWPU are set to 50%-50% and 80%-20%, 20%-80% and 50%-50%, 10%-90% and 20%-80% respectively. Table 1 presents the relevant information of the three datasets, including the number of images in each category, the number of scene categories, the total number of images, the spatial resolution size, and the size of the images.
[0201] Table 1 Data information of the four datasets
[0202]
[0203]
[0204] Experimental details: All experiments were implemented on a workstation with a GeForce RTX 3070Ti, using the Pytorch framework. We used the Adaptive Moment Estimation to optimize the model, and the parameters of the initial network were randomly initialized. The initial learning rate of the network was set to 0.001, and it was trained for 150 epochs. Cosine annealing was adopted to adjust the learning rate. The image size was resized to 224×224, and random horizontal flipping and random vertical flipping were adopted to augment the images. To reduce the influence of randomness, the main experiments in this paper were repeated five times and the average results were reported.
[0205] In the experiment, the Overall accuracy (OA) and the Confusion Matrix (CM) were used to evaluate the effectiveness of the proposed method. The definition of OA is the number of correctly classified images divided by the total number of test images, which reflects the overall performance of the classification model. CM is used to analyze the classification errors and confusion degrees between different scene categories.
[0206] Experimental results and analysis: To evaluate the method proposed in the present invention, a series of experiments were conducted on three datasets. The data for each dataset experiment are listed in the table. Table 2 lists the methods used for comparison in the present invention and the publication years.
[0207] Table 2 Different methods and publication years in the experiment
[0208]
[0209] 1) Experimental results on the UC-Merced dataset:
[0210] Methods with good classification performance on the UC-Merced dataset in recent years were selected for comparison with the proposed method. The experimental results can be seen in Table 3. When the training ratio was 80%, the classification accuracy of the proposed method reached 99.90%, exceeding all comparison methods. The proposed method had an OA 0.33% higher than that of EMTCAL, 0.61% higher than that of ViT, 0.38% higher than that of TECN, and 0.33% higher than that of VSDNet-ResNet.
[0211] When the training ratio was 50%, the classification accuracy of the proposed method reached 99.33%, exceeding all comparison methods. The proposed method had an OA 0.66% higher than that of EMTCAL, 0.58% higher than that of ViT, 0.43% higher than that of SCViT, and 0.84% higher than that of VSDNet-ResNet using the distillation method.
[0212] Table 3 Comparison between the method we proposed and the methods proposed in recent years on the UC-Merced dataset.
[0213]
[0214] The confusion matrix diagram obtained at a training ratio of 80% is as Figure 8 shown. From Figure 8 it can be seen that each category has been well classified. For the classification of the UC-Merced dataset, the method we proposed has excellent performance.
[0215] 2) Classification results on AID:
[0216] On AID, we compared the method proposed in this invention with some methods in recent years. The experimental results are shown in Table 4. When the training ratio of AID is 20%, the OA of the proposed IRCHKD is 96.10%. It is 1.2% higher than ViT, 1.41% higher than EMTCAL, 0.54% higher than SCViT, 0.65% higher than TECN which combines convolution and Transformer, and 0.93% higher than the SAGN method.
[0217] When the training ratio of AID is 50%, the OA of the proposed method reaches 97.84%, exceeding all comparison methods. It is 2.96% higher than the DFAGCN method, 1.35% higher than the ViT method, 1.43% higher than the EMTCAL method, 0.86% higher than the SCViT method, and 0.44% higher than the TECN method.
[0218] Table 4 Comparison between the method we proposed and the methods proposed in recent years on AID
[0219]
[0220]
[0221] The confusion matrix diagram at a training ratio of 50% is as Figure 9 shown. Among the 30 categories, the accuracy of 27 categories exceeds 90%, and only the accuracy of three categories does not reach 90%. These three categories are "resort", "school", and "square" respectively. Among them, the "resort" scene category is mainly misclassified as "park" and "school". The "school" scene category is mainly misclassified as "commercial area", "square", and "church". The "square" scene category is mainly misclassified as "central area" and "park". This is due to the high inter-class similarity of the scene categories. Further improving the performance of IRCHKD is our future work.
[0222] 3) Classification results on the NWPU dataset:
[0223] The comparison of the classification performance with different methods is summarized in Table 5. The overall accuracy reached 93.13% at a training ratio of 10%, and 95.29% at a training ratio of 20%. Compared with other methods in Table 5, the proposed method achieved the best results, demonstrating the effectiveness of the proposed method. At a training ratio of 10%, our IRCHKD was 2.04% higher than ACNet, 1.50% higher than EMTCAL, 1.40% higher than SAGN, and 1.0% higher than VSDNet-ResNet34. At a training ratio of 20%, our IRCHKD was 2.87% higher than ACNet, 1.64% higher than EMTCAL, 1.80% higher than SAGN, and 0.61% higher than VSDNet-ResNet34.
[0224] Table 5 Comparison of the proposed method with methods with good classification performance in recent years on the NWPU dataset
[0225]
[0226]
[0227] The confusion matrix on the NWPU dataset is as Figure 10 shown. At a training ratio of 20%, the accuracy of only two out of 45 categories did not reach 90%, and the accuracy of 31 categories reached over 95%. This once again proves the effectiveness of the proposed method. The two categories with an accuracy not reaching 90% are "church" and "palace", and their architectural styles are highly similar. The "palace" category was mainly misclassified as the "church" category. The "church" category was also misclassified as the "palace" category a lot.
[0228] Model Size Evaluation: The FLOP and number of parameters of some methods are listed, where FLOPs measure the complexity of the model and the number of parameters measures the size of the model. The experimental results are shown in Table 6. It can be seen from Table 6 that our model still has a classification accuracy 0.44% higher while the number of parameters and complexity are far lower than those of TECN and KFBNet. Compared with methods such as VGG-VD-16, Contourlet CNN, and EMTCAL, the proposed method not only has advantages in terms of the number of parameters and FLOPs but also greatly surpasses these models in terms of classification accuracy. Compared with GoogLeNet, SE-MDPMNet, DenseNet121, and LCNN-BFF, the proposed IRCHKD is slightly higher than these methods in terms of the number of parameters but greatly surpasses these models in terms of accuracy. It is worth mentioning that our model only needs to deploy the classifier network during the deployment stage, and the cross-head knowledge distillation part can be separated from the model, reducing the model complexity and computational resource requirements. It can be seen from Table 6 that the number of parameters of the model during deployment is 2.4M lower than that during training, and the FLOPs are 0.32G lower than that during training.
[0229] Table 6 Comparison of the OA, parameters, and FLOPs of the proposed method with some advanced methods on AID
[0230]
[0231] Discussion: To comprehensively evaluate the effectiveness of the proposed method, some ablation experiments were conducted.
[0232] To verify the effectiveness of the three proposed components, ablation experiments were conducted on the NWPU dataset with a training ratio of 20%. The results of the ablation experiments are shown in Table 7. In the first case, the network was MobileNetv2, and the final obtained model had the worst classification performance. In the second case, IRAM was added to MobileNetv2, and the classification accuracy was improved. In the third case, the MSSA module was added to the MobileNetv2 network, and the model classification performance increased by 1.33%. In the fourth case, CHKD was added to MobileNetv2, and after adding CHKD, the classification accuracy increased by 2.01%. These four cases indicate that the three proposed modules are effective when used alone. In the fifth case, IRAM and MSSA were added, and the accuracy was further improved compared to only adding IRAM. In the sixth case, IRAM and CHKD were added, and the accuracy was also higher than only adding IRAM. In the seventh case, IRAM, MSSA, and CHKD were added, and when the network included these three modules, the highest classification accuracy of 95.25% was obtained. This shows that the proposed modules have better effects when used in combination than when used alone. Compared with the first case, the OA in the seventh case increased by 3.29%. Therefore, the ablation experiments fully prove the effectiveness of the proposed IRCHKD method. Indicates a tick.
[0233] Table 7 Ablation experiment results of the proposed IRCHKD on the NWPU dataset
[0234]
[0235] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
[0236] References
[0237] [1] Jaiswal, R.K.; Saxena, R.; Mukherjee, S. Application of remote sensing technology for land use / land cover change analysis. J. Indian Soc. Remote Sens. 1999, 27, 123–128. [CrossRef]
[0238] [2]Chova,L.G.;Tuia,D.;Moser,G.;Valls,G.C.Multimodal classificationofremote sensing images:Areview and future directions.IEEEProc.2015,103,1560–1584.[CrossRef]
[0239] [3]Cheng,G.;Zhou,P.;Han,J.Learning Rotation-Invariant ConvolutionalNeural Networks for Object Detection in VHR OpticalRemote SensingImages.IEEETrans.Geosci.Remote Sens.2016,54,7405–7415.[CrossRef]
[0240] [4]Zhang,L.;Zhang,L.;Du,B.Deep learning for remote sensing data:Atechnical tutorial on the state-of-the-art.IEEE Geosci.Remote Sens.Mag.2016,4,22–40.[CrossRef]
[0241] [5]S.E.Grigorescu,N.Petkov,and P.Kruizinga,“Comparison oftexturefeatures based on Gabor filters,”IEEE Trans.ImageProcess.,vol.11,no.10,pp.1160–1167,Oct.2002.
[0242] [6]X.Tang,L.Jiao,andW.J.Emery,“SAR image contentretrievalbasedonfuzzysimilarityandrelevance feedback,”IEEE J.Sel.TopicsAppl.Earth Observ.RemoteSens.,vol.10,no.5,pp.1824–1842,May2017.
[0243] [7]S.Mei,J.Ji,J.Hou,X.Li,and Q.Du,“Learning sensor-specific spatial–spectral features ofhyperspectral images via convolutionalneuralnetworks,”IEEETrans.Geosci.Remote Sens.,vol.55,no.8,pp.4520–4533,Aug.2017.
[0244] [8]L.Jiao,X.Tang,B.Hou,and S.Wang,“SAR images retrieval based onsemantic classification and region-based similarity measure for Earthobservation,”IEEE J.Sel.Topics Appl.Earth Observ.Remote Sens.,vol.8,no.8,pp.3876–3891,Aug.2015.
[0245] [9]S.Sergyan,“Color histogram features based image classification incontent-based image retrieval systems,”in Proc.6th Int.Symp.Appl.Mach.Intell.Informat.,Jan.2008,pp.221–224.
[0246]
[10] X.Tang and L.Jiao,“Fusion similarity-based reranking for SARimage retrieval,”IEEE Geosci.Remote Sens.Lett.,vol.14,no.2,pp.242–246,Feb.2017.
[0247]
[11] H. Soltanian-Zadeh, F. Rafiee-Rad, and S. Pourabdollah-Nejad, “Comparison of multiwavelet, wavelet, haralick, and shape features for microcalcification classification in mammograms,” Pattern Recognit., vol. 37, no. 10, pp. 1973–1986, Oct. 2004.
[0248]
[12] X. Tang, L. Jiao, W. J. Emery, F. Liu, and D. Zhang, “Two-stage reranking for remote sensing image retrieval,” IEEE Trans. Geosci. Remote Sens., vol. 55, no. 10, pp. 5798–5817, Oct. 2017.
[0249]
[13] L. Zhang, W. Zhou, and L. Jiao, “Wavelet support vector machine,” IEEE Trans. Syst., Man, Cybern. B, Cybern., vol. 34, no. 1, pp. 34–39, Feb. 2004.
[0250]
[14] M. A. Friedl and C. E. Brodley, “Decision tree classification of land cover from remotely sensed data,” Remote Sens. Environ., vol. 61, pp. 399–409, Sep. 1997.
[0251]
[15] X. Tang, X. Zhang, F. Liu, and L. Jiao, “Unsupervised deep feature learning for remote sensing image retrieval,” Remote Sens., vol. 10, no. 8, p. 1243, Aug. 2018.
[0252]
[16] Wu, Haiyang, Cuiping Shi, Liguo Wang, and Zhan Jin. 2023. "A Cross-Channel Dense Connection and Multi-Scale Dual Aggregated Attention Network for Hyperspectral Image Classification" Remote Sensing 15, no. 9: 2367. https: / / doi.org / 10.3390 / rs15092367
[0253]
[17] C. Shi, H. Wu and L. Wang, "A Feature Complementary Attention Network Based on Adaptive Knowledge Filtering for Hyperspectral Image Classification," in IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1 - 19, 2023, Artno. 5527219, doi: 10.1109 / TGRS.2023.3321840.
[0254]
[18] Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: Transformers for image recognition at scale[J]. arXiv preprint arXiv:2010.11929, 2020.
[0255]
[19] Yu F, Koltun V. Multi-scale context aggregation by dilated convolutions[J]. arXiv preprint arXiv:1511.07122, 2015.
[0256]
[20] G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” 2015, arXiv:1503.02531.
[0257]
[21] Cuiping Shi,Tao Wang,Liguo Wang.Branch Feature Fusion ConvolutionNetwork for Remote Sensing Scene Classification[J],IEEE Journal of SelectedTopics in Applied Earth Observations and Remote Sensing,2020,13(9):5194-5210.DOI:10.1109 / JSTARS.2020.3018307.
[0258]
[22] Zhao B,Zhong Y,Xia G S,et al.Dirichlet-derived multiple topicscene classification model for high spatial resolution remote sensing imagery[J].IEEE Transactions on Geoscience and Remote Sensing,2015,54(4):2108-2123.
[0259]
[23] W.Li et al.,“Classification ofhigh-spatial-resolution remotesensing scenes method using transfer learning and deep convolutional neuralnetwork,”IEEE J.Sel.TopicsAppl.Earth Observ.Remote Sens.,vol.13,pp.1986–1995,2020.
[0260]
[24] Singh,P.;Verma,V.K.;Rai,P.;Namboodiri,V.P.HetConv:HeterogeneousKernel-Based Convolutions for Deep CNNs.arXiv 2019,arXiv:1903.04120.
[0261]
[25] Chen, Y.; Dai, X.; Liu, M.; Chen, D.; Yuan, L.; Liu, Z. Dynamic Convolution: Attention Over Convolution Kernels. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 14–19 June 2020; pp. 11030–11039.
[0262]
[26] Liu, J. J.; Hou, Q.; Cheng, M. M.; Wang, C.; Feng, J. Improving Convolutional Networks with Self-Calibrated Convolutions. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 14–19 June 2020; pp. 10096–10105.
[0263]
[27] Chen, Y.; Fan, H.; Xu, B.; Yan, Z.; Kalantidis, Y.; Rohrbach, M.; Yan, S.; Feng, J. Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution. arXiv 2019, arXiv:1904.05049.
[0264]
[28] Han, K.; Wang, Y.; Tian, Q.; Guo, J.; Xu, C.; Xu, C. GhostNet: More Features from Cheap Operations. arXiv 2020, arXiv:1911.11907.
[0265]
[29] Yang, B.; Bender, G.; Le, Q. V.; Ngiam, J. CondConv: Conditionally Parameterized Convolutions for Efficient Inference. arXiv 2019, arXiv:1904.04971 [cs.CV].
[0266]
[30] Cao, J.; Li, Y.; Sun, M.; Chen, Y.; Lischinski, D.; Cohen-Or, D.; Chen, B.; Tu, C. Depthwise Over-parameterized Convolution. arXiv 2020, arXiv:2006.12030 [cs.CV].
[0267]
[31] Cuiping Shi, Xin Zhao, Liguo Wang*. A Multi-branch Feature Fusion Strategy Based on Attention Mechanism for Remote Sensing Image Scene Classification[J]. Remote Sensing, 2021, 13(10), 1950.
[0268]
[32] X. Tang, Q. Ma, X. Zhang, F. Liu, J. Ma, and L. Jiao, “Attention consistent network for remote sensing scene classification,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 14, pp. 2030–2045, 2021.
[0269]
[33] S.-B. Chen, Q.-S. Wei, W.-Z. Wang, J. Tang, B. Luo, and Z.-Y. Wang, “Remote sensing scene classification via multi-branch local attention network,” IEEE Trans. Image Process., vol. 31, pp. 99–109, 2021.
[0270]
[34] J. Lu et al., “SOFT: SoftMax-free transformer with linear complexity,” in Proc. Adv. Neural Inf. Process. Syst., vol. 34, 2021, pp. 21297–21309.
[0271]
[35] F. Zhu, Y. Zhu, L. Zhang, C. Wu, Y. Fu, and M. Li, “A unified efficient pyramid transformer for semantic segmentation,” in Proc. IEEE / CVF Int. Conf. Comput. Vis. Workshops (ICCVW), Oct. 2021, pp. 2667–2677.
[0272]
[36] D.-J. Chen, H.-Y. Hsieh, and T.-L. Liu, “Adaptive image transformer for one-shot object detection,” in Proc. IEEE / CVF Conf. Comput. Vis. Pattern Recognit., Jun. 2021, pp. 12247–12256.
[0273]
[37] Chen C F R, Fan Q, Panda R. Crossvit: Cross-attention multi-scale vision transformer for image classification[C] / / Proceedings of the IEEE / CVF international conference on computer vision. 2021: 357-366.
[0274]
[38] Z. Liu et al., “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proc. IEEE / CVF Int. Conf. Comput. Vis., Oct. 2021, pp. 10012–10022.
[0275]
[39] D. Hong et al., “SpectralFormer: Rethinking hyperspectral image classification with transformers,” IEEE Trans. Geosci. Remote Sens., vol. 60, pp. 1–15, 2021.
[0276]
[40] Y. Bazi, L. Bashmal, M. M. A. Rahhal, R. A. Dayil, and N. A. Ajlan, “Vision transformers for remote sensing image classification,” Remote Sens., vol. 13, no. 3, p. 516, 2021.
[0277]
[41] J. Ma, M. Li, X. Tang, X. Zhang, F. Liu, and L. Jiao, “Homo–heterogeneous transformer learning framework for remote sensing scene classification,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 15, pp. 2223–2239, 2022.
[0278]
[42] X. Tang, M. Li, J. Ma, X. Zhang, F. Liu and L. Jiao, "EMTCAL: Efficient Multiscale Transformer and Cross-Level Attention Learning for Remote Sensing Scene Classification," in IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-15, 2022, Art no. 5626915, doi: 10.1109 / TGRS.2022.3194505.
[0279]
[43] S. Chen, W. Li, Y. Cao and X. Lu, "Combining the Convolution and Transformer for Classification of Smoke-Like Scenes in Remote Sensing Images," in IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-19, 2022, Art no. 4512519, doi: 10.1109 / TGRS.2022.3208120.
[0280]
[44] Heaton J. Ian Goodfellow, Yoshua Bengio, and Aaron Courville: Deep learning: The MIT Press, 2016, 800 pp, ISBN: 0262035618[J]. Genetic programming and evolvable machines, 2018, 19(1-2): 305-307.
[0281]
[45] A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, “FitNets: Hints for thin deep nets,” 2014, arXiv:1412.6550.
[0282]
[46] Park W, Kim D, Lu Y, et al. Relational knowledge distillation[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019: 3967-3976.
[0283]
[47] Heo B, Kim J, Yun S, et al. A comprehensive overhaul of feature distillation[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2019: 1921-1930.
[0284]
[48] Ahn S,Hu S X,Damianou A,et al.Variational informationdistillation for knowledge transfer[C] / / Proceedings of the IEEE / CVFConference on ComputerVision and Pattern Recognition.2019:9163-9171.
[0285]
[49] Zhang L,Song J,Gao A,et al.Be your own teacher:Improve theperformance ofconvolutional neural networks via self distillation[C] / / Proceedings ofthe IEEE / CVF International Conference on ComputerVision.2019:3713-3722.
[0286]
[50] Ji M,Shin S,Hwang S,et al.Refine myself by teaching myself:Feature refinement via self-knowledge distillation[C] / / Proceedings of theIEEE / CVF conference on computer vision and pattern recognition.2021:10664-10673.
[0287]
[51] Y.Hu et al.,"Hierarchical Self-Distilled Feature Learning forFine-Grained Visual Categorization,"in IEEE Transactions on Neural Networksand Learning Systems,doi:10.1109 / TNNLS.2021.3124135.
[0288]
[52] S. Wei, Y. Luo, X. Ma, P. Ren and C. Luo, "MSH-Net: Modality-Shared Hallucination With Joint Adaptation Distillation for Remote Sensing Image Classification Using Missing Modalities," in IEEE Transactions on Geoscience and Remote Sensing, vol. 61, pp. 1-15, 2023, Art no. 4402615, doi: 10.1109 / TGRS.2023.3265650.
[0289]
[53] Y. Hu, X. Huang, X. Luo, J. Han, X. Cao and J. Zhang, "Variational Self-Distillation for Remote Sensing Scene Classification," in IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-13, 2022, Art no. 5627313, doi: 10.1109 / TGRS.2022.3194549.
[0290]
[54] D. Li, Y. Nan and Y. Liu, "Remote Sensing Image Scene Classification Model Based on Dual Knowledge Distillation," in IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1-5, 2022, Art no. 4514305, doi: 10.1109 / LGRS.2022.3208904.
[0291]
[55] H. Liu, Y. Qu and L. Zhang, "Multispectral Scene Classification via Cross-Modal Knowledge Distillation," in IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-12, 2022, Art no. 5409912, doi: 10.1109 / TGRS.2022.3174352.
[0292]
[56] Y. Yang and S. Newsam, “Bag-of-visual-words and spatial extensions for land-use classification,” in Proc. 18th SIGSPATIAL Int. Conf. Adv. Geograph. Inf. Syst., 2010, pp. 270–279.
[0293]
[57] G. S. Xia et al., “AID: A benchmark data set for performance evaluation of aerial scene classification,” IEEE Trans. Geosci. Remote Sens., vol. 55, no. 7, pp. 3965–3981, Jul. 2017.
[0294]
[58] G. Cheng, J. Han, and X. Lu, “Remote sensing image scene classification: Benchmark and state of the art,” Proc. IEEE, vol. 105, no. 10, pp. 1865–1883, Oct. 2017.
[0295]
[59] N. He, L. Fang, S. Li, A. Plaza, and J. Plaza, “Remote sensing scene classification using multilayer stacked covariance pooling,” IEEE Trans. Geosci. Remote Sens., vol. 56, no. 12, pp. 6899–6910, Dec. 2018.
[0296]
[60] W. Zhang, P. Tang, and L. Zhao, “Remote sensing image scene classification using CNN-CapsNet,” Remote Sens., vol. 11, no. 5, p. 494, 2019.
[0297]
[61] N. He, L. Fang, S. Li, J. Plaza, and A. Plaza, “Skip-connected covariance network for remote sensing scene classification,” IEEE Trans. Neural Netw. Learn. Syst., vol. 31, no. 5, pp. 1461–1474, May 2020.
[0298]
[62] H. Sun, S. Li, X. Zheng, and X. Lu, “Remote sensing scene classification by gated bidirectional network,” IEEE Trans. Geosci. Remote Sens., vol. 58, no. 1, pp. 82–96, Jan. 2020.
[0299]
[63] S. Wang, Y. Guan, and L. Shao, “Multi-granularity canonical appearance pooling for remote sensing scene classification,” IEEE Trans. Image Process., vol. 29, pp. 5396–5407, 2020.
[0300]
[64] Q. Bi, K. Qin, Z. Li, H. Zhang, K. Xu, and G.-S. Xia, “A multiple-instance densely-connected ConvNet for aerial scene classification,” IEEE Trans. Image Process., vol. 29, pp. 4911–4926, 2020.
[0301]
[65] X. Wang, L. Duan, A. Shi, and H. Zhou, “Multilevel feature fusion networks with adaptive channel dimensionality reduction for remote sensing scene classification,” IEEE Geosci. Remote Sens. Lett., vol. 19, pp. 1–5, 2021.
[0302]
[66] G. Zhang et al., “A multiscale attention network for remote sensing scene images classification,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 14, pp. 9530–9545, 2021.
[0303]
[67] X. Wang, L. Duan, C. Ning, and H. Zhou, “Relation-attention networks for remote sensing scene classification,” IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens., vol. 15, pp. 422–439, 2021.
[0304]
[68] X. Wang, S. Wang, C. Ning, and H. Zhou, “Enhanced feature pyramid network with deep semantic embedding for remote sensing scene classification,” IEEE Trans. Geosci. Remote Sens., vol. 59, no. 9, pp. 7918–7932, Sep. 2021.
[0305]
[69] K. Xu, H. Huang, P. Deng, and Y. Li, “Deep feature aggregation framework driven by graph convolutional network for scene classification in remote sensing,” IEEE Trans. Neural Netw. Learn. Syst., early access, Apr. 15, 2021, doi: 10.1109 / TNNLS.2021.3071369.
[0306]
[70] P. Lv, W. Wu, Y. Zhong, F. Du and L. Zhang, "SCViT: A Spatial-Channel Feature Preserving Vision Transformer for Remote Sensing Image Scene Classification," in IEEE Transactions on Geoscience and Remote Sensing, vol. 60, pp. 1-12, 2022, Art no. 4409512, doi: 10.1109 / TGRS.2022.3157671.
[0307]
[71] Q. Meng, M. Zhao, L. Zhang, W. Shi, C. Su and L. Bruzzone, "Multilayer Feature Fusion Network With Spatial Attention and Gated Mechanism for Remote Sensing Scene Classification," in IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1-5, 2022, Art no. 6510105, doi: 10.1109 / LGRS.2022.3173473.
[0308]
[72] P. Deng, H. Huang and K. Xu, "A Deep Neural Network Combined With Context Features for Remote Sensing Scene Classification," in IEEE Geoscience and Remote Sensing Letters, vol. 19, pp. 1 - 5, 2022, Art no. 8000405, doi: 10.1109 / LGRS.2020.3016769.
[0309]
[73] Shi C, Zhang X, Wang L. A Lightweight Convolutional Neural Network Based on Channel Multi - Group Fusion for Remote Sensing Scene Classification. Remote Sensing. 2022;14(1):9. https: / / doi.org / 10.3390 / rs14010009
[0310]
[74] Y. Yang, X. Tang, Y.-M. Cheung, X. Zhang and L. Jiao, "SAGN: Semantic - Aware Graph Network for Remote Sensing Scene Classification," in IEEE Transactions on Image Processing, vol. 32, pp. 1011 - 1025, 2023, doi: 10.1109 / TIP.2023.3238310.
[0311]
[75] Zhang, B.; Zhang, Y.; Wang, S. A lightweight and discriminative model for remote sensing scene classifification with multidilation pooling module. IEEE J. Sel. Topics Appl. Earth Observ. Remote Sens. 2019, 12, 2636–2653. [CrossRef]
[0312]
[76] G. Huang, Z. Liu, L. Van Der Maaten, et al. Densely connected convolutional networks[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2017:4700 - 4708.
[0313]
[77] Liu, M.; Jiao, L.; Liu, X.; Li, L.; Liu, F.; Yang, S. C-CNN: Contourlet convolutional neural networks. IEEE Trans. Neural Netw. Learn. Syst. 2021, 32, 2636–2649. [CrossRef]
[0314]
[78] F. Li, R. Feng, W. Han, et al. High-resolution remote sensing image scene classification via key filter bank based on convolutional neural network[J]. IEEE Transactions on Geoscience and Remote Sensing, 2020, 58(11):8077 - 8092.
Claims
1. A method for object recognition based on an inverted residual cross-head knowledge distillation neural network, characterized by: The specific process of the method is:
1. Construct an inverted residual cross-head knowledge distillation model; the specific process is: The inverted residual cross-head knowledge distillation model includes a classifier network and a cross-head knowledge distillation structure CHKD; The classifier network includes convolution blocks, Block1, Block2, the first self-attention fused inverted residual attention module IRAM, the second self-attention fused inverted residual attention module IRAM, the multi-scale spatial attention module MSSA, and the first linear layer Linear; The cross-head knowledge distillation structure includes feature aggregation and cross-head distillation; The feature aggregation includes a first Inception former, a second Inception former, a third Inception former, a fourth Inception former, a downsampling Down, a first Fusion module, and a second Fusion module; The cross-head distillation part includes a second linear layer Linear; The classifier network is also called the student network; Feature aggregation is also called teacher network; The specific working process of the inverted residual cross-head knowledge distillation model is as follows: The feature map is sequentially input into the convolutional block and Block1 in the classifier network, and Block1 outputs the feature α; Feature α is input into Block2, and Block2 outputs feature α′; The feature α′ is input into the inverted residual attention module of the first self-attention fusion, and the inverted residual attention module of the first self-attention fusion outputs the feature α″′; The output feature α″′ of the inverted residual attention module of the first self-attention fusion is input into the inverted residual attention module of the second self-attention fusion. The output feature of the inverted residual attention module of the second self-attention fusion is feature Input the multi-scale spatial attention module, and the multi-scale spatial attention module outputs feature β; Feature α is input into the first Inception former, the first Inception former outputs feature input downsampling Down, and downsampling Down outputs feature D; Block2 outputs feature α′ which is input into the second Inception former, and the second Inception former outputs feature E; The down-sampled Down output feature D and the second Inception former output feature E are input into the first Fusion module, and the first Fusion module outputs feature F; The output feature α″′ of the inverted residual attention module fused by the first self-attention is input into the third Inception former, and the third Inception former outputs the feature G; The third Inception former outputs feature G and the first Fusion module outputs feature F which are input into the second Fusion module, and the second Fusion module outputs feature H; Output features of the inverted residual attention module after the second self-attention fusion Input the fourth Inception former, and the fourth Inception former outputs feature I; The fourth Inception former output feature I and the second Fusion module output feature H are added element by element to obtain feature J; The output feature β of the multi-scale spatial attention module is input into the first linear layer Linear, and the first linear layer Linear outputs the feature P S ; The output feature β of the multi-scale spatial attention module is input into the second linear layer, and the output feature of the first linear layer is Feature J is input into the second linear layer, and the first linear layer outputs feature P t ; Based on feature P S Calculate the student network loss function L between the student network output and the true label CE ; Feature-based and P t Calculate the loss function L for knowledge distillation between the student network output and the true label CKD ; Based on the student network loss function L CE And the loss function L of knowledge distillation CKD Calculate the total supervision loss Loss; Based on the total supervision loss Loss, the trained inverted residual cross-head knowledge distillation model is obtained; 2. Based on the trained inverse residual cross-head knowledge distillation model, the classifier network is used to identify the objects in the feature map to be tested.
2. The method for identifying objects based on inverted residual cross-head knowledge distillation neural network according to claim 1 is characterized in that: The convolution block sequentially includes a first 3×3 convolution layer, a BN and a SiLU activation function; The Block1 includes a first 1×1 convolutional layer, a first 3×3 depthwise separable convolutional layer, and a second 1×1 convolutional layer; The Block2 includes a third 1×1 convolutional layer, a second 3×3 depthwise separable convolutional layer, and a fourth 1×1 convolutional layer; The working process of Block 1 is as follows: Feature x is sequentially input into the first 1×1 convolutional layer, the first 3×3 depth-separable convolutional layer, and the second 1×1 convolutional layer. The output feature of the second 1×1 convolutional layer is added element-by-element to feature x to obtain feature α, which is used as the output feature of Block1. The working process of Block2 is as follows: The output feature α of Block1 is sequentially input into the third 1×1 convolutional layer, the second 3×3 depth-separable convolutional layer, and the fourth 1×1 convolutional layer. The output feature of the fourth 1×1 convolutional layer is added element-by-element to the feature α to obtain the feature α′, which is used as the output feature of Block2.
3. The method for identifying objects based on inverted residual cross-head knowledge distillation neural network according to claim 2 is characterized in that: The first self-attention fused inverted residual attention module IRAM includes self-attention, relative position bias, Softmax, a third 3×3 depth-separable convolutional layer, and a fifth 1×1 convolutional layer; The working process of the first self-attention fusion inverse residual attention module IRAM is as follows: 1) The feature graph passes through the third linear layer, the fourth linear layer, and the fifth linear layer respectively. The third linear layer outputs the query matrix Q, the fourth linear layer outputs the key matrix K, and the fifth linear layer outputs the value matrix V; 2) Based on the query matrix Q, key matrix K, value matrix V and relative position bias B, we get self-attention Attention(Q,K,V); the calculation process is: Where Q, K, and V are query matrix, key matrix, and value matrix respectively. d represents the dimension of Q / K, B represents the relative position offset, M 2 Represents the number of pixel blocks in the window, represents a real number; the superscript T represents the transpose; 3) Input the self-attention Attention(Q,K,V) into the third 3×3 depth-separable convolutional layer, and add the output features of the third 3×3 depth-separable convolutional layer and the self-attention Attention(Q,K,V) element by element to obtain the feature α″; The feature α″ is input into the fifth 1×1 convolutional layer, and the fifth 1×1 convolutional layer outputs the feature α″′, which is used as the output of the inverted residual attention module IRAM of the first self-attention fusion.
4. The method for identifying objects based on inverted residual cross-head knowledge distillation neural network according to claim 3 is characterized in that: The second self-attention fused inverted residual attention module IRAM includes self-attention, relative position bias, Softmax, a fourth 3×3 depth-separable convolutional layer, and a sixth 1×1 convolutional layer; The working process of the second self-attention fusion inverted residual attention module IRAM is as follows: 1) The output features of the inverse residual attention module IRAM of the first self-attention fusion pass through the sixth linear layer, the seventh linear layer, and the eighth linear layer respectively. The sixth linear layer outputs the query matrix Q, the seventh linear layer outputs the key matrix K, and the eighth linear layer outputs the value matrix V; 2) Based on the query matrix Q, key matrix K, value matrix V and relative position bias B, we get self-attention Attention(Q,K,V); the calculation process is: Where Q, K, and V are query matrix, key matrix, and value matrix respectively. d represents the dimension of Q / K, B represents the relative position offset, M 2 Represents the number of pixel blocks in the window, represents a real number; 3) Input the self-attention Attention(Q,K,V) into the fourth 3×3 depth-separable convolutional layer, and add the output features of the fourth 3×3 depth-separable convolutional layer and the self-attention Attention(Q,K,V) element by element to obtain the feature feature Input the sixth 1×1 convolution layer, the sixth 1×1 convolution layer outputs features feature The output of the inverted residual attention module IRAM as the second self-attention fusion.
5. The method for identifying objects based on inverted residual cross-head knowledge distillation neural network according to claim 4 is characterized in that: The multi-scale spatial attention module MSSA includes four branches and a 1×1 convolutional layer; The first branch includes a 1×1 convolutional layer and spatial attention in sequence; The second branch includes 1×1 convolution layer, 3×3 convolution layer, batch normalization layer BN, SiLU activation function, and spatial attention in sequence; The third branch includes 1×1 convolution layer, 3×3 convolution layer, batch normalization layer BN, SiLU activation function, 3×3 convolution layer with expansion rate 2, batch normalization layer BN, SiLU activation function, and spatial attention. The fourth branch includes 1×1 convolution layer, 3×3 convolution layer, batch normalization layer BN, SiLU activation function, 3×3 convolution layer with expansion rate 2, batch normalization layer BN, SiLU activation function, 3×3 convolution layer with expansion rate 3, batch normalization layer, SiLU activation function, and spatial attention. The working process of the multi-scale spatial attention module MSSA is as follows: Output features of the inverted residual attention module after the second self-attention fusion Input the four branches respectively, connect the features obtained from the four branches, input the connected features into the 1×1 convolution layer, and the output features of the 1×1 convolution layer are used as the output features of the multi-scale spatial attention module MSSA.
6. The method for identifying objects based on inverted residual cross-head knowledge distillation neural network according to claim 5 is characterized in that: The spatial attention includes maximum pooling, average pooling, 3×3 convolution layer, and Softmax layer in sequence; The spatial attention working process is: The feature maps are input into the maximum pooling and average pooling respectively, and the feature maps after the maximum pooling and the average pooling are connected by channel. After the connection, the feature maps pass through the 3×3 convolution layer, and the output features of the 3×3 convolution layer pass through the Softmax layer to obtain the attention map; Expressed as Ms(F)=Softmax(f 3 × 3 ([AvgPool(F);MaxPool(F)])) Where F represents the feature map, f 3 × 3 represents a 3×3 convolutional layer, M s (F) represents the attention map output after the feature map F undergoes spatial attention; AvgPool(F) means average pooling of feature map F; MaxPool(F) means to perform maximum pooling on feature map F.
7. The method for identifying objects based on inverted residual cross-head knowledge distillation neural network according to claim 6 is characterized in that: Each Inceptionformer in the first Inceptionformer, the second Inceptionformer, the third Inceptionformer, and the fourth Inception former includes a first branch, a second branch, a ninth linear layer, a GELU, and a tenth linear layer; The first branch contains 1×1 convolutional layer, batch normalization BN and SiLU activation function in sequence; The second branch contains a 3×3 convolutional layer, batch normalization BN, and SiLU activation function in sequence; The working process of each Inceptionformer in the first Inceptionformer, the second Inceptionformer, the third Inceptionformer, and the fourth Inception former is as follows: The feature maps are input into the first branch and the second branch respectively, and the output features of the first branch and the second branch are concatenated. The concatenated features are input into the ninth linear layer, GELU, and the tenth linear layer in turn. The output features of the tenth linear layer are added to the concatenated features element by element to obtain the output features of Inceptionformer.
8. The method for identifying objects based on inverted residual cross-head knowledge distillation neural network according to claim 7 is characterized in that: The downsampling Down includes a first branch, a second branch, and a 1×1 convolution layer; The first branch contains two-dimensional average pooling, 1×1 convolution, BN and SiLU activation functions in sequence; The second branch contains 3×3 convolution, BN and SiLU activation functions in sequence; The working process of downsampling Down is as follows: The feature maps are input into the first branch and the second branch respectively, the output feature maps of the first branch and the second branch are connected, and a 1×1 convolution layer is used to mix the channels of the connected feature maps to obtain the downsampled Down output features.
9. The method for identifying objects based on inverted residual cross-head knowledge distillation neural network according to claim 8 is characterized in that: Each of the first and second fusion modules includes a first branch, a second branch, a 1×1 convolution layer, a BN and a SiLU activation function; The first branch contains average pooling, 1×1 convolutional layer, BN and SiLU activation functions in sequence; The second branch contains 3×3 convolutional layers, BN and SiLU activation functions in sequence; The working process of each Fusion module in the first Fusion module and the second Fusion module is as follows: The two input feature maps are spliced along the channel dimension, the channels of the spliced feature maps are shuffled, and the shuffled feature maps are input into the first branch and the second branch respectively. The feature maps obtained from the two branches are spliced in the channel dimension. The spliced feature maps are sequentially input into the 1×1 convolutional layer, BN and SiLU activation function, and the SiLU activation function outputs the feature map.
10. The method for identifying objects based on inverted residual cross-head knowledge distillation neural network according to claim 9 is characterized in that: Based on the feature P S Calculate the student network loss function L between the student network output and the true label CE ; Feature-based and P t Calculate the loss function L for knowledge distillation between the student network output and the true label CKD ; Based on the student network loss function L CE And the loss function L of knowledge distillation CKD Calculate the total supervision loss Loss; The expression is: Student network loss function L CE The true labels are learned using the cross entropy loss, which is expressed as: Where N represents the training sample, x represents the input tensor, θs represents the parameters in the student network, and yi represents the true label. represents the output of the student network after the fully connected layer; The loss function L for knowledge distillation CKD If it is: Where T represents the temperature hyperparameter, D KL Pt and KL distance between two distributions; The total supervision loss Loss is expressed as: The inverted residual cross-head knowledge distillation model is trained based on the total supervision loss Loss to obtain a trained inverted residual cross-head knowledge distillation model.
Citation Information
Cited By
Knowledge distillation-based 1D-CNN online partial discharge identification method
CN121410481A
Image classification method based on personalized federated distillation
CN121708359A
An image classification method based on personalized federated distillation
CN121708359B