Lightweight High-Resolution Remote Sensing Scene Classification Method Based on Multi-Level Adaptive Knowledge Distillation
Through multi-level adaptive knowledge distillation technology, the output layer distillation temperature of the high-score remote sensing image scene classification model is adaptively adjusted, and spatial attention and channel correlation are increased, which solves the problems of large model calculation volume, low accuracy and poor knowledge distillation in the existing technology, and realizes a high-performance lightweight model.
Patent Information
- Application Number
- CN202210837954.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-16
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-07-16
AI Technical Summary
In the classification of high-score remote sensing image scenes, the deep neural network model is computationally large and time-consuming, while the lightweight model is fast but has low accuracy, so it cannot be directly applied to embedded devices. In addition, the use of constant temperature during knowledge distillation leads to a decrease in model effect, and the output layer semantic information of the teacher model is not sufficient to guide the generalized learning of the student model.
The multi-level adaptive knowledge distillation method is adopted to adjust the distillation temperature of the output layer through the adaptive temperature mechanism, so that the student model can selectively learn the probability distribution knowledge of the teacher model's output layer. At the same time, the spatial attention and inter-channel correlation of the features of the last convolutional layer of the model are increased, and the learning guidance of the student model is enhanced.
It realizes that on the basis of ensuring model accuracy, reduces the number of parameters and improves the overall performance of the model, is suitable for embedded devices, and improves the generalized learning ability of students' models.
Smart Images

Figure CN115187863B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of remote sensing technology, and in particular to a lightweight high-resolution remote sensing scene classification method based on multi-level adaptive knowledge distillation. Background Art
[0002] Scene classification and recognition of high-resolution remote sensing images refer to the use of semantic features at different levels to identify and label the image content of sub-regions extracted from high-resolution remote sensing images. It is a key link in the intelligent processing of remote sensing information and is of great significance for land resource management and optimization, urban planning, natural disaster assessment and detection, vegetation mapping, and promoting sustainable development. One limitation of the current remote sensing image scene classification work is that the deep neural network model has a large amount of calculation and is time-consuming, while the lightweight model is fast but has low accuracy, and neither can be directly applied to embedded devices. In this case, an effective solution is to compress the deep CNN model to obtain a more concise and effective model. Knowledge distillation is one of the current mainstream model compression algorithms. The output semantic information of the trained deep network is used as teacher knowledge to guide the training of a new shallow network, thereby improving the model learning and generalization ability of the shallow network. Finally, on the basis of ensuring the model accuracy, the number of parameters is reduced and the overall performance of the model is improved.
[0003] However, since remote sensing scene images are taken from high altitudes, cover a large area, contain more objects than ordinary images, and the object composition is more complex, this leads to an imbalance in the degree of difference between images of different categories. In the knowledge distillation process of high-resolution remote sensing image scene classification, if a constant temperature is used, the model effect will decline because the distillation temperature affects the learning degree of the student model for soft label samples. In addition, due to the rich structural information of high-resolution remote sensing, the structural information will be disrupted after the features extracted based on the deep convolutional neural network model are fed into the classifier. Therefore, using only the semantic information of the output layer of the teacher model as the guiding information for the student model limits the generalization learning of the student model. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a lightweight high-resolution remote sensing scene classification method based on multi-level adaptive knowledge distillation, which can adaptively adjust the distillation temperature of the output layer in the knowledge distillation process, so that the student model can selectively learn the probability distribution knowledge of the output layer of the teacher model.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A lightweight high-resolution remote sensing scene classification method based on multi-level adaptive knowledge distillation, including the following steps:
[0006] Step S1: Send the remote sensing image data into the teacher model and the student model respectively for feature extraction;
[0007] Step S2: The features extracted by the teacher model and the student model are respectively fed into a classification module composed of a normalization function BatchNorm, a fully connected layer, and a fully connected layer to obtain corresponding probability distribution outputs. After adjusting the temperature through an adaptive temperature mechanism, the adjusted probability distribution output of the teacher model is used for distillation learning with the probability distribution output of the student model;
[0008] Step S3: The features extracted by the teacher model and the student model are respectively used to generate corresponding spatial attentions, and distillation learning is performed from the spatial attentions of the features of the teacher model and the student model;
[0009] Step S4: The features extracted by the teacher model and the student model are respectively used to generate corresponding inter-channel correlations, and distillation learning is performed from the inter-channel correlations of the features of the teacher model and the student model.
[0010] In a preferred embodiment, the temperature is adjusted through an adaptive temperature mechanism as follows:
[0011] Set an initial temperature value T s , and adaptively adjust the temperature according to the soft labels of the teacher model:
[0012] T s
[0013] where P t_max is the maximum probability value output by the classifier of the teacher model; by adaptively adjusting the distillation temperature for different image samples, for scene category samples with high correlation, the temperature increases, and the student model can learn more probability distribution knowledge; while for scene category samples with low correlation, the temperature decreases, and the student model learns less probability distribution knowledge.
[0014] In a preferred embodiment, during the knowledge distillation process, the calculation of spatial attention and inter-channel correlation in the features of the last convolutional layer of the teacher model and the student model is introduced. The feature spatial attention calculation of the teacher model is used to guide the feature spatial region that the student model should focus on; the inter-channel correlation calculation of the teacher model's features is used to guide the student model to learn from the inter-channel correlation of the features. Assume that the intermediate feature outputs of the teacher model and the student model are respectively and It is assumed that the features of the teacher model have m-dimensional channels, and A(·) represents the conversion to convert the feature map into a single-channel feature attention with a size of h×w:
[0015]
[0016] where After obtaining the feature attentions of the teacher model and the student model, the Smooth L1 loss function is used to calculate the feature attentions. The calculation of the inter-channel correlation of the features is as follows:
[0017]
[0018] After obtaining the inter-channel correlations of the features of the teacher model and the student model, the L2 loss function is used to calculate the inter-channel correlation loss of the features.
[0019] In a preferred embodiment, the spatial attention and inter-channel correlation information of the feature layers of the output layer and the last convolutional layer of the collaborative model are used as constraints, so that the output layer and the feature layer of the student model are similar to those of the teacher model.
[0020] Compared with the prior art, the present invention has the following beneficial effects: It can adaptively adjust the distillation temperature of the output layer in the knowledge distillation process, so that the student model selectively learns the probability distribution knowledge of the output layer of the teacher model; in addition, it increases the spatial attention and inter-channel correlation of the features of the last convolutional layer of the model, enhances the guidance of the student model learning at multiple levels, and finally obtains a high-performance lightweight model. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic diagram of the principle of a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The present invention will be further described below in conjunction with the drawings and embodiments.
[0023] It should be noted that the following detailed description is illustrative and is intended to provide further description of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0024] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0025] A lightweight high-resolution remote sensing scene classification method for multi-level adaptive knowledge distillation, referring to Figure 1 , includes the following steps:
[0026] Step S1: Send the remote sensing image data into the teacher model and the student model respectively for feature extraction. Among them, the teacher model usually selects a classification model with high accuracy, including but not limited to ResNet-152, and the student model usually selects a lightweight classification model, including but not limited to MobileNetV3;
[0027] Step S2: Send the features extracted by the teacher model and the student model into the classification module composed of the normalization function BatchNorm, the fully connected layer, and the fully connected layer respectively to obtain the corresponding probability distribution outputs. After adjusting the temperature through the adaptive temperature mechanism, perform distillation learning on the probability distribution output of the adjusted teacher model and the probability distribution output of the student model; Through the output layer knowledge distillation of the adaptive temperature mechanism, the student model can better learn the negative sample knowledge from the output layer of the teacher model, while reducing the impact of label noise on scenes with low difference.
[0028] Step S3: Generate corresponding spatial attentions for the features extracted by the teacher model and the student model respectively, and perform distillation learning from the spatial attentions of the features of the teacher model and the student model; During the knowledge distillation process, increase the spatial attention and inter-channel correlation of the features of the last convolutional layer of the model to enhance the guidance for the student model to learn.
[0029] Step S4: Generate corresponding inter-channel correlations for the features extracted by the teacher model and the student model respectively, and perform distillation learning from the inter-channel correlations of the features of the teacher model and the student model.
[0030] The temperature is adjusted through the adaptive temperature mechanism as follows:
[0031] Set an initial temperature value T s , and adaptively adjust the temperature according to the soft labels of the teacher model:
[0032]
[0033] In the formula, P t_max is the maximum probability value output by the classifier of the teacher model; By adaptively adjusting the distillation temperature for different image samples, for the scene category samples with high correlation, the temperature increases, and the student model can learn more probability distribution knowledge; for the scene category samples with low correlation, the temperature decreases, and the student model learns less probability distribution knowledge.
[0034] During the knowledge distillation process, the calculation of spatial attention and inter-channel correlation in the features of the last convolutional layers of the teacher model and the student model is introduced. The spatial attention calculation of the teacher model's features guides the feature space region that the student model should focus on; the inter-channel correlation calculation of the teacher model's features guides the student model to learn from the inter-channel correlation of the features. Assume that the intermediate feature outputs of the teacher model and the student model are respectively and It is assumed that the features of the teacher model have m-dimensional channels, and A(·) represents the conversion to convert the feature map into a single-channel feature attention of size h×w:
[0035]
[0036] In the formula After obtaining the feature attentions of the teacher model and the student model, the Smooth L1 loss function is used to calculate the feature attentions. The calculation of the inter-channel correlation of the features is as follows:
[0037]
[0038] After obtaining the inter-channel correlations of the features of the teacher model and the student model, the L2 loss function is used to calculate the inter-channel correlation loss of the features.
[0039] The spatial attention and inter-channel correlation information of the output layer and the feature layer of the last convolutional layer of the collaborative model are used as constraints, making the output layer and the feature layer of the student model similar to those of the teacher model. The information of the output layer and the feature layer of the last convolutional layer of the collaborative model allows for architectural differences between the student model and the teacher model, making the model more widely applicable.
Claims
1. A lightweight high-resolution remote sensing scene classification method based on multi-level adaptive knowledge distillation, Characterized in that, It includes the following steps: Step S1: Send the remote sensing image data into the teacher model and the student model respectively for feature extraction; Step S2: Send the features extracted by the teacher model and the student model into the classification module composed of the normalization function BatchNorm, the fully connected layer, and the fully connected layer respectively; obtain the corresponding probability distribution output, and after adjusting the temperature through the adaptive temperature mechanism, perform distillation learning on the adjusted probability distribution output of the teacher model and the probability distribution output of the student model; Step S3: Generate corresponding spatial attentions for the features extracted by the teacher model and the student model respectively, and perform distillation learning from the spatial attentions of the features of the teacher model and the student model; Step S4: Generate corresponding inter-channel correlations for the features extracted by the teacher model and the student model respectively, and perform distillation learning from the inter-channel correlations of the features of the teacher model and the student model; The specific adjustment of the temperature through the adaptive temperature mechanism is as follows: Set an initial temperature value T s , and adaptively adjust the temperature according to the soft labels of the teacher model: where P t_max is the maximum probability value output by the classifier of the teacher model; by adaptively adjusting the temperature of distillation for different image samples, the temperature increases for scene category samples with high correlation, enabling the student model to learn more knowledge of probability distribution; while for scene category samples with low correlation, the temperature decreases, and the student model learns less knowledge of probability distribution; During the process of knowledge distillation, introduce the calculation of the spatial attention and the inter-channel correlation in the features of the last convolutional layer of the teacher model and the student model. Guide the feature space area that the student model should focus on through the calculation of the feature spatial attention of the teacher model; Guide the student model to learn from the inter-channel correlation of the features through the calculation of the inter-channel correlation of the features of the teacher model; Suppose the intermediate feature outputs of the teacher model and the student model are respectively and It is assumed that the features of the teacher model have m-dimensional channels, and A(·) represents the feature attention that converts the feature map into a single-channel feature with a size of h×w: In the formula After obtaining the feature attentions of the teacher model and the student model, the Smooth L1 loss function is used to calculate the feature attentions; the calculation of the inter-channel correlation of the features is as follows: After obtaining the inter-channel correlations of the features of the teacher model and the student model, use the L2 loss function to calculate the channel correlation loss of the features.
2. The lightweight high-resolution remote sensing scene classification method based on multi-level adaptive knowledge distillation according to claim 1, Characterized in that, Use the spatial attention and the inter-channel correlation information of the output layer and the feature layer of the last convolutional layer of the collaborative model as constraint conditions, so that the output layer and the feature layer of the student model are similar to those of the teacher model.
Citation Information
Patent Citations
Knowledge graph alignment method, device and equipment based on knowledge distillation
CN114328952A