Knowledge distillation-based lightweight remote sensing image target detection method
By using knowledge distillation technology in remote sensing image target detection, the knowledge of the teacher model is transferred to the lightweight student model, which solves the problem of excessively large remote sensing image target detection model size and realizes real-time and accurate target detection on edge devices.
Patent Information
- Application Number
- CN202511196885.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-18
AI Technical Summary
Existing remote sensing image target detection models are too large to be deployed on resource-constrained devices, making real-time target detection impossible on edge devices.
By employing knowledge distillation technology, a pre-trained teacher model and a lightweight student model are trained through knowledge transfer to generate a lightweight remote sensing image target detection model, which is then deployed on resource-constrained devices for target detection.
It enables accurate and convenient target detection of remote sensing images on resource-constrained devices, improving the accuracy and timeliness of detection.
Smart Images

Figure CN120976783A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a lightweight remote sensing image target detection method based on knowledge distillation. Background Technology
[0002] Remote sensing technology originated from aerial photography in the early 20th century. With advancements in space technology, the launch of satellites in 1957 ushered in the era of satellite remote sensing. Remote sensing images have also undergone a series of changes with the innovation of sensor technology, from early visible light imaging to multispectral images, hyperspectral images, and the now widely used synthetic aperture radar (SAR) images. These images have all improved humanity's ability to observe the Earth. However, with the increasing number of satellites and the implementation of Earth observation programs by various countries, satellites worldwide now generate petabytes of data daily. Such massive amounts of data cannot be processed manually. Furthermore, while the resolution of remote sensing images has gradually increased, enhancing detail capture capabilities, it has also made it more difficult to directly discern useful information with the human eye. Finally, with the intensification of climate change and the increasing frequency of global natural disasters, there is a growing need for a system capable of real-time emergency response, continuously monitoring ground conditions and providing early warnings. Based on these increasingly complex realities, the task of remote sensing image interpretation emerged and has continued to advance with the development of computer technology. The goal is to use powerful computing devices to process this remote sensing image data and directly provide results that enable further decision-making.
[0003] With the development of deep learning technology, neural networks are increasingly replacing traditional methods for interpreting remote sensing images. These networks can automatically extract features from remote sensing images and achieve better interpretation results. From the initial CNN (Convolutional Neural Network) to the now popular Vision Transformer and other large models, the predictive performance of these models has improved significantly. However, the model size and number of parameters have also increased dramatically, making deployment on edge devices such as satellites and drones difficult. Currently, the prevailing approach is "space-to-ground computing," transmitting satellite-captured images to the ground for analysis and processing. However, with the deepening deployment of satellite networks, the ability to perform preliminary analysis and processing of remote sensing images on satellites is becoming increasingly urgent and better meets the current real-time requirements. Therefore, how to significantly compress the model size with minimal loss of accuracy, enabling deployment on devices with limited computing power, remains a problem to be solved. Summary of the Invention
[0004] This invention provides a lightweight remote sensing image target detection method based on knowledge distillation, which can solve the problem that existing remote sensing image target detection models are too large to be deployed on resource-constrained devices. It can also directly and accurately detect targets in the acquired remote sensing images on resource-constrained devices where lightweight models are deployed, thereby improving the accuracy and timeliness of target detection.
[0005] In a first aspect, embodiments of the present invention provide a lightweight remote sensing image target detection method based on knowledge distillation, comprising:
[0006] Obtain a training sample set and a pre-built teacher-student model architecture; the teacher-student model architecture includes: a pre-trained target teacher model and a preset student model to be trained; the preset student model is a lightweight model compared to the target teacher model;
[0007] Based on the training sample set and the target teacher model, the preset student model is trained by knowledge distillation to obtain a trained target student model; the target student model is deployed on a resource-constrained device.
[0008] In response to the issued target remote sensing image, target detection is performed on the target remote sensing image based on the target student model to obtain the target detection result corresponding to the target remote sensing image.
[0009] Secondly, embodiments of the present invention also provide a lightweight remote sensing image target detection device based on knowledge distillation, the device comprising:
[0010] The data acquisition module is used to acquire a training sample set and a pre-built teacher-student model architecture; the teacher-student model architecture includes: a pre-trained target teacher model and a preset student model to be trained; the preset student model is a lightweight model compared to the target teacher model;
[0011] The target student model determination module is used to perform knowledge distillation training on the preset student model based on the training sample set and the target teacher model to obtain a trained target student model; the target student model is deployed on a resource-constrained device.
[0012] The target detection result determination module is used to respond to the issued target remote sensing image, perform target detection on the target remote sensing image based on the target student model, and obtain the target detection result corresponding to the target remote sensing image.
[0013] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:
[0014] One or more processors;
[0015] Memory, used to store one or more programs;
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the lightweight remote sensing image target detection method based on knowledge distillation as provided in any embodiment of the present invention.
[0017] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the lightweight remote sensing image target detection method based on knowledge distillation as provided in any embodiment of the present invention.
[0018] Fifthly, embodiments of the present invention provide a computer program product, including a computer program that, when executed by a processor, implements the lightweight remote sensing image target detection method based on knowledge distillation as provided in any embodiment of the present invention.
[0019] The technical solution of this invention involves acquiring a training sample set and a pre-built teacher-student model architecture. The teacher-student model architecture includes a pre-trained target teacher model and a preset student model to be trained. The preset student model is a lightweight model compared to the target teacher model. Based on the training sample set and the target teacher model, the preset student model is trained using knowledge distillation to obtain a trained target student model. This utilizes knowledge distillation technology to solve the problem of excessively large scale in existing remote sensing image target detection models, resulting in a lightweight remote sensing image target detection model, i.e., the target student model. This allows the target student model to be deployed on resource-constrained devices such as embedded devices, edge devices, and satellites. In response to a distributed target remote sensing image, target detection is performed on the target remote sensing image based on the target student model, obtaining the target detection result corresponding to the target remote sensing image. This enables accurate and convenient target detection of the acquired target remote sensing image on resource-constrained devices where the lightweight model is deployed, improving the accuracy and timeliness of target detection.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a lightweight remote sensing image target detection method based on knowledge distillation provided in Embodiment 1 of the present invention;
[0023] Figure 2 This is a flowchart of a lightweight remote sensing image target detection method based on knowledge distillation provided in Embodiment 2 of the present invention;
[0024] Figure 3 This is an example diagram illustrating feature distillation of a student model and a teacher model according to Embodiment 2 of the present invention;
[0025] Figure 4 This is an example diagram illustrating the calculation of characteristic distillation loss per layer according to Embodiment 2 of the present invention;
[0026] Figure 5 This is a schematic diagram of the structure of a lightweight remote sensing image target detection device based on knowledge distillation provided in Embodiment 3 of the present invention;
[0027] Figure 6 This is a schematic diagram of the structure of an electronic device that implements the lightweight remote sensing image target detection method based on knowledge distillation according to embodiments of the present invention. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] Example 1
[0031] Figure 1This document provides a flowchart of a lightweight remote sensing image target detection method based on knowledge distillation, as described in Embodiment 1 of the present invention. This embodiment is applicable to situations where the model size is reduced, particularly for models that need to be deployed on resource-constrained devices. This method can be executed by a lightweight remote sensing image target detection device based on knowledge distillation, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0032] S110. Obtain the training sample set and the pre-built teacher-student model architecture; the teacher-student model architecture includes: a pre-trained target teacher model and a preset student model to be trained; the preset student model is a lightweight model compared to the target teacher model.
[0033] In this embodiment of the disclosure, the training sample set may refer to the set of remote sensing images used to train the preset teacher model. The preset teacher model and the preset student model use the same training sample set. Each sample in the training sample set has a corresponding label for supervising model training. The preset teacher model may refer to a pre-built teacher model to be trained. The preset student model may refer to a pre-built student model to be trained. Conversely, the target teacher model is the teacher model obtained by training the preset teacher model using the training sample set.
[0034] The teacher-student model architecture comprises a pre-trained target teacher model and pre-set student models to be trained. The target teacher model employs a high-performance deep learning network structure, such as a deep convolutional neural network, to extract rich feature information from remote sensing images, achieving high-precision target detection. The pre-set student models are pre-built lightweight models, designed based on a miniaturized network structure to reduce model parameters and computational load. Lightweight models reduce model size and computational requirements while maintaining high accuracy and performance by optimizing model structure, parameter count, and computational complexity. Lightweight models enable efficient inference on resource-constrained devices (such as mobile devices, edge computing terminals, and IoT devices), reducing storage, computation, and energy costs.
[0035] For example, both the teacher and student models use YOLOv5 as the training network. The teacher model selects a larger YOLOv5m model with more parameters to ensure its accuracy, thus better guiding the training of the student model. The student model, on the other hand, selects a lighter YOLOv5s model and trains it under the guidance of the teacher model, aiming to improve its accuracy with fewer parameters and a smaller model size, thus approximating the teacher model and achieving model compression.
[0036] Specifically, various publicly available remote sensing image target detection datasets, such as the DOTA dataset, OpenSARShip dataset, and HRSC2016 dataset, were acquired online and normalized to obtain the training sample set. A pre-constructed teacher-student model architecture was obtained. YOLOv5m was pre-selected as the teacher model and pre-trained using the training sample set. YOLOv5s was pre-selected as the lightweight student model, and the pre-trained target teacher model was loaded when the pre-trained student model began training, thus obtaining the pre-constructed teacher-student model architecture. Except for the knowledge distillation process, the structure of the teacher and student models remained unchanged.
[0037] As an optional implementation of this disclosure, obtaining a training sample set may specifically include: obtaining a remote sensing image target detection dataset from a preset channel; and processing the remote sensing image target detection dataset based on a preset image preprocessing method to obtain a preprocessed training sample set.
[0038] In this embodiment of the disclosure, the preset channel can refer to a pre-selected image acquisition channel with stable remote sensing image quality. A remote sensing image is a visualized image formed by processing electromagnetic wave information reflected or radiated from the Earth or other celestial bodies by non-contact sensors (such as satellites, aircraft, drones, ground platforms, etc.) from a distance. Remote sensing images can be used to acquire information about the Earth's surface and atmosphere, and have characteristics such as wide coverage, periodicity, and objectivity. The remote sensing image target detection dataset can refer to a dataset containing target detection labels corresponding to the remote sensing images. The preset image preprocessing method can refer to pre-setting image standardization and / or data augmentation methods.
[0039] Specifically, various publicly available remote sensing image target detection datasets are acquired online, such as the DOTA dataset, the OpenSARShip dataset, and the HRSC2016 dataset. The image data in these datasets are then standardized in pixel size, for example, by using the `resize` function to convert all images to 640x640 pixels, to serve as input for both the student and teacher models. The standardized images are then subjected to a series of preprocessing operations, such as normalization and data augmentation, to obtain a preprocessed training sample set.
[0040] For example, image preprocessing operations can normalize the original remote sensing image, adjusting the pixel values to a suitable range, such as [0,1] or [-1,1], to accelerate model training convergence. Data augmentation techniques, including random flipping, rotation, scaling, and cropping, can also be used to expand the training dataset and improve the model's generalization ability.
[0041] For example, another special method in data augmentation is the Mosaic data augmentation method. Its core idea is to stitch four images together to form a new image. Generally, four images are randomly selected from the dataset, and then these four images are flipped, scaled, or cropped before being stitched together to form a new image as a training sample. This increases the diversity of the data, allowing the model to encounter more different scenes and target combinations during training, thus improving the model's generalization ability. In addition, adaptive anchor box calculation and adaptive image scaling are implemented at the input layer. This allows for the calculation of the optimal anchor box size based on different image types and reasonable scaling and padding according to the image's aspect ratio, avoiding the impact of image distortion on detection results and improving training efficiency and model adaptability. Using data augmentation techniques to expand the training dataset improves the model's generalization ability, enabling the model to maintain good detection performance in remote sensing images with complex backgrounds and targets of different scales.
[0042] S120. Based on the training sample set and the target teacher model, the preset student model is trained by knowledge distillation to obtain the trained target student model; the target student model is deployed on a resource-constrained device.
[0043] In this embodiment, the target student model can refer to a student model trained using a pre-defined knowledge distillation technique. The target student model can be used to detect the target category, category confidence, and target location information in an input image. Once the target student model is obtained, it can be deployed on a resource-constrained device. A resource-constrained device can refer to a hardware system with significant limitations in computing power, storage space, energy supply, or communication bandwidth. Such devices are typically deployed in edge computing, Internet of Things (IoT), and mobile terminal scenarios, and need to complete specific tasks under harsh conditions. For example, resource-constrained devices can be satellites and drones.
[0044] Specifically, a pre-defined student model is trained using a labeled training sample set. During training, the total distillation loss is minimized to allow the student model to learn from the target teacher model, gradually improving its detection performance until training is complete, resulting in a well-trained target student model. This model can then be deployed on resource-constrained devices. The lightweight target student model is suitable for resource-constrained scenarios, such as drones and mobile devices, enabling rapid, real-time detection of targets in remotely sensed images and showing broad application prospects.
[0045] For example, before training the pre-defined student model, a large, high-performance model, the target teacher model, needs to be trained using a training sample set. The target teacher model typically achieves or approaches the optimal performance for target detection in remotely sensed images. The trained target teacher model is then used to make predictions on the entire training sample set (or an unlabeled dataset) to obtain soft targets for each sample. Soft targets are the class probability distribution output by the target teacher model (usually the output vector processed by the Softmax function). To amplify the implicit knowledge contained in the soft targets, a temperature parameter is typically introduced into the Softmax function: softmax(z_i,T) = exp(z_i / T) / sum_j(exp(z_j / T)). When T = 1, the standard Softmax results in a relatively sharp probability distribution. When T > 1 (commonly used), increasing the temperature T makes the output probability distribution smoother. The relative sizes of classes with lower probabilities (representing classes similar to but different from the true labels) are amplified, making inter-class relationships more apparent. This is crucial for the pre-defined student model to learn the knowledge from the target teacher model.
[0046] Training a pre-defined student model requires using two types of supervision signals (soft targets and hard targets) to train a small student model. The soft target loss aims to make the output probability distribution of the pre-defined student model (using the same temperature T) as close as possible to the soft target probability distribution of the target teacher model. The difference between the two probability distributions is typically measured using the Kullback-Leibler divergence. The hard target loss aims to make the predictions of the pre-defined student model (at T=1) as close as possible to the true labels (hard targets). The difference is typically measured using cross-entropy loss. The total loss of the pre-defined student model is usually a weighted sum of the soft target loss and the hard target loss, as follows: Total Loss = β * KL Divergence (Student_Soft Target || Teacher_Soft Target) + (1-β) * Cross-Entropy (Student_Hard Target || True Label). β is a hyperparameter used to balance the importance of the two losses. In practice, sometimes the soft target loss (β close to 1) is used primarily initially, and the hard target loss is added later for fine-tuning.
[0047] S130. In response to the issued target remote sensing image, perform target detection on the target remote sensing image based on the target student model to obtain the target detection result corresponding to the target remote sensing image.
[0048] In this embodiment of the disclosure, the target remote sensing image may refer to the remote sensing image to be detected. The target detection result may refer to the target detection result for the image content in the target remote sensing image. For example, the target detection result may include at least one of the following: target category in the image, category confidence, and target location information.
[0049] Specifically, the trained lightweight target student model is used to perform the target detection task. After performing the same data preprocessing operations as in the training phase on the remote sensing image of the target to be detected, the data is input into the trained target student model, so that the target student model outputs the target detection result corresponding to the remote sensing image.
[0050] As an optional implementation of this disclosure, in response to the issued target remote sensing image, target detection is performed on the target remote sensing image based on the target student model to obtain the target detection result corresponding to the target remote sensing image. Specifically, it may include: in response to the issued target remote sensing image, processing the target remote sensing image based on a preset image preprocessing method to obtain a preprocessed standard remote sensing image; and performing target detection based on the target student model and the standard remote sensing image to obtain the target detection result corresponding to the target remote sensing image.
[0051] In this embodiment of the disclosure, a standard remote sensing image can refer to an image obtained through standardization processing. Processing the target remote sensing image using a preset image preprocessing method to obtain a standard remote sensing image is done to maintain consistency with the specifications of the model input during student model training.
[0052] Specifically, in response to the issued target remote sensing image, the target remote sensing image to be detected is preprocessed, uniformly scaled to 640x640 pixels, and subjected to corresponding data preprocessing operations such as normalization to obtain a preprocessed standard remote sensing image. The lightweight remote sensing image target detection model trained during the lightweight model training process is initialized and pre-trained parameters are loaded to obtain a target student model, preparing it for image detection. The standard remote sensing image is then input into the target student model for target detection, obtaining the target detection result corresponding to the target remote sensing image.
[0053] The technical solution of this invention involves acquiring a training sample set and a pre-built teacher-student model architecture. The teacher-student model architecture includes a pre-trained target teacher model and a pre-set student model to be trained. The pre-set student model is a lightweight model compared to the target teacher model. Based on the training sample set and the target teacher model, the pre-set student model is trained using knowledge distillation to obtain a trained target student model. This utilizes knowledge distillation technology to solve the problem of excessively large scale in existing remote sensing image target detection models, resulting in a lightweight remote sensing image target detection model, i.e., the target student model. This allows the target student model to be deployed on resource-constrained devices such as embedded devices, edge devices, and satellites. In response to the issued target remote sensing image, target detection is performed on the target remote sensing image based on the target student model, obtaining the target detection result corresponding to the target remote sensing image. This enables accurate and convenient target detection of the acquired target remote sensing image directly on resource-constrained devices where the lightweight model is deployed, improving the accuracy and timeliness of target detection.
[0054] Example 2
[0055] Figure 2 This is a flowchart of a lightweight remote sensing image target detection method based on knowledge distillation, provided in Embodiment 2 of the present invention. This embodiment describes in detail the process of training a preset student model using knowledge distillation, based on the previous embodiments. Explanations of terms that are the same as or corresponding to those in the above embodiments are not repeated here. Figure 2 As shown, the method includes:
[0056] S210. Obtain the training sample set and the pre-built teacher-student model architecture; the teacher-student model architecture includes: a pre-trained target teacher model and a preset student model to be trained; the preset student model is a lightweight model compared to the target teacher model.
[0057] S220. Based on the training sample set and the target teacher model, determine the first feature extraction result and the first detection result corresponding to the target teacher model.
[0058] In this embodiment of the disclosure, the first feature extraction result may refer to the sample features extracted by the target teacher model based on the input training samples. The first detection result may refer to the target detection result output by the target teacher model based on the input training samples.
[0059] Specifically, the training sample set is input into the target teacher model for feature extraction and target detection, and the first feature extraction result and the first detection result corresponding to the target teacher model are determined.
[0060] S230. Based on the training sample set and the preset student model, determine the second feature extraction result and the second detection result corresponding to the preset student model.
[0061] In this embodiment of the disclosure, the second feature extraction result may refer to the sample features extracted by the preset student model based on the input training samples. The second detection result may refer to the target detection result output by the preset student model based on the input training samples.
[0062] Specifically, the training sample set is input into the preset student model for feature extraction and target detection, and the second feature extraction result and the second detection result corresponding to the preset student model are determined.
[0063] S240. Based on the preset loss function, the first feature extraction result, the second feature extraction result, the first detection result, and the second detection result, determine the loss value corresponding to the preset student model.
[0064] In this embodiment of the disclosure, the preset loss function may refer to a pre-set distillation loss function. The preset loss function can be used to measure the difference between the feature extraction results of the target teacher model and the preset student model, as well as the difference between the detection results. The loss value may refer to the loss function value obtained by measuring the difference between the feature extraction results and the detection results using the preset loss function.
[0065] For example, based on the feature extraction results and detection results, the preset loss function can be divided into a first preset loss function targeting the differences between feature extraction results, and a second preset loss function targeting the differences between detection results. The preset loss functions are expressed as follows:
[0066] L all =L original +α·L distill
[0067] Among them, L all L represents the preset loss function. distill L represents the preset first loss function. original This represents the preset second loss function. α is a hyperparameter used to balance the preset first loss function and the preset second loss function.
[0068] Specifically, the first feature extraction result and the second feature extraction result are substituted into the preset first loss function, and the first detection result and the second detection result are substituted into the preset second loss function to determine the loss value corresponding to the preset student model.
[0069] As an optional implementation of this disclosure, the loss value corresponding to the preset student model is determined based on the preset loss function, the first feature extraction result, the second feature extraction result, the first detection result, and the second detection result. Specifically, this may include: determining the feature distillation loss value corresponding to the preset student model based on the preset first loss function, the first feature extraction result, and the second feature extraction result; determining the target detection loss value corresponding to the preset student model based on the preset second loss function, the first detection result, and the second detection result; and determining the loss value corresponding to the preset student model by performing a weighted summation of the feature distillation loss value and the target detection loss value.
[0070] In this embodiment of the disclosure, the feature distillation loss value may refer to the loss function value of a preset first loss function. The feature distillation loss value can be used to characterize the degree of difference between feature extraction results. The target detection loss value may refer to the loss function value of a preset second loss function. The target detection loss value can be used to characterize the degree of difference between detection results.
[0071] For example, based on the content contained in the object detection result, the preset first loss function can be divided into classification loss, bounding box regression loss, and confidence loss. The preset first loss function is expressed as follows:
[0072] L original =L cls +L box +L obj
[0073] Among them, L cls For classification loss, L box For bounding box regression loss, L obj The confidence loss is defined by these three parts, which measure the model's performance on different tasks and are combined by weight to form the final loss function.
[0074] Specifically, L cls The binary cross-entropy loss (BCE) is used to measure the accuracy of the model's predictions of the target class. In YOLOv5, for each detected target, the model needs to predict the probability that it belongs to each class. The classification loss is calculated by comparing these predicted probabilities with the true label, as shown in the following formula:
[0075] L BCE = -[ylog(p) + (1-y)log(1-p)]
[0076] Where y is the true label (0 or 1) and p is the probability predicted by the model.
[0077] Specifically, L boxThe CIoU (Complete Intersection over Union) loss is used to measure the difference between the model's predicted bounding box and the ground truth bounding box. Building upon IoU, it considers not only the overlap area between the predicted and ground truth boxes but also factors such as their center distance and aspect ratio, providing a more comprehensive evaluation of the bounding box regression quality and thus improving detection accuracy. It is a crucial part of object detection. The specific formula is as follows:
[0078]
[0079] Where IoU is the intersection-union ratio of the predicted bounding box and the ground truth bounding box, ρ 2 (b,b gt ) is the square of the Euclidean distance between the center points of the predicted box and the ground truth box, c is the diagonal length of the smallest bounding rectangle containing the predicted box and the ground truth box, α is the weight parameter, and v is a parameter that measures aspect ratio consistency.
[0080] Specifically, L obj The binary cross-entropy loss (BCE) is also used to measure the model's accuracy in determining whether a region contains a target. Each predicted bounding box has a confidence score, representing the probability that the box contains a target. When calculating the target confidence loss, positive samples (regions containing the target) and negative samples (regions not containing the target) are treated differently. For positive samples, the model's prediction confidence is expected to be close to 1; for negative samples, the model's prediction confidence is expected to be close to 0. Furthermore, a dynamic positive-negative sample balancing strategy can be used to improve training stability and efficiency.
[0081] In this embodiment of the disclosure, the feature map of the intermediate layer between the teacher model and the student model is selected, and the feature difference is measured by calculating the mean squared error (MSE) between the two feature maps to obtain the target detection loss value. Thus, the target detection loss value can prompt the student model to learn the feature representation of the target teacher model and enhance the representation ability of the student model.
[0082] As an optional implementation of this disclosure, both the first feature extraction result and the second feature extraction result are determined based on the mean square error between the feature maps extracted in the preset intermediate layer.
[0083] In this embodiment, the preset intermediate layer can refer to a specific layer that is pre-set. For example, layers 3, 5, and 7 of the target teacher model and the preset student model are pre-selected, i.e., the last activation function of the three C3 layers in the backbone of each model, and joint feature transfer is performed before them. The feature distillation loss value for each layer is calculated as follows: First, the features output by the preset student model are randomly masked using a masking algorithm. Then, the masked feature map is used to generate the features output by the teacher model through a generation module built on the Swing Transformer Block. The loss function is then calculated using the MSE distance function, thereby preserving the basic information of the shallow layers in the backbone to the greatest extent possible and improving the representation ability of the student model. This makes the intermediate layer features output by the student model converge with the features output by the teacher model, ultimately improving the target detection performance of the student model.
[0084] The formula for calculating the characteristic distillation loss value of each layer is as follows:
[0085] L distill (S,T)=MSE(F tea ,F stu )
[0086] Where S and T represent the student model and teacher model respectively, and F tea In the teacher model, the input image passes through the backbone network, where features are extracted and corresponding feature maps are obtained at this layer. stu The generated features are obtained by the student model after passing through a random mask and a generative module built on the basis of the Swing Transformer Block.
[0087] Specifically, Figure 3 An example diagram of feature distillation for a student model and a teacher model is provided. See also... Figure 3 Joint feature transfer is performed before the last activation function of the three C3 layers in the 3rd, 5th, and 7th layers of the student model and the teacher model, which are the backbone parts of the two models (Teacher Model Backbone and Student Model Backbone). The feature distillation loss of the three layers is calculated separately. Finally, all feature distillation loss values are added together and the sum is used as the overall feature distillation loss value.
[0088] Furthermore, Figure 4 An example diagram illustrating the calculation of feature distillation loss per layer is provided. See [link / reference] Figure 4In the Student model, the input image is first processed through each pre-processing layer in the backbone network for feature extraction, and feature maps F to be distilled are generated at layers 3, 5, and 7 respectively. stu Then, a random mask is used to process the obtained feature map (R). C×H×W The process involves masking a certain proportion of pixels to obtain the mask feature F. stu_mask .
[0089] Next, the masked features are normalized by layers and then doubled in number using a linear layer. They are then fed into two consecutive Swin Transformer blocks for feature generation, resulting in a new feature map F. stu_gen Finally, a 1x1 convolutional layer (Conv 1x1) is used to bring the doubled number of channels back to the normal number. The Swin Transformerblock is the core module of the Swin Transformer architecture, including: Window Multi-Head Self-Attention (W-MSA), Shifted Window Multi-Head Self-Attention (SW-MSA), Feedforward Network (FFN), and Residual Connections and LayerNorm. Window Multi-Head Self-Attention (W-MSA) can be used to compute self-attention within a local window, preserving spatial resolution and reducing computational cost. Shifted Window Multi-Head Self-Attention (SW-MSA) can be used to achieve cross-window information interaction through window shifting, enhancing global dependency modeling. The Feedforward Network (FFN) contains two layers of linear transformation and the GELU activation function, which can be used to perform non-linear transformations on features. The Residual Connections and LayerNorm adopt a ResNet-like residual structure, combined with LayerNorm to improve training stability.
[0090] In the teacher model, the input image also passes through the backbone network for feature extraction, resulting in the corresponding feature map F. tea Finally, the feature maps generated by the student model and those inferred by the teacher model are used to calculate MSE, which serves as the distillation loss function in knowledge distillation. This allows the student model to learn the knowledge of the target teacher model, thereby improving its own performance and representation ability.
[0091] S250. When the loss value or the number of training iterations corresponding to the preset student model meets the preset convergence condition, the trained target student model is obtained; the target student model is deployed on a resource-constrained device.
[0092] In this embodiment of the disclosure, the preset convergence condition can refer to a pre-set condition for the completion of training of a preset student model. The preset convergence condition can be that the loss value reaches a preset number of iteration cycles within a preset numerical range, or that the training times corresponding to the preset student model reach a preset maximum number of training times.
[0093] Specifically, when neither the loss value nor the number of training iterations corresponding to the preset student model meets the preset convergence condition, backpropagation is performed using the preset loss function to adjust the model parameters in the preset student model, and the adjusted preset student model is then used for training. Iterative training ends when either the loss value or the number of training iterations corresponding to the preset student model meets the preset convergence condition, resulting in a trained target student model that can be deployed on resource-constrained devices.
[0094] S260. In response to the issued target remote sensing image, perform target detection on the target remote sensing image based on the target student model to obtain the target detection result corresponding to the target remote sensing image.
[0095] The technical solution of this invention determines the first feature extraction result and the first detection result corresponding to the target teacher model based on a training sample set and the target teacher model; it also determines the second feature extraction result and the second detection result corresponding to the preset student model based on the training sample set and the preset student model; and it determines the loss value corresponding to the preset student model based on a preset loss function, the first feature extraction result, the second feature extraction result, the first detection result, and the second detection result. By constructing a teacher-student model architecture and designing a comprehensive loss function that includes the differences between feature extraction results and the differences between detection results, the knowledge of the target teacher model can be effectively transferred to the lightweight preset student model, thereby reducing the computational load and parameter count while ensuring high detection accuracy. When the loss value or the number of training iterations corresponding to the preset student model meets the preset convergence condition, the trained target student model is obtained.
[0096] The following are embodiments of the lightweight remote sensing image target detection device based on knowledge distillation provided in this invention. This device and the lightweight remote sensing image target detection method based on knowledge distillation in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the lightweight remote sensing image target detection device based on knowledge distillation, please refer to the embodiments of the lightweight remote sensing image target detection method based on knowledge distillation described above.
[0097] Example 3
[0098] Figure 5 This is a schematic diagram of a lightweight remote sensing image target detection device based on knowledge distillation, provided in Embodiment 3 of the present invention. Figure 5 As shown, the device includes: a data acquisition module 510, a target student model determination module 520, and a target detection result determination module 530.
[0099] The data acquisition module 510 is used to acquire a training sample set and a pre-built teacher-student model architecture. The teacher-student model architecture includes a pre-trained target teacher model and a preset student model to be trained. The preset student model is a lightweight model compared to the target teacher model. The target student model determination module 520 is used to perform knowledge distillation training on the preset student model based on the training sample set and the target teacher model to obtain a trained target student model. The target student model is deployed on a resource-constrained device. The target detection result determination module 530 is used to perform target detection on the target remote sensing image based on the target student model in response to the issued target remote sensing image to obtain the target detection result corresponding to the target remote sensing image.
[0100] The technical solution of this invention involves acquiring a training sample set and a pre-built teacher-student model architecture. The teacher-student model architecture includes a pre-trained target teacher model and a pre-set student model to be trained. The pre-set student model is a lightweight model compared to the target teacher model. Based on the training sample set and the target teacher model, the pre-set student model is trained using knowledge distillation to obtain a trained target student model. This utilizes knowledge distillation technology to solve the problem of excessively large scale in existing remote sensing image target detection models, resulting in a lightweight remote sensing image target detection model, i.e., the target student model. This allows the target student model to be deployed on resource-constrained devices such as embedded devices, edge devices, and satellites. In response to the issued target remote sensing image, target detection is performed on the target remote sensing image based on the target student model, obtaining the target detection result corresponding to the target remote sensing image. This enables accurate and convenient target detection of the acquired target remote sensing image directly on resource-constrained devices where the lightweight model is deployed, improving the accuracy and timeliness of target detection.
[0101] Based on the above technical solution, the data acquisition module 510 is specifically used to: acquire a remote sensing image target detection dataset from a preset channel; process the remote sensing image target detection dataset based on a preset image preprocessing method to obtain a preprocessed training sample set.
[0102] Based on the above technical solution, the target student model determination module 520 may include:
[0103] The first result determination submodule is used to determine the first feature extraction result and the first detection result corresponding to the target teacher model based on the training sample set and the target teacher model.
[0104] The second result determination submodule is used to determine the second feature extraction result and the second detection result corresponding to the preset student model based on the training sample set and the preset student model.
[0105] The loss value determination submodule is used to determine the loss value corresponding to the preset student model based on the preset loss function, the first feature extraction result, the second feature extraction result, the first detection result, and the second detection result.
[0106] The target student model determination submodule is used to obtain a trained target student model when the loss value or the number of training iterations corresponding to the preset student model meets the preset convergence condition.
[0107] Based on the above technical solution, the loss value determination submodule is specifically used for: determining the feature distillation loss value corresponding to the preset student model based on the preset first loss function, the first feature extraction result, and the second feature extraction result; determining the target detection loss value corresponding to the preset student model based on the preset second loss function, the first detection result, and the second detection result; and determining the loss value corresponding to the preset student model by performing a weighted summation of the feature distillation loss value and the target detection loss value.
[0108] Based on the above technical solution, both the first feature extraction result and the second feature extraction result are determined according to the mean square error between the feature maps extracted in the preset intermediate layer.
[0109] Based on the above technical solution, the target detection result determination module 530 is specifically used for: responding to the issued target remote sensing image, processing the target remote sensing image based on a preset image preprocessing method to obtain a preprocessed standard remote sensing image; and performing target detection based on the target student model and the standard remote sensing image to obtain the target detection result corresponding to the target remote sensing image.
[0110] The lightweight remote sensing image target detection device based on knowledge distillation provided in this invention can execute the lightweight remote sensing image target detection method based on knowledge distillation provided in any embodiment of this invention, and has the corresponding functional modules and beneficial effects for executing the lightweight remote sensing image target detection method based on knowledge distillation.
[0111] It is worth noting that in the above embodiments of lightweight remote sensing image target detection based on knowledge distillation, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0112] Example 4
[0113] Figure 6A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0114] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0115] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0116] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a lightweight remote sensing image target detection method based on knowledge distillation.
[0117] In some embodiments, the lightweight remote sensing image target detection method based on knowledge distillation can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the lightweight remote sensing image target detection method based on knowledge distillation described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the lightweight remote sensing image target detection method based on knowledge distillation by any other suitable means (e.g., by means of firmware).
[0118] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0119] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0120] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0121] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0122] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0123] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0124] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the lightweight remote sensing image target detection method based on knowledge distillation as provided in any embodiment of this application.
[0125] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider). This program product belongs to the same inventive concept as the lightweight remote sensing image target detection method based on knowledge distillation disclosed in the embodiments of this application, and therefore will not be described further here.
[0126] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0127] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A lightweight remote sensing image target detection method based on knowledge distillation, characterized in that, include: Obtain the training sample set and the pre-built teacher-student model architecture; The teacher-student model architecture includes: a pre-trained target teacher model and a preset student model to be trained; the preset student model is a lightweight model compared to the target teacher model. Based on the training sample set and the target teacher model, the preset student model is trained by knowledge distillation to obtain a trained target student model; the target student model is deployed on a resource-constrained device. In response to the issued target remote sensing image, target detection is performed on the target remote sensing image based on the target student model to obtain the target detection result corresponding to the target remote sensing image.
2. The method according to claim 1, characterized in that, The acquisition of the training sample set includes: Obtain a target detection dataset from remote sensing images sourced from a predefined channel; The remote sensing image target detection dataset is processed based on a preset image preprocessing method to obtain a preprocessed training sample set.
3. The method according to claim 1, characterized in that, The step of performing knowledge distillation training on the preset student model based on the training sample set and the target teacher model to obtain a trained target student model includes: Based on the training sample set and the target teacher model, determine the first feature extraction result and the first detection result corresponding to the target teacher model; Based on the training sample set and the preset student model, determine the second feature extraction result and the second detection result corresponding to the preset student model; Based on the preset loss function, the first feature extraction result, the second feature extraction result, the first detection result, and the second detection result, the loss value corresponding to the preset student model is determined; When the loss value or the number of training iterations corresponding to the preset student model meets the preset convergence condition, the trained target student model is obtained.
4. The method according to claim 3, characterized in that, The step of determining the loss value corresponding to the preset student model based on the preset loss function, the first feature extraction result, the second feature extraction result, the first detection result, and the second detection result includes: Based on the preset first loss function, the first feature extraction result, and the second feature extraction result, the feature distillation loss value corresponding to the preset student model is determined; Based on the preset second loss function, the first detection result, and the second detection result, the target detection loss value corresponding to the preset student model is determined; The loss value corresponding to the preset student model is determined by weighted summation of the feature distillation loss value and the target detection loss value.
5. The method according to claim 3, characterized in that, Both the first feature extraction result and the second feature extraction result are determined based on the mean square error between the feature maps extracted in the preset intermediate layer.
6. The method according to claim 1, characterized in that, In response to the transmitted target remote sensing image, target detection is performed on the target remote sensing image based on the target student model to obtain the target detection result corresponding to the target remote sensing image, including: In response to the target remote sensing image being sent, the target remote sensing image is processed based on a preset image preprocessing method to obtain a preprocessed standard remote sensing image; Target detection is performed based on the target student model and the standard remote sensing image to obtain the target detection result corresponding to the target remote sensing image.
7. A lightweight remote sensing image target detection device based on knowledge distillation, characterized in that, The device includes: The data acquisition module is used to acquire a training sample set and a pre-built teacher-student model architecture; the teacher-student model architecture includes: a pre-trained target teacher model and a preset student model to be trained; the preset student model is a lightweight model compared to the target teacher model; The target student model determination module is used to perform knowledge distillation training on the preset student model based on the training sample set and the target teacher model to obtain a trained target student model; the target student model is deployed on a resource-constrained device. The target detection result determination module is used to respond to the issued target remote sensing image, perform target detection on the target remote sensing image based on the target student model, and obtain the target detection result corresponding to the target remote sensing image.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the lightweight remote sensing image target detection method based on knowledge distillation as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the lightweight remote sensing image target detection method based on knowledge distillation as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the lightweight remote sensing image target detection method based on knowledge distillation as described in any one of claims 1-6.