A Remote Sensing Weak and Dense Target Detection Method Based on Feature Imitation Learning
By introducing feature imitation learning modules in the YOLO model, the problems of model optimization and real-time detection in weak remote sensing and dense object detection are solved, and more efficient object detection accuracy is achieved.
Patent Information
- Application Number
- CN202510323777.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The prior art has problems in weak remote sensing, intensive object detection, where models cannot be optimized end-to-end and cannot ensure real-time detection.
Using a method based on feature imitation learning, a feature imitation head is added to the head of the single-stage object detection model YOLO, and a task alignment learning technology is used to calculate sample quality metrics, filter high-quality and low-quality sample features, update the class center feature memory library online, build modeling and comparison losses, and train the detection model.
Without adding additional parameters or calculation overhead, the feature representation capability of weak and dense targets is improved, better detection accuracy is achieved, and real-time detection capability is ensured.
Smart Images

Figure CN119850934B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of optical remote sensing target detection, and particularly to a remote sensing weak and dense target detection method based on feature imitation learning. Background Art
[0002] Remote sensing weak and dense target detection, as a core technology in computer vision, aims to effectively find out the target of interest from images and accurately determine its category and location. Thanks to the effective balance achieved by the YOLO series of object detectors between detection performance and inference latency, it has received extensive attention in recent years, developed rapidly and achieved remarkable results. However, the research of YOLOs mainly focuses on general object detection, and there are still huge challenges in the detection of weak and dense remote sensing and aerial photography targets with unclear features, adjacent targets or overlapping targets. Different from general objects, the targets in remote sensing and aerial photography images are small, blurred and dense, which will bring unclear target features and mutual interference with each other.
[0003] Researchers have sought solutions from multiple perspectives. Wu et al. proposed SML, which mainly designed an imitation loss based on Faster R-CNN to force the feature representation of small-scale pedestrian targets to be close to that of large-scale pedestrian targets, and enhanced the representation of small-scale pedestrian targets by imitating the rich representation of large-scale pedestrian targets. Kim et al. proposed a pedestrian detection framework based on Faster R-CNN that imitates cue recall for detecting small-scale pedestrian targets. By constructing an offline large-scale pedestrian target feature repository, the features of small-scale pedestrian targets are brought closer to the features of large-scale pedestrian targets obtained from the repository. Yuan et al. proposed CRPN based on a two-stage detector to ensure obtaining higher-quality small target proposal regions, and at the same time added a feature imitation branch to the detector to promote the regional representation of instances with limited size and confusing the model in an imitation manner.
[0004] The above several methods have tried different improvements for weak targets, but they are all based on two-stage detectors and the way of constructing an offline memory bank, which hinders the model from being optimized in an end-to-end manner and cannot ensure real-time detection in practical applications. Summary of the Invention
[0005] Aiming at the above deficiencies in the prior art, the remote sensing weak and dense target detection method based on feature imitation learning provided by the present invention solves the problems that the existing methods hinder the model from being optimized in an end-to-end manner and cannot ensure real-time detection in practical applications.
[0006] In order to achieve the above invention purpose, the technical solution adopted by the present invention is: a remote sensing weak and dense target detection method based on feature imitation learning, including the following steps:
[0007] S1: Select the single-stage object detection model YOLO as the basic detection model, and add a feature imitation head to the head of the basic detection model;
[0008] S2: According to the output results of the classification head, regression head and feature imitation head in the head of the basic detection model, calculate the classification score and intersection over union between the target prediction result and the ground truth label, and use the task alignment learning technology to calculate the example quality metric of each example feature;
[0009] S3: Set the example high-quality metric threshold and example low-quality metric threshold, and based on the set thresholds, screen the example quality metrics of each example feature to obtain the corresponding high-quality example features and low-quality example features;
[0010] S4: Use the high-quality example features to update the class center feature memory bank online;
[0011] S5: Use the low-quality example features and the updated class center feature memory bank to construct an imitation contrast loss;
[0012] S6: Train the remote sensing weak and dense object detection model based on the imitation contrast loss to obtain the trained remote sensing weak and dense object detection model;
[0013] S7: Apply the trained remote sensing weak and dense object detection model to the actual scenario to complete the remote sensing weak and dense object detection based on feature imitation learning.
[0014] Further, the feature imitation head in S1 includes two convolutional modules and a non-linear mapping head connected in series in sequence;
[0015] The convolutional module includes a convolutional layer Conv2d, a normalization layer BatchNorm2d and an activation layer SiLU;
[0016] The non-linear mapping head includes a convolutional layer Conv2d, a normalization layer BatchNorm2d, an activation layer SiLU and a convolutional layer Conv2d.
[0017] Further, S2 includes the following sub-steps:
[0018] S21: According to the output results of the classification head, regression head and feature imitation head, obtain the bounding box information matrix , the class score matrix and the target feature embedding matrix , where, is the real number field, is the batch size, is the number of anchors in each image, is the maximum range of the anchor parameters, is the number of categories in the dataset, is the feature embedding dimension;
[0019] S22: Calculate the classification score between the target prediction result and the ground truth label based on the bounding box information matrix, the class score matrix, and the target feature embedding matrix and the intersection over union IoU , and the formula is:
[0020]
[0021]
[0022] where is the intersection, is the union, is the ground truth bounding box annotation;
[0023] S23: Based on the classification score and the intersection over union between the target prediction result and the ground truth label, and use the task alignment learning technique to calculate the example quality metric of each example feature. The formula is:
[0024]
[0025] where is the example quality metric, is the scale normalization function for training stability, is the classification confidence of the example, is the intersection over union of the example, and are hyperparameters.
[0026] Furthermore, in S4, the high-quality example features are used to update the class center feature memory bank online. The formula is:
[0027]
[0028]
[0029] where is the feature embedding of the -th class in the class center feature memory bank after update, is the feature embedding of the -th class in the class center feature memory bank before update, is the momentum factor, is the class center feature of the -th class calculated in the current batch, is the number of all high-quality examples, is the target feature embedding of the -th high-quality example, is the indicator function, For whether the class label of the th high-quality example is , if so, then it is 1, otherwise it is 0.
[0030] Furthermore, the S5 includes the following sub-steps:
[0031] S51: Calculate the similarity between the features of the low-quality examples and the class center features in the updated class center feature memory bank based on cosine similarity;
[0032] S52: Construct an imitation contrast loss based on the similarity.
[0033] Furthermore, the imitation contrast loss constructed in the S52 is:
[0034]
[0035] Among them, is the imitation contrast loss, is the classification loss in the basic detection model, is the regression loss in the basic detection model, is the weight parameter, is the instance-level loss, is the class-level loss;
[0036]
[0037] Among them, is the number of low-quality examples, is the cosine similarity function, is the feature embedding of the th low-quality example, and the category is , is the temperature parameter;
[0038]
[0039] Among them, is the number of categories, is the feature embedding of the th class in the class center feature memory bank before update, is the feature embedding of the th class in the class center feature memory bank after update.
[0040] The beneficial effects of the present invention are as follows: The present invention integrates a plug-and-play feature imitation learning module based on the single-stage object detector YOLO to improve feature learning, especially for weak and dense objects. The FIL paradigm proposed in this paper is based on contrastive learning and enhances the representation of weak and dense object features without adding extra parameters or computational overhead during the inference process. Compared with the original baseline detector, the improved detector of the present invention achieves better detection accuracy for weak and dense remote sensing objects with the same inference latency as the original baseline detector, and can be extended to any single-stage detector. Description of the Drawings
[0041] Figure 1 It is a schematic diagram of the overall framework of the method for detecting weak and dense remote sensing objects based on feature imitation learning of the present invention.
[0042] Figure 2 It is a flowchart of the method for detecting weak and dense remote sensing objects based on feature imitation learning of the present invention.
[0043] Figure 3 It is a schematic diagram of the detection results of the method for detecting weak and dense remote sensing objects based on feature imitation learning of the present invention. Detailed Embodiments
[0044] The present invention will be further described below with reference to the drawings and specific embodiments.
[0045] The present invention provides a method for detecting weak and dense remote sensing objects based on feature imitation learning, and the corresponding overall network structure is as Figure 1 shown, mainly including the construction and use of a feature imitation head, the selection of high / low-quality positive examples, the update of the class center memory bank, the construction of an imitation contrast loss, the comparison of detection performance, and visualization. Based on Figure 1 the overall network structure shown, the method mainly includes the following steps:
[0046] As Figure 2 shown, a method for detecting weak and dense remote sensing objects based on feature imitation learning includes the following steps:
[0047] S1: Select the single-stage object detection model YOLO as the basic detection model, and add a feature imitation head to the head of the basic detection model;
[0048] S2: According to the output results of the classification head, regression head, and feature imitation head at the head of the basic detection model, calculate the classification score and intersection over union between the target prediction result and the true label, and calculate the example quality metric of each example feature using the task alignment learning technique;
[0049] S3: Set the high-quality sample metric threshold and the low-quality sample metric threshold, and screen the sample quality metrics of each sample feature based on the set thresholds to obtain the corresponding high-quality sample features and low-quality sample features;
[0050] S4: Use the high-quality sample features to update the class center feature memory bank online;
[0051] S5: Use the low-quality sample features and the updated class center feature memory bank to construct an imitation contrast loss;
[0052] S6: Train the remote sensing weak and dense target detection model based on the imitation contrast loss to obtain the trained remote sensing weak and dense target detection model;
[0053] S7: Apply the trained remote sensing weak and dense target detection model to the actual scenario to complete the remote sensing weak and dense target detection based on feature imitation learning.
[0054] In this embodiment, two current optimal single-stage detectors, YOLOv10-N and YOLOv11-N, are selected as the basic detection models, and their weak and dense remote sensing target detections are improved.
[0055] To enhance the detector's ability to recognize weak and densely distributed targets and at the same time not increase additional computational or parameter overhead during the inference process, a Feature Imitation Head (FIH) is introduced at the head stage of the basic detection model. As Figure 1 shown in (a) below, the FIH works in parallel with the Regression Head (RH) responsible for object localization and the Classification Head (CH) responsible for classification. The FIH focuses on optimizing the target feature embedding to improve the detection performance. Its detailed architecture is as Figure 1 shown in (b) below, which is mainly composed of two convolutional modules and a non-linear mapping head connected in series. During the training stage, the FIH is jointly optimized with the RH, CH, and the rest of the network, enabling the backbone network and the neck to utilize the comprehensive supervision provided by Feature Imitation Learning (FIL), and the two convolutional modules ConvModule in the FIH and CH share weights, thus minimizing the parameter usage and computational complexity. It should be noted that during the inference stage, the FIH is discarded, and only the RH and CH are retained. This ensures that the proposed framework does not incur additional computational costs or parameter burdens compared to the baseline detector during the inference process.
[0056] The feature imitation head in the said S1 includes two convolutional modules and a non-linear mapping head connected in series in sequence;
[0057] The convolutional module includes a convolutional layer Conv2d, a normalization layer BatchNorm2d, and an activation layer SiLU;
[0058] The non - linear mapping head includes a convolutional layer Conv2d, a normalization layer BatchNorm2d, an activation layer SiLU, and a convolutional layer Conv2d.
[0059] To further optimize the high - quality and low - quality target feature embeddings, by processing the outputs of RH, CH, and FIH, three key matrices are obtained: the bounding box information matrix , the class score matrix , and the target feature embedding matrix . Based on the three output matrices and task - alignment learning, each sample is classified as a positive or negative sample according to the instance quality metric (IQM), which comprehensively evaluates the classification confidence and regression quality of the instance and is used for the evaluation of instance quality.
[0060] S2 includes the following sub - steps:
[0061] S21: According to the output results of the classification head, regression head, and feature imitation head, obtain the bounding box information matrix , the class score matrix , and the target feature embedding matrix , where is the real number field, is the batch size, is the number of anchors in each image, is the maximum range of anchor parameters, is the number of classes in the dataset, is the feature embedding dimension;
[0062] S22: Based on the bounding box information matrix, class score matrix, and target feature embedding matrix, calculate the classification score between the target prediction result and the ground truth label and the intersection - over - union IoU , and the formula is:
[0063]
[0064]
[0065] where is the intersection, is the union, is the ground truth bounding box annotation;
[0066] S23: Based on the classification score and intersection - over - union between the target prediction result and the ground truth label, and using the task - alignment learning technique, calculate the instance quality metric of each sample feature, and the formula is:
[0067]
[0068] where is the sample quality index, is the scale normalization function for training stability, is the classification confidence of the sample, is the intersection over union of the sample, and are hyperparameters. A higher value indicates that the anchor sample has a higher classification score and more accurate location prediction. For each instance, the m anchors with the highest alignment value are selected as positive samples. By setting appropriate thresholds and , high-quality and low-quality positive samples can be effectively distinguished. This supports the dynamic update of class samples and feature imitation learning, thus improving the overall quality of feature representation.
[0069] In S4, the class center feature memory bank is updated online using high-quality sample features, and the formula is:
[0070]
[0071]
[0072] where, is the feature embedding of the -th class in the class center feature memory bank after update, is the feature embedding of the -th class in the class center feature memory bank before update, is the momentum factor, is the class center feature of the -th class calculated in the current batch, is the number of all high-quality samples, is the target feature embedding of the -th high-quality sample, is the indicator function, is whether the class label of the -th high-quality sample is , if so, then is 1, otherwise is 0.
[0073] Figure 1 In (a) shows the structure of the class center feature memory bank (MBCE). The class examples of the current dataset are defined as , where corresponds to the class center feature of the -th class, and the feature dimension is . Initially, all class examples are initialized to zero. MBCE operates in the manner of an online updatable memory bank, aiming to preserve the learned knowledge and guide the model's feature learning. Specifically, the alignment metric Above the threshold Samples are selected as high-quality positive samples (HQP). The target feature embeddings and class labels corresponding to each HQP sample are and . In the memory bank, class exemplars are updated through a momentum-based mechanism. For the th class, the current class center is first calculated using the high-quality positive samples in the current batch. Then, the corresponding class exemplar in the memory bank is updated using this class center, thus ensuring adaptive and incremental improvement of the class-level representation.
[0074] The S5 includes the following sub-steps:
[0075] S51: Calculate the similarity between the low-quality example features and the class center features in the updated class center feature memory bank based on cosine similarity;
[0076] S52: Construct an imitation contrast loss based on the similarity.
[0077] The imitation contrast loss constructed in S52 is:
[0078]
[0079] where, is the imitation contrast loss, is the classification loss in the basic detection model, is the regression loss in the basic detection model, is the weight parameter, is the instance-level loss, is the class-level loss;
[0080]
[0081] where, is the number of low-quality examples, is the cosine similarity function, is the th low-quality example feature embedding, and the class is , is the temperature parameter;
[0082]
[0083] where, is the number of classes, is the feature embedding of the th class in the class center feature memory bank before update, is the feature embedding of the th class in the class center feature memory bank after update.
[0084] Based on the low-quality threshold Screen the embedding features of low-quality positive example (LQP) samples . Then calculate the features of low-quality samples based on cosine similarity and the updated class center features . Calculate the similarity between them, and construct a supervised contrastive loss based on the obtained similarity. The purpose of this process is to effectively pull the low-quality feature embeddings closer to their respective corresponding class center features while pushing them away from other class center features. To achieve this, an instance-level loss is introduced.
[0085] To address mode collapse and improve the efficiency of feature imitation learning, a class-level loss is proposed, mainly using the class center sample features before and after update. This loss ensures the consistency and diversity of class sample representations. This loss encourages consistency between class exemplars before and after update while maintaining inter-class separability. The bi-directional formula ensures that the update does not overly bias the representation, effectively alleviating mode collapse and maintaining the diversity and robustness of class examples.
[0086] By carefully designing these losses, this method not only achieves feature imitation learning but also effectively alleviates mode collapse, thereby enhancing the feature representations of weak and dense targets and improving the detection performance of the detection algorithm.
[0087] In an embodiment of the present invention, to prove the effectiveness of the method of the present invention for detecting weak and dense remote sensing targets, DIOR, AI-TOD-v2, and VisDrone2019 are used to train and validate the model. The specific description information of the dataset is shown in Table 1. Starting from the third column in the table, the numbers on the left side of each column represent the number of images in different datasets, and the numbers on the right side of each column represent the number of instances in different datasets. And the performance comparison of the detection algorithm on different datasets is given, as shown in Table 2. The detection metrics corresponding to the present invention on each data are shown with a darker background color. To understand the effectiveness of the detection algorithm based on feature imitation learning from a more intuitive perspective, some representative image data are selected for visualization of the detection results, as shown in Figure 3 . The specific performance metrics and visualization results further verify the effectiveness of the method proposed in the present invention for detecting weak and dense remote sensing targets.
[0088] Table 1 Specific description information of the dataset
[0089]
[0090] Table 2 Performance comparison of the detection algorithm on different datasets
[0091]
[0092] The present invention first attempts to use feature imitation learning in a single-stage object detection framework to solve the problem of dense object detection. Moreover, the proposed Feature Imitation Learning (FIL) is a plug-and-play module that can be seamlessly integrated into any single-stage detector, and can improve the detection accuracy of weak and dense remote sensing targets without introducing additional parameters and computational complexity during the inference process.
[0093] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention according to these technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the invention.
Claims
1. A remote sensing weak and dense target detection method based on feature imitation learning, characterized in that: The following steps are involved: S1: Select the single-stage target detection model YOLO as the basic detection model, and add a feature imitation head to the head of the basic detection model; The feature imitation head in S1 includes two convolution modules and a nonlinear mapping head connected in series; The convolution module includes a convolution layer Conv2d, a normalization layer BatchNorm2d and an activation layer SiLU; The nonlinear mapping head includes a convolution layer Conv2d, a normalization layer BatchNorm2d, an activation layer SiLU and a convolution layer Conv2d; S2: Based on the output results of the classification head, regression head and feature imitation head of the basic detection model, the classification score and intersection-over-union ratio between the target prediction result and the true label are calculated, and the sample quality metric of each sample feature is calculated using task alignment learning technology; S3: setting a sample high-quality metric threshold and a sample low-quality metric threshold, and filtering the sample quality metrics of each sample feature based on the set threshold to obtain corresponding high-quality sample features and low-quality sample features; S4: Use high-quality sample features to update the class-centric feature memory library online; In S4, high-quality sample features are used to update the class center feature memory library online, and the formula is: in, After the update The feature embedding of the class in the class-centric feature memory, Before update The feature embedding of the class in the class-centric feature memory, is the momentum factor, is the first Class center features, is the number of all high-quality examples, For the The target feature embedding of high-quality examples, is the indicator function, For the Is the class label of the high-quality examples , if yes, then is 1, otherwise is 0; S5: Use low-quality sample features and the updated class center feature memory library to construct imitation contrast loss; The S5 comprises the following sub-steps: S51: Calculate the similarity between the low-quality sample features and the class center features of the updated class center feature memory library based on cosine similarity; S52: constructing imitation contrast loss based on similarity; The simulated contrast loss constructed in S52 is: in, To mimic the contrast loss, is the classification loss in the base detection model, is the regression loss in the basic detection model, is the weight parameter, is the instance-level loss, is the class-level loss; in, is the number of low-quality samples, is the cosine similarity function, For the Low-quality sample feature embedding, and the category is , is the temperature parameter; in, is the number of categories, Before update The feature embedding of the class in the class-centric feature memory, After the update Feature embedding of classes in class-centric feature memory bank; S6: Train the remote sensing weak and dense target detection model based on the imitation contrast loss to obtain the trained remote sensing weak and dense target detection model; S7: Apply the trained remote sensing weak and dense target detection model to actual scenarios to complete remote sensing weak and dense target detection based on feature imitation learning.
2. According to claim 1, a remote sensing weak and dense target detection method based on feature imitation learning is characterized in that: The S2 includes the following sub-steps: S21: Obtain the bounding box information matrix based on the output results of the classification head, regression head, and feature imitation head , category score matrix and the target feature embedding matrix ,in, is the field of real numbers, is the batch size, is the number of anchors in each image, is the maximum range of the anchor parameter, is the number of categories in the data set, is the feature embedding dimension; S22: Calculate the classification score between the target prediction result and the true label based on the bounding box information matrix, the category score matrix and the target feature embedding matrix Compared with intersection IoU , the formula is: in, For the intersection, For the union, Annotate the ground-truth bounding box; S23: Based on the classification score and intersection-over-union ratio between the target prediction result and the true label, the sample quality metric of each sample feature is calculated using task alignment learning technology. The formula is: in, is the sample quality indicator, is the scale normalization function used for training stability, is the classification confidence of the sample, is the intersection-over-union ratio of the samples, and is a hyperparameter.
Citation Information
Patent Citations
General target detection method based on remote sensing image and medium
CN118230183A
Ground disaster remote sensing detection method and device based on YOLOv8 model and medium
CN118570663A