Small-Sample Aerial Image Small Target Detection Method Based on Unbiased Proposal Box Filtering
By using confidence-IoU collaborative suggestion box filtering processing and adding box regression loss function of small target constraints in the RPN stage of the small sample object detection method, the model's error deletion of new category suggestions box is solved, and the detection performance is improved.
Patent Information
- Application Number
- CN202211578576.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-09
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-12-09
AI Technical Summary
The existing small sample object detection method is prone to accidentally delete the suggestion box of new categories in the RPN stage, resulting in the model's detection accuracy of new categories reduced, especially when small targets are dense in aerial images, the task is more difficult.
The recommended box filtering process of confidence-IoU collaboration is adopted, and the box regression loss function of small target constraints is added in the RPN stage, and only the classification branches and box regression branches are adjusted in network training to fix the parameters of other network layers.
It effectively solves the bias of the small sample object detection model to the basic categories, prevents the suggestion box of the new category from being accidentally deleted, and improves the model's detection performance for new categories and small objects.
Smart Images

Figure CN115775360B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image target detection, and in particular relates to a small target detection method for small sample aerial images based on unbiased suggestion frame filtering. Background Art
[0002] In recent years, with the improvement of computer performance, deep learning has achieved remarkable success in aerial image target detection tasks with its powerful data processing capabilities. However, deep learning algorithms often require a large number of labeled samples to train the model, and the labeling of samples will also consume a lot of manpower and material resources. In addition, the trained model can only detect the existing training categories. For new target categories, it is necessary to collect a large amount of new labeled data and retrain the model on the integrated new data set, which is time-consuming and laborious. It is worth noting that when the new category is a scarce category (such as a new reconnaissance aircraft), the number of labeled samples that can be obtained is extremely small, which cannot meet the needs of model retraining. Therefore, it is crucial to use the existing training data and a small amount of supervision information of the new target category to improve the model's generalization ability for new categories while avoiding additional model training.
[0003] Few-shot object detection (FSOD) is a research direction that has received much attention in the field of aerial images. It can effectively alleviate the high dependence on the amount of labeled data in deep learning and the problem of easy misjudgment of new classes. The purpose of few-shot object detection is to identify and locate new class target instances by using a small amount of labeled information. Currently, most few-shot object detection methods use Faster R-CNN as the basic detection framework. They train the detection model on the base classes with sufficient labeled training data using a fine-tuning-based learning method, and then fine-tune the model on the new classes with a small amount of labeled information. In the Faster R-CNN detection framework, the proposals generated by the Region Proposal Network (RPN) are screened according to the foreground-background scores output by the RPN and the non-maximum suppression (NMS) algorithm. In the base training stage, the model only regards the base classes as the foreground and the others as the background. Through screening, more proposals belonging to the base classes will be obtained. However, in the fine-tuning stage, although the data of the novel classes participate in the model training, due to the limited amount of labeled data of the novel classes, the bias of the model still exists. The foreground-background scores output by the RPN are biased towards the base classes. That is, the model will still inevitably regard some target instances of the novel classes as the background. The proportion of the novel classes in the proposals finally generated by the RPN is very small, resulting in a reduction in the features of the novel class targets extracted in the subsequent proposal processing stage (i.e., the R-CNN stage), which greatly affects the detection accuracy of the model for the novel classes. Therefore, in the RPN stage of the few-shot object detection method, it is particularly crucial to avoid misdeleting the proposals that originally belong to the novel classes. In addition, due to the high-altitude shooting, there are a large number of small targets in the aerial images, which further exacerbates the task difficulty of few-shot object detection. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the present invention provides a few-shot small target detection method for aerial images based on unbiased proposal filtering. First, use the aerial sample data of the base classes to conduct basic training on the Faster R-CNN network. During training, add proposal filtering processing based on confidence-IoU collaboration in the RPN stage, and adopt a box regression loss function with an additional small target constraint term in the R-CNN stage. Then, use the aerial sample data of the base classes and the aerial sample data of the new classes to retrain the network. During training, only adjust the parameters of the classification branch and the box regression branch in the network, and fix the parameters of the remaining network layers. Finally, input the aerial image into the trained network to obtain the target detection result. The present invention can effectively solve the bias of the few-shot object detection model towards the base classes and improve the performance of small target detection in aerial images.
[0005] A small - sample aviation image small - target detection method based on unbiased proposal box filtering, characterized in that the steps are as follows:
[0006] Step S1, construct a training data set: Randomly divide all categories of the aviation image data set into a base category C b and a new category C n , and the category spaces of the base category and the new category do not intersect; among them, the number of labeled samples in each category of the base category is greater than or equal to 500, forming a sub - data set D b , and the number of labeled samples in each category of the new category does not exceed 20, forming a sub - data set D n ;
[0007] Step S2, basic training: Use the pre - trained model provided by the official Faster R - CNN network to train on the sub - data set D b , to obtain the initial network parameters. During training, keep the Faster R - CNN network structure unchanged, add a confidence - IoU collaborative proposal box filtering process in its RPN stage, and adopt a box regression loss function with a small - target constraint term added in the R - CNN stage;
[0008] The confidence - IoU collaborative proposal box filtering process includes the following steps:
[0009] Step a, sort the set P of proposal boxes output by the RPN in descending order according to their foreground - background scores, and select the top proposal boxes to form a set Among them,
[0010] Step b, calculate the intersection - over - union (IoU) scores of the remaining proposal boxes in the proposal box set P except the set with the ground - truth boxes, and sort them in descending order according to the IoU scores, and select the top proposal boxes to form a set Among them,
[0011] Step c, use the NMS algorithm to filter the proposal box set , and the filtered proposal boxes form a set Set the number of filtered boxes to
[0012] Step d, calculate the intersection - over - union (IoU) scores of the proposal boxes in the proposal box set with the ground - truth boxes, and sort them in descending order according to the IoU scores, and select the top proposal boxes to form a set
[0013] Step e, merge the set of proposal boxes and use the merged set as the set of proposal boxes finally generated in the RPN stage;
[0014] The box regression loss function with the addition of a small object constraint term in the R-CNN stage refers to modifying the original box regression loss function in the R-CNN stage and adding a small object constraint term L tiny , specifically:
[0015]
[0016]
[0017]
[0018] where represents the box regression loss function after adding the small object constraint term, and L GIoU represents the original box regression loss function, and φ(S g ) is a function for normalizing the area of the ground truth box; Sg represents the area size of the ground truth box; α is a hyperparameter, set α = 3; d represents the distance between the center points of the predicted box and the ground truth box; u top represents the distance between the top edges of the predicted box and the ground truth box, and u bottom represents the distance between the bottom edges of the predicted box and the ground truth box, and u left represents the distance between the leftmost edges of the predicted box and the ground truth box, and u right represents the distance between the rightmost edges of the predicted box and the ground truth box, μ represents a hyperparameter, set μ = 2; h g represents the height of the ground truth box; w g represents the width of the ground truth box; min(S g ) represents selecting the one with the smallest area from all ground truth boxes; max(S g ) represents selecting the one with the largest area from all ground truth boxes;
[0019] Step S3, network parameter adjustment: Use the method in Step S2 to perform network training again on the sub-dataset D b and the sub-dataset D n to obtain a trained network. During training, only the parameters of the classification branch and the box regression branch in the network are adjusted, and the parameters of the remaining network layers are fixed;
[0020] Step S4, object detection: Input the aviation image dataset to be processed into the trained network, and output its object detection results.
[0021] The beneficial effects of the present invention are as follows: Due to the class unbiasedness based on the Intersection over Union (IoU) score, confidence-IoU collaborative proposal box filtering processing is performed, which can effectively solve the bias of the small-sample object detection model towards the base classes, prevent proposal boxes belonging to new classes from being mistakenly deleted, and reserve more target proposal boxes of new classes for the subsequent R-CNN stage learning of the network; Since a small object constraint term is added to the box regression loss function, and this constraint term simultaneously considers the distance between the center points of the predicted box and the ground truth box, the size of the ground truth box, and the distances between the four sides of the predicted box and the ground truth box, the detection performance of the model for small objects can be improved; The present invention is simple to implement and does not change the network structure. While improving the object detection effect of the aerial image dataset, it can also be extended and applied to other two-stage small-sample object detection models. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 FIG. is a flow chart of a small-sample aerial image small object detection method based on unbiased proposal box filtering according to the present invention;
[0023] Figure 2 FIG. is a schematic diagram of the distance variables involved in the small object constraint loss term;
[0024] Figure 3 FIG. is a schematic diagram of the number of proposal boxes of different classes generated in the RPN stage using the benchmark algorithm on the DIOR dataset;
[0025] In the figure, ae - airplane, at - airport, bf - baseball field, bl - basketball court, be - bridge, cy - chimney, dm - dam, es - highway service area, et - highway toll station, hr - seaport, go - golf course, gd - track and field stadium, os - overpass, sp - ship, sm - stadium, sk - storage tank, tt - tennis court, tn - railway station, ve - vehicle, wl - windmill;
[0026] Figure 4 FIG. is a schematic diagram of the number of proposal boxes of different classes generated in the RPN stage using the method of the present invention on the DIOR dataset;
[0027] Figure 5 FIG. is an image of the detection results of different scene images in the DIOR dataset using the method of the present invention;
[0028] Figure 6 FIG. is an image of the detection results of different scene images in the AI-TOD dataset using the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The present invention will be further described below in conjunction with the drawings and embodiments. The present invention includes but is not limited to the following embodiments.
[0030] As Figure 1As shown in the figure, the present invention provides a small sample aviation image small target detection method based on unbiased proposal box filtering, and its specific implementation process is as follows:
[0031] Step S1, construct a training data set: randomly divide all categories of the aviation image data set into a basic category C b and a new category C n , and the category spaces of the basic category and the new category do not intersect, that is Among them, the number of labeled training samples in the basic category is sufficient to form a sub-data set D b , where the number of labeled training samples for each category is at least 500; each category in the new category has only a small number of labeled samples, forming a sub-data set D n , where the number of labeled training samples for each category does not exceed 20.
[0032] Step S2, basic training: The basic training of the model is carried out on the training data containing only the basic category, and its purpose is to obtain initial parameters of a good small sample target detection model. The purpose of the small sample target detection task is to quickly detect new category targets with only a small number of labeled samples. Directly using a small number of training data of the new category to train the model will cause the model to overfit. Basic training of the model is very important for small sample target detection. First, use a large number of basic categories with sufficient labeled samples to train the model to train a model with certain detection capabilities, and then fine-tune some network layers of the model to avoid the risk of model overfitting. Therefore, the present invention follows this training method of first basic training and then partial fine-tuning, that is, first use the pre-trained model provided by the official Faster R-CNN network to train on the sub-data set D b to obtain the initial network parameters.
[0033] The Faster R-CNN detection network model was proposed by Ren et al. in their 2015 work "Faster r-cnn: Towards real-time object detection with region proposal networks[C]Advances in Neural Information Processing Systems, 2015, 28". It mainly includes a backbone network (such as ResNet101), FPN (Feature Pyramid Networks), RPN, and R-CNN. Among them, ResNet101 was proposed by He et al. in their 2016 work "Deep residual learning for image recognition[C]Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770-778"; the FPN network framework was proposed by Lin et al. in their 2017 work "Feature pyramid networks for object detection[C]Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 2117-2125". The backbone network extracts the features of training samples, and RPN generates a set of proposal boxes related to the basic categories. The sample features and the set of proposal boxes are input into the R-CNN network together to obtain classification scores and localization results.
[0034] The foreground-background classification loss function in the RPN stage and the classification loss function in the R-CNN stage both adopt the Focal loss function proposed by Lin et al. in their 2017 work "Focal loss for dense object detection[C]Proceedings of the IEEE international conference on computer vision. 2017:2980-2988"; the bounding box regression loss function in the RPN stage adopts the SmoothL1 Loss function proposed by Ross Girshick in his 2015 work "Fast r-cnn[C]Proceedings of the IEEE international conference on computer vision. 2015:1440-1448"; the bounding box regression loss function in the R-CNN stage adopts the GIoU Loss function proposed by Rezatofighi et al. in their 2019 work "Generalized intersection over union: A metric and a loss for bounding box regression[C]Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2019:658-666". The network parameters are optimized using the stochastic gradient descent algorithm, with an initial learning rate of 0.002 and a maximum number of iterative training times of 110,000.
[0035] Since the basic training is carried out on the sub-dataset D that only contains basic categories b it is easy to cause the proposed bounding boxes generated by the RPN to be more inclined to basic categories, which is not conducive to aerial image target detection. Therefore, the present invention adopts a confidence-IoU collaborative proposed bounding box filtering method. At the same time, a small target constraint term is added to the bounding box regression loss function in the R-CNN stage to improve the detection performance of the model for small targets while reducing the redundancy of the proposed bounding boxes and alleviating the problem of model bias. That is, the entire Faster R-CNN network structure remains unchanged, and the following proposed bounding box filtering process based on confidence-IoU collaboration is added in the RPN stage:
[0036] (1) Sort the set P of proposed bounding boxes output by the RPN in descending order according to their foreground-background scores, and select the top proposed bounding boxes with higher scores to form the set where, can be set according to experience
[0037] (2) Calculate the intersection-over-union (IoU) scores between the remaining proposal boxes in the set \(P\) of proposal boxes and the ground-truth boxes, and sort these scores in descending order. Select the top proposal boxes with the highest scores to form a set where, According to experience, it can be set that In this way, a total of 3000 proposal boxes are selected through the foreground-background scores and the IoU scores.
[0038] (3) Use the NMS algorithm to filter the set of proposal boxes and set the number of filtered boxes to that is, 500, and finally obtain the set of proposal boxes The NMS algorithm adopted in the present invention is the same as the NMS algorithm adopted in the work of Ren et al. in 2015, Faster r-cnn: Towards real-time object detection with region proposal networks [C] Advances in Neural Information Processing Systems, 2015, 28. The parameter settings use the default values, that is, the NMS threshold is 0.7.
[0039] (4) Filter the set of proposal boxes According to the IoU scores between the proposal boxes in the set of proposal boxes and the ground-truth boxes, sort these scores in descending order, and select the top proposal boxes, that is, the top 1500 proposal boxes, and finally obtain the set of proposal boxes
[0040] (5) Merge the set of proposal boxes and and use the merged set as the set \(P\) of proposal boxes finally generated in the RPN stage final .
[0041] Input the set \(P\) of proposal boxes obtained from the above process final and the region of interest (RoI) features extracted by the backbone network into the classification branch and the target box regression branch of R-CNN, and calculate the classification loss and the box regression loss in the R-CNN stage. In the box regression loss function in the R-CNN stage of the present invention, a small object constraint term \(L\) tiny is added, that is:
[0042]
[0043]
[0044]
[0045] Among them, represents the bounding box regression loss function after adding the small target constraint term, L GIoU represents the original bounding box regression loss function, L GIoU Adopt the GIoU Loss function, that is:
[0046]
[0047]
[0048] Among them, IoU represents the ratio of the overlapping area of the predicted box and the ground truth box to the total area of these two boxes, that is, the ratio of the intersection area to the union area. B represents the predicted box, G represents the ground truth box, and C represents the area of the smallest external rectangle enclosing the two boxes B and G.
[0049] In formulas (4)-(6), φ(S g ) represents the normalization function, and the purpose is to normalize all ground truth boxes to the range of 0 to 1 according to the area size; α is a hyperparameter, and in the present invention, α = 3 is set; S g represents the area size of the ground truth box; d represents the distance between the center points of the predicted box and the ground truth box; u top represents the distance between the top edges of the predicted box and the ground truth box, u bottom represents the distance between the bottom edges of the predicted box and the ground truth box, u left represents the distance between the left edges of the predicted box and the ground truth box, u right represents the distance between the right edges of the predicted box and the ground truth box, μ represents a hyperparameter, and in the present invention, μ = 2 is set; h g represents the height of the ground truth box; w g represents the width of the ground truth box; min(S g ) represents selecting the ground truth box with the smallest area from all ground truth boxes; max(S g ) represents selecting the ground truth box with the largest area from all ground truth boxes. Figure 2 The schematic diagram of some distance variables involved in the small target constraint loss term is given.
[0050] According to the foreground-background classification loss and bounding box regression loss in the RPN stage, as well as the classification loss in the R-CNN stage and the bounding box regression loss after adding the small target constraint term as above, calculate the final loss of the network and optimize the model parameters.
[0051] Step S3, Network parameter adjustment: Using the detection network model with the initialized network parameters obtained from Step S2, perform model fine-tuning training on the training data containing the base classes and new classes. Among them, similar to Step S2, in the RPN stage, add a confidence-IoU collaborative proposal box filtering process, and in the R-CNN stage, adopt a box regression loss function with a small object constraint term added. Only fine-tune the classification branch and the box regression branch in the network, and fix the parameters of the remaining network layers to obtain a trained network.
[0052] Figure 3 The number of proposal boxes of different classes generated in the RPN stage using the baseline algorithm on the DIOR dataset is given. Figure 4 The number of proposal boxes of different classes generated in the RPN stage using the method of the present invention on the DIOR dataset is given, where bf, bl, be, cy, sp are new classes, and the rest are base classes. It can be seen that the number of proposal boxes generated for the new classes in the RPN stage by the method of the present invention has increased significantly.
[0053] Step S4, Object detection: Input the aviation image dataset to be processed into the trained network, and output the object detection result.
[0054] To verify the effectiveness of the method of the present invention, tests are carried out on two aviation image datasets, DIOR and AI-TOD, and compared with the FSCE and DeFRCN algorithms. The comparison results are shown in Tables 1 and 2. Among them, the algorithm FSCE is the algorithm proposed by Sun et al. in their 2021 work "Fsce: Few-shot object detection via contrastive proposal encoding [C] Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021: 7352-7362". The algorithm DeFRCN is the algorithm proposed by Qiao et al. in their 2021 work "Defrcn: Decoupled faster r-cnn for few-shot object detection [C] Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021: 8681-8690". The calculated performance metrics include: average precision AP, which represents the percentage of the number of correctly recognized objects to the total number of recognized objects; the AP value AP50 when the intersection over union IoU = 0.5; the AP value AP75 when the intersection over union IoU = 0.75; representing the object area at 22 to 8 2 the AP value APvt when [the object area] is between 8 2 to 16 2 the AP value APt when [the object area] is between 16 2 to 32 2 the AP value APs when [the object area] is between 32 2 to 96 2 the AP value APm when [the object area] is between 96 2 and the AP value APl when [the object area] is greater than 96. Table 1 shows the detection calculation results on the DIOR dataset. The DIOR dataset has a total of 20 categories. 15 categories are selected as the base categories, and 5 categories are new categories, which are baseball field, basketball court, bridge, chimney, and ship respectively. The number of labeled samples for the new categories is 10 (10Shot) and 20 (20Shot) respectively, and the remaining categories are base categories. It can be seen from Table 1 that the method of the present invention can effectively improve the performance of small-sample aerial image target detection. Table 2 shows the detection effect on the AI-TOD dataset. The AI-TOD dataset has a total of 8 categories, and there are many small targets. 5 categories are selected as the base categories, and 3 categories are new categories, which are airplane, swimming pool, and windmill respectively, and the remaining categories are base categories. It can be seen from Table 2 that the method of the present invention can effectively improve the performance of small-sample aerial image target detection and can alleviate the problem of difficult small-target detection.
[0055] Table 1
[0056]
[0057]
[0058] Table 2
[0059]
[0060] In summary, the present invention discloses a small-sample aviation image small-target detection method based on unbiased proposal box filtering, and particularly designs a confidence-IoU collaborative proposal box filtering and a small-target constraint loss term in the training stage. According to the class unbiasedness of the IoU score, a confidence-IoU collaborative proposal box filtering method is designed to solve the bias of the small-sample target detection model towards the base class, prevent proposal boxes belonging to new classes from being mistakenly deleted, and retain more target proposal boxes of new classes for subsequent R-CNN stage learning; aiming at the problem of small-target detection, a small-target regression constraint term is designed, taking into account the distance between the center points of the predicted box and the ground truth box, the size of the ground truth box, and the distances between the four sides of the predicted box and the ground truth box, so as to improve the detection performance of the model for small targets. At the same time, the method of the present invention does not require special design and change of the network structure, but modifies the proposal box filtering scheme in the RPN stage and adds a small-target regression constraint term. The design is simple and can be inserted into the existing two-stage small-sample target detection model without changing the model structure, achieving an obvious improvement in the detection effect on the aviation image dataset.
Claims
1. A small target detection method for small-sample aerial images based on unbiased proposal box filtering, characterized in that The steps are as follows: Step S1, constructing a training data set: randomly divide all categories of the aerial image data set into basic categories and new categories and the category spaces of the basic categories and the new categories do not intersect; among them, the number of labeled samples in each category of the basic categories is greater than or equal to 500, constituting a sub-data set The number of labeled samples in each category of the new categories does not exceed 20, constituting a sub-data set Step S2, basic training: Use the pre-trained model provided by the official Faster R-CNN network to train on the sub-dataset to obtain the initial network parameters. During training, keep the Faster R-CNN network structure unchanged, add confidence-IoU collaborative proposal box filtering processing in its RPN stage, and adopt a box regression loss function with a small target constraint term added in the R-CNN stage; The proposed box filtering process combining confidence and IoU includes the following steps: Step a: Sort the set P of proposed boxes output by RPN in descending order according to their foreground-background scores, and select the proposed boxes with higher rankings to form a set Among them, Step b: Calculate the intersection over union (IoU) scores of the remaining proposal boxes in the proposal box set P except for the set and sort them in descending order according to the IoU scores. Select the proposal boxes with higher rankings to form the set where Step c: Use the NMS algorithm to filter the set of proposed bounding boxes and the filtered proposed bounding boxes form a set Set the number of filtering bounding boxes to Step d, calculate the intersection over union (IoU) scores of the proposed bounding boxes and the ground truth bounding boxes in , and sort them in descending order according to the IoU scores. Select the top proposed bounding boxes to form a set Step e, merge the set of proposal boxes and Use the merged set as the set of proposal boxes finally generated in the RPN stage; The box regression loss function with a small object constraint term added in the R-CNN stage mentioned above refers to modifying the original box regression loss function in the R-CNN stage and adding a small object constraint term L tiny , specifically as follows: Among them, represents the bounding box regression loss function after adding the small object constraint term, L GIoU represents the original bounding box regression loss function, φ(S g ) is a function for normalizing the area of the ground truth bounding box; S g represents the area size of the ground truth bounding box; α is a hyperparameter, set α = 3; d represents the distance between the center points of the predicted bounding box and the ground truth bounding box; u top represents the distance between the top edges of the predicted bounding box and the ground truth bounding box, u bottom represents the distance between the bottom edges of the predicted bounding box and the ground truth bounding box, u left represents the distance between the left edges of the predicted bounding box and the ground truth bounding box, u right represents the distance between the right edges of the predicted bounding box and the ground truth bounding box, μ represents a hyperparameter, set μ = 2; h g represents the height of the ground truth bounding box; w g represents the width of the ground truth bounding box; min(S g ) represents selecting the one with the smallest area from all ground truth bounding boxes; max(S g ) represents selecting the one with the largest area from all ground truth bounding boxes; Step S3, network parameter adjustment: Use the method of Step S2 to perform network training again on sub-dataset and sub-dataset to obtain a trained network. During training, only adjust the parameters of the classification branch and the bounding box regression branch in the network, and fix the parameters of the remaining network layers; Step S4, object detection: Input the aerial image dataset to be processed into the trained network, and output the object detection results.
Citation Information
Patent Citations
A Faster RCNN target detection method based on refractory sample mining
CN109800778A
Small sample target detection method based on double-branch region suggestion network
CN114743045A