A Small-Sample Aerial Image Rotated Object Detection Method

By performing two-stage fine-tuning of the rotating object detector network Redet, adding angle constraints and classification reweighting module CRM, the positioning and classification problems in small sample rotation object detection are solved, and accurate detection of any rotation box is achieved.

CN115830480BActive Publication Date: 2025-07-11NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211578217.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2025-07-11
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

The existing rotation object detection algorithm is difficult to achieve accurate positioning and classification under the case of small samples, especially the detection performance of the new category is degraded, and the existing small sample rotation object detection model has rotation angle confusion problems when dealing with approximate square or square boxes.

Method used

The rotating object detector network Redet was trained using a two-stage fine-tuning method, adding angle constraint terms and classification reweighting module CRM, replacing the classification loss of RPN with Focal loss, and introducing edge vector cosine similarity loss EVCS Loss in the box regression loss of RCNN, adjusting some network parameters to improve model performance.

Benefits of technology

It improves the classification performance and positioning accuracy of small sample rotation target detection, can effectively handle the angle prediction of any rotation box, and is suitable for small sample rotation target detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115830480B_ABST
    Figure CN115830480B_ABST
Patent Text Reader

Abstract

The present invention provides a small-sample aviation image rotation target detection method. First, the rotation target detector network Redet is basically trained using basic category aviation sample data, and an angle constraint term is added to the box regression loss of RCNN; then, retraining is carried out using the basic category aviation sample data and the new category aviation sample data. During training, a classification reweighting module is added to the RPN module of the Redet network, the classification loss of the RPN is replaced with the Focal loss and a learnable loss weight term is added. Similarly, an angle constraint term is also added to the box regression loss of RCNN, and only some parameters of the network are adjusted while the rest of the parameters remain unchanged; finally, the aviation image is input into the trained network to obtain the target detection result. The present invention can effectively improve the small-sample rotation target detection performance of aviation images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image target detection, and particularly relates to a small-sample aerial image rotated target detection method. Background Art

[0002] The targets in aerial images are usually distributed densely in arbitrary directions. Using horizontal bounding boxes (HBBs) cannot accurately locate such targets in arbitrary directions. For this reason, many rotated target detection algorithms based on aerial images have emerged. These algorithms use oriented bounding boxes (OBBs) to represent and achieve accurate positioning of aerial image targets. However, the existing rotated target detection algorithms usually require a large amount of annotation information to obtain satisfactory detection performance. Compared with horizontal box annotation, the annotation of rotated boxes is more difficult. Because in addition to needing to mark the center position (x c , yc ) of the bounding box, the width w and height h of the bounding box, it is also necessary to mark the rotation angle θ of the bounding box, where θ represents the angle between the rotated box and the horizontal direction. Moreover, when some new classes appear and only a small amount of annotation information is available for these classes, the detection performance of the already trained algorithm model for these new classes will drop sharply. In short, on the one hand, the annotation difficulty of rotated boxes limits the acquisition of a large amount of annotation information and affects the model performance; on the other hand, it is also difficult for the existing rotated target detection models to generalize to targets of new classes with only a small amount of annotation information. Few-Shot learning is derived from the human's fast learning ability, similar to learning by analogy: people can quickly learn to detect a target of a new class by using only a small amount of annotation information based on a large amount of knowledge learned from base classes. Applying this learning ability to the rotated target detection task can alleviate the above-mentioned problems. Although some small-sample target detection algorithm models have been proposed currently, they are all for horizontal box detection and rarely involve rotated target detection.

[0003] Most small-sample object detection algorithms generally adopt a two-stage fine-tuning approach (TFA). The TFA training scheme mainly consists of two stages: the base training stage and the few-shot fine-tuning stage. In the base training stage, the entire object detector is trained on base classes, where each base class contains sufficient labeled training samples. In the few-shot fine-tuning stage, only the box predictor is fine-tuned on a balanced data subset, and the parameters of the remaining modules remain fixed. The balanced data subset is composed of base classes and novel classes, and each class has only a small number of labeled training samples. The class spaces of base classes and novel classes are also disjoint. The small-sample object detection model trained by TFA has certain localization ability, but the discrimination of object classes needs to be improved. Because the samples of base classes participate in both the base training stage and the few-shot fine-tuning stage, and the number of labeled training samples is much larger than that of novel classes, the trained model is easily biased towards base classes, resulting in low recognition accuracy for novel classes. Therefore, in small-sample rotated object detection, it is particularly crucial to improve the classification performance of the model.

[0004] In addition, the localization difficulty of rotated bounding boxes is higher than that of horizontal bounding boxes, and there are many challenges that are difficult to overcome by themselves. Therefore, it is still necessary to further improve the localization ability of the small-sample rotated object detection model. Currently, many excellent algorithms have emerged in the conventional rotated object detection tasks and achieved good progress. For example, the KFIoU loss function proposed by Yang et al. in their 2022 work "The KFIoU Loss for Rotated Object Detection[J]. arXiv preprint arXiv:2201.12558, 2022" effectively alleviates the long-existing boundary discontinuity and square problems and achieves excellent performance in the rotated object detection task. However, when dealing with approximately square or square boxes, the KFIoU loss still has the problem of rotation angle confusion. Because the KFIoU loss function models the predicted box and the ground truth box as 2D Gaussian distributions, shortens the distance between the centers of the two distributions, and then calculates the overlapping area through the Kalman filtering algorithm. But the Gaussian modeling of approximately square or square boxes is not an ellipse but a circle, which makes it difficult to distinguish the rotation angle of the object box. Summary of the Invention

[0005] To overcome the deficiencies of the prior art, the present invention provides a method for detecting rotated objects in small-sample aerial images. First, the rotated object detector network Redet is basically trained using the basic category aerial sample data, and an angle constraint term is added to the bounding box regression loss of RCNN; then, retraining is performed using the basic category aerial sample data and the new category aerial sample data. During training, a classification reweighting module is added to the RPN module of the Redet network, the classification loss of the RPN is replaced with the Focal loss and a learnable loss weight term is added. Similarly, an angle constraint term is also added to the bounding box regression loss of RCNN, and only some parameters of the network are adjusted while the rest remain unchanged; finally, the aerial image is input into the trained network to obtain the object detection result. The present invention can effectively improve the performance of detecting rotated objects in small-sample aerial images.

[0006] A method for detecting rotated objects in small-sample aerial images, characterized by the following steps:

[0007] Step S1, construct a training data set: all categories of the aerial image data set are randomly divided into basic categories and new categories, and the category spaces of the basic category and the new category do not intersect. Among them, the number of labeled samples in each category of the basic category is greater than or equal to 500, forming a basic category sub-data set, and the number of labeled samples in each category of the new category does not exceed 20, forming a new category sub-data set;

[0008] Step S2, basic training: Use the pre-trained model provided by the official ResNet50 network to train on the basic category sub-data set to obtain the initial network parameters. The network mentioned refers to the rotated object detector network Redet. During training, an angle constraint term is added to the bounding box regression loss of RCNN, that is:

[0009] L RCNN_reg =L KFIoU +0.04*L EVCS (1)

[0010]

[0011] where L RCNN_reg represents the bounding box regression loss function after adding the angle constraint term; L KFIoU represents the original KFIoU loss function; L EVCS represents the angle constraint term; represents the set of 8 directed vectors formed by the 4 vertices of the predicted bounding box, represents the directed vector starting from the upper left vertex 1 of the predicted bounding box and ending at the upper right vertex 2, represents the directed vector starting from the upper left vertex 1 of the predicted bounding box and ending at the lower left vertex 4, Denotes a directed vector starting from the upper right vertex 2 of the prediction box and ending at the upper left vertex 1, Denotes a directed vector starting from the upper right vertex 2 of the prediction box and ending at the lower right vertex 3, Denotes a directed vector starting from the lower right vertex 3 of the prediction box and ending at the upper right vertex 2, Denotes a directed vector starting from the lower right vertex 3 of the prediction box and ending at the lower left vertex 4, Denotes a directed vector starting from the lower left vertex 4 of the prediction box and ending at the lower right vertex 3, Denotes a directed vector starting from the lower left vertex 4 of the prediction box and ending at the upper left vertex 1; Denotes two directed vectors, both starting from the upper left vertex 1 of the ground truth box. The end point of the directed vector is the upper right vertex 2 of the ground truth box, and the end point of the directed vector is the lower left vertex 4 of the ground truth box; The vector Denotes the directed vector selected from the prediction box vector set that is closest to the two directed vectors of the ground truth box in terms of direction and vector length; exp represents the exponential function with the natural constant e as the base; Cosinesimilarity represents the calculation of cosine similarity;

[0012] Step S3, Network parameter adjustment: Re - train the network on the basic category sub - dataset and the new category sub - dataset to obtain a trained network. Among them, during training, add a classification re - weighting module CRM to the RPN module of the Redet network, replace the classification loss of RPN with Focal loss and add a learnable loss weight term, add an angle constraint term to the box regression loss of RCNN. At the same time, during training, only adjust the parameters of the classification branch, regression branch of RPN and RCNN, and the classification re - weighting module CRM, and keep the parameters of the remaining modules of the network fixed;

[0013] The described classification re - weighting module CRM is mainly composed of a convolutional layer, placed after the box localization branch, and inputs the localization information predicted by the RPN model into the CRM module to obtain an output value containing the localization information;

[0014] The classification loss function of RPN that is replaced with Focal loss and adds a learnable loss weight term is as follows:

[0015] L RPN_cls = loss_weight * [-α * (1 - p t ) γ * label * log(p t )-(1 - α) * pt γ *(1-label)*log(1-p t )] (3)

[0016]

[0017]

[0018]

[0019]

[0020] Among them, L RPN_cls represents the classification loss of the RPN module; loss_weight represents the learnable loss weight term; W represents the learnable weight matrix; α is the first hyperparameter, and its value is set to 0.25; p t represents the predicted foreground-background classification score finally output by the RPN module; γ is the second hyperparameter, and its value is set to 2; label represents the foreground-background class true label; scores cls represents the reweighted foreground-background classification score; scores_cls represents the foreground-background score output by the classification branch of the RPN module; reg2cls represents the output value of the CRM module;

[0021] Step S4, object detection: Input the aviation image dataset to be processed into the trained network, and output the object detection result thereof.

[0022] The beneficial effects of the present invention are: Since the classification reweighting module CRM is added, the predicted positioning information is used to reweight the classification score, and at the same time, the classification loss of the RPN is replaced with the Focal loss, and a learnable loss weight term is added, which can effectively improve the classification performance of small-sample rotated object detection; Since aiming at the problem that the KFIoU box loss function is confused about the rotation angles of approximate square or square boxes, an edge vector cosine similarity loss is designed and added to the KFIoU box loss function as an angle constraint term, which can achieve accurate prediction of the angles of arbitrary rotated boxes (including approximate square or square boxes); The implementation of the present invention is simple. While obtaining good small-sample rotated object detection results, it can also be extended and applied to other rotated object detection models. Description of the Drawings

[0023] Figure 1 is a flowchart of a small-sample aviation image rotated object detection method of the present invention;

[0024] Figure 2 is a schematic diagram of the training stage of the rotated object detection network of the present invention;

[0025] Figure 3 This is the detection result image of the method of the present invention on the DOTA dataset. Specific implementation manners

[0026] The present invention will be further described below in conjunction with the drawings and embodiments. The present invention includes but is not limited to the following embodiments.

[0027] As Figure 1 shown, the present invention provides a small-sample aerial image rotated target detection method, and its specific implementation process is as follows:

[0028] Step S1: Construct a training data set

[0029] Randomly divide all categories of the aerial image data set into basic categories and new categories and the category spaces of the basic categories and the new categories do not intersect Among them, the number of labeled training samples of the basic categories is sufficient to form a basic category sub-data set where the number of labeled training samples of each category is at least 500; each category in the new category has only a small number of labeled samples to form a new category sub-data set where the number of labeled training samples of each category does not exceed 20.

[0030] Step S2: Basic training

[0031] The pre-trained model provided by the official ResNet50 network is used to train on the basic category sub-dataset to obtain the initial network parameters. The network mentioned refers to the rotation target detector network Redet, which mainly includes the backbone network ResNet50, FPN (Feature Pyramid Networks), RPN, and RCNN. Among them, ResNet50 is the network structure proposed by He et al. in their 2016 work "Deep residual learning for image recognition [C] Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770-778"; FPN is the network framework proposed by Lin et al. in their 2017 work "Feature pyramid networks for object detection [C] Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 2117-2125". The object detection process of the Redet rotation target detection method is basically the same as that of Faster RCNN. The differences are as follows: First, considering that aerial targets are often distributed in any direction, and ordinary CNNs do not explicitly model direction changes, a large amount of rotation augmentation data is required to train an accurate target detector. Therefore, the Redet rotation target detection method adds a rotation-equivariant network to ResNet50 and FPN to extract rotation-equivariant features for accurately predicting directions. Second, Redet proposes rotation-invariant RoI alignment (RiRoI Align), which adaptively extracts rotation-invariant features from equivariant features according to the direction of the RoI. In summary, the backbone network extracts the features of training samples, RPN generates a set of proposal boxes related to the basic categories, and the sample features and the set of proposal boxes are input into the RCNN network together to obtain classification scores and localization results.

[0032] In addition, an angle constraint term is added to the bounding box regression loss of RCNN, namely the Edge-Vectors Cosine Similarity Loss (EVCS Loss). The bounding box regression loss of the original RCNN is the KFIoU loss function proposed by Yang et al. in their 2022 work "The KFIoU Loss for Rotated Object Detection[J]. arXiv preprint arXiv:2201.12558, 2022". The KFIoU loss function models the predicted bounding box and the ground truth bounding box as 2D Gaussian distributions, shortens the distance between the centers of the two distributions, and then calculates the overlapping area through the Kalman filtering algorithm. However, the Gaussian modeling of approximately square or square bounding boxes is not an ellipse but a circle, which makes it difficult to distinguish the rotation angle of the target bounding box. Therefore, based on the KFIoU loss, the present invention designs an angle constraint term, called the Edge-Vectors Cosine Similarity Loss (EVCS Loss), as follows:

[0033]

[0034] Where L EVCS represents the angle constraint term designed by the present invention, namely the Edge-Vectors Cosine Similarity Loss (EVCS Loss); represents the 8 directed vectors formed by the 4 vertices of the predicted bounding box. The order of the 4 vertices of the predicted bounding box (or the ground truth bounding box) is: the upper left vertex 1 of the box, the upper right vertex 2 of the box, the lower right vertex 3 of the box, and the lower left vertex 4 of the box. represents the directed vector starting from the upper left vertex 1 of the predicted bounding box and ending at the upper right vertex 2, represents the directed vector starting from the upper left vertex 1 of the predicted bounding box and ending at the lower left vertex 4, represents the directed vector starting from the upper right vertex 2 of the predicted bounding box and ending at the upper left vertex 1, represents the directed vector starting from the upper right vertex 2 of the predicted bounding box and ending at the lower right vertex 3, represents the directed vector starting from the lower right vertex 3 of the predicted bounding box and ending at the upper right vertex 2, represents the directed vector starting from the lower right vertex 3 of the predicted bounding box and ending at the lower left vertex 4, represents the directed vector starting from the lower left vertex 4 of the predicted bounding box and ending at the lower right vertex 3, Denotes a directed vector starting from the lower left vertex 4 of the prediction box and ending at the upper left vertex 1; Denotes two directed vectors, both starting from the upper left vertex 1 of the ground truth box. The directed vector ends at the upper right vertex 2 of the ground truth box, and the directed vector ends at the lower left vertex 4 of the ground truth box; The vector Denotes the directed vector selected from the prediction box vector set that is closest in direction and vector length to the two directed vectors of the ground truth box ; exp represents the exponential function with the natural constant e as the base; Cosinesimilarity represents calculating the cosine similarity. As long as the cosine similarity between the edge vectors of the prediction box and the edge vectors of the ground truth box is large, it means that the included angle between the two vectors is smaller, and the included angle range between the vectors is constantly [0, 180] degrees. Therefore, the final rotation box regression loss function L RCNN_reg is:

[0035] L RCNN_reg = L KFIoU + 0.04 * L EVCS (9)

[0036] where, L KFIoU represents the original KFIoU loss function.

[0037] The final loss of the network is the sum of the RPN loss L RPN (including the classification loss and the box regression loss) and the RCNN loss L RCNN (including the classification loss and the box regression loss).

[0038] Step S3, Network parameter adjustment

[0039] Through the basic training stage, the initial parameters of the small sample rotation target detection model can be obtained. The purpose of the small sample rotation target detection task is to be able to quickly generalize to new category targets only by using a small amount of annotation information. Therefore, the present invention adopts the strategy of this two-stage fine-tuning method TFA. After obtaining the network initial parameters, the network is trained again on the sub-dataset containing the basic category and the new category sub-dataset to obtain a trained network.

[0040] To enable the network to better perform small-sample rotated object detection, the present invention designs an RPN classification reweighting module and an edge vector cosine similarity loss. Specifically, during training, a classification reweighting module CRM is added to the RPN module of the Redet network. At the same time, the classification loss of the RPN is replaced with the Focal loss, and a learnable loss weight term is added. Similarly, an angular constraint term is added to the box regression loss of the RCNN. Meanwhile, except for the classification branch and regression branch of the RPN and RCNN, and the parameters of the CRM that need to be fine-tuned, the parameters of the remaining modules remain fixed. Figure 2 The schematic diagram of network parameter adjustment through network retraining is given.

[0041] The RPN module includes a classification head and a box localization head. The designed classification reweighting module CRM of the present invention is mainly placed after the box localization head, aiming to use the relatively reliable localization information predicted by the model to assist the classification learning of target instances. The classification reweighting module CRM is mainly composed of a convolutional layer. The predicted localization information of the model is input into the CRM, and then a weight matrix containing the localization information is output, and this weight matrix is applied to the classification scores predicted by the model. Considering that the small-sample rotated object detection model is prone to bias towards the base classes That is, it will give the base classes a higher predicted classification score, which is extremely unfavorable for new classes Therefore, the present invention maps the elements in the CRM output value reg2cls to (0, 1) to form a learnable weight matrix W, and applies it to the foreground-background classification scores scores_cls of the RPN module, that is:

[0042]

[0043] scores cls =(1 + W) * scores_cls (11)

[0044] Among them, scores cls represents the foreground-background classification scores after reweighting.

[0045] Then, this foreground-background classification scores scores after reweighting is used clsFilter the suggestion box. The purpose of this re-weighting is to re-adjust the classification scores using category-unbiased localization information, making the scores of new categories with originally low classification scores higher and increasing the model's attention to new categories. In addition, in order to better update the parameters of the classification re-weighting module CRM and alleviate the foreground-background imbalance problem, the present invention uses the Focal loss proposed by Lin et al. in their 2017 work "Focal loss for dense object detection[C]Proceedings of the IEEE international conference on computer vision. 2017:2980-2988" as the classification loss of the RPN, and on this basis, adds a learnable loss weight:

[0046] L RPN_cls = loss_weight * [-α * (1 - p t ) γ * label * log(p t ) - (1 - α) * p t γ * (1 - label) * log(1 - p t )] (12)

[0047]

[0048]

[0049] Among them, L RPN_cls represents the Focal classification loss of the RPN module after adding the learnable loss weight term. loss_weight represents the learnable loss weight term designed by the present invention; α represents the first hyperparameter in the Focal loss, and the present invention uses the default value of 0.25; p t represents the predicted foreground-background classification score finally output by the RPN module; γ represents the second hyperparameter in the Focal loss, and the present invention uses the default value of 2; label represents the foreground-background category true label; reg2cls represents the output value of the CRM module.

[0050] Similarly, an angle constraint term as described in the previous step S2, i.e., the Edge-Vectors Cosine Similarity Loss (EVCS Loss), is also added to the box regression loss function of the RCNN.

[0051] Step S4, Object detection

[0052] Input the aviation image dataset to be processed into the trained network, and the target detection result is output.

[0053] To verify the effectiveness of the method of the present invention, tests for rotated object detection are carried out on the DOTA-v1.0 rotated aviation image dataset. The DOTA-v1.0 dataset has a total of 15 categories. Randomly select 10 categories as the basic categories and 5 categories as the new categories. The new categories include airplane (PL), track and field ground (GTF), ship (SH), storage tank (ST), and swimming pool (SP). The remaining categories are all basic categories, including baseball infield (BD), bridge (BR), small vehicle (SV), large vehicle (LV), tennis court (TC), basketball court (BC), football field (SBF), roundabout (RA), seaport (HA), and helicopter (HC). Compare with the RoI Trans-KFIoU and Redet-KFIoU algorithms, and the comparison results are shown in Table 1. Among them, RoI Trans-KFIoU is an algorithm proposed by Ding et al. in their 2019 work "Learning RoI transformer for oriented object detection in aerial images[C]Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2019:2849-2858"; Redet-KFIoU is an algorithm proposed by Han et al. in their 2021 work "Redet: A rotation-equivariant detector for aerial object detection[C]Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021:2786-2795". The metric AP is the average precision, which represents the percentage of the number of correctly recognized objects in the total number of recognized objects. The larger the value, the more accurate the detection result. It can be seen from Table 1 that the method of the present invention can effectively improve the performance of rotated object detection in small-sample aviation images, and the number of labeled samples for each new category is 10.

[0054] Table 1

[0055]

[0056] Figure 3 The detection result images of the method of the present invention on the DOTA-v1.0 dataset are given.

[0057] As can be seen from Table 1 and Figure 3 it can be seen that the method of the present invention can effectively improve the performance of small-sample aerial image rotated object detection.

[0058] In summary, the present invention discloses a small-sample aerial image rotated object detection method (few-shot oriented object detection, FSO2D), and mainly designs a classification reweighting module (CRM) and an edge vector cosine similarity loss (EVCS Loss), etc. According to the predicted positioning information, a classification reweighting module is designed to improve the classification performance of small-sample rotated object detection; aiming at the problem that the KFIoU loss is confused by the rotation angles of approximate square or square boxes, an edge vector cosine similarity loss is designed and added to the KFIoU loss function as an angle constraint term to achieve accurate prediction of the angles of arbitrary rotated boxes (including approximate square or square boxes). The implementation method of the present invention is simple and can be inserted into the existing rotated object detection model, and remarkable small-sample rotated object detection results are obtained on the aerial image dataset.

Claims

1. A small-sample aviation image rotation target detection method, characterized in that The steps are as follows: Step S1, construct a training data set: Randomly divide all categories of the aerial image data set into a basic category and a new category, and the category spaces of the basic category and the new category do not intersect. Among them, the number of labeled samples in each category of the basic category is greater than or equal to 500, constituting a basic category sub-data set, and the number of labeled samples in each category of the new category does not exceed 20, constituting a new category sub-data set; Step S2, basic training: Use the pre-trained model provided by the official ResNet50 network to train on the basic category sub-data set to obtain initial network parameters. The network mentioned refers to the rotation target detector network Redet. When training, add an angle constraint term to the box regression loss in RCNN, that is: L RCNN_reg = L KFIoU + 0.04 * L EVCS (1) Among them, L RCNN_reg represents the bounding box regression loss function after adding the angular constraint term; L KFIoU represents the original KFIoU loss function; L EVCS represents the angular constraint term; represents the set of 8 directed vectors formed by the 4 vertices of the predicted bounding box, represents the directed vector starting from the upper left vertex 1 of the predicted bounding box and ending at the upper right vertex 2, represents the directed vector starting from the upper left vertex 1 of the predicted bounding box and ending at the lower left vertex 4, represents the directed vector starting from the upper right vertex 2 of the predicted bounding box and ending at the upper left vertex 1, represents the directed vector starting from the upper right vertex 2 of the predicted bounding box and ending at the lower right vertex 3, represents the directed vector starting from the lower right vertex 3 of the predicted bounding box and ending at the upper right vertex 2, represents the directed vector starting from the lower right vertex 3 of the predicted bounding box and ending at the lower left vertex 4, represents the directed vector starting from the lower left vertex 4 of the predicted bounding box and ending at the lower right vertex 3, represents the directed vector starting from the lower left vertex 4 of the predicted bounding box and ending at the upper left vertex 1; represents two directed vectors, both starting from the upper left vertex 1 of the ground truth bounding box, and the end point of the directed vector is the upper right vertex 2 of the ground truth bounding box, and the end point of the directed vector is the lower left vertex 4 of the ground truth bounding box; the vector represents the directed vector selected from the predicted bounding box vector set that is closest to the two directed vectors of the ground truth bounding box in terms of direction and vector length; exp represents the exponential function with the natural constant e as the base; Cosinesimilarity represents calculating the cosine similarity; Step S3, network parameter adjustment: Train the network again on the basic category sub-data set and the new category sub-data set to obtain a trained network. Among them, when training, add a classification reweighting module CRM to the RPN module of the Redet network, replace the classification loss of RPN with Focal loss and add a learnable loss weight term, and add an angle constraint term to the box regression loss in RCNN. At the same time, when training, only adjust the parameters of the classification branch and regression branch of RPN and RCNN and the classification reweighting module CRM, and keep the parameters of the remaining modules of the network fixed; The classification reweighting module CRM mainly consists of a convolutional layer, which is placed after the box localization branch, and inputs the localization information predicted by the RPN model into the CRM module to obtain an output value containing the localization information; The classification loss function of RPN that is replaced with Focal loss and adds a learnable loss weight term is as follows: L RPN_cls = loss_weight * [-α * (1 - p t ) γ * label * log(p t ) - (1 - α) * p t γ * (1 - label) * log(1 - p t )](3) Among them, L RPN_cls represents the classification loss of the RPN module; loss_weight represents the learnable loss weight term; W represents the learnable weight matrix; α is the first hyperparameter, and its value is set to 0.25; p t represents the predicted foreground-background classification score finally output by the RPN module; γ is the second hyperparameter, and its value is set to 2; label represents the foreground-background class true label; scores cls represents the reweighted foreground-background classification score; scores_cls represents the foreground-background score output by the classification branch of the RPN module; reg2cls represents the output value of the CRM module; Step S4, target detection: Input the aerial image data set to be processed into the trained network, and output its target detection result.

Citation Information

Patent Citations

  • Small sample target detection method based on self-supervised contrast constraint

    CN114841257A

  • Remote sensing image small sample scene classification method based on multi-task dynamic contrast learning

    CN114913379A