A weakly supervised small sample object detection system and method based on localization pre-training

Through transformer-based attention-guided positioning pre-training, positioning distillation learning, dual-factor driven optimization and hybrid loss module, the problems of low candidate box quality and border annotation dependence in weakly supervised small-sample target detection are solved, and high-precision and low-latency target detection is achieved.

CN116958742BActive Publication Date: 2025-10-10FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310832460.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-07
Publication Date
2025-10-10
Estimated Expiration
2043-07-07

AI Technical Summary

Technical Problem

Existing weakly supervised small-sample target detection methods have problems such as low candidate box quality, model performance relying on the quality of bounding box annotations, high overhead and long latency when processing high-resolution images, making it difficult to achieve high-precision and low-latency target detection.

Method used

It adopts a transformer-based attention-guided localization pre-training module, a localization distillation learning module, a dual-factor-driven asymptotic optimization module, and a hybrid loss module. Through pre-training and optimization mechanisms, it reduces the dependence on bounding box annotation and improves target detection performance.

Benefits of technology

It achieves accurate detection of targets under weak supervision conditions, reduces the cost of acquiring training data, improves the versatility and detection accuracy of the model, and reduces the overhead and delay of processing high-resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958742B_ABST
    Figure CN116958742B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of machine learning, and particularly relates to a weakly supervised small sample target detection method and system based on positioning pre-training and progressive optimization strategy. The present application introduces a weakly supervised learning mechanism into a small sample deep target detection framework, and establishes a weakly supervised small sample target detection system with high accuracy. The method framework of the present application is simple, convenient to use, strong in scalability and strong in interpretability, and the weakly supervised small sample target detection results on two mainstream visual attribute data sets all exceed those of existing methods. The present application can provide support for basic frameworks and algorithms for target detection technology in military and industrial application fields, and can also be easily extended to other small sample learning tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning, and in particular relates to a weakly supervised small sample target detection system and method. Background Art

[0002] Object detection is a fundamental task in computer vision, achieving tremendous success in many real-world scenarios. Currently, deep learning-based methods such as Faster R-CNN, YOLO, and DETR have become mainstream. These methods typically rely on large amounts of well-annotated data to train models, enabling accurate object recognition and localization. However, collecting and annotating such data is expensive and time-consuming, limiting their application. In recent years, few-shot object detection (FSOD) has received significant attention, aiming to achieve effective object detection using only a small amount of labeled data for novel classes. However, to train the models, researchers must still collect a large amount of strongly annotated training data for the base classes, including the category and bounding box of each object of each object class in each training image, resulting in significant annotation costs. Furthermore, the performance of few-shot object detection models relies heavily on the quality of bounding box annotations. However, due to the complexity of images and the diversity of object morphology, ensuring the quality of bounding box annotations is difficult, which inevitably impacts model performance. Weakly supervised learning provides a systematic framework for addressing this problem. Model training requires only image category annotations, eliminating the need for fully annotated training data, and enables accurate detection of objects in few-shot categories.

[0003] Most existing weakly supervised small-sample target detection algorithms are modified based on the WSDDN weakly supervised target detection framework, and these methods have the following limitations:

[0004] (1) Most methods based on multi-instance learning use non-parameterized methods (such as selective-search and edgebox) to generate candidate boxes. This method cannot guarantee the quality of the candidate boxes and often marks background areas without objects as foreground boxes, which introduces noise into subsequent training.

[0005] (2) Based on the class activation map method, the class activation map is thresholded to generate candidate boxes. Due to the locality of the convolutional neural network, the model often only focuses on the area of ​​the target containing significant distinguishing features. The candidate box cannot contain the entire target, resulting in the common problem of distinguishing regions in weakly supervised target detection.

[0006] (3) In general, existing methods have low accuracy and high overhead when processing high-resolution images, making them difficult to achieve practical use. Summary of the Invention

[0007] In response to the above problems in the prior art, the purpose of the present invention is to provide a weakly supervised small sample target detection system and method based on positioning pre-training, so as to reduce the cost of obtaining training data and improve the small sample target detection performance.

[0008] The weakly supervised small-sample target detection system based on positioning pre-training provided by the present invention is a new target detection system (denoted as WFS-DETR). Based on the transformer-based target detector (i.e., DETR detector), it designs and uses an attention-guided positioning pre-training module and a positioning distillation learning module to train the model's target positioning ability, designs and uses a dual-factor-driven asymptotic optimization module to train the model's image classification ability, and designs and uses a hybrid loss function module to coordinate different module components to drive overall end-to-end training; wherein:

[0009] The attention-guided positioning pre-training module is described in detail as follows:

[0010] Threshold screening of class activation maps is an important way to generate candidate boxes in weakly supervised target detection. Class activation maps are often generated by back-mapping the classifier weights back to the image features. Such methods often fall into the problem of discriminative regions, that is, the model only focuses on a part of the salient region of the target. The root of the problem lies in the locality of convolutional neural networks. In recent years, transformers (a new type of neural network structure that uses attention mechanisms) have been widely used in the field of computer vision, and their global attention mechanisms are very suitable for overall modeling of targets in images. Therefore, utilizing the characteristics of transformer global modeling, the present invention adopts an attention-guided positioning pre-training module, denoted as ALN, and uses visual transformers to accurately locate targets. Specifically, as shown in the attached figure Figure 2 As shown in , ALN is composed of multiple standard visual transformer blocks, which are connected as a whole to the third stage of the feature extraction network. Its purpose is to input the features output by the feature extraction network together with the category label into the ALN, and then use the characteristics of the visual transformer to output attention score maps of different sizes for the foreground target and background in the image. According to the attention score map, the area with a high score is predicted as the foreground target area.

[0011] The positioning distillation learning module is described in detail as follows:

[0012] After pre-training, the positioning pre-training module has a certain target positioning capability. In order to further improve the positioning accuracy and optimize the overall training process, the present invention adopts a positioning distillation learning module to distill the positioning capability of the positioning pre-training module into the DETR target detector through data enhancement and knowledge distillation. Specifically, as shown in the attached Figure 1As shown in

[15] , the localization pre-training module ALN is jointly trained with the DETR target detector. Data augmentation is applied to the output of the localization pre-training module ALN as supervision, and the DETR target detector is trained using foreground discrimination loss and bounding box regression loss. The target localization capability in the localization pre-training module ALN is distilled into the DETR target detector.

[0013] The dual-factor driven asymptotic optimization module is specifically described as follows:

[0014] In previous methods, cascade structures are often used to optimize preliminary weakly supervised detection results, using the output of the upper layer as a pseudo-label to supervise the prediction results of the lower layer. In this structure, the selection of pseudo-labels is crucial, and inaccurate pseudo-labels will bring incorrect supervision guidance to subsequent training. Most existing methods only consider classification scores as the basis for selecting pseudo-labels. The disadvantage of this approach is that the area with the highest classification score usually does not have the integrity of the foreground, and falls into the common problem of discriminative regions in weakly supervised target detection. In order to solve this problem, the present invention adopts a dual-factor driven asymptotic optimization module, which jointly considers the classification score and the target integrity score as the basis for pseudo-label selection, and reasonably selects pseudo-labels to drive subsequent model optimization. Specifically, as shown in the attached figure Figure 1 As shown in , the dual-factor driven asymptotic optimization module consists of three optimization sub-blocks, each of which consists of an optimized classifier and an optimized foreground discriminator. During the optimization training process, the classification score is combined with the foreground discriminant score as the basis for pseudo-label selection, which improves the quality of pseudo-labels and effectively drives the prediction optimization training.

[0015] The mixed loss module is specifically described as follows:

[0016] There are many problems with existing methods of weakly supervised small sample target detection, such as: 1) inaccurate positioning, the model can often only detect part of the target area, the intersection of the detection frame and the real target frame is relatively low, and it is difficult to meet the high-precision prediction requirements. 2) The method has high overhead. When processing high-resolution images, the existing method takes too long and is difficult to meet the low-latency prediction requirements. This invention adopts a hybrid loss module to fuse the positioning pre-training loss. Locating distillation losses Multiple instance learning loss Classification optimization loss and positioning optimization loss By making the above components precisely combined and coordinated, the problems existing in the existing methods can be effectively solved. Figure 1 As shown in , during the pre-training phase, the hybrid loss module is composed of the localization pre-training loss and locate distillation losses The main function of the joint structure is to give the DETR target detector a universal target positioning capability. In the optimization training phase, the hybrid loss module is composed of multiple instance learning losses. Classification optimization loss and positioning optimization loss The main purpose of the joint construction is to give the DETR object detector accurate classification capabilities through multi-instance learning and optimization training. In general, the hybrid loss module plays different roles at different stages. With the hybrid loss module, the DETR object detector can finally achieve accurate object detection using only weakly labeled data.

[0017] Based on the above detection system, the present invention also provides a weakly supervised small sample target detection method, the specific steps are as follows:

[0018] (1) Prepare pre-training data.

[0019] The ImageNet dataset, commonly used in computer vision tasks, is used as a pre-training dataset. Each image only has the image category annotation information, without the need for bounding box annotations. The elements in the dataset are composed of two-tuples such as (Image, Label).

[0020] (2) Target positioning pre-training.

[0021] The localization pre-training module ALN is trained using the pre-training dataset. ALN consists of K multi-head self-attention blocks stacked together. The original image features are converted into P*P image blocks after the first three stages of the swin transformer (visual transformer using sliding window) feature extraction network. A total of N (N=P*P) image blocks are labeled t ns With 1 additional category marker t c The splicing is sent to ALN for multi-head self-attention calculation. The calculation process is as follows, where W Q is the query transformation matrix weight, W K is the contrast transformation matrix weight, W V is the assignment conversion matrix weight, Attention multi It is a multi-head self-attention operation, is the attention score matrix:

[0022] The shape is (D is the characteristic dimension of the matrix, N is the total number of image marker blocks), take the last N columns to get (D is the feature dimension of the matrix, N is the total number of image tag blocks), that is, the image block tag enhanced by attention; then use the linear layer mapping (D is the feature dimension of the matrix, and C is the number of categories) is used to label the image block with category information, and the linear mapping is trained using the cross-entropy loss function, and the training process is as follows:

[0023]

[0024] where w c and w i is the category mapping matrix weight, T is the matrix transpose operation, and GAP is the global average pooling operation.

[0025] The trained linear layer assigns category information, and the attention map of the K visual transformer block is averaged to obtain The first row and the last N columns are taken to obtain the category-independent attention map, and the category-independent attention map is multiplied by the image block label containing category information to obtain the category-dependent activation map. The category activation map is screened, and the low response value region is set as the background, so that the high response value region is obtained. Then, the maximum and minimum values of the horizontal and vertical axes of all point coordinates in the high response region are taken, and the minimum enclosing rectangle that can enclose the high response value region is drawn according to these extreme points, and the bounding box is generated. The calculation process is as follows:

[0026]

[0027] where, is the category-independent attention map, is the image block label containing category information, Thr is the threshold value of the screening operation, and MMR refers to the process of generating the minimum enclosing rectangle.

[0028] (3) Target positioning ability distillation.

[0029] Although the ALN can locate the target, the positioning accuracy is not enough, and it cannot accurately frame the whole target. In order to solve this problem, the present application uses random jitter to data enhance the candidate box output by the pre-trained target positioning module ALN. Random jitter is applied to the minimum horizontal coordinate x1 of the candidate box, the minimum vertical coordinate y1 of the candidate box, the maximum horizontal coordinate x2 of the candidate box, and the maximum vertical coordinate y2 of the candidate box in four directions. The specific steps are as follows:

[0030] b aug [x1±α1*w,y1±α2*h,x2±α3*w,y2±α4*h], (4)

[0031] Among them, α(α1,α2,α3,α4) is the jitter coefficient, w and h are the width and height of the original candidate box, which are calculated from the original candidate box coordinates (w = x2-x1, h = y2-y1). In order to ensure the stability of data enhancement, it is necessary to limit the value of α(α1,α2,α3,α4) to the range of [0,1 / 6] to ensure that the candidate box matches the foreground target. Then, the enhanced box is optimally matched with the model output box and the loss is calculated. Specifically expressed as:

[0032]

[0033] Among them, y i Refers to the supervision information of the i-th real box and its category, Refers to the prediction information of the i-th prediction box and its category, o i Refers to the foreground integrity supervision information of the i-th real box, Refers to the foreground completeness prediction information of the i-th prediction box, b i Refers to the bounding box coordinate supervision information of the i-th real box, Refers to the bounding box coordinate prediction information of the i-th prediction box.

[0034] By jointly training the pre-trained target localization module ALN and the DETR target detector using Equation (6), the target localization performance in ALN ​​is transferred from ALN to the detector, and the random jitter enhancement technology further improves the detector's target localization performance, ensuring that the detection box can cover the entire foreground target. This allows the detector to perform accurate foreground localization regardless of category.

[0035] (4) Dual factors drive optimization.

[0036] In the general framework of weakly supervised target detection, image-level classification predictions are obtained by aggregating the results of the category score predictor and the category contribution predictor, that is, multiplying the prediction result matrices of the two and summing them in the category dimension, that is, obtaining the image-level classification prediction results of the original input. Some subsequent work applies a progressive refinement structure to optimize the detection results layer by layer, that is, using the output results of the upper layer as pseudo-labels to supervise the predictions of the lower layer. However, these methods usually only use classification scores as the basis for pseudo-label mining, which can easily lead to discriminative area problems. The present invention uses a dual-factor driven optimization strategy that takes into account both classification scores and target integrity scores, overcomes the problem of discriminative areas, and helps to obtain more accurate pseudo-labels. Specifically, for each image, a series of candidate boxes are generated using a DETR detector that has been pre-trained to obtain general target positioning capabilities:

[0037]

[0038] in, are the bounding box coordinates, is the target completeness score, is the classification score, is the detection score. Combining the instance-level classification score with the detection score, we can get the image-level classification prediction score:

[0039]

[0040] Where N is the number of candidate boxes in the image and C is the total number of predicted categories.

[0041] The image category label is used as supervision information, and the DETR detector can be trained for image classification using the cross entropy loss shown in formula (9):

[0042]

[0043] Among them, y c is the true category supervision information of the image, is the predicted image classification score.

[0044] In order to generate more accurate proposal boxes, K optimization layers are constructed, including target category score optimization predictors and target completeness score optimization predictors. The output of each layer is expressed as:

[0045]

[0046] in, is to optimize bounding box prediction, is the optimization target completeness score, is the optimization classification score. The optimization target integrity score of the K-1 level is and optimize classification scores Combined as the basis for selecting the K-1 level pseudo label:

[0047]

[0048] Next, use the pseudo-label of the K-1th level as supervision information to supervise the output of the Kth level and calculate the classification loss and target integrity loss:

[0049]

[0050] in, is the classification score prediction result of the K-1 level, is the target integrity score prediction result of the K-1 level, is the classification score prediction result of the K-th level, is the target integrity score prediction result of level K, N ris the number of all matched prediction boxes to supervision information in the K-th stage, BCE refers to binary cross-entropy loss, and CE refers to cross-entropy loss.

[0051] Under the supervision of the progressive pseudo-label, the discriminative information captured by the small candidate box is transmitted to the overlapped large candidate box, and the target integrity information of the large candidate box is transmitted to the small candidate box at the same time, so that the detection accuracy is improved in both directions.

[0052] (5) Hybrid loss function calculation.

[0053] For the training needs of different stages, the hybrid loss function is used for training. Specifically, in the pre-training stage, the loss function is:

[0054]

[0055] In the optimization stage, the loss function is:

[0056]

[0057] Wherein, λ P , λ1, λ2 are weight coefficients, which can be determined according to specific actual conditions; for example, the value of λ P may be different in different stages of pre-training, for example, in the first half of the pre-training, λ P is 1, and in the second half of the pre-training, λ P is 0.5. In the training stage, λ1 can be 1, and λ2 can be 10.

[0058] By combining these loss functions, different components can work well together.

[0059] The present application at least includes the following beneficial effects:

[0060] The present application designs a weakly supervised small sample target detection method based on positioning pre-training and progressive optimization strategy, uses the pre-training-optimization mechanism to let the model learn to locate the target, so as to get rid of the dependence on the frame label information, and only uses the image level label to train the model for accurate detection. Specifically, the present application designs an attention guided positioning pre-training module and a double factor driven asymptotic optimization module, which obtains general target positioning ability through large-scale pre-training, and avoids falling into the problem of discriminative region. The present application can complete accurate positioning of the foreground target of different data sets through only one pre-training, so that the present application has good universality.

[0061] Other advantages, objects and features of the present application will be partly embodied through the following description, and partly understood by those skilled in the art through research and practice of the application. BRIEF DESCRIPTION OF DRAWINGS

[0062] Figure 1 This is a network structure diagram.

[0063] Figure 2 This is the structural diagram of the attention-guided localization pre-training module.

[0064] Figure 3 Comparison of AP results for pre-training with different types / number ratios.

[0065] Figure 4 This is a test example for the method of the present invention and other methods. DETAILED DESCRIPTION

[0066] The present invention is further described below through specific embodiments with reference to the accompanying drawings.

[0067] 1. Method Implementation

[0068] Unless otherwise stated, the following tests all use the Swin transformer-s as the feature extraction module and use the parameters pre-trained on ImageNet as weight initialization.

[0069] The weakly supervised small sample target detection system based on positioning pre-training and progressive optimization strategy provided by the present invention includes the following modules:

[0070] (1) Attention-guided positioning pre-training module

[0071] The attention-guided localization pre-training module ALN uses a visual transformer to accurately locate the target. In this embodiment, six multi-head self-attention ViT blocks are stacked to implement the localization pre-training module. This module is added to the third stage of the Swin transformer feature extractor and works in conjunction with the feature extractor.

[0072] (2) Positioning Distillation Learning Module

[0073] In this embodiment, random jitter is applied to the candidate boxes output by the ALN as supervision information to supervise the DETR detector. The bounding box regression loss and single-category classification loss are used to distill the target localization capability from the ALN to the detector, and enable the detector to distinguish the integrity of the foreground.

[0074] (3) Dual-factor driven asymptotic optimization module

[0075] This paper uses a dual-factor driven asymptotic optimization module that combines the classification score and the target integrity score as the basis for pseudo-label selection. This allows for rational selection of pseudo-labels to drive subsequent model optimization. In this embodiment, a three-layer cascaded asymptotic optimization module is employed. Each layer of the cascade structure includes an optimized classifier and an optimized target integrity discriminator. The classification score output by the classifier is multiplied by the target integrity score output by the target discriminator to form the basis for the comprehensive score.

[0076] (4) Mixed loss module

[0077] The present invention proposes a hybrid loss module that integrates positioning pre-training loss, positioning distillation loss, multi-instance learning loss, classification and positioning optimization loss, so that the above components can be precisely combined and coordinated, which can effectively solve the problems existing in existing methods.

[0078] Based on the above detection system, the present invention also provides a weakly supervised small sample target detection method, the steps of which are as follows:

[0079] (1) Prepare pre-training data.

[0080] This step uses the ImageNet dataset commonly used in computer vision tasks as a pre-training dataset. Each image only has image category annotation information, without border annotation. The elements in the dataset are composed of two-tuples such as (Image, Label). In this embodiment, in order to ensure the rigor of the small sample setting, the categories that overlap with the pre-training set and the training and test datasets are removed through screening. After screening, more than 300 categories of overlapping categories were removed from the pre-training dataset.

[0081] (2) Target positioning pre-training.

[0082] This step uses the pre-training dataset to train the localization pre-training module ALN. ALN is composed of 6 multi-head self-attention visual transformer blocks. The original image features are converted into 14*14 image blocks after the first three stages of the swin-transformer feature extractor. A total of 196 (196=14*14) image blocks are labeled t ns With 1 additional category marker t c The splicing is sent to ALN for multi-head self-attention calculation. The calculation process is as follows:

[0083]

[0084] The shape is Take the next N columns to get That is, the image block is marked after attention enhancement, and then a linear layer mapping is used The image block label is given class information, and the linear mapping is trained using a cross-entropy loss function, and the training process is as follows:

[0085]

[0086] The linear layer after training gives class information, and the attention map of the 6 visual transformer blocks is averaged to obtain The first row and the last 196 columns are taken to obtain a class-independent attention map, and the class-independent attention map is multiplied by the image block label containing class information to obtain a class-dependent activation map. The high response value region can be obtained by threshold screening on the class activation map, and then the minimum enclosing rectangle algorithm is used to generate the candidate box. In this embodiment, the screening threshold is set to 0.15. The specific steps are as follows:

[0087]

[0088] (3) Target positioning ability distillation. Although ALN can locate the target, the positioning accuracy is not enough, and it cannot accurately frame the whole target. In order to solve this problem, the present application uses random jitter to data enhance the candidate box output by ALN. Random jitter is applied to the minimum horizontal coordinate x1 of the candidate box, the minimum vertical coordinate y1 of the candidate box, the maximum horizontal coordinate x2 of the candidate box, and the maximum vertical coordinate y2 of the candidate box in four directions. The specific steps are as follows:

[0089] b aug = [x1 ± α1 * w, y1 ± α2 * h, x2 ± α3 * w, y2 ± α4 * h]

[0090] Wherein α is the jitter coefficient, w and h are the width and height of the original candidate box, which are calculated from the original candidate box coordinates (w = x2 - x1, h = y2 - y1). In order to ensure the stability of data enhancement, the value of α (α1, α2, α3, α4) needs to be limited. Then the enhanced box is optimally matched with the model output box and the loss is calculated. The specific steps are as follows:

[0091]

[0092] The pre-trained target positioning module ALN and the DETR target detector are jointly trained by the above formula, the target positioning performance in ALN is transferred to the detector, and the random jitter enhancement technology further improves the target positioning performance of the detector, ensuring that the detection box can cover the whole foreground target. The detector can perform class-independent accurate foreground positioning.

[0093] In this embodiment, to address the discriminative region problem of weakly supervised target detection, the GIoU Loss in the optimal match and border loss is replaced by DIoU Loss. Compared with GIoU Loss, DIoU Loss can better handle the situation where two borders overlap and surround, and can effectively handle the discriminative region problem.

[0094] (4) Dual-factor driven optimization. In the general framework of weakly supervised target detection, image-level classification predictions are obtained by aggregating the results of the category score predictor and the category contribution predictor, that is, multiplying the prediction result matrices of the two and summing them in the category dimension, that is, obtaining the image-level classification prediction results of the original input. Some subsequent work applies a progressive refinement structure to optimize the detection results layer by layer, that is, using the output results of the upper layer as pseudo labels to supervise the predictions of the lower layer. However, these methods usually only use classification scores as the basis for pseudo-label mining, which can easily lead to discriminative region problems. The present invention uses a dual-factor driven optimization strategy that takes into account both the classification score and the target integrity score, overcomes the problem of discriminative regions, and helps to obtain more accurate pseudo labels. Specifically, for each image, a pre-trained detector is used to generate a series of candidate boxes. in are the bounding box coordinates, is the target completeness score, is the classification score, is the detection score. Combining the instance-level classification score with the detection score, we can get the image-level classification prediction score:

[0095]

[0096] The image category annotation is used as supervision information and the cross entropy loss is used to train the multi-instance classifier:

[0097]

[0098] In order to generate more accurate suggestion boxes, the present invention constructs three optimization layers, including an optimized target classifier and an optimized target integrity discriminator. The output of each layer is expressed as:

[0099]

[0100] in, is to optimize bounding box prediction, is the optimization target completeness score, is the optimization classification score. The optimization target integrity score of the K-1 level is and optimize classification scores Combined as the basis for selecting the K-1th level pseudo-label, in this embodiment, vector multiplication is used as the fusion method of the two types of scores:

[0101]

[0102] Next, use the pseudo-label of the K-1th level as supervision information to supervise the output of the Kth level and calculate the classification loss and target integrity loss:

[0103]

[0104] Under the supervision of progressive pseudo-labels, the discriminative information captured by the small candidate boxes will be passed to the overlapping large candidate boxes, while the target completeness information of the large candidate boxes will be simultaneously passed to the small candidate boxes, thereby improving the detection accuracy in both directions.

[0105] (5) Calculation of hybrid loss function.

[0106] To meet the training requirements at different stages, the present invention uses a hybrid loss function for training. Specifically, in the pre-training stage, the loss function is:

[0107]

[0108] In the optimization stage, the loss function is:

[0109]

[0110] By combining these loss functions, different components can work well together. P The value of λ is different in different stages of pre-training. P The value of λ1 is 1 in the first half of pre-training and 0.5 in the second half of pre-training. During the training phase, the value of λ1 is 1 and the value of λ2 is 10.

[0111] Second, testing and verification of the present invention.

[0112] This paper selects two public datasets: ImageNetLoc-FS and CUB-200, to test and verify the performance of this method on weakly supervised small sample object detection datasets;

[0113] ImageNetLoc-FS is a dataset sampled from the ImageNet dataset. It includes both category and bounding box annotations for a total of 331 categories, including 101 base classes, 214 novel classes, and 16 validation classes. All bounding box annotations are used only for model performance evaluation and do not participate in model training. Following the N-way / K-shot approach for few-shot object detection, five categories are randomly selected for the novel classes, and 1 / 5 of the samples in each category are randomly sampled as few-shot training data.

[0114] CUB-200 is a widely used bird dataset that includes both category and bounding box annotations. It contains 200 categories, 100 of which are base classes, 50 are novel classes, and 50 are used for validation. All bounding box annotations are used only for model performance evaluation and do not participate in model training. Following the N-way / K-shot approach in few-shot object detection, we randomly select five categories for the novel classes and randomly sample 1 / 5 of the samples in each category as the few-shot training data.

[0115] To verify the superiority of this method, this embodiment is compared with the following existing weakly supervised target detection methods on public datasets: MetaOpt (from “Kwonjoon Lee, Subhransu Maji, Avinash Ravichandran, and Stefano Soatto. 2019. Meta-learning with differentiable convex optimization. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 10657–10665.”), OICR (from “Peng Tang, Xinggang Wang, XiangBai, and Wenyu Liu. 2017. Multiple instance detection network with onlineinstance classifier refinement. In Proceedings of the IEEE conference on computer vision and pattern recognition. 2843–2851.”), PCL (from “Peng Tang, Xinggang Wang, Song Bai, Wei Shen, Xiang Bai, Wenyu Liu, and Alan Yuille. 2018. Pcl: Proposal clustering”). learning for weakly supervised object detection. IEEE transactions on pattern analysis and machine intelligence 42, 1 (2018), 176–191."), CAN (from "Ruibing Hou, Hong Chang, Bingpeng Ma, Shiguang Shan, and Xilin Chen. 2019. Cross Attention Network for Few-shot Classification. In Proceedings of the Conference and Workshop on Neural Information Processing Systems.4005–4016."), WSOD2 (from "Zhaoyang Zeng, Bei Liu, Jianlong Fu, Hongyang Chao, and LeiZhang. 2019. Wsod2: Learning bottom-up and top-down objectness distillation for weaklysupervised object detection. In Proceedings of the IEEE / CVFinternational conference on computer vision.8292–8300."), StarNet (from "LeonidKarlinsky,Joseph Shtok,Amit Alfassy,Moshe Lichtenstein,Sivan Harary,EliSchwartz,Sivan Doveh,Prasanna Sattigeri,Rogerio Feris,Alex Bronstein,etal.2021.Starnet:towards weakly supervised few-shot object detection.InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 35, pp. 1743–1753). This example uses the average precision (AP30) and AP50 (AP50) under different positioning accuracy settings as evaluation metrics to measure the performance of each algorithm. The experimental results are shown in Table 1.

[0116] The method WFS-DETR of the present invention outperforms all the comparison methods in any sample and all indicators, which proves the effectiveness and superiority of the method of the present invention in the weakly supervised small sample target detection task. For 1-shot, WFS-DETR exceeds StarNet by 8.4% and 18.9% in AP30 and AP50. As the number of training samples increases, the method of the present invention always maintains its advantage. For 5-shot, the method of the present invention is 4.7% and 17.9% higher than StarNet in AP30 and AP50, respectively. Specifically, the performance of the method of the present invention in the AP50 indicator is significantly better than the previous state-of-the-art method, which requires higher positioning accuracy. Experiments show that the method can accurately detect the entire object rather than the part, effectively solving the most challenging problem of distinguishing regions in weakly supervised small sample target detection.

[0117] Table 1 Performance comparison on weakly supervised small sample target detection dataset

[0118]

Claims

1. A weakly supervised small sample target detection system based on positioning pre-training, characterized in that: It is a new target detection system, denoted as WFS-DETR. Based on the transformer-based target detector, it designs and uses an attention-guided positioning pre-training module and a positioning distillation learning module to train the model's target positioning ability, designs and uses a dual-factor driven asymptotic optimization module to train the model's image classification ability, and designs and uses a hybrid loss function module to coordinate different module components and drive the overall end-to-end training. The attention-guided localization pre-training module, denoted as ALN, uses a visual transformer to accurately locate objects. Specifically, the ALN is composed of multiple standard visual transformer blocks, which are connected as a whole after the third stage of the feature extraction network. The features output by the feature extraction network are input into the ALN along with the category label. The characteristics of the visual transformer are then used to output attention score maps of different sizes for the foreground object and background in the image. Based on the attention score maps, the area with the highest score is predicted as the foreground object area. The positioning distillation learning module distills the positioning capability of the positioning pre-training module into the DETR target detector through data augmentation and knowledge distillation. Specifically, the positioning pre-training module ALN is jointly trained with the DETR target detector. Data augmentation is applied to the output of the positioning pre-training module ALN as supervision, and the DETR target detector is trained using foreground discrimination loss and bounding box regression loss. The target positioning capability in the positioning pre-training module ALN is distilled into the DETR target detector. The dual-factor driven asymptotic optimization module jointly considers the classification score and the target integrity score as the basis for pseudo-label selection, rationally selecting pseudo-labels to drive subsequent model optimization. Specifically, the dual-factor driven asymptotic optimization module consists of three optimization sub-blocks, each of which consists of an optimized classifier and an optimized foreground discriminator. During the optimization training process, the classification score and the foreground discriminant score are combined as the basis for pseudo-label selection to improve the quality of pseudo-labels and effectively drive predictive optimization training. The hybrid loss module integrates positioning pre-training loss, positioning distillation loss, multi-instance learning loss, classification optimization loss and positioning optimization loss; specifically, in the pre-training stage, the hybrid loss module is jointly composed of positioning pre-training loss and positioning distillation loss, and its main function is to give the DETR target detector a general target positioning capability; in the optimization training stage, the hybrid loss module is jointly composed of multi-instance learning loss, classification optimization loss and positioning optimization loss, and its main purpose is to give the DETR target detector accurate classification capability through multi-instance learning and optimization training; the hybrid loss module has different functions in different stages. With the help of the hybrid loss module, the DETR locator finally achieves accurate target detection using only weakly labeled data.

2. The weakly supervised small sample target detection system according to claim 1 is characterized in that The ALN is composed of K multi-head self-attention blocks stacked together. The original image features are converted into P*P image blocks after the first three stages of the visual transformer feature extraction network using a sliding window. A total of N image blocks are marked t ns With 1 additional category marker t c The splicing is sent to ALN for multi-head self-attention calculation. The calculation process is as follows: Among them, W Q is the query transformation matrix weight, W K is the contrast transformation matrix weight, W V is the assignment conversion matrix weight, Attention multi It is a multi-head self-attention operation, is the attention score matrix; The shape is Take the next N columns to get That is, the image block labeling after attention enhancement; then use the linear layer mapping Assign category information to the image block labels and use the cross entropy loss function to train the linear mapping. The training process is as follows: Among them, w c With w i is the category mapping matrix weight, T is the matrix transpose operation, GAP is the global average pooling operation; D is the feature dimension of the matrix, N is the total number of image label blocks, and C is the number of categories; After the training, the linear layer is given category information, and the attention map of K visual transformer blocks is averaged to obtain Take the first row and the last N columns to get the category-independent attention map. Multiply the category-independent attention map with the image block label containing category information to get the category-related activation map. Filter the category activation map and set the area with low response value as the background to obtain the high response value area. Then, for all the point coordinates in the high response area, take the maximum and minimum values ​​in the horizontal and vertical directions respectively. Draw the minimum enclosing rectangle surrounding the high response value area based on these extreme points, and generate the candidate box. The specific calculation is as follows: in, is a category-independent attention map, It is an image block label containing category information, Thr is a filtering operation with a threshold of 0.15, and MMR refers to the process of generating a minimum bounding rectangle.

3. The weakly supervised small sample target detection system according to claim 2, characterized in that The positioning distillation learning module distills the positioning capability of the positioning pre-training module into the DETR target detector through data enhancement and knowledge distillation. The specific process is: using random jitter to perform data enhancement on the candidate box output by the pre-trained target positioning module ALN; applying random jitter in four directions to the candidate box's minimum horizontal coordinate x1, the candidate box's minimum vertical coordinate y1, the candidate box's maximum horizontal coordinate x2, and the candidate box's maximum vertical coordinate y2, specifically expressed as: b aug [x1±α1*w,y1±α2*h,x2±α3*w,y2±α4*h], (4) Among them, α(α1,α2,α3,α4) is the jitter coefficient, w and h are the width and height of the original candidate box, which are calculated from the original candidate box coordinates: w = x2-x1, h = y2-y1; then the enhanced box is optimally matched with the model output box and the loss is calculated; specifically expressed as: Among them, y i Refers to the supervision information of the i-th real box and its category, Refers to the prediction information of the i-th prediction box and its category, o i Refers to the foreground integrity supervision information of the i-th real box, Refers to the foreground completeness prediction information of the i-th prediction box, b i Refers to the bounding box coordinate supervision information of the i-th real box, Refers to the prediction information of the bounding box coordinates of the i-th prediction box; The pre-trained target localization module ALN and the DETR target detector are jointly trained by the above formula (6). The target localization performance in ALN ​​is transferred from ALN to the target detector, so that the target detector can perform accurate foreground localization regardless of category.

4. The weakly supervised small sample target detection system according to claim 3, characterized in that The dual-factor driven asymptotic optimization module considers both the classification score and the target integrity score as the basis for pseudo-label selection, and rationally selects pseudo-labels to drive subsequent model optimization. Specifically: For each image, a set of candidate boxes is generated using a DETR detector pre-trained for general object localization capabilities: in, are the bounding box coordinates, is the target completeness score, is the classification score, is the detection score; the instance-level classification score is combined with the detection score to obtain the image-level classification prediction score: The image category label is used as supervision information, and the DETR detector is trained for image classification using cross entropy loss: Among them, y c is the true category supervision information of the image, is the predicted image classification score; In order to generate more accurate proposal boxes, K optimization layers are constructed, including target category score optimization predictors and target completeness score optimization predictors. The output of each layer is expressed as: in, is to optimize bounding box prediction, is the optimization target completeness score, is the optimization classification score; the optimization target integrity score of the K-1 level is and optimize classification scores Combined as the basis for selecting the K-1 level pseudo label: Next, use the pseudo-label of the K-1th level as supervision information to supervise the output of the Kth level and calculate the classification loss and target integrity loss: in, is the classification score prediction result of the K-1 level, is the target integrity score prediction result of the K-1 level, is the classification score prediction result of the K-th level, is the target integrity score prediction result of level K, N r is the number of all prediction boxes that match the supervision information in the K-th level, BCE refers to the two-class cross entropy loss, and CE refers to the cross entropy loss; Under the supervision of progressive pseudo-labels, the discriminative information captured by the small candidate boxes will be passed to the overlapping large candidate boxes, while the target completeness information of the large candidate boxes will be simultaneously passed to the small candidate boxes, improving the detection accuracy in both directions.

5. The weakly supervised small sample target detection system according to claim 4, characterized in that For the hybrid loss module, a hybrid loss function is used for training according to the training requirements of different stages. Specifically, in the pre-training stage, the loss function is: In the optimization stage, the loss function is: Among them, λ P ,λ1,λ2 weight coefficients are determined according to the specific actual situation; By combining these loss functions, different components can be made to work well together.

6. A weakly supervised small sample target detection method based on positioning pre-training, characterized in that: The specific steps are: (1) Prepare pre-training data The ImageNet dataset, commonly used in computer vision tasks, is used as a pre-training dataset. Each image only has image category annotation information, without bounding box annotation. The elements in the dataset are composed of two-tuples such as (Image, Label). (2) Target positioning pre-training; The positioning pre-training module ALN is trained using the pre-training dataset. ALN consists of K multi-head self-attention blocks stacked together. The original image features are converted into P*P image blocks after the first three stages of the sliding window visual transformer feature extraction network. A total of N image blocks are marked t ns With 1 additional category marker t c The splicing is sent to ALN for multi-head self-attention calculation. The calculation process is as follows: Among them, W Q is the query transformation matrix weight, W K is the contrast transformation matrix weight, W V is the assignment conversion matrix weight, Attention multi It is a multi-head self-attention operation, is the attention score matrix; The shape is Take the next N columns to get That is, the image block labeling after attention enhancement; then use the linear layer mapping Assign category information to the image block labels and use the cross entropy loss function to train the linear mapping. The training process is as follows: Among them, w c With w i is the category mapping matrix weight, T is the matrix transpose operation, GAP is the global average pooling operation; D is the feature dimension of the matrix, N is the total number of image label blocks, and C is the number of categories; After the training, the linear layer is given category information, and the attention map of K visual transformer blocks is averaged to obtain Take the first row and the last N columns to get the category-independent attention map. Multiply the category-independent attention map with the image block label containing category information to get the category-related activation map. Filter the category activation map and set the low response value area as the background to obtain the high response value area. Then, for all the point coordinates in the high response area, take the maximum and minimum values ​​in the horizontal and vertical directions respectively. Draw the minimum enclosing rectangle surrounding the high response value area based on these extreme points, that is, generate the candidate box. The specific calculation is as follows: in, is a category-independent attention map, It is the image block label containing category information, Thr is the screening operation with a threshold of 0.15, and MMR refers to the process of generating the minimum bounding rectangle; (3) Target positioning capability distillation Use random jitter to perform data enhancement on the candidate boxes output by the pre-trained target localization module ALN. Apply random jitter in four directions to the candidate box's minimum horizontal coordinate x1, the candidate box's minimum vertical coordinate y1, the candidate box's maximum horizontal coordinate x2, and the candidate box's maximum vertical coordinate y2. The specific calculation is as follows: b aug [x i ±α1*w,y1±α2*h,x2±α3*w,y2±α4*h], (4) Among them, α(α1,α2,α3,α4) is the jitter coefficient, w and h are the width and height of the original candidate box, which are calculated from the original candidate box coordinates: w = x2-x1, h = y2-y1; then the enhanced box is optimally matched with the model output box and the loss is calculated; the specific calculation is as follows: Among them, y i Refers to the supervision information of the i-th real box and its category, Refers to the prediction information of the i-th prediction box and its category, o i Refers to the foreground integrity supervision information of the i-th real box, Refers to the foreground completeness prediction information of the i-th prediction box, b i Refers to the bounding box coordinate supervision information of the i-th real box, Refers to the prediction information of the bounding box coordinates of the i-th prediction box; The pre-trained target localization module ALN and the DETR target detector are jointly trained by the above formula (6). The target localization performance in ALN ​​is transferred from ALN to the detector, so that the target detector can perform accurate foreground localization regardless of category. (4) Dual-factor driven optimization Using a dual-factor driven optimization strategy, we consider both the classification score and the target integrity score, overcome the problem of discriminative regions, and help obtain more accurate pseudo labels. Specifically, for each image, we use the DETR detector that has been pre-trained to obtain general target localization capabilities to generate a series of candidate boxes: in, are the bounding box coordinates, is the target completeness score, is the classification score, is the detection score; combining the instance-level classification score with the detection score gives the image-level classification prediction score, where N is the number of candidate boxes in the image and C is the total number of predicted categories; The image category label is used as supervision information, and the DETR detector is trained for image classification using cross entropy loss: Among them, y c is the true category supervision information of the image, is the predicted image classification score; In order to generate more accurate proposal boxes, K optimization layers are constructed, including target category score optimization predictors and target completeness score optimization predictors. The output of each layer is expressed as: in, is to optimize bounding box prediction, is the optimization target completeness score, is the optimization classification score; the optimization target integrity score of the K-1 level is and optimize classification scores Combined as the basis for selecting the K-1 level pseudo label: Next, use the pseudo-label of the K-1th level as supervision information to supervise the output of the Kth level and calculate the classification loss and target integrity loss: in, is the classification score prediction result of the K-1 level, is the target integrity score prediction result of the K-1 level, is the classification score prediction result of the K-th level, is the target integrity score prediction result of level K, N r is the number of all prediction boxes that match the supervision information in the K-th level, BCE refers to the two-class cross entropy loss, and CE refers to the cross entropy loss; Under the supervision of progressive pseudo-labels, the discriminative information captured by small candidate boxes is passed to the overlapping large candidate boxes, while the target integrity information of the large candidate boxes is simultaneously passed to the small candidate boxes, thus improving the detection accuracy in both directions. (5) Hybrid loss function calculation For the training requirements at different stages, a mixed loss function is used for training; specifically, in the pre-training stage, the loss function is: In the optimization stage, the loss function is: Among them, λ P ,λ1,λ2 weight coefficients are determined according to the specific actual situation; By combining these loss functions, different components can be made to work well together.

Citation Information

Patent Citations

  • Small sample remote sensing image target detection method based on meta-learning and collaborative attention

    CN112818903A

  • Small sample target detection method and system based on support and query samples

    CN113191359A