A small sample target detection method based on an adversarial region proposal network

CN118485889BActive Publication Date: 2026-09-29UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410647164.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-23
Publication Date
2026-09-29
Estimated Expiration
2044-05-23

AI Technical Summary

Technical Problem

通常的基于深度学习的目标检测方法需要在拥有各个待检测类别大量标注的训练样本上进行训练,模型才能具备检测这些类别目标的能力,且难以泛化地检测其他类别样本

Benefits of technology

[0021]本发明所提出的一种基于对抗式区域提议网络的小样本目标检测算法通过使用细粒度分类器作为对抗网络,提高区域提议网络的类别无关性,从而提升基类训练网络的泛化性,提高模型对新类的预测准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118485889B_ABST
    Figure CN118485889B_ABST
Patent Text Reader

Abstract

The application discloses a small sample target detection method based on an adversarial region proposal network and belongs to the fields of small sample learning and target detection.The technical problems solved by the application include that after a model is trained on a large number of labeled base class samples, the model is fine-tuned on a small number of labeled new class training samples, and finally the model can have the detection capability of the new class and the base class simultaneously.In the fine-tuning process, the feature extraction network parameters of the model are frozen, and only the predictor parameters are fine-tuned.Due to the freezing of the region proposal network parameters, the model is insufficient in the capability of detecting new class targets.The application proposes an adversarial region proposal network, the region proposal network is trained to be more general by emphasizing class-independent features in the base class learning process, and thus the detection capability of the new class targets is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of few-shot learning and object detection. Background Technology

[0002] Object detection aims to extract foreground objects from a given image and predict their location and category. Typical deep learning-based object detection methods require training on a large number of labeled training samples for each target category to be detected, and struggle to generalize to other categories. In real-world applications, due to difficulties in sample acquisition or high labeling costs, only a small number of labeled training samples are available for some categories. Few-shot object detection (FSOD) addresses how to train a network to detect new object categories with a large number of base class samples and a small number of new class samples.

[0003] Current FSOD algorithms are mainly divided into two categories: meta-learning-based and fine-tuning-based. Early algorithms primarily used meta-learning methods, learning a paradigm for detecting new samples from a small number of labeled samples to obtain a highly generalized object detection network. Recently, researchers have discovered that fine-tuning methods can also solve the small-sample object detection problem. First, the model is trained using a base class dataset, and then the model predictor is fine-tuned using a small number of new class samples to achieve good detection performance. Building on this, many researchers have proposed different methods to improve the model's knowledge transfer effect, thereby improving the model's detection performance on new classes.

[0004] Adversarial networks (ANNs) are networks that assist the main network in training a model, with the training objective being opposite to that of the main network. A typical application of AANs is Generative Adversarial Networks (GANs). GANs are widely used in image generation tasks and mainly consist of a generator and a discriminator acting as the adversarial network. The generator maps noise signals following a certain distribution to a fake image, while the discriminator is responsible for distinguishing between fake and real images. During training, the generator and discriminator are updated alternately, allowing the distribution of fake images generated by the generator to more closely approximate the distribution of real images. Besides its application in GANs, adversarial networks are also used in domain transfer tasks. Domain transfer aims to apply a network trained on labeled source domain data to unlabeled target domain data. Researchers have proposed using adversarial networks to eliminate domain correlation of intermediate features during the training of source domain data. The training objective of the adversarial network is to discriminate the data domain of the image based on the intermediate features of the main network. During backpropagation, the main network parameters are updated by reversing gradients, thereby gradually reducing the domain correlation of features during training.

[0005] Much research focuses on the fine-tuning process of models, designing methods to enable models to better learn detection of new classes based on knowledge learned from base classes. From another perspective, if the generalization ability of learned knowledge can be improved during base class training, the detection knowledge from base classes can be better utilized to learn detection of new classes during the fine-tuning phase. Region proposal networks, as an important component of two-stage object detection models, can filter anchor boxes to find possible foreground objects. Optimizing the generalization ability of region proposal networks allows the network to autonomously discover potential foreground objects, thereby optimizing the knowledge transfer ability during training for new classes. This invention introduces adversarial networks into the training process of few-shot object detection, designing an adversarial region proposal network to improve the generalization ability of the few-shot object detection model during base class training. Summary of the Invention

[0006] The technical problem addressed by this invention is as follows: After training a model on a large number of labeled base class samples, it is fine-tuned on a small number of labeled new class training samples, ultimately enabling the model to detect both new and base classes simultaneously. During fine-tuning, the feature extraction network parameters are frozen, and only the predictor parameters are fine-tuned. Due to the frozen region proposal network parameters, the model's ability to detect new class targets is insufficient. This invention proposes an adversarial region proposal network, which trains a more generalized region proposal network by emphasizing class-independent features during base class learning, thereby improving the detection capability for new class targets.

[0007] The technical solution of this invention is a few-shot target detection method based on an adversarial region proposal network, the method comprising:

[0008] Step 1: Initialize a two-stage object detector, which includes: a backbone network, an adversarial region proposal network, a RoI pooling and feature extraction module, and a predictor; the input image is first input into the backbone network, and the output of the backbone network is split into two paths, one of which is input into the adversarial region proposal network, and the other path, together with the output of the adversarial region proposal network, is input into the RoI pooling and feature extraction module. The output of the RoI pooling and feature extraction module is input into the predictor;

[0009] The region proposal network includes: a feature extraction layer F and a foreground / background classifier C. fb Location regressor R, channel mask layer M, fine-grained classifier C fine ;

[0010] Step 2: During training, the feature extraction layer F further extracts features from the feature spectrum extracted from the backbone network to obtain the feature spectrum f. This feature spectrum is weighted by the channel mask layer to obtain a new feature spectrum f′. The new feature spectrum is input into the foreground / background classifier and the position regressor for foreground / background classification and position prediction.

[0011] Step 3: The feature spectrum f′ is fed into a fine-grained classifier after passing through a gradient inversion layer to perform a specific category classification task on the anchor boxes containing foreground targets;

[0012] Step 4: During gradient backpropagation, the foreground and background classifiers C... fb The position regressor R normally sends back the gradient to the channel mask layer M, while the gradient inversion layer inverts the gradient before sending it back to the channel mask layer M.

[0013] Step 5: Train the two-stage target detector on the base class samples using the methods from Steps 1 to 4 to obtain preliminary training results;

[0014] Step 6: Freeze the backbone network, adversarial region proposal network, RoI pooling and feature extraction module of the two-stage object detector obtained in Step 5, and fine-tune the predictor using a small number of new class samples; complete the final training.

[0015] Step 7: Use the trained two-stage target detector to detect actual targets.

[0016] Furthermore, in step 2, f ′ =M(f)·f, where M(f) is the channel mask layer M that weights f, and "·" indicates element-wise multiplication.

[0017] Furthermore, the overall loss L of the adversarial region proposal network RPN for:

[0018] L RPN =L fb +L r +λL fine

[0019] Among them, L fb L represents the foreground / background classification loss. r L represents the regression loss. fine λ represents the fine-grained classification loss, and λ represents the scaling factor.

[0020] Furthermore, the L fb For cross-entropy loss, L r For L1 loss, L fine This represents the cross-entropy loss.

[0021] The proposed algorithm for few-shot target detection based on adversarial region proposal networks uses a fine-grained classifier as an adversarial network to improve the class independence of the region proposal network, thereby enhancing the generalization of the base class training network and improving the model's prediction accuracy for new classes. Attached Figure Description

[0022] Figure 1 This is a flowchart of the training process of the present invention;

[0023] Figure 2 This is a diagram of the adversarial region proposal network model framework proposed in this invention. Implementation

[0024] Step 1: Define a Faster R-CNN object detection network, where the backbone network is initialized with ImageNet pre-trained parameters, and the other parts are randomly initialized; the region proposal network adopts an adversarial region proposal network framework;

[0025] Step 2: Train the network using base class samples. After inputting the image into the network, the feature spectrum f is extracted through the backbone network. main The backbone feature spectrum is input into the region proposal network;

[0026] Step 3: The feature extraction layer of the region proposal network further extracts features to obtain the feature spectrum f. The feature spectrum is weighted by the channel masking layer M to obtain a new feature spectrum f. ′ =M(f)·f, where · represents element-wise multiplication, and the mask layer is a 1×1 convolutional layer with the number of output channels equal to the number of input channels;

[0027] Step 4: Input the feature spectrum f′ into the foreground / background classifier to calculate the foreground confidence for each anchor box, and use cross-entropy loss to calculate the foreground / background classification loss L. fb The input regressor predicts the offset between the anchor frame and the target position, and the regression loss L is calculated using L1 loss. r ;

[0028] Step 5: The feature spectrum f′ is input into the fine-grained classifier after passing through the gradient inversion layer. The fine-grained classifier is a multilayer perceptron consisting of a linear layer, a ReLU activation layer, and another linear layer. The fine-grained classification loss L is calculated using cross-entropy loss. fine The overall loss of the adversarial region proposal network is expressed as:

[0029] L RPN =L fb +L r +λL fine

[0030] Where λ is the scaling factor, which can be 0.5;

[0031] Step 6: The region proposal network generates candidate boxes based on the foreground confidence and then uses these candidate boxes to construct a new region proposal network based on the feature spectrum f. main Position pooling is used to obtain the features of the Region of Interest (RoI), and further features are extracted and the category and location are predicted, and the loss is calculated;

[0032] Step 7: Update parameters through gradient backpropagation. Except for the fine-grained classification loss, other losses are backpropagated in a regular manner. When the fine-grained classification loss is backpropagated to the gradient inversion layer, the gradient is inverted before it is backpropagated. The backpropagated gradient is only used to update the channel mask layer and is not backpropagated to the feature extraction layer.

[0033] Step 8: After the base class training is completed, freeze the model feature extraction network parameters, and fine-tune the predictor using a small number of new class samples to obtain a detection model that can detect new class targets.

Claims

1. A few-shot target detection method based on adversarial region proposal networks, the method comprising: Step 1: Initialize a two-stage object detector, which includes: a backbone network, an adversarial region proposal network, a RoI pooling and feature extraction module, and a predictor; the input image is first input into the backbone network, and the output of the backbone network is split into two paths, one of which is input into the adversarial region proposal network, and the other path, together with the output of the adversarial region proposal network, is input into the RoI pooling and feature extraction module. The output of the RoI pooling and feature extraction module is input into the predictor; The region proposal network includes: a feature extraction layer F and a foreground / background classifier C. fb Location regressor R, channel mask layer M, fine-grained classifier C fine ; Step 2: During training, the feature extraction layer F further extracts features from the feature spectrum extracted from the backbone network to obtain the feature spectrum f. This feature spectrum is weighted by the channel mask layer to obtain a new feature spectrum f′. The new feature spectrum is input into the foreground / background classifier and the position regressor for foreground / background classification and position prediction. Step 3: The new feature spectrum f′ is fed into a fine-grained classifier after passing through a gradient inversion layer to perform a specific category classification task on the anchor boxes containing foreground targets; Step 4: During gradient backpropagation, the foreground and background classifiers C... fb The position regressor R normally sends back the gradient to the channel mask layer M, while the gradient inversion layer inverts the gradient before sending it back to the channel mask layer M. Step 5: Train the two-stage target detector on the base class samples using the methods from Steps 1 to 4 to obtain preliminary training results; Step 6: Freeze the backbone network, adversarial region proposal network, RoI pooling and feature extraction module of the two-stage object detector obtained in Step 5, and fine-tune the predictor using a small number of new class samples; complete the final training. Step 7: Use the trained two-stage target detector to detect actual targets.

2. The few-sample target detection method based on adversarial region proposal networks as described in claim 1, characterized in that, In step 2, f ′ =M(f)·f, where M(f) is the channel mask layer M that weights f, and "·" indicates element-wise multiplication.

3. The few-sample target detection method based on adversarial region proposal networks as described in claim 1, characterized in that, The overall loss of the adversarial region proposal network L RPN for: THE RPN =L fb +L r +λL fine Among them, L fb L represents the foreground / background classification loss. r L represents the regression loss. fine λ represents the fine-grained classification loss, and λ represents the scaling factor.

4. The few-sample target detection method based on adversarial region proposal networks as described in claim 3, characterized in that, The L fb For cross-entropy loss, L r For L1 loss, L fine This represents the cross-entropy loss.

Citation Information

Patent Citations

  • Target detection method and device based on traffic scene, equipment and storage medium

    CN115527070A

  • Continuous few-sample target detection method based on category registration mechanism and regional contrast learning

    CN117292112A