Small-Sample Object Detection Method Based on Increment of Target Proposal Boxes

By building and optimizing the deep network model for object detection, and using clustering and data augmentation technology to increase the number of suggested boxes, the problem of low recognition accuracy in small sample object detection is solved, and higher recognition accuracy and generalization effect is achieved.

CN115439645BActive Publication Date: 2025-06-24NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210711361.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-22
Publication Date
2025-06-24
Estimated Expiration
2042-06-22

AI Technical Summary

Technical Problem

The existing two-stage fine-tuning method is used to detect and recognize small sample object. Due to the small sample object number, the recognition accuracy is low, and it is difficult to effectively improve the recognition accuracy of small sample object detection.

Method used

By building a deep network model for object detection, training the model with the basic category data set, and obtaining H1 and H2 classes through feature extraction and clustering, optimizing the parameters of the classifier and regressor modules in the deep network model for object detection, increasing the number of suggested boxes, and improving the generalization effect of the small sample object detection network model.

Benefits of technology

The generalization effect of the small sample object detection network model is improved, the number of suggestions is increased, the overall applicability to the small sample data set is improved, and the recognition accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115439645B_ABST
    Figure CN115439645B_ABST
Patent Text Reader

Abstract

The present invention discloses a small-sample object detection method based on the increment of target proposal boxes, including: Step 1, constructing and training an object detection deep network model; Step 2, feature extraction and clustering; Step 3, optimizing the object detection deep network model; Step 4, screening target proposal boxes by using the optimized object detection deep network model; Step 5, enhancing and screening the unselected positive samples to obtain supplementary proposal boxes; Step 6, obtaining the final small-sample object detection network model; Step 7, obtaining the small-sample image to be detected; Step 8, obtaining the object detection result. The present invention optimizes the small-sample object detection network model, improves the overall applicability to the small-sample data set, and fully considers the similarity between the unselected positive samples and the H1 classes obtained by the first clustering and the H2 classes obtained by the second clustering, obtains supplementary proposal boxes, increases the number of proposal boxes, and improves the generalization effect of the small-sample object detection network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a small-sample object detection method based on the increment of target proposal boxes. Background Art

[0002] The task of object detection is to determine whether there are objects of interest in an image, that is, object detection and recognition, and then accurately locate the objects of interest. As a basic computer vision problem, object detection and recognition are widely used in fields such as autonomous driving, medical imaging, image classification, image detection, image segmentation, and industrial inspection, and high detection accuracy and positioning accuracy are required in each field.

[0003] Currently, methods for object detection generally belong to machine learning-based methods or deep learning-based methods. With the rise of deep learning, the research on convolutional neural network (CNN) has been further developed. As one of the representative technologies of artificial intelligence, it has achieved unprecedented breakthroughs and achievements, demonstrating the dominant position of the convolutional neural network in pattern recognition algorithms. Deep learning generally requires a large amount of labeled data for training.

[0004] RCNN in 2014 is a classic work on object detection using deep learning, which kicked off the prelude to object detection using deep learning. Fast RCNN in 2015 achieved end-to-end detection and convolution sharing based on RCNN. Faster RCNN proposed the epoch-making idea of anchor boxes, pushing object detection to the first peak. In object detection algorithms, the object bounding box changes from none to some, and the process of bounding box change to a certain extent reflects whether the detection is first-order or second-order. Second-order algorithms usually focus on finding the positions where objects appear in the first stage to obtain proposal boxes and ensure a sufficient precision-recall rate, and then focus on classifying the proposal boxes and finding more accurate positions in the second stage. Typical algorithms include FasterRCNN.

[0005] Traditional two-stage fine-tuning methods use Faster RCNN as the basic framework for small-sample object detection, and adopt a two-stage training method - the training set in the first stage is a large amount of labeled basic category data, and the second stage uses a small amount of basic categories and new categories for fine-tuning. The TFA method has relatively great potential in small-sample object detection. The TFA method simply freezes the network parameters learned on the basic category data, and then fine-tunes the last two fully connected layers on the new category data, that is, classification and bounding box regression, and all other structures are frozen. Subsequently, it was found that good results can also be obtained without freezing the model structure and by adopting appropriate training.

[0006] Through experiments, it is observed that for small-sample target detection and recognition, the number of suggestion boxes for positive samples in the fine-tuning stage is only 1 / 4 of that in basic training. Many suggestion boxes containing new categories have low confidence and will be filtered out by NMS, resulting in a decrease in the number of foreground suggestion boxes and fewer opportunities for the network to learn new categories.

[0007] In summary, the two-stage fine-tuning method has been initially applied in the field of target detection and recognition, but it has not yet been well used to solve the problem of low recognition accuracy in small-sample target detection and recognition due to the small number of foreground proposal boxes. Improving the accuracy of small-sample target detection and recognition is an important issue that needs to be solved urgently. Summary of the invention

[0008] The technical problem to be solved by the present invention is to provide a small sample target detection method based on target suggestion box increment in response to the above-mentioned deficiencies in the prior art. The method has a simple structure and a reasonable design, optimizes the small sample target detection network model, improves the overall applicability to small sample data sets, and fully considers the similarity between the unscreened positive samples and the H1 classes obtained by the first clustering and the H2 classes obtained by the second clustering, and screens out supplementary suggestion boxes, thereby increasing the number of suggestion boxes and improving the generalization effect of the small sample target detection network model.

[0009] In order to solve the above technical problems, the technical solution adopted by the present invention is: a small sample target detection method based on target suggestion box increment, characterized in that it includes the following steps: step 1, constructing a target detection deep network model, and using a basic category data set to train the target detection deep network model;

[0010] Step 2: Use the trained target detection deep network model for feature extraction and clustering:

[0011] Step 201, extracting feature vectors from basic category data samples and clustering them into H1 categories;

[0012] Step 202: randomly crop the basic category data samples to obtain base category data blocks, extract features from the obtained base category data blocks and cluster them into H2 categories;

[0013] Step 3: Use the new category data set to optimize the parameters of the classifier and regressor modules in the target detection deep network model to obtain an optimized target detection deep network model;

[0014] Step 4: Use the optimized target detection deep network model to screen target suggestion boxes:

[0015] Obtain a sampled data set, input the sampled data set into an optimized object detection deep network model to obtain k positive samples, use the non-maximum suppression method to screen the object proposal boxes of the k positive samples, and select k1 positive sample object proposal boxes with high confidence as the object proposal boxes. The remaining k2 positive samples that are not screened, where k1 + k2 = k;

[0016] Step Five: Enhance and screen the positive samples that are not screened to obtain supplementary proposal boxes:

[0017] Step 501: Calculate the similarity P of the j-th positive sample that is not screened with H1 classes j1 , P j1 represents the first foreground score of the j-th positive sample that is not screened, where 1 ≤ j ≤ k2;

[0018] Step 502: Calculate the similarity P of the j-th positive sample that is not screened with H2 classes j2 , P j2 represents the second foreground score of the j-th positive sample that is not screened;

[0019] Step 503: The computer calculates according to the formula P j = ω1·P j1 + ω2·P j2 to perform weighted summation on the first foreground score and the second foreground score of the j-th positive sample that is not screened. P j represents the comprehensive foreground score, ω1 represents the first weight, ω2 represents the second weight, and ω1 + ω2 = 1;

[0020] Step 504: Screen P j of the comprehensive foreground score according to the comprehensive foreground score threshold, and use the j-th positive sample that is not screened corresponding to the screened comprehensive foreground score P j as the supplementary proposal box. The supplementary proposal box and the object proposal box form the prediction proposal box;

[0021] Step Six: Use the obtained prediction proposal box to update the parameters of the classifier and regressor modules in the optimized object detection deep network model in Step Three to obtain the final small-sample object detection network model;

[0022] Step Seven: Obtain the small-sample image to be detected;

[0023] Step Eight: Input the small-sample image to be detected into the final small-sample object detection network model to obtain the object detection result.

[0024] The above small-sample object detection method based on object proposal box increment is characterized in that: The specific method of Step Three includes:

[0025] Step 301: Delete the parameters of the classifier network and the regressor network module of the trained object detection deep network model in Step 1, and re-initialize the parameters of the classifier network and the regressor network module randomly;

[0026] Step 302: Obtain new category data samples, and the new category data samples and the basic category data samples constitute a balanced data set;

[0027] Step 303: Perform data augmentation on each sample in the balanced data set respectively to obtain an augmented data set;

[0028] Step 304: Use the augmented data set to optimize the parameters of the classifier and regressor modules in the object detection deep network model to obtain an optimized object detection deep network model.

[0029] The above small-sample object detection method based on object proposal box increment is characterized in that: the specific method of performing data augmentation on each sample in the balanced data set in Step 303 includes: randomly linearly combining the q-th sample and the p-th sample according to randomly generated linear combination coefficients to obtain a first input image, and performing multiple random linear combinations in a loop to obtain multiple first input images, where 1 ≤ q ≤ K, 1 ≤ p ≤ K, and p ≠ q; the multiple first input images constitute the augmented data set.

[0030] The above small-sample object detection method based on object proposal box increment is characterized in that: the specific method of performing data augmentation on each sample in the balanced data set in Step 303 includes: randomly cropping a data block on the q-th sample, and randomly superimposing the data block on the p-th sample to obtain a second input image, and performing multiple random superimpositions in a loop to obtain multiple second input images, where 1 ≤ q ≤ K, 1 ≤ p ≤ K, and p ≠ q; the multiple second input images constitute the augmented data set.

[0031] The above small-sample object detection method based on object proposal box increment is characterized in that: the specific method of performing data augmentation on each sample in the balanced data set in Step 303 includes:

[0032] Step 3031: Randomly linearly combine the q-th sample and the p-th sample according to randomly generated linear combination coefficients to obtain a first input image, and perform multiple random linear combinations in a loop to obtain multiple first input images, where 1 ≤ q ≤ K, 1 ≤ p ≤ K, and p ≠ q;

[0033] Step 3032: Randomly crop a data block on the q-th sample, and randomly superimpose the data block on the p-th sample to obtain a second input image, and perform multiple random superimpositions in a loop to obtain multiple second input images;

[0034] Step 3033: Multiple first input pictures and multiple second input pictures form an enhanced data set.

[0035] The above small-sample object detection method based on target proposal box increment is characterized in that: The specific method of step 201 includes:

[0036] Step 2011: Input n basic category data samples into the trained object detection deep network model for feature vector extraction to obtain the feature vectors of each basic category data sample;

[0037] Step 2012: Use the K-means clustering algorithm to perform clustering operations on the feature vectors of the basic category data samples, and cluster them into H1 classes.

[0038] The above small-sample object detection method based on target proposal box increment is characterized in that: The specific method of step 202 includes:

[0039] Step 2021: Randomly crop n basic category data samples L times respectively to obtain n×L basic class data blocks, and input the n×L basic class data blocks into the trained object detection deep network model for feature extraction to obtain n×L feature vectors;

[0040] Step 2022: Use the K-means clustering algorithm to perform clustering operations on the n×L feature vectors, and cluster them into H2 classes.

[0041] The present invention has the following advantages compared with the prior art:

[0042] 1. The structure of the present invention is simple, reasonably designed, and convenient to implement and use.

[0043] 2. The present invention uses the new category data set to optimize the parameters of the classifier and regressor modules in the object detection deep network model to obtain an optimized object detection deep network model, thereby improving the overall applicability of the optimized object detection deep network model to the small-sample data set, improving the recognition accuracy of the target in the case of a small number of labelable samples, and providing a basis for subsequent sample screening.

[0044] 3. The present invention fully considers the similarity between the unselected positive samples and the H1 classes obtained by the first clustering and the H2 classes obtained by the second clustering, calculates the comprehensive foreground score by weighted summation, and obtains supplementary proposal boxes through the screening of the comprehensive foreground score, thereby increasing the number of proposal boxes and improving the generalization effect of the small-sample object detection network model.

[0045] In summary, the structure of the present invention is simple and reasonably designed, optimizing the small-sample object detection network model, improving the overall applicability to small-sample data sets, and fully considering the similarity between the unselected positive samples and the H1 classes obtained by the first clustering and the H2 classes obtained by the second clustering, screening to obtain supplementary proposal boxes, thereby increasing the number of proposal boxes and improving the generalization effect of the small-sample object detection network model.

[0046] Next, through the accompanying drawings and embodiments, the technical solutions of the present invention will be further described in detail. Brief Description of the Drawings

[0047] Figure 1 It is a flowchart of the method of the present invention. Detailed Embodiments

[0048] Next, the method of the present invention will be further described in detail with reference to the accompanying drawings and embodiments of the present invention.

[0049] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the drawings and embodiments.

[0050] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0051] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0052] For ease of description, spatial relative terms such as "above", "over", "on the upper surface", "upper", etc. can be used here to describe the spatial positional relationship between a device or feature shown in the figure and other devices or features. It should be understood that spatial relative terms are intended to encompass different orientations in use or operation other than the orientation depicted in the figure for the device. For example, if the device in the attached drawing is inverted, a device described as "above" or "over" other devices or structures will then be positioned "below" or "under" other devices or structures. Thus, the exemplary term "above" can include both the orientations of "above" and "below". The device can also be positioned in other different ways (rotated 90 degrees or in other orientations), and corresponding interpretations are made for the spatial relative descriptions used here.

[0053] Embodiment 1

[0054] As Figure 1 shown, the small-sample object detection method based on the increment of target proposal boxes of the present invention includes the following steps:

[0055] Step 1: Construct an object detection deep network model and train the object detection deep network model using a basic category dataset.

[0056] To improve the representation ability of the small-sample object detection network model and reduce the training time, the small-sample object detection network model uses a general two-stage object detection algorithm in deep learning. Constructing and training the object detection deep network model belongs to the first stage of the two-stage object detection algorithm. In this embodiment, the Faster R-CNN network model is selected for the object detection deep network model, and the basic category dataset includes n basic category data samples. The Faster R-CNN network model is trained using the n basic category data samples. Faster R-CNN is a fast region convolutional neural network. The Faster R-CNN network model can be built using the Faster R-CNN structure in the prior art. The Faster R-CNN network model includes: a feature extraction network, a region proposal network, a region feature map generation network, a classifier network, and a regressor network.

[0057] Step 2: Use the trained object detection deep network model for feature extraction and clustering:

[0058] Step 201: Extract feature vectors from the basic category data samples and cluster them into H1 classes.

[0059] In actual use, Step 201 includes:

[0060] Step 2011: Input n basic category data samples into the trained object detection deep network model for feature vector extraction to obtain the feature vectors of each basic category data sample.

[0061] Step 2012: Use the K-means clustering algorithm to perform clustering operations on the feature vectors of the basic category data samples, clustering them into H1 classes.

[0062] In actual use, the Faster R-CNN network extracts images through the convolutional layer to obtain a feature map, performs object detection and precise positioning on the feature map through the RPN network to obtain candidate boxes, and performs max pooling operations on the candidate boxes through the RoI pooling layer in the Faster R-CNN network to output a feature set including multiple feature vectors, thus completing feature vector extraction and obtaining the feature vectors of each basic category data sample. The basic idea of image pixel clustering is to group similar pixels into clusters. The features in the same cluster are similar, and the features in different clusters are different. Classify the positive samples according to the clustering results to obtain H1 classes.

[0063] Step 202: Randomly crop the basic category data samples to obtain base class data blocks, perform feature extraction and clustering on the obtained base class data blocks, and cluster them into H2 classes.

[0064] Step 2021: Randomly crop n basic category data samples L times respectively to obtain n×L base class data blocks, and input the n×L base class data blocks into the trained object detection deep network model respectively for feature extraction to obtain n×L feature vectors.

[0065] In actual use, random cropping can meet the segmentation requirements of real-world scenarios. The random cropping data augmentation algorithm crops the basic category data samples in a random position and size manner, making the distribution of the basic category data samples more uniform, and at the same time alleviating the dependence of deep learning on labeled data, and having better generalization ability for images in real scenarios. Therefore, add the step of random cropping and then perform clustering.

[0066] Step 2022: Use the K-means clustering algorithm to perform clustering operations on the n×L feature vectors, clustering them into H2 classes.

[0067] In the first stage, after constructing the object detection deep network model in step one, first use the basic category data set to train the object detection deep network model to obtain the trained object detection deep network model, and then use the clustering algorithm in step two to perform two clustering operations on the feature vectors extracted from the samples in the basic category data set respectively. The two clustering operations obtain H1 classes and H2 classes respectively.

[0068] Step 3. Optimize the parameters of the classifier and regressor modules in the object detection deep network model using the new category dataset to obtain an optimized object detection deep network model:

[0069] Step 301. Delete the parameters of the classifier network and regressor network modules of the trained object detection deep network model in Step 1, and re-initialize the parameters of the classifier network and regressor network modules randomly.

[0070] In actual use, fine-tune the last layer of the Faster R-CNN network model and keep the parameters of other layers unchanged, that is, fine-tune the parameters of the classifier network and regressor network using the new category data samples.

[0071] Step 302. Obtain new category data samples, and the new category data samples and basic category data samples form a balanced dataset;

[0072] In actual use, collect new category data samples and annotate the collected new category data samples. By adding new category data samples, the data and categories of the balanced dataset are expanded. Optimize the object detection deep network model based on the balanced dataset, combine the generalization ability of the new category data samples for the small sample dataset with the ability of the Faster R-CNN network model to extract features with a small amount of sample data, so that the object detection deep network model can identify data samples that do not appear in the balanced dataset, greatly improving the recognition ability of the object detection deep network model for small sample data. Thus, the applicability of the object detection deep network model to the small sample dataset as a whole is improved, and a basis for subsequent sample screening is provided.

[0073] Step 303. Perform data augmentation on each sample in the balanced dataset respectively to obtain an augmented dataset:

[0074] Step 3031. Randomly linearly combine the q-th sample and the p-th sample according to the randomly generated linear combination coefficients to obtain the first input image. Perform the random linear combination multiple times in a loop to obtain multiple first input images, where 1 ≤ q ≤ K, 1 ≤ p ≤ K, and p ≠ q. MixUp is a simple and effective data augmentation method. MixUp is to generate new sample-label corresponding data by weighting two sample-label corresponding data according to a set ratio, that is, the first input image.

[0075] Step 3032. Randomly crop a data block on the q-th sample and randomly stack the data block on the p-th sample to obtain the second input image. Perform the random stacking multiple times in a loop to obtain multiple second input images. CutMix means regarding data augmentation as a positive sample, that is, the second input image is highly similar to the original image.

[0076] Step 3033: Multiple first input images and multiple second input images constitute an enhanced dataset.

[0077] The method of data augmentation is used to solve the problem of insufficient data volume. Data augmentation actually makes the enhanced data different from the original data to facilitate the simulation of abnormal data. Using multiple data augmentation methods is better than using a single data augmentation method. Therefore, the data augmentation method of MixUp used in step 3031 and the data augmentation method of CutMix used in step 3032 are added, resulting in more data augmentation methods and improving the generalization ability of the model.

[0078] Step 304: Use the enhanced dataset to optimize the parameters of the classifier and regressor modules in the object detection deep network model to obtain an optimized object detection deep network model.

[0079] Mixup optimizes the parameters of the classifier and regressor modules in the object detection deep network model on the first input images, which can reduce the memory of wrong labels, increase the robustness to adversarial samples, and stabilize the training process of the object detection deep network model. The second input images have two different image informations, enabling the object detection deep network model to identify two objects from a local perspective in the second input images, improving the training efficiency and having better classification performance and object localization functions. In actual use, if the parameters of the classifier network and regressor network modules randomly initialized are not much different from the parameters of the object detection deep network model trained in step one, then the combined data augmentation method of MixUp data augmentation and CutMix data augmentation is adopted to provide more training information for the training of the object detection deep network model, increase the number of negative samples, and at the same time improve the robustness of the object detection deep network model, making the optimized object detection deep network model more stable and having better classification performance.

[0080] In the second stage, first, use the new category dataset in step three for the fine-tuning operation of the object detection deep network model, and use the small sample object detection network model with a two-stage training method as the optimized object detection deep network model. Among them, the basic category dataset has a large number of labeled samples, while the new category dataset has only a small number of labeled samples. When fine-tuning and updating the parameters of the object detection deep network model, only the parameters of the classifier and regressor modules in the object detection deep network model are changed, and the feature extraction ability of the object detection deep network model trained through the basic category dataset is transferred to the new category dataset with only a small number of training samples, enabling the new category dataset to still have good expression ability even when the training data is insufficient and improving the recognition accuracy of the new category data.

[0081] Step Four: Use the optimized object detection deep network model to screen object proposal boxes:

[0082] Obtain a sampled data set, input the sampled data set into an optimized object detection deep network model to obtain k positive samples, use the non-maximum suppression method to screen the object proposal boxes of the k positive samples, select k1 positive sample object proposal boxes with high confidence as the object proposal boxes, and the remaining are k2 positive samples that are not screened, where k1 + k2 = k.

[0083] In this embodiment, obtain a sampled data set, use the sampled data set as the input of the optimized object detection deep network model, and the optimized object detection deep network model outputs the k positive samples and the confidence levels of the k positive samples obtained by screening. In actual use, the optimized object detection deep network model uses a feature extraction network to extract the feature map of the sampled data set, and inputs the feature map to a candidate region extraction network and a classification and regression network. In the candidate region extraction network, the feature map is further processed, and the further processed feature map is sent to the classification and regression network to classify the image and output the confidence levels of the candidate regions corresponding to various categories. Samples with a confidence level greater than the category confidence threshold are used as positive samples.

[0084] Step Five: Enhance and screen the positive samples that are not screened to obtain supplementary proposal boxes:

[0085] Step 501: Calculate the similarity P between the jth positive sample that is not screened and H1 classes j1 , P j1 represents the first foreground score of the jth positive sample that is not screened, where 1 ≤ j ≤ k2; it should be noted that the calculation methods of similarity include the cosine of the angle method, the Pearson correlation coefficient method, the Euclidean distance, and the DTW dynamic time warping algorithm.

[0086] Step 502: Calculate the similarity P between the jth positive sample that is not screened and H2 classes j2 , P j2 represents the second foreground score of the jth positive sample that is not screened.

[0087] Step 503: The computer calculates according to the formula P j = ω1·P j1 + ω2·P j2 to perform weighted summation on the first foreground score and the second foreground score of the jth positive sample that is not screened, where P j represents the comprehensive foreground score, ω1 represents the first weight, ω2 represents the second weight, and ω1 + ω2 = 1.

[0088] The first foreground score and the second foreground score are weightedly summed as the comprehensive foreground score. This process fully considers the similarity between the unscreened positive samples and the H1 classes obtained by the first clustering and the H2 classes obtained by the second clustering, thereby improving the accuracy of the calculation of the comprehensive foreground score. In addition, there is no front-to-back dependency between the calculation processes of the first foreground score and the second foreground score. Therefore, the first foreground score and the second foreground score can be calculated in parallel, with fast calculation speed and good use effect.

[0089] Step 504: Calculate the comprehensive foreground score P according to the comprehensive foreground score threshold. j Screening is performed and the comprehensive prospect score P that passes the screening is j The corresponding jth unscreened positive sample is used as a supplementary suggestion box, and the supplementary suggestion box and the target suggestion box constitute the predicted suggestion box. The positive samples that have not passed the screening are rescreened to obtain supplementary suggestion boxes, thereby increasing the number of suggestion boxes.

[0090] In the second stage, conventional two-stage fine-tuning is used, and the number of target proposal boxes obtained is too small, that is, the number of k1 is too small, resulting in poor generalization effect of small sample target detection network model.

[0091] Therefore, first, according to step 4, the optimized target detection deep network model is used to screen the target suggestion boxes to obtain a certain number of target suggestion boxes. Then, the H1 class and H2 class obtained by clustering the basic category data set twice in the first stage are used as the standard. According to the similarity between the unscreened positive samples and the H1 class, the first foreground score is obtained, and according to the similarity between the unscreened positive samples and the H2 class, the second foreground score is obtained. The sum of the first foreground score and the second foreground score multiplied by the corresponding weights is the comprehensive foreground score. In actual use, the weights can be randomly generated. The unscreened positive samples with high comprehensive foreground scores are used as supplementary suggestion boxes, and the target suggestion boxes obtained in step 4 are added to form the predicted suggestion boxes, so the number of predicted suggestion boxes is increased.

[0092] Step 6: Use the obtained prediction suggestion box to update the parameters of the classifier and regressor modules in the optimized target detection deep network model to obtain the final small sample target detection network model.

[0093] Step 7: Obtain a small sample image to be detected.

[0094] Step 8: Input the small sample image to be detected into the final small sample target detection network model to obtain the target detection result.

[0095] Embodiment 2

[0096] Different from the first embodiment, the specific method for data augmentation of each sample in the balanced dataset in step 303 includes: randomly linearly combining the q-th sample and the p-th sample according to randomly generated linear combination coefficients to obtain a first input image, and repeatedly performing random linear combination multiple times to obtain multiple first input images, where 1 ≤ q ≤ K, 1 ≤ p ≤ K, and p ≠ q; the multiple first input images constitute the augmented dataset.

[0097] In actual use, in the fine-tuning stage, considering the influence of re-randomly initializing the parameters of the classifier network and the regressor network module on the features of the small-sample object detection network model, the data augmentation method is correspondingly selected, which not only improves the training speed of the model but also reduces the computational amount, and has a good usage effect. If the parameters of the re-randomly initialized classifier network and regressor network module are quite different from the parameters of the small-sample object detection network model trained in step one, the MixUp data augmentation method is adopted to optimize the parameters of the classifier and regressor modules in the small-sample object detection network model, making the optimized object detection deep network model more stable and having better classification performance.

[0098] Embodiment Three

[0099] Different from the first and second embodiments, the specific method for data augmentation of each sample in the balanced dataset in step 303 includes: randomly cropping a data block on the q-th sample and randomly superimposing the data block on the p-th sample to obtain a second input image, and repeatedly performing random superimposition multiple times to obtain multiple second input images, where 1 ≤ q ≤ K, 1 ≤ p ≤ K, and p ≠ q; the multiple second input images constitute the augmented dataset.

[0100] In actual use, in the fine-tuning stage, considering the influence of re-randomly initializing the parameters of the classifier network and the regressor network module on the features of the object detection deep network model, the data augmentation method is correspondingly selected. If the parameters of the re-randomly initialized classifier network and regressor network module are very different from the parameters of the small-sample object detection network model trained in step one, the CutMix data augmentation method is adopted to optimize the parameters of the classifier and regressor modules in the small-sample object detection network model, making the optimized small-sample object detection network model more stable and having better classification performance.

[0101] It should be understood that although Figure 1 the steps in Figure 1At least some of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily executed and completed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turns with at least some of the steps or stages in other steps or other steps.

[0102] As described above, the above are only embodiments of the present invention and do not impose any limitations on the present invention. Any simple modifications, changes, and equivalent structural changes made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A small-sample object detection method based on the increment of the target proposal box, characterized in that: It includes the following steps: Step 1: Construct a target detection deep network model and train the target detection deep network model using a basic category dataset; Step 2: Use the trained target detection deep network model for feature extraction and clustering: Step 201: Extract feature vectors from basic category data samples and perform clustering, clustering into H1 classes; Step 202: Randomly crop the basic category data samples to obtain base class data blocks, extract features from the obtained base class data blocks and perform clustering, clustering into H2 classes; Step 3: Optimize the parameters of the classifier and regressor modules in the target detection deep network model using a new category dataset to obtain an optimized target detection deep network model; Step 4: Use the optimized target detection deep network model to screen target proposal boxes: Obtain a sampling dataset, input the sampling dataset into the optimized target detection deep network model to obtain k positive samples, use the non-maximum suppression method to screen the target proposal boxes of the k positive samples, select k1 positive sample target proposal boxes with high confidence as target proposal boxes, and there are k2 remaining positive samples that are not screened, where k1 + k2 = k; Step 5: Augment and screen the positive samples that are not screened to obtain supplementary proposal boxes; Step 501: Calculate the similarity P between the j-th unfiltered positive sample and H1 classes respectively j1 , P j1 represents the first foreground score of the j-th unfiltered positive sample, where 1 ≤ j ≤ k2; Step 502: Calculate the similarity P between the j-th unfiltered positive sample and H2 classes respectively j2 , P j2 represents the second foreground score of the j-th unfiltered positive sample; Step 503: The computer calculates according to the formula P j = ω1·P j1 + ω2·P j2 to perform a weighted sum of the first foreground score and the second foreground score of the j-th unfiltered positive sample. P j represents the comprehensive foreground score, ω1 represents the first weight, ω2 represents the second weight, and ω1 + ω2 = 1; Step 504. Screen the comprehensive foreground score P according to the comprehensive foreground score threshold j and take the comprehensive foreground score P that passes the screening j The corresponding j-th unscreened positive sample is used as a supplementary suggestion box, and the supplementary suggestion box and the target suggestion box form a predicted suggestion box; Step 6: Use the obtained prediction proposal boxes to update the parameters of the classifier and regressor modules in the optimized target detection deep network model in Step 3 to obtain a final small sample target detection network model; Step 7: Obtain a small sample image to be detected; Step 8: Input the small sample image to be detected into the final small sample target detection network model to obtain a target detection result.

2. The small-sample object detection method based on the increment of the target proposal box according to claim 1, characterized in that: The specific method of Step 3 includes: Step 301: Delete the parameters of the classifier network and regressor network modules of the trained target detection deep network model in Step 1 and re-initialize the parameters of the classifier network and regressor network modules randomly; Step 302: Obtain new category data samples, and the new category data samples and basic category data samples form a balanced dataset; Step 303: Augment each sample in the balanced dataset respectively to obtain an augmented dataset; Step 304: Use the augmented dataset to optimize the parameters of the classifier and regressor modules in the target detection deep network model to obtain an optimized target detection deep network model.

3. The small-sample object detection method based on the increment of the target proposal box according to claim 2, wherein: The specific method of augmenting each sample in the balanced dataset in Step 303 includes: Randomly linearly combine the q-th sample and the p-th sample according to randomly generated linear combination coefficients to obtain a first input image, and perform the random linear combination multiple times in a loop to obtain multiple first input images, where 1 ≤ q ≤ K, 1 ≤ p ≤ K, p ≠ q; The multiple first input images form the augmented dataset.

4. The small-sample object detection method based on the increment of the target proposal box according to claim 2, characterized in that: The specific method of augmenting each sample in the balanced dataset in Step 303 includes: Randomly crop a data block from the q-th sample and randomly overlay the data block on the p-th sample to obtain a second input image, and perform the random overlay multiple times in a loop to obtain multiple second input images, where 1 ≤ q ≤ K, 1 ≤ p ≤ K, p ≠ q; The multiple second input images form the augmented dataset.

5. The small-sample object detection method based on the increment of the target proposal box according to claim 2, wherein: The specific methods for data augmentation of each sample in the balanced dataset in step 303 include: Step 3031: Randomly linearly combine the q-th sample and the p-th sample according to the randomly generated linear combination coefficients to obtain the first input image. Repeat the random linear combination multiple times to obtain multiple first input images, where 1 ≤ q ≤ K, 1 ≤ p ≤ K, and p ≠ q; Step 3032: Randomly crop a data block from the q-th sample and randomly overlay the data block on the p-th sample to obtain the second input image. Repeat the random overlay multiple times to obtain multiple second input images; Step 3033: Multiple first input images and multiple second input images form the augmented dataset.

6. The small-sample object detection method based on the increment of the target proposal box according to claim 1, wherein: The specific method of step 201 includes: Step 2011: Input n basic category data samples into the trained object detection deep network model for feature vector extraction to obtain the feature vectors of each basic category data sample; Step 2012: Use the K-means clustering algorithm to perform clustering operations on the feature vectors of the basic category data samples and cluster them into H1 classes.

7. The small-sample object detection method based on the increment of the target proposal box according to claim 1, wherein: The specific method of step 202 includes: Step 2021: Randomly crop each of the n basic category data samples L times to obtain n × L basic class data blocks. Input the n × L basic class data blocks into the trained object detection deep network model for feature extraction to obtain n × L feature vectors; Step 2022: Use the K-means clustering algorithm to perform clustering operations on the n × L feature vectors and cluster them into H2 classes.

Citation Information

Patent Citations

  • Deep learning-based few-sample target detection method

    CN113705570A

  • Few-sample target detection method based on singular value decomposition feature enhancement

    CN113971815A