A small sample target detection method based on adaptive feature generation

By adopting an adaptive feature generation method in small sample object detection, and using an adaptive variational autoencoder and RoI feature generator to generate aligned synthetic features, the problems of insufficient diversity of new features and distribution offset are solved, and high-accuracy object detection is achieved.

CN119445089BActive Publication Date: 2025-05-13PLA DALIAN NAVAL ACADEMY

Patent Information

Application Number
CN202510013992.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-13
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The existing small sample object detection method is more prone to errors when dealing with new classes due to the lack of feature diversity, and fails to effectively overcome the distribution offset between synthetic features and real features.

Method used

Using a small sample object detection method based on adaptive feature generation, the adaptive variational autoencoder and RoI feature generator are trained to generate aligned synthetic features and fine-tune the Faster-RCNN detector to improve detection accuracy.

Benefits of technology

It effectively improves the detection accuracy of small sample object detectors, especially when the diversity of new types of features is insufficient, target objects can be accurately positioned and classified, and overcomes the distribution offset between synthetic features and real features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119445089B_ABST
    Figure CN119445089B_ABST
Patent Text Reader

Abstract

The present invention discloses a small sample target detection method based on adaptive feature generation. According to the number of samples in the data category, the training samples are divided into base class training samples and new class training samples. Based on this, the Faster-RCNN detector is trained and fine-tuned to detect small sample targets. The present invention divides the training set into base class training samples and new class training samples according to the number of data categories in the training samples, effectively improving the detection accuracy of the small sample target detector based on the two-stage strategy, especially after using an adaptive RoI feature generator, the detection accuracy can be further improved. The synthetic features generated by the present invention can align with the real RoI features, effectively overcome the distribution deviation of the synthetic features, achieve accuracy surpassing on all average indicators, and can accurately locate the target object and classify the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision target detection, and in particular to a small sample target detection method based on adaptive feature generation. Background Art

[0002] Object detection is a method to obtain the location and category of target objects in an image, and is one of the important applications in the field of computer vision. Different from traditional object detection methods that require a large number of labels for each class, small-sample object detection methods use very limited labeled data to detect new class samples. Given a set of disjoint base class samples with a large number of labels, the knowledge learned from the base class samples and the limited new class labels are used to accurately detect the location and category of the new class. A simple and effective small-sample object detection method is based on a two-stage strategy. In the first stage, only the base class samples are used to pre-train the Faster-RCNN detector. In the second stage, the Faster-RCNN detector is fine-tuned using class-balanced samples containing new and base classes.

[0003] This classic detection strategy mainly solves the problem of knowledge transfer between classes, but fails to solve the problem of lack of feature diversity of new classes. Due to the lack of feature diversity, the classifier layer in the detector is more prone to errors. To overcome this problem, feature generation strategies are used to enrich the feature diversity of new classes and improve the performance of the classifier layer. The performance gains obtained by these methods show that synthesizing new features is beneficial for fine-tuning the detector.

[0004] Inspired by the large model CLIP for feature generation tasks in low-shot learning, we propose to capture powerful semantic embeddings from pre-trained CLIP text encoders. These semantic embeddings can guide the generator to generate reliable region of interest (RoI) features for new classes. However, this approach ignores the distribution shift between synthetic features and real features. Overcoming distribution shift becomes a core challenge of the feature generation paradigm. Summary of the invention

[0005] The present invention discloses a small sample target detection method based on adaptive feature generation to overcome the above technical problems.

[0006] In order to achieve the above object, the technical solution of the present invention is:

[0007] A small sample target detection method based on adaptive feature generation comprises the following steps:

[0008] Step 100: dividing an annotated original data set including multiple data categories into a training set and a test set; and dividing the training set into a base class training sample and a plurality of training samples according to the number of original data in the data category in the training set; randomly obtaining base class training subsamples with the same number as the new class training samples from the base class training samples to obtain few-shot training samples including the new class training samples and the base class training subsamples;

[0009] Step 200: According to the base class training samples, based on the loss function of the Faster-RCNN detector, the Faster-RCNN detector is trained to obtain a trained Faster-RCNN detector;

[0010] Step 300: performing instance-level pruning on the base class training samples and the few-shot training samples to obtain base class instance training samples and few-shot instance training samples;

[0011] Step 400: According to the base class instance training samples, based on the loss function of the adaptive variational autoencoder, the adaptive variational autoencoder is trained to obtain the reconstruction features of the base class instance training samples according to the trained adaptive variational autoencoder;

[0012] Step 500: obtaining a first real RoI feature corresponding to the base class instance training sample according to the trained Faster-RCNN detector and the base class instance training sample, obtaining a synthetic feature of the aligned base class instance training sample according to the reconstructed feature of the base class instance training sample and the RoI feature generator, and training the RoI feature generator based on the loss function of the RoI feature generator;

[0013] Step 600: According to the trained Faster-RCNN detector and the few-shot instance training samples, a second real RoI feature corresponding to the few-shot instance training samples is obtained; and the trained RoI feature generator is fine-tuned according to the synthetic features of the aligned base class instance training samples to obtain a fine-tuned RoI feature generator;

[0014] Step 700: Obtain new synthetic features according to the base class training samples, the new class training samples, the trained adaptive variational autoencoder, and the fine-tuned RoI feature generator; fine-tune the trained Faster-RCNN detector according to the new synthetic features and the few-shot instance training samples, and evaluate the fine-tuned Faster-RCNN detector according to the test set; and then detect small sample targets according to the evaluated fine-tuned Faster-RCNN detector.

[0015] Furthermore, the loss function based on the adaptive variational autoencoder in step 400 is:

[0016] ,

[0017] Where: Represents the loss function value of the adaptive VAE; represents the reconstruction loss; The loss function represents the Kullback-Leibler divergence; Representation decoder The expected log-likelihood of ; represents the decoder of VAE; KL represents the Kullback-Leibler divergence; represents the VAE encoder; It follows the standard normal distribution The prior distribution of Represents a standard normal distribution with mean 0 and variance 1.

[0018] Furthermore, the relationship between the base class training samples and the new class training samples is: n 基类 » n 新类,

[0019] Where: n 基类 Represents the number of original data in each data category in the base class training sample; n 新类 Represents the number of original data in each data category in the new class training samples; » indicates much larger than.

[0020] Furthermore, in step 200, the composition of the Faster-RCNN detector is expressed by the following formula:

[0021] ,

[0022] Where: Represents Faster-RCNN detector; x Represents the sample image to be detected; Represents classification and regression layers; Represents function composition operation; represents the region of interest pooling network; represents the region proposal network; Represents the backbone network.

[0023] Furthermore, in step 200, the loss function of the Faster-RCNN detector is as follows: L base = L rpn + L cls + L loc

[0024] Where: L base Represents the total loss function of the Faster-RCNN detector; L rpn Representation Region Proposal Network loss; L cls represents the classification loss; L loc represents the positioning loss.

[0025] Furthermore, in step 400, the reconstruction features of the base class instance training samples are obtained. The method is as follows:

[0026] Step 401: Obtain visual features of the base class instance training sample based on the base class instance training sample and the pre-trained CLIP image encoder with fixed parameters v ;

[0027] Step 402: According to the encoder of the variational autoencoder VAE and the visual features of the base class instance training samples v , get the spatial encoding of the base class instance training sample z ; To encode the space of base class instance training samples z After three fully connected layers, local deviation is obtained r , and then obtain the prompt set after adding local deviation; the formula used is as follows:

[0028] ,

[0029] in, represents randomly initialized learnable cues; i Indicates the index of the prompt; I Indicates the total number of prompts; represents the set of prompts after adding local bias;

[0030] Step 403: Obtain the final prompt description according to the prompt set after adding the local deviation, and obtain the reconstructed features of the base class instance training sample by using the CLIP text encoder based on fixed parameters and the decoder of the variational autoencoder VAE;

[0031] The formula used to obtain the final prompt description is as follows:

[0032] ,

[0033] The formula used to obtain the reconstruction features of the base class instance training sample is as follows:

[0034] ,

[0035] in, Feature representation that represents the category; t Indicates the final prompt description; Represents the reconstructed features of the base class instance training samples; Represents a fixed-parameter CLIP text encoder.

[0036] Furthermore, the loss function of the RoI feature generator is as follows:

[0037] ,

[0038] Where: represents the loss function of the RoI feature generator; MSE represents mean square error; Represents a synthetic feature; Represents the first true RoI feature corresponding to the base class instance training sample.

[0039] Furthermore, the formula used to fine-tune the trained RoI feature generator is as follows:

[0040] ,

[0041] Where: represents the RoI feature generator; Indicates when When the minimum value is reached The value of

[0042] Furthermore, in step 700, the formula used to obtain the new synthetic feature is as follows:

[0043] .

[0044] Furthermore, in step 700, the formula used to fine-tune the trained Faster-RCNN detector is as follows:

[0045] ,

[0046] Where: represents the classification loss of fine-tuning the trained Faster-RCNN detector; One-hot encoding representing the category label; Represents the classification layer network in Faster-RCNN; Represents the softmax activation function.

[0047] Beneficial effects: The present invention discloses a small sample target detection method based on adaptive feature generation, wherein the first real RoI feature corresponding to the base class instance training sample is obtained through the trained Faster-RCNN detector through the base class training sample with a large number of samples, and the adaptive variational autoencoder is trained through the base class instance training sample to obtain the reconstruction feature of the base class instance training sample, and then the synthetic feature is obtained according to the RoI feature generator to train the RoI feature generator; according to the second real RoI feature corresponding to the few-shot instance training sample, the fine-tuned RoI feature generator is obtained to obtain the new synthetic feature to fine-tune the trained Faster-RCNN detector, and the fine-tuned Faster-RCNN detector is evaluated according to the test set; and then the small sample target is detected according to the fine-tuned Faster-RCNN detector after evaluation. The present invention divides the training set into base class training samples and new class training samples according to the number of data categories in the training sample, effectively improving the detection accuracy of the small sample target detector based on the two-stage strategy, especially after using the adaptive RoI feature generator, the detection accuracy can be further improved. The synthetic features generated by the present invention can align with the real RoI features, take into account the distribution offset between the synthetic features and the real features, effectively overcome the distribution deviation of the synthetic features, achieve accuracy surpassing in all average indicators, and can accurately locate and classify the target objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0049] Figure 1 It is a workflow diagram of the small sample target detection method based on adaptive feature generation of the present invention;

[0050] Figure 2 is an operation diagram of a small sample target detection method based on adaptive feature generation in an embodiment of the present invention;

[0051] Figure 3is a schematic diagram of a decoder of an adaptive variational autoencoder VAE in an embodiment of the present invention;

[0052] Figure 4 is a distribution diagram of new class real RoI features and synthetic features in an embodiment of the present invention;

[0053] Figure 5 is a visualization comparison diagram of the results of using and not using the adaptive RoI feature generator in an embodiment of the present invention;

[0054] Figure 6 This is another visual comparison diagram of the results of using or not using the adaptive RoI feature generator in the embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0056] This embodiment introduces a small sample target detection method based on adaptive feature generation, such as Figure 1 As shown, the following steps are included:

[0057] Step 100: Divide the labeled original data set including multiple data categories into a training set and a test set;

[0058] And according to the amount of original data in the data category in the training set, the training set is divided into a base class training sample and a plurality of training samples;

[0059] Randomly obtaining base class training sub-samples having the same number as the new class training samples from the base class training samples to obtain few-shot training samples including the new class training samples and the base class training sub-samples;

[0060] Preferably, the relationship between the base class training samples and the new class training samples is:

[0061] n 基类 » n 新类

[0062] Where: n 基类 Represents the number of original data in each data category in the base class training sample; n 新类Represents the number of original data in each data category in the new class training sample; » represents much larger than;

[0063] Specifically, in this embodiment, the original data in each data category of the base class training samples includes at least 438 target objects (such as dining table) and at most 9385 target objects (person), with an average of 1417 target objects per category; the element data in each data category of the new class training samples includes 1-10 target objects.

[0064] Specifically, the original data in this embodiment includes multiple data categories, and the original data comes from the public data set PASCAL VOC data set, which has a total of twenty categories. According to specific needs, this embodiment uses PASCAL VOC 2012 and PASCAL VOC 2007trainset as training sets, and PASCAL VOC 2007testset as test sets. According to the number of samples in each data category in the training set, the data in the training set is divided into base class training samples and labeled training samples. Among them, the number of base class training samples is much larger than the number of new class training samples.

[0065] Specifically, there are 20 data categories in this embodiment: aeroplane, bicycle, bird, boat, bottle, bus, car, cat, chair, cow, dining table, dog, horse, motorcycle, person, potted plant, sheep, sofa, train and TV / monitor. This embodiment divides the training set into base class training samples and new class training samples, with three division methods, which are divided into 15 base classes and 5 new classes respectively. In the first division method, the new classes are potted plants, sheep, sofas, trains and TV / monitors, and the remaining 15 classes are base classes. In the second division method, the new classes are airplanes, bicycles, sofas, trains and TV / monitors, and the remaining 15 classes are base classes. In the third division method, the new classes are bicycles, birds, boats, cats and TV / monitors, and the remaining 15 classes are base classes.

[0066] Specifically, the data volume and usage phase of the base class training samples and the new class training samples in this embodiment are different. The base class training samples are categories with sufficient annotations selected from the original data set, which are mainly used for the initial training of Faster-RCNN to allow Faster-RCNN to learn general feature extraction and classification capabilities; while the new class training samples are categories with sparse annotations, usually only a few annotated samples per category, which are used in the fine-tuning phase after the initial training of Faster-RCNN, with the focus on allowing Faster-RCNN to quickly learn the features of new categories under the condition of a small amount of data.

[0067] Specifically, the annotated raw data in this embodiment refers to the raw data that annotates the category and location information of the target objects in the training samples in the training set. Category annotation refers to specifying the category name (such as "ship" or "person") for the target object; location information annotation refers to using a rectangular box to indicate the location of the target in the image, usually including the coordinates of the upper left corner and the lower right corner or the center point and the width and height. These annotations provide classification and positioning information for Faster-RCNN.

[0068] In addition, the base class sub-training samples randomly sampled from the base class training samples constitute a few-shot training sample with a balanced number with the new class training samples, which is used in the fine-tuning stage of Faster-RCNN, that is, the number of base class sub-training samples is equal to that of new class training samples;

[0069] Specifically, in this embodiment, the base class training samples do not overlap with the new class training samples, the test set does not overlap with the training set samples, and the test set contains the base class training samples and the new class training samples. The number of labeled samples in the base class training samples is huge. For each new class training sample, k labeled samples are randomly selected, and k can take the values ​​of 1, 2, 3, 5, 10 to simulate a small sample learning scenario, that is, only k training samples are available for the new class. In addition, a set of subsets is randomly sampled from the base class training samples, namely the base class sub-training samples, which is equal to the number of new class training samples, and the base class training sample subset is merged with the new class training samples into a few-shot training sample.

[0070] Step 200: According to the base class training samples, based on the loss function of the Faster-RCNN detector, the Faster-RCNN detector is trained to obtain a trained Faster-RCNN detector;

[0071] Specifically, this embodiment uses the base class training samples to train the Faster-RCNN detector .

[0072] Preferably, the composition of the Faster-RCNN detector is expressed by the following formula:

[0073] (1)

[0074] Where: Represents Faster-RCNN detector; x Represents the sample image to be detected; Represents classification and regression layers; Represents function composition operation; represents the region of interest pooling network; represents the region proposal network; Represents the backbone network.

[0075] Specifically, each picture x It will be processed by the following network modules in sequence. 1. Backbone network : Input image x Through the backbone network, feature extraction is performed to generate high-level feature maps of the image. The backbone network is ResNet101; 2. Region Proposal Network : The feature map extracted by the backbone network will be input into the region proposal network , the network generates candidate regions through sliding windows and convolution operations, and assigns a confidence score to each region; 3. Region of interest pooling network : The generated candidate regions will be subjected to RoI pooling operation through the region of interest head (RoI Head) to convert candidate regions of different sizes into feature maps of uniform size; 4. Classification and regression layer : The feature map of uniform size will be input into the classification and regression layer for target category prediction and position regression to obtain the final detection result.

[0076] Preferably, during the training process of the Faster-RCNN detector, the optimization of the Faster-RCNN detector is calculated according to the following loss function:

[0077] L base = L rpn + L cls + L loc (2)

[0078] Where: L base Represents the total loss function of the Faster-RCNN detector, which combines all optimization objectives; L rpn Representation Region Proposal Network The loss is responsible for the accuracy of the generated candidate regions; Lcls Represents the classification loss, which is responsible for classifying each candidate region into categories, where the candidate region refers to a set of rectangular boxes generated by the region proposal network based on the input image, which cover the areas in the image that may contain targets; L loc Represents the positioning loss, which is responsible for the position regression of each candidate region, that is, the precise positioning of the bounding box.

[0079] Specifically, this embodiment uses the SGD optimizer to train the Faster-RCNN detector. The batch size of the SGD optimizer is 16, the initial learning rate is 0.02, and the decay rate is 1e-4. The detector trained using the base class training sample is denoted as .

[0080] Step 300: performing instance-level pruning on the base class training samples and the few-shot training samples to obtain base class instance training samples and few-shot instance training samples;

[0081] Specifically, this embodiment performs instance-level cropping on the base class training samples and the few-shot training samples according to the annotations, wherein the instance-level cropping refers to cropping the image according to the coordinates of the position information of the target object annotated in the image to obtain a set of sub-images, each of which contains a target object. Specifically, the "instance level" is different from the "image level". An image may contain many target objects, which are also called instances. After cropping the image, each sub-image obtained contains only one target object (instance), so it is called "instance-level" cropping. Instance-level cropping is a conventional operation in the field, so it will not be described in detail here. After the base class training samples are cropped at the above instance level, the base class instance training samples are obtained. Similarly, the few-shot training samples are cropped to obtain the few-shot instance training samples.

[0082] Step 400: Based on the loss function of the adaptive variational autoencoder, the adaptive variational autoencoder is trained to obtain the reconstruction features of the base class instance training samples according to the trained adaptive variational autoencoder and the base class instance training samples. ; The adaptive variational autoencoder includes a variational autoencoder VAE and a pre-trained and fixed-parameter CLIP image encoder;

[0083] Specifically, in this embodiment, the base class instance training sample is transferred to the pre-trained and fixed-parameter CLIP image encoder to obtain the visual features of the base class instance training sample. v, and then the visual features of the base class reconstruction are obtained through the adaptive variational autoencoder VAE, and this step is used to train the adaptive variational autoencoder (VAE, VariationalAutoencoder).

[0084] Preferably, the loss function based on the adaptive variational autoencoder in step 400 is:

[0085] (3)

[0086] Where: Represents the loss function value of the adaptive VAE; Represents the reconstruction loss, which is used to measure the difference between the reconstructed base class instance training sample and the original sample; The loss function represents the Kullback-Leibler divergence, which is used to measure the difference between the variational distribution and the prior distribution; Representation decoder The expected log-likelihood of , used to measure the latent space encoding z and the categories in the base class instance training samples c The reconstruction error is also called reconstruction loss. Represents the decoder of VAE, which is used to encode z and the categories in the base class instance training samples c Generate reconstruction samples; KL Represents the Kullback-Leibler divergence, which is used to measure The difference between the target prior distribution and Represents the VAE encoder, which is used to transform visual features v Encoding into a latent space z superior; It follows the standard normal distribution The prior distribution of Represents a standard normal distribution with mean 0 and variance 1.

[0087] Preferably, in step 400, the reconstructed features of the base class instance training samples are obtained according to the trained adaptive variational autoencoder and the base class instance training samples. The method is as follows:

[0088] Step 401: Obtain visual features of the base class instance training sample based on the base class instance training sample and the pre-trained CLIP image encoder with fixed parameters v ;

[0089] Step 402: According to the encoder of the variational autoencoder VAE and the visual features of the base class instance training samples v, get the spatial encoding of the base class instance training sample z ; To encode the space of base class instance training samples z After three fully connected layers, local deviation is obtained r , and then obtain the prompt set after adding local deviation; the formula used is as follows:

[0090] (4)

[0091] in, represents randomly initialized learnable cues; i Indicates the index of the prompt; I Indicates the total number of prompts; Represents the set of prompts after adding local bias.

[0092] Step 403: Obtain the final prompt description according to the prompt set after adding the local deviation, and obtain the reconstructed features of the base class instance training sample by using the CLIP text encoder based on fixed parameters and the decoder of the variational autoencoder VAE;

[0093] The formula used to obtain the final prompt description is as follows:

[0094] ,

[0095] The formula used to obtain the reconstruction features of the base class instance training sample is as follows:

[0096] ,

[0097] in, Feature representation that represents the category; t Indicates the final prompt description; Represents the reconstructed features of the base class instance training samples; Represents a fixed-parameter CLIP text encoder.

[0098] Specifically, the adaptive variational autoencoder in this embodiment is composed of a variational autoencoder VAE and a pre-trained CLIP image encoder with fixed parameters. The base class instance training sample is input into the pre-trained CLIP image encoder with fixed parameters. , and then output the visual features of the base class instance training samples Among them, the variational autoencoder VAE consists of an encoder and a decoder Composition, encoder The visual features Encoding into a latent space On, decoder Through latent space encoding and visual features Corresponding categories c To reconstruct visual features , and obtain the reconstructed features .like Figure 3 As shown, the decoder A fixed-parameter CLIP text encoder and multiple groups of prompts. Multiple groups of prompts consist of texts of multiple words, which can be learned through network training and used to guide the model to generate text information related to specific data categories. The specific form is as follows:

[0099] Get the latent space encoding After that, a set of local deviations is generated , by calculating Get, where h is a three-layer fully connected layer. r Added to a set of prompts by calculation:

[0100] ,

[0101] in, represents randomly initialized learnable cues; i Indicates the index of the prompt; I Indicates the total number of prompts; represents the set of prompts after adding local bias;

[0102] After that, extract the feature representation of the category corresponding to the category , and concatenate it with this set of prompts. Insert into the front, middle, and back of the prompt to get the final prompt description During the adaptive variational autoencoder training phase, The feature representation of the category is divided into Represents the characteristics of airplanes, bicycles, birds, boats, bottles, buses, cars, cats, chairs, cows, dining tables, dogs, horses, motorcycles, and people, divided into two The characteristics of a bird, a boat, a bottle, a bus, a car, a cat, a chair, a cow, a dining table, a dog, a horse, a motorcycle, a person, a potted plant, and a sheep are divided into three Represents the features of airplanes, bottles, buses, cars, chairs, cows, dining tables, dogs, horses, motorcycles, people, potted plants, sheep, sofas, and trains. Feature representation that represents the category; t represents the final prompt description, which is the concatenation of the prompt and category feature representations;

[0103] Finally, the CLIP text encoder with fixed parameters Get the reconstruction features .

[0104] Specifically, the adaptive variational autoencoder of this embodiment is composed of a variational autoencoder VAE and a pre-trained CLIP image encoder with fixed parameters. The specific network design is as follows: Figure 2 and Figure 3 As shown:

[0105] Step 500: Obtain the first real RoI feature corresponding to the base class instance training sample according to the trained Faster-RCNN detector and the base class instance training sample , to reconstruct features of the training samples based on the base class instances And the RoI feature generator obtains the synthetic features of the aligned base class instance training samples , and based on the loss function of the RoI feature generator, train the RoI feature generator;

[0106] Specifically, the RoI feature generator can align the reconstructed features Compared with the real RoI features, the specific design is as follows Figure 2 As shown in part (a), the RoI feature generator Use three fully connected layers to build and reconstruct the features Input to RoI feature generator In the above example, the synthetic features of the aligned base class instance training samples are obtained. .

[0107] This embodiment uses the trained Faster-RCNN detector and inputs the base class instance training sample into the trained Faster-RCNN detector to obtain the first real RoI feature corresponding to the base class instance training sample. ,use For RoI feature generator To supervise, the loss function of the RoI feature generator is calculated by the following formula:

[0108] (5)

[0109] Where: represents the loss function of the RoI feature generator; MSE represents mean square error; Represents a synthetic feature; Represents the first true RoI feature corresponding to the base class instance training sample.

[0110] In this embodiment, the total optimization function in the process of obtaining the synthetic feature is expressed as follows:

[0111] (6)

[0112] Where: represents the encoder of the variational autoencoder VAE; represents the decoder of the variational autoencoder VAE; h Indicates the local deviation. r The three fully connected layers used when represents the total optimization function; represents the RoI feature generator;

[0113] Specifically, Represents the encoder of VAE , VAE decoder , three fully connected layers h , tips to learn Optimize the parameters of the RoI feature generator to minimize the value of the optimization function. When the optimization function reaches the minimum value, the parameters of these modules are the optimal values, which can better meet the goals of model design;

[0114] Specifically, in this embodiment, the Adam optimizer is used for the overall optimization of the adaptive variational autoencoder and the RoI feature generator, and the learning rate is 1e-3.

[0115] Step 600: Obtain a second true RoI feature corresponding to the few-shot instance training sample based on the trained Faster-RCNN detector and the few-shot instance training sample ; Fine-tune the trained RoI feature generator based on the synthetic features of the aligned base class instance training samples to obtain the fine-tuned RoI feature generator. The formula used is as follows:

[0116] (7)

[0117] Where: represents the RoI feature generator; Indicates when When the minimum value is reached The value of

[0118] Specifically, this embodiment uses a few-shot instance training sample to fine-tune the trained RoI feature generator to obtain a fine-tuned RoI feature generator.

[0119] Preferably, in step 303, the trained RoI feature generator is fine-tuned through a few-shot instance training sample to align the synthetic features of the new class with the real features of the new class. The specific design is as follows: Figure 2As shown in part (b) of the diagram, the design method is as follows:

[0120] Use the trained Faster-RCNN detector to obtain a few-shot instance training sample consisting of an equal number of base class instance training samples and new class instance training samples, and obtain the corresponding second real RoI features ,use After training, the RoI feature generator Make fine adjustments.

[0121] Step 700: Obtain new synthetic features based on the base class training samples, the new class training samples, the trained adaptive variational autoencoder, and the fine-tuned RoI feature generator; The trained Faster-RCNN detector is fine-tuned based on the few-shot instance training samples, and the fine-tuned Faster-RCNN detector is evaluated based on the test set; then, the small sample target can be detected based on the evaluated fine-tuned Faster-RCNN detector.

[0122] Specifically, after fine-tuning the trained RoI feature generator, use the base class training samples and the new class training samples to generate a new set of synthetic features through the trained adaptive variational autoencoder and the fine-tuned RoI feature generator; in the three division methods, the categories composed of the base class and the new class are the following twenty categories: airplane (aeroplane), bicycle (bicycle), bird (bird), boat (boat), bottle (bottle), bus (bus), car (car), cat (cat), chair (chair), cow (cow), dining table (diningtable), dog (dog), horse (horse), motorcycle (motorbike), person (person), potted plant (pottedplant), sheep (sheep), sofa (sofa), train (train) and TV / monitor (tv / monitor). The specific design is as follows Figure 2 As shown in part (c) of the figure, twenty categories, such as aeroplane and bicycle, are used to input the category names into the adaptive variational autoencoder to obtain new synthetic features. The new features are further adjusted through the RoI feature generator to obtain generated features.

[0123] Specifically, assuming that the category of the new class is known, the category names of the base class training samples and the category names of the new class training samples are compared with the Gaussian distribution obtained by sampling from the prior distribution. Latent space encoding , expressed as , input to the decoder of the trained adaptive variational autoencoder and the fine-tuned RoI feature generator , generate new synthetic features , the formula used is as follows:

[0124] (8)

[0125] Among them, from the Gaussian distribution Randomly select multiple latent space encodings from z , so that each category generates multiple synthetic features.

[0126] Specifically, in this embodiment, the Faster-RCNN detector is fine-tuned and evaluated, and the detector is fine-tuned using synthetic features and real features. The specific design is as follows: Figure 2 As shown in part (d) of the figure, the detection accuracy of the Faster-RCNN detector is evaluated using the test set.

[0127] Preferably, in step 800, a new synthetic feature is used The Faster-RCNN detector is fine-tuned with the real RoI features of the few-shot instance training samples. The real RoI features refer to the few-shot instance training samples input into Faster-RCNN, and the region of interest pooling network of Faster-RCNN is obtained. Output RoI features. The design method is as follows:

[0128] The synthetic features of all categories The true RoI features of the few-shot training instance samples Fine-tune the detector's classifier layer together, calculated using the following formula:

[0129] (9)

[0130] Where: represents the classification loss of fine-tuning the trained Faster-RCNN detector; One-hot encoding representing the category label; Represents the classification layer network in Faster-RCNN; Represents the softmax activation function.

[0131] Specifically, in the process of fine-tuning the trained Faster-RCNN detector, this embodiment uses the existing knowledge inheritance strategy to initialize the weights of the new class classification layer.

[0132] The test set is used to evaluate the detection accuracy of the fine-tuned Faster-RCNN detector. The test set contains all twenty categories of base classes and new classes, a total of 4,952 images, and approximately 25,000 instances (i.e., labeled target objects). The image distribution of this test set is similar to that of actual application scenarios, covering complex backgrounds, target occlusions, and multi-scale targets, so as to comprehensively examine the robustness and detection capabilities of the detector in various environments. After inputting the images in the test set into the fine-tuned Faster-RCNN detector, the target category, prediction confidence, and detection box position coordinates of each test image are output to implement a small sample target detection method. Figure 4 As shown, the dots represent the real RoI feature distribution, and the cross represents the synthetic feature distribution. Figure 4 Part (a) shows the distribution offset between synthetic features and real features before using the RoI feature generator. The arrow indicates the correct alignment direction. Figure 4 Part (b) in the figure shows that the synthetic features can fall within the distribution range of the real features using the RoI feature generator and meet the correct alignment direction, which can fully illustrate that the adaptive feature generator of this embodiment can overcome the distribution deviation between the synthetic features and the real features, so that the synthetic features of the base class and the new class can align with the real RoI features of the base class and the new class. The evaluation method is a prior art in the field and will not be described in detail here.

[0133] like Figure 5 and Figure 6 Part (a) shows the detection results without using the adaptive feature generation method, part (b) shows the detection results after using the adaptive feature generation method, and part (c) is the correct annotation. Figure 5 The cow and sofa in the figure belong to the new class training samples, and the sheep belongs to the base class training sample; Figure 6 The sofa, TV / monitor, and bird in the figure belong to the new class training samples, while the bottle and chair belong to the base class training samples. Figure 5 and Figure 6 The detection results in all showed false detection and missed detection of new training samples: for example, Figure 5 The cow in the image is mistakenly detected as the sheep in the base class, and the sofa is missed; Figure 6 The sofa and bird are missed. After using the adaptive feature generation method, the problems of false detection and missed detection are effectively solved. Figure 5 As shown in part (c), the cow is correctly identified and the sofa is recognized; Figure 6As shown in part (c), sofa and bird are also correctly identified, which fully demonstrates that the adaptive feature generation method of this embodiment can effectively improve the accuracy of the detector, and the classification results for a small number of new class samples are more accurate.

[0134] In summary, this embodiment uses the adaptive feature generation method on different benchmark small sample target detection benchmarks. The experimental results show that the small sample target detection method of this embodiment can produce the best detection accuracy compared with the most advanced small sample target detection method.

[0135] Compared with other small sample target detection methods based on feature generation, this embodiment can make CLIP semantic embedding better adapt to the target domain, align synthetic features with corresponding real features, and make synthetic features participate in the second stage of the two-stage training strategy, and further improve the accuracy of small sample target detection by improving the diversity of new class features. It can overcome the distribution deviation between synthetic features and real features, effectively improve the feature diversity of new classes through feature generation, and greatly improve the accuracy of target positioning and classification.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A small sample target detection method based on adaptive feature generation, characterized in that: The steps include: Step 100: dividing an annotated original data set including multiple data categories into a training set and a test set; and dividing the training set into a base class training sample and a plurality of training samples according to the number of original data in the data category in the training set; Randomly obtaining base class training sub-samples having the same number as the new class training samples from the base class training samples to obtain few-shot training samples including the new class training samples and the base class training sub-samples; Step 200: According to the base class training samples, based on the loss function of the Faster-RCNN detector, the Faster-RCNN detector is trained to obtain a trained Faster-RCNN detector; Step 300: performing instance-level pruning on the base class training samples and the few-shot training samples to obtain base class instance training samples and few-shot instance training samples; Step 400: According to the base class instance training samples, based on the loss function of the adaptive variational autoencoder, the adaptive variational autoencoder is trained to obtain the reconstruction features of the base class instance training samples according to the trained adaptive variational autoencoder; The loss function of the adaptive variational autoencoder VAE in step 400 is: Where: L vae Represents the loss function value of adaptive VAE; L recon represents the reconstruction loss; L kld The loss function represents the Kullback-Leibler divergence; represents the expected log-likelihood of the decoder D(z, c); D(z, c) represents the decoder of VAE; KL represents the Kullback-Leibler divergence; E(v) represents the VAE encoder; p(z|c) represents the standard normal distribution The prior distribution of represents a standard normal distribution with a mean of 0 and a variance of 1; z represents spatial encoding; c represents category; v represents visual features; Step 500: obtaining a first real RoI feature corresponding to the base class instance training sample according to the trained Faster-RCNN detector and the base class instance training sample, obtaining a synthetic feature of the aligned base class instance training sample according to the reconstructed feature of the base class instance training sample and the RoI feature generator, and training the RoI feature generator based on the loss function of the RoI feature generator; Step 600: According to the trained Faster-RCNN detector and the few-shot instance training sample, obtain a second real RoI feature corresponding to the few-shot instance training sample; Fine-tune the trained RoI feature generator according to the synthetic features of the aligned base class instance training samples to obtain a fine-tuned RoI feature generator; Step 700: Obtain new synthetic features according to the base class training samples, the new class training samples, the trained adaptive variational autoencoder, and the fine-tuned RoI feature generator; fine-tune the trained Faster-RCNN detector according to the new synthetic features and the few-shot instance training samples, and evaluate the fine-tuned Faster-RCNN detector according to the test set; and then detect small sample targets according to the evaluated fine-tuned Faster-RCNN detector.

2. The small sample target detection method based on adaptive feature generation according to claim 1, characterized in that: In step 200, the composition of the Faster-RCNN detector is expressed by the following formula: Where: Φ Det (·) represents the Faster-RCNN detector; x represents the sample image to be detected; Φ Cls&Reg Represents classification and regression layers; Represents function composition operation; Φ RoI represents the region of interest pooling network; Φ RPN represents the region proposal network; Φ Backbone Represents the backbone network.

3. The small sample target detection method based on adaptive feature generation according to claim 1, characterized in that: In step 200, the loss function of the Faster-RCNN detector is as follows: L base =L rpn +L cls +L loc Where: L base represents the total loss function of the Faster-RCNN detector; L rpn Represents the region proposal network Φ RPN loss; L cls represents the classification loss; L loc represents the positioning loss.

4. The small sample target detection method based on adaptive feature generation according to claim 1, characterized in that: In step 400, the method for obtaining the reconstructed features of the base class instance training sample is as follows: Step 401: Obtain visual features v of the base class instance training samples according to the base class instance training samples and a pre-trained CLIP image encoder with fixed parameters; Step 402: According to the encoder E(v) of the variational autoencoder VAE and the visual features v of the base class instance training samples, the spatial encoding z of the base class instance training samples is obtained; the spatial encoding z of the base class instance training samples is passed through three fully connected layers to obtain the local deviation r, and then the prompt set after adding the local deviation is obtained; the formula used is as follows: p(z)=[p1+r,p2+r,...,p i +r,…,p I +r] Among them, p i represents a randomly initialized learnable prompt; i represents the index of the prompt; I represents the total number of prompts; p(z) represents the prompt set after adding local bias; Step 403: Obtain the final prompt description according to the prompt set after adding the local deviation, and obtain the reconstructed features of the base class instance training sample by using the CLIP text encoder based on fixed parameters and the decoder of the variational autoencoder VAE; The formula used to obtain the final prompt description is as follows: t={p(z),e c } The formula used to obtain the reconstruction features of the base class instance training sample is as follows: Among them, e c represents the feature representation of the category; t represents the final prompt description; Represents the reconstructed features of the base class instance training samples; Represents a fixed-parameter CLIP text encoder.

5. The small sample target detection method based on adaptive feature generation according to claim 4, characterized in that: The loss function of the RoI feature generator is as follows: Where: L cons represents the loss function of the RoI feature generator; MSE represents the mean square error; Represents a synthetic feature; Represents the first true RoI feature corresponding to the base class instance training sample.

6. The small sample target detection method based on adaptive feature generation according to claim 5, characterized in that: The formula used to fine-tune the trained RoI feature generator is as follows: Where: represents the RoI feature generator; Indicates that when L cons When the minimum value is reached The value of Represents the second true RoI feature.

7. The small sample target detection method based on adaptive feature generation according to claim 1, characterized in that: In step 700, the formula used to obtain the new synthetic feature is as follows: Where: represents the new synthetic features; D(z, c) represents the decoder of VAE, which is used to generate reconstructed samples through the latent space encoding z and the category c in the base class instance training samples; G(·) represents the RoI feature generator.

8. The small sample target detection method based on adaptive feature generation according to claim 1, characterized in that: In step 700, the formula used to fine-tune the trained Faster-RCNN detector is as follows: Where: L cls ′ represents the classification loss of fine-tuning the trained Faster-RCNN detector; t′ represents the one-hot encoding of the category annotation; Φ cls represents the classification layer network in Faster-RCNN; σ represents the softmax activation function; Represents the second real RoI feature; Represents a new synthetic feature.

Citation Information

Patent Citations

  • Few-sample target detection method, electronic equipment and computer medium

    CN114926622A

  • Small sample target detection method based on semantic enhancement feature generation and predictive optimization

    CN118135298A

Cited By

  • Small sample target detection method for robot intelligent operation

    CN121053637A