Small-sample target detection method and system based on aggregation variational prototype
By aggregating variational prototypes, using P-VAE to generate class prototypes and combining them with the MFM module, the problem of insufficient detection accuracy caused by the scarcity of new class samples is solved, and efficient target detection is achieved in scenarios with few samples.
Patent Information
- Application Number
- CN202510934296.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-11-07
AI Technical Summary
Existing target detection methods struggle to effectively utilize limited samples for feature representation and generalization when new category samples are scarce, resulting in insufficient detection accuracy and generalization ability, making them particularly difficult to implement in scenarios such as medical image analysis and rare species monitoring.
The method of aggregation variational prototyping is adopted. Class prototypes are generated through the prior variational autoencoder (P-VAE) module, and feature diversity and matching ability are enhanced by the mutual fusion module (MFM). A joint training mechanism of base class and new class is constructed, and the support set and query set are divided by the context meta-learning paradigm to achieve feature fusion and object detection.
It significantly improves the detection accuracy and generalization ability of new target categories under conditions of few samples, solves the model bias problem caused by data imbalance in traditional methods, and achieves more accurate feature extraction and target recognition.
Smart Images

Figure CN120912855A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision, in particular to a few-shot object detection method and system based on aggregated variational prototype. BACKGROUND
[0002] In the field of computer vision, object detection as a basic task has made significant progress, but its highly dependent on large-scale labeled dataset characteristics, faces severe challenges in practical applications. In the scene such as medical image analysis, rare species monitoring, the labeled samples of new class targets are extremely scarce, and the traditional detection method is prone to overfitting due to insufficient samples, or is difficult to land due to high cost of manual labeling. Few-shot object detection (FSOD) aims to detect new targets through very few training samples, but its core difficulty lies in the serious imbalance between base class and new class data, and the insufficient representation of new class features, which leads to insufficient generalization ability of the model to new classes. Therefore, it is urgent to develop a detection scheme that can effectively utilize limited new class samples and integrate prior knowledge to enhance new class features, in order to solve the problem of data scarcity in practical applications.
[0003] Traditional object detection methods are mainly divided into single-stage (such as YOLO, SSD) and two-stage (such as FasterR-CNN) methods, which rely on a large amount of labeled data to train the model. The advantages of this method are: single-stage method is fast, suitable for real-time scenarios; two-stage method has high detection accuracy through region proposal optimization. However, its significant disadvantage is the high requirement for data volume - when there is a lack of labeled data, the model is prone to overfitting, and it cannot quickly adapt to new classes. For example, FasterR-CNN requires thousands of labeled images on the PASCAL VOC dataset to achieve ideal performance, but when faced with new classes with less than 10 samples, the detection accuracy drops by more than 50%. This "sample starvation" characteristic makes it difficult to deal with scenarios with limited samples such as industrial defect detection and new target recognition for autonomous driving.
[0004] Existing FSOD technologies are mainly divided into fine-tuning paradigm and meta-learning paradigm. Fine-tuning methods (such as TFA) classify new samples through cosine similarity and fine-tune the model, but they are sensitive to class differences and prone to negative transfer; meta-learning methods (such as MetaR-CNN) detect through feature matching between support set and query set, but they rely on limited support samples to generate class prototypes, which have the problem of insufficient representation of new class prototypes. For example, the VFA method introduces a variational autoencoder (VAE) to generate class prototypes, but it only relies on the transfer of base class pre-trained models, and does not fully utilize the semantic prior of new classes, resulting in a large gap between new class prototypes and real features. In addition, existing methods generally ignore the feature interaction between support set and query set, and the diversity of support features is insufficient, which limits the matching accuracy of the query, especially in 1-shot, 2-shot and other extreme few-shot scenarios, the detection performance cannot meet the actual demand. SUMMARY
[0005] To address the aforementioned technical problems, this application discloses a few-shot target detection method and system based on a convergent variational prototype. The few-shot target detection method based on the convergent variational prototype specifically includes:
[0006] S1. Construct a dataset containing a base class and a new class, and divide the dataset into a support set and a query set. The base class contains a large number of samples, and the new class contains a small number of samples.
[0007] S2. The support set samples are processed using the prior variational autoencoder (P-VAE) module. The P-VAE module enhances the training of new classes through a feature discriminator and enriches semantic information using class priors to generate class prototypes of base classes and new classes.
[0008] S3. The features of the support set and query set are processed through the mutual fusion module MFM. The MFM module enhances the diversity of support features and promotes feature matching of query samples.
[0009] S4. Merge the query set features with the generated class prototypes to obtain the merged features;
[0010] S5. Perform target detection based on the fused features and output the detection results.
[0011] Preferably, the specific process of constructing a dataset containing base classes and new classes and dividing it into a support set and a query set in S1 is as follows: obtaining a base class dataset D containing rich data. b and its base class set C b Obtain a new class dataset D containing k instances for each category. n and its new class set C n C n With C b Disjoint, following the contextual meta-learning paradigm, in D b and D n A large number of small-sample tasks, each containing a support set and a query set, are constructed to form scenarios; the base class dataset D b For meta-training, the new class dataset D n and D b A subset of is used for meta-fine-tuning.
[0012] Preferably, the specific process of processing the support set samples in S2 by the prior variational autoencoder P-VAE module is as follows: generating semantic features of the base class and the new class by the CLIP pre-training model, fusing the semantic features of the base class with its image features to form a multi-modal feature as the base class feature, taking the semantic features of the new class as the image features of the new class in the initial training stage, generating the class prototype of the new class by the P-VAE module and reconstructing by the decoder, training the image features and the semantic features of the new class by the feature discriminator to minimize the difference between the two, and enhancing the class prototype of the new class by back propagation.
[0013] Preferably, in the prior variational autoencoder P-VAE module, the support feature S is converted into the distribution N of the base class and the new class, and the class prototype z is sampled from the distribution N. b and the new class prototype z n By using the VAE optimization model, it is assumed that the class prototype z is subject to the prior distribution p(z), the support feature S is subject to the conditional distribution p(S|z), the real posterior distribution p(z|S) is approximated by the variational distribution q(z|S) through variational inference, the Kullback-Leibler divergence is minimized, and the formula is: which is equivalent to maximizing the evidence lower bound, and the formula is: ELBO=E q(z|S) [logp(S|z)]-D KL (q(z|S)||p(z)), wherein the prior distribution p(z)=N(0,I), the posterior distribution q(z|S)=N(μ,σ), and the parameters μ and σ are determined by the feature encoder F enc , the prototype z is obtained by using the reparameterization trick, that is, z=μ+σ⊙ε, wherein ε~N(μ,σ), and the reconstruction loss is obtained by optimizing the reconstruction loss L dec and the KL divergence loss L KL The support feature S is converted into the distribution N and the class prototype is sampled.
[0014] Preferably, the specific process of enhancing the new class training by the P-VAE module in S2 and enriching the semantic information of the base class and the new class prototype by using the class prior is as follows: generating the text feature and the image feature by using the CLIP model, calculating the cosine similarity of the text embedding vector T e and the image embedding vector I e , and the formula is: wherein t is a temperature parameter, the image loss L img and the text loss L text are combined to obtain the symmetric loss L sem =L img +L textThe training process obtains semantic information about the categories. This semantic information is then input into the P-VAE module to obtain the feature distribution of the new classes. Class prototypes are sampled from the latent space and fed into the decoder to obtain reconstructed features. The semantic features of the new classes are then combined with the real image features x. n Input discriminator, loss L through discriminator dis =L real +L fake Minimize the gap between the reconstructed semantic features of the new class and the real image features, enhance the similarity between the new class prototype and the real features, and generate class prototypes for the base class and the new class, where L real For the loss function of the real samples, L fake The loss function is for fake samples.
[0015] Preferably, the specific process of processing the support set and query set features through the inter-fusion module MFM in S3 is as follows: receiving features X extracted from the query set and support set branches. q and X s The similarity between the query branch and the support branch is calculated. Based on the similarity, the support features and query features are recombined. Customized support features and query features are obtained through convolutional layers. The fusion conditions are calculated using global average pooling and cosine similarity. The reconstructed query features are integrated into the support features to enhance the feature diversity of the support branch. The reconstructed support features are integrated into the query features to promote feature matching of the query features.
[0016] Preferably, the specific process of processing the support set and query set features through the inter-fusion module MFM in S3 is as follows: representing the query set features and support set features as key-value pairs X respectively. q and X s Calculate the branch key K s And query branch value V q The compatibility function f(K) s V q ) = Softmax(k s ·V q Based on weight f(K) s V q Recombination support feature V s and query feature K q Customized support features x are obtained through convolutional layers. s =Conv(f(K) s V q )·K q ) and query feature x q =Conv(f(K) s V q )·V s The fusion condition F is calculated using global average pooling (GAP) and cosine similarity. sD(GAP(X q ), X s ) and F q =D(GAP(X s ), X q ), where D is a cosine similarity distance measure, the reconstructed query feature x q is integrated into the support feature X s to obtain X' s =F s ⊙x q +X s , the reconstructed support feature x s is integrated into the query feature X q to obtain X' q =F q ⊙x s +X q , which enhances the diversity of support features and promotes the matching of query features.
[0017] Preferably, the specific process of fusing the query set features with the generated class prototypes in S4 to obtain fused features is: the processed query set features are aggregated with the base class prototypes and new class prototypes generated by the P-VAE module, the channel attention mechanism is used to enhance the feature dimensions related to the categories, the semantic information of the class prototypes is used to guide the target region identification in the query features, the complementary fusion of the query features and the class prototypes in the semantic and spatial dimensions is realized, and the fused features containing the target category prior and the query image details are formed.
[0018] Preferably, the specific process of target detection based on the fused features and outputting the detection results in S5 is: the fused features are input into the detection head, the region proposal network RPN is used to generate foreground candidate regions, RoIAlign is used to extract and align the candidate region features, the classification head and the regression head are combined to respectively perform target category prediction and bounding box regression, the classification head uses a cross-entropy loss function to calculate the category probability, the regression head uses a smooth L1 loss function to optimize the bounding box coordinates, the target detection results are screened according to the pre-set confidence threshold, and the detection results containing the target category, the bounding box coordinates and the confidence are output.
[0019] The few-shot target detection system based on aggregated variational prototypes comprises a data set construction module, a prior variational autoencoder P-VAE module, a mutual fusion module MFM, a feature fusion module and a target detection module.
[0020] The data set construction module is used to construct a data set containing base classes and new classes, and divide the data set into a support set and a query set, wherein the base classes contain a large number of samples, and the new classes contain a small number of samples.
[0021] The prior variational autoencoder P-VAE module is connected with the data set construction module, and is used for processing support set samples, enhancing new class training through a feature discriminator, and generating class prototypes of base classes and new classes by utilizing semantic information of class priors;
[0022] The mutual fusion module MFM is connected with the data set construction module and the P-VAE module, and is used for processing features of the support set and the query set, enhancing the diversity of the support features, and promoting feature matching of the query samples;
[0023] The feature fusion module is connected with the P-VAE module and the MFM, and is used for fusing the query set features and the generated class prototypes to obtain fused features;
[0024] The target detection module is connected with the feature fusion module, and is used for target detection based on the fused features, and outputs a detection result.
[0025] Compared with the prior art, the technical scheme of the present application has the following technical effects:
[0026] By constructing a joint training mechanism of base classes and new classes, the present scheme breaks the limitation of data division between base classes and new classes in traditional few-shot detection, divides the support set and the query set by using a scenario meta-learning paradigm, and enables the model to learn rich features of base classes and scarce information of new classes in the meta-training and meta-fine-tuning stages, thereby avoiding model bias caused by a large difference in data volume. This mechanism effectively balances the training weights of different classes of data, so that the model can still capture key features in the case of extremely few new class samples, and fundamentally alleviates the inhibition of data imbalance on detection performance.
[0027] The prior variational autoencoder module of the present application introduces semantic prior knowledge of a pre-trained model, and combines a feature discriminator to strengthen the authenticity of new class features. By fusing semantic information and image features, the generated new class prototype is closer to the essential attributes of the real target, thereby solving the problem of fuzzy prototype representation caused by insufficient support samples in traditional methods. This module enables the model to extract more representative feature distribution from limited new class samples, provides a more accurate class reference for subsequent query sample matching, and significantly improves the representation ability of new class features.
[0028] The mutual fusion module of the present application realizes the complementary advantages of the support set and the query set features through a bidirectional feature interaction mechanism. The injection of support features into the query branch enhances the discrimination ability of the query features for new class targets, and the feedback of the query features to the support branch enriches the diversity of the support features. This interaction is realized through feature similarity calculation and customized recombination, so that the two types of features complement each other in semantic and spatial dimensions, effectively solving the problem of low matching accuracy caused by insufficient feature interaction in traditional methods, and improving the feature utilization efficiency in the few-shot scenario.
[0029] The present application is based on a multi-module collaborative optimization fusion feature detection process, combined with the classification and regression optimization design of the detection head. The present application realizes the comprehensive improvement of detection performance in the few-sample scene, from feature extraction to prototype generation, to cross-branch feature fusion. The synergistic effect of each link enables the model to more accurately locate and identify new class targets. Compared with traditional methods, the present application has significant progress in detection accuracy, positioning accuracy and generalization ability of new class targets under few-sample conditions, providing a more reliable technical solution for target detection of scarce samples in practical applications.
[0030] The above description is only a summary of the technical solutions of the present application. In order to more clearly understand the technical means of the present application, the contents of the specification can be implemented, and in order to make the above and other purposes, characteristics and advantages of the present application more obvious and easy to understand, the following will be described in detail with the preferred embodiments of the present application and with the help of the accompanying drawings.
[0031] According to the detailed description of the specific embodiments of the present application below in combination with the accompanying drawings, those skilled in the art will more clearly understand the above and other purposes, advantages and characteristics of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings described below are some embodiments of the present application. Those skilled in the art can obtain other drawings according to these drawings without creating any creative labor. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, each element or part is not necessarily drawn according to the actual proportion.
[0033] Figure 1 Flow chart of few-shot target detection method based on aggregated variational prototype;
[0034] Figure 2 Schematic diagram of generating semantic features through CLIP model and P-VAE module processing support set samples;
[0035] Figure 3 Flow chart of mutual fusion module MFM processing support set and query set features;
[0036] Figure 4 Overall architecture diagram of few-shot target detection model based on aggregated variational prototype;
[0037] Figure 5 Module connection diagram of few-shot target detection system based on aggregated variational prototype;
[0038] Figure 6Figure 4. Comparison of TSNE dimensionality reduction results for traditional methods and the new class prototype generated by AVP.
[0039] Figure 7 Figure 5. Illustration of the impact of hyperparameter λ on the performance of AVP in 1-shot scenario.
[0040] Figure 8 Figure 6. Visualization of detection results of AVP in complex scenarios.
[0041] Figure 9 Figure 7. Comparison of results of different methods in target detection scenarios. DETAILED DESCRIPTION
[0042] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. In the following description, specific details such as specific configurations and components are provided only to help a comprehensive understanding of the embodiments of the present application. Therefore, those skilled in the art should understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. In addition, in order to be clear and concise, the description of known functions and structures is omitted in the embodiments.
[0043] It should be understood that the "one embodiment" or "the embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "one embodiment" or "the embodiment" appearing throughout the specification does not necessarily mean the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner.
[0044] In addition, reference numerals and / or letters can be repeated in different examples in the present application. Such repetition is for the purpose of simplification and clarity, and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0045] The term "and / or" herein is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, B exists alone, and A and B exist together. The term "and" herein is a description of another association relationship of the associated objects, which means that there can be two relationships, for example, A and B can mean that A exists alone, and A and B exist together. In addition, the character " / " herein generally means that the associated objects before and after are in an "or" relationship.
[0046] The term "at least one of' as used herein is merely a descriptive association relationship of associated objects, which means that three relationships can exist, for example, at least one of A and B can mean that A exists alone, A and B exist together, and B exists alone.
[0047] It should also be noted that in this paper, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion.
[0048] Embodiment 1
[0049] This embodiment mainly describes a few-shot target detection method based on aggregated variational prototype, as shown in Figure 1 , specifically:
[0050] S1, constructing a data set containing base classes and new classes, dividing the data set into support set and query set, the base class contains a large number of samples, and the new class contains a small number of samples;
[0051] S2, processing the support set samples by using the prior variational autoencoder P-VAE module, the P-VAE module enhances the new class training through the feature discriminator, and generates the class prototype of the base class and the new class by using the category prior to enrich the semantic information;
[0052] S3, processing the features of the support set and the query set by using the mutual fusion module MFM, the MFM module enhances the diversity of the support features, and promotes the feature matching of the query samples;
[0053] S4, fusing the query set features with the generated class prototype to obtain the fusion features;
[0054] S5, performing target detection based on the fusion features, and outputting the detection result.
[0055] Further, the specific process of constructing a data set containing base classes and new classes and dividing it into support set and query set in S1 is: obtaining a base class data set D b and its base class set C b , obtaining a new class data set D n containing k instances of each category and its new class set C n , wherein C n and C b are disjoint, follow the scenario meta-learning paradigm, and construct a large number of small sample tasks containing support set and query set in D b and D n respectively to form a scenario; the base class data set D bFor meta-training, the new class dataset D n and D b A subset of is used for meta-fine-tuning.
[0056] Furthermore, such as Figure 2 As shown, the specific process of processing support set samples using the prior variational autoencoder P-VAE module in S2 is as follows: semantic features of the base class and the new class are generated through the CLIP pre-trained model. The semantic features of the base class are fused with its image features to form multimodal features as base class features. In the initial training stage, the semantic features of the new class are used as the image features of the new class. The class prototype of the new class is generated using the P-VAE module and reconstructed by the decoder. The image features and semantic features of the new class are trained using a feature discriminator to minimize the discrepancy between the two. The class prototype of the new class is enhanced through backpropagation.
[0057] Furthermore, in the prior variational autoencoder (P-VAE) module, the supporting features S are converted into distributions N of base classes and new classes, and class prototypes z are sampled from them, including the base class prototypes z. b and the new class prototype z n Using a VAE optimization model, let the class prototype z follow a prior distribution p(z), and the support features S follow a conditional distribution p(S|z). Through variational inference, the variational distribution q(z|S) is used to approximate the true posterior distribution p(z|S), minimizing the Kullback-Leibler divergence. The formula is as follows: This is equivalent to maximizing the lower bound of evidence, and the formula is: ELBO = E q(z|S) [logp(S|z)]-D KL (q(z|S)||p(z)), where the prior distribution p(z)=N(0,I) and the posterior distribution q(z|S)=N(μ,σ), and the parameters μ and σ are passed through the feature encoder F. enc Determined, the prototype z = μ + σ⊙ε is obtained using the reparameterization technique, where ε ~ N(μ,σ), and the reconstruction loss is... By optimizing the reconstruction loss L dec and KL divergence loss L KL Transform the supporting feature S into a distribution N and sample class prototypes.
[0058] Furthermore, the specific process by which the P-VAE module in S2 enhances the training of new classes through a feature discriminator and enriches semantic information using category priors to generate base and new class prototypes is as follows: Text features and image features are generated using the CLIP model, and the text embedding vector T is calculated. e and image embedding vector I e The cosine similarity is calculated using the following formula: Where t is the temperature parameter, and L is obtained through image loss. img and text loss L textThe combined symmetric loss L is obtained sem = L img + L text , wherein n represents the number of samples, p j represents the predicted probability that the j-th sample belongs to its true label category, logits[j, k] represents the actual category of the j-th sample, p i represents the predicted probability that the i-th sample belongs to its true label category, logits[i, k] represents the model output value of the i-th sample in category k, and labels[i] represents the actual category of the i-th sample. The semantic information of the category is obtained by training, the semantic information is input into the P-VAE module to obtain a new category feature distribution, a category prototype is sampled from the latent space and input into the decoder to obtain a reconstructed feature, and the new category reconstructed semantic feature and the real image feature x n is input into the discriminator, and the discriminator loss L dis = L real + L fake is minimized to reduce the gap between the new category reconstructed semantic feature and the real image feature, enhance the similarity between the new category prototype and the real feature, and generate category prototypes of the base category and the new category, wherein L real is the loss function of the real sample, and L fake is the loss function of the false sample, N is the number of samples, x n is a real sample of a new category, and G(z n ) is a false sample generated by the P-VAE module from the semantic feature of the new category.
[0059] Further, as shown in Figure 3 , the specific process of processing the support set and query set features through the mutual fusion module MFM in S3 is: receiving the features X q and X s extracted from the query set and the support set branch, calculating the similarity between the query branch and the support branch, reorganizing the support features and the query features based on the similarity, obtaining customized support features and query features through a convolution layer, using global average pooling and cosine similarity to calculate the fusion condition, integrating the reconstructed query features into the support features to enhance the feature diversity of the support branch, and integrating the reconstructed support features into the query features to promote the feature matching of the query features.
[0060] Further, the specific process of processing the support set and query set features through the mutual fusion module MFM in S3 is: the query set features and the support set features are represented as key-value pairs X q and X s , the support branch key K sand query branch value V q compatibility function f(K s ,V q )=Softmax(k s ·V q ), based on the weight f(K s ,V q ) reorganize support features V s and query features K q , through the convolution layer to get customized support features x s =Conv(f(K s ,V q )·K q ) and query features x q =Conv(f(K s ,V q )·V s ), using global average pooling GAP and cosine similarity calculation fusion conditions F s =D(GAP(X q ),X s ) and F q =D(GAP(X s ),X q ), wherein D is the cosine similarity distance measure, through the mask activation to the reconstructed query features x q integrated into the support features X s get X' s =F s ⊙x q +X s , the reconstructed support features x s integrated into the query features X q get X' q =F q ⊙x s +X q , enhance support feature diversity and promote query feature matching.
[0061] Further, the specific process of fusing the query set features with the generated class prototypes in S4 to obtain the fusion features is: the processed query set features are aggregated with the base class prototypes and new class prototypes generated by the P-VAE module, the channel attention mechanism is used to enhance the feature dimension related to the class, the semantic information of the class prototype is used to guide the target region identification in the query features, the complementary fusion of the query features and the class prototypes in the semantic and spatial dimensions is realized, and the fusion features containing the target class prior and the query image details are formed.
[0062] Further, the specific process of target detection based on the fusion features in S5 and outputting the detection result is: inputting the fusion features into a detection head, generating foreground candidate regions by using a region proposal network (RPN), extracting and aligning the candidate region features by RoIAlign, combining a classification head and a regression head to respectively perform target category prediction and boundary box regression, calculating the category probability by using a cross-entropy loss function for the classification head, optimizing the boundary box coordinates by using a smooth L1 loss function for the regression head, screening out the target detection result according to a pre-set confidence threshold, and outputting the detection result containing the target category, boundary box coordinates and confidence.
[0063] The embodiment details introducing the semantic prior of the CLIP pre-training model through a prior variational autoencoder (P-VAE), combining a feature discriminator to enhance the authenticity of the new class features; and realizing the bidirectional feature interaction of the support set and the query set by using a mutual fusion module (MFM) to improve the support feature diversity and the query matching capability.
[0064] Based on embodiment 1, the embodiment details the overall architecture of the prior variational autoencoder and the mutual fusion module and the interaction of the core modules, as shown in Figure 4 , and specifically as follows:
[0065] The joint feature optimization of the base class and the new class is realized by using a prior variational autoencoder (P-VAE) and a mutual fusion module (MFM), Figure 4 The left side is the support branch, the right side is the query branch, and both share a backbone network (such as ResNet-101) to extract basic features; the P-VAE module is located in the feature processing path of the support branch, and the input includes the features of the support image and the semantic features generated by the CLIP pre-training model, that is, the CLIP model generates a text embedding vector T e and an image embedding vector I e by using a text encoder and an image encoder respectively, after linear projection and L2 normalization, the cosine similarity is calculated , and combined with a symmetric loss function L sem = L img + L text to train the category semantic information. After the fusion of these semantic information and the support image features, the latent distribution N(μ,σ) is generated by the encoder of the P-VAE, the class prototype z is sampled, and the features are reconstructed by the decoder, and finally the feature discriminator is used to minimize the difference between the new class reconstructed features and the real features to enhance the representativeness of the new class prototype.
[0066] The bidirectional feature fusion mechanism of the MFM module is attached Figure 2 The mutual fusion module (MFM) on the right side reflects the interaction logic of the support branch and the query branch. The module receives the support features X s and the query features X q, first compute the similarity weight of both by the compatibility function (f(K s , V q ) = Softmax(k s ·V q ), where K s and V s are the key-value pairs of support branch, K q and V q are the key-value pairs of query branch. Based on the weight, the support feature V s is recombined with the query key K q after convolution layer to generate the customized support feature x s = Conv(f(K s ,V q )·K q ), and the customized query feature x q = Conv(f(K s ,V q )·V s ) is generated in the same way. Further, the fusion conditions F s = D(GAP(X q ), X s ) and F q = D(GAP(X s ), X q ) are calculated by global average pooling (GAP) and cosine similarity, and x q is integrated into the support feature in a masked activation manner (X' s = F s ⊙x q +X s ), and x s is integrated into the query feature (X' q = F q ⊙x s +X q ). This bidirectional fusion mechanism not only enhances the diversity of support features, but also preserves the details of query features through the perception of support sets, providing better feature representation for the subsequent target positioning and classification of the detection head (RPN+RoIAlign).
[0067] The embodiment describes in detail that the semantic prior and the discriminator are introduced by the P-VAE, the new class prototype representation is enhanced, the support and query features are combined by the MFM bidirectional fusion, the feature diversity and matching ability are improved, the base class and the new class are jointly trained, the data imbalance is alleviated, and the new class detection precision is significantly improved in the few-shot scene, which is better than most existing methods, and provides an effective solution for the rare sample target detection.
[0068] Embodiment 2
[0069] The embodiment is based on embodiment 1, and details of a few-shot target detection system based on a polymeric variational prototype are described, as shown in Figure 5 Specifically, as shown in the figure, it comprises a data set construction module, a prior variational autoencoder P-VAE module, a mutual fusion module MFM, a target detection module and a feature fusion module;
[0070] The data set construction module obtains a base class data set containing rich data and its base class set, and a new class data set containing a small number of instances in each class and its new class set, and a large number of small sample tasks containing support sets and query sets are constructed to form "scenarios" according to the scenario meta-learning paradigm, the base class data set is used for meta-training, and the new class data set and the base class subset are used for meta-fine-tuning;
[0071] The prior variational autoencoder P-VAE module is connected with the data set construction module, the base class and new class semantic features are generated through the CLIP pre-training model, the base class semantics and image features are fused into multi-modal features, the new class semantic features are used as image features in the initial stage, the new class prototype is generated and decoded, and the feature discriminator is used to train the new class image and semantic features to reduce the difference, and the new class prototype is enhanced through back propagation;
[0072] The mutual fusion module MFM is connected with the data set construction module and the P-VAE module, receives the features extracted by the query set and the support set branch, calculates the similarity between the two, reorganizes the features based on the similarity, obtains customized features through convolutional layers, uses global average pooling and cosine similarity calculation to fuse the conditions, integrates the reconstructed query features into the support features to enhance the diversity of the support branch, and integrates the reconstructed support features into the query features to promote the query feature matching;
[0073] The feature fusion module is connected with the P-VAE module and the MFM, the processed query set features and the base class and new class prototypes generated by the P-VAE module are aggregated, the class-related feature dimensions are enhanced through channel attention mechanism, the target region in the query feature is guided by using the semantic information of the class prototype, and the complementary fusion of semantic and spatial dimensions is realized to obtain the fusion features;
[0074] The target detection module is connected with the feature fusion module, the fusion features are input into the detection head, the region proposal network is used to generate foreground candidate regions, the candidate region features are extracted and aligned through RoIAlign, the classification head and the regression head are combined to perform target class prediction and bounding box regression respectively, the results are filtered according to the preset confidence threshold, and the detection results containing target class, bounding box coordinates and confidence are output.
[0075] The embodiment details the construction of a balanced data set through multi-module cooperation, and the fusion of semantic priori and discriminator to enhance the new class prototype by P-VAE. The bidirectional fusion of MFM features improves diversity and matching force. After feature fusion and detection head processing, the precise detection of new class targets under few-shot is realized, effectively alleviating data imbalance, improving feature representation and matching ability, and significantly improving detection accuracy and generalization in few-shot scenarios.
[0076] Embodiment 3
[0077] The embodiment based on embodiment 1 or 2 details the implementation process and effect of the few-shot target detection of the aggregated variational prototype, specifically including:
[0078] As shown in Figure 6 , the TSNE dimensionality reduction results of the traditional method and the new class prototype generated by AVP are shown. The prototype distribution of the traditional method on the left shows obvious discretization features, and there is significant overlap between different new class prototypes. For example, the feature cluster boundary of "motorcycle" and "bicycle" is blurred, reflecting the insufficient representation of the prototype due to the scarcity of new class samples. The prototype distribution generated by AVP on the right is more compact and the boundaries between classes are clear. The feature clusters of new classes such as "sofa" and "cow" are gathered together, and there is a clear separation from the base class feature clusters. This is due to the CLIP semantic priori and feature discriminator introduced by the P-VAE module. The former injects the essential attributes of the class through text-image contrastive learning, and the latter narrows the gap between reconstructed features and real features through adversarial training. As shown in Table 3, the P-VAE module alone can improve the representation ability of the new class prototype by 1.6%, and the TSNE visualization results and quantitative data provide evidence that AVP effectively solves the core problem of "fuzzy new class prototype" in traditional methods.
[0079] As shown in Figure 7 , the influence of hyperparameter λ on the performance of AVP in 1-shot scenarios is revealed, as shown in the following table.
[0080] Table 1 Influence of hyperparameter λ on the performance of AVP model
[0081]
[0082] As can be seen from the above table, when λ = 0.2, the model reaches 59.8% mAP in NovelSet1 of PASCAL VOC, which is the best in each test group; when λ = 0.05 and 0.1, the performance shows an upward trend, but does not reach the peak; when λ = 0.5 and 1, the performance decreases significantly, to 58.3% and 57.6% respectively. This phenomenon is due to the fact that λ controls the fusion weight of semantic prior and image features in P-VAE - when λ is too small, the semantic information is insufficient to guide the prototype generation, and the new class features still rely on limited samples, resulting in weak representation ability; when λ is too large, the semantic prior dominates the prototype generation too much, which may deviate from the real image features, causing a "semantic-visual" gap. Table 1 further shows that when λ = 0.2, it performs best in 1-shot, 5-shot and 10-shot scenarios, confirming the key role of this parameter in balancing semantic guidance and visual features, and also showing that the module design of AVP has good robustness to hyperparameters.
[0083] As shown in Figures 8-9 , intuitively shows the detection advantage of AVP in complex scenes, Figure 8 , the baseline method has missed detection or low confidence problem in detecting "obstructed children" and "small size ships", while AVP enhances the semantic understanding of support features to the occluded area through the bidirectional feature fusion of MFM module, successfully locates the child target and gives a confidence of more than 90%; in the detection of overlapping ships, the bounding box of AVP fits the target contour more closely, and the confidence is improved by about 15%. Figure 9 Further comparison of typical cases: in the "double festival bus" scene, the baseline method mistakenly identifies a single target as two, while AVP accurately distinguishes the target structure through the prototype constraint of P-VAE; in the "small size cat" detection, AVP breaks through the size sensitivity limitation of the baseline method due to insufficient samples, relying on semantic prior guidance. These visual results and quantitative data in Table 1 echo each other - AVP reaches 70.1% mAP in the 10-shot scenario of NovelSet1, which is 2.7% higher than the baseline, proving its generalization ability in complex scenes such as complex lighting, occlusion and scale change.
[0084] For the performance of AVP on MSCOCO dataset, as shown in Table 2 below: nAP is 16.8% when K = 10, and increases to 19.4% when K = 30, surpassing VFA, MetaFR-CNN and other methods. Although there is a gap of 23.1% with ICPE, it should be noted that ICPE adopts a query set customization prototype strategy, and the computational complexity is significantly higher than AVP. MSCOCO contains a large number of fine-grained differences in 80 categories (such as different breeds of "dog"), and AVP effectively captures the subtle commonalities of the same object through the feature interaction of the MFM module - for example, in the detection of "Labrador" and "Golden Retriever", the mAP of AVP is 9.2% higher than the baseline, while traditional methods often cause confusion due to insufficient feature diversity. In addition, although the performance improvement of AVP on COCO (about 3.2%) is smaller than that of PASCAL VOC (about 5.8%), considering that the category complexity of COCO is higher, this result still proves its effectiveness across datasets, especially in the scenario where "new classes are highly similar to base classes".
[0085] Table 2: Few-shot object detection performance comparison on MSCOCO dataset
[0086]
[0087] Ablation experiment data reveals the synergistic effect of AVP modules, as shown in Tables 3 and 4. When only P-VAE is used, the average mAP of NovelSet1 increases by 1.4%, verifying the optimization effect of semantic prior and discriminator on new class prototypes; when only MFM is used, it increases by 0.7%, indicating that bidirectional feature fusion can enhance the representation diversity of the support set; when both are combined, it increases by 2.0%, exceeding the sum of the effects of the individual modules, confirming the complementary mechanism of "prototype optimization-feature interaction". From a qualitative perspective, when P-VAE is missing, the new class prototype shows a "semantic drift" phenomenon in the TSNE graph, such as the "bus" prototype shifting towards "truck"; when MFM is missing, the matching ability of support features to query samples decreases, Figure 9 Baseline methods cannot recognize small-sized targets due to the lack of query feedback. These results form a logical closed loop with the SOTA performance in Table 1 - AVP solves the representation problem of "from nothing to something" for new class prototypes through the joint design of P-VAE and MFM, and improves the matching accuracy of "from something to optimal" through feature fusion, ultimately achieving performance breakthroughs in few-shot detection.
[0088] Table 3: Ablation experiment results of AVP core modules
[0089]
[0090] Table 4: Single ablation experiment results of P-VAE module
[0091]
[0092] The embodiment is illustrated by experimental visualization and data comparison, and the detailed description shows that AVP cooperates with P-VAE and MFM modules, significantly optimizes new class prototype representation, makes the prototype distribution more compact and the class distinction degree high, the super parameter optimization ensures the balance of semantic and visual features, enhances the model robustness, the detection result visualization shows that AVP has obvious advantages in detection of occlusion, small target and similar classes in complex scenes, and cross-dataset experiment verifies the adaptability of AVP in different complexity scenes. The ablation experiment proves the synergistic effect between the modules, effectively solves the problems of data imbalance and insufficient feature representation in few-shot detection, and improves the generalization ability of the model.
[0093] The above is only the preferred embodiment of the present application, which does not limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and variations; any changes, modifications, replacements, integrations and parameter changes of these embodiments within the spirit and principles of the present application, which can realize the same function without departing from the principles and spirit of the present application, fall within the protection scope of the present application.
Claims
1. A few-shot object detection method based on aggregated variational prototypes, characterized in that, The method comprises the following steps: S1, constructing a data set containing a base class and a new class, dividing the data set into a support set and a query set, wherein the base class contains a large number of samples, and the new class contains a small number of samples; S2, processing the support set samples by using a prior variational autoencoder P-VAE module, wherein the P-VAE module enhances the new class training through a feature discriminator, and generates class prototypes of the base class and the new class by using category prior to enrich semantic information; S3, processing the features of the support set and the query set by using a mutual fusion module MFM, wherein the MFM module enhances the diversity of the support features, and promotes the feature matching of the query samples; S4, fusing the query set features with the generated class prototypes to obtain fused features; S5, performing target detection based on the fused features, and outputting a detection result.
2. The few-shot object detection method based on poly variational prototype according to claim 1, wherein, The specific process of constructing data sets containing base classes and new classes and dividing into support sets and query sets in S1 is: obtaining a base class data set D containing rich data b and its base class set C b , obtaining a new class data set D containing k instances for each category n and its new class set C n , wherein C n is disjoint with C b , following the scenario meta-learning paradigm, constructing a large number of small sample tasks containing support sets and query sets in D b and D n respectively to form scenarios; the base class data set D b is used for meta-training, and subsets of the new class data sets D n and D b are used for meta-fine tuning.
3. The polymeric variational prototype based few-shot object detection method of claim 1, wherein, The specific process of processing the support set samples by using the prior variational autoencoder P-VAE module in S2 is as follows: generating semantic features of the base class and the new class by using a CLIP pre-training model, fusing the semantic features of the base class with image features of the base class to form multi-modal features as base class features, taking the semantic features of the new class as image features of the new class in an initial training stage, generating a class prototype of the new class by using the P-VAE module and reconstructing the class prototype by using a decoder, training the image features and the semantic features of the new class by using a feature discriminator to minimize the difference between the two features, and enhancing the class prototype of the new class through back propagation.
4. The few-shot object detection method based on poly variational prototype of claim 3, wherein, The support feature S is converted into the distribution N of base class and new class and the class prototype z is sampled from the distribution N in the prior variational autoencoder P-VAE module, including the base class prototype z b and the new class prototype z n The VAE optimization model is used, the class prototype z is subject to the prior distribution p(z), the support feature S is subject to the conditional distribution p(S|z), the true posterior distribution p(z|S) is approximated by the variational distribution q(z|S) through variational inference, the Kullback-Leibler divergence is minimized, and the formula is: which is equivalent to maximizing the evidence lower bound, the formula is: ELBO=E q(z|S) [logp(S|z)]-D KL (q(z|S)||p(z)), wherein the prior distribution p(z)=N(0,I), the posterior distribution q(z|S)=N(μ,σ), and the parameters μ and σ are determined by the feature encoder F enc , the prototype z=μ+σ⊙ε is obtained by using the reparameterization trick, wherein ε~N(μ,σ), and the reconstruction loss is obtained by optimizing the reconstruction loss L dec and the KL divergence loss L KL The support feature S is converted into the distribution N and the class prototype is sampled.
5. The polymeric variational prototype based few-shot object detection method according to claim 1 or 3, characterized in that, The specific process by which the P-VAE module in S2 enhances new class training through a feature discriminator and enriches semantic information using category priors to generate base and new class prototypes is as follows: Text and image features are generated using the CLIP model, and the text embedding vector T is calculated. e and image embedding vector I e The cosine similarity is calculated using the following formula: Where t is the temperature parameter, and L is obtained through image loss. img and text loss L text The merging yields a symmetric loss L sem =L img +L text The training process obtains semantic information about the categories. This semantic information is then input into the P-VAE module to obtain the feature distribution of the new classes. Class prototypes are sampled from the latent space and fed into the decoder to obtain reconstructed features. The semantic features of the new classes are then combined with the real image features x. n Input discriminator, loss L through discriminator dis =L real +L fake Minimize the gap between the reconstructed semantic features of the new class and the real image features, enhance the similarity between the new class prototype and the real features, and generate class prototypes for the base class and the new class, where L real For the loss function of the real samples, L fake The loss function is for fake samples.
6. The polymeric variational prototype based few-shot object detection method of claim 1, wherein, The specific process of processing support set and query set features through the mutual fusion module MFM in the S3 is: receiving features X extracted from the query set and the support set q and X s , calculating the similarity between the query branch and the support branch, reorganizing the support features and the query features based on the similarity, obtaining customized support features and query features through a convolution layer, using global average pooling and cosine similarity calculation fusion conditions, integrating the reconstructed query features into the support features to enhance the feature diversity of the support branch, and integrating the reconstructed support features into the query features to promote the feature matching of the query features.
7. The polymeric variational prototype based few-shot object detection method of claim 1 or 6, wherein, The specific process of processing support set and query set features through the inter-fusion module MFM in S3 is as follows: Query set features and support set features are represented as key-value pairs X, respectively. q and X s Calculate the branch key K s And query branch value V q The compatibility function f(K) s V q ) = Softmax(k s ·V q Based on weight f(K) s V q Recombination support feature V s and query feature K q Customized support features x are obtained through convolutional layers. s =Conv(f(K) s V q )·K q ) and query feature x q =Conv(f(K) s V q )·V s The fusion condition F is calculated using global average pooling (GAP) and cosine similarity. s =D(GAP(X) q ),X s ) and F q =D(GAP(X) s ),X q ), where D is the cosine similarity distance measurement, and the reconstructed query feature x is activated by masking. q Integration into support feature X s X' is obtained from s =F s ⊙x q +X s The reconstructed supporting feature x s Integrate into query feature X q X' is obtained from q =F q ⊙x s +X q This enhances support for feature diversity and facilitates query feature matching.
8. The polymeric variational prototype based few-shot object detection method of claim 1, wherein, The specific process of fusing the query set features with the generated class prototypes to obtain fused features in S4 is as follows: aggregating the processed query set features with the base class prototype and the new class prototype generated by the P-VAE module, enhancing the feature dimensions related to the categories by using a channel attention mechanism, guiding the target region identification in the query features by using the semantic information of the class prototype, realizing the complementary fusion of the query features and the class prototypes in the semantic and spatial dimensions, and forming fused features containing the target category prior and the query image details.
9. The polymeric variational prototype based few-shot object detection method of claim 1, wherein, The specific process of performing target detection based on the fused features and outputting a detection result in S5 is as follows: inputting the fused features into a detection head, generating foreground candidate regions by using a region proposal network RPN, extracting and aligning the candidate region features by using RoIAlign, combining a classification head and a regression head to respectively perform target category prediction and boundary box regression, calculating the category probability by using a cross-entropy loss function for the classification head, optimizing the boundary box coordinates by using a smooth L1 loss function for the regression head, screening out the target detection result according to a pre-set confidence threshold, and outputting the detection result containing the target category, the boundary box coordinates and the confidence.
10. A few-shot object detection system based on polymeric variational prototyping, suitable for any of claims 1-9, characterized in that, The method comprises a data set construction module, a prior variational autoencoder P-VAE module, a mutual fusion module MFM, a feature fusion module and a target detection module; The data set construction module is used for constructing a data set containing a base class and a new class, and dividing the data set into a support set and a query set, wherein the base class contains a large number of samples, and the new class contains a small number of samples; The prior variational autoencoder P-VAE module is connected with the data set construction module, and is used for processing the support set samples, enhancing the new class training through a feature discriminator, and generating class prototypes of the base class and the new class by using category prior to enrich semantic information; The mutual fusion module MFM is connected with the data set construction module and the P-VAE module, and is used for processing the features of the support set and the query set, enhancing the diversity of the support features, and promoting the feature matching of the query sample; The feature fusion module is connected with the P-VAE module and the MFM, and is used for fusing the query set features with the generated class prototypes to obtain fusion features; The target detection module is connected with the feature fusion module, and is used for performing target detection based on the fusion features and outputting a detection result.
Citation Information
Cited By
Deep learning hidden danger automatic identification method and system with few samples and strong generalization
CN121582551A
Metalearning-based small sample ship target detection method, system and equipment
CN122223309A
A small sample ship target detection method, system and device based on meta learning
CN122223309B