Small sample target detection method for prototype interaction enhancement
By introducing a class-aware attention mechanism and cross-class feature interaction, and optimizing the prototype vector, the problems of insufficient prototype representativeness and insufficient utilization of inter-class relationships in existing small sample target detection are solved, thereby improving the detection accuracy of new category targets.
Patent Information
- Application Number
- CN202511094059.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-14
AI Technical Summary
Existing small-sample target detection technologies have shortcomings in prototype construction and utilization. They lack sufficient integration of features at different scales of supporting samples, resulting in weak prototype representation capabilities. Furthermore, they lack utilization of shared information between classes, making it impossible to comprehensively consider the relationships between different classes during the detection process. This limits the model's detection performance for new class targets under small-sample conditions.
A class-aware attention mechanism is introduced, prototype vectors are optimized through the Support Prototype Enhancement Network (SPA-Net), and cross-class feature interaction is achieved using the Query-Support Multi-Class Aggregation Module (QSMCA), thereby improving the model's ability to detect new categories.
By fully exploring supporting sample information and utilizing inter-class correlations, the model's detection performance for new target categories was improved, the representativeness and discriminativeness of the prototype were enhanced, and the detection accuracy was increased.
Smart Images

Figure CN120953694A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of computer vision and deep learning technology, and specifically relates to a prototype interaction-enhanced few-shot target detection method, which aims to improve the detection performance of target detection models for new types of targets when training samples are extremely limited. Background Technology
[0002] Few-Shot Object Detection (FSOD) refers to the accurate detection of novel object categories with only a very small number of labeled samples. In traditional object detection, models often require a large amount of labeled data to achieve good detection results. However, in many real-world scenarios, acquiring large amounts of labeled data is both difficult and time-consuming, limiting the model's generalization ability to new or rare object categories. To address this, Few-Shot Learning (FSL) methods have emerged, aiming to enable models to quickly adapt to new tasks with only a small number of training samples. As an important branch of FSL, Few-Shot Object Detection typically borrows from meta-learning paradigms to achieve rapid adaptation to new object categories.
[0003] Meta-learning-based few-shot object detection methods typically operate through a "base class training + new class fine-tuning" approach. Specifically, an initial detection model is first trained using a base class containing a large number of samples. Then, it is fine-tuned or adapted using only a small number of support samples for new classes, enabling the model to detect these new classes of objects. To effectively utilize the information from the support samples, many existing techniques introduce a support-query feature interaction mechanism: during the model's detection phase, features from the support set (containing a small number of samples from the new class) are combined with features from the query image (the image to be detected) to guide the model's classifier to better distinguish the target class.
[0004] Currently, most meta-learning FSOD methods employ a category-related aggregation strategy when supporting-query feature interactions. Category-related aggregation refers to calculating the similarity or correlation between supporting features and query features for each category separately. This is typically achieved by obtaining a prototype vector for each category through dot product matching, and then comparing candidate region features in the query image with each category's prototype to determine the category. While this category-independent aggregation approach is simple, it has significant drawbacks: Firstly, since each category's prototype is generated from only a few supporting samples of that category, it lacks sufficient representation of features at different scales and variations in different instances, resulting in insufficient representativeness and difficulty in covering the diversity of the target within that category. Secondly, treating each category separately ignores potential correlations and differences between categories, preventing the model from acquiring shared information between prototypes of different categories. For example, different categories may share certain feature patterns, or some categories may be easily confused. If the detection process does not consider these inter-class relationships, classification errors or confusion often occur, reducing detection performance.
[0005] Some studies have recognized these problems and proposed improvements, such as optimizing prototypes by extracting shared information from similar samples in the support set, or introducing class-independent randomness during the aggregation process to increase prototype diversity. However, most of these methods are still limited to interactions between pairs of classes and do not fully utilize the global relationships of one class relative to all other classes. Simple single-class prototype matching methods perform poorly when dealing with the simultaneous occurrence of multiple classes or diverse class features, easily leading to decreased model generalization ability and a decline in the accuracy of new class detection.
[0006] In summary, existing few-sample object detection techniques still have shortcomings in the construction and utilization of prototypes: they lack sufficient fusion of features at different scales supporting samples, resulting in weak prototype representation capabilities; and they lack utilization of shared information between classes, failing to comprehensively consider the relationships between categories during the detection process. These shortcomings ultimately limit the model's detection performance for new categories of objects under few-sample conditions. Therefore, an improved technical solution is urgently needed to enhance the representativeness and discriminativeness of prototypes, fully explore inter-class correlations, and thus improve the performance of few-sample object detection. Summary of the Invention
[0007] To address the problems existing in the above-mentioned background technology, the present invention provides a prototype-enhanced small sample target detection method. By introducing a class-aware attention mechanism into the detection model, the method optimizes the supporting sample prototypes and performs multi-class fusion, thereby improving the model's ability to detect new categories.
[0008] The technical solution of this invention is as follows:
[0009] A prototype-interactive enhanced few-shot object detection method includes:
[0010] Step 1: Support set feature extraction, including: obtaining support set images and their annotation information in the few-shot task; inputting each image in the support set into a pre-trained backbone network for feature extraction; for each image in the support set, performing RoI alignment on the corresponding feature map based on the labeled ground truth candidate box target position to obtain the support RoI feature map.
[0011] Step 2, Support Prototype Enhancement: Input the extracted RoI support features into the support prototype enhancement network to generate enhanced prototype vectors for the corresponding target categories, thus obtaining semantically enhanced RoI feature maps;
[0012] Step 3: Query Image RoI Feature Extraction: Input the query image to be detected into a backbone network with the same shared parameters as the support set for feature extraction, and then obtain the query RoI feature map through the RoI alignment method;
[0013] Step 4: Query support for multiple aggregations: Input all the semantically enhanced RoI features obtained above, along with the query RoI features, into the query-support for multiple aggregations module;
[0014] Step 5, Classification and Regression: The output of Step 4 is fed into the bounding box classification and regression head to calculate the classification loss, regression loss and meta-loss respectively.
[0015] A readable storage medium storing program instructions that, when read and executed by a computing device, cause the computing device to perform a prototype-interactive enhanced few-sample target detection method.
[0016] Beneficial effects:
[0017] In summary, the method of this invention, through the cooperation of the SPA-Net and QSMCA modules, achieves full mining and utilization of supporting sample information. On the one hand, the prototype enhancement network improves the ability of prototype vectors to describe the target category; on the other hand, the query-support multi-class aggregation module incorporates comparative considerations of other categories into the discrimination decision of each category, that is, it introduces a class-aware attention fusion mechanism, which makes up for the deficiency of isolated processing of various types of information in existing methods. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating the small sample target detection method with enhanced prototype interaction of the present invention.
[0019] Figure 2 This is a schematic diagram of the structure supporting the prototype enhancement network (SPA-Net) in an embodiment of the present invention;
[0020] Figure 3This is a schematic diagram of the Query-Support Multiple Aggregation Module (QSMCA) in an embodiment of the present invention. Detailed Implementation
[0021] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0022] To further illustrate the technical solution of the present invention, the implementation methods of the present invention will be described in detail below with reference to specific embodiments.
[0023] like Figure 1 As shown, the prototype interaction-enhanced few-sample target detection method of the present invention generally includes the following steps:
[0024] Step 1: Support set feature extraction, including: obtaining support set images and their annotation information in the few-shot task; inputting each image in the support set into a pre-trained backbone network for feature extraction; for each image in the support set, performing RoI alignment on the corresponding feature map based on the labeled true candidate box target position, thereby obtaining the support RoI feature map.
[0025] The support set contains several target categories, with only a very small number (e.g., K) labeled samples for each category.
[0026] Step 2, Support Prototype Enhancement (SPA-Net): Input the extracted RoI support features obtained above into the Support Prototype Enhancement Network (SPA-Net) to generate semantically enhanced RoI support features corresponding to the target category. Specifically, as follows... Figure 2 As shown, firstly, the Support Feature Input Feature Pyramid (FPN) extracts features at four different scales. Then, based on known ground truth candidate boxes, RoI alignment is performed on these four sets of features at different scales to generate four sets of RoI support feature representations at different scales. The SPA-Net module then performs comprehensive processing on these four sets of RoI support feature representations at different scales, including:
[0027] (1) Preprocessing: For the four sets of supporting RoI features at different scales, the preprocessing step enhances the network's learning of local features at each scale. Therefore, an adjustable degree of freedom operator is introduced in this embodiment. This is used to control the contribution of features at each scale. This operator can be viewed as a set of learnable parameters used to amplify or reduce the proportion of a certain scale feature in the fusion result, thus enabling the network to adaptively select the feature scale information most beneficial to the prototype. Typically, Set as , , or The convolution kernel. According to... Figure 2 This operation can be summarized as follows:
[0028] ,
[0029] in, These are support RoI feature representations at different scales. It is a preprocessed feature map supporting RoI. It is a convolution operation, where, This refers to the number of layers output by the FPN, typically... = 1, 2, 3, 4.
[0030] (2) Feature Aggregation: The preprocessed feature maps of four different scales supporting RoIs are aggregated to obtain a set of feature maps containing rich scale information. This invention uses concatenated aggregation to better preserve the independent semantic information of each set of scale features. According to... Figure 2 This operation can be summarized as follows:
[0031] ,
[0032] in, It is a serial splicing operation. These are the aggregated supporting features. These are four different scales of supporting RoI feature maps.
[0033] (3) Refinement: Since the preprocessing step extracts single-scale local information, while the goal of SPA-Net is to make the network treat each RoI feature equally, it is necessary to further introduce a new learnable degree of freedom operator. Global refinement is performed on the supporting RoI feature maps. This allows for the extraction of shared semantics at multiple scales and produces semantically enhanced supporting RoI feature maps. Figure 2 The above operations can be summarized as follows:
[0034] ,
[0035] in, The final output is a semantically enhanced support RoI feature map.
[0036] Step 3: Query Image RoI Feature Extraction: The query image to be detected is input into a backbone network with the same shared parameters as the support set for feature extraction, resulting in a query feature map. Then, a number of candidate regions (RoIs) are generated on the query feature map using the Faster R-CNN conventional object detection mechanism (RPN, Region Proposal Network). RoI alignment is then performed to obtain the query RoI feature map. In this invention, the size of the candidate region RoI feature map for this query set is set to... These candidate region features will be used to match with supporting prototypes to determine their category and location boundaries.
[0037] Step 4: Query Support for Multi-Class Aggregation (QSMCA): Input all the semantically enhanced supporting RoI features obtained above, along with the query RoI features, into the Query Support for Multi-Class Aggregation (QSMCA) module. This module utilizes the multi-head self-attention mechanism in the Transformer architecture to achieve cross-class feature interaction. The specific process is as follows (e.g., Figure 3 As shown): According to the formula of the self-attention mechanism:
[0038] ,
[0039] in, It is the output of the multi-head attention mechanism. It is a multi-head attention mechanism. It is an activation function. It is the feature embedding dimension. It is the height of the feature of input Q. It is the width of the feature of input Q. It is the feature dimension. It is the height of the features of input K and V. It is the width of the features of the input K and V.
[0040] The query-supporting multi-class aggregation module utilizes the Transformer's self-attention mechanism to calculate a classification weight matrix. The weight matrix assigns corresponding weights to the prototype vectors of each target category based on the features of the query image, so as to generate a class-specific prototype matrix by linearly combining the prototype vectors of all categories according to their weights.
[0041] The fusion of class-specific prototype matrix and query image features includes using the class-specific prototype matrix as the weight or feature input of the classifier in the few-sample detection model, matching or concatenating it with the feature vector of the candidate region in the query image, thereby completing the category discrimination and localization regression of the target object in the query image.
[0042] This process can specifically include:
[0043] (1) Prototype encoding layer: Based on the basic structure of Transformer, such as Figure 3 In the Prototype Encoding Layer (PEL) of QSMCA, the semantically enhanced supporting RoI feature map obtained in step two is first processed... The input is processed using a multi-head attention mechanism. This mechanism refines the effective information of multiple class support features, achieves cross-attention between feature vectors, and generates more representative class prototypes. Then, in order to obtain category-related feature vectors, the HW space dimension is adjusted. Perform global average pooling to obtain (in, It is set up to expand the matrix dimensions and has no practical meaning. It refers to the batch size. (This refers to the number of channels), and the result of global average pooling. This is the supporting class prototype, which will participate in subsequent aggregations.
[0044] (2) Multi-class aggregation layer: The query RoI feature map generated in step three is first input into Figure 3 The multi-class aggregation layer (MCAL) in the middle is refined to capture global information and generate .Will The Q-value of the cross-attention mechanism in the multi-class aggregation layer is input, and the output of the prototype encoding layer is used as the input. As the cross-attention mechanism in the multi-class aggregation layer, K and V are important. Note that the input to K requires first passing through a sigmoid function to form a class matcher, and then... Through interaction, the multi-head attention mechanism can focus on regions highly correlated with the supporting class prototype (such as the target object location) in the query RoI features, while weakening the responses of regions that do not match the supporting features, generating the original matching weight matrix. Finally, the original matching weight matrix is transformed into probabilities using the softmax function to obtain the prototype matching weights. This indicates the degree of matching between the query image to be detected and each supporting prototype. By performing a weighted aggregation with V, we can obtain a supporting prototype information representation that conforms to the query characteristics. :
[0045] ,
[0046] in, This represents the cross-attention mechanism. This represents the operation of the Sigmoid function. It is a multi-head attention mechanism. It is the embedding dimension of the multi-head attention mechanism.
[0047] Then, to and The final Query RoI feature map is generated by using Hadamard product aggregation and processing through a linear feedforward network. The data is then fed into the bounding box classification and regression head for classification and regression.
[0048] Step 5, Classification and Regression: Output of Step 4 The bounding box classification and regression heads are fed into the database to calculate the classification loss, regression loss, and meta-loss, respectively. The cross-entropy loss function is used for both the classification and meta-losses, while the smoothed L1 loss function is used for the regression loss. The final results are then output.
[0049] Since this invention employs a meta-learning paradigm of N categories and K samples (N-way K-shot) for training, the five steps described above constitute an iterative training process. This process is repeated until the maximum number of iterations is reached, at which point training terminates. Specifically, in the fine-tuning stage of new classes on complex datasets (such as COCO), CLIP is used to fine-tune the bounding box classification regression head to achieve optimal results. The baseline in this implementation is set to Meta R-CNN, and the strong baseline (after CLIP fine-tuning) is named... All experiments on the COCO dataset were conducted in The above will be carried out.
[0050] The few-shot object detection network employs a meta-learning training paradigm, divided into two phases: base class training and new class fine-tuning. In the base class training phase, a large number of samples from the base classes are used to construct an N-way K-shot task to train the parameters of the prototype augmentation network module and the query-support multi-class aggregation module. Then, in the new class fine-tuning phase, new classes that do not intersect with the base class classes are used for fine-tuning, where the number of K varies depending on the validation strategy.
[0051] A readable storage medium storing program instructions is also provided, which, when read and executed by a computing device, causes the computing device to perform a prototype-interactive enhanced few-sample target detection method.
[0052] Next, this implementation (PIA) will demonstrate its experimental results on the PASCAL VOC and MS COCO datasets, and compare them with existing mainstream methods. The evaluation metric is AP50 (mean accuracy), proving its effectiveness. Table 1 below shows the validation results on the PASCAL VOC dataset, where the performance outperforms most mainstream methods:
[0053] Table 1. Validation results on the PASCAL VOC dataset.
[0054]
[0055] Continued from Table 1: Validation results on the PASCAL VOC dataset
[0056]
[0057] Secondly, this implementation method compares the validation results on the MS COCO dataset, with 10 samples and 30 samples respectively, as shown in Table 2:
[0058] Table 2 Validation results on the MS COCO dataset
[0059]
[0060] In Table 2, AP represents the average precision.
[0061] The effects of SPA-Net and QSMCA were verified through ablation experiments in this implementation method, as shown in Table 3:
[0062] Table 3 Ablation Experiment Results
[0063]
[0064] In Table 3, S represents the average accuracy of targets with bounding box pixel size less than 32 × 32, M represents the average accuracy of targets with bounding box pixel size between 32 × 32 and 96 × 96, and L represents the average accuracy of targets with bounding box pixel size greater than 96 × 96.
[0065] For the degree-of-freedom operator in SPA-Net The impact on actual effects was also verified through ablation experiments, as shown in Table 4:
[0066] Table 4 Ablation Experiment Results
[0067]
[0068] The above embodiments are preferred implementations of the present invention. In addition, the present invention can be implemented in other ways. Any obvious substitutions without departing from the concept of the present technical solution are within the protection scope of the present invention.
[0069] To facilitate understanding by those skilled in the art of the improvements of this invention over the prior art, some of the accompanying drawings and descriptions have been simplified. Parts not described in detail are well-known to those skilled in the art. Furthermore, for clarity, some other elements have been omitted from this application. Those skilled in the art should realize that these omitted elements may also constitute the content of this invention.
Claims
1. A prototype-interactive enhanced few-sample target detection method, characterized in that, include: Step 1: Support set feature extraction, including: obtaining support set images and their annotation information in the few-shot task; inputting each image in the support set into a pre-trained backbone network for feature extraction; for each image in the support set, performing RoI alignment on the corresponding feature map based on the labeled ground truth candidate box target position to obtain the support RoI feature map. Step 2, Support Prototype Enhancement: Input the extracted RoI support features into the support prototype enhancement network to generate enhanced prototype vectors for the corresponding target categories, thus obtaining semantically enhanced RoI feature maps; Step 3: Query Image RoI Feature Extraction: Input the query image to be detected into a backbone network with the same shared parameters as the support set for feature extraction, and then obtain the query RoI feature map through the RoI alignment method; Step 4: Query support for multiple aggregations: Input all the semantically enhanced RoI features obtained above, along with the query RoI features, into the query-support for multiple aggregations module; Step 5, Classification and Regression: The output of Step 4 is fed into the bounding box classification and regression head to calculate the classification loss, regression loss and meta-loss respectively.
2. The prototype-interactive enhanced few-sample target detection method according to claim 1, characterized in that, Step two includes: Preprocessing: For the RoI feature representation at each scale, the support prototype enhancement network module fuses the features at different scales to obtain four different scales of support RoI feature maps; Feature aggregation: The four preprocessed RoI feature maps at different scales are aggregated to obtain a set of RoI feature maps containing rich scale information; Refinement: A new learnable degree-of-freedom operator is introduced to globally refine the feature maps supporting RoIs.
3. The prototype-interactive enhanced few-sample target detection method according to claim 2, characterized in that, The fusion process includes combining fine-grained and coarse-grained features by using weighted summation or concatenation for feature maps of different proportions from the feature map pyramid network.
4. The prototype-interactive enhanced few-sample target detection method according to claim 2, characterized in that, The aggregation process includes: using average pooling and attention weighting methods for intra-class aggregation, including: using average aggregation to obtain the base vector of the class prototype, and then combining the attention mechanism to fine-tune the contribution weight of each instance to the prototype, so as to ensure that the important features of each supporting sample can be reflected in the final prototype.
5. The prototype-interactive enhanced few-sample target detection method according to claim 1, characterized in that, Step four includes: Prototype encoding layer: First, the enhanced support RoI feature map obtained in step two is input into the multi-head attention mechanism for processing to generate class prototypes; global average pooling is performed on the class prototypes in the HW space dimension to obtain the global average pooling result; Multi-class aggregation layer: The query RoI feature map generated in step three is first input into the multi-head attention mechanism of the multi-class aggregation layer for refinement, capturing global information. The global information is input into the Q of the cross-attention mechanism in the multi-class aggregation layer, and the output of the prototype encoding layer is used as the K and V of the cross-attention mechanism in the multi-class aggregation layer. Then, through... Through interaction, the multi-head attention mechanism can focus on regions highly correlated with supporting class prototypes in the query RoI features and weaken the responses of regions that do not match the supporting features, generating the original matching weight matrix; finally, the original matching weight matrix is transformed into probabilities through the softmax function to obtain the prototype matching weights. This indicates the degree of matching between the query image to be detected and each supporting class prototype; We perform a weighted aggregation with V to obtain a supporting prototype information representation that matches the query characteristics. ; right and The final Query RoI feature map is generated by using Hadamard product aggregation and processing through a linear feedforward network. .
6. The prototype-interactive enhanced few-sample target detection method according to claim 1, characterized in that, In step five, both the classification loss and the meta-loss use the cross-entropy loss function, while the regression loss uses the smoothed L1 loss function.
7. The prototype-interactive enhanced few-sample target detection method according to claim 1, characterized in that, The method employs a meta-learning paradigm with N categories and K samples for training.
8. The prototype-interactive enhanced few-sample target detection method according to claim 1, characterized in that: The backbone network adopts a feature pyramid network structure, thereby outputting feature maps at different scales. The prototype enhancement network module supports the introduction of operators with adjustable degrees of freedom to perform weighted fusion of features at different scales, so as to generate more representative category prototype vectors.
9. The prototype-interactive enhanced few-sample target detection method according to claim 7, characterized in that: The proposed few-sample object detection method adopts a meta-learning training paradigm, which is divided into two stages: base class training and new class fine-tuning. In the base class training stage, a large number of samples are used to construct N categories and K sample numbers for the base categories, which are used to train the parameters of the prototype enhancement network module and the query-support multi-class aggregation module. Then, in the new class fine-tuning stage, new classes that do not intersect with the base class categories are used for fine-tuning, where the number of K varies with the verification strategy.
10. A readable storage medium storing program instructions, characterized in that, When the program instructions are read and executed by the computing device, the computing device performs the prototype interaction-enhanced few-sample target detection method as described in any one of claims 1-9.
Citation Information
Cited By
Metalearning-based small sample ship target detection method, system and equipment
CN122223309A