Small sample target detection method based on dynamic potential feature guidance
By introducing dynamic latent feature guidance technology into the small sample object detection method, using latent feature reconstruction and dynamic multi-scale similarity guidance modules, the problem of low detection accuracy of small sample object in the existing technology is solved, and higher detection accuracy and model generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510049551.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-13
AI Technical Summary
The existing small sample object detection method is difficult to effectively capture the complex relationship between support features and query features when there are few samples, resulting in a decrease in detection accuracy.
A small sample object detection method based on dynamic latent feature guidance is adopted, and the latent feature reconstruction module and dynamic multi-scale similarity guidance module are used to enrich the feature representation capabilities, dynamically adjust the internal weight of the supported features, and highlight key information highly related to the query image.
The adaptability and detection accuracy of the model in small sample scenarios is improved, the prediction ability of the target is enhanced, and the overall detection effect and the generalization ability of the model are improved.
Smart Images

Figure CN119992174A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a small sample target detection method based on dynamic potential feature guidance. Background Art
[0002] In recent years, deep learning-based object detectors have achieved remarkable results, which mainly rely on a large amount of high-quality training data and accurate bounding box annotations. However, this process is costly, and data labeling is cumbersome and time-consuming. In contrast, humans can quickly learn new concepts with a small amount of data. In order to reduce the cost of manual labeling and narrow the gap between detectors and human learning ability, small sample object detection came into being. Its goal is to achieve efficient and accurate detection under the condition of scarce samples, so as to better cope with various challenges in real-world applications.
[0003] Traditional small-sample object detection methods usually follow a two-stage training paradigm. In the first stage, a large number of basic samples are used to train the model to build a general object detector. In the second stage, fine-tuning is performed on a small number of new samples in the hope of improving detection performance when data is extremely scarce.
[0004] However, due to the limited number of new class samples, the model may be insufficient in representing the supporting features, which in turn limits the performance of the detector. In addition, the imbalance of the data makes it easier for the model to learn the base class features during training, while ignoring the key information of the new class, thereby reducing the detection accuracy of the new class. To solve the above problems, some methods try to improve the representation ability by weighting and integrating the supporting features and the query features, but these methods often rely on fixed feature abstraction forms, fail to make full use of information at different scales, and have difficulty capturing the complex relationship between the supporting features and the query features, thus limiting the adaptability of the model in small sample scenarios. Summary of the invention
[0005] The present invention solves the problems existing in the prior art and provides a small sample target detection method based on dynamic potential feature guidance.
[0006] The technical solution adopted by the present invention is a small sample target detection method based on dynamic potential feature guidance, the method obtains a sample data set, constructs a small sample target detection model based on dynamic potential feature guidance; freezes the local parameters of the model after training the model with base class data in the sample data set, and adjusts (fine-tunes) the model with new class data in the sample data set;
[0007] Input the data to be tested into the adjusted model to obtain the small sample target detection results.
[0008] Preferably, the small sample target detection model includes a support branch and a query branch;
[0009] The query branch includes a backbone network, an RPN module, and a region interest alignment module which are arranged in sequence, wherein the RPN module is used to generate a candidate frame, and the output of the region interest alignment is aggregated and output with the output of the support branch;
[0010] The support branch includes a backbone network and a fusion layer. The outputs of the query branch and the backbone network of the support branch are aligned and then input into the fusion layer. A potential feature reconstruction module and a dynamic multi-scale similarity guidance module are sequentially arranged after the fusion layer.
[0011] Preferably, the query branch and the support branch share backbone network parameters.
[0012] Preferably, the output of the backbone network of the query branch is dimensionally aligned with the output of the backbone network of the support branch through a global average pooling layer and an expansion module.
[0013] Preferably, the latent feature reconstruction module comprises an encoder and two parallel convolution blocks arranged in sequence, and the outputs of the two parallel convolution blocks are input into a decoder after re-parameterization.
[0014] Preferably, the dynamic multi-scale similarity guidance module includes a multi-scale information generator for generating features of different scales, and the features of different scales are input into the multi-similarity feature guidance module, spliced and output after being processed by multiple similarity layers; the processing here includes weighted fusion and guidance.
[0015] Preferably, the first branch reduces information loss and learns high-resolution representation of features through multi-branch convolution operations; the second branch focuses on low-resolution representation of features through global maximum pooling and deconvolution operations; the third branch combines global average pooling and channel attention mechanism to aggregate feature representations of the first and second branches to enhance information representation capabilities.
[0016] Preferably, similarity scores between features of different scales are calculated based on a hybrid similarity metric, internal weights of supporting features are dynamically adjusted, and features of three different scales are concatenated to obtain dynamically optimized multi-scale supporting features.
[0017] Preferably, the model is trained with the base class data in the sample data set, and the loss function is
[0018] L=L rpn +L LFR +L meta +L cls +L reg
[0019] Among them, L rpn is the loss of RPN, L LFR is the reconstruction loss of the latent feature reconstruction module, L metais the meta-classification loss of the model (the loss in the baseline model), L cls is the classification loss (common image classification loss), L reg is the regression loss.
[0020] Preferably, after training the model with the base class data in the sample data set, the local parameters of the model are frozen as the parameters of the backbone network, the RPN module and the region of interest alignment module.
[0021] The present invention relates to a small sample target detection method based on dynamic potential feature guidance. The method comprises the following steps: obtaining a sample data set, constructing a small sample target detection model based on dynamic potential feature guidance; freezing local parameters of the model after training the model with base class data in the sample data set, and adjusting the model with new class data in the sample data set; inputting data to be detected into the adjusted model to obtain a small sample target detection result.
[0022] The beneficial effects of the present invention are:
[0023] (1) A latent feature reconstruction (LFR) module is proposed to add latent information contained in the image to the query features and support features, thereby enriching the representation ability of the features and compensating for the information loss caused by insufficient samples. This module encodes and reconstructs the features and mines the implicit information in the latent space, making the feature representation more delicate and diversified. This not only improves the feature expression ability and enables the model to effectively learn key information when faced with a small number of samples, but also provides richer and more accurate information support for subsequent detection tasks, thereby improving the overall detection effect and the generalization ability of the model.
[0024] (2) A dynamic multi-scale similarity guidance (DMSG) module is designed to dynamically adjust the internal weights of supporting features according to the needs of the query image, thereby highlighting the key information that is highly relevant to the query image and effectively suppressing useless background noise and possible occlusion interference. The DMSG module extracts multi-level information through a multi-scale information generator to enhance the model's ability to understand complex scenes and ensure that key features can be captured at different perspectives and scales. The DMSG module uses multiple similarity layers to quantify and model the similarity relationship between features at different levels, which can more accurately capture the connection between supporting features and query images and highlight key prediction information. Through the extraction of multi-scale information and the guidance of multi-similarity, the model's effective use of supporting features is significantly improved, thereby enhancing its prediction ability for the target. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic diagram of the overall process of the present invention.
[0026] Figure 2 It is a model diagram of the present invention.
[0027] Figure 3 It is a structural diagram of the multi-scale generator module of the present invention.
[0028] Figure 4 It is a structural diagram of the multi-similarity feature guiding module of the present invention. DETAILED DESCRIPTION
[0029] The present invention is further described in detail below in conjunction with embodiments, but the protection scope of the present invention is not limited thereto.
[0030] The present invention relates to a small sample target detection method based on dynamic potential feature guidance, and the method mainly comprises the following steps:
[0031] (1) Obtain a sample data set;
[0032] (2) Construct a small sample target detection model guided by dynamic latent features;
[0033] (3) Freeze the local parameters of the model after training the model with the base class data in the sample data set;
[0034] (4) Adjust the model with new class data in the sample data set;
[0035] (5) Input the data to be tested into the adjusted model to obtain the small sample target detection results.
[0036] The following is a detailed description of the steps.
[0037] (1) Obtain a sample data set;
[0038] Obtain the target detection dataset and divide it into a base class dataset and a new class dataset to create a small sample detection task;
[0039] The data in the sample data set is obtained from more than one data set.
[0040] In the implementation process of the present invention, PASCAL VOC and MS COCO datasets are selected as datasets;
[0041] For the PASCAL VOC dataset, three random split settings consistent with the TFA model are used. Each split contains 20 categories, 15 of which are used as base classes and 5 as new classes. Each category in the new class contains K = 1, 2, 3, 5, and 10 annotated instances.
[0042] For the MS COCO dataset, 20 categories overlapping with PASCAL VOC are selected from the 80 object categories as new categories, and the remaining 60 categories are used as base categories, and the K values are set to 10 and 30;
[0043] Use the divided data sets to create small sample detection tasks, each of which contains supporting sample data and query sample data.
[0044] (2) Construct a small sample target detection model guided by dynamic latent features;
[0045] The small sample target detection model includes a support branch and a query branch;
[0046] The query branch includes a backbone network, an RPN module, and a regional interest alignment module which are arranged in sequence, and the output of the regional interest alignment is aggregated and output with the output of the support branch;
[0047] The support branch includes a backbone network and a fusion layer. The outputs of the query branch and the backbone network of the support branch are aligned and then input into the fusion layer. A potential feature reconstruction module and a dynamic multi-scale similarity guidance module are sequentially arranged after the fusion layer.
[0048] In the present invention, query features (output of region of interest alignment) and multi-scale support features (output of dynamic multi-scale similarity guidance module) are aggregated, and the generated aggregated features are input into the detection head, which is used to classify the target (class head) and regress the bounding box (box head), and finally output the prediction result.
[0049] Here, ResNet-101 is used as the backbone network to extract the basic features of the image. The query branch and the support branch share the backbone network parameters. After feature extraction using the backbone network, the query features are refined using a global average pooling layer and expanded to the same dimension as the support features.
[0050] The latent feature reconstruction module comprises an encoder and two parallel convolution blocks arranged in sequence, and the outputs of the two parallel convolution blocks are input into a decoder after re-parameterization.
[0051] In the present invention, the query features and supporting features are reconstructed by the latent feature reconstruction module to add the latent information contained in the image. The latent feature reconstruction (LFR) uses variational inference to encode the query features and supporting features, maps the conditional distribution of the samples to a multivariate Gaussian distribution, and then reconstructs the latent code z from the distribution using the reparameterization technique. z contains additional information representation. The calculation formula of z is as follows:
[0052]
[0053] In the training phase, LFR follows the reparameterized setting, but in the inference phase, μ is directly used as the sampling result, where ε~N(0,1), ⊙ represents the element-by-element product, and the mean μ and variance σ of the approximate posterior are obtained through two convolution blocks, satisfying
[0054]
[0055] in, is a standard 3×3 convolution, F enc is the encoder in LFR, and finally resamples the query features and support features in z to achieve feature reconstruction. The LFR reconstruction loss L LFR The calculation formula is as follows,
[0056]
[0057] Where μ is the mean of the approximate posterior, σ is the variance of the approximate posterior, and X / s are the query features and support features extracted by the backbone network, X C / s is the feature reconstructed by LFR, but it is not the final feature used. The final feature used is the potential code z.
[0058] The dynamic multi-scale similarity guidance module includes a multi-scale information generator for generating features of different scales. The features of different scales are input into the multi-similarity feature guidance module, and are spliced and output after being processed by multiple similarity layers.
[0059] Here, features of different scales are used to characterize global features, regional features, and local features.
[0060] In the present invention, a dynamic multi-scale similarity guidance (DMSG) module is used to reduce the detector's bias towards the base class;
[0061] In the first part, the multi-scale information of features is generated by a multi-scale information generator, which consists of three branches;
[0062] The first branch High-level is mainly composed of downsampling modules, which use multi-branch convolution to reduce information loss and learn high-resolution representation of features. Here is the output of the first-level convolution (1×1conv) and the second-level convolution (3×3conv) in series, the output of the first-level convolution (1×1conv) and the second-level convolution (3×3conv) and the third-level convolution (3×3conv) in series, and then the output is connected in series through the fourth-level convolution (depthwise separable convolution) and the fifth-level convolution (1×1conv).
[0063]
[0064] The second branch, Low-level, learns the low-resolution representation of features through global maximum pooling and deconvolution. Specifically, the first-level convolution (1×1conv) is input into the global maximum pooling branch and the deconvolution branch respectively, and then the results are fused. The global maximum pooling branch includes the sequentially connected global maximum pooling layer (GMP), deconvolution layer (DeConv) and convolution layer (3×3conv), and the deconvolution branch includes the sequentially connected deconvolution layer (DeConv) and convolution layer (3×3conv); the fusion result is output to the deconvolution layer (DeConv), convolution layer (1×1conv) and then output
[0065]
[0066] The third branch Mid-level aggregates the feature representations from the other two branches. In order to improve the adaptability of regional features to the scale changes of different objects, the output of the first-level convolution is multiplied after passing through the global average pooling branch and the depth-separable convolution layer. The global average pooling branch includes the global average pooling layer (GAP), convolution layer (1×1conv) and sigmoid activation function set in sequence. The multiplication result is fused with the output of the four-level convolution (depth-separable convolution) in the first branch and the output of the fused deconvolution layer (DeConv) in the second branch, and then output through the convolution layer (1×1conv).
[0067]
[0068] Among them, DW() stands for depthwise separable convolution, De() stands for deconvolution, GAP() stands for global average pooling, GMP() stands for global maximum pooling, and S() stands for Sigmoid function.
[0069] In the second part, the similarity scores between features of different scales are calculated based on the hybrid similarity metric, the internal weights of the supporting features are dynamically adjusted, and the features of three different scales are spliced to obtain dynamically optimized multi-scale supporting features; that is, based on these multi-scale features, DMSG uses the multi-similarity feature guidance module to generate global similarity, regional similarity and local similarity (classified by global features, regional features and local features of different scales) to guide the supporting features more finely and comprehensively;
[0070] Specifically, after obtaining multi-scale feature information, the features of each scale are divided into supporting features and query features. Taking the high-resolution features output by the High-level branch as an example, the multi-similarity feature guidance module uses different similarity layers to calculate the similarity scores between them. The calculation formulas of the multi-similarity layers are as follows:
[0071]
[0072] Among them, D c is the cosine similarity, D e is the Euclidean distance, D m is the Manhattan distance, i represents the i-th image, and n is the total number of images.
[0073] Then, each similarity measure is weighted and fused to finally obtain the global similarity score as the similarity evaluation index at this scale;
[0074] Based on this, the internal weights of the supporting features are dynamically adjusted. Finally, the features of three different scales are concatenated together to obtain dynamically optimized multi-scale supporting features.
[0075] (3) Train the model with base class data in the sample dataset;
[0076] During the training process, SGD was used as the optimizer, the momentum parameter was set to 0.9, the weight decay factor was 1e-4, the batch size was set to 10, and the learning rate was set to 0.006 during the training phase to ensure the stability and efficiency of the optimization process.
[0077] The model is trained with base class data in the sample dataset, and the loss function is
[0078] L=L rpn +L LFR +L meta +L cls +L reg
[0079] Among them, L rpn is the loss of RPN, L LFR is the reconstruction loss of the latent feature reconstruction module, L meta is the meta-classification loss of the model, L cls is the classification loss, L reg is the regression loss;
[0080] The base class data is input into the model, and the model parameters are continuously optimized through the back propagation algorithm to finally obtain the base class model.
[0081] (4) Freeze the local parameters of the model and adjust the model with the new class data in the sample data set;
[0082] After training the model with the base class data in the sample dataset, the local parameters of the model are frozen as the parameters of the backbone network, RPN module and regional interest alignment module.
[0083] In the present invention, most of the parameters of the base class model, including the parameters of the backbone network, RPN module and region of interest alignment module, are frozen to retain the general feature representation ability of the base class model. Subsequently, the model is fine-tuned using the new class data. Only some trainable parameters are adjusted in the fine-tuning stage, and the learning rate is set to 0.001.
[0084] It should be noted that the new class data only contains a small number of annotated instances per category, so the fine-tuning process aims to efficiently adapt the new class features while maintaining good detection performance on the base classes.
[0085] (5) Input the data to be tested into the adjusted model to obtain the small sample target detection results.
[0086] By comprehensively evaluating the trained model on the test set, analyzing its detection accuracy on different data categories, and deeply verifying the generalization ability and practical application effect of the model, we can finally obtain an optimized model that can adapt to new types of data sets and provide reliable support for small sample target detection tasks.
[0087] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0088] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0089] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1A function specified in one or more boxes.
[0090] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0091] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0092] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A small sample target detection method based on dynamic latent feature guidance, characterized by: The method obtains a sample data set and constructs a small sample target detection model guided by dynamic potential features; after training the model with base class data in the sample data set, freezes the local parameters of the model, and adjusts the model with new class data in the sample data set; Input the data to be tested into the adjusted model to obtain the small sample target detection results.
2. The small sample target detection method based on dynamic latent feature guidance according to claim 1 is characterized by: The small sample target detection model includes a support branch and a query branch; The query branch includes a backbone network, an RPN module, and a regional interest alignment module which are arranged in sequence, and the output of the regional interest alignment is aggregated and output with the output of the support branch; The support branch includes a backbone network and a fusion layer. The outputs of the query branch and the backbone network of the support branch are aligned and then input into the fusion layer. A potential feature reconstruction module and a dynamic multi-scale similarity guidance module are sequentially arranged after the fusion layer.
3. The small sample target detection method based on dynamic latent feature guidance according to claim 2 is characterized by: Query the trunk network parameters shared by the branch and the supporting branch.
4. The small sample target detection method based on dynamic latent feature guidance according to claim 2 is characterized by: The output of the query branch backbone network is dimensionally aligned with the output of the support branch backbone network through a global average pooling layer and an expansion module.
5. The small sample target detection method based on dynamic latent feature guidance according to claim 2 is characterized in that: The latent feature reconstruction module comprises an encoder and two parallel convolution blocks arranged in sequence, and the outputs of the two parallel convolution blocks are input into a decoder after re-parameterization.
6. The small sample target detection method based on dynamic latent feature guidance according to claim 2 is characterized by: The dynamic multi-scale similarity guidance module includes a multi-scale information generator for generating features of different scales. The features of different scales are input into the multi-similarity feature guidance module, and are spliced and output after being processed by multiple similarity layers.
7. The small sample target detection method based on dynamic latent feature guidance according to claim 6 is characterized by: The first branch reduces information loss through multi-branch convolution operations and learns high-resolution representation of features; The second branch focuses on the low-resolution representation of features through global maximum pooling and deconvolution operations; the third branch combines global average pooling and channel attention mechanism to aggregate the feature representations of the first and second branches.
8. The small sample target detection method based on dynamic latent feature guidance according to claim 7 is characterized in that: Based on the hybrid similarity metric, the similarity scores between features of different scales are calculated, the internal weights of the supporting features are dynamically adjusted, and the features of three different scales are concatenated to obtain dynamically optimized multi-scale supporting features.
9. The small sample target detection method based on dynamic latent feature guidance according to claim 2 is characterized by: The model is trained with the base class data in the sample data set, and the loss function is L = L rpn +L LFR +L meta +L cls +L reg Among them, L rpn is the loss of RPN, L LFR is the reconstruction loss of the latent feature reconstruction module, L meta is the meta-classification loss of the model, L cls is the classification loss, L reg is the regression loss.
10. The small sample target detection method based on dynamic latent feature guidance according to claim 2, characterized in that: After training the model with the base class data in the sample dataset, the local parameters of the model are frozen as the parameters of the backbone network, RPN module and regional interest alignment module.
Citation Information
Patent Citations
Cross-domain pedestrian re-identification method based on attention guidance and multi-scale label generation
CN114092964A
Small target detection method based on graph attention network
CN117274744A
Two-stage generalized small sample target detection method and system based on self-supervision
CN119131366A