A small sample target detection method and device based on multi-angle optimization
Patent Information
- Application Number
- CN202510658669.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-05-21
AI Technical Summary
第三,改善模型混淆性:模型在学习小样本类别时,容易与预训练阶段的基础类别知识产生混淆,无法准确区分目标对象属于哪个类别,这种混淆性问题会严重影响检测结果的准确性,降低模型的实用性
[0039](1)本发明通过联合BLIP模型、扩散模型和CLIP模型进行图像数据增强,不需要如GPT4o一样的成本巨大的大语言模型,同时还能够实现将数据集中原有的基础类别图像以及小样本类别图像进行合理的利用,更加契合小样本目标检测任务对于小样本的要求。
Smart Images

Figure CN120472148B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of target detection technology, specifically relating to a small-sample target detection method and device based on multi-angle optimization. Background Technology
[0002] Object detection, a fundamental task in computer vision, aims to accurately identify various objects in an image and perform classification. Traditional object detection methods typically rely on large amounts of labeled data to train models. However, in practical applications, collecting large-scale, high-quality labeled data faces numerous challenges. On the one hand, the labeling process requires significant manpower and time, increasing project costs. On the other hand, data acquisition is difficult in certain specific scenarios, leading to a scarcity of labeled data. These issues limit the widespread adoption and application of traditional object detection methods in more scenarios.
[0003] To overcome the performance bottleneck of traditional object detection methods when labeled data is scarce, Few-Shot Object Detection (FSOD) has emerged. It aims to narrow the gap between machine vision models and the human visual system in handling small-sample problems, enabling models to accurately detect new object categories even with only a very small number of labeled samples. This not only reduces the dependence on large-scale labeled data but also makes it possible to apply object detection technology in scenarios where data acquisition is difficult.
[0004] Currently, few-shot object detection methods mainly improve model performance from four key aspects: First, image enhancement: In recent years, diffusion models (DM) have performed excellently in generating high-quality data. How to leverage diffusion models to achieve richer image enhancements and expand the diversity of training data has become a research hotspot in the field of few-shot object detection. Second, improving model forgetting: In few-shot object detection, the model needs to utilize the basic category knowledge learned in the pre-training stage to assist in learning few-shot categories. Therefore, maintaining the knowledge coherence of the model during the learning process is crucial. Third, improving model confusion: When learning few-shot categories, the model is prone to confusion with the basic category knowledge from the pre-training stage, making it unable to accurately distinguish which category the target object belongs to. This confusion problem will seriously affect the accuracy of the detection results and reduce the practicality of the model. Fourth, handling class imbalance: The number of basic category images in the pre-training stage is large, while the number of few-shot category images is small. This class imbalance phenomenon will cause the model to favor basic categories and ignore few-shot categories during training. Therefore, how to balance the difference in the number of images of the two classes is an important problem that few-shot object detection needs to solve.
[0005] Currently, common few-shot object detection methods employ a two-stage training strategy. Stage 1 involves pre-training on base category images with abundant labeled data, allowing the model to learn basic visual features and category knowledge. Stage 2 involves freezing the backbone network weights and then fine-tuning the model using few-shot category images, aiming to improve the model's detection capability for few-shot categories. However, existing methods still have some problems that urgently need to be addressed:
[0006] First, the data augmentation solutions are not ideal: data augmentation solutions based on diffusion models either have overly simplistic cue word designs that fail to fully leverage the advantages of the diffusion model, or they rely on powerful and costly large language models. Furthermore, existing methods are not closely connected to the original images in the dataset, but rather rely more on the capabilities of the diffusion model itself for augmentation, introducing additional semantics beyond the dataset images, which may contradict the original purpose of few-sample object detection.
[0007] Second, the problems of confusion and forgetting have not been effectively solved: Although there are various solutions, most of them start from adjusting the model or optimizing the loss function. The problems of confusion and forgetting are still prominent, which limits the improvement of model performance.
[0008] Third, the method of learning class centers directly through training is not friendly to small sample classes: the current method allows the classifier to learn the class centers of different class images by backpropagation of gradients. However, since the number of small sample class images is small, it is often difficult to learn relatively accurate class centers, which further exacerbates the confusion problem.
[0009] Therefore, it is necessary to comprehensively and deeply optimize and innovate existing small-sample target detection methods in order to promote the maturity of small-sample target detection technology and achieve wider application. Summary of the Invention
[0010] In view of the above, the purpose of this invention is to provide a small sample target detection method and apparatus based on multi-angle optimization. By improving data augmentation methods, designing feature perturbation methods, and classifier optimization methods, the data quality in small sample target detection is improved, and the problems of confusion, forgetting, and class imbalance are mitigated, thereby enhancing the performance of small sample target detection.
[0011] To achieve the above-mentioned objectives, the present invention provides the following technical solution:
[0012] In a first aspect, the present invention provides a small-sample target detection method based on multi-angle optimization, comprising the following steps:
[0013] The combined BLIP model, diffusion model, and CLIP model are used to perform data augmentation on few-sample class images to obtain few-sample class data-augmented images;
[0014] A few-shot target detection model is constructed, comprising a feature extraction network, a feature fusion network, a region candidate network, a region pooling network, a feature perturbation network, a classification network, and a regression network. The feature extraction network is used to extract preliminary image features, the feature fusion network is used to fuse the preliminary features to obtain fused features, the region candidate network is used to extract target region candidate boxes based on the fused features, the region pooling network is used to extract target features based on the fused features and target region candidate boxes and adjust the target feature scale to a specified size, the feature perturbation network is used to add perturbation based on feature category to the target features through gradient-aware feature perturbation to obtain the final features, the classification network is used to perform classification tasks based on the final features, and the regression network performs regression calculations based on the target region candidate boxes.
[0015] The classifier in the classification network is first initially trained based on the basic category features and few-sample category features obtained from the CLIP model; then the classifier parameters are frozen, and the few-sample object detection model is trained as a whole based on the few-sample category data augmented image and the basic category image to obtain the trained few-sample object detection model.
[0016] Small sample target detection is performed using a trained small sample target detection model.
[0017] Preferably, the method of combining the BLIP model, diffusion model, and CLIP model to perform data augmentation on few-shot class images to obtain few-shot class data-augmented images includes:
[0018] Extract cue words T using the BLIP model to describe the content of images in the basic category. base The prompt word T base Replace the category keywords in the text with the category names of the selected small sample category images to obtain the new cue word T. novel Generate a mask image P for the target location based on the image annotations of the base category images. mask Based on the image annotations for the minority class, a sub-image corresponding to the target location is cropped from the minority class image and resized to cover the target location in the base class image to obtain a new image P. novel ;
[0019] Based on the new image P using a diffusion model novel Mask image P mask And the new cue word T novel Image enhancement is performed to obtain several enhanced small sample images P. gen ;
[0020] Based on the enhanced few-sample image P using the CLIP model genVisual features are extracted and similarity is calculated for target sub-images extracted from image annotations and target sub-images extracted from small sample category images based on image annotations. The enhanced small sample image P corresponding to a similarity greater than a set threshold is then selected. gen Images that are qualified are added to the dataset for model training, while the remaining enhanced small sample images that are less than a set threshold are discarded as unqualified images.
[0021] Preferably, the feature perturbation network is used to add feature category-based perturbations to the target features through a gradient-aware feature perturbation method to obtain the final features, including:
[0022] In the feature perturbation network, the target feature F is extracted by fusing features and target region candidate boxes. target The perturbation magnitude F of each target feature is calculated based on its category. dis and with target feature F target Adding them together yields the final feature F. final The formula is as follows:
[0023]
[0024] F final =F dis +F target
[0025] Where α represents the hyperparameter, n represents the number of images corresponding to the current category, and k represents a positive integer. The classification loss l represents the small sample object detection model. cls For the gradient of each target feature, ||·||2 represents the L2 norm.
[0026] Preferably, the preliminary training of the classifier in the classification network based on the basic category features and few-sample category features obtained from the CLIP model includes:
[0027] The class names of the augmented images and the base class images from the small sample class data are input into the text encoder of the CLIP model to repeatedly extract several different class features as class centers. The classifier in the classification network is initially trained using these different class features. The cross-entropy loss function is used during the initial training so that the classifier can learn in advance the class centers that distinguish different classes.
[0028] Preferably, the original fully connected layer of the classifier is optimized into two fully connected layers. The first fully connected layer is used to reduce the dimensionality of the features output by the text encoder of the CLIP model, and the second fully connected layer is used to perform classification based on the dimensionality-reduced features.
[0029] Preferably, only the second fully connected layer is trained during the initial training of the classifier, and only the parameters of the second fully connected layer are frozen during the overall training of the small sample target detection model.
[0030] Secondly, in order to achieve the above-mentioned objectives, the present invention also provides a small sample target detection device based on multi-angle optimization, which is implemented using the above-mentioned small sample target detection method based on multi-angle optimization, including: a data augmentation module, a model building module, a model training module, and a target detection module;
[0031] The data augmentation module is used to combine the BLIP model, diffusion model and CLIP model to perform data augmentation on small sample class images to obtain small sample class data-augmented images;
[0032] The model building module is used to construct a small sample target detection model that includes a feature extraction network, a feature fusion network, a region candidate network, a region pooling network, a feature perturbation network, a classification network, and a regression network. The feature extraction network is used to extract preliminary image features, the feature fusion network is used to fuse the preliminary features to obtain fused features, the region candidate network is used to extract target region candidate boxes based on the fused features, the region pooling network is used to extract target features based on the fused features and target region candidate boxes and adjust the target feature scale to a specified size, the feature perturbation network is used to add perturbation based on feature category to the target features through gradient-aware feature perturbation to obtain the final features, the classification network is used to perform classification tasks based on the final features, and the regression network performs regression calculations based on the target region candidate boxes.
[0033] The model training module is used to first train the classifier in the classification network based on the basic category features and few-sample category features obtained from the CLIP model; then freeze the classifier parameters and train the few-sample object detection model as a whole based on the few-sample category data augmented image and the basic category image to obtain the trained few-sample object detection model.
[0034] The target detection module is used to perform small-sample target detection using a trained small-sample target detection model.
[0035] Thirdly, to achieve the above-mentioned objectives, embodiments of the present invention also provide an electronic device, including a memory and one or more processors, wherein the memory is used to store a computer program, and the processor is used to implement the above-mentioned small sample target detection method based on multi-angle optimization when executing the computer program.
[0036] Fourthly, to achieve the above-mentioned objectives, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a computer, implements the above-mentioned small sample target detection method based on multi-angle optimization.
[0037] Fifthly, to achieve the above-mentioned objectives, the present invention also provides a computer product comprising a computer program that, when executed by a processor, implements the above-mentioned small sample target detection method based on multi-angle optimization.
[0038] Compared with the prior art, the beneficial effects of the present invention include at least the following:
[0039] (1) This invention enhances image data by combining the BLIP model, diffusion model and CLIP model. It does not require a large language model with huge costs like GPT4o. At the same time, it can make reasonable use of the original basic category images and small sample category images in the dataset, which is more in line with the small sample target detection task's requirements for small samples.
[0040] (2) Unlike existing technologies that directly use adjustment of loss function or network to improve model confusion and forgetting, this invention adopts gradient-aware feature perturbation for optimization. Through feature perturbation, the model has stronger generalization ability, further improving confusion and forgetting. At the same time, it enables image features of small sample classes to be closer to the class center, alleviating the problem of excessive class center shift caused by class imbalance and too few small sample images.
[0041] (3) Unlike the existing technology that uses gradient backpropagation to train the model to learn the class center, which often fails to learn a relatively accurate small sample class center, this invention cleverly utilizes the sufficient semantics of the CLIP model during training, and uses the text encoder of the CLIP model to calculate the class center of different categories. The class center is used to guide the optimization of the classifier, enabling the classifier to learn to effectively distinguish different class centers, thereby guiding the features of different categories of images to be closer to the class center, further improving the class imbalance problem and improving the accuracy of target detection. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1This is a flowchart illustrating the small-sample target detection method based on multi-angle optimization provided in an embodiment of the present invention.
[0044] Figure 2 This is a schematic diagram of the data augmentation process based on the diffusion model provided in an embodiment of the present invention;
[0045] Figure 3 This is a schematic diagram of the structure of the small sample target detection model provided in an embodiment of the present invention;
[0046] Figure 4 This is a schematic diagram of the structure of a small sample target detection device based on multi-angle optimization provided in an embodiment of the present invention. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of this invention.
[0048] The inventive concept of this invention is as follows: Addressing the shortcomings of existing few-shot object detection methods, such as unsatisfactory data augmentation schemes, ineffective solutions to confusion and forgetting issues, and the unfriendly nature of category center learning methods for few-shot categories, this invention provides a few-shot object detection method and apparatus based on multi-angle optimization. It significantly improves data diversity and quality through data augmentation methods based on BLIP, diffusion, and CLIP models, while maintaining low computational costs. By constructing a gradient-aware feature perturbation strategy, it successfully enhances the model's robustness against confusion and forgetting issues. Furthermore, by employing a category center separation method based on a CLIP model text encoder, it more effectively improves the classifier's discriminative ability in few-shot scenarios compared to traditional gradient backpropagation training. These multi-angle optimization innovations collectively improve the performance of few-shot object detection tasks, providing new technical approaches for the field.
[0049] like Figure 1 As shown in the embodiment, a small-sample target detection method based on multi-angle optimization is provided, including the following steps:
[0050] S1, combining the BLIP model, diffusion model and CLIP model to perform data augmentation on small sample class images to obtain small sample class data augmented images.
[0051] Existing data augmentation based on diffusion models suffers from two significant problems. First, in terms of cue word design, some existing works use very simple cue words, such as "A photo of [class name]", where "class name" refers to the category name, such as cat or dog. Others use large language models, such as GPT4o. For the former cue word method, diffusion models often fail to generate the required image, while for the latter, using large language models incurs significant costs. This research aims to find a balance between ensuring image quality and cost savings. Therefore, the embodiment designs a joint BLIP model (specifically using the BLIP-2 model or other versions), a diffusion model (specifically using an image-inpainting version of the diffusion model or other versions), and a CLIP model to accomplish this task. The total size of the three models is much smaller than that of a large language model, effectively saving computational overhead and cost.
[0052] like Figure 2 As shown, firstly, two images are selected from the dataset: one representing the basic category and the other representing a subset category. Considering the differences arising from different categories, the selected images should be of similar categories; for example, cats and dogs are more similar than cats and cars. After obtaining the two images, the BLIP-2 model is used to extract cue words T to describe the content of the basic category image. base The prompt word T base Replace the category keywords in the text with the category names of the selected small sample category images to obtain the new cue word T. novel Next, using the image annotations in the dataset, the location of the target in the base category images is found, and a mask image P is generated based on the annotations. mask Then, using the same method, the target location in the small sample class images is found, and the sub-image corresponding to the target location is cropped and overlaid on the target location in the base class image to obtain a new image P. novel Since this is an object detection task, the traditional CutMix method cannot be used arbitrarily for image overlay; otherwise, it will lead to changes in the position of the object detection box, resulting in errors or requiring additional manual labor for annotation.
[0053] Next, the new image P is obtained. novel Mask image P mask And the new cue word T novelThe images are input together into the image inpainting version of the diffusion model. The diffusion model then performs image inpainting on a masked area of the new image based on prompts, resolving the inconsistency between foreground and background caused by image overlay. Leveraging the powerful capabilities of the diffusion model, it enhances the image and obtains more enhanced small sample images P. gen This method not only makes good use of the rich background of the basic category images, but also uses small sample category images to generate enhanced images, without allowing the diffusion model to generate images arbitrarily and deviating from the small sample setting.
[0054] Finally, the generated enhanced few-sample image P gen Both the enhanced image and the few-sample category images extract their corresponding target sub-images based on the labeled locations in the data. These sub-images are then input into the visual part of the CLIP model for feature extraction. The resulting enhanced image features and few-sample category features are compared using a cosine similarity calculation. A threshold is set, and images with a similarity score greater than or equal to the threshold are considered acceptable and added to the training dataset. The reason for using the CLIP model for filtering is that during image inpainting, the diffusion model may fail to restore the content of few samples due to issues such as image sharpness. Therefore, samples with similarity scores below the threshold need to be removed; otherwise, the training accuracy will be affected.
[0055] By combining the image inpainting versions of the BLIP-2 model and the diffusion model with the CLIP model, a set of small sample category images based on the original dataset can be generated. The generated small sample category images not only have a certain degree of generalization, but also do not deviate from the basic setting of the small sample object detection task while being generated from the original dataset. Furthermore, the CLIP model is used to perform automated image screening to remove images with insufficient quality, saving manpower and making the method more transferable.
[0056] S2 constructs a small sample target detection model that includes a feature extraction network, a feature fusion network, a region candidate network, a region pooling network, a feature perturbation network, a classification network, and a regression network. The feature perturbation network adds perturbations based on feature categories to the target features through a gradient-aware feature perturbation method.
[0057] like Figure 3As shown, the small sample target detection model includes a feature extraction network, a feature fusion network, a region candidate network, a region pooling network, a feature perturbation network, a classification network, and a regression network. The feature extraction network (specifically using ResNet-101 or another version as the backbone network) is used to extract preliminary image features. The feature fusion network (specifically using an FPN network (Feature Pyramid Network)) is used to fuse the preliminary features to obtain fused features. The region candidate network (specifically using an RPN network (Region Proposal Network)) is used to extract target region candidate boxes based on the fused features. The region pooling network (specifically using an ROI Pooling network (Region of Interest)... Pooling (Region of Interest Pooling) is used to extract target features based on fused features and target region candidate boxes and adjust the scale of the target features to a specified size. The feature perturbation network is used to add perturbation based on feature category to the target features through gradient-aware feature perturbation to obtain the final features. The classification and regression network (including classifier and bounding box regressor) is used to perform target detection based on the final features. Specifically, the classification network performs classification tasks based on the final features, and the regression network performs regression calculations based on the target region candidate boxes.
[0058] Confusion in few-shot object detection refers to the tendency for images of a few categories to be easily classified as similar basic categories during inference due to insufficient training (e.g., classifying a cat image as a dog). This confusion arises from insufficient separation between class centers of different categories. Forgetting refers to the tendency for few-shot object detection to forget some learned knowledge during training. This forgetting is caused by insufficient intra-class compactness and loose feature distribution around class centers. For few-shot object detection tasks, the lack of few-shot category images can lead to biases in class center judgments during training. Therefore, to guide the model to learn more accurate class centers for few-shot categories and to improve the model's generalization ability, this embodiment designs a gradient-aware feature perturbation method based on target features extracted through fused features and target region candidate boxes. After obtaining the target features, the perturbation network calculates the perturbation magnitude F of each target feature according to its category. dis and with target feature F target Adding them together yields the final feature F. final The formula is as follows:
[0059]
[0060] F final =F dis +F target
[0061] Where α represents the hyperparameter, n represents the number of images corresponding to the current category, and k represents a positive integer, which in the example takes the value of 2 or 4, etc., to standardize the range of feature perturbation and prevent it from being too large or too small. The classification loss l represents the small sample object detection model. cls For the gradient of each target feature; ||·||2 represents the L2 norm, which is used to normalize the gradient and prevent the gradient from being too large and having a negative impact on the feature shift, causing the model training process to fail to converge.
[0062] Since an image may contain different categories of target objects, and the feature perturbations for different target objects are different, directly applying feature perturbations to the image would cause interference between different categories of perturbations. Therefore, it is necessary to first use a region candidate network to provide region predictions for different categories, and then obtain sub-images based on the region predictions before applying feature perturbations. Finally, the final feature F... final The data is input into the corresponding classifier, and the predicted bounding boxes obtained from the region candidate network are input into the bounding box regressor for subsequent classification and bounding box regression calculations.
[0063] As can be seen from the above formula, when the total number of images in the target image's category is large, the magnitude of the feature perturbation is small; when the total number of images in the target image's category is small, the magnitude of the feature perturbation is large. This is because in categories with a sufficient number of images, the images contain ample semantics, and good results can be achieved in object detection even without feature perturbation. Perturbation for these categories is only to further enhance robustness and generalization. However, for small sample categories, although there are already many images generated by the diffusion model, due to the limitations of the diffusion model itself, the obtained images still cannot extract the rich semantics of the base category images. Therefore, applying a larger perturbation to the images of these categories will greatly promote the learning of small sample categories.
[0064] By perturbing the features in the direction that reduces the loss gradient, the features can be more clustered around the class center, while strengthening the intra-class compactness of each category's features and further expanding the inter-class separation, thereby improving the confusion and forgetting problems that exist in small sample object detection tasks.
[0065] S3. Based on the basic category features and few-sample category features obtained from the CLIP model, the classifier in the classification network is initially trained; then the classifier parameters are frozen, and the few-sample object detection model is trained as a whole based on the few-sample category data augmented image and the basic category image to obtain the trained few-sample object detection model.
[0066] In current mainstream research, classifiers for few-sample object detection are trained using gradient backpropagation. This method yields relatively good results when there are sufficient training images. However, for few-sample class images, since these images do not represent the actual distribution of all images and may have some deviation from the true class centers, direct training often fails to obtain accurate few-sample class centers.
[0067] Previous work has proposed using the Simplex Equiangular Tight Frame (Simplplex ETF) to address this challenge. Specifically, this involves forcibly limiting the separation between category centers of different categories to the same degree. However, this approach still has a significant drawback: the correlation between different categories is not uniform, making the requirement for identical separation of category centers unreasonable. For example, the similarity between images of cats and dogs is certainly greater than the similarity between images of cats and cars. Therefore, while this method can alleviate the confusion problem of small sample category centers to some extent, it also introduces new issues.
[0068] Therefore, to find a more accurate separation degree between different category centers, this invention utilizes the rich semantics learned during the CLIP model's pre-training phase. During pre-training, the CLIP model trains itself by mapping a large number of text pairs to images, thus leveraging the rich semantics learned in the past to help distinguish the separation degree between different category centers. In this embodiment, firstly, the category names of the augmented images and the base category images from the small sample category data are input into the CLIP model's text encoder. The CLIP model then calculates the features for different categories; these features are essentially the category centers. The CLIP model's text encoder generates these features 30 times. Subsequently, this data is used to initially train the classifier, enabling it to learn to distinguish the category centers of different categories. The classifier employs traditional cross-entropy loss. Then, the classifier's weight parameters are frozen, and the entire model is trained again. At this point, the classifier has learned how to distinguish the category centers of different categories and will guide the features of different category images towards the category centers.
[0069] In practice, since the features extracted from the image are 1024-dimensional, while the features calculated by the CLIP model's text encoder are 512-dimensional, the original classifier cannot be directly trained using the text encoder's features. Therefore, the classifier is optimized from a single fully connected layer to two identical fully connected layers. The first fully connected layer reduces the 1024-dimensional features to 512-dimensional features, and the second fully connected layer classifies the 512-dimensional features. Thus, the proposed method only requires training the second fully connected layer. Furthermore, after training, the weights of the second fully connected layer need to be frozen to prevent the learned class center separation knowledge from being forgotten during subsequent training.
[0070] Finally, after initial training and overall training, a small-sample target detection model was obtained.
[0071] S4, using the trained few-sample target detection model to perform few-sample target detection.
[0072] After completing the model training process and ensuring that the few-sample object detection model reaches the ideal performance state, the trained few-sample object detection model is formally used to carry out comprehensive and accurate object detection work on the given few-sample dataset.
[0073] In summary, the few-shot object detection method based on multi-angle optimization provided by this invention first utilizes advanced data augmentation methods, combining the BLIP-2 model, an image restoration version of the diffusion model, and the CLIP model to perform multiple diverse data augmentation operations on images of few-shot categories. This results in a richer and more varied collection of images of few-shot categories, which are then effectively added to the training dataset. This approach not only expands the data scale but also cleverly solves the common class imbalance problem in few-shot object detection tasks, providing a more balanced data foundation for model training. Subsequently, this method innovatively designs a feature perturbation method. By performing detailed and targeted perturbation processing on images of different categories to varying degrees, the model can obtain more robust feature representations when extracting features. This also promotes a more compact clustering of features in the feature space, significantly amplifying the feature differences between images of different categories, thereby effectively improving the model's classification ability and discriminative power. Finally, through deep optimization and careful training of the classifier, the entire few-shot object detection model has achieved significant improvements in several key indicators such as accuracy, recall, and generalization ability, providing an efficient and practical solution for the field of few-shot object detection.
[0074] Based on the same inventive concept, such as Figure 4As shown, this embodiment of the invention also provides a small sample target detection device 400 based on multi-angle optimization, including: a data augmentation module 410, a model building module 420, a model training module 430, and a target detection module 440.
[0075] The data augmentation module 410 is used to combine the BLIP model, diffusion model and CLIP model to perform data augmentation on small sample class images to obtain small sample class data augmented images.
[0076] The model building module 420 is used to build a small sample target detection model that includes a feature extraction network, a feature fusion network, a region candidate network, a region pooling network, a feature perturbation network, a classification network, and a regression network. The feature extraction network is used to extract preliminary image features, the feature fusion network is used to fuse the preliminary features to obtain fused features, the region candidate network is used to extract target region candidate boxes based on the fused features, the region pooling network is used to extract target features based on the fused features and target region candidate boxes and adjust the target feature scale to a specified size, the feature perturbation network is used to add perturbation based on feature category to the target features through gradient-aware feature perturbation to obtain the final features, the classification network is used to perform classification tasks based on the final features, and the regression network performs regression calculations based on the target region candidate boxes.
[0077] The model training module 430 is used to first train the classifier in the classification network based on the basic category features and few-sample category features obtained from the CLIP model; then freeze the classifier parameters and train the few-sample object detection model as a whole based on the few-sample category data augmented image and the basic category image to obtain the trained few-sample object detection model.
[0078] The object detection module 440 is used to perform few-sample object detection using a trained few-sample object detection model.
[0079] Based on the same inventive concept, embodiments of the present invention also provide an electronic device, including a memory and one or more processors, wherein the memory is used to store a computer program, and the processor is used to implement the above-described small sample target detection method based on multi-angle optimization when executing the computer program.
[0080] Based on the same inventive concept, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a computer, implements the above-described small sample target detection method based on multi-angle optimization.
[0081] Based on the same inventive concept, this invention also provides a computer product comprising a computer program that, when executed by a processor, implements the above-described small sample target detection method based on multi-angle optimization.
[0082] It should be noted that the small sample target detection device, electronic device, computer-readable storage medium, and computer product based on multi-angle optimization provided in the above embodiments all belong to the same inventive concept as the small sample target detection method based on multi-angle optimization. For details of their specific implementation process, please refer to the embodiments of the small sample target detection method based on multi-angle optimization, which will not be repeated here.
[0083] The specific embodiments described above illustrate the technical solution and beneficial effects of the present invention in detail. It should be understood that the above description is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A small-sample target detection method based on multi-angle optimization, characterized in that, Includes the following steps: The combined BLIP model, diffusion model, and CLIP model are used to perform data augmentation on few-sample class images to obtain few-sample class data-augmented images; Construct a few-sample object detection model that includes a feature extraction network, a feature fusion network, a region candidate network, a region pooling network, a feature perturbation network, a classification network, and a regression network; The feature extraction network is used to extract preliminary image features; features The fusion network is used to fuse preliminary features to obtain fused features; the region candidate network is used to extract target region candidate boxes based on the fused features; the region pooling network is used to extract target features based on the fused features and target region candidate boxes and adjust the scale of the target features to a specified size; the feature perturbation network is used to add feature category-based perturbations to the target features through a gradient-aware feature perturbation method to obtain the final features, including: target features extracted based on the fused features and target region candidate boxes. The perturbation magnitude of each target feature is calculated based on its category. and the size of the disturbance With target features Adding them together yields the final feature. The formula is as follows: , , in, Indicates hyperparameters, This indicates the number of images corresponding to the current category. Represents positive integers. This represents the classification loss of a small-sample object detection model. For the gradient of each target feature, The L2 norm is represented; the classification network is used for classification tasks based on the final features; the regression network performs regression calculations based on candidate boxes of the target region. The classifier in the classification network is first initially trained based on the basic category features and few-sample category features obtained from the CLIP model; then the classifier parameters are frozen, and the few-sample object detection model is trained as a whole based on the few-sample category data augmented image and the basic category image to obtain the trained few-sample object detection model. Small sample target detection is performed using a trained small sample target detection model.
2. The small sample target detection method based on multi-angle optimization according to claim 1, characterized in that, The combined BLIP model, diffusion model, and CLIP model are used to augment few-sample class images to obtain few-sample class data-augmented images, including: Extract cue words describing the content of base category images using the BLIP model. , prompt words Replace the category keywords in the text with the category names of the selected small sample category images to obtain new prompt words. Generate a mask image of the target location based on the image annotations of the base category images. Based on the image annotations for the minority class, a sub-image corresponding to the target location is cropped from the minority class image and resized to cover the target location in the base class image to obtain a new image. ; Based on the new image using a diffusion model Mask image and new prompt words Image enhancement is performed to obtain several enhanced small sample images. ; Based on the enhanced few-sample images using the CLIP model Visual features are extracted and similarity is calculated for target sub-images extracted from image annotations and target sub-images extracted from small sample category images based on image annotations. Enhanced small sample images with similarity greater than a set threshold are then selected. Images that are qualified are added to the dataset for model training, while the remaining enhanced small sample images that are less than a set threshold are discarded as unqualified images.
3. The small sample target detection method based on multi-angle optimization according to claim 1, characterized in that, The preliminary training of the classifier in the classification network based on the basic category features and few-sample category features obtained from the CLIP model includes: The class names of the augmented images and the base class images from the small sample class data are input into the text encoder of the CLIP model to repeatedly extract several different class features as class centers. The classifier in the classification network is initially trained using these different class features. The cross-entropy loss function is used during the initial training so that the classifier can learn in advance the class centers that distinguish different classes.
4. The small sample target detection method based on multi-angle optimization according to claim 1 or 3, characterized in that, The original fully connected layer of the classifier is optimized into two fully connected layers. The first fully connected layer is used to reduce the dimensionality of the image features extracted from the small sample object detection model, and the second fully connected layer is used to classify based on the dimensionality-reduced features.
5. The small sample target detection method based on multi-angle optimization according to claim 4, characterized in that, During the initial training of the classifier, only the second fully connected layer is trained, and during the overall training of the small sample object detection model, only the parameters of the second fully connected layer are frozen.
6. A small-sample target detection device based on multi-angle optimization, implemented using the small-sample target detection method based on multi-angle optimization as described in any one of claims 1 to 5, characterized in that, include: The module includes data augmentation, model building, model training, and object detection. The data augmentation module is used to combine the BLIP model, diffusion model and CLIP model to perform data augmentation on small sample class images to obtain small sample class data-augmented images; The model building module is used to construct a small sample target detection model that includes a feature extraction network, a feature fusion network, a region candidate network, a region pooling network, a feature perturbation network, a classification network, and a regression network. The feature extraction network is used to extract preliminary image features, the feature fusion network is used to fuse the preliminary features to obtain fused features, the region candidate network is used to extract target region candidate boxes based on the fused features, the region pooling network is used to extract target features based on the fused features and target region candidate boxes and adjust the target feature scale to a specified size, the feature perturbation network is used to add perturbation based on feature category to the target features through gradient-aware feature perturbation to obtain the final features, the classification network is used to perform classification tasks based on the final features, and the regression network performs regression calculations based on the target region candidate boxes. The model training module is used to first train the classifier in the classification network based on the basic category features and few-sample category features obtained from the CLIP model; then freeze the classifier parameters and train the few-sample object detection model as a whole based on the few-sample category data augmented image and the basic category image to obtain the trained few-sample object detection model. The target detection module is used to perform small-sample target detection using a trained small-sample target detection model.
7. An electronic device comprising a memory and one or more processors, the memory being used to store a computer program, characterized in that, The processor is used to implement the small sample target detection method based on multi-angle optimization as described in any one of claims 1 to 5 when executing a computer program.
8. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by a computer, it implements the small sample target detection method based on multi-angle optimization as described in any one of claims 1 to 5.
9. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the small sample target detection method based on multi-angle optimization as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Small sample target detection method based on data enhancement and distribution calibration
CN117475212A
Small sample target detection method and device based on mixed experts
CN119418041A