A Method, System, Device, and Storage Medium for Small-Sample Object Detection
The proposed method for small-sample target detection uses training triplets and a modified Faster-RCNN network to improve detection performance by enhancing feature extraction and classification, addressing the limitations of existing small-sample detection methods.
Patent Information
- Application Number
- CN202211155350.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-22
Smart Images

Figure CN115546470B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method, a system, a device and a storage medium for small-sample object detection, belonging to the technical field of object detection. Background Art
[0002] As is well known, humans can generalize from only one instance of an animal to other instances of that animal. Most existing deep learning methods are still data-driven, that is, they require training with thousands of class instances so that the model can "recognize" new instances of the class. Therefore, few shot learning, which trains from only a few instances so that the model can recognize new instances, has become a current research hotspot.
[0003] It is important to learn a model with certain generalization ability by applying semi-supervised methods with less labeled data or using weakly supervised methods with incompletely matched labeled data, which is the problem to be solved in few shot learning.
[0004] The invention patent with the patent number "CN113239980A" discloses an underwater object detection method based on few-shot machine learning and hyperparameter optimization, a few-shot object detection model based on Cascade-RCNN; a pre-training data set; using the pre-training data set to pre-train the few-shot object detection model to obtain the pre-training weights of the few-shot object detection model; constructing an object data set to be detected; dividing the object data set into a labeled support set and an unlabeled query set; preprocessing the object data set; fine-tuning the few-shot object detection model to obtain a finally trained few-shot object detection model; using a Bayesian optimization model based on TPE to optimize the hyperparameters of the trained few-shot object detection model to obtain an optimized few-shot object detection model; inputting the preprocessed query set into the optimized object detection model to obtain the object detection result.
[0005] The above-mentioned existing technology only realizes object detection in the case of a small number of samples, but the method adopted is still the existing object detection scheme and is not optimized for the characteristics of small samples. Summary of the Invention
[0006] In order to solve the problems existing in the above-mentioned existing technology, the present invention proposes a method, a system, a device and a storage medium for small-sample object detection.
[0007] The technical solution of the present invention is as follows:
[0008] On the one hand, the present invention proposes a method for small-sample object detection, including the following steps:
[0009] Establish an object detection model;
[0010] The target detection model is initially trained using a public dataset, and after the initial training is completed, a small sample dataset containing images of the target object is added to fine-tune the target detection model; wherein, there is no overlapping part between the small sample training set and the public dataset.
[0011] During fine-tuning training, several support images and query images are sampled from the overall dataset after adding the small sample dataset to generate a support sample set and a query sample set.
[0012] Randomly select a query image Qc containing target objects of category c in the query sample set, and a support image Sc containing target objects of category c, and another support image Sn containing target objects different from category c to form a training triplet. In the training triplet, mark the target objects of category c in the query image Qc as the foreground, and all other target objects as the background.
[0013] Input the training triplet into the target detection model for feature extraction, obtain the target bounding boxes of each image in the training triplet, calculate the bounding box loss of the target bounding boxes, and determine which support image the input query image Qc matches through the target detection model. Calculate the matching loss according to the matching result output by the target detection model, and iteratively train the target detection model by combining the bounding box loss and the matching loss to obtain a trained small sample target detection model.
[0014] Use the trained small sample target detection model for target detection.
[0015] As a preferred embodiment, the target detection model is established based on the Faster-RCNN network, including a backbone network, an RPN network, and an Roi pooling operation module, and a foreground enhancement network module and a spatial attention module are added between the backbone network and the RPN network; and a spatial attention module is added at the backend of the Roi pooling operation module.
[0016] The foreground enhancement network module is used to extract foreground attention features from the input image; the spatial attention module is used to extract spatial attention from the input image.
[0017] As a preferred embodiment, the step of inputting the training triplet into the target detection model for feature extraction is specifically as follows:
[0018] Use the ResNet50 pre-trained model as the backbone network, input the query image Qc, the support image Sc, and the support image Sn, extract the basic features of the input images, and obtain the support image feature map and the query image feature map.
[0019] Input the support image feature map into the foreground enhancement network module to extract foreground attention features and obtain the first feature map. Then, perform broadcast pixel addition on the first feature map and the input support image feature map to obtain the second feature map.
[0020] Input the query image feature map into the spatial attention module to extract spatial attention features and obtain a third feature map that contains negative numbers, zeros, or positive numbers. Input the third feature map into the sigmoid function for activation operation to obtain the fourth feature map, and perform weighted fusion on the fourth feature map and the input query image feature map to obtain the fifth feature map.
[0021] Input the second feature map and the fifth feature map into the RPN network to generate target candidate boxes.
[0022] Input the target candidate boxes into the Roi pooling operation module to obtain the Roi feature map.
[0023] Input the Roi feature map into the spatial attention module to extract spatial attention features and obtain the final feature map. Perform classification and coordinate regression on the final feature map to obtain the prediction result, which includes the position information of the target box and the matching result.
[0024] As a preferred implementation, during the process of calculating the matching loss based on the matching result output by the target detection model and iteratively training the target detection model by combining the bounding box loss and the matching loss:
[0025] Use the support image Sn as the negative sample and the support image Sc as the positive sample for four pairings, namely the first positive support pair formed by the positive sample and the foreground target box, the second positive support pair formed by the positive sample and the background target box, and the negative support pair formed by the negative sample and the foreground and background target boxes.
[0026] During the training process, select N pairs of the first positive support pairs, 2N pairs of the second positive support pairs, and N pairs of negative support pairs according to the matching result output by the target detection model, and calculate the matching loss of the selected pairs.
[0027] On the other hand, the present invention also proposes a small-sample object detection system, including:
[0028] A model establishment module for establishing a target detection model.
[0029] A preliminary training module for preliminarily training the target detection model using a public dataset and fine-tuning the target detection model using a small-sample dataset containing target object images after the preliminary training is completed; wherein, there is no overlapping part between the small-sample training set and the public dataset.
[0030] A sample generation module, which is used to sample a number of support images and query images from the overall dataset after adding a small sample dataset during fine-tuning training, and generate a support sample set and a query sample set;
[0031] A triplet generation module, which is used to randomly select a query image Qc containing a target object of category c in the query sample set, and a support image Sc containing a target object of category c, and another support image Sn containing a target object different from category c to form a training triplet. In the training triplet, the target object of category c in the query image Qc is marked as the foreground, and all other target objects are marked as the background;
[0032] A fine-tuning training module, which is used to input the training triplet into the object detection model for feature extraction, obtain the target bounding boxes of each image in the training triplet, calculate the bounding box loss of the target bounding boxes, and determine which support image the input query image Qc matches through the object detection model. Calculate the matching loss according to the matching result output by the object detection model, and iteratively train the object detection model by combining the bounding box loss and the matching loss to obtain a trained small sample object detection model;
[0033] A detection module, which is used to perform object detection using the trained small sample object detection model.
[0034] As a preferred implementation, the object detection model is established based on the Faster-RCNN network, including a backbone network, an RPN network, and an Roi pooling operation module, and a foreground enhancement network module and a spatial attention module are added between the backbone network and the RPN network; and a spatial attention module is added at the backend of the Roi pooling operation module;
[0035] The foreground enhancement network module is used to extract foreground attention features from the input image; the spatial attention module is used to extract spatial attention from the input image.
[0036] As a preferred implementation, the fine-tuning training module includes:
[0037] A basic feature extraction unit, which uses the ResNet50 pre-trained model as the backbone network, inputs the query image Qc, the support image Sc, and the support image Sn, extracts the basic features of the input images, and obtains the support image feature map and the query image feature map;
[0038] A foreground enhancement unit, which is used to input the support image feature map into the foreground enhancement network module, extract foreground attention features, obtain a first feature map, and then perform broadcast pixel addition on the first feature map and the input support image feature map to obtain a second feature map;
[0039] The first spatial attention extraction unit is used to input the query image feature map into the spatial attention module, extract spatial attention features, and obtain a third feature map with negative numbers, zeros, or positive numbers; input the third feature map into the sigmoid function for activation operation to obtain a fourth feature map, and perform weighted fusion of the fourth feature map with the input query image feature map to obtain a fifth feature map;
[0040] The RPN network unit is used to input the second feature map and the fifth feature map into the RPN network to generate target candidate boxes;
[0041] The pooling unit is used to input the target candidate boxes into the Roi pooling operation module to obtain the Roi feature map;
[0042] The second spatial attention extraction unit is used to input the Roi feature map into the spatial attention module, extract spatial attention features, obtain the final feature map, perform classification and coordinate regression on the final feature map to obtain the prediction result, and the prediction result includes the position information of the target box and the matching result.
[0043] As a preferred embodiment, the fine-tuning training module includes:
[0044] The support pair generation unit is used to use the support image Sn as the negative sample and the support image Sc as the positive sample to perform four pairings, namely the first positive support pair formed by the positive sample and the foreground target box, the second positive support pair formed by the positive sample and the background target box, and the negative support pair formed by the negative sample and the foreground and background target boxes;
[0045] The selection unit is used to select N pairs of the first positive support pairs, 2N pairs of the second positive support pairs, and N pairs of the negative support pairs according to the matching result output by the target detection model during the training process, and calculate the matching loss of the selected pairs.
[0046] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements a small sample target detection method as described in any embodiment of the present invention.
[0047] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and characterized in that when the program is executed by a processor, it implements a small sample target detection method as described in any embodiment of the present invention.
[0048] The present invention has the following beneficial effects:
[0049] 1. A method for small-sample object detection according to the present invention constructs training triplets to train an object detection model. During the training process, the support sample set not only contains positive samples existing in the query sample set, but also contains negative samples not existing in the query sample set, enabling the network to determine whether the objects in the query set match either of them, thereby enhancing the discrimination ability of the network.
[0050] 2. A method for small-sample object detection according to the present invention enables the trained object detection model to match each object box generated from the query image with the target object in the support image. Therefore, the model not only learns to match objects of the same category between the query image and the support image, but also learns to distinguish objects of different categories between the query image and the support image, thus having better detection performance.
[0051] 3. A method for small-sample object detection according to the present invention performs four pairings of the support image and the object box during the training process, balancing the pairing ratio between the object box and the support image, and solving the problem that the background usually dominates in the training. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic flowchart of the method according to the embodiment of the present invention;
[0053] Figure 2 It is a schematic diagram of the network structure of the small-sample object detection model according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0055] It should be understood that the step numbers used herein are only for convenient description and do not limit the execution order of the steps.
[0056] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0057] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0058] The term "and / or" refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0059] Example 1:
[0060] See Figure 1 , a method for small-sample object detection, comprising the following steps:
[0061] S100. Establish an object detection model based on a convolutional neural network;
[0062] S200. Initially train the object detection model using a public dataset. In this example, the public dataset uses the COCO2014 dataset. First, unify all the data in COCO2014 train / val as the training set, and use the data in COCO2014 test as the test set for experiments to initially train the object detection model; after the initial training is completed, add a small-sample dataset containing images of the target object to fine-tune the object detection model. In this example, the small-sample dataset uses the sea drift garbage dataset. Select N types of sea drift garbage in the sea drift garbage dataset, such as bottles, construction waste, wood, wheels, derelict ships. In this example, a total of 32 types of sea drift garbage are selected, and K images are selected for each type of sea drift garbage; wherein, there is no overlapping part between the small-sample training set and the public dataset;
[0063] S300. When performing fine-tuning training, balance the data of the public dataset and the small-sample dataset. Sample a number of support images S and query images Q from the public dataset containing basic classes and the sea drift dataset containing new sea drift garbage classes to generate a support sample set and a query sample set;
[0064] S400. In this embodiment, a two-way contrastive training strategy is adopted. First, a query image Qc containing target objects of category c in the query sample set and a support image Sc containing target objects of category c are randomly selected. Additionally, a support image Sn containing target objects different from those of category c is taken, where Sn contains target objects of category n, and c and n are not the same category, forming a training triplet <Qc, Sc, Sn>. In the training triplet, the target objects of category c in the query image Qc are marked as the foreground, and all other target objects are marked as the background. During the training process, the support sample set not only contains positive samples existing in the query sample set but also contains negative samples that do not exist in the query sample set, enabling the network to determine whether the objects in the query set match either of them to enhance the network's discrimination ability.
[0065] S500. Input the training triplet into the target detection model for feature extraction, obtain the target bounding boxes of each image in the training triplet, calculate the bounding box loss of the target bounding boxes, and determine which support image the input query image Qc matches through the target detection model. Calculate the matching loss based on the matching result output by the target detection model, and iteratively train the target detection model by combining the bounding box loss and the matching loss to obtain a trained small-sample target detection model.
[0066] S600. Use the trained small-sample target detection model for target detection.
[0067] Based on this embodiment, the trained target detection model matches each target bounding box generated from the query image Qc with the target objects in the support image. Therefore, the model not only learns to match the same-category objects between (Qc, Sc) but also learns to distinguish different-category objects between (Qc, Sn), having better detection performance.
[0068] As a preferred implementation manner of this embodiment, specifically refer to Figure 2 , in step S100, the target detection model is established based on the Faster-RCNN network, including a backbone network, an RPN network, and an Roi pooling operation module, and a Foreground Augumention Block foreground enhancement network module and a Space Attention Block spatial attention module are added between the backbone network and the RPN network; and a SpaceAttention Block spatial attention module is added to the backend of the Roi pooling operation module.
[0069] The Foreground Augumention Block is used to extract foreground attention features from the input image; the Space Attention Block is used to extract spatial attention from the input image.
[0070] As a preferred implementation manner of this embodiment, in step S400, the step of inputting the training triplet into the target detection model for feature extraction is specifically:
[0071] S401. Use the ResNet50 pre-trained model as the backbone network, input the query image Qc, the support image Sc, and the support image Sn, extract the basic features of the input image, and obtain the support image feature map and the query image feature map, which are support features and query feature respectively;
[0072] S402. Input the support image feature map into the foreground enhancement network module to extract foreground attention features, obtain the first feature map, and then perform broadcast pixel addition on the first feature map and the input support image feature map to obtain the second feature map;
[0073] S403. Input the query image feature map into the spatial attention module to extract spatial attention features, and obtain a third feature map with negative numbers, zeros, or positive numbers; input the third feature map into the sigmoid function for activation operation to obtain a fourth feature map with a range of 0 to 1, and perform weighted fusion on the fourth feature map and the input query image feature map to obtain the fifth feature map;
[0074] S404. Input the second feature map and the fifth feature map into the RPN network to generate target candidate boxes;
[0075] S405. Input the target candidate boxes into the Roi pooling operation module to obtain the Roi feature map;
[0076] S406. Input the Roi feature map into the spatial attention module to extract spatial attention features, obtain the final feature map, perform classification and coordinate regression on the final feature map, and obtain the prediction result, where the prediction result includes the position information of the target box and the matching result.
[0077] During the fine-tuning training process, a small number of samples including 80 categories of the COCO2014 dataset and 32 categories of marine floating garbage are used for fine-tuning. The network structure remains unchanged, and only a few images (1, 2, 3, 5, 10) are selected for each category. Then, iterative training is performed on a total of 112 categories to obtain a trained small-sample object detection model, which can finally obtain the location information and category information of the query target object in the query image according to the support image support img.
[0078] As a preferred implementation manner of this embodiment, during the process of calculating the matching loss according to the matching result output by the object detection model and iteratively training the object detection model by combining the bounding box loss and the matching loss, a large number of background proposals usually dominate in the training, especially the negative support image negetive support. For this reason, in order to balance the ratio of these matching pairs between the query target box query proposals and the support image objects, the following specific measures are adopted in this embodiment:
[0079] Taking the support image Sn as the negative sample and the support image Sc as the positive sample, four pairings are performed, namely the first positive support pair (pf; sp) formed by the positive sample and the foreground target box, the second positive support pair (pb; sp) formed by the positive sample and the background target box, and the negative support pair (p; sn) formed by the negative sample and the foreground and background target boxes;
[0080] During the training process, N pairs of the first positive support pairs, 2N pairs of the second positive support pairs, and N pairs of negative support pairs are selected according to the matching result output by the object detection model, and the matching loss of the selected pairs is calculated. In this embodiment, the multi-task loss for each sampled proposal is L = Lmatching + Lbox, where the bounding box loss Lbox is the same as that of faster-rcnn, and the matching loss is the binary cross-entropy loss.
[0081] Embodiment 2:
[0082] This embodiment proposes a small-sample object detection system, including:
[0083] A model establishment module for establishing an object detection model; this module is used for the functions in step S100 of the first embodiment above and will not be elaborated here;
[0084] A preliminary training module for preliminarily training the object detection model using a public dataset, and adding a small-sample dataset containing target object images to fine-tune the object detection model after the preliminary training is completed; wherein, there is no overlapping part between the small-sample training set and the public dataset; this module is used for the functions in step S200 of the first embodiment above and will not be elaborated here;
[0085] A sample generation module, which is used to sample a number of support images and query images from the overall dataset after adding a small sample dataset during fine-tuning training to generate a support sample set and a query sample set; this module is used for the function of step S300 in the first embodiment above and will not be elaborated here;
[0086] A triplet generation module, which is used to randomly select a query image Qc containing target objects of category c from the query sample set, a support image Sc containing target objects of category c, and another support image Sn containing target objects different from category c to form a training triplet. In the training triplet, the target objects of category c in the query image Qc are marked as the foreground, and all other target objects are marked as the background; this module is used for the function of step S400 in the first embodiment above and will not be elaborated here;
[0087] A fine-tuning training module, which is used to input the training triplet into the object detection model for feature extraction, obtain the bounding boxes of each image in the training triplet, calculate the bounding box loss of the bounding boxes, and determine which support image the input query image Qc matches through the object detection model, calculate the matching loss according to the matching result output by the object detection model, and iteratively train the object detection model by combining the bounding box loss and the matching loss to obtain a trained small sample object detection model; this module is used for the function of step S500 in the first embodiment above and will not be elaborated here;
[0088] A detection module, which is used to perform object detection using the trained small sample object detection model; this module is used for the function of step S600 in the first embodiment above and will not be elaborated here.
[0089] As a preferred embodiment, the object detection model is established based on the Faster-RCNN network, including a backbone network, an RPN network, and an Roi pooling operation module, and a foreground enhancement network module and a spatial attention module are added between the backbone network and the RPN network; and a spatial attention module is added at the backend of the Roi pooling operation module;
[0090] The foreground enhancement network module is used to extract foreground attention features from the input image; the spatial attention module is used to extract spatial attention from the input image.
[0091] As a preferred embodiment, the fine-tuning training module includes:
[0092] The basic feature extraction unit uses the ResNet50 pre-trained model as the backbone network, takes the query image Qc, the support image Sc, and the support image Sn as inputs, extracts the basic features of the input images, and obtains the support image feature map and the query image feature map;
[0093] The foreground enhancement unit is used to input the support image feature map into the foreground enhancement network module, extract the foreground attention features, obtain the first feature map, and then perform broadcast pixel addition on the first feature map and the input support image feature map to obtain the second feature map;
[0094] The first spatial attention extraction unit is used to input the query image feature map into the spatial attention module, extract the spatial attention features, and obtain a third feature map with negative numbers, zeros, or positive numbers; input the third feature map into the sigmoid function for activation operation to obtain the fourth feature map, and perform weighted fusion on the fourth feature map and the input query image feature map to obtain the fifth feature map;
[0095] The RPN network unit is used to input the second feature map and the fifth feature map into the RPN network to generate target candidate boxes;
[0096] The pooling unit is used to input the target candidate boxes into the Roi pooling operation module to obtain the Roi feature map;
[0097] The second spatial attention extraction unit is used to input the Roi feature map into the spatial attention module, extract the spatial attention features, obtain the final feature map, perform classification and coordinate regression on the final feature map, and obtain the prediction result, where the prediction result includes the position information of the target box and the matching result.
[0098] As a preferred embodiment, the fine-tuning training module includes:
[0099] The support pair generation unit is used to use the support image Sn as the negative sample and the support image Sc as the positive sample to perform four pairings, namely the first positive support pair formed by the positive sample and the foreground target box, the second positive support pair formed by the positive sample and the background target box, and the negative support pair formed by the negative sample and the foreground and background target boxes;
[0100] The selection unit is used to select N pairs of the first positive support pairs, 2N pairs of the second positive support pairs, and N pairs of negative support pairs according to the matching results output by the target detection model during the training process, and calculate the matching losses of the selected pairs.
[0101] Example three:
[0102] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a method for small sample object detection as described in any embodiment of the present invention.
[0103] Embodiment 4:
[0104] This embodiment provides a computer-readable storage medium, on which a computer program is stored. The program, when executed by a processor, implements a method for small sample object detection as described in any embodiment of the present invention.
[0105] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent the cases of A existing alone, A and B existing simultaneously, and B existing alone. Where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the front and rear associated objects. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.
[0106] Those of ordinary skill in the art can realize that the units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0107] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0108] In several embodiments provided by the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0109] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for small-sample object detection, characterized in that, Including the following steps: Establish a target detection model; Use a public dataset to preliminarily train the target detection model, and after the preliminary training is completed, add a small sample dataset containing target object images to fine-tune the target detection model; wherein, there is no overlapping part between the small sample training set and the public dataset; When performing fine-tuning training, sample a number of support images and query images from the overall dataset after adding the small sample dataset to generate a support sample set and a query sample set; Randomly select a query image Qc containing target objects of category c and a support image Sc containing target objects of category c from the query sample set, and take another support image Sn containing target objects different from category c to form a training triplet. In the training triplet, mark the target objects of category c in the query image Qc as the foreground, and mark all other target objects as the background; Input the training triplet into the target detection model for feature extraction, obtain the target bounding boxes of each image in the training triplet, calculate the bounding box loss of the target bounding boxes, and determine which support image the input query image Qc matches through the target detection model. Calculate the matching loss according to the matching result output by the target detection model, and iteratively train the target detection model by combining the bounding box loss and the matching loss to obtain a trained small sample target detection model; Use the trained small sample target detection model for target detection; Among them, the step of inputting the training triplet into the target detection model for feature extraction is specifically: Adopt the ResNet50 pre-trained model as the backbone network, input the query image Qc, the support image Sc, and the support image Sn, extract the basic features of the input images, and obtain the support image feature map and the query image feature map; Input the support image feature map into the foreground enhancement network module to extract foreground attention features, obtain the first feature map, and then perform broadcast pixel addition of the first feature map and the input support image feature map to obtain the second feature map; Input the query image feature map into the spatial attention module to extract spatial attention features, and obtain a third feature map with negative numbers, zeros, or positive numbers; input the third feature map into the sigmoid function for activation operation to obtain the fourth feature map, and perform weighted fusion of the fourth feature map and the input query image feature map to obtain the fifth feature map; Input the second feature map and the fifth feature map into the RPN network to generate target candidate bounding boxes; Input the target candidate bounding boxes into the Roi pooling operation module to obtain the Roi feature map; Input the Roi feature map into the spatial attention module to extract spatial attention features, obtain the final feature map, perform classification and coordinate regression on the final feature map, and obtain the prediction result. The prediction result includes the position information of the target bounding box and the matching result.
2. The method for small sample target detection according to claim 1, wherein: The target detection model is established based on the Faster-RCNN network, including a backbone network, an RPN network, and an Roi pooling operation module, and a foreground enhancement network module and a spatial attention module are added between the backbone network and the RPN network; and a spatial attention module is added to the backend of the Roi pooling operation module; The foreground enhancement network module is used to extract foreground attention features from the input image; The spatial attention module is used to extract spatial attention from the input image.
3. A method for small sample object detection according to claim 1, characterized in that In the process of calculating the matching loss according to the matching result output by the target detection model and iteratively training the target detection model by combining the bounding box loss and the matching loss: The support image Sn is used as a negative sample, and the support image Sc is used as a positive sample, and four pairs are formed, namely the first positive support pair formed by the positive sample and the foreground target box, the second positive support pair formed by the positive sample and the background target box, and the negative support pair formed by the negative sample and the foreground and background target boxes; During the training process, N pairs of the first positive support pairs, 2N pairs of the second positive support pairs, and N pairs of negative support pairs are selected according to the matching result output by the target detection model, and the matching loss of the selected pairs is calculated.
4. A system for small-sample object detection, characterized in that, Including: A model establishment module, used to establish a target detection model; A preliminary training module, used to preliminarily train the target detection model using a public data set, and add a small sample data set containing target object images to fine-tune the target detection model after the preliminary training is completed; wherein, there is no overlapping part between the small sample training set and the public data set; A sample generation module, used to sample a number of support images and query images from the overall data set after adding the small sample data set during the fine-tuning training to generate a support sample set and a query sample set; A triplet generation module, used to randomly select a query image Qc containing target objects of category c in the query sample set, and a support image Sc containing target objects of category c, and another support image Sn containing target objects different from category c to form a training triplet. In the training triplet, the target objects of category c in the query image Qc are marked as foreground, and all other target objects are marked as background; A fine-tuning training module, used to input the training triplet into the target detection model for feature extraction, obtain the target boxes of each image in the training triplet, calculate the bounding box loss of the target boxes, and judge which support image the input query image Qc matches through the target detection model, calculate the matching loss according to the matching result output by the target detection model, and iteratively train the target detection model by combining the bounding box loss and the matching loss to obtain a trained small sample target detection model; A detection module, used to perform target detection using the trained small sample target detection model; In the fine-tuning training module, it includes: A basic feature extraction unit, using the ResNet50 pre-trained model as the backbone network, inputting the query image Qc, the support image Sc, and the support image Sn, extracting the basic features of the input images, and obtaining the support image feature map and the query image feature map; A foreground enhancement unit, configured to input a support image feature map into a foreground enhancement network module, extract foreground attention features to obtain a first feature map, and then perform broadcast pixel addition on the first feature map and the input support image feature map to obtain a second feature map; A first spatial attention extraction unit, configured to input a query image feature map into a spatial attention module, extract spatial attention features to obtain a third feature map with negative numbers, zeros or positive numbers; input the third feature map into a sigmoid function for activation operation to obtain a fourth feature map, and perform weighted fusion on the fourth feature map and the input query image feature map to obtain a fifth feature map; An RPN network unit, configured to input the second feature map and the fifth feature map into an RPN network to generate target candidate boxes; A pooling unit, configured to input the target candidate boxes into an Roi pooling operation module to obtain an Roi feature map; A second spatial attention extraction unit, configured to input the Roi feature map into a spatial attention module, extract spatial attention features to obtain a final feature map, perform classification and coordinate regression on the final feature map to obtain a prediction result, and the prediction result includes the position information of the target box and the matching result.
5. A small-sample object detection system according to claim 4, wherein: The object detection model is established based on the Faster-RCNN network, includes a backbone network, an RPN network and an Roi pooling operation module, and a foreground enhancement network module and a spatial attention module are added between the backbone network and the RPN network; And a spatial attention module is added to the backend of the Roi pooling operation module; The foreground enhancement network module is configured to extract foreground attention features from the input image; The spatial attention module is configured to extract spatial attention from the input image.
6. The system for small sample object detection according to claim 4, wherein The fine-tuning training module includes: A support pair generation unit, configured to use the support image Sn as a negative sample and the support image Sc as a positive sample to perform four pairings, namely a first positive support pair formed by the positive sample and the foreground target box, a second positive support pair formed by the positive sample and the background target box, and a negative support pair formed by the negative sample and the foreground and background target boxes; A selection unit, configured to select N first positive support pairs, 2N second positive support pairs and N negative support pairs according to the matching results output by the object detection model during the training process, and calculate the matching losses of the selected pairs.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that When the processor executes the program, it implements a small-sample object detection method according to any one of claims 1 to 3.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements a small-sample object detection method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Underwater target detection method based on small sample machine learning and hyper-parameter optimization
CN113239980A
Image recognition model training method, image recognition method and related device
CN111368934A
Target detection method and device, electronic equipment and storage medium
CN112906685A