Few-Shot Object Detection Method Based on Multi-View Learning and Meta-Learning
By adopting multi-view learning and meta-learning methods in the small sample object detection method, a multi-view data set is constructed and the meta-learning model parameter training method is used, the problem of feature forgetting and overfitting in small sample object detection is solved, and the detection accuracy of the model in basic categories and small sample categories is improved.
Patent Information
- Application Number
- CN202111453576.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-01
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-12-01
AI Technical Summary
The existing small sample object detection method has feature forgetting problems in the model fine-tuning stage, resulting in a decrease in the detection ability of the model in the basic category and is prone to overfitting problems.
A small sample object detection method based on multi-view learning and meta-learning is adopted. By constructing a multi-view data set and using meta-learning model parameter training method, multi-view transfer feature information is provided in the model fine-tuning stage, the current transfer learning situation is judged, and actions to promote or suppress transfer are taken.
It effectively solves the feature forgetting problem of the basic category and the overfitting problem of small sample categories, ensuring the detection accuracy of the model on the basic category and small sample categories.
Smart Images

Figure CN114119966B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and specifically relates to a small sample target detection method based on multi-view learning and meta-learning. Background Art
[0002] Small sample target detection technology aims to detect the corresponding objects from images with a small sample size, and has important application value in the fields of maritime rescue, medical imaging, etc. Since the number of samples required to train neural networks is large, the core problem of small sample target detection is how to transfer the common features of the detected objects to the objects in the small sample category, so that the model can quickly adapt to the features of the small sample category and obtain the same level of detection results.
[0003] With the development of deep learning, the detection accuracy (mAP) of small sample target detection results has been significantly improved. However, existing methods have a serious problem of feature forgetting during the model fine-tuning stage, forgetting the features previously learned on categories with sufficient samples (basic categories). This is because neural networks tend to remember the features of the current training samples. After the model is trained on small sample categories such as medical images, the detection ability of basic categories such as people and cars that have previously performed well will be greatly reduced. At the same time, because the amount of small sample data is small, the model is prone to overfitting problems on very few small sample data sets during the fine-tuning process. The feature forgetting problem will cause the model to gradually forget the common features of the detected objects, and it will also hinder the transfer learning of small sample features to a certain extent, resulting in varying degrees of accuracy reduction in the model on both basic categories and small sample categories. Summary of the invention
[0004] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art and to provide a small sample target detection method based on multi-view learning and meta-learning. By constructing a sample multi-view data set and using a model parameter training method based on meta-learning, multi-view migration feature information is provided in the model fine-tuning stage. The current transfer learning situation is judged based on this information, and actions are taken to promote or inhibit migration, thereby effectively solving the feature forgetting problem of the basic category and the overfitting problem of the small sample category.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] The present invention provides a small sample target detection method based on multi-view learning and meta-learning, comprising the following steps:
[0007] A small sample target detection model is constructed, and a small sample target detection model with a two-stage training method is used as a target detector; the two-stage training method is divided into a pre-training stage and a model fine-tuning stage. The training sets used in the pre-training stage and the model fine-tuning stage are different. The pre-training stage uses all basic category samples to allow the model to learn the common features of the image from a large number of basic category samples; in the model fine-tuning stage, the model transfers the learned features of the basic category samples to the feature learning of the small sample category; the target detector includes a backbone network, a candidate box extractor, a candidate box pooling layer, a candidate box feature convolution layer, a regressor, a classifier, and a high-confidence feature comparison learner;
[0008] The inter-class sample pair sampling method based on multi-view learning adopts the principle of category balance and divides the basic category dataset into multiple basic category sub-datasets. The number of samples in each sub-dataset is equal to the number of small sample category samples. Each basic category sub-dataset and small sample category samples are combined respectively to obtain multiple combined single-view mixed datasets, namely multi-view datasets.
[0009] Based on the feature contrast learning method of high-confidence deep features, in the fine-tuning stage of the small sample target detection model, the multi-view dataset is input into the small sample target detection model, and the high-confidence feature contrast learner selects the high-confidence features of the basic category and the small sample category, and constructs the loss function according to the Euclidean distance between the high-confidence features to achieve intra-class and inter-class feature contrast learning of the basic category and the small sample category;
[0010] Based on the meta-learning model parameter training method, in the fine-tuning stage of the small sample target detection model, the multi-view dataset is input into the small sample target detection model to obtain the loss values of the basic category and the small sample category respectively, and the gradient corresponding to the loss value is calculated and fed back to update the parameters of the small sample target detection model.
[0011] As a preferred technical solution, the pre-trained detector uses a two-stage detector Faster-RCNN.
[0012] As a preferred technical solution, the backbone network adopts the ResNet-101 network architecture.
[0013] As a preferred technical solution, after building a small sample target detection model, the small sample target detection model is first pre-trained using a basic category dataset, and then the small sample target detection model is fine-tuned using a multi-view dataset.
[0014] As a preferred technical solution, the inter-class sample pair sampling method based on multi-view learning is specifically:
[0015] The multi-view dataset D is composed of the basic category dataset D base And the small sample category dataset Dnovel The composition is expressed as:
[0016]
[0017] in, They represent the i-th base category sample and the j-th small sample category sample, x represents the sample, i, j represents the sample number, base, novel represent the base category and small sample category respectively, N 1 , N represents the total number of basic category samples and the total number of small sample category samples, and N 1 >>N;
[0018] use and They represent the i-th basic category and the j-th small sample category, respectively. C represents the category, i and j represent the sample numbers, and base and novel represent the basic category and the small sample category, respectively.
[0019] Different N samples are sampled from the basic category and the small sample category to obtain M sub-datasets of the basic category and 1 data set of the small sample category. Each sub-dataset of the basic category is combined with the data set of the small sample category to obtain a single-view mixed data set. After sampling, a multi-view data set of M views is obtained. D all Represents a multi-view dataset, expressed as:
[0020] D all ={D 1 , D 2 , ..., D M}
[0021] Among them, D i represents the mixed dataset of the i-th view.
[0022] As a preferred technical solution, during the fine-tuning stage of the small sample target detection model, multiple single-view mixed data sets are sequentially put into the network for training.
[0023] As a preferred technical solution, the loss function is constructed according to the Euclidean distance between high-confidence features, specifically:
[0024] In the fine-tuning stage of the small sample target detection model, the images in the multi-view dataset are passed through the backbone network, the candidate box extractor, the candidate box pooling layer and the candidate box feature convolution layer to obtain the feature encoding of N candidate boxes. i and i Represent the feature code and true label of the i-th candidate box respectively. Use the fully connected layer and L2 regularization operation to process the feature code to obtain the regularized feature code of the i-th candidate box.
[0025] Match the candidate box with the real object. According to the degree of overlap between the candidate box and the real object, retain the regularized feature encoding of the high-confidence candidate box with an intersection-over-union ratio greater than 0.7. The intersection-over-union ratio IOU is expressed as:
[0026]
[0027] Among them, d 1 and d 2 Respectively represent the area of the candidate box and the area of the real object;
[0028] Constructing the contrast loss function L C , specifically expressed as:
[0029]
[0030]
[0031] Among them, u i Represents the IOU value between the i-th candidate box and the real object, represents the regularized feature encoding of the kth candidate box, represents the feature contrast learning loss function of the i-th candidate box, τ is a hyperparameter, y i represents the true label of the i-th candidate box, N yi Indicates that the true category is y i The total number of candidate boxes, II{y i =y j} represents an indicative function that determines whether the true label of the i-th candidate box is the same as the true label of the j-th candidate box. If they are the same, the value is 1, otherwise it is 0.
[0032] As a preferred technical solution, the model parameter training method based on meta-learning is specifically as follows:
[0033] In the fine-tuning stage of the small sample target detection model, the multi-view dataset obtains deep features through the backbone network and the candidate box pooling layer, and further obtains a total loss value L through the locator, classifier and high-confidence feature comparison learner;
[0034] According to the true category of the candidate box, the loss value L is divided into the loss value L of the basic category base And the loss value L of the small sample class novel ; First calculate the basic category loss value L base The gradient is returned to update the parameters of the small sample target detection model, and then the small sample category loss value L is calculated. novel The gradient of is then passed back to update the parameters of the small sample target detection model. The parameter update formula is as follows:
[0035]
[0036]
[0037] θ i =θ i-1 +γ·(θ i,2 -θ i-1 )
[0038] Among them, θ i represents the parameter value of the small sample target detection model in the i-th iteration, α and γ represent the parameter learning rate and parameter change learning rate of the small sample target detection model respectively, and θ i,1 Represents θ i-1 Passing by L base The parameters of the small-shot target detection model updated by the gradient backpropagation, θ i,2 Represents θ i,1 Passing by L novel The gradient backpropagation updates the parameters of the few-shot object detection model.
[0039] As a preferred technical solution, the total loss value L is expressed as:
[0040] L=L reg +L cls +L C
[0041] Among them, L reg , L cls Represent the loss values of the regressor and classifier respectively, L C Represents the high confidence feature contrast learning loss value.
[0042] As a preferred technical solution, when updating the parameters of the small sample target detection model, all parameters of the backbone network and the candidate box pooling layer are frozen to retain the feature distribution extracted by the small sample target detection model.
[0043] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0044] 1. Based on the inter-class sample pair sampling method of multi-view learning, the present invention constructs a multi-view dataset with balanced categories and more sufficient sample size, which alleviates the feature forgetting problem of small sample target detection model on basic categories and provides multi-view comparative learning opportunities for the features of small sample categories.
[0045] 2. The present invention further enhances the contrastive learning capability of multi-view datasets through high-confidence feature contrastive learning and parameter learning based on meta-learning strategy. It retains a large number of basic category features by freezing the parameters of the backbone network and the candidate box pooling layer and updating the parameters by alternating the return gradient. In the fine-tuning stage, the influence of the return gradient of the small sample category on the model features is considered, and the parameter update is enhanced or suppressed accordingly, thereby alleviating the problem of forgetting the basic category features and the overfitting problem of the small sample category. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0047] Figure 1 This is a flow chart of a small sample target detection method based on multi-view learning and meta-learning according to an embodiment of the present invention;
[0048] Figure 2 This is a flow chart of model parameter updating based on meta-learning in an embodiment of the present invention. DETAILED DESCRIPTION
[0049] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.
[0050] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments.
[0051] like Figure 1 As shown, in one embodiment of the present application, a small sample target detection method based on multi-view learning and meta-learning is provided, comprising the following steps:
[0052] S1. Construct a small sample target detection model, and use a small sample target detection model with a two-stage training method as a target detector; the two-stage training method is divided into a pre-training stage and a model fine-tuning stage. The pre-training stage and the model fine-tuning stage use different training sets. The pre-training stage uses all basic category samples to allow the model to learn the common features of the image from a large number of basic category samples; in the model fine-tuning stage, the model transfers the learned features of the basic category samples to the feature learning of the small sample category, including a backbone network, a candidate box extractor, a candidate box pooling layer, a candidate box feature convolution layer, a regressor, a classifier, and a high-confidence feature comparison learner;
[0053] After building the small sample target detection model, the basic category dataset is first used to pre-train the small sample target detection model, and then the multi-view dataset is used to fine-tune the small sample target detection model.
[0054] In this embodiment, the pre-trained detector uses the two-stage detector Faster-RCNN, and the backbone network uses the ResNet-101 network;
[0055] S2. The inter-class sample pair sampling method based on multi-view learning is used to increase a large number of basic category features, providing multi-view information for subsequent feature comparison learning; the principle of category balance is adopted to divide the sufficient basic category data set into multiple basic category sub-data sets, and the number of samples in each sub-data set is equal to the number of samples in the small sample category. Each basic category sub-data set and the small sample category samples are combined respectively to obtain multiple combined single-view mixed data sets. The collection of these data sets is called a multi-view data set, specifically:
[0056] The multi-view dataset D is composed of the basic category dataset D base And the small sample category dataset D novel The composition is expressed as:
[0057]
[0058] in, They represent the i-th basic category sample and the j-th small sample category sample, x represents the sample, i, j represent the sample number, base, hovel represent the basic category and small sample category respectively, N 1 , N represents the total number of basic category samples and the total number of small sample category samples. In the small sample target detection task, N 1 Much larger than N;
[0059] use and Denote the i-th basic category and the j-th small sample category, respectively, C denotes the category, and N different samples are sampled from the basic category and the small sample category to obtain M sub-datasets of the basic category and 1 data set of the small sample category. Each sub-dataset of the basic category is combined with the data set of the small sample category to form a single-view mixed data set. After sampling, a multi-view data set of M views is obtained. D all represents a multi-view dataset, then the multi-view dataset D all It is expressed as:
[0060] D all ={D 1 , D 2 , ..., D M}
[0061] Among them, D i represents the mixed dataset of the i-th view.
[0062] In this embodiment, during the fine-tuning stage of the small sample target detection model, multiple single-view mixed data sets are sequentially input into the network for training.
[0063] S3, feature contrast learning method based on high-confidence deep features, is used to further learn the spatial distribution of features and enhance the feature contrast learning ability of multi-view data sets; the high-confidence feature contrast learner selects high-confidence features of basic categories and small sample categories, and constructs a loss function based on the Euclidean distance between high-confidence features to achieve intra-class and inter-class feature contrast learning of basic categories and small sample categories, specifically:
[0064] In the model fine-tuning stage, the multi-view dataset images are passed through the backbone network, candidate box extractor, candidate box pooling layer and candidate box feature convolution layer to obtain 1024-dimensional feature encoding of N candidate boxes. i and i Represent the feature code and true label of the i-th candidate box respectively. Use the fully connected layer and L2 regularization operation to process the feature code to obtain the 128-dimensional regularized feature code of the i-th candidate box. Further reduce the feature dimension and make the feature distribution more concentrated;
[0065] Match the candidate box with the real object. According to the degree of overlap between the candidate box and the real object, retain the regularized feature encoding of the high-confidence candidate box with an intersection-over-union ratio greater than 0.7. The intersection-over-union ratio IOU is defined as:
[0066]
[0067] Among them, d 1 and d 2 Respectively represent the area of the candidate box and the area of the real object;
[0068] Constructing the contrast loss function L C , specifically expressed as:
[0069]
[0070]
[0071] Among them, u i Represents the IOU value between the i-th candidate box and the real object, represents the regularized feature encoding of the kth candidate box, represents the feature contrast learning loss function of the i-th candidate box, τ is a hyperparameter, and in this embodiment, its value is 0.2, y i represents the true label of the i-th candidate box, N yi Indicates that the true category is y i The total number of candidate boxes, II{y i =y j} represents an indicative function that determines whether the true label of the i-th candidate box is the same as the true label of the j-th candidate box. If they are the same, the value is 1, otherwise the value is 0.
[0072] In this embodiment, high-confidence candidate box features are selected based on the degree of overlap between the candidate box features and the real objects, which can better reflect the characteristics of the category. The feature contrast learning method is used to increase the feature distance of different categories and reduce the feature distance of the same category, providing multi-perspective contrast learning information for small sample features and alleviating the overfitting problem of small sample categories caused by insufficient sample size.
[0073] S4. Meta-learning-based model parameter training method is used to alleviate the feature forgetting problem that occurs in the fine-tuning process of small sample target detection models; according to the learning direction of the current model, it is judged whether the current model needs to strengthen or suppress the transfer learning ability, specifically:
[0074] like Figure 2 As shown in the figure, in the model fine-tuning stage, the multi-view dataset obtains deep features through the backbone network and the candidate box pooling layer, and further obtains a total loss value L through the locator, classifier and high-confidence feature comparison learner. The calculation method is:
[0075] L=L reg +L cls +L C
[0076] Using L reg , L cls Represent the loss values of the regressor and classifier respectively.
[0077] According to the true category of the candidate box, the loss value L can be divided into the loss value L of the basic category baseAnd the loss value L of the small sample class novel ; First calculate the basic category loss value L base The gradient is returned to update the parameters of the small sample target detection model, and then the small sample category loss value L is calculated. novel The gradient is returned to update the parameters of the small sample target detection model. During the parameter update process, all parameters of the backbone network and the candidate box pooling layer are frozen to keep the feature distribution extracted by the model relatively stable. The specific parameter update formula is:
[0078]
[0079]
[0080] θ i =θ i-1 +γ·(θ i,2 -θ i-1 )
[0081] Among them, θ i represents the parameter value of the small sample target detection model in the i-th iteration, α and γ represent the parameter learning rate and parameter change learning rate of the small sample target detection model respectively, and θ i,1 Represents θ i-1 Passing by L base The parameters of the small-shot target detection model updated by the gradient backpropagation, θ i,2 Represents θ i,1 Passing by L novel The parameters of the small sample target detection model updated by the gradient backpropagation. In this embodiment, α and γ are set to 0.002 and 1 respectively.
[0082] In this embodiment, L reg , L cls , L C According to the true category of the candidate box, it can be divided into the loss value of the basic category part and the loss value of the small sample category part, which can be expressed as:
[0083]
[0084]
[0085]
[0086]
[0087]
[0088] They represent the basic category loss value and the small sample category loss value of the m loss value respectively, so the basic category loss value L baseAnd the loss value L of the small sample class novel The sum is equivalent to the total loss value L.
[0089] The invention is based on a small sample target detection method of multi-view learning and meta-learning, uses an inter-class sample pair sampling method to construct a class-balanced multi-view data set, inputs it into a small sample target detection model for model fine-tuning, and provides a multi-view comparative learning opportunity for the features of the small sample category; the input multi-view data set image passes through a backbone network, and the feature map of the image is output from the fourth convolution group of the backbone network; then a candidate box extractor is used to perform binary classification and regression positioning of anchor points to obtain a series of candidate boxes, which are then passed through a candidate box feature convolution layer, input into a classifier, a regressor and a high-confidence feature comparative learner to calculate the loss value, and finally the total loss value is divided into a basic category loss value and a small sample category loss value, thereby further enhancing the comparative learning ability of the multi-view data set; when the backbone network and the candidate box pooling layer are frozen, the gradients of the basic category and the small sample category are successively fed back and the parameters of the small sample target detection model are updated, thereby effectively solving the problem of forgetting the features of the basic category and the problem of overfitting of the small sample category.
[0090] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited to the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously.
[0091] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0092] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. Small sample target detection method based on multi-view learning and meta-learning, It is characterized in that The steps include: A small sample target detection model is constructed, and a small sample target detection model with a two-stage training method is used as a target detector; the two-stage training method is divided into a pre-training stage and a model fine-tuning stage. The training sets used in the pre-training stage and the model fine-tuning stage are different. The pre-training stage uses all basic category samples to allow the model to learn the common features of the image from a large number of basic category samples; in the model fine-tuning stage, the model transfers the learned features of the basic category samples to the feature learning of the small sample category; the target detector includes a backbone network, a candidate box extractor, a candidate box pooling layer, a candidate box feature convolution layer, a regressor, a classifier, and a high-confidence feature comparison learner; Based on the inter-class sample pair sampling method of multi-view learning, the basic category dataset is divided into multiple basic category sub-datasets using the principle of category balance. The number of samples in each sub-dataset is equal to the number of samples in the small sample category. Each basic category sub-dataset and the small sample category samples are combined to obtain multiple combined single-view mixed datasets, namely multi-view datasets. Based on the feature contrast learning method of high-confidence deep features, in the fine-tuning stage of the small sample target detection model, the multi-view dataset is input into the small sample target detection model, and the high-confidence feature contrast learner selects the high-confidence features of the basic category and the small sample category, and constructs the loss function according to the Euclidean distance between the high-confidence features to achieve intra-class and inter-class feature contrast learning of the basic category and the small sample category; Based on the meta-learning model parameter training method, in the small sample target detection model fine-tuning stage, the multi-view dataset is input into the small sample target detection model, and the loss values of the basic category and the small sample category are obtained respectively. The gradient corresponding to the loss value is calculated and fed back to update the parameters of the small sample target detection model; The inter-class sample pair sampling method based on multi-view learning is specifically as follows: The multi-view dataset D is composed of the basic category dataset D base And the small sample category dataset D novel The composition is expressed as: in, They represent the i-th base category sample and the j-th small sample category sample, x represents the sample, i, j represents the sample number, base, novel represent the base category and small sample category respectively, N 1 , N represent the total number of basic category samples and the total number of small sample category samples, and N 1 >>N; use and They represent the i-th basic category and the j-th small sample category, respectively. C represents the category, i and j represent the sample numbers. Base and novel represent the basic category and the small sample category, respectively. Different N samples are sampled from the basic category and the small sample category to obtain M sub-datasets of the basic category and 1 data set of the small sample category. Each sub-dataset of the basic category is combined with the data set of the small sample category to obtain a single-view mixed data set. After sampling, a multi-view data set of M views is obtained. D all Represents a multi-view dataset, expressed as: D all ={D 1 ,D 2 ,…,D M } Among them, D i represents the mixed data set of the i-th view; The model parameter training method based on meta-learning is specifically as follows: In the fine-tuning stage of the small sample target detection model, the multi-view dataset obtains deep features through the backbone network and the candidate box pooling layer, and further obtains a total loss value L through the locator, classifier and high-confidence feature comparison learner; According to the true category of the candidate box, the loss value L is divided into the loss value L of the basic category base And the loss value L of the small sample class novel ; First calculate the basic category loss value L base The gradient is then passed back to update the parameters of the small sample target detection model, and then the small sample category loss value L is calculated. novel The gradient of is returned and the parameters of the small sample target detection model are updated. The parameter update formula is as follows: i i =θ i-1 +γ·(θ i,2 -θ i-1 ) Among them, θ i represents the parameter value of the small sample target detection model in the i-th iteration, α and γ represent the parameter learning rate and parameter change learning rate of the small sample target detection model respectively, and θ i,1 Represents θ i-1 Passing by L base The parameters of the small-shot target detection model updated by the gradient backpropagation, θ i,2 Represents θ i,1 Passing by L novel The gradient backpropagation updates the parameters of the few-shot object detection model.
2. According to the small sample target detection method based on multi-view learning and meta-learning according to claim 1, It is characterized in that The pre-trained detector uses a two-stage detector Faster-RCNN.
3. According to the small sample target detection method based on multi-view learning and meta-learning according to claim 1, It is characterized in that The backbone network adopts the ResNet-101 network architecture.
4. According to the small sample target detection method based on multi-view learning and meta-learning according to claim 1, It is characterized in that After building the small sample target detection model, the basic category dataset is first used to pre-train the small sample target detection model, and then the multi-view dataset is used to fine-tune the small sample target detection model.
5. According to the small sample target detection method based on multi-view learning and meta-learning according to claim 1, It is characterized in that During the fine-tuning stage of the small sample target detection model, multiple single-view mixed datasets are sequentially put into the network for training.
6. According to the small sample target detection method based on multi-view learning and meta-learning as claimed in claim 1, It is characterized in that The loss function is constructed based on the Euclidean distance between high-confidence features, specifically: In the fine-tuning stage of the small sample target detection model, the images in the multi-view dataset are passed through the backbone network, the candidate box extractor, the candidate box pooling layer and the candidate box feature convolution layer to obtain the feature encoding of N candidate boxes. i and i Represent the feature code and true label of the i-th candidate box respectively. Use the fully connected layer and L2 regularization operation to process the feature code to obtain the regularized feature code of the i-th candidate box. Match the candidate box with the real object. According to the degree of overlap between the candidate box and the real object, retain the regularized feature encoding of the high-confidence candidate box with an intersection-over-union ratio greater than 0.
7. The intersection-over-union ratio IOU is expressed as: Among them, d 1 and d 2 Respectively represent the area of the candidate box and the area of the real object; Constructing the contrast loss function L C , specifically expressed as: Among them, u i represents the IOU value between the i-th candidate box and the real object, represents the regularized feature encoding of the k-th candidate box, represents the feature contrastive learning loss function of the i-th candidate box, τ is a hyperparameter, y i represents the true label of the i-th candidate box, represents that the true category is y i the total number of candidate boxes, represents the indicator function for judging whether the true label of the i-th candidate box is the same as that of the j-th candidate box. If they are the same, the value is 1, otherwise 0.
7. According to the small sample target detection method based on multi-view learning and meta-learning as claimed in claim 1, It is characterized in that The total loss value L is expressed as: L=L reg +L cls +L C Among them, L reg ,L cls Represent the loss values of the regressor and classifier respectively, L c Represents the high confidence feature contrast learning loss value.
8. According to the small sample target detection method based on multi-view learning and meta-learning as claimed in claim 1, It is characterized in that When updating the parameters of the small sample target detection model, all parameters of the backbone network and the candidate box pooling layer are frozen to retain the feature distribution extracted by the small sample target detection model.
Citation Information
Patent Citations
Transfer learning-based multi-view commodity image retrieval and identification method
CN107908685A
Incremental small sample target detection method based on meta-learning
CN112329827A