Small-Sample Object Detection Method and Related Equipment Based on Multi-Angle Information
By using multi-angle information in small sample object detection, the initial detection model is trained using a specific network structure, and the problem of small sample size is solved, resulting in low object detection accuracy is achieved, and high-accuracy object detection in small sample situations is achieved.
Patent Information
- Application Number
- CN202211510029.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-11-29
AI Technical Summary
In some application scenarios of object detection, the small number of samples leads to poor results and low accuracy in existing object detection models, making it difficult to train a higher accuracy object detection model.
A small sample object detection method based on multi-angle information is adopted. By acquiring the image to be detected and the supported images at multiple corresponding shooting angles, the initial detection model is trained using a twin network, a relational distillation network, a candidate box generation network and a full connection layer to generate an accurate object detection model.
Even in the case of small samples, the object detection model trained by this method can significantly improve the accuracy of object detection and ensure the accuracy of predicted object detection results of the image to be detected.
Smart Images

Figure CN115731210B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a small sample target detection method based on multi-angle information and related equipment. Background Art
[0002] The target detection model is used to identify the area where the target is located and the category of the target in the image. The existing target detection model needs to be trained with a large number of pre-labeled image samples. Among them, the pre-labeled image samples can be understood as image samples with the area where the target is located and the category of the target pre-labeled.
[0003] However, in some application scenarios of target detection, the number of samples that can be collected is relatively small. For example, in the power scenario, since the probability of power system failure is very low, the number of image samples that can be collected for equipment failures such as hanging foreign objects, broken wires, and self-explosion of insulators is very small. When the number of samples is small, the target detection model trained by the existing method is less effective and the accuracy of target detection is low. Therefore, it is necessary to train a target detection model with higher accuracy when the number of samples is small to meet the target detection needs in small sample conditions. Summary of the invention
[0004] In view of this, an embodiment of the present invention provides a small sample target detection method and related equipment based on multi-angle information, so as to meet the target detection needs in small sample conditions by training a target detection model using multiple query images and support images at multiple different shooting angles with the same target category as the query images.
[0005] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0006] In a first aspect, the present application discloses a small sample target detection method based on multi-angle information, including:
[0007] Acquire the image to be detected;
[0008] Inputting the image to be detected and a plurality of support images of a support set into a target detection model, and obtaining and outputting a predicted target detection result of the image to be detected by the target detection model;
[0009] Among them, the predicted target detection result of the image to be detected is used to illustrate the target category of the predicted image to be detected and the region where the target is located in the query image; the target detection model is obtained by training an initial detection model with multiple query images and multiple support images corresponding to the query images; the multiple support images corresponding to the query image include support images at multiple different shooting angles with the same target category as the query image; the initial detection model includes: a siamese network, a relation distillation network, a candidate box generation network, and a fully connected layer; the siamese network is used to process the query image and multiple support images corresponding to the query image to obtain the feature map of the query image and the feature map of each support image; the relation distillation network is used to process the feature map of the query image and all support images corresponding to the query image to obtain an aggregated feature map; the aggregated feature map is used to illustrate the similarity between the feature map of the query image and the feature maps of all support images corresponding to the query image; the candidate box generation network is used to generate multiple candidate boxes of the aggregated feature map; the candidate box is the region where the predicted target is located; the fully connected layer is used to obtain the predicted target detection result of the query image according to each candidate box of the aggregated feature map.
[0010] Optionally, in the above small-sample target detection method based on multi-angle information, the construction process of the target detection model includes:
[0011] Construct an initial detection model;
[0012] Input the query image into the first feature extraction network of the siamese network to obtain the feature map of the query image; and input multiple support images corresponding to the query image into the second feature extraction network of the siamese network to obtain the feature map of each support image;
[0013] Input the feature map of the query image and the feature maps of all support images into the relation distillation network to obtain an aggregated feature map; wherein, the aggregated feature map is used to illustrate the similarity between the feature map of the query image and the feature maps of all support images;
[0014] Input the aggregated feature map into the candidate box generation network to generate multiple candidate boxes of the aggregated feature map;
[0015] Input each candidate box of the aggregated feature map into the fully connected layer to obtain the predicted target detection result of the query image; wherein, the predicted target detection result of the query image is used to illustrate the target category of the predicted query image and the region where the target is located in the query image;
[0016] Adjust the parameters in the initial detection model according to the error between the predicted target detection result and the actual target detection result. When the error between the predicted target detection result output by the adjusted initial detection model and the actual target detection result meets the preset convergence condition, determine the adjusted initial detection model as the target detection model.
[0017] Optionally, in the above small-sample object detection method based on multi-angle information, the step of inputting the feature map of the query image and the feature maps of all the support images into the relational distillation network to obtain an aggregated feature map includes:
[0018] Fuse the feature maps of all the support images to obtain a fused feature map;
[0019] Input the feature map of the query image and the fused feature map into the relational distillation network to obtain an aggregated feature map.
[0020] Optionally, in the above small-sample object detection method based on multi-angle information, the step of inputting the feature map of the query image and the fused feature map into the relational distillation network to obtain an aggregated feature map includes:
[0021] Input the feature map of the query image and the fused feature map into the relational distillation network. The relational distillation network performs K convolutions on the fused feature map to obtain a first feature matrix of the fused feature map; and performs N convolutions on the fused feature map to obtain a second feature matrix of the fused feature map; where N and K are two different positive integers;
[0022] Perform K convolutions on the feature map of the query image to obtain a first feature matrix of the feature map of the query image; and perform N convolutions on the feature map of the query image to obtain a second feature matrix of the feature map of the query image;
[0023] Aggregate the first feature matrix of the feature map of the query image and the first feature matrix of the fused feature map to obtain a weight matrix; where the weight matrix is used to illustrate the weight values of the features in the query image; the higher the similarity between the features in the query image and the features of the fused feature map, the greater the weight value of the features in the query image;
[0024] Multiply the weight matrix and the second feature matrix of the feature map of the query image to obtain the processed feature map of the query image;
[0025] Aggregate the processed feature map of the query image and the second feature matrix of the fused feature map to obtain an aggregated feature map.
[0026] Optionally, in the above small-sample object detection method based on multi-angle information, before inputting each candidate box of the aggregated feature map into the fully connected layer to obtain the predicted object detection result of the query image, the method further includes:
[0027] Filtering out a plurality of candidate boxes belonging to negative samples from all candidate boxes of the aggregated feature map according to the reference boxes of a plurality of support images corresponding to the query image; wherein, the reference box of the support image is the pre-annotated area where the target is located in the support image; the candidate box belonging to the negative sample is the candidate box in the non-target area;
[0028] The step of inputting each candidate box of the aggregated feature map into the fully connected layer to obtain the predicted object detection result of the query image includes:
[0029] Inputting each candidate box of the filtered aggregated feature map into the fully connected layer to obtain the predicted object detection result of the query image.
[0030] Optionally, in the above small-sample object detection method based on multi-angle information, the step of filtering out a plurality of candidate boxes belonging to negative samples from all candidate boxes of the aggregated feature map according to the reference boxes of a plurality of support images corresponding to the query image includes:
[0031] Inputting the reference boxes of a plurality of support images corresponding to the query image and all candidate boxes of the aggregated feature map into a metric learning network to match and obtain the best-angle reference box; wherein, the best-angle reference box is the reference box with the highest similarity between the reference boxes of a plurality of support images corresponding to the query image and the candidate boxes of the aggregated feature map;
[0032] Performing pooling on the best-angle reference box to obtain the depth feature vector of the best-angle reference box;
[0033] Performing convolution calculation on the depth feature vector of the best-angle reference box and each candidate box of the aggregated feature map respectively to obtain an attention candidate box corresponding to each candidate box;
[0034] Inputting each attention candidate box corresponding to each candidate box into a binary classification model to obtain the score of each attention candidate box; wherein, the score of the attention candidate box is used to indicate whether the candidate box is predicted as the target category;
[0035] Filtering out a plurality of candidate boxes belonging to negative samples from all candidate boxes of the aggregated feature map according to the score of each attention candidate box.
[0036] Optionally, in the above small-sample object detection method based on multi-angle information, the step of inputting each candidate box of the aggregated feature map into a fully connected layer to obtain the predicted object detection result of the query image includes:
[0037] Pool each candidate box of the aggregated feature map to obtain a feature vector for each candidate box;
[0038] For each candidate box, process the feature vector of the candidate box through the classification head of the fully connected layer to obtain the predicted class result of the candidate box;
[0039] Process the feature vector of the candidate box through the localization head of the fully connected layer to obtain the corresponding region of the candidate box in the query image; wherein, the predicted object detection result of the query image includes: the predicted class result of each candidate box and the corresponding region of each candidate box in the query image;
[0040] The step of adjusting the parameters in the initial detection model according to the error between the predicted object detection result and the actual object detection result includes:
[0041] Calculate a classification loss value according to the error between the predicted class result of the candidate box and the actual class result of the candidate box; and calculate a localization loss value according to the error between the corresponding region of the candidate box in the query image and the actual corresponding region of the candidate box in the query image;
[0042] Calculate the loss value of the initial detection model according to the classification loss value and the localization loss value;
[0043] Adjust the parameters in the initial detection model according to the loss value of the initial detection model.
[0044] In a second aspect, the present application discloses an object detection device, including:
[0045] An acquisition unit, configured to acquire an image to be detected;
[0046] An output unit, configured to input the image to be detected and multiple support images of a support set into an object detection model, and the object detection model obtains and outputs a predicted object detection result of the image to be detected;
[0047] Among them, the predicted object detection result of the image to be detected is used to illustrate the object category of the predicted image to be detected and the region where the object is located in the query image; the support set includes: support images of multiple object categories at multiple different shooting angles; the object detection model is obtained by training an initial detection model with multiple query images and the multiple support images corresponding to the query images; the multiple support images corresponding to the query image include support images of multiple different shooting angles with the same object category as the query image; the initial detection model includes: a siamese network, a relationship distillation network, a candidate box generation network, and a fully connected layer; the siamese network is used to process the query image and the multiple support images corresponding to the query image to obtain the feature map of the query image and the feature map of each support image; the relationship distillation network is used to process the feature map of the query image and all the support images corresponding to the query image to obtain an aggregated feature map; the aggregated feature map is used to illustrate the similarity between the feature map of the query image and the feature maps of all the support images corresponding to the query image; the candidate box generation network is used to generate multiple candidate boxes of the aggregated feature map; the candidate box is the region where the predicted object is located; the fully connected layer is used to obtain the predicted object detection result of the query image according to each candidate box of the aggregated feature map.
[0048] In a third aspect, the present application discloses a computer-readable medium, on which a computer program is stored, wherein when the program is executed by a processor, the method described in any one of the first aspects above is implemented.
[0049] In a fourth aspect, the present application discloses an object detection device, including:
[0050] One or more processors;
[0051] A storage device, on which one or more programs are stored;
[0052] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of the first aspects above.
[0053] Based on the small-sample object detection method based on multi-angle information provided by the embodiments of the present invention, by obtaining the image to be detected, and then inputting the image to be detected and multiple support images of the support set into the object detection model, the predicted object detection result of the image to be detected is obtained and output by the object detection model. Since the object detection model is trained by multiple query images and multiple support images corresponding to the query images, and the multiple support images corresponding to the query images include support images at multiple different shooting angles with the same object category as the query image, and the relationship distillation network in the initial detection model can process and obtain an aggregated feature map according to the feature map of the query image and all support images corresponding to the query image, where the aggregated feature map is used to illustrate the similarity between the feature map of the query image and the feature maps of all support images. Therefore, the aggregated feature map aggregates the features of support images and query images at multiple different shooting angles, which can make the features similar to those in the support images at multiple different shooting angles in the query image more obvious. Furthermore, the accuracy of multiple candidate boxes of the aggregated feature map generated by the candidate box generation network in the initial detection model will be relatively high, and the accuracy of the predicted object detection result of the query image obtained by the fully connected layer according to each candidate box of the aggregated feature map will also be improved. Therefore, even in the case of small samples, the accuracy of the object detection model trained by the initial detection model is still very high. Furthermore, after inputting the image to be detected into the object detection model, a predicted object detection result with high accuracy for the image to be detected can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0055] Figure 1 It is a schematic flowchart of a small-sample object detection method based on multi-angle information proposed by an embodiment of the present application;
[0056] Figure 2a It is a schematic flowchart of a method for constructing an object detection model proposed by an embodiment of the present application;
[0057] Figure 2b It is a schematic structural diagram of an object detection model proposed by an embodiment of the present application;
[0058] Figure 3 It is a schematic flowchart of a method for obtaining an aggregated feature map proposed by an embodiment of the present application;
[0059] Figure 4Schematic flowchart of another method for obtaining an aggregated feature map proposed in an embodiment of this application;
[0060] Figure 5 Schematic flowchart of a method for filtering candidate bounding boxes proposed in an embodiment of this application;
[0061] Figure 6 Schematic flowchart of a method for determining a predicted object detection result proposed in an embodiment of this application;
[0062] Figure 7 Schematic flowchart of a method for adjusting model parameters proposed in an embodiment of this application;
[0063] Figure 8 Schematic diagram of the result of an object detection device proposed in an embodiment of this application. Detailed implementation manners
[0064] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0065] In this application, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0066] Refer to Figure 1 , an embodiment of this application discloses a small-sample object detection method based on multi-angle information, and the method includes the following steps:
[0067] S101. Obtain an image to be detected.
[0068] Among them, the image to be detected refers to the image that needs to be subjected to object detection. Object detection refers to detecting the region where the object is located and the object category in the image. The object category refers to the category to which the object belongs. The object can be an entity such as a person, an animal, an object, etc. For example, in a certain image to be detected, a shell is placed on the beach, and the shell is the object that needs to be detected.
[0069] It should be noted that the present application embodiment does not limit the number, manner, etc. of obtaining the image to be detected. For the sake of more concise subsequent description, Figure 1 the shown process is introduced based on the processing process of a single image to be detected. When there are multiple images to be detected, the processing process for each image to be detected can refer to the processing process of a single image to be detected.
[0070] S102. Input the image to be detected and multiple support images of the support set into the target detection model. The target detection model obtains and outputs the predicted target detection result of the image to be detected, where the predicted target detection result of the image to be detected is used to illustrate the target category of the predicted image to be detected and the region where the target is located in the query image. The target detection model is obtained by training an initial detection model with multiple query images and multiple support images corresponding to the query images. The multiple support images corresponding to the query image include support images at multiple different shooting angles with the same target category as the query image.
[0071] The initial detection model includes: a siamese network, a relation distillation network, a candidate box generation network, and a fully connected layer. The initial detection model refers to a model for target detection of images before training.
[0072] The support image is an image sample used to provide the basis for target detection of the query image. The region where the target is located and the target category in the support image are pre-annotated. By calculating the similarity between the query image and multiple support images, the support image with similar features to the query image can be found. Furthermore, the features of the region where the target is located in the support image can be used to locate the region where the target is located in the query image and determine the target category. The query image (which can also be called the query image) is an image sample used for target detection.
[0073] In the small sample scenario, the number of support images and query images available for training the initial detection model is small. To improve the accuracy of the trained model, when collecting support images, for the same target, multiple different shooting angles are used for shooting, and thus multiple support images at different shooting angles with the same target category are obtained, enriching the number of support images. And because there are many support images at different shooting angles under the same target category, the query image can find support images with very high feature similarity. Furthermore, when training the initial detection model with multiple query images and multiple support images corresponding to the query images in the small sample case, since the multiple support images corresponding to the query image include support images at multiple different shooting angles with the same target category as the query image, a target detection model with high accuracy can be obtained.
[0074] The Siamese network is used to process and obtain the feature map of the query image and the feature maps of each of the multiple support images corresponding to the query image. The Siamese network has a feature extraction function. The Siamese network includes a first feature extraction network and a second feature extraction network. The weights are shared between the first feature extraction network and the second feature extraction network, and the operation processes of extracting features by the first feature extraction network and the second feature extraction network are similar or identical. The first feature extraction network is used to input the support image and obtain the feature map of the support image, while the second feature extraction network is used to input the query image and obtain the feature map of the query image.
[0075] The relational distillation network is used to process and obtain an aggregated feature map based on the feature map of the query image and all the support images corresponding to the query image. The aggregated feature map is used to illustrate the similarity between the feature map of the query image and the feature maps of all the support images corresponding to the query image. The relational distillation network has the function of aggregating (or fusing) features. Since the multiple support images corresponding to the query image include support images at multiple different shooting angles with the same target category as the query image, the aggregated feature map aggregates the features of the query image and the support images at multiple different shooting angles with the same target category as the query image, so that the features in the feature of the query image that are similar to the support images (i.e., the features of the support images with the same target category as the query image and the same shooting angle) will become more obvious (i.e., feature enhancement), while the features that are not similar to the support images are suppressed and become less obvious. Therefore, the aggregated feature map processed by the relational distillation network in the initial detection model can easily locate the area where the accurate target is located, and based on the target category of the support image similar to the query image, the target category of the query image can be accurately identified.
[0076] The candidate box generation network is used to generate multiple candidate boxes for the aggregated feature map. The candidate box is the area where the target predicted by the candidate box generation network is located.
[0077] The fully connected layer is used to obtain the predicted target detection result of the query image according to each candidate box of the aggregated feature map. The predicted target detection result of the query image is used to illustrate the predicted target category and the area where the target is located.
[0078] In the target detection model finally trained by the initial detection model, the image to be detected can be used as the query image. Through the trained Siamese network, relational distillation network, candidate box generation network and fully connected layer, based on the features of the support images similar to the query image in the multiple support images in the support set, the accurate predicted target detection result of the image to be detected can be obtained and output, meeting the requirement of target detection.
[0079] Specifically, all support images of the support set can be input in step S102. The support images input in step S102 can include multiple support images of different shooting angles under all collected target categories. The more shooting angles of the support images under the same target category in the support set, the more obvious the features of the region where the target of the query image is located in the aggregated feature map obtained by the relationship distillation network, and the higher the accuracy of the predicted target detection result of the image to be detected obtained by the final target detection image. Similarly, during the process of training the initial detection image, the more shooting angles of the support images under the same target category included in the multiple support images corresponding to the query image, the higher the accuracy of the finally trained target detection model.
[0080] Specifically, the execution process of step S102 is as follows: input the image to be detected into the first feature extraction network of the siamese network in the target detection model to obtain the feature map of the image to be detected. And input multiple support images of the support set into the second feature extraction network of the siamese network to obtain the feature map of each support image. Input the feature map of the image to be detected and the feature maps of all support images into the relationship distillation network to obtain an aggregated feature map. Input the aggregated feature map into the candidate box generation network to generate multiple candidate boxes of the aggregated feature map. Each candidate box of the aggregated feature map has a corresponding relationship with the candidate box of the image to be detected, that is, the candidate box of the aggregated feature map can be used to describe the region where the target in the image to be detected predicted by the candidate box generation network is located. Finally, input each candidate box of the aggregated feature map into the fully connected layer to obtain the predicted target detection result of the image to be detected.
[0081] Among them, the processing flow of the image to be detected and multiple support images inside the target detection model is the same as the processing flow of the query image and support images inside the initial detection model. For specific reference, please refer to the relevant description of the training process of the target detection model in the following Figure 2a part, and details are not described here again.
[0082] It should be noted that the subsequent usage method of the predicted target detection result of the image to be detected is not limited in the embodiments of the present application. For example, in a video processing scenario, each video frame can be used as the image to be detected, and then the predicted target detection result of the video frame can be obtained to automatically screen out the video frames where a certain target is located. There are many similar application scenarios for target detection, including but not limited to the content proposed in the embodiments of the present application.
[0083] Optionally, refer to Figure 2a , the construction process (which can also be understood as the training process) of the target detection model includes the following steps:
[0084] S201. Construct an initial detection model, where the initial detection model includes: a siamese network, a relational distillation network, a candidate box generation network, and a fully connected layer.
[0085] Optionally, the siamese network can be connected to the relational distillation network, the relational distillation network can be connected to the candidate box generation network, and the candidate box generation network can be further connected to the fully connected layer, so that the output of the previous network can be used as the input of the next network to achieve automatic training.
[0086] S202. Input the query image into the first feature extraction network of the siamese network to obtain the feature map of the query image, and input multiple support images corresponding to the query image into the second feature extraction network of the siamese network to obtain the feature map of each support image.
[0087] Input the query image into the first feature extraction network of the siamese network to obtain the feature map of the query image, and input multiple support images corresponding to the query image into the second feature extraction network of the siamese network to obtain the feature map of each support image, where the multiple support images corresponding to the query image include support images at multiple different shooting angles with the same target category as the query image.
[0088] Specifically, to ensure that the initial detection model can obtain an accurate predicted target detection result of the query image. During the training process, among the multiple support images corresponding to the query image input into the second feature extraction network, there should be at least multiple support images at multiple different shooting angles with the same target category as the query image, serving as positive samples in the support images for comparing similarity with the query image. When the multiple support images corresponding to the query image include multiple support images at multiple different shooting angles with the same target category as the query image, it can be ensured that among the support images input into the second feature extraction network, there are support images with the same target category as the query image and the same shooting angle, so that the initial detection model can accurately identify the target of the query image through the support images similar to the query image during the training process.
[0089] For example, as Figure 2b shown, input the multi-angle support images (i.e., the multiple support images at multiple different shooting angles with the same target category as the query image mentioned above) into a feature extraction network (i.e., the second feature extraction network mentioned above) to obtain the multi-angle feature maps (i.e., the feature maps of each support image corresponding to the query image). Input the query image into another feature extraction network to obtain the feature map (i.e., the feature map of the query image). Among them, the weights of the two feature extraction networks are shared to form a siamese network.
[0090] Optionally, only multiple support images at different shooting angles that are the same as the target category of the query image can be selected to perform step S202. Additionally, multiple support images at different shooting angles that are the same as the target category of the query image, as well as multiple support images at different shooting angles that are different from the target category of the query image, can be selected to perform step S202. When only using multiple support images at different shooting angles that are the same as the target category of the query image to perform step S202, it can enable the initial detection model to converge quickly during the training process and improve the training efficiency.
[0091] Exemplarily, for each query image in the dataset, multiple support images corresponding to the query image can be obtained from the dataset according to the target category in the query image, and then step S202 can be performed.
[0092] It should be noted that the process of inputting the query image into the first feature extraction network of the Siamese network to obtain the feature map of the query image can be executed in parallel with the process of inputting multiple support images corresponding to the query image into the second feature extraction network of the Siamese network. The order of execution of these two processes does not affect the implementation of the embodiments of this application.
[0093] The feature map of the query image includes the image features of the query image, and the feature map of the support image includes the image features of the support image. Since the first feature extraction network and the second feature extraction network in the Siamese network share weights, the feature extraction methods of the first feature extraction network and the second feature extraction network are basically similar. For the same feature, the feature maps processed by the first feature extraction network and the second feature extraction network will be nearly the same. Therefore, if there are very similar image features between the query image and the support image, the way of aggregating them through the subsequent distillation network in the feature maps of the query image and the support image can make the similar image features more obvious.
[0094] S203. Input the feature map of the query image and the feature maps of all support images into the relationship distillation network to obtain an aggregated feature map, where the aggregated feature map is used to illustrate the similarity between the feature map of the query image and the feature maps of all support images.
[0095] After the feature map of the query image and the feature maps of each support image corresponding to the query image are input into the relational distillation network, the relational distillation network aggregates the features in these feature maps. This process of aggregating features can be understood as a process of measuring the similarity between the features of the query image and the features of the support images. As a result, in the aggregated feature map, the features with a very high similarity between the query image and the support images will become more obvious after aggregation, while the features with a low similarity between the query image and the support images will become less obvious after aggregation. Therefore, the similarity between the feature map of the query image and the feature maps of all support images can be seen from the aggregated feature map.
[0096] Among all the support images in step S203, there are support images with multiple different shooting angles that have the same target as the query image. As a result, for the support images that have the same target as the query image and the same shooting angle, after the aggregation operation in step S203, the features related to the target in the query image will become more obvious. Compared with the training method that does not use support images with multiple different shooting angles, in the embodiment of the present application, the features of support images with multiple shooting angles are aggregated, further enhancing the similar features, so that the subsequent trained target detection model can accurately obtain the predicted target detection result of the image to be detected under different shooting angles.
[0097] Optionally, referring to Figure 3 , in a specific embodiment of the present application, an implementation manner of executing step S203 includes:
[0098] S301. Fuse the feature maps of all support images to obtain a fused feature map.
[0099] Among them, the feature maps of all support images in step S301 refer to the feature maps of each support image corresponding to the query image output by the second feature extraction network in step S202. In the fused feature map obtained by fusing the feature maps of all support images, the image features of all support images are fused, that is, all image features with the same target category as the query image are fused in the fused feature map. Furthermore, the fused feature map can be used as a positioning basis for the feature map of the query image to help the initial detection model detect the target of the query image.
[0100] S302. Input the feature map of the query image and the fused feature map into the relational distillation network to obtain an aggregated feature map.
[0101] The feature map of the query image and the fused feature map are input into the relational distillation network. The relational distillation network aggregates the feature map of the query image and the fused feature map, so that the partial features in the feature map of the query image that are similar to the fused feature map become more obvious, while the partial features that are different from the fused feature map are suppressed, reflecting the similarity between the feature map of the query image and the feature maps of all the support images corresponding to the query image. That is, the aggregated feature map specifically reflects the features of the part in the query image that is similar to the fused feature map.
[0102] For example, as Figure 2b shown, after multi-angle fusion of the multi-angle feature maps, a fused feature map is obtained. Then, after aggregating the fused feature map and the feature map (i.e., the feature map of the query image mentioned above), an aggregated feature map can be obtained.
[0103] Optionally, referring to Figure 4 , an implementation manner of executing step S302 includes:
[0104] S401. Input the feature map of the query image and the fused feature map into the relational distillation network. The relational distillation network performs K convolutions on the fused feature map to obtain the first feature matrix of the fused feature map, and performs N convolutions on the fused feature map to obtain the second feature matrix of the fused feature map, where N and K are two different positive integers.
[0105] Specifically, the fused feature map is represented as two categories. One category is Key, and the feature matrix obtained by performing K convolutions belongs to the Key part, that is, the first feature matrix of the fused feature map belongs to the Key part. One category is Value, and the feature matrix obtained by performing N convolutions belongs to the Value part, that is, the second feature matrix of the fused feature map belongs to the Value part.
[0106] S402. Perform K convolutions on the feature map of the query image to obtain the first feature matrix of the feature map of the query image, and perform N convolutions on the feature map of the query image to obtain the second feature matrix of the feature map of the query image.
[0107] Among them, the execution process and principle of step S402 are similar to those of "performing K convolutions on the fused feature map to obtain the first feature matrix of the fused feature map, and performing N convolutions on the fused feature map to obtain the second feature matrix of the fused feature map" in the foregoing step S401. The difference is only that the processed feature maps are different. Step S401 is the processing of the fused feature map, while step S402 is the processing of the feature map of the query image.
[0108] It should be noted that the present application embodiment does not limit the execution order of step S401 and step S402, and they can also be executed in parallel.
[0109] S403. Aggregate the first feature matrix of the feature map of the query image and the first feature matrix of the fused feature map to obtain a weight matrix, where the weight matrix is used to illustrate the weight values of the respective features in the query image. The higher the similarity between the features in the query image and the features in the fused feature map, the greater the weight value of the feature in the query image.
[0110] Specifically, the first feature matrix of the feature map of the query image in step S403 can be understood as the features of the Key part of the query image, and the first feature matrix of the fused feature map can be understood as the features of the Key part of the fused feature map. Step S403 can be understood as aggregating the features of the Key part of the query image and the features of the Key part of the fused feature map, thereby obtaining the weight matrix.
[0111] Aggregating the first feature matrix of the feature map of the query image and the first feature matrix of the fused feature map can be understood as measuring the similarity between the features of the Key part of the query image and the features of the Key part of the fused feature map. The higher the similarity, the greater the weight value of the corresponding feature in the weight matrix; the lower the similarity, the smaller the weight of the corresponding feature in the weight matrix. Thus, when using the weight matrix in subsequent step S404, it can enhance the features in the query image that are similar to the support image and suppress the features in the query image that are not similar to the support image.
[0112] Optionally, after executing step S403, it further includes: normalizing the weight matrix using the Softmax function to obtain a normalized weight matrix, and substituting the normalized weight matrix as the new weight matrix into the execution of step S404.
[0113] S404. Multiply the weight matrix and the second feature matrix of the feature map of the query image to obtain the processed feature map of the query image.
[0114] In the weight matrix, the higher the similarity between the features in the query image and the features in the fused feature map, the greater the weight value of the features in the query image. Because in the processed feature map of the query image obtained by multiplying the weight matrix and the second feature matrix of the feature map of the query image, the features in the query image that are similar to the fused feature map will become more obvious, while the features that are not similar to the fused feature map will be suppressed. Compared with the feature map of the query image before processing, the processed feature map of the query image can more clearly show the features similar to the fused feature map. The features similar to the fused feature map are the features in the area where the target is located, which can also be understood as the features related to the target to be detected in the query image. Those features that are not similar to the fused feature map are the features in the area of the query image where the target is not located, such as the features of the background area, or the features of the targets other than the target to be detected (i.e., the targets that do not belong to the target categories included in the public dataset).
[0115] S405. Aggregate the processed feature map of the query image and the second feature matrix of the fused feature map to obtain an aggregated feature map.
[0116] Aggregating the processed feature map of the query image and the second feature matrix of the fused feature map can be understood as calculating the similarity between the features of the processed feature map of the query image and the features in the second feature matrix of the fused feature map, so as to enhance the features in the query image that are similar to the fused feature map again and suppress the features that are not similar to the fused feature map again, further improving the effect of feature enhancement aggregation. The finally obtained aggregated feature map can be understood as the feature map of the query image after the process shown in Figure 4 the process flow shown.
[0117] It should be noted that both the first feature matrix and the second feature matrix of the fused feature map can be understood as matrices for describing the features of multiple support images corresponding to the query image. Similarly, both the first feature matrix and the second feature matrix of the feature map of the query image can be understood as matrices for describing the features of the query image.
[0118] S204. Input the aggregated feature map into the candidate box generation network to generate multiple candidate boxes of the aggregated feature map.
[0119] Among them, the candidate box is the predicted area where the target is located. Each candidate box of the aggregated feature map can correspond to an area on the query image. The corresponding area on the query image is the position area of the target predicted by the candidate box generation network on the query image.
[0120] Since in the aggregated feature map obtained in the foregoing step S203, the similar features between the query image and the support image are enhanced and the dissimilar features are made consistent, the candidate box generation network can predict the position where the candidate box is located through the relatively obvious similar features enhanced in the aggregated feature map.
[0121] Optionally, in a specific embodiment of the present application, after performing step S204, it further includes:
[0122] According to the reference boxes of multiple support images corresponding to the query image, filter out multiple candidate boxes belonging to negative samples from all candidate boxes of the aggregated feature map, where the reference box of the support image is the region where the target is located in the pre-annotated support image, and the candidate box belonging to the negative sample is the candidate box in the non-target region. Step S205 includes: inputting each candidate box of the filtered aggregated feature map into the fully connected layer to obtain the predicted target detection result of the query image.
[0123] Specifically, during the process of training the initial detection model, when the accuracy of the multiple candidate boxes obtained in step S204 is not very high, there will be many candidate boxes belonging to negative samples, and these candidate boxes belonging to negative samples are candidate boxes in non-target regions, which will lead to a very low accuracy of the predicted target detection result obtained in the subsequent step S205, and the parameters need to be continuously adjusted to train the model to gradually improve the accuracy of the model. To improve the efficiency of model training, the multiple candidate boxes obtained in step S204 can be filtered once to filter out some candidate boxes belonging to negative samples and avoid the situation where the number of candidate boxes belonging to negative samples is too large. Filtering out multiple candidate boxes belonging to negative samples can be understood as selecting and filtering out multiple candidate boxes of negative samples from all candidate boxes belonging to negative samples, and the selection method can be arbitrary selection or uniform distribution selection, and the embodiments of the present application do not limit this.
[0124] When there are too many candidate boxes belonging to negative samples, the computing power and the number of training times required for training the initial detection model will increase. After filtering out multiple candidate boxes belonging to negative samples from all candidate boxes of the aggregated feature map and only selecting each candidate box of the filtered aggregated feature map to perform the subsequent step S205, the number of training times can be reduced and the training effect can be improved. The filtered aggregated feature map has many fewer candidate boxes belonging to negative samples, which improves the training efficiency.
[0125] Optionally, refer to Figure 5 In a specific embodiment of the present application, an implementation manner of filtering out multiple candidate boxes belonging to negative samples from all candidate boxes of the aggregated feature map according to the reference boxes of multiple support images corresponding to the query image includes:
[0126] S501. Input the bounding boxes of multiple support images corresponding to the query image and all candidate boxes of the aggregated feature map into the metric learning network, and match to obtain the best-angle bounding box, where the best-angle bounding box is the bounding box among the bounding boxes of multiple support images corresponding to the query image that has the highest similarity with the candidate boxes of the aggregated feature map.
[0127] Specifically, after inputting the bounding boxes of multiple support images corresponding to the query image and all candidate boxes of the aggregated feature map into the metric learning network, the metric learning network starts to measure the similarity between the bounding boxes and the candidate boxes. This similarity can specifically be a calculated value such as the Euclidean distance that can be used to reflect the degree of feature similarity.
[0128] Among them, there can be one or more best-angle bounding boxes obtained by matching. The specific process can be that the metric learning network calculates the feature similarity between each candidate box of the aggregated feature map and the bounding boxes of each support image respectively, and then selects the bounding box with the highest feature similarity as the best bounding box, thereby obtaining multiple best bounding boxes. Or it can be to fuse all candidate boxes of the aggregated feature map to obtain a fused candidate box, then calculate the feature similarity between the fused candidate box and each bounding box respectively, and then use the bounding box with the highest feature similarity as the best bounding box.
[0129] After selecting the best-angle bounding box, then using the best-angle bounding box to perform subsequent steps can improve the accuracy and efficiency of filtering out candidate boxes belonging to negative samples.
[0130] S502. Perform pooling on the best-angle bounding box to obtain the depth feature vector of the best-angle bounding box.
[0131] Among them, the depth feature vector of the best-angle bounding box is used to describe the features in the best bounding box. There are many ways to perform pooling on the best-angle bounding box. For example, average pooling can be performed on the best-angle bounding box to obtain a 1×1 depth feature vector of the best-angle bounding box. The pooling method can be determined according to actual experience. For example, if it is found through multiple experiments that the average pooling method has a better effect on training the model, then the average pooling method is adopted.
[0132] It should be noted that the bounding boxes and candidate boxes mentioned in the embodiments of the present application can both be understood as features within a specific region.
[0133] Optionally, in another specific embodiment of the present application, step S501 may not be executed, but instead, directly perform pooling on the bounding boxes of multiple support images corresponding to the query image to obtain the depth feature vector of the bounding boxes, and then use the depth feature vector of the bounding boxes as the depth feature vector of the best-angle bounding box to perform subsequent steps S503 to S505.
[0134] S503. Convolve the depth feature vector of the optimal angle reference box with each candidate box of the aggregated feature map respectively to obtain an attention candidate box corresponding to each candidate box.
[0135] Among them, the attention candidate box corresponding to the candidate box can be understood as the candidate box after being processed.
[0136] S504. Input each attention candidate box corresponding to the candidate box into the binary classification model respectively to obtain the score of each attention candidate box, where the score of the attention candidate box is used to indicate whether the candidate box is predicted as the target category.
[0137] Specifically, the score of the attention candidate box can be used to indicate the probability that the candidate box is predicted as the target category. Exemplarily, the higher the score of the attention candidate box, the greater the probability that the target in the candidate box is predicted as the target category of this optimal reference box; if the score is lower, it indicates that the probability that the target in the candidate box is predicted as the target category of this optimal reference box is smaller.
[0138] S505. Filter out multiple candidate boxes belonging to negative samples from all candidate boxes of the aggregated feature map according to the score of each attention candidate box.
[0139] Exemplarily, a score threshold can be preset. If the score of the attention candidate box is less than this score threshold, it indicates that the candidate box corresponding to this attention candidate box is a candidate box belonging to negative samples. By this method, all candidate boxes belonging to negative samples can be found from all candidate boxes of the aggregated feature map, and then some candidate boxes belonging to negative samples are selected from all candidate boxes belonging to negative samples for filtering. The remaining candidate boxes that do not belong to negative samples and some candidate boxes that belong to negative samples are used to execute the subsequent step S205.
[0140] For Figure 5 the described process to be clearer, refer to Figure 2b , such as Figure 2b shown, extract multi-angle reference boxes (i.e., each reference box of the aforementioned support images) and candidate boxes (i.e., candidate boxes of the aggregated feature map) from the multi-angle feature map and input them into the metric learning network. The metric learning network matches the optimal angle reference box. Then use the optimal angle reference box to filter the candidate boxes so that some candidate boxes belonging to negative samples in the candidate boxes can be filtered out.
[0141] S205. Input each candidate box of the aggregated feature map into the fully connected layer to obtain the predicted target detection result of the query image, where the predicted target detection result of the query image is used to indicate the predicted target category of the query image and the region where the target is located in the query image.
[0142] In the aforementioned steps S201 to S204, by comparing the similarity between the query image and the multiple supporting images corresponding to the query image, multiple candidate boxes are finally confirmed. The aggregated feature map can be understood as the processed query image obtained after the aforementioned steps. Therefore, the candidate boxes of the aggregated feature map can confirm the area where the target in the query image is located. Therefore, after each candidate box of the aggregated feature map is input into the fully connected layer, the obtained predicted target detection result can predict the area where the target of the query image is located. Because the aforementioned steps compare the similarity between the query image and the multiple supporting images corresponding to the query image, the target category of the query image can be predicted based on the supporting images that have been pre-labeled with the target category and the target area.
[0143] Among them, the fully connected layer can be understood as a (Region of Interest, ROI) feature extractor. The fully connected layer may include one or more of a classification head, a positioning head, and a clustering head. Through the classification head, the positioning head, and the clustering head, the region where the target in the query image is located and the target category can be accurately confirmed.
[0144] For example, see Figure 6 If the fully connected layer includes a classification head and a positioning head, an implementation of step S205 includes:
[0145] S601. Pool each candidate box of the aggregated feature map to obtain a feature vector for each candidate box.
[0146] Specifically, for each candidate box, the candidate box can be input into the ROI pooling layer to obtain the feature vector of the candidate box. After pooling, the feature dimension is reduced, reducing the amount of calculation in subsequent steps.
[0147] S602: For each candidate box, the feature vector of the candidate box is processed by the classification head of the fully connected layer to obtain a predicted category result of the candidate box.
[0148] The predicted category result of the candidate box can be understood as the result of the target category to which the target in the candidate box belongs predicted by the classification head. The predicted category result of the candidate box can be in the form of a matrix, which illustrates the probability of the candidate box belonging to each target category. For example, when the small sample target detection method based on multi-angle information proposed in the embodiment of the present application is applied to an electric power scene, when the equipment fails, the image of the monitored electric power scene will have target categories such as hanging foreign objects, broken wires, and self-explosion of insulators. Therefore, the predicted category result of the candidate box includes the probability that the candidate box belongs to the target category of hanging foreign objects, the probability of belonging to the target category of broken wires, and the probability of belonging to the target category of self-explosion of insulators.
[0149] S603. Process the feature vector of the candidate box through the localization head of the fully connected layer to obtain the corresponding region of the candidate box in the query image, where the predicted object detection result of the query image includes: the predicted class result of each candidate box and the corresponding region of each candidate box in the query image.
[0150] Specifically, after the localization head processes the feature vector of the candidate box, the offset of the position of the candidate box in the aggregated feature map compared to its corresponding position in the query image can be calculated according to the position of the candidate box in the aggregated feature map, and then the corresponding region of the candidate box in the query image can be determined. The offset specifically may include the offset of the center point of the candidate box and the offsets of the length and width.
[0151] Among them, the order between performing step S602 and step S603 is not limited in the embodiments of the present application, and they may also be executed in parallel.
[0152] Optionally, after performing step S603, it may further include:
[0153] For each candidate box, process the feature vector of the candidate box through the clustering head of the fully connected layer to obtain a measure of the adjusted similar features. The clustering head is used to adjust the feature metric value in the aggregated feature map to make the similar feature part between the query image and the fused feature map closer.
[0154] Specifically, for example, as Figure 2b shown, after the filtered candidate boxes are input into the ROI pooling layer, the feature vector of each candidate box is obtained, and then the feature vector of the candidate box passes through the classification head, regression head (i.e., the aforementioned localization head) and clustering head on the ROI feature extractor to obtain the predicted object detection result of the query image.
[0155] S206. Adjust the parameters in the initial detection model according to the error between the predicted object detection result and the actual object detection result until the error between the predicted object detection result output by the adjusted initial detection model and the actual object detection result meets the preset convergence condition, and then determine the adjusted initial detection model as the target detection model.
[0156] The actual object detection result refers to the actual object category and the actual area where the object is located in the query image. During the training process of the initial detection model, there is an error between the predicted object detection result finally obtained by the initial detection model in the foregoing step S205 and the actual object detection result. By referring to the error between the predicted object detection result and the actual object detection result, the parameters in the initial detection model can be adjusted, and then the adjusted initial detection model is used to execute steps S202 to S206 again until the error between the predicted object detection result output by the adjusted initial detection model and the actual object detection result meets the preset convergence condition. It is considered that the error of the adjusted initial detection model is small enough to meet the accuracy requirement of object detection. Therefore, the adjusted initial detection model is determined as the object detection model. The determined object detection model is then applied to Figure 1 the step S102 shown in
[0157] Optionally, referring to Figure 7 , if the processes of steps S601 to S603 are executed before step S206, an implementation manner of adjusting the parameters in the initial detection model according to the error between the predicted object detection result and the actual object detection result includes:
[0158] S701. Calculate the classification loss value according to the error between the predicted category result of the candidate box and the actual category result of the candidate box, and calculate the localization loss value according to the error between the corresponding area of the candidate box in the query image and the actual corresponding area of the candidate box in the query image.
[0159] Among them, the classification loss value is used to illustrate the error of the classification head, and the localization loss value is used to illustrate the error of the localization head. Exemplarily, substitute the predicted category result of the candidate box into the classification loss function to calculate the classification loss value, and substitute the corresponding area of the candidate box predicted by the foregoing localization head in the query image into the localization loss function to obtain the localization loss value.
[0160] Optionally, if the fully connected layer further includes a clustering head, the loss of the clustering head can also be calculated to obtain the clustering loss value.
[0161] S702. Calculate the loss value of the initial detection model according to the classification loss value and the localization loss value.
[0162] Exemplarily, the classification loss value and the localization loss value can be summed to obtain the loss value of the initial detection model. The classification loss value and the localization loss value can also be weighted and summed to obtain the loss value of the initial detection model. The loss value of the initial detection model is used to illustrate the error of the initial detection model.
[0163] S703. Adjust the parameters in the initial detection model according to the loss value of the initial detection model.
[0164] In the small-sample object detection method based on multi-angle information provided by the embodiments of the present invention, by obtaining the image to be detected, and then inputting the image to be detected and multiple support images in the support set into the object detection model, the object detection model obtains and outputs the predicted object detection result of the image to be detected. Since the object detection model is trained by multiple query images and multiple support images corresponding to the query images, and the multiple support images corresponding to the query images include support images at multiple different shooting angles with the same object category as the query image, the relationship distillation network in the initial detection model can process and obtain an aggregated feature map according to the feature map of the query image and all support images corresponding to the query image, where the aggregated feature map is used to illustrate the similarity between the feature map of the query image and the feature maps of all support images. Therefore, the aggregated feature map aggregates the features of support images at multiple different shooting angles and the query image, which can make the features similar to those in the support images at multiple different shooting angles in the query image more obvious. Furthermore, the accuracy of multiple candidate boxes of the aggregated feature map generated by the candidate box generation network in the initial detection model will be relatively high, and the accuracy of the predicted object detection result of the query image obtained by the fully connected layer according to each candidate box of the aggregated feature map will also be improved. Therefore, even in the case of small samples, the accuracy of the object detection model trained by the initial detection model is still very high. Furthermore, after inputting the image to be detected into the object detection model, a predicted object detection result with high accuracy for the image to be detected can be obtained.
[0165] Refer to Figure 8 , based on the small-sample object detection method based on multi-angle information proposed in the above embodiments of the present application, the embodiments of the present application correspondingly disclose a small-sample object detection device based on multi-angle information, including: an acquisition unit 801 and an output unit 802.
[0166] The acquisition unit 801 is configured to acquire the image to be detected.
[0167] The output unit 802 is configured to input the image to be detected and multiple support images in the support set into the object detection model, and the object detection model obtains and outputs the predicted object detection result of the image to be detected.
[0168] Among them, the predicted object detection result of the image to be detected is used to illustrate the object category of the predicted image to be detected and the region where the object is located in the query image. The support set includes: support images of multiple object categories at multiple different shooting angles. The object detection model is obtained by training an initial detection model with multiple query images and the multiple support images corresponding to the query images. The multiple support images corresponding to the query image include support images of multiple different shooting angles with the same object category as the query image. The initial detection model includes: a Siamese network, a relationship distillation network, a candidate box generation network, and a fully connected layer. The Siamese network is used to process the query image and the multiple support images corresponding to the query image to obtain the feature map of the query image and the feature map of each support image. The relationship distillation network is used to process the feature map of the query image and all the support images corresponding to the query image to obtain an aggregated feature map. The aggregated feature map is used to illustrate the similarity between the feature map of the query image and the feature maps of all the support images corresponding to the query image. The candidate box generation network is used to generate multiple candidate boxes for the aggregated feature map. The candidate box is the region where the predicted object is located. The fully connected layer is used to obtain the predicted object detection result of the query image according to each candidate box of the aggregated feature map.
[0169] Optionally, in a specific embodiment of the present application, the few-shot object detection device based on multi-angle information further includes: a construction unit, a feature extraction unit, an aggregation unit, a generation unit, a prediction unit, and an adjustment unit.
[0170] The construction unit is used to construct an initial detection model.
[0171] The feature extraction unit is used to input the query image into the first feature extraction network of the Siamese network to obtain the feature map of the query image. And input the multiple support images corresponding to the query image into the second feature extraction network of the Siamese network to obtain the feature map of each support image.
[0172] The aggregation unit is used to input the feature map of the query image and the feature maps of all the support images into the relationship distillation network to obtain an aggregated feature map. Among them, the aggregated feature map is used to illustrate the similarity between the feature map of the query image and the feature maps of all the support images.
[0173] Optionally, the aggregation unit includes: a fusion subunit and a first aggregation subunit.
[0174] The fusion subunit is used to fuse the feature maps of all the support images to obtain a fused feature map.
[0175] The first aggregation subunit is used to input the feature map of the query image and the fused feature map into the relationship distillation network to obtain an aggregated feature map.
[0176] Optionally, the aggregation subunit includes: a first convolutional subunit, a second convolutional subunit, a weight calculation subunit, a processing subunit, and a second aggregation subunit.
[0177] The first convolutional subunit is configured to input the feature map of the query image and the fused feature map into the relational distillation network. The relational distillation network performs K convolutions on the fused feature map to obtain a first feature matrix of the fused feature map, and performs N convolutions on the fused feature map to obtain a second feature matrix of the fused feature map. Here, N and K are two different positive integers.
[0178] The second convolutional subunit is configured to perform K convolutions on the feature map of the query image to obtain a first feature matrix of the feature map of the query image, and perform N convolutions on the feature map of the query image to obtain a second feature matrix of the feature map of the query image.
[0179] The weight calculation subunit is configured to aggregate the first feature matrix of the feature map of the query image and the first feature matrix of the fused feature map to obtain a weight matrix. The weight matrix is used to illustrate the weight values of the respective features in the query image. The higher the similarity between the features in the query image and the features in the fused feature map, the greater the weight value of the features in the query image.
[0180] The processing subunit is configured to perform matrix multiplication on the weight matrix and the second feature matrix of the feature map of the query image to obtain a processed feature map of the query image.
[0181] The second aggregation subunit is configured to aggregate the processed feature map of the query image and the second feature matrix of the fused feature map to obtain an aggregated feature map.
[0182] The generating unit is configured to input the aggregated feature map into a candidate box generating network to generate a plurality of candidate boxes of the aggregated feature map.
[0183] The predicting unit is configured to input each candidate box of the aggregated feature map into a fully connected layer to obtain a predicted target detection result of the query image. The predicted target detection result of the query image is used to illustrate the predicted target category of the query image and the region where the target is located in the query image.
[0184] The adjusting unit is configured to adjust the parameters in the initial detection model according to the error between the predicted target detection result and the actual target detection result until the error between the predicted target detection result output by the adjusted initial detection model and the actual target detection result meets a preset convergence condition, and then determine the adjusted initial detection model as the target detection model.
[0185] Optionally, in a specific embodiment of the present application, the small-sample target detection device based on multi-angle information further includes:
[0186] A filtering unit, configured to filter out a plurality of candidate boxes belonging to negative samples from all candidate boxes of the aggregated feature map according to the reference boxes of a plurality of support images corresponding to the query image. Wherein, the reference box of the support image is the region where the target is located in the pre-annotated support image. The candidate box belonging to the negative sample is the candidate box in the region where the target is not located.
[0187] Wherein, when the prediction unit inputs each candidate box of the aggregated feature map into the fully connected layer to obtain the predicted object detection result of the query image, it is used for: inputting each candidate box of the filtered aggregated feature map into the fully connected layer to obtain the predicted object detection result of the query image.
[0188] Optionally, the filtering unit includes: a matching subunit, a first pooling subunit, a third convolutional subunit, a classification subunit, and a filtering subunit.
[0189] The matching subunit is configured to input the reference boxes of a plurality of support images corresponding to the query image and all candidate boxes of the aggregated feature map into the metric learning network to match and obtain the best-angle reference box. Wherein, the best-angle reference box is the reference box with the highest similarity between the reference boxes of a plurality of support images corresponding to the query image and the candidate boxes of the aggregated feature map.
[0190] The first pooling subunit is configured to pool the best-angle reference box to obtain the depth feature vector of the best-angle reference box.
[0191] The third convolutional subunit is configured to perform convolutional calculations on the depth feature vector of the best-angle reference box and each candidate box of the aggregated feature map respectively to obtain an attention candidate box corresponding to each candidate box.
[0192] The classification subunit is configured to input the attention candidate box corresponding to each candidate box into the binary classification model respectively to obtain the score of each attention candidate box. Wherein, the score of the attention candidate box is used to indicate whether the candidate box is predicted as the target category.
[0193] The filtering subunit is configured to filter out a plurality of candidate boxes belonging to negative samples from all candidate boxes of the aggregated feature map according to the score of each attention candidate box.
[0194] Optionally, the prediction unit includes:
[0195] The second pooling subunit is configured to pool each candidate box of the aggregated feature map to obtain the feature vector of each candidate box.
[0196] The first prediction subunit is configured to process the feature vector of each candidate box through the classification head of the fully connected layer for each candidate box to obtain the predicted class result of the candidate box.
[0197] A second prediction subunit, configured to process the feature vector of the candidate box through the localization head of the fully-connected layer to obtain the corresponding region of the candidate box in the query image. The predicted object detection result of the query image includes: the predicted class result of each candidate box and the corresponding region of each candidate box in the query image.
[0198] Wherein, when the adjustment unit performs adjustment on the parameters in the initial detection model according to the error between the predicted object detection result and the actual object detection result, it is used to: calculate the classification loss value according to the error between the predicted class result of the candidate box and the actual class result of the candidate box. And calculate the localization loss value according to the error between the corresponding region of the candidate box in the query image and the actual corresponding region of the candidate box in the query image. Calculate the loss value of the initial detection model according to the classification loss value and the localization loss value. Adjust the parameters in the initial detection model according to the loss value of the initial detection model.
[0199] The execution principles of each unit and subunit in the small-sample object detection device based on multi-angle information proposed in this application are the same as those of the small-sample object detection method based on multi-angle information mentioned above, and will not be elaborated here.
[0200] Based on the small-sample object detection device based on multi-angle information provided in the above embodiments of the present invention, the acquisition unit 801 acquires the image to be detected, and then the output unit 802 inputs the image to be detected and multiple support images of the support set into the object detection model, and the object detection model obtains and outputs the predicted object detection result of the image to be detected. Since the object detection model is trained by multiple query images and multiple support images corresponding to the query images, and the multiple support images corresponding to the query images include support images at multiple different shooting angles with the same object category as the query image, and the relationship distillation network in the initial detection model can process the feature map of the query image and all support images corresponding to the query image to obtain an aggregated feature map, where the aggregated feature map is used to illustrate the similarity between the feature map of the query image and the feature maps of all support images. Therefore, the aggregated feature map aggregates the features of the support images at multiple different shooting angles and the query image, which can make the features similar to those in the support images at multiple different shooting angles in the query image more obvious. Furthermore, the accuracy of the multiple candidate boxes of the aggregated feature map generated by the candidate box generation network in the initial detection model will be relatively high, and the accuracy of the predicted object detection result of the query image obtained by the fully-connected layer according to each candidate box of the aggregated feature map will also be improved. Therefore, even in the case of small samples, the accuracy of the object detection model trained by the initial detection model is still very high. Furthermore, after the output unit 802 inputs the image to be detected into the object detection model, a predicted object detection result with high accuracy of the image to be detected can be obtained.
[0201] The present application also discloses a computer-readable medium, on which a computer program is stored. When the program is executed by a processor, the method for small-sample object detection based on multi-angle information as described in any one of the above is implemented.
[0202] The present application also discloses a small-sample object detection device based on multi-angle information, including:
[0203] One or more processors; a storage device on which one or more programs are stored; when the one or more programs are executed by the one or more processors, the one or more processors implement the method for small-sample object detection based on multi-angle information as described in any one of the above.
[0204] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0205] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0206] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A small-sample object detection method based on multi-angle information, characterized in that Including: Obtain the image to be detected; Input the image to be detected and multiple support images of the support set into the target detection model, and the target detection model obtains and outputs the predicted target detection result of the image to be detected; Among them, the predicted target detection result of the image to be detected is used to illustrate the predicted target category of the image to be detected and the region where the target is located in the query image; the target detection model is obtained by training an initial detection model with multiple query images and multiple support images corresponding to the query images; the multiple support images corresponding to the query image include support images at multiple different shooting angles with the same target category as the query image; the initial detection model includes: a siamese network, a relation distillation network, a candidate box generation network, and a fully connected layer; the siamese network is used to process the query image and multiple support images corresponding to the query image to obtain the feature map of the query image and the feature map of each support image; the relation distillation network is used to process the feature map of the query image and all support images corresponding to the query image to obtain an aggregated feature map; the aggregated feature map is used to illustrate the similarity between the feature map of the query image and the feature maps of all support images corresponding to the query image; the candidate box generation network is used to generate multiple candidate boxes of the aggregated feature map; the candidate box is the region where the predicted target is located; the fully connected layer is used to obtain the predicted target detection result of the query image according to each candidate box of the aggregated feature map.
2. The method according to claim 1, wherein The construction process of the target detection model includes: Construct an initial detection model; Input the query image into the first feature extraction network of the siamese network to obtain the feature map of the query image; and input multiple support images corresponding to the query image into the second feature extraction network of the siamese network to obtain the feature map of each support image; Input the feature map of the query image and the feature maps of all support images into the relation distillation network to obtain an aggregated feature map; among them, the aggregated feature map is used to illustrate the similarity between the feature map of the query image and the feature maps of all support images; Input the aggregated feature map into the candidate box generation network to generate multiple candidate boxes of the aggregated feature map; Input each candidate box of the aggregated feature map into the fully connected layer to obtain the predicted target detection result of the query image; among them, the predicted target detection result of the query image is used to illustrate the predicted target category of the query image and the region where the target is located in the query image; According to the error between the predicted target detection result and the actual target detection result, adjust the parameters in the initial detection model until the error between the predicted target detection result output by the adjusted initial detection model and the actual target detection result meets the preset convergence condition, and determine the adjusted initial detection model as the target detection model.
3. The method according to claim 2, characterized in that, Inputting the feature map of the query image and the feature maps of all the support images into the relational distillation network to obtain an aggregated feature map includes: Fusing the feature maps of all the support images to obtain a fused feature map; Inputting the feature map of the query image and the fused feature map into the relational distillation network to obtain an aggregated feature map.
4. The method according to claim 3, wherein Inputting the feature map of the query image and the fused feature map into the relational distillation network to obtain an aggregated feature map includes: Inputting the feature map of the query image and the fused feature map into the relational distillation network. The relational distillation network performs K convolutions on the fused feature map to obtain a first feature matrix of the fused feature map; and performs N convolutions on the fused feature map to obtain a second feature matrix of the fused feature map; where N and K are two different positive integers; Performing K convolutions on the feature map of the query image to obtain a first feature matrix of the feature map of the query image; and performing N convolutions on the feature map of the query image to obtain a second feature matrix of the feature map of the query image; Aggregating the first feature matrix of the feature map of the query image and the first feature matrix of the fused feature map to obtain a weight matrix; where the weight matrix is used to illustrate the weight values of the features in the query image; the higher the similarity between the features in the query image and the features of the fused feature map, the greater the weight value of the features in the query image; Multiplying the weight matrix and the second feature matrix of the feature map of the query image to obtain a processed feature map of the query image; Aggregating the processed feature map of the query image and the second feature matrix of the fused feature map to obtain an aggregated feature map.
5. The method according to claim 2, characterized in that, Before inputting each candidate box of the aggregated feature map into a fully connected layer to obtain the predicted object detection result of the query image, it further includes: Filtering out a plurality of candidate boxes belonging to negative samples from all the candidate boxes of the aggregated feature map according to the ground truth boxes of the multiple support images corresponding to the query image; where the ground truth box of the support image is the pre-annotated area where the object is located in the support image; the candidate box belonging to the negative sample is the candidate box in the non-object location area; Inputting each candidate box of the aggregated feature map into a fully connected layer to obtain the predicted object detection result of the query image includes: Inputting each candidate box of the filtered aggregated feature map into a fully connected layer to obtain the predicted object detection result of the query image.
6. The method according to claim 5, characterized in that, Filtering out a plurality of candidate boxes belonging to negative samples from all the candidate boxes of the aggregated feature map according to the ground truth boxes of the multiple support images corresponding to the query image includes: Inputting the ground truth boxes of the multiple support images corresponding to the query image and all the candidate boxes of the aggregated feature map into a metric learning network to match and obtain an optimal angle ground truth box; where the optimal angle ground truth box is the ground truth box with the highest similarity between the ground truth boxes of the multiple support images corresponding to the query image and the candidate boxes of the aggregated feature map; Pool the optimal angle reference box to obtain the depth feature vector of the optimal angle reference box; Perform convolution calculations on the depth feature vector of the optimal angle reference box and each candidate box of the aggregated feature map respectively to obtain an attention candidate box corresponding to each candidate box; Input each attention candidate box corresponding to the candidate box into a binary classification model respectively to obtain the score of each attention candidate box; wherein, the score of the attention candidate box is used to indicate whether the candidate box is predicted as the target category; Filter out multiple candidate boxes belonging to negative samples from all candidate boxes of the aggregated feature map according to the score of each attention candidate box.
7. The method according to claim 2, characterized in that, The step of inputting each candidate box of the aggregated feature map into a fully connected layer to obtain the predicted target detection result of the query image includes: Pool each candidate box of the aggregated feature map to obtain the feature vector of each candidate box; For each candidate box, process the feature vector of the candidate box through the classification head of the fully connected layer to obtain the predicted class result of the candidate box; Process the feature vector of the candidate box through the localization head of the fully connected layer to obtain the corresponding region of the candidate box in the query image; wherein, the predicted target detection result of the query image includes: the predicted class result of each candidate box and the corresponding region of each candidate box in the query image; The step of adjusting the parameters in the initial detection model according to the error between the predicted target detection result and the actual target detection result includes: Calculate the classification loss value according to the error between the predicted class result of the candidate box and the actual class result of the candidate box; and calculate the localization loss value according to the error between the corresponding region of the candidate box in the query image and the actual corresponding region of the candidate box in the query image; Calculate the loss value of the initial detection model according to the classification loss value and the localization loss value; Adjust the parameters in the initial detection model according to the loss value of the initial detection model.
8. A small-sample object detection device based on multi-angle information, characterized in that, It includes: An acquisition unit, configured to acquire an image to be detected; An output unit, configured to input the image to be detected and multiple support images of the support set into a target detection model, and the target detection model obtains and outputs the predicted target detection result of the image to be detected; Among them, the predicted target detection result of the image to be detected is used to illustrate the target category of the predicted image to be detected and the region where the target is located in the query image; the support set includes: support images of multiple target categories at multiple different shooting angles; the target detection model is obtained by training an initial detection model with multiple query images and multiple support images corresponding to the query images; the multiple support images corresponding to the query image include support images of the same target category as the query image at multiple different shooting angles; the initial detection model includes: a siamese network, a relation distillation network, a candidate box generation network, and a fully connected layer; the siamese network is used to process the query image and the multiple support images corresponding to the query image to obtain the feature map of the query image and the feature map of each support image; the relation distillation network is used to process the feature map of the query image and all support images corresponding to the query image to obtain an aggregated feature map; the aggregated feature map is used to illustrate the similarity between the feature map of the query image and the feature maps of all support images corresponding to the query image; the candidate box generation network is used to generate multiple candidate boxes of the aggregated feature map; the candidate box is the region where the predicted target is located; the fully connected layer is used to obtain the predicted target detection result of the query image according to each candidate box of the aggregated feature map.
9. A computer-readable medium, characterized in that, A computer program is stored thereon, wherein when the program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.
10. A small-sample object detection device based on multi-angle information, characterized in that, Comprising: One or more processors; A storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
SAR Target Recognition Method Based on Incomplete Training Set of Twin Neural Networks
CN109508655A
Method and device for detecting part defects
CN110148130A