Small-Sample Object Detection Method and Model Based on Comparative Clustering and Template Matching

By adding a method of comparative clustering and template matching to the Faster-RCNN architecture, the problem of small sample target recognition is solved, and the recognition accuracy and adaptability of autonomous driving and autonomous navigation are improved, especially when embedded hardware resources are limited.

CN117036761BActive Publication Date: 2025-08-05UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310967550.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-02
Publication Date
2025-08-05
Estimated Expiration
2043-08-02

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify small sample targets in autonomous driving and autonomous navigation, and the embedded hardware resources are limited to support the secondary training and gradient propagation of the model, resulting in insufficient recognition accuracy, especially poor adaptability in cross-domain and multi-scene.

Method used

A small sample object detection method based on contrast clustering and template matching is adopted. Large sample type prototype vectors are trained through the Faster-RCNN architecture, combined with ROI pooling and contrast clustering modules, and a comparison loss function is added to form large sample type prototype vectors, and the recognition accuracy is improved by using small sample template matching and non-maximum suppression algorithms.

Benefits of technology

It realizes high-precision recognition of small sample targets in autonomous driving and autonomous navigation, enhances the recognition ability in complex environments, reduces false alarms, and improves the accuracy and adaptability of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117036761B_ABST
    Figure CN117036761B_ABST
Patent Text Reader

Abstract

A small sample target detection method based on contrast clustering and template matching includes the following steps: step 1. collecting large sample class images and inputting class prototype vector contrast clustering detection model for training; step 2. collecting small sample class images and inputting the large sample initial model obtained after training in step 1 to obtain a small sample class prototype vector group; step 3. inputting the image to be identified into the large sample initial model obtained in step 2, and outputting multiple target frames of the image to be identified; setting a confidence threshold for screening; step 4. screening the recognition results of the target frames that did not pass; step 5. merging the credible recognition results obtained in step 3 and the recognition results of the target frames that did not pass in step 4 as the recognition results of the image to be identified. The present invention can realize high-dimensional feature space mapping for large sample class targets, which is conducive to increasing the inter-class distance of the large sample class, reducing its intra-class distance, and enhancing the target recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of software and relates to image recognition technology, and specifically to a small sample target detection method and model based on contrast clustering and template matching. Background Art

[0002] Target recognition is an essential target perception method for applications such as autonomous driving and autonomous navigation. Relying on accurate target recognition technology, it can provide a good initial tracking template for target tracking, reducing the need for template preparation. Furthermore, in complex interference environments, target recognition can effectively recapture the target, effectively improving tracking accuracy and meeting the requirements of applications such as autonomous driving and autonomous navigation.

[0003] Traditional target recognition methods typically use a sliding window method based on manually designed features to classify target features within the sliding window, thereby identifying the target. Due to the variable scale of targets and the increasingly complex distribution of target features, traditional target recognition methods have relatively low recognition accuracy, making it difficult to meet the requirements for accurate target identification. With the development of deep learning-based target detection methods, intelligent target recognition is becoming an essential method for achieving intelligent target recognition in applications such as autonomous driving and autonomous navigation. Intelligent target recognition uses multi-layer convolutional neural networks to automatically extract salient target features in complex situations and predict the target's location and classification for accurate target recognition.

[0004] In practice, application scenarios are complex and changeable, and the characteristic distribution of targets shows huge differences. For offline trained neural network models, samples at a certain inference moment are difficult to obtain during the offline training stage of the model, or the number of samples obtained is relatively small, which cannot meet the needs of model training. However, the current intelligent target recognition method requires obtaining and training samples of small sample targets during the training stage in order to effectively identify them, which has many shortcomings.

[0005] First, small-sample target data for practical applications is often difficult to obtain during the training phase of intelligent algorithms. Even if it is available, it is difficult to meet the sample size required for training, making it impossible to recognize small-sample targets through training. Even if relatively large amounts of data can be obtained, embedded hardware in applications such as autonomous driving and autonomous navigation does not support secondary model training. This is because models running on embedded systems are usually in INT8 format due to speed requirements and cannot be well used for gradient propagation. Secondly, gradient propagation and parameter updates require huge static graph resources, which are difficult to effectively support given the limited resources of embedded systems.

[0006] Secondly, most data-driven small-sample target recognition algorithms often have good target recognition effects in scenarios similar to the training data set. However, their adaptability to general scenarios is seriously insufficient for cross-domain and multi-scenario usage.

[0007] Therefore, neither the traditional target recognition method nor the existing data-driven small-sample target recognition method can meet the target recognition accuracy requirements of applications such as autonomous driving and autonomous navigation. Summary of the Invention

[0008] The purpose of the present invention is to overcome the shortcomings of the existing technology and provide a small sample target detection method and model based on contrast clustering and template matching to solve the problem that the small sample target detection algorithm relies on small sample target data for training to improve performance, but it is difficult to obtain small sample targets in practice.

[0009] The small sample target detection method based on contrast clustering and template matching of the present invention is characterized by comprising the following steps:

[0010] Step 1. Collect a large sample of class images and input the class prototype vector comparison clustering detection model of the Faster-RCNN architecture for training;

[0011] The prototype vector comparison clustering model of the Faster-RCNN architecture includes a basic module, an FPN module, a binary classification prediction module, and a ROI pooling module connected in sequence, and also includes a comparison clustering module connected to the ROI pooling module;

[0012] The specific functions implemented by the comparative clustering module are:

[0013] The large sample class images are classified and stored in a vector group. The contrast loss is added to the training loss during each training process. After a set number of iterations, the large sample initial model and the large sample class prototype vector are obtained.

[0014] The formula of the contrast loss Lcontrastive is as follows:

[0015] Contrastive loss

[0016]

[0017] Where N is the total number of feature vectors for each training, z i ·z j is the cosine distance between features of different categories, and τ is the temperature coefficient;

[0018] Step 2. Collect small sample class images of the images to be identified and input them into the large sample initial model obtained after training in step 1. Then, perform inference through the ROI pooling module and the contrast clustering module to obtain the small sample class prototype vector group;

[0019] Step 3. Input the image to be identified into the detection model obtained in step 1, and output multiple target frames of the image to be identified. Each target frame corresponds to a target vector to be identified, a confidence score, and a target position. Target frames with low confidence scores are eliminated. The specific process is as follows:

[0020] S31. The detection model obtained by inputting the image to be identified in step 2 outputs multiple target frames of the image to be identified, and the target vector to be identified, the confidence level and the target position corresponding to each target frame;

[0021] S32. Set a first confidence threshold conf1. For target frames of the image to be identified whose confidence is higher than the first confidence threshold conf1, set a first overlap threshold IOU1 for the target frames. Use non-maximum suppression to filter out target frames with overlap higher than the overlap threshold IOU1, and proceed to step S33.

[0022] For the target frame of the image to be identified whose confidence is not higher than the first confidence threshold conf1, proceed to step S35;

[0023] S33. Sort all target frames whose confidence is higher than the first confidence threshold conf1 in descending order according to the confidence level, and calculate the overlap between the target frame corresponding to the maximum confidence and other target frames;

[0024] S34. Remove all associated target boxes whose overlap rate is greater than the first threshold;

[0025] For the remaining target frames, steps S33 and S34 are executed repeatedly until all target frames meet the requirements;

[0026] S35. For the target box whose confidence is lower than the first confidence threshold conf1 and fails, a second confidence threshold conf2 lower than the first confidence threshold is reset, and a second overlap threshold IOU2 higher than the first overlap threshold is set, and a non-maximum suppression operation is performed according to steps S33-S34;

[0027] All target frames and their corresponding categories retained after steps S32-S35 are regarded as reliable recognition results;

[0028] The target box eliminated by step 3 is input into step 4;

[0029] Step 4. Template matching, specifically:

[0030] Perform template matching on the target vectors to be identified corresponding to the target boxes eliminated in step 3 with each vector in the complete type vector group obtained in step 2, and calculate the matching degree respectively. The category of the prototype vector with the highest matching degree with the target box is taken as the category of the target box, and the template matching degree is taken as the confidence degree of the target box. The template matching degree can be represented by cosine similarity;

[0031] Set a third overlap threshold IOU3, perform non-maximum suppression on the target boxes eliminated in step 3 according to steps S33-S34, and use the target boxes and their categories remaining after non-maximum suppression in step 4 as the secondary recognition results output in step 4;

[0032] Step 5: Combine the reliable recognition result obtained in step 3 and the secondary recognition result obtained in step 4 as the recognition result of the image to be recognized.

[0033] Preferably, when comparing clusters in step 1, a class feature storage vector group Fstore=f1,f2…f c Store temporary feature vectors during training, C is the number of categories of large sample images, f1,f2…f c A distinct vector representing a vector group.

[0034] Preferably, in step 4, the template matching degree is represented by cosine similarity.

[0035] Preferably, the overlap IOU calculation formula is:

[0036] IOU=A∩B / A∪B

[0037] The numerator on the right side of the equal sign represents the area of the intersection of the two target box areas A and B, and the denominator represents the area of the union of the two target box areas A and B.

[0038] A small sample target detection model based on contrast clustering and template matching, characterized by including a basic module, an FPN module, a binary classification prediction module, a ROI pooling module, and a contrast clustering module connected to the ROI pooling module in sequence under the Faster-RCNN architecture;

[0039] The specific functions implemented by the comparative clustering module are:

[0040] The large sample class images are classified and stored in a vector group. The contrast loss is added to the training loss during each training process. After a set number of iterations, the large sample initial model and the large sample class prototype vector are obtained.

[0041] The formula for the contrast loss is as follows:

[0042] Contrastive loss

[0043]

[0044] Where N is the total number of feature vectors for each training, z i ·z j is the cosine distance between features of different categories, and τ is the temperature coefficient.

[0045] Preferably, the basic module adopts the resnet residual structure, the FPN module is a multi-scale pyramid structure, the binary classification prediction module outputs the candidate box target box prediction result and the candidate box classification prediction result to the ROI pooling module according to the feature map, and the ROI pooling module converts the ROI features of different sizes into features of the same size according to the prediction results.

[0046] During the training process of the target intelligent recognition model, the present invention adds a class prototype vector comparison clustering process to the Faster-RCNN two-stage target detection network to train large sample classes to obtain large sample class prototype vectors. The large sample class prototype vectors can realize high-dimensional feature space mapping for large sample category targets, which is beneficial to increasing the inter-class distance of large sample classes and reducing their intra-class distance.

[0047] The present invention uses small sample templates for template matching. Different from traditional template matching methods, it uses high-level features obtained through neural networks and uses non-maximum suppression to filter frames with high overlap rates to enhance the accuracy of target recognition, so as to achieve the accuracy requirements of small sample target recognition for applications such as autonomous driving and autonomous navigation. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 This is a flow chart of a specific implementation of the detection method of the present invention;

[0049] Figure 2 This is a flow chart of a specific implementation of the detection model of the present invention;

[0050] Figure 3 The figure is a schematic diagram of the detection effect of a specific embodiment of the present invention. DETAILED DESCRIPTION

[0051] The specific embodiments of the present invention are described in further detail below.

[0052] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely explained below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0053] like Figure 1 As shown, the small sample target detection method based on contrast clustering and template matching of the present invention includes the following steps:

[0054] Step 1. Collect a large sample of class images and input the class prototype vector comparison clustering detection model of the Faster-RCNN architecture for training.

[0055] In the prior art, the prototype vector contrast clustering model of the Faster-RCNN (Faster Region Convolutional Neural Networks) architecture includes a base module (BACK-BONE), an FPN module, a binary classification prediction module, and a ROI pooling module connected in sequence; in the present invention, a contrast clustering module is further added after the ROI pooling module.

[0056] The basic module uses the ResNet residual structure, and the FPN module (Feature Pyramid Networks) is a multi-scale pyramid structure. It is mainly used to obtain feature maps of different scales in the input image and input them into the binary classification prediction module. The binary classification prediction module outputs the candidate box target box prediction results and the candidate box classification prediction results based on the feature maps to the ROI pooling module. Based on the prediction results, the ROI pooling module converts ROI features (Region of Interest) of different sizes into feature outputs of the same size to ensure that they can be connected to the fully connected layer of the contrast clustering module to obtain a large sample prototype vector.

[0057] The specific functions implemented by the newly added comparative clustering module in the present invention are:

[0058] The large sample class images are classified and stored in a vector group. The contrast loss is added to the training loss during each training process. After the set number of iterations, the large sample initial model is obtained:

[0059] In contrastive clustering, let the number of categories of large sample class images be C, then the prototype vector of large sample class images can be represented as P = p1, p2…p c ,,Establish a class feature storage vector group Fstore=f1,f2…f c To store temporary feature vectors during training, the feature vectors of each class are stored in their corresponding locations. Contrastive loss is added to the training loss to reduce intra-class variance and increase inter-class variance. After NA iterations, the final large-sample class prototype vector and large-sample initial model are obtained, where NA is the set number of iterations.

[0060] The formula for contrast loss Lcontrastive is as follows:

[0061] Contrastive loss

[0062]

[0063] Where N is the total number of feature vectors for each training, z i ·z j is the cosine distance between different category features of the large sample prototype vector, and τ is the temperature coefficient. Different category features refer to the features extracted by the candidate box after ROI pooling and then passed through the large sample prototype vector after the fully connected layer.

[0064] By contrasting the clustering method, prototype vectors of different categories can be extracted more effectively, and at the same time, the extracted prototype vectors have a larger inter-class distance and a smaller intra-class distance in the feature space.

[0065] This step uses the two-stage processing idea of the Faster-RCNN network to extract the features of the target box candidate area, providing support for the formation of large-sample class prototype vectors. The ROI pooling module in the Faster-RCNN network can realize the pooling operation of the features of the target box candidate area with a uniform size, which is convenient for the progressive clustering of features. Therefore, a comparative clustering module is added afterwards to use large-sample class data for training to obtain a large-sample class prototype vector group and a large-sample initial model;

[0066] Step 2. Collect the small sample class images of the image to be identified and input them into the large sample initial model obtained after training in step 1. Accurately extract the features of the small sample target of the template image to form a small sample class prototype vector group.

[0067] The small sample class images of the image to be identified are generally several frames of images in a video containing the target to be identified, such as a new car.

[0068] The small sample class image is input into the large sample initial model, and the prototype vector group of the small sample class is obtained through reasoning through the ROI pooling module and the contrast clustering module;

[0069] Combine the small sample class prototype vector group obtained in this step with the large sample class prototype vector group obtained in step 1 to form a complete class prototype vector group;

[0070] This step can effectively realize the feature space mapping and representation of large sample classes, and can also quickly form the feature space mapping and representation of small sample classes. The feature space mapping and representation together form the class prototype vector space under effective representation, laying the foundation for subsequent template matching.

[0071] Step 3. Input the image to be identified into the detection model obtained in step 1, and output multiple target frames of the image to be identified. Each target frame corresponds to a target vector to be identified, a confidence score, and a target position. Target frames with low confidence scores are eliminated. The specific process is as follows:

[0072] S31. The detection model obtained by inputting the image to be identified in step 2 outputs multiple target frames of the image to be identified, and the target vector to be identified, the confidence level and the target position corresponding to each target frame;

[0073] S32. Set a first confidence threshold conf1. For target frames in the image to be identified with a confidence level higher than conf1, set a first overlap threshold IOU1 for the target frames. Use non-maximum suppression to filter out target frames with overlap higher than IOU1, and proceed to step S33.

[0074] For the target frame of the image to be identified whose confidence is not higher than the first confidence threshold conf1, proceed to step S35;

[0075] For example, if the target boxes are A, B, C...J, the ones above the first confidence threshold are A, B, D, E, H, and J. The non-maximum suppression operation is performed as follows:

[0076] S33. Sort all target boxes A, B, D, E, H, and J whose confidences are higher than the first confidence threshold conf1 in descending order, in the order of B, A, E, H, D, and J.

[0077] Set the first overlap rate threshold to 0.5, and calculate the overlap between the target box corresponding to the maximum confidence and other target boxes. For example, if target box B has the maximum confidence, calculate the overlap between target box B and other target boxes A, E, H, D, and J.

[0078] S34. Remove all associated target boxes whose overlap rate is greater than the first overlap rate threshold. For example, if it is found that the overlap rate of B and H is greater than the first overlap rate threshold, then remove the B and H target boxes, leaving A, E, D, and J.

[0079] For the remaining target frames, steps S33 and S34 are executed repeatedly until all target frames meet the requirements. The retained results are E and J, and A, B, D, and H are eliminated.

[0080] S35. For the failed target boxes C, F, and I whose confidence is lower than the first confidence threshold conf1, reset the second confidence threshold conf2 to 0.2, which is lower than the first confidence threshold, and set the second overlap threshold IOU2 to 0.6, which is higher than the first overlap threshold. Perform non-maximum suppression according to steps S33-S34.

[0081] For example, C and F were eliminated, and target box I was retained;

[0082] Non-maximum suppression can further improve the small sample target detection method. The non-maximum suppression operation can filter out those boxes with high overlap and enhance the accuracy of target detection.

[0083] Among them, the IOU calculation formula between the two target boxes A and B is:

[0084] IOU=A∩B / A∪B

[0085] The numerator on the right side of the equal sign represents the area of the intersection of the two target box areas A and B, and the denominator represents the area of the union of the two target box areas A and B.

[0086] After steps S32-S35, all target boxes E, I, J and their corresponding categories are retained as reliable recognition results.

[0087] The target boxes A, B, C, D, F, and H eliminated in step 3 are input into step 4;

[0088] Step 4. Template matching, specifically:

[0089] Perform template matching on the target vectors to be identified corresponding to the target boxes A, B, C, D, F, and H eliminated in step 3, one by one, with each vector in the complete type vector group obtained in step 2, and calculate the matching degree respectively. The category of the prototype vector with the highest matching degree with the target box is taken as the category of the target box, and the template matching degree is taken as the confidence degree of the target box; the template matching degree can be represented by cosine similarity.

[0090] The third overlap threshold IOU3 is set to 0.4. The third overlap threshold is usually between the first and second overlap thresholds.

[0091] The target boxes A, B, C, D, F, and H eliminated in step 3 are subjected to non-maximum suppression according to steps S33-S34, and the target boxes and their categories remaining after non-maximum suppression in step 4 are used as the secondary recognition results output by step 4.

[0092] For example, after the target boxes A, B, C, D, F, and H are subjected to non-maximum suppression according to steps S33-S34, A, B, F, and H are retained; then the A, B, F, and H target boxes and their corresponding categories are the secondary recognition results.

[0093] Step 5: Combine the reliable recognition result obtained in step 3 and the secondary recognition result obtained in step 4 as the recognition result of the image to be recognized.

[0094] In this embodiment, target frames A, B, E, F, H, I, and J and their corresponding categories are used as recognition results of the image to be recognized.

[0095] like Figure 3 As shown, the recognition results of the present invention are compared with those of manual recognition and the prior art DEFRCN (decoupled fast regional convolutional neural network) algorithm. It can be seen that the present invention has fewer false alarms and higher confidence levels than the mainstream method, and is closer to the manual recognition results.

[0096] The foregoing are the preferred embodiments of the present invention. Unless the preferred implementation modes in each preferred embodiment are obviously self-contradictory or based on a certain preferred implementation mode, each preferred implementation mode can be arbitrarily superimposed and used in combination. The embodiments and the specific parameters in the embodiments are only for the purpose of clearly describing the inventor's invention verification process, and are not intended to limit the patent protection scope of the present invention. The patent protection scope of the present invention shall still be based on its claims. Any equivalent structural changes made using the contents of the description and drawings of the present invention should also be included in the protection scope of the present invention.

Claims

1. A small sample target detection method based on contrast clustering and template matching, characterized by , including the following steps: Step 1. Collect a large sample of class images and input them into the Faster-RCNN architecture class prototype vector comparison clustering detection model for training; The prototype vector comparison clustering model of the Faster-RCNN architecture includes a basic module, an FPN module, a binary classification prediction module, and a ROI pooling module connected in sequence, and also includes a comparison clustering module connected to the ROI pooling module; The specific functions implemented by the comparative clustering module are: The large sample class images are classified and stored in a vector group. The contrast loss is added to the training loss during each training process. After a set number of iterations, the large sample initial model and the large sample class prototype vector are obtained. The formula for the contrast loss is as follows: Contrastive loss Where N is the total number of feature vectors for each training, z i ·z j is the cosine distance between features of different categories, τ is the temperature coefficient; Step 2. Collect small sample class images of the images to be identified and input them into the large sample initial model obtained after training in step 1. Then, perform inference through the ROI pooling module and the contrast clustering module to obtain the small sample class prototype vector group; Step 3. Input the image to be identified into the detection model obtained in step 1, and output multiple target frames of the image to be identified. Each target frame corresponds to a target vector to be identified, a confidence score, and a target position. Target frames with low confidence scores are eliminated. The specific process is as follows: S31. The detection model obtained by inputting the image to be identified in step 2 outputs multiple target frames of the image to be identified, and the target vector to be identified, the confidence level and the target position corresponding to each target frame; S32. Set a first confidence threshold conf1. For target frames of the image to be identified whose confidence is higher than the first confidence threshold conf1, set a first overlap threshold IOU1 for the target frames. Use non-maximum suppression to filter out target frames with overlap higher than the overlap threshold IOU1, and proceed to step S33. For the target frame of the image to be identified whose confidence is not higher than the first confidence threshold conf1, proceed to step S35; S33. Sort all target frames whose confidence is higher than the first confidence threshold conf1 in descending order according to the confidence level, and calculate the overlap between the target frame corresponding to the maximum confidence and other target frames; S34. Remove all associated target boxes whose overlap rate is greater than the first threshold; For the remaining target frames, steps S33 and S34 are executed repeatedly until all target frames meet the requirements; S35. For the target box whose confidence is lower than the first confidence threshold conf1 and fails, a second confidence threshold conf2 lower than the first confidence threshold is reset, and a second overlap threshold IOU2 higher than the first overlap threshold is set, and a non-maximum suppression operation is performed according to steps S33-S34; All target frames and their corresponding categories retained after steps S32-S35 are regarded as reliable recognition results; The target box eliminated by step 3 is input into step 4; Step 4. Template matching, specifically: Perform template matching on the target vectors to be identified corresponding to the target boxes eliminated in step 3 with each vector in the complete type vector group obtained in step 2, and calculate the matching degree respectively. The category of the prototype vector with the highest matching degree with the target box is taken as the category of the target box, and the template matching degree is taken as the confidence of the target box. Template matching can be characterized by cosine similarity; Set a third overlap threshold IOU3, perform non-maximum suppression on the target boxes eliminated in step 3 according to steps S33-S34, and use the target boxes and their categories remaining after non-maximum suppression in step 4 as the secondary recognition results output in step 4; Step 5: Combine the reliable recognition result obtained in step 3 and the secondary recognition result obtained in step 4 as the recognition result of the image to be recognized.

2. The detection method according to claim 1, wherein When comparing clusters in step 1, a class feature storage vector group Fstore=f1,f2…f is established. c Store temporary feature vectors during training, C is the number of categories of large sample images, f1,f2…f c A distinct vector representing a vector group.

3. The detection method according to claim 1, wherein In step 4, the template matching degree is represented by cosine similarity.

4. The detection method according to claim 1, wherein The calculation formula for overlap IOU is: IOU=A∩B / A∪B The numerator on the right side of the equal sign represents the area of the intersection of the two target box areas A and B, and the denominator represents the area of the union of the two target box areas A and B.

5. A small sample target detection model based on contrast clustering and template matching, characterized in that: Used to perform the detection method according to any one of claims 1 to 4, comprising a basic module, an FPN module, a binary classification prediction module, a ROI pooling module, and a comparison clustering module connected to the ROI pooling module in sequence under the Faster-RCNN architecture; The specific functions implemented by the comparative clustering module are: The large sample class images are classified and stored in a vector group. The contrast loss is added to the training loss during each training process. After a set number of iterations, the large sample initial model and the large sample class prototype vector are obtained. The formula of the contrast loss Lcontrastive is as follows: Contrastive loss Where N is the total number of feature vectors for each training, z i ·z j is the cosine distance between features of different categories, and τ is the temperature coefficient.

6. The detection model according to claim 5, wherein: The basic module adopts the resnet residual structure, the FPN module is a multi-scale pyramid structure, the binary classification prediction module outputs the candidate box target box prediction results and the candidate box classification prediction results to the ROI pooling module according to the feature map, and the ROI pooling module converts ROI features of different sizes into features of the same size according to the prediction results.

Citation Information

Patent Citations

  • Road thrown object detection method based on contrast clustering self-learning

    CN114782891A

  • Small sample target detection method, system and device and storage medium

    CN115546470A