Target detection method and device based on visual large model
Through the object detection method based on the visual big model, adaptive reference frames are generated and feature similarity comparison is performed, and the robustness and adaptability problems of object detection under small sample training are solved, achieving efficient object detection effect.
Patent Information
- Application Number
- CN202510597932.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-12
AI Technical Summary
The prior art is difficult to train an effective object detection model in the case of small samples, resulting in low robustness, inability to adapt to fine-grained features of specific defects, and detection without category prospects.
The object detection method based on the visual big model is adopted, and the object detection is detected using a small number of labeled samples by generating an adaptive reference frame and performing feature similarity comparison, combined with non-maximum suppression processing.
Achieve simple and efficient target detection under small sample conditions, meets actual project requirements, and improves the accuracy and adaptability of the detection.
Smart Images

Figure CN120472143A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target detection technology, and in particular to a target detection method and device based on a large visual model. Background Art
[0002] In actual target detection tasks, there are often many long-tail tasks. The main characteristics of long-tail tasks are: there are not many images available for training at the project site, and only a very small number of labeled samples can be obtained (usually only a few to dozens of images). Traditional target detection methods, such as those based on RPN (Region Proposal Network), rely on large-scale labeled datasets for training. If based on traditional target detection methods, it is likely that a usable target detection model will not be well trained with such a small number of samples. For example, this may lead to problems such as low robustness of the target detection model, inability to adapt to the fine-grained features of specific defects, and inability to detect unseen category prospects.
[0003] To sum up, how to use small samples to detect target objects to meet the diverse needs of actual projects has become an urgent problem that needs to be solved. Summary of the Invention
[0004] The present invention provides a target detection method and device based on a large visual model, which is used to solve the defect in the prior art that a small amount of samples cannot be used to train a usable target detection model, and realize the detection of target objects using small samples.
[0005] The present invention provides a target detection method based on a large visual model, comprising:
[0006] Input the image to be detected into a preset visual basic model to obtain the image features output by the visual basic model;
[0007] Calculating the image features and a preset positive example first feature prototype vector to generate multiple adaptive reference frames;
[0008] Performing feature conversion processing on the reference frame so that the size of the reference frame is the same as the size of the preset annotation frame, performing candidate feature similarity comparison between the reference frame and a preset positive example second feature prototype vector, and obtaining a first similarity score between each reference frame and each category;
[0009] Performing non-maximum suppression processing on each reference frame for each category, so as to retain the most accurate reference frame for each category;
[0010] Based on a score threshold preset for each category, the reference frame whose similarity score exceeds the score threshold is used as the detection result of the corresponding category.
[0011] According to a method for target detection based on a large visual model provided by the present invention, the positive example first feature prototype vector and the positive example second feature prototype vector are generated by the following method:
[0012] Inputting a sample picture set with annotated frames into the visual basic model to obtain a feature map output by the visual basic model;
[0013] Mapping each of the annotation boxes to a corresponding area on the feature map;
[0014] Performing global average pooling on the feature points within the annotation box, and using the resulting vector as the first feature prototype vector of the positive example; aligning the regional features corresponding to the annotation box through a feature region alignment module, and using the generated fixed-size feature tensor as the second feature prototype vector of the positive example;
[0015] The positive example first feature prototype vector and the positive example second feature prototype vector are clustered respectively to obtain the positive example first feature prototype vector with a dimension of [C, N, D] and the positive example second feature prototype vector with a dimension of [C, N, K*K, D], where C is the number of categories, N is the number of clusters, D is the length of the feature vector, and K is the target size for feature region alignment.
[0016] According to a method for object detection based on a large visual model provided by the present invention, the image features are calculated with a preset positive example first feature prototype vector to generate multiple adaptive reference frames, including:
[0017] Normalize the image features of dimension [D, H, W] and the first feature prototype vector of the positive example of dimension [C, N, D];
[0018] Performing a matrix multiplication operation on the normalized image features and the first feature prototype vector of the positive example to obtain an intermediate result with a dimension of [C, N, H, W];
[0019] Taking the maximum value along the second dimension of the intermediate result and taking the maximum value along the first dimension of the intermediate result to obtain a final result with a dimension of [1, 1, H, W], removing the dimensions with a value of 1 from the final result to obtain a response graph with a dimension of [H, W];
[0020] The T positions with the highest similarity are selected as candidate center points, and the reference frames of R aspect ratios and S scales are generated at each center point. The number of the reference frames is T*R*S, where T, R, and S are preset positive integers.
[0021] According to a method for object detection based on a large visual model provided by the present invention, feature conversion processing is performed on the reference frame so that the size of the reference frame is the same as the size of a preset annotation frame, and candidate feature similarity comparison is performed on the reference frame and a preset positive example second feature prototype vector to obtain a first similarity score between each reference frame and each category, including:
[0022] Aligning the reference frame by a feature region alignment module so that the size of the reference frame is the same as the size of the annotation frame, thereby obtaining the reference frame with dimensions of [T*R*S,N,K*K,D];
[0023] Perform cosine similarity calculation on the reference frame with a dimension of [T*R*S, N, K*K, D] and the second feature prototype vector of the positive example with a dimension of [C, N, K*K, D] to obtain a cosine similarity result matrix with a dimension of [T*R*S, C, N, K*K];
[0024] The cosine similarity result matrix is averaged in the fourth dimension and maximized in the third dimension to obtain a final similarity matrix with a dimension of [T*R*S, C]. The final similarity matrix represents the first similarity score between each reference box and each category.
[0025] According to a method for object detection based on a large visual model provided by the present invention, after obtaining a first similarity score between each reference frame and each category, and before performing non-maximum suppression processing on each reference frame for each category, the method further includes:
[0026] Performing a candidate feature similarity comparison between each reference frame and a preset negative feature prototype vector to obtain a second similarity score between each reference frame and each category;
[0027] When the first similarity score is less than the second similarity score, the reference frame is determined to be an invalid reference frame and is removed.
[0028] According to a target detection method based on a large visual model provided by the present invention, the negative example feature prototype vector is generated by the following method:
[0029] Inputting a sample picture set with negative example frames into the visual basic model to obtain a feature map output by the visual basic model;
[0030] Mapping each of the negative example boxes to a corresponding area on the feature map;
[0031] Aligning the regional features corresponding to the negative example frame through a feature region alignment module, and using the generated fixed-size feature tensor as the negative example feature prototype vector;
[0032] The negative example feature prototype vectors are clustered to obtain negative example feature prototype vectors with a dimension of [N, K, K, D].
[0033] According to a method for object detection based on a large visual model provided by the present invention, when the image features include image features of multiple scales, the image features are calculated with a preset positive example first feature prototype vector to generate multiple adaptive reference frames, including: for the image features of each scale, respectively calculating with the preset positive example first feature prototype vector to generate multiple adaptive reference frames of corresponding scales;
[0034] In the case where the image features include image features at multiple scales, the reference frame is compared with a preset positive example second feature prototype vector for candidate feature similarity to obtain a first similarity score between each reference frame and each category, including: at each scale, similarity calculation is performed between the reference frame at the corresponding scale and the positive example second feature prototype vector at the corresponding scale to obtain a first similarity score between each reference frame at the corresponding scale and each category at the corresponding scale.
[0035] The present invention also provides a target detection device based on a large visual model, comprising:
[0036] A feature extraction module is used to input the image to be detected into a preset visual basic model and obtain the image features output by the visual basic model;
[0037] A reference frame generation module, configured to calculate the image features and a preset positive example first feature prototype vector to generate a plurality of adaptive reference frames;
[0038] a feature comparison module, configured to perform feature conversion processing on the reference frame so that the size of the reference frame is the same as the size of a preset annotation frame, perform candidate feature similarity comparison on the reference frame and a preset positive example second feature prototype vector, and obtain a first similarity score between each reference frame and each category;
[0039] A post-processing module, configured to perform non-maximum suppression processing on each reference frame for each category, so as to retain the most accurate reference frame for each category;
[0040] The result output module is configured to, based on a score threshold preset for each category, take the reference frame whose similarity score exceeds the score threshold as the detection result of the corresponding category.
[0041] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the target detection method based on the large visual model as described above is implemented.
[0042] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for target detection based on a large visual model.
[0043] The target detection method and device based on the visual large model provided by the present invention obtain the image features output by the visual basic large model by inputting the image to be detected into the preset visual basic large model; the image features are calculated with the preset positive example first feature prototype vector to generate multiple adaptive reference frames; the reference frame is subjected to feature conversion processing so that the size of the reference frame is the same as the size of the preset annotation frame, and the reference frame is subjected to candidate feature similarity comparison with the preset positive example second feature prototype vector to obtain the first similarity score of each reference frame and each category; each reference frame is subjected to non-maximum suppression processing for each category so that the most accurate reference frame is retained for each category; based on the preset score threshold of each category, the reference frame whose similarity score exceeds the score threshold is used as the detection result of the corresponding category. Through the application of the above method and device, the approximate area of the target is obtained by adaptively generating the reference frame, which can achieve simple and efficient target detection with only a small number of labeled samples, which is more in line with the actual landing needs of the detection project. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0045] Figure 1 Schematic diagram of the process of target detection method based on large visual model provided by the present invention;
[0046] Figure 2 1 is a schematic diagram of an execution process of an optional target detection method based on a large visual model in an embodiment of the present invention;
[0047] Figure 3 It is a structural diagram of the target detection device based on the visual large model provided by the present invention;
[0048] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0049] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0050] The following combination Figure 1-Figure 2 The object detection method based on the visual large model of the present invention is described.
[0051] like Figure 1 As shown, the target detection method based on the visual large model provided by the present invention includes the following steps:
[0052] S1. Input the image to be detected into the preset visual basic model to obtain the image features output by the visual basic model.
[0053] The visual basic large model used in the present invention has a strong feature extraction capability. Before using the visual basic large model, it is necessary to load the pre-trained weight parameters first. The pre-training of the visual basic large model adopts the self-supervised pre-training training method, including but not limited to DINOV2, EVA02, etc. In an optional embodiment of the present invention, in order to continue to strengthen the feature extraction capability of the visual basic large model, it is also possible to continue to fine-tune other supervised tasks after the self-supervised pre-training is completed, and finally take out the weight parameters of the backbone network part. Since the present invention does not update the weight parameters of the visual basic large model during the entire process of target detection, the entire visual basic large model will be frozen after loading the weight parameters, so there is only inference cost but no training cost.
[0054] like Figure 2 As shown in the figure, the image features are obtained after passing through the backbone network. The dimension of the image features is [D, H, W], where D is the vector length, H and W are the height and width of the image features respectively.
[0055] S2. Calculate the image features and the preset positive example first feature prototype vector to generate multiple adaptive reference frames.
[0056] In an optional embodiment of the present invention, Figure 2 As shown, the preset positive example first feature prototype vector (positive example first prototypes feature) and the preset positive example second feature prototype vector (positive example second prototypes feature) are generated offline through the following steps:
[0057] S201: Input the sample image set with the labeled boxes into the visual basic model to obtain the feature map output by the visual basic model.
[0058] In an optional embodiment of the present invention, the sample picture set used is a small sample picture set, that is, the size of the picture set is small and the number of sample pictures included is very small.
[0059] S202: Map each annotation box to a corresponding area on the feature map.
[0060] The purpose of this step is to map the coordinates of each annotation box to the feature map, extract the features of the corresponding area, and locate the local features of the target in the feature map.
[0061] S203: Perform global average pooling on the feature points in the annotation box, and use the obtained vector as the first feature prototype vector of the positive example; align the regional features corresponding to the annotation box through the feature region alignment module, and use the fixed-size feature tensor generated by the feature region alignment module as the second feature prototype vector of the positive example.
[0062] Specifically, we perform global average pooling on the feature points within the annotated box, calculating the mean of each channel to obtain the first feature prototype vector of the positive example. This method captures the global semantic characteristics of the target while suppressing spatial details. The feature region alignment module (ROIAlign module) aligns the regional features corresponding to the annotated box into a fixed-size K*K grid and generates a feature tensor of the corresponding size. This aligned feature grid is used as the second feature prototype vector of the positive example, eliminating differences in regional shape and size to facilitate subsequent processing.
[0063] S204: Clustering the positive example first feature prototype vector and the positive example second feature prototype vector respectively.
[0064] Specifically, the first and second prototype vectors of the positive example are clustered separately. Assuming the number of clusters is preset to N, the resulting first prototype vector of the positive example has a dimension of [C, N, D] and the second prototype vector of the positive example has a dimension of [C, N, K*K, D], where C is the number of categories, D is the length of the feature vector, and K is the target size for feature region alignment (ROIAlign), which can usually be set to 3, 5, 7, etc. The clustered prototype vectors can be used as category representations for subsequent detection tasks.
[0065] Specifically, the image features are calculated with the preset positive example first feature prototype vector to generate multiple adaptive reference frames (anchor boxes), including the following steps:
[0066] S211: Normalize the image features with dimensions [D, H, W] and the first feature prototype vector of the positive example with dimensions [C, N, D].
[0067] Among them, normalizing the image features and the first feature prototype vector of the positive example (for example, L2 normalization) can ensure that the feature vectors are at the same scale and avoid deviations in subsequent calculations due to dimensional differences.
[0068] S212: Perform matrix multiplication on the normalized image features and the first feature prototype vector of the positive example to obtain an intermediate result with a dimension of [C, N, H, W].
[0069] S213: Take the maximum value along the second dimension of the intermediate result, and take the maximum value along the first dimension of the intermediate result to obtain a final result with a dimension of [1, 1, H, W]. Remove the dimensions with a value of 1 in the final result to obtain a response graph with a dimension of [H, W].
[0070] Specifically, we maximize the image along the second dimension (N): for each category C, we retain the highest response among the N most matching prototypes, resulting in [C, 1, H, W]. We also maximize the image along the first dimension (C): we take the highest response across categories, resulting in a final result of dimensions [1, 1, H, W]. After removing redundant dimensions, we obtain a response map of dimensions [H, W]. The response map marks the locations in the image that are most similar to any category prototype, with high-response areas likely indicating the center of the object.
[0071] S214: Select the T positions with the highest similarity as candidate center points, and generate reference frames with R aspect ratios and S scales at each center point. The number of reference frames is T*R*S, where T, R, and S are preset positive integers.
[0072] Specifically, the top T most responsive locations in the response map are selected as candidate center points. For each candidate center point, reference boxes with R aspect ratios (e.g., 1:1, 1:2, 2:1) and S scales (e.g., small / medium / large) are generated, totaling T×R×S reference boxes. Unlike traditional fixed reference boxes, this method dynamically determines candidate centers through feature matching and then generates reference boxes at multiple scales / ratios, making it more adaptable to complex scenarios and improving the match with real-world objects.
[0073] S3. Perform feature conversion processing on the reference frame, compare the candidate feature similarity between the reference frame and the preset positive example second feature prototype vector, and obtain a first similarity score between each reference frame and each category.
[0074] After completing the feature conversion of the reference frame, the reference frame is compared with the preset positive example second feature prototype vector for candidate feature similarity to obtain the first similarity score between each reference frame and each category, including the following steps:
[0075] S301: Align the reference frame through the feature region alignment module so that the size of the reference frame is the same as the size of the annotation frame, and obtain a reference frame with dimensions of [T*R*S,N,K*K,D].
[0076] Specifically, the reference frame is subjected to feature conversion processing (roi align processing), each reference frame is mapped to the feature map, divided into K*K sub-regions, and multiple points (such as 4) are sampled in each sub-region and the eigenvalues are calculated by interpolation. After processing, the feature dimension is still [T*R*S,N,K*K,D], but the features of each K*K grid are aligned with the size of the annotation frame in step S201. The purpose of the roialign processing is to make the size of the reference frame the same as the size of the preset annotation frame, that is, the same as the annotation frame size of the above-mentioned small sample image set with annotated frames, and add the features output by roialign to its average features along the third dimension (K*K) to enhance the feature expression capability of each spatial position, realize the fusion of local details and global context, and enhance the sensitivity to the target shape and position.
[0077] S302: Calculate the cosine similarity between the reference box with the dimension [T*R*S,N,K*K,D] and the second feature prototype vector of the positive example with the dimension [C,N,K*K,D] to obtain a cosine similarity result matrix with the dimension [T*R*S,C,N,K*K].
[0078] Specifically, for each reference box (T*R*S) and each category (C), the cosine similarity between it and the prototype vector at each point on the K*K grid is calculated, and the dimension [T*R*S,C,N,K*K] represents the similarity between each reference box and each category prototype on the K*K grid.
[0079] S303: Average the cosine similarity result matrix in the fourth dimension and maximize the third dimension to obtain a final similarity matrix of dimension [T*R*S,C]. The final similarity matrix represents the first similarity score between each reference box and each category.
[0080] Specifically, the cosine similarity matrix is averaged in the fourth dimension, that is, the similarity of each K*K grid is averaged to obtain [T*R*S,C,N], which represents the overall spatial similarity between each reference box and each category prototype. The maximum is taken in the third dimension, that is, the maximum similarity of the N prototypes of each category is taken to obtain the final matrix [T*R*S,C], which represents the highest similarity score between each reference box and each category, that is, the first similarity score.
[0081] S4. Perform non-maximum suppression on each reference box for each category.
[0082] Since the same object may be matched by multiple reference boxes (especially adjacent or overlapping reference boxes), resulting in a large number of duplicate prediction boxes being output, each reference box is subjected to non-maximum suppression (NMS) for each category. This aims to eliminate redundant prediction boxes for the same object and retain the most accurate prediction box for each category.
[0083] Specifically, in an optional embodiment of the present invention, based on the similarity matrix of dimension [T*R*S, C] obtained in step S3, for each category C, all prediction boxes are sorted by confidence, the box with the highest confidence is retained, and other boxes of the same category whose IoU (intersection over union) with the box exceeds a threshold (such as 0.5) are removed. NMS processing is performed independently for each category. For example, for the pedestrian category, only the prediction boxes marked as "pedestrian" are compared, and the most accurate box is retained; for the vehicle category, only the prediction boxes marked as "vehicle" are compared, and the most accurate box is retained. The strategy of classified NMS ensures that only the most accurate detection box is retained for each target, while avoiding mutual interference between different categories. It is a key step in target detection post-processing.
[0084] For example, suppose category A is cat and the threshold is 0.5. Get the scores of all reference boxes for "cat" from the score matrix (e.g., reference boxes 3, 5, and 10 are 0.9, 0.6, and 0.3, respectively). Keep the reference boxes with scores ≥ 0.5 (reference boxes 3 and 5), and discard the rest (reference box 10). Calculate the Intersection over Union (IoU) for the retained reference boxes. If IoU > 0.5, only keep the reference box with the highest score (e.g., if IoU between reference boxes 3 and 5 is 0.7, discard reference box 5). The final predicted box for category A is reference box 3.
[0085] S5. Based on a preset score threshold for each category, the reference frames whose similarity scores exceed the score threshold are taken as detection results of the corresponding category.
[0086] The reference frame remaining after processing in step S4 is used as the prediction frame result of the corresponding category.
[0087] In an optional embodiment of the present invention, in order to optimize the false detection problem, after obtaining the first similarity score between each reference frame and each category in step S3, and before performing non-maximum suppression on each reference frame for each category in step S4, a step of removing invalid reference frames is further included, as follows:
[0088] Step 1: Compare the candidate feature similarity of each reference box with the preset negative feature prototype vector to obtain the second similarity score of each reference box with each category.
[0089] Among them, the generation process of negative feature prototype vectors (negative prototypes) is as follows:
[0090] Input the sample picture set with negative example frames into the visual basic model to obtain the feature map output by the visual basic model. Map each negative example frame to the corresponding area on the feature map. Align the regional features corresponding to the negative example frame through the feature region alignment module (ROIAlign module), and use the generated fixed-size feature tensor as the negative example feature prototype vector. Cluster the negative example feature prototype vector to obtain a negative example feature prototype vector with a dimension of [N, K, K, D]. Among them, the number of clusters needs to be the same as the number of clusters used in step S204, and the dimension of the negative example feature prototype vector for each category is [N, K, K, D].
[0091] Specifically, the process of generating a negative feature prototype vector is similar to that of the positive second feature prototype vector described above. The difference between the two lies in the sample image set. The negative feature prototype vector is generated using a sample image set with negative boxes. The negative boxes are the prediction results based on target detection. The output prediction box is compared with the ground truth (GT) box to calculate the intersection over union (IoU). Prediction boxes with an IoU ratio less than a preset threshold are regarded as negative boxes. Negative boxes represent background or false positive areas and are used to train the model to distinguish non-target areas, thereby reducing the false detection rate.
[0092] Step 2: When the first similarity score is less than the second similarity score, the reference frame is determined to be an invalid reference frame and removed.
[0093] For example, taking category A as an example, the first similarity score between a reference frame and the second feature prototype vector of the positive example is 0.4, and the second similarity score between the reference frame and the feature prototype vector of the negative example is 0.6. Then the reference frame is judged as an invalid reference frame and removed to prevent it from participating in the processing of subsequent steps, thereby achieving the purpose of optimizing false detection.
[0094] In an optional embodiment of the present invention, if the scale range of the object to be detected varies greatly, multi-scale features may be applied. Specifically, if the image features include image features at multiple scales, step S2 calculates the image features against a preset positive example first feature prototype vector to generate multiple adaptive reference frames, including: for each scale of the image features, respectively, calculating against the preset positive example first feature prototype vector to generate multiple adaptive reference frames of the corresponding scale.
[0095] Specifically, in an optional embodiment of the present invention, three scale features are used, namely, first scale features (1 / 4 of the original image, which is conducive to detecting small targets), second scale features (1 / 8 of the original image, which is conducive to detecting medium targets), and third scale features (1 / 16 of the original image, which is conducive to detecting large targets). It should be understood that the use of multi-scale features also requires that when generating the positive first prototype features and the positive second prototype features, they are converted into multi-scale positive first prototype features and positive second prototype features. The steps are similar to the S201 to S204 process, except that for small-scale targets, the present invention only uses the first scale features as the positive first prototype features and the positive second prototype features; for medium-scale targets, the second scale features are used as the positive first prototype features and the positive second prototype features; for large-scale targets, the third scale features are used as the positive first prototype features and the positive second prototype features. After applying the multi-scale features as a whole, the model's prediction process continues as follows:
[0096] Similar to step S1, the image is input into the visual basic model to extract features. The difference is that features at multiple scales need to be extracted. Here, the visual basic model also needs to be a hierarchical network structure, such as SwinTransformer.
[0097] Similar to step S2, at the first scale feature, the second scale feature, and the third scale feature, the first prototype features of the positive example corresponding to each scale feature are calculated to obtain the candidate reference frame at each scale.
[0098] Similar to step S3, the reference frame is compared with the preset positive example second feature prototype vector for candidate feature similarity to obtain a first similarity score between each reference frame and each category. Specifically, at each scale, the reference frame at the corresponding scale is similar to the positive example second feature prototype vector at the corresponding scale to obtain a first similarity score between each reference frame at the corresponding scale and each category at the corresponding scale.
[0099] Similar to step S4, the reference frames at each scale are summarized, and then NMS processing and other operations are performed according to the process of step S4.
[0100] After obtaining the results of NMS processing, according to the score threshold of each category, the reference box whose score exceeds the threshold is output for each category, which is the predicted box result of the corresponding category.
[0101] In summary, the target detection method based on the visual large model provided by the present invention is to obtain the image features output by the visual basic large model by inputting the image to be detected into the preset visual basic large model; the image features are calculated with the preset positive example first feature prototype vector to generate multiple adaptive reference frames; the reference frame is subjected to feature conversion processing so that the size of the reference frame is the same as the size of the preset annotation frame, and the reference frame is subjected to candidate feature similarity comparison with the preset positive example second feature prototype vector to obtain the first similarity score of each reference frame and each category; each reference frame is subjected to non-maximum suppression processing for each category so that each category retains the most accurate reference frame; based on the preset score threshold of each category, the reference frame whose similarity score exceeds the score threshold is used as the detection result of the corresponding category. By applying the above method, target detection can be performed using a small sample of labeled pictures without pre-training. The approximate area of the target is obtained by adaptively generating the reference frame, and simple and efficient target detection can be achieved with only a small amount of labeled samples, which is more in line with the actual landing needs of the detection project. The present invention adopts an effective adaptive anchorbox generation method and a candidate feature similarity comparison method, and designs enhanced schemes for optimizing false detection and multi-scale problems, which can be applied to actual projects efficiently and conveniently.
[0102] Based on the same inventive concept, the present invention also provides a target detection device based on a visual large model. The target detection device based on a visual large model provided by the present invention is described below. The target detection device based on a visual large model described below and the target detection method based on a visual large model described above can be referenced to each other.
[0103] like Figure 3 As shown, the target detection device based on the visual large model provided by the present invention includes a feature extraction module 31, a reference frame generation module 32, a feature comparison module 33, a post-processing module 34 and a result output module 35.
[0104] The feature extraction module 31 is used to input the image to be detected into the preset visual basic model and obtain the image features output by the visual basic model;
[0105] A reference frame generation module 32 is used to calculate the image features and the preset positive example first feature prototype vector to generate multiple adaptive reference frames;
[0106] A feature comparison module 33 is configured to perform feature conversion processing on the reference frame so that the size of the reference frame is the same as the size of the preset annotation frame, and perform candidate feature similarity comparison between the reference frame and the preset positive example second feature prototype vector to obtain a first similarity score between each reference frame and each category;
[0107] A post-processing module 34 is configured to perform non-maximum suppression processing on each reference frame for each category, so as to retain the most accurate reference frame for each category;
[0108] The result output module 35 is configured to take, based on a preset score threshold for each category, reference frames whose similarity scores exceed the score threshold as detection results for the corresponding category.
[0109] In an optional embodiment of the present invention, the above-mentioned target detection device based on the visual large model also includes a feature prototype vector generation module for offline generating a positive example first feature prototype vector, a positive example second feature prototype vector and a negative example feature prototype vector.
[0110] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the object detection method based on the visual large model provided by the above methods, which includes:
[0111] Input the image to be detected into a preset visual basic model to obtain the image features output by the visual basic model;
[0112] Calculating the image features and a preset positive example first feature prototype vector to generate multiple adaptive reference frames;
[0113] Performing feature conversion processing on the reference frame so that the size of the reference frame is the same as the size of the preset annotation frame, performing candidate feature similarity comparison between the reference frame and a preset positive example second feature prototype vector, and obtaining a first similarity score between each reference frame and each category;
[0114] Performing non-maximum suppression processing on each reference frame for each category, so as to retain the most accurate reference frame for each category;
[0115] Based on a score threshold preset for each category, the reference frame whose similarity score exceeds the score threshold is used as the detection result of the corresponding category.
[0116] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0117] On the other hand, the present invention further provides a computer program product, comprising a computer program, which may be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the object detection method based on a large visual model provided by the above methods, which method comprises:
[0118] Input the image to be detected into a preset visual basic model to obtain the image features output by the visual basic model;
[0119] Calculating the image features and a preset positive example first feature prototype vector to generate multiple adaptive reference frames;
[0120] Performing feature conversion processing on the reference frame so that the size of the reference frame is the same as the size of the preset annotation frame, performing candidate feature similarity comparison between the reference frame and a preset positive example second feature prototype vector, and obtaining a first similarity score between each reference frame and each category;
[0121] Performing non-maximum suppression processing on each reference frame for each category, so as to retain the most accurate reference frame for each category;
[0122] Based on a score threshold preset for each category, the reference frame whose similarity score exceeds the score threshold is used as the detection result of the corresponding category.
[0123] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the object detection method based on a large visual model provided by the above methods, the method comprising:
[0124] Input the image to be detected into a preset visual basic model to obtain the image features output by the visual basic model;
[0125] Calculating the image features and a preset positive example first feature prototype vector to generate multiple adaptive reference frames;
[0126] Performing feature conversion processing on the reference frame so that the size of the reference frame is the same as the size of the preset annotation frame, performing candidate feature similarity comparison between the reference frame and a preset positive example second feature prototype vector, and obtaining a first similarity score between each reference frame and each category;
[0127] Performing non-maximum suppression processing on each reference frame for each category, so as to retain the most accurate reference frame for each category;
[0128] Based on a score threshold preset for each category, the reference frame whose similarity score exceeds the score threshold is used as the detection result of the corresponding category.
[0129] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0130] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A target detection method based on a large visual model, characterized in that: include: Input the image to be detected into a preset visual basic model to obtain the image features output by the visual basic model; Calculating the image features and a preset positive example first feature prototype vector to generate multiple adaptive reference frames; Performing feature conversion processing on the reference frame so that the size of the reference frame is the same as the size of the preset annotation frame, performing candidate feature similarity comparison between the reference frame and a preset positive example second feature prototype vector, and obtaining a first similarity score between each reference frame and each category; Performing non-maximum suppression processing on each reference frame for each category, so as to retain the most accurate reference frame for each category; Based on a score threshold preset for each category, the reference frame whose similarity score exceeds the score threshold is used as the detection result of the corresponding category.
2. The target detection method based on a large visual model according to claim 1, characterized in that: The positive example first feature prototype vector and the positive example second feature prototype vector are generated by the following method: Inputting a sample picture set with annotated frames into the visual basic model to obtain a feature map output by the visual basic model; Mapping each of the annotation boxes to a corresponding area on the feature map; Performing global average pooling processing on the feature points in the marked box, and using the obtained vector as the first feature prototype vector of the positive example; Aligning the regional features corresponding to the annotation box by a feature region alignment module, and using the generated fixed-size feature tensor as the second feature prototype vector of the positive example; The positive example first feature prototype vector and the positive example second feature prototype vector are clustered respectively to obtain the positive example first feature prototype vector with a dimension of [C, N, D] and the positive example second feature prototype vector with a dimension of [C, N, K*K, D], where C is the number of categories, N is the number of clusters, D is the length of the feature vector, and K is the target size for feature region alignment.
3. The target detection method based on a large visual model according to claim 2, characterized in that: The image features are calculated with a preset positive example first feature prototype vector to generate multiple adaptive reference frames, including: Normalize the image features of dimension [D, H, W] and the first feature prototype vector of the positive example of dimension [C, N, D]; Performing a matrix multiplication operation on the normalized image features and the first feature prototype vector of the positive example to obtain an intermediate result with a dimension of [C, N, H, W]; Taking the maximum value along the second dimension of the intermediate result and taking the maximum value along the first dimension of the intermediate result to obtain a final result with a dimension of [1, 1, H, W], removing the dimensions with a value of 1 from the final result to obtain a response graph with a dimension of [H, W]; The T positions with the highest similarity are selected as candidate center points, and the reference frames of R aspect ratios and S scales are generated at each center point. The number of the reference frames is T*R*S, where T, R, and S are preset positive integers.
4. The target detection method based on a large visual model according to claim 3, characterized in that: Performing feature conversion processing on the reference frame so that the size of the reference frame is the same as the size of the preset annotation frame, performing candidate feature similarity comparison on the reference frame and a preset positive example second feature prototype vector, and obtaining a first similarity score between each reference frame and each category, including: Aligning the reference frame by a feature region alignment module so that the size of the reference frame is the same as the size of the annotation frame, thereby obtaining the reference frame with dimensions of [T*R*S,N,K*K,D]; Perform cosine similarity calculation on the reference frame with a dimension of [T*R*S, N, K*K, D] and the second feature prototype vector of the positive example with a dimension of [C, N, K*K, D] to obtain a cosine similarity result matrix with a dimension of [T*R*S, C, N, K*K]; The cosine similarity result matrix is averaged in the fourth dimension and maximized in the third dimension to obtain a final similarity matrix with a dimension of [T*R*S, C]. The final similarity matrix represents the first similarity score between each reference box and each category.
5. The target detection method based on a large visual model according to claim 1, characterized in that: After obtaining the first similarity score between each reference frame and each category, and before performing non-maximum suppression processing on each reference frame for each category, the method further includes: Performing a candidate feature similarity comparison between each reference frame and a preset negative feature prototype vector to obtain a second similarity score between each reference frame and each category; When the first similarity score is less than the second similarity score, the reference frame is determined to be an invalid reference frame and is removed.
6. The target detection method based on a large visual model according to claim 5, characterized in that: The negative example feature prototype vector is generated by the following method: Inputting a sample picture set with negative example frames into the visual basic model to obtain a feature map output by the visual basic model; Mapping each of the negative example boxes to a corresponding area on the feature map; Aligning the regional features corresponding to the negative example frame through a feature region alignment module, and using the generated fixed-size feature tensor as the negative example feature prototype vector; The negative example feature prototype vectors are clustered to obtain negative example feature prototype vectors with a dimension of [N, K, K, D].
7. The method for target detection based on a large visual model according to any one of claims 1 to 6, characterized in that: When the image features include image features of multiple scales, the image features are calculated with a preset positive example first feature prototype vector to generate multiple adaptive reference frames, including: for each scale of the image features, respectively calculating with the preset positive example first feature prototype vector to generate multiple adaptive reference frames of corresponding scales; In the case where the image features include image features at multiple scales, the reference frame is compared with a preset positive example second feature prototype vector for candidate feature similarity to obtain a first similarity score between each reference frame and each category, including: at each scale, similarity calculation is performed between the reference frame at the corresponding scale and the positive example second feature prototype vector at the corresponding scale to obtain a first similarity score between each reference frame at the corresponding scale and each category at the corresponding scale.
8. A target detection device based on a large visual model, characterized in that: include: A feature extraction module is used to input the image to be detected into a preset visual basic model and obtain the image features output by the visual basic model; A reference frame generation module, configured to calculate the image features and a preset positive example first feature prototype vector to generate a plurality of adaptive reference frames; a feature comparison module, configured to perform feature conversion processing on the reference frame so that the size of the reference frame is the same as the size of a preset annotation frame, perform candidate feature similarity comparison on the reference frame and a preset positive example second feature prototype vector, and obtain a first similarity score between each reference frame and each category; A post-processing module, configured to perform non-maximum suppression processing on each reference frame for each category, so as to retain the most accurate reference frame for each category; The result output module is configured to, based on a score threshold preset for each category, take the reference frame whose similarity score exceeds the score threshold as the detection result of the corresponding category.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the target detection method based on the visual large model as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the target detection method based on a visual large model as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Target detection method, medium, equipment and product
CN122336448A
A target detection method, medium, device and product
CN122336448B