Class-independent object number estimation method
Through a category-independent object number estimation method, the density map, mask, semantic information and spatial clustering method are used to solve the over-prediction problem of the object detection method in the scenarios of few sample learning and no category counting, achieving higher counting accuracy and processing ability of occluded objects.
Patent Information
- Application Number
- CN202510011192.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-16
AI Technical Summary
Existing object detection methods are susceptible to target local response similarity in scenarios with few sample learning and no category counting, resulting in overprediction problems.
A class-independent object number estimation method is proposed, which estimates the number of objects by generating density maps, obtaining masks and high-resolution feature maps of candidate objects, extracting semantic information of examples and candidate objects, and filtering non-target objects using spatial clustering method.
Improve counting accuracy, enhance the processing ability of occlusion and self-similarity objects, generate density maps more accurate, and achieve higher accuracy in counting and detection tasks.
Smart Images

Figure CN120014222A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for estimating the number of category-independent objects, and belongs to the technical field of image recognition. Background Art
[0002] Unlike traditional target counting methods, category-free counting does not rely on pre-defined category labels. By learning the common features of the target, it achieves the ability to count new categories that have not been seen. Its advantage is that it can handle unlabeled or newly emerging target categories, has stronger generalization ability and lower labeling cost. Category-free counting is suitable for a variety of different target counting tasks, such as crowd counting, cell counting, animal counting, etc.
[0003] The category-free counting method achieves generalization ability between different categories by learning the common features of the target. In addition, the prompt-based image segmentation model (such as Segment Anything Model, SAM) has demonstrated its potential in image segmentation tasks with its efficient segmentation ability and flexible prompt mechanism. Its application in the field of few-shot learning and category-free counting has important research value.
[0004] Segment Anything Model (SAM) is a prompt-based image segmentation model that can generate accurate segmentation masks from simple point, box or text prompts. Its efficient segmentation ability and flexible prompt mechanism enable it to excel in various image segmentation tasks. However, the application of SAM in few-shot learning and category-free counting scenarios is still in the exploratory stage. Combining SAM with few-shot learning and category-free counting methods is expected to significantly improve the performance of the model in practical applications by leveraging its powerful segmentation ability and the generalization ability of few-shot learning. This counting method guides the model to focus on the specified category through the similarity of example features and image features, but is easily affected by the similarity of local responses of the target, resulting in over-prediction.
[0005] Therefore, it is necessary to conduct more in-depth research on existing target detection methods to solve the above problems. Summary of the invention
[0006] In order to overcome the above problems, in-depth research was conducted and a category-independent object number estimation method was proposed, which includes the following steps:
[0007] S1, generate a density map based on the original image;
[0008] S2, taking the peak point in the density map as the position of the candidate object, and obtaining the mask and high-resolution feature map of the candidate object;
[0009] S3, extracting semantic information of examples and candidate objects based on the acquired masks and high-resolution feature maps;
[0010] S4. Filter non-target objects based on the spatial clustering method to obtain the number of objects in the original image.
[0011] In a preferred embodiment, S1 comprises the following sub-steps:
[0012] S11, extract image features from the original image through convolutional network;
[0013] S12, cutting out example features and negative example features from image features;
[0014] S13, the negative example feature is used as a convolution kernel to convolve with the image feature to generate an attention map;
[0015] S14. Connect the enhanced features and the attention map to obtain the final feature map, which is passed to the regression head of the convolutional network to obtain the density map.
[0016] In a preferred embodiment, in S12, the negative examples are obtained by randomly cutting a portion of each example.
[0017] In a preferred embodiment, the loss function of the negative example is set to the sum of the mean square error loss of the negative example feature and the L1 norm of the density map.
[0018] In a preferred embodiment, the loss function of the example is set to the sum of the mean square error loss of the attention map and the mean square error loss of the predicted density map.
[0019] In a preferred embodiment, in S3, a mask of the candidate object is obtained by a hint-based image segmentation model.
[0020] In a preferred embodiment, in S2, the high-resolution feature map is obtained by bilinear interpolation transformation of image features;
[0021] The image features are extracted based on an image encoder in a prompted image segmentation model.
[0022] In a preferred embodiment, in S3, the semantic information of the example is expressed as:
[0023]
[0024] in, Example b i The semantic information of F h represents a high-resolution feature map, a mask representing an example;
[0025] The semantic information of the candidate object is expressed as:
[0026]
[0027] in, Represents candidate objects The semantic information of F h represents a high-resolution feature map, Masks representing candidate objects.
[0028] The present invention also provides an electronic device, comprising:
[0029] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform any of the above methods.
[0030] The present invention also provides a computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute any of the above-mentioned methods.
[0031] The beneficial effects of the present invention include:
[0032] (1) It not only improves the counting accuracy, but also enhances the model's ability to handle occlusion and self-similar objects;
[0033] (2) The generated density maps are more accurate and more interpretable;
[0034] (3) Higher accuracy can be achieved in counting and detection tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 The figure is a flow chart of a method for estimating the number of category-independent objects according to a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0036] The present invention will be further described in detail below through the accompanying drawings and embodiments. Through these descriptions, the characteristics and advantages of the present invention will become more clear and distinct.
[0037] The word "exemplary" is used exclusively herein to mean "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise noted.
[0038] According to a method for estimating the number of objects independent of category provided by the present invention, Figure 1 As shown, the following steps are included:
[0039] S1, generate a density map based on the original image;
[0040] S2, taking the peak point in the density map as the position of the candidate object, and obtaining the mask and high-resolution feature map of the candidate object;
[0041] S3, extracting semantic information of examples and candidate objects based on the acquired masks and high-resolution feature maps;
[0042] S4. Filter non-target objects based on the spatial clustering method to obtain the number of objects in the original image.
[0043] In S1, a local response suppression density map predictor is used to generate a density map, including the following sub-steps:
[0044] S11, extract image features from the original image through convolutional network;
[0045] S12, cutting out example features and negative example features from image features;
[0046] S13, the negative example feature is used as a convolution kernel to convolve with the image feature to generate an attention map;
[0047] S14. Connect the enhanced features and the attention map to obtain the final feature map, which is passed to the regression head of the convolutional network to obtain the density map.
[0048] In the present invention, the specific structure of the convolutional network is not limited, and those skilled in the art can freely choose according to actual needs, for example, ConvNeXT, ResNet, Swin Transformer, DenseNet, etc., preferably ConvNeXT. The specific structure of ConvNeXT can refer to Liu Z, Mao H, Wu CY, et al. A convnet for the 2020s [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022: 11976-11986, which will not be repeated in the present invention.
[0049] In S12, the example feature T is cropped from the image feature according to the example bounding box and the negative example bounding box. E and negative example features
[0050] Preferably, the negative examples are obtained by randomly cutting out a portion of each example, and the corresponding true density map is set to zero. This setting can force the model to focus on the overall information of the specified target and reduce the number of peak points predicted on the same object.
[0051] According to the present invention, in the training sample, the example bounding box and the negative example bounding box are marked, and in the recognition estimation, the example bounding box and the negative example bounding box are generated by a convolutional network.
[0052] In a preferred embodiment, the loss function of the negative example is set to the sum of the mean square error loss of the negative example feature and the L1 norm of the density map, expressed as:
[0053]
[0054] Among them, L neg represents the loss function of negative examples, λ1 is a hyperparameter, L MSE represents the mean square error loss, Denotes the negative example feature, D neg represents the density map and ||||1 represents the L1 norm.
[0055] In a preferred embodiment, the loss function of the example is set to the sum of the mean square error loss of the attention map and the mean square error loss of the prediction density map, expressed as:
[0056]
[0057] Among them, L pos represents the loss function of the example, λ2 is a hyperparameter, represents the attention map, D represents the true density map, and D pos Density plot representing the predictions.
[0058] The loss function L of the local response suppression density map predictor is the sum of the negative example loss function and the example loss function, expressed as:
[0059] L=L neg +L pos .
[0060] In S13, the attention map is used to enhance the image features. Preferably, the attention map is also globally averaged to enhance the generalization ability.
[0061] In the present invention, the attention map generated by convolving the negative example feature with the image feature as the convolution kernel and the density map obtained based on this can generate a unimodal density map for each target, which can effectively reduce the number of peak points predicted on the same object and improve the accuracy of target detection.
[0062] In S2, the peak point is a peak point of a local area in the density map, and there may be multiple peak points.
[0063] Furthermore, the mask of the candidate object is obtained by using a hint-based image segmentation model, preferably a SAM model (Segment Anything Model).
[0064] The high-resolution feature map is obtained by bilinear interpolation transformation of image features extracted by an image encoder in a hint-based image segmentation model.
[0065] In S3, the semantic information of an example is represented as:
[0066]
[0067] in, Example b i The semantic information of F h represents a high-resolution feature map, a mask representing an example;
[0068] The semantic information of the candidate object is expressed as:
[0069]
[0070] in, Represents candidate objects The semantic information of T h represents a high-resolution feature map, Masks representing candidate objects.
[0071] Through experimental verification, we surprisingly found that by introducing negative sample enhancement combined with a prompt-based image segmentation model to assist in semantic information extraction, we not only improved the counting accuracy, but also enhanced the ability to handle occluded and self-similar objects.
[0072] In S4, the candidate regions and examples are clustered by a spatial clustering method, and the candidate regions in the same cluster as the examples are marked as correct targets and output as the final counting and detection results, thereby obtaining the number of objects in the original image.
[0073] In the present invention, the specific process of the spatial clustering method is not limited. Those skilled in the art can select from existing spatial clustering methods according to actual needs. For example, the DBSCAN method is used. For details, see the literature Ester M, Kriegel HP, Sander J, et al. A density-based algorithm for discovering clusters in large spatial databases with noise [C] / / kdd. 1996, 96 (34): 226-231.
[0074] Various embodiments of the method described above in the present invention can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs, which may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general programmable processor, which may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0075] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in the present disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and this document is not limited here.
[0076] Example
[0077] Example 1
[0078] The public dataset FSC-147 dataset is used for object estimation experiments. The FSC-147 dataset contains 6135 images, of which the training set contains 89 categories, and the validation set and test set each contain 29 categories. Each image is annotated with at least three visual examples.
[0079] The experiment includes the following steps:
[0080] S1, generate a density map based on the original image;
[0081] S2, taking the peak point in the density map as the position of the candidate object, and obtaining the mask and high-resolution feature map of the candidate object;
[0082] S3, extracting semantic information of examples and candidate objects based on the acquired masks and high-resolution feature maps;
[0083] S4. Filter non-target objects based on the spatial clustering method to obtain the number of objects in the original image.
[0084] In S1, a local response suppression density map predictor is used to generate a density map, including the following sub-steps:
[0085] S11, extract image features from the original image through convolutional network;
[0086] S12, cutting out example features and negative example features from image features;
[0087] S13, the negative example feature is used as a convolution kernel to convolve with the image feature to generate an attention map;
[0088] S14. Connect the enhanced features and the attention map to obtain the final feature map, which is passed to the regression head of the convolutional network to obtain the density map.
[0089] The convolutional network uses ConvNeXT, the negative examples are obtained by randomly cutting a part of each example, the corresponding true density map is set to zero, and the loss function of the negative examples is set to:
[0090]
[0091] The loss function for the example is set to:
[0092]
[0093] The loss function of the local response suppression density map predictor is:
[0094] L=L neg +L pos .
[0095] Among them, λ1 and λ2 are set to 0.001 and 1 respectively to balance the loss contribution of positive and negative samples.
[0096] The local response inhibition density map predictor model was trained for 300 epochs using the AdamW optimizer with a learning rate of 10 -5 , weight decay is 10 -6 , with a batch size of 4.
[0097] In S2, the SAM model is used to obtain the mask of the candidate object, and the high-resolution feature map is obtained by bilinear interpolation transformation of the image features extracted by the image encoder in the SAM model.
[0098] In S3, the semantic information of an example is represented as:
[0099]
[0100] The semantic information of the candidate object is expressed as:
[0101]
[0102] In S4, the candidate regions and examples are clustered by a spatial clustering method (DBSCAN), and the candidate regions in the same cluster as the examples are marked as correct targets and output as the final count and detection results to obtain the number of objects in the original image. The parameters of DBSCAN are set to eps of 0.07 and the minimum number of samples is 2.
[0103] Comparative Example
[0104] Comparative Example 1
[0105] The same data set as in Example 1 was used to perform the same experiment as in Example 1, except that FamNet, C-DETR, BMNet+, SAFECount, PSeCo, CounTR, LOCA, and DAVE were used respectively.
[0106] Among them, the FamNet method can be found in the literature Ester M, Kriegel HP, Sander J, et al. A density-based algorithm for discovering clusters in large spatial databases with noise [C] / / kdd. 1996, 96 (34): 226-231.
[0107] For the C-DETR method, please refer to the literature Nguyen T, Pham C, Nguyen K, et al. Few-shot object counting and detection [C] / / European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022: 348-365.
[0108] For the BMNet+ method, please refer to the literature Shi M, Lu H, Feng C, et al.Represent, compare, and learn: A similarity-aware framework for class-agnostic counting[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition.2022:9529-9538.
[0109] The SAFECount method can be found in the literature: You Z, Yang K, Luo W, et al. Few-shot object counting with similarity-aware feature enhancement[C] / / Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision. 2023:6315-6324.
[0110] The PSeCo method can be found in the literature: Huang Z, Dai M, Zhang Y, et al. Point Segment and Count: A Generalized Framework for Object Counting[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2024:17067-17076.
[0111] The CounTR method can be found in the literature: Liu C, Zhong Y, Zisserman A, et al. Countr: Transformer-based generalised visual counting[J]. arXiv preprint arXiv:2208.13721, 2022.
[0112] The LOCA method can be found in the literature N, A, Zavrtanik V, et al. A low-shot object counting network with iterative prototype adaptation[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision. 2023:18872-18881.
[0113] For the DAVE method, please refer to the literature Pelhan J, Zavrtanik V, Kristan M.DAVE-A Detect-and-Verify Paradigm for Low-Shot Counting[C] / / Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition.2024:23293-23302.
[0114] The mean absolute error (MAE) and root mean square error (RMSE) are used as evaluation indicators to measure the target estimation effect. The results are shown in Table 1.
[0115] Table 1
[0116]
[0117] The average precision (AP) and average precision 50 (AP50) are used as evaluation indicators to measure the target estimation effect. The results are shown in Table II.
[0118] Table 2
[0119]
[0120]
[0121] It can be seen from Table 1 and Table 2 that the method in Example 1 can achieve the best performance in all indicators compared with the existing advanced methods.
[0122] The present invention has been described above in conjunction with preferred embodiments, but these embodiments are only exemplary and serve only as an illustration. On this basis, the present invention may be subjected to a variety of substitutions and improvements, all of which fall within the scope of protection of the present invention.
Claims
1. A method for estimating the number of objects independent of category, characterized in that: The following steps are involved: S1, generate a density map based on the original image; S2, taking the peak point in the density map as the position of the candidate object, and obtaining the mask and high-resolution feature map of the candidate object; S3, extracting semantic information of examples and candidate objects based on the acquired masks and high-resolution feature maps; S4. Filter non-target objects based on the spatial clustering method to obtain the number of objects in the original image.
2. The method for estimating the number of objects independent of category according to claim 1, characterized in that: S1 includes the following sub-steps: S11, extract image features from the original image through convolutional network; S12, cutting out example features and negative example features from image features; S13, the negative example feature is used as a convolution kernel to convolve with the image feature to generate an attention map; S14. Connect the enhanced features and the attention map to obtain the final feature map, which is passed to the regression head of the convolutional network to obtain the density map.
3. The method for estimating the number of objects independent of category according to claim 1, characterized in that: In S12, the negative examples are obtained by randomly cropping a portion of each example.
4. The method for estimating the number of objects independent of category according to claim 2, characterized in that: The loss function of the negative example is set to the sum of the mean square error loss of the negative example feature and the L1 norm of the density map.
5. The method for estimating the number of objects independent of category according to claim 2, characterized in that: The loss function of the example is set to the sum of the mean squared error loss of the attention map and the mean squared error loss of the prediction density map.
6. The method for estimating the number of objects independent of category according to claim 1, characterized in that: In S3, masks of candidate objects are obtained through a hint-based image segmentation model.
7. The method for estimating the number of objects independent of category according to claim 1, characterized in that: In S2, the high-resolution feature map is obtained by bilinear interpolation transformation of image features; The image features are extracted based on an image encoder in a prompted image segmentation model.
8. The method for estimating the number of objects independent of category according to claim 1, characterized in that: In S3, the semantic information of an example is represented as: in, Example b i The semantic information of F h represents a high-resolution feature map, a mask representing an example; The semantic information of the candidate object is expressed as: in, Represents candidate objects The semantic information of T h represents a high-resolution feature map, Masks representing candidate objects.
9. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.