Few-sample target counting method and device of example prompt strategy based on point guidance

Through point-guided paradigm prompt strategy, combined with multi-scale attention and iterative coding module, the existing methods depend on bounding box annotation is solved, and the counting accuracy and robustness is achieved, which is suitable for a variety of application scenarios.

CN120279374APending Publication Date: 2025-07-08YANSHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510339059.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing few-sample target counting methods rely on bounding box annotations, resulting in poor generalization performance on different data sets, and the high cost of manual labeling limits its application scenarios.

Method used

A point-guided paradigm prompt strategy is adopted, and a new point-guided paradigm prompt method is designed through a multi-scale attention fusion module and iterative example coding module, combining spatial and channel attention mechanisms, and multi-stage coding is used to match multi-stage coding and multi-head similarity, enhancing the model's perception ability of different scale features and the focus of key areas.

Benefits of technology

It significantly improves the counting accuracy and robustness of the model in complex environments, broadens applicable scenarios, reduces data preparation costs, and improves the generalization ability across data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279374A_ABST
    Figure CN120279374A_ABST
Patent Text Reader

Abstract

The invention provides a few-sample target counting method and device of a paradigm prompt strategy based on point guidance, and belongs to the field of computer vision, and the method comprises the steps: extracting deep and shallow layer features of a query image through a feature extractor, and obtaining the features of the query image; the multi-scale attention fusion module obtains multi-scale attention weighted features through convolution processing of different scales and space and channel attention weighting; the point-guided coding block extracts scale sensing addition paradigm features by using point labeling and combining scale prediction; the paradigm features are sent into a multi-head similarity matching block for convolution processing of a given convolution kernel, the correlation between the query image and the paradigm is obtained, and matching features are obtained; sending the matching features into a regression head to obtain a prediction density map; and carrying out pixel-by-pixel addition and summation on the prediction density map to obtain the total number of the interested targets. The point guiding strategy is introduced, the application scene of few-sample target counting is expanded, and the problem that the generalization ability is poor due to the fact that a traditional method depends on bounding box prompting is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a few-shot object counting method and device based on a point-guided exemplar prompting strategy. Background Art

[0002] Object counting is an important and challenging task in the field of computer vision, aiming to estimate the number of specific-category objects in an image or video frame. It is widely applied in fields such as traffic management, agricultural monitoring, and medical research. A large number of studies have shown that after years of in-depth research and practical verification, the object counting for certain specific categories has reached a relatively high precision level, such as crowds, vehicles, animals, and cells. However, traditional methods highly rely on large-scale labeled datasets, which will cause high manual annotation costs in limited fields such as remote sensing. In addition, the specific-category object counting assumes that the object categories in the test phase are already included in the training set, which results in significant limitations when dealing with unseen categories.

[0003] To solve the above problems, recently, the few-shot object counting task has emerged, aiming to estimate the number of objects of any category in an image based on a small number of visual exemplars. Initially, Ranjan et al. pioneered the release of the large-scale few-shot counting dataset FSC-147. They annotated a point for the approximate center of each object in the image and randomly selected three objects of the same category as visual exemplars, annotated with axis-aligned bounding boxes. This move has promoted the progress in the field of few-shot learning, enabling researchers to explore more flexible and generalizable object counting methods. Most existing FSOC methods rely on bounding box prompts and achieve object counting by regressing density maps. Specifically, these methods first extract visual features related to the bounding boxes from the query image, then match these features with the features of the entire query image, and finally use a regression model to generate a density map to predict the number of objects. Although this method has improved the effect of few-shot object counting to a certain extent, this bounding box-based exemplar prompting method has a strong dependence on the specific structure of the dataset, limiting its application in other object counting datasets. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a few-shot object counting method and device based on a point-guided exemplar prompting strategy, by proposing a point-guided exemplar prompting strategy to broaden the applicable scenarios of few-shot object counting methods and create a new few-shot object counting paradigm; in order to overcome the problem of less information provided by point prompts, a multi-scale attention fusion module is designed, which enhances the counting model's perception ability of different-scale features and attention to key regions by combining multi-scale information and introducing spatial and channel attention mechanisms; an iterative example encoding module is also designed, which innovatively uses point prompts to perform multi-stage encoding on visual exemplars for the first time; and through iterative multi-head similarity matching, features irrelevant to the exemplar are effectively suppressed, and key features are highlighted, thereby improving the accuracy of object recognition and counting.

[0005] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:

[0006] A few-shot object counting method based on a point-guided exemplar prompting strategy, comprising the following steps:

[0007] Step 1, a feature extractor extracts the shallow and deep layer features of the query image to obtain the query image features;

[0008] Step 2, the query image features obtained in Step 1 are sent into the multi-scale attention fusion module, and through convolution processing of different scales, spatial and channel attention weighting, multi-scale attention weighted features are obtained;

[0009] Step 3, the multi-scale attention weighted features obtained in Step 2 and the point prompts are input into the iterative encoding and matching module;

[0010] Step 4, the point-guided encoding block uses point annotations and combines scale prediction to extract scale-aware exemplar features;

[0011] Step 5, the exemplar features obtained in Step 4 are sent into the multi-head similarity matching block for convolution processing with a given convolution kernel to obtain the correlation between the query image and the exemplar, and finally matching features are obtained;

[0012] Step 6, the matching features obtained in Step 5 are sent into the regression head, and through a series of convolution and linear layer operations, a predicted density map is obtained;

[0013] Step 7, the predicted density maps in Step 6 are summed pixel by pixel to obtain the total number of objects of interest.

[0014] A further improvement of the technical solution of the present invention lies in: in step 1, the query image is fed into a feature extraction module composed of a convolutional neural network, and multi-scale features are extracted by the second layer, the third layer, and the fourth layer respectively, denoted as {F1, F2, F3}, and then upsampling and channel connection are performed to obtain the query image feature F q ′ .

[0015] A further improvement of the technical solution of the present invention lies in: in step 2, the implementation process of the multi-scale attention fusion module includes the following steps:

[0016] S2.1, first use a 1×1 convolution to reduce the channel dimension of the query image feature F preliminarily fused in step 1, and then use different convolution kernels to extract the feature expressions of multiple receptive fields to generate intermediate multi-scale features F q ′ ; s ′ ;

[0017] S2.2, for the spatial attention branch, perform a pooling operation on the intermediate multi-scale feature F obtained in step S2.1, and use a 7×7 convolution and a Sigmoid activation to generate a spatially weighted feature F s ′ ; The specific process of spatial attention weighting is expressed as: s ;

[0018] F s = F s ′ ⊙σ(Conv 7×7 (Pool(F s ′ ))) (1)

[0019] In the formula, F s ′ is the intermediate multi-scale feature; ⊙ is element-wise multiplication; σ(·) is the Sigmoid function;

[0020] S2.3, for the channel attention branch, use average pooling and max pooling to generate feature maps to extract information in different dimensions, and generate a spatially weighted feature F through 1×1 convolution, ReLU, and Sigmoid s ; Then multiply the spatially weighted feature F s and the channel attention weight element-wise to obtain the spatial and channel attention weight F sc ; The specific process is expressed as:

[0021] F sc = F s ⊙σ(MLP(GAP(F q ′)) + MLP(GMP(F q ′ ))) (2)

[0022] Where F q ′ is the query image feature; F s is the spatial weighted feature; MLP is a shared multi-layer perceptron, consisting of two Conv1×1 and ReLU activations; GAP is global average pooling; GMP is global max pooling;

[0023] S2.4. Weight the spatial attention and channel attention obtained in steps S2.2 and S2.3 to the original feature to obtain the multi-scale attention weighted feature F q ; The specific process is expressed as:

[0024] F q = F q ′ + F sc (3)

[0025] Where F q ′ is the query image feature; F sc is the spatial and channel attention weight.

[0026] A further improvement of the technical solution of the present invention is that in step 4, the exemplar is mapped to different feature spaces in four stages for encoding and matching to obtain the exemplar feature; the exemplar encoding process specifically includes the following steps:

[0027] S4.1. Use the coordinate points to index the point feature F q from the multi-scale attention weighted feature F p ; The specific process is expressed as:

[0028]

[0029] Where Index(·) represents the indexing operation of the corresponding position in the query feature according to the given point coordinates;

[0030] S4.2. Predict the scale information of the target of the point feature obtained in S4.1, and use the self-attention mechanism to enhance the self-correlation between the exemplar features; then generate the blurred bounding box B f ;

[0031] S4.3. Perform RoIAlign operation on the multi-scale attention weighted feature F q to extract the exemplar feature F e ′ of the blurred scale;

[0032] S4.4, use the example feature F obtained in step S4.3 e ′ as a query, and then use the multi-scale attention weighted feature F q as keys and values to input into the multi-head attention module to further optimize the example feature, and finally generate the integrated example feature F e ;

[0033] The whole process of example encoding based on point prompts and query features is expressed as:

[0034]

[0035] where MHA is the multi-head attention mechanism; F q is the multi-scale attention weighted feature; B f is the fuzzy bounding box; F p is the point feature.

[0036] A further improvement of the technical solution of the present invention lies in: in step 5, the process of correlation matching between the query image and the example specifically includes the following steps:

[0037] S5.1, perform linear mapping on the multi-scale attention weighted feature F obtained in step 2 q and the example feature F obtained in step S4.4 e ;

[0038] S5.2, use the example feature F e as a dynamic convolution to perform multi-head convolution operations on the query feature, and obtain a multi-head similarity map R';

[0039] S5.3, perform operations such as channel connection, residual structure, and element-wise addition on the similarity map R' obtained in step S5.2 to obtain a matching feature F m ;

[0040] The similarity weight calculation is defined as the following expression:

[0041]

[0042] where and represent the i-th query image feature and the j-th example feature respectively; l(·) represents the process of linear mapping.

[0043] A few-shot object counting model based on a point-guided example prompting strategy includes:

[0044] A feature extractor for extracting the deep and shallow features of the query image to obtain the query image feature;

[0045] The multi-scale attention fusion module is used to process the query image features through convolutions of different scales, and weight them with spatial and channel attention to obtain multi-scale attention weighted features;

[0046] The iterative encoding and matching module is used to map the multi-scale attention weighted features and the point prompts into different feature spaces in four stages for encoding and matching to obtain the example features;

[0047] The regression head is used to obtain the predicted density map by performing a series of convolutional and linear layer operations on the matching features.

[0048] A further improvement of the technical solution of the present invention lies in that: the iterative encoding and matching module includes a point-guided encoding block and a multi-head similarity matching block;

[0049] The point-guided encoding block is used to obtain the feature vectors of the corresponding position targets from the image features through the point prompts, and then enhance the self-correlation of the targets to be counted, the regional interest alignment and the multi-head attention mechanism through the self-attention mechanism to obtain the predicted example features;

[0050] The multi-head similarity matching block is used to measure the similarity between the example and the query image in a multi-head manner, so as to more comprehensively understand the correlation between the two.

[0051] A storage medium, the storage medium includes stored instructions, wherein when the instructions are running, the device where the storage medium is located is controlled to execute a few-shot object counting method based on a point-guided example prompt strategy.

[0052] Due to the adoption of the above technical solution, the technical progress obtained by the present invention is:

[0053] 1. The present invention innovatively extends the existing few-shot object counting method based on bounding box prompts to a more general point-guided example prompt strategy, effectively solving the problem of poor generalization performance of existing few-shot counting methods due to relying on bounding box annotations. It also significantly improves the counting accuracy and robustness of the model in complex environments by introducing innovative technologies such as multi-scale attention mechanism, point-guided example feature extraction, and dynamic convolution, and is applicable to a variety of application scenarios.

[0054] 2. The present invention first proposes a novel point-guided example prompt method, innovatively introducing a point-guided strategy to expand the application scenarios of few-shot object counting methods, providing a new idea and direction for future research.

[0055] 3. The present invention proposes a multi-scale attention fusion module, which solves the limitations of existing methods in perceiving features of different scales and the deficiencies in focusing on key regions by combining multi-scale information and introducing spatial and channel attention mechanisms.

[0056] 4. The present invention proposes an iterative example encoding module, which for the first time uses point prompts to perform multi-stage encoding on visual examples; at the same time, through iterative multi-head similarity matching, it effectively suppresses features irrelevant to the examples and highlights key features, solving the problems of excessive background interference and easy neglect of key features during the feature matching process, thereby improving the accuracy of target recognition and counting. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings;

[0058] Figure 1 It is a flowchart of a few-shot object counting method based on a point-guided example prompting strategy proposed in the embodiments of the present invention;

[0059] Figure 2 It is a schematic diagram of the overall structure of a few-shot object counting model based on a point-guided example prompting strategy proposed in the embodiments of the present invention;

[0060] Figure 3 It is a detailed schematic diagram of the multi-scale attention fusion module in the embodiments of the present invention;

[0061] Figure 4 It is a detailed schematic diagram of the iterative encoding and matching module in the embodiments of the present invention;

[0062] Figure 5 It is a schematic diagram of the specific implementation of object counting for the few-shot object counting method with a point-guided example prompting strategy provided in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0064] The following will further describe the present invention in detail with reference to the drawings and embodiments:

[0065] As Figure 1 、 Figure 2As shown in the figure, a few-shot object counting method based on a point-guided exemplar prompting strategy, whose core function is to obtain exemplar features through stages such as indexing point features, predicting scales, and obtaining blurred bounding boxes given a small number of point prompts; then train a counting model to enable it to learn the correlation between visual exemplars and objects to be counted, so as to estimate the number of objects of any category in the query image.

[0066] To achieve this goal, the few-shot object counting model based on the point-guided exemplar prompting strategy mainly integrates a feature extractor, a multi-scale attention fusion module, an iterative encoding matching module, and a regression head; as Figure 2 shown;

[0067] The feature extractor is used to extract the deep and shallow layer features of the query image to obtain the query image features;

[0068] The multi-scale attention fusion module is used to process the query image features through convolutions at different scales, and weighted by spatial and channel attention to obtain multi-scale attention weighted features;

[0069] The iterative encoding matching module is used to map the exemplar to different feature spaces for encoding and matching in four stages with the multi-scale attention weighted features and point prompts to obtain exemplar features;

[0070] The regression head is used to obtain the predicted density map by a series of convolutional and linear layer operations on the matching features.

[0071] The iterative encoding matching module includes a point-guided encoding block and a multi-head similarity matching block;

[0072] The point-guided encoding block is used to obtain the feature vectors of the objects at the corresponding positions from the image features through point prompts, and then enhance the self-correlation of the objects to be counted, align the regional interests, and obtain the predicted exemplar features through the self-attention mechanism and the multi-head attention mechanism;

[0073] The multi-head similarity matching block is used to measure the similarity between the exemplar and the query image in a multi-head manner, so as to more comprehensively understand the correlation between the two.

[0074] As Figure 1 shown, the few-shot object counting method based on the point-guided exemplar prompting strategy specifically includes the following steps:

[0075] Step 1, the feature extractor extracts the deep and shallow layer features of the query image to obtain the query image features;

[0076] Specifically, the query image is fed into a feature extraction module composed of a convolutional neural network (CNN, specifically a backbone network consisting of ResNet-50). Multi-scale features are extracted through the second, third, and fourth layers, denoted as {F1, F2, F3}, and then upsampling and channel concatenation are performed to obtain the query image feature F q ′ .

[0077] Step 2: Feed the query image feature F obtained in Step 1 q ′ into the multi-scale attention fusion module. Through convolutional processing at different scales, spatial and channel attention weighting, the multi-scale attention weighted feature F q is obtained;

[0078] The multi-scale attention fusion module integrates spatial and channel attention, adaptively highlighting key regions while capturing multi-scale features, and effectively balancing global context and local details. To facilitate understanding of the implementation process of the multi-scale attention fusion module, the following is a specific description in combination with Figure 3 it:

[0079] S2.1: First, use a 1×1 convolution to reduce the channel dimension of the query image feature F preliminarily fused in Step 1 q ′ Then, extract feature expressions of multiple receptive fields using different convolutional kernels, such as different convolutional kernel sizes (3×3, 5×5, 7×7), to generate intermediate multi-scale features F s ′ ;

[0080] S2.2: For the spatial attention branch, perform a pooling operation on the intermediate multi-scale feature F obtained in Step S2.1 s ′ and use a 7×7 convolution and a Sigmoid activation to generate the spatial weighted feature F s ; The specific process of spatial attention weighting is expressed as:

[0081] F s = F s ′ ⊙ σ(Conv 7×7 (Pool(F s ′ ))) (1)

[0082] In the formula, F s ′ is the intermediate multi-scale feature; ⊙ is element-wise multiplication; σ(·) is the Sigmoid function.

[0083] S2.3. For the channel attention branch, average pooling and max pooling are used to generate feature maps to extract information in different dimensions. Through 1×1 convolution, ReLU, and Sigmoid, the spatial weighted feature F is generated. s . Then, the spatial weighted feature F s and the channel attention weight are multiplied element by element to obtain the spatial and channel attention weight F sc ; The specific process is expressed as:

[0084] F sc = F s ⊙ σ(MLP(GAP(F q ′ )) + MLP(GMP(F q ′ ))) (2)

[0085] In the formula, F q ′ is the query image feature; F s is the spatial weighted feature; MLP is a shared multi-layer perceptron composed of two Conv1×1 and ReLU activations; GAP is global average pooling; GMP is global max pooling.

[0086] S2.4. The spatial and channel attention obtained in steps S2.2 and S2.3 are weighted to the original features to obtain the multi-scale attention weighted feature F q ; The specific process is expressed as:

[0087] F q = F q ′ + F sc (3)

[0088] Step 3. The multi-scale attention weighted feature F q in step 2 and the point prompt are input into the iterative encoding matching module;

[0089] Step 4. The point-guided encoding block extracts scale-aware exemplar features by using point annotations and combining scale prediction, maps the exemplars to different feature spaces in four stages for encoding and matching, and obtains the exemplar feature F e ;

[0090] It should be noted that due to the small amount of point prompt information, it is difficult to identify the overall target. Therefore, the size and depth of the exemplar in the process of mapping the exemplar to different feature spaces will affect the final counting accuracy. The sizes of the exemplars in the four stages are {1, 3, 5, 7} in sequence, and the depth of the exemplars is 128. The following will combine Figure 4 (a) to specifically describe the exemplar encoding process:

[0091] S4.1, Index the feature F of the indexed point from the multi-scale attention weighted feature F q using the coordinate points; p The specific process is expressed as:

[0092]

[0093] In the formula, Index(·) represents the indexing operation on the corresponding position in the query feature according to the given point coordinates.

[0094] S4.2, Predict the scale information of the target from the point features obtained in S4.1, and adopt the self-attention mechanism to enhance the self-correlation between the example features, so as to improve the integrity of the example feature expression and the resolution of the target area. Then generate the fuzzy bounding box B in combination with the coordinate points f ;

[0095] S4.3, Perform RoIAlign operation on the multi-scale attention weighted feature F q to extract the example features F of the fuzzy scale e ′ so as to expand the feature range and enhance the richness and robustness of the features;

[0096] S4.4, Use the example features F e ′ obtained in step S4.3 as the query, and then use the multi-scale attention weighted feature F q as the key and value to input into the multi-head attention mechanism to further optimize the example features, and finally generate the integrated example features F e to ensure the comprehensiveness and accuracy of the feature information.

[0097] The whole process of example encoding according to the point prompt and query feature is expressed as:

[0098]

[0099] In the formula, MHA is the multi-head attention mechanism; F q is the multi-scale attention weighted feature; B f is the fuzzy bounding box; F p is the point feature.

[0100] Step 5, Send the example features F e obtained in step 4 into the multi-head similarity matching block for convolution processing with a given convolution kernel to obtain the correlation between the query image and the example. The multi-head similarity matching uses the example features as dynamic convolution kernels in a multi-head manner for similarity matching, effectively highlighting the target area while suppressing background noise; this process also goes through four stages, and finally obtains the matching features F m .

[0101] The following will combine with Figure 4 (b) to specifically describe the process of correlation matching between the query image and the exemplar:

[0102] S5.1, linearly map the multi-scale attention weighted feature F q obtained in step 2 and the exemplar feature F e obtained in step S4.4;

[0103] S5.2, use the exemplar feature F e as a dynamic convolution to perform a multi-head convolution operation on the query feature to obtain a multi-head similarity map R';

[0104] S5.3, perform operations such as channel connection, residual structure, and element-wise addition on the similarity map R' obtained in step S5.2 to obtain a matching feature F m ;

[0105] The similarity weight calculation is defined as follows:

[0106]

[0107] In the formula, and respectively represent the i-th query image feature and the j-th exemplar feature; l(·) represents the process of linear mapping.

[0108] Step 6, send the matching feature F m obtained in step 5 into the regression head (tail network), and through a series of convolution and linear layer operations, obtain a predicted density map;

[0109] Step 7, perform pixel-wise addition and summation on the predicted density map in step 6 to obtain the total number of target objects of interest.

[0110] Figure 5An example of the few-shot object counting model based on point-guided exemplar prompting strategy in practical applications is presented, specifically taking a set of photos of grapes as an example. In this set of photos, a part provides three grape samples of a specific type, with a red dot marked at the center of each sample, and the other part is the photos of grapes to be counted. By using the method of the present invention, the counting model can quickly and accurately identify and count each grape in the grapes only through a small number of point prompts, even when the annotation information of the counting objects provided is very limited. Specifically, the counting model analyzes the query image and the several visual exemplars provided, and extracts key context features from them, including information such as the shape, color, and texture of the grapes. Then, these features are used to guide the counting process of the grape photos. In the grape photos, the counting model uses these features to identify the grapes, thus achieving accurate counting. Due to considering the context information, this method can better handle the counting tasks in complex environments, such as problems like occlusion, overlap, or background interference that the grapes may have, thus achieving efficient and accurate counting performance.

[0111] In summary, the present invention aims to solve the problem that existing few-shot counting methods generally rely on bounding box annotations, resulting in poor generalization performance on other datasets. Traditional methods are limited in their flexibility and adaptability in different application scenarios due to their excessive reliance on annotated data in a specific format. To address this challenge, the present invention proposes a more general and flexible solution, significantly improving the performance of the counting model on unseen datasets, thereby enhancing the cross-dataset generalization ability. By reducing the reliance on bounding box annotations, the present invention not only reduces the cost and complexity of data preparation but also improves the robustness and applicability of the counting model in diverse tasks.

[0112] In the present invention, by inputting a query image into a feature extraction module composed of a convolutional neural network (CNN), the features of the last three stages are first upsampled and fused to obtain the deep and shallow features of the query image. Subsequently, these query image feature maps are processed by convolutional kernels of different scales (such as 3×3, 5×5, 7×7) to generate intermediate multi-scale features, which are weighted by a spatial attention mechanism to capture the significant regions in the query image, resulting in spatially weighted features. At the same time, a residual structure is introduced and combined with a channel attention mechanism to further enhance the spatial attention features, and the final multi-scale attention weighted features are obtained through element-wise multiplication and addition operations. Next, based on the above multi-scale attention weighted features and combined with point cue coordinates, the counting model can accurately locate the feature representation corresponding to the position of the point, and through scale prediction, RoIAlign pooling, and a multi-head attention mechanism, extract point-guided exemplar features. Subsequently, the counting model performs a linear mapping on these point-guided exemplar features and the query image features, uses the exemplar features as dynamic convolutional kernels to perform convolutional operations on the query features, and generates a similarity map. To further optimize the similarity map, the model introduces a residual structure to perform weighted and element-wise addition operations on it, thereby obtaining matching features. Finally, the matching features are fed into a regression head, and after a series of complex non-linear transformations, including the application of multi-layer convolutional operations and activation functions, the predicted density map is finally regressed. By summing up the predicted density map pixel by pixel, the counting model can efficiently and accurately calculate the total predicted number of the target of interest in the query image. This method not only effectively solves the problem that the existing few-shot counting methods rely on bounding box annotations and result in poor generalization performance, but also significantly improves the counting accuracy and robustness of the model in complex environments by introducing innovative technologies such as multi-scale attention mechanism, point-guided exemplar feature extraction, and dynamic convolution, and is applicable to a variety of application scenarios.

[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A few-shot object counting method based on a point-guided exemplar prompting strategy, characterized in that, It includes the following steps: Step 1, the feature extractor extracts the shallow and deep features of the query image to obtain the query image features; Step 2, the query image features obtained in Step 1 are fed into the multi-scale attention fusion module, and through convolutional processing at different scales, spatial and channel attention weighting, multi-scale attention weighted features are obtained; Step 3, the multi-scale attention weighted features and the point prompt obtained in Step 2 are input into the iterative encoding and matching module; Step 4, the point-guided encoding block uses point annotations and combines scale prediction to extract scale-aware exemplar features; Step 5, the exemplar features obtained in Step 4 are fed into the multi-head similarity matching block for convolutional processing with a given convolutional kernel to obtain the correlation between the query image and the exemplar, and finally the matching features are obtained; Step 6, the matching features obtained in Step 5 are fed into the regression head, and through a series of convolutional and linear layer operations, a predicted density map is obtained; Step 7, the predicted density maps in Step 6 are summed pixel by pixel to obtain the total number of target objects of interest.

2. The few-shot object counting method based on a point-guided exemplar prompt strategy according to claim 1, wherein In step 1, the query image is fed into a feature extraction module composed of a convolutional neural network. Multi-scale features, denoted as {F1, F2, F3}, are extracted through the second layer, the third layer, and the fourth layer respectively, and then upsampling and channel connection are performed to obtain the query image feature F′ q .

3. The few-shot object counting method based on a point-guided exemplar prompting strategy according to claim 1, characterized in that In Step 2, the implementation process of the multi-scale attention fusion module includes the following steps: S2.1, first use a 1×1 convolution to reduce the channel dimension of the preliminarily fused query image feature F′ in step 1, and then use different convolutional kernels to extract feature expressions of multiple receptive fields to generate intermediate multi-scale features F′ q ; s ; S2.

2. For the spatial attention branch, for the intermediate multi-scale feature F obtained in step S2.1 s ', perform a pooling operation, and use a 7×7 convolution and Sigmoid activation to generate a spatially weighted feature F s ; The specific process of spatially attention weighting is expressed as: F s = F s ' ⊙ σ(Conv 7×7 (Pool(F s '))) (1) where F s ′ is the intermediate multi-scale feature; ⊙ is the element-wise multiplication; σ(·) is the Sigmoid function; S2.

3. For the channel attention branch, average pooling and max pooling are used to generate feature maps to extract information in different dimensions, and spatial weighted feature F is generated through 1×1 convolution, ReLU, and Sigmoid. s Then, the spatial weighted feature F s and the channel attention weights are multiplied element by element to obtain the spatial and channel attention weights F sc ; The specific process is expressed as: F sc = F s ⊙σ(MLP(GAP(F q ′)) + MLP(GMP(F q ′))) (2) where F q ′ is the query image feature; F s is the spatial weighted feature; MLP is a shared multi-layer perceptron composed of two Conv1×1 and ReLU activations; GAP is global average pooling; GMP is global max pooling; S2.4, weight the spatial attention and channel attention obtained in steps S2.2 and S2.3 to the original features to obtain the features F after multi-scale attention weighting q ; The specific process is expressed as: F q = F' q + F sc (3) Where, F′ q is the query image feature; F sc is the spatial and channel attention weight.

4. A few-shot object counting method based on a point-guided exemplar prompting strategy according to claim 1, characterized in that In Step 4, the exemplar is mapped to different feature spaces for encoding and matching in four stages to obtain exemplar features; the exemplar encoding process specifically includes the following steps: S4.1, using the coordinate points to index the point feature F in the multi-scale attention weighted feature F q ; specifically, the process is expressed as: p ​ Where Index(·) represents the indexing operation on the corresponding position in the query feature according to the given point coordinates; S4.

2. Scale information of the point feature prediction target obtained in S4.1 is used, and a self-attention mechanism is adopted to enhance the self-correlation between exemplar features; then a blurred bounding box B is generated by combining coordinate points. f ; S4.3, perform RoIAlign operation on the multi-scale attention weighted feature F q to extract the exemplar feature F of the fuzzy scale e ′ ; S4.4, use the exemplary feature F obtained in step S4.3 e ′ as a query, and then use the multi-scale attention weighted feature F q as the key and value to input into the multi-head attention module to further optimize the exemplary feature, and finally generate the integrated exemplary feature F e ; The whole process of exemplar encoding according to the point prompt and the query feature is expressed as: Wherein, MHA is the multi-head attention mechanism; F q is the multi-scale attention weighted feature; B f is the fuzzy bounding box; F p is the point feature.

5. The few-shot object counting method based on a point-guided exemplar prompt strategy according to claim 4, wherein In Step 5, the process of correlation matching between the query image and the exemplar specifically includes the following steps: S5.1, linearly map the multi-scale attention weighted feature F obtained in step 2 q and the exemplar feature F obtained in step S4.4 e ; S5.2, use the example feature F e to perform a multi-head convolution operation on the query feature as a dynamic convolution to obtain a multi-head similarity map R'; In S5.3, perform operations such as channel connection, residual structure, and element-wise addition on the similarity graph R' obtained in step S5.2 to obtain the matching feature F m ; The similarity weight calculation is defined as follows: In the formula, and represent the i-th query image feature and the j-th exemplar feature respectively; l(·) represents the process of linear mapping.

6. A few-shot object counting model based on a point-guided exemplar prompting strategy, characterized in that, It includes: A feature extractor for extracting the shallow and deep features of the query image to obtain the query image features The multi-scale attention fusion module is used to perform convolutional processing at different scales, spatial and channel attention weighting on the query image features to obtain multi-scale attention weighted features; The iterative encoding and matching module is used to map the multi-scale attention weighted features and the point prompt to different feature spaces for encoding and matching in four stages to obtain exemplar features; The regression head is used to perform a series of convolutional and linear layer operations on the matching features to obtain a predicted density map.

7. A few-shot object counting model based on a point-guided exemplar prompting strategy according to claim 6, wherein The iterative encoding and matching module includes a point-guided encoding block and a multi-head similarity matching block; The point-guided encoding block is used to obtain the feature vector of the target at the corresponding position from the image features through the point prompt, and then enhance the self-correlation of the target to be counted, regional interest alignment and multi-head attention mechanism through the self-attention mechanism to obtain the predicted exemplar features; The multi-head similarity matching block is used to perform similarity measurement on the exemplar and the query image in a multi-head manner, so as to more comprehensively understand the correlation between the two.

8. A storage medium, characterized in that, The storage medium includes stored instructions, wherein when the instructions are running, the device where the storage medium is located is controlled to execute the few-shot object counting method based on the point-guided exemplar prompt strategy according to any one of claims 1 to 5.